The AWS corpus, the tokenizer, and the spell corrector

22 min read~200 wpm

The previous article was about why I'm changing course from the Manning book's toy corpus to AWS documentation. This one is about the actual work: getting the corpus on disk, rewriting the tokenizer so AWS-shaped tokens survive, and fixing the spell corrector so it stops suggesting typos that happen to exist in the corpus.

visual · what the tokenizer rewrite changes
old tokenizer
strip everything non-alphanumeric
s3:GetObjects3getobject
arn:aws:s3:::my-bucketarnawss3mybucket
t2.microt2micro
HTTP/1.1http11
us-east-1useast1
s3:Get*s3get
new tokenizer
keep . - _ : / ' @ * with asymmetric edge trimming
s3:GetObjects3:getobject
arn:aws:s3:::my-bucketarn:aws:s3:::my-bucket
t2.microt2.micro
HTTP/1.1http/1.1
us-east-1us-east-1
s3:Get*s3:get*

Getting the corpus

My first instinct was to scrape docs.aws.amazon.com as HTML, parse the DOM, and extract text content. That's what most search-engine corpora look like — you pull pages, parse HTML, tag the structural elements, and lift the body text out.

While reading through AWS docs by hand to figure out which tags I'd need to extract, I noticed something. AWS publishes its documentation source on GitHub under the awsdocs/ org — the EKS user guide, the Lambda developer guide, OpenSearch, Panorama, all of them are public repositories of markdown files. The docs site renders these to HTML at build time.

Better still, the URLs are symmetric. Every page at https://docs.aws.amazon.com/.../foo.html is also available at https://docs.aws.amazon.com/.../foo.md. Swap the extension and you get the markdown source directly.

This removes a whole layer of work. There's no HTML parsing, no DOM traversal, no figuring out which <div> is the article body and which is the sidebar. The markdown already has the structure I need — # is a title, ## is a section, triple-backticks are a code block — which are identifiers I can rely on later when I extract zones.

The scraper runs two passes. The first is sitemap discovery: AWS publishes a sitemap.xml per service guide, so for each of the 18 services I care about I hit every plausible guide name (UserGuide, APIReference, DeveloperGuide, dg, monitoring, and so on), grab all <loc> URLs from the XML, swap .html for .md, fetch, and save to disk mirroring the URL path. The second pass is link-driven verification: after Pass 1, walk every saved .md file, extract every [text](path.md) link, and fetch anything that isn't already on disk. Repeat until no new files appear. This catches pages the sitemap missed.

The scraper is polite (0.5s delay between requests), resumable (skips files already on disk), logs the full run to scraper.log, and records every failure in failures.jsonl with a timestamp, the URL, the pass number, and the reason.

The final corpus is 14,266 markdown files across 18 AWS services. Total scrape time was about 1h48m. There were 4 failures, all http_404, all dead internal links in AWS's own docs — not my problem to fix.

I spot-checked 20 random files and the content looked clean: real markdown, real AWS prose, no encoding artifacts. One thing I noticed was that a handful of EKS pages start with a "Help improve this page" GitHub-contribute boilerplate before the actual content begins — flagged for the markdown adapter to strip later, since the tokenizer would tokenize the boilerplate as if it were content.

What the old tokenizer destroyed

The original cleanup.rs was 5 lines:

pub fn split_string(content: String) -> Vec<String> {
    content.split_whitespace()
        .map(|word| word.to_string().to_lowercase()
            .chars().filter(|c| c.is_alphanumeric()).collect())
        .collect()
}

Split on whitespace, lowercase, strip everything that isn't [a-z0-9]. This works fine on the 20 Newsgroups corpus, where the tokens of interest are English words, but it falls apart immediately on AWS docs. The hero visual above shows the failure modes — s3:GetObject destroyed into s3getobject, arn:aws:s3:::my-bucket mangled to arnawss3mybucket, t2.micro flattened to t2micro, and so on.

If a user searches for s3:GetObject, the tokenizer destroys their query into s3getobject. The index — which went through the same tokenizer — contains the same junk token. The query technically matches something, but the result is garbage, because s3getobject doesn't carry the semantic content of s3:GetObject. The IAM action structure has been flattened past recognition.

Fixing this is the whole reason the tokenizer is getting rewritten. The question is what to keep.

Picking what to keep, from data

The answer depends on the corpus. AWS prose has specific connector characters that appear inside tokens, and rather than guess at the right set, I scanned all 14,266 files and counted.

For every non-alphanumeric character in the corpus I counted two things. An internal occurrence is where the character has an alphanumeric on both sides — like the : in s3:getobject. A boundary occurrence is where it has whitespace or other non-alphanumeric on at least one side — like the : in **Cause:**. The intuition is that anything above ~50% internal is part of a token, anything below ~10% is punctuation, and the middle ground needs human judgment.

Top non-alphanumeric chars by total frequency:

chartotalinternalboundaryinternal %examples
.1,326,799720,179606,62054.3%t2.micro, HTTP/1.1
-1,092,133855,724236,40978.4%us-east-1, my-bucket
/960,805507,860452,94552.9%HTTP/1.1, application/json
"706,9981,052705,9460.1%(quote characters in code)
*703,089196702,8930.03%(markdown emphasis, mostly)
:646,625133,656512,96920.7%s3:GetObject, arn:aws:iam
`555,430115555,3150.02%(inline code markers)
,515,4078,553506,8541.7%(sentence punctuation)
_322,369294,54827,82191.4%API_GetSAMLProvider
#251,28220,428230,8548.1%(markdown headers)
=183,95716,490167,4679.0%key=value pairs
'86,31739,85746,46046.2%don't, it's
@9,3461,1058,24111.8%noreply@example.com

The underscore at 91.4% internal is unambiguous — it exists almost exclusively inside snake_case identifiers like API_GetSAMLProvider. The hyphen at 78.4% is similarly clear: AWS uses hyphens in resource IDs, region codes, and command flags. Both get kept.

The dot and slash sit at 54% and 53%, split roughly evenly between internal and boundary uses, but the internal cases are exactly what I want to keep: t2.micro, HTTP/1.1, file paths, version strings. The boundary cases are sentence punctuation and URL endings, which I can handle separately by trimming the edges of tokens.

The colon at 20.7% looks low at first, but the internal cases are critical — ARNs, IAM actions, HTTP headers. These are exactly the tokens I rebuilt the tokenizer to preserve. Keep.

The apostrophe at 46.2% is contractions like don't, it's, doesn't, with curly-quote normalization so it's and it’s collapse to the same form. The at-sign at 11.8% looks low but the internal cases are all emails, which are real semantic units.

The asterisk is the interesting one. At 0.03% internal it looks like an obvious discard, and most of the 196 internal occurrences are junk like aws:ecs:clusterName*Value from IAM policy edge cases. But the use case I actually care about — s3:Get* as an IAM wildcard — has the * at the end of the token with whitespace after it, which my classifier calls a boundary occurrence and not internal. So the metric doesn't capture the case that matters. The character is semantically meaningful even when it appears terminally, so I'm keeping it. The slash has the same property at the start of paths: /api/path is classified as boundary but the leading slash is semantically meaningful.

The final keep list comes out to . - _ : / ' @ * — eight characters.

Everything else splits. Markdown formatting characters like **, [, ], (, ), >, <, |, = all split, which means the tokenizer never has to know anything about markdown syntax — those characters either carry no meaning on their own or are already accounted for in the keep list.

Edges aren't symmetric

Once * and / are in the keep list, a new problem appears. Consider four cases:

What this means is that not all keep-chars deserve the same treatment at edges — some carry meaning at the start or end of a token, others are just dangling. The tokenizer ends up with two sets: a KEEP set of characters allowed inside tokens (. - _ : / ' @ *), and a TRIM_AT_EDGES set of characters stripped if they appear at a token boundary (. - _ : ' @). The asterisk and slash are in KEEP but not in TRIM_AT_EDGES, so they survive at the edges of tokens, while everything else in KEEP gets trimmed if it ends up there.

This is asymmetric on purpose: cause: trims to cause because the trailing : is just markdown punctuation, but s3:Get* stays as s3:get* because the trailing * is the wildcard.

The .airforce case isn't a problem because the same tokenizer runs during indexing and querying. If someone searches for .airforce, the leading . gets trimmed and the query becomes airforce. The doc was indexed as airforce too. They still match.

Normalization

On top of the keep/trim logic, the tokenizer also normalizes a few things. Curly quotes (’, ‘, “, ”) collapse to their ASCII equivalents so the same word in different encodings produces the same token. Latin ligatures (fi, fl, ffi) expand to their component letters (fi, fl, ffi) — there are only 19 of these in the corpus, which is tiny, but they were enough to crash the indexer the first time one slipped through. Everything gets lowercased. And I cap max token length at 64 bytes, since real AWS tokens (full ARNs, snake_case API names) top out around 50 characters and anything longer is base64, binary garbage, or stray data that should be dropped.

The character analysis already told me to exclude characters with very low internal counts, but I went one step further on the alphabetic side. Instead of Rust's is_alphanumeric() (which returns true for Unicode "Number" category characters like ¹ U+00B9), I'm using is_ascii_alphanumeric(). The corpus has ¹ appearing as a footnote marker fused to words: messages¹. With the looser check, the tokenizer produces tokens like numberofmessagesdeleted¹, which is exactly the kind of garbage that has zero IDF value and confuses the spell corrector downstream. With ASCII-only, the same input becomes numberofmessagesdeleted — same word, no garbage, no Unicode-induced crashes downstream either.

I'll get to the Unicode crash in a minute.

Results: tokenizer

26 unit tests pass. The ones that matter:

s3:GetObject                  → ["s3:getobject"]
arn:aws:s3:::my-bucket        → ["arn:aws:s3:::my-bucket"]
t2.micro                      → ["t2.micro"]
HTTP/1.1                      → ["http/1.1"]
us-east-1                     → ["us-east-1"]
API_AssociateEncryptionConfig → ["api_associateencryptionconfig"]
noreply@example.com           → ["noreply@example.com"]
s3:Get*                       → ["s3:get*"]
/2013-04-01/foo               → ["/2013-04-01/foo"]
don't                         → ["don't"]
don’t  (curly)                   → ["don't"]
cause:                        → ["cause"]
**Note**                      → ["**note**"]   (markdown adapter strips ** later)
"!!! ??? ,,,"                 → []             (pure punctuation)

Indexing the full 14,266-doc corpus took 187 seconds in a debug build, and the final index has 266,624 terms.

What changed in live queries: with the old tokenizer, s3:GetObject got destroyed into s3getobject and matched a single junk term across the index. With the new tokenizer, s3:GetObject stays as s3:getobject, which matches 10 real docs — all of them about S3 IAM permissions — with batch-ops-iam-role-policies.md at the top. The query is preserved with its semantic structure intact, and the index has the same structure on disk. Same story for arn:aws:iam, which finds 27 docs across S3, DynamoDB, KMS, and CloudFormation, all of them real ARN-bearing pages. RESOURCE:CPU doesn't show up in tier 0 or tier 1, but tier 2 returns 3 docs and the top one is AmazonECS/.../api_failures_messages.md — exactly where this error message appears in the corpus. None of these queries return anything useful with the old tokenizer.

A digression about Unicode that cost me half an hour

The first run of the new engine on the AWS corpus crashed during indexing:

thread 'main' panicked at src/three_gram_index.rs:11:26:
byte index 7 is not a char boundary; it is inside 'fi' (bytes 6..9) of `$significant$`

The corpus contains a few pages with the Latin ligature fi (U+FB01) instead of fi. The tokenizer (before I added ligature normalization) accepted it because Rust's is_alphanumeric() considers it alphanumeric. The trigram indexer then tried to slice the token significant into 3-byte windows: [0..3], [1..4], and so on.

fi is one Unicode codepoint, but in UTF-8 it's 3 bytes. When the slicing landed at [7..10], it tried to cut into the middle of the ligature's byte sequence. Rust's String slicing requires char boundaries, so the program panicked rather than silently producing garbage.

Two fixes were needed. First, normalize ligatures in the tokenizer (fi → fi, fl → fl, ffi → ffi) so the trigram indexer never sees a multibyte character. Second, rewrite the trigram indexer itself to use char-iteration instead of byte-slicing, as defense in depth — even if a Unicode character somehow slips through, the indexer should walk characters, not bytes. I replaced padded[i..i+3].to_string() with padded.chars().collect::<Vec<_>>().windows(3).map(...). In theory this is ~2–3x slower per character because of UTF-8 boundary checks, but in practice it's irrelevant because trigram indexing happens once per term at index time.

The same pattern repeated in the spell corrector — same byte-slicing bug, same fix.

The general principle: when working with text, walk by character, not by byte. The performance cost is negligible for offline operations and the robustness benefit is that the indexer doesn't crash on real-world data.

The spell corrector and the instnace problem

The Manning-style spell corrector pipeline is straightforward:

  1. Tokenize the misspelled query word into trigrams.
  2. Look up each trigram in the trigram index. Every term that shares at least one trigram with the query is a candidate.
  3. Compute Jaccard similarity (trigram set overlap) for each candidate. Keep candidates above a threshold (I use 0.3).
  4. Compute Levenshtein edit distance for each surviving candidate. Keep candidates whose edit distance is below term_length / 2.
  5. Sort by edit distance ascending and return the top candidate.

This works fine on a small clean corpus, but on AWS docs it fails in a subtle way that's worth showing.

A user searches EC2 instnace metadata service. The corrector processes instnace. The trigram lookup returns 20,826 raw candidates. The Jaccard filter narrows that down to ~3, and the Levenshtein filter keeps 2: instance (edit distance 1) and instace (also edit distance 1).

When two candidates have the same edit distance, which one wins? The original sort returned them in candidate order, which is essentially random. On the AWS corpus the random winner was instace — a typo that appears in exactly one AWS doc as a corpus-internal misspelling. The query gets rewritten to EC2 instace metadata service. The term instace returns 1 doc. The intersection of all 4 query terms returns zero. The user sees nothing.

The right answer was obvious in retrospect: when edit distances tie, prefer the more frequent term. instance appears in roughly 5,000 AWS docs, and instace appears in 1. The probability that a user typing instnace meant instance is overwhelmingly higher than the probability they meant instace. The frequency information was already in the index — the TermEntry.doc_freq field, used elsewhere for IDF — but the corrector wasn't reading it.

The fix is two pieces.

1. Frequency tiebreaker on the final sort.

// before: sort by edit distance only
candidates.sort_by_key(|c| c.edit_distance);

// after: sort by (edit_distance ASC, doc_freq DESC)
candidates.sort_by(|a, b|
    a.edit_distance.cmp(&b.edit_distance)
        .then(b.doc_freq.cmp(&a.doc_freq))
);

Edit distance is still the primary sort key, so a close match always wins over a popular one. But when matches tie on edit distance, the popular one wins.

2. Minimum doc-frequency threshold for candidates.

Even with the tiebreaker, the candidate pool still includes terms like instace (doc_freq = 1), which lose at the sort but consume CPU during Jaccard and Levenshtein computation along the way. There's also a case where the correct answer happens to appear only once in the corpus, and the threshold prevents that from leaking through as a suggestion either.

const MIN_CANDIDATE_DOC_FREQ: u32 = 3;

candidates.retain(|c| c.doc_freq >= MIN_CANDIDATE_DOC_FREQ);

A term appearing in fewer than 3 docs is almost certainly itself a typo or rare junk token, and suggesting it as a correction misleads the user. The cutoff of 3 is judgment — I'd want to validate it against a held-out test set when I build the evaluation harness — but it's close enough to ship.

The effect on the candidate pool is substantial. Here's instnace walking through the funnel:

spell-correction funnel · query "instnace"
trigram lookup
20,826
after freq filter
3,823
after Jaccard
1
after edit distance
1
→ instance ✓

Three more queries through the same funnel:

'kubernetis' → 12,613 raw → 10,515 dropped by freq → 2,098 → 16 after Jaccard →  3 after edit distance → kubernetes ✓
'scailing'   → 30,659 raw → 25,306 dropped by freq → 5,353 → 12 after Jaccard → 10 after edit distance → scaling    ✓
'lamda'      →  4,501 raw →  3,656 dropped by freq →   845 →  1 after Jaccard →  1 after edit distance → lambda     ✓

80%+ of trigram candidates are corpus typos. The frequency filter cuts them before Jaccard and Levenshtein ever run on them, which had a side benefit I wasn't expecting: spell-check latency dropped 2-3x because the expensive operations were operating on fewer candidates. The kubernetis correction went from 562ms to 165ms, and the clouwatch correction went from 379ms to 140ms. The accuracy fix happened to be a speed fix too.

Live results

Eight test queries with typos, top result reported:

querycorrected totop result
kubernetiskuberneteseks/.../Welcome.md
dynmodb partishun keydynamodb [dropped] keyamazondynamodb/.../HowItWorks.Partitions.md
clouwatch alarmcloudwatch alarmAmazonRDS/.../creating_alarms.md
lamda layerlambda layerlambda/.../creating-deleting-layers.md
auto scailing groupauto scaling groupeks/.../enable-asg-metrics.md
vpc routng tablevpc routing tablevpc/.../intra-vpc-route.md
EC2 instnace metadata serviceec2 instance metadata serviceAWSEC2/.../instancedata-data-retrieval.md
s3 buckts3 bucketRoute53/.../troubleshooting-s3-bucket-website-hosting.md

A few things worth noting in those results. The lamda layer case is the kind of result I'm aiming for everywhere: lamda corrects to lambda, and the top 4 results are all canonical Lambda layers documentation — creating-deleting-layers.md, chapter-layers.md, adding-layers.md, packaging-layers.md.

The dynmodb partishun key case is more interesting because partishun got dropped entirely. No candidate within the edit-distance cap (len/2 = 4) had a doc_freq above the threshold — partition is 4 edits away from partishun, right at the cap. The engine warns the user that a term was dropped and proceeds with dynamodb key. The top result is still HowItWorks.Partitions.md, which is the right doc anyway, because dynamodb carried enough specificity to surface it.

The s3 buckt case shows a problem that the spell corrector can't fix. The correction itself works fine (s3 bucket), but the top result is a Route53 troubleshooting page that mentions "s3 bucket" frequently in a short doc, rather than the canonical S3 user guide. This is a scoring issue, not a tokenizer or spell-corrector issue — TF-IDF with cosine normalization rewards short docs with high term frequency. BM25 handles this differently with its b parameter for length normalization, which is what the next article is about.

What I haven't done yet

This article covers the corpus, the tokenizer, and the spell corrector. The engine's ranking is still TF-IDF + cosine normalization (with the proximity boost from the earlier article). The known failure modes that remain:

  1. Short tangential pages outrank long authoritative pages on common queries. Cosine normalization divides by doc length too aggressively. BM25 will fix this.
  2. Long natural-language queries like how to create vpc peering connection return zero results because strict AND-intersection of 6+ terms is fragile. BM25's partial-match scoring helps, and dropping common terms like how and to before intersection helps more.
  3. No zones yet. A token in a page title scores the same as a token in body text. BM25F adds per-zone weights so titles, section headers, and code blocks get distinct weights. The zones are sitting in the markdown structure waiting to be extracted (#, ##, code fences, the ** [fieldName] ** pattern in API references) — same approach I used on Wikipedia film infoboxes for a previous project.
  4. No PageRank yet. AWS docs cross-link extensively, and a doc that 50 other docs link to is more authoritative than one nothing links to. The link graph is already implicit in the corpus from Pass 2 of the scraper. PageRank on this graph combined with BM25F gives the final ranking.
  5. No evaluation harness. Until I have 30 hand-judged queries with relevance labels, every "is this better?" claim is a vibe check, and BM25 tuning without one is a way to spend a week and have no idea whether anything actually improved.

All of these are on the path. BM25 is next, and the evaluation harness gets built in parallel because the BM25 tuning has nothing to optimize against without it.

What I have today is a corpus of 14,266 AWS docs indexed with a tokenizer that preserves AWS-specific syntax, queried by a spell corrector that handles typos on AWS jargon and resolves the instnace/instance ambiguity correctly. It answers queries like s3:GetObject that no general-purpose tokenizer can parse, and queries like kubernetis that the AWS docs site itself would fail on, in under 200ms. The math is the same math from the Manning chapters — the work was in making it operate on a corpus that real developers actually search.


Code: github.com/sreenish27/Search_engine_rs