Reference Sheets
Reference Data
Sources and generation rules for reference-sheets-oanc-leipzig-wikidata-2026-08-03.
English statistics
English letters, digrams, trigrams, contacts, word boundaries, doubled letters, words, and index of coincidence come from the written component of the Open American National Corpus 1.0.1. The release contains 6,424 written documents; Decipher counted 11,405,230 annotated word tokens.
The live archive endpoint had an expired TLS certificate when this bundle was built. Decipher recovered the exact official archive through a dated Internet Archive capture record, matched its archived SHA-1, and pinned SHA-256 1a5c37b0227c176d2c5a00f23721e70e7489782e31ba14e58e47f56eff04d24f.
The OANC public use statement permits unrestricted use and redistribution. The 1.0.1 archive also contains acknowledgement wording for unmodified and transparently modified redistribution; it does not expressly classify aggregate frequency tables. Decipher publishes only aggregate counts and retains this discrepancy in the source record.
Language comparison
English, French, German, Italian, Portuguese, and Spanish profiles use the one-million-sentence Wikipedia packages from the Leipzig Corpora Collection, University of Leipzig. Each profile passed an exact one-million-row gate. Leipzig publishes the downloadable corpora under CC BY 4.0 terms; the source text remains subject to Wikimedia licensing.
Prefixes and suffixes
The affix inventory uses the Wikidata Lexeme dump from 29 July 2026, released as CC0. Only explicit “combines lexemes” relationships are joined to OANC word counts. This editor-maintained inventory is useful but incomplete; it is not a complete account of English morphology.
Generation contract
Text is normalized with Unicode NFKD, folded to lowercase Latin A–Z, and counted with one shared token policy. The approved publication minimums are 25 observations for trigrams, 50 for words, and 25 tokens across at least 2 member word types for affixes.
Corpus irregularities are repaired only when token spans provide an unambiguous sentence boundary or malformed bytes are confined to unused feature metadata. Every repair is retained in the reproducibility artifact; malformed structural data stops generation.