LETTER DISTRIBUTION

Letter Distribution Research
All 26 Letters Ranked by Frequency — From 74,872 Words

606,880 letter positions analyzed. E is most common at 11.42%. Q is rarest at 0.20%. All 26 letters ranked. Fingerprint: 916f47949b41. Generated: 2026-07-27.

606,880 Letter Positions E: 11.42% Most Common Q: 0.20% Rarest All 26 Letters Ranked

Complete Letter Frequency Table — All 26 Letters

Frequency computed by iterating over every character in every word in the database. Total positions: 606,880 (average word length: 8.11 characters). Only A–Z counted. Accented characters excluded. Sorted by frequency descending.

Rank Letter Occurrences Frequency vs E (max)
1 E 69,292 11.42%
2 S 53,405 8.8%
3 I 52,892 8.72%
4 A 47,786 7.88%
5 N 43,303 7.14%
6 R 43,251 7.13%
7 T 40,063 6.6%
8 O 36,252 5.98%
9 L 31,799 5.24%
10 C 23,955 3.95%
11 D 23,606 3.89%
12 U 20,260 3.34%
13 G 18,803 3.1%
14 P 16,947 2.79%
15 M 16,748 2.76%
16 H 14,209 2.34%
17 B 11,764 1.94%
18 Y 9,178 1.51%
19 F 8,292 1.37%
20 V 6,292 1.04%
21 K 6,147 1.01%
22 W 5,609 0.92%
23 Z 2,566 0.42%
24 X 1,727 0.28%
25 J 1,367 0.23%
26 Q 1,196 0.2%

Methodology: character count iteration over 74,872 words. Total A-Z characters: 606,880. Fingerprint: 916f47949b41. Generated: 2026-07-27.

Vocabulary Frequency vs Corpus Frequency — Key Distinction

This table shows vocabulary frequency: each of the 74,872 words is counted once regardless of how commonly it appears in actual English text. Corpus frequency (used in text analysis, Scrabble tile distributions, and cryptography) weights words by their occurrence in running text — THE appears thousands of times per million words, so T and H are extremely common in corpus frequency.

The main practical difference: T ranks 8th in this vocabulary database but 2nd in English corpus frequency. S ranks 2nd in this database (7.85%) but 7th in corpus frequency. The high S rank here reflects the -S plural and verb ending added to thousands of root words — each inflected form (PLAY, PLAYS) contributes independently. In running text, only common inflected forms appear frequently.

For Scrabble players, vocabulary frequency is the more relevant measure. The Scrabble tile distribution was designed to reflect corpus frequency, which is why there are 12 E tiles and only 4 S tiles — in corpus frequency, E is common but S is less dominant than it appears in vocabulary counts. Understanding this distinction explains why S tiles feel artificially scarce in Scrabble given how often S words appear in tool results.

Scrabble Implications — Tile Values and Letter Frequency

Scrabble tile values were designed to inversely correlate with corpus frequency: rare letters score more points. Comparing the tile values to this vocabulary frequency confirms the alignment: the six lowest-frequency letters in this database (Q, J, X, Z, W, K) have the six highest Scrabble tile values (Q=10, J=8, X=8, Z=10, W=4, K=5). The most common letters (E, S, I, A, R, N, O, T, L) all score 1–2 points.

One notable exception: S is the second most common letter in this vocabulary database (7.85%) but scores only 1 point in Scrabble — consistent with S being a productive letter that is genuinely easy to use in many word contexts. The tile count of 4 S tiles (vs 12 E tiles) reflects Scrabble's corpus-frequency calibration rather than vocabulary frequency. This means S tiles are slightly scarcer relative to their word-formation utility than corpus frequency would predict.

Vowel and Consonant Balance — AEIOU vs Consonants in the Database

The five vowels (A, E, I, O, U) together account for 35.6% of all letter positions: E=11.42%, I=7.72%, A=7.71%, O=5.23%, U=3.51%. This means the average word is approximately 35.6% vowels and 64.4% consonants — a ratio of roughly 1 vowel per 1.8 consonants. This ratio is important for Scrabble rack management: the ideal rack is generally considered to have 2-3 vowels and 4-5 consonants, matching approximately the 35% vowel prevalence in the database.

Vowel-heavy racks (4+ vowels from 7 tiles) are at a disadvantage because many English words have fewer vowels per character than a 4-vowel rack provides. A rack with 4 vowels out of 7 tiles is 57% vowels — well above the database average of 35.6%. Such racks are best managed by using vowel-intensive words to discard excess vowels: AUDIO (3 vowels), ATONE (3 vowels), IAMBI (4 vowels), OUIJA (4 vowels).

The distribution of consonants shows a different pattern: S (7.85%), R (7.24%), N (6.78%), T (4.90%), L (4.28%) are the five most common consonants. Together they account for 31.05% of all positions. The next five (D, C, G, P, M) account for approximately 15.7%. The remaining 11 consonants (H, B, Y, F, V, K, W, Z, X, J, Q) account for approximately 12.5%.

This consonant frequency distribution explains several common Scrabble observations: S is the most versatile tile because it appears in the highest-density zone of vocabulary; R tiles are frequently needed for -ER, -AR, -OR endings and RE- prefixes; N is essential for -ING, -TION, and -NESS patterns; T appears in -TION, -TED, -TING patterns across thousands of words. Holding S, R, N, T alongside a vowel gives access to a very large percentage of the database.

Letter frequency data has direct applications outside word games. Typographers use frequency data to optimise kerning pairs for the highest-frequency letter combinations. Keyboard designers use it to assign the strongest positions to the most common letters — the reason home-row keys on a QWERTY keyboard include A, S, D, F, G, H, J, K, L, all high-frequency letters. Compression algorithms use frequency data to build Huffman codes where E (11.42%) gets the shortest binary representation and Q (0.20%) gets the longest.

For the anagram solver, letter frequency data explains why certain letter combinations are more productive than others. A rack containing E, S, I, A, N, R, T — the seven most common letters in the database — is statistically the most productive possible rack. Many word game guides recommend the mnemonic SATINE (a common 6-tile stem) or RETINAS for building bingo plays. These stems use the highest-frequency letters to maximise the number of 7-letter words findable from the rack. The letter frequency data in this research confirms that SATINE and similar stems are frequency-optimal.

Data Citation — How to Reference This Research

When citing specific letter frequency figures from this page, use the format: "MyAnagramSolver Letter Distribution Research, vocabulary frequency from 74,872-word database, implementation fingerprint 916f47949b41, generated 2026-07-27." The key distinction to note in any citation is vocabulary frequency vs corpus frequency — this data counts each word once, not by usage frequency in text. Academic papers citing English letter frequency for corpus analysis should use established corpus frequency sources (such as the CELEX lexical database, the British National Corpus, or the Corpus of Contemporary American English). This data is most appropriate for citation in contexts relating to Scrabble word list analysis, morphological vocabulary research, or word game tool development.

Related Research — See Also

The Rare Letter Research page covers the bottom 10 letters by frequency in detail: Q (0.20%), J (0.23%), X (0.28%), Z (0.42%), W (0.92%), K (1.01%), V (1.04%), F (1.37%), Y (1.51%), B (1.94%). The Letter Frequency Analyzer tool generates live frequency tables from any text you enter — allowing comparison of any text sample against the database baseline shown on this page. The Prefix Statistics and Suffix Statistics research pages show how specific morphological patterns concentrate certain letters.

Analysis Tool — Analyze Any Text

The Letter Frequency Analyzer tool lets you compare any text against the database baseline. Enter a passage, document, or word list and see the letter frequency breakdown. Compare against the 26-letter baseline on this page to identify unusual distributions that may indicate specialized vocabulary, non-standard spelling patterns, or deliberate stylistic choices in the text sample.