Cadmeo

Word Length Analyser

Longest word

disproportionate: 16 letters

15
Words
14
Distinct words
6.13
Average length
5
Median length
Ten longest
  1. disproportionate (16)
  2. extraordinarily (15)
  3. hippopotamus (12)
  4. watches (7)
  5. brown (5)
  6. jumps (5)
  7. quick (5)
  8. while (5)
  9. lazy (4)
  10. over (4)
Ten shortest
  1. an (2)
  2. The (3)
  3. fox (3)
  4. dog (3)
  5. over (4)
  6. lazy (4)
  7. while (5)
  8. quick (5)
  9. jumps (5)
  10. brown (5)
Length distribution
21
34
42
54
71
121
151
161

The analyser finds the ten longest and ten shortest distinct words in a text and reports the average and median word length with a distribution chart. Average word length is one of the two inputs to almost every readability formula, so it is a quick way to see whether writing is denser than intended.

How it works

A word is a run of letters or digits, with apostrophes and internal hyphens kept as part of the word. That means "don't" is one word of five characters and "well-known" is one word of ten, rather than two words each.

  • The longest and shortest lists show distinct words, so a short word repeated fifty times appears once rather than filling the list.
  • The average counts every occurrence, so common short words pull it down, which is the honest measure of how the text reads.
  • The median is shown alongside because one very long technical term can lift an average without changing how the text feels.
  • The distribution chart shows how many words of each length there are. English prose usually peaks at three to five characters.

Average word length feeds directly into readability scoring. The Flesch Reading Ease formula uses syllables rather than characters, but the two track each other closely enough that a rising average here means denser text.

Examples

A sentence with one long word

Text

The quick brown fox jumps over the lazy dog while an extraordinarily disproportionate hippopotamus watches.

Result

Longest: disproportionate (16) · average 6.13 · median 5

The gap between the mean of 6.13 and the median of 5 is the signature of a few long words in otherwise plain text.

A hyphenated word

Text

A well-known problem

Result

well-known counted as one word of 10 characters

Splitting on the hyphen would give two words of four and five, and a different average. Keeping it whole matches how a reader treats it.

Distinct versus total

Text

"the" repeated 40 times plus 10 other words

Result

the appears once in the shortest list; the word count is 50

The lists are distinct words, so repetition does not crowd them out. Repetition is what the keyword frequency tool measures.

Frequently asked questions

How is a word defined here?

A run of letters or digits, keeping apostrophes and internal hyphens. So "don't" is one five-character word and "well-known" is one ten-character word. Punctuation around a word is stripped before it is measured.

Why is the average different from the median?

Because a handful of long words pull the average up without moving the median. A large gap between the two means the text is mostly short words with occasional technical terms. Worth knowing, because it reads more easily than the average alone suggests.

What is a normal average word length?

English prose typically runs about 4.5 to 5.5 characters. Technical and academic writing runs higher, often past 6. There is no target to hit. It is a signal, and a rising number simply means denser text.

Why does the shortest list contain single letters?

Because "a" and "I" are genuine one-letter words, and initials or list markers in your text count as words too. If the single letters are artefacts rather than words, clean the text first.

How does this differ from the word counter?

The word counter gives totals, words, characters, sentences, reading time. This looks at the shape of the vocabulary: which words are long, which are short, and how lengths are distributed.