โ† Lexical Diversity Explorer

METHOD 1.0.0 ยท 30 AUGUST 2026

Measure repetition in context.

TOKENISEWINDOWMAP

This is an editing lens, not a quality grade.

Repetition can make prose monotonous, expose template residue or reveal an overused keyword. It can also be necessary: product names, legal terms, technical concepts and deliberate rhetorical parallelism should often stay consistent.

The explorer therefore provides a profile and evidence map instead of declaring that high diversity is always good.

What counts as a word.

Text is lowercased and split into Unicode letter-or-number tokens. Internal apostrophes are retained. Punctuation and capitalisation do not create new word types. Unique-word counts include function words; the concentrated-term list excludes a disclosed English stop-word set and tokens shorter than three characters.

MATTR corrects the length problem.

Raw type-token ratio is unique tokens divided by all tokens. It naturally falls as a document gets longer, so it is displayed only as a descriptive figure.

The primary signal is moving-average type-token ratio (MATTR): the unique-token ratio is calculated for every overlapping 50-word window, then averaged. Text shorter than 50 words is not analysed. Text between 50 and 99 words is valid but less stable than a longer sample.

Terms, phrases and openings.

Content terms
Repeated non-stopwords, ordered by count.
Phrase echoes
Repeated two- and three-token sequences containing at least two content words.
Sentence openings
The first three tokens of each sentence, listed when repeated.
Echo map
The text is divided into eight sequential segments and dominant-term frequency is shown in each.

Descriptive thresholds.

Broad, lightly repeated requires MATTR of at least 0.82 and the three most frequent content terms below 9% of content tokens. Balanced and consistent requires MATTR of at least 0.68 and top-three dominance below 16%. Other samples are labelled concentrated vocabulary.

These are editorial heuristics, not linguistic norms, ranking factors or AI-detection signals. Genre, audience, language, subject and document size all affect the numbers.

Your text stays in the browser.

Analysis runs locally in page JavaScript. The tool does not submit or save pasted text. Standard site infrastructure may still log the page visit, but not the textarea contents through this tool.