Word Frequency Counter

Paste text or upload a .txt file and get a ranked word table with counts, share and first position, plus stats, a top-15 chart and CSV export.

Text and files are tokenized in this page; nothing is uploaded, stored, or sent over the network at any point.

The stopword list is English-only — switch it off for other languages.

How It Works

Every time you type, change a control or load a file, the tool re-tokenizes the text with a Unicode-aware regular expression, applies your filters (minimum length, case handling, stopword list), counts the surviving tokens in a map and records each word's first position in the stream. The results feed four places at once: the summary stat cards, the top-15 bar chart on canvas, the ranked table (capped at the first 50 rows for readability) and the CSV/clipboard exports, which always reflect the current filters rather than the raw text.

The tokenizer regex
Words are found with /p{L}+'?p{L}*/gu: one or more Unicode letters of any script, optionally with one apostrophe inside so English contractions ("don't", "l'été") survive as single tokens. Because the class is letters-only, digits, hyphens, underscores and punctuation act as separators — "2026", "e-mail" and "$1,000" contribute nothing whole. The u flag makes the pattern surrogate-pair safe for emoji and rare CJK code points.
Why stopword filtering matters
Zipf's law guarantees the top of any raw English list is occupied by the, of, and, to — words that describe grammar, not topic. For the usual jobs this counter gets (checking whether an article actually covers its promised keywords, auditing page copy, spotting filler), you want the content-word profile, so the ~200-word English stopword list is on by default. Turn it off to see the true token distribution, which is also the honest way to compare two texts of different lengths.
Reading the table
The % column is the word's share of all tokens that passed the current filters (the percentages sum to 100 only if you show every unique word). A 3% share is genuinely high for a content word — English prose rarely gives a topic noun more than 1–2%. The first-position column is the word's index in the token stream, useful for spotting which terms a piece introduces early (a real signal in search-intent analysis).

Frequently Asked Questions

Does it count phrases or word pairs?

No — the analyzer is strictly single-word. Each run of letters (with one optional internal apostrophe, so "don't" stays whole) is one token, and multi-word phrases are never assembled. A bigram or n-gram view needs different statistics, because phrase frequency behaves quite differently from word frequency; this page keeps to the one thing a simple counter can get exactly right.

Are numbers counted as words?

No. The tokenizer keeps only letters (any alphabet) plus an optional embedded apostrophe, so "42", "$5.50" and "2026" are skipped entirely — they would otherwise flood financial or date-heavy text with junk rows. If you need numeric tokens counted, this is the wrong tool for that job.

Why is my top list nothing but "the" and "and"?

Because in any English text the most frequent words are function words — the, and, of, to. That is exactly what the stopword filter (on by default) removes: a built-in list of about 200 common English function words. With it off you see the true raw distribution; with it on you see the words that actually carry the text's subject matter. The list is English-only, so turn it off for other languages.

Can I analyze non-English text?

Yes. The tokenizer is Unicode-aware (the \p{L} letter class covers Latin accents, Greek, Cyrillic, Arabic, Hebrew, and it also splits CJK scripts character-group by character-group since those have no spaces). What is English-only is the stopword list — for other languages, switch the filter off and judge the top words yourself.

Is anything uploaded to a server?

No. Pasted text and uploaded .txt files are read and tokenized entirely by JavaScript in this page — the file never leaves your disk, and there is no network request of any kind. The CSV download and clipboard copy are generated locally too.