Paste text or upload a .txt file and get a ranked word table with counts, share and first position, plus stats, a top-15 chart and CSV export.
Text and files are tokenized in this page; nothing is uploaded, stored, or sent over the network at any point.
Every time you type, change a control or load a file, the tool re-tokenizes the text with a Unicode-aware regular expression, applies your filters (minimum length, case handling, stopword list), counts the surviving tokens in a map and records each word's first position in the stream. The results feed four places at once: the summary stat cards, the top-15 bar chart on canvas, the ranked table (capped at the first 50 rows for readability) and the CSV/clipboard exports, which always reflect the current filters rather than the raw text.
/p{L}+'?p{L}*/gu: one or more Unicode letters of any script, optionally with one apostrophe inside so English contractions ("don't", "l'été") survive as single tokens. Because the class is letters-only, digits, hyphens, underscores and punctuation act as separators — "2026", "e-mail" and "$1,000" contribute nothing whole. The u flag makes the pattern surrogate-pair safe for emoji and rare CJK code points.No — the analyzer is strictly single-word. Each run of letters (with one optional internal apostrophe, so "don't" stays whole) is one token, and multi-word phrases are never assembled. A bigram or n-gram view needs different statistics, because phrase frequency behaves quite differently from word frequency; this page keeps to the one thing a simple counter can get exactly right.
No. The tokenizer keeps only letters (any alphabet) plus an optional embedded apostrophe, so "42", "$5.50" and "2026" are skipped entirely — they would otherwise flood financial or date-heavy text with junk rows. If you need numeric tokens counted, this is the wrong tool for that job.
Because in any English text the most frequent words are function words — the, and, of, to. That is exactly what the stopword filter (on by default) removes: a built-in list of about 200 common English function words. With it off you see the true raw distribution; with it on you see the words that actually carry the text's subject matter. The list is English-only, so turn it off for other languages.
Yes. The tokenizer is Unicode-aware (the \p{L} letter class covers Latin accents, Greek, Cyrillic, Arabic, Hebrew, and it also splits CJK scripts character-group by character-group since those have no spaces). What is English-only is the stopword list — for other languages, switch the filter off and judge the top words yourself.
No. Pasted text and uploaded .txt files are read and tokenized entirely by JavaScript in this page — the file never leaves your disk, and there is no network request of any kind. The CSV download and clipboard copy are generated locally too.
If you just finished with Word Frequency Counter, the natural next steps are Readability Calculator, Mock JSON Data Generator, Word & Character Counter, or browse every tool in Daily Tools.