PDF Text Extractor

Extract the text layer from any PDF, page by page. Check each page, then copy the text or download it as a .txt file. Files up to 20 MB.

Your PDF is read and processed entirely in your browser. Only the pdf.js library is fetched from a CDN — the file itself is never uploaded.

Drop a PDF here

PDF · up to 20 MB · processed locally

Load a PDF to extract its text.

How It Works

A digitally created PDF stores real text — positions, fonts and glyph runs — in a text layer that sits underneath the rendered page image. This tool opens that layer with the open-source pdf.js engine and reads it out page by page. Only the pdf.js library is downloaded, from the jsDelivr CDN; your PDF file is read straight from your disk and processed in a Web Worker on your machine.

The text layer
PDFs generated from Word, Google Docs, code or printers' digital output embed genuine text. Scanned PDFs are just images of pages and have no such layer — that is why extraction from a scan returns nothing. This tool does not run OCR.
Page-by-page extraction
pdf.js walks through every page in order and collects its text items, which carry end-of-line markers. Items are joined into lines, each page becomes a collapsible block so you can check the content, and the full document is concatenated for copying or downloading.
Where does the code run?
On first use the browser fetches pdf.js from jsDelivr (cdn.jsdelivr.net) and caches it for later visits. The PDF bytes themselves never leave your device — they are read via the File API and parsed locally in a Web Worker. No upload, no server, nothing stored.

Frequently Asked Questions

Does my PDF get uploaded to a server?

No. The PDF is read locally from your disk and processed entirely in your browser. The only thing fetched from the internet is the pdf.js library itself, loaded from the jsDelivr CDN on first use — your document's bytes are never sent anywhere.

Why is the pdf.js library loaded from a CDN?

pdf.js is the open-source PDF engine behind this tool. It is fetched from jsDelivr (https://cdn.jsdelivr.net) the first time you load a file, then cached by your browser for subsequent visits. Loading the library from a CDN keeps the tool lightweight; your PDF file itself is never uploaded — only the library code crosses the network.

Why do I get no text from my PDF?

Scanned or image-only PDFs have no text layer — the page is a photograph, not text. This tool extracts the embedded text layer but does not run OCR, so image-only PDFs return nothing. Digitally created PDFs (from Word, Google Docs, etc.) extract cleanly.

What is the maximum file size?

20 MB. Everything runs in your browser's memory, so very large documents can be slow or hit browser memory limits. For bigger files, split them first.

How is the extracted text organized?

Each page becomes its own collapsible block in reading order. Use 'Copy all' to grab the entire document, or 'Download .txt' to save it as a plain-text file named after your PDF. Character, word and page counts are shown above the results.

If a page has no text layer because it is a scan, pull the words out with Image OCR, convert Word documents over first with DOCX to PDF, or check how many pages you are dealing with using the PDF Page Counter.