Image OCR — Extract Text from Photos & Screenshots

Pull text out of photos and screenshots with a WASM OCR engine that runs entirely in your browser — no uploads, with English, Spanish, Portuguese and Indonesian models.

Your image is processed entirely in your browser by Tesseract.js — it is never uploaded. Only the public language model file is downloaded from a CDN on first use.

Extract text from an image

Drop an image here or click to browse — PNG, JPG or WebP · processed locally

First run downloads the language model (a few MB); afterwards it is cached.
Image preview
Recognized text

How It Works

Recognition is done by Tesseract, an open-source OCR engine compiled to WebAssembly. The image, the text it extracts and every intermediate buffer stay inside your browser tab — there is no upload endpoint, no account and no request that carries your pixels.

A WASM engine, fully in-browser
Tesseract.js runs the same C++ engine native apps use, compiled to WebAssembly and executed in a background worker, so the page stays responsive while text is recognised. The whole pipeline — layout analysis, word segmentation, LSTM character classification — happens on your device.
First run downloads the language model
Each language needs its traineddata file (roughly 2–15 MB, English about 10 MB). It is fetched once from a public CDN and then cached by your browser, so repeat visits and, within the same tab session, fully offline use work the same way. Switching language triggers one equivalent download for that language.
Screenshots vs photos
Screenshots are the easy case: sharp glyphs on a flat background, near-perfect accuracy out of the box. Photos of documents demand more — even lighting, minimal skew, text filling the frame and a steady camera. Small fonts under ~10 px, curved pages, shadows and motion blur are what degrade results, so crop and straighten before expecting magic.
Not for handwriting
The models are trained on printed type. Cursive, personal handwriting and stylised lettering are outside what it can read reliably — for those, printed-first workflows or manual transcription remain the honest options.

Frequently Asked Questions

Can it read handwriting?

No. Tesseract is trained on printed type — books, documents, UI screenshots and signs. Cursive and messy handwriting read poorly or not at all, so don't expect it to digitise old letters or doctor's notes. Printed text at a reasonable size is where it shines.

Does it read PDFs?

This page handles image files only — PNG, JPG and WebP. For a PDF, use its sibling tool, the PDF Text Extractor, which reads the embedded text layer of a document page by page without any OCR at all.

Where does the language model come from, and does it work offline afterwards?

On first use the browser downloads a small trained model for the selected language (a few MB) from the Tesseract CDN, along with its WASM engine. Your browser caches these files, so subsequent visits — and even an offline tab after one load — recognise text without fetching anything.

How accurate is it, and how do I get the best results?

Screenshots are near-perfect: they are pixel-sharp text on a flat background. Photos depend on capture quality — use even lighting, hold the page flat and square to the camera, and fill the frame with the text. Blurry, skewed or very small type is the main cause of mistakes, so crop tightly to the region you want.

Is my image uploaded to a server?

No. The OCR runs 100% in your browser with WebAssembly — the image itself is never sent anywhere. The only network traffic is the public, pre-built model file that every visitor downloads from the same CDN URL once.

Need text out of a PDF rather than an image? The sibling PDF Text Extractor reads a document's embedded text layer page by page — no OCR needed.