Got a scan or a photo of a document and need the text out of it? Run free OCR that works right in your browser. The recognition engine downloads once, then reads your pages locally — the file itself is never uploaded.
A text-based PDF already contains real, selectable text — for those, the faster Extract text tool is all you need. A scanned PDF is really a picture of a page, so selecting or copying gives you nothing. OCR (optical character recognition) looks at the picture and works out the letters. If Extract text comes back empty, your PDF is scanned — use OCR.
No. The engine downloads to your browser; your file is read on your device and never uploaded. See our privacy policy.
English, Greek, Russian, French, German, Spanish and Italian.
Your PDF is scanned (a picture of text). Run OCR to turn the picture into text.
When you run OCR on meldpdf, a recognition engine called Tesseract.js loads in your browser. It is a WebAssembly (WASM) port of the open-source Tesseract OCR engine, the same engine that powers Google's document scanning. The first time you use it, a language model file (typically 2 to 12 MB depending on the language) downloads and is cached by your browser for future use.
The engine works in stages. First, it analyses the image to find blocks of text, separating them from pictures, lines and whitespace. Then it segments each block into individual lines and characters. Finally, it runs pattern recognition against the language model — comparing the shapes it found to known letter forms — and outputs its best guess for every character. Modern Tesseract also uses a neural-network recogniser (LSTM) that reads sequences of characters in context, which improves accuracy on real-world text considerably compared to older shape-matching alone.
All of this runs on your device. The scanned image is never sent to a server, and the recognised text exists only in your browser until you download it.
OCR is remarkably good at reading clean, printed text — typed letters on a white background at 300 DPI or higher will typically produce near-perfect results. But it has limits worth knowing about:
In short, OCR is excellent for extracting body text from scanned documents, receipts, letters and book pages. It is not a substitute for manual transcription of complex or handwritten material.
Once you have the recognised text, there are several useful next steps depending on what you need:
It depends on the number of pages and your device's processing power. A single page typically takes a few seconds on a modern computer or phone. A ten-page document might take 20 to 40 seconds. The first run is slower because the language model needs to download; subsequent runs use the cached model and start faster.
Yes. Add a multi-page PDF and the engine processes each page in sequence. Longer documents take proportionally longer, but there is no hard page limit — it depends on your device's memory and patience. For very large documents (50+ pages), you may want to split them first using the Split PDF tool and OCR in batches.
Last updated: 3 September 2026.