OCR Reader & Image Text Extractor

Extract text from images or scanned PDF pages. Processing happens in-browser with OCR; files are never uploaded.
OCR
Image Text Extractor
PNG, JPG, WebP screenshots and scans
PDF
PDF OCR Reader
Renders pages locally, then OCRs them
Extracted Text
Open Notes →
Tesseract OCR runs in the browser. The library may be fetched from CDN, but your selected files and extracted text do not upload.

Pull text out of any image or scanned document

This OCR tool reads text from photos, screenshots and scanned PDF pages using Tesseract running entirely in your browser. Nothing is uploaded — the recognition engine downloads once from a CDN and then processes your files locally.

What it's good for

Tips for better accuracy

OCR works best on clear, well-lit, high-contrast text. Straighten skewed photos before scanning where possible, and choose the matching language option — mixing languages without selecting the combined option can reduce accuracy.

How accurate is browser-based OCR?
Accuracy depends on image quality — clean, high-resolution scans of printed text perform very well, while handwriting or low-resolution photos are less reliable.
Can I OCR multiple PDF pages at once?
Yes — enter a page range like 1-3,5 in the PDF OCR field, or leave it blank to process every page.
Does this support languages other than English and Hindi?
This tool currently ships with English and Hindi recognition. Let us know if you need another language and we'll consider adding it.
Is there a size limit for images or scanned PDFs?
There's no fixed server-side limit since recognition runs in your browser — very large images or long page ranges simply take longer depending on your device's processing power.

Using this in a Central Government office

A very large share of the instructions a government office relies on exist only as scans — older Office Memoranda, circulars photocopied and rescanned several times over, and service records digitised years after they were written. None of it is searchable, and none of it can be quoted without retyping.

Where this comes up

Always check the output against the original

This matters more in government work than in most contexts. OCR is reliable on ordinary prose and much less reliable on exactly the things a file turns on: file numbers, dates, rupee amounts, rule and sub-rule numbers, and initials. The characters most often confused are 0 and O, 1 and l, 5 and S, and 8 and B — every one of which can appear in an OM number. Read the extracted text against the scan before you paste it into a noting, and never quote a rule number from OCR output without verifying it.

What affects accuracy

Straight, high-contrast, flat scans read best. Skewed pages, shadows from phone photography, faint carbon copies, handwriting and text printed over a stamp or seal all reduce accuracy sharply. If the source is a photograph, straighten and clean it in the Image Editor's document-scanner mode first — the improvement is usually substantial.