This OCR tool reads text from photos, screenshots and scanned PDF pages using Tesseract running entirely in your browser. Nothing is uploaded — the recognition engine downloads once from a CDN and then processes your files locally.
OCR works best on clear, well-lit, high-contrast text. Straighten skewed photos before scanning where possible, and choose the matching language option — mixing languages without selecting the combined option can reduce accuracy.
A very large share of the instructions a government office relies on exist only as scans — older Office Memoranda, circulars photocopied and rescanned several times over, and service records digitised years after they were written. None of it is searchable, and none of it can be quoted without retyping.
This matters more in government work than in most contexts. OCR is reliable on ordinary prose and much less reliable on exactly the things a file turns on: file numbers, dates, rupee amounts, rule and sub-rule numbers, and initials. The characters most often confused are 0 and O, 1 and l, 5 and S, and 8 and B — every one of which can appear in an OM number. Read the extracted text against the scan before you paste it into a noting, and never quote a rule number from OCR output without verifying it.
Straight, high-contrast, flat scans read best. Skewed pages, shadows from phone photography, faint carbon copies, handwriting and text printed over a stamp or seal all reduce accuracy sharply. If the source is a photograph, straighten and clean it in the Image Editor's document-scanner mode first — the improvement is usually substantial.