Klarfile

OCR a scanned PDF without uploading it

A scan is an image: nothing to search, nothing to select. So we read each page and lay the recognised text OVER it, invisibly. The page looks exactly as it did; what changes is that you can finally search it for a name.

What this tool does

What it does not do: Text recognition gets things wrong, and it does so silently: a 1 becomes an l, a proper noun loses its accent. This is not a certified transcription, and it should not be used for anything you will not read back. Handwriting is not read at all. Columns are not untangled and tables are not rebuilt as tables. Words a standard font cannot write — Greek, Cyrillic, the diacritics of Central Europe — are not placed at all rather than placed by halves, and the screen counts them. Finally, the engine and the language model weigh a few megabytes, downloaded from this site on the first read and then kept by your browser: the first page takes longer than the rest.

Nothing is uploaded. Cut your network connection once this page has loaded: the tool keeps working. How to check it yourself.

The other tools