filetity

Image / scan -> TXT · Local tool

OCR

Read the text out of an image or a scanned PDF, on your own device.

  1. 01 / Your device

    the file stays here

  2. 02 / Browser memory

    the work runs locally

  3. 03 / Your downloads

    the result arrives here

How it works / no cloud

A conversion engine, not an upload form.

filetity runs PP-OCRv6, an open-source text recognition model from the PaddleOCR project, inside your browser on ONNX Runtime compiled to WebAssembly. The model finds each line of text in the picture, reads it character by character, and reports how sure it is about every line. The engine is about 9 MB and downloads once, from this site, the first time you press the button. Your file is never uploaded; there is no server to upload it to.

  • Offline ready
  • Installable app
  • No file limit queue

Straight answers

No accordion. Nothing hidden.

Is it free?
Yes, with no watermark and no page limit.
Does it need an account?
No. There is no account system.
Is my file processed on my device?
Yes. The file is read by your browser and never uploaded. There is no server to send it to.
What formats are supported?
PNG, JPG, WEBP or PDF in. Plain text out, on screen and as a .txt download. A PDF is read page by page.
What are the limits?
Printed text in Latin alphabets and Chinese. It does not read Arabic, Cyrillic, Korean, Japanese kana or Devanagari, and handwriting is hit and miss. Lines the model itself scored low are flagged, though a confident misread can still slip past the flags, so the output is shown for checking rather than downloaded blind.
Does it work on mobile?
Yes, on any modern mobile browser. Large files are limited by the memory the phone gives the browser.
Does it work offline?
Yes, after the first use. The engine is about 9 MB and downloads once, then your browser keeps it, so later visits work with no network connection.
How accurate is it?
On a clean screenshot or scan of printed text, close to perfect: our own rendered-page test came back 232 of 233 characters right, and the one miss was a space. On a deliberately noisy, tilted, low-quality photo of the same text it read 95 percent of the characters, and the mistakes it did make were confident ones the flags did not catch. That is the honest shape of OCR: very good, never guaranteed, so check numbers and names against the original.
Can it read a scanned PDF?
Yes. Each page is drawn to an image in your browser and the text is read out of the picture, which is exactly what a scan needs. A PDF that already has a text layer is better served by the PDF to Text tool, which copies the real characters instead of recognising them.
Why is the download so big?
The engine is a neural network, about 9 MB with its runtime, and reading letters out of pixels genuinely needs one. It downloads the first time you use the tool, from this site, and your browser caches it, so it is a one-time cost. Nobody who visits any other page on this site downloads a byte of it.
Does it read handwriting?
Sometimes, badly. The model is trained on printed text; neat block capitals often come out usable, cursive rarely does. Anything it does read from handwriting deserves a closer look than usual, and the confidence flags will usually tell you that themselves.

Keep working locally

4 available · more coming