pdfcosmos

OCR PDF

Make a scanned PDF searchable (or pull out its text) with OCR that runs in your browser.

Free, fast, and privacy-first.

OCR a PDF

or drag & drop files here

PDF files onlyone file at a timeup to 100 MB each

OCR a PDF: make scanned documents searchable, privately

A scanned PDF is just pictures of pages. You can't select or search the words. OCR (optical character recognition) reads those pictures and adds a real text layer. pdfcosmos does it entirely in your browser, so your document never leaves your device. The OCR engine and language data (English, Spanish, German, French, Hindi) are served from this site, not a third-party cloud.

Choose a searchable PDF (the same pages, but now you can select, copy, and search the text) or a plain .txt file of the recognised words. Recognition is English-only for now and runs page by page with a limit per run, because OCR is slow, so clean, high-resolution scans give the best results. It also complements PDF to Text and PDF to Word, which can only read PDFs that already have a text layer.

  • Searchable PDF or text

    Add an invisible text layer to your scan, or export the recognised words as .txt.

  • 100% in your browser

    The OCR engine and language data are self-hosted. Nothing is uploaded, nothing fetched from a third party.

  • Pick the pages

    All pages, the first or last, or a custom range, with a per-run page limit so long files stay responsive.

  • Honest about accuracy

    OCR isn't perfect. Clean, high-resolution scans read best. You get a strong starting point.

How to OCR a PDF

  1. Add your PDF

    Click “Add PDF”, or drag a single scanned or image-only file into the upload area.

  2. Pick the output

    A searchable PDF (looks the same, but the text is selectable) or a plain .txt file, and which pages.

  3. Run OCR and download

    Click “Run OCR”. Recognition happens in your browser and can take a little while on long documents.

Frequently asked questions

Is my PDF uploaded to a server?
No. Optical character recognition runs entirely in your browser using a self-hosted Tesseract engine. Your file never leaves your device. You can confirm it by opening your browser's Network tab while it runs: no file data is sent, and even the OCR engine and language data load from this site, not a third party.
What is a searchable PDF?
It looks identical to your scan, but an invisible text layer is placed over each page, so you (and search tools) can select, copy, and find the words. It's the standard way to make scanned documents usable without changing how they look.
Which languages are supported?
English, Spanish, German, French, and Hindi. Pick the document's main language before running OCR for the best accuracy. Each language's trained data is served from this site (not a third party) and loads on first use.
Why is it slower than other tools?
OCR is genuinely heavy work, and to keep the rest of the site (including ads and analytics) working we run the engine single-threaded. A few pages are quick. A long document takes a while. It shows progress page by page, and there's a page limit per run so a huge file can't hang your browser.
The recognised text has mistakes. Why?
Accuracy depends on the scan: clean, straight, high-resolution pages read well. Faint, skewed, handwritten, or low-resolution pages read worse. OCR is never perfect, so treat the output as a strong starting point, not a flawless copy.
Is OCR PDF free?
Yes, it's completely free to use, with no sign-up and nothing to install.