TL;DR
- Scanned PDFs and documents are image files — OCR must convert the image to machine-readable text before any translation can happen
- OCR accuracy (and therefore translation quality) depends heavily on scan quality: 300 DPI minimum for reliable results per university archival standards
- Typical accuracy by source type: screenshots 95%+, scanned documents 85–95%, phone photos 70–90%
- Two approaches: two-step (OCR separately, then translate) or one-step pipeline (both in one tool)
- One-step pipelines preserve layout better because structural information isn't lost at the handoff between tools
Scanned documents aren't text files. They're images. And that one fact explains why most translation tools fail on them — and what you need to do differently.
Quick answer: A scanned PDF is an image file - there's no text for a translator to read directly. OCR must run first to extract the text, then translation happens. Tools that do both in one pipeline (like Smallpdf or Doctranslate.io ) are faster than the two-step approach. The section below explains how to pick the right one.
How We Evaluated
We assessed each tool on three questions:
- Does OCR run automatically, or does the user need to pre-process the file?
- How well does the output preserve the original layout after both OCR and translation?
- What happens with scanned files in non-Latin scripts (CJK, Arabic)?
Tools were tested with a scanned invoice, a scanned two-column academic paper, and a phone photo of a printed document.
Why You Can't Translate a Scanned PDF Directly
Upload a scanned PDF to Google Translate and you'll get one of two results: nothing, or garbled output.
The reason is straightforward. A scanned PDF is a photograph of a page — a grid of pixels, not a text file. When a translation tool receives it, there's no text to read. This is fundamentally different from a native PDF (one created digitally in Word, InDesign, or any software that exports to PDF), which has actual text data embedded in its structure.
The workflow that actually works requires two steps:
- OCR: Convert the image of text into machine-readable characters
- Translation: Pass that extracted text through a translation engine
The quality of step one determines everything about step two. A translation engine can only work with what OCR gives it.
Note: If your PDF is native (not scanned) but gets rejected because of its size, that's a different problem - see our guide on how to translate a PDF larger than 10MB .
What Determines OCR Accuracy
Professional OCR systems targeting high-quality printed documents can reach 99%+ character accuracy. That number drops with poor scan quality:
- Standard business documents, photocopies: 95–98%
- Older documents, faded text, complex layouts: 85–94%
- Handwritten text, severely degraded documents: below 85%
Scan resolution (DPI). The recommended minimum for OCR is 300 DPI, per the University of Illinois archival standards. Below 200 DPI, text starts to blur and OCR misreads increase sharply. For small fonts (under 10pt), 400–600 DPI is recommended.
Image contrast. High contrast between text and background is essential. Faded ink, yellowed paper, or low-contrast printing reduces OCR accuracy.
Page alignment. Even a 1–2 degree tilt reduces OCR accuracy. Most modern OCR tools include automatic deskewing, but heavily skewed documents may need manual correction.
What OCR cannot handle reliably: Handwritten text (requires specialized HTR technology), very small fonts (under 6pt), heavily stylized fonts, documents with dense graphical overlays.
Method 1: Two-Step Workflow (OCR First, Then Translate)
Use this when you're working with separate tools for each step.
Step 1: Run OCR
Google Drive (free):
- Upload your scanned PDF to Google Drive
- Right-click → Open with → Google Docs
- Google Docs runs OCR and opens the extracted text
- Copy the text
Adobe Acrobat:
- Open the scanned PDF
- Go to Tools → Scan & OCR → Recognize Text
- Select language and run OCR
- Save as searchable PDF or export to Word
Smallpdf OCR (free, browser-based):
- Upload your scanned PDF
- Smallpdf converts the scanned file to a searchable PDF
- Download, then translate separately
Step 2: Translate the OCR output
Once you have machine-readable text, use any translation tool:
- Google Translate — free, 249 languages
- DeepL — better quality for European language pairs
- Doctranslate.io Document Translation — layout preservation
Limitation of the two-step approach: OCR tools that export to plain text strip the original layout. The translated output is readable but not usable for professional purposes without reformatting.
Method 2: One-Step OCR + Translation Pipeline (Recommended)
Better approach: a tool that runs OCR and translation in the same pipeline, preserving document structure throughout.
How to translate a scanned PDF with Doctranslate.io :
- Go to doctranslate.io/translation/document
- Upload your scanned PDF — files up to 300MB are supported, so multi-hundred-page scans don't need splitting
- Select source and target language from 100+ supported pairs
- Turn on Scanned PDF (OCR) if it isn't detected automatically, then click Translate
- Download the translated PDF with the original layout preserved
For scanned images (JPG, PNG) rather than PDFs, Doctranslate.io Visual Translation handles those directly.
Tool Comparison
| Tool | OCR built-in? | Translation built-in? | Layout preserved after translation? | Free option? |
|---|---|---|---|---|
| Google Translate | No | Yes (249 languages) | Poor | Yes |
| Google Drive / Docs | Yes (basic) | No (separate step) | No | Yes |
| Smallpdf | Yes | Yes (separate tool) | Good for standard layouts | Yes (2 tasks/day) |
| Adobe Acrobat | Yes (advanced) | No (separate step) | Yes - exports editable PDF | No ($19.99+/mo) |
| iLovePDF | Yes (120+ source languages) | Yes (25 target languages) | Good | Yes (1 task/day) |
| Doctranslate.io | Yes (auto) | Yes (100+ languages) | Strong - coordinate-level | Free trial (15 credits) |
The Tools in Detail
1. Google Translate
- What it does: Translation only, 249 languages, free — no OCR at all
- Limits: A scanned PDF won't work; layout is lost even on native PDFs
- Best for: Quick gist translations after running OCR elsewhere
2. Google Drive / Docs
- What it does: Free OCR — upload, open with Google Docs, text is extracted ( steps above )
- Limits: All layout discarded (tables → flat text); translation is a separate step
- Best for: Simple text-only pages on zero budget
3. Smallpdf
- What it does: Browser-based OCR to searchable PDF, plus a separate translate tool
- Limits: Free tier is 2 tasks/day — fine for one-offs, not batch work
- Best for: Occasional scanned files with standard layouts
4. Adobe Acrobat
- What it does: The most mature OCR engine here — fine-grained language settings, editable-PDF export
- Limits: No built-in translation; from $19.99/mo
- Best for: Difficult scans when you already have an Acrobat license
5. iLovePDF
- What it does: OCR (120+ source languages) + translation (25 target languages) in one flow
- Limits: 1 task/day free; the 25-language target list may not cover your pair
- Best for: Occasional users with common language pairs
6. Doctranslate.io
- What it does: OCR + translation in one pipeline, coordinate-level layout preservation, 100+ languages, files up to 300MB
- Limits: Free trial is 15 credits, then paid
- Best for: Business documents where the output must look like the original — invoices, contracts, reports
Tips to Improve OCR and Translation Quality
- Before scanning: Scan at 300 DPI minimum. Most scanner apps default to 150–200 DPI — check the setting. Maximize contrast. Scan pages flat.
- When using a phone camera: Use a scanner app (Adobe Scan, Microsoft Lens, Google PhotoScan) rather than the native camera. These apps apply automatic deskewing and contrast enhancement.
- After OCR, before translation: Review the OCR output for obvious errors before translating. A misread character in the source text becomes a mistranslation in the output.
- Specify the source language if your tool supports it. Auto-detection works for common single-language documents. Mixed-language documents benefit from manual language selection.
When One-Step Translation Isn't Enough
- Handwritten documents. OCR is not designed for handwriting. For handwritten legal records or correspondence, professional transcription before translation is more reliable.
- Documents with heavy graphical content. OCR translates the text around graphics but cannot interpret the graphics themselves.
- Certified or sworn translations. For immigration documents, court submissions, or official certificates, automated OCR + AI translation is insufficient. Most jurisdictions require human translators for certified work.
What to Do
If your scanned document is clean and high-resolution, most tools with built-in OCR will produce usable results. Start with a free tier and test on your actual document before committing.
If your document is older, lower quality, or has complex formatting, the OCR quality and layout preservation matter more. A one-step pipeline tool like Doctranslate.io saves significant reformatting time compared to a two-step approach.
Discussion
No comments yet