Scanned documents aren't text files. They're images. And that one fact explains why most translation tools fail on them — and what you need to do differently.

Quick answer: Scanned PDFs are image files — there's no text for a translator to read directly. OCR must run first to extract the text, then translation happens. Tools that do both in one pipeline (like Smallpdf or Doctranslate.io ) are faster than the two-step approach. The section below explains how to pick the right one.

How We Evaluated

We assessed each tool on three questions:

  1. Does OCR run automatically, or does the user need to pre-process the file?
  2. How well does the output preserve the original layout after both OCR and translation?
  3. What happens with scanned files in non-Latin scripts (CJK, Arabic)?

Tools were tested with a scanned invoice, a scanned two-column academic paper, and a phone photo of a printed document.

Why You Can't Translate a Scanned Document Directly

Upload a scanned PDF to Google Translate and you'll get one of two results: nothing, or garbled output.

The reason is straightforward. A scanned document is a photograph of a page — a grid of pixels, not a text file. When a translation tool receives it, there's no text to read. This is fundamentally different from a native PDF (one created digitally in Word, InDesign, or any software that exports to PDF), which has actual text data embedded in its structure.

The workflow that actually works requires two steps:

  1. OCR: Convert the image of text into machine-readable characters
  2. Translation: Pass that extracted text through a translation engine

The quality of step one determines everything about step two. A translation engine can only work with what OCR gives it.

What Determines OCR Accuracy

Professional OCR systems targeting high-quality printed documents can reach 99%+ character accuracy . That number drops with poor scan quality:

  • Standard business documents, photocopies: 95–98%
  • Older documents, faded text, complex layouts: 85–94%
  • Handwritten text, severely degraded documents: below 85%

Scan resolution (DPI). The recommended minimum for OCR is 300 DPI , per the University of Illinois archival standards. Below 200 DPI, text starts to blur and OCR misreads increase sharply. For small fonts (under 10pt), 400–600 DPI is recommended.

Image contrast. High contrast between text and background is essential. Faded ink, yellowed paper, or low-contrast printing reduces OCR accuracy.

Page alignment. Even a 1–2 degree tilt reduces OCR accuracy. Most modern OCR tools include automatic deskewing, but heavily skewed documents may need manual correction.

What OCR cannot handle reliably: Handwritten text (requires specialized HTR technology), very small fonts (under 6pt), heavily stylized fonts, documents with dense graphical overlays.

Method 1: Two-Step Workflow (OCR First, Then Translate)

Use this when you're working with separate tools for each step.

Step 1: Run OCR

Google Drive (free):

  1. Upload your scanned PDF to Google Drive
  2. Right-click → Open withGoogle Docs
  3. Google Docs runs OCR and opens the extracted text
  4. Copy the text

Adobe Acrobat:

  1. Open the scanned PDF
  2. Go to ToolsScan & OCRRecognize Text
  3. Select language and run OCR
  4. Save as searchable PDF or export to Word

Smallpdf OCR (free, browser-based):

  1. Upload your scanned PDF
  2. Smallpdf converts the scanned file to a searchable PDF
  3. Download, then translate separately

Step 2: Translate the OCR output

Once you have machine-readable text, use any translation tool:

Limitation of the two-step approach: OCR tools that export to plain text strip the original layout. The translated output is readable but not usable for professional purposes without reformatting.

Method 2: One-Step OCR + Translation Pipeline (Recommended)

Better approach: a tool that runs OCR and translation in the same pipeline, preserving document structure throughout.

How to translate a scanned document with Doctranslate.io :

  1. Go to doctranslate.io/translation/document
  2. Upload your scanned PDF
  3. Select source and target language from 100+ supported pairs
  4. Click Translate — OCR runs automatically
  5. Download the translated PDF with the original layout preserved

For scanned images (JPG, PNG) rather than PDFs, Doctranslate.io Visual Translation handles those directly.

Tool Comparison

ToolOCR built-in?Translation built-in?Layout preserved after translation?Free option?
Google TranslateNoYes (249 languages)PoorYes
Google Drive / DocsYes (basic)No (separate step)NoYes
SmallpdfYesYes (separate tool)Good for standard layoutsYes (2 tasks/day)
Adobe AcrobatYes (advanced)No (separate step)Yes — exports editable PDFNo ($19.99+/mo)
iLovePDFYes (120+ source languages)Yes (25 target languages)GoodYes (1 task/day)
Doctranslate.ioYes (auto)Yes (100+ languages)Strong — coordinate-levelFree trial (15 credits)

Tips to Improve OCR and Translation Quality

Before scanning: Scan at 300 DPI minimum. Most scanner apps default to 150–200 DPI — check the setting. Maximize contrast. Scan pages flat.

When using a phone camera: Use a scanner app (Adobe Scan, Microsoft Lens, Google PhotoScan) rather than the native camera. These apps apply automatic deskewing and contrast enhancement.

After OCR, before translation: Review the OCR output for obvious errors before translating. A misread character in the source text becomes a mistranslation in the output.

Specify the source language if your tool supports it. Auto-detection works for common single-language documents. Mixed-language documents benefit from manual language selection.

When One-Step Translation Isn't Enough

Handwritten documents. OCR is not designed for handwriting. For handwritten legal records or correspondence, professional transcription before translation is more reliable.

Documents with heavy graphical content. OCR translates the text around graphics but cannot interpret the graphics themselves.

Certified or sworn translations. For immigration documents, court submissions, or official certificates, automated OCR + AI translation is insufficient. Most jurisdictions require human translators for certified work.

Frequently Asked Questions

Can Google Translate translate a scanned PDF? No. Google Translate's document upload works on native PDFs but cannot read scanned PDFs. To use Google Translate on a scanned document, first run OCR using Google Drive (open in Google Docs), then translate.

What is the best free tool to translate a scanned PDF? Smallpdf (2 tasks/day free) handles both OCR and translation. iLovePDF (1 task/day) offers OCR for 120+ source languages. Both are browser-based with no installation.

Why is my scanned document translation full of errors? Almost always a scan quality problem, not a translation quality problem. OCR errors from low-resolution scans carry directly into the translation output. Re-scan at 300 DPI+ with good contrast and re-run before translating.

Does the OCR step preserve the original layout? It depends on the tool. Basic OCR exports text in reading order, losing table structure and column layout. Advanced OCR integrated into a translation pipeline preserves spatial coordinates.

What to Do

If your scanned document is clean and high-resolution, most tools with built-in OCR will produce usable results. Start with a free tier and test on your actual document before committing.

If your document is older, lower quality, or has complex formatting, the OCR quality and layout preservation matter more. A one-step pipeline tool saves significant reformatting time compared to a two-step approach.