Guides
6 min read

How to Translate a PDF Without Losing Formatting

PDFs are notoriously hard to translate because the format was designed for display, not editing. Here is what actually works — and why most tools get it wrong.


PDFs are the most common format for sharing business documents — contracts, reports, whitepapers, product sheets — and they are also the hardest document type to translate. The Portable Document Format was invented by Adobe in 1993 to solve a specific problem: make a document look the same on any device. It solved that problem brilliantly. Unfortunately, that same design makes PDFs a nightmare to edit or translate.

Why PDFs Are Hard to Translate

A PDF is not a structured document like a DOCX or PPTX. It is a stream of drawing commands: "place this glyph at position (142, 380), in font Arial 11pt, color #000000." There is no concept of a paragraph, a sentence, or even a word boundary in the base format. When a PDF parser extracts text, it is reconstructing sentences from those positional drawing commands — a process that works well for simple documents but fails on multi-column layouts, tables, text-over-image designs, and documents with complex typographic elements.

To make things harder, many PDFs are not text-based at all. Scanned documents and photos of documents are rasterized images — there is literally no text to extract. Translating these requires OCR (optical character recognition) before any translation can happen.

Common Approaches and Their Limitations

Copy-paste into a translation tool

Selecting text in a PDF viewer and pasting it into Google Translate or ChatGPT is the most common approach. For simple, single-column PDFs this works adequately. For anything with columns, headers, footers, tables, or floating images, the extracted text arrives scrambled. Column text gets interleaved; table cells lose their grid relationship; footnotes appear mid-sentence. You can get a rough translation, but recreating the formatted document from that is hours of manual work.

Google Translate file upload

Google Translate accepts PDF uploads but the output is almost always plain text with minimal formatting retained. For a simple one-column report this is workable. For anything with a designed layout — a brochure, a product datasheet, a multi-column research paper — you get back a wall of translated text with no resemblance to the original layout.

DeepL file translation

DeepL does not currently offer PDF-to-PDF translation with layout preservation as a supported workflow. It can process simple text-heavy PDFs, but complex layouts are not reliably handled. The output is typically a DOCX with extracted text, which still requires reformatting.

Manual translation services

Human translation agencies handle PDF translation by converting the PDF to an editable format (InDesign, DOCX, or XML), translating the text, and typesetting the result. This produces the best possible output quality and handles edge cases well. The tradeoff: cost is $0.10–$0.25 per word for common language pairs, turnaround is 3–5 business days, and the process involves sending sensitive documents to a third party.

How Layout-Preserving PDF Translation Works

The most robust technical approach to PDF translation with formatting preservation follows a three-step pipeline:

  1. 1Convert PDF to a structured intermediate format. For text-based PDFs, extract the text with position and style metadata. For scanned PDFs, run OCR to get a text layer. For complex layouts, convert each page to a high-resolution image to preserve the visual structure.
  2. 2Translate the text content while preserving positional metadata. This is where AI model quality matters most — the translation needs to fit within approximately the same text box dimensions as the original.
  3. 3Reconstruct the output document, placing translated text back at the original positions and regenerating a PDF that matches the source layout as closely as possible.

This is exactly how Vernacia handles PDF translation. It converts each page to a structured representation, passes text through GPT-4.1-mini (or GPT-4.1 for Enterprise jobs), and reassembles the output PDF with translated content in the original positions. The result is a PDF that is visually close to the source document — same column layout, same header/footer positions, same image placements.

Tip

For best results with PDF translation, start with a text-based PDF rather than a scanned one. If your document was originally created in Word, InDesign, or a similar tool, the PDF will contain real text data and will translate with much higher fidelity than a scanned image. You can check by trying to select text in your PDF viewer — if you can highlight individual words, the PDF is text-based.

Handling Scanned PDFs

If your PDF was scanned or exported from an image-based source, Vernacia runs an OCR pass before translation. OCR quality depends on scan resolution and clarity — 300 DPI or higher produces good results, while low-resolution or skewed scans may have recognition errors that carry through into the translation. For critical documents, review the OCR extraction before committing to the full translation job.

Special Cases

Legal and financial documents

For contracts, financial statements, and regulatory filings, you need more than layout preservation — you need terminological precision. Vernacia supports custom glossaries: upload a list of defined terms (in CSV format) and the AI will apply your preferred translations for those terms consistently throughout the document. This is important for legal documents where a mistranslation of a defined term can have real consequences.

Multi-language PDFs

Some PDFs contain text in multiple languages — a bilingual contract, a product manual with an embedded glossary in the original language. Vernacia handles these by auto-detecting the dominant language of each text block, translating accordingly, and giving you control over which blocks to translate and which to leave in the source language.

PDFs with charts and infographics

Text inside embedded charts, vector graphics, and infographic elements is a known challenge. If the PDF embeds these as rasterized images, the text within them cannot be extracted for translation. In these cases, Vernacia clearly marks untranslated image regions in the output so you know exactly what needs manual follow-up.

Getting the Best Results

  • Use the original PDF, not a print-to-PDF of a print-to-PDF chain — each rasterization degrades text quality.
  • For documents with precise terminology, build a glossary before translating.
  • For long reports (50+ pages), use the "sections" feature to translate and review in segments.
  • For RTL target languages (Arabic, Hebrew, Urdu), Vernacia adjusts text direction in output columns automatically.
  • Download the output and check it in a PDF viewer before sending to clients — a 5-minute review catches any layout issues.

Ready to translate your documents without losing formatting?

Translate your first PDF free — 25 page credits on signup