Back to all work

Certified translation bureau

From a scanned certificate to a finished translation

Client
Certified translation bureau, Lithuania
Timeline
First result in week one, live after a month
Status
Live, stage 2 in progress

Birth certificates, diplomas and passports from five countries arrive as photocopies. Translators used to retype them. Now the system reads the scan, fills in what the bureau has translated before and hands the translator a draft in the bureau’s own form.

The problem.

Every incoming document was a scan from an office copier, with no text layer. Translators retyped civil documents line by line, and years of finished translations sat in folders that nobody could search.

What we built.

  • The scanner sends to a mailbox; the system picks up each letter and opens a job for it.
  • Text recognition that keeps line positions and a confidence score for every word, so doubtful lines are marked for the translator.
  • A translation memory in the industry TMX format. Anything translated before is filled in from it; only new text goes to machine translation with the bureau’s glossary.
  • An automatic check that names, numbers and dates in the translation match the original.
  • Assembly into the bureau’s Word form, PDF export and a review screen. Every approved line goes back into memory.
  • A tool that aligns the bureau’s old archive with its originals, to grow the memory from past work.

Decisions worth explaining.

Cloud OCR instead of a desktop licence

About $1.5 per thousand pages, no page cap and no machine lock-in, hosted in an EU region because the documents contain passport data.

No model training on 100 documents

With this little data a curated memory is more reliable than fine-tuning, and the bureau can see and correct every pair in it.

Seals and names stay with a human

Seal text is never guessed, a placeholder goes in its place. Any change that touches a name or a number needs the translator’s confirmation.

In numbers.

1 537automated tests, run in 17 seconds
19 / 19documents of every type reached a finished file in the regression run
5source languages: UA, RU, BY, AZ, UZ
365checked translation pairs in memory after the first month

Stack

  • Python 3.12
  • FastAPI
  • Azure Document Intelligence
  • DeepL API
  • python-docx
  • PyMuPDF
  • OpenCV
  • LibreOffice
  • Windows Server
  • Caddy

Need something like this?

Describe your process. We will tell you which parts of this system fit it and what it would cost.

I want the same

Have a process that eats your evenings?