40 languages automatically: how our AI translation handles technical terminology

40 languages automatically: how our AI translation handles technical terminology

A look behind the scenes of our automatic product-data translation - and why technical terminology has to be treated differently from a novel.

Machine translation today is so good that in many cases you can no longer tell it apart from human translation. Translation services work fluently, idiomatically, with a feel for register. Then you translate a DPP data set - and suddenly “rear lock fibre closure” becomes “Hinterschloss-Faserverschluss”.

The problem is called technical terminology. Here we explain why product data is not to be treated like novels and which tools Transpareo provides so that your 40 language versions stay comprehensible.

The fundamental problem: one word, several meanings

“Seal” in the DPP of an outdoor jacket: a watertight closure. “Seal” in a laboratory: a seal animal or a gasket, depending on context. “Seal” in a maintenance log: possibly a stamp.

A general translation model chooses based on the statistical context. With a flowing text this works - the novel provides plenty of context. With a data field primary_closure: seal there is barely any context. The model guesses.

A data field gives no context, and without context the model guesses.

The result is subtle errors. Not as dramatic as “Hinterschloss-Faserverschluss”, but consequential: a component that is called “Dichtung” in German is suddenly called “sigillo” instead of “guarnizione” in an Italian DPP. A buyer can no longer find the spare part.

What Transpareo delivers today

Our translation system transfers every new content automatically into all active languages. Four properties characterise it:

  • Markdown and variable preservation: placeholders such as <a href="/en/register">Pro-Mitgliedschaft</a> and Markdown structures are extracted before the translation, the pure text is translated, and afterwards the structures are reinserted unchanged. This keeps links, forms and layout consistent across all languages.
  • Central translation entries: translations are not stored in the data set itself but in a shared layer. Several data sets with the same source text share a translation. This saves translation costs and unifies terms automatically across the data model.
  • Automatic re-translation on change: if the source text is changed, the translations are regenerated in all languages. A correction in German - 39 other language versions follow automatically.
  • Per-data-set markers: content can be excluded from the automatic run or existing translations can be locked - for international product names or manual corrections, for example.

Where the customer fills the gaps

The automatic translation delivers mostly correct results for descriptive texts, marketing texts and care instructions. With critical technical terminology - the “seal”/“guarnizione” case - a residual set of errors remains that the customer’s admin has to correct.

Here the admin has three levers:

  1. Manual override per language and key: every translation entry can be opened in the application manager and adjusted per language. With the lock marker this manual translation is retained at the next automatic run.
  2. Glossary import: existing terminologies from translator tools or PDF glossaries can be imported as CSV and directly create written translation entries.
  3. Per-language corrections in ongoing operation: an Italian sales team notices an error and corrects it in the application manager - the correction takes effect immediately, the other translations are left untouched.

The application manager speaks the same 40 languages

It is not only the product passports that are multilingual - so is the interface you maintain them in. The application manager is translated into all 40 languages, and you can enter content in any of them. The system translates it automatically into all the others.

For multilingual teams this changes the way they work together: product management in Milan writes the care instructions in Italian, purchasing in Warsaw adds the material data in Polish, quality assurance in Hamburg reviews in German.

Everyone works in their own language, everyone sees the same data set, and whoever corrects an entry corrects it for all language versions at once.

The caveat from above applies here too: for critical technical terminology, a human should check in the end what the machine has chosen.

The EU language reality

24 official EU languages sounds like a lot. In practice there are three layers:

  • core markets: DE, EN, FR, IT, ES, NL - here every consumer expects perfection
  • significant markets: PT, PL, SV, DA, FI - a good level, occasionally you notice the machine
  • rare languages: MT, GA, ET, LV, LT - sometimes you have a DPP in Maltese without an end consumer in Malta ever scanning it. Mandatory nonetheless.

The obligation is not optional. The ESPR requires DPP content in the language of the member state in which the product is sold. Anyone serving 27 states therefore has 24 languages in play (some share languages).

From Maltese to Bengali

Transpareo translates into 40 languages - all 24 official EU languages and 16 more for global reach:

  • Europe: Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Irish, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Spanish and Swedish - plus Albanian, Bosnian, Icelandic, Macedonian, Norwegian, Russian, Serbian, Turkish and Ukrainian.
  • Worldwide: Bengali, Chinese, Hindi, Indonesian, Japanese, Korean and Vietnamese.

For the upcoming textile DPP, Bengali and Vietnamese cover the big production countries - a supplier in Dhaka reads the same passport as a buyer in Paris.

Why a centralised localisation layer

Most platforms store translations as additional fields on the data set: description_de, description_en, … 40 fields per translatable attribute. Sounds simple, but it has three disadvantages:

  • duplicated text. Two products with the same material note produce 40 + 40 translations instead of 40 once
  • hard to scale. Adding a 41st language means: a schema migration across all translatable models
  • corrections hard to apply globally. If “guarnizione” is corrected everywhere, all data sets would have to be edited individually

The shared translation layer solves this: one entry, many references. One correction, all data sets benefit.

What we do not yet have

A customer-specific terminology database with automatic suggestion detection is in the development plan but is not shipped today. Anyone who starts today gets far with the existing tools: manual overrides, glossary imports and the lock marker cover the most common use cases.

We believe that machines should do the bulk of the work and humans should intervene only where it is really necessary. Until the automatic terminology detection is available, the manual lever is transparent - and that is more honest than a promise that is not kept.

Questions on this article

In which languages does a product passport have to be available?

The ESPR requires the passport content in the language of the member state in which the product is sold. Anyone supplying the whole Union therefore has the 24 official EU languages in play, some of which are shared between states. Transpareo delivers those 24 and 16 further languages, and every new or changed text is transferred into all of them automatically. In practice that means language stops being a launch criterion - you write once, in the language you work in, and the market versions follow.

Who is liable if a machine translation gets a technical term wrong?

You are. The duty to keep the passport content correct sits with the economic operator placing the product on the market, and it passes neither to a translation system nor to us. That is why we say plainly where the machine carries and where it does not - descriptions, care instructions and marketing copy come out clean, while narrow technical terms such as component names, closures or coatings need a human eye. Treat the automatic run as a complete first draft in 40 languages and put your review time into the few hundred terms that actually matter to repairers and buyers.

How do I correct a term so the next automatic run does not overwrite it?

Open the translation entry in the application manager, adjust the value for that language and set the lock marker. The marker takes the entry out of the automatic run, so it survives every later change to the source text. Without the marker your correction counts as an ordinary translation and is regenerated as soon as the source text changes. The same two steps apply to a single word and to a whole paragraph.

Can we import a terminology list we already have?

Yes. Terminology you maintain anyway, whether it comes from a translator tool or from the PDF glossary your agency wrote, can be imported as CSV and lands as a written translation entry. Those entries then behave like manual corrections. The cheapest moment for this is before the first large import - the terms are in place before thousands of records are translated, and nothing has to be reworked afterwards.

Does a correction apply to every product or only to the record I opened?

To every record with the same source text. Translations live in a shared layer rather than on the individual product, so records with identical wording share one entry and a single correction applies to all of them. That is also why the correction effort does not grow with the catalogue - a term used on 4,000 products is one entry, not 4,000.

Do colleagues abroad have to work in German or English?

No. The application manager itself is available in all 40 languages, and content can be entered in any of them. Product management in Milan writes the care instructions in Italian, purchasing in Warsaw adds the material data in Polish, and everyone sees the same record. Whoever corrects an entry corrects it for every language version at once, which removes the usual loop through a central translation desk.

Is there automatic terminology detection?

Not today. A customer-specific terminology database that proposes terms on its own is in the development plan and is not shipped. Until it is, the levers are the manual override, the lock marker and the glossary import. We would rather say that clearly than have you plan around a feature that does not exist. For the common cases those three tools cover the ground, and the number of terms you have to pin is far smaller than the number of terms that need translating.

Updates on multilingualism and DPP practice

New languages, data quality and product features - curated once a month to your inbox.