2025-12-09 · Baduno Editorial Team · 28 blog.readMin · Blog & Knowledge
Controlling AI Translation with Terminology: Glossaries That Work
A well-thought-out glossary is the key to controlling AI translations with precision and consistency. Learn how to build terminology databases, integrate them into MT workflows, and avoid common pitfalls. Practical tips for translators, project managers, and companies looking to optimize their multilingual communication.

Terminology in AI Translation: Fundamentals and Terms
The integration of terminology into machine translation (MT) systems is a key lever for achieving consistent and technically accurate results. Unlike purely statistical or neural translation, modern AI translation models work with contextual patterns. However, a terminology database (also known as a termbase) forces the system to follow your specifications in case of doubt. Generally, a distinction is made between static glossaries (lists with fixed translations) and dynamic termbases, which contain additional information such as word class, gender, context examples, or usage restrictions.
In practice, this means: a glossary is not a dictionary, but a set of rules for domain-specific terms. Example: in mechanical engineering, "Zugspannung" must always be translated as "tensile stress," not "tension" or "pull stress." Without terminology support, the MT system will choose the most probable variant based on training data—which often leads to inconsistencies. Also important is the difference between preference and obligation: in most MT systems, you can specify for each entry whether the translation should be preferred or enforced. The latter can lead to grammatically awkward sentences if the term does not fit the sentence structure.
Another fundamental concept is morphology: entries in the base form (e.g., "Schraube") often need to be supplemented with inflected forms, as MT systems do not automatically decline. Therefore, depending on the language pair, you should also include plural forms and conjugated verb forms. Otherwise, terminology only applies when there is an exact match. In practice, it has proven effective to have a maximum of 5,000 to 10,000 entries per language, prioritized by frequency and domain relevance. A well-maintained glossary significantly reduces post-editing effort—especially for technical documentation or legal texts.
Recommendation for action: start with a core set of 200–500 terms from your product or domain area. Decide whether the terminology should be applied "hard" (enforced) or "soft" (preferred). Test with 10 representative sentences to ensure translations remain fluent. Also document why a term was included—this facilitates later maintenance. Note: the more specific the domain, the more effective terminology-based control is.
Building a Terminology Database: Structure and Maintenance
An effective terminology database (TDB) thrives on thoughtful structure and regular maintenance. The foundation is selecting the right fields: minimum required are source term, target term, language code (e.g., DE-DE, EN-US), and status (e.g., "approved", "preliminary", "obsolete"). In practice, it has proven useful to also specify part of speech, subject area, and a short definition context. Example: For "Laufzeit" in IT, distinguish between "runtime" (program execution) and "term" (period). Without context, the MT system cannot correctly map.
Maintenance should be organized as a continuous process, not a one-time action. A central terminology management tool (e.g., T-Manager or a corresponding module in your translation memory system) is recommended. Define responsibilities: a subject expert reviews new terms, a translator or localizer maintains the entries. Ensure the TDB is language-neutral—each entry for a language gets its own record. Otherwise, problems arise with multiple meanings.
A common mistake is overloading with rare terms. Focus on terms that appear repeatedly in your texts or are particularly sensitive. Suitable sources for initial filling include: existing customer glossaries, terminology from translation memories (extracted via frequency analysis), standards and norms (e.g., ISO terminology), and product descriptions. Ensure each entry is unambiguous—synonyms should be marked as cross-references or with attributes like "preferred"/"permitted".
Recommendation: Create a maintenance log with a monthly rhythm. Query for unused entries (older than 12 months) and decide whether to delete or archive. Update entries from current projects at least quarterly. Test the TDB regularly with a sample of 50 sentences—if more than 10% of expected terms are not matched, check morphology or system configuration. A well-maintained glossary is not a static document, but a living tool that grows with your content.

Glossary Formats and Interfaces to MT Systems
The technical connection of a glossary to an AI translation system is crucial for actual terminology usage. Common MT platforms support different import formats. The most frequent is CSV (Comma-Separated Values) with a header line defining the fields. Example: "source_language","target_language","source_term","target_term","pos","domain". Important: Use UTF-8 encoding to correctly transmit special characters. Some systems also expect a specific column order—check the documentation. Alternatively, XML formats like TBX (TermBase eXchange) are used, which are ISO-standardized and allow more complex metadata. XLIFF (XML Localisation Interchange Format) can also contain terminology, but usually as annotations.
How do you now control terminology in the translation process? Modern MT systems offer two main variants: static glossaries (lists before translation) and dynamic glossaries (prompt-based integration). For platforms with an API (e.g., DeepL, Google Cloud Translation), you can pass a glossary in real time during the API call. Ensure each glossary is tailored to the specific language pair and domain—a general glossary for all cases dilutes its effect. In practice, we recommend one separate glossary per language pair and subject area with a maximum of 1000 entries.
Post-checking glossary effectiveness is often neglected. After a translation, you should spot-check whether the defined terms were correctly translated. Many MT systems do not log whether a glossary entry was actually applied. Therefore, an automated comparison is recommended: export the translation and search for glossary terms in the target text using a script. If a term is not translated as specified, check the cause: incorrect morphology, missing context, or overwriting by sentence structure. Limitations of control become apparent especially with highly polysemous terms or sentences that activate multiple glossary entries simultaneously—here the system can run into conflicts.
Recommendation: Start with CSV in UTF-8 format, as this is accepted by most MT systems. Use the provider's API to directly test glossaries. After each glossary update, perform a regression test with 20–30 test segments. Document the exact settings (e.g., priority "force" or "prefer") for each glossary. If you encounter unexpected deviations, often reducing the entries or adding context examples helps. The technical interface is only as good as the quality of the data—so invest in clean, uniform glossaries.
Integration of Termbases into Machine Translation Workflows
Integrating a terminology database into the MT workflow requires careful technical and organizational implementation. Modern MT systems like DeepL, Google Translate, or specialized platforms offer interfaces to import glossaries as separate files (CSV, TBX, XLSX) or via APIs. It is crucial that the term base is in a format supported by the system: TBX (TermBase eXchange) is an ISO standard suitable for exchange between different tools. CSV files are easier to maintain but require a clean column structure with term, translation, optional definition, and grammar information.
In practice, it has proven effective to host the term base directly within the MT system if offered, rather than manually uploading it each time. For example, some CAT tools like memoQ or Trados allow linking term databases with the MT engine. For cloud-based services like DeepL Pro, you can store a glossary in the customer portal. Note that the maximum number of entries may be limited – for DeepL, it is 5,000 per glossary. Therefore, prioritize the most important technical terms.
A common mistake is assuming the MT engine automatically applies the glossary to every sentence variant. In fact, many systems only consider terminology if the term appears exactly in the source text. Inflections, compounds, or synonyms are often ignored. To circumvent this, you can define "protected terms" that are recognized even in inflected forms, if the system offers this feature. Test before productive use to ensure your term entries are actually applied.
Recommendation: Run a test translation with 50–100 sentences containing your critical terms. Check the output for correct rendering. If the glossary does not take effect, verify the format, spelling (case sensitivity), and language direction. Document the workflow so that the term base can be easily re-imported when the MT engine is updated.
Prompt Engineering for Terminology-Driven Translation
With large language models (LLMs) like GPT-4 or Claude used for translations via API, you can control terminology through prompts. Instead of a traditional glossary, you define in the prompt how specific terms should be translated. A proven approach is to specify "Translation Rules" in the system prompt: "Always translate the following technical terms as indicated: 'data warehouse' → 'Data-Warehouse', 'machine learning' → 'maschinelles Lernen'." Ensure the rules are precise and context-free, otherwise the model may apply its own interpretations.
The effectiveness depends heavily on the model and prompt structure. Practice shows that explicit examples in a few-shot prompt work better than pure instructions. Provide 2–3 sample pairs with source and target text containing the desired terminology. Then add the text to be translated. Avoid the model treating the examples as part of the text – separate them clearly using formatting like """Examples""" and """To translate""".
A disadvantage of the prompt method is lack of persistence: each prompt must contain the terminology again, which is cumbersome for many requests. Moreover, LLMs are sensitive to minor changes in the prompt – a missing comma can alter the output. Therefore, for recurring translations, it is advisable to program an API request that automatically generates the prompt and loads the terminology from an external database.
Recommendation: Test different prompt variants with the same 20 test terms and compare results. Note whether the model reliably follows the rules or makes exceptions. For productive workflows, version your prompts and re-validate them after model updates. Use prompt engineering only if your infrastructure allows dynamic prompt generation; otherwise, classical glossary integration into MT systems is more robust.
Testing and Validating Glossary Entries in MT Output
Before integrating a glossary into the productive translation workflow, systematic testing is essential. Create a test suite with sentences containing your most important terms in various contexts – e.g., singular, plural, compounds, and different sentence positions. Translate these with the glossary enabled and disabled to isolate its effect. Automate this step if possible via API: compare the output with a reference corpus or extract specific term occurrences.
Validation should not only check the correct translation of the term itself, but also its grammatical embedding. A glossary that translates "Datenbank" as "database" is useless if the German sentence does not correctly form the dative or accusative. Some MT systems adapt glossary entries to the sentence context (e.g., inflection), others do not. Therefore, explicitly test difficult cases: "mit der Datenbank" vs. "die Datenbanken". If you find deviations, you can enrich glossary entries with extended attributes (e.g., POS tags) if the system supports this.
Another check point is completeness: does your glossary cover all relevant terms for the current text? Conduct a coverage analysis by searching the source text for glossary terms and calculating the hit rate. Fill gaps with additional entries. However, be aware that an overloaded glossary can overwhelm the MT engine – some systems prioritize the first entries or break off with too many rules. Keep the number per language pair at 200–500 entries, unless the system explicitly allows more.
Recommendation: Create a validation protocol that records the expected and actual results for each test case. Re-run the tests after every glossary update or MT system update. Involve subject matter experts to assess technical accuracy. Only when the glossary consistently enforces the desired terminology in tests should it be adopted for production. Otherwise, revise the entries or optimize the integration technique.

Review and Quality Assurance: Check Terminology Consistency
Controlling terminology through glossaries is a powerful tool, but actual compliance in the translation output must be verified. Relying solely on AI is not sufficient. In practice, a multi-stage review process has proven effective: first, run automatic checks using CAT tools or specialized QA scripts. These compare the target text against your termbase and flag deviations or missing translations for specific terms. A practical example: if your glossary specifies "specification document" for "Lastenheft," but a translator uses "requirements document," this will be listed in the QA report.
Next follows manual review by a subject matter editor or a second translator. This person reads the target text and pays close attention to the terms defined in the glossary. A helpful approach is to use search functions within the document or translation environment to locate and verify all occurrences of the terminological terms. Alternatively, you can perform a spot check: select ten to twenty key terms from the glossary and verify whether they have been consistently translated throughout the document. This is particularly efficient for large projects with many repetitions.
Another aspect is ensuring consistency across multiple files or project versions. Regular terminology reconciliation, where all translations from the current project are compared against the termbase, is recommended. A practical example from technical documentation: in a machine manual, the term "Sicherheitsabschaltung" appears. The glossary prescribes "safety shutdown." If a later revision uses "emergency stop," this is a case for follow-up review. The decision on whether it is an error or whether the term must be translated differently depending on context should be documented.
Finally, we recommend systematically recording the results of the follow-up review and feeding them back into the glossary in a feedback loop. If a glossary entry leads to incorrect translations, the entry should be adjusted or supplemented. This continuously improves the termbase. However, note that checking terminology compliance can have legal implications—especially in regulated areas such as medicine or legal technology. Consult your legal department or an external advisor on this matter.
Limitations of Terminology Control and Handling Exceptions
Even with carefully maintained glossaries, terminology control reaches its limits. Machine translation systems often interpret glossaries strictly, which can lead to undesirable results when context or polysemy are not taken into account. A common problem: a term has different translations depending on the sentence or subject area. If the glossary specifies only one variant, the MT translates it compulsively, even when the context requires a different meaning. Example: the English word "bank" can mean both "Bank" (financial institution) and "Ufer" (river bank) in German. A glossary defining "Bank" as a financial institution would produce an incorrect translation for "river bank." Here, exceptions must be allowed.
A pragmatic approach is to define glossaries not as rigid rules but as preferred translations. Many MT systems allow setting a priority: the glossary is considered but can be overruled by context (e.g., via a confidence level). In practice, it has proven effective to store separate glossary entries with conditions for ambiguous terms—for example, by specifying the subject area or an example phrase. This way, the system can select the translation "Ufer" for "river bank" when the term appears in a geographical context.
Another limitation arises with neologisms or proper names not yet included in the termbase. The MT may either leave them untranslated or provide a creative but incorrect translation. Manual post-editing is indispensable here. A practical example: the product name "SpeedMaster 3000" should not be translated into English. If the term is not in the glossary, the system risks translating it as "Geschwindigkeitsmeister 3000." To avoid this, proper names should be explicitly marked as non-translatable.
Finally, overregulation through too many or too detailed glossary entries can impair translation quality. If every word receives a fixed specification, the MT loses its ability to generate fluent, natural text. The solution: prioritize key terms and give the system free rein for less critical terms. After each major project, review which glossary entries actually led to improvements and which were more harmful. The question of liability for terminology errors can be legally relevant—consult a lawyer on this matter.
Automatic Term Extraction as a Basis for Glossaries
Building a glossary manually is time-consuming. An efficient alternative is automatic term extraction from existing reference texts. Using software such as TAUS, Sketch Engine, or integrated tools in CAT systems, a list of potential domain-specific terms is generated from a corpus. The extraction is based on statistical and linguistic methods: the system looks for recurring word groups (collocations) or rare words typical of the subject area. A typical approach: upload a collection of 10 to 20 carefully translated documents into the software and have it perform a frequency analysis. Terms that appear frequently and in various contexts are marked as term candidates.
However, the candidates obtained in this way must be validated manually. Not every frequent term is a relevant term—common words like "work" or "system" can appear as noise. A practical example: when extracting from a technical manual, the algorithm might classify the term "screw" as important, but it is actually an everyday term. This requires a subject-matter expert to review the list and remove irrelevant entries. A two-phase review has proven effective: first, an automatic prescreening based on frequency and statistical significance (e.g., using TF-IDF), and second, a manual review of the top 100 candidates by a domain expert.
Another advantage of automatic term extraction is the ability to generate multilingual glossaries. If you have parallel reference texts in the source and target languages, the software can also provide translation suggestions for the extracted terms. This is done through alignment methods that recognize sentence or word pairs. The quality of these suggestions varies: with good parallel data (e.g., from consistent translations), the results are often usable; with poor data, they are error-prone. Therefore, each automatically generated translation suggestion should be reviewed by a native speaker before being added to the glossary.
Automatic term extraction is particularly suitable as a first step for building a glossary or updating existing term bases. It saves time and uncovers terms that might be overlooked in manual creation. However, it does not replace human quality control. A hybrid approach—automatic prescreening plus manual review—delivers the best results in practice. When using term extraction services, also consider data protection aspects, especially if your reference texts contain confidential information. Clarify this with your legal department in advance.
A well-thought-out glossary is the key to controlling AI translations with precision and consistency. Learn how to build terminology databases, integrate them into MT workflows, and avoid common pitfalls. Practical tips for translators, project managers, and companies looking to optimize their multilingual communication.
Terminology Variants: Recognizing Ambiguity and Context
In practice, translators and project managers often encounter terms that must be translated differently depending on context. A classic example is the English word "lead": in marketing it can mean "lead" (potential customer), but in a technical manual it may mean "cable" or "lead glass." Without context recognition, machine translation can get this wrong. The challenge is to systematically capture such ambiguities and store them in the glossary with context-dependent rules.
An effective approach is to use context attributes in the term base. Instead of a single entry for "lead," create multiple entries, each specifying the domain (e.g., marketing, electrical engineering, medicine). In your MT system, you can then define rules that select the appropriate translation depending on the source document category or even the sentence environment. In practice, this means maintaining fields in the glossary such as "context," "source example," and "target example." This allows the engine to recognize "lead generation" immediately as a marketing context and select "Lead-Generierung." While this granularity requires more maintenance, it significantly reduces post-editing.
To limit effort, prioritize the most frequent or critical ambiguities. Create a list of the top 50 terms that are repeatedly mistranslated in your texts. Analyze existing translations and note the contexts in which errors occur. Then create context-sensitive entries for these terms. Leverage the capabilities of your MT platform: many systems allow conditional translation rules based on part of speech, neighboring terms, or document metadata.
Test these entries deliberately: create a short sentence for each context and check the output. For example, for "lead": "The lead is 2 cm long" (electrical engineering) vs. "The lead clicked on the CTA" (marketing). Adjust the rules iteratively. Document the decisions in the glossary so that all team members can understand the logic. Over time, this creates a finely tuned set of rules that noticeably improves translation quality in your domain.

Cost-Benefit Analysis: Effort of Glossary Maintenance vs. Quality Gain
Maintaining a terminology database requires resources: time for research, team alignment, technical integration, and regular updates. At the same time, consistent terminology reduces post-editing correction efforts and increases reader satisfaction. A blanket ROI statement is not possible, as it depends heavily on text volume, error frequency, and the consequences of incorrect translations. In practice, once you translate more than 50,000 words per month or operate in highly regulated fields (medical, legal, technical), the effort usually pays for itself within a few months.
Conduct a cost estimate: note how many hours you currently spend correcting terminology errors. Measure the time spent on post-editing over two to three projects. Then estimate how many of these errors could be avoided with a well-maintained glossary – typically 30–50%. Compare this with the estimated maintenance effort: a basic glossary with 200 entries requires about 20–40 hours initially, and monthly maintenance (for 10–20 changes) about 2–4 hours. Calculate whether the saved correction time exceeds this investment.
Also consider soft factors: consistent terminology strengthens brand perception and prevents customer misunderstandings. In technical documents, incorrect terms can lead to malfunctions or safety risks – the damage then outweighs any glossary effort. Start with a minimal but focused glossary: only include terms that are truly frequent or critical. Expand it gradually based on the most common corrections.
To make the benefit measurable, define metrics: terminology errors per 1,000 words in the MT output before and after glossary introduction. Measure this over a period of three months. This will objectively show whether quality improves. If the effort exceeds the benefit, check whether you can automate maintenance – for example, through term extraction tools or integration into your CMS. In many cases, the investment is worthwhile if you think long-term and establish glossary maintenance as a fixed part of your translation process.
Practical Checklist: Introducing a Glossary to the Team
Introducing a glossary for AI translation only succeeds if everyone pulls together. Use the following checklist to approach the process systematically.
1. Inventory and define goals: Analyze your most common terminology errors from recent projects. Define concrete goals, e.g., “Reduce terminology errors in the MT output by 30% within three months.” Determine which subject areas and languages should be prioritized.
2. Assemble the team and assign roles: Appoint a terminology manager who maintains the glossary and coordinates changes. Involve subject matter experts (e.g., engineers, legal professionals) to decide on disputed terms. The translator/editor checks entries for practical suitability.
3. Define glossary structure: Decide which fields your glossary should contain: source term, target term, definition, context, subject area, status (approved/obsolete), validity date. Keep the structure simple – too many fields make maintenance difficult.
4. Collect and approve initial entries: Start with 50–100 critical terms. Each entry should be reviewed by at least two team members. Document decisions including rationale to avoid later discussions.
5. Test technical integration: Integrate the glossary into your MT system. Test with representative texts from your inventory. Check whether the rules work as expected. Make adjustments until results are satisfactory.
6. Train the team and establish processes: Train all translators, editors, and project managers on how to use the glossary. Define how new terms are proposed (e.g., via a form) and who approves them. Establish a monthly review cycle for glossary updates.
7. Measure success and iterate: Regularly measure the terminology error rate. Collect feedback from the team and adjust the glossary accordingly. Celebrate small successes to maintain motivation.
With this checklist, you create a solid foundation for glossary introduction. The key lies in consistent maintenance and the involvement of all stakeholders. Start small and expand the glossary step by step – this keeps the effort manageable and the quality gain tangible.
Tools and Platforms for Terminology Management
Choosing the right terminology management solution depends largely on company size, number of languages, and integration with existing translation tools. For beginners, cloud-based solutions offer a low-threshold entry: they provide centralized access, version control, and user permissions. A typical setup includes a web interface for maintaining entries with fields such as term, definition, language, status (e.g., "approved" or "under review"), as well as synonyms and invalidation notes. Export functionality to standard formats like TBX (TermBase eXchange) or CSV is crucial so that data can be imported into CAT tools or MT platforms.
For companies with high translation volumes, platforms that support direct API integration with common MT systems are advisable. The glossary is queried live during translation: relevant terms are passed as context to the MT engine. Effectiveness depends on prompt design – experience shows that glossary entries should include source references and context examples to avoid misinterpretations. Some tools also allow priority rules: if a term appears in multiple glossaries, the ranking determines which one is applied. When selecting a tool, ensure that translators in the CAT tool can mark terms that are not correctly followed and make direct changes to the glossary.
Integration into the QA process is another key criterion. Modern platforms offer automated checks: after translation, the output is validated against the glossary, and discrepancies are listed. In practice, a two-step workflow has proven effective: first, a machine consistency check, and second, a random manual review by the terminologist. For global teams, collaborative features such as comments or change suggestions are recommended so that subject matter experts can also provide feedback. Be careful with too much freedom: define clearly who is allowed to approve entries to avoid unregulated growth.
Finally, keep an eye on costs: basic solutions are often free up to a certain volume, while enterprise features incur monthly fees. Check whether a one-time license or a subscription model is more suitable for your budget. Also consider the training time: the more interactive the interface, the faster editors can work. Before making a decision, request a proof of concept with your real data – only then will you see how well the terminology performs in the MT output. Please consult your own legal advisors regarding data storage and GDPR compliance.
Outlook: Adaptive Terminology and Continuous Learning
The next stage of terminology-driven translation is adaptive terminology. These are systems that learn from user corrections and automatically update their glossary. Imagine a translator changes a term in the CAT tool that was incorrectly translated by the MT system. The model remembers this intervention and applies it in similar contexts. This continuous learning significantly reduces manual maintenance effort. Some providers already integrate such feedback loops: after each approved translation, the term translation is adopted into the MT model's knowledge base. However, practice shows that the quality of these automatic adoptions varies – too many unchecked corrections can lead to inconsistencies.
The technical foundation is models using Retrieval-Augmented Generation (RAG): for each translation, not only the glossary is queried, but also the context from the current sentence and from previously corrected examples. This creates a dynamic profile per client or domain. The challenge lies in balancing freshness and stability: a glossary entry that has been learned incorrectly once can be hard to correct. A two-stage process is recommended: in learning mode, suggestions are collected but only adopted into the productive glossary after manual approval. Companies should regularly export the learned term inventory to compare it with the authoritative glossary.
Another trend is context-dependent terminology: not every term is translated the same way in every field. Adaptive systems can recognize whether a text comes from the legal or technical domain and automatically activate the appropriate sub-glossary. This requires that terminology is tagged with metadata such as domain, client, or document type. In practice, this requires clean classification of source texts. For many companies, a pragmatic approach is sensible: start with a central glossary and add domain markers as errors accumulate in specific areas.
The limits of adaptive methods lie in transparency and control. If a system learns independently, it is not always traceable why a particular translation was chosen. Therefore, especially in regulated industries (medical, legal), the final decision should always rest with a human. Looking ahead: in the future, glossaries could be directly incorporated into the fine-tuning of AI models, rather than just passed via prompts. This promises more consistent results but requires high computational capacity and regular updates. Companies that invest early in structured terminology have a clear advantage here. Have your MT provider explain their roadmap for adaptive terminology – and always test new features in a protected environment before deploying them productively.
Pitfalls and Common Mistakes in Glossary Work
Even a carefully created glossary can miss its mark if typical pitfalls are not heeded. A common mistake is overloading the glossary with too many entries. In practice, 100 to 200 well-maintained terms are sufficient for most projects. More entries often lead to inconsistencies and increase maintenance effort without a proportional improvement in quality. Focus on terms critical to your field where mistranslations are particularly serious, such as legal or technical terminology.
Another pitfall is insufficient context. An entry like "Kopf" -> "head" without distinguishing between "head of a screw," "head of a team," or "head of a list" leads to errors. Each entry should include at least a brief definition or an example sentence. Ignoring parts of speech is also problematic: a glossary that lists "überweisen" only as a verb will not correctly translate the noun "Überweisung." Therefore, include the relevant part of speech and, if necessary, inflection forms for each term.
Updating the glossary is often neglected. As soon as new products are introduced or naming conventions change, the glossary must be updated promptly. Plan fixed review intervals, e.g., quarterly. Without this maintenance, the glossary ends up in a drawer and is no longer used. Also ensure that the glossary is actually activated in the MT system. Some systems allow multiple glossaries with prioritization. After each update, check whether the changes are visible in the output.
A classic misconception is that a glossary alone solves all terminology problems. It cannot clarify grammar or style issues and reaches its limits with highly context-dependent terms. Therefore, supplement the glossary with translation rules in the form of "if-then" conditions, as far as the system supports it. Finally, gather feedback from translators. They can often provide practical insights into which entries are missing or incorrect. Only a living glossary that is regularly reviewed and adapted fulfills its purpose.
Collaboration with Vendors: Briefing and Review
When outsourcing glossary maintenance or terminology-driven translation to external vendors, a clear briefing is crucial. Define in advance which terms are non-negotiable for your company. Create a priority list: must-have terms that must be translated exactly, and nice-to-have terms where slight variation is acceptable. Do not hand over a raw glossary; instead, provide a cleaned version with clear field definitions (e.g., "Only in mechanical engineering context"). Without this clarity, vendors will translate at their own discretion.
A proven approach is to provide the vendor with a sample set of 200–300 segments in advance, on which they must demonstrate their terminology application. Request confirmation that the glossary has been correctly integrated into their MT workflow. Ask whether the glossary can be imported as TBX or XLSX file – many vendors use standard formats. After the first delivery, perform spot checks on terminology adherence. In practice, a 10% sample of the output, cross-referencing glossary entries, has proven effective. If errors occur, request corrections before the vendor processes the remaining volume.
Control mechanisms must be agreed upon from the start. Request a report on the number of glossary entries used and their match rate. Some MT platforms provide standardized logs showing which terms were triggered how often. This report should be submitted monthly or per project. If the vendor maintains their own terminology databases, clarify whether they will be overwritten or supplemented by your glossary. Misunderstandings can otherwise lead to parallel, conflicting glossaries.
Pay attention to the legal side: contractually agree on ownership rights to the glossary. Make clear that the glossary remains your intellectual property and the vendor may only use it for your project. Also discuss how changes are handled: who updates the glossary when new terms arise? How are correction loops billed? A transparent communication channel, e.g., a ticket system for terminology questions, prevents misunderstandings. For larger projects, a joint terminology workshop at the start is recommended – this investment pays off by avoiding later friction. For legal details, always consult your own legal advisor.
blog.faqT
How large should a glossary be for AI translation?
The optimal size depends on the subject domain and the number of target languages. For a specific project, 50 to 200 entries are often sufficient. The key is to select the most relevant terms with high impact on consistency and correctness. An overly extensive glossary can impair the performance of MT systems. Experts recommend starting with a core glossary and expanding it iteratively.
Which formats are suitable for exchanging glossaries with MT systems?
Common formats are CSV, TBX (TermBase eXchange) and XLSX. CSV is universally applicable, while TBX was specifically developed for terminology management and supports complex metadata. Many MT platforms also accept JSON or proprietary formats. Pay attention to correct encoding (UTF-8) and consistent column naming. Test before production to ensure all terms are imported correctly.
How often should a glossary be updated?
Ideally, updating should be continuous: every translation that contains new or deviating terminology should be reviewed. In dynamic fields such as technology or medicine, a monthly review is recommended. For more stable areas, a quarterly update suffices. It is important that changes are documented and communicated with the team. A designated person responsible for glossary maintenance increases sustainability.