What is TQA?
TQA (Translation Quality Assessment) is the systematic process of evaluating translations against defined quality criteria. It involves identifying, categorizing, and measuring translation errors to provide objective quality scores and actionable feedback.
Modern TQA has evolved from subjective reviewer opinions to data-driven assessment using standardized frameworks like MQM (Multidimensional Quality Metrics), enabling consistent quality measurement across projects, languages, and vendors.
TQA vs QA
| Aspect | TQA | QA (Quality Assurance) |
|---|---|---|
| Focus | Quality evaluation | Quality prevention |
| Timing | Post-translation | Throughout process |
| Output | Quality scores, reports | Process improvements |
| Tools | Evaluation frameworks | Automated checks |
Best practice: Use both — QA to prevent errors, TQA to measure results.
Key TQA Frameworks
MQM (Multidimensional Quality Metrics)
The industry-standard framework with hierarchical error categories:
Major Categories:
- Accuracy: Meaning transfer errors
- Fluency: Language quality issues
- Terminology: Term usage problems
- Style: Stylistic inconsistencies
- Design: Format/layout issues
- Locale: Localization errors
DQF (Dynamic Quality Framework)
TAUS-developed framework focusing on:
- Content type-specific evaluation
- Productivity-quality balance
- Industry benchmarking
LISA QA Model
Traditional model with:
- Minor/Major/Critical severity levels
- Pass/Fail thresholds
- Sample-based evaluation
Error Categories
Accuracy Errors
| Error Type | Description | Severity |
|---|---|---|
| Mistranslation | Incorrect meaning | Major/Critical |
| Addition | Unnecessary content added | Minor/Major |
| Omission | Content missing | Major/Critical |
| Untranslated | Left in source language | Major |
| Over-translation | Excessive interpretation | Minor/Major |
| Under-translation | Insufficient interpretation | Minor/Major |
Fluency Errors
| Error Type | Description | Severity |
|---|---|---|
| Grammar | Grammatical mistakes | Minor/Major |
| Spelling | Spelling errors | Minor |
| Punctuation | Punctuation issues | Minor |
| Typography | Typos, formatting | Minor |
| Coherence | Logical flow issues | Major |
| Register | Inappropriate tone | Minor/Major |
Terminology Errors
| Error Type | Description | Severity |
|---|---|---|
| Wrong term | Incorrect terminology | Major |
| Inconsistent | Same term translated differently | Minor/Major |
| Non-standard | Not using approved terms | Minor |
TQA Process
1. Define Quality Model
Project: Technical Documentation
Quality Model: MQM Full
Severity Weights:
- Critical: 10 points
- Major: 5 points
- Minor: 1 point
Pass Threshold: 98.5% (max 1.5 penalty points per 1000 words)
2. Select Sample
| Method | Use Case |
|---|---|
| Random sample | Large volumes |
| Full review | Critical content |
| Risk-based | Known problem areas |
| Statistical | Quality certification |
3. Evaluate Content
For each error identified:
- Category (Accuracy, Fluency, etc.)
- Subcategory (Mistranslation, Grammar, etc.)
- Severity (Critical, Major, Minor)
- Comment (explanation, correction)
4. Calculate Score
Quality Score = 100 - (Penalty Points / Word Count × 1000)
Example:
- 5000 words reviewed
- Errors: 2 Critical (20pts) + 5 Major (25pts) + 10 Minor (10pts) = 55pts
- Score: 100 - (55/5000 × 1000) = 100 - 11 = 89%
5. Report and Act
- Generate quality reports
- Identify error patterns
- Provide feedback to translators
- Implement improvements
TQA Metrics
Quality Score Calculation
Penalty-based scoring:
Score = 100 - Σ(Error_Severity × Error_Count) / Word_Count × Multiplier
Common thresholds:
| Score | Quality Level |
|---|---|
| 99%+ | Excellent |
| 97-99% | Good |
| 95-97% | Acceptable |
| 93-95% | Needs improvement |
| <93% | Fail |
Additional Metrics
| Metric | Description |
|---|---|
| Error density | Errors per 1000 words |
| Category distribution | % of errors by type |
| Severity distribution | % by severity level |
| Trend analysis | Quality over time |
Automated TQA
AI-Powered Quality Estimation
Modern TQA tools use AI for:
- Automatic error detection: Grammar, spelling, terminology
- Quality estimation: Predictive quality scoring
- Consistency checking: Cross-document validation
- Pattern recognition: Identifying systematic issues
Automated Checks
| Check Type | Detection |
|---|---|
| Spelling | Misspelled words |
| Grammar | Rule-based grammar |
| Terminology | Term compliance |
| Consistency | Segment variation |
| Numbers | Numeric accuracy |
| Formatting | Tag integrity |
Human + AI Approach
Workflow:
1. Automated QA checks → Fix obvious errors
2. AI quality estimation → Flag high-risk segments
3. Human TQA review → Evaluate sample
4. Feedback loop → Improve AI models
TQA Best Practices
For Evaluators
Use consistent criteria
- Same framework across projects
- Clear error definitions
- Documented severity guidelines
Be objective
- Focus on errors, not preferences
- Separate critical from minor issues
- Provide constructive feedback
Consider context
- Content type matters
- Target audience expectations
- Time/budget constraints
For Managers
Set clear expectations
- Define quality levels upfront
- Communicate thresholds
- Align with business goals
Enable improvement
- Share feedback with translators
- Track quality trends
- Reward quality achievement
Calibrate regularly
- Inter-evaluator agreement checks
- Framework updates
- Benchmark against industry
TQA Tools
Evaluation Platforms
| Tool | Focus |
|---|---|
| KTTC | AI-powered TQA platform |
| Memsource QA | Integrated CAT tool QA |
| XBench | Standalone QA checker |
| Verifika | Automated QA |
| QA Distiller | Error analysis |
Features to Look For
- MQM/DQF framework support
- Customizable error categories
- Automated checking
- Reporting and analytics
- Integration with CAT tools
- Collaborative review
FAQ
What is a good TQA score?
Industry standards typically consider 97%+ as good quality. Critical content (medical, legal) often requires 99%+. The appropriate threshold depends on content type, risk level, and business requirements.
How many words should be evaluated?
Sample size depends on project size and requirements. Common approaches: 10-20% for large projects, 100% for critical content, or statistical sampling for certification (e.g., ISO 17100 guidelines).
Should TQA be done by translators or separate reviewers?
Ideally, by qualified reviewers who didn't translate the content. This provides objectivity. However, translator self-review is valuable as a first pass before formal TQA.
How do you handle TQA disagreements?
Calibration sessions help align evaluators. For specific disputes: document both perspectives, escalate if needed, and update guidelines based on decisions. Consistency is more important than any single ruling.
Can AI replace human TQA?
AI enhances TQA but doesn't replace human judgment for nuanced quality assessment. AI excels at consistency checks and pattern detection; humans are essential for meaning, style, and cultural appropriateness.
How often should TQA be performed?
Every project should include TQA. Frequency within projects depends on risk: continuous for high-risk content, sample-based for routine work. Regular cadence enables trend tracking and improvement.
