Pages

Thursday, March 22, 2012

Exploring Issues Related to Post-Editing MT Compensation

As the practice of post-editing MT continues to gain momentum and perhaps even some acceptance as a legitimate practice, there continue to be questions raised about how to do this in a way that is equitable and beneficial to all the stakeholders.  There was an interesting discussion in LinkedIn on this subject where it is possible to see the perspectives of tools developers, LSPs and clients, and even some translators in their own words. Some of the things that stand out from this discussion are the general sense of the lack of trust between constituents in the translation production chain, the inability to share and take operational risk between stakeholders, and the difficulty in defining critical elements in the process e.g. MT/final translation quality and accurate productivity implications.


The issue of equitable compensation for the post-editors is an important one, and it is essential to understand the issues related to post-editing, that many translators find to be a source of great pain and inequity.  MT can often fail or backfire if the human factors underlying work are not properly considered and addressed. 

From my vantage point, it is clear that those who understand these various issues and take steps to address them are most likely to find the greatest success with MT deployments. These practitioners will perhaps pave the way for others in the industry and “show you how to do it right” as Frank Zappa says. Many of the problems with PEMT are related to ignorance about critical elements, “lazy” strategies, and lack of clarity on what really matters, or just simply using MT where it does not make sense. These factors result in the many examples of poor PEMT implementations that antagonize translators. 

Some of the key elements that need to be understood or implemented to maximize the probability of successful PEMT outcomes include:
  • Customize your MT engine for your domain requirements and generally MT engines make the most sense if you do repeat/ongoing work in the same domain and language. And be wary of any MT vendor or LSP who assures you that “for a nominal service charge you could reach nirvana tonight. If you do this properly there are no instant solutions and few shortcuts.
  • Use objective and mostly transparent and repeatable measurements of efficiency and quality that are trusted by key stakeholders e.g. SAE J2450
  • A good understanding of the cost structure and efficiency of the pre-MT translation production process (Human TEP = Translate, Edit, Proof). If you don’t understand where you are, how will you know what direction is forward? It makes little sense to deploy MT if you cannot improve upon the old process in some meaningful way i.e. timeliness, cost, and quality.  Tradeoffs will need to be made as it is not possible to improve all three of these elements.
  • An understanding of the “average translation quality” of the MT engine output. This can be determined at the outset by sample-based tests, and are useful input to establishing fair rates for the full project. It should be understood that MT engines that produce higher quality will require less effort to get to a level where the final delivered translation is equivalent to that produced from a standard TEP process. Really good engines can produce an average output segment that looks like an 85% fuzzy match or better from translation memory. This kind of system will also produce a large number of 100% matches for new segments, which still need to be verified and editors need to be compensated appropriately for this validation. Learn how to interpret and link measures like J2450, BLEU, TER, and Edit Distance to create your own unique measurements so that you can quickly understand what you are dealing with. Badly done, these metrics are a black hole of misunderstanding and wrong conclusions remember that automated metrics are only as good as the users' understanding of them. Human assessments are ALWAYS needed and always used in successful PEMT case studies.  If a survey of 90 interested users in March 2012 is to be believed, “over 80% of MT users have no reliable way of measuring MT quality”. If this is indeed true, it surely explains why translators are so outraged and why so much of PEMT yields less than satisfactory results.
  • An understanding of the target translation quality level. Interestingly, the easiest case to define is one where a client requires the same quality as they have from a standard TEP process. MT in this case is a draft version of the translation step of the TEP process and will still require the EP, or edit and proofing steps. Expect your EP costs to rise as your T costs fall. It is much harder to define the quality level when MT output will only be “slightly” edited for “understandability” in very high-volume knowledge-based projects. “Slightly” or “lightly” is very hard to define and even harder for a translator to understand. Studies (see the Sharon O’Brien links below) have shown that translator productivity is lower with this type of task than with one where the target is TEP quality. It is important to provide many examples of what is required. In these high-volume cases, it may be more useful to follow the 80/20 rule and focus 80% of the post-edit efforts on the 20% of content that is most important. Often this is best done through corpus analysis to define the human focus, and then compensating editors for the corrections they make or at a fair hourly rate, i.e. at a rate they would make on average for TEP work.
  • An understanding of the effort required to raise the MT to the target level. Once you understand your average MT output quality and have a clear target, it is possible to make an estimate of the post-editing effort. This should be the key determinant of what post-editor compensation is. If you wish to build a long-term relationship with post-editors it would be wise to compensate post-editors fairly. Thus if a system raises translation production efficiency consistently, I would recommend that you compensate editors at a rate to ensure their net income is higher than it would be in a TEP process. (So easy for me to say.) The proper rate can only be learned through experience so there are few useful generalizations to be made. The quality of your measurement systems really matters here and can help you get to the “right” (win-win scenario) rate faster. Also, it would probably be better to err on over-paying rather than under-paying as shown in these completely hypothetical examples e.g.
    • Average TEP rate 15 Cents, Average Daily Translation Output 2,500 words  =  $375  per day 
    • MT Engine 1: Average Post Edit Translation Output 7,000 words, Average Rate 7.5 cents = $525 per day
    • MT Engine 2: Average Post Edit Translation Output 5,000 words, Average Rate 10 cents = $500 per day
    • MT Engine 3: Average Post Edit Translation Output 4,000 words, Average Rate 12 cents = $480 per day
Omnilingua is an example of a company that has long term experience (5 years +) with PEMT and has developed sophisticated processes and methodology to understand this gap and human effort with rare precision. They are committed users of SAE J2450 for many years now, and thus understand quality and productivity with distinctive precision. You can see a video presentation of the Omnilingua approach in their own words starting at 31:00. It is my opinion that very few LSPs can make this PEMT effort assessment with any precision. This is where superior LSPs will excel and this competence should become a clear differentiator in future. (Ask your LSP the question: “How much effort is needed to make the MT output indistinguishable from standard TEP?” and watch them fidget around a bit as FZ says.)
 Remember also that most often, starting with a good professionally developed engine will produce better ROI, than starting with quick and dirty DIY options that require much more post-MT labor to raise the output to target levels.
  • Expect and plan for a learning curve and develop post-editor training materials. MT requires an investment in engine development and process learning as well, as measurement systems fall into place. However, once this new process is understood it is possible to have success, even with tough languages like Hungarian as Hunnect as shown with their training program. Not all translators are interested in post-editing and it is important to determine this early and then provide guidance to those who are interested and best suited to this kind of translation work.
  • While accurate quality measurements are important it is also critical to understand productivity impacts in as much detail as possible over time. Best practice suggests that it is important to monitor the use of MT through the various learning stages to best understand the financial and productivity impact. This may not be the same for every language as MT does not work equally well in all languages. Some MT systems will continuously improve and some will not. LSPs will need to decide where they should invest: MT technology, measurement systems and processes, PEMT training and new workflow, and/or solving new translation problems like customer support chat. It is unlikely that many will be able to do it all and the overall complexity and time taken to achieve mastery of all these new initiatives should not be underestimated.
  • Involve some translators in the MT engine steering process to identify major error patterns. This action has been shown to produce much more useful systems and higher productivity when you go into production. They can also help to establish meaningful and trusted measurements between the raw MT quality and establishing reasonable translator productivity expectations.
It will still require collaboration and trust (that rarely exists) between corporate customers, tool vendors, LSPs, and translators. The stakeholders will also all need to understand that the nature of MT requires a higher tolerance for “outcome uncertainty” than most are accustomed to. Though it is increasingly clear that domain-focused systems in Romance languages are more likely to succeed with MT, it is not clear very often how good an MT engine will be a priori, and investments need to be made to get to a point to understand this. The stakeholders all need to understand this and work together and make concessions and contributions to make this happen in a mutually beneficial way. This is of course easier said than done as somebody has to usually put some money down to begin this process. The reward is long-term production efficiency so hopefully, buyers are willing to fund this, rather than go the fast and dirty MT route as some have been doing.

I hope we have all reached a point where we understand that arbitrarily setting lower pay rates for MT-related cleanup work is unwise and that the lowest initial cost of building MT engines is rarely the best TCO (total cost of ownership) with MT technology. MT in 2012 is still very complex and requires real financial and intellectual investment to build real competency. 

I suspect that the most compelling evidence of the value and possibilities of PEMT will come from LSPs who have teams of in-house editors/translators who are on fixed salaries and are thus less concerned about the word vs. hourly compensation issues. For these companies, it will only be necessary to prove that first of all MT is producing high enough quality to raise productivity and then ensure that everybody is working as efficiently as possible. (i.e not "over-correcting"). I would bet that these initiatives will outperform any in-house corporate MT initiative in quality and efficiency.

I have seen that there are LSPs out there that know how to build the eco-system to make this a win-win scenario for all stakeholders so I know it is possible, even though it is not very common in 2012. In these win-win examples, the customer and the LSP understand the risks, and post-editors are paid more when the engine is not great and less when it is. Quality and productivity-related information flows freely in the production chain and is trusted, and often translators are compensated for higher productivity. Thus, I think there are three basic principles to keep in mind in developing fair and equitable compensation practices:
  1. Measure productivity and quality accurately, frequently, and objectively and share critical findings. Ensure that the links between MT quality and productivity expectations are realistic.
  2. Train and select the right people for post-editing work and monitor progress carefully, especially in the early stages.
  3. Link the compensation to the effort required to complete the job which means you need to have some understanding of this effort. Not all PEMT work is equal, when uncertain about the correct rates, initially err on the side of overpaying rather than underpaying to build a loyal workforce.
image
The LinkedIn discussion goes into many more details and is worth a look to get a broader and varied perspective on the post-editor compensation issue. It would be wonderful to hear other perspectives on this in the comments. Practitioners, especially LSPs, should understand that the real benefit of making these investments is long-term cost and productivity advantages that are sustainable and defensible. This, however, requires “hard work” as George Bush said, apart from the time and money investment and has a learning curve. Finally, I would warn you that we live in a time of Moses Madness and many yearn for quick fixes that cost nothing. These quick fixes can often backfire and we should heed the wise words of Frank Zappa in the song Cosmik Debris:
The Mystery Man came over
And he said: "I'm outa-site!"
He said, for a nominal service charge,
I could reach nirvana tonight

If I was ready, willing 'n able
To pay him his regular fee
He would drop all the rest of his pressing affairs
And devote his attention to me
But I said . . .
Look here brother,
Who you jivin' with that Cosmik Debris?

For those interested, here are some other references that may be useful to those trying to understand PEMT issues from other perspectives :

Wednesday, February 29, 2012

Highlights from Recent Coverage on MT Related Subjects

This is a summary of what I think are some interesting recent articles on the web on subjects relating to MT.

The Big Wave, an Italian initiative that focuses on the changes happening in language technology released details and proceeding papers from their conference held in Rome in the summer of 2011. There are many interesting papers related to MT, controlled language and collaborative translation related issues. These papers provide a balance of practitioner, academic and user perspectives on these subjects and are worth a close examination.
Some highlights include:

Linguistic resources and MT trends for the Italian language by Isabella Chiari discusses the implications of various kinds of data and their value for building data-driven MT systems and provides some specifics for EN <> IT MT systems. The paper is a great overview on the kinds of data that can be used and also provides insight on what data to use and where to use it with summary implications. It also makes a great case for the inevitability of corpus driven approaches in MT (without meaning to) by providing the theoretical rationale for this and points to rising momentum of the data driven approach.

Productivity and quality in MT post-editing – by Ana Guerberof provides specific evidence of the productivity advantage of MT over TM and new segments in a translation workflow.

“In this context, it seems logical to think that if prices, quality and times are already established for TMs according to different level of fuzzy matches then we just need to compare MT segments with TM segments, rather than comparing MT to human translation. “

This study also helps to establish that in reality MT is just a new kind of TM fuzzy match. Even though the test only involved a small number of translators and a small amount of work, it was done with care to ensure the translators saw a mixture or MT, TM and new segments in a way that was “blind” and then carefully measured the productivity of the translators in processing these different segments.

 

The results show that MT had higher productivity than TM or New segments and that on average MT produced higher productivity. (We are certain these results would have been more pronounced with an Asia Online customized system). Interestingly this study also shows that weaker translators seem to benefit more from MT and TM than the “best” translators. There are some interesting observations about the error analysis which showed that TM produced the greatest amount of final errors.

 

I would hypothesize that a test with more translators in the pool, and a bigger set of test data would be useful to do, as the results would establish the benefits of the use of customized MT much more clearly. It may even be useful to include “bad” or free MT to show how differently translators react to a segment that looks like it is an 85% match and to one that looks obviously like raw free MT or instant customization (50% TM match) that some use today.

 

 Why Machine Translation Matters: Trends & Best Practices 

This article summarizes the forces driving the increasing use of MT which can be summarized as:

External Forces in the World at Large :-

  • The digital data explosion and its impact on new content that begs to be to translated quickly

  • The global thirst for knowledge and information

  • The growing online population that does not speak English or FIGS but represents a major commercial opportunity for global enterprises

Internal Forces affecting Global Enterprises :-

  • The growing importance of customer conversations and user generated content which affects purchase decisions and impacts customer loyalty

  • The growing importance of open collaboration in B2C relationships

  • The Rise of Asia and BRICI which requires huge amounts of new content in new languages

These forces, together amount to a shift towards more dynamic content, and increase the need to handle streaming flows of information that simply cannot be done without more automation and MT.

 

MT: the new 'lingua franca' is a fascinating perspective by Nicholas Oster, a historian of world languages on how MT is enabling linguistic diversity on the Internet.

“Between 2000 and 2009, Arabic on the internet grew twentyfold, Chinese x20, Portuguese x9, Spanish X7 and French x6, while content in English ‘only’ tripled. Proportionally, then, English is declining in importance relatively quickly. “The main story of growth on the Internet … is of linguistic diversity, not concentration.”
Ostler sees a key role for MT in this new environment. Just as the print revolution changed the ‘ground rules of communication’ in 16th century Europe, he expects that language and translation technology will revolutionize global communications tomorrow, removing the need for a ‘single lingua franca for all who wish to participate directly in the main international conversation.’

Translation errors or nuances in both humans and computers can naturally have an important impact. But there is no point in dismissing MT by judging it by some presumed norm of ‘perfect’ human translation. MT is a revolutionary tool that can help the world communicate better. TAUS will be welcoming Nicholas Ostler as a speaker at the upcoming TAUS European Summit on May 31 – June 1 in Paris.

When Machine Translation Usefulness Is Higher Than Quality:  

This article provides some interesting feedback for those who insist that MT only has value when it approaches human quality, and since MT rarely reaches human quality it has very limited value. In this study, English news was translated into FIGS by MT, but users were always given access to the English source. The study measures the usefulness of the MT in the context of assessed translation quality as shown below and interestingly MT is considered useful even when the quality falls short of excellence. Since this study was performed some time ago we would assume that the usefulness curve continues to shift upwards, driven by improving MT quality, whatever some translators may think about the quality.

clip_image001

The graph shows that although the machine translation quality was evaluated as being far from perfect, the translation’s usefulness was regarded as higher than its quality. However, this applies only when translation quality is above certain threshold. Bad or poor quality machine translations are naturally deemed as useless.

 

This result confirms what many MT proponents have themselves experienced. Pure MT can be rough – often obscure, frequently humorous – but it can be useful. If one really has little facility in the source language, pure MT translations, however clumsy, can be a boon to understanding and, by extension, to productivity.

 

The graph below illustrates the breakdown of responses to the question, “How would you rate the overall quality of the newsletter translation?” by language group. Note that Germans felt the quality was more lacking, possibly because the MT was poorer in quality or possibly because they had higher expectations. It is actually well known in the MT community that German <> English is more difficult than English <> Romance languages.

clip_image002

When we segment answers to the question, “How would you rate the usefulness of the newsletter translation?” by the respondents’ English ability, we see an even stronger vote in favor of MT by the two lower groups. Thus users who had a self-measured poorer English ability, found the MT much more useful. In fact even many who responded has having “Good” English ability found the MT very useful or essential.
clip_image003


There have also been some interesting discussions in LinkedIn that cover the dialogue and tension between translators and MT advocates and also expose some of the hyperbole that some MT enthusiasts are prone to. While the discussion does meander between translator emotions about plans to “eliminate” them and less than scrupulous business practices by some MT vendors, it is an interesting thread. In their rush to get on the technology bandwagon some LSPs may overlook the privacy and data security issues that they inadvertently agree to when they use instant Moses and DIY kits.  So caveat emptor.


In Machine Translation in the European Union : Renato provides some summary coverage from  a recent conference of the ever expanding use of MT in the European Union internal administration.

 

Interview with Translator David Bellos: author and award-winning translator David Bellos knows a thing or two about translation would be an understatement. With over 40 years of experience, he has achieved international recognition for his works as a translator and biographer and has an impressive list of acclaimed publications to his name.

Some interesting excerpts from the interview:

“What I expect is that machines will allow the demand for translation to carry on growing, and for translation to become an ever more integral part of the world we live in.

However, since there are almost 49 million translation directions between all the languages in the world and there is never going to be a 49-million-fold community of translators, machines might well be a useful adjunct to actual translation for many of the under served directions that exist.

Wednesday, January 25, 2012

A Short Guide to Measuring and Comparing Machine Translation Engines

This is an article from the Asia Online November 2011 Newsletter that provides useful advice for meaningful comparisons of  MT engines and is authored by Dion Wiggins, CEO of Asia Online. So the next time somebody promises you a BLEU of 60, be skeptical, and make sure you get the proper context and assurances that it was properly done. And if they say they have a BLEU score of 90 you know that you are clearly in the bullshit zone.


“What is your BLEU score?” This is the single most irrelevant question relating to translation quality, yet one of the most frequently asked. BLEU scores and other translation quality metrics greatly depend on many factors that must be understood in order for a score to be meaningful. A BLEU score of 20 in some cases can be better than a BLEU score of 50 or vice versa. Without understanding how a test set was measured and other details such as language pair and domain complexity, a BLEU score without context is not much more than a meaningless number. (For a primer on BLEU look here.)


BLEU scores and other translation quality metrics will vary based upon:
  • The test set being measured: Different test sets will give very different scores. A test set that is out of domain will usually score lower than a test set that is in the domain of the translation engine being tested. The quality of the segments in the test set should be gold standard (i.e. validated as correct by humans). Lower quality test set data will give a less meaningful score.
  • How many human reference translations were used: If there is more than one human reference translation, the resulting BLEU score will be higher as there are more opportunities for the machine translation to match part of the reference.
  • The complexity of the language pair: Spanish is a simpler language in terms of grammar and structure than Finnish or Chinese relative to English. Typically if the source or target language is relatively more complex,  the BLEU score will be lower.
  • The complexity of the domain: A patent has far more complex text and structure than a children’s story book. Very different metric scores will be calculated based on the complexity of the domain. It is not practical to compare two different test sets and conclude that one translation engine is better than the other.
  • The capitalization of the segments being measured: When comparing metrics, the most common form of measurement is Case Insensitive. However when publishing, Case Sensitive is also important and may also be measured.
  • The measurement software: There are many measurement tools for translation quality. Each may vary slightly with respect to how a score is calculated, or the settings for the measure tools may not be set the same. The same measurement software should be used for all measurements. Asia Online provides Language Studio™ Pro free of charge and this software measures the scores, for a given test set, for a variety of quality metrics.
It is clear from the above list of variables that a BLEU score number by itself has no real meaning.
How BLEU scores and other translation metrics are measured
With BLEU scores, a higher score indicates higher quality. A BLEU score is not a linear metric. A 2 BLEU point increase from 20 to 22 will be considerably more noticeable than the same increase from 50 to 52. F-Measure and METEOR also work in this manner where a higher score is also better. For Translation Error Rate (TER), a lower score is a better score. Language Studio™ Pro supports all of these metrics and can be downloaded for free.
Basic Test Set Criteria Checklist
The criteria specified by this checklist are absolute. Not complying with any of the checklist items will result in a score that is unreliable and less meaningful.
  • Test Set Data should be very high quality: If the test set data are of low quality, then the measurement delivered will not be reliable.
  • Test set should be in domain: The test set should represent the type of information that you are going to translate. The domain, writing style and vocabulary should be representative of what you intend to translate. Testing on out-of-domain text will not result in a useful metric.
  • Test Set Data must not be included in the training Data: If you are creating an SMT engine, then you must make sure that the data you are testing with or very similar data are not in the data that the engine was trained with. If the test data are in the training data the scores will be artificially high and will not represent the level of quality that will be output when other "blind" data are translated.
  • Test Set Data should be data that can be translated: Test set segments should have a minimal amount of dates, times, numbers and names. While a valid part a segment, they are not parts of the segment that are translated; they are usually transformed or mapped. The focus for a test set should be on words that are to be translated.
  • Test Set Data should have segments that are at between 8 and 15 words in length: Short segments will artificially raise the quality scores as most metrics do not take into account segment length. Short segments are more likely to get a perfect match of the entire phrase, which is not a translation and is more like 100% match with a translation memory. The longer the segment, the more opportunity there is for variations on what is being translated. This will result in artificially lower scores, even if the translation is good. A small number of segments shorter than 8 words or longer than 15 words are acceptable, but these should be limited.
  • Test set should be at least 1,000 segments: While it is possible to get a metric from shorter test sets, a reasonable statistic representation of the metric can only be created when there are sufficient segments to build statistics from. When there are only a low number of segments, small anomalies in one or two segments can raise or reduce the test set score artificially.Be skeptical of scores from test sets that only contain a few hundred sentences.
Comparing Translation Engines - Initial Assessment Checklist
Language Studio™ can be used for calculating BLEU, TER, F-Measure and METEOR scores.
  • All conditions of the Basic Test Set Criteria must be met: If any condition is not met, then the results of the test could be flawed and not meaningful or reliable.
  • Test set must be consistent: The exact same test set must be used for comparison across all translation engines. Do not use different test sets for different engines.
  • Test sets should be “blind”: If the MT engine has seen the test set before or included the test set data in the training data, then the quality of the output will be artificially high and not represent the true quality of the system.
  • Tests must be carried out transparently: Where possible, submit the data yourself to the MT engine and get it back immediately. Do not rely on a third party to submit the data. If there are no tools or APIs for test set submission, the test set should be returned within 10 minutes of being submitted to the vendor via email. This removes any possibility of the MT vendor tampering with the output or fine tuning the engine based on the output.
  • Word Segmentation and Tokenization must be consistent: If Word Segmentation is required (i.e. for languages such as Chinese, Japanese and Thai) then the same word segmentation tool should be used on the reference translations and all the machine translation outputs. The same tokenization should also be used. Language Studio™ Pro provides a simple means to ensure all tokenization is consistent with its embedded tokenization technology.
Ability to Improve is More Important than Initial Translation Engine Quality
The initial scores of a machine translation engine, while indicative of initial quality, should be viewed as a starting point for rapid improvement which is measured by the test set and BLEU scores. Depending on the volume and quality of data provided to the SMT vendor for training, the quality may be lower or higher. Most often, more important than the initial quality is how quickly the translation engine quality improves

Frequently a new translation engine will have gaps in vocabulary and grammatical coverage. Other machine translation vendors’ engines do not improve at all or merely improve very little unless huge volumes of data are added to the initial training data. Most vendors recommend retraining once you have gathered a volume of additional data that is at least 20% of the size of the initial training data that the engine was trained on. Even when this volume of data is added, only a small improvement is achieved. As a result, very few translation engines evolve in quality much further than their initial quality.

In stark contrast, Language Studio™ translation engines are created with millions of sentences of data that Asia Online has prepared in addition to the data that the customer provides. The translation engines improve rapidly with a very small amount of feedback. It is not uncommon to get a 1-2 BLEU score improvement with as little as a few thousand post-edited sentences. Language Studio has a unique 4 step approach that leverages the benefits of Clean Data SMT and manufactures additional learning data by directly analyzing the edits made to the machine translated output.
Consequently, only a small amount of post-edited feedback can improve Language Studio™ translation engine quality quite considerably, and it can do so at speeds much faster and with far less effort than with other machine translation vendors. Asia Online provides complimentary Incremental Improvement Trainings to encourage rapid translation engine quality improvement with every full customization and also offers additional complimentary Incremental Improvement Trainings when word packages are purchased, greatly reducing Total Cost of Ownership (TCO). 
 
An investment in quality at the development stages of a translation engine impacts and reduces the cost of post editing directly, while increasing post editing productivity. While the development of some rules, normalization, glossary and non-translatable term work will assist in the rate of improvement, the fastest and most efficient way to improve Language Studio™ engines is to post edit the translations and feed them back into Language Studio™ for processing. The edits will be analyzed and new training data will be generated, directly addressing the primary cause of most errors. In other words, just post editing as part of a normal project will result in an immediate improvement. Little or no other extra effort is needed. By leveraging the standard post editing process, the effort and cost of improvement as well as the volume of data required in order to improve is greatly reduced. 

Depending on the initial training data provided by the client, a small number of Incremental Improvement Trainings are usually sufficient for most Language Studio™ translation engines to improve to a quality level approaching near-human quality. 

Other machine translation vendors are now also claiming to build systems based on Clean Data SMT. Closer investigation reveals that their definition of “cleaning” is not the same as Asia Online. Removing formatting tags is not cleaning data. Language Studio™ analyzes translation memories and other training data and ensures that only the highest quality in domain data from trusted sources is included in the creation of your custom engine. The result is that improvements are rapid. Even with just a few thousand segments edited, the improvements are notable. When combined with Language Studio™ hybrid rules and an SMT approach to machine translation the quality of the translation output can increase by as much as 10, 20 or even 30 BLEU points between versions.
Comparing Translation Engines – Translation Quality Improvement Assessment
  • Comparing Versions: When comparing improvements between versions of a translation engine from a single vendor, it is possible to work with just one test set, but the vendor must ensure that the test set remains “blind” and that the scores are not biased towards the test set. Only then can a meaningful representation of quality improvement be achieved.
  • Comparing Machine Translation Vendors: When comparing translation engine output from different vendors, a second “blind” test set is often needed to measure improvement. While you can use the first test set, it is often difficult to ensure that the vendor did not adapt its system to better suit and be biased towards the test set and in doing so delivering an artificially high score. It is also possible for the test set data to be added to engines training data which will also bias the score.
As a general rule, if you cannot be 100% certain that the vendor has not included the first test set data or adapted the engine to suit the test set, then a second “blind” test set is required. When a second test set is used, a measurement should be taken from the original translation engine and compared to the improved translation engine to give a meaningful result that can be trusted and relied upon.
Bringing It All Together
The table below shows a real world example of a version 1 translation engine from Asia Online and an improved version after feedback. Additional rules were added to the translation to meet specific client requirements, which resulted in considerable improvement in translation quality. This is part of Asia Online’s standard customization process. Language Studio™ puts a very high level of control in the customer’s hands where rules, runtime glossaries, non-translatable terms and other customization features ensure the quality of the output is as close to human quality and requires the least amount of editing possible. 


BLEU Score
Comparisons
Case Sensitive
Asia Online  

V1
SMT
V2
SMT
V2
SMT +
Rules
Google Bing Systran
Reference 1 36.05 45.96 56.59 30.58 29.64 21.01
Reference 2 35.80 39.31 48.85 32.05 29.94 22.56
Reference 3 38.65 52.31 65.03 35.51 33.17 24.68
Combined References 50.45 66.52 80.48 44.58 41.65 30.26
Case Insensitive            
Reference 1 41.30 52.65 59.25 32.18 31.49 22.49
Reference 2 41.01 45.32 51.24 33.67 31.64 23.88
Reference 3 43.99 58.97 67.49 37.15 35.01 25.92
Combined References 56.83 74.35 82.89 46.26 43.68 31.68
*Language Pair: English into French.     Domain: Information Technology.           

It can be seen clearly from the scores above that when all three human reference translations are combined the BLEU score is significantly higher and that the BLEU scores vary considerably between each of the human reference translations. The impact of the improvement and the application of client specific rules can also be seen, raising the case sensitive BLEU score from 50.45 to 80.48 (an increase of 30.03 in just one improvement iteration). One interesting side effect of having multiple human references is that it is often possible to judge the quality of the human reference also. In the example above, the machine translation output is much closer to human reference 3, indicating a higher quality reference. The client later confirmed that the editor who prepared the reference was a senior editor and more skilled than the other 2 editors who prepared human reference 1 and 2. 


A BLEU score, as with other translation metrics, is just a meaningless number unless it is established in a controlled environment. Asking “What is your BLEU score?” could result in any one of the above scores being given. When controls are applied, translation metrics can be used both to measure improvements in a translation engine and compare translation engines from different vendors. However, while automated metrics are useful, the ultimate measurement is still a human assessment. Language Studio™ Pro also provides tools to assist in delivering balanced, repeatable and meaningful metrics for human quality assessment.