Pages

Wednesday, October 24, 2012

Effective Determination of PEMT Compensation

An issue that continues to be a source of great confusion and dissatisfaction in the translation industry is related to the determination of the appropriate compensation rate for post-editing work. Much of the dissatisfaction with MT is related to this being done badly or unfairly.

It is important that the translation industry develop a means to properly determine this compensation issue in a way that is acceptable to all stakeholders. Thus, developing a scheme that is considered fair and reasonable by the post-editor, the LSP and the final enterprise customer is valuable to all in the industry. It is my feeling that economic systems that provide equitable benefits to all stakeholders are the ones most likely to succeed in the long term. Achieving consensus on this issue would enable the professional translation industry to reach higher levels of productivity and also increase the scope and reach of business translation as enterprises start translating new kinds of content with higher-quality, mature, domain-focused MT engines.

While it took many years for TM compensation rates to reach general consensus within the industry, there is some consensus today on how TM fuzzy match rates relate to the compensation rate, even though there is still dissatisfaction amongst some about the methodology and commoditization of TM and fuzzy-match based compensation schemes cause to the art of translation. Basically, today it is understood that 100% matches are compensated at a lower rate than fuzzy matches and that the higher the fuzzy match level the greater the value of the segment in the new translation task. Today fuzzy match ratings provided by the different tools in the market are roughly equivalent and for the most part, trusted. There are some (or many) who complain about how this approach commoditizes translation work but for the most part, translators work with an approach that says they should get paid less for projects that contain a lot of the exact same phrases, i.e. 100% matches in the TM that is provided to do new projects. 


However, in the world of MT, it is quite different. Many are just beginning to understand that all MT systems are not equal and that all MT output does not necessarily equate to what is available on Google and Bing. Some systems are better and many are worse (especially the instant Moses kind) and to apply the same rates to any and all MT editing work is not an intelligent approach. Thus, the quality assessment of the MT output to be edited is a critical task that should precede any large project involving post-editing MT output. The first wave of many MT projects just applied an arbitrarily lower (e.g. 60%) word rate to any and all MT post-editing work with no regard to the actual quality of the MT output. This has led many to protest the nature of the work and the compensation. Many still fail to understand that MT should only be used if it does indeed improve productivity and this is a key measure of value and thus should drive compensation calculations.

The first fact to understand and have in hand before you implement any kind of MT is the production rate BEFORE MT is implemented. It is important to know what your translation production throughput is before you use any MT. The better you understand this, the higher the probability that you will be able to measure the impact of MT on your production process.  This was pointed out very clearly in this post. (I should state that many of my comments here apply to PEMT use in localization TEP type projects only).



Many now understand that the key to production efficiency with MT is to customize it for a specific domain and tune it for a specific purpose. This results in higher quality MT output but only if done with skill and expertise. We are now seeing some practitioners making attempts to make quality assessments prior to undertaking post-editing projects, but there is a lot of confusion since the quality metrics being used are not well understood. In general, any metric used, automated or human assessment based, requires extended use and use experience before they can produce useful input to rate-setting practices. BLEU is possibly the one metric that is most misunderstood and has the least value in helping to establish the correct rates for PEMT work, mostly because it is usually misused. There is one MT vendor making outlandish claims of getting a BLEU of .9 (90) or better. (This is clearly a bullshit alert!) This is somewhat ridiculous since it is quite typical for two competent human translations to score no higher than .7 when their translations are compared unless they use exactly the same phrasing and vocabulary to translate the same source material. The value of BLEU in establishing PEMT rates is limited unless the practitioner has long-term experience and a deep understanding of the many flaws of BLEU. 

Another popular approach is to use human-based quality assessment metrics like SAE J2450 or Edit Distance. They work best for those companies that have used them over a long period and understand how the metric measurements relate to past historical project experience. These are better and more reliable than most automated metrics but are much more expensive to deploy and also their link to setting correct compensation levels is not clear. There is much room for misinterpretation and like BLEU, they too can be dangerous in the hands of those with little understanding or expertise with extended use of these metrics. It is important that whatever metric is used should be trusted and easily understood by editors to build efficient and effective production systems.
While all these measurements of quality provide valuable information, I think the only metric that should matter is productivity. It is useful to use MT only if the translation production process is more efficient and more productive with the use of MT. This means that the same work is done faster and at a lower cost. This can be stated very simply in terms of average productivity as follows (I chose a number that can be easily divided by 8 and stay with round numbers):

Translator Productivity before MT 2400 Words / Day or 300 Words / Hour 

Any MT system that cannot produce translated output and related productivity that beats this throughput, is of negative value to your production efficiency, and you should stay with your old pre-MT process or find a better MT system. MT systems must beat this level of productivity to be economically useful to the production goals and to be useful in general. (BTW most Moses and Instant MT attempts often do not meet this requirement.)

Thus it is important to measure the productivity impact of the specific MT system that you are dealing with and measure the productivity implications of the very specific MT output your editors will be dealing with. To ensure that post-editors feel that compensation rates have been fairly set it is wise to use trusted and competent translators in the rate-setting process. It would also be good to be able to do this reliably with a sample or have a reconciliation process after the whole job is done to ensure that the rate was fair. The simplest way to do this could be as follows:
1. Identify a “trusted” translator and have this person do 2 hours of PEMT work that is directly related to the material that will be post edited.
2. Measure the productivity carefully both before and after the use of MT.
3. Establish the PEMT rates based on this productivity rate and err on the side of over paying editors initially to ensure that they are motivated.
Good MT systems will produce output that is very much like high fuzzy match TM. The better the system, the higher the average level of fuzzy match. This still means that you will get occasional low matches and make sure you understand what average means in the statistical sampling sense. Thus, if a system produces output that the trusted translator can edit at a rate of 750 words an hour, we can see that this is 2.5X the productivity rate without MT. Based on this data point, there is justification to reduce the rate paid to 40% of the regular rate, but since this is a small sample it would be wiser to adjust this upwards to a level that will accommodate more variance in the MT output. Thus perhaps the optimal rate would be to set the PEMT rate at 50% of the regular rate in this specific case based on this trusted measurement. It may also be advisable to offer incentives for the highest productivity to ensure that editors focus only on necessary modification and avoid excessive correction. Other editors should be informed that the rates were set based on actual measured work throughput. And at least in the early days, it would be wise to measure the productivity as often and as much as possible on larger data sets. In time, editors will learn to trust these measurements and will remain motivated to work on ongoing projects assuming the initial measurements are accurate and fair.

It is of course possible to do a larger sample or test where more translators are used and a longer test period is measured, e.g. 3 translators working for 8 hours. Though based on experiential evidence across multiple customers we have seen at Asia Online, a 2 hour test with a trusted translator provides a very accurate estimate of the productivity and can help establish a rate that is considered fair and reasonable for the work involved. I am sure there are other opinions on this and it would be interesting to hear them, but I would opt for an approach where trusted partners and actual direct production data experience are the key drivers to setting rates over metrics that may or may not be properly implemented. 

I continue to see more and more examples of good MT systems that produce output that clearly leverages production efficiency, and I hope that we will see more examples of good compensation practices in which translators and editors find that they actually make more money as I pointed out in my last post, than they would in typical TEP scenarios using just TM. 

Whatever you may think of this approach, the issue of post-editing MT needs to be linked to an accurate assessment of the quality of the MT output and the resultant productivity benefit. It is in everybody’s interest to do this accurately, fairly, and in a way that builds trust and helps drive translation into new kinds of areas. This quality assessment and productivity measurement process may be an area that translators can take a lead in and help to establish useful procedures and measurement methodology that the industry could adopt. 

I have written previously on post-editing compensation and there are several links to other research material and opinions on this issue in that posting.  I would recommend it to anybody interested in the subject.

Friday, July 13, 2012

The Relationship Between Productivity and Effective Use of Translation Technology

As machine translation continues to gain momentum, we are seeing many more instances of LSPs and some enterprise users exploring the potential use of the technology in core production work. MT today is still unfortunately quite complex and there are few universally accurate truisms or rules of thumb that replace the need for at least some minimal amount of expertise and understanding. Expertise and knowledge are key requirements for those who wish to use MT successfully in a translation production context. However, there are still many misconceptions about the effective use of the technology.

Some of the most common misconceptions include:

All MT systems are about the same. Not really, some MT systems that have undergone expert-managed customization and domain-focused training can produce dramatically better results than generic systems. This also means that you are not likely to get a very good understanding of the capabilities of an MT technology without doing a real pilot project that involves customization. Yet I often see people trying to make judgments about which MT system to use based on running a few paragraphs through a generic engine.

All MT applications are the same. Some MT applications that are focused on localization (documentation, core website content) need much higher quality to be useful, than other applications like making customer support forum content multilingual where good gisting quality is adequate. Translator productivity applications are the most difficult to do successfully and one where naïve users (e.g. your average LSP with Moses) are likely to fail.

Post-editors should be paid the same lower rate for all MT post-editing work. CSA states that this magic rate is 61% of the full rate in 2010. However, setting a fixed rate without understanding the reality of the MT output quality can often be unfair to editors and cause resentment that can undermine any attempt to build production leverage.  Compensation needs to be linked to productivity and effort expended to “fix MT” and the most successful users are respectful and careful to do this well to ensure a stable and motivated workforce.
MT is responsible for falling translation rates

This is a digression, but I wanted to highlight some interesting analysis and opinion by Luigi Muzii on why this is NOT true and he provides very interesting analysis and opinion on this matter in this article and also in a post called “Changes Ahead” that was characterized as follows by Rob Vandenberg.



I will address the first three issues in this post and provide some more context to clarify these misconceptions.
MT systems can vary and produce very different type and quality of output depending on all of the following factors:
  • Methodology used (RbMT, SMT, Hybrid which can also mean many different things)
  • The skill and knowledge of the practitioners working with the technology and building the systems. MT is still quite complex and needs skills that take time to develop and refine, to get output quality that surpasses the quality produced by public MT engines from Google and Bing.
  • Increasingly the quality and the volume of the “training data” are important determinants of the quality of the system as SMT approaches increasingly lead the way.
  • The language pair: It is much easier to get “good” systems with FIGS than with CJK relative to English. Languages like Hungarian, Finnish, and Turkish are just tough in general (relative to English).
  • The ability of the system to respond to small amounts of strategic corrective feedback. This is critical to building real business leverage. While some systems may improve slightly when many millions of new words are added to train them, very few can respond favorably to small volumes of additional data. MT system development is evolutionary and one should enter into development with this mindset.

MT can be useful in many different scenarios but it should be understood that the expected usable quality for different uses is very different. We live in a world today, where MT translates billions of words each day for internet users who are trying to understand the content of interest on the web or communicate with others across the world. There are also many corporate and business applications where the sheer volume and volatility of the information could not justify anything but MT, e.g. technical knowledge base content, customer forum discussions, hotel reviews where “good enough” is good enough. Much of this information has little or no value over time e.g. configuration guidance on DOS 5.0/Windows XP or a 3-year-old hotel review but could have great value and enhance global customer satisfaction for a brief window in time even in an imperfect linguistic-quality form. MT use for traditional LSP applications is the most demanding of all MT applications and requires the deepest knowledge and expertise and skill. MT in this context can only add value if the output produced is of sufficient quality, that it actually enhances the productivity of translators and makes the business translation process more cost-efficient. It is not a replacement for human translation and thus needs to be at a quality level that humans acknowledge its utility and actually want to use it.

Much of the early dissatisfaction with MT in the professional translation world is a result of asking translators to edit poor quality output for much lower rates in a relatively arbitrary fashion, that did not accurately reflect the level of effort that was involved. The task of post-editing MT to publication-quality levels needs an understanding of the average level of effort needed and very few in the professional translation world have figured this out.

Omnilingua is an example of how to do it right, with a very clear and trusted quality measurement profile of the MT output which then also helps to define productivity and fair compensation for editors. This task of accurate measurement of MT output quality and then a determination of the correct compensation structure is key to successful MT deployment and is quite possible in high-trust scenarios but much harder to implement when trust is less prevalent.

In the following largely hypothetical example (which is based on a generalization of actual experiences) I have summarized the possibilities to show how MT system output quality and productivity are related. I have also taken the additional step of showing how lower word rates can often make sense with “good” MT systems, and hopefully demonstrate that it is in the interests of both LSPs and translator/post-editors to figure out the key quality/productivity metrics accurately. Once the productivity is clearly established lower rates make sense because the throughput is trusted. Both parties need to be willing to make adjustments when the numbers don’t properly balance out.
In this hypothetical comparison, we will assume that there are 3 MT systems all focused on the same production task. These systems are of differing quality and their related productivity impact is characterized below. The objective in every case is to produce final output that cannot be discerned from a pure human TEP production effort:
  1. Good Instant MT/Moses System – A large majority of these systems do not produce output better than the free generic engines on the internet. I am assuming that perhaps 5% to 10% of these systems can reach a state where they can outperform Google. TAUS has highlighted several case studies where this is documented and where it is clear this is difficult. Typically productivity for a very successful effort will range from 3,000 words per day and slightly higher.
  2. Average Expert System – A product of a reasonable amount of data and expertise and experience that enables productivity over 5,000 words/day to as much as 7,000 words per day for editors who work on correcting the MT output.
  3. Excellent Expert System – This is possible with data-rich systems developed by experts that have gone through several iterations of improvement and corrective feedback. I have seen systems that enable 9,000 words/day to as much as 12,000 words/day throughput. Some exceptional systems are even higher!
In the following table, these 3 systems are profiled to compare the overall time and cost implications for a 500,000-word project. This clearly shows (fabricated though it is) that higher quality MT systems will provide the best overall production benefits. This also implies that it is worth investing in developing this better quality up front, rather than opting for a low initial cost option that provides less benefit.


image


Thursday, June 14, 2012

Thoughts on an MT technology presentation at ALC New Orleans, May 2012

This is a guest post by Huiping Iler who I had the pleasure to meet in New Orleans last month who made a very interesting presentation on how to increase the intrisnic value of an LSP firm. She runs a language services firm that is one of the growing fold of LSPs who have direct experience with post-editing MT output, and see an increasing role for MT in the future of her business.  I should add that while her own feedback on my presentation here is quite flattering, there were also others who commented through the regular feedback process that my slides were too dense and information filled, and one who even felt that my presentation was a “thinly disguised sales pitch”. (I assure you Sir, it was not.) It is difficult to find a balance that makes sense to everybody and all feedback is valuable. The pictures below come from the wonderful photographic eye of Rina Ne’eman taken during her visit to New Orleans.

--------------------------------------------------
558695_10150900208229885_1009527471_n
It was a real delight listening to Kirti Vashee from Asia Online presenting on the ROI of Machine Translation – Scoping and Measuring MT. It took place at the most recent annual Association of Language Companies conference in New Orleans between May 16-20, 2012.

Kirti pointed out that:

  • Much of the today’s business content is dynamic and continuously flowing.
  • The need for real time international -language content cannot be met by human translators alone due to cost and time restraints.
  • Machine translation (MT), especially statistical machine translation is gaining traction among enterprises that have large amounts of data to translate.
  • IT companies and travel review sites are examples of early adopters of statistical MT.
  • Compared to any general or free MT tools out there such as Google, an enterprise MT tool and service like Asia Online is highly customizable and adaptable to unique customer needs
  • It gives clients much more control on terminology, non-translatable terms, vocabulary choice and writing style. As a result, it produces much higher accuracy and translation quality, especially in highly specialized and focused domains.
This echoes the feedback I heard from one of wintranslation’s enterprise clients who has been using statistical MT for the last few years. Our translation team have been tasked with post editing, providing corrective feedback to the client’s MT engineering team for continuous improvements.

According to translators who have mastered the art of editing machine translation, post editing raw output requires a different skill set than the traditional editing of human translations. 

As a starter, text selected for MT often tends to be “low visibility.” Kirti gave an example that for a travel review site, the four or five star hotel reviews are human translated while the lower star hotel reviews are machine translated with some or no human post editing. 

Other low visibility text examples include car service manuals that not everybody reads, or web-based support content. High visibility (and typically low volume) text such as marketing communications, rarely if ever, gets selected for machine translation. 

In the situation of translating low visibility text, particularly in technical communication, it is more important for the text to be technically accurate than stylish. It is a case where the translation might sound awkward but technically correct IS acceptable, as long as translation efficiency is maximized without hurting accuracy.
521315_10150909353029885_507237146_n
But translators new to post editing may be tempted to edit the text for not only accuracy but also flow and style. It leads them to spend more time than necessary on the text and they are also more likely to complain about the quality of MT output. After all style and flow is not the strength of MT but speed and consistency is. It is important to have an agreement with the human post editors what is good enough (i.e. technical accuracy only, not style). Improved productivity and lower cost are very important to clients using MT. The best post editors understand this and can deliver a high number of edited words per hour that meet quality standards. 

One of wintranslation’s MT post editors commented, “When I have to review a translation, either done by a human or by a machine, I do not try to make it sound like if I wrote it. I mostly correct errors, terminology inconsistencies, awkward style, problems with conveying the intended meaning and issues that really bother me. If we are able to have that mindset, then it will be less cumbersome to review machine-translated text. If we have the tendency to rewrite the translation, then the editing will be time-consuming and cumbersome.” It sums up the ideal attitude a post editor should have.

Consistency is one of machine translation’s core strengths. When set up properly, non-translatable text, like numbers, acronyms and product names are reliably consistent throughout the translation. It is an area that MT can outperform human translators.
For example,
Source: Migration information for JKJ 5.x
MT Target: Información sobre migraciones para JKJ 5.x
When a post editor reviews this text, she/he knows that for sure “JKJ 5.x” is correct and she/he doesn’t have to worry about it being translated as “JKJ 6.x” or “JKJ 5.s.” This is not always the case when reviewing human translations, because the editor will always have to double check the product name and version etc.

The absence of spelling errors in machine translated text is a distinct advantage that saves time. But it is a good practice to spellcheck the translation before delivery, because post editors could have introduced typos while inputting corrections.

When a post editor finds an error pattern, communicating it with the client will help training the engine and improving results for the future. For example, in one of the MT text, the term “wireless” is always translated into Spanish as “productos inalámbricos,” which in most cases is wrong. The post editor quickly identifies and fixes the error. Because this error happens often enough to be a pattern, it is submitted to the client for dictionary updating. This and other types of pattern based corrective work can greatly enhance the overall production efficiency of post-editing work.

When words are not in the right order in the translated text, it is best the post editor just drag them to the right place and that way he/she doesn’t have to retype them and delete them from the wrong place. This saves time.

Example
Source: Where to buy ABC Anti-Theft Service related products.
MT Target: Dónde comprar ABC contra robo servicio productos relacionados.
In this case, “productos relacionados” needs to be moved towards the beginning of the sentence, the post editor just highlights the two words and drag them to their right place. She also needed to move the word “servicio” and make a few quick fixes.
Final Target: Dónde comprar productos relacionados con el servicio ABC contra robo.
When the upfront linguistic set up work has been inadequate or there is a lack of ongoing communication between the post editing team and the MT engineers, it produces a lot of frustrations for MT editors and creates unnecessary delay.

For example, a Brazilian Portuguese translator noticed the MT software was often using European Portuguese vocabulary even though the text is intended for the Brazilian market (there are significant spelling differences between Brazilian Portuguese and European Portuguese).

For instance, acção (should be ação), gestores de projectos (should be gerenciadores de projetos), etc.

She asked why the machine translation software “was not told” about that. This ability to provide feedback to the MT system is a key ingredient to getting better results and raising editor productivity and satisfaction. The best results of MT come from close collaboration between the engineering and linguistic post editing team.

The translator mentioned above also found inconsistencies in the translation of key terms such as product names. “Green Power Management” was translated as Energia verde Management, Verdes gerenciamento de energia, and Verdes poder Management. Some editing of the translation memory to reduce such inconsistency would speed up the posting editing process a lot.

In terms of productivity gains, it varies from language to language. In Spanish and Portuguese for example where MT has made more inroads, one can expect as high as 50% productivity increase in terms of number of words translated per hour assuming the MT engine has been properly set up and trained. But gains are harder to come by in Asian languages. 

There is a real and imminent opportunity for translation companies to offer real-time translation services for select type of content that is out of reach for human translations due to time and cost. The linguistic training of statistical translation engines and developing post MT editors are key pieces in realizing that opportunity.
544765_10150912191064885_1060226544_n
On a side note, I cannot help but noticing Mr. Vashee’s passion and sharing of MT expertise is contagious. He is one of the finest craftsmen in the sales and marketing field of technology and translation; an empathetic communicator, he is always able to see things from his clients’ eyes; when in the company of translation company owners, he presents possibilities to use a tool like Asia Online to generate new revenue and create differentiation (ask which translation company owner doesn’t like to hear that); he satisfies the data driven analytical types with numbers and return on investment measured in quality metrics and dollars; he has an amazing ability to stay insightful and relevant in a conversation while sticking to his value proposition; he is an outstanding marketer and an entrepreneur’s dream pitch man. 


About Huiping Iler:
Huiping Iler is the president of wintranslationTM, a Canadian based translation company *specializing in information technology and financial services. wintranslation has been coordinating post editing of machine translated text for the last several years.
clip_image001
Canadian translation company wintranslation