Pages

Wednesday, April 3, 2013

PEMT Case Study - Advanced Language Translation

The most active advocates of machine translation today are Fortune 100 companies especially in the IT industry and the translation agencies that serve them.  The large IT companies have used MT more widely than any other group. However, MT can also be used by smaller LSPs outside of this sphere, especially when they collaborate with experts. This is an example of one such case study which provides many specifics that might be illustrative and educational for others.

Corporate Translation & Localization Services

Advanced Language Translation (www.advancedlanguage.com) is a Rochester, NY based Language Service Provider (LSP) which has skillfully incorporated MT (machine translation) into its production process, after years of resisting the technology. CEO Scott Bass admits that this anti-MT stance caused them to miss out on some larger projects, as customers increasingly looked for service providers with a coherent automation strategy. Customers were looking for a partner who understood how to deploy machine translation in order to output cost-effective and high volume translation projects. After much debate, the company finally decided to jump on the MT train.

ALT began the process by identifying certain customers who were open to a collaborative PEMT (post-edited machine translation) production model. They then began to work with Asia Online in the summer of 2012 to develop MT engines for the selected clients. For ALT’s first MT project, engines were simultaneously developed for French, Spanish, Russian and Japanese; however, there were some issues that needed addressing in order to ensure successful completion of the project. The greatest challenge initially was the scarcity of data available to build and train the MT systems; and in fact, data volume was so limited that the likelihood of producing usable systems with raw SMT (statistical machine translation) approaches like Moses, was nil. The other challenge was building an engine for Japanese, as it is considered an especially difficult language for MT.

To remedy these issues, ALT collaborated with Asia Online to develop a terminology-driven data manufacturing strategy. They worked to build up critical data resources that enabled productivity enhancing systems to be developed, and they leveraged relevant monolingual data that was readily available to boost the engines’ capabilities in the domain of interest. ALT relied on the broad and deep experience of the Asia Online team to maximize and leverage their limited data assets and resources.

Additionally, ALT focused on using translators who had previous PEMT experience rather than using ones who either had no PEMT experience, or were not interested in working with PEMT output. Prior to establishing production deadlines and appropriate compensation rates, ALT sent several samples of MT output to the post-editors to ensure that the scope and difficulty of the work was well understood. Bass notes, “Many companies rush into ‘instant MT’ solutions, overlooking the fact that MT takes time to develop, and coordination among all parties. While it is possible to leverage MT systems once they have been built, practitioners must understand that there is a direct relationship between this initial effort and ongoing success with the engine.” He adds, “This outlook is critical to successfully leveraging MT in the long-run, and lack of it, is one of the main reasons why MT initiatives fail.”

ALT also allowed post-editors to set their own throughput rates based on their experience with the MT output samples produced by the customized systems. They discovered that on average this process resulted in throughput rates of 750 words per hour (6,000 words per day). For Japanese, the rate was lowered to 500 words per hour, as the MT systems produced lower quality output when translating between English and Japanese. After the throughput and MT quality issues were resolved, compensation was addressed by giving the editors a 25% premium over standard human editing rates. These parameters were established to the satisfaction of all parties for this initial “test” project; and it turned out to be successful on all accounts due to cooperation and skilled implementation.

image

PEMT Best Practices

Scott Bass summarizes lessons learned and gives advice for others undertaking MT initiatives:

  • Do not rush MT engine development. A higher quality engine takes longer to develop and may require multiple iterations to build it into a usable engine.
  • Pro-actively manage the expectations of all the people involved, including clients, project managers, post-editors and LSP sales and marketing personnel.
  • Ensure that post-editors understand the very specific nature of the work.
  • Ensure that MT output levels reach a quality level similar to a light to moderate cleanup of a human translated segment.
  • Collect as much data as possible including TMs, in-domain monolingual data in the target language and core terminology. (ALT used MemoQ LiveDocs to quickly build corpora.)
  • Test the MT engines and benchmark them prior to starting actual production work.
  • Give the post-editors insight into the kinds of edits they will have to make by producing examples with smaller representative test data sets.
  • Focus on minimizing the most frequent errors first and understand that dumb repetition can kill enthusiasm faster than anything else.
  • Ensure that the MT engine is improving through feedback from post-editors. Ask for their feedback often and give them plenty of time and attention.
  • Retune and retrain the MT engine quickly and as frequently as possible. 
  • Make sure that the strengths of MT are clearly understood, and manage any weaknesses throughout the process.

Overall Benefits

ALT is a fantastic example of a company who has leveraged MT properly. The company has demonstrated that when MT is used with skill and when human factors are carefully managed, the benefits go beyond mere increases in productivity.  ALT has found that overall business with accounts who ventured into MT has increased by over 75%. Bass notes, “In many cases, we gained preferred vendor status because we added MT to our service mix.”

Bass also emphasizes that sitting on the fence with regard to machine translation enabled ALT to deny the possible benefits of an MT-HT production model for far too long. Tackling the business and human challenges first were actually the most difficult facets of shifting ALT’s production model. In fact, Bass comments, “The process of customizing an MT engine is not that much different than undertaking formal terminology development or managing high-quality translation memories. Extending our toolset to include MT has been a natural extension of skills we already had in place as an LSP.”

To hear an online presentation of this case study you can also go to the Asia Online website.

 

Tuesday, February 26, 2013

Dispelling MT Misconceptions

MT in 2013 is still a complex affair requiring many skills, expertise and understanding that are not commonplace, to enable successful deployment as a productivity enhancing technology for business translation needs. While it has become much easier to build basic custom engines using a variety of Instant Moses solutions or by creating a dictionary for a RbMT, there are still very few who know how to coax MT system output to consistent productivity enhancing levels. Getting some kind of a basic engine up and running is NOT the same thing as having a production-ready post-editor friendly system. There are even fewer who know what to do if the first MT attempt does not work, or is lackluster. Most of these basic/instant MT systems are inferior to basically free online MT from Microsoft and Google. Building long-term productivity and strategic production advantage require much more skill, expertise and experimentation than most LSPs or users have access to, or care to invest in.   While it is sometimes possible for a user to get usable MT output after throwing some data into an instant MT/Moses engine, it is not common, even for “easy” languages like Spanish as several TAUS case studies reveal. 

It is my sense that MT is still complex enough that meaningful expertise can only be built around one methodology i.e. RbMT or SMT and that anybody who tells you that they can do both should be viewed with some skepticism. It is almost certain that they cannot do both well, and also quite likely they cannot do either well if they claim expertise in both, since very different kinds of skills are required. Specialization and long-term experience is necessary to build real competence with either approach.

We have reached a point today, where many more MT systems are successful, but we also have many mediocre systems that do not provide any long-term production/productivity leverage and can easily be duplicated by any competitor with minimal investment. Today it is quite easy to find many (usually bad) examples of free/instant MT but the best custom systems are still not widely known or commonplace. Good MT system development takes work and ongoing investment and require overall process modifications, communication and expectation management, not only technology investments.

Recently we have seen some articles in the blogosphere and even the mainstream professional translation press that continues to provide what I believe is a lop-sided and even a somewhat disingenuous view of the verifiable use and known best practices of various MT technologies. (This link gets you to full article). In this particular case it is somewhat clear that the author has a preference and a bias favoring an RbMT approach where value-add is generally limited to building dictionaries. 

The misinformation is typically around the following concepts:
  • Rules-Based vs. Statistical MT Comparisons
  • The scope and extent of possibilities with instant MT customization
  • The degree of expertise and experience required to develop skills in any of these  approaches

Firstly let me state my own biases:
  1. I think the Rules-based MT vs. Statistical MT arguments are largely irrelevant, even though I think it is increasingly evident that SMT is becoming the preferred approach, especially as more linguistics are added to the data-driven approach. To a great extent most systems out there except for raw Moses systems are all hybrids of some sort.  Recently MT technology has evolved to a point where SMT and RBMT concepts are being merged into a single ”hybrid” approach. While there is some overlap in these approaches, there are two primary hybrid models in use today.
  2. a) RbMT with SMT smoothing tacked on after the RBMT translation is completed, such as with Systran to help improve the fluency and quality of the often clumsy raw RbMT output and,
    b) Linguistically informed rules that modify source text before SMT processes it and that guides the SMT processes and additional rules after SMT processing takes place to perform normalization and adjustments to translation output where required. Or the newer syntax and morpho-syntactic SMT approaches which have shown limited success and are still emerging.
     
    Finally, what really matters is how much productivity does an MT system offer, and the RbMT vs. SMT issue is largely irrelevant. The objective is to get translation work done faster and more cost effectively.

  1. In the right hands, both approaches (RbMT or SMT) can work for projects where MT is suitable. However, there are many more user controls and much simpler options available to tune MT systems in the SMT world.
  2. In general I would say that it makes sense to specialize in one MT (SMT or RbMT) approach and go deep to understand what you can control and how it works rather than do shallow and instant approaches. It takes work and extensive experimentation to develop real expertise in either approach and there is nobody I know in the industry who can do both well.  So choose RbMT  or  SMT and figure out what it takes to make it REALLY work rather than do the kind of shallow tests that Lexcelera does and draws definitive conclusions on these results as described in the article. Many of the conclusions drawn in the article are more a reflection of the quality of their effort than the actual possibilities of the technology in more skillful hands.
Some of the specific claims made and disinformation in the Multilingual article referenced above that I would challenge and dispute are as follows:

“In our experience, languages such as Japanese and German perform best with an RBMT approach” This was actually true in the early SMT days (~2005-2007) but is simply not an accurate truism anymore. I have seen custom SMT (if done right) outperform customized RbMT systems in both these languages even when large amounts of data are not available.

if you do not have enough data — we're talking millions of segments of in-domain bilingual and monolingual segments — you may not have enough corpora to train an SMT engine This seems to me to be a statement often made by people who have little or very shallow experience with SMT. In the large majority of SMT systems I have been involved with this amount of training data volume was simply not available. However, it is possible to get productivity enhancing SMT engines with even just 50,000 segments if you know what you are doing. This is possible even for languages like Japanese and Russian as Scott Bass of Advanced Language Translation points out in this webinar, where this was done with a fraction of the data mentioned in this misleading statement. A large majority of Moses MT engines, especially those of the instant kind, produce MT systems that are inferior to the free MT provided by Google and Microsoft. This is more likely to be related to a lack of understanding about the technology rather than any fundamental deficiency in the basic technology or the data as the Multilingual article suggests. If data privacy or copyright is not an issue, most LSPs would probably be better of using the Microsoft Hub option over using some generic instant MT option or some LSP managed Moses effort. 

“If the terminology is fixed in a narrow domain such as automotive or software documentation, RBMT or a hybrid is generally the best choice. This is because the rules component protects terminology better”  While this may be true for systems developed by naïve Moses users, many SMT experts like Asia Online have figured out that terminology really matters and know how to use it. Most of the corporate SMT systems out there focus exactly on automotive and IT product user documentation of various kinds, in addition to unstructured content. It is in fact possible to build a single Automotive engine (at Asia Online) and then tune it for different clients (Toyota, Honda etc..) and have the preferred terminology dominate IF you know what you are doing. See the diagram below for example.
image

“Wild West content where the terminology runs all over the map and would be impossible to train for, such as patents, works better with SMT. ”  This again suggests the authors lack of experience with patent domain and basic unfamiliarity with SMT technology. The largest terminology effort I have seen was with a patent engine where tens of thousands of scientific and technical terms were carefully translated to ensure accurate and useful translation of patent material. SMT benefits greatly from good, consistent terminology work and we have several customers (e.g. Sajan) who have gone on record to say that terminology consistency was one of the major benefits of an Asia Online engine. In fact the strategy deployed by Asia Online in data scarce situations usually begins with a tightly focused terminological foundation. 

“However, if there are metadata tags, you should be aware that SMT doesn't preserve tags well, so RBMT or hybrid technology will save you some headaches.”  While this  may be true for many Moses efforts made by technically naïve and unskilled users, any SMT developer worth his/her salt knows how to easily resolve this problem. Asia Online handles all the formatting tags in XLIFF and TMX automatically and also provides a variety of tools that allow power users to do sophisticated handling of different kinds of formatting.

“Today's SMT systems are still hampered by a lack of predictability, which means that translators waste a lot of time verifying terminology that already ought to be automatically verified.”  Asia Online ran an experiment a few years ago using TDA data from multiple sources. It was discovered that combining data or using noisy data of any kind produces much lower quality MT systems.Understanding how to get the data clean and building a quality foundation makes on-going maintenance and update of the engine much easier and largely eliminates this unpredictability. We also discovered that consistent terminology in the TM ensures much higher quality results and thus at Asia Online we now have tools to ensure this. Again, if you know what you are doing this is a manageable issue and after you have built a few thousand engines you realize that unpredictability can be managed by data cleaning and ensuring terminological consistency. Kevin Nelson, Managing Director of Omnilingua, stated in a webinar that the terminology and writing style produced by his Asia Online MT system was even more consistent than a human only approach. This was specifically noticed by his end-client who contacted Omnilingua directly without prompting to discuss how they had accomplished recent improvements in qualityimage
“When post-editing SMT, that next training cycle may be six months or a year away because you usually want a fair bit of new data accumulated before you begin the process of retraining. In this case, the post-editors are not empowered to make lasting changes and it typically takes until the next training cycle to see any progress at all”.  This may actually be true for many Moses systems and for most naïve users of instant MT solutions. But for the higher value-add systems like the ones produced by Asia Online this is not true. There are two ways that SMT based systems can incorporate corrective feedback:
  1. Real-time corrections that are used on each job and can easily be done by translators every single time they run a translation. Since there is no additional cost for retranslating the same content at Asia Online, users are encouraged to resubmit the translation until it is in better shape to hand over to a post-editor. Many dumb and high-frequency error patterns can be corrected instantly by some simple analysis and corrections based on small test translation runs.
  2. Periodic retraining which is done when sufficient corrective feedback is available. Incremental Trainings with Asia Online can be performed in just a few days and can be performed with just a few thousand segments to show meaningful improvements especially with terminology and high-frequency phrases.
image
image

Perhaps the biggest misconception of all is that More Data is Always Better.  We now have much more evidence that this is frequently not true. Even Google, the high priest of big data, admitted this some time ago: "We are now at this limit where there isn't that much more data in the world that we can use"

So be careful to not believe everything you read (including on this blog) and if you take more than a glancing look at MT technology today you will probably understand that while it is becoming much simpler to play and experiment with MT, it is still a long way from being easy to produce production-quality systems that provide long-term business leverage. Do not underestimate the expertise requirements to be successful with MT, and realize that even after jumping in with Asia Online or others it will  take ongoing changes in process and human factor management to really achieve long-term cost advantages and build sustainable business leverage.  The reward for those who figure this out will be clear differentiation and long-term production cost advantages that others with instant MT or home-brewed Moses systems will never be able to match.

MT is messy and not quite as predictable as most want it to be yet. You have to have a stomach for uncertainty and are probably better off with "real experts" than people who say they can do it all and are "technology agnostic". And the next time you see an article that says they have all the answers for you and that for a nominal service charge you could reach nirvana tonight just tell them: "Don't you jive me with that cosmic debris!". 

Watch this video and feel your face melt at 4:50 when the guitar solo happens.




Saturday, December 29, 2012

Annual Review–Most Popular Posts of 2012

“Blogs are about sharing with authenticity. A good blog can help you really connect deeply with your audience in a meaningful way because the content is not only relevant but insightful and personal. I think most enterprises miss that point. When you do it right, your customers will walk away not only having learned something new but will also feel much more connected to your brand.David Armano EVP, Global Innovation & Integration at Edelman Digital
It seems like it was a just a moment ago that I summarized the most interesting blog posts of 2011 but here we are again and the world has not ended. I was not as active writing in 2012 as I was in 2011 as I felt that I had said much of what I had to say, and really there is only so much one can really write about machine translation without being repetitive. The topic has had more coverage across the industry and is perhaps slightly better understood now than it was last year. I am limiting the list to the top 6 since I had fewer new posts this year.  Since Google has killed the PostRank service I am now reduced to only providing the most popular list of blog posts. PostRank used to give us much better insight into the broader influence of any web content and helped identify seminal and influential rather than simply popular content. I resolve to be more active in the new year if I have ideas for new material and I am always open to suggestions. There are still many misconceptions about MT and I think that it would be useful to cover this in more detail and perhaps I will delve into that in 2013. 

Here is the list of most popular posts in order of popularity:

  1. Exploring Issues Related to Post-Editing MT Compensation This article continues to get attention today even though it was written early in the year and it still shows up regularly in the top 3 for every week. The post has links to several interesting comments on post-editing and I think this is possibly one of the reasons why it continues to be popular as it gathers different opinions and viewpoints in a useful and unbiased way. The popularity of this post suggests that this is an important issue to resolve in a fair and equitable way to enable broader MT adoption. All parties involved need to work together to establish trusted and equitable compensation for this process. I hope that others will step forward to share opinions and approaches that might further the dialogue. It would be useful for translators especially to step forward and suggest ways to do this more efficiently and accurately. For example this post by Jason Hall shows that simply equating MT output quality to TM matches may not make sense, and that leveraging MT is entirely different from leveraging TM.

  2. The Moses Madness and Dead Flowers This post was written very late in 2011 and thus it’s popularity was not reflected in the 2011 list. But it is another post that has continued to see regular traffic as more people wade through the Moses technology and realize that “free” and “DIY” is a still really a pipe dream with MT. Being able to whip up some sort of an MT system by throwing data into a computer has become very easy but the technology is still very complex and hairy, and requires at least "some" fundamental knowledge for any real success. I remain very skeptical about any instant MT approaches and I think we will continue to see a market where you get what you pay for. I would avoid any LSP whose strategy is based around instant MT solutions.

  3. Emerging Language Industry & Language Technology Trends This was a post that seemed to strike a chord and it very rapidly rose to being one of the most popular posts of the year. Thanks to all those who shared their opinions to provide broader context. In case you missed it you may also wish to take a look at Translation Guy’s humorous take on the post. You may also find the Asia Online Trends and Translation Industry predictions interesting and you can access the webinar and slides through the link provided.

  4. A Short Guide to Measuring and Comparing Machine Translation Engines This post provided specific and constructive advice on using BLEU scores correctly to assess your MT systems in a fair and accurate way. I see BLEU scores continually being used to mislead gullible users on a regular basis and there were even some presentations at the AMTA 2012 conference that claimed systems having .90 or 90 which to my mind is only possible if you cheat. In short BLEU measures the quality of MT system output against one or more human reference translations of the same material. It needs to be done carefully if you want meaningful and accurate results. It is possible to calculate BLEU scores on two human translations of the same material, and even there I have never seen a score higher than .7 or 70 since humans do things quite differently. There is a great discussion on the many issues with BLEU in this article and I recommend it so that you can understand the increasing number of discussions where it is referenced today.

  5. The Relationship Between Productivity and Effective Use of Translation Technology MT should only be used when it actually provides measurable productivity advantages. Higher quality MT systems generally provide much higher return on investment (ROI) and this post explores this issue in some detail. MT is a means to build long-term production advantage, but only when you do it well and if you are going to invest in this technology my advice is to do it as well as possible. Most of the short cuts will lead to dead-ends and remember that with MT, you are competing with smart people at Microsoft and Google who are doing the best they can for a general internet user population. Most translators will likely prefer to use these "free" engines to crappy LSP produced Moses and RbMT engines.

  6. Understanding Post-Editing  This is one of several posts on the subject of post-editing. This is a subject that is worth exploring more as there are also many misconceptions about the nature of the process and it would be useful for more voices to air both good and bad post-editing experiences so others can learn. Jost Zetsche has written about this in some detail in his newsletter but the scope and understanding of the role of language experts is still evolving and it is a worthwhile discussion to continue. I have not seen anything really useful coming out of conferences so I suspect the best stuff on the subject will happen in blogs and LinkedIn discussion forums.
I once again invite any interested guest authors who might wish to use this blog as a way to share an idea or an opinion on the translation industry. (There is a good blend of buyers, LSPs and translators who watch this blog). I do not seek only those who agree with me to apply to do this, and in fact I hope that some who disagree will also step forward. I have always thought that it is useful to hear many different opinions to better understand a subject. So please don’t hesitate to send me contributions that you think might be interesting to the audience that has been following this blog. I thank you for your support and I hope that the content here will continue to earn your interest and comments to extend the discussion beyond my thoughts on key translation automation related issues.

It is also interesting to note that some older posts continue to strike a chord with readers and remain active in terms of visibility because the themes are longer lived and also perhaps because they ring true. The original post on standards, the analysis of why Google changed the use model of their MT systems and some of the posts that discuss the reaction to automation or industry disintermediation were also posts that generate continuing interest and continue to show up high in the list in Google Analytics.

I found a very interesting blog post that I think is worth a read, as it points to the changes that widespread information availability and ease of access creates to traditional commerce by socially engaged human beings. There is also a link to the research data from Mary Meeker on the changing online world that is worth at least a quick look. I think we are heading back to world where it is more important to understand how people connect rather than assume that technology and data will solve every problem known to man. I have always preferred the emphasis on Why? rather than  How?

Commerce in 2013 is about integrating the whole experience around the customer -- social, local, and mobile, bricks and clicks, in real life, in real time, and over time.  
 
Finally, I want to share a beautiful piece of music by Mercedes Bahleda that I discovered through Pandora  - the video is also very evocative and sublime with scenes of inter-species communication and a langorous swim dance. Those of you who find the sight of a female human breast offensive (there are unfortunately many in America who actually do) may wish to avoid actually looking at the video. I suggest you turn the volume up and play this on good speakers for maximum effect.




Happy New Year – I wish you health, happiness and joy

I don’t tell the murky world
To turn pure.
I purify myself
And check my reflection
In the water of the valley brook.

Zen Master Ryokan

“If the light’s not in you, you’re in the dark.”
Marty Rubin