Pages

Friday, March 26, 2010

Dispelling some Popular Misconceptions about Statistical MT

Recently some RbMT enthusiasts have been speaking up about what they consider are SMT deficiencies in the LinkedIn MT group and also in a March 2010 article in Multilingual magazine.

I covered the SMT vs RbMT debate in a previous blog entry but I thought it would be useful to clarify some of the key issues that get brought up, more specifically. It is clear that Google and Microsoft who had years of experience with Systran RbMT have decided that an SMT-based approach is preferable. Probably because of better quality and more control. I even decided to put a Microsoft MT widget on this blog since it allows readers to both translate and correct the translations. How cool is that? This dynamic feedback loop will become a hallmark of all SMT in future.

I have noticed that much of the criticism comes from people who have had little or no direct experience with SMT (beyond Google Translate) and have not had experience customizing a commercial SMT system. Or perhaps, their information is somewhat dated and based on old research data; e.g. much is made about the blue cat being translated as le bleu chat rather than as le chat bleu. My attempts to recreate this “problem” on 4 different SMT engines showed that all the SMT engines in the market today understand how to do this without issue.Go ahead, try it.

Some of the key criticisms are as follows:

-- SMT Is unpredictable:
This is the most common criticism from RbMT “experts”. I think this may be true to some extent if your training data is noisy or not properly cleaned and normalized. An experiment that I was involved with showed exactly this, but only with dirty data (Garbage In Garbage Out). Since many of the early SMT systems were built by customizing baselines built from web-scraped data, there is risk and evidence of unpredictable behavior. At Asia Online we have seen that this unpredictability essentially disappears as the data is cleaned and normalized.
Clean data reduces Unpredictability
Another related criticism I have seen repeated often, is that SMT systems are inconsistent in terminology use. SMT systems do by definition tend to choose phrase patterns that have higher statistical density, but this is easily corrected by using glossaries (much simpler than dictionaries) which provide an over-ride function and encourage the engine to use very specific user-specified terminology.

-- SMT is harder to setup and customize:
The bulk of the effort in building custom SMT systems is focused on data cleaning, data gathering and data preparation for “training”. This should include both bilingual parallel text as well as monolingual target language text. Once data is cleaned, SMT engines can usually be built in days. In future we can expect that this could be reduced to hours for most systems.

TAUS has reported that Moses has been downloaded 4,000 times in the last 12 months. This suggests that many more will be working with SMT in the near future. Historically, the single commercial SMT vendor (US government focused) has been prohibitively expensive and provided very little user control. This is changing and as more people jump into open systems SMT, and I think we will see this become a non-issue. What many fail to understand is that RbMT systems give you the ability to control quality by adding dictionaries and that this is really the only control variable one has. There is generally no ability to modify or add rules without huge expense and vendor support.

We already see companies like Autodesk and Sun (Oracle) and even some LSPs building their own SMT systems, probably, because it is not that difficult. At Asia Online we have clean foundation data for 500+ language combinations that allow our users to quickly build custom systems by combining their data with ours. We will continue to build our clean data resources to enable more people to easily build custom SMT systems.
Data Path to Qlty

-- SMT requires large volumes of reliable and clean data:
It is true that SMT does need data to produce good results. But this includes both bilingual TMs as well as monolingual target language which can be easily harvested from the web. For many European languages it is possible to get significantly better results than Google with as little as 100,000 segments in a focused domain. Of course you get better results with more data. Initiatives like PANACEA, LDC, OPUS, Meedan  and entities like the UN and the EC are making an increasing amount of training data available to the open source community and we will see this issue also get easier and easier to solve. The TDA may also be helpful if you have money to spare. There are already several SMT-based start-ups in Spain, who have considerable experience in RbMT but choose an SMT-based approach for their future because they see higher quality results are possible with less effort. This shift in focus by competent RbMT practitioners is very telling and I think suggests more will follow.

-- RbMT is a better foundation for hybrid systems:
Hybrid systems are now increasingly seen as the wave of the future. We see that the original data-only mindset amongst SMT developers has changed and increasingly they incorporate more linguistics or syntax and grammar rules into their methodology. RbMT systems after 50 years of development (in some cases) realize that statistical methods can improve long term structural problems like the fluency of their output. I think the openness of the SMT community will out distance initiatives with RbMT foundations, especially since it is so difficult to go into the rules engines and make changes. Much of the advances over the next few years will involve pre- and post-processing around one of these approaches. There will probably be another 5,000 people download and play with Moses this year. This broad collective effort will generate knowledge that my instincts tell me will move faster to higher quality than the much lower investment on the RbMT side.
AO Hybrid
While there are some language combinations where RbMT systems outperform SMT-based systems, I think this too will change. RbMT can still make sense if you have no data but then why not just use the free online MT? Recent anecdotal surveys by the mainstream press are only the beginning of the coming SMT tidal wave. Most of us can remember how much worse the free online MT experience was a few years ago, when Google and Microsoft were still RbMT based. I now notice continuous improvements in these data-driven systems. As the internet naturally generates more data they will continue to improve. Ultimately this competition keeps everybody honest and hopefully on their toes. Whichever approach produces better quality is likely to gain increasing momentum and acceptance.  Apart from the fact that it is getting easier to get data and SMT's open source relationship, the major driving force I think is user control. SMT simply gives one more control across the board, even the learning algorithms are modifiable. My bet is on the SMT + Linguistics + Active Human Feedback loop as the clear quality leader in the very near future.

This debate is important as the types of skills that active users need to develop for each approach are quite different. For SMT:  Data Cleaning, Data Analysis, Linguistic Steering  and Linguistic Pattern Analysis and even Linguistic Rules Development for some languages. For RbMT: Dictionary Building and recently Statistical Post-Editing. But it is clear that rapid, efficient post-editing is important in either case.

Please chime in and share your opinion or experience on this issue.

Monday, March 22, 2010

Why Machine Translation Matters -- Part II

I just spent the last few days at the ATA-TCD Conference in Scottsdale AZ. You can read highlights from the Twitter stream by searching on #TCD11. While it always nice to be in the sun for a few days, it is encouraging to see people in the industry focused on change along key dimensions like standards, technology and automation as well as the impact of social media on business strategy. I enjoyed several thought provoking presentations and discussions I had with many attendees during the conference.

One of the sessions I did was with Alon Lavie and Mike Dillinger who (both represent AMTA leadership) gave a very useful overview for LSPs to get a better, more realistic sense about MT and provided a basic primer on the subject. I thought that since my original blog post on this subject is my most popular post it might be useful to further develop this theme.

My original post focused on how MT could help address information poverty.Here are some of my new comments on why MT matters from the presentation at TCD11. The issue of growth in the sheer volume of information is increasingly clear to most but it is worth restating with some specific projections from IDC and EMC who monitor this very closely. The following chart shows projections just on enterprise content volume.
Enterprise Data Growth

In actual fact the fastest growth is actually in user generated content (UGC) e.g. blogs, FB, Youtube, Flickr and community forums. It is estimated that 70% of the content on the web is UGC and much of that is very pertinent and useful to enterprises. This content is now influencing consumer behavior all over the world and is often referred to as word-of-mouth-marketing (WOMM). Consumer reviews are often more trusted than corporate marketing-speak and even “expert” reviews.We all have experienced Amazon, travel sites, C-Net and other user rating sites. It is useful for both global consumers and global enterprises to make this multilingual. Given the speed at which this information emerges, MT has to be part of the translation solution though involving humans in the process will produce better quality.
UGC Importance

So if this is going on, it also means that what used to be the primary focus for the professional industry, needs to change from the static content of yesteryear to the more dynamic and much higher volume user generated content of today. This is often where product opinions are formed and this is also where customer loyalty or disloyalty can form as the customer support experience shows. This is what I call high value content. The following chart shows that MT will play a critical role in making this content more visible because it is high value and because of the sheer volume.
Shift to Dynamic

I also found another powerful argument for any multicultural society like the US and UK in this paper by Julia Alanen. She points out that language barriers keep 25 million non-English speakers deprived of critical government services (in the US) and that this also affects the rest of the population. While she focuses on the need for translators and interpreters, the content explosion is hitting this sector too. She point out:
Deprivation of plenary language access undermines human dignity, exacerbates many immigrants’ innate vulnerabilities, and harms society at large by impeding the efficacy of the healthcare and justice systems.
Getting back to the TCD conference, I was glad to see that several people (LSP leaders) asked me how they could learn more about MT and get more engaged with the technology. AMTA is proactively reaching out and trying to connect to the ATA by timing their conference to expand collaboration with the ATA. This is heartening to see and quite a contrast to negativity and the dueling conferences we see in other parts of the localization industry.

I also saw a quote from June Cohen, Executive Producer of TED Media at SXSW that I think is pretty wonderful (even though it may be naive and idealistic) when she was asked "What technology would you like invented? Or uninvented?"
"Instantaneous, accurate translation online. Nothing would do more to promote peace on this planet." 

Change is coming, and what are initially seen as threats can often be opportunities when one changes one’s own viewpoint. So here’s to change that creates more opportunity. Cheers.

Monday, March 15, 2010

The Ongoing Quest for “Best” MT Translation Quality

MT has been in the news a lot of late and professionals are probably getting tired of this new hype wave. Major stories in The New York Times and the Los Angeles Times have been circulating endlessly – please don’t send them to me, I have seen them. 

There is also another initiative by Gabble On which asks volunteers to evaluate Google Translate, Microsoft Bing, and Yahoo Babel Fish translations. And bloggers like John Yunker and many others have posted the preliminary results to that perennial question “Which Engine Translates Best?” on their blogs.

This certainly shows that inquiring minds want to know and that this is a question that will not go away. It is probably useful to have a general sense from this kind of news but does this kind of coverage really leave you any wiser and more informed?

Without looking at a single article or any of the results, I can tell you that the results are quite predictable, based on my very basic knowledge of statistics. Google is likely to be seen as the best simply because they have greater coverage of what the engines will be tested on and have probably crawled more bilingual parallel data than everybody else added together. I think the NYT comparison clearly suggests this. But does this actually mean that they have the best quality?

I thought it would be useful to share “more informed” opinions on what these types of tests really mean. Much of what I gathered can be found scattered around the ALT Group in LinkedIn so as usual I am just organizing and repurposing.

My personal sense is that this a pretty meaningless exercise unless one has some upfront clarity on why you are doing this. It depends on what you measure, how you measure, for what objective and when you measure. On any given day, any one of these engines could be the best for what you specifically want to translate. Measuring random snippet translations on baseline capabilities will only provide the crudest measure that may or may not be useful to a casual internet user but completely useless to understanding the possibilities that exist for professional enterprise use where you hopefully have a much more directed purpose. In the professional context knowledge about customization strategies and key control parameters are much more important. The more important question for the professional is: Can I make it do what I want relatively well and relatively easily?
FreevsCustom

The following are some selected comments from the LinkedIn MT group that provides an interesting and more informed (I think so anyway) professional perspective of this news.
Maghi King said: “The only really good test would have to take into account the particular user's needs and the environment in which the MT is going to be used - one size does not really fit all.”
Tex Texin said: “Identifying the best MT by voting will only determine which company encouraged the largest number of its employees to vote.”
Craig Myers said: “MT processes must be "trained" to provide desired outputs through creating a solid feedback loop to optimize accurate outcomes over time. Benchmarking one TM system against another is a fairly ridiculous endeavor unless you accept a very limited range of languages, content, and metrics upon which to base the competition upon - but then a limited scope negates any "real world" conclusions that might be drawn about languages and/or content areas outside of those upon which the competition is based.”
Alon Lavie AMTA President & Associate Research Professor at Carnegie Mellon University has, I think some of the most useful and informed things to say (follow the link to read the full thread):
The side by side comparison in the NY Times article is NOT an evaluation of MT. These are anecdotal examples. You could legitimately claim that the examples are not representative (of anything) and that casual users may draw unwarranted conclusions from them. I too think that they were poorly chosen. But any serious translation professional should know better. I can't imagine anyone considering using MT professionally drawing any kind of definite conclusions from these particular examples.

The specific choice of examples is not only biased, but also very naive. Take the first snippet from "The Little Prince". Those of us working with SMT should quickly suspect that the human translation of the book is very likely part of Google's training data. Why? The Google translation is simply way too close to the human reference translation. Translators - imagine that the sentences in this passage were in your TM... and Google fundamentally just retrieved the human translations.”

Ethan Shen of Gabble On “is hoping to be able to detect predictive patterns in the data that he could use to predict future engine performance. But he has no control over the input data (participants choose to translate anything they want), and he's collecting just about no real extrinsic information about the data. So beyond very basic things such as language-pair and length of source, he's unlikely to find any characteristics that are predictive of any future performance with any certainty whatsoever. 

What can be done (but Ethan is not doing) is to use intrinsic properties of the MT translations themselves (for example, word and sequence agreement between the MT translations) to identify the better translation. In MT research, that's called "hypothesis selection". My students and I work extensively on a more ambitious problem than that - we do MT system combination, where we attempt to create a new and improved translation by combining pieces from the various original MT translations. Rather than select which translation is best, we leverage all of them. We have had some significant success with this. At the NIST 2009 evaluation, we (and others working on this) were able to get improvements of about six BLEU points beyond the best MT system for Arabic-to-English. That was about a 10% relative improvement. That was a particularly effective setting. Strong but diverse MT engines that each produce good but different translations are the best input to system combination.”
So while these kinds of anecdotal surveys are interesting and can get MT some news buzz, be wary of using them as any real indication of quality. They will also clearly establish that humans/professionals are needed to get real quality. The professional translation industry has hopefully learned that the “translation quality” question needs to be approached with care or you end up with conflation at best and a lot of  mostly irrelevant data. 

My best MT engine would be the system that does the best job on content I am interested in on that day. So I will try 2 at least. The best for professional use has to be the system that gives users steering control, and the ability to tune an engine to their very specific business needs as easily (and cost-effectively) as possible and helps enterprises build long term leverage in communicating with global customers.