Pages

Thursday, February 11, 2010

Translation Humor & Mocking Machine Translation

I often run into blogs by translators and LSPs or just regular people who suggest that machine translation is not quite ready. In fact some people actually, believe it or not, mock MT. So while I do believe that MT is going be very much a part of the translation landscape in the very near future, I thought it would be fun to pick some of my favorite examples of MT gone awry. 

While MT mishaps can be funny I still think that humans, especially silly humans can do better, and my first example is by Ben who translated this Bollywood song and just put down what he thought he heard. I speak the language (Hindi) she is singing in and I laughed till I cried. In fact I can’t stop smiling as I type this. For those who want to know, “meheboob mereh” actually means my beloved.  

These are Ben’s own words on what he was trying to do:
My translation of an Indian music video. This is what I think the words sound like. www.bugben.com

Translation Party is a popular site that uses a familiar technique used to make MT look bad. You keep translating the same phrase back and forth and perhaps even across various languages to make sure that you make MT output that is really bad. Interestingly there are some “MT consultants” who also use this technique to test MT technology. A pointless exercise if you are serious, but can be great fun if you are just playing. So in my test,  <practice makes perfect and there is no substitute for hard work> was translated as <after working really hard Substitute>. Interestingly it was 100% accurate on <please do not poop on my knee> and gave me the same phrase back. I think that shows that when it really matters it can get things right.

Another personal favorite of mine is from Jill Sommer who had this little gem on her blog. Here is a tiny movie with a dialog developed completely from MT round tripping. As she describes it:
This fine little film by Matt Sloan capitalizes on Babelfish for its dialog. It translates to and from English, French and German. It was filmed on location in Trouville, France. Enjoy!

Mark Liberman in his Language Log blog shows this little furry iPod docking station gadget and with the following description which is suspected to be machine translation:
iMini is built in the rhythm decoding chip MJ1191 of the programming embedded system, and to integrate the HIPS skeleton; No matter you play any kind of music, MJ1191 always make your pet in dancing for you at once.

Another site that is always good for a laugh is Engrish.com. These are examples of mostly Chinese and Japanese attempts at translation into English. And this restaurant sign is one I often use in my presentations to show what MT is without human translator involvement. If you have not ever looked at the site, it is quite funny http://engrish.com/ . Here is one that is fun. I am told that there is a site in Japan with funny Japanese phrases from foreigners and I am sure the Chinese are laughing at us too. Just take a look at some of the strange Chinese character tattoos.

Here is a blog that specializes in finding strange translation examples from across the world. http://www.lostintranslationbook.com/. Here are some examples we collected from around the world and put on our website (in the left column).  

Anyway while I do laugh at these examples, I do believe the technology is improving all the time and as they say, he who laughs last, laughs the loudest. 

Let me know if you find other fun stuff and if I like it I will add it to this entry or create another entry with the best examples that people find.  Let's focus on really funny and not just wrong, since that would be like laughing at Sarah Palin.



P.S. The Huffington Post found some funny subtitles: Lost In Translation: When Subtitles Go Wrong

I also found another site of mostly human translation gaffes but I thought I would continue to add the best links I find over time to this entry.

 
Thanks to The Full Blog

And a few more from the Globalization Group and here is an explanation on why the translation industry is "hella lame".

And of course Monty Python with their Dirty Hungarian Phrasebook.

And for those of you who don't speak hip-hop, here is an excellent translation of the song My Hump.

This is a late insertion and shows you how human beings are always  SOOOOO much funnier than anything that MT could dream up. 60 Unintentionally Offensive Business and Product Names - Anybody want to buy some Asshoe shoes, or try some of that tasty Fart juice that goes really well with JussiPussi rolls and Shitto sauce?
 



Wednesday, February 10, 2010

Making Customer Support Content Multilingual

One of the largest new opportunities for the professional translation industry is in the Customer Support departments of high technology or industrial engineering global companies. I have briefly described the reasons why, but I thought that it would be worth elaborating on this further. 

Many global companies, especially those that are members of the Consortium for Service Innovation realize that a major new trend that they face is the growing power of the community and self-service in the Web 2.0 age. We see that already 98% of customer support interactions of an average global high tech company happen in self-service and the community forums. This is a major shift. However, the focus and much of the resource allocation in companies is still on the call support center and relatively static documentation and content which the professional translation industry is involved with. This makes less sense every day as evidence suggests that the customer experience is often formed by how support problems are handled. As the CSI points out, this is done mostly by self-service content and the “community” outside of corporate control.

GCSE Non-Anglo

If English speaking customers choose to solve their problems in this way, it follows that most global customers will also want to do the same. However, the content available to the global customer is often a fraction of what an Anglophile can get to and so non-English speakers are often left frustrated. There is now clear evidence of the following:

-- The customer support experience is increasingly formed outside the call center and customers strongly prefer self-service and community support.
-- The support experience is often critical in forming customer perceptions and developing brand loyalty.
-- Global customers do not have as much local language information access and thus probably have a less satisfying support experience.
-- Making much more knowledge and product support content available is a key to generating a better support experience.
-- Good self-service knowledge base content and greater visibility to high quality community content can greatly enhance the customer experience.
-- There is a clear relationship between customer loyalty and increased revenue and probably repeat purchase.

This situation presents a significant opportunity for the professional translation industry. However, given the huge volumes of content that need to be made multilingual it is important and necessary that automation be a key component of the multilingual content development strategy.  Microsoft was a pioneer in doing this, and they showed that hundreds of millions of customers were willing to use machine translated knowledge base content. Until recently they had all the knowledge base content available in at least 9 languages and are expected to expand this to more languages in future.

The benefits of global enterprise making large amounts of support content multilingual are significant, both financially and in terms of positive customer as the following graphic shows. Not only is self-service content a HUGE cost saver it can also create real positive brand perceptions. 

Call Deflection Benefit

The highest quality content production process will always need a high degree of human steering and expert linguistic guidance. Machine translation without humans may not provide the translation quality necessary to provide a positive support experience. The professional translation industry has a major opportunity ahead as major global corporations begin to act on this trend.

The role of customer support is shifting from answering questions and solving customer problems to facilitating a network of people and content. The Consortium’s research shows that the majority of the customer support experience is with content; not with people. The benefit of offering that content in the language of the customer is huge.

I will continue on this theme for at least another blog entry. I strongly recommend that you take a look at the CSI website as it is filled with great information and research on what is going on in the world of Customer Support.

Quote from the CSI web site:
Rather than continuing to invest in doing what we do faster, better, and cheaper, maybe we need to look at doing something altogether different... maybe there is a lesson for the localization industry in this.

Monday, February 1, 2010

The Impact of “Clean Data” on SMT

This is a summary of a study on translation memory consolidation that I was involved with and a continued examination of the issue of “clean data” which I believe is essential to long-term success with data-driven MT initiatives.The Asia Online rating system rates data that is deemed to be best suited for SMT training and is not a judgment on TM quality for TM purposes. The study conducted by Asia Online took a relatively small set of TM data with the kind facilitation of TAUS and data provided by three members and attempted to answer three questions:
  1. Is there a benefit to sharing TM data for the purpose of building SMT engines?
  2. What are some practical guidelines to help enhance serious data sharing attempts ?
  3. What do best practices look like?
As many people continue to believe that sheer data volume alone is enough to solve many problems with SMT I thought it would be useful to provide an overview of the Asia Online data consolidation study in this blog. Apart from the simple common sense of the "Garbage In Garbage Out" principal which is important in any data processing applications and perhaps even more so in SMT, does it not make sense that if SMT engines learn from parallel corpus, it would be wise and efficient to clean this corpus first?


The Google paper that has also been referenced by TAUS/TDA as foundational justification for “more data is always better” has been criticized by many.  Jaap van Der Meer suggests that Norvig said “forget trying to come up with elegant theories and embrace the unreasonable effectiveness of data.” A more careful examination of the paper reveals that many of the examples in the paper are related to graphical image examples where erroneous pixels are much more tolerable.In fact, Norvig himself, has stated in his own blog that his comments were misinterpreted:


To set the record straight: That's a silly statement, I didn't say it, and I disagree with it. … Peter Norvig


So I maintain that both data quality and algorithms matter, and unless we are talking about huge magnitudes of order differences, clean data will produce better SMT engines and respond more easily to corrective feedback. I have seen this proven over and over again. We are all aware that TM tends to get messy over time and that it is wise to scrub and clean it periodically for best results in any translation automation endeavor.


Basically the study found that some TM is better suited for SMT and that it is important to understand this BEFORE you consolidate data from multiple sources. The graphic below shows the details of the data in question. Additionally we also found that Datasets A and C were more consistent in their use of terminology.


DataOvw
The key findings from the study are as follows:
  • -- Data quality matters and all translation memory data is not equally good for SMT engine development.
  • -- Data quality assessment should be an important first step in any data consolidation exercise for SMT engine development purposes.
  • -- MORE DATA IS NOT ALWAYS BETTER and smaller amounts of high quality data can produce better results than large amounts of dirty data.
  • -- Terminological consistency is an important driver for better quality and success with SMT. Efforts made to standardize terminology across multiple TM datasets will likely yield significant improvements in SMT engine quality.
  • -- Introducing “known dirty data” into the system decreases the quality of the system and increases the unpredictability of the results 
  • -- Systems built with clean data and consistent terminology tend to perform better and improve faster

BLEUchrt

  • -- Data cleaning and normalization and terminology analysis and standardization is a critical first step to having success with any project that combines TM for developing SMT engines
In the noise about soft censorship we have gotten distracted from two additional questions that also are worth our attention.


What is the best way to store data so that it is useful for both TM and SMT leverage purposes?
What are the best practices for consolidating TM? What tools are necessary to maximize benefits?


Common Source of Data Problems in TM
Encoding problems and inconsistencies in the data.
Large volumes of formatting tags and other metadata that have no linguistic value but which can impact system learning.
Punctuation differences and diacritics are inconsistent across the TM
English words appear in the French translation. This may be valid in translation memory, but will result in English being embedded in the French training data and make it possible for the SMT engine to think that English is French!
Excessive formatting tags and HTML (often representing images) embedded in segments
Frequently the French translations had bracketed terms that were not present in the English source.
Frequently the capitalization does not match.
Terminology was sometimes inconsistent, with different terms being used through the data for the same concept or meaning
Large number of variables in many different forms embedded in the text. Variable forms and formats are inconsistent.
Multiple sentences in one segment. While this is valid, the job of word alignment becomes more complex. Higher quality SMT word alignment can be achieved when these are broken out into their individual segments.
In French text, there are frequently abbreviations when there should be a complete word.
Words missing on either side.


These are early days and we are all still learning, but the tools are getting better and the dialogue is getting more useful and pragmatic as we move away from naive views, that any random pile of data is better than one that has been carefully considered and prepared.


Without understanding the relative cleanliness and quality of the data, data sharing is not necessarily beneficial.


While TM data may often be problematic for SMT in its raw state, some of what is considered “dirt” to SMT can be cleaned through automated tools used by Asia Online and others. However, these tools cannot correct situations when the translations themselves are of a lower quality. This issue has also been highlighted by Don DePalma in a recent article referring to this study where he said: “Our recent MT research contended that many organizations will find that their TMs are not up to snuff — these manually created memories often carve into stone the aggregated work of lots of people of random capabilities, passed back and forth among LSPs over the years with little oversight or management.”


Lets hope that the TDA too will let this taboo subject (data quality) out into the open. I am sure the community will come together and help develop strategies to cope and overcome the initial problems that we identify when we try and share TM data resources
.