Pages

Monday, June 27, 2016

MT Options for the Individual Translator


This is yet another post triggered by conversations in Rio at the ABRATES conference in early June. As I mentioned in my initial conference post, the level of interest in MT was unusually high and there seemed to be real and serious interest in finding out ways to engage with MT beyond just the brute force and numbing corrective work that is typical of most PEMT projects with LSPs.

MT has been around in some form for decades, and I am aware that there have always been a few translators who found some use for the technology. But since I was asked by so many translators about the options available today I thought it would be useful to write about it. 

The situation today for many translators is to work with low quality MT output produced by LSP/enterprise MT practitioners with very limited MT engine development experience, and typically they have no say in how the MT engines evolve, since they are so far down the production line. Sometimes they may work with expert developed MT systems where there is some limited feedback and steering possible, but generally the PEMT experience involves: 

1.    Static MT engines that are developed offline somewhere, and that may have periodic if any updates at all to improve the engine very marginally if at all.
2.    Post-editors work on batches of MT output and provide periodic feedback to MT developers.

This is beginning to change in the very recent past with innovative new MT technology that is described as Adaptive Interactive Dynamic Learning MT (quite a mouthful). The most visible and elegant implementation of this approach is from a startup called Lilt. This allows the translator-editor to tune and adjust the engine dynamically in real time, and thus make all subsequent MT predictions more intelligent, informed and accurate. This kind of an MT implementation is something that has to be cloud based to allow the core engine to be updated in real time. Additionally, when used in workgroups, this technology can also leverage the individual efforts of translators by spreading the benefit of a more intelligent MT engine with the whole team of translators. Each user benefits from the previous edits and corrective actions of every other translator-editor and user as this case study shows. This allows a team to build a kind of communal edit synergy in real time and theoretically allows 1+1+1 to be 5 or 7 or even higher. The user interface is also much more translator friendly and is INTERACTIVE so it changes moment to moment as the editor makes changes. Thus you have a real-time virtuous cycle which is in essence an intelligent learning TM Engine that learns with every single corrective interaction. 

CSA tells us that the SDL Language Cloud also has similar abilities but my initial perusal suggests it is definitely less real-time, and less dynamic than Lilt i.e. it is not updating phrase tables in real time. There are several little videos that explain it and superficially it looks like an equivalent but I am going to bet it is not yet at this point in time anyway.

So for a translator who wants to get hands on experience with MT and understand the technology better, what are the options? The following chart provides a very rough overview of the options available ranked by my estimation of the best options to learn valuable new skills. The simplest MT option for an individual translator has always been a desktop RbMT system like Systran or ProMT, and it still is a viable option for many, especially with Romance languages. But there has never been much that could be done to tune these older systems beyond building dictionaries, a skill that in my opinion will have low value in the future.

 
Good expert developed MT systems will involve frequent interaction with translators in the early engine development phases to ensure that engines get pattern based corrective feedback to rapidly improve output quality. The more organized and structured this feedback and improvement process, the better the engine and the more interesting the work for the linguist.

Of the “free” generic MT engines Microsoft offers much more customization capabilities and thus are a good platform to learn how different sets of data can affect an MT engine and MT output. This of course means the user needs to have organized data available and an understanding of the technology learning process. MT can be trained if you have some understanding of how it learns. This is why most Moses experiments fail I think, too much effort focused on the low value mechanics, and too little on what and why you do what you do. I remain skeptical about ignorant Moses experimentation because getting good SMT engines requires good data + understanding of how your training corpus is similar to or different from the new source that you want to translate, and a variety of tools to help keep things synced and aligned when you see differences. I am also skeptical that these DIY efforts are likely to get as engines as good as the free generic engines, and I wonder why one would bother with the whole Moses effort if you could get it at higher quality for free from Microsoft or Google. There are some translators who claim some benefit from working with Moses and its desktop implementations like Slate. 

All these options will provide some experience and insight into MT technology, but I think it is useful to have some sense for how they might impact you from two key perspectives shown below:

1.    What options are likely to give you the fastest way to get to improved personal productivity?
·         I would expect that an Adaptive MT solution is most likely to do this the fastest, but you need have good clean training data – TM, Glossaries (the more the better). Also you should see your edit experience improve rapidly as your corrective feedback modifies the engine in real time.
·         Followed by the Microsoft Translator Hub (if you do not have concern about privacy issues), SDL Language Cloud and some Expert systems which are more proactive in engaging translators but this will also involve an LSP middleman typically.
·         Generic Google & Microsoft and Desktop RbMT (Romance languages and going to English tend to have better results in general).
·         DIY Moses is the hardest way to get to productivity IMO but there is some evidence of success with Slate.

2.    What options are most likely to help develop new skills that could have long-term value?
·         My bet is that the SMT options are all going to help skills related to corpus analysis, working with n-grams, large scale corpus editing and data normalization. Even Neural MT will train on existing data so all those skills remain valuable.
·         Source analysis before you do anything is always wise and yields better results as your strategy can be micro-tuned for the specific scenario.
·         Both SMT and the future Neural MT models are based on something called machine learning. It is useful to have at least a basic understanding of this as this is how the computer “learns”. It is growing in importance and worth long-term attention.

There are tools like Matecat that show promise, but given that edits don’t directly change the underlying MT engine, I would opt for a true Adaptive MT option like Lilt instead.  


  What do we mean by PEMT in the larger context?

The traditional understanding of PEMT can be summarized in the graphic below and this is the most common kind of interaction that most translators have with MT if they have any at all.

 
However the problem definition and the skills needed to solve them are quite different when you consider an MT engine from a larger overall process perspective. It generally makes sense to address corpus level issues before going to the segment level so that many error patterns can be eliminated and resolved at a high frequency pattern level. It may also be useful to use web crawlers to gather patterns to guide the language model in SMT and get more fluent target language translations. 

The most interesting MT problems which almost always happen outside the “language services industry”, require a view of the whole user experience with translated content from beginning to end. This paper describes this holistic user experience view and the resultant corpus analysis perspective for an Ebay use case. Solving problems often involves data acquisition of the right kind of new data, normalization of disparate data, and focus on handling high frequency word patterns in a corpus to drive MT engine output quality improvements. This type of deeper analysis may happen at an MT savvy LSP like SDL, but otherwise is almost never done by LSP MT practitioners in the core CSA defined translation industry. This kind of deep analysis is also often limited at MT vendors because customers are in a hurry, and not willing to invest the time and money to do it right. Only the most committed will venture into this kind of detail, but this kind of work is necessary to get an MT system to work at optimal levels. Enterprises like Facebook, Microsoft and EBay understand the importance of doing all the pre and post data and system analysis and thus develop systems that are much more closely tuned to their very very specific needs.





MT use makes sense for a translator only if there is a productivity benefit. Sometimes this is possible right out of the gate with generic systems, but most often it takes some effort and skill to get an MT system to this point. It is important that translators have a basic understanding of three key elements before MT makes sense:
1.    Their individual translation work throughput without MT in words per hour.
2.    The quality of the MT system output and the ability of the translator to improve this with corrective feedback.
3.    The individual work throughput with MT after some effort has been made to tune it for specific use.

Obviously 3 has to be greater than 1 for MT use to make sense. I have heard that many translators use MT as way to speed up the typing or to look up individual words. I think we are at a point where the number of MT options will increase and more translators will find value. I would love to hear real feedback from anybody who reads this blog as actual shared experience is still the best way to understand the possibilities.

 

Monday, June 20, 2016

The Larger Context Translation Market

This is another post inspired by my recent visit to Rio and dialogue at the ABRATES conference, where translators were eager to engage with MT in a meaningful way, and asked many questions about where the most interesting translation challenges were. In several conversations about “the translation market” I had a very clear sense of how there really needs to be a larger perspective on what this means, as the most interesting opportunities with MT tend to lie outside what is generally understood as the translation industry. 

For most of the people who attend “translation industry” and localization conferences, the most trusted description of the industry is the market that Common Sense Advisory (CSA) describes as the Language Services Market. A market where translation agencies provide translation, localization and interpreting services to buyers for a fee. This is a market that is estimated by CSA to be $38.16 Billion in 2015, with Lionbridge proudly claiming to be a “perennial list-topper” and the largest language service provider (LSP).  Their PR piece provides a clear description of what the CSA market definition covers. This link lists the Top 20 LSPs globally, measured by total revenue. Here is another view of the CSA market summary that shows geographical concentration of the industry.
*The Language Services Market 2015, Common Sense Advisory Research, June 2015
  This CSA sourced graphic from the Lionbridge blog describes the fastest growth segments in the language services market. Clearly, translation, software and website globalization are at the top of the growth list for this type of paid translation service.


The following graphic looks more closely at what the focus of the LSP Translation industry is, and we see the kind of content they focus on, and the tools and skills most relevant to addressing translation of this type of content. Thus, project management is the core business function and the most important tools are TMS systems and TM.

While some LSPs do use MT, it is generally not a mission critical tool. MT is used if the LSP is able to pull together an MT system that improves productivity and reduces costs, and is largely reactive i.e. often because the client insisted. But, it is important to understand that the focus is still on the same kind of content shown above. The translation industry has mixed success with MT, and maybe a few systems do become integral to the overall translation production process. But most do-it-yourself LSP MT initiatives fail or wallow in a kind of confused and isolated geekdom with mediocre results.  MT systems that consistently produce excellent output are the hardest to develop, so it is somewhat ironic that those that understand the least about how the technology works, try and build the most capable and efficient MT systems. The MT experience that Jost and other translators describe in blogs,  described as PEMT, MpT, MT+PE etc.. is presented as the great evil by IAPTI, or horrid commoditization of the work of translation by many others. Most often they are working with MT systems at arms length, and have no ability to steer or guide the  MT system development to make it more useful.  Hopefully, the notion of Moses-as-instant-magic is now widely understood as a limited success strategy, and the more savvy enterprises and LSPs leave it to experts, who also struggle to meet these consistent high quality output goals. Good MT systems will always take time, expertise and articulate linguistic feedback to develop.   

However, I think the new Adaptive Dynamically Learning MT that Lilt is producing has a very bright future with smart LSPs, and provides a platform to transform the MT experience into a much more predictable and worthwhile endeavor and will also allow translators to be much more engaged and involved in steering the MT system.

The Larger Context for Value Added Translation

 

Though the “translation industry” MT experience is mixed, I would argue that MT has been responsible for driving revenue or definable economic value in a variety of non-traditional scenarios, on a scale that dwarfs “the translation industry” as defined above. Interestingly, the most successful implementations of MT are done by global enterprises who often still work with LSPs for the static structured content, but for the higher value unstructured dynamic content are largely choosing to do it themselves (sometimes with some expert help) or build internal teams to address what they see as a long-term and strategic need to enhance international business initiatives.  It is useful to consider some case studies of these strategic MT initiatives by enterprises, to  understand this better.
 
The graphic above shows the volume of words that Google translates every single day with their MT systems as reported at Google I/O in April 2016. To put this in context, I saw a Lionbridge presentation a few years ago, where the CFO said they translated just over a billion words that year (2009). SDL who is probably the most MT savvy and active with MT LSP, recently claimed they are doing 20B words per month through MT. So, it is quite possible that Google alone, translates more words a week than the whole “translation industry” does in a year. When one considers that perhaps over 90% of Google’s revenue (~$78B/year) is generated from advertising linked to key words, it is quite possible that Google derives tens of billions of dollars from their MT technology initiative! It also gives them very specific intelligence on what matters to people across the world and what cross language content is the most sought after. The economic value of this knowledge is significant, and hidden in the advertising revenue they report. This knowledge of what matters across the globe is something the “translation industry” and SEO experts would love to know.

is another example of high value derived from MT. Their initial entry into MT was like Google related to search, but additionally they had a massive knowledge base in English for their software products that was difficult for their substantial global customer base to efficiently access. Thus, while Microsoft probably spent hundreds of millions on “translation industry” services for static content, this only covered a tiny fraction of what they needed to translate. Given that Microsoft gets as much as 70% of their revenue from non-English speaking countries, translation of all kinds of product related content is important. Making more technical support and customer care content rapidly multilingual was an imperative for executives who cared about the customer experience, and also generated huge savings in support costs and dramatically improved the user experience for the non-English speaking customer. The software industry measures the value of self-service content by something called deflection cost. So, if they can deflect a call to the support center, by making more knowledge base content available in more languages,  using MT, they can save possibly as much as a $100M+ per day given the size of their user base and actual volumes of knowledge base access. Add maybe another 50B words/day that their Bing MT does for the random internet user, and we have another stream of economic value coming from search words that generate advertising revenue across the globe. Their recent Skype STS initiative also will likely yield great benefit and new ways to monetize their translation technology expertise.

When you consider that both Intel and Adobe also use the Microsoft Hub MT to translate knowledge base support content, the deflected cost savings impact is easily worth hundreds of millions of dollars a day. This is not even considering the many other IT companies doing this on their own using other MT technology e.g. Symantec. The “translation industry” has a very small footprint in this kind of translation activity,which is now often considered mission-critical and probably involves several billion words per month.

The online eCommerce market is another example of economic value generated by competent MT efforts that is off-the-books of the “translation industry”. EBay decided some years ago that emerging economies were a huge opportunity worth strategic attention. So they acquired MT technology and built a competent MT team that had astrong linguistic collaborative component in the team. Based on presentations they made at the AMTA 2014 conference it was clear that there was a huge growth impact in the Russian market from their initial efforts to to make more Russian content available. It would be safe to say that the value of the impact is probably in the hundreds of millions of dollars of new revenue, from all the new markets that they have been addressing using MT. It is also really worth taking a look at what is involved in doing this. It takes focus on solving new kinds of translation problems and making sure the translation problems you solve do enhance the value of your MT efforts. This last link shows the special issues related just to Brazilian Portuguese. We should note that most competent MT efforts of any scale, move carefully, one language at time rather than trying to do 20 or 30 in a single go. We also see that Amazon acquired Safaba in 2015 and possibly have similar plans to make catalogue content multilingual to drive bigger volumes of international business. Alibaba and Baidu also have eCommerce focused MT efforts well underway but fewer details are available. The net economic value of all these type of MT initiatives: Probably in excess of $20B per year by my very rough estimates.

Recently Facebook surprised the world by announcing that they have a substantial MT effort underway after using the Microsoft Bing MT technology for several years. When asked why they did this Alan Packer said: Scale is one reason Facebook has invested in its own MT technology. The other reason is adaptability, they wanted technology that was optimized for their very specific needs. Facebook is now serving 2 billion text translations per day. The problem they had with Bing they claimed was, it was built to translate properly written website text and did not do well with the slang, metaphor and idiom typical in Facebook comments. Packer described Facebook language as “extremely informal. It’s full of slang, it’s very regional.” He said it is also laden with metaphors, idiomatic expressions, and is riddled with misspellings (most of them intentional). Additionally, as in the rest of the world, there is a marked difference in the way different age groups communicate on Facebook. They know that already 50% of Facebook users regularly use auto translation. This user group will only  grow as more people come online. Packer says that access to the translation product leads users to “have more friends, more friends of friends, and get exposed to more concepts and cultures.” The more people across the world that Facebook users can connect with, the longer they’ll spend on the social network, and the more revenue-earning ads they’ll see. The economic value of this is probably several billion dollars a year. Emerging social networks that are global will need to address the same problem. 


So what we see is that a select few companies are generating more economic value from solving very specialized and much more challenging translation problems than the whole gross revenue output of the “translation industry”.   Solving large-scale translation problems using MT is apparently a very high value proposition, and none of the global enterprises mentioned above considered going to the “translation industry” to help solve these really challenging and complex translation problems. Probably because it is very clear to any strategic observer at these companies, that most LSPs lack the vision, skill, interest and competence to solve these types of translation problems.  We can perhaps even generalize the core requirements for a larger set of global enterprises as shown below. There is a real mismatch in terms of skills and focus between the broader translation needs of global enterprises and the service focus of the “translation industry”. While the static content will likely remain important as a mandated requirement, it is not where long-term corporate value is built either for the enterprise buyer or the LSP in my opinion.
To illustrate this further let us consider the investor sentiment on value that can be gleaned from stock market data. While this might be a stretch of logic to some, I think we can fairly assume that investors value solutions to certain kinds of translation problems more than others. Facebook has seen a huge growth in mobile ad revenues and it seems that they are taking ad share away from Google recently, and so these Market Value/Sales numbers reflect very active market trends. The investor sentiment is that Facebook is probably better poised to gain $$$s from the next wave of internet adoption than any of the companies listed in the chart as they climb beyond 2B users, very few who speak English or French or German. As one analyst says: "Advertising budgets are moving towards Facebook, and it seems to be a winner in the online advertising world with measurable results." Contrast this with the investor sentiment for large LSPs, surely, it has something to do with long-term promise and potential. To me this suggests that investors in general view the LSP focus as lower in value but understand that translation can produce huge leverage in the right hands.

So to those wonderful translators at ABRATES who asked me what kinds of MT projects to get involved with, I would say the following:
  • Focus on companies who are solving interesting translation problems. They will have the most rewarding work and it might involve stepping out of the translation industry.
  • Stay away from LSPs who sell the Moses Mirage, this is likely to be the worst PEMT experience. MT systems that don't want translator feedback at a pattern level are not likely to be a professionally satisfying experience.
  • Work with people (LSP/Enterprise) who allow and want you (translator/editor) to provide feedback and interact with the MT development process.
  • Learn about Machine Learning and AI in other domains, and develop skills with Regex, Corpus Analysis & Corpus Editing and Pattern Identification skills to be considered valuable.
  • Explore Adaptive Dynamic Learning MT like Lilt (maybe others will appear soon)  to understand how MT can work with you and for you while you wait for the right opportunity. This is truly a paradigm shift that is worth at least some experimentation to see how the translator desktop could evolve.
  • Ease up on the need to have everything on the desktop. The future of Machine Intelligence solutions will require big data and big computing, so the best and most sophisticated tools will by definition only be available in the cloud. Lilt is a first generation example of this, others are coming. The cloud makes sense and allows new, more effective ways to solve old problems and is not a bad thing.
And for those who think the four companies above are the exception rather than the rule, it is worth noting that we are just beginning with what is possible to do with Machine Learning and Artificial Intelligence. Neural MT is just beginning and could drive a whole new wave of higher quality and more adaptive MT. The Machine Intelligence market is still nascent and we will very likely see big data + big computing + smart algorithms come together, to solve problems that we thought were beyond the scope of computers just last year.

Tuesday, June 7, 2016

The ABRATES Conference in Rio: Translators focusing on MT

I had the honor of participating in the 7th ABRATES International Translation and Interpreting Conference in Rio de Janeiro last week. An event that had over 500 attendees, based on my casual observation. A large portion of the attendees were translators, but there were also some LSPs and Enterprise representatives. As much of the information was presented in Portuguese I had direct experience with simultaneous translation via a headset which was also kind of cool, and it was fun to switch around when I was less interested in the actual subject matter.


The formidable, emotion packed sign language interpretation by Paloma Bueno, intensely focused simultaneous interpreter volunteers in the booth, and the abounding loveliness of Rio.

I found the conference surprisingly refreshing for several reasons including:
  • The high level of understanding that many translators had about MT, Post-Editing practice and their general attitude that it is better to understand and use translation technology than fight it or fear it.
  • The beautiful location, as Rio is a naturally scenic and inviting spot.
  • An emotionally powerful sign language interpretation of the keynote session by Paloma Bueno who I cannot believe was doing this in real time.
  • The eagerness and openness of many translators present, to try and understand how they as translators could engage and work with MT and develop meaningful expertise in MT related skills.
  • The willingness to explore and understand how translation technology will continue to evolve and possibly impact their professional work.
  • Several conversations with translators who had long term experience with MT and thus had direct knowledge of MT systems that improved over time and had also seen both good and bad MT engines over the years, so were much more coherent in their criticism.
  • The shared experience of many different kinds of MT encounters from a variety of translators, ranging from DIY horror, experts systems that slowly evolved in quality gradually over years, and some proprietary efforts that produce astonishing quality.  
  • The presence of several very competent presentation sessions on developing MT related skills including:
    • Corpus Preparation for MT training
    • Working with the varying quality of PEMT output that translators get from LSPs
    • Using REGEX (Regular Expression) to develop more powerful text based editing skills when deal with corpora
    • PEMT best practices and tools and shared experiences
I also found this conference special, because I personally had no corporate allegiance at the event and was truly just an independent spokesperson with some knowledge of MT technology and it’s potential and place within the context of many of the attendees professional lives. As I am no longer employed or affiliated with Asia Online I felt very comfortable sharing my opinions, with no concern about persuading anybody to go one way or another. My opinions were all truly independent and the truest expression of what I try and do in this blog, i.e. provide useful and relevant information to inquiring minds. So while I am indeed looking for professional work, I am really enjoying this independence and focus on what really matters.

It was interesting to find that when one has this kind of openness and lack of bias as a presenter, there is an opening of the perception, and I was able to see much of what I was saying with a new and fresh eye. It was like playing improvised music to a keen and attentive audience, the shared attention of the musician and the audience creates a new, more evolved, version of an existing musical idea. I will share some of those insights in upcoming posts.

I also understood much more clearly that most often, translators have very little control of the content they are given to translate, because of the current structure of the professional translation business which is usually: Enterprise > MLV (Big Agency) > SLV (Small Agency) > Translator. Thus translators are often left to deal with poor quality source which cannot by contract be corrected or changed, work with crappy MT output produced by DIY practitioners who do not know how to actually do it themselves, or have no say in how the MT engines evolve since they are so far down the production line. Thus we have the current situation of unnecessarily mind numbing PEMT work, rather than evolving and rapidly evolving MT technology from more efficient production processes. And very often the extremely valuable linguistic feedback that translators provide is lost or ignored. An MT paradigm that organizes and collects valuable translator feedback will surely be more competitive and produce higher quality and benefit to all concerned. Not to mention that it will be personally rewarding for the many translators who will need to be involved, as the nature of the problems they solve will evolve in value and impact from the typical LSP project.


Plenary session on MT
 
I had an interesting experience during a plenary session panel on MT where all the other speakers were speaking in Portuguese, so I had to have a headset to understand what they were saying. When I started speaking, the interpreters of course started speaking in Portuguese, and I found it very strange and unsettling to hear a voice saying everything I was saying in English in Portuguese in real time. Somebody once said that MT is magic which I felt deserved some scorn, but to me this act of listening and translating into another language in the instant, not knowing what I was going to say, was surely closer to magic.

If this conference is an indicator of what is happening in the professional translation world, it is very promising for several kinds of translation technology initiatives. I have always felt, much to the chagrin of my former employers, that the real promise of MT will be seen when translators seek it out and learn to steer, drive and enhance the ongoing evolution of this technology. If this conference is really only representative of the Brazilian reality with translation technology, then I predict that the most exciting advances in MT will come from those working with Portuguese. This community is primed for the most interesting new Adaptive MT initiatives like Lilt which can empower motivated and technically savvy translators.

You can find some Twitter coverage of the event by searching on the hashtag #abrates16  or if you look up the following accounts:

http://twitter.com/AlberoniTrans 
http://twitter.com/oscarcurros 
http://twitter.com/sidney_barros

https://storify.com/oscarcurros/abrates16-rio  

  

Ipanema Street Market