Pages

Tuesday, March 8, 2011

The Changing Face of Localization (Professional Translation)

I was at a party recently where somebody asked a Language Service Provider, what they did in their professional work. It was amazing to witness how completely mystified the questioner was by the response which included the word “localization” several times. It took several minutes of conversation before the person (who admittedly was a little slow) gathered that it involved translation for business purposes in some way. To my view “localization” is not a great word, to get the general-world-out-there, engaged and interested in, or even just understand what you do. Looking at the localization entry in Wikipedia explains the confusion felt by the average guy on the street; the word has different meanings in translation, psychology, medicine, physics, mathematics and more recently even in location based services like FourSquare and Facebook Places.

Does it really matter? (I think it does, especially in interactions with people outside the translation industry, = the real world?) I have always felt that it is important to be able to communicate what you do, quickly and easily in casual social settings to enhance your professional life. Good casual social interaction can often lead to useful professional references and interactions, but only if people actually understand what you do. It matters even more when you as an industry are trying to increase your visibility to the world out there. I think the word localization may have made sense when the focus was only “software and documentation localization” (SDL), but this view of what we do is increasingly being questioned in terms of overall value to generating and facilitating international business.(BTW the word “transcreation”, to my mind is even worse in terms of obfuscation and classic HUYA-ness.)

Ironically, I had a brief Twitter exchange with Ultan O’Broin (aka @localization) discussing this shift.

(Unfortunately the service I used to show the conversation is now defunct. And since Twitter makes it so hard to get old conversations it is pretty hard to retrieve those snippets.) 


We were basically discussing data interchange standards in the translation industry (TMX, XLIFF) and Ultan said something that I thought was very insightful about the old SDL view of the business:”people don't get "structure". Obsessed with formatting, still”. This helps to explain the relatively low status of localization professionals in most global enterprises. The view is that, the localizers handle the translation production of carefully formatted material that goes into product packaging, and some pro-corporate, self-congratulating, mostly irrelevant content on the corporate web site. Thus, it is not surprising that localization professionals have kind of a secretarial status in most internationally focused business groups. They provide basic services and assistance to international business initiatives. As Ultan said, they have an administrative assistant view rather than a system administrator view on information flows related to international business initiatives.

As I have stated in previous posts the world is changing, and to stay relevant we need to also change what we do, how we do it and why we do it. At the executive level of global enterprises, it is increasingly becoming clear that customer decision making processes have changed, largely due to open and free access to more information. This information is increasingly created outside the global enterprise and is not easily controlled by stakeholders within the global enterprise. In many industries global customer conversations are MORE influential in driving customer behavior (and corporate sales) than corporate content. To be relevant, we need a new mindset that looks at the flow from information creation (internal and external) to information consumption and has an honest and real focus on the final customer. Real conversations with real users matter more than corporate content and some are beginning to realize this. Value needs to be defined by how useful a customer finds the information, not by how many translation and formatting errors there are in a user manual that few are likely to read. Ultan is at the leading edge of this new focus in an area called User Experience (UX) and thus we should all be listening to what he and others like him have to say.

Here is a more detailed overview on these broad changes from my viewpoint at a recent Localization ;-) Technology Roundtable seminar in Palo Alto:


An interesting aside: I was informed by an SDL Plc marketing representative that I would not be welcome at their recent SDL Innovate  event in Palo Alto because of “my position at Asia Online”, however they did admit that, “we will look forward to seeing you at future industry events.” To be honest, I did apply as Kirti Vashee, CEO of Maya Acoustics (which I truly am involved with). But unfortunately the expert marketing department sleuths there tracked me down as the author of this blog post. (Hmm, could it have been my name?)  Or perhaps because I think that associating SDL with Innovation is oxymoronic, or perhaps because I represent competition that is feared and formidable. Apparently Renato and people from TermWiki were also denied admission into the compound.

Interestingly the keynotes brought forth new supporting data for many of the themes I have been writing about in the last year (I think so anyway): Openness, Customer Focus & Collaboration, Standards and Information Flow. Maybe I am biased, but do you really see these as themes that resonate and receive meaningful commitment at SDL?

Some highlights (gathered from Twitter, thank you @scottabel) from Toby Bell of the Gartner Group:
  • People are becoming brands
  • More stuff is uploaded to YouTube in 60 days than all television networks combined have created in 60 years (Yes indeed, UGC is for real)
  • Everything is interactive. If you're not polling or offering live interactive contact with customers, you're missing out on the engagement opportunity
  • Manufacturers, retailers now allow customers to create documentation and interestingly this content is often better than what their own employees create
  • You must tune the experience (with the "right" content and information about your products and services) to the users goals
  • Corporate leadership still views web content management as a publishing function. It's not. It's really about customer experience.
Some highlights from Marcia Metz, EMC on Information Liquidity:
  • Information is an asset that can be turned into revenue
  • We can't keep pace with the growth of the volume of information and speed and efficiency are becoming more critical to business success
  • We are working to provide information as a service. Content creators and consumers should be able to collaborate for best results
  • Information liquidity requires a comprehensive platform that is standards-based, organized and manages the full information life cycle from co-creation to consumption
There was also another interesting Twitter based discussion that focused on the dubious value of TMS systems. While I do understand that translation projects have been messy historically, and that some level of automation is required to make things more efficient, I have many doubts about the "solution" that many have chosen. Most of my doubts are about relative value not absolute value. Why are there so many TMS systems? Why do they all have such small installed bases? Why does every LSP and Corporate Localization department think that their translation project management process is so unique that it can only be properly automated by creating a new TMS system?  Could this not be better accomplished by using more mainstream (= installed base of hundreds or thousands) collaboration and database tools? Jaap van der Meer of TAUS stated at the now defunct LISA's final standards summit event: GMS/TMS will disappear over time, in favor of plug-ins to other systems. Adam Blau surprised me at our technology roundtable meeting in Palo Alto when he said that milengo does not use or believe in TMS systems. The reason: Too much investment for too little return and a reduction in overall flexibility. You can see him say it in his own words here. He also provides some interesting observations on best-of-breed tools that they use, and the issues related to developing a specific technology agnostic strategy in his talk.

Also, the future of translation I think will see more deployments of collaborative communities or crowdsourcing. The Monterey Institute of International Studies (MIIS), announced recently that they’re deploying Lingotek’s hosted Collaborative Translation Platform. As we head into more 10 million and 100 million+ word translation initiatives, this kind of collaboration facilitating infrastructure becomes more and more necessary. It is interesting to see that few if any universities ever adopted translation memory and TMS tools into their curriculum in the past. Tools that enable and facilitate open collaboration and easily integrate into mainstream content management software infrastructure become ever more important. The people that lead the charge in these new initiatives are often not from the localization community and they seem to understand that data must flow freely for the tools to be useful.


We are at a point in time where it can be recognized that professional translation efforts focused on customer conversations are actually impacting overall international business success. Quite possibly we are at a point where what we do (enable and facilitate global customer conversations), is seen as driving customer satisfaction and customer loyalty and thus international revenues across the globe.

And if there is somebody out there who could educate me about personalization, I would love to learn more, as I think that personalization and mobile will also grow in importance as drivers for building international business. The coming shift to mobile is hopefully obvious to everybody, there are already 4X as many mobile (cell phone) users in the world as there are online users and we should expect that a lot of information consumption will shift to these devices as they get more powerful. This does not mean that PCs go away, but rather, that the conversation becomes more mobile and free form.

So if the future is about more free flowing conversations with customers and much more dynamic internal and external content, we should be thinking about new ways to describe what we do. I think our new description will likely include terms like professional translation, collaboration, global customer engagement and effective use of translation technology. While traditional localization work is unlikely to disappear, I think the best is yet to come and some of us will be lucky enough to be involved with world changing initiatives.


For a cool response on "what do you do?" see what that other localization expert had to say: "We are building technology that facilitates serendipity." --Dennis Crowley, Foursquare co-founder, as quoted by the Los Angeles Times. For me personally, I like the sound of: "We develop technology that enables global enterprises to talk to their customers across the world and we also help to address and alleviate information poverty in South East Asia". 

Tuesday, February 15, 2011

An Exploration of Post-Editing MT – Part I

The topic of post-editing MT (PEMT) (yes, somebody has already come up with an acronym) continues to gain prominence in the professional translation world and is also a subject of heated and sometimes disdainful discussion amongst many translators in the blogosphere. There is a lot of confusion and conflation in the discussion, and this post is an attempt to define and clarify the issues around this growing phenomenon, to see if it is possible to have a more constructive dialogue. 

This is a subject of some importance, so I think it is worth exploring over a series of entries, but we can hopefully frame the key issues in this post. I should also say at the outset, that I maintain a perspective that assumes that much of the investment in MT is considered as creation of long-term translation production infrastructure rather than as something used for a single one-time project. This perspective I think makes MT at least somewhat different from most other “CAT Tools”.  

There are at least three (and probably more) areas of general concern on PEMT:
  1. A clear definition of what PEMT actually is
  2. Compensation for PEMT work
  3. The quality of the work experience

    What is the task of post-editing machine translation?

    The most commonly understood definition of PEMT in the localization world is described graphically below. It is usually understood to be the linguistic work needed to correct MT output to a linguistic quality level that is close to, or indiscernible from a standard TEP process. At it’s worst, this corrective work can be very tedious and perhaps even a mind-numbing process.
    image 
    However, this view neglects several factors that affect the actual PEMT process, e.g.
    • MT system output varies, and some MT systems produce much better output than others, and thus the PEMT experience can be very different and is highly dependent on the actual quality produced by each individual MT system or engine. “Good” MT engines (systems) can noticeably enhance the productivity of translators and vice versa.
    • Google Translate is often not the example of the best that MT can do, and  customized, domain focused system usually outperform generic engines.
    • Increasingly, MT is used to translate content that would never be translated through a standard TEP process. Usually this involves much larger volumes of content that is also frequently updated and changed.
    • There are many MT applications where the editing work is only focused on making sure the content is understandable and accurate in meaning, even if it is grammatically imperfect.
    • It is possible to do linguistic analysis and a priori terminology and linguistic work to enhance the ability of the MT system to produce better output, and this can also be considered a kind of PEMT activity.
    • Many equate PEMT to “janitorial” work (often accurately so) that no self respecting translator would ever resort to, but there are many different bilingual skill levels required in the development process of an MT engine and very few of the critics seem to realize this.
     Thus, I would expand the definition of PEMT to:

    All tasks that are intended to improve the linguistic quality of output produced by an MT engine. This includes both the a priori analysis and post-editing and structural linguistic analysis work involved in the development of an MT engine. The objective always being to improve translator productivity (as well as reducing cost and time) on every translation project in that domain in future.

    image

    The experience at Asia Online is more closely characterized by the following graphic which shows a rapidly (weeks) evolving MT system that produces continuously improving MT quality as corrective feedback is fed back into the MT engine and error patterns are identified and gradually eliminated. This is a process that has been in use in the Asia Online project to translate the English Wikipedia to Thai. The processes are transferable to other languages. There is a highly collaborative process underlying the MT systems here where linguists are looking to eliminate as many linguistic error patterns as possible and thus also enhance the ongoing PEMT experience and process, and expedite the translation of a billion word corpus and bring it up to human quality levels in the quickest time possible.  

    Reducing Post-Editing Efforts

    It is also often assumed that translators are the only people who are capable of doing post-editing work. However, many translators do not care for the “janitorial” aspect of the work so it is not for everybody. MT systems that produce generally “high-quality” output could also be edited by monolingual speakers of the target language with expertise in the subject domain. Thus, it is possible to use humans who are less skilled than your average translator to accomplish business objectives, especially  where there are really large volumes involved and/or when grammatical perfection matters less than immediate access. There are many students and housewives that may actually find the flexibility and money offered by PEMT tasks attractive, and for very large ongoing projects they may indeed have a role. Problems arise when it is assumed that the work done by professionals can easily be done by non-professionals, so it is wise to be clear about your objectives and understand the nature of the work being offered and the skills required for competent performance. In general, open-minded professionals will always have an edge in producing better quality, but bringing in hostile or reluctant translators is also a sure way to fail.

    PEMT Skills Hierarchy

    The issues that get the most attention in the broader discussions on PEMT are the other two issues which are stated in brief below. These two issues will be examined in more detail in future posts.

    Compensation for PEMT: The early experience in the industry has been to arbitrarily reduce standard rates for PEMT work as it is assumed that to some extent the translation is already done. This practice causes a lot of resistance amongst translators who are expected to actually produce a usable translation (at low rates) when the starting point MT output is not useful or viable. It is already true that the unfortunate state of affairs today is to link payment of per word rates to TM matching rates. This has caused commoditization of the actual translation work, and translators today are expected to track baskets of words and get paid different rates in schemes such as the hypothetical one shown below:
    Matching Rate Pay Rate Per Word
    Repetitions 25%
    100% 25%
    85% to 99% 45%
    75% to 84% 60%
    0 to 74% 100%

    MT is sometimes used as a way to push these rates even lower. Thus, to be fair to translators there needs to be an assessment of the scope of the PEMT task. An MT system that produces 50% usable output should be compensated differently from one the produces 75% usable output, assuming the content has to be raised to the same target quality level. Logically, the greater the quality gap that needs to be filled, the greater the pay rate for doing the work. To do this accurately and fairly, we require rapid and widely accepted MT engine quality measures which do not exist today.  As our understanding of this issue evolves, I would not be surprised to see more hourly, consulting fee and project based payment schemes develop in future, as linguists and tech-savvy translators get more involved in steering MT engine development initiatives.

    The Nature Of PEMT Work: Another common complaint about PEMT is about the drudgery and mind-numbing nature of error correction work. (This seems surprising to me ;-) as I notice that many translators love to correct and point out linguistic, typo and errors of expression errors in blog comments and social networking discussions.) But this does however suggest that not all translators want to do this kind of work or are well suited to it. I have seen very positive and very negative translator feedback,  but often the implementation of this technology has as much to do with it, as the work process itself. Success with this technology often requires close collaboration with key translators who often provide great value in the process. And of course they must be compensated for these contributions. For very large projects it will become necessary to use less skilled workers or the “crowd”. I think that the tools will improve and possibly even make translation work more fun, whole and organic. (The fragmentation caused by the current TM matching rate approach has got to be a nightmare for many.) TAUS offers some guidelines for PEMT and this is a subject of great and growing interest to many. There are many levels of human steering interaction possible with MT and as some translators engage more often with MT systems they will start to understand where they have a long-term role to play as quality drivers. I expect that they will become critical members of teams who undertake to make massive content repositories increasingly multilingual.

    There are in fact environments where MT systems have been developed in close collaboration with translators, and thus offer clearly understood productivity advantages. In these cases the translators really want to use the systems  and would see not having access to the MT system as a disadvantage e.g. PAHO, Asia Online, Andovar. Most of the time, MT has yet to really rise to a level where it is a must-have tool for a professional translator but I think that is the direction we are heading in. The content deluge shows no sign of slowing down, and any global enterprise worth anything realizes, that they need to translate a whole lot more than user manuals and spec sheets if they want to build strong international businesses in future. So this is worth some attention and worth doing well.

    Tuesday, January 18, 2011

    Has Google Translate Reached the Limits of its Ongoing Improvement?

    The End of the Road for “The More Data The Better” Assumption?

    It has been commonly understood by many in the world of statistical MT, that success with SMT is based almost completely on the volume of data that you have. The experts at Google, ISI and TAUS have all been saying this for years. There has been some discussion that questioned this “only data matters” assumption but largely many in the SMT world continue to believe the statement below, because to some extent it has actually been true. Many of us have witnessed the steady quality improvements at Google Translate in our casual use to read an occasional web page (especially after they switched to SMT), but for the most part these MT engines rarely rise above gisting quality.

     

    "The more data we feed into the system, the better it gets..." Franz Och, Head of SMT at Google


    However, in an interesting review of the challenges of Google’s MT efforts in the Guardian, we begin to see some recognition that MT is a REALLY TOUGH problem to solve with machines, data and science alone. The article also quotes Douglas Hofstader who questions whether MT will ever work as a human replacement, since language is the most human of human activities. He is very skeptical and suggests that this quest to create accurate MT (as a total replacement for human translators), is basically impossible. While I too have serious doubts whether machines will ever learn meaning and nuance at a level that compares with competent humans, I think we should focus on the real discovery here, i.e. more data is not always better and/or that computers and data alone are not enough.  MT is still a valuable tool, and if used correctly can provide great value in many different situations. The Google admission according to this article is as follows:

    “Each doubling of the amount of translated data input led to about a 0.5% improvement in the quality of the output,”  and "We are now at this limit where there isn't that much more data in the world that we can use." Andreas Zollmann of Google Translate.

     

    But, Google is hardly throwing in the towel on MT, they will try “to add on different approaches and (explore) rules-based models."

    Interestingly, the “more data is better” issue is also being challenged in the search arena. In their zeal to index the world’s information, Google attempts to crawl and index as many sites as possible (because more data is better, right?). However, spammers are creating SEO focused “crap content” that increasingly shows up at the top of Google searches. (I experienced this first hand myself, when I searched for  widgets to enhance this blog. I gave up after going through page after page of SEO focused crap.) This article describes the impact of this low-quality content created by companies like Demand Media and are summarized succinctly in the quote below.

    Searching Google is now like asking a question in a crowded flea market of hungry, desperate, sleazy salesmen who all claim to have the answer to every question you ask.    Marco Arment

     

    But getting back to the issue of data volume and MT engine improvements, have we reached the end of the road? I think this is possibly true for some languages, i.e. data-rich languages like French, Spanish and Portuguese, where it is quite possible that tens of billions of words underlie the MT systems already. It is not necessarily true for sparse-data, or less present languages on the net (pretty much anything other than FIGS and maybe CJK), and we will hopefully see these other languages continue to improve as more data becomes available. In the graphic below we can see a very rough and generalized relationship between data volume and engine quality. I have a very rough estimate of the Google scale on top, and a lower data volume scale for customized systems at the bottom (that are generally focused on a single domain) where less is often more.
    Data
    Ultan O’Broin provides an important clue (I think anyway) for continued progress: “There's a message about information quality there, surely.” At Asia Online we have always been skeptical of the “the more data the better” view and we have ALWAYS claimed that data quality is more important than volume. One of the problems created by large scale automated data-scraping is that it is more than possible to pick-up large amounts of noise and digital dirt or just plain crap through this approach. Early SMT developers all use crawler based web-scraping techniques to acquire the training data to build their baseline systems. We have all learned by now I hope, that it is very very difficult to identify and remove noise from a large corpus, since by definition noise is random and unidentifiable through automated cleaning routines which can usually only target known patterns. (It is interesting to see that “crap content” also undermines the search algorithms, since machines (i.e.spider programs) don’t make quality judgments on the data they crawl. Thus Google can, and does easily identify crap content as the most relevant and important content for all the wrong reasons as Arment points out above.)

    Though corporate translation memories (TM) can be of higher quality than web-scraped data sometimes, TM also tends to gather digital debris over time. This noise comes from a) tools vendors who try to create lock-in situations by adding proprietary meta-data to the basic linguistic data, b) the lack of uniformity between human translators and c) poor standards that make consolidation and data sharing highly problematic. In a blog article describing a study of TAUS TM data consolidation, Common Sense Advisory describes this problem quite clearly: “Our recent MT research contended that many organizations will find that their TMs are not up to snuff — these manually created memories often carve into stone the aggregated work of lots of people of random capabilities, passed back and forth among LSPs over the years with little oversight or management.”
     

    So what is the way forward, if we still want to see ongoing improvements?

     

    I cannot really speak to what Google should do (they have lots of people smarter than me thinking about this), but I can share the basic elements of a strategy that I see is clearly working in producing continuously improving  customized MT systems developed by Asia Online. It is much easier to improve customer specific systems than a universal baseline.
    • Make sure that your foundation data is squeaky clean and of good linguistic quality (which means that linguistically competent humans are involved in assessing and approving all the data that is used in developing these systems).
    • Normalize, clean and standardize your data on an ongoing and regular basis.
    • Focus 75% of your development effort on data analysis and data preparation.
    • Focus on a single domain.
    • Understand that dealing with MT is more akin to interaction with an idiot-savant than with a competent and intelligent human translator.
    • Involve competent linguists through various stages of the process to ensure that the right quality focused decisions are being made.
    • Use linguistically informed development strategies as pure data based strategies are only likely to work to a point.
    • For language pairs with very different  syntax, morphology and grammar it will probably be necessary to add linguistic rules.
    • Use linguists to identify error patterns and develop corrective strategies.
    • Understand the content that you are going to translate and understand the quality that you need to deliver.
    • Clean and simplify the source content before translation.
    • And if quality really matters always use human validation and review.
    Clean data reduces Unpredictability
    All of this could be summarized simply as, make sure that your data is of high quality and use competent human linguists throughout the development process to improve the quality. This is true today and will be true tomorrow. 

    I suspect that effective man-machine collaborations will outperform pure data-driven approaches in future, as we are already seeing with both MT and search, and I would not be so quick to write off Google. I am sure that they can still find many ways to continue to improve. As long as the 6 billion people not working in the professional translation industry care about getting access to multilingual content, people will continue to try and improve MT. And if somebody tells you that machines can generally outperform or replace human translators (in 5 years no less), don’t believe them, (but understand there is great value in learning how to use MT technology more effectively anyway). We have quite a ways to go yet till we get there, if ever at all.
    universal_translator
    I recall a conversation with somebody at DARPA a few years ago, who said that the universal translator in Star Trek was the single most complex piece of technology on the Starship Enterprise, and that mankind was likely to invent everything else on the ship, before they had anything close to the translator that Captain Kirk used

    MT is actually still making great progress but it is wise to be always be skeptical of the hype. As we have seen lately, huge hype does not necessarily lead to success, as Google Buzz and Wave have shown.


    PS: Just a few days after this post was originally published Google admitted that they need to address the search SPAM problem, and this was further reinforced by a story in the Wall St. Journal.