Pages

Tuesday, February 15, 2011

An Exploration of Post-Editing MT – Part I

The topic of post-editing MT (PEMT) (yes, somebody has already come up with an acronym) continues to gain prominence in the professional translation world and is also a subject of heated and sometimes disdainful discussion amongst many translators in the blogosphere. There is a lot of confusion and conflation in the discussion, and this post is an attempt to define and clarify the issues around this growing phenomenon, to see if it is possible to have a more constructive dialogue. 

This is a subject of some importance, so I think it is worth exploring over a series of entries, but we can hopefully frame the key issues in this post. I should also say at the outset, that I maintain a perspective that assumes that much of the investment in MT is considered as creation of long-term translation production infrastructure rather than as something used for a single one-time project. This perspective I think makes MT at least somewhat different from most other “CAT Tools”.  

There are at least three (and probably more) areas of general concern on PEMT:
  1. A clear definition of what PEMT actually is
  2. Compensation for PEMT work
  3. The quality of the work experience

    What is the task of post-editing machine translation?

    The most commonly understood definition of PEMT in the localization world is described graphically below. It is usually understood to be the linguistic work needed to correct MT output to a linguistic quality level that is close to, or indiscernible from a standard TEP process. At it’s worst, this corrective work can be very tedious and perhaps even a mind-numbing process.
    image 
    However, this view neglects several factors that affect the actual PEMT process, e.g.
    • MT system output varies, and some MT systems produce much better output than others, and thus the PEMT experience can be very different and is highly dependent on the actual quality produced by each individual MT system or engine. “Good” MT engines (systems) can noticeably enhance the productivity of translators and vice versa.
    • Google Translate is often not the example of the best that MT can do, and  customized, domain focused system usually outperform generic engines.
    • Increasingly, MT is used to translate content that would never be translated through a standard TEP process. Usually this involves much larger volumes of content that is also frequently updated and changed.
    • There are many MT applications where the editing work is only focused on making sure the content is understandable and accurate in meaning, even if it is grammatically imperfect.
    • It is possible to do linguistic analysis and a priori terminology and linguistic work to enhance the ability of the MT system to produce better output, and this can also be considered a kind of PEMT activity.
    • Many equate PEMT to “janitorial” work (often accurately so) that no self respecting translator would ever resort to, but there are many different bilingual skill levels required in the development process of an MT engine and very few of the critics seem to realize this.
     Thus, I would expand the definition of PEMT to:

    All tasks that are intended to improve the linguistic quality of output produced by an MT engine. This includes both the a priori analysis and post-editing and structural linguistic analysis work involved in the development of an MT engine. The objective always being to improve translator productivity (as well as reducing cost and time) on every translation project in that domain in future.

    image

    The experience at Asia Online is more closely characterized by the following graphic which shows a rapidly (weeks) evolving MT system that produces continuously improving MT quality as corrective feedback is fed back into the MT engine and error patterns are identified and gradually eliminated. This is a process that has been in use in the Asia Online project to translate the English Wikipedia to Thai. The processes are transferable to other languages. There is a highly collaborative process underlying the MT systems here where linguists are looking to eliminate as many linguistic error patterns as possible and thus also enhance the ongoing PEMT experience and process, and expedite the translation of a billion word corpus and bring it up to human quality levels in the quickest time possible.  

    Reducing Post-Editing Efforts

    It is also often assumed that translators are the only people who are capable of doing post-editing work. However, many translators do not care for the “janitorial” aspect of the work so it is not for everybody. MT systems that produce generally “high-quality” output could also be edited by monolingual speakers of the target language with expertise in the subject domain. Thus, it is possible to use humans who are less skilled than your average translator to accomplish business objectives, especially  where there are really large volumes involved and/or when grammatical perfection matters less than immediate access. There are many students and housewives that may actually find the flexibility and money offered by PEMT tasks attractive, and for very large ongoing projects they may indeed have a role. Problems arise when it is assumed that the work done by professionals can easily be done by non-professionals, so it is wise to be clear about your objectives and understand the nature of the work being offered and the skills required for competent performance. In general, open-minded professionals will always have an edge in producing better quality, but bringing in hostile or reluctant translators is also a sure way to fail.

    PEMT Skills Hierarchy

    The issues that get the most attention in the broader discussions on PEMT are the other two issues which are stated in brief below. These two issues will be examined in more detail in future posts.

    Compensation for PEMT: The early experience in the industry has been to arbitrarily reduce standard rates for PEMT work as it is assumed that to some extent the translation is already done. This practice causes a lot of resistance amongst translators who are expected to actually produce a usable translation (at low rates) when the starting point MT output is not useful or viable. It is already true that the unfortunate state of affairs today is to link payment of per word rates to TM matching rates. This has caused commoditization of the actual translation work, and translators today are expected to track baskets of words and get paid different rates in schemes such as the hypothetical one shown below:
    Matching Rate Pay Rate Per Word
    Repetitions 25%
    100% 25%
    85% to 99% 45%
    75% to 84% 60%
    0 to 74% 100%

    MT is sometimes used as a way to push these rates even lower. Thus, to be fair to translators there needs to be an assessment of the scope of the PEMT task. An MT system that produces 50% usable output should be compensated differently from one the produces 75% usable output, assuming the content has to be raised to the same target quality level. Logically, the greater the quality gap that needs to be filled, the greater the pay rate for doing the work. To do this accurately and fairly, we require rapid and widely accepted MT engine quality measures which do not exist today.  As our understanding of this issue evolves, I would not be surprised to see more hourly, consulting fee and project based payment schemes develop in future, as linguists and tech-savvy translators get more involved in steering MT engine development initiatives.

    The Nature Of PEMT Work: Another common complaint about PEMT is about the drudgery and mind-numbing nature of error correction work. (This seems surprising to me ;-) as I notice that many translators love to correct and point out linguistic, typo and errors of expression errors in blog comments and social networking discussions.) But this does however suggest that not all translators want to do this kind of work or are well suited to it. I have seen very positive and very negative translator feedback,  but often the implementation of this technology has as much to do with it, as the work process itself. Success with this technology often requires close collaboration with key translators who often provide great value in the process. And of course they must be compensated for these contributions. For very large projects it will become necessary to use less skilled workers or the “crowd”. I think that the tools will improve and possibly even make translation work more fun, whole and organic. (The fragmentation caused by the current TM matching rate approach has got to be a nightmare for many.) TAUS offers some guidelines for PEMT and this is a subject of great and growing interest to many. There are many levels of human steering interaction possible with MT and as some translators engage more often with MT systems they will start to understand where they have a long-term role to play as quality drivers. I expect that they will become critical members of teams who undertake to make massive content repositories increasingly multilingual.

    There are in fact environments where MT systems have been developed in close collaboration with translators, and thus offer clearly understood productivity advantages. In these cases the translators really want to use the systems  and would see not having access to the MT system as a disadvantage e.g. PAHO, Asia Online, Andovar. Most of the time, MT has yet to really rise to a level where it is a must-have tool for a professional translator but I think that is the direction we are heading in. The content deluge shows no sign of slowing down, and any global enterprise worth anything realizes, that they need to translate a whole lot more than user manuals and spec sheets if they want to build strong international businesses in future. So this is worth some attention and worth doing well.

    Tuesday, January 18, 2011

    Has Google Translate Reached the Limits of its Ongoing Improvement?

    The End of the Road for “The More Data The Better” Assumption?

    It has been commonly understood by many in the world of statistical MT, that success with SMT is based almost completely on the volume of data that you have. The experts at Google, ISI and TAUS have all been saying this for years. There has been some discussion that questioned this “only data matters” assumption but largely many in the SMT world continue to believe the statement below, because to some extent it has actually been true. Many of us have witnessed the steady quality improvements at Google Translate in our casual use to read an occasional web page (especially after they switched to SMT), but for the most part these MT engines rarely rise above gisting quality.

     

    "The more data we feed into the system, the better it gets..." Franz Och, Head of SMT at Google


    However, in an interesting review of the challenges of Google’s MT efforts in the Guardian, we begin to see some recognition that MT is a REALLY TOUGH problem to solve with machines, data and science alone. The article also quotes Douglas Hofstader who questions whether MT will ever work as a human replacement, since language is the most human of human activities. He is very skeptical and suggests that this quest to create accurate MT (as a total replacement for human translators), is basically impossible. While I too have serious doubts whether machines will ever learn meaning and nuance at a level that compares with competent humans, I think we should focus on the real discovery here, i.e. more data is not always better and/or that computers and data alone are not enough.  MT is still a valuable tool, and if used correctly can provide great value in many different situations. The Google admission according to this article is as follows:

    “Each doubling of the amount of translated data input led to about a 0.5% improvement in the quality of the output,”  and "We are now at this limit where there isn't that much more data in the world that we can use." Andreas Zollmann of Google Translate.

     

    But, Google is hardly throwing in the towel on MT, they will try “to add on different approaches and (explore) rules-based models."

    Interestingly, the “more data is better” issue is also being challenged in the search arena. In their zeal to index the world’s information, Google attempts to crawl and index as many sites as possible (because more data is better, right?). However, spammers are creating SEO focused “crap content” that increasingly shows up at the top of Google searches. (I experienced this first hand myself, when I searched for  widgets to enhance this blog. I gave up after going through page after page of SEO focused crap.) This article describes the impact of this low-quality content created by companies like Demand Media and are summarized succinctly in the quote below.

    Searching Google is now like asking a question in a crowded flea market of hungry, desperate, sleazy salesmen who all claim to have the answer to every question you ask.    Marco Arment

     

    But getting back to the issue of data volume and MT engine improvements, have we reached the end of the road? I think this is possibly true for some languages, i.e. data-rich languages like French, Spanish and Portuguese, where it is quite possible that tens of billions of words underlie the MT systems already. It is not necessarily true for sparse-data, or less present languages on the net (pretty much anything other than FIGS and maybe CJK), and we will hopefully see these other languages continue to improve as more data becomes available. In the graphic below we can see a very rough and generalized relationship between data volume and engine quality. I have a very rough estimate of the Google scale on top, and a lower data volume scale for customized systems at the bottom (that are generally focused on a single domain) where less is often more.
    Data
    Ultan O’Broin provides an important clue (I think anyway) for continued progress: “There's a message about information quality there, surely.” At Asia Online we have always been skeptical of the “the more data the better” view and we have ALWAYS claimed that data quality is more important than volume. One of the problems created by large scale automated data-scraping is that it is more than possible to pick-up large amounts of noise and digital dirt or just plain crap through this approach. Early SMT developers all use crawler based web-scraping techniques to acquire the training data to build their baseline systems. We have all learned by now I hope, that it is very very difficult to identify and remove noise from a large corpus, since by definition noise is random and unidentifiable through automated cleaning routines which can usually only target known patterns. (It is interesting to see that “crap content” also undermines the search algorithms, since machines (i.e.spider programs) don’t make quality judgments on the data they crawl. Thus Google can, and does easily identify crap content as the most relevant and important content for all the wrong reasons as Arment points out above.)

    Though corporate translation memories (TM) can be of higher quality than web-scraped data sometimes, TM also tends to gather digital debris over time. This noise comes from a) tools vendors who try to create lock-in situations by adding proprietary meta-data to the basic linguistic data, b) the lack of uniformity between human translators and c) poor standards that make consolidation and data sharing highly problematic. In a blog article describing a study of TAUS TM data consolidation, Common Sense Advisory describes this problem quite clearly: “Our recent MT research contended that many organizations will find that their TMs are not up to snuff — these manually created memories often carve into stone the aggregated work of lots of people of random capabilities, passed back and forth among LSPs over the years with little oversight or management.”
     

    So what is the way forward, if we still want to see ongoing improvements?

     

    I cannot really speak to what Google should do (they have lots of people smarter than me thinking about this), but I can share the basic elements of a strategy that I see is clearly working in producing continuously improving  customized MT systems developed by Asia Online. It is much easier to improve customer specific systems than a universal baseline.
    • Make sure that your foundation data is squeaky clean and of good linguistic quality (which means that linguistically competent humans are involved in assessing and approving all the data that is used in developing these systems).
    • Normalize, clean and standardize your data on an ongoing and regular basis.
    • Focus 75% of your development effort on data analysis and data preparation.
    • Focus on a single domain.
    • Understand that dealing with MT is more akin to interaction with an idiot-savant than with a competent and intelligent human translator.
    • Involve competent linguists through various stages of the process to ensure that the right quality focused decisions are being made.
    • Use linguistically informed development strategies as pure data based strategies are only likely to work to a point.
    • For language pairs with very different  syntax, morphology and grammar it will probably be necessary to add linguistic rules.
    • Use linguists to identify error patterns and develop corrective strategies.
    • Understand the content that you are going to translate and understand the quality that you need to deliver.
    • Clean and simplify the source content before translation.
    • And if quality really matters always use human validation and review.
    Clean data reduces Unpredictability
    All of this could be summarized simply as, make sure that your data is of high quality and use competent human linguists throughout the development process to improve the quality. This is true today and will be true tomorrow. 

    I suspect that effective man-machine collaborations will outperform pure data-driven approaches in future, as we are already seeing with both MT and search, and I would not be so quick to write off Google. I am sure that they can still find many ways to continue to improve. As long as the 6 billion people not working in the professional translation industry care about getting access to multilingual content, people will continue to try and improve MT. And if somebody tells you that machines can generally outperform or replace human translators (in 5 years no less), don’t believe them, (but understand there is great value in learning how to use MT technology more effectively anyway). We have quite a ways to go yet till we get there, if ever at all.
    universal_translator
    I recall a conversation with somebody at DARPA a few years ago, who said that the universal translator in Star Trek was the single most complex piece of technology on the Starship Enterprise, and that mankind was likely to invent everything else on the ship, before they had anything close to the translator that Captain Kirk used. 

    MT is actually still making great progress but it is wise to be always be skeptical of the hype. As we have seen lately, huge hype does not necessarily lead to success, as Google Buzz and Wave have shown.


    PS: Just a few days after this post was originally published Google admitted that they need to address the search SPAM problem, and this was further reinforced by a story in the Wall St. Journal.





    Monday, January 10, 2011

    The Most Worthwhile Conferences to Attend in 2011 and Finding the Real Customer

    This is an expansion of a conversation I had with Renato Beninatto which he also blogged on and was also video taped here. I am sharing my opinion here not so much because I am endorsing one event or the other, rather this list is just my personal list of preferences and nothing more. I do not claim there is any special status to my personal preferences and you may notice I tend to like conferences where translation technology is emphasized.

    We live in an age, where increasingly marketing and corporate-speak is challenged, undermined and sometimes even seen as disingenuous and false. (Raise your hand if you trust and respect corporate press releases).  Today we see customer voices rise above the din of corporate messaging, and taking control of branding and corporate reputations with their own “authentic” discussions of actual customer experiences, while marketing departments look on haplessly. I think this phenomenon is happening on many fronts, including conferences in the localization industry. There are too many events in the L10N industry that seem formulaic, routine, repetitive and engineered based on the same old viewpoints. This, I think affects the ability of these events to really spark dialogue, excitement and generate vital learning experiences that make these conferences must-attend events. While these events remain useful for “face-time”, they often have little value for really engaging attendees at a professional level and providing insights that drive new action plans.

    tekomacross
    What makes for a great conference or professional industry event? To my mind: high quality content, interactive and engaged audiences in sessions that broaden one’s horizons, interesting people who continue the professional dialogue outside of the sessions and share learning experiences and of course a good location. And if you can offer all of this at a reasonable cost, even better. A great professional event is characterized by learning, the more intensive the learning experience, the better. The best ones leave you thinking for awhile after the event.  Intense learning rarely happens at “really big” events because it is hard to scale this, but hopefully you have a few intense one-on-one interactions.

    I also really like events that really focus on the customer: the real customer. The real customer would be the management team that runs, handles and is held accountable for success and failure of international initiatives (rarely the localization department of that company IMO). So the real customer would be senior sales, marketing, product management and customer support people who may also fund and direct the localization department mission in the global enterprise. (We rarely see them at any localization conferences because localization is rarely a central focus for them.) The real customer is more likely to focus on market share trends and customer satisfaction / loyalty  rather than word rates, fuzzy matching rates, TM ownership, SimShip or vendor management.
    The question that Ultan O'Broin posed most recently in Quora was:

    He also presents a categorization of these conferences as follows Generalist, Specialty and Geography-focused events. He said that he liked to attend one or two of each category and his preferences are stated in the links above. 

    I still see myself as more of a pragmatic technology guy, trying to solve meaningful and useful translation problems (hopefully) with technology, rather than an industry insider (not quite a localization professional), so I would organize this a little bit differently but it still has much in common with Renato’s view. For MT especially it is all about understanding and learning how to use it at this point in time.  

    I think there are 3 or more categories of conferences that touch localization and translation. The following is my very crude categorization. (Hopefully somebody can suggest a better categorization scheme. Please feel free to tear this apart).
     
    1) Traditional Corporate L10N & Translation Focused Conferences
    2) Translation Technology & L10N Research Focused Conferences
    3) Special Focus & Miscellaneous  and

    New Opportunities and Events/Industries/Markets to Explore 
    If I limit myself to three or so in each category I would select the following events.
    Category 1: Localization World and ELIA are the best in terms of content quality and networking value IMO. Localization World is the largest industry insider event (and the only one where I have seen a real customer view occasionally) and ELIA is a great example of sharing and collaboration between peers and competitors. This is the most crowded conference category and my recommendation would be for people to choose carefully amongst the available options (using location and content as a guide).  There are some who prefer the LISA and GALA versions of this category and they can also be good and sometimes slightly different like the LISA event at UC Berkeley.

    Category 2:  TAUS Annual User Conference for MT focus from enterprise customer perspectives, (not the regional meetups held around the world) even though this event has some overlap with the first category.
    AMTA for a deep dive on machine translation related issues that covers both the gory details of the technology and its use in public and private sectors, but this event has a very strong US focus.
    LRC for broad and innovative localization research and thinking that is truly focused on the next generation of needs.
    Translingual Europe 2010:  This was a free event held in Berlin that I think shows promise had much about MT and broader language technology initiatives across the world but especially in the EU. 

    Category 3: IMTT events have great content and a wonderful collaborative and sharing culture and I also think are one of the few events where you really get to see both LSP and translator perspectives engage together.
    AGIS to understand the issues in the non-profit world where translation is often linked to national development priorities or alleviation of information poverty. A different perspective and much more ambitious initiatives that involve national policy oftentimes. I am willing to bet that the leaders in the non-profit arena will also be the first to really use technology well and drive standards forward. I suspect that the most interesting crowdsourcing initiatives will also come from the non-profit area and passionate community leaders rather than global enterprises.
    tekom is an opportunity for  localization professionals to connect to a broader customer community and I hope that more of this happens in 2011. The industry can only grow and gain momentum by becoming more involved with larger broader and vertical market focused shows which are important to real customers.
    I think some of the smaller events can also be very interesting e.g. The last LISA Crowdsourcing round table was informative and showed potential and promise but lost momentum because of weak follow-up. I am told that some of the smaller events in Eastern & Southern Europe also have very high quality content and great engagement. Web based events are growing in popularity but very few have found the right mix of content and engagement.

    I hope that we will see more events focused on resolving issues around data interchange and exchange standards so that translation data flows much more easily and fewer on process standards.

    New Opportunities and Events/Industries/Markets to Explore
    Possibly the best and most exciting business opportunities and ability to learn about new long-term strategic opportunities are shows that have been off the beaten path and ignored by most industry insiders.
     
    Some possibilities for translation industry collaborations in future and where I think the best opportunities lie for emerging translation and localization demand are listed below. Again just some suggestions and not a complete list by any means. It would be worth finding out which are the best conferences to meet customers, providers and thought leaders in the following areas and develop marketing communications that interest, educate and engage these attendees, on localization and translation issues.
    I think video content will be a major new opportunity, and will likely cover all of the above segments. I have seen that there are video subtitling/dubbing focused conferences but have not attended one yet and I suspect that this is an area worth exploring as a long-term opportunity. Video content translation is likely to be the fastest growing new sector for the industry as Cisco estimates that 90% of IP traffic will be video related by 2013. If that is where the end-customer is, it makes sense to focus on it. I am sure mobile will also be a growing and strategic area.
     
    Finally, I think we should all be exploring how to get more connected in to BRICI global commerce focused events. (This is more complicated than holding an event in China and/or India.) It is very likely that as the export/import sectors in these fast growing countries expands there will be events and conferences, that will be worth tapping into to really get access into these markets. I would bet that the best events will be organized by locals in these regions who are interested in globalization issues, possibly even government sponsored events.

    Please join the discussion on Quora or comment on this blog as this is by no means a definitive or authoritative list. What do you think? And lets hope that we all find “a real customer” at the events we go to.
    dialogue