Pages

Showing posts with label SMT. Show all posts
Showing posts with label SMT. Show all posts

Thursday, September 22, 2016

Comparing Neural MT, SMT and RBMT – The SYSTRAN Perspective

This is the second part of an interview with Jean Senellart (JAS) , Global CTO and SYSTRAN SAS Director General. The first part can be found here: A Deep Dive into SYSTRAN’s Neural Machine Translation (NMT) Technology .

Those translators who accuse MT vendors of stealing or undermining their jobs should take note that SYSTRAN is the largest independent MT vendor. A position it has held for most of its existence, but has never generated more than €20 million in annual revenue. Which to me suggests that MT is mostly being used for different kinds of  translation tasks and is hardly taking any jobs away. The problem has been much more related to unscrupulous or incompetent LSPs who used MT improperly in rate negotiations. MT is hugely successful for those companies who find other ways to monetize the technology, as I pointed out in  The Larger Context Translation Market.This does suggest that MT has huge value in enabling global communication and commerce, and even in its less than perfect state is considered valuable by many who might otherwise need to acquire human translation services. If anything, MT vendors are the ones that are trying hardest to develop technology that is actually useful to professional translators as the Lilt offering shows and as this new SYSTRAN offering also promises to be. The new MT technology is reaching a point where it is becoming rapidly responsive to corrective feedback and thus much more useful to professional translation use case scenarios. The forces affecting translator jobs and work quality are much more complex and harder to pin down as I have also mentioned in the past.

In my conversations with Jean, I realized that he is one of the few people around in the "MT industry", who has deep knowledge and production use experience with all three MT paradigms. Thus, I tried to get him to share both his practical experience-based and philosophical perspectives about the three approaches. I found his comments fascinating and thought that it would be worth highlighting them separately in this post. Jean (through SYSTRAN) is unique in being one of the rare practitioners around, who has produced commercial release versions, of say French <> English MT systems, using all three MT paradigms. More if you count the SPE and NPE “hybrid” variants where more than one approach is used in a single production process.

Again in this post, I have tried as much as possible to keep Jean Senellart’s direct quotes in here to avoid any misinterpretation. My comments are in italics when they do occur within his quotes.

 

Comparing The Three Approaches; The Practical View

Some interesting comparative comments made by Jean about his actual implementation experience with the three methodologies (RBMT, SMT, NMT):

“The NMT approach is extremely smart to learn the language structure but is not as good at memorizing long lists of terminology as RBMT or SMT was. With RBMT, the terminology coverage was right, but the structure was clumsy – with SMT or SPE, we had an “in-between” situation where we got the illusion of fluency, but sometimes at the price of a complete mistranslation, and a strange ability to memorize huge lists but without any consideration of their linguistic nature. With NMT, the language structure is excellent, as if the neural network really deeply understands the grammar of the language – and introducing [greater] support of terminology was the missing link to the previous technologies.”

(This inability of NMT to handle large vocabulary lists is considered one of the main weakness of the technology currently. Here is another reference discussing this issue. However, it appears that SYSTRAN has developed some kind of a solution to address this issue.)

What is interesting with NMT is it seems far more tolerant than PBSMT (Phrase-Based SMT) to noisy data. However, it does need far less data than PBSMT to learn – so we can afford to provide only data for which we know we have a good alignment quality. Regarding the domain or the quality [of MT output of] the languages, we are for the moment trying to be as broad as possible [
rather than focusing on specialized domains].”

In terms of training data volume, JAS said: “This is still very empirical, but we can outperform Google or Microsoft MT on their best language pairs using only 5M translation units – and we have a very good quality (BLEU score about 45** ) for languages like EN>FR with only 1M TU. I would say we need 1/5 of the data necessary to train SMT. Also, generic translation engines like Google or Bing Translate are using billions of words for their language models, here we need probably less than 1/100th.”

(**I think it bears saying that I fully expect that the BLEU here is measured with great care and competence, unlike what we see so often with Moses practitioners and LSPs in general who assume scores of 75+ are needed for the technology to be usable.

The ability of the new MT technology to improve rapidly with small amounts of good quality training data and small amounts of corrective feedback suggests that we may be approaching new thresholds in the use of MT for professional use.)
 

Comparing RbMT, SMT, and NMT: The Philosophical View

When I probed more deeply into the differences between these MT approaches, (since really SYSTRAN is the only company who has real experience in all 3), JAS said: “I would more compare them in terms of what they are trying to do, and on their ability to learn.” His explanation reflects his long-term experience and expertise and is worth careful reading. I have left the response in his own words as much as possible.

RBMT, fundamentally (unlike the other two), has an ulterior motive: it attempts to describe the actual translation process. And by doing that, it has been trying to solve a far more complicated challenge than just machine translation; it tries to decompose [and deconstruct and analyze] and explain how a translation is produced. I still believe that this goal is the ultimate goal of any [automated translation] system. For many applications, in particular, language learning, but also for post-editing, the [automated ] system would be far more valuable if it could produce not only the translation but also explain the translation.”

“In addition, RBMT systems are facing three main limitations in language [modeling] which are:

1) the intrinsic ambiguity of language for a machine, which is not the same for a human who has access to meaning [and a sense for the underlying semantics],

2) the exception-based grammar system, and

3) the huge, contextual and always expanding volume of terminology units.”

“Technically, RBMT systems might have different levels of complexity depending on the linguistic formalism being used, making it hard to compare with the others (SMT, NMT), so I would rather say that one of the main reasons for the limitations of a pure RBMT system lies in its [higher reaching] goal. The fact is that fully describing a language is a very complicated matter, and there is no language today for which we can claim a full linguistic description.”

“SMT came with an extraordinary ability to memorize from its exposure to existing translations – and with this ability, it brought a partial solution to the first challenge mentioned above, and a very good solution to the third one – the handling of terminology – however, it mostly ignores the modeling of the grammar. Technically, I think SMT is the most difficult of the 3 approaches, it combines many algorithms to optimize the MT process, and it is hard work to deal with the huge database of [relevant training ] corpus.”

“NMT has access to the meaning and is dealing well with modeling of human language grammar. In terms of difficulty, NMT engines are probably the simplest to implement, a full training of an NMT engine involve only several thousands of lines of code. The simplicity of implementation is, however, hiding the fact that we/nobody knows why it is so effective.”

“I would use an analogy of a human learning to drive a car to explain this more fully:

- The Rule-based approach will attempt to provide a full modeling of the car dynamic, on how the engine is connected to the wheel, on the effect of acceleration in the trajectory, etc. This is very complicated. (And possibly impossible to model in totality.)

- The Statistical approach, will use data from past experience and will try to compare a new situation with a past situation and will decide on the action based on this large database [of known experience]. This is a huge task and very difficult to implement. (And can only be as good as the database it learns from.)

- The Neural approach, with a limited access to the phenomenon involved, or with limited ability to remember, will build its own “thinking” system to optimize the driving experience, it will actually learn to drive the car, build reflexes – but will not be able to explain why and how such decisions are being made, and will not be able to leverage local knowledge – for instance that at a specific bend on the road in very specific weather condition, it had to anticipate [braking] because it is particularly dangerous, etc... This approach is surprisingly very simple and thanks to computation power evolution have become more accessible.”

“Today, this last approach (NMT) is clearly the most promising but will need to integrate the second (SMT) to be more robust, and eventually to deal with the first one (RBMT) to be able to not only make choices but also explain them.”

 

Comparing NMT to Adaptive MT

When probed about this, JAS said: “Adaptive MT is an innovative concept based on the current SMT paradigm – it is, however, a concept that is quite naturally embedded in the NMT paradigm; of course, there will be work needed to be done to make it work as nicely as the Lilt product does it. But, my point is that NMT (and not just from Systran) will bring a far more intuitive solution to this issue of continuous adaptive learning, because it is built for that: on a trained model, we can tune the model without any tricks with feedback of one single sentence – and produce a translation which immediately adapts to user input.”

The latest generation MT technology, especially NMT and Adaptive MT look like a major step forward to enabling the expanding use of MT in professional translation settings. With continuing exploration and discovery in the fields of NLP, artificial intelligence and machine intelligence, I think we may be in for some exciting times ahead as these discoveries benefit MT research. Hopefully, the focus will shift to making new content multilingual and solve new kinds of translation challenges, especially in speech and video. I believe that we will see more of the kinds of linguistic steering activities we are seeing in motion at eBay and that there will always be a role for competent linguists and translators.

Jean Senellart, CEO, SYSTRAN SA



The first part of the SYSTRAN interview can be found at A Deep Dive into SYSTRAN’s Neural Machine Translation (NMT) Technology 

Post Script: There is a detailed article that describes the differences between these approaches on the SYSTRAN website  released a few weeks after this post was published.

Tuesday, August 30, 2016

De Domo Sua – The Rationale for Hiring a Consultant when Working with MT

I have always found Luigi Muzii's aka Il barbaro (The Barbarian) opinions on the "translation industry" interesting and have often felt that his blog never got the attention it deserved. Possibly because he often wrote with frequent references to Roman historical precedence, in what I have been told is an Italian scholarly style, and perhaps because English is not his preferred language to communicate complex thoughts. 
 
His willingness to state the obvious (to common sense) sometimes makes him unpopular, especially with the MT naysayers e.g. "Moreover, translation data - i.e. project data - has a limited lifespan and, at some point in time, it becomes outdated, possibly inaccurate, and definitely irrelevant." I think that many in the "translation industry" still fail to realize that the bulk of the material that they translate;  i.e. the manuals, documentation and generic business content on websites focused on the wonder and excellence of corporate goods and services, become less relevant, and less valuable with each passing day no matter how well it has been translated. They clearly don't like to hear someone suggest that this might be so. And it is only natural that the buyers of business translation services might ask: "Is there a cheaper way to do this?" since they are acutely aware of the rapidly declining relevance and low value of a lot of the corporate content they produce.

His willingness to raise fundamental questions I think makes his voice worthy of attention, for me at least. Anyway, I am glad that he sent me this brief overview which I am told will be expanded in future, and I hope that he will be back in future on this blog with other observations on the business of translation related to MT or not. 
 
Judge a man by his questions, not his answersVoltaire

----------------

On his house [*]

[*] De domo sua is a famous speech that Marcus Tullius Cicero gave in theRoman Senate in 57 B.C. against tribune Publius Clodius Pulcher, to have his house rebuilt at public cost, after it was destroyed by Clodius' supporters. 

"On his house" is the equivalent of "De domo sua" in English, which we largely use in Italy to label a cause one pleads for his own benefit. In this case, "domus mea" (my home) is the rationale for hiring a consultant. 


Implementing machine translation (MT) effectively is not an easy task for a translation agency, or even for a large translation buyer whose core business is not translation. Still, with 99,9% of translations performed daily by machines, it is more and more perfectly reasonable that businesses are eager to have more and more of their content translated, quickly and cost-effectively using the technology that is available. 

Speed and cost-effectiveness are crucial to attain competitive advantage in global market competition, and we should understand that not all content is equal in terms of how it should and can be properly translated. And not all content deserves the same TEP (Translate>Edit>Proof) consideration and production process. 

Most enterprises, especially SMEs seeking a foothold in international markets, see machine translation as a viable solution, but they do not have the necessary knowledge and skills needed to deal with the challenging effort of implementing MT properly. Most logically, then, a SME would seek help from translation agency vendors assuming they would know how to use and implement MT technology. 

Unfortunately, most translation agencies often do not have these skills either. Considering that the translation technology market - which includes MT - is noticeably smaller than the broader linguistic services market, and is somewhat under-served in terms of availability of competent service providers, more and more language service vendors unfortunately have been adding MT consulting to their service offerings as a means to generate new forms of revenue. This can only result in a situation somewhat like that shown below. 


The following very brief tips are intended as basic practical guidelines for enterprises interested in taking advantage of MT. 

Tip #1

Hire an independent consultant. If you are a language service vendor, being supported by an independent advisor will spare you the cost and the hassle of an internal department while offering your customers the peace of mind of an unbiased opinion and professional help. 

Tip #2

Either ask for or run a preliminary analysis. Through interviews with staff and examination of the current modus operandi (MO), an independent consultant can help you objectively assess your processes and facilities in view of the potential implementation and integration of a machine translation platform, as well as provide guidance on the possibility of the provision of any ancillary services. 

Tip #3

Either ask for or draft a report that provides the following:
  • Review of your goals,
  • The outcomes of the preliminary analysis
  • Appraisal of the suitability of current modus operandi (MO) for machine translation, and
  • An outline of an exploratory program.
Use flowcharts and/or sequence diagrams to draw your modus operandi (MO), starting from future MO (FMO), then present MO (PMO) and finally transitional MO (TMO.) By overlapping the three diagrams you should be able to track the differences and be able to make adjustments to streamline the transitional process and ensure a smooth update of the existing processes. 

Tip #4

Write down the requirements that have been collected during the interviews for the preliminary analysis. Prepare a grid with the technical specifications (TS) and the statement of work (SoW) detailing what is to be done. 

Tip #5

Write a clear template for a request for proposals (RFP) including the TS and the SoW and send it out to a group of selected MT vendor candidates, asking them to fill in the grid. 

Tip #6

When hiring the advisor, consider that the advisor should have comprehensive insight and understanding of the general MT market and its players to help identify the best candidates. 

Do not forget 

 

Discuss the report and the proposals with the advisor and the customer, if you are a translation agency. Provide a program, a configuration plan and a training/recruiting program including all tasks, from benchmarking to data qualification and preparation, from vendor selections to proposal evaluation, from negotiations to implementation, from testing to training. 



Luigi Muzii's profile photo

Luigi Muzii has been in the "translation business" since 1982 and has been a business consultant since 2002, in the translation and localization industry through his firm . He focuses on helping customers choose and implement best-suited technologies and redesign their business processes for the greatest effectiveness of translation and localization related work.

This link provides access to his other blog posts.

Monday, August 8, 2016

Overview of Expert MT Systems - Iconic Translation Machines

This is a third post in a series on expert MT systems developers. Iconic fits my definition of expert much more closely (as does tayou) as they do-it-for-you rather than give you a technology platform to do-it-yourself. (DIY does not work well for those who do not really know what they are doing.) They bring deep expertise to bear on the very specific translation problem that you bring to them, and may create a unique MT solution for each customer. In the case of Iconic, the components used to build an MT solution, may also vary from customer to customer, as they assemble sub-components to address unique customer problems, which can vary by content, language and quality expectation. In general, the interaction with an expert is going to much more consultative and customer specific and less off-the-shelf. As MT is an evolving technology, it is especially useful for an expert developer to have close ties with an academic institution that is engaged in ongoing research in MT, and related technology like machine learning and artificial intelligence. This allows them to continue to enhance the technology and most expert developers will have this kind of a relationship in place. This is especially critical at this juncture, as there is a lot of leverage able work being done in the machine intelligence and neural network based learning field. This post developed from a conversation with John Tinsley. Most of the emphasis in the post below is mine.
-------------------------------------
 

From the lab to the market
The concept and technology behind Iconic originated from a large scale research and development project at Dublin City University (DCU). The state-of-the-art MT technology developed by a number of MT PhDs, who are still with the company today, was adapted to meet the translation requirements at the European Patent Office. On the back of this, the company spun-out of the university and our first commercial MT engine, for English to Portuguese patent translation, went into production in July 2010. We haven’t looked back since. 

Our original focus was exclusively on patent translation and we worked with a number of LSPs and information providers who specialised in this field. We’ve since evolved into a turnkey machine translation software and solutions provider that specializes in custom solutions tailored with subject matter expertise for specific industry sectors and content types. These are areas where our sophisticated MT technology has significant value to add over off-the-shelf solutions.

MT: a complex technology
There are two established paradigms of MT: rule based, and statistical (though Neural MT is clearly on the horizon, more on that later). Statistical MT is by far the predominant approach, and extensions to this include the incorporation of rules or linguistic information which are often referred to as “hybrid” MT. Despite this, there is no single approach or configuration that works best across all language pairs, content types, and writing styles.

Our approach, from a technological perspective, tries to address the fact that there’s no “one-size-fits-all” machine translation solution. The best approach completely depends on the job at hand. We have developed an Ensemble ArchitectureTM which combines 100’s of processes across different paradigms that can be combined to produce the best translation for a given task. 

Some of these processes are specific to a particular language, to a certain content type (e.g. legal, financial), or to a particular writing style (e.g. patents, contracts, annual reports). Depending on a given input document or client project, our platform takes an on-the-fly decision as to the most appropriate combination of processes to use. In fact, in many cases, different sections of a single document can be translated using completely different processes. The goal is to use the most effective set of tools at our disposal for any given task, so that our end users get the best possible translation.

The Ensemble Architecture: a sample of just a few of the processes available to enhance translation quality


From a client customisation perspective, this model is super extensible. It allows us to develop and add new processes to the ensemble architecture as needed on a case by case basis, making it a continuously evolving architecture.

One of the biggest challenges in maintaining such a complex technology is the need for deep expertise in the underlying MT methodologies. This places a lot of importance on the team, which is something we’ve put a lot of resources into building at Iconic. Having a mix of experts in the science underpinning MT technology, and language experts, combined with software development skill in the team allows us to transfer the subject matter expertise in our team, directly into our clients’ applications.

This is crucial when it comes to ongoing engine development and customisation. When adding data to an engine, quality can hit a ceiling quite quickly. End users will highlight critical issues in the output that cannot be resolved by simply “retraining” and adding more data. Even if we have valuable post-edited data from translators, it is still often not enough. Having the expertise in the team, in a) knowing where to look “under the hood” to identify the cause of an issue, and b) having the skill to implement a fix, is crucial to success. In fact, we recently ran a video series introducing some of the members of our team which you can watch here

Applying MT successfully across industry sectors
Iconic’s first product was IPTranslator, a suite of MT engines adapted for patent and intellectual property related content. Patents are challenging enough to read in your native language, never mind trying to develop MT software to handle them. The complexity of the documents was one of the key motivators in developing our Ensemble Architecture - we needed the capability for the MT engine to be able to dynamically adapt to the content domain, be it a chemical, engineering, or electronic patents. Similarly, it needs to be able to treat the different patent sections in different ways, e.g. the stunted telegraphic style of the title vs. the long-winded “legalese” of the claims. 

Since then, this architecture has grown to inherently cover other areas, and verticals that pose similar technological challenges to MT, such aseDiscovery, the financial and life sciences industries, and, in particular, user-generated content in the hospitality and e-commerce industries. 

From the language perspective, like most approaches to MT, our technology is language independent and can be applied to any combination of languages. However, given the extensibility and the ability to combine domain-adapted approaches with language-specific processes, we’ve been able to tackle harder projects had have great success with traditionally more challenging languages such as Chinese and German, to name just two. 

These are areas we’ve chosen to focus on because we have significant value to add with our technology and expertise. MT has been rolled out generally with relative success in industries such as IT and for languages like French, which tend to lend themselves better to language technology in general. But when it comes to more difficult languages and content types, a smarter approach is needed and that’s often where we come into play. 

MT for multiple buyers
We are the MT partner of choice for some of the world’s largest translation companies, information providers, and government and enterprise organisations, each of whom can have very different requirements from an MT technology solution. We have worked very closely with Welocalize over a number of years, in particular Park IP, to bring post-edited MT into their workflow for patent translation across a number of languages. Similarly, we have worked extensively with RWS in the same field. 

Over the past 18 months, in collaboration with the ADAPT research centre, we have been providing English to Irish MT to the Irish government, also for post editing. To the best of our knowledge, is the first successful commercial deployment of MT for this language combination.

Outside of traditional language services use cases, Iconic’s MT is used in a number of areas where translation is just one part of a much bigger picture. For instance, over the course of the last year we have been working on large scale digitization projects, taking millions of documents of archive content stretching back over the past 150 years (in various challenging formats as you can imagine!), converting them, cleaning them, machine translating them, and then making the available in multilingual searchable databases. Billions of words of historical analog content have passed through our engines and are now accessible to a global audience. 

In an increasing number of these cases, MT is producing quality output that is fit for purpose as is without the need for post editing and this trend is only going to continue as the volume of content being produced grows further.

Business models for complex technology
In the same way that there’s no single “best” approach to MT, it is also not an “out of the box” type of software. Yes, there are use cases and instances where it can work in this way, but they are the arguably the exception rather than the rule. It can work very well for anecdotal or spot usage, but when it comes to complex real-world problems, a deeper level of expertise and experience is required.

This has probably been distorted by (fantastic, ground-breaking) tools like Moses which have done a great job of lowering the bar to entry to machine translation from a software development perspective. But doing things in this way will only get you so far and, ultimately, without understanding the underlying complexities of the software, of the languages being translated, and of how to compliment and extend upon Moses with additional natural language processing (NLP) techniques, performance will plateau. 

It’s possible to provide interfaces for users to upload their own training data, terms, post-editing rules, and so on, which has been done extensively and often to good effect. Ultimately, however, that’s really over-simplifying the technology and basically taking it out of expert hands. Doing that, as a software provider you have no recourse when someone misuses your technology and it doesn’t meet their needs.
5 steps to adopting machine translation

Software with a Service
To overcome this, we provide MT as a fully managed cloud-service with a different take on the standard business model. We call this model Software with a Service (as opposed to Software as a Service) and what it does is combine technology automation (the software) with specialized expert labor (the service) to deliver a complete solution to a business problem. It’s as much about expert people-powered customer service as it is about code-powered efficiency. The result is awesome customer experience that can only be delivered with a human touch.
MT engine development and maintenance is costed on a professional services basis as required, which serves the purpose of ruling out unnecessary “retraining” when it is not going to bring any value. Usage of the engines, cloud-based or otherwise, is costed as normal based on usage either on a subscription basis or pay-as-you-go. This can change to standard license model when the software needs to be installed on local servers, which is typically done for security reasons.
Buyers of MT are smart, and they are increasingly aware that assessing MT based on output samples or bake-offs (comparisons) between providers doesn’t make sense. MT is never going to be as good as it can be straight out-of-the-box. It takes time, and requires expert guidance to develop, incorporate, and maintain. This is why we’ve had a good response to this model - it just makes sense. 

When should MT be used in the first place?
As I mentioned, MT is not a one-size-fits-all solution and, in fact, in some cases it’s not suitable at all. That’s why the first step in all of our engagements is to assess whether it’s feasible to use MT in the first place. Never mind other business models; the worst business model is accepting projects that are destined to fail from the outset! The industry has been plagued by false dawns and lack of expectation management for a long time, so MT providers need to become more transparent about when their solutions should and should not be used. 

How do we assess feasibility? At Iconic, we look at 8 factors that can influence things in different ways when it comes to MT. In order for us to have a clear picture in our mind as to the best way to proceed in each case, we need to understand these factors, interpret them, and then apply our experiential-driven expertise in order to communicate to the customer how they impact things. We do this even if it means turning down a project because it wasn’t suitable. We discussed these factors in a webinar recently and we’re currently running a blog series on our website where we’re diving even deeper on each factor. 


The 8 factors that influence machine translation projects


The Next Frontier? The Neural Frontier?
So, where do we go from here? For all MT providers, there’s no hiding from the fact that neural MT is not just a fad and is most likely here to stay, and Iconic is no exception. That being said, there has been a lot of talk and hype about it, but relatively little action outside of academia (yet), in part due to the fact there are a still a few hurdles to overcome before it becomes a practicable commercial solution.

We pride ourselves on our deep MT expertise and our ability to develop complex language solutions, so we are very keen and excited to begin working with neural MT. To that end, we’ve been working on an exciting project with ADAPT research centre in Dublin, who have been pioneers in MT research over the last 15 years. We will be in a position to reveal more later this year, so stay tuned! 

Thursday, July 14, 2016

When MT does not take translators' jobs away - and may create more jobs

This is a guest post by Silvio Picinini who works in a team at eBay that provides linguistic feedback and addresses linguistic issues, specifically to enhance large scale MT projects underway at eBay. To my mind this is an example of best practices in MT, where you have NLP and MT experts working together with linguists to solve large scale translation problems in a collaborative way.  

The eBay linguistic team has actually been producing a number of articles that describe various kinds of linguistic tasks that are increasingly needed to add value and quality to large scale MT efforts. I think these articles are worth greater attention, as they have a high SNR (signal to noise ratio.) They are educating and informing readers of very specific things that IMO together add up to examples of best practice. I am hoping that Silvio and his colleagues become regular contributors to this blog so that more people get access to this valuable information.
------------------------------------------------------------------------------------

I was honored to be invited by Kirti to write for this blog. I hope to deserve it, by sharing my experiences as a translator working with machine translation. Recently I was really impressed by Kirti's post on how a lot of content is being translated outside of the translation services industry. I would like to add a few thoughts to that.
I work with User-Generated Content for eBay. Users all over the world describe what they are selling, creating titles and descriptions for their items. In the millions. We need to translate the information on these items so that users that speak other languages can buy them. So this is the job, translate millions of items quickly, almost instantly. A new initiative at eBay is structuring data in a different way, and making it easier to create product reviews. In a short period, we accumulated millions of reviews. A review written in English about a digital camera (a product sold globally) is probably very useful for a buyer in Germany or in Mexico. So we need these reviews translated for these buyers. Could we do this hiring human translators? No. It is easy to see that given the volume, time and cost involved, human intervention is out of the question. Virtually anything that is open to users, allowing them to create their own content, will generate volumes that are not feasible to be translated by hand. These are real scenarios from eBay, but also Facebook recently announced the translation of posts with their own MT engine, and Amazon is working on MT

In addition to what is already happening, we live in a world where new forms of content created by users appear every day. This is of interest to a lot of people, and that will require translation. So here are some types of User-Generated Content that, in my opinion, seem that will be of interest beyond their original language. I am guessing that their companies may be interested in translating this in the (near) future:
  • Rental Homes reviews on Airbnb
  • TripAdvisor reviews of places to see, eat and stay
  • Netflix movie reviews
  • LinkedIn articles
  • Tweets
  • How-to guides
  • Knowledge bases
  • Even Yelp reviews that seem local can be of interest to visitors from other countries or speakers of a second language in the same country (French in Canada, Spanish in the US)
  • In e-commerce: Product titles and descriptions, product reviews, messaging and user searches.

So this is what I meant with the title of the post: Translators would never be offered User-Generated Content translations, so when these jobs go to machine translation engines, they are not really affected by it in any way. MT is not taking any translator's jobs if there was no job in the first place. But maybe translators would like to affect this enormous translation market. Kirti has been posting guidance on how translators can prepare to participate in this opportunity. 
From my experience at eBay, here are a few thoughts about the role that translators may play.
  • MT engines will need to be trained. The specific content needed for training may not be available to be harvested. Therefore, companies will need to create training data for their engines. This training data will be post-edited from the MT output, and this is a job that requires the human intervention of post-editors and reviewers. The quality of the MT output needs to be measured, and the measurement requires (in the case of BLEU) a human translated reference. So there is also a role for translators, instead of post-editors, in creating references for MT measurements.
  • The importance of the pattern over the individual error: the usual mindset for translators and reviewers is to focus on every error that they see, correct them and then produce perfect quality. For MT, the mindset should focus on patterns of errors. Translators will be trying to make a bigger impact by finding patterns of errors that will improve the quality on a larger scale, on every better translation that the MT engine produces.

Translators have the linguistic ability to see these patterns. In this paper at AMTA 2014, I presented a few patterns found in Brazilian Portuguese:
  • Diminutives are widely used by users in informal language, and are not commonly present in the training data, which is usually in a more formal language.
  • The lack of diacritical marks is common among users, both for accents and for marks that modify letters such as ç, ã and õ. The usual training data is usually written in a more formal language and will contain all the diacritical marks. The MT will have to deal with these differences, such as "relogio" vs. "relógio" and "calca" vs. "calça".
  • Some words are intended for the target language but are also words in the source language, causing issues. "Costumes" is a word in English, but also in Portuguese.
  • Some words are misspelled because certain letters have the same sound, causing issues for MT. For example, "engraçados" spelled as "engrassados" (ç and ss have the same sound).
  • Some words are spelled as people pronounce them, and this is different from the correct written pronunciation. For example, "roupa" spelled as "ropa". MT needs to deal with that.
  • Some English words are spelled as they would be written with Portuguese language rules. So "Michael Jordan" would become "Maico Jordam". 
 

There are MT companies, academic experts and customer engineering teams working with MT. It may be time for the language experts to play a role. 


Silvio Picinini is a Machine Translation Language Specialist at eBay since 2013. With over 20 years in Localization, he worked on quality-focused positions for an LSP, and as an in-house English into Brazilian Portuguese translator for Oracle. A former quality engineer, Silvio holds degrees in electrical and electronic/software engineering. 
 
LinkedIn: https://www.linkedin.com/in/silviopicinini

Thursday, April 21, 2016

The Concept of MT Maturity

I would like to introduce a series of posts that looks at and discusses the concept of MT Maturity. I hope to illustrate that getting real business advantage from MT requires alignment with a variety of other related business processes. I also plan to look at how the MT technology continues to evolve, especially as related to use in professional translation and corporate use to make large amounts of content multilingual. 

To facilitate the discussion I have created my own very rough evaluation and analysis model, which is an adaptation of the CSA Localization Maturity Model which is an adaptation of the software industry’s capability maturity model (CMM). Basically, it is a way of assessing if the technology user understands the technology, and is also using it in an efficient and effective manner by properly linking it to other organizational processes. 

Thus, companies that are at a higher stage of the CSA LMM model referenced above, will tend to have much more responsive and adaptive localization practices, that enable them to be much more nimble and efficient and effective in international market initiatives. These maturity models define very specific “stages” to characterize the efficiency and effectiveness of the business process as shown below. Thus an LSP at a higher stage is probably much better to work with, and an Enterprise at a higher level is also much more nimble and international market savvy and superior localization practices.





While this is interesting, it is somewhat theoretical, and I am going to attempt to make it less so in my own analysis of MT technology. This will make my comments less academically robust, but since all I do here is just speak my opinions on things it does not really matter. I do NOT represent any corporate opinion here, and these comments are purely personal observations though hopefully still valid professional assessments. My intention is to highlight best practices and point out what I think are interesting trends.

MT (machine translation) or Automated Language Translation (ALT) or Machine Pseudo Translation (MpT) use can also be described to some extent using these maturity model stages, however, I am going to try and simplify this further, as I only use the maturity model perspective to better structure my comments and analysis, of how this might apply to the business use of MT technology to further international initiatives.

So while there are still many naysayers who will never see any value to professional business use of MT, there is a growing community of users that are working through the vagaries and complexity of MT with varying degrees of success. The Gartner Hype Cycle curve below provides a useful graphic to describe the stages from a very common expectations cycle, that so many, if not most users of this technology seem to go through. I will describe the various maturity stages in more detail in upcoming posts though my analysis may not be quite as tightly defined as the CSA analysis.



It may be useful at the outset, to list some of the most common pitfalls, which are the opposite of best practices, as the continual recurrence of these factors seem to plague many business use cases of MT technology. They are:
  • Looking for cheap, fast and “easy” approaches like Moses and Instant Do-it-yourself solutions.
  • Expecting to do better than the really decent generic engines provided by Microsoft and Google without investments in time and core skill building.
  • Looking to use MT as a wholesale replacement for human translation.
  • Using MT for one-off projects and/or for small volume requirements (LT 500,000 words). MT makes most sense for very large projects that involve millions of words, especially those that would not make sense to do using only human translators for time and cost reasons.
It is worth re-stating that MT technology in the right hands and right use cases can generate long-term sustainable competitive efficiencies, but this generally needs a strategic focus, and a long term commitment to the technology deployment to achieve this.


Emerging MT Trends

While the “translation industry” only focuses on the older SMT approaches most often based around Moses, we are seeing an increasing and building momentum around newer approaches to building MT engines. These involve techniques like deep learning, neural nets and artificial intelligence to improve on the results of the current approaches. These new approaches are apparently yielding much better results than ideas like morpho-syntactic approaches to SMT which have also had small imporvements.

The presentation here presents some of the new perspectives in AI and Neural MT versus the traditional SMT NLP views in a relatively understandable way.

I have also noted that Microsoft appears to have seriously stepped up their MT technology of late. They are attempting to solve the most difficult automated translation challenges in the world today, specifically:
  • Facebook comments (so I now understand what Clio Schils, Renato and my Russian Facebook friends are talking about)
  • Skype based multilingual voice conferences
  • Customization with limited sets of training data using AI and allowing four levels of customization.
Also of interest, a new kind of interactive MT solution that I think holds great promise, especially for individual translators, are adaptive MT solutions like Lilt that take real time corrective feedback and improve dynamically while also leveraging TM in multiple ways. They are also built on post-Moses technology, so provide much better foundation engine quality which means less post-editing. While only available for a limited set of languages, I think it is one of the few MT technology initiatives that actually generate enthusiasm from translators. For an individual translator, I think that using something like Lilt is a much better approach than trying to build your own Moses engine since you have real expertise building your foundations. This article provides a good overview of what adaptive MT looks like and tries to do. 

I found out that Prince died today and while I was never a real fan, I think his guitar solo here is quite exceptional, and should leave no doubt to his amazing musical ability. Theatrics aside, it is the notes and varied textures that he plays that make it so special, and wins the respect and admiration of the other band members.  It is worth a listen. May he rest in peace.



Peace.

Friday, May 30, 2014

Monolithic MT or 50 Shades of Grey?

In the many discussions by different parties in the professional translation world involving machine translation, we see a great deal of conflation and confusion because most people assume that all MT is equivalent and that any MT under discussion is largely identical in all aspects. Here is a slightly modified description of what conflation is from the Wikipedia.
Conflation occurs when the identities of two or more implementations, concepts, or products, sharing some characteristics of one another, seem to be a single identity — the differences appear to become lost.[1] In logic, it is the practice of treating two distinct MT variants as if they were one, which produces errors or misunderstandings as a fusion of distinct subjects tends to obscure analysis of relationships which are emphasized by contrasts.
However, there are many reasons to question this “all MT is the same” assumption, as there are in fact many variants of MT, and it is useful to have some general understanding of the core characteristics of each of these variants so that a meaningful and more productive dialogue can be had when discussing how the technology can be used. This is particularly true in discussions with translators as the general understanding is that all the variants are essentially the same. This can be seen clearly in the comments to the last post about improving the dialogue with translators. Misunderstandings are common when people use the same words to mean  very different things.

There may be some who view my characterizations as opinionated and biased, and perhaps they are, but I do feel that in general these characterizations are fair and reasonable and most who have been examining the possibilities of this technology for a while, will likely agree with some if not all of my characterizations.

The broadest characterization that can be made about MT is around the methodology used in developing the MT systems i.e. Rule-based MT (RbMT) and Statistical MT (SMT) or some kind of hybrid as today users of both of these methodologies claim to have a hybrid approach. If you know what you are doing both can work for you but for the most part the world has definitely moved away from RbMT, and towards statistically based approaches and the greatest amount of commercial and research activity is around evolving SMT technology. I have written previously about this but we continue to see misleading information about this often, even from alleged experts. For practitioners the technology you use has a definite impact on the kind and degree of control you have over the MT output during the system development process so one should care what technology is used. What are considered valuable skills and expertise in SMT may not be as useful with RbMT and vice versa, and they are both complex enough that real expertise only comes from a continuing focus and deep exposure and long-term experience. 

The next level of MT categorization that I think is useful is the following:
  • Free Online MT (Google, Bing Translate etc..)
  • Open Source MT Toolkits (Moses & Apertium)
  • Expert Proprietary MT Systems
The toughest challenge in machine translation is the one that online MT providers like Google and Bing Translate attempt to address. They want to translate anything that anybody wants to translate instantly across thousands of language pairs. Historically, Systran and some other RbMT systems also addressed this challenge on a smaller scale, but the SMT based solutions have easily surpassed the output quality of these older RbMT systems in a few short years. The quality of these MT systems varies by language, with the best output produced in Romance languages (FR, IT, ES, PT) and the worst quality in languages like Korean, Turkish and Hungarian and of course most African, Indic and lesser Asian languages. Thus the Spanish experience with “MT” is significantly different to the Korean one or the Hindi one. This is the “MT” that is most visible, and most widely used translation technology across the globe. This is also what most translators mean and reference when they complain about “poor MT quality”. For a professional translator user, there are very limited customization and tuning capabilities, but even the generic system output can be very useful to translators working with romance languages and save typing time if nothing else. Microsoft does allow some level of customization depending on user data availability. This type of generic MT is the most widely used “MT” today, and in fact is where most of the translation done on the planet today is done. The number of users numbers in the hundreds of millions per month. We should note that in the many discussions about MT in the professional translation world most people are referring to these generic online MT capabilities when they make a reference to “MT”.

Open Source MT Toolkits (Moses & Apertium)

I will confine the bulk of my comments to Moses, mostly because I pretty much know nothing about Apertium other than it being an open source RbMT tool. Moses is an open source SMT toolkit that allows anybody with a little bit of translation memory data to experiment and develop a personal MT system. This system can only be as good as the data and the expertise of the people using the system and tools, and I think it is quite fair to say that the bulk of Moses systems produce lesser/worse output quality than the major online generic MT systems. This does not mean that Moses users/developers cannot develop superior domain-focused systems but the data,skills and ancillary tools needed to do so are not easily acquired and I believe definitely missing in any instant DIY MT scenario. There is a growing suite of instant Moses based MT solutions that make it easy to produce an engine of some kind, but do not necessarily make it easy produce MT systems that meet professional use standards. For successful professional use the system output quality and standards requirements are generally higher than what is acceptable for the average user of Google or Bing Translate. 

While many know how to upload data into a web portal to build an MT engine of some sort, very few know what to do if the system underperforms (as many initially do) as it requires diagnostic, corpus analysis and identification skills to get to the source of the problem, and then knowledge on what to fix and how to fix it as not everything can be fixed. It is after all machine translation and more akin to a data transformation than a real human translation process.  Unfortunately, many translators have been subjected to “fixing” the output from these low quality MT systems and thus the outcry within the translator community about the horrors of “MT”. Most professional translation agencies that attempt to use these instant MT system toolkits underestimate the complexity and skills needed to produce good quality systems and thus we have a situation today where much of the “MT” experience is either generic online MT or low quality do-it-yourself (DIY) implementations.  DIY only makes sense if you really do know what you are doing and why you are doing it, otherwise it is just a gamble or a rough reference on what is possible with “MT”, with no skill required beyond getting data into an up loadable data format.



Expert Proprietary MT Systems
 
Given the complexity, suite of support tools and very deep skill requirements of getting MT output to quality levels that provide real business leverage in professional situations I think it is safe to say that this kind of “MT” is the exception rather than the rule. Here is a link to a detailed overview of how an expert MT development process would differ from a typical DIY scenario. I have seen a few expert MT development scenarios from the inside and here are some characteristics of the Asia Online MT development environment:
  • The ability to actively steer and enhance the quality of translation output produced by the MT system to critical business requirements and needs.
  • The degree of control over final translation output using the core engine together with linguist managed pre processing and post-processing rules in highly efficient translation production pipelines.
  • Improved terminological consistency with many tools and controls and feedback mechanisms to ensure this.
  • Guidance from experts who have built thousands of MT systems and who have learned and overcome the hundreds of different errors that developers can make that undermine output quality.
  • Improved predictability and consistency in the MT output, thus much more control over the kinds of errors and corrective strategies employed in professional use settings.
  • The ability to continuously improve the output produced by an MT system with small amounts of strategic corrective feedback.
  • Automatic identification and resolution of many fundamental problems that plague any MT development effort.
  • The ability to produce useful MT systems even in scarce data situations by leveraging proprietary data resources and strategically manufacturing the optimal kind of data to improve the post-editing experience.
   So while we observe many discussions about “MT” in the social and professional social web, they are most often referring to the translator experience with generic MT as this is the most easy to access MT. In translator forums and blogs the reference can also often be a failed DIY attempt. The best expert MT systems are only used in very specific client constrained situations and thus rarely get any visibility, except in some kind of raw form like support knowledge base content where the production goal is always understandability over linguistic excellence. The very best MT systems that are very domain focused and used by post editors who are going through projects at 10,000+ words/day are usually very client specific and for private use only and are rarely seen by anybody outside the involvement of these large production projects. 

It is important to understand that if any (LSP) competitor can reproduce your MT capabilities by simply throwing some TM data into an instant MT solution, then the business leverage and value of that MT solution is very limited. Having the best MT system in a domain can mean long-term production cost and quality advantage and this can provide meaningful competitive advantage and provide both business leverage and definite barriers to competition.

In the context of the use of "MT" in a professional context, the critical element for success is demonstrated and repeatable skill and a real understanding of how the technology works. The technology can only be as good as the skill, competence and expertise of the developers building these systems. In the right hands many of the MT variants can work, but the technology is complex and sophisticated enough that it is also true that non-informed use and ignorant development strategies (e.g. upload and pray) can only lead to problems and a very negative experience for those who come down the line to clean up the mess. Usually the cleaners are translators or post-editors and they need to learn and insist that they are working with competent developers who can assimilate and respond to their feedback before they engage in PEMT projects. I hope that in future they will exercise this power more frequently. 

So the next time you read about “MT”, think about what are they actually referring to and maybe I should start saying Language Studio MT or Google MT or Bing MT or Expert Moses or Instant Moses or Dumb Moses rather than just "MT". 

Addendum: added on June 20

This was a post that I just saw, and I think provides a similar perspective on the MT variants from a vendor independent point of view. Perhaps we are now getting to a point where more people realize that competence with MT requires more than dumping data into the DIY hopper and expect it to produce useful results.

Machine translation: separating fact from fiction

Wednesday, September 18, 2013

Understanding MT Customization

While we have reached a point in time where many more people realize that machine translation (MT) produces the best results when it is properly customized, what customization actually means is still not well understood. 

There is a significant difference between shallow customization and deep customization in terms of the impact on the MT system’s output quality. The quality of output in turn has a direct impact on the potential business leverage and return on investment. There are a growing number of MT vendors, but very few real MT developers in the market today, and deep expertise is the key differentiator that leads directly to better output and better productivity. It is important for anyone considering purchasing an MT solution to understand the difference between the two types of vendors.

Generally, MT developers have created either Rules Based Machine Translation (RBMT) or Statistical Machine Translation (SMT) systems with hands-on coding at the deepest levels of the core MT engine and its surrounding technologies. Thus they are likely to have insight into how and why an MT engine works the way it does. They are also more likely to be able to coax an engine to produce better quality output by applying the optimal corrective actions to improve on initial results.

In contrast, most Do-It-Yourself (DIY) MT vendors provide little, if any, real innovation and focus on simplifying and packaging a collection of open source tools into a web based offering. Their primary emphasis is on simplifying the interface to these open source tools and enabling a user to build a basic MT system with user data instantly. I would characterize this approach as a shallow customization. When real understanding of the engine technology and data is required, few have the necessary skills needed to make this initial MT engine quality better on an ongoing basis and even less ability to make it reach levels of quality that provides real competitive advantage.

When evaluating MT vendors, there are a few simple things that anyone considering purchasing an MT offering should understand:

Is your MT vendor a serious developer of MT technology or do they simply provide/package other third-party or open source MT technology?
There are only a very small number of companies developing commercial enterprise class MT. Most MT vendors are users or packagers of third-party technology. Many do not have the depth of understanding to do anything but the simplest and shallowest MT customization tasks. These vendors will often present themselves as experts and sometimes claim to be technology agnostic. Some Language Service Providers (LSPs) that have a few years’ experience using open source or third party RBMT systems are presenting themselves as MT experts. Be wary of any vendor that claims deep experience in multiple MT technologies. Advanced skills in any MT technology require long-term investment and long-term experience to get to any kind of distinctive expertise. To get good results from any of these approaches require very different skill-sets and independent and unique expertise must be developed for each approach. The notion that a standard set of MT development skills that work anywhere and everywhere is a myth.

Any MT vendor that does not have a strong and experienced human skill and human steering component in the customization process will always deliver lower quality results. 

Does your MT vendor use a Clean Data SMT or Dirty Data SMT strategy?
The Clean Data SMT approach was pioneered by Asia Online in 2008. Most MT vendors do not have the technology or rigorous data analysis and data cleaning processes to deliver a Clean Data SMT approach, and so take the easier Dirty Data SMT approach. Clean Data SMT has many benefits such as more rapid improvement from post editing and provides the ability to manage and control terminology so that it is consistent. Dirty Data SMT by its very nature is unpredictable and inconsistent and therefore is difficult to manage and much slower to improve with corrective feedback.

Does your MT vendor claim that MT is easy?


3 Monkeys
Some MT vendors claim that MT is not complex. One DIY MT vendor even likens those who claim MT to be complex to be monkeys. The reality is that running an open source MT solution or using a “upload and pray” solution like that of many DIY MT vendors has become very easy. 

Building an instant MT engine is not the same as delivering a production quality MT system that provides production efficiency. Indeed, a significant number of DIY custom MT engines deliver translation quality well below the quality of Google.
 
Delivering high quality MT requires skill, a deep understanding of the different approaches to MT and the inner workings of the technology, a deep understanding of the data used to engineer the customization process and a range of tools, skills and knowledge that permit optimization to deliver the highest possible quality. There will be unique requirements to each and every engine – after all, the point of customizing is to match the translation output to a particular customer writing style and audience. This can only be achieved with human cognition and guidance and cannot be fully automated. 

The impact of real expertise is clear. Asia Online customers can speak on the record of achieving productivity gains greater than 300%, while DIY MT vendors typically claim that they can deliver productivity gains between 20-40% if any at all.

Does your MT vendor give you control of the data and the process?
Many MT vendors today provide very limited control of core data elements and typically rely on a simple “upload and pray” web interface that promises instant results. They generally lack the ability to manage, control and normalize data used to customize an MT engine and generally do not have any data analysis and data manufacturing capabilities. A developer like Asia Online provides multiple levels of control, both during the development and translation process that enable much better output quality and thus higher productivity.

What is expected of you as a user in order to customize and MT engine?
If the answer is nothing more than uploading your translation memories (TMs) then a red flag should already be raised. Machine translation can be very high quality when managed with expertise, but expecting good results without any knowledge investment and real expertise is not realistic. 

Just as in any high quality focused human translation project management, special tools, processes and expertise are required to get better results. 

Any custom MT technology that does not require your involvement in steering the customization process will deliver considerably lower quality output - often worse than anybody could do with Google or Bing. MT systems that produce good quality output require human steering, guidance and control. This is possible with today's technology, but does require more effort than just uploading some translation memories.
 
How much effort does it take and how quickly can the customized engine improve after the first version?
Dirty Data SMT systems offered by DIY MT vendors require significant amounts of new data to improve the system after an initial system is in place, usually around 25% of the total training data that the custom MT engine is built on. So if your engine has 3 million segments provided by your MT vendor and 200,000 segments provided by you, to see any improvement you will need at least 640,000 new segments to see a noticeable improvement in quality. Getting this much additional data is usually beyond the reach for nearly all users of MT systems. As the customization approach is Dirty Data SMT, errors are very difficult to trace and correct. The standard means to correct issues is to add more data and hope that the problem is resolved. 

Clean Data SMT systems such as Asia Online’s Language Studio™ can learn and improve with just a few thousand edits and every edit counts. Terminology is consistent, and there are tools to identify common problems ahead of time and means to automatically resolve them. Data manufacturing is also applied to amplify edits and corrective feedback and ensures they are applied to the engine in a broader set of contexts. The cause of errors can quickly be traced and the errors can be rectified using a number of problem analysis tools. The resulting improvement is rapid and noticeable even with a very small effort by a single person. 

Bottom Line: Creating a high quality custom MT engine requires deep expertise, control and broad experience, elements that are usually not present in the "upload and pray" approach provided in a DIY MT model. Developing high quality MT is complex and in 2013 still an expertise based affair. 

To simply upload a translation memory and expect high MT quality to come out is wishful thinking. A computer cannot automatically know your preferred terminology, vocabulary choices, writing style, target audience and purpose. Just like a human translation project, achieving quality requires effort, time, management and skill. 

Customizing an MT engine to produce “near-human” output quality levels is possible and there are many proof points where raw MT output has been able to produce 50% or more of the MT translated segments requiring no editing at all - i.e. they were “perfect”, with many of the remaining segments having minor issues that could be quickly edited. A fully customized MT engine built on the Clean Data SMT approach consistently deliver 150%-300% (sometimes even greater) productivity gains. The long-term ROI impact is clear relative to the meager productivity that instant MT approaches sometimes produce. 

MT in 2013 is still a complex affair that requires deep expertise and collaboration with experts if your intention to build long-term business leverage through more efficient translation production processes. There is no advantage to a system that any of your competitors could create instantly and there is no value or business advantage to just dabbling with MT.

“When conceiving the idea of Moses, the primary goal was to foster research and advance the state of MT in academia by providing a de facto base from which to innovate from.

Currently the vast majority of interesting MT research and advancements still takes place in academia. Without open source toolkits such as Moses, all the exciting work would be done by the Google’s and Microsoft’s of the world, as is the case in related fields such as information retrieval or automatic speech recognition
.
Philipp Koehn

As a platform for academic research, Moses provides a strong foundation. However, Moses was not intended to be a commercial MT offering. There are considerable amounts of additional functionality, beyond providing a web based user interface for Moses, that are not included in Moses that are essential in order to offer a strong and innovative commercial MT platform. “ 
Professor Philipp Koehn, University of Edinburgh, Chief Scientist, Asia Online

Addendum: This post triggered a strong reaction from Manual Herranz at Pangeanic and I am including a response I made on his blog in case it does not make it through the approval process there. 


My primary point in my blog posting is that expertise, long-term experience and a real understanding of how the technology works is necessary and critical to get the best results. Most DIY users do not have these characteristics and thus are very likely to get much lower quality results. Pointing this out, to my mind is not equivalent to "bad mouthing competition",  I am simply comparing approaches and pointing out the value implications.

Also, while I claim that expertise does matter, I do not suggest that Asia Online is the only company with this expertise. There are several other MT experts including RbMT developers like Systran and a specialist like Tayou in Spain.

I do believe that  MT technology is complex enough that it does require specialization, and that developing real competence with MT is difficult enough that it is unlikely to be successfully done by a company  whose primary business is being a translation agency. It is clear that you disagree.  I am also pointing out that the value received by a customer is very likely to be lower for a DIY user. I can understand that you may have a different opinion to mine and assure you that my observations are not borne of virulence.

Historically we saw many LSPs develop their own TMS systems too, but most people in the industry would concede that  the best TMS systems have come from companies that focus and specialize in the development of these tools e.g. MemoQ, Memsource, Across, XTM etc.. We have also seen the SDL acquisitions of software companies like Idiom, LW and Trados result in what most perceive as reduced customer responsiveness, quality and commitment to these products. Buying critical production infrastructure from a competitor generally does not make sense in any industry and thus we have seen the momentum slow down on all the SDL software acquisitions. IMO Specialization matters and with technology this complex, one will get the best results using technology developed and managed by specialists for the foreseeable future.

Anyway, I wish you peace and health.



Kirti