Pages

Tuesday, March 27, 2018

Sharing Efforts to get the most from MT and Post-Editing


This is a guest post from Luigi Muzii which is basically made up primarily of the speaker notes of his presentation at the ELIA Together 2018 conference. I believe that Luigi has wise words and represent best practices thinking and thus offer it on this blog forum.

While some of his advice might seem obvious to some, I am always struck by often these very commonsensical recommendations are either overlooked or simply ignored willfully.

As generic MT improves, I think it becomes more and more important for enterprises to consider bringing rampant, uncontrolled MT use by employees and outsourced translators under control. 

As always the emphasis of bold text in the post is my doing.


====





The presentation was designed to provide some practical advice about tackling the challenges that freelancers, project managers, and translation buyers face when approaching MT, implementing MT or running MT and post-editing projects.


Three target groups can be identified for three kinds of task:
  1. In the making; machine translation is used for “suggestions” while processing a document by a human translator; this is probably the most common approach today;
  2. Downstream; machine translation is used to fully process a document that will be possibly post-edited; this approach is typically adopted by larger clients with established experience in the field;
  3. On constraints; machine translation is used by an LSP to finalize a translation job by asking translators to work on suggestions coming from a specialized engine.
While scenarios two and three might meet the customer’s need for confidentiality and IP protection through an in-house engine using only the client’s own data, scenario one is becoming more and more general among professional translators, given the astonishing improvement of online engines. At the same time, scenario three is slowly but constantly applying to LSPs who try to escape price and volume pressures through machine translation and post-editing.

Laying foundations

The three scenarios just outlined require different strategies to be devised. The first one involves the machine-translation method.

Given the circumstances, the premises and the many reservations about it—basically, the hype—a question must be asked: Is PEMT already in the past?

In this paper, the reference method is PB-SMT (Phrase-Based Statistical Machine Translation) because PB-SMT engines are inexpensive and effective, whereas customizable NMT (Neural Machine Translation) engines are still quite pricey and challenging as to technical requirements and operational complexity, thus out of range for most customers. Translators working on unrestricted documents (scenario one applies here,) would generally choose an online NMT engine. For customers requiring confidentiality and IP protection and willing to leverage their own language data (scenario two and three apply here,) a highly customizable PB-SMT engine might be a valid option, especially where no major investment is envisaged in the field, vast and suitable data is available and/or limited volumes are processed.

In general, the main drivers in the adoption of MT are productivity (speed and volumes) and usefulness (consistency and marginal cost) especially for large projects otherwise involving many translators. Unfortunately, MT is not exactly child’s play. MT engines are complex and challenging applications requiring a combination of tools, data, and knowledge. This is a rare commodity, especially in a single person, on both sides of a translation project.

Also, in the future, while MT will continue to proliferate, the shortage of good translators will increase.

Join forces

Therefore, joining forces is important to explore and vet as many solutions as possible. According to a popular saying, everyone’s needed, no one’s indispensable, and can be easily replaced by anyone else with a similar profile in the same role. This also means, though, that, to be valuable, everyone’s effort is needed, on the highest level of performance all the time. For quite some time now, translation is no longer a solitary feat, but a collaborative effort. This is especially true with the current level of technology.

In this respect, three steps should be completed before starting any MT project.
  1. Recap your goals and requirements
    What you expect from MT and how much you are willing to rely on it.
  2. Check your MT readiness;
    Realistically analyze your knowledge, tools, and data.
  3. Plan for assistance
    Never venture into an unfamiliar territory without a guide.

Defining requirements

When planning to implement MT, keep scope, goals, and expectations clearly distinct.
Identify one or more items within your scope of business for MT, possibly picking those where the amount of data available is larger, and the quality higher.

Clearly define and prioritize your goals. Major goals may be reducing labor, boosting productivity and keeping consistency, especially in larger projects.

Be realistic in your expectations. Therefore, familiarize with technology and strengthen your expertise to make the best of it. Tackle any security issues for confidentiality and data integrity and protection of intellectual property. Don’t forget to scrub your data if you plan to train an engine and to plan for any relevant support. Finally, revise your pricing model to encompass MT-related tasks.

Building a platform

When building an MT platform, never forget that MT engines are not all equal, for different environments, configurations, and data. Therefore, although the output could be considered someway predictable, raw output quality can vary across systems and language combinations, and error may not follow a consistent pattern. Performances of MT engines also vary.

In data-driven MT, data maintenance is crucial, and it is the first task when setting up an MT platform. Data must be organized, cleaned, and fixed for terminology and style. For a customized engine, at least 100,000 paired segments are necessary and the cleaner and healthier the better.

Another important factor to the effectiveness of an MT strategy are the tools used for data preparation, pre-translation, and post-editing. Special attention must be given to translation tool settings to allow for sub-segment recall and fuzzy match repair.

Finally, when choosing the engine, the items to consider are:
  • The total cost of ownership (TCO)
  • Integration
  • Expertise
  • Security (especially as to intellectual property and confidentiality)

 

Best practices: Running projects

When running MT projects, best practices may be different depending on whether you are a translation buyer or vendor.

In general, knowing your data and mastering quality metrics is a must. As to post-editing, devise your own compensation scheme.

A common mistake is to consider all content as equal and then mess with data. In the same way, absolutely avoid relying only on your capacities or on vendors. In the end, everything can be summed up in the simple invitation to not expect any miracles.

Never forget to agree with the customer and the content owner about using MT, especially if you are using a SaaS/online platform to prevent being sued.

Also, remember that data is the fuel to any SMT/NMT engine and that the output is only as good as the data used. In fact, these engines perform statistical predictions by inference, and when the amount and quality of data increase, an engine improves.

Collect as much data as possible, but always make sure it comes from few reliable sources in a restricted domain, that it contains correct translations with no errors, it is real and recent, and terminologically and stylistically consistent.

At this point, you must accept that MT output is unpredictable. For this reason, MT quality assessment should be run in such a way as to prevent any subjectivity.

For the same reason, post-editing is and will remain a critical, integrated part of MT usage, and it is expected to be fast, unchallenging, and flowing.

Anyway, the amount of post-editing required can be hard to assess. To plan deadlines and allocate a budget for the job, two different measures can be used, the edit time and the post-editing effort. The first is the time required to get a raw MT output to the desired standard, and the latter is the percentage of edits to be applied to raw MT output to attain the desired standard.

The main problem with edit time is that it can only be computed downstream, assuming that the time spent has been entirely devoted to editing.

The post-editing effort can be estimated through probabilistic forecasts based on automatic metrics as a reverse projection of the productivity boost. In fact, translation productivity is measured as the throughput expressed in the number of words per hour, and MT is supposed to improve it by reducing the turnaround time and increasing the workable volumes. However, the post-editing effort and the turnaround time are hard to predict, especially for translation of publishable quality and/or data for incremental training of engines. In fact, it depends on diverse factors such as the quality expectations for the finalized output, the volume of content to process, and the allotted time for the task. Also, the effort required depends on the post-editing level.

The post-editing level is generally restricted to:
  1. Gisting;
  2. Light;
  3. Full.
Gisting consists in fixing recurring errors in raw MT output with automatic scripts. It is used for volatile content, e.g. messaging, conversations, etc. Light post-editing consists of fixing capitalization and punctuation errors, replacing unknown words, removing redundant words or inserting missing ones, and ignoring all stylistic issues. It is generally used for continuous delivery and reprocessing. Full post-editing consists in fixing meaning distortion, grammar, and syntax, translating untranslated terms (possibly new terms), and adjusting fluency. It is reserved for publishing and engine training.
Finally, always follow a few basic rules before boarding on a post-editing project:
  • Test before operating;
  • Ask for MT samples for negotiation;
  • Negotiate throughput rates;
  • Ask for glossary with the list of DNT words;
  • Ask for instructions;
  • Be open to feedback.
Similarly,
  • Never use MT to sustain price competition;
  • Never process poor MT outputs;
  • Never treat post-editing as fuzzy matches.
Remember that different engines, domains, and language pairs produce different outputs, involve different post-editing efforts, and require different post-editing instructions. These should address tool selection criteria and environment setup guidelines, as well as a style guide, and a comprehensive term base. They should also address conventions and the type and number of project details as well as the general pricing model and the actual operating instructions.

As to pricing and compensation, for light post-editing of very good output when speed is a major concern and the first requirement, a model should be settled prior to any assessment of the actual MT output based on a clear-cut predictive scheme. However, do not follow any translation-memory fuzzy matches scheme. In fact, while fuzzy matches over 85% are inherently correct and generally require minor changes, machine-translated segments may contain errors and inaccuracies, and even a supposedly light post-editing may prove challenging. A downstream computation scheme might also be devised in full post-editing for an accurate measurement of the actual work performed. This is usually made by computing the edit distance and then inferring the percentage on the hourly rate.

A negotiation grid can be helpful to cross-reference type and nature of engines, quality of raw output, and all the relevant technical requirements with compensation based on productivity, type of performance, bid (hourly rate) and ancillary services (e.g. filling in QA forms for ongoing training of engine.)

In any case, a strong and clear “No, thanks!” is more than reasonable when a considerably low pay rate is offered that is unrelated to language pair and MT output quality and/or MT output quality is lower than a generic free online service.

Lastly, raw MT output should be processed before a post-editing job for automatic removal of empty and/or untranslated segments and duplications, fixing of diacritics, punctuation marks, extra spaces, noise, and numbers; terminology should also be checked for consistency and a spellcheck should be run. A post-processing stage should also be envisaged involving encoding, normalization, formatting (especially tag injection,) a terminology check and, obviously, a spellcheck.

=======================
Luigi Muzii's profile photo


Luigi Muzii has been in the "translation business" since 1982 and has been a business consultant since 2002, in the translation and localization industry through his firm. He focuses on helping customers choose and implement best-suited technologies and redesign their business processes for the greatest effectiveness of translation and localization related work.

This link provides access to his other blog posts.



Thursday, March 1, 2018

Machine Translation Maturity Model (MTMM)

This is a guest post by Valeria Cannavina, Project Coordinator at Donnelley Language Solutions, adopting the Common Sense Advisory’s Localization Maturity Model (LMM) which is itself an adaptation of the software industry’s Capability Maturity Model (CMM). The resulting Machine Translation Maturity Model (MTMM) is a way of assessing the users’ understanding of the technology, and whether they are using it in an efficient and effective manner, properly linking it to other organizational processes. Valeria provides a framework for businesses to “identify where they are and what they can do to either significantly or modestly improve their existing production model to maximize the value that MT can provide to their organizations.” 

This is, however, a perspective that is quite localization-centric, and process alignment for a global Enterprise MT service that might be used by thousands of users, across an enterprise to translate hundreds of millions of words could be quite different.

As we head into 2018, we continue to see excitement and hype around Neural MT (Machine Translation), which is a breakthrough approach on the verge of providing a wealth of possibilities for the creation and management of business content. But, because the technology is relatively new, many players in the translation industry are overlooking the importance of implementing aligned procedures to guide the use of MT.

Neural MT, or any other kind of MT on its own, is not a magic wand that can solve any and every translation problem. Production and work procedures need to be aligned in an informed and competent way to enable the technology to provide maximum benefits and also minimize risks and data security issues.

It is useful to always ask fundamental questions before embarking on new technology deployment initiatives. The most fundamental question for businesses looking to invest in MT may be:

Why do we want to translate the content at all? 

Content translation only makes sense if it furthers overall business objectives and improves the global customer experience in some way. Today’s markets are massively global and that means communication and collaboration need to happen at scale and in volumes that were inconceivable just a decade ago. Today, any business that seeks to have even a moderately global footprint must understand how MT can provide increasing volumes of relevant content to their customer base.

Customers all over the world expect relevant information at their fingertips as quickly as possible, and this information is increasingly more dynamic and also short-lived. Business information is very important for a brief instant and of very limited value after that. This ability to deliver the right content quickly and effectively is often critical to the impression the customer forms of the business and its product offerings.

In addition, businesses are rightly concerned about data security and privacy. Improperly implemented MT deployments, where key processes and systems are not properly aligned, can expose private and confidential data. As the sheer volume of information continues to increase, businesses need to ensure that security is not compromised when content flows through these new translation processes. This may be especially critical with new product/service developments, sensitive employment, credit-related, medical, and financial data.

As MT becomes much more pervasive, it is wise for us to understand the bigger picture. In this paper, Valeria provides a unique and valuable perspective on assessing organizational alignment with new technology deployments. I hope you find it a useful guide for assessing your business needs around MT.

This post was originally published last month and then removed so that Donnelly could prepare the more complete document that is referenced at the bottom. 

 
 ====

5 different approaches to succeed with machine translation

Introduction

With the ever-increasing volume and pace of global trade, the need to communicate to multiple markets simultaneously has never been greater. Add to this the huge technological advances of recent years, and it’s easy to see why machine translation (MT) has emerged as a translation tool of choice for high volume, high-speed translations, with ever-improving quality.

MT delivers big time-saving and money-saving benefits, plus big gains in productivity. But as is often the case when technology moves at speed, many businesses are lagging behind. While the demand for MT is growing very fast, there are still some basic challenges that clients are not aware of. For example, not all language combinations, documents, and formats types are ideal for MT. In fact, the quality of the output can change considerably based on these criteria which could affect your workflow, the time to market, quality and business targets.

This is where partnering with a specialized Language Service Provider (LSP) can give you the upper hand. A professional LSP will not only help you understand the MT landscape together with the latest developments but also advise on how it can be best used to optimize your processes, productivity and profits.

MT engines, its output, and training require skilled professionals and solid technologies to support automated workflows. Simply put, MT is currently far from being just a plug-and-play technology.

This approach is paramount, especially when confidentiality is key to the process. Assessing the risks of publishing data and securing processes is not a standard practice for all language service providers, so while clients and regulators are setting up very strict measures for data breaches, vendors are struggling to create processes to ensure top quality processes and services for MT.

Whether you are a large or a small business; whether you have a little or a lot of knowledge of MT, this paper will show you how to take full advantage of it. We follow a Machine Translation Maturity Model (MTMM) which is based on the Localization Maturity Model[1] created by Common Sense Advisory, an independent market research company for the language services industry. This paper is a guide to help businesses identify where they are and what they can do to either significantly or modestly improve their existing model to maximize the value that MT can provide to their organizations.

[1] The Localization Maturity Model was created by Common Sense Advisory, an independent Massachusetts-based market research company for the language services industry. It is based on the Capability Maturity Model (CMM), a development model informed by a study of data collected from organizations contracted with the US Department of Defense, who funded the research. The term "maturity" relates to the degree of formality and optimization of processes – from ad hoc practices, to formally defined steps, to managed result metrics, to active optimization of the processes. The model's aim is to improve existing software development processes, but it can also be applied to other processes.

Machine Translation Maturity Model (MTMM)

The model has five maturity levels, each divided into different areas which we encourage companies to evaluate individually on their own merits, allowing you to freely move from one level to another, not necessarily in successive steps, although an ideal path is shown below.

We will now walk through them so you can identify where you are and what you need to do to move to a different level  - and enjoy the associated benefits.


Level 1: Initial


If you're at this level, you are requesting MT only when absolutely necessary. This could be due to time or budget constraints, or because your communication is internal only. It may be that you are not satisfied with the service that you are getting. Or it could be that you have no MT investment, resources or best practices in place. This may be because you don't have a localization department in place, or because it is not directly affecting a core activity within your organization.

This model may work for some organizations with a very limited usage of MT, but who may benefit from taking the following steps to increase its value within their organization:

  • Governance: make a case to management for investment in the maturity process.
  • Organization: appoint a dedicated MT resource in the localization/translation/marketing department who will set the basis for the process investigation.
  • Process: document the main tasks of the outsourcing process so you can track those that can be repeated and those that can be deleted because they don't generate any added value. For example, to track which department translation requests come from, the type of content you receive, and the turnaround times. This will help you start setting best practice for your process.



Level 2: Repeatable


The first step to maturity through process improvement is two-fold: describing what you do, and doing what you've described. If you are at this level, you have started documenting some tasks of the process which are repeatable - for example, your criteria for identifying texts that can be sent out to MT and how to store them by category. You view terminology management as a relatively low priority, whereas in reality, it's an investment that will pay dividends by ensuring consistency.
This is where most organizations may be at and would benefit from taking the following steps to evolve their existing processes:

  • Organization: define an internal process to track feedback on the source text from your LSP. For example, you might have received a comment saying that the text wasn't suitable for MT because it was too creative or overly complex in structure. Make a note of this in the process documentation.
  • Process:
    • analyze your existing content in order to understand exactly which documents can be translated with MT and how the source text is structured.
    • organize linguistic reviews on the translated material involving internal country reviewers, so that terminology management starts to become part of your internal process.
  • Governance:
    • track the costs of the process improvements you are implementing (expectations/forecasts vs. reality).
    • define KPIs for this process to track the ROI of activities involved at this level
  • Automation: investigate available tools for automating some tasks. For example, look for tools to help you build a repository of texts that have already been sent to MT and decide a naming convention. You will then be able to identify similar content and remove text that's not suitable for MT. The automation will be run parallel to the process of analyzing the texts.

 

Level 3: Defined


At this level, you will have clear goals around integrating MT into your business, in the form of a roadmap of tasks aimed at continuous improvement through collaboration with your LSP.
Your processes will be documented and fully executed. The internal process of outsourcing MT is defined, repeatable and managed. You have best practices in place - a process to identify the MT content, a process to collect and implement feedback, a process for internal translation review, a process to track costs etc.. and can now measure your process.

For example, to track productivity you might want to measure word output per hour. Before a process improvement is introduced, a baseline measurement is taken. At the end of the project, the process is measured again to show whether the change resulted in more words produced per hour.

Terminology management is no longer seen as a secondary task, but as a fundamental step which adds value to your business and your brand. The benefits are now showing in your ROI. For example, you can identify which content has already been sent to MT, which means fewer words to process, fewer man-hours for both you and your LSP, reduced time to market and reduced costs.
This is the stage where most organizations should aim to reach in order to optimize their supplier relationship with their language service provider and maximize return on investment. That said, some organizations take additional steps to further mature their MT procurement strategy as follows:
  • Organization: supported by your LSP, hire specialized terminology management staff who will work closely with other departments to:
    • organize feedback received on source text for fields of application
    • pass the feedback to internal technical writers
    • check the feedback has been implemented
    • incorporate the feedback into your CMS or whichever tool you're using to store the source and target documents.
  • Process:
    • define the internal process to combat source linguistic inconsistency. For example, give clear guidelines on what to do if a word has more than one meaning, who is the decision maker, how many review cycles the work will go through, and how this will be implemented in the CMS
    • plan internal review cycles so you can send feedback to your LSP and implement it in the CMS.
  • Governance:
    • establish the budget for multilingual projects based on the forecast for previous MT projects in terms of volumes and languages
    • establish decision making to prioritize languages and markets.
  • Automation: identify a tool to automatically apply correct terminology to source content in the CMS.

Level 4: Managed


MT is now tied to your corporate goals and part of your production process. The different departments rely on the MT department to prepare documents before sending the files to translate. The idea of 'department' here is fluid; for example, it could take the form of resources performing MT tasks alongside non-MT tasks. Alternatively, if you don't have your own MT department you can contact your trusted LSP to help you manage this internally or serve as the department itself.

The size of your business will dictate its scope.

The MT department has its own budget and schedule for incoming projects during the year, with automated checks in place for managing terminology and producing the source documents. The focus at this level is on automation; you will be working closely with the technology department to improve the source i.e., a style guide for writing source text to ensure consistency.

If your organization is in this camp, consider taking the following steps to improve the maturity of your sourcing model:
  • Organization: define roles within the process; for example, a project manager (PM) to handle requests from different departments, dedicated engineers for automation, and internal reviewers.
  • Processes:
    • define delivery parameters around new products/documents. For example, you're issuing a new set of letters to shareholders and, based on previous experience, you know they will need to go through X reviews. This knowledge enables you to set realistic deadlines, review cycles, and delivery volumes
    • define text structure rules (short sentences typically translate much better than long, complicated sentences).
  • Governance: measure business benefits vis-a-vis strategic use of MT budgets.
  • Automation: technology resources work on a roadmap to automation which integrates all the previous stages: from terminology management to check that the source text follows the defined rules.


Level 5: Optimized


At level 5 you have a team of engineers, terminologists, internal reviewers and project managers running the content creation process for MT on a daily basis. You have rules in place for content editors writing source text, ad hoc terminology and an internal tool to check and select the right kind of source text for MT, extracting only the new parts to be translated. You are now looking at new ways to get the same quality output while trying to keep costs and content creation time to a minimum.


What to do to achieve continuous improvement beyond level 5?
  • Organization:
    • prepare training material for staff and build a career path in the MT office
    • plan to offer a global service 24/7/365.
  • Process
    • content creators work to minimize content sent to MT: less content = lower costs
    • customize writing rules based on target language to minimize linguistic disparity between source and target. You will already have writing rules for new content creation, but with your accrued experience, constant LSP feedback, and the help of internal reviewers with in-depth knowledge of the target languages, the rules can be customized further to optimize MT for each target market
    • allocate LSP resources strategically, based on the language combination that best fits your quality expectations.
  • Governance: based on your long-term business goals plan how to support everyone involved in the roadmap to continuous improvement.
  • Automation: connect your CMS with your LSP's Translation Management System (TMS) to:
    • speed up the process
    • send requests and import the MT content automatically
    • centralize review cycles in a familiar environment
    • ensure consistency across content with a shared repository of linguistic assets.
     

Conclusion

Even if localization is not central to your organization, it can and does have an impact on your business. To reap the maximum rewards from MT, working with a trusted LSP is key to strengthening your supply chain, improving ROI and protecting your brand and reputation.

The MTMM was designed for any size or type of organization to use to make the most of MT. It is an ideal path to maturity because it has built-in flexibility – each activity can be performed as a stand-alone step as well as in a sequential way.

We believe that MT is not just about getting the technology right. It’s also about having a strong relationship with your LSP; a partnership that is characterized by collaboration and constant feedback between the whole team. Only by analyzing your process and implementing some of the suggested tasks will you arrive at an MT roadmap which delivers against your expectations.

It is also possible to get a more complete  version of this post directly from  the RRD website at this  link:



 ======*****=====



Valeria Cannavina holds a degree in language and culture mediation, and a master's degree in technical and scientific translation from Libera Università degli Studi “San Pio V” in Rome. She has spent 10 year in the GILT industry as Project Manager and while working for companies like SAP and Xerox she was involved in quite a few research projects on new processes implementation and Machine Translation. At present she works for Donnelley Language Solutions a Project Manager.




Wednesday, February 14, 2018

A Change in Status & The Larger Translation Market Beyond Localization

The Emerging MT-Driven Translation Market Opportunity


As we roll on into the New Year, it is clear that machine translation is now pervasive, universally available, and easily accessible to millions across the globe on all kinds of digital devices. It is often even accessible when we are not connected to the web. This widespread access to, and use of MT is true across the globe and recent advances in translation quality from Neural MT have raised the profile of MT all over again. And while Google is a large provider of generic MT, it is not dominant across the globe, and we see that regional MT portals dominate local markets, e.g. Baidu in China, Naver in Korea, and Yandex in Russia.

By my very conservative estimates, the sheer volume of translation done is astonishing, and I think the global daily use of MT is easily in excess of 500 billion words a day. This amounts to approximately 185 trillion words a year, and this is if MT use does not continue to grow even further as new populations become more digital and connected! While the bulk of this use is by the consumer-at-large on the global internet, I think it is quite possible and even likely that a growing portion of this usage is uncontrolled use by a growing population of enterprise workers and many translators. The need for instant, largely accurate translation has now become a fundamental and often critical business requirement.

MT is much bigger than what we understand as the reality of the total translation need from the localization industry perspective. The quality of MT can vary dramatically, be significantly less than "perfect", and still be hugely valuable. There are many business applications where the output quality requirements are a long way from perfect post-editability for MT to add value to the business mission. It is quite likely that the highest quality MT output today is seen in systems developed by experts in localization, but the evidence provided by the huge demand for raw MT use suggests that the real translation need is greater and broader than understood by localization experts. Based on data provided by Common Sense Advisory of the translation volumes in the localization industry, I estimate that currently the localization industry services less than 1% of the translation needs on the globe on any given day. Perhaps for business content that number is slightly higher, but ever so slightly.

This suggests that the market opportunity for business translation can grow substantially. This growth can only come from a different strategy to what is translated and how it is done than it has been in the historical approach and focus. Thus, possibly the new leaders will create and move to a new industry business model where MT drives other translation activity, rather than the other way around. While addressing this larger market opportunity is only a potential strategy for a select few in the translation industry, I think few are as well positioned to lead the way and define this as SDL is today. Thus, I have decided to join SDL in a senior marketing strategy role as a “Language Technology Evangelist”.

The need for translation goes far beyond the focus of the localization industry

 Why SDL?

I have had the good fortune to closely examine the MT offerings and strategies of many of the leading players and users in the industry over the years, and recently even compare the various solutions and capabilities from both vendor and buyer perspectives. Having an honest and objective outlook is fundamental to seeing accurately, and my role in maintaining this blog has enabled this clarity, as all claims need to be validated and reviewed. Based on my personal assessment over some time, I believe that SDL is in a unique position to capture much of this new MT market opportunity that lies beyond traditional localization and expand the worldview of the business translation market in general. Those who have interacted with me in my consulting role, have already seen that I have been advocating SDL as a superior MT solution (amongst others) for enterprise and professional translation use. There are three primary reasons that I am excited about my new role at SDL:
  1. The Executive Management commitment to the role of MT in the organization and the future business strategy of the company which I touched upon in a post early last year. Their vision and understanding of the role and possibilities of using MT technology to drive deeper and broader engagement with global enterprises, I find is a refreshing contrast to the narrow localization MT perspective that still pervades much of the thinking in the industry. The relatively new management team has a strikingly different perspective from the prior management, and I believe they see the potential of using SDL’s substantial MT competence to drive a larger presence for business translation in the broader enterprise translation market that lies beyond localization.
  2. The depth and overall NLP (Natural Language Processing) and ML (Machine Learning) competence that exists within the technology team is, in my opinion, a significant long-term competitive advantage. In contrast to the teams at Google, FB, and Microsoft, the SDL team is focused on developing MT solutions that are optimized around global enterprise needs and use-cases. It is my opinion that the capabilities and competence of the SDL MT team are beyond the reach of other large LSPs, and also quite probably of many smaller MT vendors. This ability to tune and customize the technology to specialized business use-cases requires deep knowledge of the underlying technology, and those who work with open source toolkits are clearly at a disadvantage. SDL has regularly beat out competitors in customer-driven quality evaluations, but much of this is under the radar and not well known. SDL also has a long history of deploying their MT technology in Government National Security user environments where system robustness, scalability, elasticity and overall manageability REALLY matter. This experience is excellent preparation for the larger enterprise MT market which has many of the same organizational needs. Today, SDL is the only MT technology vendor that has enterprise-grade SMT, Adaptive MT, and Neural MT solutions in place. This competence allows ongoing comparisons internally of how different data sets can be optimized across these technology alternatives to optimize for the business purpose in ways that are not easily possible by competitors. It matters more that the business purpose is achieved and much less what technology is used. The engineering team here has many patents to their credit and have published hundreds of papers on ML strategies, SMT techniques, NMT and Adaptive MT.
  3. The experience of Internal Translation teams working with MT over an extended period has built unique capabilities and a deep understanding of successful MT technology deployment. SDL is one of the few language service companies that have a large team of internal translators. This fact enhances the possibilities of building a more comprehensive Linguistic Steering & MT Development collaboration when the highest quality output is required. This is one of the reasons why SDL has such a robust Adaptive MT offering. The deep collaboration between expert NLP engineers and expert linguists has only just begun, and I believe that as this engagement expands, SDL will have the potential to build distinctively and consistently superior MT engines that are optimized for enterprise use. Enterprise MT solutions also require specialized process infrastructure that surrounds MT in a business use scenario, and the long-term internal MT use experience that SDL has had provides great insight into building this process infrastructure for clients whether they are global enterprises or LSPs. The quality and robustness of the surrounding processes around MT deployment are key to maximizing MT benefits. This interaction between computational linguists and human linguists where needed could also enable new kinds of AI and machine learning driven process improvements at several different points in new business translation workflows. Having competent linguists involved with your MT systems also results in data refinement over time. Many are now beginning to realize that the data is the highest value asset in a world where the best machine learning technology is only as good as the data it learns from. The existence of a linguist team that understands the learning implications of "good data" vs. "bad data" creates the possibility of having the most valuable, cleanest and most relevant data for the new AI frontiers.

When these three factors are placed in the core competency portfolio and context of a company that already has deep penetration into Global 1000 Enterprises and Major Government Agencies around the world, the potential for further success is indeed great. The initial engagement with the Global 1000 that SDL already has makes the likelihood for a broader and deeper business relationship to develop around new international initiatives and needs even more likely. From my historical vantage point as an independent industry analyst, it is easy to see that the SDL MT offerings are already formidable and unique, and hopefully, with my involvement this will only become more so.

The many paths to driving SDL MT quality above generic public MT solutions


Personal Factors


While it is important to consider these relatively objective assessments of organizational strengths and characteristics in choosing one’s employment, I think the personal factors around the choice are possibly even more important. The following are the ones that immediately come to mind.

The People: I have enjoyed my interactions with many SDL employees in my various roles as an analyst, collaborator, and potential employee. The straightforward and clear communication styles and professional personal manner of the key executives I dealt with encouraged me to consider taking on this new role with minimal hesitation and consideration. In a way this was also a homecoming, as I originally entered the world of translation in a leadership position with Language Weaver (acquired by SDL in 2010), except now it has better management, better focus, and more resources. And also hopefully I too am wiser and smarter, as my new grey beard suggests.

I plan to keep this blog going and will continue to produce or curate what I consider to be high-value content on the subjects of Translation, MT, AI, ML, and Collaboration. This is done with SDL Executive permission (and possibly even their blessings) and again points to the openness of the SDL culture and new management. SDL does not necessarily agree with everything that I say on this blog but sees the value of serious content sharing forums. I will, of course, endeavor to keep this blog from becoming a marketing mouthpiece for SDL MT alone, and will try my best to keep it objective and issue focused. There is great value in intelligent dialogue and discussion, and I will strive to maintain this here. While this post clearly shows where my allegiances and biases lie from this point on, I will always strive to be fair and allow different viewpoints to be heard. I have always considered it a great and important skill to listen, to listen to everyone, especially those with a different worldview. I have been a long-term MT and translation technology evolution advocate, but I have also found that it is possible to communicate respectfully and meaningfully with translators and opponents of MT. The regular contributions by translators as guest writers or in the comments on this blog is a testimony to that. I hope to maintain that open dialogue quality and the overall respectful cultural characteristic of this forum. I invite guest writers who have an interest in sharing insights and new perspectives on the primary themes of this blog.

There is also great value in being an insider and seeing the challenges from the inside in contrast to the view that is offered to a consultant. As an insider one has the opportunity to shape and change the outcomes in a much more substantial and organic way. I look forward to this new insider role in helping to shape the future of SDL MT. This role will also allow me to engage with many more actual MT deployments and possibly have a role in helping to drive many more successful outcomes.

My first public-facing SDL event will be a webinar on February 27th. I look forward to seeing you there and I look forward to ongoing honest and forthright dialogue.

Peace.