Pages

Showing posts with label translation technology. Show all posts
Showing posts with label translation technology. Show all posts

Tuesday, September 20, 2022

The Localization Tech Stack Evolution

 As the world moves increasingly to a “digital-first” approach across the business and government spectrum, it has become increasingly clear to any enterprise interested in reaching a larger digital population, that providing more multilingual content matters, and that the demand for more translated content will only grow, which also means that enterprise translation capabilities will need to be pervasive and scalable.

The challenge for globalization managers is further complicated by the increasing focus on customer experience (CX) which means that the content can vary greatly, in volume, velocity, and value to customers and internal stakeholders. All content does not need to go through traditional localization production and quality validation processes.

Content that is focused on understanding, communication, and listening does not require the same linguistic quality assurance, and in the digital space, user-generated and other external content is now often the most impactful content to consider.

Modern-era globalization managers need to understand what matters most to customers and balance their focus on the “mandatory” legally required content that localization has typically focused on, against the non-corporate content customers find most useful.

Though the value and business benefit of large-scale translation and localization are now well understood, globalization and localization managers tasked with making the global customer outreach happen, struggle with this objective.

They are faced with a fragmented, inconsistent, and fractured technology landscape and many sub-optimal tools currently exist in the language technology marketplace shown in the graphic below.

These tools are needed both to perform the many specific tasks involved in any globalization effort and to help establish structured processes that enable ongoing and emerging global customer-focused needs to be efficiently serviced.

Given the wide variety of tools and the diversity of the people needed to effectively execute the multiple globalization processes involved, straightforward and efficient data flow from sub-system to sub-system is desirable.

Successful globalization outcomes are often directly linked to enabling fast-flowing, unhindered data flows through a variety of translation-related processes. This is a necessary condition for success in a digital-first world.

However, what many Loc Buyers find is that some of the systems and tools they use impede and obstruct this smooth data flow, and thus the digital globalization initiative is often undermined and overly focused on repairing broken and problematic data flows.

If we look at the TMS part of the tech stack more closely, we can understand the challenge that globalization managers have when making long-term decisions on what their tech stack should look like. There are many choices, and identifying the specific characteristics of a superior system is not so clear.

We have learned that the best AI outcomes are driven by high-quality data above all else, and thus selecting technology that facilitates ongoing data access as technology changes, should be a prime concern and strategy for any forward-thinking globalization manager.

The critical technology components for most localization managers include the following categories:

1.      CAT Tools used by translators

2.      Translation Management Systems

3.      Language Quality Assurance (LQA) Tools

4.      Enterprise-capable MT

5.      Audiovisual Translation tools are growing in importance

In general, it can be said that the better the integration between these key components is, the more successful the globalization outcomes and the more efficient the enterprise will be in providing a high-quality global customer experience.

Unfortunately, the reality for many localization teams is focused on repairing broken or non-existent connections between sub-systems, and building better data sharing between the various components in their back-end localization tech stack to power the customer-pleasing expanded multilingual CX.


Avoiding Lock-In with Proprietary Systems

As the enterprise matures in localization and globalization sophistication, it will likely develop and build valuable linguistic assets over time. These linguistic assets need to be easily accessed and available to be shared with emerging new tools and platforms that provide business leverage, powered by new AI capabilities, as customer needs and CX imperatives dictate.

The tech stack complexity challenge for globalization managers is further exacerbated by the fact that much of the technology is still evolving.

Thus, any technology component that creates lock-in and prevents the straightforward transfer of linguistic assets to new superior language technology tools or platforms as they become available IS TO BE AVOIDED.

These siloed systems create what is called Tech Debt. This refers to the off-balance-sheet accumulation of all the technology work a company needs to do in the future. Tech debt results from software entropy and a lack of integration between different systems and data.

And it’s not just a minor inconvenience. A majority of businesses say that tech debt is slowing their pace of development, and resulting in real-world losses in sales and productivity.

Tech debt can produce several negative consequences for businesses:

The benefits of reducing tech debt are also significant and include:

Buyers should demand that enabling straightforward API access to client linguistic data without restriction or restraint should be a basic and critical requirement for any modern enterprise software solution or TMS.

Sophisticated new capabilities emerging from NLP research in Large Language Models, Responsive MT, and other emerging Language AI research will be almost useless to those companies that cannot quickly move relevant enterprise linguistic data to these new applications. They will be unable to properly explore possibilities of providing better CX with emerging linguistic AI capabilities.

This data lock-in is especially true for some of the current TMS systems that create multiple layers of technical, and even legal obstacles to straightforward data sharing. These obstacles invariably trap some Loc Buyers into sub-optimal workflows and solutions.

It is surprising that more Loc Buyers do not understand the importance of free and easy access to all linguistic data over the long-term, and suggests that Loc Buyers are extremely naïve in terms of making technology evaluations and selections that stand the test of time.

Sub-optimal initial choices will require regular overhauls in the technology stack to overcome obstacles created by proprietary lock-in technology.

Data is the lifeblood of any organization and the backbone that supports the creation of market-leading CX. The AI-driven world of tomorrow will be increasingly data-driven.

However, it is nearly impossible to make this data actionable in marketing activations and other business processes without data centrality and shareability. To avoid this, globalization managers should look for partners who can help them democratize their data sets so that they’re integrated and accessible by all.


The Growing Importance of Integration with Enterprise IT

Translation technology has reached an inflection point, as translation connects to the major trends affecting every industry: big data, cloud computing, and artificial intelligence (AI). Language platforms that can scale from millions to billions of words of translated content per month are being created as enterprise buyers and innovative language service providers seek to align their language systems with the technology stacks of globally focused enterprises.

Implementing many different systems and sources that don’t speak to each other will make it harder for the business to enact data-backed decisions and integrate with core IT functionality. Straightforward integration with core enterprise IT is a key requirement to enable successful global CX outcomes.

The ability to quickly import, clean, and use data from countless sources is critical to marketers and globalization managers, but rarely easy.

One of the barriers to realizing this data actionability stems from rigid data structures that can’t onboard, transport, and unify both structured and unstructured data from different sources. Flexible data architecture and a scalable hygiene framework can speed up the timeline for data activation and value creation.

The chart above shows the relationship between translation quality and content volume. It also shows that the highest returns on investments in translation technology will come from those areas focused on global CX and eCommerce.

The collaboration between localization teams and enterprise IT teams are growing in sophistication and now increasingly both internal corporate data and external data from social media and customer reviews are being mingled and merged to provide better CX.

This often requires handling large volumes of user-generated content (UGC) and monitoring social media brand impressions which are so voluminous that traditional localization workflows are not valid.

UGC is a dominant element of the eCommerce content landscape and even presents special challenges for MT technology. UGC content is often written by non-native speakers and, most likely, by non-professional content writers and thus needs specialized treatment and a different approach from typical localization content. But we see today that global market leaders learn to do this at scale, with new techniques that assume and drive evolutionary quality improvements.

Tech-savvy localization managers who understand this “start now and improve gradually approach“ on massive content volumes are now being seen as vital partners in global growth strategies. Best practices suggest that the most effective strategy is to have MT and Human translators working together to build a continuous improvement cycle.

The strategy to translate a billion new words across multiple use cases every month has to be different than a typical localization translate-edit-proof (TEP) process. This is made difficult or even impossible with TMS systems that do not allow easy access to ALL linguistic assets.

Airbnb is an example of emerging localization leadership where the localization team is seen as a vital partner in enabling global growth and works closely with IT, Legal, and Product teams to deliver better global customer experiences.

The Airbnb localization team oversees both typical localization content and user-generated content (UGC), across the organization, which means they oversee billions of words a month being translated across 60+ languages using a combined human plus continuously improving MT translation model. The localization team enables Airbnb to translate customer-related content across the organization at scale. High-value external content is often given the same attention as internally produced marketing content.

When dealing with CX-focused translation scenarios, the business requirements direct globalization managers to focus on optimizing the translation production mode to the volume, speed, quality requirements, and the value of the content to customers.

This is a clear shift away from the traditional LQA-focused localization workflows where TMS systems have traditionally been useful.

TMS systems have been most useful in relatively low-volume, complex workflows that involve multiple levels of human touch on the translated content. This is the top left-hand corner of the chart above. TMS systems add little to no value in scenarios with high volume fast flowing CX data where data flow straight from MT to dissemination.

Dated monolithic translation management systems (TMS) are giving way to micro-service and cloud-based architectures, with machine learning driving systems toward enterprise-scale automation where speed, scale, and the value of the content to the global customer matter more than achieving perfect linguistic quality.

Thus, increasingly we see that TMS systems are completely bypassed or irrelevant, and there is greater use of “raw MT” or carefully pre-tuned MT rather than fully post-edited MT.



The Emerging Requirements for a Language Platform

As more senior executives in the global enterprise ask questions like:

  • How do we integrate our international strategy with our overall corporate strategy?
  • What will this take in terms of people, process, and technology?

We should expect a shift to language as a feature at the platform level wherein language is designed, delivered, and optimized as a feature of a product and/or service from the beginning.

Language accessibility is integrated into content and procedural workflows that affect almost everyone within the organization at some point. Something that analysts call a "language platform."

Rebecca Ray of CSA describes the impact of producing relevant content for modern eCommerce marketplaces at scale and touches upon the key requirements of a Language Platform.

The success of globalization leaders like Airbnb demonstrates the value of developing comprehensive and collaborative capabilities with a more globally embedded and pervasive translation-focused ecosystem. The Airbnb deployment is a pioneering example in global CX best practice that shows how extensive and deep-reaching translation workflows can be integrated into corporate IT when the value is understood at executive levels.

A CX-focused and capable Language Platform that is ready for digital-first globalization and localization challenges would need all of the following key components working together in a highly integrated and seamless manner.

  • Essential TMS capabilities to monitor translation projects, and linguistic quality, and generate and manage critical translation workflows for different content types as needed.
  • An adaptive and continuously improving MT system that automates personalization and performance optimization for each enterprise customer and manages the collection of corrective feedback across dozens of enterprise use cases. This element is increasingly becoming the most important element of the back-end tech stack and the heart of the Translation Engine to provide superior global CX.
  • Computer-assisted translation (CAT) tools to enhance translator productivity, simplify project management, and share corporate linguistic assets like translation memories (TMs), glossaries, and terminology. In the modern era, CAT tools would also need to handle video, audio, and other social media-focused multimedia data.
  • Open and service-based architecture to allow continuing evolution of translation processes and addition of new core functions powered by machine learning with speed and agility. Linguistic assets are maintained in a continuously leverageable state so that these assets can be connected to emerging linguistic AI technology without hindrance or restraint.
  • Connectors: As the need for an Enterprise Translation Engine becomes more apparent the Language Platform will need to connect to CMS, Marketing Automation, Customer Data Platforms (CDP), CRM, ERP, Messaging, and Customer Support & CX platforms, in addition to leading social media to listen to and monitor customer conversations.

Look for a partner rather than a vendor, that can help you simplify and rationalize your tech stack and the increasing amounts of data you’re ingesting. Instead of logging in to different systems repeatedly according to content type and purpose, look for a vendor who can consolidate these into one simplified view. Additionally, ensure that data can be shared across vendors in this centralized hub so that you can leverage the power of these insights across the scope of your audience.


The Translated Tech Stack

TranslationOS is a hyper-scalable translation platform, that directly connects clients with translators that also provides management access to translation-related KPIs. It also provides the technical foundations to build a next-generation Language Platform.  It provides customizable dashboards that give globalization managers access to KPIs, project status, quality performance, and linguist profiles.

TranslationOS is a technology platform that allows straightforward access to client data whenever it is required for other downstream applications, or just for internal archival purposes. Client linguistic assets always remain within easy reach of the client's developers to support other valued added processes that can arise over time.

TranslationOS is also the overarching technology that tightly ties together enterprise translation memory, adaptive MT and corrective feedback management, CAT, and multimedia data management tools.

TranslationOS has a growing set of content connectors enabling external data ingestion and export to enterprise IT infrastructure.

TranslationOS includes an AI-driven translator matching tool (T-Rank) to ensure optimal selection from a qualified, and continuously verified pool of 400,000 translators for different projects using 30+ factors (e.g., availability, historical performance, subject matter experience, qualifications) to drive objective rankings to ensure the identification of the best-suited resources.

ModernMT is an adaptive MT system that is highly flexible, responsive, easy to manage and maintain, continuously learning, and able to incorporate ongoing human corrective feedback to ensure better MT output.

It consistently shows up as a top-performing MT system in independent third-party MT quality evaluations even before it has been adapted and tuned to specific enterprise content. It seamlessly integrates into TranslationOS and leading CAT tools like MateCat, Trados, and MemoQ.

In 2022 it has also been integrated with MateSub and MateDub to enable the automated translation of corporate multimedia content.

MateCat is a free, open-source, performance-oriented online CAT tool that allows translators to easily share TMs, and glossaries and interact dynamically with ModernMT to ensure continuously improving MT suggestions. It is integrated with MyMemory, a massive, yet clean TM gathered over 20 years, to augment and increase TM matching possibilities.

MateSub is a CAT tool optimized for subtitling tasks. It combines state-of-the-art AI (auto-spotting, auto-transcription, auto-translation) with a powerful and easy-to-use editor to let you create higher quality subtitles, dramatically faster.

MateDub is an AI-powered tool to assist in voice-over dubbing projects which can add digital voices synthesized from human voice-actor sampling. It allows users to dub videos by simply editing text.



As we move more deeply into the "digital-first" age, we will also move beyond the reach of traditional language technology like TM and TMS systems for more of our translation needs.

We are going to see much more focus and discussion on Language Platforms, Translation Layers, and Translation Operating Systems for fast-flowing CX-related data that are also built on much more open, new integrations-friendly, and transparent technology stacks.


Friday, March 11, 2022

The Evolving Relationship of MT with the Translator

 Machine translation is pervasive today and even the most conservative estimates say that MT is “translating” trillions of words a month across multiple large public MT portals and is used by hundreds of millions of internet users daily at virtually no cost.

As more of the global population comes online, people need MT to access the content that interests them even if only in a gist-sense, and today we see that there is growing momentum in the development and advancement of the state-of-the-art (SOTA) on “low-resource” (languages with limited or scarce data) languages to further accelerate global MT use.

MT technology has been around in some form for the last 70 years and unfortunately has a long history of over-promising and under-delivering. A history of eMpTy promises as it were. However, the more recent history of data-driven MT has been especially troubling for translators, as SMT and NMT pioneers have repeatedly claimed to have reached human parity.

These over-exuberant claims about the accomplishment of MT technology, have driven translator compensation down and have made many would-be translators reconsider their career choices.

It does not help that a more careful examination of the human parity claims by experts shows that these claims are not true, or perhaps only true for a tiny sample of test sentences.

Many say, that the market perception of exaggerated MT capabilities has damaged translator livelihood and there is often great frustration by many who use MT in production environments where the high-quality human equivalent translation is expected but never delivered, without significant additional effort and expense.

To add insult to injury, the overly optimistic MT performance claims have also resulted in many technology-incompetent LSPs attempting to use MT to reduce costs by forcing translators to post-edit low-quality MT output at low rates.

It does not seem to matter that most LSPs have yet to properly learn to use MT in localization production work, according to a survey of MT use by LSPs done by Common Sense Advisory last year.

It is also very telling that the author wrote a blog post on MT post-editing compensation in March 2012 that has had the widest readership of any post he has written ever, and continues even in 2022 to be an actively read post!

Thus, often "monolithic MT" is considered a dark, unuseful,  and unwelcome factor in the lives of translators. However, this state of affairs is often a result of incompetent and unethical use of the technology rather than a core technology characteristic.


The Content and Demand Explosion

However, the news on MT is not all doom and gloom from the translator's perspective. There is a huge demand for language translation as evidenced by the volume of use of public MT, and by the digital transformation imperatives for global enterprises driving the need for better professional MT.

Both public MT and enterprise MT are building momentum. The demand for content from across the globe is exponential which means that translation volumes will also likely explode. And, while much of it can be handled with carefully optimized Enterprise MT, it will also need an ever-growing pool of tech-savvy translators to drive continuously improving MT technology.

World Bank estimates say that by 2022, yearly total internet traffic is projected to increase by about 50 percent from 2020 levels, reaching 4.8 zettabytes, equal to 150,000 GB per second. The growth in global internet traffic is as dazzling as the volume. Personal data are expected to represent a significant share of the total volume of data being transferred cross-border.


It is estimated that the amount of digital data created over the next five years will be more than twice the amount created since the advent of digital storage. Global data creation and replication will experience a compound annual growth of 23 percent in the 2020–2025 forecast (IDC, 2021a). Data traffic trends are related to economic development, value creation, and prosperity.

The sheer volume and explosion in content volumes driven by these trends are already creating an increasing awareness of the supply shortage of translators. The furor around the poor quality of the translation of the Korean hit show “Squid Games” is a telling example of this changing scene.

LSPs and translators are critical to the distribution of that local content on a global scale. But because of a labor shortage and no viable automated solution, the translation industry is being pushed to its limits.

“I can tell you literally, this industry will be out of supply over demand for the upcoming two to three years,” David Lee, the CEO of Iyuno-SDI, one of the industry’s largest subtitling and dubbing providers, said recently. “Nobody to translate, nobody to dub, nobody to mix –– the industry just doesn’t have enough resources to do it.” Interviews with industry leaders reveal most streaming platforms are now at an inflection point, left to decide how much they are willing to sacrifice on quality to subtitle their streaming roster.

So while it is true that as we enter 2022 most LSPs have yet to learn how to use MT efficiently for production use, and that translator compensation at the word level has been decreasing over the last five years, there are also positive changes.

The Translated Srl experience with ModernMT shows that it is possible to use MT effectively for production localization work as Translated uses MT in 95% of their production workload, mainly because the technology is flexible, easy to set up, highly responsive, and agile enough to handle the variations typical in production work.

This is the result of superior architecture, better process integration, and sensitivity to human factors, refined over decades, to ensure sustainable and increasing productivity improvements.


The Translated Srl experience is also direct proof that MT can be a valuable assistive technology tool for serious, i.e. professional human translation work.

The ModernMT technology is perhaps the only MT technology optimized for production localization work and is already in the process of being extended to work with video content (MateDub & MateSub). Video adds time synchronization challenges to the basic translation tasks.


The Importance of the Human-In-The-Loop

The exploding content and enterprise CX demands to provide more relevant content to their customers also suggests that there is a potential for rates to rise as more enterprises begin to understand that improving translation quality has to be linked to an increased role of humans-in-the-loop to make MT perform better on the specific content that matters to the enterprise.

As we consider the possibility of MT achieving human parity on language translation at production scale we need to remind ourselves of the following. Language is the cornerstone of human intelligence.

The emergence of language was the most important intellectual development in our species’ history. It is what separates us from all other species on the planet. It is through language that we formulate thoughts and communicate them to one another. Language enables us to reason abstractly, to develop complex ideas about what the world is and could be, and to build on these ideas across generations and geographies. Almost nothing in modern civilization would be possible without language.

Building machines that can “understand” language has thus been a central goal of the field of artificial intelligence dating back to its earliest days, but this has proven to be maddeningly elusive. The current state of MT is the result of 70 years of effort, and having a machine master language may either be impossible or simply much farther out in the future than the ML-focused singularity-is-nigh fanboys can envision.

This is because mastering language is what is known as an “AI-complete” problem: that is, an AI that can understand language the way a human can, would by implication be capable of any other human-level intellectual activity. Put simply, to solve the language challenge is to create human-equivalent machine intelligence.

Competent linguistic feedback is needed to improve the state of MT technology, and humans are needed to improve the quality of MT output for enterprise use.

We see today that machine translation is ubiquitous, and by many estimates is responsible for 99.5% or more of all language translation done on the planet on any given day. But we also see that MT is used mostly to translate material that is voluminous, short-lived, transitory and that would never get translated if the machine were not available to help.

Trillions of words are being translated by MT weekly, yet when it matters, there is always human oversight on translations that may have a high impact, or when there is great potential risk or liability from mistranslation.

While machine learning use-cases continue to expand dramatically, there is also an increasing awareness that a human-in-the-loop is necessary since the machine lacks comprehension, cognition, and common sense, all elements that constitute “understanding”.

As Rodney Brooks, the co-founder of iRobot said in a post entitled - An Inconvenient Truth About AI: "Just about every successful deployment of AI has either one of two expedients: It has a person somewhere in the loop, or the cost of failure, should the system blunder, is very low."

As the use of machine learning proliferates, there is an increasing awareness that humans working together with machines in an active learning contribution mode can often outperform the possibilities of machines or humans alone.

Many of the public generic MT engines already have billions of sentence pairs that underlie and “train” the model. Yet, we see an increasing acknowledgment from the AI community that language is indeed a hard problem. One that cannot necessarily be solved by using more data and algorithms alone, and a growing awareness that other strategies will need to be employed.

This does not mean that these systems cannot be useful, but we are beginning to understand that while language AI tools are useful, they have to be used with care and human oversight, at least until machines have more robust comprehension and common sense.

Effective human-in-the-loop (HITL) implementations allow the machine to capture an increasing amount of highly relevant knowledge and enhance the core application as ModernMT does with MT.

Another way to look at this is to see the Language AI or MT model as a prediction system, rather than as a representative model of a human translator.

Very simply put, we are using information that we do have to generate information that we don’t have.

MT models are built primarily with translation memory (a.k.a training data) and are most successful with material that is most similar to this training data. MT models take the new source material and produce a prediction of this material into a target language based on what it knows from what it has been explicitly trained with.

With deep learning, pattern detection and prediction have gotten more sophisticated, but we are still, quite some distance from actual understanding, comprehension, and cognition.

A human translation cognitive flow within the brain of a competent human translator has significantly more sophisticated capabilities around the many translation-related sub-tasks that require and involve actual intelligence, gathered from multisensorial life experience and common sense.

Human translators understand the relevant document, historical, and situational context even though it may not be explicitly stated. They identify semantic intent, and add cultural context into the translation, reading between the lines to ensure overall accuracy, guided by common sense, on what may not be stated but can be “understood” from life experience, insight, and deep comprehension.

This is in stark contrast to just performing the literal conversion of word strings and patterns from the source language to a target language that MT systems are limited to. Systems trained on billions of example "training" sentences have yet to capture what humans do. More data is not enough.

To restate, it is more accurate to see the MT model as a prediction system rather than an understanding system. Much of the recent success with AI and machine learning is a result of converting problems that were not historically prediction problems into prediction problems e.g. self-driving cars, fraud detection, and automated email replies.

MT systems are most useful when they produce a large number of useful predictions, even if these are not "perfect". It is as useful for a translator as TM, maybe even more so, when MT is responsive, continuously learning, and a true assistant.

The overview of the development and deployment of the prediction model can be seen in this generic graphic overview which is true for MT and many other ML use cases.


Once a model has been deployed ongoing improvement in its prediction ability can be driven by more data, better learning algorithms, more computing power, and ongoing corrective feedback that becomes increasingly important as an ML model evolves in competence and performance.

After 70 years of MT research, it is increasingly clear that the efficient incorporation of human corrective feedback is one of the fastest and most useful ways available to improve an MT system's performance.

The following chart shows what happens at the monitor stage where human judgment and active corrective feedback on model outputs begin to drive improvements on the specific material in focus. The best systems will take feedback and process, learn, update, and incorporate new learning quickly to improve the predictions of the model in real-time.


The speed and ease with which new learning can be incorporated into an MT system are critical determinants of the value of the MT system to an individual translator. There is great value for all stakeholders in improving the predictive capabilities of an MT system.


ModernMT: An MT system designed for the translator

The modern era translator work experience often involves the use of translation memory (TM). Since it improves translator productivity when the TM is related and relevant to any new translation work that a translator may undertake.

MT is used less often by professional translators in general because of the following reasons:

  • Generic MT output is of limited value.
  • Most MT systems have a very limited ability to customize and adapt the generic system to the translator's area of focus and specialization.
  • The typically complex customization process often requires that translators have skills that are typically outside of the scope of translator education.
  • A large volume of data (more than most translators can summon) is needed to have any impact on generic engine performance. This also makes it difficult for most LSPs to also customize an MT engine as most of the MT models in the market require tens of thousands or more segments of training data to have an impact.
  • The very slow rate of improvement of most MT engines means that translators must correct the same errors over and over again. The whole improvement process can itself be a significant engineering undertaking and task.
  • The open admission of MT use is often penalized with lower compensation and lower word rates.
  • The inability to control and improve MT output predictably means that translators themselves have a higher level of uncertainty about the utility of MT given project deadlines and thus fallback to traditional approaches.

For MT to be useful to a translator it needs the following attributes:

  • Tight integration with CAT tools that are the primary work environment for translators.
  • Easy to start using without geeky technical preparation and ML-customization-related work.
  • Rapid learning of new material and incorporation of any corrective feedback so that the MT system is continuously improving, by the day or even the hour.
  • The ability to handle project-related terminology with ease.
  • Keep translator data private and secure.
ModernMT is an MT system that is designed to adapt to the unique needs and focus of an individual translator in essentially the same way that TM does. In many ways, it is a next-generation TM technology that has predictive capabilities.

ModernMT is a translator-focused  MT architecture that has been built and refined over a decade with active feedback and learning from a close collaboration between translators and MT researchers.

ModernMT has been used intensively in all the production translation work done by Translated Srl for over 15 years and was a functioning human-in-the-loop (HITL) machine learning system before the term was even coined.

ModernMT is perhaps the only MT system that was designed by translators for translators rather than by pure technologists working in isolation with data and algorithms.

This long-term engagement with translators and continuous feedback-driven improvement process also results in creating a superior training data set over the years. This superior training data enables users to have an efficiency and quality advantage that is not easily or rapidly replicated.

This is also the reason why ModernMT does so consistently well in third-party MT system comparisons, even though evaluators do not always measure its performance optimally. ModernMT simply has more informed translator feedback built into the system.

The following is a summary of features in a well-designed Human-in-the-loop (HITL) system, such as the one underlying ModernMT:

  • Easy setup and startup process for any and every new adapted MT system that allows even a single translator to build hundreds of domain-focused systems.
  • Responsive: Active and continuous corrective feedback is rapidly processed so that translators can see the impact of corrections in real-time and the system improves continuously without requiring the translator to set up a data collection and re-training workflow.
  • An MT system that is continuously training and improving with this feedback (by the minute, day, week, month). Small volumes of correction can improve the ongoing MT performance.
  • Tightly integrated into the foundational CAT tools used by translators who provide the most valuable system-enhancing feedback.
  • Different engagement and interaction with MT than a typical PEMT experience. 

I recently interviewed several translators who are active ModernMT users and have summarized their comments (+ve and -ve) below. Their comments contain pearls of wisdom and anecdotal experience that may be useful to other translators who are still considering MT.

Subject focus by those who shared their usage patterns with me included accounting/finance, legal contracts, complex engineering equipment-related content, marketing content, product manuals, newsletters & press releases, medical information for patients, and even Buddhism & meditation-related content. Many simply provided categories like Law, Medical, Technical.

The extent of use: Used in the large majority of work they did, except for DTP or very specialized domain content that they did on an infrequent basis. Many said that the real benefits start to accrue after one builds up some TM and that over time ModernMT learns to support your primary workload.

How is MT engaged: CAT Tools (Trados), ModernMT GUI, and MateCat

Why: Work volumes and turnaround requirements and high-level data privacy and availability of TM to enable adaptation.

Competitive systems evaluated: Google, DeepL, Systran, Kantan

“I have used DeepL and Google, which can be very useful, although I still find ModernMT to have better overall accuracy compared to both of them. DeepL is a good alternative for comparing output, although it is much less consistent compared to ModernMT when working on large documents e.g. consistency of terminology etc.”

“I can tell you this with peace in my mind that nothing can replace ModernMT. ModernMT has magic that no one can describe. It really adapts to contexts and stores my previous translations and yields me 99% accurate translations.

Improvements needed: Word case handling for acronyms and abbreviations, handling of short phrases and titles, the lack of persistence of terms across documents, better format preservation, better dashboard.

Desirable New features: Glossary and terminology handling, a dashboard on data and usage, more robust punctuation handling, real-time predictive capabilities, pre-translation quality assessment.

”I consider MT as a development tool, making our job easier, but not a tool that gives the final product. It is like an advanced medical tool used by a surgeon during surgery, which helps the surgeon to make fewer mistakes, to save time, and to save the life of the patient.”

A strong positive comment by a translator who also provided constructive areas of improvement content: “I have noticed incredible improvement [in the MT quality] as if it is my roommate who was trying to get to know me and my translation style and way of constructing the sentences.”

Many were surprised to find out that glossary and terminology terms are best introduced to ModernMT in sentence form rather than as short phrases as the context and variants shown in sentence-context ensures a faster pick-up and learning.

Several expressed surprise that more translators did not realize cost/benefit and productivity advantages to be gained by using a responsive MT system like ModernMT and also mentioned that success with ModernMT required investment in one or all of the following: time, corrective feedback, and personal TM but can yield surprisingly good results in as little as a few weeks.


To close this post I include a podcast done with ProZ last year, that I got very positive feedback on, from many translators.

Conversation with Paul Urwin of Proz on MT

Paul talks with machine translation expert Kirti Vashee about interactive-adaptive MT, linguistic assets, freelance positioning, how to add value in explosive content situations, e-commerce translation, and the Starship Enterprise.


Paul continues the fascinating discussion with Kirti on machine translation. In this episode, they talk about how much better MT can get, which languages it works well for, data, content, pivot languages, and machine interpreting.

Thursday, November 11, 2021

The Challenge of Using MT in Localization

We live in an era where MT is translating more than 99% of all the translation being done on the planet on any given day.

However, the adoption of MT by the enterprise is still nascent and still building momentum. Business enterprises have been slower to adopt MT even though national security and global surveillance-focused government agencies have used MT heavily. This adoption delay has mostly been because MT has to be adapted and tuned to perform better with very specific language used in specialized enterprise content.

Early enterprise adoption of MT was focused on eCommerce and customer support use-cases (IT, Auto, Aerospace) where huge volumes of technical support content made it a necessity to use MT technology to allow any possibility of translating the voluminous content in a timely and cost-effective manner to improve the global customer experience.

Microsoft was a pioneer who translated its widely used technical knowledge base to support an increasingly global customer base. The positive customer feedback for doing this has led to many other large IT and consumer electronics firms doing the same.

The adaptation of the MT system to perform better on enterprise content is a critical requirement in producing successful outcomes. In most of these early use-cases we see that MT is used to manage translation challenges when the content volumes were huge, i.e., millions of words a day or week. These were “either use MT or provide nothing” knowledge-sharing scenarios.

These enterprise-optimized MT systems have to adapt to the special terminology and linguistic style of the content they translate, and this customization has been a key element of success with any enterprise use of MT.

eBay was an early MT adopter in eCommerce and has stated often that MT is key in promoting cross-border trade. It was understood that “Machine translation can connect global customers, enabling on-demand translation of messages and other communications between sellers and buyers, and helps them solve problems and have the best possible experiences on eBay.”

A study by an MIT economist showed that after eBay improved its automatic translation program in 2014, commerce shot up by 10.9 percent among pairs of countries where people could use the new system.

Today we see that MT is a critical element of the global strategy for Alibaba, Amazon, eBay, and many other eCommerce giants.

Even in the COVID-ravaged travel market segment, MT is critical as we see with Airbnb, which now translates billions of words a month to enhance the international customer experience on their platform. In November 2021 Airbnb announced a major update to the translation capabilities of their platform in response to rapidly growing cross-border bookings and increasingly varied WFH activity.

“The real challenge of global strategy isn’t how big you can get, but how small you can get.”
Dennis Goedegebuure, former head of Global SEO at Airbnb.

However, MT use for localization use cases has trailed far behind these leading-edge examples, and even in 2021, we find that the adoption and active use of MT by Language Service Providers (LSPs) is still low. Much of the reason lies in the fact that LSPs work on hundreds or thousands of small projects rather than a few very large ones.

Early MT adopters tend to focus on large-volume projects to justify the investments needed to build adapted systems capable of handling the high-volume translation challenge.

What options may be available to increase adoption in the localization and professional business translation sectors?

At the MT Summit conference in August 2021, CSA's Arle Lommel shared survey data on MT use in the localization sector in his keynote presentation. He noted that while there has been an ongoing increase in adoption by LSPs there is considerable room to grow.

Arle specifically pointed out that a large number of LSPs who currently have MT capacity only use it for less than 15% of their customer workload and, “our survey reveals that LSPs, in general, process less than one-quarter of their [total] volume with MT.”

The CSA survey polled a cross-section of 170 LSPs (from their "Ranked 191" set of largest global LSPs) on their MT use and MT-related challenges. The quality of the sample is high and thus these findings are compelling.

The graphic below highlights the survey findings.

CSA Survey of MT Use at LSPs


When they probed further into the reasons behind the relatively low use of MT in the LSP sector they discovered the following:

  • 72% of LSPs report difficulty in meeting quality expectations with MT
  • 62% of LSPs struggle with estimating effort and cost with MT

Both of these causes point to the difficulty that most LSPs face with the predictability of outcomes with an MT project.

Arle reported that in addition to LSPs, many enterprises also struggle with meeting quality expectations and are often under pressure to use MT in inappropriate situations or face unrealistic ROI expectations from management. Thus, CSA concluded that while current-generation MT does well relative to historical practice, it does not (yet) consistently meet stakeholder requirements.

This apparent market reality validated by this representative sample is in stark contrast to what happens at Translated Srl, where 95% of all projects and all client work use MT (ModernMT) since it is a proven way to expedite and accelerate translation productivity.

Adaptive, continuously learning ModernMT has been proven to work effectively over thousands of projects with tens of thousands of translators.

This ability to properly use MT in an effective and efficient assistive role in production translation work has resulted in Translated being one of the most efficient LSPs in the industry, with the highest revenue per employee and high margins.



Another example of the typical LSP experience: a recent study by Charles University done with only 30 translators using 13 engines (EN>CS) concludes: "the previously assumed link between MT quality and post-editing time is weak and not straightforward." They also found that these translators had “a clear preference for using even imprecise TM matches (85–94%) over MT output."

This is hardly surprising, as getting MT to work effectively in production scenarios requires more than choosing the system with the best BLEU score.

Understanding The Localization Use Case For MT

Why is MT so difficult for LSPs to deploy in a consistently effective and efficient manner?

There are at least four primary reasons:

  1. The localization use case requires the highest quality MT output to drive productivity which is only possible with specialized expertise and effort,
  2. Most LSPs work on hundreds/thousands of smallish projects (relative to MT scale) that can vary greatly in scope and focus,
  3. Effective MT adaptation is complex,
  4. MT system development skills are not typically found in an LSP team.

MT Output Expectations

As the CSA survey showed, getting MT to consistently produce output quality to enable use in production work is difficult. While using generic MT is quite straightforward, most LSPs have discovered that rapidly adapting and optimizing MT for production use is extremely difficult.

It is a matter of both MT system development competence and workflow/process efficiency. 

Many LSPs feel that success requires the development of multiple engines for multiple domains for each client, which is challenging since they don't have a clear sense of the effort and cost needed to achieve positive ROI.

If you don’t know how good your MT output will be, how do you plan for staffing PEMT work and calculate PEMT costs?

Thus, we see MT is only used when very large volumes of content are focused around a single subject domain or when a client demands it.

A corollary to this is that it requires deep expertise and understanding of NMT models to acquire the skills and data needed to raise MT output to useful high-quality levels consistently.

Project Variety & Focus

Most LSPs handle a large and varied range of projects that cover many subject domains, content types, and user groups on an ongoing basis. The translation industry has evolved around a Translate>Edit>Proof (TEP) model that has multiple tiers of human interaction and evaluation in a workflow.

Most LSPs struggle to adapt this historical people-intensive approach to an effective PEMT model which requires a deeper understanding of the interactions between data, process, and technology.

The biggest roadblock I have seen is that many LSPs get entangled in opaque linguistic quality assessment and estimation exercises, and completely miss the business value implications created by making more content multilingual. Localization is only one of several use-cases where translation can add value to the global enterprise's mission.

Typically, there is not enough revenue concentration around individual client subject domains, thus, it is difficult for LSPs to invest in building MT systems that would quickly add productivity to client projects. 

MT development is considered a long-term investment that can take years to yield consistently positive returns.

This perceived requirement for the development of multiple engines for many domains for each client requires an investment that cannot be justified with short-term revenue potential. MT projects, in general, need a higher level of comfort with outcome uncertainty, and, handling hundreds of MT projects concurrently to service the business is too demanding a requirement for most LSPs.

MT is Complex

Many LSPs have dabbled with open-source MT (Moses, OpenNMT) or AutoML and Microsoft Translator Hub only to find that everything from data preparation to model tuning, and quality measurement is complicated, and requires deep expertise that is uncommon in the language industry.

While it is not difficult to get a rudimentary MT model built, it is a very different matter to produce an MT engine that consistently works in production use. For most LSPs, open-source and DIY MT is the path to a failed project graveyard.

Neural MT technology evolution is happening at a significantly faster pace than Statistical MT. To stay abreast with the state-of-the-art (SOTA) requires a serious commitment, both in manpower and computing resources.

LSPs are familiar with translation memory technology that has barely changed in 25 years, but MT has changed dramatically over the same period. In recent years the neural network-based revolution has driven multiple open-source platforms to the forefront and keeping abreast with the change is difficult.

NMT requires expertise not only around "big data", NMT algorithms, and open-source platform alternatives but also around understanding parallel processing hardware.

Today AI and Machine Learning (ML) are synonymous, and engineers with ML expertise are in high demand.

MT requires long-term commitment and investment before consistent positive ROI is available and few LSPs have an appetite for such investments.

Some say that an MT development team might be ready for prime-time production work only after they have built a thousand engines and have this experience to draw from. This competence-building experience seems to be a requirement for sustainable success.

Talent Shortage

Even if LSP executives are willing to make these strategic long-term investments, finding the right people has gotten increasingly harder. According to a recent survey by Gartner, executives see the talent shortage not just as a major hurdle to progressing organizational goals and business objectives, but it is also preventing many companies from adopting emerging technologies.

The Gartner research, which is built on a peer-based view of the adoption plans of 111 emerging technologies from 437 IT global organizations over a 12- to 24-month time period, shows that talent shortage is the most significant adoption barrier to 64% of emerging technologies, compared with just 4% in 2020.

IT executives cited talent availability as the main adoption risk factor for the majority of IT automation technologies (75%) and nearly half of digital workplace technologies (41%).

But using technology early and effectively creates a competitive advantage. Bain estimates that “born-tech” companies have captured 54% of the total market growth since 2015. “Born-tech” companies are those with a tech-led strategy. Think Tesla in automobiles, Netflix in media, and Amazon in retail.

Technology has emerged as the primary disruptor and value creator across all sectors. The demand for data scientists and machine learning engineers is at an all-time high.

LSPs need to compete with the global 2000 enterprises who offer more money and resources to the same scarce talent. Thus, we even see technical talent migrating out of translation services to the “mainstream” industry.

There is a gold rush happening around well-funded ML-driven startups and enterprise AI initiatives. ML skills are being seen as critical to the next major evolution in value creation in the overall economy as the chart below shows.

This perception is driving huge demand for data scientists, ML engineers, and computational linguists who are all necessary to build momentum and produce successful AI project outcomes. The talent shortage will only get worse as more people realize that deep learning technology is fueling most of the value growth across the global economy.


Thus, it appears that MT is likely to remain an insurmountable challenge for most LSPs. The option for an LSP to start building robust state-of-the-art MT capabilities in 2021 is increasingly unlikely. 

Even the largest LSPs today have to use “best-of-breed” public systems rather than build internal MT competence. Strategies employed to do this typically depend on selecting MT systems based on BLEU, hLepor, TER, Edit Distance, or some other score-of-the-day, which again explains why there is <15% MT-in-production-use.

As CSA has discovered, LSP MT use has been largely unsuccessful because a good Edit Distance/hLepor/Comet score does not necessarily translate to responsiveness, ease of use, adaptability of the MT system to the production localization use-case needs.

For MT to be useable on 95%+ of the production translation work done by an LSP, it needs to be reliable, flexible, manageable, rapidly adaptive, and continuously learning. MT needs to produce predictably useful output and be truly assistive technology for it to work in localization production work.

The contrast of the MT experience at Translated Srl is striking. ModernMT was designed from the outset to be useful to translators and created to collect the right kind of data needed to rapidly improve and assist in localization project-focused systems.

ModernMT is a blend of the right data, deep expertise in both localization processes and machine learning, and a respectful and collaborative relationship between translators and MT technologists. It is more than just an adaptive MT engine.

Translated has been able to overcome all of the challenges listed above using ModernMT, which today is possibly the only viable MT technology solution that is optimized for the core localization-focused business of LSPs.

ModernMT is the creation of an MT system optimized for LSP use. It could be used quickly and successfully by any LSP as there is no startup setup and training needed, it is a simple "load TM and immediately use" model.

ModernMT Overview & Suitability for Localization

ModernMT is an MT system that is responsive, adaptable, and manageable in the typical localization production work scenario. It is an MT system architecture that is optimized for the most demanding MT use-case: localization. And it is thus able to handle many other use-cases which may have more volume but are less demanding on the output quality requirements.

ModernMT is a context-aware, incremental, and responsive general-purpose MT technology that is price competitive to the big MT portals (Google, Microsoft, Amazon) and is uniquely optimized for LSPs and any translation service provider, including individual translators.

It can be kept completely secure and private for those willing to make the hardware investments for an on-premise installation. It is also possible to develop a secure and private cloud instance for those who wish to avoid making hardware investments.

ModernMT overcomes technology barriers that hinder the wider adoption of currently available MT software by enterprise users and language service providers:

  • ModernMT is a ready-to-run application that does not require any initial training phase. It incorporates user-supplied resources immediately without needing upfront model training.
  • ModernMT learns continuously and instantly from user feedback and corrections made to MT output as production work is being done. It produces output that improves by the day and even the hour in active-use scenarios.
  • ModernMT is context-sensitive.
  • The ModernMT system manages context automatically and does not require building domain-specific systems.
  • ModernMT is easy to use and rapidly scales across varying domains, data, and user scenarios.
  • ModernMT has a data collection infrastructure that accelerates the process of filling the data gap between large web companies and the machine translation industry.
  • Driven easily by the source sentence to be translated and optionally small amounts of contextual text or translation memory.

ModernMT’s goal is to deliver the quality of multiple custom engines by adapting to the provided context on the fly. This fluidity makes it much easier to manage on an ongoing basis as only a single engine is needed.

The translation process in ModernMT is quite different from common, non-adapting MT technologies. The models created with this tool do not merge all the parallel data into a single indistinguishable heap; separate containers for each data source are created instead and this is how it maintains the ability to adapt to hundreds of different contextual use scenarios instantly.

ModernMT consistently outperforms the big portals in MT quality comparisons done by independent third-party researchers, even on the static baseline versions of their systems.

ModernMT systems can easily outperform competitive systems once adaptation begins, and active corrective feedback immediately generates quality-improving momentum.

The following charts show how ModernMT is a consistent superior performer even as the quality measurement metrics change over multiple independent third-party evaluations conducted over the last three years. 

None of these metrics capture the ongoing and continuous improvements in output quality that is the daily experience of translators who work with dynamically improving ModernMT at Translated Srl.

Independent evaluations confirm ModernMT quality improves faster with COVID data set on English > German in the chart below.

ModernMT was also the "top performer" on several other languages tested with COVID data.

In Q4 2021 the COMET metric is widely being considered a "better" score because it is more aligned with human assessments and also incorporates semantic similarity, and again ModernMT shines.


 
If the predictions about the transformative impact of the deep learning-driven revolution are true, DL will likely disrupt many industries including the translation industry. MT is a prime example of an opportunity lost by almost all the Top 20 LSPs.

While it is challenging to get MT working consistently in localization scenarios, ModernMT and Translated show that it is possible and that there are significant benefits when you do.

This success also shows that when you get MT properly working in professional translation work, you create competitive advantages that provide long-term business leverage.  The future of business translation increasingly demands collaborative working models with human services integrated with responsive adapted MT. The future for LSPs that do not learn to use MT effectively will not be rosy.

A detailed overview of ModernMT is provided here. It is easy to test it against other competitive MT alternatives, as the rapid adaptation capabilities can be easily seen by working with MateCat/Trados or with a supported TMS product (MemoQ) if that is preferred.

ModernMT is an example of an MT system that can work for both the LSP and the Translator. The ease of the "instant start experience" with Matecat + ModernMT is striking when compared to the typical plodding, laborious MT customization process we see elsewhere today. Try it and see.

Tuesday, February 16, 2021

Building Equity In The Translation Workflow With Blockchain

This is a guest post by Bob Kuhns on the subject of blockchain use in the translation industry. He presents a very "simple" model where he shows how a blockchain could enable an ongoing,  robust, and trusted Buyer to Translator business connection that could quite possibly reduce the role of LSP middlemen whose primary value-add in business translation work today is project management services and coordination. Though this is a valuable service, it often significantly increases the cost of translation, and also sometimes creates discord, disgruntlement, and enmity amongst the freelance translation-service suppliers who ultimately do the work. 

A blockchain solution is not just about technology, it’s about solving business problems that have been insolvable before due to the inability of the ecosystem to share information in a transparent, immutable, and trusted manner.  

 LSPs continue to struggle in their communications on translation quality and the value of ongoing project services and thus it theoretically seems that a blockchain that did in fact deliver direct-to-buyer translation services that are trusted, reliable and predictable would indeed be a great step forward for the business translation industry. This is also exactly the primary reason why some in the LSP sector would not want blockchain to succeed. However, there is also a role for a more enlightened LSP in a functioning translation blockchain, one committed to transparency, equitable sharing of business value and benefits, and ultimate customer success in all their globalization initiatives.    

Blockchains show potential to address key concerns of our digitally-driven lives, such as a lack of transparency, accountability, verifiable identity, and control of dataBlockchain has the potential to enable defined quality to be delivered at a defined price in a defined timeframe with minimal administration overhead. The extent that it is able to do this in a clean and trusted way will likely drive its adoption. Blockchains are poised to catalyze new business models by cutting the costs of verifying the truth.  

Explanations of blockchains tend to get complicated quickly. However, in very basic language, blockchains help us certify that something is true, without someone in the middle doing checks and balances.  But recent translation market activity perhaps points to some of the vital building blocks to making blockchain more real in the translation industry.     

What might some of the elements to make a functional blockchain in the translation industry possible be? 

  • A robust TMS platform that could enable ANY buyer (not just localization buyers) to engage ANY translator to perform a necessary translation task.
  • Assistive translation technology that can be easily connected into a blockchain workflow (MT, TM, NLP Tools)
  • The ability to create self-sovereign data that would allow more equitable sharing of data. In the vision of many pioneers in the blockchain space, the “ownership” of data would switch from the organization that gathers it to the individual or organization that contributed it.
  • An Independent Translator rating, certification, ranking, identification database to enable competent resources to be identified and selected. 
  • Smart Contracts
  • And....

For those who think blockchain is still a distant dream, there is evidence that it is making meaningful headway in improving efficiency, accuracy, and transparency in some areas that have historically been project-management nightmares. The members of Tradelens, a blockchain joint venture between IBM and the shipper Maersk, control more than 60% of the world’s containership capacity. Seriously, do we really believe that a translation task has more variance and unplanned changes than these goods and trade flows? Watch the short movie clip at the link above to see it at work. Having trusted food quality from all over the world is perhaps an even more challenging scenario. Food Trust is a blockchain ecosystem that covers more than 100 organizations, including Carrefour and the top four grocery retailers in the U.S.     

It is quite likely that the old guard (executives, managers, and localization teams) will not be at the forefront of translation blockchain if it ever does become a reality. Mostly because they are "just too old, too tired, and too blind " as the movie says. Change is most often driven by the young who see the new potential and have the motivation in solving old enduring problems. 

"Disruption could also be spurred by an even younger generation. New York Times writer David Brooks traveled to college campuses to understand how students see the world. In a story, he wrote after the experience, starkly titled “A Generation Emerging from the Wreckage,” Brooks describes a cohort with diminished expectations. Their lived experience includes the Iraq war, the financial crisis, police brutality, political fragmentation, and the advent of fake news as a social force. In short, an entire series of important moments in which “big institutions failed to provide basic security, competence, and accountability.” To this cohort, in particular, blockchains’ promise of decentralization, with its built-in ability to ensure trust, is tantalizing. To circumvent and disintermediate institutions that have failed them is a ray of hope—as is establishing trust, accountability, and veracity through technology, or even the potential to forge new connections across fragmented societies.

This latent demand is well aligned with the promise of blockchains. While it’s a long road to maturity, these social forces provide a receptive environment, primed and ready for the moment entrepreneurs strike the right formula. "             

Alison McCauley, author of Unblocked 




===============


News recently has shed light on unfair work arrangements for freelancers or “gig workers” with Uber being a prominent example, and now more industries have begun to examine their relationships with freelancers. Though on a modest scale, this self-examination has come to the translation industry as well. The proposed translation model is grounded on the idea that blockchain can bring equity to translators while streamlining the translation workflow.

The Realization For Change

Even before the pandemic, the translation industry was changing and self-reflecting. NMT took center stage and the less-than-equitable working relationships for translators gained notice [1]. In “The State of the Linguist Supply Chain,” Common Sense Advisory examined the translation supply chain from the perspective of over 7,000 linguists, 75% of whom are freelancers representing 178 countries and 155 language pairs [2]. This data-rich survey identified many discrepancies between translators/linguists and their clients - the Buyers of translations and LSPs.

Several illustrative takeaways are:

  • Over half (54%) of the respondents could not live solely on their translation income.
  • Linguists are attracted to the translation profession because of flexible hours (91%) and the diversity of projects (75%). Only 33% rated their pay as being a plus.
  • The frustrations of linguists include fluctuating income (65%), irregularity of work (57%), and lack of respect (25%).
  • There is a preference for working for clients (65%) because translators (80%) earn more, have more job flexibility (56%), and quicker payments. Clients (76%) pay in less than 30 days compared with only 32% with LSPs.
  • The largest benefit of working for LSPs is more work (71%).
  • Translators perceived that cost (40%) and speed (33%) are more important than

Some of the largest challenges facing translators are finding clients (55%), negotiating prices (50%), dealing with tight deadlines (35%). Just before the pandemic, linguists felt the market changing to lower prices (64%) and faster turnaround times (56%).

Though not included here, linguists’ comments found in the report provide a more human context of their jobs than the raw percentages.


The Standard Model and Its Beneficiaries

The current translation workflow is one where translation Buyers hire LSPs to provide translators and manage the translation process including day-to-day project management, translation reviews, source-target file transfers, and handling of invoicing/payments between Buyers and translators. Even with translation management systems (TMSs), the tasks of LSPs are still mostly manual with continual updates to project status and translation reviews. In short, LSPs orchestrate the translation process and relieve Buyers from needing dedicated localization departments.

Who Benefits From the Standard Model

The primary beneficiaries of the Standard Model are Buyers which can have their content translated without investing heavily in a localization team and the LSPs. While LSPs do provide a much-needed project management function, they are in control of the purse strings and, like other businesses, they exist to maximize their share of the purse.

Weaknesses of the Standard Model

Inadequacies range from workflow issues to fairness.

Project management overhead is a glaring inefficiency. Despite the use of TMSs, there is simply too much human administration and intervention throughout a project.

Translation delays can result from time differences between LSPs and their geographically-dispersed pool of translators especially when issues can not be resolved promptly.

Security is a major problem for the Standard Model. LSPs do not know who is actually doing the translating. Online MT engines have been used to translate texts risking exposure of propriety material. [3]

The inequity of the Standard Model, where the Buyer wants to minimize translation costs and the LSPs want to maximize their profits, leaves [especially freelance] translators at the bottom of the food chain.


A Blockchain-based Translation Workflow

Breaking with the Standard Model, the proposed translation workflow reduces human administration and improves translation workflows with blockchain as the backbone [4]

Blockchain, smart contracts and oracles are the key pieces of the proposed workflow. A blockchain is a decentralized ledger of immutable, encrypted records (blocks) securely denoting asset transfers such as source/target files. Since each block on a blockchain contains the identifiers of a provider and recipient of an asset, the provenance of an asset is traceable.

A smart contract is a computer protocol that is intended to enforce the execution of a contract without a third party. Smart contracts could facilitate direct source-target file transfers between Buyers and translators and quicker payment for translators when projects are completed. These transactions are trackable and irreversible.

For a wide set of applications, a blockchain, actually, a smart contract, might require real-world information. Fulfilling that need, a blockchain oracle is an entity providing network-external data through an external transaction. Linguists and MT engines are examples of oracles receiving data (source files) and sending data (target files) to a translation blockchain workflow.

Figure 1: Blockchain Translation Schematic

A skip through the workflow

While the SkipThrough of the blockchain translation schematic (Figure 1.) glosses over many details of the translation process, it points to where the human workflow administration performed by LSPs is completed by smart contracts with a blockchain recording the handoffs of files and payments.

  1. As with the Standard model, a Buyer assembles source documents and project requirements including target languages, linguistic assets (terminologies and TMs), budgets, and schedules.
  1. The source content, linguistic assets, and requirements are recorded on a blockchain.
  1. Smart contracts execute throughout the workflow directing texts to TMs, then to MT engines, or directly to MT engines or translators. Each file transfer is recorded on the blockchain.
  1. Based on MT review acceptability, smart contracts initiate transfers of translations to Translators/MTPEs for review. The blockchain records the translation transfers.
  1. Once a reviewer approves the translations, a smart contract executes, thereby sending translated material to the Buyer. The blockchain is updated.
  1. The buyer’s acceptance of the completed translations invokes a smart contract that sends payment to the Translators/MTPEs. The blockchain records the details of payments.


Who Benefits from the Blockchain Model?

The Buyers and Linguists are the primary beneficiaries. That is not to say that there is no role for LSPs in the translation industry. Until there is widespread adoption of a blockchain or some other nearly fully-automated model of translation, LSPs will co-exist with automation. Also, Buyers may turn to LSPs for projects when they lack the availability of PMs.

Buyers

Buyers want quality translations with tight deadlines and a limited budget.

The streamlining of the translation workflow with blockchain replacing much of the managerial overhead will reduce costs. The savings could be used for other localization projects and ideally for fairly compensating the translators.

Because every file transfer or handoff is recorded on the blockchain, the Buyer is fully aware of the state of the project at any time. They now have access to real-time project management.

The Buyer’s material is secure. Inherent features of blockchains in recording asset transfers are transparency, traceability, and security. With the appropriate service agreements with any of the oracles involved, security weak points can be tightened.

Increased automation leads to faster translations. Since files are transferred and tracked upon execution of smart contracts, time differences are eliminated and translations can be produced 365/24/7.

Linguists

As noted, linguists prefer to work directly for clients with better pay, less time pressure, and recognition.

In the proposed workflow, the work of translators can be tracked. Those producing quality translations would gain recognition and could be compensated in a fair, transparent, and consistent way. There is another side effect of traceability as well. Consistent errors either by humans or an MT engine could be identified and corrections made to improve quality on future projects.

Without the intermediate LSPs, linguists will be working and communicating directly with the Buyer. Yes, the workflow will be automated, but there would be no obstacles to human communication between the Buyer and the translators.

Elimination of time-difference delays and the human management level could allow for more time for actual translation and should lessen the time pressure felt by translators today.


Obstacles to the Blockchain Workflow

While there is much to be gained from the blockchain workflow, there are three broad hurdles for its success. One is technical and the others are due to industry resistance.

  • The technical obstacles are huge. Despite all the hype and predictions, blockchain technology is in its infancy and its future remains uncertain. Nevertheless, blockchain technology is viable and scales with Bitcoin as the most visible example. So the utility of blockchain cannot be dismissed a priori.
  • Industries do not usually embrace change, especially when it changes their business models. The Standard Model represents business as usual and it has taken time and effort to put the infrastructure into place. LSPs, who play a valuable role in today’s translation process, would be most resistant to change. However, the overhead of LSPs drives up translation costs, perhaps at the expense of the translators at the far end of the supply chain.
  • Another major barrier is the human one. With the diminished roles of LSPs, localization managers would need to adapt to a new work environment, undoubtedly a stressful situation. Translators, especially those who have well-established relationships with LSPs, would lose a conduit for work and would also have to adapt. However, there could be monetary rewards as their work and expertise are recognized via blockchain.

Change is not easy!

A Few Final Remarks

Blockchain can provide the backbone for a supply chain that brings equity, improved work arrangements, and recognition to translators. At the same time, blockchain with smart contracts streamlines the workflow by automating much of the current managerial tasks. Granted, the blockchain model is a heavy lift and faces opposition from stakeholders that control much of the workflow today. In any case, whether technical, legal, or social pressures bring about change to the supply chain, solutions for a more equitable translation environment are being discussed and concrete solutions are being proposed. Change seems inevitable.


[1] See TAUS Webinar “Blockchain: When the Token Economy Meets the Translation Industry” - https://blog.taus.net/blockchain-when-the-token-economy-meets-the-translation-industry; For fair pay, see: TAUS blog “Fair Pay for the translators and data-keepers!” - https://blog.taus.net/2021-according-to-taus

[2] Pielmeier, Hélène, and Paul O’Mara, “The State of the Linguist Supply Chain,” CSA Research, January 2020.

[3] See: https://slator.com/technology/translate-com-exposes-highly-sensitive-information-massive-privacy-breach/ and https://www.nrk.no/urix/warning-about-translation-web-site_-passwords-and-contracts-accessible-on-the-internet-1.13670874

[4] For other blockchain architectures, see: Exfluency - https://www.exfluency.com and Kuhns, Bob, “The Pros and Cons of Blockchains and L10N Workflows,” TAUS White Paper, March 2019 - https://www.taus.net/insights/reports/the-pros-and-cons-of-blockchains-and-l10n-workflows-white-paper






Bob Kuhns is an independent consultant specializing in language technologies. His clients have included the Knowledge Technology Group in the Sun Microsystems Labs and Sun’s Globalization group. In the Labs, Bob was part of a team developing a conceptual indexing system and for the Globalization group, he was the project manager and lead translation technology designer for a controlled language checker, a terminology management system, and a hybrid MT system. He was also responsible for developing translation metrics and leading a competitive MT evaluation. Bob has also conducted research and published reports with Common Sense Advisory, TAUS, and MediaLocate on a variety of topics including managed authoring, advanced leveraging, MT, blockchain, and L10n workflows, and global social media.

Bob’s email is: kuhns@rcn.com