Pages

Showing posts with label Controlled Language. Show all posts
Showing posts with label Controlled Language. Show all posts

Friday, November 9, 2012

Understanding Post-Editing

MT continues to build momentum because the need for large global enterprises to make more information available faster, continues relentlessly. There are still some who question the “content tsunami”, and we are now getting some data points that define this for industry players in very specific terms, for those who are still doubtful.  For example, last week at AMTA 2012, a senior Dell localization professional gave us a specific data point: Dell has increased its volume of business and product-related translation from 30 million words to 60 million words in two years. This was done without any increase in the translation budget. This situation is mirrored across the information technology industry, and now many with information-intensive products apparently do realize that translating large amounts of product/service-related information enhances global business success. Given the speed and volume of information creation, it is often necessary and perhaps even imperative to use technology like MT. 

While much of the discussion about MT tends to get gets bogged down around linguistic quality issues, we should all remember that finally, the whole point of business translation and the whole localization industry is to facilitate cross-language trade and commerce. We now have many examples where we see that the final customers of Dell, Microsoft, Apple customers say that machine-translated content is more than acceptable even though this same content would fail a linguistic review in a typical Translate-Edit-Proof (TEP) process.  We see terms like linguistic usability and readability being applied to translated content which is often short of the TEP quality that many of us have grown accustomed to or expect. Customer expectations change and free online MT has made MT more acceptable, also as we understand that the content that is translated is being created by writers who are not really writers, for readers who do not have scholarly expectations on this content. There is content that requires TEP rigor and there is some that can be raw MT and there is much in between with various shades of grey.  This is not acquiescence to crappy quality across the board, rather, it is understanding that for a lot of business translation, MT or PEMT does produce quality that helps to accomplish business goals of getting information to customers in a cost-effective and timely manner.

Thus, we see the growing use of MT in business translation contexts, but there is still a lot of misinformation and it is useful to share more information about successful practices so that the use and adoption of this technology is more informed and the discussions can become more dispassionate and pragmatic. 

There were some recent sessions exploring what post-editing MT is about in the ProZ Virtual Conference that I participated in, that I thought might be interesting to highlight in this post.

The integrated audio/video session is available on the ProZ site by clicking on the link to the left and playing the presentation back on the “low-bandwidth” image towards the bottom of the page. (Unfortunately, the live session had many video resolution issues but the recording is fine.) I have also included the Slideshare below for those who just want to see the basic content of the presentation. Hopefully, this presentation does provide a more realistic perspective on what is and is not possible with MT.



A second session included a panel discussion on post-editing with speakers from various translation agencies talking about their direct experiences with post-editing MT (PEMT).
PEMT is an important issue to understand as there are very strongly felt opinions on this (many based on actual bad experiences) but the signal-to-noise ratio is still very poor. Many translators feel that the work is demeaning and are not interested in doing it and practitioners should understand this. However, much of the negative feedback is based on early practices where the MT quality was very bad and translator/editors were paid unfairly for the effort involved. Some recent feedback from TAUS even suggests that many translators are considering leaving the profession because they do not enjoy this type of work. Better MT and fair compensation practices can address some of this dissatisfaction.  While early experiences often only focus on the most mechanical aspects of PEMT, I think there is an opportunity for the professional translation industry to get more engaged in solving different kinds multilingual problems e.g. Support Chat, Customer Forum discussions where translation could greatly enhance global customer satisfaction and increase dialogue and engagement.
 
I think that as we as an industry could further our prospects and greatly reduce the emotional content in the debate and discussion by getting better definitions of quality across the spectrum of content that is worth translating to facilitate commerce.  Competent TEP and raw online free MT are two opposite ends of the quality spectrum and it would be useful to get better definitions of the useful quality levels for the variety of grey shades in-between. Preferably in terms that are meaningful to the consumers of that content rather than in terms of linguistic errors.
In the PEMT context, it would useful for both translation agencies and translator/editors to understand the specific MT output quality involved better, so that compensation structures can be set more rationally and equitably.  This quality assessment, I believe is an opportunity for translators to develop measures that link the quality of specific MT output to their compensation on a project-by-project basis.  My previous post suggests one such approach but I am sure there are many other ways that translators can rapidly assess the scope and difficulty of a PEMT task and help the agencies and the buyers understand equitable compensation structures based on trusted measurements of the scope of work.
There is also a growing discussion on what an ideal PEMT environment looks like and Jost Zetsche provided some clues in the 213th edition of his newsletter. But basically, we need tools that provide some different context since MT errors are not quite the same as TM fuzzy match errors. Perhaps some of the frustration that translators have, stems from expecting to see the same type of errors as they see in low fuzzy matches.   I would suggest the following for an ideal PEMT environment:
  • Rapid Error Detection (Grammar and Spelling Checkers)
  • Rapid Error Correction (e.g. Move word order, global correct and replace)
  • Dictionary and Terminology DB links
  • Error Pattern Identification so that hundreds of strings can be corrected by correcting a pattern
  • Quality measurement utilities to assess specific and unique MT output
  • Productivity measurement tools
  • Context as well as individual segment handling
  • Tight integrations with TM
  • Linguistic data manufacturing capabilities to create corrective data
  • Regex and Word Macro-like capabilities

Tuesday, May 24, 2011

A Case Study on the Use & Benefits of Controlled Language

This is a post by guest writer Anna Fellet who I met in Rome last month. This further explores the theme of Controlled Language and Process Standards and continues on the themes presented Valeria Cannavina in her posting earlier this month. This slide presentation provides additional background on this case study.

=================================================================================
This article presents a case study to show the benefits of Controlled Language strategies, and highlights the key lessons learnt in the pilot project on a dedicated MT workflow created for ARREX Le Cucine, a leading Italian furniture company. This post also contains a reply to Laura Rossi’s comments on Valeria Cannavina’s previous posting on standards and the application of the CMMI model to translation.

A Business Case for Controlled Language
The goal of the ARREX project was the development of a corporate controlled language for Italian to be used in a customized authoring and (machine) translation workflow.

Why a Controlled Language?
 A CL was chosen to eliminate ambiguity and complexity in product data sheets, installation and maintenance instructions (for support), catalogs, price lists, orders, reports, memos, documents for compliance. We chose to test our CL with both RbMT provided by Synthema and with SMT by Asia Online. We improved the repetitiveness of ARREX texts, and the result with RbMT was successful

 As for SMT, which is a typical brute-force data-driven computing application, most of the difficulties come from the high degree of unpredictability in searching through a massive set of possible options of even in the simplest word combinations. As the number of words in a sequence increases, the precision score decreases because longer matching word sequences are more difficult to find. A controlled language increases predictability, increases statistical density and thus improves probability and boosts SMT success.

So, the single most powerful rule for authors/writers still holds its validity: one idea per sentence makes text that is easier for humans to understand also easier for MT engines to understand.

Poor source quality can lead to low quality target language content (e.g. SAP translations often result in hardly translatable/comprehensible Italian), however technical documents are ideally all written in the same “language”, even though with different idioms. Setting up terminology resources and developing writing rules enables the Italian text to be more easily handled by the MT system.

Moreover, language combinations with English are more commonly implemented, so by translating Italian into a terminologically coherent and syntactically simple English target we could use it as a starting point for other potentially successful combinations.

At the end of our preliminary investigations, we found that ARREX CL adds value to technical documentation as it allows:
  • Increase in the perceived value of the product and of the whole brand: consistent, stylistically uniform, and controllable documentation (user-targeted material) created for a user/client to understand and thus helps to build customer loyalty;
  • More efficient communication with clients/distribution partners/maintenance staff, thus reducing customer support calls and general costs associated with customer service;
  • Reduction of translation costs (see table below);
 
Source text
Without CL
With CL
Difference
Words to be translated
70.000
64.000
-8%
Repetitions
39.800
38.300
+3%
Words to be translated from scratch
32.100
25.700
-5%
Human translation costs (250 words/hour)
280 hours
255 hours
-9%





Human translation costs
MT post editing cost
Difference
Translation Costs
280 hours
50 hours
-80%

We moved beyond the study on Italian CL for customized MT, and discovered that much can be done with a holistic approach to authoring workflows.

We became convinced that by adopting an ad hoc CL, by creating reusable corporate specific terminology resources, training corporate internal staff on authoring strategies for MT we could influence the company authoring workflow at a greater extent. By ad hoc CL we mean that rules are created specifically for ARREX. These rules may be valid for other domains, as well, but we had to focus on and improve inefficient writing practices unique to ARREX’s own internal corporate-speak. We were brought to focus on critical aspects that the company may not have had clear at the beginning of the project e.g. improve ARREX terminology standards by extracting most frequently used terms, and analyze synonyms and non-standard/irregular uses of terms that ARREX had already implemented.

Not only did this project help us understand how to clean data for MT, create resources for MT (glossaries, TM, post-editing guidelines), highlight costly and time consuming translation processes, i.e. outsourcing translation/editing and publishing, it also helped us in seamlessly adapting our work to the existing corporate strategy by addressing the internal staff’s needs directly.

WORKFLOW & VALUE

The graphic below shows how we changed the company’s workflow and the results we achieved.

Activity


WORKFLOW ANALYSIS
Goal
Before
After

Exhaustive map of how ARREX processes are organized and who is in charge for what.
Texts were translated externally or (sometimes) internally.
MT (internal); monolingual review with support of term base and glossaries approved by ARREX.

Value
Faster, cheaper and more accurate translations and reviewer’s feedback for continuous improvement of CL and MT.




Activity


RESOURCES ANALYSIS


Goal
Before
After

Measuring ARREX staff performance.
No trained technical writers and no unique point of reference for the production of technical material and translated material.

Training of staff to repeat and manage the process; new professionals (pre-editing, post editing).

Value
Involvement through requalification of internal staff in charge for documentation.
Resources are the only feature of the whole process that cannot be cut or reduced. People will always be the key element to deliver ‘quality’. Internal staff is the best people to talk to, to understand the quality level expected. We are only providing the right tools and the right knowledge to achieve such ‘quality’ (e.g. glossaries, CL style guide, MT workflow, QA report). We offer an improvement of the quality of the process, not of the product. Products can always change.


PLANNING
Goal
Before
After

Address the workflow step by step to build long term relationships with valid collaborators.
Undocumented processes, undefined organization of roles for technical writing and translation.
Ad hoc procedures for each phase of the process, from technical writing to delivery of translated material.

Value
The only possible way to deal with planning is to set a common framework to communicate with ARREX to find the appropriate strategy for text editing and MT.

Activity
ACTION
Goal
Before
After

Write a protocol of requirements suitable for new requests.
Fragmented process.
Independent and autonomous management.
Value
Flexible processes become repeatable.

Repeatable processes


  • Process documentation;
  • Roles definition and (re)qualification;
  • Building of internal writing team;
  • Internal terminology approval procedure;
  • Target Language Monolingual reviewer selection;
  • Target Language Monolingual reviewer feedback.

Q&A to questions posed by Laura Rossi in comments of previous post.

Laura Rossi: Will translation software developers be ready to provide their customers just with what they need, instead of trying to ‘impose’ an overall comprehensive solution, which, in fact, force them to follow a specific process and workflow?

Will the definition of a standard model not be another reason for them to justify this rigidity?

ARREX was anchored to old trusted but imperfect and inefficient processes, and the change we introduced was sometimes shocking for internal technical writers. “The difficulty lies, not in the new ideas, but in escaping from the old ones” (John Maynard Keynes), but if one sees the new idea as a means to improve one’s work (and save time), participation will be natural. In this sense, ARREX drove its own change.

Laura Rossi: As long as translation and localization will be considered as an accessory activity and a cost by the customers, more than a possibly business-driving and revenue-generating task, there won't be much interest from side of the customers to rethink their internal processes and organization, as well as from side of the LSPs and translation software companies to really act as part of their customers' development and production cycle.
 
I fully agree with Laura’s response to Valeria’s Post, it’s time for “translation (software) companies to really act as part of their customers' development and production cycle”, but I do not agree that service providers should “teach customers to involve LSPs in an early stage of development”. I think that it’s the other way around.

Providers should be able to integrate seamlessly in a company strategy for content, and detect processes that can be improved. This is what we could define as a holistic approach to content creation, where translation is only one piece of a broader internal and external corporate communication puzzle.

Laura Rossi: I think the landscape is actually changing, but the change is still quite slow, especially from the side of the customers, and I wonder what will be able to cause the shift on a massive scale from the traditional way of seeing translation as a 'service industry' to consider it an essential part of a business.

I might be wrong, but I suspect this shift is driven by economic imperatives, and MT offers terrific time and cost savings. Now MT is the right technology, handling repetitive tasks to let humans do what they are best at, but technology can be applied to processes, not to outcomes. One useful approach to a realistic, sustainable translation market is to explicitly differentiate between processes and outcomes.
This is why we focus on the quality of the process, instead of the quality of the product/output, and think of a Customer Centered Business Model, with single services satisfying multiple needs.

“Quality in a product or service is not what the supplier puts in. It is what the customer gets out and is willing to pay for. A product is not quality because it is hard to make and costs a lot of money, as manufacturers typically believe. Customers pay only for what is of use to them and gives them value. Nothing else constitutes quality” (Peter Drucker).

We see that: 

  • translation students are not trained on MT (in Italy), and mostly they don’t have a sense for the realities of the professional translation workplace after graduation;
  • professional translators are very suspicious of MT, and generally do not welcome new ways of approaching the job, perhaps because they don’t have direct access to the client company, due to agencies (LSPs) intermediation;
  • Agencies (LSPs) see MT only as a means to pay translators much less than what they pay them presently.
In this scenario, translators with the skills required to offer premium services would just abandon the industry. Underpaid and undervalued, they will simply disappear with no one to replace them. Those hard-to-acquire skills will be transferred to other areas.
We believe there are possibilities for new approaches to content creation, and translation management, and that companies wishing to change the way they write and translate their content, like ARREX did, will drive this change, not LSPs.

Laura Rossi: Can we avoid the trap ‘we-are-following-a-standard-or-model-therefore-we-are-good’, which, as you say, can ‘hook’ the customers, but, in my view, does not necessarily ensure their satisfaction?

How can we make sure that a possible specific translation industry standard process model will be flexible and modular enough so as to avoid the risk of LSPs and customers ‘anchoring’ to that as a ‘given’ and a ‘must’?

During the ARREX project we saw that it was hard for clients to ask for specific services since they see translation as marginal and take it for granted, and as Renato Beninatto often says, “translation is really like toilet paper, it’s only important when it’s not there when you need it.” 

We ended seeing translation as a product, and not as a service provided with methods akin to those of industrial production. With the commoditizatioin of translation, i.e. with almost no difference between suppliers, there is an undue and ineffective emphasis on prescriptive standards and the ‘we-are-following-a-standard-or-model-therefore-we-are-good’ scenario. Prescriptive standards, though, are useless for those LSPs wishing to differentiate, and to adapt their service to the client’s needs, because they may have to change their service and approach for the unique requirements of each client. Process Standards are not, cannot be, and should not be laws, not even strict regulations because every company is different. The general state of information asymmetry between the LSP and client, make a process standard useful only if it implies transparency and flexibility. Reiterative and rigid procedures, instead, lead to static monolithic workflows.

This is why a scalable and transparent path like CMMI is useful. Only if client and customer are transparent in processes, can they find the most adequate actions to interact. The client will explain (and understand) its own level of maturity (requirements) to leverage the service provided, and the provider will be able to address the unique needs of the customer.

When it comes to adopting standards, a company does not know exactly where the ensuing process changes will lead. It can also lead to other changes that were not originally envisioned. This happened to ARREX as well: when they saw that along with improving their authoring and writing strategy, they could also act on other issues, i.e. translation, they did not hesitate in considering that, as well.

Laura Rossi: Is it really possible to capture in a standard something as ‘subjective’ as quality?

One could wonder what standards really mean to customers. Are they all concerned about “quality”? Quality is subjective in this sense that it is subjective and dependent on each customer.

A list of ‘ad hoc’ requirements to measure the level of adequacy of the service for the particular customer is useful both to assess customer satisfaction, service improvements, and to define client’s profile and demands in different domains. It is also true that if the quality of the product is subject to the assessment of the client, you can not say the same for the evaluation of the process. In fact, process standards should aim at increasing and improving the quality of the process. This can only be done with transparency in client-vendor relationships.

In our experience, we saw that processes such as post-editing can be measured either by the customer, in terms of satisfaction with the final result (via a series of requirements that must be met by the output of the MT output), and by the monolingual reviewer in a questionnaire on ‘linguistic’ aspects of translation and measurements of price/time/productivity. In this sense, requirements can be defined differently as: MT output for publication, pre-translation or internal use.

What I think would be extremely useful, and hope to see promoted in the industry, is a framework for pricing, especially for post-editing, to help customer-vendor relationships be more transparent. Crowdsourcing, as well, could be better and more widely accepted and used with a clear, simple, and common standards framework. Repeatable processes are worth sharing.
Mark Zuckerberg said “By giving people the power to share, we're making the world more transparent.”  It has proved to be a very profitable strategy, as well.
 
Anna graduated in 2007 in modern languages and cultures at the University of Padua, and in 2009 in technical and scientific translation at LUSPIO of Rome with a final dissertation on '‘Machine Translation: productivity, quality, customer satisfaction.’ At present she works as a freelance translator and subtitle translator and on a pilot project on Italian Controlled Language and Machine Translation with LUSPIO University, Asia Online, Synthema and ARREX Le Cucine. She can be reached at anna@s-quid.it  http://www.s-quid.it



Wednesday, May 18, 2011

Can a Controlled Language Help Machine Translation?

Here is another guest posting initiated from the LTAC conference at LUSPIO in Rome.This post authored by Orlando Chiarello, provides an overview of the benefits of controlled language, or in a broader sense improved / standardized  source material to any translation process. Historically CL has often been associated with RbMT but the benefit of cleaning and standardizing source material is beneficial to SMT as well, as the example below shows. Any efforts made to improve and/or standardize source material are very likely to result in better MT quality and help any ongoing translation automation process. While the degree of control suggested by CL is not always possible with dynamic customer content, this post presents some examples of where this approach does make sense.
 
For a more complete set of links and further discussion on this subject, some may also wish to refer to the old but still relevant discussion in the LinkedIn Automated Translation Group (requires membership) Discussion on the use of Controlled Language in SMT vs RbMT This link has a detailed discussion on how CL or source language simplification can improve the results obtained from MT. 
==================================================

The ASD* Simplified Technical English Maintenance Group, or STEMG (www.asd-ste100.org) is having its Spring Meeting these days (17 – 20 May) at Airbus in Toulouse. I am the Chair of this group and I would like to take this opportunity to give a brief overview of ASD Simplified Technical English, ASD-STE100 (STE).

STE is an international specification for the preparation of maintenance documentation in a controlled language.
 
It was developed in the early Eighties (as AECMA Simplified English) to help the users of English-language documentation to understand what they read. The STE provides a set of Writing Rules and a Dictionary of controlled vocabulary.


The Writing Rules cover aspects of grammar and style; the Dictionary specifies the general words that can be used. These words were chosen for their simplicity and ease of recognition. In general, there is only one word for one meaning, and one part of speech for one word. In addition to the specified general vocabulary, STE accepts the use of company-specific or project-oriented technical words (Technical Names and  Technical Verbs), provided that they fit into one of the categories listed in the specification.
 
The international language of many industries and specifically of the aviation industry is English and English is the language most used for technical documentation. However, it is often the native language neither of the readers nor of the authors of such documentation. Many readers have knowledge of English that is limited, and are easily confused by complex sentence structures and by the number of meanings and synonyms which English words can have.
 
The controlled grammatical structures and vocabulary – on which STE is based – have the purpose of producing texts that are easily understandable and, consequently, STE reduces errors during the maintenance tasks.
 
Although this controlled language was originally designed for the aviation  industry, companies from other industries and domains use it to standardize their documentation in an easy, understandable and unambiguous way. As an example, in March, I gave a two day training course on STE to a company located in Munich producing medical devices. The course turned out to be a great success.
 
Also, the LUSPIO University in Rome was involved in a project with an Italian company producing furniture for the development of a controlled Italian to be used by that company in all their documentation. The STE principles and rules have been the primary basis for the creation of this Controlled Italian. The results of this project were presented at the LTAC (LUSPIO Translation Automation Conference)  on 5 and 6 April where I was also invited and made a presentation of STE.
 
STE can really help also Machine Translation, which was one of the primary objectives when this Controlled Language was developed. As an example, the following is a paragraph in STE (taken from a component maintenance manual) translated into Spanish by simply copying the text in the Web Google Translator and run it:
 
Original text in STE:
The procedures in this manual are a guide to do the correct maintenance of the component. Some equivalent procedures - that come from the experience and skills of the maintenance personnel - are also satisfactory.
 
Text translated by Google:
Los procedimientos en este manual son una guía para hacer el mantenimiento correcto de los componentes. Algunos procedimientos equivalentes - que vienen de la experiencia y habilidades del personal de mantenimiento - son también satisfactorios.

As we can see, the result is quite impressive.
 
To conclude, the above example proves that if the "source text" is English and the text is written in STE, Machine Translation can be dramatically helped by the principle of "one word = one meaning". A further help to Machine Translation could be the availability of a "mirror" Controlled Language based on STE. For example, the French Aviation Industries (GIFAS) in the Eighties created the "Rationalized French" based on STE. They actually used the same structure of the Writing Rules and Dictionary and adapted them to French. The result was exceptionally good with benefit to translations in both ways. Other attempts were made and other are currently in progress with other Languages including Swedish, German, Spanish, Chinese and Italian.

“ASD represents the aeronautics, space, and defense industries in Europe. ASD has 28 member associations in 20 countries, representing over 2000 companies with a further 80 000 suppliers, many of which are SMEs. Total annual industry turnover is over €137 billion



Orlando 1
Orlando Chiarello is the Product Support Manager of Secondo Mona, an Italian aerospace equipment manufacturer. He is responsible for the aftermarket support of the company products.
 
He is also the Chairman of the ASD Simplified Technical English Maintenance Group (STEMG), responsible for the development and maintenance of the ASD-STE100 Specification.