Pages

Showing posts with label China. Show all posts
Showing posts with label China. Show all posts

Thursday, September 24, 2020

NiuTrans: An Emerging Enterprise MT Provider from China

 This post highlights a Chinese MT vendor who I suspect is not well known in the US or Europe currently, but who I expect will become better known over the coming years. While the US giants (FAAMG) still dominate the MT landscape around the world today, I think it is increasingly possible that other players from around the world, especially from China may become much more recognized in the future. 

One indicator that has been historically reliable to forecast and predict emerging economic power is the volume of patent filings in a country. This has been true for Japan and Germany historically where we saw voluminous patent activity precede the economic rise of these countries, and recently we see that this predictor is also aligned with the rise of S. Korea and China as economic powerhouses. However, the sheer volume of filings is not necessarily a lead indicator of true innovation, and some experts say that the volume of patents filed and granted abroad is a better indicator of innovation and patent quality. But today we see emerging giants from Asia in consumer electronics, automobiles, eCommerce, internet services, and nobody questions the building innovation momentum happening in Asia today. 


Artificial Intelligence (AI) is heralded by many as a key driver of wealth creation for the next 50 years. To build momentum with AI requires a combination of access to large volumes of "good" data, computing resources, and deep expertise in machine learning, NLP, and other closely related technologies. Today, the US and China look poised to be the dominant players in the wider application of AI and machine learning-based technologies with a few others close behind. And here too deep knowledge and clout are indicated by the volume of influential papers published and referenced by the global community. A recent analysis, by the Allen Institute for Artificial Intelligence in Seattle, Washington found that China has steadily increased its share of authorship of the top 10% most-cited papers. The researchers found that America’s share of the most-cited 10 percent of papers declined from a high of 47 percent in 1982 to a low of 29 percent in 2018. China’s share, meanwhile, has been “rising steeply,” reaching a high of 26.5 percent last year, Though the US still has significant advantages with the relative supply of expert manpower and dominance in manufacture of AI semiconductor chip technology, this too is slowly changing even though most experts expect the US to maintain leadership for other reasons. 

Credit: Allen Institute for Artificial Intelligence

These trends also impact the translation industry and they change the relative benefit and economic value of different languages. The global market is slowly changing from a FIGS-centric view of the world to one where both the most important source language (ZH, KO, HI) and target languages are changing.  The fastest-growing economies today are in Africa and Asia and are not likely to be well served by a FIGS-centric view though it appears that English will remain a critical world language for knowledge sharing for at least another 25 years. These changes create an opportunity for agile and skillful Asian technology entrepreneurs like NiuTrans who are much more tuned-in to this rapidly evolving world.  I have noted that some of the most capable new MT initiatives I have seen in the last few years were based in China. India has lagged far behind with MT, even though the need there is much greater, because of the myth that English matters more, and possibly because of the lack of governmental support and sponsorship of NLP research.


The Chinese MT Market: A Quick Overview

I recently sat down with Chungliang Zhang from NiuTrans, an emerging enterprise MT vendor in China, to discuss the Chinese MT market and his company’s own MT offerings. He pointed out that China is the second-largest global economy today, and it is now increasingly commonplace for both Chinese individuals and enterprises to have active global interactions. The economic momentum naturally drives the demand for automated translation services.

Some examples, he pointed out:

In 2019, China’s outbound tourist traffic totaled 155M people, up 3.3% from the previous year. This massive volume of traveler traffic results in a concomitant demand for language translation. Chungliang pointed out that this travel momentum significantly drives the need for voice translation devices in the consumer market like those produced by Sougou, iFlyTek, and others, which have been very much in demand in the last few years.

There is also a growing interest by Chinese enterprises, both state-owned or privately owned, to build and expand their business presence in global markets. For example, Alibaba, China’s largest eCommerce company, is listed on the NYSE and has established an international B2B portal (Alibaba.com) where 20 million enterprises gather and work to “Buy Global, Sell Global.” Currently, the Alibaba MT team builds the largest eCommerce MT systems globally, often reaching volumes of 1.79 billion translation calls per day, which is a larger transaction volume than either Google or Amazon.

“All in all, as we can see it, there is a clear trend that MT is increasingly being used in more and more industries, such as language service industries, intellectual property services, pharmaceutical industries, and information analysis services.”

While it is clear that consumers and individuals worldwide are regularly using MT, the primary enterprise users of MT in China are government agencies and internet-based businesses like eCommerce. This need for translation is now expanding to more enterprises who seek to increase their international business presence and realize that MT can enable and accelerate these initiatives.

The Chinese MT technology leaders in terms of volume and regular user base are the internet services giants (such as Baidu, Tencent, Alibaba, Sogou, Netease) or the AI tech giants (such as iFlyTek). Google Translate and Microsoft Bing Translator are also popular in China since they are free, but they don’t have a large share of the total use if the focus is strictly on MT technology.

When asked to comment on the characteristics and changes in the Chinese MT market, Chungliang said:

“In our understanding, Sogou and iFlytek's primary business focus is the B2C market, and thus both of them develop consumer hardware like personal voice translators. Sogou was recently (July 29, 2020) purchased by Tencent (a major social media player), so we don’t know what will happen next. iFlytek is famous for its Speech-To-Speech technology capabilities. Thus it is natural for them to develop MT, to get the two technologies integrated and grab a larger share of the market.

As for the other important MT players in China, Alibaba MT mainly serves its own global focused eCommerce business, and Tencent Translate focuses on providing the translation needs of its users in social networking use scenarios. Like Google Translate, Baidu Translate is a portal to attract individual users who might need translation during a search. It also serves to expand Baidu’s influence as a whole. While Netease Youdao focuses on the education industry, and the Youdao Team integrates the Youdao online dictionary, direct MT, and human translation.

What are the main languages that people/customers translate? As far as we know, the most translated language is English, Japanese is second, followed by Arabic, Korean, Thai, Russian, German, and Spanish.” Of course, this is all direct to and from Chinese.”


NiuTrans Focus: The Enterprise

The NiuTrans team learned very early in their operational history and during their startup phase that their business survival was linked to providing MT services for the enterprise rather than for individual users and consumers. The market for individuals is dominated by offerings like Google Translate and Baidu Translate that offer virtually-free services. In contrast, NiuTrans is focused on meeting the enterprise demands for MT, which often means deploying on-premise MT engines and the development of custom engines. These enterprises tend to be concentrated around Intellectual Property and Patent services, Pharmaceuticals, Vehicle Manufacturing, IT, Education, and AI companies. For example, NiuTrans builds customized patent-domain MT engines for the China Patent Information Center (CNPAT, a branch of the China National Intellectual Property Administration, a large-scale patent information service based in Beijing.)

CNPAT has the largest collections of multilingual parallel data for patents, and services ongoing and substantial demands for patent-related MT needs in various use scenarios such as patent application filing and examination, patent-related transactions, and patent-based lawsuits. Given the scale of the client’s needs, NiuTrans sends an R&D team on-site to work with CNPAT’s technical team for data processing and data cleaning. This data is then used in the NiuTrans.NMT training module to develop patent-domain NMT engines on CNPAT’s on-premise servers. The on-site team also develops custom MT APIs on-demand to fit into CNPAT’s current workflow and customer servicing needs.


Besides powering and enabling the specialized translation needs of services like CNPAT, NiuTrans also provides back-end MT services for industrial leaders, including iFlyTek (also an early investor in NiuTrans), JD.com (the No. 2 eCommerce business in China), Tencent (the largest social networking company in China), Xiaomi (a leader of smart devices OEMs in China), and Kingsoft (a leader of office software in China).

NiuTrans has an online cloud API that also attracts 100,000+ small and medium enterprises interested in expanding their international operations and business presence. The pricing for these smaller users are based on the volume of characters these users translate and is much lower than Google Translate and Baidu Translate prices.

NiuTran’ Online Cloud User Locations

You can visit the NiuTrans Translate portal at https://niutrans.com

NiuTrans write and maintain their own NMT code-base rather than use open source options for NiuTrans.NMT and claim that they achieve comparable, if not better, quality performance with their competitors. Their comparative performance at the WMT19 evaluations suggests that they actually do better than most of their competitors. They are not dependent on TensorFlow, PyTorch, or OpenNMT to build their systems. Today, NiuTrans is a key MT technology provider, especially for enterprises in China.

NiuTrans.NMT is a lightweight and efficient Transformer-based neural machine translation system. Its main features are:

  • Few dependencies. It is implemented with pure C++, and all dependencies are optional.
  • Fast decoding. It supports various decoding acceleration strategies, such as batch pruning and dynamic batch size.
  • Advanced NMT models, such as Deep Transformer.
  • Flexible running modes. The system can be run on various systems and devices (Linux vs. Windows, CPUs vs. GPUs, FP32 vs. FP16, etc.).
  • Framework agnostic. It supports various models trained with other tools, e.g., Fairseq models.
  • The code is simple and friendly to beginners.

When I probed into why NiuTrans had chosen to develop their own NMT technology rather than use the widely accepted open-source solutions, I was provided with a history of the company and its evolution through various approaches to developing MT technology.

The NiuTrans team originated in the NLP Lab at Northeastern University, China (NEUNLP Lab), a machine translation research leader in the Chinese academic world going as far back as 1980. Like many elsewhere in the world, the team initially studied rule-based MT from 1980 to 2005. In 2006 Professor Jingbo Zhu (the current Chairman of NiuTrans) returned from a year-long visit to ISI-USC and decided to switch to statistical MT research working together with Tong Xiao, who was a fresh graduate student at the time and is now the CEO of NiuTrans. They made rapid strides in SMT research, releasing the first version of NiuTrans SMT open source in 2011. At that time, Chinese academia primarily used Moses to conduct MT-related research and develop MT engines. The development of the NiuTrans.SMT open-source proved that Chinese engineers could do the same as, or even better than Moses, and also helped to showcase the strength and competence of the NiuTrans team. Thus, in 2012, confident with their MT technology and armed with a dream to expand the potential of this technology to connect the world with MT, the NiuTrans team decided to form an MT company, converting the 30+ years’ of MT research work to developing MT software for industrial use.

Given their origins in academia, they kept a close watch on MT research and breakthroughs worldwide and noticed in 2014 that there was a growing base of research being done with neural network-based deep learning models. Therefore, the NiuTrans team started studying deep learning technologies in 2015 and released its first version of NiuTrans.NMT in December 2016, just three months after Google announced the release of its first NMT engines.

NiuTrans prefers to avoid using open source MT platforms like TensorFlow, PyTorch, or OpenNMT as they have developed deep competence in MT technology gathered over 40 years of engagement. The leadership believes there are specific advantages to building the whole technology stack for MT and intend to continue with this basic development strategy. As an example, Chunliang pointed me to the release of NiuTensor, their own deep learning tool: (https://github.com/NiuTrans/NiuTensor) and NiuTrans.NMT Open Source (https://github.com/NiuTrans/NiuTrans.NMT). They are confident that they can keep pace with continuous improvements in open source with support from the NEUNLP Lab, which has eight permanent staff and 40+ Ph.D./MS students focusing on MT issues of relevance and interest for their overall mission. This group also allows NiuTrans to stay abreast of the worldwide research being done elsewhere.

NiuTrans understands that a critical requirement for an enterprise user is to adapt and customize the MT system to enterprise-specific terminology or use. Thus, it provides both a user terminology module to introduce user terminology into the MT system and a user translation memory module to introduce the users’ sentence pairs to tune the MT system. Another more sophisticated solution is incremental training. They incorporate user data to modify the NiuTrans model parameters to get the MT model better adjusted to user data features.

NiuTrans also gathers post-editing feedback on critical language pairs like ZH <> EN and ZH <> JP on an ongoing basis, then analyze error patterns to develop continuing engine performance improvements.


Quality Improvement, Data Security, and Deployment

NiuTrans evaluates MT system performance using BLEU and a human evaluation technique that ranks relative systems. They prefer not to use the widely used 5-point scale to assign an absolute value to a translation. Thus if they were comparing NiuTrans, Google, and DeepL, they would use a combination of BLEU and have humans rank the same blind test set for the three systems.

NiuTrans also has an ongoing program to improve its MT engines continually. They do this in three different ways:

  1. Firstly, as the company has a strong research team that is continually experimenting and evaluating new research, the impact of this research is continuously tested to determine if it can be incorporated into the existing model framework. This kind of significant technical innovation is added into the model two or three times a year.
  2. Secondly, customer feedback, ongoing error analysis, or specialized human evaluation feedback also trigger regular updates to the most important MT systems (e.g. ZH<>EN) at least once a month.
  3. Thirdly, engines will be updated as new data is discovered, gathered, or provided by new clients. High-quality training data is always sought after and considered valuable to drive ongoing MT system improvements.

NiuTrans has performed well in comparative evaluations of their MT systems against other academic and large online MT solutions. Here is a summary of the results from WMT19. They report that their performance in WMT20 is also excellent, but final results have not yet been published.

NiuTrans training data comes mainly from two sources: data crawling and data purchase from reliable vendors.

NiuTrans uses crawlers to collect the parallel texts from the websites that do not prohibit or prevent this, e.g., some Chinese government agencies’ websites that often provide data in several languages. They also buy parallel sentences (TM) and dictionaries from specific data provider companies, who might require signing an agreement, specifying that the data provider retains the intellectual property rights of the data.

NiuTrans gets the bulk of its revenue from data-security concerned customers who deploy their MT systems on On-premise systems. However, NiuTrans is also working on an Open Cloud https://niutrans.com offering, allowing customers to access an online API and avoid installing the infrastructure needed to set up on-premise systems. The Open Cloud is a more cost-effective option for smaller SME companies, and NiuTrans has seen rapid adoption of this new deployment in specific market segments.

International customers, especially the larger ones, much prefer to deploy their NiuTrans MT systems on-premise. For those international customers who cannot afford on-premise systems, the NiuTrans Open Cloud solution is an option. This system is deployed on the Alibaba Cloud that is governed by Chinese internet security laws that require that user data be kept for six months before deletion. The company plans to build another cloud service on the Amazon Cloud for international customers who have data security concerns. This new capability will allow users to encrypt their data locally, transfer the data securely to the Amazon Cloud. NiuTrans will then decrypt the source data on their servers, translate it, and finally delete all the user data and the corresponding translation results once the source data has been translated.


NiuTrans currently has 100+ employees, directed by Dr. Jjingbo Zhu and Dr. Tong Xiao, two leading MT scientists in China. Shenyang is the seat of the company’s headquarters and R&D team as well. Technical support and services are available in Beijing, Shanghai, Hangzhou, Chendu, and Shenzhen currently, but the company is now exploring entering the Japanese market, with the assistance of partners in Tokyo and Osaka. While NiuTrans is not a well-known name in the US/EU translation industry today, I suspect that they will become an increasingly better-known provider of enterprise MT technology in the future.


Friday, April 6, 2018

UTH - Another Chinese Translation Memory Data Utility

This is a guest post by Henry Wang of UTH. I include a brief interview I conducted before Henry wrote this post. I think this focus on developing a data marketplace is interesting as I happen to believe that the data used to train the machine learning systems is often more important than the algorithms themselves. The number of open source toolkits available for building Neural MT system is now almost 10. 

 I do not have a sense of whether the quality of the UTH data is better than other data utilities that exist and this post is not an endorsement of UTH by me. They, however, appear to be investing much more effort in cleaning the data, but I still feel that the metadata is still sorely lacking for real value to come from this data. And metadata is not just about domain classification. It will be interesting to see the quality of the MT systems that are built using this data, and that evidence will be the best indicator of the quality and value of this data to the MT community.

These data initiatives in China also reflect the building AI momentum in China. If you have the right data you can learn to develop high-value narrow purpose focused machine learning solutions. 


  1. What are the primary sources of your data?
Henry: The primary sources of our data include LSPs(language service providers), freelance translators, language service buyers, and several big data organizations.
  1. Can you describe the metadata that you allow users to access to extract the most meaningful subsets for their purposes? Can you provide an overview of your detailed data taxonomy?
Henry: We created a three-tier pyramid structure of the data with 15 top-tier domains, 41 intermediate domains, and 178 bottom-level domains. Users can extract the subsets by choosing domain names (among the three tiers), language combinations, and other items that we provided and are going to provide on our product UIs.
  1. Who are your primary customers?
Henry: MT companies/labs, LSPs, AI companies, e-commerce companies and universities
  1. Do you price differently for LSP who might use less data than for MT developers who need much more data?
Henry: Yes
  1. Do you plan to provide an English interface so that users across the world can also access your data?
Henry: Yes, we have launched several products with English UIs, including Sesame Search (www.zhimasousuo.com).
  1. Do you have your own MT solution? How does it compare with Google for some key languages?
Henry: We are working on that. We also partner with Sogou and several MT labs in China for different language combinations. We believe we will do better than Google in China-related language pairs, and this will come true within 2 years.
  1. Do you see an increasing interest in the use of this kind of language data? What other applications beyond translation?
Henry: Yes, an increasing number of leading AI, e-commerce, MT, and cross-border business companies are reaching out to us for cooperation. Also, we see a big potential in the education/e-learning field. Sesame Lingo is one of our innovative products for language teaching and training with the language data in the core database. Other applications include smart writing and pure data mining that might be applicable to many industries.
  1. What are some of the most interesting research applications of your data from the academic sector?
Henry: Corpus-based studies, and a lot of others.
  1. What are the most urgent data needs that you have by language where there is not enough data?
Henry: Southeast Asian languages, and South Asian languages.
  1. Are you trying to create new combinations of parallel data from existing data? e.g. If there is English to Hindi and English to Chinese in the same domain and subject – could you align the data to create Hindi <> Chinese data?
Henry: Yes, we already mastered that technology years ago, thus an increasing number of language combinations and an increasing amount of data.
  1. What is your feeling about the general usefulness of this kind of data in future?
Henry: With the development of data mining technologies, it will be applied to many more industries for sure. We are currently working very hard on in-context data and comparable data, which will be even more useful.

========


UTH, a Shanghai-based company, is a pioneer in the language service industries. UTH’s mission is to deliver innovative solutions to overcome challenges in the language services with petabyte translation data. Since 2012 when it was founded, UTH has accumulated more than 15 billion translation units across over 220 languages, including Arabic, Bulgarian, Chinese-Simplified, Chinese-Traditional, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Japanese, Korean, Polish, Portuguese, Romania, Russian, Slovenian, Spanish, Swedish, Thai, Lao, and Khmer, which enables it to secure a strong foot holding in China’s Belt and Road Initiative with a majority coverage of the languages used in participating countries and helps it win the support and cooperation of research institutions, language service buyers and providers, IT giants, e-commerce companies, government agencies as well as the investment support from venture capitals. Last year, it had successfully completed its Series B investment from Sogou, the second largest search engine by mobile queries in China. Sogou has completed its own IPO last year and posted $908.36 million of revenue in FY17.


UTH enhances its translation data business with the diversification in the handy tools in MT, language teaching and learning, and corpus research, which in turn sharpens its insights in the exploitation of big language data and artificial intelligence. Sesame Lingo is one of its products used for language teaching and training with parallel corpora data in the core database, and Sesame Search is an online corpus platform featuring multiple dimensional data classification, search, intuitive data presentation and patented language processing technologies. Recently, UTH has completed several acquisitions to expand its business territory in e-learning, smart writing, data mining and language services. With the strong alliances of Sogou, 5 mid-sized LSPs, 2 AI companies, and much more in 2018, UTH has already established an initial eco-system and become the largest translation database in China. UTH has seen an increasing number of leading AI, e-commerce, MT, and cross-border business companies worldwide reaching out to it for potential collaboration opportunities.


UTH embarks on a pioneering road similar to TAUS, yet UTH possesses a uniquely different advantage. TDA from TAUS is based on a data-sharing mechanism, and the control of data quality is largely determined by data-owners’ integrity and their internal quality control process. When accumulating the language data, TAUS exploits a data-breeding technology, in which it cross-selects the translation units from different languages but with a common translation in the third language to form new pairs. At UTH, more than 50 in-house corpus linguists and engineers, supported by around 400 contracted linguists, are working meticulously in language data sourcing, collection, alignment, and annotation, overseen by the trained testers under rigorous internal quality rules. UTH has formulated a relatively complete set of language quality management practices with reference to LQA models and ISO standards, embedded in the in-house tools for higher efficiency.




 UTH’s close cooperation with academia sectors imbues the company with a unique perspective on the potentials of language data. Its data are classified into a unique three-tier pyramid (15 Level I domains, 41 Level II domains and 178 Level III domains) for the purpose of mapping the requirements in LSPs to the academic disciples in Chinese universities and making the data easily accessible to teachers and students on campus, which have won wide acclaims from education experts. In addition, the company launched its education cooperation initiatives in 2017, building several internship bases and joint research programs with prestigious universities in China and overseas, including Southeast University, the University of International Business and Economics, and Nanyang Technological University.




UTH’s focus on in-domain and in-context data is currently its priority and its major differentiator. As the largest repository of parallel articles in China, UTH is cooperating with LSPs (language service providers), freelance translators, language service buyers, and several big data organizations, orchestrating the high-quality data flow among these organizations and turning the immobile language data into flowing values. As a hub in the data exchange, filtering, and processing, UTH becomes an indispensable part and a booster in this trade.

Nowadays, with the increasingly wide application of NMT in Twitter, Facebook, WeChat, QQ and UGC platforms as well as the industrial application of MT in interpretation that help people connect each other across language barriers, translation is growing into a crucial business energizer. However, the technology edges of the forerunners such as Google is diminishing, resulting in a closer gap in the translation quality among NMT vendors, including Bing, SYSTRAN, SDL, DeepL and Baidu, Sogou, NetEase, and iFlytek in China. NMT is a data-hungry application, where data is fed into neural networks to improve its intelligence. Therefore, good quality and fine-tuned translation data will become a crucial part of this fierce competition.




As a trailblazer in China, UTH is now feeding its translation data to several MT companies and MT labs, and together improving the final products, in the hope that it will do better than Google in Chinese-related language pairs in the very near future.

Sunday, January 17, 2010

Rising Asia and its Implications


Somebody whose opinion I really value, just told me that my blog is about "men fighting" after seeing my first three posts. While I really do care about responsible free speech and will always aggressively defend it; enough about censorship and moderator abuses and let's move on.

I find the notion of "A Rising Asia" interesting,  and I think it presents a major opportunity for the localization and translation industry in the years to come. I was first introduced to this concept many years ago by Nick Kristof and Sheryl WuDunn in their excellent book, Thunder from the East.  (They were also among the first to point out the promise of China/Rising Asia, several years before the "experts" did.)  I really liked the book because it also quickly provided an economic history of the world and is a very easy read.

I thought it would be good to update and expand upon an article I wrote for GALA a little while ago. I have been surprised how little real awareness there is even in the localization industry where leaders often equate Asia with China & Japan (CJK).  This was really brought home at the #LTBKK conference, when I saw what a revelation Biraj Rath's excellent presentation on the Indian localization market opportunity was, to many industry experts. LISA, to their credit announced an India Forum shortly after the conference.


So I thought it might be useful to provide a basic primer on the broader Asian market opportunity. For some in the L10N industry this might all be obvious but here goes anyway. I am interested because:

- There are a lot of people living in Asia (maybe 95% of the next billion Internet users)

- The Internet has very low penetration thus far. Multilingual content will very likely play a key role in driving increasing  penetration and commercial opportunity.

- They will need a lot of information quickly (huge opportunity for automated translation technology)

- Largest concentration of young people in the world (also the Middle East and Brazil) 

 

Asia is extremely diverse, economically and culturally, but yet there are some strong common elements. It is also much less connected than Europe. Today, we are aware of the current economic momentum that India and China have, historically they both also had a deep and lasting cultural influence on much of Asia. An awareness of this history is very useful in developing effective business strategies for different countries. The internet is only just beginning to take root in much of Asia (18% vs. 73% for North America), however, it is expected that almost half of all Internet users will be Asian by 2013. Already, China has more people online than the US. Asia could be a major opportunity for companies that learn to tap into this new emerging online population. But this will require an understanding of the diversity and characteristics of the various segments and will also need new approaches in communication and marketing. Asian economies continue to rise in importance and growth, as both a supplier and consumer. Today China and India are the largest mobile phone markets in the world.

Some interesting and perhaps less known facts about Asia that provide a useful contrast to Europe are shown below. They also give one a sense for the different type of opportunities available and the differing reality of Asia.


-GDP per Capita in Asia (~$15,000) is less than half of the EU average and there is a much wider standard distribution and a large population living in poverty throughout the continent.
-While India and China are among the fastest growing economies in the world, the GDP per Capita is $2,800 for India and $6,000 for China and they should still be considered developing economies.
-The top GDP/Capita countries (2008) in Asia are: Singapore ($52K), HK, Japan, Taiwan, South Korea ($23K), Malaysia, Kazakhstan and Thailand ($8.5K).
-India has 22 official languages that are as distinct and different as the 23 EU languages, and also include at least 6 different scripts. English is only spoken by about 7% of the people in India. However, it is possible to get deep penetration into the Indian market with 5 key languages.
-There is very little local language content for Asian languages on the web in general. Based on a survey done by Asia Online in 2007, less than 15% of the total content on the web is in Asian languages. Almost 90% of the Asian language content is in Chinese and Japanese. There is huge need for more local language content all over SE Asia.
-Mandarin is beginning to edge out English as the preferred 2nd language in Asia
-China is now the fastest growing patent office in the world. The WIPO and others state that China is clearly an emerging scientific and technological power.
-The share of Asian country based patent filings is now in excess of 50% of all patents filed across the world.
-India has more gifted and talented students in high school than the total school student population in the US.
-China has more students in Science and Technology college degree programs than India and the US combined.
-McKinsey has identified a “Rising Asia” as a stable long term trend that will fundamentally change consumption patterns. Gartner suggests using IT to reach the market. They suggest that global companies use IT to ‘lighten’ their Asian business model to address the specific cultural, geographic reach, and supply chain considerations.
-The wealthy Asians are concentrated in major cities like Shanghai, Beijing, Hong Kong, Singapore, Kuala Lumpur, Mumbai, Delhi, Seoul, Manila and Bangkok.
-China is now the fastest growing market for Bentley and BMW.
-More cars are now sold in China than in America.
-Even countries like Laos, Nepal, Pakistan, Sri Lanka, Myanmar, and Cambodia which have very low GDP/Capita are interesting markets for cell phones and basic commodities.
-An understanding of Buddhism, Hinduism and Confucianism cultural perspectives can dramatically enhance your communications strategy into most parts of Asia.
-The fastest growing FaceBook markets in 2H2009 are Taiwan, Indonesia, Philippines and Thailand.
-Google is not dominant in key Asian markets, in Korea they have less than 2% search market share and they are a distant second in China and Japan. Maybe even completely out of China soon. Local companies dominate because of better understanding of local content, language and customer preferences. This suggests that standard US approaches may not work as well in many Asian markets.
-Chinese social networking startups have produced many innovations that have led to them becoming profitable much faster than US equivalents like MySpace and Facebook.  We are now seeing Asian innovation gradually making its way to the west.
-Most of Asia has been relatively unscathed by the global financial and real estate market collapse.
-India is increasingly considered a "soft power". Influential culturally way beyond it's direct sphere of influence.
-The venture capital markets in India and China are rapidly developing with help from "returning" entrepreneurs and hostile US immigration policies.

    But simple strategies like simply making your web content available in the local language may not work. Asian cultures may look superficially similar and even western on the surface, but can have deep cultural differences. The localization market is estimated to be $1.5B in 2010 and could grow dramatically. My sense is that those numbers miss much of the impact of recent growth as the Facebook trends show, mobile computing and successful bottom of the pyramid marketing strategies.


    All of these factors point to fundamental shifts in the global economy and indicate that many of these trends will accelerate further. Asia is a significant opportunity for informed globalization managers -- and probably key for long-term leadership for many global enterprises.

    Global companies need to develop broad and unique country-specific strategies to be able to prosper and thrive in this rapidly changing world. Localization and translation will be key elements of any successful globalization plan and should present significant opportunities to vendors that prepare for this change.

    It's wise to remember that the Chinese ideogram for "change" can also mean "opportunity."

    Thursday, January 14, 2010

    Censorship in the News

    It is interesting that my first blog entry which talked about censorship coincided with the Google news storm in China.

    In a blog entry that rocked the world they said: "We have decided we are no longer willing to continue censoring our results on Google.cn" and "we have evidence to suggest that a primary goal of the attackers was accessing the Gmail accounts of Chinese human rights activists." Apparently they are willing to pull out of China if necessary.

    I was heartened to see this, as I have often felt that Google (and others) really had a policy that was more accurately described , "Don't be evil (except if it's inconvenient)".

    The best coverage of this issue that I have seen comes from Rebecca Mackinnon who is sympathetic, Imagethief who provides some analysis, Techcrunch who is skeptical about Google's real motivation and James Fallows who looks at the big picture political implications. The WSJ suggests that the China issue was a major moral dilemma for Russian founder Sergey Brin who felt strongly about not supporting censorship (unlike somebody else we know). I also found a Chinese perspective at China Youren interesting. Ars Technica (love that name!) suggests that Chinese hackers infiltrated automated systems set up to provide information to law enforcement in the US government. Ahh, the plot thickens.... we are being watched too, huh?

    I REALLY like that I can gather different perspectives to get a "real" sense for what this means. The adult world (where people are allowed to speak) is so cool sometimes, but so messy and muddy isn't it?

    Thus, I have decided to take @localization's advice and just air some of my comments from the LinkedIn censorship firestorm "out there" and make them visible in my blog, outside the Papal Conclave so to speak. There is too much in all to put everything here, so these are just my preferred highlights. So back to our little world.

    Selected comments (mine in small font) from the 80 comment original Linked In discussion:

    I happened by chance to read the original post that Renato made and while I saw that he clearly had a different opinion and viewpoint, I saw nothing offensive in language, style or intent of his deleted posting. I wish I had copied it so that others could see how sober and tempered it actually was.

    It is unfortunate that you (Serge) chose to delete it based on your sole judgment. To my view this is an abuse of privilege and "power". There was nothing in his post that was not already stated in his blog and the original discussion was launched and stated in a very positive way.

    I guess we should all be wary of making a statement that holds opinions different from yours lest we be judged and deleted

    Serges Response: Kirti, thank you for joining in. There's a fine line between strongly opposing views and strongly aggressive offending views. ...........

    Serge Response after several opposing views: What keeps me going here are private comments (one shown below ==) that I am getting (of course people understandbly (sic) are not too willing to get Jussara-messages
    (another villain with an opposing opinion from the discussion ...KV
    ) addressed to themselves):
    ==

    I really do not have time to write a formal comment on this thread to support you. But I think you are doing the right thing. I have been to a few forums where people keep bashing each other, leading to the demise of the group. Just stand firm.

    ==


    My response: It would be wonderful to actually hear from one of your "supporters" as I am not sure how anybody could support your action without actually seeing the content that was deleted, unless they have already decided that anything Renato has to say is irrelevant.

    I actually saw and read the actual posting that was deleted and it was clear to me that a lot of thought and care had gone into writing it.

    Also you seem to forget that your opening comments were less than respectful to an initial very benevolent statement encouraging comments and discussion on how associations can raise money through means other than events. I think the tenor and tone of your first comments pretty much speaks for itself.

    I don't necessarily have to agree with everything he said but I certainly respect the right of a member to state an opinion, especially one that to me seemed to have been done with great care, even if it is different from mine.

    Also, perhaps you overestimate the "support" you have - it certainly does not seem to be strong enough to get somebody to actually step up and say it out aloud in public.

    Censorship always works best when it dark, veiled and hidden. You don't have to justify it then.

    I for one will always be suspicious of deletions in this group from this point on.

    And the basic point he (Renato) made at the outset is still valid: how can we get fewer, higher quality events and more collaboration between the associations to build a better future for every Localization Professional?

    --- and me later again after reading the recreated version of the deleted post:

    Having read the original post, I think the recreation of the original posting on Renato's blog is a very close if not an exact replica of his posting that was deleted, especially in terms of tone, tenor and substance.

    And again, I have to say from any reasonable moderation stance that I can think of, I cannot see a good reason for deleting it simply because he may/may not have said that "XXX has lost its direction or leadership". Since it was deleted we will never even know if this is true.

    Surely, we as professionals can discern and handle this level of comment without getting defensive and resorting to suppression and arbitrary censorship .

    I fundamentally question the judgment made by the moderator that characterized this as negative enough to be suppressed and deleted. In my opinion this was clearly a use of "excessive force".

    The news is filled with anti-Obama (or any current administration) comments constantly and this is (unfortunately) very much a part if not the very essence of democracy - the haters and ugly voices are quickly identified and mostly dismissed by most reasonable people.

    I am glad that there are other forums within LinkedIn and the web where openness and free-speech are seen as less threatening.


    Frank Wang said.. (now the conversation moves to Renato's blog)
    Kirti, We are adults. We all know the politically correct statement kids pick up in middle schools. Rubbing it time and again into the readers' face does not make your argument any stronger. Each organization/group has its own rules and policies, to keep them running the way the organizer sees fit. When you join a group/organization, you accept the terms. Or you can choose not to join or to leave. It's that simple. Don't complicate the issue with politically charged terms which we all know by heart. There is enough negativity/bashing both in the LinkedIn thread and here, from those defending Renato. "USSR", "dictatorship", "dark, veiled and hidden" (from yourself). If those are not bad enough, how about this one: "Following the thread of the dispute there, I wondered at times whether his command of English was really up to the task of moderation and actually understanding a point being made in all but the simplest language." May I paraphrase it into "You moron, you don't know what you are reading or doing"? Has Renato or any of his defenders said anything about this? I would definitely quit a group that allows this attack on a colleague.
    Frank,

    I guess I continue to speak because I did NOT accept or expect that a comment like the one that was deleted could or would be removed without some kind of due process.

    And precisely because we are adults, we do not need to make examples of personally-oriented disrespectful remarks. They speak for themselves, don't they?

    You should watch CSPAN or British parliamentary proceedings sometimes if you think this discussion was not civil.

    Also, you misrepresent my comment which was: "Censorship always works best when it dark, veiled and hidden. You don't have to justify it then."

    In functioning censorship nobody protests because nobody knows.

    In this case the censorship was visible and questionable and that was made it disturbing and worth a little bit of furor.

    Maybe we do need to pick up our middle school textbooks and refresh our minds on why it is important to speak up when you see something that you believe is just plain wrong.

    I hope that the furor will cause some change and raise the level of transparency and accountability in all the groups in LinkedIn. As a community member I reserve the right to be heard, especially when I speak with civility and respect.

    The group does not belong to Serge, it is only what it is and has value, because the community has decided to trust it's integrity and moderation.

    Perhaps I am overly sensitive as I grew up in South Africa under apartheid = institutionalized racism. One thing I learnt very clearly from middle school there: nothing will change if you do not speak up.

    I am glad that we are able to have this discussion and I do appreciate the point you make about how some people do make personal and ethnically based insulting remarks. I agree it is not appropriate.

    But as CSPAN shows, democracy is messy but still worthwhile.

    And back to the GALA group forum in Linked In now:

    Serge to me: Kirti, exactly what change do you want to bring? I am not sure that this is clearly verbalized.

    My response: The change I would like to see is that anybody who writes the kind of posting that Renato wrote BE ALLOWED TO DO SO. There was nothing in his post that warranted a unilateral deletion. If you felt that there was something was offensive you needed to at least talk to him BEFORE you deleted it. I read it initially and will continue to defend his right to say what he did even though I actually don't agree with him.

    As a moderator I delete blatant ads posted as discussions all the time. Once I deleted somebody who posted an ad for Used Cars and also blocked this person from the group. But I always leave any serious MT focused issue alone even if I think it is wrong. The community decides what is interesting and what is not.

    These forums are not the personal playgrounds of the moderators, and the community members can and should hold them accountable in exchange for their trust, involvement and engagement.

    Leadership in the forums is most clearly demonstrated by giving voice to opinions that are different in a fair and equitable way.

    I think an apology would be in order.

    Having said that, I actually feel that GALA produced a better conference than Localization World in terms of content. Miserable in terms of location and attendance. It makes sense to me that they take a greater leadership role in defining the agenda of major conferences (Buyers, Vendors AND Freelancers) and perhaps make money from it too, but I would love to see more accommodation for freelance translators, lower rates for them to attend etc...

    And I agree it that it would be great if there were fewer more collaborative events held. I am still surprised that the associations don't work together more often to get collaboration happening, even at the individual member level. Isn't feudalism over?

    I hope we see more collaboration happening in 2010.
    -----------------------------------------------

    Messy isn't it? And I said that this blog was going to be about MT didn't I?