Jun 13, 2014

A ride in the sun, rough roads and

"So what happened after he climbed up the tower and rescued her? She rescued him right back."

Mobility again at last! With power assistance for cobblestoned hills from Hell.
Got up after an hour and a half sleep, more or less, thought about work and decided instead to explore the bairros a bit before the morning's physical therapy. A sweet ride in the morning sun, past the aqueduct, the hay-harvested fields and ruins to an ATM, where I picked up cash and decided to enjoy pasteis nata and galão at Café Ebora before gymnastics, soothing fingers and hot packs on the knotted web of my back.


Alas, my Portuguese failed me for a simple order of coffee and pastry, where it worked so well for many hours of language private tuition, long chats to end the day and greet the morning sun, bake cookies, drink sangria, whiskey, learn the secrets of sopa de beldroegas and the right brew of bachalhau and spinach in my only pot, discussing facas, tendeiros and their murderous habits, the quality of local schools, the revolution to be finished, coca, heroin, speedballs and methadone, primos, sobrinhas, filhos, putas, baratas, children buying time for their mother to jump out the second story window and escape another beating, and when a knife in the leg can teach a husband to behave before he needs his throat cut. Still, as I sipped coffee (once the misunderstanding was resolved), I reflected on how, after nearly 15 years of living in cold places with too many cold people, Portugal, its bairros and people, their ability to get on with life and live it, their persistent search for another vein in which to inject some hope after all the blood has been drawn by German banks, have restored my will to live.


Jun 4, 2014

OmegaT’s Growing Place in the Language Services Industry

Guest post by John Moran

As both a translator and a software developer, I have much respect for the sophistication of the well-known proprietary standalone CAT tools like memoQ, Trados, DejaVu and Wordfast. I started with Trados 2.0 and have seen it evolve over the years. To greater and lesser extents these software publishers do a reasonable job at remaining interoperable and innovating on behalf of their main customers - us translators. Kudos in particular to Kilgray for using interoperability standards to topple the once mighty Trados from its monopolistic throne and forcing SDL to improve their famously shoddy customer support. Rotten tomatoes to Across for being a non-interoperable island and having a CAT tool that is unpopular with most (but curiously not all) of the freelance translators I work with in Transpiral.

But this piece is about OmegaT. Unlike some of the other participants in the OmegaT project, I became involved with OmegaT for purely selfish reasons. I am currently in the hopefully final stage of a Ph.D. in computer science with an Irish research institute called the Centre for Next Generation Localisation (www.cngl.ie). I wanted to gather activity data from translators working in a CAT tool for my research in a manner similar to a translation process research tool called TransLog. My first thought was to do this in Trados as that was the tool I knew best as a translator but Trados’ Application Programming Interface did not let me communicate with the editor.

Thus, I was forced to look for an open-source CAT tool. After looking at a few alternatives like the excellent Virtaal editor and a really buggy Japanese one called Benten I decided on OmegaT. 

Aside from the fact that it was programmed in Java, a language I have worked with for about ten years as a freelancer programmer, it had most of the features I was used to working with in Trados.  I felt it must be reliable if translators are downloading it 4000 times every month. That was in 2010. Four years later that number is about to reach 10,000. Even if most of those downloads are updates, it should be a worrying trend for the proprietary CAT tools. Considering SDL report having 135,000 paid Trados licenses in total - that is a significant number.

Having downloaded the code, I added a logging feature to it called instrumentation (the “i” in iOmegaT) and programmed a small replayer prototype. Imagine pressing a record button in Trados and later replaying the mechanical act of crafting the translation as a video, character-by-character or segment-by-segment, and you will get the picture. So far we use the XML it generates mainly to measure the impact of machine translation on translation speed relative to not having MT. Funnily enough, when I developed it I assumed it would show me that MT was bunk. I was wrong. It can aid productivity, and my bias was caused by the fact that I had never worked with useful trained MT. My dreams of standing ovations at translator association meetings turned to dust.

If I can’t beat MT I might as well join it. About a year and a half ago, using a government research commercialization feasibility grant, I was joined by my friend Christian Saam on the iOmegaT project. We studied computational linguistics in Ireland and Germany on opposite sides of an Erasmus exchange programme, so we share a deep interest in language technology and a common vocabulary. We set about turning the software I developed in collaboration with Welocalize into a commercial data analysis application for large companies that use MT to reduce their translation costs.

However, MT post-editing is just one use case. We hope to be able to use the same technique to measure the impact of predictive typing and Automatic Speech Recognition on translators. I believe these technologies are more interesting to most translators as they impose less on word order.

At this point I should point out that CNGL is a really big research project with over 150 paid  researchers in areas like speech and language technology. Localization is big business in Ireland. My idea is to funnel less commercially sensitive translator user activity data securely, legally, transparently and, in most cases anonymously from translators using instrumented CAT tools into a research environment to develop and, most importantly, test algorithms to help improve translation productivity. Someone once called it telemetry for offline CAT tools. My hope is that though translation companies take NDAs very seriously, it is also a fact that many modern content types like User Generated Content and technical support responses appear on websites almost as soon as they are written in the source language, so a controlled but automated data flow may be feasible. In the future it may also be possible to test algorithms for technologies like predictive typing without uploading any linguistic data from a working translator’s PC. Our bet is that researchers are data-tropic. If we build it they will come.

We have good cause to be optimistic. Welocalize, our industrial partner, is an enlightened kind of large translation company. They have a tendency to want to break down the walls of walled gardens. Many companies don’t trust anything that is free, but they know the dynamics of open-source. They had developed a complex but powerful open-source translation memory system called GlobalSight, and its timing was precipitous.

It was released around the same time SDL announced they were mothballing their newly acquired Idiom WorldServer systemtheir system to replace it with the newly acquired Idiom WorldServer (now SDL WorldServer). This panicked a number of corporate translation buyers, who suddenly realized how deeply networked their translation department was via its web services and how strategically important the SDL TMS system was. As the song goes, "you don’t know what you’ve got till its gone" – or, in this case, nearly gone.

SDL ultimately reversed the decision to mothball TMS WorldServer and began to reinvest in its development, but that came too late for many some corporates who migrated en-masse to GlobalSight. It is now one of the most implemented translation management systems in the world in technology companies and Fortune 500’s. A lot of people think open-source is for hippies, but for large companies open-source can be an easy sell. They can afford engineering support, department managers won’t be caught with their pants down if the company doing the development ceases to exist, and most importantly their reliance on SDL’s famously expensive professional services division is reduced to zero. If they need a new web-service, they can program it themselves. GlobalSight is now used in many companies who are both customers of Welocalize and companies like Intel who are not. Across should pay heed. At a C-Suite level corporates don’t like risk.

However, GlobalSight had a weakness. Unlike Idiom WorldServer it didn’t have its own free CAT tool. Translators had a choice of download formats and could use Trados but Trados licenses are expensive and many translators are slow to upgrade. Smart big companies like to have as much technical control of their supply-chain as possible so Welocalize were on the lookout for a good open-source CAT tool. OpenTM2 was a runner for a while but it proved unsuitable. In 2012 they began an integration effort to make OmegaT compatible with GlobalSight. When I worked with Welocalize as an intern I saw wireframes for an XLIFF editor on the wall but work had not yet started. Armed with data from our productivity tests and Didier Briel, the OmegaT project manager, who was in Dublin to give a talk on OmegaT, I made the case for integrating OmegaT with GlobalSight. It was a lucky guess. Two years later it works smoothly and both applications benefit from each other.

What did I have to gain from this? Data.

So why this blog? Next week I plan to present our instrumentation work at the LocWorld tradeshow and I want Kilgray to pay heed. OmegaT is a threat to their memoQ Translator Pro sales and that threat is not going to reduce with time. Christian and I have implemented a sexy prototype of a two-column working grid, and we can do the same trick importing SDL packages with OmegaT as they do with memoQ. Other large LSPs are beginning to take note of OmegaT and GlobalSight.

However, I am a fan of memoQ, and even though the poison pill has been watered down to homeopathic levels, I also like Kilgray’s style. The translator community has nothing to gain if a developer of a good CAT tool suffers poor sales. This reduces manpower for new and innovative features. Segment-level A/B testing using time data is a neat trick. The recent editing time feature is a step in the right direction, but it could be so much better. The problem is that CAT tools waste inordinate amounts of translator time, and the recent trend towards CAT tools connected to servers makes that even worse. Slow servers that are based on request-response protocols instead of synchronization protocols, slow fuzzy matches, bad MT, bad predictive typing suggestions, hours wasted fixing automatic QA to catch a few double spaces. These are the problems I want to see fixed using instrumentation and independent reporting.

So here is my point in the second person singular. Kilgray – I know you read this blog. Listen! Implement instrumentation and support it as a standard. You can use the web platform Language Terminal to report on the data or do it in memoQ directly. On our side, we plan to implement an offline application and web-application that lets translators analyse that data by manually importing it so they can see exactly how much they earn per hour for each client in any CAT tools that implement that standard. €10 says Trados will be last. A wise man once said you get the behavior you incentivize, and the per-word pricing model incentivizes agencies to not give a damn about how much a translator earns per hour. The important thing is to keep the choice about sharing translation speed data with the translator but let them share it with clients if they want to.  Web-based CAT tools don’t give them that choice, so play to your strengths. Instrumentation is a powerful form of telemetry and software QA.

So to summarize: OmegaT’s place in the language services industry is to keep proprietary CAT tool publishers on their toes!


*******


See also the CNGL interview with Mr. Moran....

May 18, 2014

memoQ 2014: a first look

I couldn't make it to memoQfest this year - the first one in Budapest that I have missed since the event began in 2009. But the first family visit since that same year took priority, so my exposure to the upcoming memoQ 2014 version was strictly second hand until today.

I wasn't too happy with thing I heard on the Yahoogroups user list. In fact, when I read one message describing how the new transcription feature for bitmap graphics in some files required the Product Manager version, I was quite annoyed. The reality - a whole month before the official release - is very good for both freelancers and corporate outsourcers, and I think by the time this version makes its official debut in June there will be many good reasons to smile. I'm frankly amazed at how much Kilgray seems to be getting its act together and balancing the needs of users at all levels.

This afternoon I downloaded the first test release (alpha??) of memoQ 2014, installed it and began to take a cautious tour. My first impression was that it looked the same. And then, bit by bit, subtle and excellent small differences began to emerge. I looked for and found major new features I had heard about and discovered many interesting things not mentioned along the way.

The grammar checking feature seems to be implemented in a sensible way, though it actually doesn't work at all right now for me. But I can see where it's headed, and it is going in a good direction.

I had a quick look at the new plug-ins, particularly TaaS, and made notes about testing the potential for teamwork. What I have seen of TaaS for its much-advertised terminology extraction is a huge disappointment, and those who have followed my comments on Twitter will know I have nothing good to say about this EU boondoggle, but I see potential for other possibilities that nobody has really talked about, and if my instinct is right, this could be really useful. But I will need to invest a lot of testing time for the approach I have in mind.

The Project home view has gotten even more impossibly cluttered with the addition of "People", a rather sensible reworking of role assignments that even in the Translator Pro version clearly acknowledges that most freelance translators are not, in fact, 'islands' in their work.


This will surely make the small screen (netbook) usage problems worse if Kilgray does not redesign the view a bit, but in every other respect I see this as a significant improvement of project workflow, emphasizing the relationships between project participants in a better way.

One little bit that I stumbled across was the new way of handling the export of unfinished translations. This is a nice way of recognizing the frequent pressure in some projects to export incomplete stages of work.


I have had ways of dealing with this need for years in memoQ, but this new approach will make things simpler and obvious for all users.

There is a nice little feature for tracking time too:


This will facilitate record keeping for some jobs involving time charges.

The feature I have looked at in some depth so far, which makes me very happy, is Kilgray's very sophisticated handling of embedded objects and graphics, which sets new standards in many ways. I think there is still a key feature missing to make it the equal of OmegaT for handling charts with data stored as XML in the MS Office file (though I have not had time to check this yet), but what I have seen so far goes way beyond similar features I have seen in STAR Transit and Déjà Vu X2.


Embedded objects and images are imported as separate files from within the media and embeddings folders of the Microsoft Office file. I see a few potential problems with the current way of displaying a file and its objects and media. I've had projects with multiple files having embedded Excel spreadsheets, PowerPoint slides and other objects as well as any number of pictures needing to be localized. One recent project had 59 spreadsheets embedded in a DOCX file. Without an accordion or tree structure to collapse the subordinate structure view and show the embedded content again, the overview will be lost quickly. But this is a very good start. Note how the main file includes a count of the segments in the subordinate objects and graphics. (And take note of the new progress bar with different colors for different process stages like translation and proofreading.)

Bitmap texts can be recorded with a new transcription feature, which is also compatible with voice recognition. I dictated my German source texts with Dragon Naturally Speaking set to German, then switched to English for the translation. And of course the bitmap transcriptions are included in the word counts of the Statistics functions and the translations are written to the translation memory. I believe this is utterly unique in translation environment tools. Fluency has a transcription module too, of course, but its purpose and application are very different.

The exported translations with translated objects will look like they are not done at present, because the difficult refresh problem has not been solved by Kilgray. Each translated spreadsheet, slide, etc. will need to be opened in the document before the translation will become visible. This is much easier using the macro I published two years ago, and I am certain that by release time or soon thereafter Kilgray will find an elegant way of dealing with this difficulty. Atril handles the same problem by distributing macros as I recall.

In the recent Kilgray blog post on the six reasons to upgrade to memoQ 2014, the only overlap with the above points is the image localization. Peter Reynolds talks instead about other good stuff, such as the long-awaited project templates and Language Terminal. There are so many nice things ahead with this upgrade that we'll all just have to take it slowly, one bit at a time.

Of course the usual precautions for any new software version apply. The new version can be installed in parallel to your current version, and it can be tested while you continue to do the bulk of your work in the older, stable version. Typically it takes a few months for any new version to get the kinks out, but this allows plenty of time for planning the transition and preparing to take full advantage of the new features relevant to you. Migration is also not a trivial matter in many cases, but this time around there may be a little more help with that. More on that another time!