Showing posts with label Wordfast Pro. Show all posts
Showing posts with label Wordfast Pro. Show all posts

Jan 28, 2020

Another look at Windows 10 speech recognition


A few years ago while on "holiday", I returned from dinner to find that my laptop had bluescreened. Panic time! It was Saturday night, and I still had quite a lot of text to translate and deliver on Monday morning. And up on the highest mountain in Portugal, I wasn't sure where I could find a replacement to finish the project, which was, at least, not utterly lost, because I had put it on a memoQ Cloud server for testing. The next day I got lucky: about 50 km away there was a Worten, where I picked up a gamer laptop with lots of RAM and an SSD. Well, not so lucky, as it was a Hewlett Packard Omen, with a fan prone to failure, but that's another story....

This new laptop was my first encounter with Windows 10. I had heard that this operating system offered improved speech recognition capabilities, and since I prefer to dictate my translations and downloading the 3 GB installation file for Dragon NaturallySpeaking (DNS) from my server at the office was going to take forever, I thought I would give Windows 10 speech recognition a try. I hadn't installed my CAT tool of choice yet, so I fired up Microsoft Word and began dictating. "Not bad," I thought. Then I tried it in my translation environment, and the results were a complete disaster. So I put that mess out of my mind.

Since then there have been some notable advances in speech-to-text capabilities on a number of platforms. But the best solution for my languages (German and English) with DNS became increasingly cranky thanks to neglect of the product by Nuance. Every week I read new reports of trouble with DNS in a variety of environments in which it used to perform very well. Apple's iOS 13 was a great leap forward of sorts for speech recognition and voice-controlled editing, but the new features are only available in English, and having Voice Control activated totally screws up my otherwise rather good dictation in German and Portuguese (or any other language). And don't get me started on the crappy vocabulary addition feature, which uses text entry alone with no link to actual pronunciation. Good luck with that garbage. It's not a bad solution in Hey memoQ with the additional command features added, but iOS dictation is not completely up to reasonable professional standards yet.

I probably would have given no further thought to Windows 10's speech-to-text features if it weren't for Anthony Rudd. We've corresponded a bit since I bought his excellent book on regular expressions for translators (and there's another practical guide for us coming soon from him!), and in a recent discussion he alluded to the use of Unicode with regex as a simple way of dealing with some things another colleague was struggling with. I was intrigued by this, and so for about half a day, I ran down a rabbit hole, testing Unicode subscripts and superscripts for a variety of purposes like fixing bad OCR of footnote markers and empirical formulae, autocorrecting common expressions for subscripted variables and chemical terms, including subscripts and superscripts in term bases and much more. Fascinating and useful stuff on the whole, even if some fonts don't support it well.

And of course I looked at using these special Unicode characters in speech-to-text applications. DNS had some funky quirks (not allowing numbers in the "spoken" version of terms, for example), but it worked rather well, so I can now say "calcium nitrate formula" and get Ca(NO₃)₂ without much ado. And for some reason it occurred to me to give Windows 10 speech recognition a try, just because I was curious whether vocabulary could in fact be trained. Indeed it can, and that feature is better than iOS 13 or DNS by far.

But first I had to remember how to activate speech recognition for Windows on my laptop again. When in doubt, type what you're looking for in the search box....

Notice I've pinned Windows Speech Recognition to my taskbar on the right, which is good for quick tasks.

Gesucht, gefunden
. Unlike other speech recognition solutions, the one in Windows 10 works only for the language set for the operating system. And options there are limited to English (United States, United Kingdom, Canada, India, and Australia), French, German, Japanese, Mandarin (Chinese Simplified and Chinese Traditional) and Spanish.

I put on my trusty Plantronics earset (the best microphone I've used for dictation tasks or audio in my occasional webinars in the past year) and began to dictate, first in Microsoft Word, which had shown acceptable results in my tests long ago. I found that adding vocabulary in the Speech Dictionary (accessed via the context menu in the dictation control element shown as a graphic at the top of this post) was dead simple.

The option to record pronunciation enabled me to record non-English names and words in several languages. And sure enough, the Unicode subscripts and superscripts worked, so I can now say CO₂ (I just dictated that) to my heart's content.

I was expecting a mess when I tried to use Windows 10 speech-to-text in a CAT tool, but it was not to be. It was brilliant, actually. I tried it in my copy of SDL Trados Studio, and with the scratchpad disabled so I could dictate directly into the target it worked well. No voice-controlled editing like I'm used to with DNS in memoQ, but that DNS feature does not work in SDL Trados Studio anyway, so this is no worse. But with the scratchpad box enabled (see the screenshot below), I could use voice commands to select and correct text or perform other operations. Brilliant!

After clicking or speaking "Insert", the text will be written to the target field with the proper formatting
So users of SDL Trados Studio who translate to a target language supported by Windows 10 speech recognition are probably better off not giving their money to Nuance, which I'm told can't even be bothered to make a 64-bit version of DNS now (which probably accounts for a lot of the trouble people have with that program.

I tested Wordfast Pro 5, which seems to confuse the speech recognition tool horribly, with source text displayed in the floating bar for some odd reason. But my earlier tests of Wordfast with DNS were equally unhappy, so somehow I'm not surprised. And I didn't test the Memsource desktop editor, which took the price a few years ago for the worst-ever DNS dictation results with a CAT tool. I'll leave that to someone with a much wider masochistic streak.

But what about memoQ, my personal environment of choice for most translation work? Equally brilliant, works just the same as SDL Trados Studio. No voice control for editing without the dictation scratchpad enabled (there, DNS has an advantage in memoQ), but with the scratchpad you can use the voice commands to edit before inserting in the target text field.


Wanna see this in action? Have a look at this short demo video:


I hope that the future will bring us more language support for Windows 10 dictation (Portuguese, Russian and Arabic, please!) and that other providers (like Google, if you're listening, and Apple, which never listens to anyone anymore except to spy on them with Siri) will expand the speech-to-text features offered, particularly to include sound-linked vocabulary training and better adaptation to individual users' speech. Five years ago when I began to investigate alternatives for non-DNS languages, I expected we would have more by now, and we do, but professional needs require all providers to raise their game.

Addendum: Someone asked me if Windows Speech Recognition is a cloud resource or a locally installed one which will work without an Internet connection. It's definitely the latter. So if you have lousy bandwidth or find yourself disconnected from the Internet, you can still use speech-to-text features.

And more: I use a lot of spoken commands for keyboard shortcuts when I work, so I did a little research and testing. It seems that Windows 10 speech recognition gives full access to an application's keyboard shortcuts via voice. So in memoQ, for example, I can dictate the insertion of tags, items from the Translation Results pane and a lot more. Watch out, Nuance. Windows 10 is going to kick your Dragon's scaly butt!

Jan 13, 2019

A second look at Wordfast Pro


The generally good impression made by Wordfast Anywhere in my recent tests inspired me to take a new look at the premium environment for freelance translators: Wordfast Pro 5. A lot has changed with Wordfast Pro since its early days, and much of what I found troublesome with early versions has been corrected. A new look has also been on my agenda for a while since I realized that two new formats were introduced (TXLF, an XLIFF format, and GLP, a zipped project package format for Wordfast), which can be handled by my usual translation environment but (currently) with a few extra steps required compared to the old TXML format.

The installation took about a minute and started off with a good impression from the warning about cloud drive synchronization:


I've seen a number of people come to grief with other tools when their projects, translation memories or other resources are stored in Dropbox or similar configurations so they can be shared by installations on different computers, and I appreciate Wordfast's attempt to warn people off from this dodgy practice. If you want to share resources, play it safe and stick them in Wordfast Anywhere.

At first, the program is in demo mode, which limits translation memories to 500 translation units (TUs) and does not allow access to remote resources such as Wordfast Anywhere. Fortunately, there is a fully functional 30-day trial available, and it took all of about two minutes to fill out the simple request form, receive the mail with the trial license key and activate it in Wordfast Pro 5.

I was really enthusiastic about the clean, uncluttered feel of the interface. There's a lot more functionality in SDL Trados Studio or memoQ, but all the myriad features of those environments can be intimidating to some, and even for experienced users navigation can be confusing at times to locate some obscure setting or feature. Not in Wordfast Pro 5: the features mostly aren't there, and what is there can be found without much ado. Given the limited scope of mastery and inclination to learn on the part of many hamsters running on the freelance translation wheel, this can be a definite advantage.

On the Help ribbon I saw a Feedback icon. I don't know why, but this inspired a weird enthusiasm in me, so I clicked it, and when the dialog appeared, I wrote a quick note to the development team to say what a great impression the new user interface was making before I had even started to do anything useful with it. I noticed that the feedback dialog also had options to include files and projects in case of a problem, which I also thought was really cool. Something like that in other tools would be very helpful to their users and probably encourage more suggestions and interaction.

It was really easy to navigate through the ribbon menus and explore the configuration options. I was pleased to see that different sets of keyboard shortcuts were available to make the ergonomics easier for users of some other tools.


But SDLX? Huh? That's kind of Jurassic. No memoQ shortcuts, but no problem. I can customize, right? Yes... but I soon discovered that I apparently had no way to save my customized shortcuts as "memoQ style" or whatever else I might want to call them. And then I noticed that I probably can't save the configuration to move it onto a second computer where the terms of the license agreement allow private individuals to install another copy. And, hmmmm, no option to print a cheat sheet I can refer to as I learn the keyboard shortcuts. memoQ users are kind of spoiled on both counts, I guess.

One thing I was very eager to try was the connection to my translation memories and glossaries in my Wordfast Anywhere account. That proved to be quite straightforward: it worked exactly as the clear instructions of the Wordfast Pro Help described the process.

So I was ready to try out some translation, maybe a little dictation with Dragon NaturallySpeaking. I imported a little text file to get started:


WTF??? Now I know what the problem is here, but importing the same file to Wordfast Anywhere gives this result:


And in memoQ:


The import with the simple text filter of Wordfast Pro 5 (version 5.7) does not map the characters correctly. I had to change the source text file from ANSI to UTF-8: not a big deal for me, but a lot of translators I know will be over their heads right there.

The choice of import filters available is fairly good as one might expect from most professional translation environments these days, but two important things were missing for me. There seems to be no option to cascade filters, useful for example if you have a Microsoft Excel file containing HTML text to translate, and there is also no facility for configuring custom regex-based text filters or tagging text content which needs protection (such as placeholder text). This won't be an issue for a lot of translators, but for those who deal with challenging, often unexpected formatting issues in customers' files it could be a real pain in the neck.

On to dictation... Dragon NaturallySpeaking (DNS) seemed to perform well. I had to turn off the DNS dictation box by unmarking thew checkbox in its dialog. Text was then transcribed well into the target field, and my spoken keyboard shortcut to confirm a segment and go on to the next one worked perfectly. Then I misspoke and used a spoken editing command to correct my error. Nothing happened. I tried several different spoken selection and editing commands that I use every day in memoQ. Nothing worked. Shit. What we have here is a failure of compatibility. The full potential of Dragon NaturallySpeaking cannot be used in Wordfast Pro 5.

I explored the settings further... quality assurance. That looked pretty good; the options were easy to understand and I could set them as I wanted to check my work. But the QA settings I need vary in many projects, and sometimes I want to do a QA check on just one aspect like tags or maybe terminology. Wordfast Pro 5 offered no facility to save a QA configuration or profile and load it as one might do in SDL Trados Studio or memoQ. This too would be a deal-breaker for me, alas. I depend on a full hand of memoQ quality assurance profiles for selective checking of important quality parameters in my jobs. Toggling settings back and forth in Wordfast would drive me nuts. Still, this wouldn't disturb many CAT tool users who can barely be bothered to run a spelling check on their work, much less run a check or missing or mismatched tags.

In contrast to my conclusions years ago, I can now say that Wordfast Pro is "ready for prime time". It has a nice, clean, easy to navigate interface, and the Help descriptions are clear, if somewhat idiosyncratic in their spelling at times. The options are limited compared to other professional tools I use which have comparable costs of use, but that may be perceived as an advantage by many... until they need what's not there, which is probably inevitable if they work at translation in a full-time freelance capacity. Over the years I have heard many good things about Wordfast support, so I expect that users will at least find help and advice when they need it.

The integration with the online Wordfast Anywhere resources is also simple and good. That's a major point in favor of this tool and should be very helpful for collaboration.

Overall, I think that users who invest in a Wordfast Pro license will get their money's worth. A three-year license costs €400, with three-year renewals costing half the list price after that. If you aren't willing to pay after the three years, your license will stop working (unlike SDL Trados or memoQ, where the current license models allow you to keep working with the software long after your claim to support and upgrades has lapsed - basically "forever" if nothing strange happens with newer operating systems).

The possibilities for collaboration between Wordfast users and those who work with other environments are much better than they used to be, and in just a short time I was able to see how I can prepare projects for a colleague using Wordfast Pro 5. (SDL Trados packages can apparently be handled, though that's not the case for memoQ project packages prepared with the PM Edition - I would have to make MQXLIFF files and export TM and term base resources.) And I hope that this situation will only get better, with more environments offering various kinds of Wordfast resource integration and Wordfast acquiring new capacities to work with other formats and resources.

Jan 12, 2019

Another look at Wordfast Anywhere

The Wordfast suite of applications has a long history, and through much of it I've had my eye on the tools but up to now never really found them up to the demands of my work. Wordfast Classic (back when it was the only Wordfast app) was brought to my attention by an enthusiastic manager of a German bank's translation team more than 15 years ago; he found that the "blacklist" feature for terminology (since adopted by others - for example in memoQ's "forbidden" terms) was extremely helpful to his translators in avoiding terms which might provoke branding controversies or which were simply inappropriate in a particular specialist context.

When Wordfast Pro came along, I was disappointed in the interoperability of its early versions and it being late to the party for supporting XLIFF formats (as were some other popular tools). That issue is solved in the meantime, so I suspect I might not be quite so unhappy were I to revisit the application.

But really, Wordfast doesn't come onto my radar very often, and when it does, it's not so much the application suite itself as it is the Wordfast creator - Yves Champollion, who follows in a way the family tradition of the famous French Egyptologist, Jean-François Champollion, translator of the Rosetta Stone, and who has earned his own fair share of praise for his many years of support for individual translators and their professional organizations. It would not surprise me if much of the loyalty I find among users of Wordfast is inspired by the personal qualities of Yves as much as by any technical features of his tools.

The least among these tools was, in my consideration, the web-based Wordfast Anywhere (WFA). I looked at it briefly in the early days and was unimpressed: too limited, I thought. And the idea of translating in a browser seemed dubious to me, and it remains so in many scenarios that are relevant to my work. WFA was a bit ahead of its time, before the scamming Gold Rush that targeted corporate clients for web-based solutions designed to wrest data and control away from translators. WFA wasn't welcome in that party: its focus on empowering individual translators is anathema to most of the web CAT solutions ones sees today.

My interest in Wordfast generally was revived recently when I saw that memoQ has integration plug-ins for Wordfast term bases and translation memories on servers. This inspired the thought that perhaps Wordfast Anywhere might function as a collaboration server here, sort of like some had hoped for the Language Terminal resources, but one that actually works perhaps. Alas no, or not yet at least; the memoQ plug-in cannot "see" the WFA server and an individual account. Oh, but if it could....

Collaboration and interoperability between translation environments have been topics of great interest for me since I began to use specialist tools for organizing translation resources some 19 years ago. And on those occasions when I want to share resources with someone who does not have a professional suite of desktop translation resources, I'm always a little uncomfortable with my default recommendations, because they are just a little too nerdy to work well with everyone. So I wondered... how well might WFA work with resources I prepare in SDL Trados Studio or memoQ and pass on to a colleague unequipped with those tools or other desktop solutions. I thought I remembered limits that would restrict such an effort, but either my memory is wrong or these limits changed.

WFA can accept files to translate which are up to 20 MB in size. I receive files that are sometimes larger than this, but not routinely, so this is not much of a restriction. But then I thought the limit on translation memory size would be the stumbling block, and indeed, when I tried to upload a 390 MB TM with about 330,000 translation units, I got an error message telling me that 300 MB (or rather 300000000 with no indication of units!) was the limit. Looking in the online documentation I found that 100,000 TUs is the limit for an individual translation memory in WFA. But you can attach multiple TMs and term bases (which can be much larger as I saw from the 800,000+ entry IATE termbase supplied by the environment). And most TMs that I see for mid-size companies are well under that size limit.

So I spent some time kicking the virtual tires again. Uploaded some damned big EU directives in various formats, including bilingual alignments in an XLIFF. No problem. Loaded a big memoQ XLIFF file: the *.mqxliff extension wasn't recognized, but I fixed that the usual way by changing it to *.xlf and it worked well, roundtripping perfectly back to memoQ and confirming that interoperability would work well enough for collaboration.

Indeed, the range of original file formats handled by this free online translation environment is impressive.

As I browsed through the options and customizing features of the WFA environment, my respect for its capabilities increased further. The thought occurred to me at one point that this might even be suited as an environment for a small company with limited translation needs to manage its language resources and make them available for in-house or external translators. With the several exchange formats available, translators and reviewers could easily perform their work with other translation environment tools or even word processors, and the results could be merged with the master records in the WFA account. This is probably the least expensive, secure way for a company to take its first steps toward central management of its translations and terminology resources. No big server investments needed, and later all resources can be migrated easily to more sophisticated environments, such as a memoQ Server, if necessary.

Some years ago, I opposed the use of Wordfast Anywhere in a local university program, arguing instead that more established professional tools like SDL Trados Studio and/or memoQ should be used instead, especially as the cost of doing so is negligible in teaching curricula. I take that back now. And my impression is that WFA is better suited to a teaching program than other, perhaps slicker web-based tools, because of the underlying philosophy of its design, which leaves translators and their partners in control of the data, not some third-party provider inclined to carry out dubious data mining and use the results to sell more dodgy commercial solutions.

Wordfast users also know that their desktop software can access translation memories and term bases on a WFA account as remote resources. My last look at Wordfast Pro showed me that the tool had come a long, long way since I last dealt with it to clean up some messes a French translator inflicted on an agency client of mine. It's been on my list to look at further for some time; I know it will likely not meet my criteria for the broad range of translation, quality assurance and consulting tasks I do, but it does do a good job of covering the real, practical needs of many colleagues, and it is important to me to understand other translation environments to facilitate collaboration with people who use them.

And for these cases of working together with a mix of environments, it seems to me that Wordfast Anywhere can be a productive bridge to bring partners together. To create a free account and start testing Wordfast Anywhere, click here.

Jun 14, 2018

Translating Wordfast GLP packages... elsewhere.


One reason to keep  translation environment tool licenses up to date is that new formats continue to appear. New formats for translatable files as well as new file formats for the tools that help to process files for translation. Very often I have heard some "professional" say "I'm a translator, not a [fill in the blank]. If the client wants this translated, I'll have to get it in a Microsoft Word file." Or something like that.

Let's get real for a moment.

  • That attitude is simply lazy and disrespectful toward translation consumers who would like to make use of one's services and
  • a lot of money is being left on the table here in many cases. I built a huge clientele at the start of the last decade, because my use of translation environment tools like Trados, Déja Vu, STAR Transit and Wordfast enabled me as an individual to tackle translation challenges that many agencies at the time had no concept of how to cope with.
As translation agencies have acquired more technical tools, most of them still remain unfortunately unaware of how to use them properly or plan more than the simplest workflows well, but that's a subject for another day. Also...
  • ... by using tools and techniques that are compatible with what your clients require for a final format, you can save your client a lot of time and money for further layout work - and probably avoid the introduction of errors in your translation work in its final format as well.
  • And in my experience, showing technical and process competence to benefit clients usually leads to greater trust and better work together.
So what has all this got to do with Wordfast?

Well... I didn't like the Wordfast brand for a very long time. Its various incarnations were perhaps the weakest of the popular tools in a technical sense, and inevitably when agency friends called me, desperate to fix some massive translator screw-up (usually by somebody in France), Wordfast "Pro" was often involved in the disaster.

I looked at the "newer" Wordfast versions a number of times over the years, and honestly they always seemed like lobotomized wannabe tools. This was about the time that many other toolmakers were trying to decide if they should support XLIFF.

Well, a lot has changed since then. I became aware of the changes the other day when somebody posted a question in a social media forum for memoQ asking how to handle Wordfast Pro 5 GLP packages. I had never heard of these, so of course I was curious and decided to take a look. This finally led me to download a 30-day trial of the latest Wordfast Pro software to evaluate its potential for interoperable work with other translation environments. I see a lot of changes since my last look, and so far I think they are all positive, and along the way I had good cause to look at Wordfast Anywhere, the free web-based CAT tool that I talked some university colleagues into not wasting their time with a while ago. Well, my recommendation in that regard might change, but that and commentary on the latest incarnation of WF Pro will have to wait for another day.

About those GLP packages....


Yes, those. This was the question:


Someone pointed out that GLP files - like every other translation "package" one finds from all the tool providers - are merely ZIP files with particular structure inside and the extension re-named. 


Gotta love Facebook. You'll always get an answer in some group, usually a wrong one. That's why I keep a blog. Good information gets buried in social media noise too often, and good luck finding it in any kind of search. In this case... we don' have no steenkeen TXML files as I learned... that's the old Wordfast Pro....

A colleague in Germany kindly provided me with a little GLP package to examine, which I promptly unzipped. I noticed that at least one tool (7-Zip) sees through the renamed extension nonsense and saved me the usual trouble of renaming it before unpacking.


So far, so good... inside the folder for the unpacked GLP file I found the following:


The test package was an English to Portuguese project. But source? Hello? Let's have a look there!


Very interesting. The original source files (English) came along for the ride. This is good, because I often like to translate source files in memoQ - taking advantage of the preview there for many file types - and then use the translation memory to translate the file that is created by other other tool (usually SDL Trados SDLXLIFF files in my work). Now let's have a look inside the pt target folder. There's actually another folder named txlf inside that one. And there I found:


No TXML files! TXLF is a new instance of the rather ubiquitous XLIFF files one finds in the translation world, some of which have some rather bothersome "extensions" that may require special handling in the translation process. In the simple test I performed, none of that was apparent; an ordinary XLIFF filter seemed to work well. Future tests will show me if there are any quirks I hope, but so far, so good.

So one strategy, with pretty much any CAT tool, would be to unpack the GLP file, get at those TXLF files and then bring them into another working environment using an XLIFF filter. Maybe also use my approach with the source files too, which will ensure that you can deliver a good target file even if quirky tags in the XLIFF lead you to produce less than an optimal result there. 


The current version of memoQ (8.4) does not recognize the TXLF extension, so as in all such cases, the All files option must be used and the correct filter applied in a later dialog. Unlike with some other tools, memoQ cannot be "trained" by the user to recognize new extensions as far as I know.

But what about importing the GLP files directly to memoQ? Wouldn't that be nice? And I thought it might be possible using the ZIP file filter recently introduced (and the same All files trick to get the GLP file and apply the ZIP filter later). Well...


It looked promising.


So much so that I even optimistically named and saved a custom configuration for the ZIP filter. All I need to do now is cascade an XLIFF filter!


Ack. Sooooo close. I've been here before. There are more things in heaven and down-to-earth cascading formats, Kilgray, than are dreamt of in your philosophy! Please, please expand the list of possible cascaded formats sensibly to make better use of this lovely new ZIP filter!

So for now, that's a no-go, but soon? Who knows? If you bother support@kilgray.com and tell the memoQ team how helpful it would be, maybe this and similar problems can be solved with relative ease.

In any case, for now it seems that the unpack-and-do-the-XLIFF approach will work for most anyone with a modern CAT tool. And that's good news, because in today's fast-changing technology environment for translation, interoperability of CAT tools is increasingly important. It is a foolish waste of time to translate in a large number of CAT tools and probably a bad idea to do so in two or three according to my old research. I've usually found that such JOATs are, professionally, often stupid goats who lack the depth in a single major environment or two, which could allow them to get the most out of their tools and serve their clients in the best way with their linguistic skills and subject matter knowledge.

So is the latest Wordfast a tool worth checking out? I don't know yet. But it may be used by colleagues and clients with whom I like to work, and understanding how to share projects and project resources in painless ways will benefit all of us, no matter what our tool preferences may be. Wordfast seems to be developing very much in that spirit, so I will revisit it for more collaboration scenarios in the future.


Jun 16, 2012

memoQuickie: footnote, cross-reference & index entry segmentation in Microsoft Word files

If you have a Microsoft Word DOC file or RTF to translate, it is important to be aware of the different behaviors of the memoQ import filter options you can use. If there are footnotes, cross-references or index entries, it is far better to use the option to import the DOC or RTF file as DOCX.

The DOC file shown below has a footnote, a cross-reference and an index entry:


Adding it to a memoQ project with the default filter for Microsoft Word in memoQ 5


gives the following segmentation result:


Importing the same document with the DOCX option of the filter


yields much cleaner segmentation and better tags to work with:


Compare what some other programs do with this file:

WordFast Pro
DVX2 (DOC)
DVX2 (DOCX)

TagEditor salad (partial)

SDL Trados Studio 2009 segmentation

SDL Trados Studio 2011

There is room for improvement with most tools.


Jan 7, 2012

Translation tool concordances compared

A recent experience when tutoring a new memoQ user started me thinking about the way concordance searches work in various translation environment tools and how the results are displayed. The user, who was quite experienced with OmegaT, kept telling me that memoQ could not find examples of a term's use in the TM and she had to do all her searches in OmegaT. I was somewhat puzzled by that, and when I looked at her screen with the memoQ concordance dialog, I saw something like this:

The memoQ version 5 concordance dialog
Looks like the term ("Inverkehrbringen") was found. So what was the problem? For years she had looked at this concordance view:

The OmegaT concordance dialog
The differences in layout and the lack of highlighting of the key term (which was aligned in the center of the memoQ concordance window in the ancient KWIC display tradition) were unexpected and confusing to the new user.

This inspired me to have a look at how various other tools display concordance results. I was not very happy with some of what I discovered, especially with some of today's leading commercial tools. I took a look at the TWB translation memories in SDL Trados 2007, concordancing in SDL Trados Studio 2009, Wordfast Pro (very limited test due to a demo license and my inability to load my TMX test data), memoQ and OmegaT.

In terms of overall performance, the best results were obtained with OmegaT and "Trados Classic" (2007). Searching a huge TM gave results in a flash. Concordance searches with SDL Trados Studio 2009, on the other hand, really sucked with a big TM (EU data, about 400,000 TUs). I vacuumed my entire apartment and fed the dog while I waited for the result, and I wasn't even told how many hits were found. Unfortunately, my favorite working environment, memoQ, performed worst with the same big data set: it simply gave an error message. Further testing revealed that this error was due to the very large number of hits. (This would have been obvious had I paid enough attention to read the dialog title in the first place.)

memoQ error message from too many concordance hits
So it looks like some development attention may need to be directed here. (Update: Kilgray's develops are actively working to remove this restriction.) Of all the tools I was able to test with a large concordance, memoQ was the only one to fail this way. My personal TM with about 10 years of my work in it is nearly as long as my German/English EU legal test database, but concordance searches in it using memoQ are not unduly slow.

Other concordance views looked like this:

The concordance in SDL Trados 2007 - hits limited compared to OmegaT (see above)
SDL Trados 2009 - perhaps the easiest to read, but slower than molasses
Wordfast Pro - format not bad, but the test was limited due to the demo license
The Déjá Vu X concordance hasn't changed significantly in appearance in the latest version (DVX2). Once again, Victor Dewsbery was kind enough to provide me with screenshots of the two "scan" options for searching the translation memory. The initial scan produces only fairly close matches, while the "power scan" is more like the usual concordance with the term embedded in a larger body of text (the non-matching parts being crossed out)

DVX2 scan (first click)



DVX2 Power Scan (second click)

I do have a license for the older version of DVX, but I didn't attempt any stress testing. While its performance with large TMs has always been good (my personal "Big Mama" is about 330,000 TUs), import and export of such data volumes are painfully slow. We're talking overnight. I hope the new version is better in that respect. There I must really give kudos to the OmegaT developer: loading the TM was even faster than with Trados Workbench, which for me has always been a benchmark of speed to aspire to. All you have to do to add a TMX file to the TM of an OmegaT project is to drop it in the "TM" folder of the project. Very nice :-)

I also received a screenshot of a search in Transit NXT from colleague Hans Lenting in the Netherlands. He searched the term "Inverkehrbringen" in the German/Dutch EU dataset from the DGT:

STAR Transit NXT concordance search
As you can see, there are many ways to display data from a concordance search. Which do you find easiest to deal with? Personally, I love the insertion features of the memoQ concordance, but for readability I think some of the other tools are better. And I do like to know how many results I can expect from my data, and I might even want to view them all.

Jan 2, 2012

ODT files in translation environment tools

After an interesting afternoon with a friend who was a bit frustrated with the behavior of her translation assistance technology with an ODT (Open Office text) source file, I decided to have a look at how a variety of common tools handle this format. I created a small test file which contained some of the troublesome elements and saved it as *.odt for testing. The test file looked like this:

The ordered list was created using the numbering feature.

When the file was imported to OmegaT, the segmentation looked as follows:

Fairly clean, though the segmentation is a bit off due to the encoding of the space after the end of the sentence in the second block of text. Nine segments where there should have been ten.

With memoQ, the result was:

Altogether there were a dozen segments after import. The part with the hyperlink was segmented incorrectly in three parts instead of one. However, memoQ did handle the space tag after "tool." correctly and start a new segment at "Here". Once can, of course, use the segment joining function to correct the segmentation until Kilgray gets around to fixing the segmentation on the hyperlink tag:

Update 9 January 2012: The developers at Kilgray have informed me now that this quirk in the ODT filter has been corrected and will be included in the next build released.

When I tried to test my SDL Trados Studio 2009 license, at first it refused to joint the party:

Never a dull moment with SDL as we all know. Of course SDL Trados 2007 was in fact installed, but when I upgraded to Studio 2009, of course it trashed my 2007 installation, and I had been too irritated to do anything about it for over half a year since I don't use Trados for anything more than file preparation and compatibility testing anymore, and I was still able to do that for my projects with the damaged installation. However, when I discovered that the ODT file caused TagEditor to run and hide without even saying goodbye, I sighed deeply and wasted half an hour reinstalling SDL Trados 2007. At least I didn't have to go through that insane check-in/check-out license procedure online. I trusted in God and my Windows Registry entries, and the location of my license file was remembered, so all was well.

The second attempt at SDL Trados Studio 2009 was much better:

Same segmentation problem as OmegaT, and examining the tags reveals where the issue might be addressed in a tweak of the filter.

I haven't got the latest upgrade, but someone was kind enough to run my test file through SDL Trados Studio 2011, which appears to offer the best results for filtering ODT (the settings were slightly different, with the URL included, but that is also possible with some other tools):


SDL Trados TagEditor also worked after re-installation. The results were:

Oh dear. Well, it works, but if I still used TagEditor, I would run, not walk, to the much cleaner interface of OmegaT for this sort of thing if I didn't have the good sense to upgrade to Studio or something else commercial. Note the same segmentation issue and the need for filter modification.

Victor Dewsbery was kind enough to import my test file to the original Atril DVX and the newer DVX2 and send me the results:
DVX import of the test file
DVX2 import of the test file.
I also tried to test SDLX, Wordfast Pro and Wordfast Anywhere. The first two tools don't support ODT. Wordfast Anywhere claims too, but went nowhere, with the following status message displayed in my browser for about half an hour before I gave up and went to lunch:

Of course I canceled. I had a blog post to write and a New Year to get on with. Anyone who wants to try the test file in another tool (to compare apples with apples) can get it here.