Showing posts with label Dragon Naturally Speaking. Show all posts
Showing posts with label Dragon Naturally Speaking. Show all posts

Feb 6, 2020

Speech-to-text in language services and learning: an update (rescheduled)


This presentation has been rescheduled due to unanticipated conflicts. On March 4th at 4:00 pm Central European Time (10 am Eastern Standard Time), I'll be presenting an overview of some popular and/or possible platforms for generating text from spoken words for professional work and language learning. As those who have followed this blog for years know, I have written quite a bit about this in the past and done a number of videos for demonstration and instruction using various platforms, but this is a field subject to frequent change and many new developments, so it is difficult sometimes to understand the value of one tool versus another for different applications.

The webinar is available free to anyone interested, and there will be time for questions afterward. We will compare and contrast Dragon NaturallySpeaking, iOS-based applications (including Hey memoQ), Google Chrome and Windows 10 for speech recognition work in translation and transcription, discussing some of the advantages and trade-offs with each platform working in translation environments and text-editing software, and the range of languages covered by each. Join us, and see if there are good fits for your speech recognition needs!

You can register here for the discussion.

Aug 20, 2019

Dragon NaturallySpeaking tip: killing the "please say that again" message


One annoying default setting for dictation using Nuance's Dragon NaturallySpeaking is the display of a message which appears if speech is unclear or - more usually - when there is some background noise. Depending on the active settings, the message box may persist and cover up text so that it cannot be read easily or at all. This is particularly annoying when one edits while dictation is active.

It is possible to limit the amount of time this message is displayed or eliminate it altogether.

The solution to the problem is found in the DNS options:


Under Appearance...


The auto-hide delay settings in the Results Box section are the key. Set to "Never show" (as in the screenshot above), that annoying message will never appear on your screen. If you want to see it for some period of time, choose the desired time to display (the delay before hiding) and when the message appears, right click on it and enable Auto-hide:


Then the message will display and disappear again after the specified time. In my work, I find it an unnecessary interference that can be triggered by a noisy laptop fan or background chatter in Portuguese (when the doutora has visitors), so I turn off the message entirely as shown above.


(Many thanks to David Hardisty for making me aware of this possibility!)

Feb 4, 2019

Review: the Plantronics Voyager Legend monoaural headset for translation

Ergonomics are often a challenge with the microphones used for dictated translation work. I've used quite a few over the years, usually with USB connections to my computer, though I've also had a few Logitech wireless headphones with integrated mikes that performed well. However, all of them have had some disadvantages.

The country where I live (Portugal) has a rather warm climate for more than a few months of the year. Wearing headphones can get rather uncomfortable on a hot day, and even on a cold one, the pressure on my ears starts to drive me nuts after an hour or so.

Desktop microphones seem like a good solution, and I get good results with my Blue Yeti. But sometimes, when I turn my head to look at something, the pickup is not so good, and my dictation is transcribed incorrectly.

The Hey memoQ app released by memoQ Translation Technologies Ltd. underscored the ergonomic challenges of dictation for me; the app uses iOS devices to access their speech recognition features, and positioning a phone well in such a way that one can still make use of a keyboard is not easy. And trying to connect a microphone or headset by cable to the dodgy Lightning port on my iPhone 7 is usually not a good experience.

So I was intrigued by a recommendation of Plantronics headsets from Dragos Ciobanu of Leeds University (also the author of the E-learning Bakery blog). A specific model mentioned by someone who had attended a dictation workshop with him recently was the Plantronics Voyager Legend, though when I asked Dragos about his experience, he spoke mostly about the Plantronics Voyager 5200, which is a little more expensive. I decided to go "cheap" for my first experience with this sort of equipment and ordered the Voyager Legend from Amazon in Spain. I did so with some trepidation, because the reviews I read were not entirely positive.


The product arrived in simple packaging which led me to think that the Amazon review which suggested the "new" products sold might in fact be refurbished. But in the EU, all electronic gear comes with a two-year warranty, so I don't worry too much about that.

Complaints I read in the reviews about a short charger cable seem ridiculous; the cable I received was over half a meter long, and like anyone who has computers these days, I have more USB extension cords than I know what to do with should I require a longer cable for charging. The magnetic coupler for charging has a female mini-USB port, so it can be attached to another cable as well. Power connections include the most common EU two-pronged charger, the 3-pole UK charger and one for a car's cigarette lighter.

The package also included extra earpieces and covers of different sizes to customize the fit on one's ear.

I tested the microphone first with my laptop; the device was recognized easily, and the results with Dragon NaturallySpeaking were excellent. Getting the connection to my iPhone 7 proved more difficult, however. I read the Getting Started instructions carefully, tried updating the firmware (not necessary - everything was current) and tried various switching and reboot tricks, all to no avail.

Finally, I called the technical support line in the US in total frustration. I didn't expect an answer since it was still the wee hours of the morning in the US, but someone at a support call center did answer the phone. He instructed me to press and hold the "call" button on the device until its LED begins to flash blue and red.


I did that, and when the LED began flashing, "PLT_Legend" appeared in the list of available devices on my iPhone. Then I was ready to test the Voyager Legend for dictated translation with Hey memoQ.

Because I work with German and English, I rely on Dragon NaturallySpeaking for my dictation, and the iOS-based dictation of Hey memoQ will never compete with that. But I am very interested in testing and demonstrating the integrated memoQ app, because many other languages, such as Portuguese, are not available for speech recognition in Dragon NaturallySpeaking or any other readily accessible speech recognition solution of its class.

As I suspected, my dictation in Hey memoQ (and other iOS applications) was easier with the Voyager Legend. This is the first hardware configuration I have tested that really seems like it would offer acceptable ergonomics for Hey memoQ with my phone. And I can use it for Skype calls, listening to my audio books and other things, so I consider the Plantronics Voyager Legend to be money well spent. Now I'll see how it holds up for long sessions of dictated legal translation. The product literature and a little voice in my ear both claim that the device can operate for seven hours of speaking time on a battery charge, and the 90 minutes required for a full recharge will work well enough with the breaks I take in that time anyway.

Of course there are many Bluetooth microphone devices which can be used with speech recognition applications, but what distinguishes this one is its great comfort of wear and the secure fit on my ear. I look forward to a closer acquaintance.

Dec 10, 2018

"Hey memoQ" command tests

In my last post on the release of memoQ 8.7 with its new, integrated speech recognition feature I included a link to a long, boring video record of my first tests of the speech recognition facility, most of which consisted of testing various spoken iOS commands to generate text symbols, change capitalization, etc. I tested some of the integrated commands that are specific to memoQ, but not in an organized way really.

In a new testing video, I attempt to show all the memoQ-specific spoken command types and how the commands are affected by the environment (in this case I mean whether the cursor is on the target text side or the source text side or in some other place in the concordance, for example).

Most of the spoken commands work rather well, except for insertion from the concordance, which I could not get to work at all. When the cursor is in a source text cell, commands have to be given in the source text language currently, which is sure to prove interesting for people who don't speak their source language with a clean accent. Right now it's even more interesting, because English is the only language with a ready-made command list; other languages have to "roll their own" for now, which is a bit of a trial-and-error thing. I don't even want to think how this is going to work if the source language isn't supported at all; I think some thought had to be given to how to use commands with source text. I assume if it's copied to the target side it will be difficult to select unless, with butchered pronunciation, the text also happens to make sense in the target language.


It's best to watch this video on YouTube (start it, then click "YouTube" at the bottom of the running video). There you'll find a time code index in the description (after you click SEE MORE) which will enable you to navigate to specific commands or other things shown in the test video.

My ongoing work with Hey memoQ make it clear that what I call "mixed mode" (dictation with concurrent use of the keyboard) is the best and (actually) necessary way to use this feature. The style for successful dictation is also quite different than the style I need to use with Dragon NaturallySpeaking for best results. I have to discipline myself to speak more in short phrases, less in longer ones, much less in long sentences, which may cause some text to be dropped.

There is also an issue with Translation Results insertions and the lack of spaces before them; the command to insert a space ("spacebar" in English) is dodgy, so I usually have to speak it twice and end up with a superfluous space. The video shows my workaround for this in one part: I speak a filler word (in one case I tried "dummy" which was rendered as "dumb he") and then select it later and insert an entry from the Translation Results pane over the selected text. This is in fact how we can deal with specialist terminology not recognized by the current speech dictionary until it becomes possible to train new words some day.

The sound in the video (spoken commands) is also of variable quality; with some commands I had to turn my head toward the iPhone on its little tripod next to my laptop, which caused the pickup of that speech to be bad on the built-in microphone on the laptop's screen. So this isn't a Hollywood-class recording; it's simply a slightly edited record of some of my tests to give other memoQ users some idea of what they can expect from the feature right now.

Those who will be dictating in supported languages other than English need some patience right now. It's not always easy coming up with commands that will be recognized easily but which are unlikely to occur as words to be transcribed in typical dictation work. During the beta test of Hey memoQ I used some bizarre and unusual German words which just happened to be recognized. I'm developing a set of more normal-sounding commands right now, but it's a work in progress.

The difficulties I am encountering making up new command phrases (or changing the English ones in some cases) simply reinforce my belief that these command lists should be made into portable light resources as soon as possible.

I am organizing summary tables of the memoQ-specific commands and useful iOS commands for symbols, capitals, spacing, etc. comparing their performance in other iOS apps with what we see right now in Hey memoQ.

Update: the summary file for English is available here. I will post links here for any other languages I can prepare later.

Oct 21, 2016

A day in the life....


One of the things I enjoy most about professional translation is the range of activities and subject matters that one can encounter, even as a specialist in a few domains. I can't say the work is never boring, but when it does drift that way, very suddenly it isn't any more. Quite unpredictably.

Yesterday I typed translations. A bit more than expected after two sets of PowerPoint slides - a small one to translate from German and another to edit the rather acceptable English - turned out to have about 8,000 words of highly specialized slide notes about military command and control structures and the technology of fighting forest fires. (Note to self: no matter how busy you are, always import those presentations into memoQ with the options set to extract every kind of text as well as the bitmap graphics if you have to translate those too. Then do a word count! Appearances can be deceiving.)

Yesterday I dictated translations. The job started out as a bunch of text fragments from slides, where context über alles was the rule, lots of terminology required research, and voice recognition offered no particular advantages, then suddenly it became the translation of a rather long lecture using all that new terminology, and the deadline was tighter than thumbscrews operated by an angry ex-girlfriend. Dragon NaturallySpeaking to the rescue. Not only was this necessary to finish the text in a long workday rather than most of a week, but the more natural style of translation by dictation suited the purpose of the translated presentation particularly well. I could imagine myself in the room with equipment vendors, military commanders, firefighting specialists and freight forwarders, talking about the challenges faced and the technology required to avoid the tragedies of an out-of-control firestorm. And the words came out, transcribed from my voice directly into the target text fields of memoQ, exactly as they should be spoken to that audience. And at the end of that long day my hands still had feeling in them, which would not have been the case if I had typed even a third of the text.

Yesterday I made a specialized glossary to share with a presenter who will travel halfway around the world to lecture with the slides I translated for his talk. Long ago I discovered that the way I produce translations has the potential to provide additional benefits for those who will use my work. Sales representatives might need to write letters to their prospects, discussing their products in a language not mastered as a native, and the vocabulary from my work may help them to improve communication and avoid confusion that might result from using incorrect or simply different words to describe the same stuff. Or an attorney might need a quick overview of the language I used to translate the pleading she intends to file, to ensure that it is consistent with previous efforts and will not complicate discussions with her client. The terminology I research and record for each translation can be exported and reformatted quickly to produce glossaries or more complex dictionaries in a variety of formats suited for purpose. Little time and often a lot of benefits for my clients.

Yesterday I translated bitmap graphics and not only had to deal with the editing tools for that but also had to consider the best strategy for transforming the original German graphics into English ones. Would those charts be translated again into other languages? Would the graphics be re-used in other types of documents, so that I should consider ease of portability in my approach to the translation? And how the Hell do I actually use that new bitmap graphics transcription and substitution for Microsoft Office files which was added to memoQ some time ago and sort out the five charts to translate from the fifty to ignore? (Maybe I should blog the solutions some day.)

And yesterday I was asked to write summaries of large, badly scanned articles so that the equipment manufacturer would understand how its latest technology was discussed by German reviewers. As a kid I had a silly fantasy about getting paid to read, and this is just one of the many ways it unexpectedly came true. But before I get that far, these scanned files needed to be reworked so that they could be read and searched on the screen, so as I described in a guest post on another blog some years ago, I converted them to searchable PDF/A with ABBYY FineReader, which in this case also reduced their size by about 75%. The video below also shows how this works. Strangely, when I describe this procedure to other translators, many of them don't get it, and they go on about converting PDF files into editable MS Word files or plain text, or, God help them, something really stupid like importing PDF files directly into a CAT tool for translation, though none of this really relates to my purpose. Conversions often contain errors, and many texts are harder to interpret when the context of an accurate layout is lost. So "text-on-image" PDF files for translation reference to the original source files are often critical, and for files to summarize or consult sporadically for reference (with many pages to look at and essentially nothing to translate), a searchable PDF is the gold standard for efficient work.


In the course of that day I had to work with two computers linked by remote access using four networks at various time, working in German, English and Portuguese (the latter mostly involving questions to the housekeeper on how to do an online pizza delivery order so I could stay in the office and keep working). I used well over a dozen software applications for necessary tasks. These, and the environments in which they operate must be balanced carefully for efficient work. And even after some months in my new office, the balance isn't quite as good as I've had it before, and more attention to ergonomics is required.

Some colleagues are nostalgic for the "good old days" when they received a stack of paper to translate and sent off another stack of paper when the work was done, and they had a filing cabinet or a shelf of notebooks full of old work to use as reference material, and boxes of index cards stuffed full of scribbled notes on terminology next to seldom-dusty specialist dictionaries prepared by presumed experts, often full of marginalia commenting on errors or omissions and stuffed with papers bearing other scribbled notes. Not me. Since the day 30 years ago when I laboriously typed a text file full of file folder numbers and content descriptions for my research work and personal papers I have been a big believer in electronic retrieval of information wherever possible, and I miss retyping botched pages just as little as I miss the lines in the post office or the stress of dealing with delivery services.

I suspect that some feel a loss of control with the advent of new technologies in an old profession, and certainly the changes in the business environment for translation since the days of the typewriter often require a very different mentality to survive and thrive. What that mentality is, exactly, is a matter of healthy debate and often misunderstanding - again, because of the great diversity of the profession and the professions and unprofessionals in it.

The greatest challenges of new technologies that I find are the same as those faced in many other kinds of work and in modern life in general. Filtering the overabundance of input for the few things that are truly of use or interest and maintaining focus and calm amidst omnipresent distractions. Not relying too much on technologies that are far more fallible than most people, even experts, realize or acknowledge. And remembering that a fool with a tool, however many features and failsafes it may offer, remains a fool.

Aug 18, 2015

Enter the Dragon, Anywhere!

Today Nuance made a presentation of a new product to be released this autumn for mobile devices using iOS and Android operating systems: Dragon Anywhere.
 

The mobile app will allow secure transcription with a WiFi or cellular data connection as well as synchronization of custom vocabulary with Dragon desktop computer applications wor Windows and MacOS.

The initial presentation made no mention of which languages will be available in Dragon Anywhere; the synchronization feature makes me worry that it might be restricted to the current seven or eight languages available for Dragon NaturallySpeaking (Windows) or Dragon Dictate (Mac), but perhaps the standalone applications for desktop computers and laptops will finally be upgraded to offer the 40+ languages currently available for mobile devices with apps such as Dragon Dictation for iOS or Swype + Dragon Dictation for Android.

It is also unclear at this point whether the new Dragon Anywhere app will allow direct dictation or transfer to the cursor location on a linked computer, as one can do with myEcho or using Swype + Dragon Dictation in conjunction with Chrome Remote Desktop. But with the addition of extensive voice-controlled editing features to the mobile app, Dragon Anywhere represents significant progress toward better ergonomics for writing and translation!

Jan 25, 2015

SDL conquers translation at Universidade Nova in Lisbon


The day started inauspiciously for me, with a TomTom navigation system determined to keep me from the day planned at Lisbon's New University to discuss SDL Trados Studio and its place in the translation technology ecosphere. When the fourth GPS location almost proved a charm, and I hiked the last kilometer on an arthritic foot, swearing furiously that this was my last visit to the Big City, I found the lecture hall at last, an hour and a half late, and managed to arrive just after Paul Filkin's presentation of the SDL OpenExchange, an underused, but rather interesting and helpful resource center for plug-ins and other resources for SDL Trados Studio victims to bridge the gap between its out-of-the-box configurations and what particular users or workflows might require. There are a lot of good things to be found there - the memoQ XLIFF definition and Glossary Converter are my particular favorites. Paul talked about many interesting things, I was told, and there is even a plug-in created for SDL Trados Studio by a major governmental organization which has functionality much like memoQ's LiveDocs (discussed afterward but not shown in the talk I missed, however). In the course of the day, Paul also disclosed an exciting new feature for SDL Trados Studio which many memoQ users have been missing in the latest version, memoQ 2014 R2 (see the video at the end).

I arrived just in time for the highlight of the day, the demonstration of Portuguese speech recognition by David Hardisty and two of his masters students, Isabel Rocha and Joana Bernardo. Speech recognition is perhaps one of the most interesting, useful and exciting technologies applied to translation today, but its application is limited to the languages available, which are not so many with the popular Dragon Naturally Speaking application from Nuance. Portuguese is curiously absent from the current offerings despite its far more important role in the world than minor languages like German or French.

Professor Hardisty led off with an overview of the equipment and software used and recommended (slides available here); the solution for Portuguese uses the integrated voice recognition features of the Macintosh operating system. With Parallels Desktop 10 for Mac it can be used for Windows applications such as SDL Trados Studio and memoQ as well. Nuance provides the voice recognition technology to Apple, and Brazilian and European Portuguese are among the languages provided to Apple which are not part of Nuance's commercial products for consumers (Dragon Naturally Speaking and Dragon Dictate).

Information from the Apple web site states that
Dictation lets you talk where you would type — and it now works in over 40 languages. So you can reply to an email, search the web or write a report using just your voice. Navigate to any text field, activate Dictation, then say what you want to write. Dictation converts your words into text. OS X Yosemite also adds more than 50 editing and formatting commands to Dictation. So you can turn on Dictation and tell your Mac to bold a paragraph, delete a sentence or replace a word. You can also use Automator workflows to create your own Dictation commands.
Portuguese was among the languages added with OS X Yosemite.

Ms. Bernardo began her demonstration by showing her typing speed - somewhat less than optimal due to the effects of disability from cerebral palsy. I was told that this had led to some difficulties during a professional internship, where her typing speed was not sufficient to keep up with the expectations for translation output in the company. However, I saw for myself how the integrated speech recognition features enable her to lay down text in a word processor or translation environment tool as quickly as or faster than most of us can type. In Portuguese, a language I had thought not available for work by my colleagues translating into that language.

A week before I had visited Professor Hardisty's evening class, where after my lecture on interoperability for CAT tools, Ms. Rocha had shown me how she works with Portuguese speech recognition as I do, in "mixed mode" using a fluid work style of dictation, typing, and pointing technology. She said that her own work is not much faster than when she types, but that the physical and mental strain of the work is far less than when she types and the quality of her translation tends to be better, because she is more focused on the text. This greater concentration on words, meaning and good communication matches my own experience, but I don't necessarily believe her about the speed. I don't think she has actually measured her throughput. My observation after the evening class and again at the event with SDL was that she works as fast as I do with dictation, and when I have a need for speed that can go to triple my typing rate or more per hour.

In any case, I am very excited that speech recognition is now available to a wider circle of professionals, and with integrated dictation features in the upcoming Windows 10 (a free upgrade for Windows 8 users), I expect this situation will only improve. I cannot emphasize enough the importance of this technology for improving the ergonomics of our work. It's more than just leveling the field for gifted colleagues like Joana Bernardo, who can now bring to bear her linguistic skills and subject knowledge at a working speed on par with other professionals - or faster - but for someone like me who often works with pain and numbness in the hands from strain injuries, or all the rest of you banging away happily on keyboards, with an addiction to pain meds in your future perhaps, speech recognition offers a better future. Some are perhaps put off by the unhelpful, boastful emphasis of others on high output, which anyone familiar with speech recognition and human-assisted machine pseudo-translation (HAMPsTr) editing knows is faster and better than what any processes involving human revision of computer-generated linguistic sausage can produce, but it's really about working better and doing better work with better personal health. It's not about silly "Hendzel Units".

It has been pointed out a few times that Mac dictation or other speech recognition implementations lack the full range of command features found in an application like Dragon Naturally Speaking. That's really irrelevant. The most efficient speech recognition users I know do not use a lot of voice-controlled command for menu options, etc. I don't bother with that stuff at all but work instead very comfortably with a mix of voice, keyboard and mouse as I learned from a colleague who can knock off over 8,000 words of top-quality translation per short, restful day before taking the afternoon off to play with her cats or go shopping and spend some of that 6-figure translation income that she had even before learning to charge better rates.

Professor Hardisty also gave me a useful surprise in his talk - a well-articulated suggestion for a much more productive way to integrate machine translation in translation workflows:
David Hardisty's "pre-editing" approach for MpT output
The approach he suggested is actually one of the techniques I use with multiple TM matches in the working translation grid where I dictate - look at a match or TM suggestion displayed in a second pane and cherry-pick any useful phrases or sentence fragments and simply speak them along with selected term suggestions from glossaries, etc. and do it right the first time, faster than post-editing. This does work, much better than the sort of nonsense pushed too often into university curricula now by the greedy technotwits and Linguistic Sausage Purveyors, who in their desire for better margins and general disrespect of human service providers and employees fail to understand that good people, well-treated and empowered with the right tools, will beat the software and hardware software of "MT" and its hamsterized process extensions every time. Hardisty's approach is the most credible suggestion I have seen yet for possibly useful application of machine pseudo-translation in good work. Don't dump the MpT sewage directly into the target text stream like so many do as they inevitably and ignorantly diminish the level of achievable output quality.

After the lunch break, Paul Filkin gave an excellent Q&A clinic on Trados Studio features, showing solutions for challenges faced by users at all levels. It's always a pleasure to see him bring his encyclopedic knowledge of that difficult environment to bear in poised, useful ways to make it almost seem easy to work with the tools. I've sent many people to Paul and his team for help over the years, and none have been disappointed according to the feedback I have heard. The Trados Studio "clinic" at Universidade Nova reminded me why.

Finally, in the last hour of the day, I presented my perspective on how the SDL Trados Studio suite can integrate usefully in teamwork involving colleagues and customers with other technology and how over the years as a user of Déja Vu and later memoQ as my primary tool, the Trados suite has often made my work easier and significantly improved my earnings, for example with the excellent output management options for terminology in SDL Trados MultiTerm.


I spoke about the different levels of information exchange in interoperable translation workflows. I have done so often in the past from a memoQ perspective, but on this day I took the SDL Trados angle and showed very specifically, using screenshots from the latest build of SDL Trados Studio 2014, how this software can integrate beautifully and reliably as the hub or a spoke in the wheel of work collaboration.

The examples I presented using involved specifics of interoperability with memoQ or OmegaT, but they work with any good, professional tool. (Please note that Across is neither good nor a professional translation tool.) Those present also left with interoperability knowledge that no others in the field of translation have as far as I know - a simple way to access all the data in a memoQ Handoff package for translation in other environments like SDL Trados Studio, including how to move bilingual LiveDocs content easily into the other tool's translation memory.


Working in a single translation environment for actual translation is ergonomically critical to productivity and full focus on producing good content of the best linguistic character and subject presentation without the time- and quality-killing distractions of "CAT hopping", switching between environments such as SDL Trados Studio, memoQ, Wordfast, memSource, etc. Busy translators who learn the principles of interoperability and how to move the work in and out of their sole translation tool (using competitive tools for other tasks at which they may excel, such as preparing certain project types, extracting or outputting terminology, etc.) will very likely see a bigger increase in earnings than they can by price increases in the next decade. On those rare occasions where it might be desirable to use a different tool or to cope with the stress of change from one tool to another, harmonization of customizable features such as keyboard shortcuts can be very helpful.

I ended my talk with a demonstration of how translation files (SDLXLIFF) and project packages (SDLPPX) from SDL Trados Studio can be brought easily into memoQ for translation in that ergonomic environment, with all the TMs and terminology resources, returning exactly the content required in an SDLRPX file. Throughout the presentation there was some discussion of where SDL and its competitors can and should strive to go beyond the current and occasionally dubious levels of "compatibility" for even better collaboration between professionals and customers in the future.

One of the attendees, Steve Dyson, also published an interesting summary of the day on his blog.


Oct 9, 2014

Dragon Naturally Speaking Version 13 - Review!

I've had a number of people ask me recently whether I have upgraded to Dragon Naturally Speaking version 13 for my dictation work in translation. I have not; I am still using the German version 12.5 (which includes English - I sometimes dictate poorly legible source texts in German rather than waste my time with OCR if I want to work with translation environment tools, so I need the bilingual edition). However, a colleague was kind enough to point me to this review of the new version, which gives me more than sufficient reason to upgrade soon:


I have a few YouTube videos demonstrating the use of version 11.5 in memoQ and a word processor, which seem to have generated some excitement because of the ridiculously high speed at which I can translate by dictation (and many others are much faster). However, the point of voice recognition for me is not speed and the possibly higher earnings which can result if my editing afterward is not excessive (dictation requires a completely different approach to checking your work, and there is a significant learning curve here). Also (or really more) important are:

  • greater engagement with the text on the screen, in my case leaving my hands free to point at various parts of long, complex sentences to help me sequence the translated text better as I work;
  • less physical and mental strain during my translation work (I am less tired during and after);
  • relief for hands damaged by too many years of working with vibrating power equipment (tillers and chainsaws), riding bicycles on rough ground and typing, typing, typing (some days I have to wash dishes in very hot water for an hour and load up on pain meds just to use a keyboard and mouse without tears - there may be surgery for that in my future and tools like DNS can give others relief or help prevent the sort of strain injuries too common in this profession and others which involve a lot of keyboard work). 
Dragon Naturally Speaking is currently available for U.S. English, UK English, German, French, Italian, Spanish, Dutch, and Japanese. Given the importance of voice recognition for relieving or avoiding strain injuries as well as for productivity (translators working with voice recognition routinely have much higher outputs of good quality than the best realistic claims for crap produced by post-editing machine pseudo-translations), I sincerely hope that Nuance and others will pursue the development of speech technology for other major languages such as Portuguese, Arabic, Chinese and Greek. Such an investment is likely to produce far greater benefits all-round than any money flushed down the machine pseudo-translation toilet, and speech recognition could probably also improve the working conditions of some stressed post-editors in the HAMPsTr world.

Apr 18, 2014

10 Steps to Determine CAT Tool Compatibility with the Dragon

Guest post by Jim Wardell

The following steps can be used to investigate the degree to which a CAT tool is compatible with Dragon NaturallySpeaking.
  1. In Dragon NaturallySpeaking, go to Tools>Options>Miscellaneous and check the box Use the Dictation Box for unsupported applications.
  2. Open a project in the CAT tool you are testing, place the cursor in the translation target cell and start dictating with Dragon. If the Dictation Box opens immediately in Dragon, that means that Dragon has determined that your cat tool is not fully compatible with Dragon and that you must first dictate your translation in the Dictation box that it has just opened.
    You can then transfer your dictated translation from the Dictation Box into the translation target cell of your CAT tool by typing Ctrl+T or by saying “Transfer”. This extra step does not reduce your productivity too much if your source segment contains little or no formatting, tags, auto translate items or placeables. But if your source text does have a lot of these sorts of things, you’re going to have to add the extra step of first copying all of the source text segment into the Translation Box. When you do this, though, you’re going to lose any tags and formatting. Once you’ve translated the raw text in the Dragon Translation Box and have transferred the contents of the Translation Box into your translation target cell, you’re still going to have to reformat, add tags, and generally mess about a bit, maybe a lot. I went through this process for many years when using CAT tools that were not fully compliant with Dragon and eventually realized that all of this extra work was completely wiping out the productivity and income boost I was getting from dictating. Plus, it was more fatiguing because I was making my workflow more complicated.
  3. The next thing to try if the Transfer Box comes up automatically is to go back to the above Dragon NaturallySpeaking setting and uncheck the Use the Dictation Box for unsupported applications setting.
  4. Now go back to a target translation cell and try to dictate a medium-length sentence. With Tools that are completely incompatible, nothing will happen and perhaps your system will hang. In most cases, though, you will be able to dictate something. Unfortunately, this does not necessarily mean that your CAT tool is compliant with Dragon.
  5. Now, using Dragon, say:  “Select <a string of 2 or 3 words from the sentence you just dictated>. If those two or three words now get marked in your CAT tool, you at least have partial compatibility with Dragon.
  6. Now comes the acid test: Say “Correct that”. The Dragon Correction Menu should appear and give you a list of possible alternatives. If the Correction Menu does not appear, your CAT tool is not sufficiently compatible with Dragon and you should not do your translation work directly in your CAT tool target cell. Instead, you and will need to use the Dictation Box workaround if you really “must” use the CAT tool you are testing.
    I call this the “acid test” because if you can’t correct words that are not recognized correctly when you dictate, your incorrect dictations are probably being fed into your Dragon language profile. This will degrade the accuracy of speech recognition over time! It’s hard to know for sure what’s really going on in such cases, but the best that you can hope for is that incorrectly recognized words are simply being ignored by Dragon. However, anyone who uses Dragon a lot knows that the Correction process is one of the ways that Dragon becomes more and more accurate over time. Not correcting misrecognized speech is bad Dragon practice, and an application that doesn’t allow you to correct misrecognized speech should not be used or should at least always be used with the Dictation Box workaround. The most important reason to always use the Dictation Box with noncompliant software is that this permits the correction of speech recognition errors and hence ensures that Dragon will keep getting better and better at recognizing your speech and vocabulary.
    However, my recommendation to anyone who is serious about using Dragon to increase productivity and income is to use a CAT tool like memoQ or Déjà Vu that works flawlessly and seamlessly WITHIN the CAT tool target cell, i.e. with correction working properly directly in the CAT tool.
  7. If your CAT tool fails the acid test, then you might just as well stop here. But if it passes the acid test, then try a few other commonly used Dragon commands in your target cell. Use “Select …” to mark some other text strings. Then say “Capitalize that” or “Make that bold” or “Underline that”. Also try using various Dragon cut-and-paste commands. If all this works, then congratulations! Your CAT tool has a high level of Dragon compatibility!
  8. The next level of compatibility is achieved when you can also use Dragon in other functions in your CAT tool - ideally all functions. So now open a “comment” or “note” for a translation segment. Dictate something into the comment, repeating steps 4, 5, 6 and 7. If your CAT tool passes these tests, you can celebrate big-time because being able to dictate comments quickly and effectively can add significant value to your translation when it comes to working in translation teams and providing feedback to clients, or querying clients articulately about terminology questions.
  9. Being able to dictate using Dragon is also absolutely invaluable when it comes to quickly creating really good terminology base entries. So create or go to a terminology base entry and dictate something into one of the fields in which you are able to enter free-form text. In memoQ, for example, this might be the “Note” field or one of the “Definition” fields. Once again repeat steps 4, 5, 6 and 7.
  10. Now test your CAT tool’s ability to recognize dictated CAT commands. First, you will have to make sure that a configuration setting  in Dragon is set right. Go to Tools>Options>Miscellaneous and check Voice-enable menus, buttons, and other controls, excluding: … Then open the pull-down menu and make sure that your CAT tool is NOT checked.
    Now look to see what your CAT software calls the command that is used to confirm a segment and enter the segment content into translation memory. Place your cursor in a translation segment and try saying this command to see what happens. In memoQ, for example, I simply dictate “Confirm” and memoQ operates just as if I had pressed Ctrl+Enter. It is not essential to have this level of compatibility, but it is great to use when working in tight spaces on planes and car seats. In memoQ I can also dictate any of the names of the main menus and can then dictate submenu names as well and essentially “menu down” by voice commands. Again, this is not essential but it sometimes comes in handy.
If you work INTO more than one language, you’ll want to test whether you can work in the user interface of these target languages in your CAT tool and whether your CAT tool responds to voice commands from Dragon in the same target language.  This is really having your cake and eating it too! If you find this level of Dragon compatibility in your CAT tool, it means that the developer was really committed to getting maximum Dragon compatibility for working translators like you. Send them a nice thank-you note and publicize their commitment to working-stiff translators every chance you get!

***


Jim Wardell will be presenting optimized work methods for speech recognition once again at this year's memoQfest in Budapest, Hungary.

Apr 15, 2014

Speech recognition for translators: microphone tips

Guest post by Jim Wardell

Mark Myworts is heavy into a major procrastination project, pawing through moldy old copies of Popular Mechanics in Grandpa’s basement, when a small classified ad in the back pages catches his eye. He blows off the dust:
“Translators! Double your income overnight with amazing new technology! Works wonders for English, German, Spanish, French, Italian, Dutch. No obligation. Call ... Full confidentiality guaranteed.”
[Fade in “Twilight Zone” theme music.]

[Cut to Mark talking intently to his computer.] “... should be used instead of the diminutive terms ‘pud’ and ‘loser’” ...

[Fade to Mark and kiddies.] “Well daddy, can we? can we? Can we go to the circus tonight?” “Sure kids, I’m knocking off early today,” says Mark nonchalantly, getting a kiss and one of those sexy “well-what-about-after-the circus” looks from his admiring wife.

Science fiction? 1950s social mythology? Perhaps.

But the simple fact remains – at least now in 2014 – that translators into English, German, Spanish, French, Italian and Dutch really can double their productivity on average using some amazing, although not quite so new technology: speech recognition.


***

I’ve been using speech recognition to translate from German into English for nearly 20 years. But it was not until about seven or eight years ago that computing power and speech recognition software had improved to the point where serious productivity gains became possible. It was at that point that it became imperative for me to find a CAT tool that was totally compatible with Dragon NaturallySpeaking. At the time, only two products met this requirement: Déjà Vu and memoQ. For various reasons, which I won’t get into now, I decided to go with memoQ, a decision I have never once regretted.

Most any CAT tool can be made to work with Dragon by using the little “Dictation Box” text buffer that’s provided as a workaround in Dragon for software that is not truly Dragon compliant. The procedure needed to translate average moderately sophisticated technical documents in noncompliant CAT tools can often be cumbersome and inefficient: one copies the contents of the current source segment text into the Dictation Box so any strings that do not need to be translated can be left as is or moved around as desired and so that sections of text that need to be translated can be overwritten by marking them and then dictating the new translation “over the top”. Once the source segment has been duly massaged in the Dictation Box, the contents of the box are then transferred to the target box in the noncompliant CAT tool. Of course, various tags and formatting that might have been present in the source segment are often lost when source text pasted into the Dictation Box. So they need to be put back in again after the contents of the Dictation Box have been pasted into the CAT target box. If this sounds gruesome, it is.

So why not just dictate straight into the noncompliant target box and fix the messes as they occur? The answer is simple: there are often too many messes, and, worse still, any incorrect speech recognition that occurs cannot be corrected in a way that will ensure that correct and not erroneous data will be fed and saved in the Dragon speech recognition engine. Over time, this would degrade speech recognition accuracy! I spent some years trying to address this issue with publishers of CAT tools other than memoQ and Déjà Vu ... with zero success. So if you’re not already using Dragon and want to use it with a CAT tool, make sure that the CAT tool that you are thinking of using is really fully compatible with Dragon and that you can get your money back if it’s not. Do not trust and do verify.[1]

At one point before I switched memoQ, I was compelled to do a good bit of this acrobatics moving text into an out of the Dragon Dictation Box. I began to have the feeling that the cutting edge that I was working on in a number of well-known CAT tools was so dull that I might just as well have been typing in my translation in the old-fashioned way. So I collected some statistics discovered that that was indeed the case. My output was the same in noncompliant CAT tools and Dragon as with touch typing without Dragon.

All that changed with the memoQ’s full Dragon compatibility! Incidentally, memoQ has full Dragon compatibility throughout the interface and not just in the translation grid. So if you want to dictate notes or definitions in term base entries, you can, and you can still use all of the selection and correction features you are accustomed to using in Dragon. Want to write a longish note to a client in a memoQ Comment box? No problem. Dictate away.

* * *

Anyone who has dealt with integrated technologies, any process in fact, knows that the old saying “A chain is only as strong as its weakest link” is totally true. So to get really great speech recognition results, not only does one’s CAT tool have to be compatible and outstandingly good, one’s computer needs to be sufficiently powerful, and one needs to use the best available microphones. I heartily recommend KnowBrainer.com as a source of top quality microphones for speech recognition. To my knowledge, KnowBrainer is the expert in the USA, probably the world, when it comes to speech recognition products. People who want to achieve maximum accuracy with speech recognition software should make their first stop KnowBrainer’s Microphone Comparison page.[2] For many years, I used KnowBrainer’s top-rated Samson Airline 77 microphone. This microphone was vastly superior to anything I had ever used in the past and came delightfully close to delivering 100% speech recognition accuracy. Earlier this year, however, I learned that the wireless channel used in my old Airline 77, which I bought while I was still located in the United States, was being shifted to use by mobile phones in Europe and would no longer be legal. So I checked out KnowBrainer again, and learned about a relatively new microphone being produced specifically for speech recognition by a Belgian company: SpeechWare. Upon consulting with KnowBrainer’s Lunis Orcutt (Mr. Speech Recognition in my book!), I ordered the SpeechWare 3-1 TableMike. This desktop mic is a great product and just as good as my old Airline 77. It’s the mic that Lunis himself uses.

However, after using it for a week or so, I realized that it was not for me because I had to keep my mouth relatively close to the microphone and couldn’t move around like I was used to in the past in order to relax my back muscles and stay fresh. I then ordered a FlexyMike headset mic from SpeechWare that basically uses the same technology but allows one to move around freely. SpeechWare has three models of the FlexyMike: the FlexyMike Basic (FMK01), the Single Ear (SE) and the Dual Ear. I chose the Dual Ear on the principle that distributing the weight of the mic over two ears would be more comfortable and stable for hours and hours of continuous use.

When I was still using the tabletop TableMike, I found that I had a tendency to move a little too far away from the microphone over time, which occasionally reduced speech recognition performance. The TableMike has two settings: a long-range setting, which allows one to have one’s mouth as far as 30 cm (12 inches) from the microphone, and a “normal and VoIP” setting (with a maximum distance of 15 cm / 6 inches). The greatest accuracy is achieved with the closer distance. Lunis says he likes the TableMike because he moves around in the office a lot and doesn’t need to fumble around with headset whenever he leaves his desk. For this reason, I would recommend the TableMike as the best choice for project managers and administrators who may frequently have to leave their desks and who mainly use Dragon at brief stretches to dictate e-mail messages or enter data in translation business management software. For hard-core translation work, the FlexyMike is the way to go. I find the accuracy with the FlexyMike to be perhaps a tad better than that of the Airline 77, which is saying a lot. I can use the FlexyMike while listening to the radio a moderate volume levels, so the noise cancellation is also quite good. KnowBrainer gives its noise cancellation a score of 9, which is better than that of the Sennheiser ME3 KB headset mic (gets an 8), which I have used with good success for years in automobiles, trains and airplanes! All the same, if anyone has to dictate in an extremely noisy environment, one might want to check out “theBoom v4 KB”, which gets a high accuracy rating and a 10 for noise cancellation (but only a 9 for comfort!) from KnowBrainer. My experience is that KnowBrainer is pretty fanatical about these evaluations and that they are quite reliable.

For the average translator, who works long hours in a relatively quiet environment, accuracy and comfort are the two most important factors, more important than noise cancellation. I don’t need speakers on my headset, which means that the headset can be as light as a feather and can be worn comfortably all day long. If need be, Skype calls or music can be played through normal computer speakers. On the other hand, if one is working in an open office setting with a number of other translators close by, one might want to have a headset with speakers covering both ears to block out distracting voices so one can concentrate better. In such cases, I’d consider the mono Umevoice “theBoom Pro-2 KB” or the stereo hi-fi equivalent “... 3 KB” if you want to block out room noise and also want to listen to music while you translate. (I translate very complicated, detailed stuff and usually extremely distracting to listen to music while translating, but not all material that gets translated requires extreme concentration. I could also easily imagine listening to a high-bandwidth feed from, say, jazzradio.com premium (unabashed plug) to make routine administrative work more pleasant.

Getting back to the FlexyMikes: SpeechWare was kind enough to also send me a single-ear model to test and evaluate, so I have used both versions extensively. Both the single-ear and the double-ear mics are extremely comfortable and both are very easy to adjust to get a custom fit that is secure and comfortable. The materials used in both mics are of exceptionally high quality and should provide many years of reliable service.

Both FlexyMikes connect to a computer USB port across a “SpeechMatic MultiAdapter”, which has been especially configured for high accuracy with speech recognition. I am convinced that the special design of the MultiAdapter is one of the main reasons why the FlexyMikes work so well.[3] Be sure to buy this along with your FlexyMike. The same circuitry that’s in the MultiAdapter is integrated into the TableMike units. So if you already have a TableMike, you don’t need to buy a MultiAdapter, unless of course you want something really small and light to use with a notebook computer when traveling.

I did not test the basic version of the FlexyMike, from the pictures it didn’t look as comfortable as the other models.

KnowBrainer.com ships internationally. SpeechWare microphones are also available directly from SpeechWare in Europe.

[1] If you want to see what “fully compatible” means, have a look at http://kilgray.com/news/once-upon-time-there-was-dragon.
[2] http://www.knowbrainer.com/core/pages/miccompare.cfm
[3] So is KnowBrainer: See http://www.knowbrainer.com/NewStore/pc/viewPrd.asp?idproduct=464



***


Jim Wardell will be presenting optimized work methods for speech recognition once again at this year's memoQfest in Budapest, Hungary.

Jan 4, 2014

TeamViewer and Dragon Naturally Speaking: currently a bad mix

On New Year's Eve I took delivery of new hardware to support my translation work. This will be the first time in more than a decade that the bulk of my work will not be done on a laptop, but the demands I've put on my hardware in recent years are a bit much for any laptop I'm willing to invest in. Now set with 32 GB RAM, a few SSD drives, souped-up video and other features to make my work go with a bit less hassle, I decided it was time to try the remote access solutions that some of my friends and colleagues have relied on for the past few years. I've been particularly impressed with what one of them does running all the applications on his home system with excellent performance from his desk at work or other remote locations. At last I am ready to do the same.

The new dream machine is still being configured, but I've got memoQ and other useful tools loaded, even SDL Trados Studio 2014 carefully isolated in a well-configured VMware machine to avoid trashing my main system as SDL software always has in the past.

There are, of course, many possibilities for remote access. Because I use TeamViewer sometimes for remote assistance to clients and colleagues and impromptu mini-webinars of an informal nature, I thought I would try the new, improved access in version 9 that one colleague mentioned. Things have looked quite good on the whole.

The only major failure I have experienced has been with voice recognition. I use Dragon Naturally Speaking sometimes for my translation work, and out of curiosity I decided to try it with a text I had in SDL Trados Studio on the virtual machine on the remote computer. Typing worked just fine in this configuration.

Dictation with DNS was another matter altogether. Sentences were not capitalized at the beginning, and small pauses in my voice caused spaces to be dropped in the text on the remote virtual machine. Now I know that even Trados isn't this bad with dictation, so I repeated the test in a simple word processor on the VM and repeated it in the same word processor in the remote host system. In each case, the problem was the same: failures to capitalize the beginning of sentences and frequent dropped spaces. I had to discipline myself to speak capitalization commands and insert spaces by voice after any pause. Editing by voice was also impossible and had to be done manually with the mouse and keyboard. Word accuracy was as good as ever, but that's not surprising, as that processing all occurs locally. The difficulties are in transmission to the remote system.

I suspect this is a problem to be addressed by TeamViewer rather than Nuance. I am very curious to see whether other remote access solutions have similar difficulties. If anyone else has relevant experience with this, please share it.

Nov 2, 2013

Games freelancers translate

ames are no longer a big part of my world despite years spent collecting, playing and developing them ages ago. In the world of translation, I am an interested observer, fascinated a little by the technical peculiarities I hear of in that domain as well as what appears to be a diversity of opinion and working methods even greater than one finds in my familiar areas of work.

I always enjoy a close look at the working processes of colleagues and clients; often I learn new things from the observation, and I like to ask myself as I see each stage what approach I might take or whether there are changes in the available tools which might make a process more efficient.

An Italian freelance team (leader?) put together a series of seven YouTube videos showing how jobs are prepared and distributed, as well as some particulars of their translation process and QA. The main working tool is Kilgray's memoQ - one of the 6-series versions it seems - as well as the Italian version of Dragon Naturally Speaking and Apsic Xbench, which also make a brief appearance. Altogether 22 minutes of show and tell, which I find mostly interesting and recommend as a nice little process overview.

I've made a YouTube playlist compilation here so it is easier to view the clips in sequence, since I had a little trouble navigating them myself in the somewhat random YouTube suggestion menus. I'm not embedding them here, because the interface for navigating a playlist is much easier to cope with on YouTube itself.

I wish there were more overviews like this available for common translation workflows in other areas as well, such as patent translation, financial report translation in the midst of the "silly season", web site translation or just about anything else. It's doubtful that any of these would betray great trade secrets, but they might offer clients and prospects a little more realistic view of what some might think involves little more than "retyping in another language".

Some content notes on the individual videos of the playlist:

#1 Discusses background research and style guides in the team's approach

#2 Covers the import of the source files (Excel) and the selection of ranges

#3 Term extraction

#4 Statistics, handoff packages and sending out the jobs with the project management system

#5 Creating views of multiple files, voice recognition in Italian, concordances and term lookups

#6 Receiving translated project packages; text to speech reviewing!

#7 QA in memoQ, export to XLIFF for final QA in Apsic Xbench

Aug 16, 2013

SDL Trados Studio vs. memoQ: Translating Text Columns in Excel

Paul Filkin of SDL recently showed a few "little known gems" of SDL Trados Studio 2011 in one of his recent blog posts, which is quite useful for learning how to approach some not-so-rare project challenges with Trados Studio. Here I would like to share his video tutorial about one of those gems - how to translate multiple columns of text in an Excel file. (HINT: these embedded videos are easier to watch if you do that in full screen mode by toggling the icon at the lower right of the play window.)



memoQ can also translate Excel files with a similar structure, and here's how to do that, with a little bit of Dragon Naturally Speaking thrown in just for fun:




Aug 9, 2013

Translating against the clock with Dragon Naturally Speaking

After a discussion on voice recognition, which began in the comments for a totally unrelated post about a scatologically bad Linguistic Sausage Producer (LSP), I attempted a timed test of translation with voice transcription in my usual working environment as a demonstration of the productivity gains to be achieved. The Gods were against me that day or my microphone was adjusted wrong; the results were somewhere around 2000 miserable words an hour with a lot of edits as I dictated.

Tonight I resolved to do better with the deck stacked against me. I waited until about 3 am (easy to do on a hot day in a Mediterranean climate) and picked a relatively easy but unfamiliar German text from Wikipedia, took out most of the Greek words I can't pronounce and fired up The Dragon to burn the translation.

No CAT tool this time, nothing but me, the text about snakes and a timer... how many words of draft translation would you expect to do with an easy text in an hour in the middle of the night without coffee? (Of course it will need some revision later, but I wanted a crappy translation for editing demos anyway.)

Jul 26, 2013

The trouble with voice recognition in translation environment tools....


I had not planned to make a video on voice recognition tools any time soon, but a few remarks by my American colleague Kevin Hendzel well down in the many comments about thepigturd's letter to translators sort of goaded me into it. I thought, "What the heck, I'll just grab some text from Wikipedia, record a bit of the work with Camtasia, and post a quick demo of how easy it is to work with Dragon Naturally Speaking." So I got a text about chickens. And activated the screencast recorder. And then the trouble started.

It really sucked. Working with Dragon in memoQ is usually a fairly painless process, but tonight the dogs were anxious and kept poking me in the ribs, and I never did get the microphone adjusted quite right. Some days, microphone position is everything to my scaly transcriptionist. So I suffered with a lot more editing than usual, as anyone watching the video above will see. I worked in my usual "mixed mode" manner, with both keyboard and voice control. Some colleagues who swear by DNS like to do everything by voice and would probably wipe their backsides in the WC that way as well if they could, but that's way too geeky for me. After watching my copywriting partner fly through some 10,000 words of legal translation - and edit it - in a short working day while I slogged through my 3,000 and finished long after she called it a day, I realized that I could work in the relaxed way she did with thoughtful stares at the screen, muttered bursts and the occasional keyboard touch.


But today was a bad day with the Dragon. I might have gone a bit faster with the text. After all, chickens aren't rocket science or even chemistry, with its tag-ridden notation. I could have just dictated in a word processor and everything would have one faster. And if I really want a TM or want to check the terminology, alignment is fast and also a good environment for editing my first draft. I know a number of translators who work that way now. Even with a dictaphone.

In his comments on the other post, Kevin Hendzel expressed a similar feeling to mine when translating with voice recognition: greater engagement and concentration on the text and its structure and meaning. But these tools are not without risk: any errors will in fact pass muster with a spelling checker, so proofreading workflows may have to be very different to be effective. I have noticed this myself - reading my text soon after I have translated it, I am very likely to overlook a missing or switched article or a homophone. Perhaps dictating into a word processor or - since I often look to the glossary hits and other hints on the right of my working window - exporting my text and re-aligning it in the CAT tool after an external rewrite may force my eyes to see things a little differently. In the two years that I have been making serious use of voice recognition I have not yet found the "perfect" workflow.

There are a lot of ways I can tease better results out of this work. But even on a bad day like today, things aren't all that awful. In fact, those familiar with some of the more honest estimates of output in optimized machine translation and post-editing scenarios will realize that today's lousy results (see the end of the video), maintained over the course of a working day, meet or beat the expectations for post-editing in a highly optimized scenario. Without the brain rot typically caused by PEMT! Now that's an advantage. Why don't we stop wasting time with machine translation and instead increase output by more research into the best ways of using voice recognition technology? Ah, but voice recognition is not yet optimized for every language! Ha ha ha... like MT is or ever will be. The millions that get flushed down the toilet with machine translation could and should buy a lot of improvement with voice recognition.

The real trouble with voice recognition is that you may not want your competition to use it. With or without CAT tools. Unlike machine translation.

Apr 21, 2012

TM Follies


A recent comment by Iwan Davies on Twitter revealing a reviewer's rather odd notions of the requirements imposed on translation by the use of a translation memory tool led me to reflect with a friend on some of the very strange and wrong ideas that persist in some minds with regard to such technology. In the case of the twitstream discussion, the reviewer's stupid notion that each segment in a translation must stand on its own without context provoked an interesting flurry of responses, ranging from the astute observation from @PaulAppleyard that "if you wanted to translate segments as 'standalone', then you would work in a random segment file, not a text that flows..." to some rather disturbing remarks from a few to the effect of "this is why I don't like to use such tools". Various people pointed out that modern translation environment tools such as SDL Trados Studio, OmegaT and memoQ make use of context in their translation memories to avoid the problems of more primitive systems which, in the hands of translating monkeys, too often result in matches being used in very inappropriate ways.

The list I could compile of wrong-headed ideas about TMs is a long one, and I would probably only capture ten percent of the foolishness on a lucky day. A few highlights in my memory include:
  • A statement by an otherwise respected colleague some years ago that translators must not sacrifice potential "leverage" by combining segments to make sense in the translation. This included cases where someone inserts
    carriage
    returns and line
    breaks into the sentence to
    make it fit in some odd space. In a source language like German, where word order is often very different than in a good English translation, this can quickly pollute a TM to the point of being worst than worthless. This in fact describes the real state of many "promiscuous" agency TMs that I have seen over the years. Fortunately, advanced features in modern translation memories, like memoQ's "TM-driven segmentation" encourage much better practice among smart service providers today.
  • The widespread notion that translation memory systems are only useful if one works on repetitive texts. I've got news for you: much of the repetitive stuff was outsourced to King Louie & Co. years ago. And yet I still find great value in working with good TMs. Why? A friend of mine summarized it nicely the other day when she talked about how she spent two hours researching a very obscure term for roadworks equipment in a minor European language: "The next time this comes up, I can find it right away and see the context." Indeed. I am amazed sometimes at the obscure technical terminology that comes out of my personal TM with its 12 year record of my work. Sometimes that amazement is even positive. An hour invested in researching a term and saving it in a TM (or much better: a proper termbase with metadata including domain´, source and examples of use) is probably several more hours saved over the next few years. At least.
  • The idea that a translation memory is a reliable source of terminology and obviates the need to create and maintain termbases or proper glossaries. Wrong, wrong, wrong. Particular offenders in this regard are agencies with their brothel-like practices of letting any number of translators screw the end customers' texts. Do a concordance search to find the right term in one of those TMs? Riiiiiiiiiight. Even agencies I've worked with for years who have made a real effort to keep TMs clean can't keep the terms in them on the straight and narrow. And using TMs to replace a real termbase, even a limited one, sacrifices the enormous potential benefits of automated terminology QA procedures offered by some modern translation environments.
  • King Louie & Co. as well as many other agencies in the race to the bottom of the quality barrel truly believe that once a good TM has been established by top translators, the second- or third-tier team can take over at lower cost and keep the customer happy. Well... at the moment, the lock on my Volvo's rear hatch is broken. I could get it fixed by a mechanic on Monday, or I could follow my neighbor's suggestion, and just hold it shut with a bungee cord. And the next time a tail light cover gets broken, I could just tape some red or yellow plastic film over it. Replace the hubcap that flew off when I hit that pothole? Naw. But sooner or later, people will notice the difference and draw their own conclusions. Will those be good for business? Can a jobbing student equipped with a good TM really produce the quality of legal translation you can rely on before the court? Trägt er auch 'nen gold'nen Ring, der Affe bleibt....
I have noted over the years, that most of the best clients never ask about translation memories or the tools related to them, even though a good number of them are aware of the technology and many of these use it. But these are the ones who understand that the monkeys who rely slavishly on CAT tools without the use of BAT* too often produce stale, stilted text unsuited for its communication purpose. At best. And all the king's machine translation engines won't change that.

Nonetheless, I believe there is great value for nearly all translators, even "creatives", in using advanced translation environment tools. But that value will not be in the same methods nor in the same features necessarily. Calls to "throw away your TMs" with the introduction of advanced alignment technologies like Kilgray's LiveDocs in memoQ, which allow final edited versions of past documents to be incorporated quickly when a new versions are to be translated, may be a bit premature, but they are often appropriate in my recent experience. And combinations of that with voice recognition technologies, term QA tools and other features offer a wealth of creative possibilities for taking the best and leaving the rest in our quest for better results and working conditions.


* brain-assisted translation

Nov 6, 2011

Transcribing with Dragon

One idea that fascinated me with Dragon Naturally Speaking is the potential for using the software to transcribe dictated messages on recording devices. I have often been interested in possibilities for translation away from the computer. In fact, the possibility of translating my by dictation is one that has a long history and which is of great personal interest to me, because I think that it may be possible to achieve a very different flow and quality from that which is typically achieved when writing on a computer. So, while visiting a friend, I purchased an inexpensive Olympus recording device for about €50, then I trained DNS to use it and began my experiment. 
The text above is my first attempt at transcribing from a small, portable recording device to Dragon Naturally Speaking. One small correction (marked). My expectations were not high when I decided to try this. The microphone in my headset is of rather modest quality, and the errors, as noted in my last post, can be devilish. But the recording device gave excellent results. A few other texts I have tested look equally good. So it seems my dream of dictating translation drafts off original printed texts or originals migrated to my Kindle may become true soon. In combination with memoQ LiveDocs, I see some very interesting potential translation and revision workflows. Stay tuned.

Nov 2, 2011

Enter the Dragon

For a number of years now, various colleagues of mine have sung the praises of Nuance's Dragon Naturally Speaking, making productivity claims that I occasionally suspected had their basis in illegal substance abuse. The work results I saw on one project a few years ago convinced me that the fellow had in fact been stoned.

I used the software briefly seven years ago during a rather unpleasant bout of RSI while I was traveling abroad and some of the keys on my laptop's keyboard began to fail, but Windows XP Service Pack 2 soon rendered the application unusable, and my memories of it weren't so great that I was inspired to have another look any time soon.

And so it remained until I had occasion to visit another colleague and see a "mixed mode" way of working with The Beast on complex legal texts. That got my attention. Particularly the quality of the results and an output well beyond my usual capacity. So I had another look.

First of all, one must be aware that DNS is dangerous. Even with good training and a high-quality microphone, it produces a number of errors, some bizarre, some very subtle, which may require a significant change in one's review workflow. For someone like me who is a miserable proofreader, this can be quite a challenge. Reading texts at top speed aloud doing a Donald Duck imitation seems to help. But for God's sake, don't rely on a casual silent read and a spellchecker.

I discovered that overall, my working speed, even with the considerable increase in review effort, improved significantly. It is faster to read terminology from my TEnT hit list than it is to insert it with a keyboard shortcut, and I think that looking at the source text on the screen more and thinking about it as I dictate at a slow, relaxed pace gives me a better, more natural text faster. The improvement in output isn't a matter of "typing speed" so much but rather that I spend more time reading the text. I am a hunt and peck typist, albeit a rather fast one, and I can keep pace with the fastest touch typists I have seen when I am working on a translation. The real bottleneck has never been typing time but rather thinking time, and I do not have the impression that a touch typist does any more thinking or does it faster.

When I shared my findings with a Dutch friend, an agency owner with nearly 30 years of experience as a translator, he told me about how the colleagues he knew long ago had mocked him for using the first generation of word processors, because they could dictate so much faster. However, back then a great deal of time was lost sending tapes to a typist and revising, and he could often deliver faster, though his personal time investment was perhaps greater. Now, he said, it seems that technology and tradition can be combined to produce the best result.

And now, just for laughs, the raw results of dictating the text above with a medium-quality microphone and an ignorance of various useful commands for quotation marks, etc.:

For a number of years now, various colleagues have sung the praises of nuances Dragon naturally speaking, making productivity claims I occasionally suspected had their basis in your legal substance abuse. The work results I saw on one project a few years ago convinced me at hello had in fact been stoned.
I used the software briefly seven years ago during a rather unpleasant out of RSI while I was traveling abroad and some of the keys on my laptop's keyboard began to fail, but windows XP service pack two soon rendered the application unusable, and my memories of it weren't so great that I was inspired to have another look anytime soon.
And so it remained until I had occasion to visit another call to and see a mixed mode way of working with the beast on complex legal texts. That got my attention. Particularly the quality of the results and an output well beyond my usual capacity. So I had another look.
First of all, one must be aware that DNS is dangerous. Even with good training and a high-quality microphone, it produces a number of errors, some bizarre, some very subtle, which may require a significant change in one's review work. For someone like me who is a miserable proofreader, this can be quite a challenge. Reading texts at top speed allowed doing a Donald Duck imitation seems to help. But for gods sake, don't rely on a casual site and read and spell check.
I discovered that overall, my working speed, even with the considerable increase in review effort, improved significantly. It is faster to read terminology from my tent yet list than it is to insert it with a keyboard shortcut, and I think that looking at the source text on the screen more and thinking about it as I dictate at a slow, relaxed pace gives me a better, more natural text faster. The improvement in output isn't a matter of typing speed so much but Robert that I spend more time reading the text. I am a hunt and peck typist, albeit a rather fast one, and I can keep pace with the fastest touch typist I have seen when I am working on a translation. The real bottleneck has never been typing time but rather thinking time, and I do not have the impression that a touch typist does anymore thinking or does it faster.
When I shared my findings with a Dutch friend, an agency owner with nearly 30 years of experience as a translator, he told me about how the colleagues he knew long ago had mocked him for using the first generation wordprocessed, because they could dictate so much faster. However, back then a great deal of time was lost sending tapes to a typist and revising, and he could often deliver faster, though his personal time investment was perhaps greater. Now, he said, it seems that technology and tradition can be combined to produce the best result.