Showing posts with label dictation. Show all posts
Showing posts with label dictation. Show all posts

Jan 28, 2020

Another look at Windows 10 speech recognition


A few years ago while on "holiday", I returned from dinner to find that my laptop had bluescreened. Panic time! It was Saturday night, and I still had quite a lot of text to translate and deliver on Monday morning. And up on the highest mountain in Portugal, I wasn't sure where I could find a replacement to finish the project, which was, at least, not utterly lost, because I had put it on a memoQ Cloud server for testing. The next day I got lucky: about 50 km away there was a Worten, where I picked up a gamer laptop with lots of RAM and an SSD. Well, not so lucky, as it was a Hewlett Packard Omen, with a fan prone to failure, but that's another story....

This new laptop was my first encounter with Windows 10. I had heard that this operating system offered improved speech recognition capabilities, and since I prefer to dictate my translations and downloading the 3 GB installation file for Dragon NaturallySpeaking (DNS) from my server at the office was going to take forever, I thought I would give Windows 10 speech recognition a try. I hadn't installed my CAT tool of choice yet, so I fired up Microsoft Word and began dictating. "Not bad," I thought. Then I tried it in my translation environment, and the results were a complete disaster. So I put that mess out of my mind.

Since then there have been some notable advances in speech-to-text capabilities on a number of platforms. But the best solution for my languages (German and English) with DNS became increasingly cranky thanks to neglect of the product by Nuance. Every week I read new reports of trouble with DNS in a variety of environments in which it used to perform very well. Apple's iOS 13 was a great leap forward of sorts for speech recognition and voice-controlled editing, but the new features are only available in English, and having Voice Control activated totally screws up my otherwise rather good dictation in German and Portuguese (or any other language). And don't get me started on the crappy vocabulary addition feature, which uses text entry alone with no link to actual pronunciation. Good luck with that garbage. It's not a bad solution in Hey memoQ with the additional command features added, but iOS dictation is not completely up to reasonable professional standards yet.

I probably would have given no further thought to Windows 10's speech-to-text features if it weren't for Anthony Rudd. We've corresponded a bit since I bought his excellent book on regular expressions for translators (and there's another practical guide for us coming soon from him!), and in a recent discussion he alluded to the use of Unicode with regex as a simple way of dealing with some things another colleague was struggling with. I was intrigued by this, and so for about half a day, I ran down a rabbit hole, testing Unicode subscripts and superscripts for a variety of purposes like fixing bad OCR of footnote markers and empirical formulae, autocorrecting common expressions for subscripted variables and chemical terms, including subscripts and superscripts in term bases and much more. Fascinating and useful stuff on the whole, even if some fonts don't support it well.

And of course I looked at using these special Unicode characters in speech-to-text applications. DNS had some funky quirks (not allowing numbers in the "spoken" version of terms, for example), but it worked rather well, so I can now say "calcium nitrate formula" and get Ca(NO₃)₂ without much ado. And for some reason it occurred to me to give Windows 10 speech recognition a try, just because I was curious whether vocabulary could in fact be trained. Indeed it can, and that feature is better than iOS 13 or DNS by far.

But first I had to remember how to activate speech recognition for Windows on my laptop again. When in doubt, type what you're looking for in the search box....

Notice I've pinned Windows Speech Recognition to my taskbar on the right, which is good for quick tasks.

Gesucht, gefunden
. Unlike other speech recognition solutions, the one in Windows 10 works only for the language set for the operating system. And options there are limited to English (United States, United Kingdom, Canada, India, and Australia), French, German, Japanese, Mandarin (Chinese Simplified and Chinese Traditional) and Spanish.

I put on my trusty Plantronics earset (the best microphone I've used for dictation tasks or audio in my occasional webinars in the past year) and began to dictate, first in Microsoft Word, which had shown acceptable results in my tests long ago. I found that adding vocabulary in the Speech Dictionary (accessed via the context menu in the dictation control element shown as a graphic at the top of this post) was dead simple.

The option to record pronunciation enabled me to record non-English names and words in several languages. And sure enough, the Unicode subscripts and superscripts worked, so I can now say CO₂ (I just dictated that) to my heart's content.

I was expecting a mess when I tried to use Windows 10 speech-to-text in a CAT tool, but it was not to be. It was brilliant, actually. I tried it in my copy of SDL Trados Studio, and with the scratchpad disabled so I could dictate directly into the target it worked well. No voice-controlled editing like I'm used to with DNS in memoQ, but that DNS feature does not work in SDL Trados Studio anyway, so this is no worse. But with the scratchpad box enabled (see the screenshot below), I could use voice commands to select and correct text or perform other operations. Brilliant!

After clicking or speaking "Insert", the text will be written to the target field with the proper formatting
So users of SDL Trados Studio who translate to a target language supported by Windows 10 speech recognition are probably better off not giving their money to Nuance, which I'm told can't even be bothered to make a 64-bit version of DNS now (which probably accounts for a lot of the trouble people have with that program.

I tested Wordfast Pro 5, which seems to confuse the speech recognition tool horribly, with source text displayed in the floating bar for some odd reason. But my earlier tests of Wordfast with DNS were equally unhappy, so somehow I'm not surprised. And I didn't test the Memsource desktop editor, which took the price a few years ago for the worst-ever DNS dictation results with a CAT tool. I'll leave that to someone with a much wider masochistic streak.

But what about memoQ, my personal environment of choice for most translation work? Equally brilliant, works just the same as SDL Trados Studio. No voice control for editing without the dictation scratchpad enabled (there, DNS has an advantage in memoQ), but with the scratchpad you can use the voice commands to edit before inserting in the target text field.


Wanna see this in action? Have a look at this short demo video:


I hope that the future will bring us more language support for Windows 10 dictation (Portuguese, Russian and Arabic, please!) and that other providers (like Google, if you're listening, and Apple, which never listens to anyone anymore except to spy on them with Siri) will expand the speech-to-text features offered, particularly to include sound-linked vocabulary training and better adaptation to individual users' speech. Five years ago when I began to investigate alternatives for non-DNS languages, I expected we would have more by now, and we do, but professional needs require all providers to raise their game.

Addendum: Someone asked me if Windows Speech Recognition is a cloud resource or a locally installed one which will work without an Internet connection. It's definitely the latter. So if you have lousy bandwidth or find yourself disconnected from the Internet, you can still use speech-to-text features.

And more: I use a lot of spoken commands for keyboard shortcuts when I work, so I did a little research and testing. It seems that Windows 10 speech recognition gives full access to an application's keyboard shortcuts via voice. So in memoQ, for example, I can dictate the insertion of tags, items from the Translation Results pane and a lot more. Watch out, Nuance. Windows 10 is going to kick your Dragon's scaly butt!

Jul 11, 2019

iOS 13: interesting options for dictators


Given the deteriorating political situation of many countries in the world today, the title of this post may seem ominous to some; however, the actual situation for those who use Apple's iOS operating system seems to call for some optimism in the months ahead. Among all the myriad feature changes in the upcoming Apple iOS 13 (now in the Public Beta 2 phase), there are a few which may be of particular value to writers and translators who dictate using their iOS devices.

Attention Awareness
This is 2019, and not only is Big Brother watching you, but your iPhone will as well. The rear-facing camera on some models will detect when you look away from the phone  perhaps to tell your dog to get off the couch  and switch off voice control. The scope of application for this feature isn't clear yet, and I have my doubts whether this would be relevant to more ergonomic ways of working with applications like Hey memoQ (which involve Bluetooth headsets or earsets to avoid directionality problems as the head may turn to examine references, etc.), but for some modes of text dictation work, this could prove useful. I have lost track of how often I've been interrupted by people and found my responses transcribed in one way or another, often as an amusing salad of errors when I switch languages.

Automatic language selection in Dictation
The iOS 13 features preview from Apple states, "Dictation automatically detects which language a user is speaking. The language will be chosen from the keyboard languages enabled on the device, up to a maximum of four." Well, well. I wonder how it will handle isolated sentences or paragraphs quoted in another language  or individual foreign words. I'm betting probably not. But I'll have great fun pushing this feature around with three or four spoken languages to find its limits.

Add custom words
This is what I have wanted for years. Custom audio recognition vocabulary  words and phrases  to ensure that unusual or specialist terms are recognized and transcribed correctly. BINGO!

On-device processing
All audio processing will be handled locally (on your iPhone or iPad), ensuring privacy if you believe the NSA and/or the Russians or other parties aren't tapped into your equipment.

Enhancements to voice editing and voice-driven app control
There are a lot of these. Read about them in the Accessibility section of the features description from Apple. My first impression of these possibilities is that editing and correcting text may become much easier on iOS devices, and the attractiveness of the three-stage dictation/alignment/pretranslation workflow may increase for some translators. (An old example of this is in an old YouTube video I prepared years ago for a remote conference presentation, but the procedure works with any speech-to-text options and has the advantage of at least two revision steps.)

It's even more interesting to consider how some of these new features might be harnessed by apps designed to work with translation assistance environments. And if Google responds - as I believe the company is likely to do - with new features for Chrome speech recognition and voice control features in Android and desktop computers, then there could be some very, very interesting things ahead for wordworkers in the next year or two. Vamos ver!


Feb 4, 2019

Review: the Plantronics Voyager Legend monoaural headset for translation

Ergonomics are often a challenge with the microphones used for dictated translation work. I've used quite a few over the years, usually with USB connections to my computer, though I've also had a few Logitech wireless headphones with integrated mikes that performed well. However, all of them have had some disadvantages.

The country where I live (Portugal) has a rather warm climate for more than a few months of the year. Wearing headphones can get rather uncomfortable on a hot day, and even on a cold one, the pressure on my ears starts to drive me nuts after an hour or so.

Desktop microphones seem like a good solution, and I get good results with my Blue Yeti. But sometimes, when I turn my head to look at something, the pickup is not so good, and my dictation is transcribed incorrectly.

The Hey memoQ app released by memoQ Translation Technologies Ltd. underscored the ergonomic challenges of dictation for me; the app uses iOS devices to access their speech recognition features, and positioning a phone well in such a way that one can still make use of a keyboard is not easy. And trying to connect a microphone or headset by cable to the dodgy Lightning port on my iPhone 7 is usually not a good experience.

So I was intrigued by a recommendation of Plantronics headsets from Dragos Ciobanu of Leeds University (also the author of the E-learning Bakery blog). A specific model mentioned by someone who had attended a dictation workshop with him recently was the Plantronics Voyager Legend, though when I asked Dragos about his experience, he spoke mostly about the Plantronics Voyager 5200, which is a little more expensive. I decided to go "cheap" for my first experience with this sort of equipment and ordered the Voyager Legend from Amazon in Spain. I did so with some trepidation, because the reviews I read were not entirely positive.


The product arrived in simple packaging which led me to think that the Amazon review which suggested the "new" products sold might in fact be refurbished. But in the EU, all electronic gear comes with a two-year warranty, so I don't worry too much about that.

Complaints I read in the reviews about a short charger cable seem ridiculous; the cable I received was over half a meter long, and like anyone who has computers these days, I have more USB extension cords than I know what to do with should I require a longer cable for charging. The magnetic coupler for charging has a female mini-USB port, so it can be attached to another cable as well. Power connections include the most common EU two-pronged charger, the 3-pole UK charger and one for a car's cigarette lighter.

The package also included extra earpieces and covers of different sizes to customize the fit on one's ear.

I tested the microphone first with my laptop; the device was recognized easily, and the results with Dragon NaturallySpeaking were excellent. Getting the connection to my iPhone 7 proved more difficult, however. I read the Getting Started instructions carefully, tried updating the firmware (not necessary - everything was current) and tried various switching and reboot tricks, all to no avail.

Finally, I called the technical support line in the US in total frustration. I didn't expect an answer since it was still the wee hours of the morning in the US, but someone at a support call center did answer the phone. He instructed me to press and hold the "call" button on the device until its LED begins to flash blue and red.


I did that, and when the LED began flashing, "PLT_Legend" appeared in the list of available devices on my iPhone. Then I was ready to test the Voyager Legend for dictated translation with Hey memoQ.

Because I work with German and English, I rely on Dragon NaturallySpeaking for my dictation, and the iOS-based dictation of Hey memoQ will never compete with that. But I am very interested in testing and demonstrating the integrated memoQ app, because many other languages, such as Portuguese, are not available for speech recognition in Dragon NaturallySpeaking or any other readily accessible speech recognition solution of its class.

As I suspected, my dictation in Hey memoQ (and other iOS applications) was easier with the Voyager Legend. This is the first hardware configuration I have tested that really seems like it would offer acceptable ergonomics for Hey memoQ with my phone. And I can use it for Skype calls, listening to my audio books and other things, so I consider the Plantronics Voyager Legend to be money well spent. Now I'll see how it holds up for long sessions of dictated legal translation. The product literature and a little voice in my ear both claim that the device can operate for seven hours of speaking time on a battery charge, and the 90 minutes required for a full recharge will work well enough with the breaks I take in that time anyway.

Of course there are many Bluetooth microphone devices which can be used with speech recognition applications, but what distinguishes this one is its great comfort of wear and the secure fit on my ear. I look forward to a closer acquaintance.

Jan 6, 2019

A voice-activated recorder for iOS


Some years ago I picked up an Olympus hand-held digital recorder which served me well for some translation work as well as for evidentiary purposes during assaults by a drunken psychotic. I also used it on many occasions to record ideas for projects and other things while out and about.

But juggling several electronic devices at once has never been easy for me, and I kept losing the little recorder in jacket pockets, suitcases, desk drawers, etc. I still have it, but it's been some weeks since I could tell you where.

I use the Voice Memos app on my iPhone, but for me its continuous recording feature is inconvenient, and I don't like pausing and resuming the recording frequently. It's too distracting. The Olympus device and an old tape recorder I used decades ago both had a convenient voice activation feature which avoided excessive dead space.

Being the technosaurus that I am, it took a long time to realize that there is probably an app for that these days. And indeed there is. Several in fact. I downloaded the iOS app depicted above, and in the days ahead I'll be using it for a few translation projects to evaluate what I consider to be an improved version of the three-step translation workflow I demonstrated a few years ago in a remote conference lecture in Buenos Aires using the now-defunct Dragon Dictation app from Nuance. Stay tuned.

Jan 3, 2019

Using analog microphones with newer iPhones


Microphone quality makes a great difference in the quality of speech recognition results. And although the microphones integrated in iOS devices are generally good and give decent results, positioning the device in a way that is ergonomically viable for efficient dictated translation - and concurrent keyboard use - is not always so easy. This is a potential barrier to the effective use of Hey memoQ speech recognition.

So a good external microphone may be needed. But with recent iPhone models lacking a separate microphone jack and using the lightning port for both charging and microphone input, connecting that external microphone might not be as simple as one assumes. Especially not someone like me, who is rather ignorant of the different kinds of 3.5 mm audio connections. I have had a few failures so far trying to link my good headset to the iPhone 7.

Colleague Jim Wardell is not only the most experienced speech recognition expert for translation whom I am privileged to know; he is also a musician with extensive experience in microphones of all kinds and their connections. And recently he was kind enough to share the video below with me to clear up some misunderstandings about how to connect some good analog equipment to use with Hey memoQ on an iPhone 7 or later:


Jan 2, 2019

Hacking the "Hey memoQ" dictation commands


In the initial release of the Hey memoQ dictation feature in memoQ version 8.7.3, it's a bit inconvenient to deal with command configuration. Unlike most configurations in memoQ, the dictation commands cannot yet be exported as a light resource and shared with other users, nor can a configuration for generic German, for example, be easily transferred to a desired variant such as "ger-DE" or "ger-CH". Surely this will be addressed soon, but at the moment it's a bit of a nuisance.

But fear not... there is usually a backdoor to hack memoQ configurations, and this is no exception.


The screenshot above shows the path to the current configuration file for dictation commands. The XML file contains all the configured commands for all the memoQ languages and variants, including those of no interest whatsoever.

Deep inside the Hey memoQ dictation command file with Notepad++

A peek inside the XML file reveals that the dictation commands are structured as key-value pairs. And here it is possible to enter the text for dictation commands, simply by typing the desired text between the string tags inside the Value tags.

A configuration (Commands set) for one variant of a language - such as generic Portuguese - can also be copied to other variants - such as Brazilian or European Portuguese, saving the trouble of re-entering everything laboriously in the configuration dialog within memoQ.

I made a copy of the XML configuration and edited it to have only the variants of English and German that were of interest to me. Then I copied this file over the one in the memoQ configuration directory shown in the screenshot above. When I restarted memoQ, the file bloated a bit; upon examining it, I saw that all the deleted languages had been restored after the ones I had left in the edited file, but the new file was still only 247 KB in size because the senseless copying of English commands to the other languages was gone.

A customized XML file can be shared with other users, who can use it to replace the existing configuration file and probably save time configuring their languages and variants of interest. My file with generic English, EN-US, EN-UK, generic German and DE-DE is here.


Dec 11, 2018

Your language in Hey memoQ: recognition information for speech

There are quite a number of issues facing memoQ users who wish to make use of the new speech recognition feature – Hey memoQ – released recently with memoQ version 8.7. Some of these are of a temporary nature (workarounds and efforts to deal with bugs or shortcomings in the current release which can reasonably be expected to change soon), others – like basic information on commands for iOS dictation and what options have been implemented for your language – might not be so easy to work out. My own research in this area for English, German and Portuguese has revealed a lot of errors in some of the information sources, so often I have to take what I find and try it out in chat dictation, e-mail messages or the Notes app (my favorite record-keeping tool for such things) on the iOS device. This is the "baseline" for evaluating how Hey memoQ should transcribe text in a given language.

But where do you find this information? One of the best way might be a Google Advanced Search on Apple's support site. Like this one, for example:


The same search (or another) can be made by adding the site specification after your search terms in an ordinary Google search:


The results lists from these searches reveal quite a number of relevant articles about iOS dictation in English. And by hacking the URLs on certain pages and substituting the language code desired, one can get to the information page on commands available for that language. Examples include:
All the same page, with slightly modified URLs.

The Mac OS information pages are also a source of information on possible iOS commands that one might not find so easily otherwise. An English page with a lot of information on punctaution and symbols is here: https://support.apple.com/en-us/HT202584

The same information (if available) for other languages is found just by tweaking the URL:

and so on. Some guidance on Apple's choice of codes for language variants is here, but I often end up getting to where I want to go by guesswork. The Microsoft Azure page for speech API support might be more helpful to figure out how to tweak the Apple Support URLs.

When you edit the commands list, you should be aware of a few things to avoid errors.
  • The current command lists in the first release may contain errors, such as mistakenly typing "phrase" in angular brackets as shown in the first example above; on editing, the commands that are followed by a phrase do not show the placeholder for that phrase, as you see in the example marked "2".
  • Commands must be entered without quotation marks! Compare the marked examples 1 and 2 above. If quotes are typed when editing a command, this will not be revealed by the appearance of the command; it will look OK but won't work at all until the quote marks are removed by editing.
  • Command creation is an iterative process that may entail a lot of frustrating failures. When I created my German command set, I started by copying some commands used for editing by Dragon NaturallySpeaking, but often the results were better if I chose other words. Sometimes iOS stubbornly insists on transcribing some other common expression, sometimes it just insists on interpreting your command as a word to transcribe. Just be patient and try something else.
The difficulties involved in command development at this stage are surely why only one finished command set (for the English variants) for memoQ-specific commands was released at first. But that makes it all the more important to make command sets "light resources" in memoQ, which can be easily exported and exchanged with others.

At the present stage, I see the need for developing and/or fixing the Hey memoQ app in the following ways:
  • Fix obvious bugs, which include: 
  • The apparently non-functional concordance insertions. In general, more voice control would be helpful in the memoQ Concordance.
  • Capitalization errors which may affect a variety of commands, like Roman numerals, ALL CAPS, title capitalization (if the first word of the title is not at the start of the segment), etc.
  • Dodgy responses to the commands to insert spaces, where it is often necessary to say the command twice and get stuck with two spaces, because a single command never responds properly by inserting a space. Why is that needed? Well, otherwise you have to type a space on the keyboard if you are going to use a Translation Results insertion command to insert specialized terminology, auto-translation rule results, etc. into your text. 
  • Address some potentially complicated issues, like considering what to do about source language text handling if there is no iOS support for the source language or the translator cannot dictate commands effectively in that language. I can manage in German or Portuguese, but I would be really screwed these days if I had to give commands in Russian or Japanese.
  • Expand dictation functionality in environments like the QA resolution lists, term entry dialog, alignment editor and other editors.
  • Look for simple ideas that could maximize returns for programming effort invested, like the "Press" command in Dragon NaturallySpeaking, which enables me to insert tags, for example, by saying "Press F9". This would eliminate the need for some commands (like confirmation and all the Translation Results insertion commands) and open up a host of possibilities by making keyboard shortcuts in any context controllable by voice. I've been thinking a lot about that since talking to a colleague with some pretty tough physical disabilities recently.
Overall, I think that Hey memoQ represents a great start in making speech recognition available in a useful way in a desktop translation environment tool and making the case for more extensive investments in speech recognition technology to improve accessibility and ergonomics for working translators.

Of course, speech recognition brings with it a number of different challenges for reviewing work: mistakes (or "dictos" as they are sometimes called, a riff on keyboard "typos") are often harder to catch, especially if one is reviewing directly after translating and the memory of intended text is perhaps fresh enough to override in perception what the eye actually sees. So maybe before long we'll see an integrated read-back feature in memoQ, which could also benefit people who don't work with speech recognition. 

Since I began using speech recognition a lot for my work (to cope with occasionally unbearable pain from gout), I have had to adopt the habit of reading everything out loud after I translate, because I have found this to be the best way to catch my errors or to recognize where the text could use a rhetorical makeover. (The read-back function of Dragon NaturallySpeaking in English is a nightmare, randomly confusing definite and indefinite articles, but other tools might be usable now for external review and should probably be applied to target columns in an exported RTF bilingual file to facilitate re-import of corrections to the memoQ environment, though the monolingual review feature for importing edited target text files and keeping project resources up-to-date is also a good option.)

As I have worked with the first release of Hey memoQ, I have noticed quite a few little details where small refinements or extensions to the app could help my workflow. And the same will be true, I am sure, with most others who use this tool. It is particularly important at this stage that those of us who are using and/or testing this early version communicate with the development team (in the form of e-mail to memoQ Support - support@memoq.com - with suggestions or observations). This will be the fastest way to see improvements I think.

In the future, I would be surprised if applications like this did not develop to cover other input methods (besides an iOS device like an iPhone or iPad). But I think it's important to focus on taking this initial platform as far as it can go so that we can all see the working functionality that is missing, so that as the APIs for relevant operating systems develop further to support speech recognition (especially the Holy Grail for many of us, trainable vocabulary like we have in Dragon NaturallySpeaking and a very few other applications). Some of what we are looking for may be in the Nuance software development kits (SDKs) for speech recognition, which I suggested using some years ago because they offer customizable vocabularies at higher levels of licensing, but this would represent a much greater and more speculative investment in an area of technology that is still subject to a lot of misunderstanding and misrepresentation.

Nov 9, 2018

Chrome speech recognition in all your Windows and Linux applications

In a recent social media discussion, a Slovenian colleague was asking me about the upcoming hey memoQ feature that I've been testing, and I found that iOS apparently doesn't support that language (nor does MacOS for that matter). But then she commented
I use Chrome's voice notebook plugin with memoQ. It works somehow for a while, then it gets laggy and I have to refresh Chrome or restart it. I miss the consistency and learning ability of DNS. But yes, the paid version allows you to use it with any app, including memoQ. The free version does not have this functionality. I love translating with dictation, I am not a fast typist and I rather hate typing...
I had no idea what she was talking about, but a few more questions and a little investigation cleared up my confusion. Some years ago when Chrome's speech recognition feature was introduced, it seemed to me that it should be possible to adapt it for use in other (non-browser) applications, and I think this was even stated as a possibility. But at the time I could not find any application to do this, and I'm too out of practice these days to program my own.

Well, it seems that someone has addressed this deficiency.


The voice to text notebook extension of Chrome has additional tools available on the creator's website which enable the speech recognition functions to be used in any other application. This additional functionality is a service with fees, but at USD 2.50 per month or USD 16.00 per year (via PayPal), it's unlikely to break the bank. And a free trial of two days can be activated once you have registered. I'm testing it now, and it's rather interesting. Not perfect (as noted by the colleague who made me aware of this tool), but it may be an option for those wanting to use speech recognition in languages not currently supported by other applications.

Jun 3, 2018

Survey for Translation Transcription and Dictation

The website with the survey and short explainer video is http://www.sightcat.net
The idea is to build a human transcription service. We just need a few translators per language that want to work with a transcriptionist due to RSI, productivity etc. and we can use that data to build an ASR system for that language. There is also a good chance the ASR system will be accurate for domain-specific terminology and accents as it will be adaptive and use source language context. 
Take the Sight CAT survey - click here
Click on the graphic to go to the survey
memoQ Fest 2018 was, among other things, a good opportunity as always to spend time discussing things with some of the best and most interesting consultants, teachers, creative developers and brainstormers I know in the translation profession. One of these was my friend and colleague, John Moran, whose work on iOmegaT introduced me to the idea that properly designed, translator-controlled (voluntary) data logging could be a great boon to feature research and software development investment decisions. Sort of like SpyGate in translation, except that it isn't.

John and I have been talking, brainstorming and arguing about many aspects of translation technology for years now, dictation (voice recognition, ASR, whatever you want to call it) foremost among the topics. So I was very pleased to see him at the conference in Budapest last week, where he spoke about logging as a research tool in the program and a lot about speech recognition before and after in the breaks, bars, coffee houses and social event venues.

I think that one of the most memorable things about memoQ Fest 2018 was the introduction of the dictation tool currently called hey memoQ, which covers a lot of what John and I have discussed until the wee hours over the past four years or so and which also makes what I believe will be the first commercial use of source text guidance for target text dictation (not to mention switching to source text dictation when editing source texts!). John introduced that to me years ago based on some research that he follows. Fascinating stuff.

One of the things he has been interested in for a while for commercial, academic and ergonomic reasons is support for minor languages. Understandable for a guy who speaks Gaelic (I think) and has quite a lot of Gaelic resources which might contribute to a dictation solution some day. So while I'm excited about the coming memoQ release which will facilitate dictation in a CAT tool in 40 languages (more or less, probably a lot more in the future), John is thinking about smaller, underserved or unserved languages and those who rely on them in their working lives.

That's what his survey is about, and I hope you'll take the time to give him a piece of your mind... uh, share your thoughts I mean :-)

The Great Dictator in Translation.

I have no need for words. memoQ will have that covered in quite a few languages.


This is not your grandfather's memoQ!

Jun 24, 2017

Germany needs Porsches! And Microsoft has the Final Solution....


I hear that Germany is suffering from a shortage of Porsches. Odd, given that the cars are made there and should be readily available, but it's true, because my friend who lives there told me. He owns a large, successful LSP (Linguistic Sausage Production) company, and to celebrate its rise in revenues, he decided to get everyone on the sales staff a new Porsche as a company car. The problem is that he can't find any for €5000 euros.

So he was left with no choice but to cut overhead using the latest technologies. Microsoft to the rescue! With Microsoft Dictate, his crew of  intern sausage technologists now speak customer texts into high-quality microphones attached to their Windows 10 service stations, and these are translated instantly into sixty target languages. As part of the company's ISO 9001-certified process, the translated texts are then sent for review to experts who actually speak and perhaps even read the respective languages before the final, perfected result is returned to the customer. This Linguistic Inspection and Accurate Revision process is what distinguishes the value delivered by Globelinguatrans GmbHaha from the TEPid offerings of freelance "translators" who won't get with the program.

But his true process engineering genius is revealed in Stage Two: the Final Acquisition and Revision Technology Solution. There the fallible human element has been eliminated for tighter quality control: texts are extracted automatically from the attached documents in client e-mails or transferred by wireless network from the Automated Scanning Service department, where they are then read aloud by the latest text-to-speech solutions, captured by microphone and then rendered in the desired target language. Where customers require multiple languages, a circle of microphones is placed around the speaker, with each microphone attached to an independent, dedicated processing computer for the target language. Eliminating the error-prone human speakers prevents contamination of the text by ums, ahs and unedited interruptions by mobile phone calls from friends and lovers, so the downstream review processes are no longer needed and the text can be transferred electronically to the payment portal, with customer notification ensuing automatically via data extracted from the original e-mail.

Major buyers at leading corporations have expressed excitement over this innovative, 24/7 solution for globalized business and its potential for cost savings and quality improvements, and there are predictions that further applications of the Goldberg Principle will continue to disrupt and advance critical communications processes worldwide.

Articles have appeared in The Guardian, The Huffington Post, The Wall Street Journal, Forbes and other media extolling the potential and benefits of the LIAR process and FARTS. And the best part? With all that free publicity, my friend no longer needs his sales staff, so they are being laid off and he has upgraded his purchase plans to a Maserati.

Aug 18, 2015

Enter the Dragon, Anywhere!

Today Nuance made a presentation of a new product to be released this autumn for mobile devices using iOS and Android operating systems: Dragon Anywhere.
 

The mobile app will allow secure transcription with a WiFi or cellular data connection as well as synchronization of custom vocabulary with Dragon desktop computer applications wor Windows and MacOS.

The initial presentation made no mention of which languages will be available in Dragon Anywhere; the synchronization feature makes me worry that it might be restricted to the current seven or eight languages available for Dragon NaturallySpeaking (Windows) or Dragon Dictate (Mac), but perhaps the standalone applications for desktop computers and laptops will finally be upgraded to offer the 40+ languages currently available for mobile devices with apps such as Dragon Dictation for iOS or Swype + Dragon Dictation for Android.

It is also unclear at this point whether the new Dragon Anywhere app will allow direct dictation or transfer to the cursor location on a linked computer, as one can do with myEcho or using Swype + Dragon Dictation in conjunction with Chrome Remote Desktop. But with the addition of extensive voice-controlled editing features to the mobile app, Dragon Anywhere represents significant progress toward better ergonomics for writing and translation!

May 31, 2015

Authoring and Editing with memoQ (webinar)

Last February I described my initial work with translation tools as environments for authoring and editing documents in a single language. Some people have been doing this quietly for a while; occasionally I would hear puzzled comments from a trainer who had held a class on SDL Trados Studio, OmegaT or memoQ which had been attended by a technical writer or someone with other professional writing interests not related to translation. But to my knowledge there has been no systematic approach to this.

Some weeks later I began to discuss and present some new possibilities for speech recognition in 38 languages which go well beyond the limitations of Dragon NaturallySpeaking for automated speech transcription in the eight languages for which it is available. These possibilities include a number of mobile solutions which are quickly gaining traction among translators and other professional writers.

On Tuesday, June 2nd (two days from now), I will be presenting a one-hour introduction to "MemoQ for Single-language Authoring and Editing" in the eCPD Webinar series. The registration page is here.

This presentation will be an update of the talk I gave earlier this year which discussed CAT tools in general as authoring and editing tools. Although any tool works in principle (and even a user of SDL Trados Studio, for example, can probably draw enough ideas from the upcoming eCPD talk to make good use of the approach), memoQ has some particular advantages, not the least due to its corpus-handling features in LiveDocs and its superior predictive typing facilities, including "Muses" (which are like SDL's AutoSuggest with more flexibility and without the onerously high data quantity requirements).

The presentation will include an overview of some of the latest advances in speech recognition in 38 languages for ergonomically superior writing by automated transcription as well as discussions of version management and dictation workflows which can be applied for greater ease in editing monolingual documents or even translations, including post-editing of machine pseudo-translation (PEMpT by the "Hardisty Method"). I've been fairly quiet on this blog in recent months due to conference organization and travels and the considerable time put in to researching improved work ergonomics for translation, writing and editing processes. (In fact I didn't even find time to blog the memoQ Day on April 22nd in Lisbon yet!) Elements of all these efforts, which have sparked no little interest at recent conferences and workshops I have presented at in Europe, will be part of Tuesday's talk, which will include Q&A afterward to explore the interests of those participating.

So if you are a translator involved in a lot of revision or editing work (bilingual or monolingual, a technical writer or other professional writing in a single language for publication, someone working on a thesis or authoring for other purposes, the eCPD presentation may help you to do this with better organized resources and greater efficiency. As one friend of mine who wrote a thesis just before I developed this approach put it, with this she would at least have been able to keep track of the feedback on her work from its five or so reviewers without going completely nuts.

Apr 20, 2015

Máquinas virtuais para traduzir e ditar no Apple OS/X para Totós (ou como os Macs são os melhores amigos dos tradutores Portugueses)


RESUMO

O OS/X, o sistema operativo da Apple, inclui nas suas últimas versões a aplicação “Ditado”, que permite ditar para o computador numa panóplia de diferentes idiomas, com a dita máquina a transcrever o texto ditado para o ecrã.

Neste momento, poderão eventualmente pensar “Ah, pois, e tal, isso é o que o Dragon NaturallySpeaking faz!”. Pois é. Mas a versão actual do DNS já não suporta Português, e a aplicação do OS/X, com base no software da mesma Nuance que produz o DNS, suporta Português (deste e do outro lado do Atlântico), Alemão (Suíço e o da Alemanha ), Checo, Chinês (três variantes), Coreano (da Coreia do Sul - assumo que quem estava a desenvolver a versão do Norte foi raptado pelo regime local), Dinamarquês, Eslovaco, Espanhol (três variantes), Finlandês, Francês (três variantes), Grego, Holandês, Húngaro, Indonésio, Inglês (quatro variantes), Italiano (duas variantes), Japonês, Malaio, Norueguês, Polaco, Romeno, Russo, Sueco, Tailandês, Turco, Ucraniano e Vietnamita.

O sistema operativo faz isto descarregando dicionários acústicos com informação acerca do sotaque e idioma esperado, que são processados localmente com base nessa informação. A melhor parte é esta: podem instalar os que quiserem.

Trabalhando eu essencialmente com Português e Inglês, e sendo detentor de uma licença do Dragon Naturally Speaking, pude comparar a funcionalidade de um e outro. No entanto, a maioria das ferramentas CAT não funciona em ambiente OS/X, pelo que se torna necessário recorrer a máquinas virtuais a correrem o Windows para beneficiar simultaneamente das potencialidades acrescidas oferecidas pela Apple e da base comum oferecida pela plataforma da Microsoft.

Assim, estabeleceu-se Inicialmente, uma base comum para avaliar o software de visualização que permita beneficiar do ditar juntamente com aplicações CAT.  Foram avaliados dois pacotes de software de virtualização com exactamente a mesma máquina virtual, e testado o desempenho e funcionalidade em ambos.

INTRODUÇÃO, OU COMO RAIO ACABEI A FORMATAR O MEU PORTÁTIL E A FAZER EXPERIÊNCIAS COM ELE EM NOME DA CIÊNCIA

Há alguns meses atrás, o Kevin Lossner falou-me de um desenvolvimento levado a cabo por William Hardisty e Joana Bernardo, uma aluna sua, que estavam a utilizar a transcrição automática de texto ditado  incluída no OS/X como uma ferramenta de acessibilidade e aumento da produtividade num contexto de tradução. Aliás, esta mesma abordagem foi descrita e discutida num evento na Faculdade de Letras da Universidade de Lisboa, o SDL Day, em 22 de Janeiro de 2015.

A utilização deste recurso não é nada de novo -  a Nuance construiu a reputação da sua empresa com base na família de produtos “Dragon”. Infelizmente, na última versão (que utilizo num PC para a escrita de documentos originais em Inglês, relacionados com os meus objectivos académicos), pude constatar que o Português tinha sido abandonado.

Felizmente, a Apple assegurou que o seu sistema operativo tinha esta funcionalidade. Agora a questão era como implementar essa funcionalidade com as ferramentas disponíveis, a maioria em ambiente Windows.

David Hardisty e Joana Bernardo relataram a sua utilização recorrendo ao Parallels Desktop, uma ferramenta de virtualização para o OS/X, que permitia utilizar este recurso num ambiente instalado na máquina virtual - que pode correr o Windows, entre outros sistemas operativos.

Este artigo debruça-se assim sobre a instalação dos dois  pacotes de software - VMWare Fusion e Parallels Desktop - o seu desempenho e outras observações relevantes.

1 - QUAL O SOFTWARE DE VIRTUALIZAÇÃO A UTILIZAR?


Como referido anteriormente, este artigo debruça-se essencialmente sobre o VMWare Fusion 7 e Parallels Desktop 10, as versões actuais à data da escrita destas palavras. Existem diversas outras aplicações de virtualização, tendo eu utilizado o VirtualBox da Oracle desde o momento em que adquiri a máquina utilizada neste teste comparativo.

Infelizmente, esta última aplicação não permitia a utilização destas funcionalidades pelo que foi imediatamente excluída. É uma pena, porque ao contrário dos outros, é gratuito. Para além disso, serviu-me fielmente até ao momento em que limpei o portátil.

Quer o software da VMWare quer o da Parallels são produtos comerciais, e ambos mencionam a integração da funcionalidade de ditado com o seu sistema de virtualização como argumentos para a sua utilização.

Portanto a questão que se coloca é: funcionam realmente, e qual dos dois é o melhor para este trabalho?


2 - INSTALAÇÃO


Não vos vou maçar com pormenores aborrecidos - digamos que aproveitei a oportunidade para limpar o portátil, tendo-o formatado e feito uma instalação do OS/X Yosemite. A máquina em que foi testado foi o meu fiel Macbook Air de 2013, com 8GB RAM, i5 1.3 GHz e um SSD de 120 GB.

A instalação de ambas as aplicações é bastante simples, e o diálogo para a criação da máquina virtual é praticamente idêntico. Ambos pedem para registarem/comprarem a aplicação, e ambos permitem testá-la - por 14 dias no caso do Parallels Desktop, por 30 no caso do VMware Fusion.


 Como estava a criar uma máquina nova, foi esta a opção que escolhi com o Parallels
 Apontei para um ficheiro ISO (os Macbook Air não têm leitor de CD...)...
 ... e escolhi a opção "Produtividade" para predefinir a configuração da máquina virtual.
Nota importante: Se quiserem configurar a instalação do Windows, desmarquem "Instalação expressa". 
Dei comigo a instalar o Windows em Alemão antes de cancelar a coisa. Eu fiz a activação do Windows no interior da máquina virtual.  
 
Eh... é o Windows 7, pá!
 
E cá está ele a instalar...
... e com o Windows a funcionar, a ditar para o Bloco de Notas!
A instalação do VMWare Fusion é muito semelhante. Eu clonei a máquina Virtual do Parallels para garantir uma igualdade entre ambos. Em todo o caso não registem a aplicação! Eu esperei uns dias e comecei a ver publicidade a oferecer descontos no Parallels, na loja online da empresa :)

3 - CONFIGURAÇÃO DA MÁQUINA VIRTUAL

Ambos os sistemas utilizaram exactamente a mesma configuração: 30 GB de disco, 4 GB RAM, Windows 7 Ultimate, Office 2010 (é a versão que tenho licenciada para PC), Mozilla Firefox e memoQ 2014R2. Foi ainda instalada a aplicação Novabench para medir o desempenho de ambas as máquinas virtuais, a correrem sob definições de desempenho máximo, com o portátil ligado à corrente.

4 - DESEMPENHO E RESULTADOS

Logo após a instalação de qualquer uma das máquinas, o ditado fica imediatamente disponível. Basta utilizar a combinação de teclas predefinida (ou outra) e este começa imediatamente a funcionar, com pouco mais atraso que quando utilizado nativamente no OS/X. Todos os comandos estão disponíveis, e pode-se mudar de idioma através do menu flutuante que acompanha o cursor da escrita. Uma vez que o reconhecimento e voz é nativo do OS/X e transposto para a máquina virtual, não há diferenças imediatamente evidentes entre ambas, embora as medições de desempenho contem outra história.

A avaliação de desempenho do Windows 7 dá uma pontuação mais alta ao Parallels Desktop, em virtude essencialmente do desempenho gráfico.

Parallels Desktop...
...e o VMWare.
Já o Novabench dá o resultado exactamente oposto, com o VMware Fusion a ter um desempenho gráfico marcadamente melhor, embora pior em tudo o resto, especialmente no desempenho do processador.

Parallels Desktop... 
...e VMWare Fusion.

Na prática, durante o período em que o testei, não notei diferenças - admito no entanto que isso possa ser diferente à medida que se utilizam mais e mais recursos do sistema através da máquina virtual.

Quanto ao desempenho da aplicação “Ditado” - é boa. Muito boa, na verdade. Consigo velocidades superiores às do DNS em Inglês, embora tenha de editar mais frequentemente o texto. O “Ditado” requer uma dicção calma, monocórdica e pausada. No seu melhor, resulta em frases de 40, 50 palavras com pouquíssimos erros (um ou outro do dicionário, e costuma ter problemas com artigos que soem ao fim ou início das palavras antecedentes ou seguintes. Também é aborrecido colocar a primeira letra sempre em maiúsculas.

Em termos de velocidade pura, para quem está habituado a utilizar o Dragon juntamente com o teclado e rato (a que o Kevin chama “mixed mode”) a diferença em velocidade e ergonomia é gritante. Este passou a ser o meu equipamento de eleição para tradução.

Descobri que numa casa em que duas crianças pequenas, dois cães e uma gata correm, em números constantemente variáveis entre si, atrás uns dos outros, o velho auricular do meu iPhone 4 me permitia uma melhor qualidade de som ao colocar o microfone mais perto da minha boca e mais longe da algazarra doméstica.

Descobri ainda que existe uma variedade interessante de comandos, não na aplicação “Ditado”, mas na opção Acessibilidade das Preferências do Sistema, sob a opção “Ditado”. Podem activar a opção de Comandos Avançados (que é pelo menos tão detalhada quanto a do DNS), ou criar os vossos próprios comandos.

Localização das opções para a ferramenta Ditado e configuração de comandos da mesma.
Localização da opção "Ditado" no menu de Acessibilidade... 
 ... e a lista de comandos disponíveis (são muitos, mesmo muitos). Ao clicarem no "+" podem criar os vossos próprios comandos!

CONCLUSÃO

Bem, tem que ter um fim, não? Pessoalmente, não notei diferenças entre o VMWare e o Parallels quanto ao ditar o texto e em velocidade, apesar dos resultados dos testes de desempenho.
No entanto, o Paralells oferece algumas coisas interessantes, tais como integração directa com o ambiente do Yosemite (podem copiar ficheiros de e para a Máquina Virtual de dentro dela ou fora).
Isto é-me particularmente útil porque assim só preciso de ter uma cópia do conteúdo local da minha Google Drive, ao invés de uma em cada máquina, virtual ou verdadeira.

Gostei ainda da possibilidade de lançar directamente as aplicações da máquina virtual a partir do Launchpad.


Dito isto, e considerando o desconto, decidi-me pelo Parallels Desktop. Considerando a minha utilização da Google Drive para questões administrativas, dá-me um jeitaço ter apenas de sincronizar uma cópia da mesma no OS/X, a qual é acedida pela máquina virtual.

Contudo, o VMWare é uma boa opção, e a diferença de desempenho não me foi notória, pelo que suponho seja uma questão de escolherem conforme a vossa preferência. O VMware é mais barato (o preço fica parecido com os códigos de desconto), mas permite a instalação em três máquinas, para além de outras diferenças. O Parallels apenas permite instalar numa.

Divirtam-se, e bom trabalho!

(Este bocado de prosa foi escrito no mais absoluto desrespeito pelo Acordo Ortográfico. Assim se escreve em bom Português.)


******

This guest post was submitted by Tiago Neto, a German/English to European Portuguese translator and veterinarian, who has written what I believe to be the first article comparing the efficiency of different virtual machines for Windows-based translation environments running on Apple Macintosh hardware. Thank you, Tiago!