An exploration of language technologies, translation education, practice and politics, ethical market strategies, workflow optimization, resource reviews, controversies, coffee and other topics of possible interest to the language services community and those who associate with it. Service hours: Thursdays, GMT 09:00 to 13:00.
The last of the planned office hour discussions for the self-guided online course "memoQuickies Resource Camp" will be held on November 27th at 17:00 CET (8:00 PST). For those not already registered for the Zoom chats, the link to do that is HERE. If you have done so previously, the access URL is the same. The chat is open to everyone, regardless of whether they are registered or not in the online course.
We'll begin with open Q&A time on any of the course sections in the term base unit or any other presentations of the subject matter by me on this blog or my YouTube channel. Afterward, I will show my general method for editing or updating term base content via Microsoft Excel exports brought in to the memoQ working grid, which facilitates certain kinds of changes or actions involving regular expression use. This goes beyond the possibilities of the integrated memoQ term base editor, which will also be presented briefly.
A recording of the talk will be made available later in the course structure.
In December the course will move on to discussions of QA profiles and other aspects of quality assurance in memoQ, with some live discussion possibilities to be announced. All course material will remain online for access until the end of March. More information on the memoQuickies Resource Camp can be found HERE.
A few years ago while on "holiday", I returned from dinner to find that my laptop had bluescreened. Panic time! It was Saturday night, and I still had quite a lot of text to translate and deliver on Monday morning. And up on the highest mountain in Portugal, I wasn't sure where I could find a replacement to finish the project, which was, at least, not utterly lost, because I had put it on a memoQ Cloud server for testing. The next day I got lucky: about 50 km away there was a Worten, where I picked up a gamer laptop with lots of RAM and an SSD. Well, not so lucky, as it was a Hewlett Packard Omen, with a fan prone to failure, but that's another story....
This new laptop was my first encounter with Windows 10. I had heard that this operating system offered improved speech recognition capabilities, and since I prefer to dictate my translations and downloading the 3 GB installation file for Dragon NaturallySpeaking (DNS) from my server at the office was going to take forever, I thought I would give Windows 10 speech recognition a try. I hadn't installed my CAT tool of choice yet, so I fired up Microsoft Word and began dictating. "Not bad," I thought. Then I tried it in my translation environment, and the results were a complete disaster. So I put that mess out of my mind.
Since then there have been some notable advances in speech-to-text capabilities on a number of platforms. But the best solution for my languages (German and English) with DNS became increasingly cranky thanks to neglect of the product by Nuance. Every week I read new reports of trouble with DNS in a variety of environments in which it used to perform very well. Apple's iOS 13 was a great leap forward of sorts for speech recognition and voice-controlled editing, but the new features are only available in English, and having Voice Control activated totally screws up my otherwise rather good dictation in German and Portuguese (or any other language). And don't get me started on the crappy vocabulary addition feature, which uses text entry alone with no link to actual pronunciation. Good luck with that garbage. It's not a bad solution in Hey memoQ with the additional command features added, but iOS dictation is not completely up to reasonable professional standards yet.
I probably would have given no further thought to Windows 10's speech-to-text features if it weren't for Anthony Rudd. We've corresponded a bit since I bought his excellent book on regular expressions for translators (and there's another practical guide for us coming soon from him!), and in a recent discussion he alluded to the use of Unicode with regex as a simple way of dealing with some things another colleague was struggling with. I was intrigued by this, and so for about half a day, I ran down a rabbit hole, testing Unicode subscripts and superscripts for a variety of purposes like fixing bad OCR of footnote markers and empirical formulae, autocorrecting common expressions for subscripted variables and chemical terms, including subscripts and superscripts in term bases and much more. Fascinating and useful stuff on the whole, even if some fonts don't support it well.
And of course I looked at using these special Unicode characters in speech-to-text applications. DNS had some funky quirks (not allowing numbers in the "spoken" version of terms, for example), but it worked rather well, so I can now say "calcium nitrate formula" and get Ca(NO₃)₂ without much ado. And for some reason it occurred to me to give Windows 10 speech recognition a try, just because I was curious whether vocabulary could in fact be trained. Indeed it can, and that feature is better than iOS 13 or DNS by far.
But first I had to remember how to activate speech recognition for Windows on my laptop again. When in doubt, type what you're looking for in the search box....
Notice I've pinned Windows Speech Recognition to my taskbar on the right, which is good for quick tasks.
Gesucht, gefunden. Unlike other speech recognition solutions, the one in Windows 10 works only for the language set for the operating system. And options there are limited to English (United States, United Kingdom, Canada, India, and Australia), French, German, Japanese, Mandarin (Chinese Simplified and Chinese Traditional) and Spanish.
I put on my trusty Plantronics earset (the best microphone I've used for dictation tasks or audio in my occasional webinars in the past year) and began to dictate, first in Microsoft Word, which had shown acceptable results in my tests long ago. I found that adding vocabulary in the Speech Dictionary (accessed via the context menu in the dictation control element shown as a graphic at the top of this post) was dead simple.
The option to record pronunciation enabled me to record non-English names and words in several languages. And sure enough, the Unicode subscripts and superscripts worked, so I can now say CO₂ (I just dictated that) to my heart's content.
I was expecting a mess when I tried to use Windows 10 speech-to-text in a CAT tool, but it was not to be. It was brilliant, actually. I tried it in my copy of SDL Trados Studio, and with the scratchpad disabled so I could dictate directly into the target it worked well. No voice-controlled editing like I'm used to with DNS in memoQ, but that DNS feature does not work in SDL Trados Studio anyway, so this is no worse. But with the scratchpad box enabled (see the screenshot below), I could use voice commands to select and correct text or perform other operations. Brilliant!
After clicking or speaking "Insert", the text will be written to the target field with the proper formatting
So users of SDL Trados Studio who translate to a target language supported by Windows 10 speech recognition are probably better off not giving their money to Nuance, which I'm told can't even be bothered to make a 64-bit version of DNS now (which probably accounts for a lot of the trouble people have with that program.
I tested Wordfast Pro 5, which seems to confuse the speech recognition tool horribly, with source text displayed in the floating bar for some odd reason. But my earlier tests of Wordfast with DNS were equally unhappy, so somehow I'm not surprised. And I didn't test the Memsource desktop editor, which took the price a few years ago for the worst-ever DNS dictation results with a CAT tool. I'll leave that to someone with a much wider masochistic streak.
But what about memoQ, my personal environment of choice for most translation work? Equally brilliant, works just the same as SDL Trados Studio. No voice control for editing without the dictation scratchpad enabled (there, DNS has an advantage in memoQ), but with the scratchpad you can use the voice commands to edit before inserting in the target text field.
Wanna see this in action? Have a look at this short demo video:
I hope that the future will bring us more language support for Windows 10 dictation (Portuguese, Russian and Arabic, please!) and that other providers (like Google, if you're listening, and Apple, which never listens to anyone anymore except to spy on them with Siri) will expand the speech-to-text features offered, particularly to include sound-linked vocabulary training and better adaptation to individual users' speech. Five years ago when I began to investigate alternatives for non-DNS languages, I expected we would have more by now, and we do, but professional needs require all providers to raise their game.
Addendum: Someone asked me if Windows Speech Recognition is a cloud resource or a locally installed one which will work without an Internet connection. It's definitely the latter. So if you have lousy bandwidth or find yourself disconnected from the Internet, you can still use speech-to-text features.
And more: I use a lot of spoken commands for keyboard shortcuts when I work, so I did a little research and testing. It seems that Windows 10 speech recognition gives full access to an application's keyboard shortcuts via voice. So in memoQ, for example, I can dictate the insertion of tags, items from the Translation Results pane and a lot more. Watch out, Nuance. Windows 10 is going to kick your Dragon's scaly butt!
In my last post on the release of memoQ 8.7 with its new, integrated speech recognition feature I included a link to a long, boring video record of my first tests of the speech recognition facility, most of which consisted of testing various spoken iOS commands to generate text symbols, change capitalization, etc. I tested some of the integrated commands that are specific to memoQ, but not in an organized way really.
In a new testing video, I attempt to show all the memoQ-specific spoken command types and how the commands are affected by the environment (in this case I mean whether the cursor is on the target text side or the source text side or in some other place in the concordance, for example).
Most of the spoken commands work rather well, except for insertion from the concordance, which I could not get to work at all. When the cursor is in a source text cell, commands have to be given in the source text language currently, which is sure to prove interesting for people who don't speak their source language with a clean accent. Right now it's even more interesting, because English is the only language with a ready-made command list; other languages have to "roll their own" for now, which is a bit of a trial-and-error thing. I don't even want to think how this is going to work if the source language isn't supported at all; I think some thought had to be given to how to use commands with source text. I assume if it's copied to the target side it will be difficult to select unless, with butchered pronunciation, the text also happens to make sense in the target language.
It's best to watch this video on YouTube (start it, then click "YouTube" at the bottom of the running video). There you'll find a time code index in the description (after you click SEE MORE) which will enable you to navigate to specific commands or other things shown in the test video.
My ongoing work with Hey memoQ make it clear that what I call "mixed mode" (dictation with concurrent use of the keyboard) is the best and (actually) necessary way to use this feature. The style for successful dictation is also quite different than the style I need to use with Dragon NaturallySpeaking for best results. I have to discipline myself to speak more in short phrases, less in longer ones, much less in long sentences, which may cause some text to be dropped.
There is also an issue with Translation Results insertions and the lack of spaces before them; the command to insert a space ("spacebar" in English) is dodgy, so I usually have to speak it twice and end up with a superfluous space. The video shows my workaround for this in one part: I speak a filler word (in one case I tried "dummy" which was rendered as "dumb he") and then select it later and insert an entry from the Translation Results pane over the selected text. This is in fact how we can deal with specialist terminology not recognized by the current speech dictionary until it becomes possible to train new words some day.
The sound in the video (spoken commands) is also of variable quality; with some commands I had to turn my head toward the iPhone on its little tripod next to my laptop, which caused the pickup of that speech to be bad on the built-in microphone on the laptop's screen. So this isn't a Hollywood-class recording; it's simply a slightly edited record of some of my tests to give other memoQ users some idea of what they can expect from the feature right now.
Those who will be dictating in supported languages other than English need some patience right now. It's not always easy coming up with commands that will be recognized easily but which are unlikely to occur as words to be transcribed in typical dictation work. During the beta test of Hey memoQ I used some bizarre and unusual German words which just happened to be recognized. I'm developing a set of more normal-sounding commands right now, but it's a work in progress.
The difficulties I am encountering making up new command phrases (or changing the English ones in some cases) simply reinforce my belief that these command lists should be made into portable light resources as soon as possible.
I am organizing summary tables of the memoQ-specific commands and useful iOS commands for symbols, capitals, spacing, etc. comparing their performance in other iOS apps with what we see right now in Hey memoQ.
Update: the summary file for English is available here. I will post links here for any other languages I can prepare later.
The topic of accessing and editing translatable text in tags comes up from time to time. I thought I had published instructions on this topic some time ago, but when a tech-savvy colleague who always does a proper search before asking questions couldn't find it, nor could yours truly, I concluded that it was time for another tutorial video. So here it is:
The video post on YouTube includes a hot-linked table of contents that will enable you to jump to key parts of the tutorial. This is a very simple function to implement with "markers" in Camtasia, and I recommend that those who make tutorials of any significant length or who post recorded webinars consider implementing such tables of contents to facilitate finding particular parts of interest without endless hit-and-miss searching in a long video.
There's been a bit of a buzz lately in professional language service circles regarding a recent ruling by California's Supreme Court, which establishes a new, simplified standard to determine independent contractor status. For example, the corporate interest blog Slator reported on the panic among large language brokerage firms sometimes known for predatory and abusive practices with the companies and individuals whom they contact to provide services, while independent interpreter Tony Rosadooffered an interesting perspective on how the ruling can have a positive impact on independent professionals in his field and, I dare say, independent translators as well.
The Court's decision established the "ABC" criteria as the new standard for distinguishing employees from independent contractors:
A) The individual must be free from the control and direction by the hiring entity with regard to the performance of the work, under the terms of the contract for the performance of this work and in fact.
B) The individual must perform work which is outside the usual course of the hiring entity’s business.
C) The individual is customarily engaged in an independently established trade, occupation, or business of the same nature as the work performed for the hiring entity.
Failure to meet all three criteria will lead to a finding that the individual is an employee and is therefore not an independent contractor.
Now you might say – correctly – that California is a long way from New York, London, Paris, Berlin and Rome, what has that got to do contractor conditions there? A lot.
Over the past three and a half decades, since the rape and pillage of unions and workers rights in general began by the rapacious disciples of Ronald Reagan, a system of work and service practices has evolved which I think could fairly be called social strip-mining. (After I wrote this term, I wondered if others have used it as well, and found, unsurprisingly, that this obvious analogy has occurred to quite a number of people.) Very few of the present practices by large corporate providers of language services (interpreting, translation, writing and editing) are in fact sustainable.
Like strip mining companies tearing down a mountain and utterly destroying its ecology of flora and fauna, polluting waters underground and on the surface nearby as well as other ecosystems, companies like Lionbridge, thebigword, TransPerfect, RWS/Moravia and others and their downstream companies in the service sewer put the squeeze on individuals at the end of their corporate digestive system to extract maximum resources for minimum benefit in a manner which often cannot sustain the living of the writers, editors, translator, interpreters and other service workers at the butt end of things. These individuals may cling for some time to their desperate situation as an alternative to prostitution or delivering pizzas, but there is very little incentive and few resources provided for them to develop as professionals and acquire greater skills to deliver greater value with time.
Some, like abused children, learn the lesson of what the abusers can get away with and continue the cycle on small and large scales, outsourcing or even founding new companies with similar practices.
In truth, nobody is well-served by these practices on the end customer side – the individuals, companies and government bodies who contract with the intermediaries for services provided by individual interpreters, translators, editors, etc. – nor on the end-provider side – those very interpreters, translators, editors, etc. And in the middle? "Growth" seems to derive largely from acquisition and from refinement of their marketing deceptions (many in the bulk market bog of language services have SEO-tuned web pages designed to capture searches for independent individual service providers), not so much from actual organic growth of internal service and quality offered. Price dumping practices are also common; many small service companies are unable to compete with "loss leader" prices to end customers which are lower than those they pay to real independent service providers of good professional business standing.
None of these abusive practices are new; they have been known in many forms throughout the modern history of labor starting more or less in the early 19th century. These practices ultimately led to the rise of unions and bodies of protective legislation in the past, so it is not surprising that some have called for "unionization" of international service providers. ("Workers of the world, unite!", anyone?) But these well-meaning calls for unions of interpreters and translators are not really practical in most situations. So many say there is nothing to be done.
Wrong. The California Supreme Court decision points the way toward ending the worst of corporate abuses of individuals providing service by creating a situation in which the true costs of these services are emphasized in the relationship with the service provider. There is nothing standing in the way of companies like Lionbridge or much smaller companies from increasing salaried staff to write, translate, interpret, etc. under local statutory conditions for ordinary employment. To the extent they find this impractical, these companies can contract with other companies or with individuals who meet the ABC criteria.
Such a requirement would also give a fair break to those companies who do in fact invest in the socioeconomic maintenance and professional development of their employees providing service to end customers. Under current practices, these good companies are unfairly disadvantaged by laws and regulatory practices which now permit these service strip-miners to operate as they do. Local and national governments would also benefit from and increase of benefit payments from registered employees or from taxes assessed on work transactions which fail the ABC test.
In the environment of expanding globalized trade and sophisticated corporate shell games to avoid tax liabilities, enforcement of necessary and proper good social practice at the "source" – the intermediate provider level or at the end customer level where there is a direct relationship between a company or government body with a presumed independent service provider – is perhaps the most practical way to accomplish some of the reforms needed on a global scale.
Lately I've been doing a bit of custom filter development for some translation agency clients. Most of it has been relatively simple stuff, like chaining an HTML filter after an Excel filter to protect HTML tags around the text in the Excel cells, but some of it is more involved; in a few cases, three levels of filters had to be combined using memoQ's cascading filter feature.
And sometimes things go too far....
A client had quite a number of JSON files, which were the basis for some online programming tutorials. There was quite a lot of non-translatable content that made it past memoQ's default JSON filter, much of which - if modified in any way - would mess up the functionality of the translated content and require a lot of troublesome post-editing and correction. In the example above, Seconds in a day: is clearly translatable text, but the special rules used with the Regex Tagger turned that text (and others) into protected tags. And unfortunately the rules could not be edited efficiently to avoid this without leaving a lot of untranslatable content unprotected and driving up the cost (due to increased word count) for the client.
In situations like this, there is only one proper thing to do in memoQ: edit the tags!
There are two ways to do this:
use the inline tag editing features of memoQ or
edit the tag on the target side of a memoQ RTF bilingual review file.
The second approach can be carried out by someone (like the client) in any reasonable text editor; tags in an RTF bilingual are represented as red text:
If, however, you go the RTF bilingual route, it's important to specify that the full text of the tags is to be exported, or all you'll get are numbers in brackets as placeholders:
Editing tags in the memoQ working environment is also straightforward:
On the Edit ribbon, select Tag Commands and chose the option Edit Inline Tag.
When you change the tag content as required, remember to click the Save button in the editing dialog each time, or your changes will be lost.
These methods can be applied to cases such as HTML or XML attribute text which needs to be translated but which instead has been embedded in a tag due to an incorrectly configured filter. I've seen that rather often unfortunately.
The effort involved here is greater than the typical word- or character-based compensation schemes can justly compensate and should be charged at a decent hourly rate or be included in project management fees.
A lot of translators are rather "tag-phobic", but the reality of translation today is that tags are an essential part of the translatable content, serving to format translatable content in some cases and containing (unfortunately) embedded text which needs to be translated in other (fortunately less common) cases. Correct handling of tags by translation service providers delivers considerable value to end clients by enabling translations to be produced directly in the file formats needed, saving a great deal of time and money for the client in many cases.
One reasonable objection that many translators have is that the flawed compensation models typically used in the bulk market bog do not fairly include the extra effort of working with tags. In simple cases where the tags are simply part of the format (or are residual garbage from a poorly prepared OCR file, for example), a fair way of dealing with this is to count the tags as words or as an average character equivalent. This is what I usually do, but in the case of tags which need editing, this is not enough, and an hourly charge would apply.
In the filter development project for the JSON files received by my agency client, the text used was initially analyzed at
14,985 words; 111,085 characters; 65 tags
and after proper tagging of the coded content to be protected it was
8766 words; 46,949 characters; 2718 tags.
The reduction in text count more than covered the cost of the few hours needed to produce the cascading filter needed for this client's case and largely ensured that the translator could not alter text which would impair the function of the product.
Twenty Years behind QA for The Chicago Manual of Style
Carol Saller is best known in her role as an editor at the University of Chicago Press, where she was head copyeditor for the 16th edition of The Chicago Manual of Style, and as the author of her often hilarious and enormously helpful book, The Subversive Copyeditor. The book is an eminently cogent response to the thousands of questions that Ms. Saller reads each year from writers and copyeditors in her role as Editor for the Q&A page of The Chicago Manual of Style Online. Many of these writers and editors have reached a stand-off with each other over prickly and sometimes humorous questions of grammar and style. To wit, “My author wants his preface to come at the end of the book. This just seems ridiculous to me. I mean, it’s not a post-face.”
Carol Saller surprises a lot of hardline editors by stressing flexibility when it comes to supposedly hard and fast “rules.” The focus, she seems to feel, should be on clarity for the reader and on a good and useful working relationship between writers and editors… As well as translators and editors!
September 9, 2017 at 4 PM UTC –This webinar will be held in English
Colleague Simon Berrill, who works from Catalan, Spanish and French into English, writes a rather thoughtful blog which I have enjoyed very much in recent months. His latest post, "Better Together", describes a rather interesting mutual review arrangement, a linguists' triangle, which I thing could be an interesting and beneficial thing for translators at any stage of their careers.
The particular arrangement between Simon and his revision partners, Tim and Victoria, is striking for me because of its flexibility and the fact that it is not linked to specific work assignments. My first thought while reading his post was something like "Gee, I could really benefit from something like this!", and my mind began to drift to all of the interesting things that could be learned in a swap like this with the right people.
No more spoilers... go read the post yourselves and comment there. Thank you, Simon, for another thought-provoking contribution.
Although tracked changes have been part of memoQ since the distant days of memoQ 5.0, many users are still confused about how to use these features and how to navigate marked changes in a translation.
The confusion starts with the menu for activating the tracked changes, which in recent versions of memoQ is found on the Review ribbon. What most people do not realize is that the first two options - Against Last Received Version and Against Last Delivered Version - are not relevant to the usual workflows of an individual translator working in a local project created on his or her computer. Often I have caught myself selecting the option Against Last Delivered Version for the tracked changes to show, because I want to compare against the last version I delivered to my client by exporting and e-mailing the document, because I forget that this refers to the actual Deliver function in a server project.
If I am working locally in my own projects, the only track changes option that is relevant is Custom, with which I can show comparisons to specific minor versions:
In the present example, I've selected a comparison with a "snapshot" I made before an editing session. A snapshot creates a record of the status of a translation at a given time and makes rollbacks possible. Use the submenu of the Versions icon on the Documents ribbon to make a snapshot of your translation:
Once the tracking of changes for the current translation compared to a previous minor version has been activated, the relevant changes will be marked in red in the translation grid. If changes have been made to the source text (correcting OCR errors, for example, by editing the source text with F2), these will be shown as well.
Changes can be rejected by choosing Revert To Earlier Version on the Review ribbon, in the context menu (right-click) or with the corresponding keyboard shortcut. Or a version of a target text not shown in the markup can be recalled with the Row History and restored by copying it from the dialog (Ctrl+C) and pasting in the target cell and editing out extraneous information.
But how can one navigate many tracked changes in a larger document? Many users think this is not possible, though in fact it's rather simple with memoQ's filtering features.
Clicking the filter icon above the target text column opens a dialog to specify filter criteria for the working view. On the Status tab under Other properties... the option Change tracked can be selected to show only those segments with tracked changes.
Alternatively, the Go to next segment settings (Shift+Ctrl+G) can be configured in the same way with Change tracked on the Status tab, so choosing Go to next (Ctrl+G) or confirming a segment (if the option Automatically jump after confirmation is selected in the Go to next segment settings dialog) will take you to the next segment with tracked changes.
Last February I described my initial work with translation tools as environments for authoring and editing documents in a single language. Some people have been doing this quietly for a while; occasionally I would hear puzzled comments from a trainer who had held a class on SDL Trados Studio, OmegaT or memoQ which had been attended by a technical writer or someone with other professional writing interests not related to translation. But to my knowledge there has been no systematic approach to this.
Some weeks later I began to discuss and present some new possibilities for speech recognition in 38 languages which go well beyond the limitations of Dragon NaturallySpeaking for automated speech transcription in the eight languages for which it is available. These possibilities include a number of mobile solutions which are quickly gaining traction among translators and other professional writers.
On Tuesday, June 2nd (two days from now), I will be presenting a one-hour introduction to "MemoQ for Single-language Authoring and Editing" in the eCPD Webinar series.The registration page is here.
This presentation will be an update of the talk I gave earlier this year which discussed CAT tools in general as authoring and editing tools. Although any tool works in principle (and even a user of SDL Trados Studio, for example, can probably draw enough ideas from the upcoming eCPD talk to make good use of the approach), memoQ has some particular advantages, not the least due to its corpus-handling features in LiveDocs and its superior predictive typing facilities, including "Muses" (which are like SDL's AutoSuggest with more flexibility and without the onerously high data quantity requirements).
The presentation will include an overview of some of the latest advances in speech recognition in 38 languages for ergonomically superior writing by automated transcription as well as discussions of version management and dictation workflows which can be applied for greater ease in editing monolingual documents or even translations, including post-editing of machine pseudo-translation (PEMpT by the "Hardisty Method"). I've been fairly quiet on this blog in recent months due to conference organization and travels and the considerable time put in to researching improved work ergonomics for translation, writing and editing processes. (In fact I didn't even find time to blog the memoQ Day on April 22nd in Lisbon yet!) Elements of all these efforts, which have sparked no little interest at recent conferences and workshops I have presented at in Europe, will be part of Tuesday's talk, which will include Q&A afterward to explore the interests of those participating.
So if you are a translator involved in a lot of revision or editing work (bilingual or monolingual, a technical writer or other professional writing in a single language for publication, someone working on a thesis or authoring for other purposes, the eCPD presentation may help you to do this with better organized resources and greater efficiency. As one friend of mine who wrote a thesis just before I developed this approach put it, with this she would at least have been able to keep track of the feedback on her work from its five or so reviewers without going completely nuts.
Last night from 6:00 to 9:30 I enjoyed a "memoQ&A Evening" at the Porto Bagel Café as a reward for surviving the long bus ride to Porto/Gaia from Évora to attend the JABA Partner Summit. About 25 local colleagues attended to hear my not-as-short-as-promised presentation and discuss approaches to memoQ and other translation technologies as our working tools. The evening was part of the Translators in Residence initiative and a good start to my second visit to the area after my whirlwind tour last month to investigate venues for teaching events. Many thanks to the sponsors. the International Association of Professional Translators and Interpreters and Chip7 of Évora for providing the funding and tools (an excellent LCD projector - thank you, Carlos!) to do this.
I very much appreciate IAPTI's commitment to the professional education and continuing development of my good colleagues in Portugal, particularly in difficult economic times when many findit difficult to attend translators' events in faraway places. The evening was free for all attendees, who only had to pay for whatever they drank (great coffee - I had my usual galão) and ate (the best bagels in Portugal!).
After an initial hour of snacks, coffee and chat, the evening began with a discussion of the game-changing implications of speech recognition technologies for our working lives. Not only is it now possible for colleagues to use high-quality speech recognition on desktop computer and laptops in languages such as Hungarian and Portuguese, which are not currently supported by Dragon NaturallySpeaking (using, for example, the integrated recognition tools in the Mac Yosemite OS, as demonstrated with SDL Trados Studio and memoQ in Lisbon the day that SDL conquered Portuguese translation), smartphones are part of the game now too. Since picking up an older iPhone model (4S) for a few hundred euros about a month ago, I have had excellent results testing it with English, German, Russian and Portuguese and e-mailing texts to myself with just a few taps on the phone's screen. Once transferred as an e-mail, the text is then aligned in a CAT tool such as memoQ and subjected to tagging, QA and other procedures of the usual virtual translation working environments.
The use of memoQ and other CAT tools for single-language original authoring and text revision was also discussed. This flexible workflow extends the relevance of translation environment tools well beyond the usual limits within which translators and translation companies live and operate and offers interesting prospects for collaboration and re-use of creative resources. This topic willalso be covered next week in a lecture and workshop at Universidade de Évora and in an eCPD webinar on June 2, 2015.
Interoperability is another important topic for translators; I discussed different ways in which I use SDL Trados Studio and other tools to prepare projects to work in memoQ and vice versa as well as mz highly profitable use of SDL Multiterm to enhance customer loyalty and my professional image with this terminology management ssystem's excellent output management features.
Other tips and tricks in the memoQ&A included the untapped potential of LiveDocs, tracked changes and row histories in memoQ, dealing with embedded objects, graphics and transcription, PDF 3-ways and new tricks for nasty and/or illegible image PDFs, versioning and a concept for transforming translation memory concordancing into something much, much more useful and less prone to errors in editing and translation.
Copies of the slides from the evening's presentation are available here. It is, however, merely a palimpsest of the evening.
Many thanks also to colleague and translation tools teacher Felix do Carmo for kindly chauffeuring me around town and for the interesting tour of the training and production facilities at his company, TIPS.
I am often asked about the monolingual editing workflows I have used for some 15 years now to improve texts which were written originally in English, not created by translation from another language. And I have discussed various corpus linguistics approaches, such as to learn the language of a new specialty or the NIFTY method often presented by colleague Juliette Scott.
However, on a recent blitz tour of northern Portugal to test the fuel performance of the diesel wheels which may take me to the BP15 and memoQfest conferences in Zagreb and Budapest respectively later this year, I stopped off in Vila Real to meet a couple of veterinarians, one of whom is also a translator. During a lunch chat with typically excellent Portuguese cuisine, the subject of corpus research as an aid for authoring a review paper came up. I began to explain my (not so unusual) methods of editing and existing document when I was asked how the tools of translation technology might be applied to authoring original content.
The other translator at the table said, "It's a shame that I cannot use my translation memories to look things up while I write", and I replied that of course he could do this, for example with the memoQ TM Search Tool or similar solutions from other providers. And then he said, "And what about my term bases and LiveDocs corpora?", and I said I would sleep on it and get back to him. In the days that followed, other friends (coincidentally also veterinarians) asked my advice about editing the English of the Ph.D. theses and other works they will author in English as non-native speakers of that language. One of them noted that it would be "nice" if she could refer to corrections made by various persons and compare them more easily. I said I would sleep on that one too.
A few days after that the pain in my hands and feet from repetitive strain injuries and arthritis was unbearable, aggravated by a rope burn accident while stopping an attack on sheep by my over-eager hunting dog and by driving over 1000 km in a day. I doubled down on the pain meds, made a big jug of toxically potent sangria and otherwise ensured that I was comfortably numb and could enjoy a night of solid sleep.
It was not meant to be. Two hours later I woke up, stone sober, with a song in my head and the solution to the problem of my Portuguese friends writing in English and Tiago wanting to author his work in memoQ for the convenience of using its filters to review content. Since then the concept has continued to evolve and improve as others suggest ways of accommodating their writing or language learning needs.
After about a week of testing I scheduled one of my "huddle" presentation classes, an intimate TeamViewer training session to discuss the approach and elicit new ideas for adapting it better to the needs of monolingual authors. The recording of that session is available for download by clicking on the image of the title slide at the top of this post. (The free TeamViewer software is needed to watch the TVS file downloaded; double-click it, and the 67-minute lecture and Q&A will play.)
I'm currently building Moodle courses which provide more details and templates for this approach to authoring and editing, and it will be incorporated in parts of the many talks and workshops planned this year.
I am aware that SDL killed their authoring product, the Author Assistant, and that Acrolinx offers interesting tools in this area, as do others. But I'm usually hesitant to recommend commercial tools in an academic environment, because their often rapid pace of development (such as we see with memoQ) can play serious havoc with teaching plans and threaten the stability of an instructional program, which is usually best focused on concepts and not on fast-changing details. So I actually started out my work and testing of this idea using the Open Source tool OmegaT, the features of which are more limited but also more stable in most cases than the commercial solutions from SDL, Kilgray and others. But as I worked, I noticed that my greater familiarity with memoQ's features made it an advantageous platform for developing an approach, which in principle works with almost every translation environment tool.
Part of my motivation in creating this presentation was to encourage improvements in the transcription features available in some translation environments. But the more I work with this idea, the more possibilities I see for extending the reach of translation technology into source text authoring and making all the resources needed for help available in better ways. I hope that you may see some possibilities for your own work or learning needs and can contribute these to the discussion.
It's been nearly three weeks since I returned from my first meeting of the Mediterranean Editors and Translators association, METM14. In that time I've wondered how to share all that I brought back from the gathering, and I'm afraid I still don't know how to put it all into words. I've become increasingly reluctant to participate in translation conferences in the past few years, took a break of about a year and a half from them, because too many had become venues for pushing a corporatist agenda which I feel has little to do professional language service of real social value. The unexpectedly excellent IAPTI conference in Athens last September marked my return, and what pleased me most there was the clear focus of the event program on the professional practice of freelance translators and interpreters. No sales pitches, no Linguistic Sausage Producers explaining their multiphase chop-and-grind workflows to redefine quality as a complex function of engineering incompetence, empty promises, pseudoscientific LQA scores and extrusion rate. MET's outstanding 10th annual meeting in San Lorenzo de El Escorial near Madrid was an exquisite dessert after the feast in Athens: once again I had the great pleasure to meet a number of professional peers with experience and competence well beyond my own, some who have been mentors of mine for years with their contributions for practical corpus linguistics in translation and other topics dear to me.
And once again I had the delight of a program that was, for me, truly something completely different. In all the professional conferences I have attended in the past 15 years, this was the first one where a substantial number of attendees were serious, professional editors. Many translators take on "editing jobs" for better or worse, but at METM14 I was surrounded by people who pursue this activity with a professional seriousness and rigor which was quite frankly new to me. And their perspectives on some matters which are routine in my own work seemed more than a little weird at first.
I was definitely out of my comfort zone on the pre-conference workshop day as I sat in an excellent session by Mary Ellen Kerans on corpus-guided decision-making. A paper that she co-authored years ago with two other colleagues gave me my first exposure to the effective use of corpora for my translation work, but I had always applied these techniques from my own rather settled perspective as a translator. On that day, I saw how editors use the same techniques for very different purposes, and after about a half hour of considerable confusion, I enjoyed surprising new insights in how I might improve my own work by considering these other perspectives.
The editors' perspectives continued to alternately confuse and inspire me for the next two days. I learn something at most conferences, but usually what I walk away with are ideas that are not too far from my usual professional comfort zone. Here I was challenged in new and different ways, and I really loved that. I had been aware of MET for a number of years because of a few colleagues in Stridonium who were members, and I've looked at the conference program off and on for about five years and was always impressed by their focus on peer-to-peer teaching, but what I found was really well beyond my good expectations.
I attended the event with another colleague from Portugal who is relatively new to translation; she was a little nervous about her first professional conference, and although I expected she would gain some useful insights, I did not really know what would await a new professional at METM14. Any concerns were quickly dispelled; I was extremely pleased to see how many new professionals were welcomed and encouraged to participate by so many with more experience than I am likely to gain still in what remains of my professional life.
My friend was thoroughly inspired by the people she met and the presentations and workshops she attended, and on the long drive home after the last day she put together puzzle pieces from a number of talks and hit me with new ideas for teaching translation support technology to new users that still have my head spinning and will be the foundation for my next book, which I hope to release by early next year. This was just one of many occasions where I have found that newcomers to a profession can contribute some of the most important insights for improvement.
Defenders of the Portuguese language at METM14
Next year's annual meeting (METM15) will be held at the end of October in Coimbra, Portugal at the university there. If you are getting tired of the same old topics presented by the same old suspects and programs clearly driven by agendas at odds with the ethics and interests of freelancers and staff professionals who put quality first, then you may want to join me next October in Portugal for another healthy serving of professional dessert.
Emma Goldsmith has blogged a good overview of the sessions she attended at METM14, which can give you a feel for some of what you may have missed. The conference program offers more, less personal information. But don't rely on the impressions of others; come next year to a great event in a great country and then tell others yourself what this unusual mix of extreme professional competence has to offer.
I couldn't make it to memoQfest this year - the first one in Budapest that I have missed since the event began in 2009. But the first family visit since that same year took priority, so my exposure to the upcoming memoQ 2014 version was strictly second hand until today.
I wasn't too happy with thing I heard on the Yahoogroups user list. In fact, when I read one message describing how the new transcription feature for bitmap graphics in some files required the Product Manager version, I was quite annoyed. The reality - a whole month before the official release - is very good for both freelancers and corporate outsourcers, and I think by the time this version makes its official debut in June there will be many good reasons to smile. I'm frankly amazed at how much Kilgray seems to be getting its act together and balancing the needs of users at all levels.
This afternoon I downloaded the first test release (alpha??) of memoQ 2014, installed it and began to take a cautious tour. My first impression was that it looked the same. And then, bit by bit, subtle and excellent small differences began to emerge. I looked for and found major new features I had heard about and discovered many interesting things not mentioned along the way.
The grammar checking feature seems to be implemented in a sensible way, though it actually doesn't work at all right now for me. But I can see where it's headed, and it is going in a good direction.
I had a quick look at the new plug-ins, particularly TaaS, and made notes about testing the potential for teamwork. What I have seen of TaaS for its much-advertised terminology extraction is a huge disappointment, and those who have followed my comments on Twitter will know I have nothing good to say about this EU boondoggle, but I see potential for other possibilities that nobody has really talked about, and if my instinct is right, this could be really useful. But I will need to invest a lot of testing time for the approach I have in mind.
The Project home view has gotten even more impossibly cluttered with the addition of "People", a rather sensible reworking of role assignments that even in the Translator Pro version clearly acknowledges that most freelance translators are not, in fact, 'islands' in their work.
This will surely make the small screen (netbook) usage problems worse if Kilgray does not redesign the view a bit, but in every other respect I see this as a significant improvement of project workflow, emphasizing the relationships between project participants in a better way.
One little bit that I stumbled across was the new way of handling the export of unfinished translations. This is a nice way of recognizing the frequent pressure in some projects to export incomplete stages of work.
I have had ways of dealing with this need for years in memoQ, but this new approach will make things simpler and obvious for all users.
There is a nice little feature for tracking time too:
This will facilitate record keeping for some jobs involving time charges.
The feature I have looked at in some depth so far, which makes me very happy, is Kilgray's very sophisticated handling of embedded objects and graphics, which sets new standards in many ways. I think there is still a key feature missing to make it the equal of OmegaT for handling charts with data stored as XML in the MS Office file (though I have not had time to check this yet), but what I have seen so far goes way beyond similar features I have seen in STAR Transit and Déjà Vu X2.
Embedded objects and images are imported as separate files from within the media and embeddings folders of the Microsoft Office file. I see a few potential problems with the current way of displaying a file and its objects and media. I've had projects with multiple files having embedded Excel spreadsheets, PowerPoint slides and other objects as well as any number of pictures needing to be localized. One recent project had 59 spreadsheets embedded in a DOCX file. Without an accordion or tree structure to collapse the subordinate structure view and show the embedded content again, the overview will be lost quickly. But this is a very good start. Note how the main file includes a count of the segments in the subordinate objects and graphics. (And take note of the new progress bar with different colors for different process stages like translation and proofreading.)
Bitmap texts can be recorded with a new transcription feature, which is also compatible with voice recognition. I dictated my German source texts with Dragon Naturally Speaking set to German, then switched to English for the translation. And of course the bitmap transcriptions are included in the word counts of the Statistics functions and the translations are written to the translation memory. I believe this is utterly unique in translation environment tools. Fluency has a transcription module too, of course, but its purpose and application are very different.
The exported translations with translated objects will look like they are not done at present, because the difficult refresh problem has not been solved by Kilgray. Each translated spreadsheet, slide, etc. will need to be opened in the document before the translation will become visible. This is much easier using the macro I published two years ago, and I am certain that by release time or soon thereafter Kilgray will find an elegant way of dealing with this difficulty. Atril handles the same problem by distributing macros as I recall.
In the recent Kilgray blog post on the six reasons to upgrade to memoQ 2014, the only overlap with the above points is the image localization. Peter Reynolds talks instead about other good stuff, such as the long-awaited project templates and Language Terminal. There are so many nice things ahead with this upgrade that we'll all just have to take it slowly, one bit at a time.
Of course the usual precautions for any new software version apply. The new version can be installed in parallel to your current version, and it can be tested while you continue to do the bulk of your work in the older, stable version. Typically it takes a few months for any new version to get the kinks out, but this allows plenty of time for planning the transition and preparing to take full advantage of the new features relevant to you. Migration is also not a trivial matter in many cases, but this time around there may be a little more help with that. More on that another time!
It's a sad fact in the professional work of translators that a lack of understanding on how to deal effectively with various PDF formats causes enormous loss of productivity and results which are not really fit for purpose. The aggressive insistence of many colleagues possessed of a dangerous Halbwissen on using half-baked methods and inappropriate tools contributes to the problem, but, bowing to the wisdom about arguing with fools, I now mostly sit back with a bemused and amused smile and watch the tribulations of those who believe in salvation by PDF import filters and cheap or free OCR. "TANSTAAFL" is a true as it ever was.
Just before the weekend I got an inquiry from an agency client I rather like. Nice people, good attitude, but struggling sometimes trying to find their way with technology despite some in-country "expert" training. This inquiry looked a bit like ripe fish at first glance. The smell got stronger after I was told that because the corporate end client had converted the PDF for their annual report and begun to edit the mess (and comment it heavily too) in the OCR file that this would be all there was to work with. It was a thoroughly appetizing sight when imported into a translation environment:
There are so many issues in that tossed salad of translation terror that I don't even know where to start describing them.
The screenshot above was in memoQ. How does it look in SDL Trados Studio? Often just as messy. In this case, this was the result in an older version of Studio:
SDL Trados Studio choked and refused to import the file!
I do have the latest version of SDL Trados Studio 2014, but unfortunately it's on a system that does not yet Microsoft Office, because I refuse to bow to Microsoft's insistence that I must buy a Portuguese version of that software. No MS Office, no file import in this case with SDL Trados Studio. memoQ fortunately has not needed MS Office to import its old file formats since the release of memoQ 6.0.
Ugly OCR trash like this file is all too common at this time of year, and as I am busy compiling the syllabus for the workshop I want to do on better living with well-used technology for legal and financial translators, I felt obliged to take this one on as a teaching example. It's actually not as bad as it looks. On the other hand, the best approach may not always be obvious, and the best solution for one document may not apply as well or at all to another.
My first approach was to use Dave Turner's CodeZapper macros. This isn't as straightforward as it used to be since I downgraded from Microsoft Office 2003 to later versions; for some reason the toolbar refuses to stay loaded between work sessions, and there's no way I can keep track of all the abbreviations for macros on it.
I can't deal with anything more complicated than clicking the "CZL" option for "Code Zapper lite", which did a rather decent job on the heavy mess above:
But all was not quite as well as it seemed:
Text in the header and footer remained trashed, and the heavy use of comments and tabbed lists meant that there were plenty of legitimate tags to deal with which were just too confusing with the DVX-like mess of memoQ's default import and display for an RTF file.
So I went for a kinder, gentler approach. I changed my import filter settings in memoQ:
There is actually seldom any good reason to import an RTF or DOC file into memoQ using the default filter settings. And marking those two little checkboxes at the bottom often accomplishes much of what CodeZapper does. Sometimes less. A bit more in this case.
The header and footer texts were absolutely clean. Don't let the extra tags in this sample fool you: overall, there were fewer than in the code-zapped file. Now there are still a number of issues to be seen in the screenshot above, including paragraph breaks in the middle of a sentence and awful manual hyphenation (many instances of that in the whole text) and joys like badly placed comments and links which mess up the text and prevent term identification by the software:
Source editing features of memoQ (F2) enable issues like the two above to be dealt with easily:
After a bit of repair like this in the memoQ environment (where it is really much, much easier to fix the problems of bad comment and link placement), I copied the entire source text to the target to enable me to export a cleaner source text file. I then opened this file in Microsoft Word and used various search and replace operations to fix the bad hyphenation and other problems like excess spaces. Replacing the hyphens had to be done occurrence-by-occurrence, because the style of writing in German meant that there were many legitimate instances of hyphens followed by spaces.
After all was done, the "before and after" looked like this:
BEFORE
AFTER
The remaining tags were all legitimate formatting tags for comments, hyperlinks, tabs after section numbering, etc. These do, of course, require attention and add complexity to the work still, so they must be included in the charges for the job. memoQ makes this calculation particularly simple by allowing weighting factors to be specified in the analysis. These are the settings I typically use for a German source text:
I find this usually represents a fair minimum for the additional effort in translation and quality assurance that tags require. In this case, of course, time charges for the cleanup apply, but as you can probably guess from comparing the two analysis tables above, the customer is actually saving a lot of money by paying me to clean up the mess, and the results will be a lot more usable. My cleaned-up version of the source text will also be returned in case the authors intend to make more revisions in the source - this will save more time and money by avoiding redundant cleanup in that case.