Showing posts with label quality. Show all posts
Showing posts with label quality. Show all posts

Jun 5, 2021

Get better dates with memoQ

There's hope for all those incels stuck with RWS Trados Studio, Memsource, Wordfast, OmegaT and a host of other horrors if they are willing to make a change....


Not surprisingly, dates can be a real nuisance to translate and check, depending on the client's specifications. Target language specifications that include the use of elements such as non-breaking spaces can be particularly troublesome. But even apparently simple tasks like writing all dates in the target language as DDMonth-AbbreviatedYYYY or the like can go wrong far more than expected, and skilled reviewers easily get caught up in the flow of the text and overlook details of format (and often even correct content) in dates. This proved to be a shock to one LSP client who found hundreds of overlooked date errors in a large volume of recently reviewed text.

What's the solution to these inevitable human errors? Proper automation of the monkey work so professionals can concentrate on what they are good at: a fluent text that accurately reflects the intent of the original.


In the case of the QA horror of checking four source languages to ensure the dog's breakfast of date input formats (which also included day+month and month+year entries) regardless of capitalization, memoQ enabled a simple auto-translation ruleset to be created (a few hours' work, including testing and documentation); this was then attached to a project using a QA profile configured to check only against the enabled auto-translation rules, and BOOM! after about a minute, all the date errors in something like 100,000 translated words were revealed. The only false positives found were a few instances where times were written after the date, and the rule can be updated easily to avoid this issue.

I do a lot of date rule development for many languages, and I've published some of this in simplified forms on this blog. But interesting new tricks come up all the time. And I've found it useful when developing rules for others, who usually understand little about writing proper specifications that capture all the likely source input, to create special screening rules like the one shown in the first screenshot, which can be used to examine an entire large TM imported to the memoQ working grid, and see how the input and target texts vary. I used that expression on two large TMs in a view, and in just a few seconds, my laptop screen showed me all renderings of every English date in those TMs into Portuguese. While researching target formats for some new rules I also found quite a number of errors in the TM which I could have corrected had I cared to.

Having rules like this available in the translation phase can prevent quite a few errors to start with. I've found too many cases of dates in March translated as May and overlooked by both the translator and the reviewers. memoQ is - as far as I know - the only tool which will offer such conversions in a results table from which they can be inserted just like any other "terminology hit".

Other tools like RWS TradoZe Studio will allow you to use regular expressions for quality assurance (checking the text), but I'm not aware of any tool other than memoQ which allows you not only to include these checks in customized QA profiles but which can also provide them for on-the-fly review from a library of named expressions. That's what memoQ now does in version 9.8 (scheduled for release in June, within a few weeks) as shown in the first screenshot above.

This new rules library feature of memoQ makes it possible for the first time for users who have a life which does not involve wasting brain cells learning to program regular expressions to use the efforts of those who actually like that sort of thing and do it well. So with this tool, anyone can easily check for things like date errors and a lot more without knowing a single bit of regex syntax. That's some progress :-)

Stuff like dates just keeps getting better for memoQ users, leaving them more time for life and better stuff... like other kinds of dates.

If this kind of thing interests you, I think my friend Marek Pawelec may be teaching an in-person course in regular expressions for memoQ in July of this year, and he, I and others (including memoQ's Business Services unit) are available to help you with turnkey solutions to project challenges like those described here.

Mar 4, 2019

What evil lurks in the results from your language service provider?

Let me start by disclosing that although I have a registered limited company through which I provide translation, training and technical consulting services for translation processes, I am essentially a sole trader who is not unreasonably, though not correctly, referred to as a freelancer much of the time. I have a long history of friendship and consulting support with the honorable owners of quite a few small and medium language service companies and of a few large ones. I vigorously dispute any foolish claims that there is no such need for such companies, and I see a natural alliance and many shared interests between the best of them and the best of independent professionals in the same sector.

But as Sturgeon's Law states so well, "ninety percent of everything is crap", and that would apply in equal measure to translation brokers and translators I suspect, though of course this is influenced by context. But what context can justify this translation of a data privacy statement from German to English? Only the section headers are shown here to protect against sensory overload and blown mental circuits:

The rest of the text is actually worse. This is the kind of thing some unscrupulous agencies take money for these days.

Why, pray tell, was the section numbering translated so variously into English? Well, if you know anything about the mix-and-match statistical crapshoot that is SMpT (statistical machine pseudotranslation) and its not-as-good-as-you-think wannabe alternatives, it's easy to guess the frequency of certain correlations in English with German numbers followed by a period.

And clearly, the agency could not even be bothered to make corrections, and the robotic webmaster put the text up, noticing nothing, where it remained for about a year to embarrass a rather good company which I hold in high esteem.

What's the moral of this story? Take your pick from the many reasonable options. "Reasonable" does not include doing business with the liars and thieves who will try to sell you on the "value proposition" of machine translation to cut costs.

A skilled translator knowledgeable in the subject matter and trained in dictation techniques paired with a good speech recognition solution or transcriptionist can beat any human post-edited machine translation process for both volume and quality. And a skilled summarizer reading source texts and dictating summaries in another language can blow them both away as a "value proposition".

One thing that is too often forgotten in the fool's gold rush to cheap language (dis)service solutions is - as noted by Bevan et alia - exposure to machine-translated output over any significant period of time has unfortunate effects one the language skills (reading, writing and comprehension) of the victims working with it. This has been confirmed time and again by translation company owners, slavelancers and other word workers. Serious occupational health measures are called for, but to date little or nothing has been done in this regard.

And when human intelligence is taken out of play or impaired by an automated linguistic lobotomy, the results inevitable gall in the lower quartile of the aforementioned 90%. Really crappy crap.

As another of my favorite fiction authors used to comment: TANSTAAFL. There ain't no such thing as a free lunch. And trust is always good, but these days you need to verify that your service providers really give you what you have paid for and don't pass off crap like you see in the example above.

Jun 24, 2017

The multilingual toolkit for getting a date in Swahili


Some time ago, I was asked by IAPTI to provide some technical support for a developing effort to assist professional translators in various African regions. The flame of the Translators Without Borders center established a few years ago in Kenya has apparently sputtered out due to an incredibly silly anti-business model which undermined local professionals, so various initiatives were launched to help translators in the region grow stronger together and improve their professional practice.

Since memoQ is perhaps the best tool for managing the challenges of expert translation under the widest range of languages and conditions, I considered how I might contribute to solving some of these and reduce the frustrations of language barriers in Africa. I thought of all the business travelers there, as well as the NGOs and representatives of governments around the world who want a piece of what's there. All alone, strangers in a strange land, sweltering in some Nairobi hotel, how can these people even get a date in Swahili?

Once again, it's Kilgray to the rescue... with memoQ's auto-translation rules!

Using the various methods I have developed and published for planning and specifying auto-translation rules, I assembled an expert team for translation in Swahili, Arabic, Hebrew, English, German, Portuguese, Spanish, French, Russian, Hungarian, Dutch, Finnish, Polish and Greek to draft the rules for getting long dates in Swahili.

And using the Cretinously Uncomplicated Process for Identifying Dates (CUPID), these results can be transmogrified quickly to support lonely translators working from German, French and English into Arabic or from German, French, English and Spanish into Portuguese, for example, or in any combination of the languages applied for Swahili dates or others as needed.

With memoQ and regex-based auto-translation, you'll never be stuck for a quality-controlled date in any language!

Jun 5, 2017

Technology for Legal Translation

Last April I was a guest at the Buenos Aires University Facultad de Derecho, where I had an opportunity to meet students and staff from the law school's integrated degree program for certified public translators and to speak about my use of various technologies to assist my work in legal translation. This post is based loosely on that presentation and a subsequent workshop at the Universidade de Évora.

Useful ideas seldom develop in isolation, and to the extent that I can claim good practice in the use of assistive technologies for my translation work in legal and other domains it is largely the product of my interactions with many colleagues over the past seventeen years of commercial translation activity. These fine people have served as mentors, giving me my first exposure to the concepts of platform interoperability for translation tools, and as inspirations by sharing the many challenges they face in their work and clearly articulating the desired outcomes they hoped to achieve as professionals. They have also generously and frequently shared with me the solutions that they have found and have often unselfishly shared their ideas on how and why we should do better in our daily practice. And I am grateful that I can continue to learn with them, work better, and help others to do so as well.

A variety of tools for information management and transformation can benefit the work of a legal translator in areas which include but are not limited to:
  • corpus utilization,
  • text conversion,
  • terminology management,
  • diverse information retrieval,
  • assisted drafting,
  • dictated speech to text,
  • quality assurance,
  • version control and comparison, and
  • source and target text review.
Though not exhaustive, the list above can provide a fairly comprehensive basis for education of future colleagues and continued professional development for those already active as legal translators. But with any of the technologies discussed below, it is important to remember that the driving force is not the hardware and software we use in technical devices but rather the human mind and its understanding of subject matter and the needs of the particular task or work process in the legal domain. No matter how great our experience, there is always something more and useful to be learned, and often the best way to do this is to discuss the challenges of technology and workflow with others and keep an open mind for new approaches with promise.


Reference texts of many kinds are important in legal translation work (and in other types of translation too, of course). These may be monolingual or multilingual texts, and they provide a wealth of information on subject matter, terminology and typical usage in particular contexts. These collections of text – or corpora – are most useful when the information found in them can be read in context rather than isolation. Translation memories – used by many in our work – are also corpora of a kind, but they are seriously flawed in their usual implementations, because only short segments of text are displayed in a bilingual format, and the meaning and context of these retrieved snippets are too often obscure.

An excerpt from a parallel corpus showing a treaty text in English, Portuguese and Spanish

The best corpus tools for translation work allow concordance searches in multiple selected corpora and provide access to the full context of the information found. Currently, the best example of integrated document context with information searches in a translation environment tool is found in the LiveDocs module of Kilgray's memoQ.

A memoQ concordance search with a link to an "aligned" translation
A past translation and its preview stored in a memoQ LiveDocs corpus, accessed via concordance search
A memoQ LiveDocs corpus has all the advantages of the familiar "translation memory" but can include other information, such as previews of the translated work as well. It is always clear in which document the information "hit" was found, and corpora can also include any number of monolingual documents in source and target languages, something which is not possible with a traditional translation memory.

In many cases, however, much context can be restored to a traditional translation memory by transforming it into a "document" in a LiveDocs corpus. This is because in most cases, substantial portions of the translation memory will have its individual segment records stored in document order; if the content is exported as a TMX file or tab-delimited text file and then imported as a bilingual document in a LiveDocs corpus, the result will be almost as if the original translations had been aligned and saved, and from a concordance hit one can open the bilingual content directly and read the parts before and after the text found in the concordance search.


Legal translation can involve text conversion in a broad sense in many ways. Legal translators must often deal with hardcopy or faxed material or scanned files created from these. Often documents to translate and reference documents are provided in portable document format (PDF), in which finding and editing information can be difficult. Using special software, these texts can be converted into documents which can be edited, and portions can be copied, pasted and overwritten easily, or they can be imported in translation assistance platforms such as SDL Trados Studio, Wordfast or memoQ. (Some of these environments include integrated facilities for converting PDF texts, but the results are seldom as suitable for work as PDF or scanned files converted with optical character recognition software such as ABBYY FineReader or OmniPage.)


Software tools like ABBYY FineReader can also convert "dead" scanned text images into searchable documents. This will even work with bad contrast or color images in the background, making it easier, for example, to look for information in mountains of scanned documents used in legal discovery. Text-on-image files like the example shown above completely preserve the layout and image context of the text to be read in the best way. I first discovered and used this option while writing a report for a client in which I had to reference sections of a very long, scanned policy document from the European Parliament. It was driving me crazy to page through the scanned document to find information I wanted to cite but where I had failed to make notes during my first reading. Converting that scanned policy to a searchable PDF made it easy to find what I needed in seconds and accurately cite its page number, etc. Where there is text on pictures, difficult contrast and other features this is often far better for reference purposes than converting to an MS Word document, for example, where the layouts are likely to become garbled.


Software tools for translation can also make text in many other original formats accessible to translators in an ergonomically simpler form, also ensuring, where necessary, that no text is overlooked because of a complicated layout or because it is in an easily overlooked footnote or margin note. Text import filters in translation environments make it easy to read and translate the words in a uniform working environment, with many reference tools and other help available, and then render the translated text back into its original format or some more useful bilingual format.

An excerpt of translated patent claims exported as a bilingual table for review

Technology also offers many possibilities for identifying, recording and controlling relevant terminology in legal translation work.


Large quantities of text can be analyzed quickly to find the most frequent special vocabulary likely to be relevant to the translation work and save these in project glossaries, often enabling that work to be organized better with much of the clarification of terms taking place prior to translation.  This is particularly valuable in large projects where it may be advisable to ensure that a team of translators all use the same terms in the target language to avoid possible confusion and misunderstanding.

Glossaries created in translation assistance tools can provide terminology hints during work and even save keystrokes when linked to predictive, "intelligent" writing features.


Integrated quality checking features in translation environments enable possible deviations of terminology or other issues to be identified and corrected quickly.


Technical features in working software for translation allow not only desirable terms to be identified and elaborated; they also enable undesired terms to be recorded and avoided. Barred terms can be marked as such while translating or automatically identified in a quality check.

A patent glossary exported from memoQ and then made into a PDF dictionary via SDL Trados MultiTerm
Technical tools enable terminology to be shared in many different ways. Glossaries in appropriate formats can be moved easily between different environments to share them with others on a team which uses diverse technologies; they can also be output as spreadsheets, web pages or even formatted dictionaries (as shown in the example above). This can help to ensure consistency over time in the terms used by translators and attorneys involved in a particular case.

There are also many different ways that terminology can be shared dynamically in a team. Various terminology servers available usually suffer from being restricted to particular platforms, but freely available tools like Google Sheets coupled with web look-up interfaces and linked spreadsheets customized for importing into particular environments can be set up quickly and easily, with access restricted to a selected team.


The links in the screenshot above show a simple example using some data from SAP. There is a master spreadsheet where the data is maintained and several "slavesheets" designed for simple importing into particular translation environment tools. Forms can also be used for simplified data entry and maintenance.


If Google Sheets do not meet the confidentiality requirements of a particular situation, similar solutions can be designed using intranets, extranets, VPNs, etc.


Technical tools for translators can help to locate information in a great variety of environments and media in ways that usually integrate smoothly with their workflow. Some available tools enable glossaries and bilingual corpora to be accessed in any application, including word processors, presentation software and web pages.


Corpus information in translation memories, memoQ LiveDocs or external sources can be looked up automatically or in concordance searches based on whole or partial content matches or specified search terms, and then useful parts can be inserted into the target text to assist translation. In some cases, differences between a current source text and archived information is highlighted to assist in identifying and incorporating changes.


Structured information such as dates, currency expressions, legal citations and bibliographical references can also be prepared for simple keystroke insertion in the translated text or automated quality checking. This can save many frustrating hours of typing and copy revision. In this regard, memoQ currently offers the best options for translation with its "auto-translation" rulesets, but many tools offer rules-based QA facilities for checking structured information.


Voice recognition technologies offer ergonomically superior options for transcription in many languages and can often enable heavy translation workloads with short deadlines to be handled with greater ease, maintaining or even improving text quality. Experienced translators with good subject matter knowledge and voice recognition software skills can typically produce more finished text in a day than the best post-editing operations for machine pseudo-translation, with the exception that the text produced by human voice transcription is actually usable in most situations, while the "gloss" added to machine "translations" is at best lipstick on a pig.


Reviewing a text for errors is hard work, and a pressing deadline to file a brief doesn't make the job easier. Technical tools for translation enable tens of thousands of words of text to be scanned for particular errors in seconds or minutes, ensuring that dates and references are correct and consistent, that correct terminology has been used, et cetera.

The best tools even offer sophisticated tools for tracking changes, differences in source and target text versions, even historical revisions to a translation at the sentence level. And tools like SDL Trados Studio or memoQ enable a translation and its reference corpora to be updated quickly and easily by importing a modified (monolingual) target text.

When time is short and new versions of a source text may follow in quick succession, technology offers possibilities to identify differences quickly, automatically process the parts which remain unchanged and keep everything on track and on schedule.


For all its myriad features, good translation technology cannot replace human knowledge of language and subject matter. Those claiming the contrary are either ignorant or often have a Trumpian disregard for the truth and common sense and are all too eager to relieve their victims of the burdens of excess cash without giving the expected value in exchange.

Technologies which do not assist translation experts to work more efficiently or with less stress in the wide range of challenges found in legal translation work are largely useless. This really does include machine pseudo-translation (MpT). The best “parts” of that swindle are essentially the corpus matching for translation memory archives and corpora found in CAT tools like memoQ or SDL Trados Studio, and what is added is often incorrect and dangerously liable to lead to errors and misinterpretations. There are also documented, damaging effects on one’s use of language when exposed to machine pseudo-translation for extended periods.

Legal translation professionals today can benefit in many ways from technology to work better and faster, but the basis for this remains what it was ten, twenty, forty or a hundred years ago: language skill and an understanding of the law and legal procedure. And a good, sound, well-rested mind.

*******

Further references

Speech recognition 

Dragon NaturallySpeaking: https://www.nuance.com/dragon.html
Tiago Neto on applications: https://tiagoneto.com/tag/speech-recognition
Translation Tribulations – free mobile for many languages: http://www.translationtribulations.com/2015/04/free-good-quality-speech-recognition.html
Circuit Magazine - The Speech Recognition Revolution: http://www.circuitmagazine.org/chroniques-128/des-techniques
The Chronicle - Speech Recognition to Go: http://www.atanet.org/chronicle-online/highlights/speech-recognition-to-go/
The Chronicle - Speech Recognition Is in Your Back Pocket (or Wherever You Keep Your Mobile Phone): http://www.atanet.org/chronicle-online/none/speech-recognition-is-in-your-back-pocket-or-wherever-you-keep-your-mobile-phone/

Document indexing, search tools and techniques

Archivarius 3000: http://www.likasoft.com/document-search/
Copernic Desktop Search: https://www.copernic.com/en/products/desktop-search/
AntConc concordance: http://www.laurenceanthony.net/software/antconc/
Multiple, separate concordances with memoQ: http://www.translationtribulations.com/2014/01/multiple-separate-concordances-with.html
memoQ TM Search Tool: http://www.translationtribulations.com/2014/01/the-memoq-tm-search-tool.html
memoQ web search for images: http://www.translationtribulations.com/2016/12/getting-picture-with-automated-web.html
Upgrading translation memories for document context: http://www.translationtribulations.com/2015/08/upgrading-translation-memories-for.html
Free shareable, searchable glossaries with Google Sheets: http://www.translationtribulations.com/2016/12/free-shareable-searchable-glossaries.html

Auto-translation rules for formatted text (dates, citations, etc.)

Translation Tribulations, various articles on specifications, dealing with abbreviations & more:
http://www.translationtribulations.com/search/label/autotranslatables
Marek Pawelec, regular expressions in memoQ: http://wasaty.pl/blog/2012/05/17/regular-expressions-in-memoq/

Authoring original texts in CAT tools

Translation Tribulations: http://www.translationtribulations.com/2015/02/cat-tools-re-imagined-approach-to.html

Autocorrection for typing in memoQ

Translation Tribulations: http://www.translationtribulations.com/2014/01/memoq-autocorrect-update-ms-word-export.html

Jun 1, 2017

"In termo qualitas" – IAPTI terminology talk

On June 17, 2017 Dr. Rodolfo Maslias, an expert terminologist with the European Parliament, will discuss the provision and improvement of terminology in today's challenging translation and communication environments. To register, send e-mail to info.request@iapti.org.



Aug 24, 2016

memoQ autotranslatables: a partial antidote for drudgery

I'm currently working on a stack of legal pleadings for a patent nullity suit – lots of "urgent" words to churn by the end of the week. And after 10,000 or so of them, I got pretty damned tired of typing out the translation of text citations of the form "Spalte 7, Zeilen 34 bis 45" as "Column 7, Lines 34 to 45".

In fact, it was really starting to piss me off. In such situations, I try not to get mad but to get an autotranslatable ruleset instead. This is perhaps one of the most under-utilized productivity tools in memoQ.


So the next time I ran into a text that fit that format, the translation was offered as an autocompletable phrase as soon as I typed the first letter:


Of course life isn't usually that simple, at least not life with technology. And authors? Well, they seem to believe firmly in the old saying that "consistency is the hobgoblin of little minds". So of course the text also includes lots of references in the form "Spalte 7, Zeilen 34 - 45", with or without spaces around the hyphen. No problem, just add a rule for that (or if you are more clever, edit the single rule to cover the variations):



Now I am not one to advocate that the unwashed masses of translators – or even the washed ones – run out and learn to write regular expressions. I've programmed more computer languages and systems than I can possibly remember for about 45 years now, and I can't keep most of the autotranslatable rules in my head if I don't use them for a week or more after yet-another-refresher, so it would be stupid and hypocritical of me (or just bloody naive) to expect most people to mess with nerdy shit like this. But....

... a few simple rules and a couple of nice "recipe templates" to start can go a long way. And sometimes it pays not to be too clever; I have one highly sophisticated set of rules for complex legal citations that was written by a professional programmer, and it's unusable. Takes minutes to load even on a very fast computer, which is a huge pain in the backside every time a project is opened in memoQ. My more verbose, brute force approach to legal reference autotranslation may not be elegant, but it loads much faster and covers 90% or more of what I encounter. Maybe a case of where it's smart to be a little stupid.

There are lots of good tutorials out there on regex (regular expressions), including a few YouTube webinar videos from Kilgray, the memoQ Help, a few chapters in old books of mine, discussions in the Yahoogroups lists and more.

The examples above require the knowledge of only a few rules:
  • Chunks of the source text to be analyzed are grouped in parentheses. In the examples shown, those groups are merely where numbers occur.
  • Numbers are represented by the escape code "\d". If there might be more than one digit, add a plus sign: \d+.
  • Spaces are represented by the escape code "\s". In the rules you can usually just type a space instead, but if you have to cover cases where it might be missing or where more than one might have been typed (usual sloppiness), then use the escape code, followed by an asterisk, which means "zero or more" of whatever it is put after: \s*.
  • For the rest of the text to match, you can usually type it just the way it occurs as I have done above. For the target translation rules, you can usually just type the literal text you want, with the groups represents by the numerical order in which they occur, preceded by a dollar sign. So the first group (parentheses set) in the source is $1, the second is $2, etc. Of course the order can be changed in the target; it's just not necessary in this case, but in autotranslatable rules for dates this happens rather often.
Not only will the little rules I wrote for this big job save me a lot of typing, I can also use them in a QA profile to check that I have made no errors by switching numbers, missing a space or anything else in my translation. That is done by marking the appropriate checkbox on the first tab of the QA profile you plan to use:


Perhaps such things are worth a little effort in your projects once in a while....

Feb 23, 2014

Cleaning up a crappy OCR job for translation

It's a sad fact in the professional work of translators that a lack of understanding on how to deal effectively with various PDF formats causes enormous loss of productivity and results which are not really fit for purpose. The aggressive insistence of many colleagues possessed of a dangerous Halbwissen on using half-baked methods and inappropriate tools contributes to the problem, but, bowing to the wisdom about arguing with fools, I now mostly sit back with a bemused and amused smile and watch the tribulations of those who believe in salvation by PDF import filters and cheap or free OCR. "TANSTAAFL" is a true as it ever was.

Just before the weekend I got an inquiry from an agency client I rather like. Nice people, good attitude, but struggling sometimes trying to find their way with technology despite some in-country "expert" training. This inquiry looked a bit like ripe fish at first glance. The smell got stronger after I was told that because the corporate end client had converted the PDF for their annual report and begun to edit the mess (and comment it heavily too) in the OCR file that this would be all there was to work with. It was a thoroughly appetizing sight when imported into a translation environment:


There are so many issues in that tossed salad of translation terror that I don't even know where to start describing them.

The screenshot above was in memoQ. How does it look in SDL Trados Studio? Often just as messy. In this case, this was the result in an older version of Studio:

SDL Trados Studio choked and refused to import the file!

I do have the latest version of SDL Trados Studio 2014, but unfortunately it's on a system that does not yet Microsoft Office, because I refuse to bow to Microsoft's insistence that I must buy a Portuguese version of that software. No MS Office, no file import in this case with SDL Trados Studio. memoQ fortunately has not needed MS Office to import its old file formats since the release of memoQ 6.0.

Ugly OCR trash like this file is all too common at this time of year, and as I am busy compiling the syllabus for the workshop I want to do on better living with well-used technology for legal and financial translators, I felt obliged to take this one on as a teaching example. It's actually not as bad as it looks. On the other hand, the best approach may not always be obvious, and the best solution for one document may not apply as well or at all to another.

My first approach was to use Dave Turner's CodeZapper macros. This isn't as straightforward as it used to be since I downgraded from Microsoft Office 2003 to later versions; for some reason the toolbar refuses to stay loaded between work sessions, and there's no way I can keep track of all the abbreviations for macros on it.


I can't deal with anything more complicated than clicking the "CZL" option for "Code Zapper lite", which did a rather decent job on the heavy mess above:


But all was not quite as well as it seemed:


Text in the header and footer remained trashed, and the heavy use of comments and tabbed lists meant that there were plenty of legitimate tags to deal with which were just too confusing with the DVX-like mess of memoQ's default import and display for an RTF file.

So I went for a kinder, gentler approach. I changed my import filter settings in memoQ:


There is actually seldom any good reason to import an RTF or DOC file into memoQ using the default filter settings. And marking those two little checkboxes at the bottom often accomplishes much of what CodeZapper does. Sometimes less. A bit more in this case.


The header and footer texts were absolutely clean. Don't let the extra tags in this sample fool you: overall, there were fewer than in the code-zapped file. Now there are still a number of issues to be seen in the screenshot above, including paragraph breaks in the middle of a sentence and awful manual hyphenation (many instances of that in the whole text) and joys like badly placed comments and links which mess up the text and prevent term identification by the software:



Source editing features of memoQ (F2) enable issues like the two above to be dealt with easily:



After a bit of repair like this in the memoQ environment (where it is really much, much easier to fix the problems of bad comment and link placement), I copied the entire source text to the target to enable me to export a cleaner source text file. I then opened this file in Microsoft Word and used various search and replace operations to fix the bad hyphenation and other problems like excess spaces. Replacing the hyphens had to be done occurrence-by-occurrence, because the style of writing in German meant that there were many legitimate instances of hyphens followed by spaces.

After all was done, the "before and after" looked like this:

BEFORE

AFTER

The remaining tags were all legitimate formatting tags for comments, hyperlinks, tabs after section numbering, etc. These do, of course, require attention and add complexity to the work still, so they must be included in the charges for the job. memoQ makes this calculation particularly simple by allowing weighting factors to be specified in the analysis. These are the settings I typically use for a German source text:


I find this usually represents a fair minimum for the additional effort in translation and quality assurance that tags require. In this case, of course, time charges for the cleanup apply, but as you can probably guess from comparing the two analysis tables above, the customer is actually saving a lot of money by paying me to clean up the mess, and the results will be a lot more usable. My cleaned-up version of the source text will also be returned in case the authors intend to make more revisions in the source - this will save more time and money by avoiding redundant cleanup in that case.


Feb 15, 2014

Indestructible Italian quality!

Language service providers like yours truly spend a lot of time talking about the high cost of crap quality. But that applies to most things, really. An old friend of mine who struggled at the start of his professional life paying his way with handyman work while creating beautiful custom furniture knew that he could afford only the finest tools though not always meat to go with the bread on his table.

I drink a lot of coffee, and I enjoy it in a variety of ways. As those who have visited my kitchen can attest, it's almost like a fetish with the various pots, presses and filters, each teasing different qualities of taste from a well-roasted bean. I thought by now I would have learned most of what I needed to know to ensure good results in the "coffee kitchen".

Distraction has taught me - or reminded me - of a few things lately. Three times now I have become involved in work or gone off shopping with friends and forgotten a moka pot on the fire. The first time, I returned to a house filled with toxic smoke from the burned plastic handles and top knob, and I was grateful the house had not burned down and the dogs were still alive. More recently I left my favorite Bialetti Brikka pot on the flame for several hours. Twice.

If anyone wants to say nasty things about Italian engineering, I will have to plead for the defense. In contrast to the complete destruction of the cheap moka pot after an hour, the Bialetti pot just needed a bit of scrubbing. Even the gasket was OK, which amazed me. The durability of the pots that cost me about €30 is so much beyond that of a €7 pot that to compare them is almost a crude joke. Oh yes, and I don't know any other manufacturer who offers such an excellent pressure valve for crema at that price.

Image from Alexandre Enkerli, Creative Commons Attribution-Share Alike 2.0 Generic license
There are obvious analogies to our work as translators, of course. But I don't need to insult anyone's intelligence by explaining them. I'll just go enjoy another shot of coffee from my indestructible Italian-engineered pot.

Aug 31, 2013

The Miracle of Late Language Translation and Modern Quality

This morning a colleague was kind enough to share a link to a thoughtful essay by Italian to English translator Wendell Ricketts, whom I have long considered to be one of the professionals among us with the clearest insight into some of the practical and philosophical problems of current translation practice. In his article Please mind the gap: defending English against 'passive' translation, Wendell describes the train wreck of unprofessional translation into one's later acquired languages quite well.

I particularly like how he distinguishes the cases of limited distribution languages and French, Italian, German and Spanish (FIGS, but be careful before you eat them... dogs have urinated on the low-hanging fruit). There are also some interesting statistics shared, such as
"An average of 59% of ProZ translators into English from Spanish, German, or French are not native-English speakers."
There's an old saying in German about how even with a golden ring an ape remains an ugly creature, and those translating into a language acquired in school or in bed who assume that revisions by a "native speaker" will make the result "all better" would do well to remember that. I've seen what they consider to be acceptable results often enough, and I often see no ugly ape there but at best a gut pile from one which has been too long in the sun :-)

Of course, in an age when charlatans work the conference routes selling snake oil and machine translation "miracles" and conjure quality in a bottle with metrics received on faith from technocrats with little understanding of real language, and presumed reputable purveyors of translation assistance technologies sprinkle themselves liberally with these elixirs and exude their perfume, Wendell's commentary is most likely to be dismissed with a knowing smirk and a learnèd discourse on fit-for-purpose quality and customer satisfaction.

I smile with some pleasure when I hear such Wise Words as low-grade linguistic sausage purveyors and their suppliers would have us believe. There is a certain economic dynamic in the circles of fantasy fandom which may lead them to conclude that they are on the right path, as an individual lemming might think with the comfort of the crowd around him.

Until there comes the cliff on which the skeptical competition will stand and look down on the remains of those who trusted in Common Sense Advice and translators without native language mastery of their target language for critical communication with customers and prospects.

Aug 16, 2013

memoQ AutoCorrect: mysteries revealed

Actually, AutoCorrect isn't that mysterious to those familiar with it. Many Microsoft Office users love it or hate it. I usually love it when I type English, but when I switch between languages in the same document, strange mutations occur in my words and I often wonder how I could possibly have typed some of the things I seem to have typed and of course did not.

Last December when I started the research to update my book of memoQ tips (which is still in progress, because the software is a fast-moving target to describe), I found a way to migrate the AutoCorrect lists from Microsoft Word to memoQ (and vice versa). This was a happy day for me, as a Dutch partner had been asking for exactly that for a very long time, and Kilgray's Support had not been able to offer a solution. I never did get around to blogging my findings, but a few months later, a similar solution was published in the Kilgray Knowledgebase. It states that it's perhaps only for migrating AutoCorrect lists from MS Word 2003, but I used an old macro from MS Word 98 when I worked out the problem, and if that still functions for MS Word 2010, then I'm sure Kilgray's posted solution must be fine for new versions. (Just be careful to use UTF-8 as the code page of text files you transfer or there may be trouble.)

But the best solution was actually published a few years earlier by Val Ivonica. In Portuguese. She included the macro code, and I like her macro (or the one she got from someplace) better. For some strange reason, the only really good information available on memoQ AutoCorrect up to now that I could find is in Portuguese. There are some nice examples of useful AutoCorrect shortcuts for periods of a year from William Cassemiro on the Janela Tradutória blog.

I was quite surprised to learn that many users of memoQ have no idea what AutoCorrect is; Déjà Vu offers the same feature, but I think it's missing in the various Trados versions, possibly because of the history of Trados Workbench as an application used primarily in the MS Word environment. The Kilgray documentation I could find was rather skimpy and seemed entirely focused on typing shortcuts. The idea of correcting spelling or vocabulary differences between language variants wasn't anywhere I could find it.

So I put together this "little" overview of how AutoCorrect works in memoQ and how and where to manage the AutoCorrect list resources there. It's a start... perhaps Kilgray or someone else can fill in the missing bits.


Time  Description
0:38  Activating AutoCorrect in an open project
1:47  AutoCorrect in action while typing
3:30  How the "primary" AutoCorrect list "rules"
3:55  Slide show: overview of AutoCorrect
4:59  Slide show: Three places to manage AutoCorrect

Jul 24, 2013

What good is memoQ fuzzy term matching?


When Kilgray introduced fuzzy term matching with the release of memoQ 2013, I was first concerned with how it worked after a few puzzling tests of the feature. Discussions with the development team soon cleared up that mystery, and I wrote an article describing the current fuzzy state of term matching technology in the translation environment tool that has done such a fine job of waking SDL and others from the long slumber of innovation that prevailed in the last decade.

But questions still remained in the minds of most users as they asked why they should care about this feature and what good it would really do for them.

The answer to that has become clearer for me as I have used the feature in recent weeks and noticed certain things. Like the fact that crappy spelling in my source texts is not as much of a burden for term matching any more:


This actually applies to more than just bad spelling. Those who translate from English will benefit from the fact that fuzzy term matching will help them if the UK source term is in the glossary but the author of the text used an American spelling. I cope with problems caused by old and new spelling conventions in German as well as the fact that a great many Germans cannot agree on how their compound words should be glued together. And my Portuguese friends tell me every week about the hassles of the spelling reform in progress in that linguistic corner.

Fuzzy term matches is currently not implemented for QA checking in memoQ, but I think it would make sense for Kilgray to add this feature to allow fuzzy term matching for QA on the source side. It could be a bit of a disaster to have it on the target side, however, for reasons I will leave readers to guess.

For those who want to set their termbases to use fuzzy matching by default in a particular language, here is a short video that shows how to change the termbase properties and how to change to match settings for legacy terms to "fuzzy":


I was initially a bit skeptical of the latest version of memoQ, but as this feature and a few others have begun to "sink in", while I still don't feel comfortable with the company's hyperbole over new features like LQA, which is largely pointless for freelance translators, I do feel confident in saying that fuzzy term matching is a reason for most of us to seriously consider upgrading to memoQ 2013. This will be even more the case if it is added to the QA features.

Ah, but what about the change to the comments function, Kevin? You really hated that!

There's more to say on that topic now. Some of it is even good.

Jun 3, 2013

Innovation? No comment. A black mark for memoQ 2013.

Update August 2013: problem solved.

"This comment thing is the last straw. I am definitely not upgrading to the new version of memoQ!"
The response from a friend who asked me to show how the comment function in memoQ had devolved in the latest version (memoQ 2013, aka version 6.5) revealed the frustration of a user whose main interest in new product features for much of the past year has been how to disable or avoid them. Unfortunately for her, there appears to be no escape from the latest innovation, which one member of the memoQ Yahoogroups list suggested would become known as Commentgate. This may be a good example of the old adage "if it ain't broke, don't fix it".

The old comment function has been at the heart of my use of memoQ for years, and the ability to export comments to share with my clients was one of the main reasons I pushed Kilgray to introduce the bilingual RTF table exports which are so popular for editing and translation. Comments added in a word processor (such as Microsoft Word) could be re-imported to the memoQ project and reviewed. It was all very quick and simple.

The old memoQ comment dialog with a personal note on edits required
If the comments included questions about a text or a term, the response in the comment column of the RTF table would be included when the bilingual file was re-imported, which was often convenient for editing purposes. A record of questions and answers could be maintained fairly easily in the editable comment field.

All that has changed with memoQ 2013. Disastrously so in the initial release on May 31. In their eagerness to implement an LQA quality assurance model relevant to only a minority of users, mostly the sort of agencies who prefer metrics in place of actual quality, Kilgray dynamited the previous straightforward, robust comment function, adding dropdown selection fields for "severity level" and scope.


That wouldn't be quite so bad despite the extra steps needed to add a correctly classified comment now if it were possible to edit the comments. It is not possible to edit comments in memoQ 2013 in the current release. Nor are all the comments included in an RTF table export. Only the last comment is included; all others are lost:


If a comment is altered in the bilingual RTF file, when the bilingual is re-imported to the project, a new comment is created and classified as "Information" applicable to the entire row:


This is all really a shame. In the effort to push server-based workflows and cater to a limited special interest group, memoQ's architects have managed to sabotage one of the tools (or two, depending on how you count) which have contributed to their great success in recent years. And unfortunately, unlike other recent, often irritating "innovations", such as the various target autotext options that drive many users of dictation software batty, this new type of comment can't be switched off so that we can work in the old, accustomed way.

Many users have already objected strenuously to this broken functionality, and some compromise solutions have been suggested by Kilgray. One of these, which involves a delimited export of all the comments in the RTF bilingual export, would be reasonable. Whatever changes are made to commentary in memoQ 2013, I hope this will include making comments editable very soon with appropriate rights.

The two selection field in the new comments dialog have a default behavior that is probably useful in some cases. Once a comment has been created, the next comment made in the text will assume the same "severity" rating applies (not necessarily a valid assumption, but if I am going through a text marking things of the same type this can be useful). All comments are assumed to apply to the target text by default. This is a bit of a nuisance to me, as most of the comments I make in a file refer to errors or unclear expressions in the source text. But really, for the way I work, the two classification steps with the dropdown menus are just extra work and additional sources of possible errors and/or confusion, so I would be quite happy to bypass these altogether.

I do like the idea of a comment history for a text. This would be relevant and useful for the way I work. But overall, the current implementation of the comment function in memoQ 2013 does not serve my interests at all and creates unnecessary complications for me and many other users.