Showing posts with label autotranslatables. Show all posts
Showing posts with label autotranslatables. Show all posts

Sep 11, 2023

memoQ "Auto-translation Roundup": 14 September 2023 at 15:00 CET

 


This week's public lecture for the "memoQuickies Resource Camp" on Thursday, September 14, 2023, at 3:00 p.m. Central European Time (2:00 p.m. Lisbon time) will be a summary of currently available auto-translation rules on the course pages which are open to everyone (enrolled or not) and those restricted to registered participants. This is all about getting to work now with stuff that is ready to go, and how to adapt that stuff for clients with different requirements.

So if you want to get right  down to productive work using available memoQ auto-translation rulesets in translation, review and quality assurance without wasting time on learning bloody regex, this is for you.


A recording of the talk will be available to all registered course participants afterward.

Last week's lecture, "Auto-translation Rules for Everyone", is available here.

Oh, and next week at the same time we'll be talking about the memoQ Regex Assistant and all the cool libraries available for QA, filtering, find & replace and other tasks.....

Icons of the resources covered im the memoQuickies Resource Camp


Sep 4, 2023

New online course: "memoQuickies Resource Camp"

Summer is almost over, but technically, "camping season" will continue in memoQ World until November 30th. Or maybe January 31st, depending on how you count.

Today, a three-month journey of exploration begins, covering six important kinds of resources to make work with the memoQ translation desktop and server environments more pleasant and efficient... and profitable. This self-guided online course will give participants full access to my 14 years of cumulative experience as a memoQ user translating, managing projects and developing hundreds of solutions with this world-leading productivity tool.

Click here or on the icon bar above to have a look at the course description and to see (and maybe download) some of the publicly available information and resources for better work in many language pairs. 

The emphasis of teaching will shift to a new resource every two weeks (with auto-translation rules as the main topic for the first two weeks), but throughout the course, information will be added continuously to all topic sections as I trawl through, sort, upgrade and publish the best or most interesting stuff from my archives. And course participants have access to open virtual office hours each week on Thursdays and some other occasions, where any questions can be asked and special requests made.

A special enrollment discount of 40% is available for the first week (code: HALFOFFLAUNCH) until September 10th, but you can join at any time and work with any of the material posted, ask questions and receive feedback. Learning material and downloadable, ready-to-use and -adapt resources will continue to be added until the end of November, and the full course will remain online through January 2024. Enrollment fees and content are subject to change without notice.

Addendum 1: On Thursday afternoon, September 7th, 2023, a presentation was made to introduce the first course topic - "Auto-translation Rules for Everyone". The recording and slides can be found here.

Addendum 2: Payment options for groups and monthly budgets have been introduced now. These options enable teams, departments and organizations to obtain blocks of passes for their members to receive continuing professional education in translation workflow tools. The host site applies VAT and other taxes where relevant and generates appropriate invoices. All relevant information can be found at the bottom of the information and enrollment page.



Jun 5, 2021

Get better dates with memoQ

There's hope for all those incels stuck with RWS Trados Studio, Memsource, Wordfast, OmegaT and a host of other horrors if they are willing to make a change....


Not surprisingly, dates can be a real nuisance to translate and check, depending on the client's specifications. Target language specifications that include the use of elements such as non-breaking spaces can be particularly troublesome. But even apparently simple tasks like writing all dates in the target language as DDMonth-AbbreviatedYYYY or the like can go wrong far more than expected, and skilled reviewers easily get caught up in the flow of the text and overlook details of format (and often even correct content) in dates. This proved to be a shock to one LSP client who found hundreds of overlooked date errors in a large volume of recently reviewed text.

What's the solution to these inevitable human errors? Proper automation of the monkey work so professionals can concentrate on what they are good at: a fluent text that accurately reflects the intent of the original.


In the case of the QA horror of checking four source languages to ensure the dog's breakfast of date input formats (which also included day+month and month+year entries) regardless of capitalization, memoQ enabled a simple auto-translation ruleset to be created (a few hours' work, including testing and documentation); this was then attached to a project using a QA profile configured to check only against the enabled auto-translation rules, and BOOM! after about a minute, all the date errors in something like 100,000 translated words were revealed. The only false positives found were a few instances where times were written after the date, and the rule can be updated easily to avoid this issue.

I do a lot of date rule development for many languages, and I've published some of this in simplified forms on this blog. But interesting new tricks come up all the time. And I've found it useful when developing rules for others, who usually understand little about writing proper specifications that capture all the likely source input, to create special screening rules like the one shown in the first screenshot, which can be used to examine an entire large TM imported to the memoQ working grid, and see how the input and target texts vary. I used that expression on two large TMs in a view, and in just a few seconds, my laptop screen showed me all renderings of every English date in those TMs into Portuguese. While researching target formats for some new rules I also found quite a number of errors in the TM which I could have corrected had I cared to.

Having rules like this available in the translation phase can prevent quite a few errors to start with. I've found too many cases of dates in March translated as May and overlooked by both the translator and the reviewers. memoQ is - as far as I know - the only tool which will offer such conversions in a results table from which they can be inserted just like any other "terminology hit".

Other tools like RWS TradoZe Studio will allow you to use regular expressions for quality assurance (checking the text), but I'm not aware of any tool other than memoQ which allows you not only to include these checks in customized QA profiles but which can also provide them for on-the-fly review from a library of named expressions. That's what memoQ now does in version 9.8 (scheduled for release in June, within a few weeks) as shown in the first screenshot above.

This new rules library feature of memoQ makes it possible for the first time for users who have a life which does not involve wasting brain cells learning to program regular expressions to use the efforts of those who actually like that sort of thing and do it well. So with this tool, anyone can easily check for things like date errors and a lot more without knowing a single bit of regex syntax. That's some progress :-)

Stuff like dates just keeps getting better for memoQ users, leaving them more time for life and better stuff... like other kinds of dates.

If this kind of thing interests you, I think my friend Marek Pawelec may be teaching an in-person course in regular expressions for memoQ in July of this year, and he, I and others (including memoQ's Business Services unit) are available to help you with turnkey solutions to project challenges like those described here.

Jun 3, 2021

A Hebrew abbreviations "hint base" points the way for other languages

Years ago I published a guideline for how to create something like a term base for memoQ that can handle the irregularities one might find in the way German attorneys on tight deadlines might type the many abbreviations they use in crazy ways. The memoQ term base model can't cope with punctuation and many special characters, so it's basically impossible to use it to map something like "US-$" to the standard currency code "USD". But regular expressions in an auto-translation rule can do that, of course.

The same principle can be used simply to map abbreviations to their full expression so the translator can decode the abbreviation and decide how to render it. Here's an example of that in Hebrew:

This can, of course, be done in other languages, but the fellow who had this idea and asked me about it happens to be a Hebrew translator working into several target languages. I'm tempted to adapt one of my German abbreviation sets to map to the full German expression in the target to serve as an aid to translators who might not be as familiar with the abbreviations as I am and who are also not bound strictly to a particular target language expression. A cheat sheet, basically, or a "hint base" if there is such a thing.

The code for this is particularly simple. Here's a quick look at the resource in an external editor:


The basic "engine" is just a list (#abbreviations#). And the resource was created quickly using search and replace on a list of over 600 abbreviations in an Excel spreadsheet.

In the awful memoQ rule editor it looks like this:


Those who know Hebrew may note that some periods are out of place. I'm not an RTL expert, so I had a few issues with punctuation migrating as I moved data from one format to another, but someone familiar with issues like that can fix things without much ado. This was just a quick prototype to demonstrate feasibility. And a few minutes of search and replace work in a text editor beats entering more than 600 pairs manually in the built-in editor for memoQ. It would be nice if that damned editor included list import features that would read Excel files directly!

As with other auto-translation rules, certain characters may need to be represented by entities or uuencoding. The simple rule shown above can also be made more robust by dealing with variable punctuation, for example. Complexity can always be added. 

Many thanks to the translator colleague who shared his challenge and gave me something fun to do after a grueling day of mapping many messed-up date formats from a lot of different source languages I mostly don't know :-)

Jul 22, 2019

No comment... on memoQ "light" resources and their editors

The memoQ working environment includes a number of editing functions and modules, some better developed and more useful than others. Unfortunately, the terms "better" and "more useful" cannot be applied to many of those functions for creating and maintaining resources for auto-translation and many other functions. And, at the present time, integrated facilities for documenting the purpose and function of resources are usually limited to a short description field. The consequence of this can be, in a better case, some confusion, and if you are unlucky you might lose and or accidentally delete resources or apply a version unfit for use.

Auto-translation rule set editor with inadequate space for reading and writing the rules.

The auto-translation rule set editor is a particular headache for me. Scrolling back and forth in a tiny field to read and edit a long match rule (here, in the example of the number-matching auto-translation rules provided with the memoQ installation) is difficult and error-prone.

Even with the dialog-based editors which don't look much like editors (such as the example for QA options configuration below), it's hard to get an overview.


Unless I can compress all the relevant information into the description field for the resource, the only way I am going to get an overview of which functions are enabled is to go through all eight tabs for the QA profile in that dialog. Yuck. And the "why"? Maybe recorded in a notebook buried in the paper pile of my work table if I'm lucky.

There is a better way. And while that way is followed by a number of technically adept colleagues and developers I know, unfortunately it is not usually discussed and taught in workshops and other training venues, nor is it promoted by the software providers for memoQ in any way of which I am aware.

I typically recommend the use of specialized text editors, such as the free text and source code editor Notepad++ for most development, maintenance and documentation tasks involving memoQ light resources. It is available at no cost to everyone and offers simple functions to help you get an overview, edit and document your resources. Only a tiny bit of arcane knowledge is required.

Armed with such a tool, or even with the simple Windows Notepad application, there are a number of useful things that can be done, such as:

Add <!-- comments --> to the text of an MQRES light resource file saved from memoQ
The markers shown in red in the previous line can be added to a line in the file to provide explanatory comments or maintenance instructions. In files with regex content, I often use comments to explain to myself the use of syntax that I will otherwise forget and be unable to understand in a matter of hours or at best weeks.
Example of an auto-translation rule set with comments added
Note that these comments are stripped when resources are imported into memoQ, remaining only in the original external file. Thus, a workflow involving external development and documentation in the resource files, with imports to memoQ for testing and use, is highly desirable. If deficiencies are found in the resource, it should be corrected externally in Notepad++, etc. and re-imported, not fixed in-situ in memoQ, where no information will be present regarding the resource, its purpose and mysterious details.
Comments of this sort night be added to a QA profile, for example, to give a quick overview of the resource and its purpose in more detail than the description field allows (and I often forget to update that description field, because it is in the header of the file, which is usually not of much interest for developing and testing configurations, except to note a version number and a few details). 
Edit the resource more sensibly, using standard text editor features like search and replace
Often I'll decide to add nonbreaking spaces to a date or currency expression (or to the French output numbers in the "French Group" numbers auto-translation rule set provided with memoQ, which unfortunately probably still uses ordinary spaces as separators for thousands, millions, etc.), and this can be totally tedious in the internal editors of memoQ. Such tasks are much simpler in Notepad or Notepad++, for example.
It's also much simpler to find multiple instances of a word or structure that needs amendment or to do just about anything else when you can see all the content in a larger display.
Where resources are in fact simpler to develop inside memoQ, it is still worthwhile to export MQRES files as security. Comments added to these are a form of internal documentation which can avoid confusion and mistakes later when sorted file messes on a hard drive.
Teach and practice resource development and maintenance more effectively
A heavily commented resource file can be thought of as an easily portable "textbook" which includes a functional, importable example of its teaching. And when another person receives a copy of such a file as an example, it will be much easier to understand its structure and purpose and make any necessary changes.

Aug 24, 2018

Webinar: Auto-Übersetzungsregeln in memoQ - Planung, Anwendung und Pflege (am 12.09.2018)

Der Vortrag wurde aufgenommen und ist hier verfügbar.


Auto-Übersetzungsregeln gehören zu den nützlichsten, kaum benutzten Aspekten der memoQ-Arbeitsumgebung. Mit Hilfe dieser Regeln kann man unter anderem viel Zeit bei der Gestaltung und Qualitätssicherung musterbasierter Texte sparen, zum Beispiel bei Datumsangaben, bibliografischen Informationen, rechtliche Referenzangabe,Währungsausdrücken usw.

Diese Regeln basieren auf regulären Ausdrücken, aber solche Kompetenzen sind für ihren effektiven Einsatz nicht unbedingt vorausgesetzt. Viel wichtiger ist es, die geeignete Erfassung der Basisinformationen zu verstehen, sowie mögliche Variationen im Ausgangstext, damit technische Ressourcen für Erstellung und Pflege gezielt und richtig anzuwenden werden können.

Die geplanten zwei Stunden dieser kostenlosen Präsentation beinhalten Beispiele häufiger Anwendungsgebiete für den praktischen Einsatz dieser Technologie in der Übersetzung mit memoQ. Um die vorgeführten Inhalte üben und anwenden zu können, sind folgende Elemente notwendig bzw. empfohlen:
  • eine aktuelle memoQ-Lizenz (ohne funtioniert nichts!)
  • die kostenlose Software Notepad++ (stark empfohlen)
  • ein tabellenfähiges Textprogramm, wie z.B. Microsoft Word, Google Docs oder der OpenOffice-Editor 

Das Webinar findet am 12. September 2018 um 15 Uhr MEZ statt und läuft bis zu 2 Stunden. Die Teilnahme ist kostenlos, aber registrierungspflichtig. Der Erwerb brauchbarer Exemplare der vorgeführten Anwendungsbeispiele bzw. persönlich angepasster Versionen ist auch nachher möglich.

Falls Sie sich für weitere memoQ-Onlineschulungen interessieren, geht es hier zu der relevanten Umfrage.

May 21, 2018

Best Practices in Translation Technology: summer course in Lisbon July 16-21

As usual each year, the summer school at Universidade Nova de Lisboa is offering quite a variety of inexpensive, excellent intensive courses, including some for the practice of translation. This year includes a reprise of last year's Best Practices in Translation Technology from July 16th to 21st, with some different topics and approaches.

Centre for English, Translation and Anglo-Portuguese Studies

The course will be taught by the same team as last year – yours truly, Marco Neves and David Hardisty – and cover the following areas:
  • Good translation workflows.
  • Using voice recognition in translation.
  • Using machine translation in a humane, intelligent way.
  • Using checklists to improve communication in translation.
  • Using glossaries, bilingual texts and other references in multiplatform environments.
  • Good practices for using terminology and reference texts in the target language.
  • Planning and creating lists for auto-translation rules and the basics of regular expressions for filters.

Some knowledge of the memoQ translation environment and translation experience are required.

The course is offered in the evening from 6 pm to 10 pm Monday (July 16th) through Friday (July 20th), with a Saturday (July 21st) session for review and exams from 9 am to 2 pm. This allows free days to explore Lisbon and the surrounding region and get to know Portugal and its culture.

Tuition costs for the general public are €130 for the 25 hours of instruction. The university certainly can't be accused of price-gouging :-) Summer course registration instructions are here (currently available only in Portuguese; I'm not sure if/when an English version will be available, but the instructors can be contacted for assistance if necessary).

Two other courses offered this summer at Uni Nova with similar schedules and cost are: Introduction to memoQ (taught by David and Marco – a good place to get a solid grounding in memoQ prior to the Best Practices course) from  July 9–14, 2018 and Translation Project Management Tools from September 3–8, 2018.

All courses are taught in English and Portuguese in a mix suitable for the participants in the individual courses.

Jun 24, 2017

The multilingual toolkit for getting a date in Swahili


Some time ago, I was asked by IAPTI to provide some technical support for a developing effort to assist professional translators in various African regions. The flame of the Translators Without Borders center established a few years ago in Kenya has apparently sputtered out due to an incredibly silly anti-business model which undermined local professionals, so various initiatives were launched to help translators in the region grow stronger together and improve their professional practice.

Since memoQ is perhaps the best tool for managing the challenges of expert translation under the widest range of languages and conditions, I considered how I might contribute to solving some of these and reduce the frustrations of language barriers in Africa. I thought of all the business travelers there, as well as the NGOs and representatives of governments around the world who want a piece of what's there. All alone, strangers in a strange land, sweltering in some Nairobi hotel, how can these people even get a date in Swahili?

Once again, it's Kilgray to the rescue... with memoQ's auto-translation rules!

Using the various methods I have developed and published for planning and specifying auto-translation rules, I assembled an expert team for translation in Swahili, Arabic, Hebrew, English, German, Portuguese, Spanish, French, Russian, Hungarian, Dutch, Finnish, Polish and Greek to draft the rules for getting long dates in Swahili.

And using the Cretinously Uncomplicated Process for Identifying Dates (CUPID), these results can be transmogrified quickly to support lonely translators working from German, French and English into Arabic or from German, French, English and Spanish into Portuguese, for example, or in any combination of the languages applied for Swahili dates or others as needed.

With memoQ and regex-based auto-translation, you'll never be stuck for a quality-controlled date in any language!

Jun 16, 2017

Troubleshooting memoQ light resource import problems

The other day I sent a friend some updated auto-translation rules for currency expressions; a short time later I received a message that they would not import into memoQ. The error message displayed was the following:


Now the problem here might seem obvious, but the name of the file I sent was nothing like any rule she already had installed,


In the example shown below, the source of the trouble is more obvious, but if there are a lot of resources in the list shown in the Resource Console or elsewhere, the redundancy of the name in the import dialog and an existing resource name in the list might not stand out so clearly....


In the MQRES file (for the memoQ light resource), the "trouble  spot" is in the XML header at the top of the file. This can be seen by opening it in any text editor (in this case I used Notepad++ to show line numbering);


In this case, the fifth line contains the name that will be applied to the resource after it is imported. The <Name> tags are found in all kinds of memoQ light resources, and the same problem will occur if a redundancy is found during import. Here is an example from a memoQ ignore list (used to exclude certain words from error indications by spellchecking functions):


There are a couple of ways to avoid or correct these problems:
  • First of all, when a ruleset is edited, the text enclosed by the Name tags should be altered. It's probably a good idea to update the Description as well. The FileName  is actually ignored and need not be updated; a difference with the real name of the MQRES file will not cause any trouble with an import.
  • When importing a light resource, you can always change the information read from the Name and Description tags of the MQRES file. This avoids the conflict.

              
      
  • The name and description of an existing light resource can be edited via the Properties of the resource in the Resource Console or Project Home > Settings, Accessing the resource via memoQ Options will currently (as of version 8.1) not show the Properties.

               
memoQ's "light resources" - the portable configurations and information lists to assist various translation tasks - are one of the environment's greatest strengths, but the generally bad state of the associated editing tools and unhelpful error handling continue to cause a lot of unnecessary confusion among users. Key people at Kilgray are not unaware of this problem, and for years there has been a debate regarding new features versus actual usability of the features already present. When you encounter difficulties like the one described above - or other troubles using this generally excellent, leading translation assistance tool - it is important to communicate your concerns to Kilgray Support (support@kilgray.com). 

Without appropriate feedback from the wordface, there is often really no way for the designers and product engineers to understand and prioritize the challenges of usability. I can understand the reluctance of those who have used other tools for many years, where it was clear that their requests for bug fixes or other improvements were largely ignored, to take such action, but it really does make a difference, though not always on a time scale of hours or days. Weeks, months, sometimes years may pass before important changes are made, but usually this is because the urgency of the matter has not been communicated with sufficient clarity, or there are in fact, more pressing matters which require attention. But in fact no serious matters are seldom ignored by those responsible, as nine years as a satisfied user have shown me.

Jun 11, 2017

On a TEUR with German financial translation


Currency expressions occur in great variety in German financial translation, and it is often a great nuisance to type and check the corresponding expressions, correctly formatted, in the target language. One group of such expressions are those involving thousands of euros, typically written in German as "TEUR". However, depending on the proclivities of the source text author, other forms such as T€, kEUR or k€ may be encountered.

On the target side, clients might want to see figures like "TEUR 1.352" rendered in a number of ways: perhaps EUR 1,352 thousand, perhaps €1,352k, perhaps something else.

I have described before how to map out source and target equivalents for developing auto-translation rules or regex-based quality checking instruments to use as the basis for development specifications and case testing as well as how to document the structure and reasoning of the respective rules.

Here you can download an example of possible solutions to the specific problem described above. The downloadable ZIP archive contains two different rulesets for each of the English target text formats cited above; these may be adapted to fit the particular requirements of a client as needed.

There are, of course, quite a number of other currency expressions one routinely encounters when translating financial texts or other business documents, and the diversity of client preferences for target language formats can be considerable. In many cases, it is worthwhile to document which rulesets correspond to which client's preference, perhaps even including client names in the filename to keep things straight. Thus "KPMG_TEUR-to-English" might be the  ruleset name for the client KPMG's preference for how to translate those particular expressions to English.

Busy financial translators who use memoQ and who have discovered the benefits of rulesets like these tell me time and again how many hours or days of effort are saved routinely by using tools like these in translation and subsequent quality checks. They are a "secret weapon" in an often competitive environment with a lot of short, stressful deadlines.

Those who wish to have rulesets of their own to handle the specific requirements of their clients can turn to a number of sources for help. Kilgray's Professional Services department can develop custom rules, as can competent consultants such as Marek Pawelec or yours truly. One caveat: in hiring development experts for memoQ tools based on regular expressions (regex), it is generally a good idea to work with consultants whose primary focus is memoQ. Regular expressions are used in many other environments, such as Apsic Xbench and SDL Trados Studio (as well as many others having nothing to do with translation), but without an intimate, daily working acquaintance with memoQ, developers are often unable to understand the best approaches for working with the memoQ environment and it is all too possible to spend a lot of money on custom work which proves to be unusable, for example because the complex rules take many minutes to load each time a project or document is opened, because the developer did not break the problem down efficiently into its component parts. But done right, these rulesets are an investment which can pay enormous dividends for many specialist translators.

Mar 4, 2017

Documenting auto-translation rule development for memoQ

In an recent article, I described my simple method of recording examples of structured information like dates, financial expressions or legal references to help developers plan auto-translation rules (or other features using regular expressions, such as Regex Tagger rules) in memoQ and other applications. These are a sort of simplified performance specification - a table of examples showing how the rules should "perform", what they should do: what patterned source language expressions are to be transformed into particular structured expressions in the target language.

The need for proper documentation of such efforts does not end there, however. It is very important, especially for more complex sets of rules, that there be clear documentation of the purpose and logic of the rules developed, and that this documentation be present
  • in the rules themselves (as comments) and
  • in external documents to be used as references for troubleshooting, maintenance and further development.
Auto-translation rules and other resources using regular expressions should not be scripted and maintained for the long run in memoQ itself or in any other environment which does not allow thorough commenting of the regular expressions used. Without comments, it is simply too easy to destroy functioning rules by forgetting why they were written a certain way once-upon-a-time, and an environment able to use comments also allows old rules to be "commented out" (disabled, but still available for reference or later re-use) while new versions are tested. That is basically impossible with memoQ's internal resource editors at the present time. And to make matters worse, if auto-translation rules are edited inside memoQ, their order changes, sometimes with dire consequences if functionality depends on the rule order. Try sorting out problems like that in a set of 70 or so rules.

Excerpt from a large set of currency format rules with extensive comments. These comments are stripped when
the rules are imported into memoQ, so all maintenance should be done externally in a tool like
Notepad++.
As I began to revise and improve old rules that I created years ago for dates and currency expressions, I found that it was helpful to create a record of what changes I had made - and why I made them - and keep this information in a tabular form for easy reference and re-use.
Click to access a PDF sample of my rule development record (2 pages)
The graphic above is one example of how I maintain my personal records of some work developing regular expressions. I usually include
  • descriptions of all information recorded
  • a specific example on which I will base the general rule
  • a simple ("fragile") version of the rule part (source input and target output) with only the most essential elements; this is not error-tolerant, but it is the easiest to understand and the first place to look if something isn't working as I would like it to
  • more robust variations which take into account differences in spacing, punctuation, etc. or include things like non-breaking spaces that might be desired in the output (this can get cluttered and hard to read)
  • color-marking for easier identification of some elements
  • comments about why things are written as they are or about possible improvements or problems
This record is a template of sorts from which rules can be assembled very quickly or rules can be re-purposed for other languages or formats in a way that is easy to follow and catch mistakes. Such records are also helpful if the rules are to be shared with other developers or maintained by someone else.

My example is certainly not the final word in project documentation for such efforts; it is simply part of a set of personal tools to help me work more efficiently with the limited time I have. Professional development and consulting organizations often have far more extensive and detailed systems of project documentation; when I was part of one such shop nearly 20 years ago, my (downloadable) 2-page example might easily have filled twenty pages of very important-looking professional technobabble. Life's too short for shit like that anymore.

But if you value your time as a developer or your investment as one who hires others to develop such useful rules, it pays big dividends in most cases to demand some sort of clear, systematic and accurate record of how your special rules, filters, etc. were developed so that they can be maintained and improved in the future.

Feb 27, 2017

Planning special rules for structured "expressions" and multi-word abbreviations

Translators and editors often deal with what I'll call "structured expressions" or "patterned data" in many forms, which include:
  • long and short dates (2016-01-13; 1/13/16; 13.01.2016; January 13, 2016; 13th January 2016; etc.
  • time expressions (14:35; 2:35 pm; 2:35 PM; 2:35 p.m.; etc.
  • currency expressions (EUR 2.3 million; € 2,300,000; €2.3m; etc.) 
  • legal references (Section 14a paragraph 3 line 2; section 14a (3) line 2; etc.
  • bibliographical references for chapters, pages, margin notes, etc.
  • and much more.
There is also a wealth of abbreviations for multiple word expressions in some categories of text; favorites in German include:
  • in Verbindung mit (variously written as i.V.m., i. V. m., iVm or some typoed hybrid of the aforementioned with spaces and periods included or forgotten depending on the authors' preferences and degree of care)
  • im Sinne des (i.S.d., i. S. d., iSd, etc.)
These can be devilishly hard to check efficiently for consistency or other quality factors in a long text, and for the translation, there is often no single "right" way to format the target text equivalents, with many individual preferences to be found with translation buyers. Even with a good style guide (all too rare anyway), these issues can be challenging time-wasters.

Translation assistance tools such as Apsic Xbench, SDL Trados Studio and others, even memoQ, have various approaches to making life easier for a translator or editor faced with these challenges. Unfortunately for most people, these approaches usually involve the use of "regular expressions" or "regex" as nerds affectionately call it. Not an easy thing even for many hardcore techies!

On past occasions when I have written about the use of regex in translation tools, I have usually stated clearly that the best approach for the best, most reliable results is to have the regex "rules" for handling the text developed by a knowledgeable third party. The experts who deal with this stuff routinely can often reduce a task that would take a semi-skilled person like myself hours or even days to the time for a coffee break, and even if a task takes a while and runs up a bit of a bill, it's much more likely to be done right the first or second time.

But... there's a catch usually. Most of these regex fireaters are not skilled in mind reading, many are not translators, and even those familiar with translation challenges might not be familiar with your working languages or your particular subject areas and their possibly unique challenges. So effective communication is really, really important (it always is, of course, but here even more so if you are dealing with a verbally challenged, monolingual math freak who might be your local expert for regex).

Even for areas I know reasonably well and languages I more or less master, I am often frustrated by help requests from colleagues and clients who need special rulesets developed for a client's preferences for date and currency information, because the request is not clear in its scope and detail, and many important cases are left out, so the end result is not fully satisfactory.

Over the years and with a lot of back and forth (sometimes inside my own head with yours truly as my nightmare of a "client"), I have developed a system of simple documentation for planning and testing rules to help translate and quality check patterned information or multi-word abbreviations. This system provides an easy structure for non-techies (or even hardcore techies) to organize the help request for most efficient handling. Here is an example of part of such a planning sheet for a recent project involving Arabic:


When the time comes to test, just copy the source text column into a separate file, add whatever variations you want to the examples to test your accomodation of typos, etc. and then load that file as a "translation text" for testing in your working environment. If you have the same information for another, overlapping language pair, such as German and English, it is easy to couple that to make a ruleset which maps multiple source languages to a target language. An example of such a result is a memoQ auto-translation ruleset for mapping long dates and month-plus-day dates from German, English, French and Spanish into Portuguese which can be obtained here.

This simple, tabular approach to data collection to plan regular expression rules has made me a lot more efficient at such tasks and faciulitated the re-use of data to make new rulesets for clients and colleagues (or myself) as needs arise. The liberal commenting of examples can be very helpful; information to include which could affect rule structure might involve capitalization, location in a sentence, variations or differences in particular contexts, etc.

For my own work, rulesets include a series for dates, currency and legal reference formats from German to English for generic and client-specific use for US and UK English. With the help of these tabular planning sheets, I can adapt any of these quickly for most other languages.

For tracking the development of rules and their improvement history I have another set of templates which I use for systematic planning and identification of areas to improve. That will be discussed on another occasion.

Feb 20, 2017

Building a regex-savvy "termbase" in memoQ


For years I have been frustrated by and dissatisfied with how abbreviations are handled in the current memoQ termbase model. The crux of the problem is the handling of the periods in the expressions. This can be seen with termbase entries like the following, for example:


If the abbreviation "Art." appears in the source text, only the second source entry - the one without the period - will give a match result in memoQ. The first entry is simply ignored.

An additional problem which one would face, even if the terminal period character in the term did not pose a problem, is that authors are often notoriously variable in the way they write abbreviations. Take, for example, the abbreviation for the German expression "in Verbindung mit", usually written as "i.V.m."

In recent legal translation work, I have encountered this expression written as above, but also as "i. V. m." (with spaces), "iVm" (no spaces no periods) and sloppily typed variations like "iV.m" or "i. V.m." What's a poor wordworker to do?

The answer came to me while refining a set of auto-translation rules for bibliography formatting and legal references. These, too, can suffer from similar troubles: "page 7" might be abbreviated as "p. 7", but in the sloppy chaos of source texts poorly edited one might find "p.7", "p 7", "p7" or even variations with the letter capitalized, like "P.7". If you are translating nearly 1000 references in a bibliography, robust shortcuts are very helpful and save a lot of time, and if those shortcuts are based on memoQ auto-translation rules, they can also be used in a QA profile to ensure that every bit matches correctly.

As the screen capture from a memoQ Facebook group above suggests, the way to go about this is to identify which parts of the expression might vary with different deliberate and accidental typing. These are usually spaces and periods in the case of abbreviations; sometimes, particularly with German legal abbreviations, capitalization and dashes may play roles as well. (I tore my hair out not long ago trying to understand an Austrian legal text referring to two laws, which differed in their three-letter abbreviations only by a dash inserted after the first letter of one.)

In regular expressions, the question mark character means "zero or one" of whatever character precedes the question mark. So if I want a rule that acts in the case of one or no periods, I put a question mark after the period character. And because in the language of regular expressions, a period is shorthand for any character, if I want to talk about an actual period ("."), I have to precede that character by a backslash ("\."). In the technical jargon of Nerdworld that is known as "escaping the period" and there is no escaping such syntax if you want a regular expression rule about periods, period.

Spaces (normal or non-breaking ones) are represented by an escaped lowercase "s": "\s". So a matching rule for the English abbreviation "e.g" which catches a lot of typing variations might be

e\.?\s?g\.?

And in German, the target replacement rule might be

d.h.

Of course, if a typist is sloppy, there might be more than one space, or a comma might be typed accidentally instead of a period (the keys are adjacent, and if your screen is as dirty as mine gets sometimes, your eyes might not notice); capitalization might also differ accidentally or based on context. The regular expressions for matching can be adapted to handle all these cases if need be.

Rules of this type are not particularly difficult to construct, but refining them to accommodate all the variations you are likely to encounter may require an expert hand. Thus, as I have suggested before,. the average user should focus on documenting all the possible source variations clearly in a table which includes the desired target equivalents, and this table should be given to an expert (Kilgray support, a qualified consultant like Marek Pawelec or a technical programmer familiar with regular expressions and their use in memoQ). Trust me, this will save a lot of frayed nerves and probably significant time and money as well.

So now I am building a few memoQ auto-translation rulesets which are essentially fault-tolerant abbreviation glossaries. These, together with the similar rulesets for formatting bibliographical references and references to sections, paragraphs, lines, margin notes, etc. in laws, have been very helpful in reducing the time spent translating messy legal source texts, and the accuracy of the work has been improved significantly. Give it a try for your translation challenges!

Dec 28, 2016

Go Figure (with memoQ!)

When translating patents, legal briefs, reports, manuals and many other kinds of documents I inevitably encounter figure references to photographs and illustrations in the text as well as the labeled captions for these. In this morning's translation of a petition in a nullity suit, one such reference takes the form in Verbindung mit Figur 1,  but it might just as well appear as

Fig. 1
Fig 1
Abb. 1
or
Abbildung 1

in this or some other text; in documents with multiple and/or sloppy authors I might even find a mix of all these in the same text.

As I value consistency in writing even when the client might not care, I try to translate all of these to the same form in English where it makes sense to do so. That might be Figure 1 or Fig. 1 depending on the situation and the styleguide stipulated for the project.

But when I finish the 10,000 or so words for this job and need to do my final check before sending it to the client, I expect to be a little tired, and I want to use my attention and energy to focus on the accuracy and reading comfort of my translation. In doing so I tend to miss little details like the occurrence of "Fig. 1" on page 32 as opposed to "Figure 1" on the other 40 pages. That is why I use the QA feature of memoQ to check the consistency with which I have translated the figure references as well as other matters such as the accurate use of special terminology for the project.

The specific feature I use here for quality assurance is


an auto-translation rule set (aka "autotranslatables"), which is highlighted and selected in the screenshot of the project's settings above.

As I have stated many times before, autotranslatables should be used, but not created by the average translator. Aside from the fact that the regular expressions involved are not particularly easy even for most of the nerds among us, there are a lot of little subtleties that make the difference between a well-functioning rule set and annoying garbage, and even the "experts" struggle with this for sophisticated rules.

But the present example of Figure mapping is a comparatively simple case which can illustrate the principles and some of the "risks" to mere mortals.



My rule set for mapping figures from many German forms to a particular English form consists of a single rule.

All of the possibilities that I expect in German are compiled in a list, along with the English expression for each, and this translation pair list is named #figurelist# and is found on the corresponding dialog tab in the memoQ rule set editor for autotranslatables. (I usually edit rules externally in Notepad++ where I can comment them liberally, but in this case I felt no need to do so.) This named list is used as a variable in the regular expression for the rule to describe a source text match.

(#figurelist#)\.?\s+?\b(\d+)\b

Jeepers. That regex for the source text looks complicated, doesn't it? Wouldn't (#figurelist#) \d+ be just as good? After all, it seems to work just fine. Well, except that the list would need a few extra entries to account for abbreviations with and without periods.

No. "(#figurelist#) \d+" is total, incompetent crap. Here are some reasons why:
  • It is more efficient to express the possibility of a period after the text for "Figure" with the regex "\.?",  because you'll never have to worry about abbreviations with or without periods in your lists. Mine will get longer, as I'll probably expand these rules to cover Portuguese as well and use the same rule for both Portuguese and German sources.
  • There may or may not be a space or even extra spaces after the Figure expression. Simply typing a standard space after the (#figurelist#) group means that it must be present and it must be an ordinary space to match. If it's missing or someone typed a non-breaking space (a reasonable thing to do to keep both parts of "Figure 1" on the same line), the rule will not work! Using \s+? to express the possibility of 0 to n spaces after "Fig." or whatever is in fact the right way to go.
  • If you test the "simple" crappy regex, you'll also find that "Abb. 14" gives to results: Figure 1 and Figure 14. That is because the rule does not stipulate that the second part must be a whole "word", so the substring match with the first character also gives a result. Bad, bad, bad. The chaos that this sort of mistake can cause with more complex rules like currency expressions used in important financial translations is frightening.
The regex for the result also appears more complex than it should be, but there is a reason behind that as well. Instead of the simple $1 $2 (first group followed by a space followed by the second group), I specified output with a non-breaking space, because it looks rather unfortunate to have a line wrap in the middle of the expression for a figure. One sees that a lot, because it's a nuisance to remember to type non-breaking spaces all the time on the keyboard. This rule can also be used to check the use of the non-breaking space; an ordinary space will generate a warning when the memoQ QA profile is run with the autotranslatables check activated.

There are many ways in which regular expression rule sets can enhance the user experience and the quality of translation results when working in memoQ. It is not hard to use these rules, but it is beyond most users to create and maintain their own rule sets. Therefore
  • Kilgray should include more useful examples of rule sets (in addition to the very helpful number rules) in future releases of memoQ
  • The average user should ask the help of Kilgray Support for simple rules they need (in most cases this would fall under the usual commitment of paid support and maintenance for the year)
  • memoQ users should work with Kilgray's Professional Services department or other competent consultants to devise robust rule sets to boost their translation and quality assurance productivity. Beware of casual advice found in forums or social media; much of it does not consider issues like the problems described above despite the aggressive insistence one might see for a particular "solution". Truly, you get what you pay for :-)

Post scriptum:
An yet ye hack by night and sun, the work of regex be never done.
Of course something was forgotten in the example here. The myriad styles and customs of source text authors will inevitably offer up challenging variants to break your well-crafted rules. Today's is a text full of figure references like Abbildung 4.12, which would refer to the twelfth figure in the fourth chapter. For this the modified rule might be 

(#figurelist#)\.?\s+?(\b\d+\.?\d+?\b) 

Or perhaps not quite. Try it and you'll see a few problems. This is just another example of why it is good to make use of professional resources to help you with these challenges and to have a systematic way of recording and elaborating them. I'll explain more about such an effective system for planning and documentation in a future article. I've noticed that the "experts" in the translation field often care little for the usual standards of project specification, perhaps because they are sick and tired of translation projects with so many specification documents for those who know better.

Dec 15, 2016

Validating Roman numerals in translation QA

The issue of Roman numerals in my translation work has been at the back of my mind for a few years now, but the pain level had not been such that I got around to dealing with it. It comes up time and again in legal translation work: references to the "X. Senat" or the like which mess up segmentation (and require a bit of regex to do a new segmentation rule); references to "Art. VII" of some law (I need to catch the typos like "VIII"); source text errors like "VIIII"; and of course dates like MCMXXIV, etc. and century references.

For simple matters I used regex which would capture and reproduce "Roman numerals", but erroneous data using the right letters would also be accepted:

[MDCLXVI]+

That is, of course, rather useless for QA which checks the correctness of the expression in the source text. So with a bit of thought I came up with:


Without the word border syntax ("\b"), non-standard expressions like "VIIII" might appear to be validated in the interface of memoQ, for example, because the whole express would be marked green in the source text, and one might not notice that it was resolved into "VIII" and "I".

These expressions can be used in various ways in any CAT tool that supports regular expressions, such as SDL Trados Studio or memoQ.

If you want this typing aid and QA tool as a memoQ autotranslatable (along with a little demo data file), you can get it here.

Aug 24, 2016

memoQ autotranslatables: a partial antidote for drudgery

I'm currently working on a stack of legal pleadings for a patent nullity suit – lots of "urgent" words to churn by the end of the week. And after 10,000 or so of them, I got pretty damned tired of typing out the translation of text citations of the form "Spalte 7, Zeilen 34 bis 45" as "Column 7, Lines 34 to 45".

In fact, it was really starting to piss me off. In such situations, I try not to get mad but to get an autotranslatable ruleset instead. This is perhaps one of the most under-utilized productivity tools in memoQ.


So the next time I ran into a text that fit that format, the translation was offered as an autocompletable phrase as soon as I typed the first letter:


Of course life isn't usually that simple, at least not life with technology. And authors? Well, they seem to believe firmly in the old saying that "consistency is the hobgoblin of little minds". So of course the text also includes lots of references in the form "Spalte 7, Zeilen 34 - 45", with or without spaces around the hyphen. No problem, just add a rule for that (or if you are more clever, edit the single rule to cover the variations):



Now I am not one to advocate that the unwashed masses of translators – or even the washed ones – run out and learn to write regular expressions. I've programmed more computer languages and systems than I can possibly remember for about 45 years now, and I can't keep most of the autotranslatable rules in my head if I don't use them for a week or more after yet-another-refresher, so it would be stupid and hypocritical of me (or just bloody naive) to expect most people to mess with nerdy shit like this. But....

... a few simple rules and a couple of nice "recipe templates" to start can go a long way. And sometimes it pays not to be too clever; I have one highly sophisticated set of rules for complex legal citations that was written by a professional programmer, and it's unusable. Takes minutes to load even on a very fast computer, which is a huge pain in the backside every time a project is opened in memoQ. My more verbose, brute force approach to legal reference autotranslation may not be elegant, but it loads much faster and covers 90% or more of what I encounter. Maybe a case of where it's smart to be a little stupid.

There are lots of good tutorials out there on regex (regular expressions), including a few YouTube webinar videos from Kilgray, the memoQ Help, a few chapters in old books of mine, discussions in the Yahoogroups lists and more.

The examples above require the knowledge of only a few rules:
  • Chunks of the source text to be analyzed are grouped in parentheses. In the examples shown, those groups are merely where numbers occur.
  • Numbers are represented by the escape code "\d". If there might be more than one digit, add a plus sign: \d+.
  • Spaces are represented by the escape code "\s". In the rules you can usually just type a space instead, but if you have to cover cases where it might be missing or where more than one might have been typed (usual sloppiness), then use the escape code, followed by an asterisk, which means "zero or more" of whatever it is put after: \s*.
  • For the rest of the text to match, you can usually type it just the way it occurs as I have done above. For the target translation rules, you can usually just type the literal text you want, with the groups represents by the numerical order in which they occur, preceded by a dollar sign. So the first group (parentheses set) in the source is $1, the second is $2, etc. Of course the order can be changed in the target; it's just not necessary in this case, but in autotranslatable rules for dates this happens rather often.
Not only will the little rules I wrote for this big job save me a lot of typing, I can also use them in a QA profile to check that I have made no errors by switching numbers, missing a space or anything else in my translation. That is done by marking the appropriate checkbox on the first tab of the QA profile you plan to use:


Perhaps such things are worth a little effort in your projects once in a while....