Oct 29, 2019

Bilingual EU legislation the easy way in #xl8


Translators of European languages based in the EU and many others deal often with citations of EU legislation or need to consult relevant EU legislation for terminology in their translations. One popular source of information for that is the EUR-LEX website, which provides a convenient archive of legislation and related information, with the possibility of multilingual text displays, as seen here:


Some years ago, I published a description of how data from these multilingual EUR-LEX displays can be transferred to translation memories or other corpora for reference purposes, and more recently I produced a video showing this same procedure. But some people don't like the paragraph-level alignment format of the EUR-LEX displays, and these can also occasionally be seriously out of sync for some reason, as in this example (or worse):


Now I don't find that much of a nuisance when I use memoQ LiveDocs, because I can simply view the full bilingual document context and see where the corresponding information really is (kind of like leaving alignments in memoQ uncorrected until you actually find a use for the data and determine that the effort is worthwhile), but if you plan to feed that aligned data to a translation memory, it's a bit of a disaster. And many people prefer data aligned at the sentence level anyway.

Well, there is a simple way to get the EU legislation texts you want, aligned at the sentence level, with the individual bitexts ready to import into a translation memory, LiveDocs corpus or other reference tool. See that document number above with the large red arrow pointing to it? That's where you start....

Did you know that much of the information available in EUR-LEX is also available in the publicly available DGT translation memories? These are sentence-level alignments. But most people go about using this data in a rather klutzy and unhelpful way. The "big data" craze some years ago had a lot of people trying to load this information into translation memories and other places, usually with miserable results. These include:

  • the inability to load such enormous data quantities in a CAT tool's TM without having far more computer RAM than most translators ever think they'll need;
  • very slow imports, some apparently proceeding on a geological time scale; 
  • data overload - so many concordance hits that users simply can't find the focused information they need; and
  • system performance degradation, with extremely sluggish responses in a wide variety of tasks.
Bulk data is for monkeys and those who haven't evolved professionally much beyond that stage. Precision data selection makes more sense, and enables better use of the resources available. But how can you achieve that precision? If I want the full bilingual text of EU Regulation No. 575/2013 in some language pair, for example, with sentence-level alignment, how can I find that quickly in the vast swamp of DGT data?

Years ago, I published an article describing how it is better to load the individual TMX files found in the downloadable ZIP archives from the DGT into LiveDocs so that the full document context can be seen from the concordance searches. What I didn't mention in that article is that the names of those individual TMX files correspond to the document numbers in EUR-LEX

Armed with that knowledge, you can be very selective in what data and how much you load from the DGT collection. For example, if you organize the data releases in folders by year...


... and simply unpack the ZIP files in each year's folder...


... each folder will contain TMX files...


... the names of which correspond to the document number found in EUR-LEX. So a quick search in Windows Explorer or by other means can locate the exact document you want as a TMX file ready to import into your CAT tool:


These TMX files typically contain 24 EU languages now, but most CAT tools will filter just the language pair you want. So the same file can usually give you Polish+French, German+English, Portuguese+Greek or whatever combination you need among the languages present.

I still prefer to import my TMX data into a LiveDocs corpus in memoQ, and there I can use the feature to import a folder structure, and in the import dialog, I simply write the name of the file I want, and all other files (thousands of them) are promptly excluded:


After I enter the file name in the Include files field, I click the Update button to refresh the view and confirm that only the file I want has been selected. Depending on where in memoQ you do the import, you may have to specify the languages (Resource Console) to extract or not (in a project, where the languages are already set). Of course, the data can also be imported to a translation memory in memoQ, but that is an inferior option, because then it is not possible to read the reference document in a bilingual view as you can in a LiveDocs corpus; only isolated segments can be viewed in the Concordance or Translation results pane.

How you work with these data and with what tools is up to you, but this procedure will provide you with a number of options for better data selection and improved access to the reference data you may need for EU legislation without getting stuck in the morass of millions of translation units in a performance-killing megabomb TM.

Sep 26, 2019

10 Tips to Term Base Mastery in memoQ! (online course)

Note: the pilot phase for this training course has passed, free enrollment has been closed, and the content is being revised and expanded for re-release soon... available courses can be seen at my online teaching site: https://transtrib-tech.teachable.com/
In the past few years I have done a number of long webinars in English and German to help translators and those involved in translation processes using the memoQ environment work more effectively with terminology. These are available on my YouTube channel (subscribe!), and I think all of them have extensive hotlinked indexes to enable viewers to skip to exactly the parts that are relevant to them. A playlist of the terminology tutorial videos in English is available here.

I've also written quite a few blog posts - big and small - teaching various aspects of terminology handling for translation with or without memoQ. These can be found with the search function on the left side of this blog or using the rather sumptuous keyword list.

But sometimes just a few little things can get you rather far, rather quickly toward the goal of using terminology more effectively in memoQ, and it isn't always easy to find those tidbits in the hours of video or the mass of blog posts (now approaching 1000). So I'm trying a new teaching format, inspired in part by my old memoQuickie blog posts and past tutorial books. I have created a free course using the Teachable platform, which I find easier to use than Moodle (I have a server on my domain that I use for mentoring projects), Udemy and other tools I've looked at over the years.

This new course - "memoQuickies: On Better Terms with memoQ! 10 Tips toward Term Base Mastery" - is currently designed to give you one tip on using memoQ term bases or related functions each day for 10 days. Much of the content is currently shared as an e-mail message, but all the released content can be viewed in the online course at any time, and some tips may have additional information or resources, such as videos or relevant links, practice files, quality assurance profiles or custom keyboard settings you can import to your memoQ installation.

These are the tips (in sequence) that are part of this first course version:
  1. Setting Default Term Bases for New Terms
  2. Importing and Exporting Terms in Microsoft Excel Files
  3. Getting a Grip on Term Entry Properties in memoQ
  4. "Fixing" Term Base Default Properties
  5. Changing the Properties of Many Term Entries in a Term Base
  6. Sharing and Updating Term Bases with Google Sheets
  7. Sending New Terms to Only a Specific Ranked Term Base
  8. Succeeding with Term QA
  9. Fixing Terminology in a Translation Memory
  10. Mining Words with memoQ
There is also a summary webinar recorded to go over the 10 tips and provide additional information.
I have a number of courses which have been developed (and may or may not be publicly visible depending on when you read this) and others under development in which I try to tie together the many learning resources available for various professional translation technology subjects, because I think this approach may offer the most flexibility and likelihood of success in communicating necessary skills and knowledge to an audience wider than I can serve with the hours available for consulting and training in my often too busy days.

I would also like to thank the professional colleagues and clients who have provided so much (often unsolicited) support to enable me to focus more on helping translators, other translation project participants and translation consumers work more effectively and reduce the frustrations too often experienced with technology.


Aug 28, 2019

The challenge of light resource updates with many projects in memoQ

"Templates" take two forms in memoQ: the configuration option for equipping new projects with relevant resources in an automated way to save time and avoid forgetting important references or other information, which was introduced several years ago, and the older sort of "template" - a configured, existing project for a particular client or subject area - where new documents are simply added and old ones archived or deleted as time goes by. I use both approaches and still tend to rely more on the latter practice, as many do.

One colleague who is a frequent source of inspiration for new workflow approaches often mentions that her projects and support resources number in the hundreds, so many ideas I have for managing my own more limited set these days are not practical for her work. But recently someone mentioned casually that it was going to be difficult to update segmentation rules in the 1500 or so projects that her team maintains to support in-house translation needs in their firm. Oh, my God. Yes, that would take some time following the usual approach of going to Project home > Settings and selecting a new resource in even 10% of that number of projects.

There is a better way. In fact, this way will work with the desktop editions typically used by individuals as well as with memoQ server installations of any size, and a "mass update" of project light resources can be performed in very little time - less than it usually takes me to finish a cup of coffee in the morning. My recent article on memoQ light resource defaults and how to change them essentially points the way, but more details, now tested with a memoQ server as well, are given here.

Light resources in a desktop edition are typically stored in the paths
C:\ProgramData\MemoQ\Resources\Defaults for default resources and
C:\ProgramData\MemoQ\Resources\Local for customized (user-created) resources, unless that path was changed (as many do if they deal with a lot of files with long names and need to shorten paths to avoid errors when the file and path names together approach the 256 character limit imposed by Microsoft Windows).

Server installations follow more or less the same logic:
C:\ProgramData\MemoQ Server\Resources\Defaults for default resources and
C:\ProgramData\MemoQ Server\Resources\Local for custom ones, unless changed as noted above.

Note that the ProgramData folder is a hidden folder by default in Windows, so you may need to change your folder settings to view it.


Light resources stored in both the Defaults and custom (Local) folders are saved without the MemoQResource header one sees in an exported light resource file. Compare the following two screenshots of the same resource in the external editing and maintenance file (with XML comments to help me keep track of what things mean) and the stored file after importing it into memoQ:

My master resource file for German segmentation, maintained with comments in Notepad++
Imported custom resource file for German segmentation. All comments are stripped by memoQ.

So, what should you do if you have 200 projects for your personal work with a memoQ desktop edition or 1500 projects on your memoQ server, and you have a new segmentation rules file, for example, which you want to apply to all of your projects? Simply
  1. Copy all the text beginning with the XML declaration

    all the way down to the end of the file.
  2. Find the default or local resource to update in the paths described above (or in your own custom path) and open the file in a text editor.
  3. Select all the text in the installed resource file, and paste the text of the new resource over it, replacing the content completely.
  4. Save and close the file.
The changes to your default or custom resource will be active immediately. No need to restart memoQ or close and re-open any projects.

In my case, with the German segmentation file given as an example, I would paste that new content into the resource files for generic German (ger), as well as German from Germany (ger-DE), Austria (ger-AT) and Switzerland (ger-CH). No need to mess with the awful integrated resource editors in memoQ, because I keep a master resource file with explanatory comments to help me maintain it better outside of memoQ, and the segmentation I want will be the same for all language variants.

This method does nothing to disturb the content of existing projects nor does it affect their stability in any way. This should work with any memoQ light resources. Thus, for example, an IT department could plan bulk updates even of local resources like keyboard shortcut or web search settings given the necessary access to user drives on a network.

The approach that many follow of deleting old resources and importing new ones with the same names won't work; this can play Hell with project settings, because memoQ notices that a resource used in projects has been removed, and it does not replace the old assignment with a new one, even if that new one has the same name. I played that game with many variations to see if I could trick memoQ into substituting the same-named resource in my project and had no success at all. Don't go there. Use the process I described above.