Several years ago a legal translator in my circles began to use memoQ for her work, and I was asked to help with the migration of data from her old environment. When she was introduced to memoQ LiveDocs, she was delighted to learn that she was able to view the original document text or bitext of concordance hits for content saved in a LiveDocs corpus.
Because her work involved a lot of references to EU directives and other information sources from the EU, the parallel corpora from the DGT had great value to her work. These are enormous bodies of data, totaling several million translation units and growing constantly. Many translators in the EU use this data, but the sheer bulk of it tends to be burdensome to many translation environments, and the lack of context often limits the value of information retrieved from these corpora when stored in translation memories.
So she decided that LiveDocs was the medium in which the DGT data were to be stored, and because the DGT translation memories contain their data in sequential document order, the document context of any concordance hits can be viewed using the context menu in the memoQ concordance:
Thanks to the expansion of file types which can be included in LiveDocs since that time, it is easier than ever to import data from parallel corpora like the EU DGT and use these to support translation work. Using the LiveDocs approach, the extraction of a single large bilingual TMX file from the many zipped data collections is also completely unnecessary (in fact, the extreme quantity of data in those single files inevitably causes memory problems). To build reference corpora for concordancing or the construction of predictive typing resources such as Muses in memoQ, it is simply necessary to unpack the individual zip files into folders full of small TMX files and then import these folder structures into memoQ:
Include only TMX files in the LiveDocs corpus import:
Selecting the desired languages extracts the bilingual data from the individual TMX files, which contain data in all the official EU languages. If a particular file does not contain the desired pairing a corresponding message will be displayed. Don't worry about it.
This approach of loading smaller TMX files into LiveDocs overcomes the memory problems which may occur with gigantic files. And once these smaller files are in a LiveDocs corpus, they can be selected en masse and exported to one or more translation memories.
In fact, this approach is useful to get around the current inability of memoQ translation memories to import more than one TMX file directly at a time. This may be helpful, for example, to OmegaT users who want to migrate their many TMX translation memories (one from each project!) if they start using memoQ.
An exploration of language technologies, translation education, practice and politics, ethical market strategies, workflow optimization, resource reviews, controversies, coffee and other topics of possible interest to the language services community and those who associate with it. Service hours: Thursdays, GMT 09:00 to 13:00.
Showing posts with label EU TMs. Show all posts
Showing posts with label EU TMs. Show all posts
Aug 19, 2015
Jan 8, 2014
Multiple, separate concordances with memoQ
In the comments of my recent post on the memoQ TM search tool, I mentioned a possibility for using that feature to "de-junk" and simplify concordance searches.
In the example above, for example, I am searching the 2 million translation unit EU DGT TM using text selected in a memoQ 6.2 project. Working this way offers me the following advantages:
I remember an argument with a translation agency owner about a year and half ago. The man told me quite insistently about his intent to force even translators with memoQ to use the web translation interface so that he could restrict them to the use of the client-specific TMs he maintained. With the use of the TM search tool, a reasonable compromise is achieved for TM data at least. (LiveDocs and termbase access remains a bit more cumbersome, however, though by setting up a dummy project with termbases, corpora and particular TMs attached, one could actually use three separate concordance sets. That could be interesting.)
In any case, the possibility of a separate concordance for handling large data volumes separately from one's main TMs and the possibility of doing this even while using older versions of memoQ may be a reason why those who do not yet want to do their routine work in the latest version or cannot do so can still benefit from upgrading now and installing the latest version alongside the old version(s).
In the example above, for example, I am searching the 2 million translation unit EU DGT TM using text selected in a memoQ 6.2 project. Working this way offers me the following advantages:
- I can separate the concordances for my project from a big reference dataset I only need for certain lookups.
- A simple copy command (Ctrl+C) automatically looks up text in either language in the TM search tool.
- If I want to avoid any possibility of unintended "leakage" of data from certain TMs in the project, selecting them for use in the memoQ TM search tool ensures that their content will never be "accidentally" inserted as an ordinary TM match as I work.
I remember an argument with a translation agency owner about a year and half ago. The man told me quite insistently about his intent to force even translators with memoQ to use the web translation interface so that he could restrict them to the use of the client-specific TMs he maintained. With the use of the TM search tool, a reasonable compromise is achieved for TM data at least. (LiveDocs and termbase access remains a bit more cumbersome, however, though by setting up a dummy project with termbases, corpora and particular TMs attached, one could actually use three separate concordance sets. That could be interesting.)
In any case, the possibility of a separate concordance for handling large data volumes separately from one's main TMs and the possibility of doing this even while using older versions of memoQ may be a reason why those who do not yet want to do their routine work in the latest version or cannot do so can still benefit from upgrading now and installing the latest version alongside the old version(s).
Jan 4, 2012
EU (DGT) translation memories for DE-EN, DE-NL and NL-EN
As part of some testing efforts, I had occasion to download the gigabyte+ of translation memory data the EU's Directorate General for Translation (DGT) made publicly accessible from the body of EU law. Aside from being massive (hundreds of thousands of translation units for most pairs) and good for stressing a translation environment tool, I find such data useful for looking up official names of directives and organizations. The English (and possibly other language text such as German, Dutch, etc.) is, as some of my British friends might say, very courageous, but if one needs to access the actual text in a target language for an EU law being quoted in the source language, this data source is perhaps a time-saver.
If anyone has a use for the German/English, German/Dutch and Dutch/English language pair TM data from this organization, here are some links to the extracted data I prepared for my tests:
If anyone has a use for the German/English, German/Dutch and Dutch/English language pair TM data from this organization, here are some links to the extracted data I prepared for my tests:
- DGT TM data for German and English (about 530,000 TUs)
- DGT TM data for German and Dutch (about 306,000 TUs)
- DGT TM data for Dutch and English (about 500,000 TUs)
Subscribe to:
Posts (Atom)





