Showing posts with label DGT. Show all posts
Showing posts with label DGT. Show all posts

Oct 29, 2019

Bilingual EU legislation the easy way in #xl8


Translators of European languages based in the EU and many others deal often with citations of EU legislation or need to consult relevant EU legislation for terminology in their translations. One popular source of information for that is the EUR-LEX website, which provides a convenient archive of legislation and related information, with the possibility of multilingual text displays, as seen here:


Some years ago, I published a description of how data from these multilingual EUR-LEX displays can be transferred to translation memories or other corpora for reference purposes, and more recently I produced a video showing this same procedure. But some people don't like the paragraph-level alignment format of the EUR-LEX displays, and these can also occasionally be seriously out of sync for some reason, as in this example (or worse):


Now I don't find that much of a nuisance when I use memoQ LiveDocs, because I can simply view the full bilingual document context and see where the corresponding information really is (kind of like leaving alignments in memoQ uncorrected until you actually find a use for the data and determine that the effort is worthwhile), but if you plan to feed that aligned data to a translation memory, it's a bit of a disaster. And many people prefer data aligned at the sentence level anyway.

Well, there is a simple way to get the EU legislation texts you want, aligned at the sentence level, with the individual bitexts ready to import into a translation memory, LiveDocs corpus or other reference tool. See that document number above with the large red arrow pointing to it? That's where you start....

Did you know that much of the information available in EUR-LEX is also available in the publicly available DGT translation memories? These are sentence-level alignments. But most people go about using this data in a rather klutzy and unhelpful way. The "big data" craze some years ago had a lot of people trying to load this information into translation memories and other places, usually with miserable results. These include:

  • the inability to load such enormous data quantities in a CAT tool's TM without having far more computer RAM than most translators ever think they'll need;
  • very slow imports, some apparently proceeding on a geological time scale; 
  • data overload - so many concordance hits that users simply can't find the focused information they need; and
  • system performance degradation, with extremely sluggish responses in a wide variety of tasks.
Bulk data is for monkeys and those who haven't evolved professionally much beyond that stage. Precision data selection makes more sense, and enables better use of the resources available. But how can you achieve that precision? If I want the full bilingual text of EU Regulation No. 575/2013 in some language pair, for example, with sentence-level alignment, how can I find that quickly in the vast swamp of DGT data?

Years ago, I published an article describing how it is better to load the individual TMX files found in the downloadable ZIP archives from the DGT into LiveDocs so that the full document context can be seen from the concordance searches. What I didn't mention in that article is that the names of those individual TMX files correspond to the document numbers in EUR-LEX

Armed with that knowledge, you can be very selective in what data and how much you load from the DGT collection. For example, if you organize the data releases in folders by year...


... and simply unpack the ZIP files in each year's folder...


... each folder will contain TMX files...


... the names of which correspond to the document number found in EUR-LEX. So a quick search in Windows Explorer or by other means can locate the exact document you want as a TMX file ready to import into your CAT tool:


These TMX files typically contain 24 EU languages now, but most CAT tools will filter just the language pair you want. So the same file can usually give you Polish+French, German+English, Portuguese+Greek or whatever combination you need among the languages present.

I still prefer to import my TMX data into a LiveDocs corpus in memoQ, and there I can use the feature to import a folder structure, and in the import dialog, I simply write the name of the file I want, and all other files (thousands of them) are promptly excluded:


After I enter the file name in the Include files field, I click the Update button to refresh the view and confirm that only the file I want has been selected. Depending on where in memoQ you do the import, you may have to specify the languages (Resource Console) to extract or not (in a project, where the languages are already set). Of course, the data can also be imported to a translation memory in memoQ, but that is an inferior option, because then it is not possible to read the reference document in a bilingual view as you can in a LiveDocs corpus; only isolated segments can be viewed in the Concordance or Translation results pane.

How you work with these data and with what tools is up to you, but this procedure will provide you with a number of options for better data selection and improved access to the reference data you may need for EU legislation without getting stuck in the morass of millions of translation units in a performance-killing megabomb TM.

Oct 25, 2015

European Commission Workshop - Contracts for translation services


What the Linguistic Sausage Producers don't want you to know:
Did you know that tenders for work with the European Commission are not just for the big Wortwurstläden but can be submitted by individual translators who are EU citizens - and that these individuals have equal standing before the Directorate General for Translation? The DGT does not differentiate and many of its best external contractors are individuals, either self-employed persons or dynamic teams of two or three professionals.

The DGT uses taxpayers’ money and must be transparent, with fair and equal treatment for each candidate. Reading their specifications may appear daunting at first, but taking a closer look is worthwhile! Questions may be submitted and are answered during the weeks when the call for tender is open; this can be done in three languages, almost in real time, with all questions and replies made public on the DGT web site.

Quality pays and they will pay for quality: decisions are based on a quality/price ratio of 70/30, in favor of quality. For each job done, a quality note with feedback is sent to facilitate ongoing improvement.

But to get this far, you must first submit a persuasive offer to the selection board.

On November 28, 2015 from noon to 4 pm, IAPTI's UK chapter is hosting a workshop in Manchester (UK) to inform you of what it takes to tender and win at Europe's highest public level for translation. Profit from this important business event at yet another iconic venue! Registration information is available here.

The beautiful Manchester Central Library, venue for the EC tender workshop!

*******

The speaker: Monica Garcia-Soriano started her EU career as a lawyer linguist 24 years ago at the Court of Justice in Luxembourg. She later joined the Spanish Translation Unit at the European Commission in Brussels and for the last 8 years she has been in charge of procurement at the Commission's External Translation Unit.

Aug 19, 2015

Doing the deed with the DGT

Several years ago a legal translator in my circles began to use memoQ for her work, and I was asked to help with the migration of data from her old environment. When she was introduced to memoQ LiveDocs, she was delighted to learn that she was able to view the original document text or bitext of concordance hits for content saved in a LiveDocs corpus.

Because her work involved a lot of references to EU directives and other information sources from the EU, the parallel corpora from the DGT had great value to her work. These are enormous bodies of data, totaling several million translation units and growing constantly. Many translators in the EU use this data, but the sheer bulk of it tends to be burdensome to many translation environments, and the lack of context often limits the value of information retrieved from these corpora when stored in translation memories.

So she decided that LiveDocs was the medium in which the DGT data were to be stored, and because the DGT translation memories contain their data in sequential document order, the document context of any concordance hits can be viewed using the context menu in the memoQ concordance:

Thanks to the expansion of file types which can be included in LiveDocs since that time, it is easier than ever to import data from parallel corpora like the EU DGT and use these to support translation work. Using the LiveDocs approach, the extraction of a single large bilingual TMX file from the many zipped data collections is also completely unnecessary (in fact, the extreme quantity of data in those single files inevitably causes memory problems). To build reference corpora for concordancing or the construction of predictive typing resources such as Muses in memoQ, it is simply necessary to unpack the individual zip files into folders full of small TMX files and then import these folder structures into memoQ:



Include only TMX files in the LiveDocs corpus import:


Selecting the desired languages extracts the bilingual data from the individual TMX files, which contain data in all the official EU languages. If a particular file does not contain the desired pairing a corresponding message will be displayed. Don't worry about it.


This approach of loading smaller TMX files into LiveDocs overcomes the memory problems which may occur with gigantic files. And once these smaller files are in a LiveDocs corpus, they can be selected en masse and exported to one or more translation memories.

In fact, this approach is useful to get around the current inability of memoQ translation memories to import more than one TMX file directly at a time. This may be helpful, for example, to OmegaT users who want to migrate their many TMX translation memories (one from each project!) if they start using memoQ.

Jan 7, 2012

Translation tool concordances compared

A recent experience when tutoring a new memoQ user started me thinking about the way concordance searches work in various translation environment tools and how the results are displayed. The user, who was quite experienced with OmegaT, kept telling me that memoQ could not find examples of a term's use in the TM and she had to do all her searches in OmegaT. I was somewhat puzzled by that, and when I looked at her screen with the memoQ concordance dialog, I saw something like this:

The memoQ version 5 concordance dialog
Looks like the term ("Inverkehrbringen") was found. So what was the problem? For years she had looked at this concordance view:

The OmegaT concordance dialog
The differences in layout and the lack of highlighting of the key term (which was aligned in the center of the memoQ concordance window in the ancient KWIC display tradition) were unexpected and confusing to the new user.

This inspired me to have a look at how various other tools display concordance results. I was not very happy with some of what I discovered, especially with some of today's leading commercial tools. I took a look at the TWB translation memories in SDL Trados 2007, concordancing in SDL Trados Studio 2009, Wordfast Pro (very limited test due to a demo license and my inability to load my TMX test data), memoQ and OmegaT.

In terms of overall performance, the best results were obtained with OmegaT and "Trados Classic" (2007). Searching a huge TM gave results in a flash. Concordance searches with SDL Trados Studio 2009, on the other hand, really sucked with a big TM (EU data, about 400,000 TUs). I vacuumed my entire apartment and fed the dog while I waited for the result, and I wasn't even told how many hits were found. Unfortunately, my favorite working environment, memoQ, performed worst with the same big data set: it simply gave an error message. Further testing revealed that this error was due to the very large number of hits. (This would have been obvious had I paid enough attention to read the dialog title in the first place.)

memoQ error message from too many concordance hits
So it looks like some development attention may need to be directed here. (Update: Kilgray's develops are actively working to remove this restriction.) Of all the tools I was able to test with a large concordance, memoQ was the only one to fail this way. My personal TM with about 10 years of my work in it is nearly as long as my German/English EU legal test database, but concordance searches in it using memoQ are not unduly slow.

Other concordance views looked like this:

The concordance in SDL Trados 2007 - hits limited compared to OmegaT (see above)
SDL Trados 2009 - perhaps the easiest to read, but slower than molasses
Wordfast Pro - format not bad, but the test was limited due to the demo license
The Déjá Vu X concordance hasn't changed significantly in appearance in the latest version (DVX2). Once again, Victor Dewsbery was kind enough to provide me with screenshots of the two "scan" options for searching the translation memory. The initial scan produces only fairly close matches, while the "power scan" is more like the usual concordance with the term embedded in a larger body of text (the non-matching parts being crossed out)

DVX2 scan (first click)



DVX2 Power Scan (second click)

I do have a license for the older version of DVX, but I didn't attempt any stress testing. While its performance with large TMs has always been good (my personal "Big Mama" is about 330,000 TUs), import and export of such data volumes are painfully slow. We're talking overnight. I hope the new version is better in that respect. There I must really give kudos to the OmegaT developer: loading the TM was even faster than with Trados Workbench, which for me has always been a benchmark of speed to aspire to. All you have to do to add a TMX file to the TM of an OmegaT project is to drop it in the "TM" folder of the project. Very nice :-)

I also received a screenshot of a search in Transit NXT from colleague Hans Lenting in the Netherlands. He searched the term "Inverkehrbringen" in the German/Dutch EU dataset from the DGT:

STAR Transit NXT concordance search
As you can see, there are many ways to display data from a concordance search. Which do you find easiest to deal with? Personally, I love the insertion features of the memoQ concordance, but for readability I think some of the other tools are better. And I do like to know how many results I can expect from my data, and I might even want to view them all.

Jan 4, 2012

EU (DGT) translation memories for DE-EN, DE-NL and NL-EN

As part of some testing efforts, I had occasion to download the gigabyte+ of translation memory data the EU's Directorate General for Translation (DGT) made publicly accessible from the body of EU law. Aside from being massive (hundreds of thousands of translation units for most pairs) and good for stressing a translation environment tool, I find such data useful for looking up official names of directives and organizations. The English (and possibly other language text such as German, Dutch, etc.) is, as some of my British friends might say, very courageous, but if one needs to access the actual text in a target language for an EU law being quoted in the source language, this data source is perhaps a time-saver.
If anyone has a use for the German/English, German/Dutch and Dutch/English language pair TM data from this organization, here are some links to the extracted data I prepared for my tests:
And in conclusion, to honor the EU's commitment to quality in what they call English, I offer the following Frank Zappa classic: