Showing posts with label alignment. Show all posts
Showing posts with label alignment. Show all posts

Jun 22, 2017

Translation alignment: the memoQ advantage

The basic features for aligning translated texts in memoQ are straightforward and can be learned easily from Kilgray documentation such as the guides or memoQ Help. However, there are three aspects of alignment in memoQ which I think are worth particular attention and which distinguish it in important ways from alignments performed with other translation environment tools or aligners.

Aligning the content of two folders with source and target documents; automatic pairing by name
The first is memoQ’s ability to determine alignment pairs automatically based on the similarity of names. This has the advantage, for example, that large numbers of files can be aligned automatically, with the source and target documents matched based upon the filenames. This can be done with individual files or with entire folders with perhaps hundreds of files. Thus if source files are contained in one folder and the translated files in the target language are in a different folder and the source and target file names are similar, the alignment process for a great number of files can be set up and run in a matter of minutes. Note in the example screenshot above that different file types may be aligned with each other.

The second important difference with alignment in memoQ is that it is really not necessary to feed the aligned content to a translation memory. memoQ LiveDocs alignments essentially function as a translation memory in the LiveDocs corpus, with one important difference: by right-clicking matches in the translation results pane or a concordance hit list, the aligned document can be opened directly and the full context of the content match can be read. A match or concordance hit found in a traditional translation memory is an isolated segment, divorced from its original context, which can be critical to understanding that translated segment. LiveDocs overcomes this problem.

A third advantage of alignment in memoQ is that, unlike environments in which aligned content can only be used after it is fed to a translation memory, a great deal of time can be saved by not “improving” the alignment unless its content has been determined to be relevant to a new source text for translation. If an analysis shows that there are significant matches to be found in a crude/bulk alignment, the specific relevant alignments can be determined and the contents of these finalized while leaving irrelevant aligned documents in an unimproved state. Should these unimproved alignments in fact contain relevant vocabulary for concordance searches, and if a concordance hit from them appears to be misaligned, opening the document via the context menu usually reveals the desired target text in a nearby segment.

Concordance lookup in memoQ with direct access to an aligned document in a LiveDocs corpus

Sep 15, 2015

A quick trip to LiveDocs for EUR-Lex bilingual texts

Quite a number of friends and respected colleagues use EUR-Lex as a reference source for EU legislation. Being generally sensible people, some of them have backed away from the overfull slopbucket of bulk DGT data and built more selective corpora of the legislation which they actually need for their work.

However, the issue of how to get the data into a usable form with a minimum of effort has caused no little trouble at times. The various texts can be copied out or downloaded in the languages of interest and aligned, but depending on the quality of the alignment tool, the results are often unsatisfactory. I've been told that AlignFactory does a better job than most, but then the question of how best to deal with the HTML bitexts from AlignFactory remains.

memoQ LiveDocs is of course rather helpful for quick and sometimes dirty alignment, but if the synchronization of the texts is too many segments off, it is sometimes difficult to find the information one needs even when the (bilingual) document is opened from the context menu in a concordance window.

EUR-Lex offers bi- or tri-lingual views of most documents in a web page. The alignments are often imperfect, but the synchronization is usually off by only one or two segments, so finding the right text in a document's context is not terribly difficult. So these often imperfect alignments are usually quite adequate for use as references in a memoQ LiveDocs corpus. Here is a procedure one might follow to get the EUR-Lex data there.


The bilingual text of a view such as the one above can be selected by dragging the cursor to select the first part of the information, then scrolling to the bottom of the window and Shift+clicking to select all the text in both columns:


Copy this text, then paste it into Excel:


Then import the Excel file as a file for "translation" in a memoQ project with the right language settings. Because of quirks with data access in LiveDocs if the target language variants are specified and possibly not matched, I have created a "data conversion project" with generic language settings (DE + EN in my case as opposed to my usual DE-DE + EN-US project settings) to ensure that data stored in LiveDocs will be accessed without trouble from any project. (This irritating issue of language variants in LiveDocs was introduced a few version ago by Kilgray in an attempt to placate some large agencies, but it has caused enormous headaches for professional translators who work with multiple sublanguage settings. We hope that urgent attention will be given to this problem soon, and until then, keep your LiveDocs language data settings generic to ensure trouble-free data access!)


When the Excel file is added to the Translations file list, there are two important changes to make in the import options. First, the filter must be changed from Microsoft Excel to "multilingual delimited text" (which also handles multilingual Excel files!). Second, the filter configuration must be "changed" to specify which data is in the columns of interest.


The screenshot above shows the import settings that were appropriate for the data I copied from EUR-Lex. Your settings will likely differ, but in each case the values need to be checked or set in the fields near the arrows ("Source language" particularly at the top and the three dropdown menus by the second arrow below).


Once the data are imported, some adjustments can be made by splitting or joining segments, but I don't think the effort is generally worth it, because in the cases I have seen, data are not far out of sync if they are mismatched, and the synchronization is usually corrected after a short interval.

In the Translations list of the Project home, the bilingual text can be selected and added to a LiveDocs corpus using the menus or ribbons.


The screenshot below shows the worst location of badly synchronized data in the text I copied here:


This minor dislocation does not pose a significant barrier to finding the information I might need to read and understand when using this judgment as a reference. The document context is available from the context menu in the memoQ Concordance as well as the context menu of the entry appearing in the Translation results pane.

A similar data migration procedure can be implemented for most bilingual tables in HTML files, word processing files or other data sources by copying the data into Excel and using the multilingual delimited text filter.

Nov 5, 2013

Proofreading LiveDocs bilinguals, recycling versions in memoQ

(These tests were performed with memoQ 2013 R2 but should, in principle, work the same in any version of memoQ 6.0 or later.)

I really like memoQ's versioning features, and I use the X-Translate function fairly often when a document I'm translating has been updated or a new version comes sometime later. However, I don't keep documents in my projects forever. I use "container projects" for particular clients or subject domains so that I don't have to keep reattaching the same translation memories, terminologies and LiveDocs corpora and various light resources (non-translatables, autocorrect lists, segmentation rules, etc.) all the time. These can get rather full, so I send my old translations off to a LiveDocs corpus after a while. And then if a new version of a document shows up, well... I'm sort of out of luck if I want to use the X-Translate function with the previous version.

Or so I thought. And then a friend rang me and asked how she can export a LiveDocs alignment she did to an RTF bilingual file to make it more convenient to proofread in Microsoft Word and pass on to one of her partners with tracked changes. With that the answer to both problems was clear.

Select the corpus and the file in it to export:


Click Export and choose a location in which to save the MQXLZ file:


In the Translations window of any memoQ project with the correct source and target language settings, select Import and choose your MQXLZ file:


After the file exported from LiveDocs has been imported as a translation file, it can serve as "version 1" for a new file version to be translated using the Reimport document and X-translate features. It does not matter that the file types are different. A bilingual RTF file can also be exported for external correction and commentary.


Here is an example of an exported bilingual RTF file with changes tracked. The changes do not have to be accepted before the bilingual file is re-imported to update the translation using the Import command.


Here is the updated translation:


Changed are marked with blue arrows. Only text changes were implemented in this case, no format changes such as italic text, because the XLIFF file does not support WYSIWYG text formatting. (MQXLZ is a Kilgray-renamed ZIP-package with XLIFF and a chocolate surprise inside.)

Now I've got a new version of my text to translate in a DOCX file. I use the Reimport document function, answer No to the dialog so I can select the new version at a different location:


I'm curious what the differences from the original text are, so I use the History/reports command in the Translations window to find that out:




Then using Operations > X-Translate in the working window, followed by pretranslation to get the changed "exact matches" (like "Aussehen und Gewicht" above) and the fuzzy matches, I end up with this:


If you make it a point to store your important versions in a LiveDocs corpus, this procedure will allow you to recover your archived texts and re-use them for more controlled, reference-based translation. It would be nice, of course, if some day Kilgray would enable specific LiveDocs files to be used as the basis of a reference translation, perhaps even scanning a corpus or a set of corpora to identify the best-matching document or documents. It would also be nice if bilinguals stored in LiveDocs could be exported to other formats and perhaps even be updated with something like an exported bilingual RTF. However, those bilinguals can simply be imported directly to a LiveDocs corpus as new documents, and any corrections made to a document in the Translations list can be sent back to LiveDocs using the relevant command in the Translations window.

Aug 10, 2013

Translation editing in memoQ with LiveDocs and a QA check

Recently I showed how Dragon Naturally Speaking can be used to dictate translations without a translation environment tool. Quite a number of translators work this way (or simply type in a word processor without a translation tool, because they feel they work faster that way). But what about the advantages lost by not having the integrated terminologies, translation memories, etc. in Trados, memoQ and others similar tools?

Well, that's really not a problem. You can have the best of both approaches. Some people do their first drafts as they like without a CAT tool. The translation can then be imported very quickly to the integrated working environment, for example by a quick memoQ LiveDocs alignment, and then the LiveDocs alignment can be used for pretranslation, with any relevant translation memories or terminologies used to check and correct the draft.

I've created a little demonstration video in which I use my dictated translation on snakes in a quick alignment, pretranslation and editing procedure with a subsequent QA check for key terms. This is just one example of the many ways we can creatively combine tools to get the flexibility we need while taking full advantage of the technologies we want to assure quality.


Time Description
0:14
  Adding the alignment pair to LiveDocs
1:17  Importing the source document to pretranslate with the LiveDocs alignment
1:28  Quick pretranslation
1:40  Examining the pretranslation result
2:07  Ad hoc improvement of the LiveDocs alignment
3:05  Inserting a match from the improved LiveDocs alignment
3:20  Beginning corrections of the translation
4:06  QA check for terminology (check the settings)
4:42  Running the QA check
5:00  Examining the QA check results, correcting terminology errors

Jul 26, 2013

The trouble with voice recognition in translation environment tools....


I had not planned to make a video on voice recognition tools any time soon, but a few remarks by my American colleague Kevin Hendzel well down in the many comments about thepigturd's letter to translators sort of goaded me into it. I thought, "What the heck, I'll just grab some text from Wikipedia, record a bit of the work with Camtasia, and post a quick demo of how easy it is to work with Dragon Naturally Speaking." So I got a text about chickens. And activated the screencast recorder. And then the trouble started.

It really sucked. Working with Dragon in memoQ is usually a fairly painless process, but tonight the dogs were anxious and kept poking me in the ribs, and I never did get the microphone adjusted quite right. Some days, microphone position is everything to my scaly transcriptionist. So I suffered with a lot more editing than usual, as anyone watching the video above will see. I worked in my usual "mixed mode" manner, with both keyboard and voice control. Some colleagues who swear by DNS like to do everything by voice and would probably wipe their backsides in the WC that way as well if they could, but that's way too geeky for me. After watching my copywriting partner fly through some 10,000 words of legal translation - and edit it - in a short working day while I slogged through my 3,000 and finished long after she called it a day, I realized that I could work in the relaxed way she did with thoughtful stares at the screen, muttered bursts and the occasional keyboard touch.


But today was a bad day with the Dragon. I might have gone a bit faster with the text. After all, chickens aren't rocket science or even chemistry, with its tag-ridden notation. I could have just dictated in a word processor and everything would have one faster. And if I really want a TM or want to check the terminology, alignment is fast and also a good environment for editing my first draft. I know a number of translators who work that way now. Even with a dictaphone.

In his comments on the other post, Kevin Hendzel expressed a similar feeling to mine when translating with voice recognition: greater engagement and concentration on the text and its structure and meaning. But these tools are not without risk: any errors will in fact pass muster with a spelling checker, so proofreading workflows may have to be very different to be effective. I have noticed this myself - reading my text soon after I have translated it, I am very likely to overlook a missing or switched article or a homophone. Perhaps dictating into a word processor or - since I often look to the glossary hits and other hints on the right of my working window - exporting my text and re-aligning it in the CAT tool after an external rewrite may force my eyes to see things a little differently. In the two years that I have been making serious use of voice recognition I have not yet found the "perfect" workflow.

There are a lot of ways I can tease better results out of this work. But even on a bad day like today, things aren't all that awful. In fact, those familiar with some of the more honest estimates of output in optimized machine translation and post-editing scenarios will realize that today's lousy results (see the end of the video), maintained over the course of a working day, meet or beat the expectations for post-editing in a highly optimized scenario. Without the brain rot typically caused by PEMT! Now that's an advantage. Why don't we stop wasting time with machine translation and instead increase output by more research into the best ways of using voice recognition technology? Ah, but voice recognition is not yet optimized for every language! Ha ha ha... like MT is or ever will be. The millions that get flushed down the toilet with machine translation could and should buy a lot of improvement with voice recognition.

The real trouble with voice recognition is that you may not want your competition to use it. With or without CAT tools. Unlike machine translation.