Showing posts with label translation memory. Show all posts
Showing posts with label translation memory. Show all posts

Dec 28, 2018

SDLTM to TMX conversion update (without SDL Trados Studio)

From time to time I hear from colleagues or read social media posts from people who have been given particular SDL reference resources to use (like an SDLTM file for translation memory or an SDLTB termbase file). But without a license to the corresponding SDL software, it can be troubling to deal with these formats and convert them into files which are easily imported into other environments, such as memoQ, WordFast, OmegaT or whatever.

There are a number of solutions to this problem published on YouTube, and memoQ Translation Technologies Ltd. (mQtech, the solution artists formerly known as "Kilgray") even has a nice help page in their knowledgebase, but all the solutions I have seen so far are either a bit dated, or they omit some information which might help to avoid problems.

The best free solution I have seen is the one described on the mQtech knowledgebase page, but the current description still leaves a few potential stumbling blocks.

There are two pieces to the solution:

Sep 21, 2018

Free TM source file data information utility

Just yesterday I was chatting with an Egyptian colleague about an interesting conference to be held in Cairo next April, and he told me how his wife sometimes gets annoyed with him because he gives away so much information. (I am a big beneficiary of his generosity, and some of the best improvements in recent presentations I've given are techniques I have taken directly from him.)

I've been criticized the same way for most of my life, but I've usually found that information shared freely in the right spirit can often feed more people than a bit of bread and some fish, and the occasional dividends that come back are often delightful surprises.

So it was today. I received a nice e-mail from a reader of this blog, who wanted to share a custom tool for which he had commissioned the development to solve a particular troublesome challenge. His letter is posted below along with a download link for this tool and an explanation of what it does. I hope that some will derive unexpected benefits from this.

*******

Hi Kevin,

I've been using memoQ for a year, and some of your posts on Translation Tribulations have helped me do things and solve problems with memoQ that I wouldn't have been able to solve otherwise. So I want to give back to you and all your readers.

I commissioned from the great Stanislav Ohkvat, the author of TransTools, a program to automatically extract the names of all documents contained in a TM. Add ".exe" to the end of the link below to download.

http://stasokhvat.s3.amazonaws.com/MemoqTmxUtilities

My particular use case is that my colleague reviews my work and sends it back to me for adding to the Master TM, while also adding it to his own Master TM. He also sends me all documents he translates himself, and I review them, adding them to my own TM. However, I recently noticed our Master TMs differed by around 7k segments, meaning we forgot to share a few documents between us.

Rather than tediously sifting through tens of thousands of segments and manually copying the document names, the script does it for us.

I give you full permission to post it on your blog as you see fit.

Cheers,

Érico Carvalho
Pharmacist and translator-subtitler for BNN Medical Translations
Working languages: Brazilian Portuguese to English & vice-versa, Spanish to English, Spanish to Brazilian Portuguese
Specializes in: Clinical Protocols, Informed Consent Forms, Investigator's Brochures, Video Subtitling

Jun 6, 2017

Build your own online reference TM for a team or anyone!


In the past, I have published several articles describing the use of free Google Sheets as a means of providing searchable glossaries on the Internet. This concept has continued to evolve, with current efforts focused on the use of forms and Google's spreadsheet service API to provide even more free, useful functionality.

On a number of occasions I have also mentioned that the same approaches can be used for translation memories to be shared with people having different translation environments, including those working with no CAT tools at all. However, the path to get there with a TM might not be obvious to everyone, and the effort of finding good tools to handle the necessary data conversions can be frustrating.

I've put up a demonstration TM in Portuguese and English here: https://goo.gl/LXXgmf

Here is a selection from the same data collection, selecting for matches of the Portuguese word 'cachorro':  https://goo.gl/9KJils
This uses the same parameterized URL search technique described in my article on searchable glossaries.

A translation memory in a Google Sheet has a few advantages:
  • It can be made accessible to anyone or to a selected group (using Google's permission scheme)
  • It can be downloaded in many formats for adding to a TM or other reference source on a local computer
  • Hits can also be read in context if the TM content is in the order it occurs in the translated documents. This is an advantage currently offered in commercial translation environment tools only by memoQ LiveDocs corpora.
Web search tools of many kinds can be configured easily to find data in these online Google Sheet "translation memories" - SDL Trados Studio, OmegaT and memoQ are among those tools with such facilities integrated, and IntelliWebSearch can bridge the gap for any environment that lacks such a thing.

But... how do you go from a translation memory in a CAT tool to the same content in a Google Sheet? This can be confusing, because many tools do not offer an option to export a TM to a spreadsheet or delimited text file. Some suggestions are found in an old PrAdZ thread, but I found a more satisfactory way of dealing with the problem.

A few years ago, the Heartsome Translation Studio went free and Open Source. It contains some excellent conversion tools. I downloaded a copy of the Heartsome TMX Editor (the available installers for Windows, Mac and Linux are here) and used it to convert my TMX file.




The result was then uploaded to a public directory on my personal Google Drive, and the URL was noted for building queries. Fairly straightforward.

The Heartsome TMX Editor seems like it might be a useful tool to replace Olifant as my TMX editor. While the TM editor in my tool of choice (memoQ) has improved in recent years, it still does not do many things I require, and some of this functionality is available in Heartsome.

Jun 15, 2016

Better ways to search the Internet while translating

When Kilgray introduced memoQ Web Search a few years ago, I was unimpressed, because I was fairly efficient at working with the several tabs of my favorite sites to search for information during translation projects and I couldn't imagine much value to be had from an "integrated" search in a stripped-down custom browser. And the buggy example templates shipped with the memoQ release (several search setups are incorrect) didn't help much. It was only when I began to take a careful, systematic look at this feature to document it for my memoQuickies guide to configuration that I realized how straightforward it really was, and since then, despite ongoing bugs in the feature which sometimes lead to crashes, it has become one of the most important practical features of the product for me. Being able to select a text in the source or target and hit a hot key to search multiple sites at once really does save me time. Lots of it.

There has, of course, been another product around which does that too, which works in more or less any Windows environment and which is far more configurable. But when I took my first look at IntelliWebSearch (IWS) years ago, I was put off and confused by the nerdishness of the presentation, and I just wasn't ready to be told how many damned options I had when I was trying to get my head around the simplest basics. Recently, however, I have beome increasingly irritated at little things with memoQ Web Search, including the impossibility of adding my user credentials to turn off ads on some sites I use, and I began to wonder if IWS might not be worth another look. And indeed it is.

 Go to the IWS web site and learn more!

I am a big fan of multiple concordance search sets in integrated translation environments- this isn't a feature in any of them as far as I know, but it is accomplished easily enough. In memoQ I use the integrated concordance feature to search all the translation memories and corpora attached to a project for my "primary" concordance search set. The memoQ TM Search Tool is configured to search another set of TMs, including some with languages other than those in my project. And for blockbuster TM concordance searches with massive resources like the DGT data sets I have TMLookup. That is three differently configured concordance searches available at a keystroke in my working environment, and you can do the same thing in your favorite CAT tool as well.

So why not try this with web searches? Searching too many tabs at once tends to be slow, which is why I generally recommend no more than a few favorites be used with the integrated memoQ feature. IWS, unlike my usual tool, offers the possibility of configuring multiple "search groups", all of which can be accessed from anywhere with a hot key combination you assign. So I tried this for a few special sites that I usually don't want in my memoQ Web Search but which are a nuisance to deal with manually when I need them. It took me just a few minutes to install IWS, and with the help of a couple of short tutorial videos on the tool's web site, I had my custom searches set up(the configuration wizard to do this is dead easy and user friendly), and in about 10 minutes I was happily invoking special web searches on multiple tabs of my default browser (where I have configured some sites to shut off the damned ads) while I worked in the memoQ translation and editing grid.

So IntelliWebSearch today isn't nearly as difficult to figure out as it seemed years ago. Whether the product has improved or my head is just a little less cluttered now I'm not sure. But it's a very useful tool and a good extension of my working environment which I can recommend with more confidence. I can do useful things without drowning in the depth of its product features.

Nonetheless, I wouldn't mind a guided tour from someone with more of a clue than I have. And later this summer I have exactly such an opportunity. Colleague Michael Farrell, the Italian to English translator who created IntelliWebSearch for all of us, is giving two webinars sponsored by IAPTI, which offer a thorough grounding in the basics of this productivity booster. Information with links (click the images!) is below. I'll be there and hope you'll join me and learn useful things to help in your research of monolingual and multiplingual information sources on the Internet.

 Registration for the webinar

 Registration for the second webinar

Jun 10, 2016

memoQ 2015 - big improvements, ongoing issues


The pace of life is slow in the Algarve, but memoQ users would generally agree that big TMX imports to memoQ are a lot slower. Until now.

While on holiday I got a note from an occasional client, asking if I could take on the translation of a little Interessenabgleich for the next day. Not a problem I thought, though I had not yet imported the backup of my Big Mama TM with legal content onto the pokey Portuguese laptop I had dragged with me. With only 4 GB of RAM and the megatrojan known as Windows 10 it isn't much for crunching big data, but it is satisfyingly slow for demos and teaching classes, giving me plenty of time for questions and explanations while things load or refresh.

Nearly three hours later, the import was finished. So was I by that time; it was much too late to think of work any more. And the next morning, as I furiously worked to meet the deadline, a new update for memoQ appeared, which I was too muddle-headed to put off until I had delivered. After updating, not only was my license no longer recognized, I could not even boot the program to re-enter it! Welcome to Build 152! On another machine updated remotely, the update worked, but there was a funny conflict with Kilgray's standard template designations, causing the templates to be re-named.

The problems were sorted out by re-installing the build with a downloaded installation file from memoQ.com, and the job made it out on time. More glitches with term base editing and addition showed up later that day as I worked on another text. Grr.

Fortunately, Kilgray responds quickly to most issues, and the automated bug reporting features recently introduced are probably a great help in pinpointing troubles which users are too often reluctant to complain about. The next morning a patched version - Build 153 - was released. So far, so good with it.

The latest build introduces some interesting new features: among them are content extraction of more Microsoft Office object types, including equations (I've waited years for this!!!), for translation, faster TM imports, word count weighting, better tag handling, improved QA checks and some other things. Looking at all this will take some time, but it will be time well spent for my work purposes.

I was intrigued by the statement that "importing a TM with really many segments has just gotten up to 70 times faster", so I decided to test that with the file that bedeviled my evening recently. The problem with marketing hype is that it emphasizes exceptional cases and too seldom reflects what a typical work experience might be. And here too this is the case. The import of approximately 330,000 segments was completed in just under half an hour, about six times faster. Perhaps on my souped-up, RAM-rich machine in the office the difference might be more dramatic, but that is still a great improvement which will make me less hesitant to do maintenance on my large data sets.

memoQ 2015 has had a rockier road than any other version since the infamous memoQ Version 6 for which the code base was rewritten; a bit over a year since its release it still surprises me with inconvenient errors at inconvenient times, in contrast to my usual expectation of stability and reliability about four months after a major release. Actually, memoQ 2015 seemed to get to that point early in my working experience until about five months after its release, when in late autumn last year things went to Hell with SDL compatibility in many projects. Interesting new features, such as regex filters for the working grid and find/replace functions, continued to be added, but it all had a bit of a bleeding edge feel.

Now that memoQ has survived the awful transition to the me-too ribbon interface, it really is time to focus on stability and process reliability. The tool has become a major part of major corporate translation operations as well as the daily work of countless independent translators and we need stability to support our business more than we need new technotunes to whistle. I appreciate many of the features introduced in the past year very much, and I use a lot of them (the new keyboard shortcuts are my favorite for managing writes to termbases properly when I work), but I might trade it all if I can just be sure again that the software won't crash and burn me just before a translation is to be sent to a client.

I am confident that these matters will be addressed, but the matter of when is of critical importance for many. Issues of novelty versus reliability are hardly new in the software world; I have seen every variation of this for about 40 years now, and Kilgray generally, though not always, earns better marks than the competition. And even with the troubles that have beset memoQ 2015 (which have made it very desirable to keep older versions installed too) it is still the best - and most reliable - tool I have seen for the ways I work.

Aug 19, 2015

Doing the deed with the DGT

Several years ago a legal translator in my circles began to use memoQ for her work, and I was asked to help with the migration of data from her old environment. When she was introduced to memoQ LiveDocs, she was delighted to learn that she was able to view the original document text or bitext of concordance hits for content saved in a LiveDocs corpus.

Because her work involved a lot of references to EU directives and other information sources from the EU, the parallel corpora from the DGT had great value to her work. These are enormous bodies of data, totaling several million translation units and growing constantly. Many translators in the EU use this data, but the sheer bulk of it tends to be burdensome to many translation environments, and the lack of context often limits the value of information retrieved from these corpora when stored in translation memories.

So she decided that LiveDocs was the medium in which the DGT data were to be stored, and because the DGT translation memories contain their data in sequential document order, the document context of any concordance hits can be viewed using the context menu in the memoQ concordance:

Thanks to the expansion of file types which can be included in LiveDocs since that time, it is easier than ever to import data from parallel corpora like the EU DGT and use these to support translation work. Using the LiveDocs approach, the extraction of a single large bilingual TMX file from the many zipped data collections is also completely unnecessary (in fact, the extreme quantity of data in those single files inevitably causes memory problems). To build reference corpora for concordancing or the construction of predictive typing resources such as Muses in memoQ, it is simply necessary to unpack the individual zip files into folders full of small TMX files and then import these folder structures into memoQ:



Include only TMX files in the LiveDocs corpus import:


Selecting the desired languages extracts the bilingual data from the individual TMX files, which contain data in all the official EU languages. If a particular file does not contain the desired pairing a corresponding message will be displayed. Don't worry about it.


This approach of loading smaller TMX files into LiveDocs overcomes the memory problems which may occur with gigantic files. And once these smaller files are in a LiveDocs corpus, they can be selected en masse and exported to one or more translation memories.

In fact, this approach is useful to get around the current inability of memoQ translation memories to import more than one TMX file directly at a time. This may be helpful, for example, to OmegaT users who want to migrate their many TMX translation memories (one from each project!) if they start using memoQ.

Aug 12, 2015

"Upgrading" translation memories for document context

Translation memories are in a way the Zombie Apocalypse of our profession. Dead data walking, with rotting bits falling off and lying out of context in a concordance for the reader to puzzle over, wondering where that bit of wordflesh one fit in the whole.

How often in my Dark Days as a Trados User did I look at some concordance hit and wonder just how stoned the translator was, corrected the "mistake" and discovered later it was quite right in context? Context matters, but with translation memories it's like the Invisible Hand at best, with a few missing fingers.

I once was lost, but now I found memoQ LiveDocs, and I
  1. export my context-free or -invisible TMs to TMX, then
  2. import the TMX to LiveDocs
where I now can go directly from a concordance hit to the translation unit with LiveDocs, which can be read in document context if that chunk of text was in fact written to the TM in the sequential order of the document, which is often the case.




Jan 15, 2014

A ferramenta de pesquisa, Memória de Tradução, do memoQ

Esta funcionalidade permite usar as TM do memoQ para procurar texto numa outra janela, por exemplo, num documento do Microsoft Word ou de uma outra ferramenta de Tradução. Também é possível utilizar esta funcionalidade em ambientes de tradução que não tenham memórias de tradução, ou cujo acesso a estas esteja restrito.


À medida que copia o texto da sua janela de trabalho, o conteúdo que está na sua área de transferência é automaticamente transferido para a janela de pesquisa e as combinações aparecem.

Ctrl+Shift+Q     Inicia a ferramenta de pesquisa da memoria de tradução do memoQ, que vai, imediatamente, procurar por qualquer texto na área de transferência do Windows.
Ctrl+C     Copia o novo texto para a janela de pesquisa da memória de tradução, e executa uma procura.
Ctrl+Alt+C     Copia, o texto de destino de uma correspondência selecionada, para a área de transferência.
Ctrl+Shift+C     Copia, o texto de origem de uma correspondência selecionada, para a área de transferência.
Ctrl+V     Cola o texto da área de transferência noutra aplicação.

Na ferramenta de pesquisa da TM pode selecionar qualquer uma das suas memórias de tradução para usar noutras aplicações. No menu Configurações de Pesquisa, desta janela, também é possível definir outros parâmetros, tais como, a percentagem mínima para que existam correspondências, as penalidades para com os alinhamentos ou o rigor a ter para com as tags existentes.

Uma memória de tradução selecionada na ferramenta de pesquisa da TM não pode ser usada nem aberta no memoQ, enquanto a ferramenta de pesquisa estiver ativa, nem a ferramenta de pesquisa consegue aceder à TM que está associada a um projecto em aberto no memoQ. Um raio laranja é exibido na lista das memórias de tradução para indicar esta condição. Depois da ferramenta de pesquisa estar fechada, a memória de tradução ficará disponível novamente para ser usada no memoQ. E, quando fechar o projeto, a ou as TM correspondentes ficarão outra vez acessíveis para a ferramenta de pesquisa.

A ferramenta de pesquisa também pode ser útil como uma segunda concordância no memoQ, para procurar um conjunto em particular de memórias de tradução, que não estejam associadas ao projeto em aberto.

*****

Excerpt from memoQ em Pequenos Passos, the Portuguese version of the second edition of my memoQ tips book. Translated by Cátea Caleço Murta.

Jan 2, 2014

The memoQ TM search tool

Release 2 of memoQ 2013 included a new utility which allows memoQ translation memories to be used for lookups, the TM search tool:

When working in other translation environment tools such as SDL Trados Studio or Wordfast, translating text in a word processor or reading PDF files and web pages, selected text can be looked up directly in chosen translation memories and text from the source or target of a translation match can be put in the Clipboard for pasting into the other application. Relevant keyboard shortcuts are:

Ctrl+Shift+Q     Starts the memoQ TM search tool, immediately searches for any text on the Windows Clipboard.
Ctrl+C     Copies new text to the TM search window and executes a search.
Ctrl+Alt+C     Copies the target text of a selected match to the Clipboard.
Ctrl+Shift+C     Copies the source text of a selected match to the Clipboard.
Ctrl+V     Pastes the Clipboard text into another application.

A translation memory selected in the TM search tool cannot be opened or used in memoQ while the search tool is active. An orange lightning bolt is displayed in the TM list of the Search settings to indicate this status. After the search tool is closed, the TM is available again for use in memoQ.

Although the initial version of this tool is quite useful, many users have realized that further refinements of its features would make its application more flexible and effective. Some suggestions so far include
  • selecting/deselecting all TMs
  • filtering TMs by metadata
  • saving and loading profiles (collections of particular TMs and settings)
  • indicating match sources (i.e. TM, preferably with metadata)
A number of other quirks, like the ability to launch multiple instances of the tool, also still need to be sorted out as of Build 52.

I hope that Kilgray will take the further development of this tool seriously and consider how to improve and expand it, perhaps to include remote translation memories as well. The current version of the TM search tool requires a memoQ license on the computer where it is used, but separate licensing could also be quite interesting. This could be useful, for example, in collaborative projects with partners who use different tools and working methods or for those who want to use memoQ translation memories as bilingual concordances. I see the potential for a value-added service here if I can provide such a concordance (for a fee) to an end client, perhaps with some sort of protective encapsulation for the memories provided. Inclusion of termbases and LiveDocs corpora in future versions of the tool could also prove interesting. memoQ could become a reference information packaging platform to create additional communication services for our clients. There are interesting possibilities for mobile applications here as well. But in the meantime I'll settle for the modest improvements in the bullet points above.

Further information on the memoQ search tool can be found in the Kilgray knowledgebase.

Dec 11, 2013

General settings for memoQ TMs

memoQ TM settings are found in the Resource Console, the Options and a project's Settings.
This is a very useful "light resource" which is well worth nearly every user's time.
To define the TM settings to be used in new projects, select a settings configuration under Tools > Options... >  Default resources > TM settings (in the row of icons) by marking its checkbox.

To define the default TM settings to be used in the project you have opened, go to Project home > Settings > TM settings (in the row of icons) and mark the checkbox for the desired project default.

Different settings for individual TMs in a project (for example to set higher or lower match criteria) may be applied by going to Project home > Translation memories, selecting the TM of interest, clicking the Settings command at the right of the window and choosing the settings to apply instead of the project's standard TM settings.

The General settings tab is the same for all currently supported versions of memoQ. Role options are included on another tab in memoQ 2013 R2, and the Project Manager editions of memoQ offer additional possibilities for filtering and/or applying penalties to content on a Filters tab.


Match thresholds
The first value here (minimum) controls the fuzzy percentage below which a match will not be displayed in the translation results pane at the upper right of the working translation window.

The "good match" threshold is relevant to pretranslation (though this is unfortunately not made obvious in the dialog). The default value of 95% is really too high and would only apply to matches with small differences in tags or numbers; since any small difference in words is penalized significantly in memoQ (something I find very helpful, as I can understand more quickly what differences to look for compared to working in Trados). I usually set my "good matches" to 80%.

Not a "good match" according to the memoQ TM default setting
Penalties
In my work, an alignment penalty, which is a deduction from the match rate of a translation unit created by feeding an alignment to a translation memory, does not make a lot of sense. This is because
  • I almost never send alignments to a TM. Why bother? LiveDocs may be slower in pretranslation, but it provides context matching just like a TM, and you can actually read what you find in a concordance search in its original document context. TMs suck because you do not get the full context for your matching segment and are thus at greater risk for missing information which may be important for a translation. This is especially the case with short match segments.
  • if I happen to be aligning a dodgy translation and want to send it to a TM, I'll put it in a "quarantine TM" which already has its own penalty.
  • on those rare occasions when I might feed an alignment to a TM, it's because the content is going to a user of another CAT tool, and if that person uses Trados or another tool that can read XLIFF files or other available bilingual formats, I'll send the data as that instad, so it can be reviewed and modified more easily before feeding to a TM. This also gives the other person a bilingual reference with document context.
  • alignment for TMs is soooooo 1990s!
User penalties: If you have the misfortune to share a TM with someone whose work you do not trust completely and you want to avoid letting that person's 100% and context match segments slip past you unnoticed, apply a suitable penalty for the level of "risk" that person represents. If you want to be sure that user's content never gets used in a pretranslation and never appears in the translation results pane, apply a whopping big penalty like 80%. Those segments not be shown or inserted but will still be there in a concordance search if you want them.

TM penalties: Sometimes a client provides you with a TM you do not trust completely, or you may have a "quarantine TM" with content of dubious quality. Or I might have a TM with good content in British English but need to deliver a translation in American English. Applying penalties to such TMs will reduce the priority of their matches and prevent 100% matches with inappropriate language from slipping past without more careful inspection. As in the case of user penalties, you can also apply a very large penalty to ensure that matches will never be displayed in the translation results pane or used in a pretranslation but still have the TM content available for concordance searches.

Adjustments
It seems to be a good idea generally to enable the adjustment of fuzzy hits and inline tags. In many (but not all) cases, this will correct small differences in numbers, punctuation, cases and inline tags.

The only significant effect I was able to determine in adjusting the inline tag strictness in my tests was that more permissive settings might count a match with different tags as a full match. While this might meet the requirements of some clients hoping to impose discount schemes, from a quality assurance perspective, this does not seem like a good idea, and I believe it is better to have a strict setting here to draw attention to differences and reduce the chance that errors might be overlooked.

Dec 8, 2013

memoQ TM settings: beware the Kilgray defaults!

memoQ 2013 R2 introduced a very significant change in the management of translation memory data which most users are likely not aware of. However, because the default behavior for information storage in translation memories was changed, it is important to be aware of this difference and what to do before your data are unacceptably compromised.


The screenshot above shows several different translations stored in my TM for the sentence in the second segment. In previous versions of memoQ, only one translation would be stored with the way this translation memory was configured. However, in memoQ 2013 R2, the role of the person editing the translation becomes an important part of the "context", and as a result, multiple translations can be stored for different roles. Personally, I find this a rather useless feature, because if I want to know previous translations for a segment, I consult the row history using the context menu. But I understand how in some processes, it may be desirable to maintain a record of translations entered by the translator and the first and second reviewer.

I have no use for these older translations, especially as these may contain errors (as seen in the example of the third entry in the screenshot). If I am proofreading my translation in a "reviewer" role and make changes, I want to overwrite the original entry in my TM and avoid the chance that its errors will be propagated in later work.

To avoid the problems that can result from this redundancy and preservation of errors in the translation memory, as of build 6.8.6 it is necessary for users to explicitly opt out of the current Kilgray TM settings defaults and create their own custom settings.

TM settings are "light resources" which can be managed in four places:
  • The Resource Console,where settings can be created, edited, imported, exported, etc.
  • The Options (Tools > Options... > Default resources > TM settings), where the default for new projects can also be set
  • Project Settings (Project home > Settings > TM settings) in a specific project, where the default settings for the current project can be set
  • Project home > Translation memories > (TM) > Settings where alternative TM settings can be specified for a particular translation memory selected in  project. This would be the case there you want to apply a special set of penalties to the content of that TM, for example.
The last tab of the default TM settings dialog looks like this:


To avoid the trouble of multiple, role-based entries being written to a TM, settings must be created in which the option to Store modifying user's role in the TM entries in not selected, and these custom settings must be applied to the primary translation memory in the project (by default or explicit selection).

Here's the "fast path" for staying out of trouble:
  1. Go to Tools > Options > Default resources > TM settings and if you do not already have custom TM settings to edit, select and clone the default settings. Give them a suitable name like "My Own TM Settings".
  2. Click the Roles tab and unmark the setting to store the user's role in TM entries.
  3. Click OK.
  4. Ensure that the checkbox next to these custom settings is marked so they will be applied to all new projects. Then click OK to exit the options.
  5. In any currently open project to which the desired settings have not been applied, go to Project home > Settings > TM settings and select the desired settings as the default by marking the corresponding checkbox.
Multiple entries written to the TM when the roles are included will not be eliminated after the TM settings are corrected. They must be explicitly removed by editing the translation memory.

I hope that in the future Kilgray will reconsider these troublesome new default settings and make the new possibilities "opt-in" values in custom TM settings. But for now, users must actively change their settings and defaults if they want to avoid role-based additional TM entries. (The current version of the memoQ Help describes roles as being disabled here by default. Would that this were so!)

You can, of course, make other useful adjustments to your custom TM settings, such as defining what a "good" match is (for pre-translation) or adjusting the tag matching behavior or applying various kinds of penalties to reduce match values for content which might have quality problems. The memoQ Help offers guidance on these options.

Postscript:
Even after the settings are "fixed", "existing damage" in a TM caused by the storage of unwanted, role-based information is not repaired. Any messes will have to be cleaned up in the rather inadequate TM editor in memoQ or in an external TMX maintenance tool. At the present time, there is no "easy option" to clean up a large number of erroneous or redundant translations stored because of this role setting. This case unfortunately underscores the woefully inadequate maintenance facilities for translation memory resources in the current version of memoQ. Perhaps some of the sophisticated options developed for Kilgray's TM Repository will finally trickle down in some way in an integrated option with Language Terminal or some sophisticated filtering and editing options will be added directly to the desktop product so that users can finally maintain their TM data in a reasonable way. memoQ is, overall, the best option available to us for project work in most cases, and I recommend it to colleagues because I know they will be able to do most ordinary tasks with a minimum of grief and calls for help (or expressions of anger) directed to me. But in 2013 it is ridiculous that my ability to manage my TM in my tool of choice is inferior to what I could do when I started using Déjà Vu as my CAT tool 13 years ago. Please join me in encouraging Kilgray to raise their game - soon - with respect to translation memory maintenance by writing to support@kilgray.com and expressing your need for better data management! (And more sensible default TM settings, of course.)

Update:
It looks like this default problem may end with the 6.8.6 build. One of the key people involved with memoQ and its features has stated that "After the [next] update, the default TM settings resource will have 'Store modifying user’s name in TM' unticked." Excellent.

For cases where there may already be data problems from older, erroneous entries being retained, the following workaround was suggested:
  • Export to TMX
  • Start up 6.5 and import into an empty TM
In the process, memoQ 2013 (version 6.5) will ignore the role information in the TMX, and entries with the same source will not create duplicates; translations with a later timestamp will be preserved in the TM if there are duplicates in the TMX.

This still doesn't change the fact that we need better means of maintaining our data in memoQ, but it is good that once again, Kilgray has responded quickly to important concerns of its users and is on the way to solving the problem.

Dec 2, 2013

Segmentation in memoQ server projects

Segmentation difficulties are often one of the most troublesome aspects of working with translation environment tools. Learning to configure segmentation rules correctly and applying that knowledge can save many hours of wasted time in alignments and translations and avoid filling translation memory resources with garbage from fractured translations of partial sentences with missing verbs, subjects and whatnot.

The usual alternative remedy for inadequately configured segmentation rules which lack the segmentation exceptions needed for abbreviations, for example, is to use the "join" function (Ctrl+J), and sometimes the split function (to manage very long, unwieldy clauses such as one might find in a patent text, and the join the parts again later).


There are situations where joining and splitting of segments is blocked. This is the case with any file which is part of a view, for example; the view must be deleted before segments can once again be joined or split. Segmentation changes are also not possible in a server project which has not been set up to allow them.

There are several options or documents available to project managers when setting up  memoQ server project. But to enable translators to correct unfavorable segmentation, there is really only one choice:


If Desktop documents (no web translation) is selected, then on the dialog page which follows, changes in segmentation can be enabled:


If a project manager does not configure a project to allow this, for example because a document is being split between multiple translators (which does not allow for segmentation changes for technical reasons), the full responsibility must be assumed by the project manager for any segmentation issues. The imported documents should be examined carefully, and if any problems are observed, the segmentation rules should be modified and the documents re-imported. Doing otherwise may unavoidably result in garbage being written to the project's translation memory.

This is a very important point for memoQ trainers to emphasize when they are teaching users of the memoQ server to set up projects. Segmentation topics should be covered thoroughly, and the potential for bad results should be understood clearly if translators are given badly segmented documents they cannot fix. Project managers should also be encouraged to avoid restricting translators options in ways which are likely to harm the quality of the results and make parts of the translation unfit for later re-use.

A good rule of thumb is to choose the desktop documents option for projects always unless there are very urgent reasons not to do so. In this way, you will avoid upsetting your translators by forcing unmanageable, fractured sentence fragments on them, and you will be assured of better quality translation memory resources.

Oct 28, 2013

Want a revolution? Try memoQ 2013 Release 2.


OK, so I'm exaggerating a bit. And even though the new version of memoQ was officially released today by Kilgray, it really is still beta software. But damned good beta. I expect that there will be more of interest to individual translators added in this version of memoQ than in any other version I've seen up to now. Lots of T's to cross and i's to dot still, but there is great promise, and it's worth having a look now at the future of memoQ.

I'm not talking about changes to the memoQ Server. There are lots of those in this version, and for a change many of them actually seem to be helpful to translators working on the server and less focused on slicing and stuffing linguistic sausage faster like many of the 6.x server features introduced. The rollout webinar with István Lengyel and Florian Sachse of Kilgray showed enough of why memoQ Server users should be pleased. But they could have filled the hour and three quarters with nothing but presentations of new or improved functions for the rest of us and still not run out of material. Since I still have a project to finish tonight, I'll just hit a few of the highlights that I'll probably return to later as the features stabilize and are truly ready for productive work.

Language recognition
memoQ now intelligently recognizes the language(s) of the source text. This is a small convenience in setting up projects perhaps, but for those occasions when a source language has many passages in another language or more than one other language, these other language segments can be identified automatically, copied source to target and locked. I can think of more than a few patent dispute translations where this would have been helpful.

Startup Wizard
A new feature under the Help menu gives a quick, friendly guided tour of important settings that are often overlooked that are hard to find for new users and many experienced ones. This is actually one of my favorite new features and possibly the best help I've seen yet for making a better start with the software.

Better Microsoft Word spelling integration
Custom dictionaries can now be imported from Microsoft Word with greater ease. Users can now also choose Microsoft Word for dynamic marking of possible spelling errors (unknown words). This is a good thing for those of us who hate Hunspell. Oh, and those pesky doubled words are caught now.

More stuff with Microsoft Word...
like exporting tracked changes between translation versions to a DOCX file (sans formatting I think), exporting target comment to a DOCX file (alas! in writing the specification Kilgray failed to consider that one might want to select which comments get exported and possibly suppress all the comments, but I'm told this will be remedied quickly), font substitution in DOCX files (this was a major WTF feature for me, but if I understood correctly, there is some way I can use this to protect text formatted a certain way, such as code in a programming guide - if that's true, this is cool) and...

the TM lookup tool,
an external application which runs in Microsoft Word and any other environment and allows you to look up text copied to the Clipboard in selected memoQ TMs. Too bad they didn't include termbases in this new feature. Yet.

New filters and processes
like direct import of InDesign files with a preview using the free online Language Terminal integration, Adobe InCopy and some file formats that must be pretty damned geeky because I've never heard of them.

Why am I excited about
a plain text view which is about as exciting as lukewarm, unspiced pea soup. Well, because it's absence has been driving me nuts for years now. It's in this version.

Meanwhile, back at the termbase
great things are happening with new import options that are still a wee bit buggy but will get very good very soon. Until now memoQ could only import terms as TMX and delimited text. New options include Excel (at last!), MultiTerm XML and TBX. It was child's play for me to tweak a couple of TermStar MARTIF exports from STAR Transit to import those terms, because TBX is a dialect of MARTIF and STAR's MARTIF is very close to TBX. Extra effort? About 2 minutes of search and replace so I'm hoping Kilgray will go the extra five yards and touch this import option down.

The addition of the MultiTerm XML import option means that memoQ users can now roundtrip data from memoQ to partners using SDL MultiTerm and back for termbase updates. Unfortunately at the moment, the only meta data transferred in the import is the definition field, but efforts are in progress to support at least the MultiTerm fields memoQ exports to XML with Kilgray's own definition. That was simply forgotten at specification time (oops). But still, this will be serious headache relief for those of us who work in teams with SDL Trados users and want to share terminology in the most effective ways.

Is that all?
No. This new version of memoQ is like a very messy Christmas where one can easily lose the overview of hat's under the tree with all the wrapping paper and bits of ribbon cluttering the floor. As it gets cleaned up, we'll all notice a good bit more, and I suspect that Santa's Hungarian and German helpers will be slipping a few more things under the tree that they might forget themselves until some user trips over them. There has been so much effort put into consolidation and improvement of existing features that it's simply too much to keep track of. I've made a list and checked it more than twice and still find things to add. But I'll end with another look at something I've already blogged about, that groundbreaking

Monolingual Project and TM Update
with edited files in any target format. It still has a lot of little quirks, especially with some formats, but here I expect a lot of improvements. I've made a little demonstration video and put it on YouTube; it shows the reimport of edited translations to update the translated file and the TM in memoQ, and it shows two different ways to look at tracked changes before revealing the dark secret of Row History Recovery which I think Kilgray didn't realize was possible. Well, damnit, they should have made it a feature with a button anyway.

 
(View this in full screen mode by clicking the icon at the lower right of the video window.) 

Oh yes, and one more cool little thing about this release that I forgot to mention...

... the quickstart shortcut to creating memoQ projects
in the context menu by right-clicking on a file. I'm not much into single-file projects any more and prefer to use "container" projects for customers or categories instead, but it's still a nice little addition that can save time once in a while:



Oct 24, 2013

The Next Big CAT Feature To Copy?

Thanks to the persistent disbelief of some users, particularly cranky financial and legal translators who don't understand the challenges of programming, that it really is "impossible" to enable a project or TM update based on an edited monolingual target text, Kilgray has decided to just do it and make this feature available to the masses with the next release, memoQ 2013 R2. It's not a perfect solution, but a great start, and when the rest of the specification is implemented later, it will be really, really good.

Here's an example of the first step reimporting a short text in which every sentence except the first was rewritten and rearranged (click the pic to see it full-sized):

memoQ monolingual document import and alignment for translation memory update

This monolingual alignment and the matches it assigned was totally automated - no adjustments by me. When applied to the translation document in memoQ, it doesn't change the order of the translation, but it does make updating the translation memory much easier. I can also use the tracked changes feature for versioning to look at the edits in more detail, and the view can be filtered to show the changes in a large document more clearly to ensure that nothing was missed.

What good is this? Well, so far Kilgray is rather fixated on the idea that one might receive edited documents from a proofreader or a client at some later date and import the changes to the project to update the TM. Maybe, but in many cases, this won't really happen in my workflow. If I finish I project, now I very often send the documents to a LiveDocs corpus, where they make a marvelous pseudo-TM (if a bit slow or slower sometimes) and an excellent context reference for concordance searches. I then delete the documents from the project, because I use projects as "containers" for repeated business, so unless some day the monolingual updates are made available for a document in LiveDocs, I may often not be able to take advantage of it. One could, of course, apply the same principle of monolingual alignment to a translation memory or even a TMX file, and I am sure somebody will do that before long if it doesn't already exist in some flaky academic freeware for supernerds somewhere.

So why am I so excited about this feature? Because I already use it every day. It saves me time and puts a big smile on my face. When I get ready to deliver  translation, the last step for me is to look at it in the source application - Microsoft Word, PowerPoint, etc. There I make last-minute adjustments, change words, combine sentences, split sentences, delete things, etc. A lot of these changes never make it back to the TM, because that maintenance can be a major pain in the backside, especially when there's a lot happening on my desk and in my e-mail inbox. This feature is a big step toward reducing that stress.

Kilgray isn't done with this by any means. Plans for source text merges to allow combined sentences in an edit to be handled easily have not been implemented yet, but I'm told that will follow no later than the next version. I hope so. The current beta version is also a bit dodgy with many formats I've tested so far beside DOCX and TXT, and changes of target text file type have been forbidden in the current version, as have any edits except segmentation adjustment and linking during the monolingual alignment.I plan a lot more testing to understand the limits of the current implementation, and I expect there will be many improvements in this area. But this is an excellent start.

When the specification was developed, Kilgray was unaware that something like this was already available from SDL in one of that company's pricy "regulatory" licenses (I wonder if the idea came from watching the video of the public argument about this option at memoQfest two years ago - I know that remarks about the SDL OpenExchange were followed closely). However, SDL has taken no great advantage of this to offer such a feature to a wider user base yet, so I have no idea how well it looks. But mark my words - in a few years, this innovation will be one of those things that users of any good CAT tool should take for granted!

Aug 24, 2013

Games agencies play, Part 1: "basic information" for Google searches

The other day a fellow freelancer called me up, irate to learn that a translation agency owner in her country had appropriated one of her ideas and literally taken her words from an online interview to represent his services as something they probably are not.

This started quite a discussion about some of the tactics we and others have observed in the trench warfare being waged for business in some segments of the market. It's not often I see egregious cases of plagiarism like the one I was shown two days ago, but there are plenty of sneaky moves and not-so-sneaky ones employed to try to get those little prospect fish in a world wide net online.

So I've decided to offer a little "series" featuring a few of these tactics and how freelancers might respond appropriately to them to promote their own business and support the real interests of translation consumers in an ethical way. And maybe we'll find a few candidates for the Hall of Shame on our tour.

In Twitspace and in my YouTube content research, I have encountered quite a number of pages, video clips and other material describing basic concepts of translation technology or translation processes, such as translation memory, terminology mining and management, etc. Rarely these are quite good presentations of academic value, which are prepared by knowledgeable colleagues or university instructors, and which are so clearly presented that I am pleased to share them with my clients and others when they are needed. More often these are superficial overviews stuffed with keywords, which communicate too little of the real information needed to assess a technology and its relevance to certain types of problems. Some of the self-serving crap I've seen on YouTube by agency representatives pushing their services is so bad that I'm sure their competitors must smile. What are they trying to accomplish?

no translation agency spam
Quite a number of different things most likely. Being noticed first by potential prospects searching for information is certainly high on the list for some. Influencing the discussion of certain technologies and approaches to translation is probably another goal of the more clever content presenters and one which freelance service providers should consider.

I am not going to give links to any of these pages I've found. Today's gem - "Guide to Translation Memory (TM)" - found on a US agency site, is fairly typical. It's superficial, but not particularly toxic, but it really makes no useful contribution to helping clients understand the advantages and limits of translation memories. The description of fuzzy matching is actually wrong in some contexts and by no means describes what a fuzzy match might be in some major translation environment tools.

But as more translators discuss their tools and working methods (which may or may not interest translation buyers and consumers), a lot of jargon gets used without adequate definition or appropriate references to consult for more information. And important concepts related to the limitations and risks of these technologies or their most appropriate uses are also very often missing.

A good response to this on our own freelance business web pages would be to present such information on secondary pages of our own sites or provide links to balanced, reputable information sources that can help our clients deal confidently with the concepts if they have an interest in them and avoid some of the traps. Linking good information pages on association web sites, university information pages or even Wikipedia may also help to raise the ranking of these pages and bury the SEO spammers in agencies or propaganda organizations like TAUS or the CSA in the far-back pages of a search where they belong. If an agency like Lionbridge, TransPerfect, thepigturd and others offers such "infopages", for God's sake don't be stupid enough to link them or tweet them; there are more ethical alternatives to be found for sharing such information.

I'm not calling for a boycott of any page associated with a translation agency. There are plenty of good ones out there with highly knowledgeable people who should be the first stop to find information on some topics. There was a guide to preparing source documents for optimal translation prepared by a small agency in Australia years ago; its original link is dead now unfortunately, but I would add the new one to a page of mine without hesitation, and if the author draws in a little extra business for that, I consider it a well-earned reward for sharing his hard-won expertise. But pages like the Guide cited above offer no real expertise; they are just more spam in a too-polluted online world.

Aug 9, 2013

memoQuickie: Exporting TMX from memoQ

Translators are often asked to export translation memory data (TMX files typically) to deliver with their work. Although this is fairly simple to do in memoQ, too often more than just the required data is sent by mistake.

Select the TM from which to export via Tools > Resource Console... > Translation memories or Project home > Translation memories.

Click Export to TMX to export all the data in the chosen TM.

If a selective export (just some of the data in the TM) is desired, click Edit, enter the filter criteria in the dialog that appears and click OK. It is also possible to filter in the editor view. Only the data shown will be included in the TMX file created.

The memoQ TM editor with filtered records shown. Click to enlarge.


Here is a "video tour" of the process:


Ten brownie points to anyone who can figure out where I "cheated" with the translation information shown in the video.

Aug 2, 2013

Translating SDL Trados Studio SDLXLIFF files & more in memoQ!



My latest demonstration video actually covers a number of memoQ features so that I would have an excuse to create this video index:
Time  Description
0:32
  Importing the first SDLXLIFF file to memoQ
1:12  Exporting the finished translation
1:27  Viewing the translation in SDL Trados Studio 2009
1:40  Re-importing the edited translation for a TM update
3:24  Saving the translation in a LiveDocs corpus for later reference
3:55  Importing a new version of the text in an SDLXLIFF source file
4:25  Comparing source text versions
5:55  Document-based pretranslation ("X-Translate")
7:11  Examining a "warning" for forgotten tags
7:46  Results of the second translation in SDL Trados Studio

That is the sort of thing I was talking about in a recent blog post about new approaches for online instruction. Many times I have wished for just such an index for long webinars or even much shorter reference videos like this one.

This tutorial was inspired by a Skype chat with a colleague in the US a few days ago. She uses memoQ but works with a number of others who use various versions of SDL Trados Studio, and there were some questions about about how one might deal with TM updates after a translation as well as the inevitable new versions that legal and financial translators often encounter. 

I have also noticed that quite a number of people are not up to date on SDLXLIFF compatibility with memoQ; this video also shows that former issues with preserving segment status have been taken care of, and everything now works well.

What is not obvious in the video is that one can also change the segmentation of the SDLXLIFF in memoQ; this happens only in the memoQ environment to allow better translation and more sensible translation memory content, and when the SDLXLIFF file is exported from memoQ, the original segmentation from Trados is preserved in the Trados environment.

Also not shown in the video is how I imported a third version of the source text, this time as a Microsoft Word file, not an SDLXLIFF. The document-based pre-translation (X-Translate) worked perfectly, and the target file was exported in the proper format (DOCX).

There are, of course, many other ways one could handle a "project" like this, but the procedure shown is not unlike what I sometimes do in projects myself.

********

I apologize for the quirky click animation in this tutorial; Camstudio had some problems I have never encountered before, and I'll have to get to the bottom of that if I keep using that tool. Otherwise, the video quality is probably the best I have achieved so far, and I would like to thank the friend who revealed the "secret" of better quality video for YouTube.

Jun 1, 2013

Translating multilingual Excel files in memoQ

Some weeks ago on a Friday, in the late afternoon, I received one of those typical Friday project inquiries: a request for a fast response on whether I would like to translate some 15 to 20 Excel files distributed in three folders, with file names redundant between the folders and over half the 50,000 source words already translated. My translations were, of course, to take the previous work into consideration and remain consistent with it. No translation memory resources were available. Fortunately for my blood pressure, I was offline that afternoon until after business hours. When I saw the request later that evening, I considered what sane approach there might be to such a project, and when none occurred to me at the time, I wrote a note to the project manager requesting more information about the source data, received no response and forgot the whole business as the usual Friday nonsense.

About a week later, while I was engaged in something completely different, it occurred to me that it would have been a fairly straightforward matter to translate the remaining text scattered through those files and build a reference translation memory from the existing translations. In fact, I could even use the available translations from other languages as references in a preview. How? By using a multilingual filter option that Kilgray added to memoQ version 6.2 (with build 15).

Finding that option is not exactly intuitive. I had heard about it but had not followed the discussions closely in the online lists, nor could I remember it from the online demonstrations I had viewed in recent months. But I knew that it worked with Excel files, so I started to import such a file and looked for the proper settings to import the source and target columns. And found nothing.

Fortunately, I used to be a software developer, so I put on my old developer's thinking cap and considered how I might best mystify users with a new feature. Aha! Name the feature something completely different! So I looked again at the list of import filters for my Excel file and found a likely candidate, the multilingual delimited text filter. (To use this filter, you must Import with options.)


The first page of the settings dialog for that filter offers Excel as one of the base formats to import:


The columns can be specified by marking the Simple bilingual configuration option, or with somewhat less confusion by examining the options on the Columns tab of the dialog. For the following test file with English source text and German as the desired target language


I used the following settings for the import:


After a little experimentation, I found that I could specify the third language (Portuguese) as a translation even though it was not a language indicated in the project. (Additional target languages are only possible with the PM edition of memoQ, but this information can be designated as a comment if needed in the Translator Pro edition.) This added the Portuguese translations (where available) to the preview in my working window:


Some odd property of the import filter in the version of memoQ tested caused the source text to be copied to the view of unpopulated translations of the project's target language in the preview, but that is of no real consequence. The preview, unlike a typical preview of an Excel file, bears no resemblance to the layout of the source file, but is instead organized by the source text grouped with other specified columns.

Considering how often I have encountered Excel files and other sources structured like this in the past decade, I would say this is probably one of the most useful filters that has been added to memoQ recently. More complex data structures may require cell background colors to be used to exclude unwanted parts of a spreadsheet (colors can be added while configuring the import). It's a shame that the current version of the filter doesn't support ranges or conditions for exclusion, but perhaps that will come later.

Making a translation memory from the existing translations in the file (which were locked and confirmed upon import in the example shown above) is a simple matter of using the command Operations > Confirm and update rows... and selecting the appropriate options. For the example shown here, selecting the locked rows would write all these to the primary translation memory:


Kilgray has a blog post and a recorded webinar (47 minutes long) with further details about using this filter. They state that "This webinar was designed for language service providers and enterprise users managing multilingual projects." However, given the frequency with which many freelancers encounter such formats and their desire to use other language information, comments, context data, etc. in their translations, I think this feature is just as relevant to freelance translators.

Update 2013-07-25: After a series of recent tests involving imports and segmentation, I wanted to see how the multilingual Excel filter would import data in which individual cells contain multiple sentences or line breaks. Theoretically the segments should correspond to the cell structures, but would they in fact? I decided to import one of my Excel files that I use for segmentation demos. To keep all the content of a cell in one segment with the regular Excel filter, I have to use a "paragraph segmentation" ruleset and set "soft breaks" as inline tags in the filter's import settings. But the default settings of the multilingual Excel filter achieve the same result:


This showed me that the "multilingual" filter might in fact save me time and trouble for importing files from certain customers where I want to avoid segmentation inside the cells altogether. And of course, the multilingual filter is an obvious quick way to load data from an Excel file and, as mentioned above for partial data, send it to a TM - a process which used to involve saving as a CSV file from Excel, worrying about saving as UTF-8, etc. That process might not even work with the test file shown here (I'm not really inclined to try it).