Showing posts with label tagging. Show all posts
Showing posts with label tagging. Show all posts

May 5, 2022

Understanding and mastering tags... with memoQ!

Everything you need to know... in 36 pages!

Following up on the success of his excellent guide to machine translation functions in memoQ, Marek Pawelec (Twitter: @wasaty) has now published his definitive guide to tag mastery in that translation environment. In a mere 36 pages of clearly written, engaging text, he has distilled more than a decade of personal expertise and exchanges with other top professionals in language services technology into simple recipes and strategies for success with situations which are often so messy that even experienced project managers and tech support gurus wail in despair. Garbage like this, for example:


This screenshot is taken from the import of The PPTX from Hell, which a frustrated PM asked for help with just as I began reviewing the draft of Marek's book about a month ago. It contained nearly 32,000 superfluous spacing tags and was such a mess that it choked all the best professional macros usually deployed to deal with such things. Last year, I had developed my own way of dealing with these things that involved RTF bilingual exports and some search and replace magic in Microsoft Word, but when I shared it with Marek, he said "There's a better way", and indeed there is. On page 23 of this book. It was much cleaner and faster, and in a few minutes I was able to produce a clean slide set that was much easier to read and translate in the CAT tool. A page that costs 50 cents (of the €18 purchase price of the guide) earned me a 140x return and saved hours of working frustration for the translation team.

The book covers a lot more than just the esoterica of really messed up source files. It is a superb introduction to dealing with tags and markup for students at university and for those new to the translation profession and its endemic technologies, and it has sober, engaging guidance at every level for experienced professionals. I consider it an essential troubleshooting work for those in support roles of internal translation departments and, quite honestly, for my esteemed colleagues in First Level Support at memoQ. Marek is a superb trainer and an articulate teacher, with a humility that masks expertise which very often surprises, delights and informs those of us who are sometimes thought to be experts.

I am also particularly pleased that in the final version of his text he addresses the seldom discussed matter of how to factor markup into cost quotations and service charges for translations. memoQ is particularly well designed to address these problems, because weighting factors equivalent to word or character counts can be incorporated in file statistics, offering a simple, transparent and fair way of dealing with the frustrations that too often leave project managers screaming and crying in frustration shortly before... or after planned deliveries.

Whatever aspect of tags may interest you in translation technology and most particularly in memoQ, this book will give you the concise, clear answers you need to understand the best actions to take.

The PDF e-book is available for purchase here: https://payhip.com/b/tHUDx


Apr 3, 2018

Dealing with tagged translatable text in memoQ

Lately I've been doing a bit of custom filter development for some translation agency clients. Most of it has been relatively simple stuff, like chaining an HTML filter after an Excel filter to protect HTML tags around the text in the Excel cells, but some of it is more involved; in a few cases, three levels of filters had to be combined using memoQ's cascading filter feature.

And sometimes things go too far....


A client had quite a number of JSON files, which were the basis for some online programming tutorials. There was quite a lot of non-translatable content that made it past memoQ's default JSON filter, much of which - if modified in any way - would mess up the functionality of the translated content and require a lot of troublesome post-editing and correction. In the example above, Seconds in a day: is clearly translatable text, but the special rules used with the Regex Tagger turned that text (and others) into protected tags. And unfortunately the rules could not be edited efficiently to avoid this without leaving a lot of untranslatable content unprotected and driving up the cost (due to increased word count) for the client.

In situations like this, there is only one proper thing to do in memoQ: edit the tags!

There are two ways to do this:

  • use the inline tag editing features of memoQ or
  • edit the tag on the target side of a memoQ RTF bilingual review file.
The second approach can be carried out by someone (like the client) in any reasonable text editor; tags in an RTF bilingual are represented as red text:


If, however, you go the RTF bilingual route, it's important to specify that the full text of the tags is to be exported, or all you'll get are numbers in brackets as placeholders:


Editing tags in the memoQ working environment is also straightforward:


On the Edit ribbon, select Tag Commands and chose the option Edit Inline Tag


When you change the tag content as required, remember to click the Save button in the editing dialog each time, or your changes will be lost.

These methods can be applied to cases such as HTML or XML attribute text which needs to be translated but which instead has been embedded in a tag due to an incorrectly configured filter. I've seen that rather often unfortunately.

The effort involved here is greater than the typical word- or character-based compensation schemes can justly compensate and should be charged at a decent hourly rate or be included in project management fees. 

A lot of translators are rather "tag-phobic", but the reality of translation today is that tags are an essential part of the translatable content, serving to format translatable content in some cases and containing (unfortunately) embedded text which needs to be translated in other (fortunately less common) cases. Correct handling of tags by translation service providers delivers considerable value to end clients by enabling translations to be produced directly in the file formats needed, saving a great deal of time and money for the client in many cases.

One reasonable objection that many translators have is that the flawed compensation models typically used in the bulk market bog do not fairly include the extra effort of working with tags. In simple cases where the tags are simply part of the format (or are residual garbage from a poorly prepared OCR file, for example), a fair way of dealing with this is to count the tags as words or as an average character equivalent. This is what I usually do, but in the case of tags which need editing, this is not enough, and an hourly charge would apply.

In the filter development project for the JSON files received by my agency client, the text used was initially analyzed at
14,985 words; 111,085 characters; 65 tags
and after proper tagging of the coded content to be protected it was
8766 words; 46,949 characters; 2718 tags.
The reduction in text count more than covered the cost of the few hours needed to produce the cascading filter needed for this client's case and largely ensured that the translator could not alter text which would impair the function of the product.




Jul 27, 2012

Translating "foreign" bilingual tables in memoQ

--- In memoQ@yahoogroups.com, Liset Nyland wrote:
> A client has sent me a 2-column rtf-file export from DVX.
> It looks similar to the MemoQ export but not quite.
>
> The target column is full of fuzzy matches, so I need to recover these.
...
> ... do you know if there's a bilingual format exported from DVX that can
> be loaded and translated directly in MemoQ?
There is one way to deal more-or-less directly with the DVX bilingual RTF tables - or any others being introduced by other providers or bilingual tables that some customers are fond of using to store translation strings or other content. I would love to see a general import routine from Kilgray that allows selection of source and target columns of various file types in a dialog, but until then...
1. Get a copy of the PlusToyZ macros by German/English to Ukrainian/Russian translator Arkady Vysotsky.
2. Copy the source and column targets into a separate RTF or MS Word file.
3. Run the PlusToyZ macro to convert that to a Trados-like bilingual (the old Wordfast/Trados RTF/DOC bilingual)
4. Import the converted file to memoQ using the default filter, which is intelligent enough to recognize that you are dealing with Trados-compatible bilingual DOC/RTF.
5. Translate, edit, feed the TM, etc.
6. Export the processed file.
7. Use the appropriate conversion macro in PlusToyZ to turn the data back into a table.
8. Paste the data back into the original bilingual table from DVX or whatever tool it came from.
This is the preferred method to use when your bilingual table is partially pretranslated, or you have a translated table you want to edit while having a better look at the source text. This would also be a useful method for jobs I've had where customers have string or terminology lists in Excel to translate that are in some cases incomplete.

Once you get to Step 3, you can translate that bilingual format in any tool which works with the old Trados RTF/Word segmentation, such as WordFast Classic.I think that was actually the reason Arkady wrote those macros in the first place.

If you want to protect the DVX codes (or similar structures, including placeholders) or store them in the TM as proper tags, run the Regex tagger or use a cascading filter a described in my other blog post about regular expressions for DVX external table translation in memoQ. Of course, for content other than DVX tags, a different regular expression will be needed.

Jul 12, 2012

RegEx for translating DVX external view tables in memoQ

Atril's Dejà Vu was the first translation environment tool I am aware of to offer a means of exchanging translation content for review, correction and translation using an ordinary word processor. These "external views" were the original inspiration for memoQ's RTF bilingual tables, which are used in many interoperable workflows not only with people using a word processor but with many other CAT tools as well.

As with memoQ RTF bilinguals, the content in the "external view" which is not to be translated can be selected and hidden with a word processor, leaving only a target column into which the source text has been copied. But these steps alone with the standard RTF filter pose a problem:


The DVX "codes" (tags), which are represented by curly brackets enclosing a number, are not protected. Erasing parts of them can damage the content. It is also not possible to perform a tag check using the memoQ QA functions.

The solution is to use the Regex tagger in memoQ. There are two ways to do this.

If the document has already been imported,


the tagger can be run from the Format menu.

Enter the appropriate regular expression to convert the DVX code to a protected tag: \{(\d+)\}


This expression describes the pattern of the text to protect: a curly bracket (with a backslash in front of it to indicate that this is to be interpreted literally as a character, not as a bracket for grouping something), one or more digits (\d indicates a digit as opposed to d, which is just the letter d, and the plus sign means one or more) and a closing curly bracket ("escaped" with a backslash so it is understood literally as the bracket character in the DVX code.)

Click Add to put the rule in the list, then click Run tagger now.


The result is protected tags in the translation grid of memoQ. These can also be verified with a QA tag check after the translation is completed.

Your regular expression rules can be saved in the dialog above and re-used, or exported from the list under Tools > Resource console... > Filter configurations and shared with others.

The regular expression tagger can also be used as a cascading filter when the RTF file for the external view is imported:



Here the configuration can also be saved or another one loaded.

Jun 24, 2011

My first look the new custom tagger in memoQ 5.0

Many months ago while I was doing some localization updates for the Online Translation Manager (OTM) from LSP.net, the project's editor asked me if there wasn't some easy way to protect the many placeholders used for standard customer correspondence and other parts of the application. These typically looked something like [% variable %], where in the case of a variable for a company name, the placeholder might be [% COMPANY_NAME %]. In this case, during translation, care had to be taken not to omit the spaces around the variable name, mistype it or accidentally edit the characters. This usually meant copying the source to target as a precaution, but this approach has some disadvantages in efficiency, as does copying placeholders from the source to insert into a fuzzy match.

When I asked the support team at Kilgray if there was some way in memoQ to protect these placeholders, Gábor Ugray, the head of development, told me "not now" but that a solution would be at hand with the release of memoQ 5.0 and its custom tagger. I passed on that bit of news and promptly forgot it.

More recently I had an irritating small translation with a lot of markup like [B]for boldface type[/B], [U]for underline[/U] and so on. The markup played havoc with the spellchecker and was generally a nuisance. Only a few hours after I sent the finished job to my customer, I saw the solution to my problem in the introductory webinar for memoQ 5.0. "Cascading" filters and the custom tagger using regular expressions.

A few days later I had my first opportunity to try the technology myself. By then I had forgotten the work sequence from the demo and tried an approach which had not yet been fully debugged (but now works perfectly in the current build), but some generous hints and good application examples from the developers soon put me on the right track.

I imported the files like I did earlier using the Microsoft Word filter. Then I opened a file which contained the tags that concerned me and selected the command from the Format menu to run the regular expressions tagger:

In the dialog that appeared, I tested expressions for the bracketed content I wanted to convert to tags and viewed the results (saving my configuration for future use once I had what I wanted):


When I ran the tagger, the text in the working area then appeared with the markup protected as tags:


I then made a view with the rest of the files in the project and ran the tagger configuration I had saved so that all the files were properly tagged. I should have made a view of everything in the first place and tested the tagger with it, but I only thought of this later.

Pretty slick. I don't encounter this sort of challenge every day, but it comes up about once a month or more in some job, and this will make those projects much easier. Once a custom tagging filter has been configured, it can be chained ("cascaded") with other filters to form exactly the configuration you need for your file import.

Addendum / June 27, 2011: Other users' reaction to this technology:

Jun 20, 2011

memoQ 5.0: great things ahead for freelancers, LSPs and enterprises

I wasn't much of a Boy Scout today; having taken an elderly neighbor to a doctor's appointment that was supposed to take a few hours, I hurried home to catch Kilgray's introductory webinar for memoQ 5.0, and when I got a call to tell me that the treatment was complete, I left the patient stranded until Gábor Ugray, the head of development at Kilgray, had finished his fascinating presentation and answered the last question. I think it's fair to say that the upcoming version is the greatest advance in memoQ technology in the company's history. And the most exciting innovations aren't even the ones I've been looking forward to the most for anybody's product for many years.

In his presentation, Gábor did a good job of describing the benefits of the new features for the three major target groups: individual translators, language service providers (agencies) and enterprises. The major innovations in version 5.0 include:
  • change tracking and versioning
  • terminology extraction and a project lexicon
  • an "on-the-fly" tagging feature (using regular expressions) for improved content filtering and protection
  • cascading (sequential) filters for mixed context such as XML embedded in Excel files
  • a source content connector interface for CMS integration, etc.
The major benefits of the new version for each target group were described as follows:

Individual translators
  • initial analysis of project terminology with term extraction
  • term precedence in fragment assembly with the project lexicon
  • better handling of complex, mixed formats with the regex tagger and cascading filters
  • selection of the termbase to which terms are added (supports better QA procedures)
  • customized inclusion of tags in word and character counts to enable appropriate compensation for the extra effort involved in translating heavily tagged content
  • viewing of corrections using the tracked changes features
LSPs (agencies)
  • the new X-translate feature leverages previous work better and more accurately based on source documents, not TMs
  • more accurate feedback with segment histories
  • improved audit trail with reporting features
  • automation via content connectors
Enterprises
  • open interface for content management system (CMS) integration with memoQ
The new web-based technologies, qTerm (the advanced server-based terminology management tool released last spring) and WebTranslate (a soon-to-be released product offering the features of the memoQ translation client in a browser) also offer functionality of interest to many agencies and enterprises. Here the Kilgray products are behind the market introduction of similar products offered by SDL and others, but a lead in time does not necessarily translate to a lead in value or function. Kilgray's server products offer innovative features that may be decisive for many clients in the targeted markets, features which are not offered by the more expensive competitive server products.

The regex tagger essentially enables custom tagging of any content, and the configurations can be saved and used sequentially with other filters. This would have been very useful in a number of projects I have done over the years, where HTML content was embedded in Excel or XML files and particular care was needed to avoid damaging the tags. I can also use this feature to protect placeholders in an ongoing localization project I do based on Java Properties files. Placeholders like [% COMPANY_NAME %] can be turned into tags and protected from damage. The tagger and the ability to chain file filters together are the features I find most personally interesting in the new version, though it is the terminology features I have wanted the most and for the longest time and shall benefit from in nearly every project. The combinations of sequential filters can also be saved for re-use.

The new memoQ statistical term extraction feature will not replace SDL's MultiTerm Extract for bilingual term mining of TMs, an application I am rather fond of. It is, however, a clearly superior tool for monolingual term extraction from source texts, and the integration with term bases looks promising at first glance. I am also quite excited by the inclusion of a feature I suggested some time ago: combination of term "hits" by stemming. If, for example, I want to combine various forms of a German adjective like säurefest, säurefeste, säürefestes, säurefestem, säürefesten, etc., I merely set a pipe character after the word root (säurefest|) and all the other entries and their statistics will be combined. I don't know at this point if it is possible to define exceptions, though I do see this as necessary, for example for German superlatives. But it's a good start in any case.

The tracked changes feature and versioning will be extremely useful to those who have to deal with projects where new versions come fast and furious and it's easy to lose the overview. It also has a lot of potential to improve editing and review workflows. Changes can be displayed between any two document versions, and "snapshots" of work in progress can also be saved for purposes of work and comparison.

The new features I've described are by no means all of what we can expect in memoQ 5.0. Gabor mentioned a number of other things in passing, such as "watched folders", which I assume will enable some sort of file import automation in projects, but frankly there was so much to absorb, and I found myself dwelling on a number of personally exciting points, so that I missed other features which may also be of interest.

Version 5.0 will be published at the end of June as a release candidate (i.e. presumably table beta version), with the official release a few weeks later after the inevitable initial bugs are ironed out. During that transition period, a parallel installer will be available so that users can work safely with the current version 4.5 while still making the acquaintance of the new version.

Congratulations are due to the Kilgray team for the many important advances in the upcoming release. I also look forward to the effect this may have on the market: by setting the bar of innovation and usability higher than ever, Kilgray should further inspire its competitors to respond with interesting new features and variations on these new themes. And that will be good for all of us.