Showing posts with label fuzzy matches. Show all posts
Showing posts with label fuzzy matches. Show all posts

Jan 22, 2019

The Ultimate Comparative Screwjob Calculator for translation rates

Some years ago I put out a number of little spreadsheet tools to help independent translators and some friends with small agencies to sort out the new concepts of "discount" created by the poisonous and unethical marketing tactics of Trados GmbH in the 1990s and adopted by many others since then. One of these was the Target Price Defense Tool (which I also released in German).

The basic idea behind that spreadsheet was the rate to charge on what looked to be a one-off job with a new client who came out of nowhere proposing some silly scale of rate reductions based on (often bogus and unusable) matches. So, for example, if your usual rate was USD 0.28 per word and that's what you wanted to make after all the "discounts" were applied, you could plug in the figures from the match analysis and determine that the rate to quote should be USD 0.35, for example.

Click on the graphic to view and download the Excel spreadsheet
Fast forward 11 years. Most of the sensible small agencies run by translators who understand the qualities needed for good text translation are gone, their owners retired, dead or hiding somewhere after their businesses were bought up and/or destroyed by unscrupulous and largely incompetent bulk market bog "leaders" with their Walmart-like tactics. Good at sales to C-level folk, with perhaps a few entertaining "inducements" on the side, but good at delivering the promised value? Not so much in cases I hear. And many of the good translators who haven't simply walked away from the bullshit have agreed to some sort of rate scale based on matching (despite the fact that there is no standard whatsoever on how different tools calculate these "matches" and now with various kinds of new and nonsensical "stealth" matches being sneaked in with little or no discussion).

So now, it's not so much whether a translator will deal with a given rate scale for a one-off job, but more often what the response should be to a new and usually more abusive rate scale proposed by some cost- and throat-cutting bogster who really cares enough to shave every cent that an independent translator can be intimidated to yield, thus destroying whatever remaining incentive there might be to go the extra mile in solving the inevitable unexpected problems one might find in many a text to translate.

And this, in fact, was the question I woke up to this morning. I told the friend who asked to go look for my ancient Target Price Defense Tool, but I was told that it wasn't helpful for the case at hand. (It actually was, but because of the different perspective that wasn't immediately obvious.)

Click on the graphic to view and download the Excel spreadsheet
So I built a new calculation tool quickly before breakfast which did the same calculations but in a little different layout with a somewhat different perspective: the Comparative Screwjob Calculator (screenshot above), because really, the point of these match scales is to screw somebody.

Shortly after that, I was asked to include the calculations of "internal matches" from SDL Trados (which are referred to as "homogeneity" in the memoQ world, stuff that is not in the translation memory but where portions of text in the document or collection of documents have some similarity based on their character strings - NOT their linguistic sense). And of course there are other creatively imagined matches in some calculation grids - for subsegments in larger sentences (expect to get screwed if an author writes "for example" a lot) or based on some sort of loser's machine pseudo-translation algorithms that some monolingual algorithm developer has decided without evidence might save the translator a little effort - cut that rate to the bone!). So I expanded the spreadsheet to allow for additional nonsense match rate types ("internal/other") and to compare a third grid which can be used, for example, to develop a counterproposal if you are currently billing based on an agreed rate scale and a new one is proposed (all the time keeping in view how much you are losing versus the full rate which might very well be getting charged to the end customer anyway).

Click on the graphic to view and download the Excel spreadsheet
The result was the Ultimate Comparative Screwjob Calculator (screenshot above). Now that's probably too optimistic a name for it, because surely those who think only of translators as providers of bulk material to be ground up for linguistic sausage have other ways to take their kilos of flesh for the delivery mix.

If this all sounds a bit ludicrous, that's because it is. I am a big fan of well-managed processes myself; I began my career as a research chemist with a knowledge of multivariate statistical optimization of industrial processes and used this knowledge to save - and make - countless millions for my employers or client companies and save hundreds of jobs for ordinary people. I get it that cost can be a variable in the equation, because starting some 34 years ago I began plugging it in to my equations along with resin mix components and whatnot.

But the objective I never lost sight of was to deliver real value. And that included minimizing defects (applying the Taguchi method or some other modeling technique or just bloody common sense). And ensuring that expectations are met, with all stakeholders (don't you hate that word? it reminds me of a Dracula movie in my dreams where I hold the bit of holly wood in my hand as we open the coffin of thebigword's CEO) protected. That is something too few slick salesfolk in the bulk market bog understand. They talk a lot of nonsense about quality (Vashinatto: "doesn't matter"; Bog Diddley: "no complaints from my clients who don't understand the target language", etc.). But they are unwilling to admit the unsustainable nature of their business models and the abusive toll it takes on so many linguistic service providers.

So use these spreadsheets I made - one and all - if you like. But think about the processes with which you are involved and the rates you need to provide the kind of service you can put your name to. The kind where you won't have to say desperately and mendaciously "It wasn't me!" because economic and time pressures meant that you were unable to deliver your best work. That goes as much for respectable translation companies (there are some left) as well as for independent service professionals who want to commit to helping all their clientele receive what they need and deserve for the long run.


Jun 5, 2017

Optimizing term properties for many entries in a memoQ termbase

Terminology: On my wishlist: an easier way to deal with termbases imported into MemoQ in Studio packages. Especially annoying: the habit of many Studio users capitalizing termbase entries & thus torpedoing recognition. It would seem the default setting in MultiTerm is fuzzy matching.

memoQ is noted for its compatibility with SDL Trados Studio files and projects; with the latest release of memoQ (version 8.1) there is apparently full compatibility with tracked changes in SDLXLIFF files and with Studio's translation quality assurance. However, there are apparently a few little points remaining to satisfy some.

The opening comment is from a colleague who seems to experience less than optional matching for terminology which is imported as part of a memoQ project using an SDL Trados Studio package (SDLPPX). The solution to this person's frustration is fairly simple, however, and it is useful in many other cases where the properties of terms in a memoQ glossary are not well optimized.

Many people are unaware of the fact that it is possible to change any of the term properties for a large number of terms at once. To do this, simply open the memoQ termbase for editing and select the terms to change. Multiple selections can be made by holding down the Shift key and clicking on the desired range of rows or by using the control key to mark individual selections. Then simply set the desired property (such as fuzzy matching) and the change will be applied to all of the selected terms.

About four years when fuzzy term matching was introduced by Kilgray I made a short video about this. The memoQ interface is a little different since then but the procedure works just as well today:

Nov 26, 2015

Fuzzy match of the month - WTF?!


Experienced translators using translation environment tool technology are quite familiar with the ludicrous results often obtained by so-called "fuzzy" matches in translation. For some 20 years now, the lie has been propagated that such matches usually help translators to work faster and that such "matches" therefore obligate one to offer discounts.

I will not rehash the familiar arguments and evidence that even truly close matches with the difference of a little word or two can cost more time that translation from scratch with no reference text or the fact that modern translation tools are useful primarily as a guide to facilitate consistency and not necessarily speed of work, especially if the translator is a real one with strong language and subject matter skills. Of course there are monkey-level jobs where a fuzzy match can usually be expected to save time, but once one ventures into fields such as legal or financial translation this is not the case as often as the linguistic sausage providers (aka LSPs) might claim.

I just wanted to share this little screenshot from my "daily bread", because it truly is worthy of sewer disposal.

All fuzzy matches are not created equal; every tool on the market will spew nonsense, and these nonsensical "values" are not even close to consistent between tools. It's time to cut the crap with fuzzies as a real means of evaluating work effort. Or at least share some of what the believers are smoking to reach such conclusions.

Dec 11, 2013

General settings for memoQ TMs

memoQ TM settings are found in the Resource Console, the Options and a project's Settings.
This is a very useful "light resource" which is well worth nearly every user's time.
To define the TM settings to be used in new projects, select a settings configuration under Tools > Options... >  Default resources > TM settings (in the row of icons) by marking its checkbox.

To define the default TM settings to be used in the project you have opened, go to Project home > Settings > TM settings (in the row of icons) and mark the checkbox for the desired project default.

Different settings for individual TMs in a project (for example to set higher or lower match criteria) may be applied by going to Project home > Translation memories, selecting the TM of interest, clicking the Settings command at the right of the window and choosing the settings to apply instead of the project's standard TM settings.

The General settings tab is the same for all currently supported versions of memoQ. Role options are included on another tab in memoQ 2013 R2, and the Project Manager editions of memoQ offer additional possibilities for filtering and/or applying penalties to content on a Filters tab.


Match thresholds
The first value here (minimum) controls the fuzzy percentage below which a match will not be displayed in the translation results pane at the upper right of the working translation window.

The "good match" threshold is relevant to pretranslation (though this is unfortunately not made obvious in the dialog). The default value of 95% is really too high and would only apply to matches with small differences in tags or numbers; since any small difference in words is penalized significantly in memoQ (something I find very helpful, as I can understand more quickly what differences to look for compared to working in Trados). I usually set my "good matches" to 80%.

Not a "good match" according to the memoQ TM default setting
Penalties
In my work, an alignment penalty, which is a deduction from the match rate of a translation unit created by feeding an alignment to a translation memory, does not make a lot of sense. This is because
  • I almost never send alignments to a TM. Why bother? LiveDocs may be slower in pretranslation, but it provides context matching just like a TM, and you can actually read what you find in a concordance search in its original document context. TMs suck because you do not get the full context for your matching segment and are thus at greater risk for missing information which may be important for a translation. This is especially the case with short match segments.
  • if I happen to be aligning a dodgy translation and want to send it to a TM, I'll put it in a "quarantine TM" which already has its own penalty.
  • on those rare occasions when I might feed an alignment to a TM, it's because the content is going to a user of another CAT tool, and if that person uses Trados or another tool that can read XLIFF files or other available bilingual formats, I'll send the data as that instad, so it can be reviewed and modified more easily before feeding to a TM. This also gives the other person a bilingual reference with document context.
  • alignment for TMs is soooooo 1990s!
User penalties: If you have the misfortune to share a TM with someone whose work you do not trust completely and you want to avoid letting that person's 100% and context match segments slip past you unnoticed, apply a suitable penalty for the level of "risk" that person represents. If you want to be sure that user's content never gets used in a pretranslation and never appears in the translation results pane, apply a whopping big penalty like 80%. Those segments not be shown or inserted but will still be there in a concordance search if you want them.

TM penalties: Sometimes a client provides you with a TM you do not trust completely, or you may have a "quarantine TM" with content of dubious quality. Or I might have a TM with good content in British English but need to deliver a translation in American English. Applying penalties to such TMs will reduce the priority of their matches and prevent 100% matches with inappropriate language from slipping past without more careful inspection. As in the case of user penalties, you can also apply a very large penalty to ensure that matches will never be displayed in the translation results pane or used in a pretranslation but still have the TM content available for concordance searches.

Adjustments
It seems to be a good idea generally to enable the adjustment of fuzzy hits and inline tags. In many (but not all) cases, this will correct small differences in numbers, punctuation, cases and inline tags.

The only significant effect I was able to determine in adjusting the inline tag strictness in my tests was that more permissive settings might count a match with different tags as a full match. While this might meet the requirements of some clients hoping to impose discount schemes, from a quality assurance perspective, this does not seem like a good idea, and I believe it is better to have a strict setting here to draw attention to differences and reduce the chance that errors might be overlooked.

Oct 16, 2013

Small caveats for memoQ fuzzy term matching

In the months since it was introduced this year, the terminology fuzzy match feature of memoQ has proved to be a great help in my work. The authors of the German texts I translate are sometimes particularly challenged with respect to spelling, and I might find the same source word spelled five or six different ways in a text, with some or all of the variations repeated frequently: Scheidungsurteil, Schaidungsurteil, Scheidungurteil, Scheidungs-Urteil, Scheidung Urteil and so on. It can be a real nuisance trying to keep terms in the translated text consistent when the source text is out of control this way, and often I've make frustrated searches for a term I know I put in the termbase, only to find that it was spelled a little differently. And then some of the changes between plural and singular forms could contribute to the difficulties of consistency, particularly in large texts with the terms thinly sown.

For cases such as these, the fuzzy term matches have been enormously helpful. I no longer have to make a catalog of crappy spelling and enjoy my bit of Schadenfreude as I share the hard-won terminology with the client in a pretty PDF dictionary that proudly displays all the misspelled variants in the source mapped to a single clean target term. Now I can maintain cleaner termbases that may actually be useful for reversed application (with German as the target language, not loaded with garbage spelling just to catch the matches with German as the source language).

But there's a dark side to this too. In some cases, I am misled by how fuzzy term matches are highlighted. An example of this can be seen here:


Look at the highlighted fuzzy term match in segment 3. The prefix un- is a negative, so in this case we're talking about fake raccoon skin underwear. Not the real thing. The problem is compounded in the QA term check:


The translation in this case actually correct, but it is flagged as an error because of the fuzzy match. I'm not actually sure there is anything to be done about this except perhaps identify troublesome cases like this and change the term entries to custom with appropriate wildcards or some other setting less likely to report an false error or overlook a real one. However, I think it might be a help for visual checking if the blue highlighting for fuzzy matches could be set not to extend to prefixes or portions at the end which go beyond the length of the term entry. Of course I do not know what the implications of this may be for other languages, so changes of any kind require careful thought.

Addendum:
Right after I posted this, I was contacted by a friend who had the same frustration with the misleading matches with some financial terms. This translator said that even adding the correct term for the mismatch did not correct the problem, and the proper match would not be displayed. That sounded very strange to me, so I had to have a look. I added the "fake raccoon underwear":


But my translation results pane showed both matches. What really bothered me, however, was that the worse match still took precedence for insertion as the tool tip indicates:


Oops. This doesn't change my positive opinion of the fuzzy matching for terms. It's still extremely helpful and overall helps me maintain better term consistency. But there are some things in the current version (6.5.15) which need a little tuning - like this goofy precedence problem - but even after any bugs are fixed there will still be a few inherent risks of which one may need to be aware and for which some particular QA strategies may need to be considered.


Jul 24, 2013

What good is memoQ fuzzy term matching?


When Kilgray introduced fuzzy term matching with the release of memoQ 2013, I was first concerned with how it worked after a few puzzling tests of the feature. Discussions with the development team soon cleared up that mystery, and I wrote an article describing the current fuzzy state of term matching technology in the translation environment tool that has done such a fine job of waking SDL and others from the long slumber of innovation that prevailed in the last decade.

But questions still remained in the minds of most users as they asked why they should care about this feature and what good it would really do for them.

The answer to that has become clearer for me as I have used the feature in recent weeks and noticed certain things. Like the fact that crappy spelling in my source texts is not as much of a burden for term matching any more:


This actually applies to more than just bad spelling. Those who translate from English will benefit from the fact that fuzzy term matching will help them if the UK source term is in the glossary but the author of the text used an American spelling. I cope with problems caused by old and new spelling conventions in German as well as the fact that a great many Germans cannot agree on how their compound words should be glued together. And my Portuguese friends tell me every week about the hassles of the spelling reform in progress in that linguistic corner.

Fuzzy term matches is currently not implemented for QA checking in memoQ, but I think it would make sense for Kilgray to add this feature to allow fuzzy term matching for QA on the source side. It could be a bit of a disaster to have it on the target side, however, for reasons I will leave readers to guess.

For those who want to set their termbases to use fuzzy matching by default in a particular language, here is a short video that shows how to change the termbase properties and how to change to match settings for legacy terms to "fuzzy":


I was initially a bit skeptical of the latest version of memoQ, but as this feature and a few others have begun to "sink in", while I still don't feel comfortable with the company's hyperbole over new features like LQA, which is largely pointless for freelance translators, I do feel confident in saying that fuzzy term matching is a reason for most of us to seriously consider upgrading to memoQ 2013. This will be even more the case if it is added to the QA features.

Ah, but what about the change to the comments function, Kevin? You really hated that!

There's more to say on that topic now. Some of it is even good.

Jun 7, 2013

Understanding fuzzy term matching in memoQ 2013

Perhaps the most interesting and potentially useful feature for me in the recently released memoQ 2013 is fuzzy term matching. I have wanted something like this for several years, and several efforts at harmonizing terminology in a large, collaborative project last year made it clear that this might be very helpful in identifying deviations from agreed terminology in cases where that terminology appears as part of a compound word (as it sometimes tends to do in German, my source language).

So when I finally downloaded the latest version of memoQ last weekend and began testing, fuzzy terminology was the second thing I looked at (after the current comment mess). My initial tests left me very, very confused. Each example I created gave different results, and it was not easy to discern a pattern from examining just a few terms. The explanation of why this feature works as it does can be difficult to follow, at least as far as I am able to explain it, so many readers may be better off to read my conclusions in the next paragraph and skip everything below it (except maybe the last graphic).

Fuzzy term matching in memoQ 2013 is a real  improvement for terminology matching and quality assurance involving terms, at least for my language pair. This is not an easy challenge that the developers have taken on, but some useful results have been achieved and no harm has been done to previous functionality. And I expect that this feature will be the subject of further refinement and improvement for other languages as users make the case for these.

My first quick test of the feature involved a verb, the German word for "to wait" ("warten"). I put it into a test termbase and then imported a translation text consisting of various sentences that used forms of the verb. I noticed that there was a term hit for "warte", but nothing for "gewartet". After adding "warte" to the termbase for fuzzy matching, there was still no match for "gewartet", although it contained that character sequence.

Then I tried another example with "Gesetz" (law). There I seemed to hit the jackpot. There were hits with Unweltgesetz (a typo, but typical of many source texts I see), Umweltgesetze, Gesetzentwurf and Umweltgesetzentwurf with blue background highlighting of the character sequence matching the termbase entry.

A third term produced more confusion: with "Ausführung" in the termbase, there was no match for "Farbausführungungen", but there was a match for "Farbausführungsbeispiel". Clearly this is not a simple matching function.

A question to Kilgray Support brought an answer that explained the match behavior. The current implementation of fuzzy term matching in memoQ uses a combination of rules which depend on the index, the "edit distance" (calculated differences between the entry string and the characters in the term to match) and, depending on the language, character maps and a threshold length for possible compound words.

German, it seems, is a privileged language, the only one for which compound word recognition rules are currently active. Apparently five characters are the minimum to be recognized as a word, so the "Farb-" in "Farbanwendungen" wasn't enough, but "-beispiel" triggered the compound recognition rule that caused "-anwendung-" to be matched in the middle of "Farbanwendungsbeispiel". I imagine that compound matching would be useful for Dutch and some other languages, and the developer suggested that expanding coverage to other languages as needed would not be difficult.

Character mapping - defined equivalence between letters - is implemented for German, Hungarian, Italian and Spanish to allow matching in cases where letters may change with plural formation, for example. Thus the German word "Bratapfel" in the termbase would yield a hit for the plural form "Bratäpfel".

Edit distance is calculated by dividing the number of deviating characters by the number of total characters in the fuzzy term entry. A match is currently assumed if the edit distance is 0.2 or less. The term "warte" differs from the six-letter entry "warten" by one character; 1/6 is less than 0.2, so memoQ 2013 reports a term match. In the example of "Farbausführungen" above, the calculated edit is 6/10 (because 6 letters - four on the left, two on the right - are added to the term entry), and because this is larger than 0.2 no match is indicated. If, on the other hand, "Farbtonausführungen" occurs, a match will be found, because the added segment "Farbton-" meets the requirement of five or more letters for a compound word.

How relevant can this feature be to your language? What changes might be required to the matching behavior to obtain useful results in your source language(s)? Your feedback to Kilgray's support team and feedback from others working with your languages are the best way to help improve the usefulness of fuzzy term matching for your language. So speak up.

If you find this feature useful and want fuzzy term matching as the default for new entries in a termbase, this can be set in the properties for the termbase under New term defaults... for termbases created in memoQ 2013. Older termbases will also display this option, but it won't actually work in practice. To use this new feature with old term collections, these will need to be migrated to a memoQ 2013 termbase.


Mar 3, 2011

Homogeneity: another "secret" competitive weapon with memoQ

Earlier today I received an e-mail with the following question:

At the moment we are wrestling with an analysis issue that should be solvable but we don't know how to. As I always see your posts about all kinds of TM issues, I was hoping you might be able to provide some advice.


The case is as follows:
From one of our clients we have received what is basically a list of tools in Excel for translation (NL-FR). My colleague made an initial quotation for the project based on the Trados analysis, which revealed 23% repetitions in the file. However, the client received a much lower quote from a different provider. The reason for this, according to him, is that there are a lot of high fuzzy matches in the file which the other provider has counted but Trados doesn't (for example, "... metaalzaagbeugel 12 inch zwaar model met D-greep" and "... metaalzaagbeugel 12 inch zwaar model met rechte greep".)


Do you know whether there is a way (or tool other than Trados) that does count these fuzziess when performing an analysis?
To me, this sounds an awful lot like my PM acquaintance has been blind-sided by Kilgray's homogeneity analysis, which has been a feature of memoQ for a very long time. It's a feature about which I personally have mixed feelings. Used in the wrong way by unscrupulous agencies or ignorant persons, it can be yet another club with which to clobber translators and their rates to the ground and bring about the Hobbesian state of being so many fear is in our future, if not our present. But I approach it as a valuable information tool for helping me estimate how much time a rush project might actually take. Or in the case of my correspondent's competitor, it can be used judiciously to calculate a competitive rate that might not land you in the poorhouse.

Classic Trados and most other CAT tools calculate fuzzy matches based on the content of a translation memory. If these sentences:
The cat is black and white.
The dog is black and white.
The rabbit is black and white.
do not have something similar in a TM used for analysis, they will all be counted as "No Match" segments. However, with a good tool like Atril's Déjà Vu X and it's functional "assembly" technology, similar sentences like these are handled almost like 100% matches from a TM. But DVX still won't tell you about the time you might save.

Kilgray's memoQ analyzes a text for internal redundancies and "fuzzy redundancies", the latter being referred to as having a degree of "homogeneity". But as anyone who works with CAT software knows, even high fuzzy matches can be utterly useless and cost more time than content with no statistical similarities. Translation is about meaning, not statistics, and the price assassins at Trados and other tool pimps of the past sold everyone a lousy bill of goods with nonsense marketing lies like "You'll never have to translate the same sentence again." Well, guess what? If you do successive versions of an information brochure or technical manual and don't start to update your language after a while, your text will soon sound like it was written for an age long past and might not communicate as clearly as it should. Those who can read German should have a look at the various editions of the classic cookbook Die Süddeutsche Küche by Katharina Prato, which was popular from the mid-19th century until the 1930s for truly dramatic examples of the changes in a language. (These are available online via Google Books and various libraries online. They are also a good source of offal recipes - people ate all manner of interesting things back then.) But this happens on a much shorter time scale as well: my eight-to-ten-year-old texts for the AOK social insurance brochure and various IT manuals sound rather awful and dated, though they were quite acceptable at the time they were written.

Used as a planning tool, however, the homogeneity function in memoQ can give you valuable information and help you compete more effectively in difficult times and markets.