An exploration of language technologies, translation education, practice and politics, ethical market strategies, workflow optimization, resource reviews, controversies, coffee and other topics of possible interest to the language services community and those who associate with it. Service hours: Thursdays, GMT 09:00 to 13:00.
Showing posts with label XLIFF. Show all posts
Showing posts with label XLIFF. Show all posts
Aug 26, 2019
Exporting compatible XLIFF (XLF) bilingual files from memoQ
Here we go again. Although memoQ is the undisputed leader for compatibility and interoperability among translation environment tools, users still encounter problems exchanging files, particularly XLIFF of some sort, with users of other tools. This is not because of any actual difficulty producing compatible XLIFF files, but rather a matter of deficient tool training and the failure to date by memoQ product designers to make the ease of interoperability a little more obvious. Some other tools, like recent versions of SDL Trados Studio, come pre-configured on installation to recognize the proprietary file extensions for memoQ's flavor of XLIFF ("MQXLIFF") and renamed ZIP packages (MQXLZ) containing XLIFF files, but others (or versions of SDL Trados Studio from many years ago) need to be configured to recognize those extensions, or someone simply has to change the MQXLIFF file extension to an extension that will be recognized by any tool: *.xliff or *.xlf are the choices.
The two-step solution is shown here:
On the Documents ribbon in memoQ, click on the tiny arrow under the Export icon and choose the option to export a bilingual file. There is some blue text which, if clicked, will allow a compatible XLIFF file to be exported, albeit with the MQXLIFF extension that some other programs might not recognize.
When the Export button in the dialog (marked 1, above) is clicked, the Save As dialog (marked 2, above) appears, simply change the file extension (the part after the period) to "xlf", for example. Then any program that reads XLIFF files can work with the file you export from memoQ. Despite the change of extension, memoQ will still recognize the file it produced, so it is possible to re-import it, for example if another person has made corrections to the XLIFF file that you want to use to update your translation or reference resources.
In some much older versions of memoQ, it does not work to change the extension in the export dialog; this has to be done directly to the exported file in whatever folder you save it in.
Of course, all of this will be rather difficult if you are one of those users who has not fixed the awful Microsoft Windows default to hide the extensions of known file types. Fixing that particular stupidity requires slightly different measures in different versions of Windows, but in Windows 10 you can do that on the View ribbon of Windows Explorer by marking the choice to show file name extensions:
Jan 8, 2019
Translating "smarter"
In response to my recent piece on the use of Fluency to translate Microsoft Publisher files, I received the following comments from the former's technical support:
It appears that you flat out ignored (or disabled) the warning message that pops up every time you try to open a Publisher file (attached).
There is no problem with Fluency with regards to Publisher. The issue lies in the fact that the Publisher interop is unstable. This (and the fact that professionals don’t use Publisher) are the reasons that CAT tools don’t support Publisher.
Are you hoping to ridicule us to encourage us to fix Publisher support? We’d just as soon remove support for it altogether, but we do have a few users who are grateful for it and understand that they might have to restart their computer a few times or kill the Publisher process that is frozen in the background. Another option that works, sometimes, is to try and save multiple times, which you discovered, but incorrectly attributed to resizing the view of the text (which doesn’t affect the output).
So we let you know before you start that Publisher is unstable and you’ve just spent a bunch of time documenting how it is unstable. To what end?
As for our XLIFF files, yeah they aren’t great, but extremely rarely is someone exporting something from Fluency to another CAT tool.
Regards,
Richard Tregaskis
Western Standard Support
Em: support@westernstandard.com
Ph: 801-224-7404
I'm not sure about the part that "professionals don't use Publisher". Certainly graphics professionals don't; when I earned my bread that way I usually used software like PageMaker, Quark Xpress or FrameMaker, nowadays Adobe InDesign seems to be the tool of choice. But I wouldn't think of some engineer in a technical department who uses Microsoft Publisher to write a manual as being unprofessional. Just foolish maybe, but no more so than the ones who use CorelDraw or even MS PowerPoint (!!!) in the same crazy way. People make the choices they do in the circumstances they work in, and professionals try to meet them at least halfway where possible to accomplish the necessary objectives. Fluency does that in the case of Microsoft Publisher, but one would do a service to the customer to suggest that another publishing platform might suit their needs better in the same way that responsible translation consultants often suggest that PDF is perhaps not the ideal format to provide for translation, and that the original format (if it isn't a Microsoft Publisher file or paper) might work better for everyone.
What concerns me about the response of Fluency's technical support, however, is the apparent lack of concern for the compatibility of their XLIFF files. If these cannot be exchanged readily with other platforms, one must ask what actual purpose they serve. Indeed, what would that be? And, perhaps, whether Fluency is really to be taken seriously as a professional platform for translation work.
Some years ago in a period where it looked a bit dark for my platform of choice, I thought that Fluency showed promise as a working platform, and I made a serious effort to investigate its suitability for my routine work as a translator of legal and scientific material. I was charmed by the generally functional approach to transcription, but the translation side of things was less encouraging, riddled with bugs at nearly every stage. After a few days I ran screaming back to more stable, well supported work platforms. The handling of SDLPPX (SDL Trados Studio package files) in particular was the sort of disaster one doesn't easily forget; even with products whose developers care about functionality and compatibility there are issues time and again as the SDL Trados platform evolves. I can only imagine what would happen with Fluency Now if I tried out one of those test files that a friend at SDL likes to play tricks on me with.
XLIFF is serious business. These days it is often the basis not only of interoperable processes with CAT tools but for all manner of bilingual exchange processes. And thus, until the technical support and/or development department of a tool takes this format seriously and makes a reasonable effort to ensure at least basic interoperability, that tool cannot be taken seriously for professional work.
![]() |
| That should be a question mark, not a period :-) |
Jan 4, 2019
Translating Microsoft Publisher files
Every few months or so I run across a question in social media or am confronted with a project like this:
Some time ago, Paul Filkin published an interesting discussion of an Open Exchange application that enables SDL Trados Studio users to deal with the Microsoft Publisher format with some limitations; in the article, he also discussed other approaches, including one I have known about for some time: the use of Western Standard's Fluency.
![]() |
Some time ago, Paul Filkin published an interesting discussion of an Open Exchange application that enables SDL Trados Studio users to deal with the Microsoft Publisher format with some limitations; in the article, he also discussed other approaches, including one I have known about for some time: the use of Western Standard's Fluency.
I looked at Fluency some years ago, and while I found some interesting things there, such as its transcription module, on the whole the application never seemed ready for prime time with its sloppy programming of details. I spent some time trying to persuade its underfunded team to correct some of the problems I saw, but after a while it became clear that the company and its product were not able to cope with the demanding technical challenges routinely faced by language service providers today.
The discussion which followed the posted question suggested a number of approaches, but if the colleague's client expected to receive a translated PUB file instead of some other format, the only realistic option for this possibly one-off job would be to use Fluency in some way. I assumed (and suggested) that a workflow involving
*.pub <-> Fluency <-> (exchange format) <-> memoQ
might do the trick (with the exchange format probably being XLIFF, but otherwise the bilingual RTF format that I remembered from my tests of Fluency long ago.)
And so it proved to be. But the Devil is in the details.
The first sign of trouble came from a colleague - a professor at a local university who is known for his technical curiosity and flexibility in translation courses - who told me that Fluency does indeed offer an XLIFF export but that memoQ experienced problems importing it. His description of the error message sounded a lot to me like the typical mistakes that CAT tool programmers who are XLIFF newbies make when implementing a spec that they are probably too lazy to read and test. (I found the same error myself and submitted it to memoQ Support for comment a few hours ago.) He said that he had then tried the RTF export, but it wasn't clear to me what the result was and he was under time pressure, so I didn't press the matter but resolved to have a look myself.
I used a modified English template file for an invitation as my PUB file to test. The file imported easily into Fluency:
The Fluency user interface offered a sort of WYSIWYG representation for the text, which makes it appear not bad for work, though appearances are deceiving. In fact, this proved to be a source of some trouble later.
As mentioned, the XLIFF export could not be used in memoQ, and although I am capable enough of analyzing structure problems in a tagged file, I wasn't in the mood to clean up someone else's mess, so I exported a "Fluency Work File" as my next attempt. That is app jargon for a bilingual RTF file similar to that found in other applications.
The difference with Fluency RTFs is that they include the WYSIWYG text representation. Nice, really, and this makes the work in another environment a little easier. I copied the source text column and pasted it into a new file (DOCX), then imported that to memoQ for translation:
Afterward, the translation exported from memoQ was pasted into the target column of the Fluency Work File (bilingual RTF exchange file). I imported that bilingual file back into Fluency and then exported a translated PUB file using the File / Save As command. I got a strange error message saying that there had been some trouble with the export and that some manual adjustment might be needed in Microsoft publisher.
At first glance I thought, "Looks OK" and then... WTF??? Everything was OK except the title. Not only was the text cut off, it was not even the text I had translated in German. When I copied the text out of the field and pasted it into Notepad, this is what I saw:
In my nearly 5 decades of casual and occasionally professional programming I have seen almost every stupidity imaginable, so in this case I imagined that somehow the problem lay in sloppy programming associated with text that is longer than the space provided in the field. Interestingly, Fluency enabled me to change the size of the target text in the translation window, so I reduced it by about half and tried to export a new target PUB file.
That worked in fact. So Fluency can indeed be used as a sort of filter for Microsoft Publisher files to be translated in other tools such as memoQ, but the process is not without trouble on the Fluency side, at least when text overruns the field size available, as one might expect to happen with some frequency.
Western Standard offers a 15-day trial of Fluency Now, their desktop tool for freelance translators, and the application can be paid on a monthly subscription of only 15 US dollars. So perhaps for the occasional project or client that requires work with PUB files that is an option. Microsoft Publisher is not taken seriously as a layout and publishing tool by graphics professionals and CAT tool providers, but because it is part of the Microsoft Office suite, one will find it in use from time to time, and this imperfect solution may be the best option for helping such clients.
And so it proved to be. But the Devil is in the details.
The first sign of trouble came from a colleague - a professor at a local university who is known for his technical curiosity and flexibility in translation courses - who told me that Fluency does indeed offer an XLIFF export but that memoQ experienced problems importing it. His description of the error message sounded a lot to me like the typical mistakes that CAT tool programmers who are XLIFF newbies make when implementing a spec that they are probably too lazy to read and test. (I found the same error myself and submitted it to memoQ Support for comment a few hours ago.) He said that he had then tried the RTF export, but it wasn't clear to me what the result was and he was under time pressure, so I didn't press the matter but resolved to have a look myself.
I used a modified English template file for an invitation as my PUB file to test. The file imported easily into Fluency:
![]() |
| I assume that "terminology" download is some silly, unhelpful public domain dictionary I would never use. |
The Fluency user interface offered a sort of WYSIWYG representation for the text, which makes it appear not bad for work, though appearances are deceiving. In fact, this proved to be a source of some trouble later.
As mentioned, the XLIFF export could not be used in memoQ, and although I am capable enough of analyzing structure problems in a tagged file, I wasn't in the mood to clean up someone else's mess, so I exported a "Fluency Work File" as my next attempt. That is app jargon for a bilingual RTF file similar to that found in other applications.
The difference with Fluency RTFs is that they include the WYSIWYG text representation. Nice, really, and this makes the work in another environment a little easier. I copied the source text column and pasted it into a new file (DOCX), then imported that to memoQ for translation:
Afterward, the translation exported from memoQ was pasted into the target column of the Fluency Work File (bilingual RTF exchange file). I imported that bilingual file back into Fluency and then exported a translated PUB file using the File / Save As command. I got a strange error message saying that there had been some trouble with the export and that some manual adjustment might be needed in Microsoft publisher.
At first glance I thought, "Looks OK" and then... WTF??? Everything was OK except the title. Not only was the text cut off, it was not even the text I had translated in German. When I copied the text out of the field and pasted it into Notepad, this is what I saw:
Tag der Tag der kulturellen VielfaltNo joke. Fluency somehow went berserk exporting the text of the title field, and sliced, diced and multiplied the whole mess in a truly bizarre way.
kulturellen Vielfalt
Vielfalt
kulturellen Vielfalt
kulturellen Vielfalt
Vielfalt
kulturellen Vielfalt
kulturellen Vielfalt
Vielfalt
In my nearly 5 decades of casual and occasionally professional programming I have seen almost every stupidity imaginable, so in this case I imagined that somehow the problem lay in sloppy programming associated with text that is longer than the space provided in the field. Interestingly, Fluency enabled me to change the size of the target text in the translation window, so I reduced it by about half and tried to export a new target PUB file.
That worked in fact. So Fluency can indeed be used as a sort of filter for Microsoft Publisher files to be translated in other tools such as memoQ, but the process is not without trouble on the Fluency side, at least when text overruns the field size available, as one might expect to happen with some frequency.
Western Standard offers a 15-day trial of Fluency Now, their desktop tool for freelance translators, and the application can be paid on a monthly subscription of only 15 US dollars. So perhaps for the occasional project or client that requires work with PUB files that is an option. Microsoft Publisher is not taken seriously as a layout and publishing tool by graphics professionals and CAT tool providers, but because it is part of the Microsoft Office suite, one will find it in use from time to time, and this imperfect solution may be the best option for helping such clients.
Jun 15, 2018
Better WordPress translation with memoQ
Translating websites is mostly a royal pain in the tush. And I avoid it most of the time. Why? Several reasons.
So for translations of WordPress websites properly configured to use WPML technology, the new memoQ filter looks like a winner!
- Those who request website translations often have no idea what platform is used nor do they really know how much content is present.
- They have very little understanding of the technical details or importance of translatable information hidden in tag attributes, selection lists, etc. and so there are often misunderstandings about the true volume to be translated.
- There are a lot of sloppy cowboys slogging through the bog, glibly bidding low rates to translate sites they neither understand nor truly care about, and their victims... uh, prospects, customers, whatever... usually lack the expertise or the patience to understand the difference between a wild-ass lowball guess from someone lacking the skills and tools to do the job right and a carefully researched, reasonably accurate estimate of time and effort from a professional.
Shopping for "quotes" when neither you nor the one submitting a "bid" actually understand the technical basis of the project is a process with no guarantee of a satisfactory outcome. And too often this process turns out badly.
These days, many small companies use the popular content management system WordPress to manage their web sites. It may not be the best by some technical standards, but sometimes it is better to define "best" according to the likelihood of finding someone to provide services involving a platform and of there being such experts available not only now but for a reasonable amount of time in the future. I think it is fair to say that WordPress has met that standard for some time and will probably do so for some time more.
I have had a good number of requests for translating WordPress content in the past, but none of the estimates given were accepted, because typically the content to be translated was an order of magnitude greater than the client realized or nobody could commit to a clear decision on what parts were to be translated and what parts were unimportant. And then we have the problem that many sites use themes which are poorly designed as multilingual structures.
The WordPress Multilingual Plug-in (WPML) makes sensible, professional management of websites with content in more than one language much easier. When I learned about this technology more than a year ago, I suggested its use to the person who requested a quotation for translation services, but that suggestion is probably still echoing somewhere out there in the Void.
At memoQ Fest 2018 this year in Budapest, I had the pleasure to attend a superb presentation by Stefan Weimar on how to cope with the translation of Wordpress sites and some of what you need to know to use the WPML technology right. I was inspired and hoped to have an opportunity to look at things more closely some day.
That day turned out to be a week later. Funny how that goes.
Three or four years ago I translated a small web site for a friend's company. At the time, the site used the Typo3 content management system, which proved to be troublesome. Not so much because of the technology, but because of the service provider using it, who rejected any suggestions for providing the content to be translated in a form that would not require his manual intervention at the text level. He copied, pasted and improved (German: verschlimmbesserte) my translation as only a German with full confidence in his grade school English skills could. It was... not what anyone had hoped for, and I never found the heart to mention all the mistakes in the final result.
So now, when someone asked me to have a look at their new site, I felt a bit queasy. Nunca mais, I thought. No way, José. Or Wolfgang as it were. But in the meantime, unbeknownst to me, he had switched service providers and CMS platforms, and the new provider managing his web content is a professional with a professional understanding of sites for international clients in many languages. And he uses WPML. The right way!
So now it was up to me to figure out what's what in memoQ. So first I used the memoQ XLIFF filter on all the little XLIFFs supported by the plug-in. I quickly saw that a few other things were needed, like a cascaded HTML filter...
Somewhat messy, but doable once the HTML tags get properly protected by a chained filter.
Then I tried again, this time applying memoQ's WordPress (WPML) filter. And this was the result:
That was easy. Hmm. I think I know which method I prefer.
Three or four years ago I translated a small web site for a friend's company. At the time, the site used the Typo3 content management system, which proved to be troublesome. Not so much because of the technology, but because of the service provider using it, who rejected any suggestions for providing the content to be translated in a form that would not require his manual intervention at the text level. He copied, pasted and improved (German: verschlimmbesserte) my translation as only a German with full confidence in his grade school English skills could. It was... not what anyone had hoped for, and I never found the heart to mention all the mistakes in the final result.
So now, when someone asked me to have a look at their new site, I felt a bit queasy. Nunca mais, I thought. No way, José. Or Wolfgang as it were. But in the meantime, unbeknownst to me, he had switched service providers and CMS platforms, and the new provider managing his web content is a professional with a professional understanding of sites for international clients in many languages. And he uses WPML. The right way!
So now it was up to me to figure out what's what in memoQ. So first I used the memoQ XLIFF filter on all the little XLIFFs supported by the plug-in. I quickly saw that a few other things were needed, like a cascaded HTML filter...
Somewhat messy, but doable once the HTML tags get properly protected by a chained filter.
Then I tried again, this time applying memoQ's WordPress (WPML) filter. And this was the result:
That was easy. Hmm. I think I know which method I prefer.
So for translations of WordPress websites properly configured to use WPML technology, the new memoQ filter looks like a winner!
Jun 14, 2018
Translating Wordfast GLP packages... elsewhere.
One reason to keep translation environment tool licenses up to date is that new formats continue to appear. New formats for translatable files as well as new file formats for the tools that help to process files for translation. Very often I have heard some "professional" say "I'm a translator, not a [fill in the blank]. If the client wants this translated, I'll have to get it in a Microsoft Word file." Or something like that.
Let's get real for a moment.
- That attitude is simply lazy and disrespectful toward translation consumers who would like to make use of one's services and
- a lot of money is being left on the table here in many cases. I built a huge clientele at the start of the last decade, because my use of translation environment tools like Trados, Déja Vu, STAR Transit and Wordfast enabled me as an individual to tackle translation challenges that many agencies at the time had no concept of how to cope with.
As translation agencies have acquired more technical tools, most of them still remain unfortunately unaware of how to use them properly or plan more than the simplest workflows well, but that's a subject for another day. Also...
- ... by using tools and techniques that are compatible with what your clients require for a final format, you can save your client a lot of time and money for further layout work - and probably avoid the introduction of errors in your translation work in its final format as well.
- And in my experience, showing technical and process competence to benefit clients usually leads to greater trust and better work together.
So what has all this got to do with Wordfast?
Well... I didn't like the Wordfast brand for a very long time. Its various incarnations were perhaps the weakest of the popular tools in a technical sense, and inevitably when agency friends called me, desperate to fix some massive translator screw-up (usually by somebody in France), Wordfast "Pro" was often involved in the disaster.
I looked at the "newer" Wordfast versions a number of times over the years, and honestly they always seemed like lobotomized wannabe tools. This was about the time that many other toolmakers were trying to decide if they should support XLIFF.
Well, a lot has changed since then. I became aware of the changes the other day when somebody posted a question in a social media forum for memoQ asking how to handle Wordfast Pro 5 GLP packages. I had never heard of these, so of course I was curious and decided to take a look. This finally led me to download a 30-day trial of the latest Wordfast Pro software to evaluate its potential for interoperable work with other translation environments. I see a lot of changes since my last look, and so far I think they are all positive, and along the way I had good cause to look at Wordfast Anywhere, the free web-based CAT tool that I talked some university colleagues into not wasting their time with a while ago. Well, my recommendation in that regard might change, but that and commentary on the latest incarnation of WF Pro will have to wait for another day.
About those GLP packages....
Yes, those. This was the question:
Someone pointed out that GLP files - like every other translation "package" one finds from all the tool providers - are merely ZIP files with particular structure inside and the extension re-named.
Gotta love Facebook. You'll always get an answer in some group, usually a wrong one. That's why I keep a blog. Good information gets buried in social media noise too often, and good luck finding it in any kind of search. In this case... we don' have no steenkeen TXML files as I learned... that's the old Wordfast Pro....
A colleague in Germany kindly provided me with a little GLP package to examine, which I promptly unzipped. I noticed that at least one tool (7-Zip) sees through the renamed extension nonsense and saved me the usual trouble of renaming it before unpacking.
So far, so good... inside the folder for the unpacked GLP file I found the following:
The test package was an English to Portuguese project. But source? Hello? Let's have a look there!
Very interesting. The original source files (English) came along for the ride. This is good, because I often like to translate source files in memoQ - taking advantage of the preview there for many file types - and then use the translation memory to translate the file that is created by other other tool (usually SDL Trados SDLXLIFF files in my work). Now let's have a look inside the pt target folder. There's actually another folder named txlf inside that one. And there I found:
No TXML files! TXLF is a new instance of the rather ubiquitous XLIFF files one finds in the translation world, some of which have some rather bothersome "extensions" that may require special handling in the translation process. In the simple test I performed, none of that was apparent; an ordinary XLIFF filter seemed to work well. Future tests will show me if there are any quirks I hope, but so far, so good.
So one strategy, with pretty much any CAT tool, would be to unpack the GLP file, get at those TXLF files and then bring them into another working environment using an XLIFF filter. Maybe also use my approach with the source files too, which will ensure that you can deliver a good target file even if quirky tags in the XLIFF lead you to produce less than an optimal result there.
The current version of memoQ (8.4) does not recognize the TXLF extension, so as in all such cases, the All files option must be used and the correct filter applied in a later dialog. Unlike with some other tools, memoQ cannot be "trained" by the user to recognize new extensions as far as I know.
But what about importing the GLP files directly to memoQ? Wouldn't that be nice? And I thought it might be possible using the ZIP file filter recently introduced (and the same All files trick to get the GLP file and apply the ZIP filter later). Well...
It looked promising.
So much so that I even optimistically named and saved a custom configuration for the ZIP filter. All I need to do now is cascade an XLIFF filter!
Ack. Sooooo close. I've been here before. There are more things in heaven and down-to-earth cascading formats, Kilgray, than are dreamt of in your philosophy! Please, please expand the list of possible cascaded formats sensibly to make better use of this lovely new ZIP filter!
So for now, that's a no-go, but soon? Who knows? If you bother support@kilgray.com and tell the memoQ team how helpful it would be, maybe this and similar problems can be solved with relative ease.
In any case, for now it seems that the unpack-and-do-the-XLIFF approach will work for most anyone with a modern CAT tool. And that's good news, because in today's fast-changing technology environment for translation, interoperability of CAT tools is increasingly important. It is a foolish waste of time to translate in a large number of CAT tools and probably a bad idea to do so in two or three according to my old research. I've usually found that such JOATs are, professionally, often stupid goats who lack the depth in a single major environment or two, which could allow them to get the most out of their tools and serve their clients in the best way with their linguistic skills and subject matter knowledge.
So is the latest Wordfast a tool worth checking out? I don't know yet. But it may be used by colleagues and clients with whom I like to work, and understanding how to share projects and project resources in painless ways will benefit all of us, no matter what our tool preferences may be. Wordfast seems to be developing very much in that spirit, so I will revisit it for more collaboration scenarios in the future.
Apr 4, 2018
Complicated XML in memoQ: a filtering case example
Most of the time when I deal with XML files in memoQ things are rather simple. Most of the time, in fact, I can use the default settings of the standard XML import filter, and everything works fine. (Maybe that's because a lot of my XML imports are extracted from PDF files using iceni InFix, which is the alternative to the TransPDF XLIFF exports using iceni's online service; this overcomes any confidentiality issues by keeping everything local.)
Sometimes, however, things are not so simple. Like with this XML file a client sent recently:
Now if you look at the file, you might think the XLIFF filter should be used. But if you do that, the following error message would result in memoQ:
That is because the monkey who programmed the "XLIFF" export from the CMS system where the text resides was one of those fools who don't concern themselves with actual file format specifications. A number of the tags and attributes in the file simply do not conform to the XLIFF standards. There is a lot of that kind of stupidity to be found.
Fear not, however: one can work with this file using a modified XML filter in memoQ. But which one?
At first I thought to use the "Multilingual XML" filter that I have heard about and never used, but this turned out to be a dead end. It is language-pair specific, and really not the best option in this case. I was concerned that there might be more files like this in the future involving other language pairs, and I did not want to be bothered with customizing for each possible case.
So I looked a little closer... and noted that this export has the source text copied exactly to the "target". So I concentrated on building a customized XML filter configuration that would just pull the text to translate from between the target tags. A custom configuration of the XML filter was created after populating the tags by excluding the "source" tag content:
That worked, but not well enough. In the screenshot below, the excluded source content is shown with a gray background, but the imported content has a lot of HTML, for which the tags must be protected:
The next step is to do the import again, but this time including an HTML filter after the customized XML filter. In memoQ jargon, this sort of configuration is known as a "cascading filter" - where various filters are sequenced to handle compounded formats. Make sure, however, that the customized XML filter configuration has been saved first:
Then choose that custom configuration when you import the file using Import with Options:
This cascaded configuration can also be saved using the corresponding icon button.
This saved custom cascading filter configuration is available for later use, and like any memoQ "!light resource", it can be exported to other memoQ installations.
The final import looks much better, and the segmentation is also correct now that the HTML tags have been properly filtered:
If you encounter a "special" XML case to translate, the actual format will surely be different, and the specific steps needed may differ somewhat as well. But by breaking the problem down in stages and considering what more might need to be done at each stage to get a workable result with all the non-translatable content protected, you or your technical support associates can almost always build a customized, re-usable import filter in reasonable time, giving you an advantage over those who lack the proper tools and knowledge and ensuring that your client's content can be translated without undue technical risks.
Sometimes, however, things are not so simple. Like with this XML file a client sent recently:
Now if you look at the file, you might think the XLIFF filter should be used. But if you do that, the following error message would result in memoQ:
That is because the monkey who programmed the "XLIFF" export from the CMS system where the text resides was one of those fools who don't concern themselves with actual file format specifications. A number of the tags and attributes in the file simply do not conform to the XLIFF standards. There is a lot of that kind of stupidity to be found.
Fear not, however: one can work with this file using a modified XML filter in memoQ. But which one?
At first I thought to use the "Multilingual XML" filter that I have heard about and never used, but this turned out to be a dead end. It is language-pair specific, and really not the best option in this case. I was concerned that there might be more files like this in the future involving other language pairs, and I did not want to be bothered with customizing for each possible case.
So I looked a little closer... and noted that this export has the source text copied exactly to the "target". So I concentrated on building a customized XML filter configuration that would just pull the text to translate from between the target tags. A custom configuration of the XML filter was created after populating the tags by excluding the "source" tag content:
That worked, but not well enough. In the screenshot below, the excluded source content is shown with a gray background, but the imported content has a lot of HTML, for which the tags must be protected:
The next step is to do the import again, but this time including an HTML filter after the customized XML filter. In memoQ jargon, this sort of configuration is known as a "cascading filter" - where various filters are sequenced to handle compounded formats. Make sure, however, that the customized XML filter configuration has been saved first:
Then choose that custom configuration when you import the file using Import with Options:
This cascaded configuration can also be saved using the corresponding icon button.
This saved custom cascading filter configuration is available for later use, and like any memoQ "!light resource", it can be exported to other memoQ installations.
The final import looks much better, and the segmentation is also correct now that the HTML tags have been properly filtered:
If you encounter a "special" XML case to translate, the actual format will surely be different, and the specific steps needed may differ somewhat as well. But by breaking the problem down in stages and considering what more might need to be done at each stage to get a workable result with all the non-translatable content protected, you or your technical support associates can almost always build a customized, re-usable import filter in reasonable time, giving you an advantage over those who lack the proper tools and knowledge and ensuring that your client's content can be translated without undue technical risks.
Labels:
cascading filter,
filters,
HTML,
Iceni Infix,
MemoQ,
XLIFF,
XML
Jun 24, 2017
The other sides of Iceni in Translation
The integration of the online TransPDF service from Iceni in memoQ 8.1 has raised the profile of an interesting company whose product, the Infix PDF Editor, has been reviewed before on this blog. TransPDF is a free service which extracts text content from PDF files, converts it to XLIFF for translation in common translation environments, and then re-integrates the target text from the translated XLIFF to create a PDF file in the target language.
This is a nice thing, though its applicability to my personal work is rather limited, as not many of my clients would be enthusiastic if I were to send PDF files as my translation results. Sometimes that fits, sometimes not. And of course, some have raised the question of whether using this online service is compatible with some non-disclosure restrictions.
I think it's a good thing that Kilgray has provided this integration, and I hope others follow suit, but for the cases where TransPDF doesn't meet the requirements of the job, it is useful to remember Iceni's other options for preparing text for translation.
Translatable XML or marked-up text export
As long as I can remember, the Infix PDF Editor has offered the option to export text on your local computer (avoiding potential non-disclosure agreement violations) so that it can be translated and then re-imported later to make a PDF in the target language. Only the location of this option in the menus has changed: the menu choices for the current version 7 are shown below.
This solution suffers from the same problem as the TransPDF service: not everyone will be happy with the translation in PDF, as this complicates editing a little. However, I find the XML extract very useful to put the content of PDF files into a LiveDocs corpus for reference or term extraction. The fact that Infix also ignores password protection on PDFs is also helpful sometimes.
"Article" export
The Article Tool of the Iceni Infix PDF Editor enables various text blocks on different pages of a PDF file to be marked, linked and extracted in various translatable formats such as RTF or HTML. The quality of the results varies according to the format.
Once "articles" are defined, they are exported via the command in the File menu:
The RTF export has some problems, as this view in Microsoft Word with the format characters made visible reveals:
However, the Simple HTML export opened in Microsoft Word shows no such troubles (and can be saved in RTF, DOCX or other formats):
Use of the article export feature requires a license for the Infix PDF editor, unlike the XML or marked-up text exports for translation. In demo mode, random characters are replaced by an "X" so that one can see how the function works but not receive any unjust enrichment from it. However, this feature has significant value for the work of translators and is well worth an investment, as the results are typically better than using OCR software on a "live" (text-accessible) PDF file.
But wait... there's more!
Version 7 also has an OCR feature:
I tested it briefly on some scanned Portuguese Help Wanted ads that I'll probably use for a corpus linguistics lesson this summer; the results didn't look too awful all considered. This feature is worth a closer look as time permits, though it is unlikely to replace ABBYY FineReader as my tool of choice for "dead" PDFs.
Sep 16, 2015
Getting around language variant issues in memoQ LiveDocs
I was told by some other users that a fundamental change had been made in the way language data are accessed in LiveDocs. It was said that until a few versions ago it had been possible to use documents for reference in LiveDocs regardless of their sublanguage settings. So I was told. The truth is more complicated than that.
According to my tests, memoQ 2015 is the first version of memoQ to have a logically consistent treatment of language variants for both bilingual and monolingual documents in corpora. All the other versions tested (memoQ 2013R2, 2014, 2014R2) are equally screwed up and show the same results.
The "visibility" of a monolingual or bilingual document when viewed in a corpus attached to a project running under memoQ 2015 follows these rules:
According to my tests, memoQ 2015 is the first version of memoQ to have a logically consistent treatment of language variants for both bilingual and monolingual documents in corpora. All the other versions tested (memoQ 2013R2, 2014, 2014R2) are equally screwed up and show the same results.
The "visibility" of a monolingual or bilingual document when viewed in a corpus attached to a project running under memoQ 2015 follows these rules:
- the sublanguage (language variant) settings for source and target (of the document or the project) must match the project
- or the language setting (of the document or the project) must be generic.
Two rules. Pretty simple. It doesn't matter what version of memoQ the project or corpus was created in, only which version is actively running.
I created a test corpus with the following document mix:
The corpus contained 11 documents, both bilingual and monolingual with a mix of generic language settings and settings with language variants specified (such as German for Germany, Switzerland and Liechtenstein and English for Zimbabwe, the US and UK).
In a project running under memoQ 2015 with the languages set to generic German and generic English, all 11 documents in the corpus were accessible.
So if you want access to all LiveDocs corpus data for the major languages of your project, it is necessary to use generic language settings, either when you load the data into LiveDocs (difficult unless you always use the resource console, since adding documents to a corpus from within a project automatically applies the project's language settings!) or in the languages specified for the project itself. And this will only work with memoQ 2015. If you want to apply penalties to particular language variants this can be done using keyword markers (as seen in the screenshot above) and configuring the More penalties tab of the LiveDocs settings file applied to that corpus.
If the same corpus is attached to a project running under memoQ 2015 with language settings for Swiss German and generic English, the documents available from the corpus are these:
For a Swiss German and UK English project under memoQ 2015, this is the picture:
And for a Germany's German and US English:
All the screenshots above can be predicted based on the two rules stated. Work it out.
"But what happens with earlier versions of memoQ?" you might wonder. It's messy. Here is a look at a Swiss German and UK English project under memoQ 2013 R2, 2014 and 2014 R2:
And here's a project with generic German and Generic English under memoQ 2013 R2, 2014 and 2014 R2:
In each case the five bilingual documents are visible no matter what the project's language settings are. However, there is strict adherence to language variants and the generic language setting for monolingual documents! In my opinion, that's for the birds. I see no good reason to follow a different rule for data availability in bilingual versus monolingual documents. So in a sense, Kilgray has cleaned up this inconsistency in the latest version of memoQ.
Some have expressed a desire for a "switch" setting to allow language variant settings to be ignored. And perhaps Kilgray will provide such a feature in the future. But the best way to get there now is simply to make your project's language settings generic.
Changing the language settings for bilingual data in an existing LiveDocs corpus
If you have a corpus with a mix of language settings and you want to convert these to generic settings or a particular variant, this can be done as follows currently only for bilingual documents:
- Select the bilingual documents to export from the corpus and export them to a folder. (If you choose to zip them all together, unpack the *.zip file later to make a folder of the exported *.mqxlz files.
- Re-import the *.mqxlz files to the LiveDocs corpus via the Resource Console so you are able to specify the exact language settings you want. In the import dialog, you'll have to change the filter setting manually from "binary" to "XLIFF". These *.mqxlz files are not the same as bilingual files from a translation document in a project and are not recognized automatically.
Unfortunately, there is no way to change the language settings of a monolingual document except to re-import it in the Resource Console in its original form and set the language variant (or generic value) there.
So really, for now, the best way to go seems to be to use memoQ 2015 with generic project language settings.
Jan 22, 2014
memoQ cloud: a team server "on tap"
This afternoon, Kilgray CEO István Lengyel held one of the best webinars I've seen him do yet to describe the convenient new hosted server facilities known as memoQ cloud, which I reviewed recently.
In the webinar, he explained the company's evolution of thought for online computing and how concerns about security were finally resolved to create a more sustainable offering than the more support-intensive "honeymoon" server solution.
He made it clear how existing desktop licenses for the Project Manager and Translator Pro editions can be used in combination with concurrent access licenses (CALs) for the server, as well as how cloud services can be suspended for periods in which they are not needed, saving considerable costs for those with only occasional needs to work in a coordinated online team.
Backing up the server configuration can be done quickly and easily from a Language Terminal account, so if cloud service is dormant for more than three months (after which data are deleted from the server), everything can be restored quickly when needed.
The webinar also included a demonstration of the integrated translation in web browsers, memoQ WebTrans. This is one way of providing access to the server for others who do not have installed copies of memoQ or working on your server when using other computers. Of course this interface also works in web browsers under other operating systems, such as MacOS or Linux. (Click on the graphic below to get a full-sized view of the web translation interface.)
Access to Kilgray's premium terminology server qTerm and memoQ server APIs is also available for an additional subscription fee. Subscribed services can be changed at any time as your needs evolve.
In the webinar, István showed how in about the same time it takes to enjoy a cup of coffee, one can get a free Kilgray Language Terminal account and register with a credit card for a month's trial of the memoQ cloud server (with any services available) for just €1/$1. If you are trying out services which you will not want beyond the trial period (like the API, qTerm or extra licenses), these can be set to cancel at the end of the trial period to avoid unwanted charges.
The embedded video below is a 20-minute tour of how simple it is to set up and manage projects in memoQ cloud. Use the icon at the lower right of the video frame to watch this on your full screen.
This is a good overview of the process, although the licenses aren't explained very well, and the project type recommendation is bad advice in many cases, as I pointed out in my post on server projects on segmentation and projects with desktop documents. Everything else in the video is good, but it's often very important to allow segmentation to be changed or corrected, particularly if the segmentation rules used in the project do not cover abbreviations which may split sentences in very unfortunate ways. If you need to have instantaneous access to work from other team members by using online documents, the the segmentation will need to be checked very carefully and corrected before the project begins to avoid difficulties.
Those testing the memoQ cloud server or using desktop editions of memoQ may also want to check out various free configuration resources on Language Terminal. These include special QA profiles, AutoCorrect files, import filters that are not part of the shipping product and auto-translation rules for easier translation of number and date formats, etc. Language Terminal offers other facilities which may be of interest even to those who do not use memoQ, such as the free InDesign server, which can create PDF previews of InDesign documents (very useful for reviews before delivery) or convert InDesign files of any type to XLIFF for translation in many different environments.
UPDATE:
The memoQ cloud webinar is now available to watch on Kilgray's page for recorded webinars; it can be accessed directly here or viewed in the embedded video below.
In the webinar, he explained the company's evolution of thought for online computing and how concerns about security were finally resolved to create a more sustainable offering than the more support-intensive "honeymoon" server solution.
He made it clear how existing desktop licenses for the Project Manager and Translator Pro editions can be used in combination with concurrent access licenses (CALs) for the server, as well as how cloud services can be suspended for periods in which they are not needed, saving considerable costs for those with only occasional needs to work in a coordinated online team.
Backing up the server configuration can be done quickly and easily from a Language Terminal account, so if cloud service is dormant for more than three months (after which data are deleted from the server), everything can be restored quickly when needed.
The webinar also included a demonstration of the integrated translation in web browsers, memoQ WebTrans. This is one way of providing access to the server for others who do not have installed copies of memoQ or working on your server when using other computers. Of course this interface also works in web browsers under other operating systems, such as MacOS or Linux. (Click on the graphic below to get a full-sized view of the web translation interface.)
Access to Kilgray's premium terminology server qTerm and memoQ server APIs is also available for an additional subscription fee. Subscribed services can be changed at any time as your needs evolve.
In the webinar, István showed how in about the same time it takes to enjoy a cup of coffee, one can get a free Kilgray Language Terminal account and register with a credit card for a month's trial of the memoQ cloud server (with any services available) for just €1/$1. If you are trying out services which you will not want beyond the trial period (like the API, qTerm or extra licenses), these can be set to cancel at the end of the trial period to avoid unwanted charges.
The embedded video below is a 20-minute tour of how simple it is to set up and manage projects in memoQ cloud. Use the icon at the lower right of the video frame to watch this on your full screen.
This is a good overview of the process, although the licenses aren't explained very well, and the project type recommendation is bad advice in many cases, as I pointed out in my post on server projects on segmentation and projects with desktop documents. Everything else in the video is good, but it's often very important to allow segmentation to be changed or corrected, particularly if the segmentation rules used in the project do not cover abbreviations which may split sentences in very unfortunate ways. If you need to have instantaneous access to work from other team members by using online documents, the the segmentation will need to be checked very carefully and corrected before the project begins to avoid difficulties.
Those testing the memoQ cloud server or using desktop editions of memoQ may also want to check out various free configuration resources on Language Terminal. These include special QA profiles, AutoCorrect files, import filters that are not part of the shipping product and auto-translation rules for easier translation of number and date formats, etc. Language Terminal offers other facilities which may be of interest even to those who do not use memoQ, such as the free InDesign server, which can create PDF previews of InDesign documents (very useful for reviews before delivery) or convert InDesign files of any type to XLIFF for translation in many different environments.
UPDATE:
The memoQ cloud webinar is now available to watch on Kilgray's page for recorded webinars; it can be accessed directly here or viewed in the embedded video below.
Subscribe to:
Posts (Atom)



















































