Showing posts with label DOC. Show all posts
Showing posts with label DOC. Show all posts

Nov 16, 2012

Trados 2007 goes to the guillotine - or not?


A recent Twitter exchange reinforced the impression of confusion I had regarding SDL's intentions with the older Trados technology. Many translators, corporate users and language service brokers continue to use the 1990s technology of Trados 2007 (which is the current translation technology of many EU institutions until it is finally phased out starting in the coming year), and recent troubles with the loss of "bilingual DOC" exports in memoQ 6.0.64 brought the matter of the old technology to a very uncomfortable head. (Earlier builds of memoQ must be used, or one must be patient until after the version 6.2 release, when this feature will be re-developed.)

The exchange with the colleague on Twitter as well as the frequent contradictions in ongoing discussions among my friends and clients in the translation world made it clear that definitive answers were needed to abate unnecessary fears and allow people to plan the future of their processes with proper information. So I talked to Paul Filkin, Client Communities Director at SDL,whose Multifarious blog is my favorite resource for reliable information about the technical arcana of Trados.

[KSL]: Judging from a recent Twitter traffic, there seems to be some confusion regarding SDL’s plans to discontinue support for SDL Trados 2007. So tell me – are Trados Workbench and TagEditor going away for good at last?

[PF] It’s worth clarifying that we are talking about SDL Trados 2007 and not SDL Trados 2007 Suite.  The difference is that the Suite contains the latest version of SDL Trados 2007 (as well as various other applications) which is 8.3.0.363.  But to answer your question specifically… no, Trados Workbench and TagEditor are not going away for good just yet.  I imagine there will still be users working with the older version of these tools for some time yet, but over time they will of course become obsolete – we just need to allow for that time.  The driving forces for anyone hanging onto these old versions will be development of hardware and new operating systems as well as upgrades to authoring systems that the retired versions will no longer be able to support.

[KSL]: What exactly ARE the difference between those two versions?

[PF] The best place to look for all the technical differences is the SDL knowledgebase where you can find a nice article called “What is new in SDL Trados 2007 Suite”:


[KSL]: If I am using SDL Trados Studio 2011 and my client expects T2007-style “uncleaned” files, what can I do?

[PF] The safest approach, because of differences between the old Trados versions is to ask your client to provide you with a fully segmented bilingual file, whether they are after TTX or Bilingual Doc.  SDL Trados Studio 2011 supports TTX and Bilingual Doc as a file type without the need for SDL Trados 2007 at all.  Your client should be able to provide these files for you because they have the appropriate software already.

The other alternative, if you don't have a copy of SDL Trados 2007 Suite which you can still purchase with SDL Trados Studio 2011 today, is to use a free application from the SDL OpenExchange called the SDLXLIFF to Legacy Converter.  This application can convert your Studio bilingual file to a Bilingual Doc or a TTX.  This process caters for two parts in this workflow.  First your client can edit these files in SDL Trados 2007 Suite and clean them into their Translation Memory, and second you can use the application to import the changes back into your SDLXLIFF so that you have the updated and approved version in your own Translation Memory.  You can get this application here:


[KSL]: What are the “dangers” in this approach? Where might it go wrong for my client?

[PF] You still have to provide your client with the “cleaned” file from Studio however because the Bilingual Doc or TTX created will not “clean up” into the fully formatted document you started with.  This is because the Bilingual Doc or TTX is created from the SDLXLIFF and not from the original source file.  

This also means that the SDLXLIFF has been segmented using the new file types in Studio and not with the old file types in Trados 2007 Suite or earlier so even though your client will be able to clean the file into their Translation Memory they may lose some ability to fully pretranslate the same source file using Trados 2007 Suite or earlier.  This is actually the same problem that could occur when converting the file using memoQ or WordFast for example but as those clients only provide the translator with the source file and not a pretranslated bilingual file in the first place this doesn't seem to be an issue for them.

So all in all both approaches seem to work… the important thing is to understand what your client wants to do with the file when they get it.

[KSL]: Can these formats be edited and “cleaned” by the client to create a properly formatted target (translated) file?

[PF] Only if they were prepared using SDL Trados 2007 in the first place.  There is no substitute for SDL Trados 2007 if the client wants a properly formatted target file and future leverage from their Translation Memory.

[KSL]: At what point can we expect support for TWB and TagEditor formats to be discontinued?

[PF] I think it’s likely that when we release the next version of the software SDL Trados 2007 Suite and SDL Trados Studio 2009 will be retired.  However, the important thing to note is that we have the Trados 2007 infrastructure built into Studio and this allows users to upgrade Translation Memories, handle legacy bilingual files and more importantly use the SDL OpenExchange to develop applications that will support workflows using the older tools.  We are already seeing developers looking at ways of improving their older solutions with Studio since we were awarded the EU contract last month.

[KSL]: Does SDL Trados Studio 2011 still include a version of TWB and TagEditor?

[PF] It’s not included automatically but you can still purchase it when you buy SDL Trados 2011.  It’s not sold as a separate piece of software anymore.

[KSL]:  That's good to know. Will this continue to be the case with the next release (Studio 2013???)?

[PF] The honest answer is we haven’t made a decision on this yet.  SDL Trados 2007 Suite is really only needed by people who have create, rather than use, these legacy files.  So in reality these people probably already have it… all they have to do is make sure they always prepare files for those who are translating them.  This may be better for them and for the translator.

Jul 27, 2012

Translating embedded objects in Microsoft Office documents

Yesterday a colleague sent me a note to say he had been searching my blog for information about translating compound Microsoft Office documents (that is documents with embedded objects) in memoQ and couldn't find any. I presume he was referring to the article about how often one CAT tool is not enough - combined workflows with other tools can frequently help solve many tricky translation problems, and DVX2 or STAR TRANSIT are definitely useful options for preparing compound Microsoft Office documents for translation in memoQ. Some time ago I recommended using STAR TRANSIT as a pre-processing tool to one of my agency friends, and he carried out a very large, complex project successfully using memoQ's excellent integration features for STAR TRANSIT projects.

There is, of course, another simple way to translate the embedded objects in a Microsoft Office document that does not involve purchasing other software licenses. I don't usually talk about it, because there are a few limitations, and until recently I had not figured out how to avoid corrupting the files when I tried to do things the "easy" way. This approach is not limited to memoQ and will actually work with most CAT tools - so SDL Trados Studio users can do this as well, for example.

It is useful to know that the Microsoft Office 2007/2010 file formats (DOCX, PPTX, XLSX) are really just ZIP files containing XML and a bunch of other stuff. That stuff includes a folder with the embedded objects in formats that can be dealt with directly.

If you have an older, binary MS Office document (DOC, PPT, XLS) with embedded objects, convert it to a 2007/2010 format.

If you rename the file extension DOCX, PPTX or XSLX to ZIP and unpack the ZIP file, inside the folder you will find a folder called "embeddings". The files in that folder can be copied elsewhere and usually handled directly in your CAT tool. But problems usually arise when you put them back, rezip the folder and change back to the original extension. The compression gets screwed up, and the Microsoft Office file is corrupted and won't open.

The only reliable method I have found for avoiding this is to use the Windows Explorer (under Windows 7) to open the ZIP file:



Here's what the "guts" of one DOCX file with a bunch of embedded Excel tables looks like:

Inside the word folder you'll find the embeddings folder:

The contents of the embeddings folder look like this:


Simply copy the embeddings folder somewhere safe, translate its contents, then copy them back to the ZIP file using Windows Explorer. Then rename the ZIP extension to the original extension for the file.

If you open the file and look at it, you'll get a shock. When you see all the objects in their original language, you might think something went wrong. Nothing bad has happened; you merely need to refresh the objects. This can be done by opening each briefly to edit or using a macro to open each object and close it again quickly. In a job with dozens of embedded objects in a long file, this macro is a helpful shortcut.

Given how easily accessible this embedded content actually is, one has to wonder why other major CAT tool providers like SDL and Kilgray have failed to offer the option of importing embedded content in their filters up to now. Let's hope they do soon. In the meantime, this workaround should enable many people to deal with this complex and irritating file format challenge.

Here's a summary of the procedure once again:
  1. Rename the *.???x file to *.zip 
  2. Under Windows 7, right-click on the ZIP file and open it using the Windows Explorer. Using ZIP tools of any kind risks corruption by changing the compression ratios. 
  3. Find the embeddings folder inside the ZIP structure. Copy this elsewhere and use it as the source for translation. It will contain all the embedded objects as single files. 
  4. Copy the translated content back into the embeddings folder in the ZIP structure.
  5. Rename the ZIP file to its original extension. 
  6. Open the file and refresh each embedded object (which will initially appear not to have been translated) by right-clicking and opening it from the context menu or running a macro to do that.

Jun 16, 2012

memoQuickie: footnote, cross-reference & index entry segmentation in Microsoft Word files

If you have a Microsoft Word DOC file or RTF to translate, it is important to be aware of the different behaviors of the memoQ import filter options you can use. If there are footnotes, cross-references or index entries, it is far better to use the option to import the DOC or RTF file as DOCX.

The DOC file shown below has a footnote, a cross-reference and an index entry:


Adding it to a memoQ project with the default filter for Microsoft Word in memoQ 5


gives the following segmentation result:


Importing the same document with the DOCX option of the filter


yields much cleaner segmentation and better tags to work with:


Compare what some other programs do with this file:

WordFast Pro
DVX2 (DOC)
DVX2 (DOCX)

TagEditor salad (partial)

SDL Trados Studio 2009 segmentation

SDL Trados Studio 2011

There is room for improvement with most tools.