Showing posts with label Iceni Infix. Show all posts
Showing posts with label Iceni Infix. Show all posts

Jun 24, 2018

What's missing in most training videos: found!

Five years ago more or less I wrote a post in which I optimistically declared that if I ever did a one-hour webinar I would edit it down to perhaps twenty minutes. The real problem for me was that there was an ever-growing catalog of video instructional material from Kilgray, SDL and other sources but that it was virtually impossible to find parts of a video with specific points of interest without wasting a lot of time. Long teaching videos need a time index.

In a few blog posts after that, I created some manual indices for some of my videos, but all of these required manual scrolling to get to a particular point. And then, while using TechSmith Camtasia to touch up the recording of a recent webinar I did on PDF handling with iceni InFix, I stumbled across a menu item I had not noticed before:


Timeline markers? Hmmmm. Why would something like that be needed? Unless maybe one could build an index with them? And indeed that is the case.

When exporting a local MP4 video file or uploading a video production to YouTube, for example, the dialogs contain options to use these markers and their labels (the blue texts seen in the screenshot above on the video editing timeline) to build a table of contents.


Wow. This is exactly what I wanted to do for years. And Camtasia is used by a lot of people I know, so I wonder why nobody ever mentioned this possibility or how they could all overlook it. The result looked like this on YouTube:


All the blue number codes are hotlinks that jump the video to exactly that play time. This makes it easy to refer quickly to some important point of interest and skip the rest. Now I'm not going to go back and rework all of my old translation tool tutorial videos, but I'll use this feature for any new recordings, and I hope others do the same.

The video of the PDF talk is embedded below, but as you can see, the TOC isn't available with embeddings.


But there is sort of a workaround for that problem, using sharing links that include the starting time:

Click this graphic to go to 21:42 in the video

But that won't control an embedded video in a web page - like the one above. If anybody has a solution for that, I would love to hear it.

May 10, 2018

Zooming inside iceni InFix for PDF translation: web meeting on 21 June 2018


Over the course of the last nine years, I have published a few articles about ways that I have found the PDF editor iceni InFix useful for my translation and terminology research work. Throughout that time iceni has continued to improve that product as well as develop other technologies for PDF translation assistance, such as the online TransPDF service now integrated with memoQ.
It's one thing to have a tool and in many cases quite another thing to know how to make the best use of it. This situation is further complicated by the very wide range of scenarios in which an editor like iceni InFix might be useful and the great differences one often finds in the needs and expectations of the clientele from one translator to another. In the product's early days I followed the commentaries of José Henrique Lamensdorf, a Brazilian engineer with long experience in technical translation, desktop publishing and other fields, and while I consider him to be among the most useful sources of good technical information for me in my early days as a commercial translator, his project needs were very different from mine, and most of the things he mentioned a decade or more ago, though very relevant to people heavily involved with publishing, weren't a fit for my clientele.
That changed as iceni expanded the feature set over the years and I began to encounter many cases where OCR and a full Adobe Acrobat license did not quite do what I needed in a simple way.


Some weeks ago I had an online meeting scheduled with a client company to discuss the advantages of certain support technologies with that company's translation and project management staff. We tried to use TeamViewer for the discussion, but unfortunately my license could not accommodate the 6+ people involved, and I was reluctant to fork over the extra cash needed for a 15 or 25 participant license, especially because some other clients had issues with TeamViewer which I never clearly understood, leading their IT departments to ban it. And the TVS recording files, while generally quite decent for viewing and of manageable size due to an excellent compression CODEC, are a nightmare to convert cleanly to MP4 or other common video formats. Just as I was caught in this dilemma, my esteemed Portuguese to UK English translation colleague and gifted instructor at Universidade Nova de Lisboa, David Hardisty, enthusiastically re-introduced me to Zoom videoconferencing.

I had seen Zoom before briefly when IAPTI decided to ditch the Citrix conferencing solutions and use it for webinars and staff meetings, but at the time I was too distracted by other matters to remember the name or notice the details. And, as we know, there one finds the Devil.

Zoom is powerful and flexible. For about €13 a month for my Pro license, I can invite up to 100 people for a web meeting, with quite a few useful options that I am still getting a grip on. Being used to the relative simplicity of TeamViewer, I am a little overwhelmed sometimes, and I have had a few recorded client meetings where the video was flawed because I got the screen sharing options mixed up. But the basics are actually dead simple if one pays a bit of attention.

A Zoom "web meeting", by the way, is what I would call a webinar, but that term means something else in Zoomworld, involving up to 50 speakers and something like 10,000 participants for some monthly premium. Not my thing. If the crowd is bigger than 10 in an online or a face-to-face class, I start to feel the constrictions of time and individual attention like an unruly anaconda around me.

But in any case, for someone who has spent many years looking for better teaching tools, Zoom is looking pretty good right now. And it enables me to share what I hope is useful professional information without dealing with the organizational nonsense and politics often associated with platforms licensed by some companies and professional associations. All for the monthly price of a cheap lunch.

So I've decided to do a series of free public talks using Zoom, not only to share some of a considerable backlog of new and exciting technical matters for translators, translation project managers and support staff and language service consumers, but also to get a better handle on how I can use this tool to support friends, colleagues and students around the world. Previously I announced a terminology talk (on May 24th, mostly about memoQ); now I have decided to share some of the ways that iceni InFix helps me in my work and what it might do for you too.

Soon Thursday, June 21st at 16:00 Central European Time (15:00 Lisbon time) I'll be talking about how you can get your fix of useful PDF handling for a variety of challenging situations. You are welcome to join me for this.

The registration link is here.


Apr 4, 2018

Complicated XML in memoQ: a filtering case example

Most of the time when I deal with XML files in memoQ things are rather simple. Most of the time, in fact, I can use the default settings of the standard XML import filter, and everything works fine. (Maybe that's because a lot of my XML imports are extracted from PDF files using iceni InFix, which is the alternative to the TransPDF XLIFF exports using iceni's online service; this overcomes any confidentiality issues by keeping everything local.)

Sometimes, however, things are not so simple. Like with this XML file a client sent recently:


Now if you look at the file, you might think the XLIFF filter should be used. But if you do that, the following error message would result in memoQ:


That is because the monkey who programmed the "XLIFF" export from the CMS system where the text resides was one of those fools who don't concern themselves with actual file format specifications. A number of the tags and attributes in the file simply do not conform to the XLIFF standards. There is a lot of that kind of stupidity to be found.

Fear not, however: one can work with this file using a modified XML filter in memoQ. But which one?

At first I thought to use the "Multilingual XML" filter that I have heard about and never used, but this turned out to be a dead end. It is language-pair specific, and really not the best option in this case. I was concerned that there might be more files like this in the future involving other language pairs, and I did not want to be bothered with customizing for each possible case.

So I looked a little closer... and noted that this export has the source text copied exactly to the "target". So I concentrated on building a customized XML filter configuration that would just pull the text to translate from between the target tags. A custom configuration of the XML filter was created after populating the tags by excluding the "source" tag content:



That worked, but not well enough. In the screenshot below, the excluded source content is shown with a gray background, but the imported content has a lot of HTML, for which the tags must be protected:


The next step is to do the import again, but this time including an HTML filter after the customized XML filter. In memoQ jargon, this sort of configuration is known as a "cascading filter" - where various filters are sequenced to handle compounded formats. Make sure, however, that the customized XML filter configuration has been saved first:


Then choose that custom configuration when you import the file using Import with Options:


This cascaded configuration can also be saved using the corresponding icon button.


This saved custom cascading filter configuration is available for later use, and like any memoQ "!light resource", it can be exported to other memoQ installations.

The final import looks much better, and the segmentation is also correct now that the HTML tags have been properly filtered:



If you encounter a "special" XML case to translate, the actual format will surely be different, and the specific steps needed may differ somewhat as well. But by breaking the problem down in stages and considering what more might need to be done at each stage to get a workable result with all the non-translatable content protected, you or your technical support associates can almost always build a customized, re-usable import filter in reasonable time, giving you an advantage over those who lack the proper tools and knowledge and ensuring that your client's content can be translated without undue technical risks.

Jun 24, 2017

The other sides of Iceni in Translation


The integration of the online TransPDF service from Iceni in memoQ 8.1 has raised the profile of an interesting company whose product, the Infix PDF Editorhas been reviewed before on this blog. TransPDF is a free service which extracts text content from PDF files, converts it to XLIFF for translation in common translation environments, and then re-integrates the target text from the translated XLIFF to create a PDF file in the target language.

This is a nice thing, though its applicability to my personal work is rather limited, as not many of my clients would be enthusiastic if I were to send PDF files as my translation results. Sometimes that fits, sometimes not. And of course, some have raised the question of whether using this online service is compatible with some non-disclosure restrictions.

I think it's a good thing that Kilgray has provided this integration, and I hope others follow suit, but for the cases where TransPDF doesn't meet the requirements of the job, it is useful to remember Iceni's other options for preparing text for translation.

Translatable XML or marked-up text export
As long as I can remember, the Infix PDF Editor has offered the option to export text on your local computer (avoiding potential non-disclosure agreement violations) so that it can be translated and then re-imported later to make a PDF in the target language. Only the location of this option in the menus has changed: the menu choices for the current version 7 are shown below.



This solution suffers from the same problem as the TransPDF service: not everyone will be happy with the translation in PDF, as this complicates editing a little. However, I find the XML extract very useful to put the content of PDF files into a LiveDocs corpus for reference or term extraction. The fact that Infix also ignores password protection on PDFs is also helpful sometimes.

"Article" export
The Article Tool of  the Iceni Infix PDF Editor enables various text blocks on different pages of a PDF file to be marked, linked and extracted in various translatable formats such as RTF or HTML. The quality of the results varies according to the format.


Once "articles" are defined, they are exported via the command in the File menu:


The RTF export has some problems, as this view in Microsoft Word with the format characters made visible reveals:


However, the Simple HTML export opened in Microsoft Word shows no such troubles (and can be saved in RTF, DOCX or other formats):


Use of the article export feature requires a license for the Infix PDF editor, unlike the XML or marked-up text exports for translation. In demo mode, random characters are replaced by an "X" so that one can see how the function works but not receive any unjust enrichment from it. However, this feature has significant value for the work of translators and is well worth an investment, as the results are typically better than using OCR software on a "live" (text-accessible) PDF file.

But wait... there's more!
Version 7 also has an OCR feature:


I tested it briefly on some scanned Portuguese Help Wanted ads that I'll probably use for a corpus linguistics lesson this summer; the results didn't look too awful all considered. This feature is worth a closer look as time permits, though it is unlikely to replace ABBYY FineReader as my tool of choice for "dead" PDFs.

Feb 1, 2014

The fix is in for PDF charts

Over four years ago, I reviewed Iceni Infix after I began working with it. I'm not as strong a fan as some, because I generally have little enthusiasm for direct editing of PDFs and dealing with frequent problems such as missing unusual fonts and having to play the guess-my-optimum-font-substitution game, but I do find it useful in many situations. I found another one of those today.

A new client of a friend works with a horrible German program to produce reports full of charts. The main body of the text is written in Microsoft Word and is available as a reasonable DOCX file, but the charts are a problem, as they are available only in the specific, oddball tool or PDF format. Nobody wants to deal with that software, really. It is supported by no translation tools vendor I am aware of, and like another example of incompatible German software, Across, it enjoys the obscurity it deserves.

After thinking about the approach needed in this case, I realized that if the graphics could be isolated conveniently on pages, the XML export from the PDF document would contain only information from the graphics. After translation, the format could be touched up with Infix before making bitmap screenshots at an enlargement which would yield decent resolution when sized in  layout. Of course, in projects involving multiple languages the XML files could be used with great convenience.

Selecting and deleting the text on the pages with Iceni Infix is really a no-brainer. The time charge for such work will be quite reasonable. And exporting the XML or marked-up text to translate is also quite straightforward:


The exports can be handled in nearly any CAT tool, so TMS and terminology resources can be put to full use. Or you can edit in a simple, free tool like Notepad++ or an XML-savvy editor.



The screenshot above shows the XML in memoQ. No customization of the default filter is required. Reports from other users who have worked in a similar way indicate that OmegaT and other environments generally have few, if any, problems. In one case there was trouble re-integrating the graphics in a project that also had 50 pages of text, but there may have been other issues I am not aware of in that case.


With the content in the TM, if the chart data are made available in another format, the translations can be transferred quickly to that for even better results. The same approach can be used for a very wide variety of other electronically generated graphic formats (except some of the really insane ones I've seen where the text is broken up; I don't know if Iceni sanitizes such messes or not).

I think this is an approach which can benefit many of us in a variety of projects. It is not really suited for cases of bitmap graphics, but I have other approaches there in which Iceni Infix may also play a useful role and allow CAT integration. Licenses for the tool are quite reasonably priced, and the trial version (in Pro mode) is entirely suited for testing and learning this process.

Sep 23, 2009

Infix PDF Editor: useful for some jobs

A few years ago a Brazilian colleague of mine recommended the Infix PDF Editor from Iceni Technology with great enthusiasm. I tried it at the time, but it was obvious that we had very different PDF types that we worked with and different philosophies of translation, so I concluded that the tool was of little value to the translator except in special situations, and I set it aside. What were those situations? For me, it seemed a decent tool for minor touch-ups of a PDF about to go to print (but I don't do a lot of pre-press proofreading), and for translating simple flyers where work with a TEnT makes little sense it also seems useful. For larger documents I did not see the value, because typeovers of large blocks of text in a PDF document, where possible, run just as contrary to my methods of working as typing over large chunks of text in MS Word or another environment. There is too much risk of content being skipped or deleted, and there is no access to integrated TM and terminology tools.

Recently, however, I discovered another area in which PDF Infix Editor makes a very useful addition to my toolbox. Occasionally I am asked to translate large batches of engineering drawings which have been scanned and reduced to A4 in PDF files. These do not contain editable text. The quality is also usually so miserable that one can forget OCR. Years ago I developed a procedure which the client is fond of: I save the individual pages out as graphics, embed them in an MS Word document and then overlay the portions to translate with text boxes in Word (with appropriate opacity settings and rotation). This can be quite time-consuming.

When another batch of some 120+ of these awful documents arrived recently, I thought of trying the Infix PDF Editor, because I remembered that it had a decent text tool. By overlaying white rectangles followed by text boxes, I was able to accomplish the translation task in the PDF document with less effort than required for my previous work in MS Word. Here's an example of the change:

This saves me time and irritation, and the client will save money. Everybody wins.

Since my initial brush-off of the tool I have also discovered a few other positive aspects for translators that overcome some of my initial objections: text copied and pasted from the Infix interface to MS Word or another environment does not contain all the awful line breaks that one usually sees when copying from Acrobat reader. And text pasted into the Infix interface takes on the properties of the text segment over which it is pasted. This has the advantage that I can copy awful, small text out of a PDF in Infix, paste it into a DOC or RTF file, change the font size to something my old eyes can read and translate it there or in another environment such as DVX or MemoQ (where the font change is unnecessary). Then I can paste chunks back into the original PDF and the fonts and formatting will be largely correct. Mind you, this can be an awful way to work in a large document, but it has advantages in some situations. And I can, in fact, by copying out and pasting in, make use of by favored TEnT methods, with the benefit of translation memory and terminology management.

So now I would say that the Infix PDF Editor is in fact a useful addition to the toolbox for handling PDF, which still must include a first-class OCR tool, such as ABBYY FineReader or OmniPage. PDF is by no means a uniform format but rather a "wrapper" for making many formats accessible to readers, and it will continue to be a time-consuming challenge to translators which must be reflected in appropriate service charges for the extra effort involved.