Showing posts with label counting text. Show all posts
Showing posts with label counting text. Show all posts

Oct 21, 2012

Put OCR in Your Business Model

This article originally appeared on an online translators portal four years ago and was long overdue for removal there. Here is an update.

*****
Optical character recognition (OCR) software is discussed often online and at translators' events, usually in the context of how to deal with PDF files. Hector Calabia, Peter Linton and others have made a useful technical contributions on this subject in articles and forums and at various conferences. However, it is useful to consider OCR software in a broader translation business context. Document conversion is often very useful for translation purposes and greatly facilitates automated quality checks of the draft, for example, but OCR can also generate additional income for your business and reduce quotation risk.
OCR for translation
There are a number of programs available for this purpose, and which one is best for your purposes may depend on the language combinations you deal with and other factors. For years now I have used Abbyy FineReader, because years ago it gave the best test results for the particular set of European languages one of our clients offered. It is also relatively inexpensive (I paid about 100 euros for FineReader 11) and easy to use.

Many OCR conversions of TIFF, JPEG and PDF documents which I receive from agencies are difficult to use for translation purposes and require significant modification - if they can be used at all. Particularly in cases where TM tools are to be used or target texts differ significantly in length (especially when they are longer) there may be problems. The best ways to avoid these problems are
  • avoid automatic settings for OCR conversions; use zone definitions instead
  • avoid saving the converted texts with full formatting in most cases
  • use a suitable post-OCR workflow to clean up the converted document by joining broken sentences, removing superfluous characters, fixing conversion errors, etc.
If the idea of doing individual zone definitions on each page of a 100 page document is intimidating, take heart. Programs such as Abbyy FineReader often allow you to define layout templates, speeding up the work considerably. One translator I know became so skilled at the use of these OCR templates and was so good with his conversions that agencies hire him just to do high-quality OCR work for them. Which brings me to….

OCR as an income-generating activity for the translator or agency
Hardcopy, scanned documents, faxes and PDF documents generally require more work for translators than electronically editable documents and require different, sometimes more fallible quality control measures than a typical workflow for a translator using original electronic documents in a translation memory system. If no conversion is performed, it is more time-consuming to check terminology or use concordances during the translation, and it is also unfortunately too easy for eyes to skip over bits of text. Under time pressure this can lead to very serious problems. Even with conversion, the OCR text requires careful checking against the original document to identify and correct any errors introduced (and there will be some at times with even the best OCR software). So it is not at all unreasonable for a translator to charge a higher rate for dealing with hardcopy, scanned documents, faxes and PDF documents.

There are a number of ways to incorporate these higher charges into your business model. The two obvious ways are a premium (surcharged) word/line/page rate and hourly service charges. I usually offer both options to my clients, with the word/line rate surcharge representing the “fixed” rate and the hourly rate the “flexible” rate where I make an non-binding estimate and they may end up paying more or less according to the actual effort. For pure OCR conversion jobs where I am not doing the translating, I charge a typical proofreading rate or a bit more, because I go through the entire document and see that it is correctly formatted for translation work and that obvious errors are fixed (i.e. basic spellcheck, etc.).

Sometimes I hear that “the client doesn’t want to pay for that”. Well, that’s OK, too. The client has the option of doing the work and doing it right and saving me the effort. The recognition that there is additional effort involved and that this effort should be compensated is important. But usually there is a way to sugar-coat the "bitter" cost pill, and this is where your marketing savvy comes into play. Some win-win arguments you might present include:
  • the availability of an editable source text the client can use for future versions;
  • the ability to create TM resources using the OCR text (which can save time/money later);
  • potentially better quality assurance, especially with tight deadlines. 
Returning a clean, nicely formatted OCR of the source document is often good "advertising". End clients may appreciate how this saves time and allows them to use the original text in a variety of ways (attorneys may like to quote arguments from the opposing side, and copy/paste beats retyping). Discriminating agencies may recognize your skill at creating documents that don’t go crazy when edited (because of screwy text boxes, bad font definitions and other format errors) and offer you more work. If your language pair is in low demand or is very competitive, this may be one more way of distinguishing yourself from the pack.
I got started doing OCR work and charging for it after suffering through the conversion of several long PDF documents by more manual methods. I finally wised up, bought FineReader and started to use it with most of the hardcopy, scanned documents, faxes and PDF documents I received simply because it enabled me to use my TM tools and do better quality checks. I started sending the cleaner-looking source texts converted with OCR along with the target text translations, and soon I started getting requests for paid OCR work. A number of my agency clients then began to buy OCR tols and use them with varying degrees of success. Even if they do all the conversion work, I still win if they do it right, because I save time for what I enjoy more – the translation.

OCR as tool for quotation
Some people I know still haven’t learned to do a high-quality OCR (or they don’t care to), but they still use the software effectively in a very important area of their business: quotation and risk limitation.

There are lots of good tools out there for text counting, which is important to many methods of costing and time planning in the translation business. Some people even still do it manually, which, though time consuming, is not a bad way of checking the numbers from an electronic estimate. A number of factors can result in text counts being too low – embedded objects, such Excel tables or PowerPoint slides in a Microsoft Word documents, or graphics with text - or even too high (as is the case with at least one CAT tool counting RTF and MS Word files). Keep using whichever method you prefer - I won't try to persuade you that any one approach is best. I use a number of methods myself.

When translating larger documents, however, or documents with a complex structure, it is often useful to have a “sanity check” for your text counts. On a number of occasions I have received translation jobs from agency clients where the text count was given a X words, where in fact there were quite a few more words embedded in Excel objects, bitmap graphics, Visio charts, etc. which had not been measured by the method used. In a few cases these clients had to take a loss on the job after giving a fixed price bid to the end client. Using OCR to check your estimates can prevent such an unfortunate scenario.

To do this, print the document (whatever it is) to a PDF file. Then run the PDF file through an OCR program with automatic settings (to save time – you don’t need to translate this OCR). Save the text and count it. There will probably be a bit more text due to headers or footers or perhaps garbage from graphics, but the results should be close to your other estimate. (You can always subtract an appropriate factor for the text count in headers and footers to improve your OCR estimate.) If there is a major deviation, this is a clear sign that you should take a much closer look at the document(s) before quoting the job.

Searchable scanned documents
Another use I have found for OCR in recent years is creating searchable "text-on-image" documents from scanned PDFs, TIFF files and other bitmap formats. Although I have used these searchable PDFs mostly for reference while I work (searching for bits of text while viewing the original, unadulterated context) and supplied them to clients on only a few occasions, the potential for an additional value-added service is fairly obvious in this case.

Conclusion
OCR software is an essential tool for the work of many translators today, even more so than CAT software in many cases. Not just a tool for recovering “lost” electronic documents or making legacy typed material more accessible for translation work, it also offers possibilities for generating additional projects and income, differentiating one’s services and reducing risks when quoting large jobs. Key features of whatever OCR you choose should include the ability to select text areas for conversion and to determine their sequence in the converted text (using user-defined zones). Various options for saving the converted text (full page format, limited text formatting and no formatting) are also very helpful. Most important of all, though, is a good quality-checking workflow for your OCR documents (possibly including formatting) to avoid difficulties in the translation process and ensure that your work has a polished, professional appearance.

OCR software is another good tool for improving your visibility with clients and making your work processes easier in an age when many archiving and ERP systems are focused on the retention of PDF documents or TIFFs and even actively discourage saving original formats. The major providers of this software often have free, functional demonstration versions to use before making a purchase decision. Try several options and choose the best one for you. You won’t be sorry.

Oct 22, 2011

Compatibility workflows with the memoQ Translator Pro edition (Part 1)

Yesterday I had the privilege to present the first of a series of workshops intended to convey my ideas for small-scale outsourcing management with the version of memoQ typically purchased by freelance translators. The participants were project managers at a translation agency that has begun to test the waters for using memoQ to overcome long-term compatibility issues between Trados versions in their accustomed workflows. I have been supporting them occasionally as a consultant over the past two years to deal with sticky issues of text encoding and translators who can't follow directions while working with a disturbing range of tools they often haven't mastered. It's been fun, and I've learned a lot from the infinite human capacity to instinctively ferret out the weaknesses of software and processes.

So I decided to put together a personal overview of the compatibility interfaces for the memoQ Translator Pro edition and my own thoughts on best practice and share it with my colleagues. I wanted to avoid burying everyone in technical detail but instead present the material in a way that most anyone can understand and apply. I don't believe in silly notions such as expecting the average intelligent user to learn and remember the use of regular expressions and other arcana that I, despite four decades of IT experience, continue to struggle with myself too often.

The presentation
  • referred to memoQ version 5 Translator Pro edition
  • focused on facilitating project workflows with different platforms rather than actual translation
  • was intended for anyone outsourcing on a small scale for a single target language in a project (multiple target languages require the memoQ Project Manager or server editions)
It was delivered in two parts in a three hour period, with a long break for coffee, chat, snacks and checking e-mail or testing ideas learned in the first part. Participants were provided with screenshots of the main application screens and did not sit in front of computers but engaged in the discussion. The actual project experience and understanding of the participants was polled at appropriate intervals to make sure that the delivery was relevant and the information was understood and able to be applied.

The goal was to achieve an understanding of memoQ as a central platform for
  • translation project input - files, translation memories, terminology and reference material
  • format conversion to facilitate work with different translation environment techniques and tools
  • translation
  • editing and quality assurance
  • creation of deliverable target files and other resources such as term lists, special review formats and commentaries
memoQ is "compatible" with
  • SDL Trados in all versions (though it is important to choose the right compatibility workflow!)
  • Star Transit
  • quite a number of other commercial and Open Source translation environment tools
  • various content management systems (CMS)
  • translators who decline to use any tool other than a word processor
  • and of course memoQ!
so except in the case of projects requiring live, direct work on a third-party translation server platform (such as one from SDL), some reasonable workflow can be found to collaborate with almost anyone using other tools.

memoQ is sort of like the Swiss Army knife of translation environment tools when it comes to compatibility. Only better. Some say it's more compatible with Trados than Trados. And in many cases they're right.

Output formats for translators
memoQ can prepare content for translation in
  • optimized formats for memoQ users
  • Trados-compatible bilingual DOC files
  • XLIFF, a standard used by many environments
  • RTF tables for those without special translation tools or for others to review, comment and answer questions using only a word processor
TM & terminology data
memoQ reads translation memory data in TMX and delimited text formats and outputs it to TMX. Term data is read in the same formats as TM data but output only to delimited text formats and a particular SDL Trados MultiTerm XML format.

memoQ can also integrate with external termbases, TM sources and machine translation engines.

I sometimes think of memoQ as the hub of a wheel with translators, reviewers and customers working with many different environments as the "spokes".

Basic project management steps with memoQ

These typically involve:

1. Reading in the data after it is properly prepared
  • files to translate in whatever source format
  • translation memory data or reference corpora
  • terminology data
  • special segmentation rules (SRX files and segmentation exceptions) or other configuration data for optimized workflows
In this step it is important to choose the best method´s of data import and the appropriate filter or combination of filters. In memoQ, filters can be cascaded to convert and protect sensitive data as tags. Thus HTML and placeholder tokens contained in cells of an Excel file might be protected by "chaining" an HTML filter and a custom filter using regular expressions after the usual filter for Microsoft's Excel format.

2. Analyzing the data


Many options are available here, including the weighting of tags to compensate the extra effort involved with complex formats and determining internal similarities in a text (aka "homogeneity" or "fuzzy repetitions") to facilitate better project planning.

3. Extracting terminology (particularly useful for large projects for one or more translators)


4. Preparing and exporting files for translators


Projects can also be sent to translators using memoQ as handoff packages or complete backups with all attached TMs, termbases and corpora.

5. Receiving and re-importing translated content


6. Review, QA and feedback workflows



7. Generating target files and other information for delivery, final statistics


Recommendations for best practices in choosing formats for translators, reviewers and others using a variety of tools will be covered in the second part of this summary. Those interested in a live presentation or relevant materials are welcome to contact me privately.

Aug 7, 2011

Format surcharging in translation

One of the strategies I pursued early on in my career as a commercial translator was to equip myself with tools able to handle a great variety of source formats and then learn (mostly by nail-biting troubleshooting for my projects and those of others) to cope with the exceptions, typical problems for a given format and the insoluble and unexpected disasters of files with Hotel California workflows in various CAT tools. At a time when a majority of translators in my language combination were perceived as incorrigible technophobes and many agencies struggled to deal with the technical intricacies of IT and data exchange in translation, this was a path to appallingly rapid business growth.

How times have changed. Or haven't. Translation environment tools have evolved more in the past decade than some industry critics will admit, making more formats reliably accessible and data exchange between environments less likely to trigger calls to the local suicide hotline. Lots and lots of translators now "have" Trados or some other tool, or the tools have them. But it's a tenuous relationship in the majority of cases. More tenuous than many realize until suddenly the simply formatted Word file won't "save as target" from the TagEditor horror chamber, or they get the idea of actually using "integrated" terminology features in some tools and learn a new definition of despair.

My Luddite friends are right to speak of the complexities that can lurk in even the "simplest" translation environments, but I believe that dealing with these complexities in simple, rational ways and sharing the information will go farther toward simplifying our lives and enhancing our professional status than desperately clinging to outmoded ways that will increasingly restrict the flow of business. However, as we adopt new ways, we must think more about these complexities, the real effort involved and how to offset this effort in simple, economic terms.

Take file formats, for example. Over the years I have heard many suggestions from colleagues for how to charge different formats. Many of these seem rather arbitrary and not necessarily sensible to me, such as all the myriad ways people charge for the translation of PowerPoint slides. Work with PowerPoint can be simple and straightforward or a hideous nightmare requiring complex, creative combined strategies of pre-translation repair, filtering, dissection and reassembly and much more. Microsoft Word is seen as a simple format, but add a hundred footnotes, cross-references, formulae, "Word Art", embedded Excel and Visio abjects objects, a rainbow of colors for coding and some massive, uncompressed images for good measure and you often face quite a challenge.

How do you deal with such complexity, plan for it in your schedule and charge it in a manner which is fair to the persons performing the service and those paying for it? The answer is not easy, but the typical response to the question - ignore it and charge "usual" rates or hocus pocus some percentage mark-up - is not very satisfactory.

Discrete, pre- or post-translation tasks such as OCR, format repairs, extraction and re-embedding of translatable objects or the transfer of these to separate documents for "key pair" translation are all fairly easy to handle in an acceptable, transparent way with hourly fees for the effort. When I deal with such matters, I occasionally provide the client with detailed work instructions for how to go about performing these tasks cleanly to "save" money with the caveat that if it isn't right, the work will be re-done and charged.

I have yet to come up with a standard way of coping with files that are simply so big that they choke the tools I use or tie up my resources for an hour while exporting a translated file. Here, technological aikido is usually the most effective strategy: at various stages in the past decade, for example, I have converted graphics-laden RTF or DOC files to HTML, TTX and now DOCX to minimize troubles and speed up processing. Once I have worked out a way of avoiding those big resource tie-ups (often at the cost of hours or days of thought and experimentation), I feel I don't have to consider the charge issue (but of course I'm really wrong). However, the risks of format failure are so great in my experience that "round trip" tests must  be performed to ensure that once a translation has taken place the results can be transferred to their deliverable form without much ado. If I forget to do this under pressure, I very often regret it. Think of round trip workflow testing for the files you translate as a possibly life-saving pre-flight safety check. You might not die in a crash, but business relationships will.

One issue I have meditated on for a very long time and mentioned at intervals in translators' forums without finding a reasonable answer is that of markup tags. Many tools didn't even used to count them; at one point Déjà Vu was the only one I was aware of that did. The best answers that colleagues seemed to offer for tag-laden documents, which inevitably require more work and frequently lead to stability problems, was to "charge more" or "run like Hell". Both good answers, really, but lacking in the quantitative rigor my background in science leads me to prefer.

The solution arrived somewhat unexpectedly with the beta version of memoQ 5. At first I thought that SDL Trados Studio 2009 offers no solution here, but with that tool the context in which you view the statistics is important. Look at the analysis under "Reports", not "Files". Any counting tool that reports tag frequency can be used to calculate this solution, if need be with a spreadsheet if the factors cannot be added the tool's internal statistics for words or characters.



The solution is obvious and really wasn't far from the discussion which has taken place over the years: simple word or character weighting for the tags. However, it was not until I saw the fields in the new memoQ count statistics window that I really began to think about what those factors should be.

In SDL Studio 2009 the tag statistics are found in the analysis under "Reports" as mentioned and look like this (thank you to Paul Filkin for the technical update and the graphic):



I thought about my own experience with tags in files over the years and the actual extra effort of inserting them for formatting at the beginning and end of segments or somewhere inline. For reasons I won't try to explain in an over-long post, I figure that a single tag costs me the effort of about half a word, or given the average word length in my source language, about 3.5 characters. So I put in "0.5" words and "3.5" characters as the weight factors in memoQ, and my count statistics are increased to compensate for the additional effort involved.

Now you may disagree on the appropriate factor, saying it should be more or perhaps less. That's OK. I consider this a matter open to negotiation with clients. The important thing for me is that we have a quantitative basis for discussion and negotiation which anyone may check. It's important that this and other issues relevant to project planning and compensation be brought out of the closet and discussed rationally. Not just to get "fair" compensation, but to educate those involved in the processes about the effort involved and to set more realistic project goals as well.

For some of the OCR trash that clients produce ineptly and try to foist off on translators as source files, this "tag penalty" may encourage better practice or at least offset the effort of using Dave Turner's CodeZapper and other methods to clean up the mess. (However basic structural problems caused by automated settings in OCR tools will never be overcome this way.)

In any case, this is a technique which I hope will inspire discussion and study to find its best application in various environments. And I do hope to see widespread adoption of such options in modern translation environment tools to further offset the grief occasionally encountered in modern translation.




Nov 15, 2010

Counting text in Microsoft Word 2010 (and 2007 apparently)

A few weeks ago I had a call from a new client regarding a small job, and when I was asked about my rates, I tried to explain briefly how translators in Germany often calculate these and how he might estimate costs himself. Unfortunately, the explanation got "stuck" at the time, because we were using different versions of Microsoft Office. I was still enjoying the old Office 2003 package with a few upgrades to enable me to deal with Office 2007 files, but he had a shiny new computer with the latest MS Office 2010. When I referred to the "Tools" menu ("Extras" in German) and said to find the word count function under it, he informed me that this menu didn't exist in that version. Score another one for Microsoft in its 24-year effort to keep its users of Word teetering on the brink of frustrated insanity as the interface cards get remixed and the rules changed with every new version.

Last Friday I finally got my long-awaited new laptop to replace my utterly decrepit Toshiba with its troublesome keyboard that my local repair shop was unable or unwilling to replace. With it I got the latest MS Office version, so I too have made the Great Leap Forward into the abyss of the new interface. And although it may be a very obvious thing for many readers, I want to take this opportunity to show graphically how to find the new word count function in Microsoft Word 2010. If you are using an older version of Word and need to explain this to a client who has the latest version, perhaps this will help:


Addendum: Another alternative in Word 2007 & 2010, which was kindly pointed out by Victor Dewsbery in the comments for this post, is to use the function at the left of the bottom bar of the Word document window:
Double-clicking the count on the bar will open the word count dialog with the full statistics.

It is also interesting to note that text in text boxes is apparently counted, which was not the case in my old 2003 version of Microsoft Word. Here I created a small text file with 12 words distributed in the ordinary document body flow, a table and a text box. Then I selected three words in the table. The count shows both the selection (3 words) and the total (12 words):