Showing posts with label terminology. Show all posts
Showing posts with label terminology. Show all posts

Jun 14, 2024

Substack in #xl8 & #l10n

 As related in my last post, I've been testing the Substack platform as an alternative means for distributing and managing the kind of information I've published for many years now on this blog, professional sites and social media, course authoring platforms, YouTube, Xitter, Mastodon and probably a few other channels I've forgotten. And I'm quite surprised to find any expectations I had to be exceeded at this point. On the whole I can do a better job of creating accessible tutorials and other reference information there, and there is far less effort involved with maintenance, style sheet fiddling and whatnot.

All my various projects there can be found here on my profile page.

Currently, these are the
Other plans include The Diary and Letters of Charles Berry Senior, a US Civil War veteran from Yorkshire, England, who participated in Sherman's march to the sea and who is my great-great grandfather. Some of his records exist in very deteriorated form in a university archive in the United States and can be found online, but much is missing, and I am in possession of the complete transcript prepared by one of the man's daughters more than 100 years ago when the family feared the information would be lost to the forces of material decay. I'll be preparing clean text from her handwritten record (the typescript done by a cousin about 50 years ago was lost and probably contains more errors) and probably an audio reading. There's some hard stuff there, as well as some surprises and interesting lessons in how our world has changed in the past 160 years.

And at some point I'll probably share some of my culinary obsessions as well as what life was like traveling on the other side of the Iron Curtain sans papers or in dingy Paris bookshops and refugee hotels 40 years ago and more.

This past week, I've done a big blitz on memoQ LiveDocs, for which there are still another dozen or so drafts to be finished, and a lot of stuff from my CAT tools resource online class from last year should be appearing there in updated form in due course. There isn't a lot of translation-related activity that I've found on Substack, at least not for the technical side of things, but a lot of historians and authors I follow are very present there, so I'm hoping the better half of my translation technologist friends will join the Substack party at some point.

Most of my new text, video and teaching content will appear on those Substack channels. It's simply far easier to manage, and there you won't have the same RSS headaches you might have here. And the damned editor of Google Blogger just keeps accumulating bugs I don't have to cope with in Substack. And don't get me started on bloody Wordpress!

I hope to see you in my Substack channels soon! I think there will be something like a memoQ QA course there before the summer is over....

Nov 25, 2023

Book review: "Terminology Extraction for Translation and Interpretation Made Easy"


A few months ago, I received a pre-release copy of this book as a courtesy from the author, terminologist Uwe Muegge, with a request to give a quick language check to the English used by its native German author. As I expected, there wasn't much to complain about, because he has lived in the US for a long time and taught at university there as well as been involved in important corporate roles. I was particularly pleased by his disciplined style of writing, the plain, consistent English of the text and the overall clarity of the presentation. Anyone with good basic English skills should have no difficulty understanding and applying the material.

At the time I read the draft, I was completely focused on language use and style, but I found his approach and suggestions interesting, so I looked forward to "field testing" and direct comparisons with my usual approach to terminology mining with that feature in memoQ. About a day's worth of tests shows very interesting potential for applying the ChatGPT section of the book and also made the context and relevance of the other two sections clearer, I will discuss those sections first before getting to the part that interests me the most.

Uwe presents three approaches:
I wouldn't really call these three approaches alternatives as the book does, because all three operate in very different ways and are fit for different purposes. That didn't register fully in my mind when I was in "editor mode", although the first part of the book made the differences, advantages and disadvantages clear enough, but as soon as I began using each of the sites, the differences were quite apparent as were the similarities to more familiar tools like memoQ's term extraction module.

Wordlist from Webcorp is functionally similar to Laurence Anthony's AntConc or memoQ's term extraction. It's essentially useful for getting frequency lists of words, but the inability to use my own stopword lists for filtering out uninteresting common vocabulary makes me prefer my accustomed desktop tools. However, the barriers to first acquaintance or use are lower than for AntConc or memoQ, so this would probably be a better classroom tool for introducing concepts of word frequencies and identifying possible useful terminology on that basis.

OneClick Terms was interesting to me mostly because friends and acquaintances in academia talk about Sketch Engine a lot. The results were similar to what I get with memoQ, including similar multiword "trash terms". I found the feature for term extraction from bilingual texts particularly interesting, and the fact that it can work well on the TMX files distributed by the Directorate-General for Translation (DGT) of the European Commission suggests that it could be an efficient tool for building glossaries to support translation referencing EU legislation, for example, though I expect only slight advantages over my usual routine with memoQ. These advantages are not worth the monthly subscription fee to me. However, for purposes of teaching and comparison, the inclusion of this platform in the book is helpful. I see more value for academic institutions and those rare large volume translation companies that do a lot of work with EU resources. 

ChatGPT was an interesting surprise. I have a very low opinion of its use as a writing tool (mediocre on its best day, clumsy and boring in nearly all its output) or for regex composition (hopelessly incompetent for what I need, and anything it does right for regex is newbie stuff for which I need no support). However, as a terminology research tool I have found excellent potential, though formatting the results can be problematic.

My testing was done with ChatGPT 3.5, not a Professional subscription with access to version 4.0. However, I am sorely tempted to try the subscription version to see if it is able to handle some formatting instructions (avoiding unnecessary capitalization) more efficiently. No matter how carefully I try to stipulate no default capitalization of the first letter of every expression, I inevitably have to repeat the instruction after a list of improperly capitalized candidate terms is created.

I keep an e-book copy of Uwe's book in the Kindle app on my laptop, so I can simply copy and paste his suggested prompts, then add whatever additional instructions I want.

The prompt
Please examine the text below carefully and list words or expressions which may be difficult to translate, but when writing the list, do not capitalize any words or expressions which don't require capitalization.

is too long, and only the part marked red is executed correctly, but this follow-up prompt will fix the capitalization in the list:

Please re-examine that text and this time when writing the list, do not capitalize any words or expressions which do not require capitalization.

Further tests involved suggesting translations for the expressions, with or without a translated text and building tables with example sentences:

Other prompt variations, for example to write terms bold in the example sentences, worked without complications.

What about the quality of the selections? Well, I used memoQ's term extraction module on the same text I submitted to ChatGPT for term extraction in order to compare something with which I am quite familiar with this new process. 

memoQ identified a few terms based on frequency, which ChatGPT ignored, but these were arguably terms that a qualified specialist would have known anyway. And ChatGPT did a superior job of selecting multi-word expressions with no "noise". It also selected some very relevant single-occurrence phrases which might be expected to arise more in later, similar texts.

Split screen review of memoQ extraction vs. ChatGPT results

The split-screenshot is an intermediate result from one of my many tests. The overlayed red box was intended to show a conversation partner the limits of ChatGPT's "alphabetizing skill", and the capitalization of the German is not correct after a prompt to correct the capitalization of adjectives misfired. It is not always trivial to get formatting exactly as I want it. However, looking at the results of each program side-by-side like this showed me that ChatGPT had in fact identified the nearly all the most relevant single words and phrases in my text. And for other texts with dates or citation formats, these were also collected by ChatGPT as "relevant terms", giving me an indication of what legislation I might want to use as reference documents and what auto-translation rules might also be helpful.

I also found that the split view as above helped me to work my way through the noise in the memoQ term candidate list much faster and make decisions about which terms to accept. The terms of interest found in memoQ but not selected by ChatGPT were few enough that I am not at all tempted to suggest people follow my traditional approach with the memoQ term extraction module and skip the work with ChatGPT.

My preferred approach would be to do a quick screening in ChatGPT, import the results into a provisional (?) term base and then, as time permits, use that resource in a memoQ term extraction to populate the target fields in the extraction grid. With those populated terms in place, I think the review of the remaining candidates would proceed much more efficiently.

All in all, I found Uwe's book to be a useful reference for teaching and for my personal work; it is one of the few texts I have seen on LLM use which is sober and modest enough in its claims that I was inspired to test them. The sale price is also well within anyone's means: about $10 for the e-book and $16 for the paperback on Amazon. For the "term curious" without access to professional grade tools, it's a great place to get started building better glossaries and for more seasoned wordworkers it offers interesting, probably useful suggestions.

The book is available HERE from Amazon.

Jan 8, 2021

memoQ Courses, Resources & Consulting at Translation Tribulations Tech

The new online school offers a variety of resources for new and experienced users of desktop and server editions

For many years now, I have advocated for better professional education for users of translation process support software at every level. I have tested curriculum delivery platforms, better ways of making information more accessible to those who need it, and more. In a limited scope, this has been a successful effort.

My greatest hope in these efforts was to encourage professional associations, technology providers and universities to do better by their clientele. I would judge the success there as mixed, at best. The wind of change discussed, for many of them, could fill one's sails... for a voyage off the edge of their flat Earth. Their reluctance to provide even minimal indexes for navigating copious video content is simply baffling, as an example.

Even with the current pandemic, I have seen little progress, though that may be as much for reasons such as those which kept me largely silent last year. It's hard to think about doing things better when you have to ask honestly which of the people you care for will be lost because of the refusal of so many national governments to do so.

In any case, I've always been one to advocate more personal involvement. If a person says they're hungry, give them food and listen to their stories. Cash may not be the answer. The courses, consulting and resources offered through my license of the Teachable platform will cover much of issues and assistance for which I have been an advocate in the translation sector for two decades. I also hope to involve other language service educators to offer their unique and valuable approaches in this venue. This is not to compete with any existing associations or companies, but rather to continue to show them how we can all work together to help users develop the competence and confidence so often needed and not found.

This, like all of us, is a work in progress. Check out Translation Tribulations Tech (here, or by clicking the school graphic at the top) and see if anything there provides missing elements for your professional toolkit.

Some of the initial offerings include:
Additional courses, consulting and tools for
  • regular expressions as an aid for translation of patterned information like currency expressions, dates, legal citations, coded information, etc.
  • better source document segmentation in projects
  • memoQ server basics for collaborating groups and small companies
  • memoQ and other technology for legal translation
will be available soon.

This platform provides a long-needed mechanism for providing more detailed learning assistance than I have enjoyed with this blog and my YouTube channel, and future publication habits on my part will reflect that. I'm excited about many ideas for moving ahead in quick and quicker steps with memoQ and so many other resources that many of us depend on for professional relief and productivity.


Jan 6, 2021

Tweeting away....

Got up this morning to not altogether unexpected good news that the Empire of MAGATs has fallen:

 

Yeah. Life is starting to feel normal again despite the usual continued death and destruction. But what does one do with babies if not put them in cages? 

A course announcement for terminology users in memoQ (i.e. any sensible user):
... which leads one to ask: How do I get there? Well, try this:

Sep 26, 2019

10 Tips to Term Base Mastery in memoQ! (online course)

Note: the pilot phase for this training course has passed, free enrollment has been closed, and the content is being revised and expanded for re-release soon... available courses can be seen at my online teaching site: https://transtrib-tech.teachable.com/
In the past few years I have done a number of long webinars in English and German to help translators and those involved in translation processes using the memoQ environment work more effectively with terminology. These are available on my YouTube channel (subscribe!), and I think all of them have extensive hotlinked indexes to enable viewers to skip to exactly the parts that are relevant to them. A playlist of the terminology tutorial videos in English is available here.

I've also written quite a few blog posts - big and small - teaching various aspects of terminology handling for translation with or without memoQ. These can be found with the search function on the left side of this blog or using the rather sumptuous keyword list.

But sometimes just a few little things can get you rather far, rather quickly toward the goal of using terminology more effectively in memoQ, and it isn't always easy to find those tidbits in the hours of video or the mass of blog posts (now approaching 1000). So I'm trying a new teaching format, inspired in part by my old memoQuickie blog posts and past tutorial books. I have created a free course using the Teachable platform, which I find easier to use than Moodle (I have a server on my domain that I use for mentoring projects), Udemy and other tools I've looked at over the years.

This new course - "memoQuickies: On Better Terms with memoQ! 10 Tips toward Term Base Mastery" - is currently designed to give you one tip on using memoQ term bases or related functions each day for 10 days. Much of the content is currently shared as an e-mail message, but all the released content can be viewed in the online course at any time, and some tips may have additional information or resources, such as videos or relevant links, practice files, quality assurance profiles or custom keyboard settings you can import to your memoQ installation.

These are the tips (in sequence) that are part of this first course version:
  1. Setting Default Term Bases for New Terms
  2. Importing and Exporting Terms in Microsoft Excel Files
  3. Getting a Grip on Term Entry Properties in memoQ
  4. "Fixing" Term Base Default Properties
  5. Changing the Properties of Many Term Entries in a Term Base
  6. Sharing and Updating Term Bases with Google Sheets
  7. Sending New Terms to Only a Specific Ranked Term Base
  8. Succeeding with Term QA
  9. Fixing Terminology in a Translation Memory
  10. Mining Words with memoQ
There is also a summary webinar recorded to go over the 10 tips and provide additional information.
I have a number of courses which have been developed (and may or may not be publicly visible depending on when you read this) and others under development in which I try to tie together the many learning resources available for various professional translation technology subjects, because I think this approach may offer the most flexibility and likelihood of success in communicating necessary skills and knowledge to an audience wider than I can serve with the hours available for consulting and training in my often too busy days.

I would also like to thank the professional colleagues and clients who have provided so much (often unsolicited) support to enable me to focus more on helping translators, other translation project participants and translation consumers work more effectively and reduce the frustrations too often experienced with technology.


Jan 14, 2019

Specialist terminology taxonomies from Cologne Technical University

Click and thou shalt go there!

Early in the last decade when I lived near Düsseldorf and began translating full time, the nearby technical university in Cologne had an excellent terminology studies program run by Prof. Klaus-Dirk Schmitz, who also had a long history in Saarbrücken back in my exchange student days there. I had the pleasure of meeting this gentleman at various professional events for Passolo (before it was swallowed by SDL) or other occasions, and I remain impressed by the professional qualities of some of the colleagues he helped to educate. At some point he or one of his students pointed me to an interesting online collection of specialist terminologies created by students at the university as part of their degree work. While student work must be viewed carefully, on the whole I found these collections to be of better quality than quite a few put together by "professionals", and their structured taxonomies were also interesting to people like me who enjoy such things. And occasionally the terminologies were rather helpful for certain technical topics I translate.

But over the years I simply forgot about them for the most part, and when they did come to mind I assumed that the old MultiTerm engine used to handle the data on the site would no longer work. That latter assumption may be partly correct; I found the collection again, noted that the most recent addition to the term library was a bit over a decade ago and that the search functions don't seem to work with Chrome, though I am able to browse the structured taxonomies without difficulty.


Looking through the list of term collections, I saw one that would be particularly useful for a current personal effort: beekeeping. One of my projects for the year ahead is to add some hives to the garden to see if I can improve some of the vegetable, fruit and nut yields. A local Portuguese beekeeper and I have been trading poultry, and he kindly provided me with a copy of his thesis on apiculture and offered assistance to get me started. So I am reading up on the subject in several languages, thinking to put together a good terminology to make cross-referencing the concepts between English, German and Portuguese a little easier.

One thing I never tried to do before was to extract data from the FH Köln (Cologne Technical University) site into any sort of terminology management tool. I don't think they were ever intended to be used that way, and at the time most of the collections were put together, translation environment tools were much less widely used by professionals and university study programs than they are today. But after a little thought and experimentation, mining the pages proved to be quite simple.

Here's how I did it:
  1. Opened a collection of interest and expanded the folder tree for a particular language completely, then selected and copied all the text in that frame:

  2. Pasted the copied content as plain text (no formatting) into Microsoft Word. The numerical codes were followed directly by the text entries.
  3. Removed parentheses by searching and replacing with nothing.
  4. Inserted a tab between the number codes using search and replace with wildcards (regex of a sort):

  5. Switched to the other language in the term collection and repeated steps 1 through 4.
  6. Transferred the contents to Excel (various ways to do this).
  7. Imported the Excel file with the specialist terms into a term base in my translation environment tool of choice.


Jan 11, 2019

Do you know Anki?

It began with a short DM last night, which I misunderstood at first:


Oh? Gábor must have read my mind. Just recently someone introduced me to Fotografia de Aves em Portugal, a public Facebook group for bird photography in Portugal, and I was thinking about making some sort of flashcard set to learn the bird names in Portuguese and English and maybe to do the same for all the mushrooms that I encounter at the quinta and out in the fields hunting. In fact, when I looked up the description of the desktop computer program Anki and its iOS mobile app companion, I realized that this is really what I have hoped to find for quite a long time for various learning tasks.

The computer app developed by Damien Elmes and its online server and synchronization site AnkiWeb are free; the charge for the iOS mobile app helps to support the development of all platforms. There is also a free compatible Android app by a different author. I like the idea of being able to coordinate my "learning decks" between devices and access them from anywhere. And a quick look at the import features tells me that it's not hard to send the content of some of my personal study term bases in memoQ to this application.

I assume this is an app he came across in his quest to learn Chinese; in fact, he did say that it's a tool that helps give one a fighting chance to learn all the myriad characters needed for basic literacy. And further research on my part showed that this is a popular tool for review in medical school and many other areas.

I downloaded and installed the app; the initial view was a little puzzling:


But within a few minutes I got my bearings and downloaded a few of the many "shared decks" online to familiarize myself with how the app works:


Pretty simple, really. Thank you, Gábor!

Dec 29, 2018

memoQ Terminology Extraction and Management

Recent versions of memoQ (8.4+) have seen quite a few significant improvements in recording and managing significant terminology in translation and review projects. These include:
  • Easier inclusion of context examples for use (though this means that term information like source should be placed in the definition field so it is not accidentally lost)
  • Microsoft Excel import/export capabilities which include forbidden terminology marking with red text - very handy for term review workflows with colleagues and clients!
  • Improved stopword list management generally, and the inclusion of new basic stopword lists for Spanish, Hungarian, Portuguese and Russian
  • Prefix merging and hiding for extracted terms
  • Improved features for graphics in term entries - more formats and better portability
Since the introduction of direct keyboard shortcuts for writing to the first nine ranked term bases in a memoQ project (as part of the keyboard shortcuts overhaul in version 7.8), memoQ has offered perhaps the most powerful and flexible integrated term management capabilities of any translation environment despite some persistent shortcomings in its somewhat dated and rigid term model. But although I appreciate the ability of some other tools to create customized data structures that may better reflect sophisticated needs, nothing I have seen beats the ease of use and simple power of memoQ-managed terminology in practical, everyday project use.

An important part of that use throughout my nearly two decades of activity as a commercial translator has been the ability to examine collections of documents - including but not limited to those I am supposed to translate - to identify significant subject matter terminology in order to clarify these expressions with clients or coordinate their consistent translations with members of a project team. The introduction of the terminology extraction features in memoQ version 5 long ago was a significant boost to my personal productivity, but that prototype module remained unimproved for quite a long time, posing significant usability barriers for the average user.

Within the past year, those barriers have largely fallen, though sometimes in ways that may not be immediately obvious. And now practical examples to make the exploration of terminology more accessible to everyone have good ground in which to take root. So in two recent webinars, I shared my approach - in German and in English - to how I apply terminology extraction in various client projects or to assist colleagues. The German talk included some of the general advice on term management in memoQ which I shared in my talk last spring, Getting on Better Terms with memoQ. That talk included a discussion of term extraction (aka "term mining"), but more details are available here:


Due to unforeseen circumstances, I didn't make it to the office (where my notes were) to deliver the talk, so I forgot to show the convenience of access to the memoQ concordance search of translation memories and LiveDocs corpora during term extraction, which often greatly facilitates the identification of possible translations for a term candidate in an extraction session. This was covered in the German talk.

All my recent webinar recordings - and shorter videos, like playing multiple term bases in memoQ to best advantage - are best viewed directly on YouTube rather than in the embedded frames on my blog pages. This is because all of them since earlier in 2018 include time indexes that make it easier to navigate the content and review specific points rather than listen to long stretches of video and search for a long time to find some little thing. this is really quite a simple thing to do as I pointed out in a blog post earlier this year, and it's really a shame that more of the often useful video content produced by individuals, associations and commercial companies to help translators is not indexed this way to make it more useful for learning.

There is still work to be done to improve term management and extraction in memoQ, of course. Some low-hanging fruit here might be expanded access to the memoQ web search feature in the term extraction as well as in other modules; this need can, of course, be covered very well by excellent third-party tools such as Michael Farrell's IntelliWebSearch. And the memoQ Concordance search is long overdue for an overhaul to allow proper filtering of concordance hits (by source, metadata, etc.), more targeted exploration of collocation proximities and more. But my observations of the progress made by the memoQ planning and development team in the past year give me confidence that many good things are ahead, and perhaps not so far away.

Dec 4, 2018

Optimizing memoQ terminology extraction

On December 28, 2018 from 2:00 to 3:30 pm Lisbon time (3:00 to 4:30 pm CET, 9:00 to 10:30 am EST), I'll be giving a talk on terminology extraction in the latest version of memoQ. Recent versions of this tool have included many improvements to its terminology features, and it's time for an update on how to get the most out of the term extraction features of memoQ among other things.

Topics to be covered include the creation of new stopword lists or the extension of existing ones, customer-, project- or topic-specific stopword lists, criteria for corpora, term mining strategies and the subsequent maintenance and use of term bases in projects. Participants will be equipped with all the information needed to use this memoQ feature confidently, reliably and profitably in their professional work.

The webinar is free, but registration is required. To register, go to:
https://zoom.us/meeting/register/cfd1a47cd5c54114d746f627e8486654

The same presentation (more or less) will be held in German on December 21 at the same time for those who prefer to hear and discuss the topic in that language.

Dec 3, 2018

Terminologieextraktion mit memoQ: die neuesten Möglichkeiten


Am 21. Dezember um 15:00 Uhr bis 16:30 Uhr MEZ findet wieder eine deutschsprachige memoQ-Schulung online statt. Thema: Optimierung der Terminologieextraktion. Der Vortrag bietet eine Übersicht der Möglichkeiten für effizientes Arbeiten mit dem Extraktionsmodul für Terminologie in memoQ. Von der Neuerstellung bzw. Erweiterung der Stoppwortlisten, kunden-, projekt- oder themenspezifische Stoppwortlisten, Korpuskriterien und Extraktionsstrategien bis zu der anschließenden Pflege der Terminologiedatenbanken und dem Einsatz im Projekt werden Sie mit den notwendigen Informationen gerüstet, diese Funktion bei Ihren professionellen Tätigkeiten sicher, zuverlässig und gewinnbringend einzusetzen. Teilnahme ist kostenlos aber registrierungspflichtig: https://zoom.us/meeting/register/d68e024c63ad506f7c24e00bf0acd2b8 Ein inhaltsgleicher Vortrag in englischer Sprache findet eine Woche (am 28.12.2018) später statt: https://zoom.us/meeting/register/cfd1a47cd5c54114d746f627e8486654

Nov 26, 2018

memoQ as an instrument of bowdlerization



Recently in a memoQ user forum, someone asked if the environment could be used to ensure that certain words would be kept out of a target text. 


The question wasn't clearly understood at first: some thought the asker wanted to ensure that certain words were not translated and suggested the use of non-translatable lists, others pointed out that the memoQ term bases have an option to mark certain term translations as forbidden, which would be indicated by a black color in the Translation Results list, for example:


But no... what was wanted was indeed a monolingual list to ensure that the target text did not contain certain words, regardless of what the source text said.

This is indeed possible in memoQ:


The red box in the screenshot above marks a word on such a forbidden list, which is in an English to English (US) revision project I set up to clean up Mark Twain's 1601 and make it fit for teaching in Sunday school. The presence of a no-no is indicated by a little lightning bolt icon for a QA warning. Actually running a QA check for terminology gave the following result, a list of forbidden expressions (with some alternatives suggested):


How is this done? With an ordinary memoQ term base. One can make a term entry only for a single language - the target language in this instance - and mark the Forbidden term checkbox on the Usage tab.

If there is no source term (or in my example above, where the source and target language are variants of the same language), there will be nothing shown in the translation results list, but if there is a termbase entry marked forbidden which is found on the target side, then a warning will be displayed in the translation and editing grid if the QA profile currently selected for the project includes terminology. (If the QA profile does not include appropriate term checking, no warning will be displayed for segments in the grid, nor will the forbidden words be indicated when QA is run. The Default QA profile does include term checking, but I use a lot of different profiles focused on fewer issues, such as just tag checking, so I have to pay attention to this detail.)


Building such a list word-by-word with manual entries is tedious. So it's probably easier to import your monolingual list of words to avoid from a text file or an Excel sheet, and then in the memoQ Term Base Editor, select a range of terms (like all of them) and set the desired forbidden status for the entire selection:


Bulk changes to any term properties are possible this way as I showed some time ago in a short video tutorial.

The critical setting for the scenario described here is marked with a red box in the QA profile below:


As you can see, the source text can also be checked for forbidden expressions if that option is also selected.

I have created a dedicated termbase to track forbidden expressions in three languages. Lists - monolingual or otherwise - of forbidden expressions can be maintained in one or more term bases or the expressions can be kept in a more ordinary term base. If barring certain common expressions in one or more languages is important to you, it might be convenient to maintain this information in a dedicated term base.

And for recent versions of memoQ (8.4 and later), make sure that the term bases you want to use for quality assurance are marked on the Term bases page of Project home. Here it's not a good idea to use your larger translation term bases for QA, because these may result in rather large numbers of false positives. Optimum term properties settings for translation and quality assurance are often not the same.


Thus memoQ can be used as a powerful tool to avoid embarrassment from an unfortunate choice of words and to adapt the target language to fit a particular audience better. Thinking back to the time, years ago, when a friend who ran the translation department at a conservative German company nearly got fired for writing that a certain software operation could be performed at the touch of a penis (a Freudian slip after he and some other translators were joking about how sick they were of a certain phrase in the user documentation) and remembering the sensitivity to terms I have seen with some clients (at the same software company, the term FAQ was banned, because executive management was afraid that it might be pronounced like fuck), I can see how this somewhat unusual approach to terminology in memoQ could be a job-saver for some.

May 25, 2018

Getting on Better Terms with memoQ

The pre-event webinar for the terminology workshop in Amsterdam on June 30th was held yesterday; for those who missed it, the recording is here, with a few call-outs added toward the end to make it easier to find information on other matters mentioned:


On the whole, I think the new presentation platform I'm using Zoom – works well. I was particularly happy to discover when my Internet router died suddenly and mysteriously about 50 minutes into the talk that the recording was not lost, and when the talk resumed a few minutes later with a different router connection, a recording of that part was made in a separate folder, so the parts could be joined later in a video editor without much ado.

I've held perhaps a dozen online meetings with clients since I licensed Zoom recently, and I'm very pleased with its flexibility, even though the many options have tripped me up in embarrassing ways a few times. So, time permitting, I'll try to do at least one talk like this each month on some aspect of translation technology in the interests of promoting better practice. The next will be held on June 21st and will cover various PDF workflows using iceni technology. Suggestions for later presentations are welcome.

The talk yesterday on terminology in memoQ was just a quick (ha ha - one hour) overview of possibilities; much more detail on these matters will be provided in the Amsterdam workshop and summer courses at Universidade Nova de Lisboa, which are open to the public at very reasonable rates (about €130 for 25 hours of instruction). There will be a lot more in this topic area in the future in various venues; right now there are some very interesting developments afoot with memoQ and other matters at Kilgray, and other providers also have good things in the works. So stay tuned.

May 21, 2018

Best Practices in Translation Technology: summer course in Lisbon July 16-21

As usual each year, the summer school at Universidade Nova de Lisboa is offering quite a variety of inexpensive, excellent intensive courses, including some for the practice of translation. This year includes a reprise of last year's Best Practices in Translation Technology from July 16th to 21st, with some different topics and approaches.

Centre for English, Translation and Anglo-Portuguese Studies

The course will be taught by the same team as last year – yours truly, Marco Neves and David Hardisty – and cover the following areas:
  • Good translation workflows.
  • Using voice recognition in translation.
  • Using machine translation in a humane, intelligent way.
  • Using checklists to improve communication in translation.
  • Using glossaries, bilingual texts and other references in multiplatform environments.
  • Good practices for using terminology and reference texts in the target language.
  • Planning and creating lists for auto-translation rules and the basics of regular expressions for filters.

Some knowledge of the memoQ translation environment and translation experience are required.

The course is offered in the evening from 6 pm to 10 pm Monday (July 16th) through Friday (July 20th), with a Saturday (July 21st) session for review and exams from 9 am to 2 pm. This allows free days to explore Lisbon and the surrounding region and get to know Portugal and its culture.

Tuition costs for the general public are €130 for the 25 hours of instruction. The university certainly can't be accused of price-gouging :-) Summer course registration instructions are here (currently available only in Portuguese; I'm not sure if/when an English version will be available, but the instructors can be contacted for assistance if necessary).

Two other courses offered this summer at Uni Nova with similar schedules and cost are: Introduction to memoQ (taught by David and Marco – a good place to get a solid grounding in memoQ prior to the Best Practices course) from  July 9–14, 2018 and Translation Project Management Tools from September 3–8, 2018.

All courses are taught in English and Portuguese in a mix suitable for the participants in the individual courses.

May 1, 2018

All-round Translator Terminology Workshop Pre-event Webinar

Link to registration for the webinar

As previously announced, on June 30th in Amsterdam, the All-round Translator (ART) is offering the workshop "Coming to Terms: Mining & Management" covering a range of practical topics for applied corpus linguistics, optimizing terminology management and efficient sharing of terms in teams. This technical workshop will cover a range of tools and techniques as described on the ART event page.

A month before the Amsterdam workshop I will be presenting a free webinar offering an overview of some of the material planned for June as well as related topics with a particular focus on one of the tools I use most - memoQ - with highlights of recent improvements in its terminology features with memoQ versions 8.3 and 8.4.

This webinar is open to anyone interested regardless of whether or not they plan to attend the June workshop. The talk will use Zoom, which I adopted for remote teaching of corporate clients and others because of its greater versatility and superior recording facilities compared to my old favorite, TeamViewer. (Technically it's also a "meeting", not a "webinar" in Zoom-speak, but that's a distinction without a difference for people who don't feel up to 50 simultaneous speakers and 10,000 viewers.) The platform is free for participants to use, and if I'm not mistaken, a web browser can also be used, though interaction is more limited via that medium. (Don't ask me how, I am still gathering experience with this tool and its myriad options.)

The presentation (approximately one hour, starting at 4 pm Central European Time = 3 pm Lisbon time on May 24th) is free, but registration is required: the link for that is here.

Apr 4, 2018

New in memoQ 8.4: easy stopword list creation!

This wasn't really on Kilgray's plan, but hey - it's now possible, and that makes my life easier. An accidental "feature".

Four years ago, frustrated by the inability of memoQ to import stopword lists obtained from other sources to memoQ, I published a somewhat complex workaround, which I have used in workshops and classes when I teach terminology mining techniques. For years I had suggested that adding and merging such lists be facilitated in some way, because the memoQ stopword list editor really sucks (and still does). Alas, the suggestion was not taken up, so translators of most source languages were left high and dry if they wanted to do term extraction in memoQ and avoid the noise of common, uninteresting words.

Enter memoQ version 8.4... with a lot of very nice improvements in terminology management features, which will be the subject of other posts in the future. I've had a lot of very interesting discussions with the Kilgray team since last autumn, and the directions they've indicated for terminology in memoQ have been very encouraging. The most recent versions (8.3 and 8.4) have delivered on quite a number of those promises.

I have used memoQ's term extraction module since it was first introduced in version 5, but it was really a prototype, not a properly finished tool despite its superiority over many others in a lot of ways. One of its biggest weaknesses was the handling of stopwords (used to filter out unwanted "word noise". It was difficult to build lists that did not already exist, and it was also difficult to add words to the list, because both the editor and the term extraction module allowed only one word to be added at a time. Quite a nuisance.

In memoQ 8.4, however, we can now add any number of selected words in an extraction session to the stopword list. This eliminates my main gripe with the term extraction module. And this afternoon, while I was chatting with Kilgray's Peter Reynolds about what I like about terminology in memoQ 8.4, a remark from him inspired the realization that it is now very easy to create a memoQ stopword list from any old stopword lists for any language.

How? Let me show you with a couple of Dutch stopword lists I pulled off the Internet :-)


I've been collecting stopword lists for friends and colleagues for years; I probably have 40 or 50 languages covered by now. I use these when I teach about AntConc for term extraction, but the manual process of converting these to use in memoQ has simply been too intimidating for most people.

But now we can import and combine these lists easily with a bogus term extraction session!

First I create a project in memoQ, setting the source language to the one for which I want to build or expand a stopword list. The target language does not matter. Then I import the stopword lists into that project as "translation documents".


On the Preparation ribbon in the open project, I then choose Extract Terms and tell the program to use the stopword lists I imported as "translation documents". Some special settings are required for this extraction:


The two areas marked with red boxes are critical. Change all the values there to "1" to ensure that every word is included. Ordinarily, these values are higher, because the term extraction module in memoQ is designed to pick words based on their frequencies, and a typical minimum frequency used is 3 or 4 occurrences. Some stopword lists I have seen include multiple word expressions, but memoQ stopword lists work with single words, so the maximum length in words needs to be one.


Select all the words in the list (by selecting the first entry, scrolling to the bottom and then clicking on the last entry while holding down the Shift key to get everything), and then select the command from the ribbon to add the selected candidates to the stopword list.

But we don't have a Dutch stopword list! No matter:


Just create a new one when the dialog appears!


After the OK button is clicked to create the list, the new list appears with all the selected candidates included. When you close that dialog, be sure to click Yes to save the changes or the words will not be added!


Now my Dutch stopword list is available for term extraction in Dutch documents in the future and will appear in the dropdown menu of the term extraction session's settings dialog when a session is created or restarted. And with the new features in memoQ 8.4, it's a very simple matter to select and add more words to the list in the future, including all "dropped" terms if you want to do that.

More sophisticated use of your new list would include changing the 3-digit codes which are used with stopwords in memoQ to allow certain words to appear at the beginning, in the middle, or at the end of phrases. If anyone is interested in that, they can read about it in my blog post from six years ago. But even without all that, the new stopword lists should be a great help for more efficient term extractions for your source languages in the future.

And, of course, like all memoQ light resources, these lists can be exported and shared with other memoQ users who work with the same source language.

Mar 14, 2018

Come to Terms in Amsterdam, June 30th


At end of June this year I'll be doing an expanded, in-person reboot of my occasional terminology workshop with new material and workflows for those who want to do more to control quality and improve communicative vocabularies in interpreting, translation and review projects.

Space is limited at the All-Round Translator event, but I hope you can join us to learn about
  • Better teamwork through timely terminology sharing
  • Faster, more effective discovery of frequently occurring specialist terminology
  • Better access to critical terminology in many environments
  • More efficient and accurate QA for terminology
  • More accurate, efficient and fault-tolerant term use when translating with memoQ
  • Greater flexibility to meet client terminology needs
The Early Bird rate for the workshop is €99 + VAT until the end of April, €120 + VAT  thereafter.

The content is applicable to work with many translation environments, but some segments will share particular tips for maximum productivity using the unbeatably practical memoQ environments.

Nov 26, 2017

MS Word Macros to Speed up Translation-Related Terminology Research

Guest post by Tanya Harvey Ciampi, English translator (DE/FR/IT>EN)

Is your terminology research slowing you down?


When we translate Microsoft Word documents, we often find ourselves having to leave Word to look up terms online, for example in monolingual dictionaries for definitions, in bilingual dictionaries or translation memory databases for translations, on specific reputable websites (such as newspaper websites) to double-check usage or frequency of use, or on clients’ own multilingual websites to check how certain terms have been translated in the past to ensure consistent use of terminology.

This sort of research involves switching to a browser, copying and pasting or retyping our term into a search box, possibly adding specific search criteria, and finally launching a search: all that typing and clicking can be time-consuming and easily cause us to become lost among the many windows opened.

Macros to the rescue!
This is where macros come in. A macro is essentially a short sequence of commands that automates repetitive tasks. Macros cost nothing to create and can be tweaked to do exactly what you need them to do, based on your specific language combinations and favourite online terminology resources, providing these lend themselves to this sort of querying.

How do macros work?
A macros consists of code, which you simply need to copy and paste into the Macros section of Word. That done, you then need to assign an icon to the macro and add it to your toolbar to launch the macro with a single click every time you need it. If you wish, you may also assign a specific key combination to the macro (for example CTRL plus a key of your choice) so that you can launch the macro from your keyboard, too.

From now on, when translating a text in Word, all you need to do is place your cursor on a word that you wish to look up and click on the corresponding icon in your toolbar (or use the assigned key combination) to launch the search. That’s all there is to it!

A few examples of macros and what they can do for you:

SCENARIO: Imagine...SOLUTION... with a single click!
...you need to look up a term in the bilingual dictionaries www.leo.org and www.dict.cc but this requires opening your browser, browsing to both dictionaries separately and pasting in or retyping your search term on each website... quite time-consuming! A macro to search both dictionaries at once taking your word from MS Word and inserting it automatically in both dictionaries for you... with a single click from within Word.
(This macro can be adapted to all sorts and any number of websites)
What this macro does essentially is launch a Google search from within Word, adding specific search criteria, in this case:
“your search term” inurl:leo.org or inurl:dict.cc

...you wish to run a search in the online translation memory database www.linguee.com (or linguee.de, linguee.fr, linguee.it etc.) to check how other translators have translated a certain term or expression. A macro to search Linguee taking your word from MS Word and inserting it directly in the Linguee search engine with a single click from within Word.
This macro produces a list of source- and target-language sentences containing your search term along with context.

...you are translating a text and need to check how a particular expression is used. You decide to search reputable sources such as high-quality newspapers to check usage and/or frequency of use of a specific term or expression. Where do you look? A macro to search specific newspaper websites which you consider reputable sources from within Word.
(This macro can be adapted to all sorts and any number of websites.)
This macro essentially launches a Google search from within Word, adding specific search criteria to target a specific website, for example:
“your search term” inurl:guardian.co.uk

...you are translating for a company that has a multilingual website and you need to check how a specific term has been translated in the past. A macro to search for the term on a specific multilingual website from within Word.
This macro can be extended to cover various related multilingual websites. In banking, for example, these might include the following:
www.ubs.com
www.credit-suisse.com
www.raiffeisen.ch
This macro essentially launches a Google search from within Word, adding specific search criteria, for example:
“your search term” site:www.ubs.com or site:www.credit-suisse.com or site:www.raiffeisen.ch

...you are translating a text and can't find an appropriate translation of an expression or technical term in any dictionary. A macro to search for your term on a large multilingual website such as that of the European Union from within Word. This macro targets the section of the EU website containing translations side by side (“parallel texts”) on the same page, saving you precious time.
This macro essentially launches a Google search from within Word, adding specific search criteria, for example:
“your search term” inurl:eur-lex.europa.eu
Once you have opened a page on the EU website, all you need to do is specify your target language under “Multilingual display” to view source and target language side by side.

See a couple of these macros in action:
https://www.youtube.com/watch?v=XlvBLgJPaFk

These and more macros are available for free at https://www.facebook.com/groups/TranslatorsSwitzerland/

The macros themselves are written by a translator with translators' needs in mind and can be adapted to your specific requirements.

Macros may also be created to automate the web-based terminology research techniques for translators found at
http://www.multilingual.ch/Search_Interfaces.htm
... reducing them, too, to a single click in Word!

The original search techniques on which these macros are based were featured in the book entitled “Google Hacks” (“Hack #19: Google Interface for Translators”) by Tara Calishain, Rael Dornfest

*******

Tanya Harvey Ciampi, Dipl. DOZ (Zurich)
English translator (DE/FR/IT>EN)
6673 Maggia, Switzerland, www.multilingual.ch

Tanya grew up in Buckinghamshire, England, and went on to study in Zurich, where she obtained her diploma in translation. She now lives in the Ticino, the Italian-speaking region of Switzerland, where she works as an English translator (from Italian, German and French) and proofreader.

Aug 3, 2017

"Coming to Terms" workshop materials for terminology mining



I recently put together a two-hour online workshop to teach some practical aspects of terminology mining and the creation and management of stopword lists to filter out unwanted word "noise" and get to interesting specialist terminology faster.

A recording of the talk as well as the slides and a folder of diverse resources usable with a variety of tools are available at this short URL: https://goo.gl/qvwJbf. The TVS recording file can be opened and played by the free TeamViewer application.

The discussion focuses primarily on Laurence Anthony's AntConc and the terminology extraction module of Kilgray's memoQ.