This wasn't really on Kilgray's plan, but hey - it's now possible, and that makes my life easier. An accidental "feature".
Four years ago, frustrated by the inability of memoQ to import stopword lists obtained from other sources to memoQ, I published a somewhat complex workaround, which I have used in workshops and classes when I teach terminology mining techniques. For years I had suggested that adding and merging such lists be facilitated in some way, because the memoQ stopword list editor really sucks (and still does). Alas, the suggestion was not taken up, so translators of most source languages were left high and dry if they wanted to do term extraction in memoQ and avoid the noise of common, uninteresting words.
Enter memoQ version 8.4... with a lot of very nice improvements in terminology management features, which will be the subject of other posts in the future. I've had a lot of very interesting discussions with the Kilgray team since last autumn, and the directions they've indicated for terminology in memoQ have been very encouraging. The most recent versions (8.3 and 8.4) have delivered on quite a number of those promises.
I have used memoQ's term extraction module since it was first introduced in version 5, but it was really a prototype, not a properly finished tool despite its superiority over many others in a lot of ways. One of its biggest weaknesses was the handling of stopwords (used to filter out unwanted "word noise". It was difficult to build lists that did not already exist, and it was also difficult to add words to the list, because both the editor and the term extraction module allowed only one word to be added at a time. Quite a nuisance.
In memoQ 8.4, however, we can now add any number of selected words in an extraction session to the stopword list. This eliminates my main gripe with the term extraction module. And this afternoon, while I was chatting with Kilgray's Peter Reynolds about what I like about terminology in memoQ 8.4, a remark from him inspired the realization that it is now very easy to create a memoQ stopword list from any old stopword lists for any language.
How? Let me show you with a couple of Dutch stopword lists I pulled off the Internet :-)
I've been collecting stopword lists for friends and colleagues for years; I probably have 40 or 50 languages covered by now. I use these when I teach about AntConc for term extraction, but the manual process of converting these to use in memoQ has simply been too intimidating for most people.
But now we can import and combine these lists easily with a bogus term extraction session!
First I create a project in memoQ, setting the source language to the one for which I want to build or expand a stopword list. The target language does not matter. Then I import the stopword lists into that project as "translation documents".
On the Preparation ribbon in the open project, I then choose Extract Terms and tell the program to use the stopword lists I imported as "translation documents". Some special settings are required for this extraction:
The two areas marked with red boxes are critical. Change all the values there to "1" to ensure that every word is included. Ordinarily, these values are higher, because the term extraction module in memoQ is designed to pick words based on their frequencies, and a typical minimum frequency used is 3 or 4 occurrences. Some stopword lists I have seen include multiple word expressions, but memoQ stopword lists work with single words, so the maximum length in words needs to be one.
Select all the words in the list (by selecting the first entry, scrolling to the bottom and then clicking on the last entry while holding down the Shift key to get everything), and then select the command from the ribbon to add the selected candidates to the stopword list.
But we don't have a Dutch stopword list! No matter:
Just create a new one when the dialog appears!
After the OK button is clicked to create the list, the new list appears with all the selected candidates included. When you close that dialog, be sure to click Yes to save the changes or the words will not be added!
Now my Dutch stopword list is available for term extraction in Dutch documents in the future and will appear in the dropdown menu of the term extraction session's settings dialog when a session is created or restarted. And with the new features in memoQ 8.4, it's a very simple matter to select and add more words to the list in the future, including all "dropped" terms if you want to do that.
More sophisticated use of your new list would include changing the 3-digit codes which are used with stopwords in memoQ to allow certain words to appear at the beginning, in the middle, or at the end of phrases. If anyone is interested in that, they can read about it in my blog post from six years ago. But even without all that, the new stopword lists should be a great help for more efficient term extractions for your source languages in the future.
And, of course, like all memoQ light resources, these lists can be exported and shared with other memoQ users who work with the same source language.
An exploration of language technologies, translation education, practice and politics, ethical market strategies, workflow optimization, resource reviews, controversies, coffee and other topics of possible interest to the language services community and those who associate with it. Service hours: Thursdays, GMT 09:00 to 13:00.
Showing posts with label stopwords. Show all posts
Showing posts with label stopwords. Show all posts
Apr 4, 2018
Aug 3, 2017
"Coming to Terms" workshop materials for terminology mining
I recently put together a two-hour online workshop to teach some practical aspects of terminology mining and the creation and management of stopword lists to filter out unwanted word "noise" and get to interesting specialist terminology faster.
A recording of the talk as well as the slides and a folder of diverse resources usable with a variety of tools are available at this short URL: https://goo.gl/qvwJbf. The TVS recording file can be opened and played by the free TeamViewer application.
The discussion focuses primarily on Laurence Anthony's AntConc and the terminology extraction module of Kilgray's memoQ.
Apr 11, 2014
memoQ: stopwords for term extraction
The recent Kilgray blog post about the "terminology as a service" (TaaS) project reminded me of the considerable unfinished business with the term extraction extraction module introduced three years ago with memoQ 5.0. It's a very useful feature that I apply frequently to my projects, and my prediction years ago that it would not replace SDL's MultiTerm Extract in my workflows was wrong. Overall it proved to be more convenient, and after the shock of discovering that the defective logic of MultiTerm Extract created new German "words" that neither existed nor were in my text sources, I dumped that dodgy tool and stuck to memoQ's extractor. But sometimes its rough edges are irritating, and I wish Kilgray would finally pick up the ball that was dropped after a great start in the game.
One of the major weaknesses (aside from never remembering my changes to the options or my preferred settings for extractions) is the management of stopword lists.
Overall, Kilgray's approach to stopwords is reasonably sophisticated; some time ago I published a rather incomprehensible post in which I demonstrated how the "stopword codes" - those three digit binary numbers which appear stopwords - control whether a word can appear at the start, the end or the middle of a phrase even if it is excluded as a single word. These codes are quite useful in some cases. However, they also complicate the use of stopword data for most users.
memoQ includes only a few stopword lists for a few languages in its shipping configuration - German, English, French and Italian I think. Not even all the user interface languages are included. That's rather sad, because there are quite a few public domain stopword lists available on the Internet. However, most memoQ users have absolutely no idea how to incorporate these in memoQ, and Kilgray offers no features or information that I am aware of to facilitate this process.
I was reminded of this problem when I was asked to discuss terminology mining with masters students at my local university in Portugal. I thought it might be nice for the students to be able to make use of the stopword lists they could find on the Internet for their target languages (Spanish and Portuguese for that group). These lists are typically just text files of single words. When I build my own master stopword list for German a few years ago, I gathered half a dozen or more large, mostly redundant lists for a start. Then I carried out the following steps:
During an extraction session, words can be added to the chosen stopword list (one at a time unfortunately - I've been asking for multiple addition as a productivity measure for 3 years now so far) by selecting a term and Clicking the Add as stopword command or pressing Ctrl+W.
You might be a little confused if you look for words you've added to a stopword list from the term extraction interface. They are not inserted in alphabetical order, but instead at the end of the words starting with a given letter. Thus, for example, the red box in the screen clipping below shows all the words I've added to my previously alphabetized list since I created it:
One of the major weaknesses (aside from never remembering my changes to the options or my preferred settings for extractions) is the management of stopword lists.
Overall, Kilgray's approach to stopwords is reasonably sophisticated; some time ago I published a rather incomprehensible post in which I demonstrated how the "stopword codes" - those three digit binary numbers which appear stopwords - control whether a word can appear at the start, the end or the middle of a phrase even if it is excluded as a single word. These codes are quite useful in some cases. However, they also complicate the use of stopword data for most users.
memoQ includes only a few stopword lists for a few languages in its shipping configuration - German, English, French and Italian I think. Not even all the user interface languages are included. That's rather sad, because there are quite a few public domain stopword lists available on the Internet. However, most memoQ users have absolutely no idea how to incorporate these in memoQ, and Kilgray offers no features or information that I am aware of to facilitate this process.
I was reminded of this problem when I was asked to discuss terminology mining with masters students at my local university in Portugal. I thought it might be nice for the students to be able to make use of the stopword lists they could find on the Internet for their target languages (Spanish and Portuguese for that group). These lists are typically just text files of single words. When I build my own master stopword list for German a few years ago, I gathered half a dozen or more large, mostly redundant lists for a start. Then I carried out the following steps:
- Combine all the stopword lists for a language from various sources into one big text file.
- Open that text file in a spreadsheet program such as Microsoft Excel.
- Sort the list and use the integrated function to eliminate duplicates.
- Fill the number "111" in the second column of the spreadsheet. (This will completely exclude the term from phrases as well; if you want to make individual exceptions according to the scheme I described in an earlier blog post, you can do so now or at any time later after the list has been imported to memoQ.)
- Save the data as tab-delimited Unicode text.
- Open the text file and paste in this XML header, adapting the red parts to your particular list:
<MemoQResource ResourceType="Stopwords" Version="1.0">
<Resource>
<Guid>dc7006ad-8db8-4724-b22d-7acfd600fd9f</Guid>
<FileName>ger#KSL_stopwords-DE.mqres</FileName>
<Name>KSL_stopwords-DE</Name>
<Description>Combined lists from many sources</Description>
<Language>ger</Language>
</Resource>
</MemoQResource>
- Save the file, change the file extension to MQRES, and import the file as a stopword list in the memoQ Resource Console.
During an extraction session, words can be added to the chosen stopword list (one at a time unfortunately - I've been asking for multiple addition as a productivity measure for 3 years now so far) by selecting a term and Clicking the Add as stopword command or pressing Ctrl+W.
You might be a little confused if you look for words you've added to a stopword list from the term extraction interface. They are not inserted in alphabetical order, but instead at the end of the words starting with a given letter. Thus, for example, the red box in the screen clipping below shows all the words I've added to my previously alphabetized list since I created it:
Jan 29, 2014
Finding resources on Kilgray's Language Terminal
Kilgray’s online platform for translation, Language Terminal at https://www.languageterminal.com/, may be a game-changer in many ways. Not only does it offer affordable, on-demand memoQ translation server capacity for small teams on demand, it provides free InDesign server availability to users of any tool for converting InDesign formats to XLIFF and PDF for translation and review, back-up features fully integrated with recent versions of memoQ, some evolving project management and invoicing tools and a growing library of light resources shared by users. This post discusses how to find and use these resources, which can be useful in all supported versions of memoQ.
Accessing your account
The user menus of Language Terminal can be accessed in two ways: in a web browser from the URL above or from the link on your memoQ Dashboard.
If you are not already a Language Terminal user, a free account can be set up in just a few minutes.
Looking for resources
The current user interface for finding resources on Language Terminal is confusing to some users. The Resource menu link in the orange navigation bar shows a list of resources you have uploaded yourself to Language Terminal. The dropdown list indicated by the arrow filters your own resources. To find resources from other people, click the Advanced Search button.
There is nothing “advanced” about this search. It simply allows you to use four fields to find resources which are publicly available on the site. Be careful of your selection criteria for language as some resources (like auto-translation rules) are not language-specific by definition even if they might have been created for use with a particular language.
The result of the search for English stopword resources to be used in terminology extraction to filter out “noise” words (like prepositions, pronouns, articles and common vocabulary) looked like this at the time I performed the search:
Download the resources you want by clicking on their names in the Resource column. The shared library of filters, QA profiles, auto-translation rules, stopword lists and more on Language Terminal continues to grow. Why not contribute something yourself?
In any case, Language Terminal is a useful place to archive one’s valuable light resources, such as segmentation rules developed over time with great effort, and these are not shared with others unless you specifically release them. Given the occasional unfortunate “disappearances” of light resources known to occur with some memoQ upgrades, this is a very useful backup option to have, and it would be nice if future integration of Language Terminal and memoQ were to facilitate more complete, automated resource backups from desktop systems.
Accessing your account
The user menus of Language Terminal can be accessed in two ways: in a web browser from the URL above or from the link on your memoQ Dashboard.
If you are not already a Language Terminal user, a free account can be set up in just a few minutes.
Looking for resources
The current user interface for finding resources on Language Terminal is confusing to some users. The Resource menu link in the orange navigation bar shows a list of resources you have uploaded yourself to Language Terminal. The dropdown list indicated by the arrow filters your own resources. To find resources from other people, click the Advanced Search button.
There is nothing “advanced” about this search. It simply allows you to use four fields to find resources which are publicly available on the site. Be careful of your selection criteria for language as some resources (like auto-translation rules) are not language-specific by definition even if they might have been created for use with a particular language.
The result of the search for English stopword resources to be used in terminology extraction to filter out “noise” words (like prepositions, pronouns, articles and common vocabulary) looked like this at the time I performed the search:
Download the resources you want by clicking on their names in the Resource column. The shared library of filters, QA profiles, auto-translation rules, stopword lists and more on Language Terminal continues to grow. Why not contribute something yourself?
In any case, Language Terminal is a useful place to archive one’s valuable light resources, such as segmentation rules developed over time with great effort, and these are not shared with others unless you specifically release them. Given the occasional unfortunate “disappearances” of light resources known to occur with some memoQ upgrades, this is a very useful backup option to have, and it would be nice if future integration of Language Terminal and memoQ were to facilitate more complete, automated resource backups from desktop systems.
Jan 7, 2012
Understanding memoQ's term extraction stopword codes
Recently I shared a link to a small stopword list for a minor language, which I had set up as a memoQ resource for a friend, and another translator questioned why I had coded the stopwords as I did. My answer was truthful: no good reason. I had simply copied the practice in Kilgray's default files for other languages. As I looked further into discussions of term extraction and stopwords on the memoQ Yahoogroups list, I realized that I was not the only one who had a hard time getting a clear picture of how things actually work. So I decided to learn by experiment.
First I created a stopword list with nonsense words having every possible coding combination. A memoQ stopword list is a test file with an XML header and *.mqres extension, with a structure that looks like this:
A "1" means yes, "0" means no. So "011" means
First I created a stopword list with nonsense words having every possible coding combination. A memoQ stopword list is a test file with an XML header and *.mqres extension, with a structure that looks like this:
<memoqresource resourcetype="Stopwords" version="1.0">The entries in the stopword list (here the nonsense words gak through bla) are each followed by a tab and a three digit binary code. The first digit of this code controls whether a phrase is excluded from the list of candidates if it begins with this entry. (Kilgray calls this "blocks as first".) The second digit of the code controls whether a phrase is excluded if the entry occurs within it (not at the beginning nor at the end, Kilgray calls this "blocks inside"). The third digit controls whether a phrase is excluded if the entry occurs at its end ("blocks as last").
<resource>
<guid>2b077cde-8c10-4ee1-86db-14eb42f010cc</guid>
<filename>KSL_test-stopwords_EN.mqres</filename>
<name>KSL_test-stopwords-EN</name>
<description>For testing only</description>
<language>eng</language>
</resource>
</memoqresource>
gak 111
unga 101
munga 011
kunga 110
fra 000
blu 100
bly 001
bla 010
A "1" means yes, "0" means no. So "011" means
- allowed at the start of the phrase,
- not allowed inside the phrase
- not allowed at the end of a phrase
My test file contained the sentence
The quick brown fox jumped over the lazy dog
repeated four times in three blocks for each test stopword, with the stopword substituted at the beginning, inside and at the end of "over the lazy dog":
The quick brown fox jumped unga the lazy dog. The quick brown fox jumped unga the lazy dog. The quick brown fox jumped unga the lazy dog. The quick brown fox jumped unga the lazy dog.The quick brown fox jumped over unga lazy dog. The quick brown fox jumped over unga lazy dog. The quick brown fox jumped over unga lazy dog. The quick brown fox jumped over unga lazy dog.The quick brown fox jumped over the lazy unga. The quick brown fox jumped over the lazy unga. The quick brown fox jumped over the lazy unga. The quick brown fox jumped over the lazy unga.
After the term extraction, the following four-word phrases from the text chunk of interest were found with the stopwords:
fra The lazy dog
bly The lazy dog
bla The lazy dog
munga The lazy dog
over unga lazy dog
over fra lazy dog
over blu lazy dog
over bly lazy dog
over The lazy kunga
over The lazy fra
over The lazy blu
over The lazy bla
All these occurrences follow the defined rules as you can see from the stopword list above. None of the stopwords occurred singly in the extraction candidates, of course. So entering "000" as the code for a stopword will exclude that stopword alone but not in any phrase.
How is this relevant in practice? In English, for example, words like in, the and first are uninteresting by themselves and belong in a stopword list. But a phrase containing them, like "in the first instance" might indeed be of interest. In cases like that, the proper code for these stopwords might be "001" or "101" (allowing inside in both cases, at the beginning as well in the first case) might be appropriate. These are matters of judgment that will differ for each language. One user commented that he finds it more useful to be very restrictive in the extraction ("111") and add phrases during the actual translation, and I am inclined to follow this practice as well. Where one discovers exceptions, the stopword rules can always be edited in various places in memoQ.
Subscribe to:
Posts (Atom)












