OmegaWiki aims to include all words in all languages and provide both lexical, terminological and ontological information. As the discussion of what makes a language is an endless one, those languages that are included in the ISO-639 codes are the ones that are supported.
Having chosen the ISO-639-3 to start of with has proven a great start. It did however not provide the granularity needed to categorize words to their linguistic entity. How to deal with languages that are written in several scripts, how to deal with regional differences? This is what this iteration of the ISO-639 standard does not deal with.
The implication is that this standard on its own does not suffice. By combining the data with other standards with other codes it is possible to provide more granularity, but how to deal with dialects like Westfries, that is spoken in the area where I grew up?
As the OmegaWiki project was evolving and getting traction, I got into contact with Debbie Garside. She is heading Geolang an organisation that has been preparing for a long time the next iteration of the ISO-639 standard, the ISO-639-6. The aim is to include at least 25.000 linguistic entities in a hierarchical structure. Adopting this data would allow OmegaWiki to better achieve its aim; include all words of all languages.
When a standard is published there is a prescribed period in which the public is invited to comment on a standard. So far this has been done using e-mail. Experience shows that when the amount of subject is too big, e-mail is not a tool to cope. Geolang had explored the option of using Wiki technology before, this sadly did not lead to the right synergy. In OmegaWiki however, there was both an active interest in language standards, it included not only the Wiki methodology, it even allows for the inclusion of the data in a true hierarchical way.
By publishing the data in a wiki, in essence everybody with an interest in orthographies and dialects is invited to comment, modify and add to the hierarchical data. To make this into a standard, there will be a need to assess the community generated data and assert the validity of the information provided. This is where the World Language Documentation Centre will play its role. As its name implies, it documents languages and it will do so in the broadest sense of the word. Obviously an organisation like this will only function well when it is as an organisation an inclusive organisation. The make-up of the current board reflects many specialities that make up linguistics and the language industry.
It is with a fair amount of satisfaction that I can announce that Sean Burke, one of the volunteers of OmegaWiki has imported the first batch of the ISO-DIS-639-6 data in time for the inaugural meeting of the World Language Documentation Centre. Both OmegaWiki and the WLDC will rely on collaboration, to get the necessary work done. Our challenge will be to provide the infrastructure and the minimal organisation to start and sustain our projects.
With the inaugural meeting, the WLDC it is proclaimed to the world that as an organisation the WLDC is ready for business. With the first data available in OmegaWiki, the first request to the world to collaborate on the languages that are spoken the orthographies that are written goes out. it is the start of acquiring the meta data that helps us understand the data that is already out there and consequently make from all this data information because we will become better able to parse the data.
Thanks,
GerardM
Showing posts with label wiki. Show all posts
Showing posts with label wiki. Show all posts
Friday, May 11, 2007
Wednesday, February 28, 2007
What is a Word?
First, let me make it perfectly clear that this is a discussion that has raged for centuries. I know full well that everybody has his or her own opinion on this matter, and that I am not going to resolve this issue today. This is an overview and a bit of personal opinion, as relates to online dictionaries.
Intuitively, we all know the answer. A word is a unit of language conveying some meaning. But how do we decide what is a real word? We look in a dictionary, of course. What do we do if we're writing a dictionary?
We are caught between cataloging what is "right" (prescriptivism) and what is actually done (descriptivism). The pendulum has lately swung towards descriptivism, and I would say that there are some good reasons for that trend. The language that is spoken on the streets is not the same language that is written in academia. Somebody learning a language may genuinely need help sorting out the less proper terms in it.
Take, for instance, colloquialisms such as "irrespective" and "humongous", and all the phrases that have gotten squished together into amalgams like "gotcha" and "woulda". Most people would readily agree that these words do not belong in a college thesis paper.
What is the scrupulous lexicographer to do? Fortunately, it is not a strict either-or question, especially in a work not substantially limited by size. In an electronic resource, we can put them in, anyway. To satisfy the formal sorts, the perscriptivists, we can then place a prominent usage note in the entry, explaining just why a writer might wish to use caution with the term: ginormous is a colloquial term, regarded by many to be something less than a proper word. Thus, the reader is both informed and cautioned.
That's fine for most of the slang and jargon, but we have another problem. People keep making up new words. My sister in law coined the term "muskaroon" to mean generically any small, furry creature that scurries past too quickly to identify. Squirrels, chipmunks, gophers, and presumably rabbits would all qualify. So we have a unit of language with a symbol and a meaning. The trouble is, if you walked up to people on the street and inquired whether there were muskaroons in the area, nobody would be able to answer who hadn't talked lately to my sister in law, and that is a small minority of people, indeed.
The test here is usage. Can we demonstrate that the word is in common use? Now, depending on the character of the dictionary, we can define the rules various ways. Was it used by so many independent sources? Did anybody important (such as Shakespeare or a prominent academic journal) publish the word?
Generally, we also try to find and present examples of the term in what is called "running text". That means that it is in a paragraph, and isn't only used as somebody's nickname, say. The edge is still a fuzzy one. Are the citations in traditional print sources, such and books and journals, or are they sprinkled in a couple blogs and forums? Was the word used in only one limited context, or in a variety of sources and over a period of years? These sorts of tests can help to weed out many of the more questionable entries. At some point, though, it may yet come down to a judgment call, if not on whether a word is real, then on how to apply the rules. In these cases, I advise the users of a dictionary to bring a healthy dose of skepticism with them, to recall that even dictionaries are not infallible, and to trust at the very least that these decisions are made by real people who care for the project.
If, knowing all that, you find you don't like the way "they" are running the place, you are invited to do a better job.
Intuitively, we all know the answer. A word is a unit of language conveying some meaning. But how do we decide what is a real word? We look in a dictionary, of course. What do we do if we're writing a dictionary?
We are caught between cataloging what is "right" (prescriptivism) and what is actually done (descriptivism). The pendulum has lately swung towards descriptivism, and I would say that there are some good reasons for that trend. The language that is spoken on the streets is not the same language that is written in academia. Somebody learning a language may genuinely need help sorting out the less proper terms in it.
Take, for instance, colloquialisms such as "irrespective" and "humongous", and all the phrases that have gotten squished together into amalgams like "gotcha" and "woulda". Most people would readily agree that these words do not belong in a college thesis paper.
What is the scrupulous lexicographer to do? Fortunately, it is not a strict either-or question, especially in a work not substantially limited by size. In an electronic resource, we can put them in, anyway. To satisfy the formal sorts, the perscriptivists, we can then place a prominent usage note in the entry, explaining just why a writer might wish to use caution with the term: ginormous is a colloquial term, regarded by many to be something less than a proper word. Thus, the reader is both informed and cautioned.
That's fine for most of the slang and jargon, but we have another problem. People keep making up new words. My sister in law coined the term "muskaroon" to mean generically any small, furry creature that scurries past too quickly to identify. Squirrels, chipmunks, gophers, and presumably rabbits would all qualify. So we have a unit of language with a symbol and a meaning. The trouble is, if you walked up to people on the street and inquired whether there were muskaroons in the area, nobody would be able to answer who hadn't talked lately to my sister in law, and that is a small minority of people, indeed.
The test here is usage. Can we demonstrate that the word is in common use? Now, depending on the character of the dictionary, we can define the rules various ways. Was it used by so many independent sources? Did anybody important (such as Shakespeare or a prominent academic journal) publish the word?
Generally, we also try to find and present examples of the term in what is called "running text". That means that it is in a paragraph, and isn't only used as somebody's nickname, say. The edge is still a fuzzy one. Are the citations in traditional print sources, such and books and journals, or are they sprinkled in a couple blogs and forums? Was the word used in only one limited context, or in a variety of sources and over a period of years? These sorts of tests can help to weed out many of the more questionable entries. At some point, though, it may yet come down to a judgment call, if not on whether a word is real, then on how to apply the rules. In these cases, I advise the users of a dictionary to bring a healthy dose of skepticism with them, to recall that even dictionaries are not infallible, and to trust at the very least that these decisions are made by real people who care for the project.
If, knowing all that, you find you don't like the way "they" are running the place, you are invited to do a better job.
Labels:
dictionary,
nonce,
protologism,
verification,
wiki,
wiktionary
Monday, February 19, 2007
Anyone may edit
Written in response to this project.
Somebody asked me today about what happens to a dictionary when anybody can edit it. As anybody who has ever edited a wiki knows, the openness is a mixed blessing.
It is a great thing, because many hands make light work. Dictionaries need to be every bit as large as the languages they catalog, so the process of gathering and maintaining the data is a huge one. As we start to add translations between languages, rather than simply defining a term, that task becomes orders of magnitude bigger. To capture all words in all languages is something that will take nothing less than a wiki and a worldwide community. It is a monumental task, but in a wiki, we can conceive of creating a resource on such a scale.
It is a great thing because so many regions and cultures can be represented. An American may understand most of the English spoken in South Africa or New Zealand, but both of those regions have slang all their own. Chile speaks Spanish far differently than Spain. All those variants can have their place.
It is a great thing because a wiki can evolve with a language. New terms come into use all the time, and a freely editable electronic resource is not limited in its capacity to store data or to accommodate a large, diverse set of editors.
The big trouble is this: if just anybody can edit, how on earth do we know it is right? I'd like to explore a few approaches here. Of course, any of these approaches could be considered as a barrier to entry, but these things are always trade-offs.
Somebody asked me today about what happens to a dictionary when anybody can edit it. As anybody who has ever edited a wiki knows, the openness is a mixed blessing.
It is a great thing, because many hands make light work. Dictionaries need to be every bit as large as the languages they catalog, so the process of gathering and maintaining the data is a huge one. As we start to add translations between languages, rather than simply defining a term, that task becomes orders of magnitude bigger. To capture all words in all languages is something that will take nothing less than a wiki and a worldwide community. It is a monumental task, but in a wiki, we can conceive of creating a resource on such a scale.
It is a great thing because so many regions and cultures can be represented. An American may understand most of the English spoken in South Africa or New Zealand, but both of those regions have slang all their own. Chile speaks Spanish far differently than Spain. All those variants can have their place.
It is a great thing because a wiki can evolve with a language. New terms come into use all the time, and a freely editable electronic resource is not limited in its capacity to store data or to accommodate a large, diverse set of editors.
The big trouble is this: if just anybody can edit, how on earth do we know it is right? I'd like to explore a few approaches here. Of course, any of these approaches could be considered as a barrier to entry, but these things are always trade-offs.
- Appoint trusted users to do the housekeeping. These are the sysops, administrators, bureaucrats, the librarians, or the janitors, depending on your point of view. These somebodies keep watch and undo the damage that some of the just-anybodies can do. If somebody writes an article containing typical vandalism, such as "asdfasdf" or "Dave is a dork!", an administrator can delete or undo it. Much vandalism is so predictable that even a bot can detect and remove it. Unfortunately, a select group of administrators, however well trusted or well-read, cannot be everywhere at once, and they cannot know everything. Things get missed, even with a checklist system such as patrolled edits. It is likewise impossible for an administrator or small group of administrators to know everything. Misinformation, intentional or otherwise, is not so easy to spot as out-and-out nonsense.
- Hold people accountable. Articles have histories, so you can see who did what. Even pseudonymous users develop reputations. Anonymous users tend to attract the most scrutiny. An active, healthy wiki often develops into a meritocracy, with leaders having sway (though not necessarily authority) based on reputation, seniority, and trust in the community. This effect generally works to improve content, but even a well-known, trusted user may make mistakes. If he or she is trusted well enough, there is a risk that an error or oversight may go unnoticed.
- Allow anybody and everybody to scrutinize and correct or flag the content. The process is not foolproof, especially in larger projects, but wikis have a remarkable capacity for self-cleaning. Of course, this approach can tend to result in a sort of groupthink effect: if enough people believe it, then it must be so.
- Demand credentials. Don't just let in any old riffraff. Wikipedia has clearly shown the power of amateurs and volunteers to create great content, but it is certainly possible to limit the users in a project, or part of the project, to a certain group. This approach is most appropriate to a wiki serving a closed community, such as a professional or academic group, especially one dedicated to a particularly narrow or specialized topic.
- Make the messes behind the scenes, and publish only the good stuff, with some review process. The German Wikipedia published a paper book containing selected articles. Online, there have been proposals for a "Stable Versions" system, where a mature article would be reviewed and locked, and any additional changes would go through a separate editing or discussion page.
- Demand references. There is a movement within Wikipedia to reference the articles and the claims made in them. In the context of a dictionary, references may be other dictionaries. Is the word recognized by RAE or OED (whom we trust to have done the requisite homework)? They may be other works about words. Or, they may be citations. Citations are quotations including the word in question. They show context and provide evidence that the word is or was in use. Of course, we must still question the validity of the evidence. Are the 400 Google hits because somebody prolific uses that nonsense word as a handle? Is a word more valid if it was used by a blogger or two, or by Thornton Wilder? Is an etymology known with reasonable certainty or is it apocryphal? Depending on the size and resources of the wiki, efforts to verify and reference articles may be systematic, or they may be requested when a given entry or fact is questioned.
Subscribe to:
Posts (Atom)