First, let me make it perfectly clear that this is a discussion that has raged for centuries. I know full well that everybody has his or her own opinion on this matter, and that I am not going to resolve this issue today. This is an overview and a bit of personal opinion, as relates to online dictionaries.
Intuitively, we all know the answer. A word is a unit of language conveying some meaning. But how do we decide what is a real word? We look in a dictionary, of course. What do we do if we're writing a dictionary?
We are caught between cataloging what is "right" (prescriptivism) and what is actually done (descriptivism). The pendulum has lately swung towards descriptivism, and I would say that there are some good reasons for that trend. The language that is spoken on the streets is not the same language that is written in academia. Somebody learning a language may genuinely need help sorting out the less proper terms in it.
Take, for instance, colloquialisms such as "irrespective" and "humongous", and all the phrases that have gotten squished together into amalgams like "gotcha" and "woulda". Most people would readily agree that these words do not belong in a college thesis paper.
What is the scrupulous lexicographer to do? Fortunately, it is not a strict either-or question, especially in a work not substantially limited by size. In an electronic resource, we can put them in, anyway. To satisfy the formal sorts, the perscriptivists, we can then place a prominent usage note in the entry, explaining just why a writer might wish to use caution with the term: ginormous is a colloquial term, regarded by many to be something less than a proper word. Thus, the reader is both informed and cautioned.
That's fine for most of the slang and jargon, but we have another problem. People keep making up new words. My sister in law coined the term "muskaroon" to mean generically any small, furry creature that scurries past too quickly to identify. Squirrels, chipmunks, gophers, and presumably rabbits would all qualify. So we have a unit of language with a symbol and a meaning. The trouble is, if you walked up to people on the street and inquired whether there were muskaroons in the area, nobody would be able to answer who hadn't talked lately to my sister in law, and that is a small minority of people, indeed.
The test here is usage. Can we demonstrate that the word is in common use? Now, depending on the character of the dictionary, we can define the rules various ways. Was it used by so many independent sources? Did anybody important (such as Shakespeare or a prominent academic journal) publish the word?
Generally, we also try to find and present examples of the term in what is called "running text". That means that it is in a paragraph, and isn't only used as somebody's nickname, say. The edge is still a fuzzy one. Are the citations in traditional print sources, such and books and journals, or are they sprinkled in a couple blogs and forums? Was the word used in only one limited context, or in a variety of sources and over a period of years? These sorts of tests can help to weed out many of the more questionable entries. At some point, though, it may yet come down to a judgment call, if not on whether a word is real, then on how to apply the rules. In these cases, I advise the users of a dictionary to bring a healthy dose of skepticism with them, to recall that even dictionaries are not infallible, and to trust at the very least that these decisions are made by real people who care for the project.
If, knowing all that, you find you don't like the way "they" are running the place, you are invited to do a better job.
Wednesday, February 28, 2007
What is a Word?
Labels:
dictionary,
nonce,
protologism,
verification,
wiki,
wiktionary
Monday, February 19, 2007
Anyone may edit
Written in response to this project.
Somebody asked me today about what happens to a dictionary when anybody can edit it. As anybody who has ever edited a wiki knows, the openness is a mixed blessing.
It is a great thing, because many hands make light work. Dictionaries need to be every bit as large as the languages they catalog, so the process of gathering and maintaining the data is a huge one. As we start to add translations between languages, rather than simply defining a term, that task becomes orders of magnitude bigger. To capture all words in all languages is something that will take nothing less than a wiki and a worldwide community. It is a monumental task, but in a wiki, we can conceive of creating a resource on such a scale.
It is a great thing because so many regions and cultures can be represented. An American may understand most of the English spoken in South Africa or New Zealand, but both of those regions have slang all their own. Chile speaks Spanish far differently than Spain. All those variants can have their place.
It is a great thing because a wiki can evolve with a language. New terms come into use all the time, and a freely editable electronic resource is not limited in its capacity to store data or to accommodate a large, diverse set of editors.
The big trouble is this: if just anybody can edit, how on earth do we know it is right? I'd like to explore a few approaches here. Of course, any of these approaches could be considered as a barrier to entry, but these things are always trade-offs.
Somebody asked me today about what happens to a dictionary when anybody can edit it. As anybody who has ever edited a wiki knows, the openness is a mixed blessing.
It is a great thing, because many hands make light work. Dictionaries need to be every bit as large as the languages they catalog, so the process of gathering and maintaining the data is a huge one. As we start to add translations between languages, rather than simply defining a term, that task becomes orders of magnitude bigger. To capture all words in all languages is something that will take nothing less than a wiki and a worldwide community. It is a monumental task, but in a wiki, we can conceive of creating a resource on such a scale.
It is a great thing because so many regions and cultures can be represented. An American may understand most of the English spoken in South Africa or New Zealand, but both of those regions have slang all their own. Chile speaks Spanish far differently than Spain. All those variants can have their place.
It is a great thing because a wiki can evolve with a language. New terms come into use all the time, and a freely editable electronic resource is not limited in its capacity to store data or to accommodate a large, diverse set of editors.
The big trouble is this: if just anybody can edit, how on earth do we know it is right? I'd like to explore a few approaches here. Of course, any of these approaches could be considered as a barrier to entry, but these things are always trade-offs.
- Appoint trusted users to do the housekeeping. These are the sysops, administrators, bureaucrats, the librarians, or the janitors, depending on your point of view. These somebodies keep watch and undo the damage that some of the just-anybodies can do. If somebody writes an article containing typical vandalism, such as "asdfasdf" or "Dave is a dork!", an administrator can delete or undo it. Much vandalism is so predictable that even a bot can detect and remove it. Unfortunately, a select group of administrators, however well trusted or well-read, cannot be everywhere at once, and they cannot know everything. Things get missed, even with a checklist system such as patrolled edits. It is likewise impossible for an administrator or small group of administrators to know everything. Misinformation, intentional or otherwise, is not so easy to spot as out-and-out nonsense.
- Hold people accountable. Articles have histories, so you can see who did what. Even pseudonymous users develop reputations. Anonymous users tend to attract the most scrutiny. An active, healthy wiki often develops into a meritocracy, with leaders having sway (though not necessarily authority) based on reputation, seniority, and trust in the community. This effect generally works to improve content, but even a well-known, trusted user may make mistakes. If he or she is trusted well enough, there is a risk that an error or oversight may go unnoticed.
- Allow anybody and everybody to scrutinize and correct or flag the content. The process is not foolproof, especially in larger projects, but wikis have a remarkable capacity for self-cleaning. Of course, this approach can tend to result in a sort of groupthink effect: if enough people believe it, then it must be so.
- Demand credentials. Don't just let in any old riffraff. Wikipedia has clearly shown the power of amateurs and volunteers to create great content, but it is certainly possible to limit the users in a project, or part of the project, to a certain group. This approach is most appropriate to a wiki serving a closed community, such as a professional or academic group, especially one dedicated to a particularly narrow or specialized topic.
- Make the messes behind the scenes, and publish only the good stuff, with some review process. The German Wikipedia published a paper book containing selected articles. Online, there have been proposals for a "Stable Versions" system, where a mature article would be reviewed and locked, and any additional changes would go through a separate editing or discussion page.
- Demand references. There is a movement within Wikipedia to reference the articles and the claims made in them. In the context of a dictionary, references may be other dictionaries. Is the word recognized by RAE or OED (whom we trust to have done the requisite homework)? They may be other works about words. Or, they may be citations. Citations are quotations including the word in question. They show context and provide evidence that the word is or was in use. Of course, we must still question the validity of the evidence. Are the 400 Google hits because somebody prolific uses that nonsense word as a handle? Is a word more valid if it was used by a blogger or two, or by Thornton Wilder? Is an etymology known with reasonable certainty or is it apocryphal? Depending on the size and resources of the wiki, efforts to verify and reference articles may be systematic, or they may be requested when a given entry or fact is questioned.
Friday, February 16, 2007
An article in Nature ...
I am absolutely thrilled with the article that was published last Wednesday in Nature... It is a great article and it explains really well what we hope to achieve with relational data in MediaWiki. The only thing that is a bit sad is, that you have to pay $30 for the privilege of reading it.
The article is great, and what makes it special is the great presentation that Knewco has created to explain what we hope to achieve; this demo available at wikiprofessional.info. It presents some really impressive figures; it indicates the work done to integrate several important resources of the bio-medical domain, the numbers involved.
For me the most important point is that this is likely to be a very important stimulus to the Open Access movement. It indicates that it is possible to bring what was divided together. It allows people to work with the terminology of their field and also add data that is very specific. Information that goes much further than what was envisioned in what was once called the "Ultimate Wiktionary".
The whole notion of a resource that because of its roots already merged lexicology, terminology and ontology is really special. With the integration of such specialised data from different domains like the bio-medical, another really interested experiment will be under way when the data gets imported and merged. There is a nascent community for the bio-medical domain and, it will find that it will co-exist with the existing OmegaWiki community.
Both communities have everything to gain from collaboration; much of what the existing OmegaWiki community cares about will be seen as a fringe benefit. On the other hand, the translations that exist for concepts like malaria will prove to be of value when scientific articles are considered that were not published in English.
I am convinced that a bright future is ahead of us. We have this vision of what may come, I wish I could look into the future and see what it will be like. :)
Thanks,
GerardM
The article is great, and what makes it special is the great presentation that Knewco has created to explain what we hope to achieve; this demo available at wikiprofessional.info. It presents some really impressive figures; it indicates the work done to integrate several important resources of the bio-medical domain, the numbers involved.
For me the most important point is that this is likely to be a very important stimulus to the Open Access movement. It indicates that it is possible to bring what was divided together. It allows people to work with the terminology of their field and also add data that is very specific. Information that goes much further than what was envisioned in what was once called the "Ultimate Wiktionary".
The whole notion of a resource that because of its roots already merged lexicology, terminology and ontology is really special. With the integration of such specialised data from different domains like the bio-medical, another really interested experiment will be under way when the data gets imported and merged. There is a nascent community for the bio-medical domain and, it will find that it will co-exist with the existing OmegaWiki community.
Both communities have everything to gain from collaboration; much of what the existing OmegaWiki community cares about will be seen as a fringe benefit. On the other hand, the translations that exist for concepts like malaria will prove to be of value when scientific articles are considered that were not published in English.
I am convinced that a bright future is ahead of us. We have this vision of what may come, I wish I could look into the future and see what it will be like. :)
Thanks,
GerardM
Sunday, February 11, 2007
Why compete when you can collaborate ?
All words of all languages of the world.. that is what we eventually aim to include in OmegaWiki. This aim is of such a magnitude that you have to be certifiable to come up with such a project. The functional design for the project includes much more; everything including the kitchen sink..
When everything is to be included in one project, it is easy to suggest that people contribute to the project. When the project includes everything why have another?
In an Open Source / Open Content environment this is not necessarily how it works. Why should the others be seen as competitors? They do their own thing, sure. You may want to achieve the same thing, also true. It is however much possible to find the synergy between projects. This way you can build on each others accomplishments.
The Shtooka project is something I learned about the other day. The one thing it does really well is the way they make recording pronunciations easy. You can record a string of words and it will save them for you one at a time.
Wiktionarians saw this and they are working an upload facility so that it will also be saved automatically to Commons. I warned that the files should not only be saved as .ogg files. In order to make sure they are relevant for scientists there should also be a .wav file. The current thinking is that the flac file format will work as well and the benefit is that it provides a loss-less compression. To make sure that this is the case, the praat software, software that is also available under a GPL license, was analysed and it was considered that it is easy to incorporate this flac file format.
People from effectively five different communities are now working together. It will be even possible to include links to OmegaWiki in the Shtooka meta data. This will be possible even though both projects do their own thing. Both the data and the functionality can be shared.
I may be certifiable, but this kind of collaboration is awesome and, it is why there may be method to this madness.. :)
Thanks,
GerardM
When everything is to be included in one project, it is easy to suggest that people contribute to the project. When the project includes everything why have another?
In an Open Source / Open Content environment this is not necessarily how it works. Why should the others be seen as competitors? They do their own thing, sure. You may want to achieve the same thing, also true. It is however much possible to find the synergy between projects. This way you can build on each others accomplishments.
The Shtooka project is something I learned about the other day. The one thing it does really well is the way they make recording pronunciations easy. You can record a string of words and it will save them for you one at a time.
Wiktionarians saw this and they are working an upload facility so that it will also be saved automatically to Commons. I warned that the files should not only be saved as .ogg files. In order to make sure they are relevant for scientists there should also be a .wav file. The current thinking is that the flac file format will work as well and the benefit is that it provides a loss-less compression. To make sure that this is the case, the praat software, software that is also available under a GPL license, was analysed and it was considered that it is easy to incorporate this flac file format.
People from effectively five different communities are now working together. It will be even possible to include links to OmegaWiki in the Shtooka meta data. This will be possible even though both projects do their own thing. Both the data and the functionality can be shared.
I may be certifiable, but this kind of collaboration is awesome and, it is why there may be method to this madness.. :)
Thanks,
GerardM
Thursday, February 08, 2007
Become an OmegaWiki developer
OmegaWiki is now running the latest version of the MediaWiki software used by Wikipedia. This is a major milestone, as it also makes it a lot easier for anyone to join in the fun of developing the open source OmegaWiki/Wikidata software. To give credit where credit is due, these are the people who have contributed to the code so far:
- Peter-Jan Roes
- Karsten Uil
- Sean Burke
- Rod A. Smith (sticky tree expansion via cookies)
- Ævar Arnfjörð Bjarmason (namespace code installer)
- Charles Pritchard (Multilingual MediaWiki development, ongoing)
- Jelte Zeilstra (untranslated meaning script, under review)
- Zdenek Broz (statistical scripts, under review)
- Paa-Kwesi Imbeah (Wikimedia Commons support, under review)
- Marc Carmen (TBX export, incomplete)
- myself
Wednesday, January 31, 2007
Greek languages
At OmegaWiki, we saw that Lou started to change the capitalisation of language names... A few days ago I was surprised that the Georgian names for languages were incorrectly spelled. Now it is Greek.
It is really powerful to see that by having the languages corrected, it will be available for everybody who wants to know about Greek. This reason for using OmegaWiki proves itself again.
Thanks,
GerardM
It is really powerful to see that by having the languages corrected, it will be available for everybody who wants to know about Greek. This reason for using OmegaWiki proves itself again.
Thanks,
GerardM
Sunday, January 28, 2007
Latin roots etc.
Well yesterday one thing came into mind - a dictionary a teacher of mine at the language school had. It was a dictionary that listed Latin words with many translations into other languages and one thing is obvious: all these words of course were similar in all languages. If you knew one of them and studied the other language it would have been easy to create the relative words following a set of rules for most of them.
So one thing should be obvious: to insert these words with their translations into OmegaWiki ... but well, there is one problem with Latin - the "normal" Latin language should not be mixed with the taxonomical Latin that is used in science ... so we need to create two languages: Latin and taxonomical Latin ... who knows if the relative language codes exist somewhere in the ISO 639 standards.
Technorati: language, Latin, ISO 639, standards, translation, dictionary
So one thing should be obvious: to insert these words with their translations into OmegaWiki ... but well, there is one problem with Latin - the "normal" Latin language should not be mixed with the taxonomical Latin that is used in science ... so we need to create two languages: Latin and taxonomical Latin ... who knows if the relative language codes exist somewhere in the ISO 639 standards.
Technorati: language, Latin, ISO 639, standards, translation, dictionary
Labels:
dictionary,
ISO 639,
language,
Latin,
standards,
translation
Friday, January 26, 2007
OLPC needs a dictionary viewer
I had a word with the director of content for the OLPC, the One Laptop Per Child Project. As you know OmegaWiki is the project that works on providing the OLPC with dictionary content. We are working on all these words, and while we are making steady progress, there is so much still left to do. We are getting more Expressions in many languages, the definitions are lagging and while we do our best, it is still very much the difference between there being nothing and there being next to nothing. It does however show that things are getting under way...
As the moment when kids are exposed to the systems is drawing closer, it is relevant that the data can be used. So we need a dictionary viewer. It needs to run on Linux and, it should have a small footprint. As we will provide all these languages, it will be interesting to see how the rich tapestry that OmegaWiki tries to weave will materialise on these nifty systems.
When you have a suggestion, please let us know :)
Thanks,
GerardM
As the moment when kids are exposed to the systems is drawing closer, it is relevant that the data can be used. So we need a dictionary viewer. It needs to run on Linux and, it should have a small footprint. As we will provide all these languages, it will be interesting to see how the rich tapestry that OmegaWiki tries to weave will materialise on these nifty systems.
When you have a suggestion, please let us know :)
Thanks,
GerardM
Wednesday, January 24, 2007
Georgian names for languages
In the past I got permission to copy content from a resource with the names of languages. I am still grateful for the data. It got the Dutch Wiktionary going really nicely and, as we needed at the time those names of languages for the user interface.
With OmegaWiki we had the same issue; we needed language names again for the user interface. This was to make it possible for people to see the labels of translations in their own language. From the moment the data became available we have learned a lot, for instance that language in languages like Danish and Italian do not capitalise the names of languages.
Today I was told that many of the names of languages in Georgian were found to be in error and had been corrected. The great news for OmegaWiki is, that we only have to do this once and it is good everywhere. The sad thing is that it is probably wrong in many, many Wiktionaries. There were two types of errors; it was just wrong or it was the name of someone from a country in stead of the name of the language.
The best I can do for the Wiktionaries is notify in this way as I do not really now what needs doing.
Thanks,
GerardM
With OmegaWiki we had the same issue; we needed language names again for the user interface. This was to make it possible for people to see the labels of translations in their own language. From the moment the data became available we have learned a lot, for instance that language in languages like Danish and Italian do not capitalise the names of languages.
Today I was told that many of the names of languages in Georgian were found to be in error and had been corrected. The great news for OmegaWiki is, that we only have to do this once and it is good everywhere. The sad thing is that it is probably wrong in many, many Wiktionaries. There were two types of errors; it was just wrong or it was the name of someone from a country in stead of the name of the language.
The best I can do for the Wiktionaries is notify in this way as I do not really now what needs doing.
Thanks,
GerardM
Tuesday, January 23, 2007
Stichting Open Progress
Stichting Open Progress is the Dutch not for profit organisation that is the legal organisation behind OmegaWiki. As OmegaWiki is growing to the extend where we have to consider contracts for hosting, grants and the like, we had a need for an organisation.
The need for an organisation was also felt as we already had some projects where we would have been better able to do things when there was a legal entity backing up the activities. Some of these projects are quite substantial.
Open Progress aims to develop both Open Source/Free Software and Open Content/Free Content projects. As part of its mission it gives room for projects that are aligned with the aims of the stichting. Obviously OmegaWiki is the first; from an organisational point of view, the OmegaWiki commission decides on the issues that arise. Resolution will be enacted for the project by the stichting provided they are in line with the Dutch law and, provided they do not circumvent the aims of the stichting. This way Open Progress hopes to make OmegaWiki a safe haven where people and organisations work in the understanding that the aims of the project will be respected.
There are two websites for OpenProgress; in line with the experiences of the Wikimedia Foundation, we have both an internal and an external wiki. The internal will use Semantic MediaWiki to leverage as much as possible the information that we will include. As the information will include both personal information and confidential project information, the internal will be invite only.
Thanks,
Gerard Meijssen
voorzitter Stichting Open Progress
The need for an organisation was also felt as we already had some projects where we would have been better able to do things when there was a legal entity backing up the activities. Some of these projects are quite substantial.
Open Progress aims to develop both Open Source/Free Software and Open Content/Free Content projects. As part of its mission it gives room for projects that are aligned with the aims of the stichting. Obviously OmegaWiki is the first; from an organisational point of view, the OmegaWiki commission decides on the issues that arise. Resolution will be enacted for the project by the stichting provided they are in line with the Dutch law and, provided they do not circumvent the aims of the stichting. This way Open Progress hopes to make OmegaWiki a safe haven where people and organisations work in the understanding that the aims of the project will be respected.
There are two websites for OpenProgress; in line with the experiences of the Wikimedia Foundation, we have both an internal and an external wiki. The internal will use Semantic MediaWiki to leverage as much as possible the information that we will include. As the information will include both personal information and confidential project information, the internal will be invite only.
Thanks,
Gerard Meijssen
voorzitter Stichting Open Progress
Monday, January 15, 2007
Destinazione Italia
Destinazione Italia is a project of the University of Bamberg. It provides training for people learning an advanced level of Italian. Bamberg is a German University and many of its students are German. Many of the students do have a different mother tongue. Learning a third language based on the knowledge of a second language is less effective than learning based on the knowledge of the mother tongue.
I am really proud to announce that OmegaWiki has been selected by the University of Bamberg as the platform that will host the lexicological information for "Destinazione Italia". The initial phase of the project will create a lot of Italian based DefinedMeanings. In the second phase we will translate these words to English, German and Spanish. The third phase is to find translations in as many other languages as we can get.
Research done by Zdenek Broz learned, that when the combination of quality translations of German, English and Spanish is found, it will allow the inclusion of translations of other languages when these translations are shared in a different resource. According to Zdenek's figures this will get us an accuracy of around, probably better than 95%.
There is a budget to get us many translations in other languages. The sweet thing is, when we are able to provide quality translations, the budget can be used for other things. This can be to improve the OmegaWiki usability, it can also be to spend money on a language that is not part of the initial list of languages "Destinazione Italia" supports.
The challenge is therefore, how much can we do with a limited budget. What will be the added value of creating content in a Wiki environment. When will OmegaWiki reach the tipping point where collaboration in OmegaWiki is the obvious thing to do, "Destinazione Italia" will help us reach that point. :)
Thanks,
GerardM
I am really proud to announce that OmegaWiki has been selected by the University of Bamberg as the platform that will host the lexicological information for "Destinazione Italia". The initial phase of the project will create a lot of Italian based DefinedMeanings. In the second phase we will translate these words to English, German and Spanish. The third phase is to find translations in as many other languages as we can get.
Research done by Zdenek Broz learned, that when the combination of quality translations of German, English and Spanish is found, it will allow the inclusion of translations of other languages when these translations are shared in a different resource. According to Zdenek's figures this will get us an accuracy of around, probably better than 95%.
There is a budget to get us many translations in other languages. The sweet thing is, when we are able to provide quality translations, the budget can be used for other things. This can be to improve the OmegaWiki usability, it can also be to spend money on a language that is not part of the initial list of languages "Destinazione Italia" supports.
The challenge is therefore, how much can we do with a limited budget. What will be the added value of creating content in a Wiki environment. When will OmegaWiki reach the tipping point where collaboration in OmegaWiki is the obvious thing to do, "Destinazione Italia" will help us reach that point. :)
Thanks,
GerardM
Saturday, January 13, 2007
Alexa and statistics again
As you may know, Alexa is a company that tries to divine the relative ranking of one website when it comes to traffic. It uses some functionality that is associated with the Internet Explorer browser. This is a browser that comes standard with the Windows operating system and it used to be absolutely dominate the global market. This market was eroded by the Firefox browser, Firefox has carved out a niche for itself and has a worldwide use of more than 15 percent.
The distribution of the use is different from country to country; in Germany Firefox is much more popular. The distribution is also different from website to website; for OmegaWiki a big percentage of people use Firefox; in November the traffic from both browsers was evenly split. This means that the reporting of Alexa has a bias that severely affects its accuracy.
For OmegaWiki Alexa provides the best statistics we currently can provide you with. We have a technical issue that limits the usefulness of our webaliser statistics. Because of the name change from WiktionaryZ, you do find that there are separate OmegaWiki and WiktionaryZ statistics at Alexa.. I wrote them about it, and Alexa will merge the data in one to two weeks.
Any way, better statistics in two weeks, both from Alexa and from our webaliser. It will be interesting to see more realistically if and to what extend our traffic is evolving.
Thanks,
GerardM
The distribution of the use is different from country to country; in Germany Firefox is much more popular. The distribution is also different from website to website; for OmegaWiki a big percentage of people use Firefox; in November the traffic from both browsers was evenly split. This means that the reporting of Alexa has a bias that severely affects its accuracy.
For OmegaWiki Alexa provides the best statistics we currently can provide you with. We have a technical issue that limits the usefulness of our webaliser statistics. Because of the name change from WiktionaryZ, you do find that there are separate OmegaWiki and WiktionaryZ statistics at Alexa.. I wrote them about it, and Alexa will merge the data in one to two weeks.
Any way, better statistics in two weeks, both from Alexa and from our webaliser. It will be interesting to see more realistically if and to what extend our traffic is evolving.
Thanks,
GerardM
Tuesday, January 09, 2007
Relation types
OmegaWiki includes one thesaurus at the moment. The GEMET thesaurus was a boon, having it demonstrated really well that what was then WiktionaryZ is able to include a thesaurus and does a good job showing relations.
The next step will be to demonstrate that we can reliably include multiple thesauri. This is a lot more complicated. The problem has to do with the relation types used and what they mean. The issue is that you cannot infer that what is meant by a particular phrase like "is part of" in one thesaurus means the same in an other.
This means that you have to tread carefully. The first thing that you can do is treat a collection as a self contained unit. The relation types would as a consequence be only available and applicable to those DefinedMeanings that are part of the collection.
When a collection is to be integrated, there will be a need to merge those DefinedMeanings that are conceptually the same. This may merge pre existing relations and collection relations. In effect this may demonstrate that certain relation types are indeed the same and consequently the collection relations may now get a relation type that is of an higher level.
The higher level relations are based on domains. You will agree with me that only organisms include proteins. The consequence is that both parts of such a relation will have to be either an organism or a protein.
The last level would be the universal relation types. They will be true never mind the domain. Currently ALL relation types can be universally applied. The current GEMET relation types will not remain that way. They will prove to be quite arbitrary and I expect that we will at some stage restrict their usage. This will likely be offset by functionality that will offset the pain of losing a tool that is quiet popular.
Thanks,
GerardM
The next step will be to demonstrate that we can reliably include multiple thesauri. This is a lot more complicated. The problem has to do with the relation types used and what they mean. The issue is that you cannot infer that what is meant by a particular phrase like "is part of" in one thesaurus means the same in an other.
This means that you have to tread carefully. The first thing that you can do is treat a collection as a self contained unit. The relation types would as a consequence be only available and applicable to those DefinedMeanings that are part of the collection.
When a collection is to be integrated, there will be a need to merge those DefinedMeanings that are conceptually the same. This may merge pre existing relations and collection relations. In effect this may demonstrate that certain relation types are indeed the same and consequently the collection relations may now get a relation type that is of an higher level.
The higher level relations are based on domains. You will agree with me that only organisms include proteins. The consequence is that both parts of such a relation will have to be either an organism or a protein.
The last level would be the universal relation types. They will be true never mind the domain. Currently ALL relation types can be universally applied. The current GEMET relation types will not remain that way. They will prove to be quite arbitrary and I expect that we will at some stage restrict their usage. This will likely be offset by functionality that will offset the pain of losing a tool that is quiet popular.
Thanks,
GerardM
Sunday, January 07, 2007
IPA or the International Phonetic Alphabet
I woke up having dreamt of OmegaWiki. I had a brainwave; it is easy to include IPA into OmegaWiki. I sprinted out of bed, asked Leftmost if it could be done, it could and then I started to look into IPA for the first time. I have started reading and I do admit understanding the text is beyond me. Some facts that can be found:
The result is that, yes we can enter IPA transcriptions when we enable it. The problem however is that we should not allow IPA transcriptions optimised for the English speakers. OmegaWiki is to be used by people with ANY language as a background.
I think that before we enable IPA transcriptions there should be some more discussion.
Thanks,
GerardM
- IPA has approximately 107 base symbols and 55 modifiers.
- There is a specific chart for the sounds of English.
- Different resources use IPA in different ways.
- There are even different symbols used for British English :(
The result is that, yes we can enter IPA transcriptions when we enable it. The problem however is that we should not allow IPA transcriptions optimised for the English speakers. OmegaWiki is to be used by people with ANY language as a background.
I think that before we enable IPA transcriptions there should be some more discussion.
Thanks,
GerardM
Wednesday, January 03, 2007
Tiny Winy Buggy Bugs ... that take loads of time to explain ...
Every now and again I have that situation ... I send people to OmegaWiki or before WiktionaryZ and they tell me: but the dictionary is completely wrong ... the reason?
Well there is that already known bug where you have a wrong page title like here with the title heavy having the contents for small. In the meantime I explained it more and more often ... all this takes time - the first impression we give when people look at such pages is: OmegaWiki is full of errors - even simple stuff seems to be wrong. They don't know that this is a bug and I am wondering how many went away being deluded by our product.
Is it so difficult to get this fixed? How valuable is our time that we need to explain the same thing over and over again?
Well: for sure it must be done before going officially live, otherwise our credibility will be very much questioned and we will need loads of hours to explain the same bug over and over again.
Thanks!
Well there is that already known bug where you have a wrong page title like here with the title heavy having the contents for small. In the meantime I explained it more and more often ... all this takes time - the first impression we give when people look at such pages is: OmegaWiki is full of errors - even simple stuff seems to be wrong. They don't know that this is a bug and I am wondering how many went away being deluded by our product.
Is it so difficult to get this fixed? How valuable is our time that we need to explain the same thing over and over again?
Well: for sure it must be done before going officially live, otherwise our credibility will be very much questioned and we will need loads of hours to explain the same bug over and over again.
Thanks!
Sunday, December 31, 2006
When a year ends, it is an opportune time to reflect and to look forward. The past year was pretty amazing; WiktionaryZ went from nothing but a proof of concept to OmegaWiki with functionality that includes language depended part of speech annotation. I want to express gratitude to all the people that made this possible, I want to particularly thank Knewco for their growing belief and understanding what Open Source, Open Content and Open Access. They have not only been instrumental in the development of OmegaWiki, my prediction is that their contribution will help the evolution immensely in the thinking how to make all this sustainable.
OmegaWiki is very much an experiment; it brings organisations and communities together. In the development of OmegaWiki we have seen how immensly valuable both are. It is because of this that I find it so disappointing that the Wikimedia Foundation finds it so difficult to entertain the possibilities that such cooperation brings.
As the WMF decided that they were not in a position to host OmegaWiki, we are now in a position with Open Progress, the Dutch not for profit organisation we had to set up, to do things different where we think it makes a difference. This notion of "eating our own dog food" has always been an important part of the way OmegaWiki works; in the same way we hope to implement the ideas that we agree on in our community.
The most important notion of OmegaWiki is in its definition of success; "success is when people find an application for our data that we did not think of". The implication is that the data of OmegaWiki is there to be used and that collaboration is what OmegaWiki is about. Collaboration in a way that recognises that the success of OmegaWiki is integral to the success of everyone who helps OmegaWiki to be a success.
Technically OmegaWiki is near the turning point where we have sufficient capability to host several ontologies together. This will coincide with the arrival of "collection relation types" and "domain relation types". With the arrival of language dependent attributes we are on the threshold of being able to include the information that differentiates one linguistic entity from another.
Both these two capabilities will be really important in 2007. They will bring a lot of relevant data to OmegaWiki and is likely to bring relevance to the project. I envision that we will indicate which UNICODE characters make up the standard characters for a language. It will show among other things where more characters are needed, it will also help define what the proper sorting order is for a linguistic entity.
In a mail to the Yahoo aphrophonewikis group, Don Osborn asked for a "Year of Unicode in Africa", I hope that OmegaWiki will help make this happen. In a mail to the Wiktionary mailing list, Javier Carro suggest to collaborate on what he calls "Schemes". Both are two mails of the last week, I expect the implementation of both will be feasible.
I am sure 2007 will be great .. Prosit Neujahr :)
Thanks,
GerardM
OmegaWiki is very much an experiment; it brings organisations and communities together. In the development of OmegaWiki we have seen how immensly valuable both are. It is because of this that I find it so disappointing that the Wikimedia Foundation finds it so difficult to entertain the possibilities that such cooperation brings.
As the WMF decided that they were not in a position to host OmegaWiki, we are now in a position with Open Progress, the Dutch not for profit organisation we had to set up, to do things different where we think it makes a difference. This notion of "eating our own dog food" has always been an important part of the way OmegaWiki works; in the same way we hope to implement the ideas that we agree on in our community.
The most important notion of OmegaWiki is in its definition of success; "success is when people find an application for our data that we did not think of". The implication is that the data of OmegaWiki is there to be used and that collaboration is what OmegaWiki is about. Collaboration in a way that recognises that the success of OmegaWiki is integral to the success of everyone who helps OmegaWiki to be a success.
Technically OmegaWiki is near the turning point where we have sufficient capability to host several ontologies together. This will coincide with the arrival of "collection relation types" and "domain relation types". With the arrival of language dependent attributes we are on the threshold of being able to include the information that differentiates one linguistic entity from another.
Both these two capabilities will be really important in 2007. They will bring a lot of relevant data to OmegaWiki and is likely to bring relevance to the project. I envision that we will indicate which UNICODE characters make up the standard characters for a language. It will show among other things where more characters are needed, it will also help define what the proper sorting order is for a linguistic entity.
In a mail to the Yahoo aphrophonewikis group, Don Osborn asked for a "Year of Unicode in Africa", I hope that OmegaWiki will help make this happen. In a mail to the Wiktionary mailing list, Javier Carro suggest to collaborate on what he calls "Schemes". Both are two mails of the last week, I expect the implementation of both will be feasible.
I am sure 2007 will be great .. Prosit Neujahr :)
Thanks,
GerardM
Saturday, December 30, 2006
Relevancy
Today it was in the news that Saddam Hussein was executed. He was no choir boy, there was a trial. Many people are happy with his death, many people are unhappy with his death. I do not want to express my opinion; it is not relevant.
With occasions like this, it is important that the words that can be associated with such an event are understood and available in resources like OmegaWiki. The words that are in news items are the words that need to be explained. The figures as they exist for Wikipedia show that the articles that are most visited are to do with sex, sport and news. I expect that this is also true for a dictionary. There are no statistics that I am aware of.
The word of the day for tomorrow could be gallows but I think that such a word on the last day of the year is a bit much.. Justice is much more appropriate.
Thanks,
GerardM
With occasions like this, it is important that the words that can be associated with such an event are understood and available in resources like OmegaWiki. The words that are in news items are the words that need to be explained. The figures as they exist for Wikipedia show that the articles that are most visited are to do with sex, sport and news. I expect that this is also true for a dictionary. There are no statistics that I am aware of.
The word of the day for tomorrow could be gallows but I think that such a word on the last day of the year is a bit much.. Justice is much more appropriate.
Thanks,
GerardM
Tuesday, December 26, 2006
A flurry of activities
With the new part of speech functionality, OmegaWiki sees a lot of activity of another kind. For the first languages some of the parts of speech have been indentified. The system is by necessity laborious; all the parts of speech have to be identified for all languages because we do not assume that a particular part of speech exists in a language.
Siebrand did a lot of work on identifying the parts of speech for the Dutch language. As a consequence we do not only have the "verb" but also the "copula". The idea is that when people know how to identify a verb as a verb, they can and may. When someone is able to identify more precisely, they can. In the mean time, when we get functionality for inflecting verbs, it should work on both.
Having functionality come on-line in small bytes, is in line with the motto of Open Source/Free Software; publish often. It really helps. I can imagine the many refinements and expansions on what we have at the moment. It is relevant to realize that our software is still very much pre-alpha. It is not complete, but it demonstrates how the functionality is growing making our dream a reality.
Thanks,
GerardM
Siebrand did a lot of work on identifying the parts of speech for the Dutch language. As a consequence we do not only have the "verb" but also the "copula". The idea is that when people know how to identify a verb as a verb, they can and may. When someone is able to identify more precisely, they can. In the mean time, when we get functionality for inflecting verbs, it should work on both.
Having functionality come on-line in small bytes, is in line with the motto of Open Source/Free Software; publish often. It really helps. I can imagine the many refinements and expansions on what we have at the moment. It is relevant to realize that our software is still very much pre-alpha. It is not complete, but it demonstrates how the functionality is growing making our dream a reality.
Thanks,
GerardM
Sunday, December 24, 2006
It is the night before Christmas
On the night before Christmas many people are full of anticipation of the presents that Christmas will bring. For OmegaWiki, we hope / expect that we will be able to have part of speech support. The software has been coded. The waiting is for the final touches and to see it enabled in OmegaWiki. Leftmost does a sterling job for us..
In order to get here, many hurdles were taken. First there was a need to have default behaviour. This led to the sample sentence functionality. With the part of speech functionality, we can have a list of values, we have functionality that is dependent on the language it applies to.
With the implementation of the functionality, there will be a need to identify what parts of speech exist in a language. This meta data needs boot strapping, so we hope people will add the parts of speech to the languages they know well. We hope that they do this well because correcting meta data is problematic.
عيد الميلاد السعي geseënde kerfees Sretan Božić Καλά Χριστούγεννα
Thanks,
GerardM
In order to get here, many hurdles were taken. First there was a need to have default behaviour. This led to the sample sentence functionality. With the part of speech functionality, we can have a list of values, we have functionality that is dependent on the language it applies to.
With the implementation of the functionality, there will be a need to identify what parts of speech exist in a language. This meta data needs boot strapping, so we hope people will add the parts of speech to the languages they know well. We hope that they do this well because correcting meta data is problematic.
عيد الميلاد السعي geseënde kerfees Sretan Božić Καλά Χριστούγεννα
Thanks,
GerardM
Thursday, December 21, 2006
OmegaWiki now supports Cebuano
Cebuano is one of the languages spoken in the Philippines. Some 20 million people speak it as their first language and some 11 million speak it as a second language. It is good that Cebuano is now enabled for editing.
There is a Wikipedia in Cebuano, as far as I can tell it has been localised in MediaWiki. This means that when the names of languages are translated, OmegaWiki will start to look attractive for the people who speak Cebuano.
Out of interest I checked if Open Office supports Cebuano. Googling learns that people use OO in Cebuan, but there is no official localisation for Cebuano and there isn't one for Tagalog either. Open Office does not support many languages and it will be hard work to get the user interface localised in more languages.
This does however not mean that Open Office cannot create content in Cebuano. Of importance is that OO is able to indicate what language people use. It is not clear to me how to do this; it seems that OO only allows for the use of languages that it fully supports. This is in my opinion not the way to approach it.
If Open Office allows for people to select the languages that they edit in, and the languages are everything that ISO-639-3 supports that is written, it should be possible for people to select the user interface that suits them best and even allow for the use of spell checkers that are created for these languages. With the correct tagging that is implicit in using the language tags, it will become easier to support the documents that are produced because it is then possible to explicitly know what language a text is in.
The questions for me are:
GerardM
There is a Wikipedia in Cebuano, as far as I can tell it has been localised in MediaWiki. This means that when the names of languages are translated, OmegaWiki will start to look attractive for the people who speak Cebuano.
Out of interest I checked if Open Office supports Cebuano. Googling learns that people use OO in Cebuan, but there is no official localisation for Cebuano and there isn't one for Tagalog either. Open Office does not support many languages and it will be hard work to get the user interface localised in more languages.
This does however not mean that Open Office cannot create content in Cebuano. Of importance is that OO is able to indicate what language people use. It is not clear to me how to do this; it seems that OO only allows for the use of languages that it fully supports. This is in my opinion not the way to approach it.
If Open Office allows for people to select the languages that they edit in, and the languages are everything that ISO-639-3 supports that is written, it should be possible for people to select the user interface that suits them best and even allow for the use of spell checkers that are created for these languages. With the correct tagging that is implicit in using the language tags, it will become easier to support the documents that are produced because it is then possible to explicitly know what language a text is in.
The questions for me are:
- Is my analysis correct .. please tell me it is not ..
- How to convince the OO people to support the use of all recognised languages that can be written
- Get support for spell checking in those languages as well.
GerardM
Subscribe to:
Posts (Atom)