The last post of the OmegaWiki blog was about statistics. Kipcool indicated that there was an issue with the numbers, in the query the deleted Expressions were not considered.
Today we have the opportunity to celebrate anew one of the milestones that went before. Today we have a cool 10.000 expressions in Italian. :)
I want to thank Kipcool, Kim Bruning and Zdenek Broz.
GerardM
Friday, April 27, 2007
Monday, April 16, 2007
Milestone
The nice thing of statistics is that when there is a milestone, it can be celebrated. It is with pleasure that I can announce that OmegaWiki now has one language with 15.000 words. It is German that has currently the most expressions. There are currently 213.550 Expressions in 133 languages.
For OmegaWiki it demonstrates that we have a nice autonomous growth. Of the languages that we started after the import of the GEMET data, Japanese currently has the highest number with almost 4.000 words. Esperanto is with some 1.500 the biggest artificial language.
You can help us with translations for the language names that we have translations in. This will improve the usability of our data when the user interface is selected in the language of people's mother tongue.
OmegaWiki is still in need of more functionality. This is something that we try to achieve in any way possible. It is however rewarding to notice what difference the existing functionality makes.
I am very happy with what we have achieved so far. It bodes well for the future :)
Thanks,
GerardM
For OmegaWiki it demonstrates that we have a nice autonomous growth. Of the languages that we started after the import of the GEMET data, Japanese currently has the highest number with almost 4.000 words. Esperanto is with some 1.500 the biggest artificial language.
You can help us with translations for the language names that we have translations in. This will improve the usability of our data when the user interface is selected in the language of people's mother tongue.
OmegaWiki is still in need of more functionality. This is something that we try to achieve in any way possible. It is however rewarding to notice what difference the existing functionality makes.
I am very happy with what we have achieved so far. It bodes well for the future :)
Thanks,
GerardM
Monday, April 09, 2007
Fall back languages
OmegaWiki has a relevant bug fix; it is now obvious what option you add a part of speech you will see "part of speech" in your own language and when we do not have it, you will get it English as the fall back language. This is a huge improvement from having the same text in some 24 other languages.
The next improvement will become possible when the Multilingual MediaWiki is finished; this will allow you to select the languages you are interested in. These languages are more personal than just falling back to English.
Another option would be to identify a fall back language for a language itself. This makes sense for languages that exist in a space where another language is well known and better supported in MediaWiki. Languages like French, German, Italian, Spanish, Portuguese and Mandarin come to mind.. I am sure there are more..
For the Incubator we want to define fall back languages to make it easier to localise the MediaWiki messages.. We hope to get answers what the best choice is.. otherwise we will have to guess..
Thanks,
GerardM
The next improvement will become possible when the Multilingual MediaWiki is finished; this will allow you to select the languages you are interested in. These languages are more personal than just falling back to English.
Another option would be to identify a fall back language for a language itself. This makes sense for languages that exist in a space where another language is well known and better supported in MediaWiki. Languages like French, German, Italian, Spanish, Portuguese and Mandarin come to mind.. I am sure there are more..
For the Incubator we want to define fall back languages to make it easier to localise the MediaWiki messages.. We hope to get answers what the best choice is.. otherwise we will have to guess..
Thanks,
GerardM
Friday, March 30, 2007
Protecting people from poisonous personalities
The BBC website has an article about the culture of abuse on line that allows people to be noxious, abusive, threatening to people and in this instance to women. The thing that sparked this off are death threats to a prominent blogger, Ms Kathy Sierra. I read a follow up article on oReillynet called "Open season on women".
We have been extremely lucky and fortunate on OmegaWiki so far with abuse and vandalism. We have also been extremely lucky with our community. I do not exaggerate when I say that we have some great females being extremely relevant to what we do. If both articles need an official reaction, it would be that we will not tolerate such things on OmegaWiki and, that I am still appalled by the way the Wikichix were driven away from the Wikimedia Foundation for all the "right" reasons.
At this stage in the life cycle of OmegaWiki, I can say that people who are deemed to be poisonous in what they do will find that there is little toleration for them. When we get to the stage where it becomes an issue if we should tolerate poisonous people, I can tell you know that I am prepared to fight tooth and nail to get such people out.
Thanks,
GerardM
We have been extremely lucky and fortunate on OmegaWiki so far with abuse and vandalism. We have also been extremely lucky with our community. I do not exaggerate when I say that we have some great females being extremely relevant to what we do. If both articles need an official reaction, it would be that we will not tolerate such things on OmegaWiki and, that I am still appalled by the way the Wikichix were driven away from the Wikimedia Foundation for all the "right" reasons.
At this stage in the life cycle of OmegaWiki, I can say that people who are deemed to be poisonous in what they do will find that there is little toleration for them. When we get to the stage where it becomes an issue if we should tolerate poisonous people, I can tell you know that I am prepared to fight tooth and nail to get such people out.
Thanks,
GerardM
Saturday, March 24, 2007
More new functionality & some stats
Lies, damned lies and statistics.. I always thought that OmegaWiki had more English words than German words. This turns out not to be the case, then I thought it may be because we make something English (United Kingdom) or English (United States) when the word is not shared between the two.. There are only 326 UK words; that does not fill the 441 gap.
These and more pondering are possible because of our new statistics functionality we have courtesy of Zdenek Broz. It shows you the actual numbers of Expressions in OmegaWiki.
Some more factoids, we now support 130 languages, 28 of these have less than 10 words.. When you compare us to Wiktionary, we would be the fourth in size when an article is considered the equivalent of an Expression or we would we the twenty-fifth in size when a DefinedMeaning is considered in this way. When you consider that all of our growth so far has been autonomously, I think we are doing well.
In the mean time, Alexa considers us the 480,601 website without statistics for the week.. I do not know what goes on there. I do like their T-shirts though.
To quote the Alexa T-shirts: "Will dance for better Alexa rank " :)
Thanks,
GerardM
These and more pondering are possible because of our new statistics functionality we have courtesy of Zdenek Broz. It shows you the actual numbers of Expressions in OmegaWiki.
Some more factoids, we now support 130 languages, 28 of these have less than 10 words.. When you compare us to Wiktionary, we would be the fourth in size when an article is considered the equivalent of an Expression or we would we the twenty-fifth in size when a DefinedMeaning is considered in this way. When you consider that all of our growth so far has been autonomously, I think we are doing well.
In the mean time, Alexa considers us the 480,601 website without statistics for the week.. I do not know what goes on there. I do like their T-shirts though.
To quote the Alexa T-shirts: "Will dance for better Alexa rank " :)
Thanks,
GerardM
Thursday, March 22, 2007
New functionality
OmegaWiki has new functionality, there are two bits of functionality. One bit is not obvious but really helpful; a function to (re)build the indexes of the database. What it does is ensure that the indexes are build in the right order. This improves performance considerably.
The other bit, is much more spectacular. It sorts the tables in the HTML based on the content of the first column. The order depends very much on the language selected in the user preferences. This leads me to my question; is it sorted well when your language is Persian, Arab or Hebrew. These are right to left languages and I have no clue if the sorting is done well.
Thanks,
GerardM
The other bit, is much more spectacular. It sorts the tables in the HTML based on the content of the first column. The order depends very much on the language selected in the user preferences. This leads me to my question; is it sorted well when your language is Persian, Arab or Hebrew. These are right to left languages and I have no clue if the sorting is done well.
Thanks,
GerardM
Monday, March 19, 2007
How do the four freedoms apply to one database?
The FSF defines four freedoms when it comes to software. What kind of freedoms applies to a database like OmegaWiki.
The OmegaWiki software is licensed, like MediaWiki what it is an integral part of, with a GPL license. This means that you can use the software as is. As OmegaWiki uses specific data structures that can be licensed separately for completeness sake, the database design is also available under a GPL license.
The data that is contained in OmegaWiki is licensed under a combined GFDL/CC-by license. Many people insist that these licenses are not compatible. At issue is that the data are just facts, it is only possible to copyright facts as a collection. We want people to make use of our collection. For us success is: "when people find a use for our data we did not think off".
We invite people to collaborate on our data, when they enter some Babel templates on their user page, we give them edit rights to the data. We invite organisations to collaborate on our data because there is so much data that organisations can share, there is so much labour invested in the type of data OmegaWiki can be a home for.
So the data and the software is Free. How about OmegaWiki itself ..
As there is only one OmegaWiki.org, the room to do whatever is limited. The data has to be useful to everyone and it has to fit in with the notion of the DefinedMeaning. When domain specific data is added, it needs to be domain specific and, there has to be agreement that this data provides a suitable extension for people involved in this domain.
When people find that this is not enough, they can have their own database. This does not mean that they cannot cooperate. Much of what they need in terms of extra functionality will be shared. It means that even when the OmegaWiki database is forked, there is still plenty of scope to improve all the things we do agree on and collaborate on those.
There is plenty of Freedom. All the Freedoms we provide. I think however we achieve the most success when we find that there is more that binds us than that drives us apart. It means that we have to work hard in understanding what our common needs are.
Thanks,
GerardM
The OmegaWiki software is licensed, like MediaWiki what it is an integral part of, with a GPL license. This means that you can use the software as is. As OmegaWiki uses specific data structures that can be licensed separately for completeness sake, the database design is also available under a GPL license.
The data that is contained in OmegaWiki is licensed under a combined GFDL/CC-by license. Many people insist that these licenses are not compatible. At issue is that the data are just facts, it is only possible to copyright facts as a collection. We want people to make use of our collection. For us success is: "when people find a use for our data we did not think off".
We invite people to collaborate on our data, when they enter some Babel templates on their user page, we give them edit rights to the data. We invite organisations to collaborate on our data because there is so much data that organisations can share, there is so much labour invested in the type of data OmegaWiki can be a home for.
So the data and the software is Free. How about OmegaWiki itself ..
As there is only one OmegaWiki.org, the room to do whatever is limited. The data has to be useful to everyone and it has to fit in with the notion of the DefinedMeaning. When domain specific data is added, it needs to be domain specific and, there has to be agreement that this data provides a suitable extension for people involved in this domain.
When people find that this is not enough, they can have their own database. This does not mean that they cannot cooperate. Much of what they need in terms of extra functionality will be shared. It means that even when the OmegaWiki database is forked, there is still plenty of scope to improve all the things we do agree on and collaborate on those.
There is plenty of Freedom. All the Freedoms we provide. I think however we achieve the most success when we find that there is more that binds us than that drives us apart. It means that we have to work hard in understanding what our common needs are.
Thanks,
GerardM
Saturday, March 10, 2007
Managing data with some SQL
The great thing about OmegaWiki is that the data is in a database. You might say that this is not that special, every wiki uses a database. Today, we have as a first time done some curation on the data; everywhere where a word in en-US was written exactly the same as in English, we have deleted the English. One example is the word "competition", in the history you will find the deletions.
I am really grateful that Leftmost has started to use SQL to fix things for us. It saves us what is most valuable; the time of our editors.
There are other things that we can do, I have asked to have all Bulgarian words that are capitalised changed to lower case where the Russian words are lower case. This is to fix something that is done consistently this way in the GEMET database. With these improvements, the GEMET data becomes usable for other purposes; things like data mining .. :)
Thanks,
GerardM
I am really grateful that Leftmost has started to use SQL to fix things for us. It saves us what is most valuable; the time of our editors.
There are other things that we can do, I have asked to have all Bulgarian words that are capitalised changed to lower case where the Russian words are lower case. This is to fix something that is done consistently this way in the GEMET database. With these improvements, the GEMET data becomes usable for other purposes; things like data mining .. :)
Thanks,
GerardM
Sunday, March 04, 2007
About reputation and education
Wikipedia has a new scandal. There are several issues here and, several interested parties. The scandal is about someone who is known by a nickname and who claimed that he had academic qualifications. He claimed to be a doctor of theology.
One of the interested parties who obviously relished this occasion was Dr Sanger of Citizendium fame. For Dr Sanger it is one of those occasions where he can sing the praises of his project. He did, I did not read anything new there. In the mean time, his project was announced in Nature and now, so many months later, there is still nothing to be seen. That does wonders for the credibility of that project ...
Given that scientific credentials will be very relevant on OmegaWiki, I have given it some thought. When people want to claim professional credentials, they would have to provide us at least with their real name and their e-mail address. These would be the requirements on the level of OmegaWiki.
For the Wikis for professionals, different rules might apply. When medical credentials are claimed, wrong information can kill. It is for such reasons that much more identifiable information is likely to be required. This does not mean that credentials are as important as Dr Sanger says. Relevant is the quality of the information provided. Relevant is the role a person plays in the community. This means that a person can become relevant in OmegaWiki by building a reputation. For this you do not need scientific credentials. Science has, like every part of society, its fair share of miserable people and I am sure Citizendium will learn that as well.
If this incident is to be a lesson, the lesson is that it does not pay to assume credentials and have a virtual reality meet real life. It is therefore sad that Essjay is now "retired". If this means that the person who did a lot of good work will do no more, than it is a sad outcome. If this is the only time that the issue of assumed credentials raises its ugly head, it has a silver lining. If this incident is sufficiently public that it will stop this phenomena, then I will be glad.
Thanks,
GerardM
One of the interested parties who obviously relished this occasion was Dr Sanger of Citizendium fame. For Dr Sanger it is one of those occasions where he can sing the praises of his project. He did, I did not read anything new there. In the mean time, his project was announced in Nature and now, so many months later, there is still nothing to be seen. That does wonders for the credibility of that project ...
Given that scientific credentials will be very relevant on OmegaWiki, I have given it some thought. When people want to claim professional credentials, they would have to provide us at least with their real name and their e-mail address. These would be the requirements on the level of OmegaWiki.
For the Wikis for professionals, different rules might apply. When medical credentials are claimed, wrong information can kill. It is for such reasons that much more identifiable information is likely to be required. This does not mean that credentials are as important as Dr Sanger says. Relevant is the quality of the information provided. Relevant is the role a person plays in the community. This means that a person can become relevant in OmegaWiki by building a reputation. For this you do not need scientific credentials. Science has, like every part of society, its fair share of miserable people and I am sure Citizendium will learn that as well.
If this incident is to be a lesson, the lesson is that it does not pay to assume credentials and have a virtual reality meet real life. It is therefore sad that Essjay is now "retired". If this means that the person who did a lot of good work will do no more, than it is a sad outcome. If this is the only time that the issue of assumed credentials raises its ugly head, it has a silver lining. If this incident is sufficiently public that it will stop this phenomena, then I will be glad.
Thanks,
GerardM
Wednesday, February 28, 2007
What is a Word?
First, let me make it perfectly clear that this is a discussion that has raged for centuries. I know full well that everybody has his or her own opinion on this matter, and that I am not going to resolve this issue today. This is an overview and a bit of personal opinion, as relates to online dictionaries.
Intuitively, we all know the answer. A word is a unit of language conveying some meaning. But how do we decide what is a real word? We look in a dictionary, of course. What do we do if we're writing a dictionary?
We are caught between cataloging what is "right" (prescriptivism) and what is actually done (descriptivism). The pendulum has lately swung towards descriptivism, and I would say that there are some good reasons for that trend. The language that is spoken on the streets is not the same language that is written in academia. Somebody learning a language may genuinely need help sorting out the less proper terms in it.
Take, for instance, colloquialisms such as "irrespective" and "humongous", and all the phrases that have gotten squished together into amalgams like "gotcha" and "woulda". Most people would readily agree that these words do not belong in a college thesis paper.
What is the scrupulous lexicographer to do? Fortunately, it is not a strict either-or question, especially in a work not substantially limited by size. In an electronic resource, we can put them in, anyway. To satisfy the formal sorts, the perscriptivists, we can then place a prominent usage note in the entry, explaining just why a writer might wish to use caution with the term: ginormous is a colloquial term, regarded by many to be something less than a proper word. Thus, the reader is both informed and cautioned.
That's fine for most of the slang and jargon, but we have another problem. People keep making up new words. My sister in law coined the term "muskaroon" to mean generically any small, furry creature that scurries past too quickly to identify. Squirrels, chipmunks, gophers, and presumably rabbits would all qualify. So we have a unit of language with a symbol and a meaning. The trouble is, if you walked up to people on the street and inquired whether there were muskaroons in the area, nobody would be able to answer who hadn't talked lately to my sister in law, and that is a small minority of people, indeed.
The test here is usage. Can we demonstrate that the word is in common use? Now, depending on the character of the dictionary, we can define the rules various ways. Was it used by so many independent sources? Did anybody important (such as Shakespeare or a prominent academic journal) publish the word?
Generally, we also try to find and present examples of the term in what is called "running text". That means that it is in a paragraph, and isn't only used as somebody's nickname, say. The edge is still a fuzzy one. Are the citations in traditional print sources, such and books and journals, or are they sprinkled in a couple blogs and forums? Was the word used in only one limited context, or in a variety of sources and over a period of years? These sorts of tests can help to weed out many of the more questionable entries. At some point, though, it may yet come down to a judgment call, if not on whether a word is real, then on how to apply the rules. In these cases, I advise the users of a dictionary to bring a healthy dose of skepticism with them, to recall that even dictionaries are not infallible, and to trust at the very least that these decisions are made by real people who care for the project.
If, knowing all that, you find you don't like the way "they" are running the place, you are invited to do a better job.
Intuitively, we all know the answer. A word is a unit of language conveying some meaning. But how do we decide what is a real word? We look in a dictionary, of course. What do we do if we're writing a dictionary?
We are caught between cataloging what is "right" (prescriptivism) and what is actually done (descriptivism). The pendulum has lately swung towards descriptivism, and I would say that there are some good reasons for that trend. The language that is spoken on the streets is not the same language that is written in academia. Somebody learning a language may genuinely need help sorting out the less proper terms in it.
Take, for instance, colloquialisms such as "irrespective" and "humongous", and all the phrases that have gotten squished together into amalgams like "gotcha" and "woulda". Most people would readily agree that these words do not belong in a college thesis paper.
What is the scrupulous lexicographer to do? Fortunately, it is not a strict either-or question, especially in a work not substantially limited by size. In an electronic resource, we can put them in, anyway. To satisfy the formal sorts, the perscriptivists, we can then place a prominent usage note in the entry, explaining just why a writer might wish to use caution with the term: ginormous is a colloquial term, regarded by many to be something less than a proper word. Thus, the reader is both informed and cautioned.
That's fine for most of the slang and jargon, but we have another problem. People keep making up new words. My sister in law coined the term "muskaroon" to mean generically any small, furry creature that scurries past too quickly to identify. Squirrels, chipmunks, gophers, and presumably rabbits would all qualify. So we have a unit of language with a symbol and a meaning. The trouble is, if you walked up to people on the street and inquired whether there were muskaroons in the area, nobody would be able to answer who hadn't talked lately to my sister in law, and that is a small minority of people, indeed.
The test here is usage. Can we demonstrate that the word is in common use? Now, depending on the character of the dictionary, we can define the rules various ways. Was it used by so many independent sources? Did anybody important (such as Shakespeare or a prominent academic journal) publish the word?
Generally, we also try to find and present examples of the term in what is called "running text". That means that it is in a paragraph, and isn't only used as somebody's nickname, say. The edge is still a fuzzy one. Are the citations in traditional print sources, such and books and journals, or are they sprinkled in a couple blogs and forums? Was the word used in only one limited context, or in a variety of sources and over a period of years? These sorts of tests can help to weed out many of the more questionable entries. At some point, though, it may yet come down to a judgment call, if not on whether a word is real, then on how to apply the rules. In these cases, I advise the users of a dictionary to bring a healthy dose of skepticism with them, to recall that even dictionaries are not infallible, and to trust at the very least that these decisions are made by real people who care for the project.
If, knowing all that, you find you don't like the way "they" are running the place, you are invited to do a better job.
Labels:
dictionary,
nonce,
protologism,
verification,
wiki,
wiktionary
Monday, February 19, 2007
Anyone may edit
Written in response to this project.
Somebody asked me today about what happens to a dictionary when anybody can edit it. As anybody who has ever edited a wiki knows, the openness is a mixed blessing.
It is a great thing, because many hands make light work. Dictionaries need to be every bit as large as the languages they catalog, so the process of gathering and maintaining the data is a huge one. As we start to add translations between languages, rather than simply defining a term, that task becomes orders of magnitude bigger. To capture all words in all languages is something that will take nothing less than a wiki and a worldwide community. It is a monumental task, but in a wiki, we can conceive of creating a resource on such a scale.
It is a great thing because so many regions and cultures can be represented. An American may understand most of the English spoken in South Africa or New Zealand, but both of those regions have slang all their own. Chile speaks Spanish far differently than Spain. All those variants can have their place.
It is a great thing because a wiki can evolve with a language. New terms come into use all the time, and a freely editable electronic resource is not limited in its capacity to store data or to accommodate a large, diverse set of editors.
The big trouble is this: if just anybody can edit, how on earth do we know it is right? I'd like to explore a few approaches here. Of course, any of these approaches could be considered as a barrier to entry, but these things are always trade-offs.
Somebody asked me today about what happens to a dictionary when anybody can edit it. As anybody who has ever edited a wiki knows, the openness is a mixed blessing.
It is a great thing, because many hands make light work. Dictionaries need to be every bit as large as the languages they catalog, so the process of gathering and maintaining the data is a huge one. As we start to add translations between languages, rather than simply defining a term, that task becomes orders of magnitude bigger. To capture all words in all languages is something that will take nothing less than a wiki and a worldwide community. It is a monumental task, but in a wiki, we can conceive of creating a resource on such a scale.
It is a great thing because so many regions and cultures can be represented. An American may understand most of the English spoken in South Africa or New Zealand, but both of those regions have slang all their own. Chile speaks Spanish far differently than Spain. All those variants can have their place.
It is a great thing because a wiki can evolve with a language. New terms come into use all the time, and a freely editable electronic resource is not limited in its capacity to store data or to accommodate a large, diverse set of editors.
The big trouble is this: if just anybody can edit, how on earth do we know it is right? I'd like to explore a few approaches here. Of course, any of these approaches could be considered as a barrier to entry, but these things are always trade-offs.
- Appoint trusted users to do the housekeeping. These are the sysops, administrators, bureaucrats, the librarians, or the janitors, depending on your point of view. These somebodies keep watch and undo the damage that some of the just-anybodies can do. If somebody writes an article containing typical vandalism, such as "asdfasdf" or "Dave is a dork!", an administrator can delete or undo it. Much vandalism is so predictable that even a bot can detect and remove it. Unfortunately, a select group of administrators, however well trusted or well-read, cannot be everywhere at once, and they cannot know everything. Things get missed, even with a checklist system such as patrolled edits. It is likewise impossible for an administrator or small group of administrators to know everything. Misinformation, intentional or otherwise, is not so easy to spot as out-and-out nonsense.
- Hold people accountable. Articles have histories, so you can see who did what. Even pseudonymous users develop reputations. Anonymous users tend to attract the most scrutiny. An active, healthy wiki often develops into a meritocracy, with leaders having sway (though not necessarily authority) based on reputation, seniority, and trust in the community. This effect generally works to improve content, but even a well-known, trusted user may make mistakes. If he or she is trusted well enough, there is a risk that an error or oversight may go unnoticed.
- Allow anybody and everybody to scrutinize and correct or flag the content. The process is not foolproof, especially in larger projects, but wikis have a remarkable capacity for self-cleaning. Of course, this approach can tend to result in a sort of groupthink effect: if enough people believe it, then it must be so.
- Demand credentials. Don't just let in any old riffraff. Wikipedia has clearly shown the power of amateurs and volunteers to create great content, but it is certainly possible to limit the users in a project, or part of the project, to a certain group. This approach is most appropriate to a wiki serving a closed community, such as a professional or academic group, especially one dedicated to a particularly narrow or specialized topic.
- Make the messes behind the scenes, and publish only the good stuff, with some review process. The German Wikipedia published a paper book containing selected articles. Online, there have been proposals for a "Stable Versions" system, where a mature article would be reviewed and locked, and any additional changes would go through a separate editing or discussion page.
- Demand references. There is a movement within Wikipedia to reference the articles and the claims made in them. In the context of a dictionary, references may be other dictionaries. Is the word recognized by RAE or OED (whom we trust to have done the requisite homework)? They may be other works about words. Or, they may be citations. Citations are quotations including the word in question. They show context and provide evidence that the word is or was in use. Of course, we must still question the validity of the evidence. Are the 400 Google hits because somebody prolific uses that nonsense word as a handle? Is a word more valid if it was used by a blogger or two, or by Thornton Wilder? Is an etymology known with reasonable certainty or is it apocryphal? Depending on the size and resources of the wiki, efforts to verify and reference articles may be systematic, or they may be requested when a given entry or fact is questioned.
Friday, February 16, 2007
An article in Nature ...
I am absolutely thrilled with the article that was published last Wednesday in Nature... It is a great article and it explains really well what we hope to achieve with relational data in MediaWiki. The only thing that is a bit sad is, that you have to pay $30 for the privilege of reading it.
The article is great, and what makes it special is the great presentation that Knewco has created to explain what we hope to achieve; this demo available at wikiprofessional.info. It presents some really impressive figures; it indicates the work done to integrate several important resources of the bio-medical domain, the numbers involved.
For me the most important point is that this is likely to be a very important stimulus to the Open Access movement. It indicates that it is possible to bring what was divided together. It allows people to work with the terminology of their field and also add data that is very specific. Information that goes much further than what was envisioned in what was once called the "Ultimate Wiktionary".
The whole notion of a resource that because of its roots already merged lexicology, terminology and ontology is really special. With the integration of such specialised data from different domains like the bio-medical, another really interested experiment will be under way when the data gets imported and merged. There is a nascent community for the bio-medical domain and, it will find that it will co-exist with the existing OmegaWiki community.
Both communities have everything to gain from collaboration; much of what the existing OmegaWiki community cares about will be seen as a fringe benefit. On the other hand, the translations that exist for concepts like malaria will prove to be of value when scientific articles are considered that were not published in English.
I am convinced that a bright future is ahead of us. We have this vision of what may come, I wish I could look into the future and see what it will be like. :)
Thanks,
GerardM
The article is great, and what makes it special is the great presentation that Knewco has created to explain what we hope to achieve; this demo available at wikiprofessional.info. It presents some really impressive figures; it indicates the work done to integrate several important resources of the bio-medical domain, the numbers involved.
For me the most important point is that this is likely to be a very important stimulus to the Open Access movement. It indicates that it is possible to bring what was divided together. It allows people to work with the terminology of their field and also add data that is very specific. Information that goes much further than what was envisioned in what was once called the "Ultimate Wiktionary".
The whole notion of a resource that because of its roots already merged lexicology, terminology and ontology is really special. With the integration of such specialised data from different domains like the bio-medical, another really interested experiment will be under way when the data gets imported and merged. There is a nascent community for the bio-medical domain and, it will find that it will co-exist with the existing OmegaWiki community.
Both communities have everything to gain from collaboration; much of what the existing OmegaWiki community cares about will be seen as a fringe benefit. On the other hand, the translations that exist for concepts like malaria will prove to be of value when scientific articles are considered that were not published in English.
I am convinced that a bright future is ahead of us. We have this vision of what may come, I wish I could look into the future and see what it will be like. :)
Thanks,
GerardM
Sunday, February 11, 2007
Why compete when you can collaborate ?
All words of all languages of the world.. that is what we eventually aim to include in OmegaWiki. This aim is of such a magnitude that you have to be certifiable to come up with such a project. The functional design for the project includes much more; everything including the kitchen sink..
When everything is to be included in one project, it is easy to suggest that people contribute to the project. When the project includes everything why have another?
In an Open Source / Open Content environment this is not necessarily how it works. Why should the others be seen as competitors? They do their own thing, sure. You may want to achieve the same thing, also true. It is however much possible to find the synergy between projects. This way you can build on each others accomplishments.
The Shtooka project is something I learned about the other day. The one thing it does really well is the way they make recording pronunciations easy. You can record a string of words and it will save them for you one at a time.
Wiktionarians saw this and they are working an upload facility so that it will also be saved automatically to Commons. I warned that the files should not only be saved as .ogg files. In order to make sure they are relevant for scientists there should also be a .wav file. The current thinking is that the flac file format will work as well and the benefit is that it provides a loss-less compression. To make sure that this is the case, the praat software, software that is also available under a GPL license, was analysed and it was considered that it is easy to incorporate this flac file format.
People from effectively five different communities are now working together. It will be even possible to include links to OmegaWiki in the Shtooka meta data. This will be possible even though both projects do their own thing. Both the data and the functionality can be shared.
I may be certifiable, but this kind of collaboration is awesome and, it is why there may be method to this madness.. :)
Thanks,
GerardM
When everything is to be included in one project, it is easy to suggest that people contribute to the project. When the project includes everything why have another?
In an Open Source / Open Content environment this is not necessarily how it works. Why should the others be seen as competitors? They do their own thing, sure. You may want to achieve the same thing, also true. It is however much possible to find the synergy between projects. This way you can build on each others accomplishments.
The Shtooka project is something I learned about the other day. The one thing it does really well is the way they make recording pronunciations easy. You can record a string of words and it will save them for you one at a time.
Wiktionarians saw this and they are working an upload facility so that it will also be saved automatically to Commons. I warned that the files should not only be saved as .ogg files. In order to make sure they are relevant for scientists there should also be a .wav file. The current thinking is that the flac file format will work as well and the benefit is that it provides a loss-less compression. To make sure that this is the case, the praat software, software that is also available under a GPL license, was analysed and it was considered that it is easy to incorporate this flac file format.
People from effectively five different communities are now working together. It will be even possible to include links to OmegaWiki in the Shtooka meta data. This will be possible even though both projects do their own thing. Both the data and the functionality can be shared.
I may be certifiable, but this kind of collaboration is awesome and, it is why there may be method to this madness.. :)
Thanks,
GerardM
Thursday, February 08, 2007
Become an OmegaWiki developer
OmegaWiki is now running the latest version of the MediaWiki software used by Wikipedia. This is a major milestone, as it also makes it a lot easier for anyone to join in the fun of developing the open source OmegaWiki/Wikidata software. To give credit where credit is due, these are the people who have contributed to the code so far:
- Peter-Jan Roes
- Karsten Uil
- Sean Burke
- Rod A. Smith (sticky tree expansion via cookies)
- Ævar Arnfjörð Bjarmason (namespace code installer)
- Charles Pritchard (Multilingual MediaWiki development, ongoing)
- Jelte Zeilstra (untranslated meaning script, under review)
- Zdenek Broz (statistical scripts, under review)
- Paa-Kwesi Imbeah (Wikimedia Commons support, under review)
- Marc Carmen (TBX export, incomplete)
- myself
Wednesday, January 31, 2007
Greek languages
At OmegaWiki, we saw that Lou started to change the capitalisation of language names... A few days ago I was surprised that the Georgian names for languages were incorrectly spelled. Now it is Greek.
It is really powerful to see that by having the languages corrected, it will be available for everybody who wants to know about Greek. This reason for using OmegaWiki proves itself again.
Thanks,
GerardM
It is really powerful to see that by having the languages corrected, it will be available for everybody who wants to know about Greek. This reason for using OmegaWiki proves itself again.
Thanks,
GerardM
Sunday, January 28, 2007
Latin roots etc.
Well yesterday one thing came into mind - a dictionary a teacher of mine at the language school had. It was a dictionary that listed Latin words with many translations into other languages and one thing is obvious: all these words of course were similar in all languages. If you knew one of them and studied the other language it would have been easy to create the relative words following a set of rules for most of them.
So one thing should be obvious: to insert these words with their translations into OmegaWiki ... but well, there is one problem with Latin - the "normal" Latin language should not be mixed with the taxonomical Latin that is used in science ... so we need to create two languages: Latin and taxonomical Latin ... who knows if the relative language codes exist somewhere in the ISO 639 standards.
Technorati: language, Latin, ISO 639, standards, translation, dictionary
So one thing should be obvious: to insert these words with their translations into OmegaWiki ... but well, there is one problem with Latin - the "normal" Latin language should not be mixed with the taxonomical Latin that is used in science ... so we need to create two languages: Latin and taxonomical Latin ... who knows if the relative language codes exist somewhere in the ISO 639 standards.
Technorati: language, Latin, ISO 639, standards, translation, dictionary
Labels:
dictionary,
ISO 639,
language,
Latin,
standards,
translation
Friday, January 26, 2007
OLPC needs a dictionary viewer
I had a word with the director of content for the OLPC, the One Laptop Per Child Project. As you know OmegaWiki is the project that works on providing the OLPC with dictionary content. We are working on all these words, and while we are making steady progress, there is so much still left to do. We are getting more Expressions in many languages, the definitions are lagging and while we do our best, it is still very much the difference between there being nothing and there being next to nothing. It does however show that things are getting under way...
As the moment when kids are exposed to the systems is drawing closer, it is relevant that the data can be used. So we need a dictionary viewer. It needs to run on Linux and, it should have a small footprint. As we will provide all these languages, it will be interesting to see how the rich tapestry that OmegaWiki tries to weave will materialise on these nifty systems.
When you have a suggestion, please let us know :)
Thanks,
GerardM
As the moment when kids are exposed to the systems is drawing closer, it is relevant that the data can be used. So we need a dictionary viewer. It needs to run on Linux and, it should have a small footprint. As we will provide all these languages, it will be interesting to see how the rich tapestry that OmegaWiki tries to weave will materialise on these nifty systems.
When you have a suggestion, please let us know :)
Thanks,
GerardM
Wednesday, January 24, 2007
Georgian names for languages
In the past I got permission to copy content from a resource with the names of languages. I am still grateful for the data. It got the Dutch Wiktionary going really nicely and, as we needed at the time those names of languages for the user interface.
With OmegaWiki we had the same issue; we needed language names again for the user interface. This was to make it possible for people to see the labels of translations in their own language. From the moment the data became available we have learned a lot, for instance that language in languages like Danish and Italian do not capitalise the names of languages.
Today I was told that many of the names of languages in Georgian were found to be in error and had been corrected. The great news for OmegaWiki is, that we only have to do this once and it is good everywhere. The sad thing is that it is probably wrong in many, many Wiktionaries. There were two types of errors; it was just wrong or it was the name of someone from a country in stead of the name of the language.
The best I can do for the Wiktionaries is notify in this way as I do not really now what needs doing.
Thanks,
GerardM
With OmegaWiki we had the same issue; we needed language names again for the user interface. This was to make it possible for people to see the labels of translations in their own language. From the moment the data became available we have learned a lot, for instance that language in languages like Danish and Italian do not capitalise the names of languages.
Today I was told that many of the names of languages in Georgian were found to be in error and had been corrected. The great news for OmegaWiki is, that we only have to do this once and it is good everywhere. The sad thing is that it is probably wrong in many, many Wiktionaries. There were two types of errors; it was just wrong or it was the name of someone from a country in stead of the name of the language.
The best I can do for the Wiktionaries is notify in this way as I do not really now what needs doing.
Thanks,
GerardM
Tuesday, January 23, 2007
Stichting Open Progress
Stichting Open Progress is the Dutch not for profit organisation that is the legal organisation behind OmegaWiki. As OmegaWiki is growing to the extend where we have to consider contracts for hosting, grants and the like, we had a need for an organisation.
The need for an organisation was also felt as we already had some projects where we would have been better able to do things when there was a legal entity backing up the activities. Some of these projects are quite substantial.
Open Progress aims to develop both Open Source/Free Software and Open Content/Free Content projects. As part of its mission it gives room for projects that are aligned with the aims of the stichting. Obviously OmegaWiki is the first; from an organisational point of view, the OmegaWiki commission decides on the issues that arise. Resolution will be enacted for the project by the stichting provided they are in line with the Dutch law and, provided they do not circumvent the aims of the stichting. This way Open Progress hopes to make OmegaWiki a safe haven where people and organisations work in the understanding that the aims of the project will be respected.
There are two websites for OpenProgress; in line with the experiences of the Wikimedia Foundation, we have both an internal and an external wiki. The internal will use Semantic MediaWiki to leverage as much as possible the information that we will include. As the information will include both personal information and confidential project information, the internal will be invite only.
Thanks,
Gerard Meijssen
voorzitter Stichting Open Progress
The need for an organisation was also felt as we already had some projects where we would have been better able to do things when there was a legal entity backing up the activities. Some of these projects are quite substantial.
Open Progress aims to develop both Open Source/Free Software and Open Content/Free Content projects. As part of its mission it gives room for projects that are aligned with the aims of the stichting. Obviously OmegaWiki is the first; from an organisational point of view, the OmegaWiki commission decides on the issues that arise. Resolution will be enacted for the project by the stichting provided they are in line with the Dutch law and, provided they do not circumvent the aims of the stichting. This way Open Progress hopes to make OmegaWiki a safe haven where people and organisations work in the understanding that the aims of the project will be respected.
There are two websites for OpenProgress; in line with the experiences of the Wikimedia Foundation, we have both an internal and an external wiki. The internal will use Semantic MediaWiki to leverage as much as possible the information that we will include. As the information will include both personal information and confidential project information, the internal will be invite only.
Thanks,
Gerard Meijssen
voorzitter Stichting Open Progress
Monday, January 15, 2007
Destinazione Italia
Destinazione Italia is a project of the University of Bamberg. It provides training for people learning an advanced level of Italian. Bamberg is a German University and many of its students are German. Many of the students do have a different mother tongue. Learning a third language based on the knowledge of a second language is less effective than learning based on the knowledge of the mother tongue.
I am really proud to announce that OmegaWiki has been selected by the University of Bamberg as the platform that will host the lexicological information for "Destinazione Italia". The initial phase of the project will create a lot of Italian based DefinedMeanings. In the second phase we will translate these words to English, German and Spanish. The third phase is to find translations in as many other languages as we can get.
Research done by Zdenek Broz learned, that when the combination of quality translations of German, English and Spanish is found, it will allow the inclusion of translations of other languages when these translations are shared in a different resource. According to Zdenek's figures this will get us an accuracy of around, probably better than 95%.
There is a budget to get us many translations in other languages. The sweet thing is, when we are able to provide quality translations, the budget can be used for other things. This can be to improve the OmegaWiki usability, it can also be to spend money on a language that is not part of the initial list of languages "Destinazione Italia" supports.
The challenge is therefore, how much can we do with a limited budget. What will be the added value of creating content in a Wiki environment. When will OmegaWiki reach the tipping point where collaboration in OmegaWiki is the obvious thing to do, "Destinazione Italia" will help us reach that point. :)
Thanks,
GerardM
I am really proud to announce that OmegaWiki has been selected by the University of Bamberg as the platform that will host the lexicological information for "Destinazione Italia". The initial phase of the project will create a lot of Italian based DefinedMeanings. In the second phase we will translate these words to English, German and Spanish. The third phase is to find translations in as many other languages as we can get.
Research done by Zdenek Broz learned, that when the combination of quality translations of German, English and Spanish is found, it will allow the inclusion of translations of other languages when these translations are shared in a different resource. According to Zdenek's figures this will get us an accuracy of around, probably better than 95%.
There is a budget to get us many translations in other languages. The sweet thing is, when we are able to provide quality translations, the budget can be used for other things. This can be to improve the OmegaWiki usability, it can also be to spend money on a language that is not part of the initial list of languages "Destinazione Italia" supports.
The challenge is therefore, how much can we do with a limited budget. What will be the added value of creating content in a Wiki environment. When will OmegaWiki reach the tipping point where collaboration in OmegaWiki is the obvious thing to do, "Destinazione Italia" will help us reach that point. :)
Thanks,
GerardM
Subscribe to:
Posts (Atom)