Kipcool had written an improvement on the Expression needing translation special page. The page now shows the number of expressions that exist in one language and have no translation in the other language.
I have had fun playing with it, you find that there are words in Hindi that have no equivalent in Dutch.
Thanks,
GerardM
Wednesday, August 19, 2009
Wednesday, July 22, 2009
Multi lingual support
When you support languages like OmegaWiki does, then there are two levels where this support is provided.
All in all, the latest major upgrade to OmegaWiki has been an important and most necessary boost to our code base. It adds a new sparkle to our project.
- Localisation of the user interface
- Localisation of the data
The complete user interface is now in Russian. We will be able to do even better in the future. We will implement the "LocalisationUpdate" extension; this allow us to continuously adopt all relevant new localisations as committed to the code repository from translatewiki.net.
Making OmegaWiki a resource that is continuously updated for its multi language support is a dream come true. Really important in this has been the perseverence of RobertL. He got to grips with the complicated even convoluted code that makes OmegaWiki. His work included a lot of sanitising of the code base, this will make it not that hard for us to upgrade in the future.
One other upgrade is planned; this is the implementation of the Babel extension. We will be saying goodbye to all the current templates. They have served us well but they have been an absolute pain to maintain. Smaller project like OmegaWiki are best served by getting the localisations for this functionality from one central place. It makes thousands of templates redundant and, new localisations will now be provided on a daily basis through the LocalisationUpdate.
All in all, the latest major upgrade to OmegaWiki has been an important and most necessary boost to our code base. It adds a new sparkle to our project.
Thanks,
GerardM
Monday, July 20, 2009
Upgrading the system
OmegaWiki is expected to go completely offline on Monday 20th July at 14:00 UTC for a major upgrade lasting several hour
Friday, June 19, 2009
ISBN 978-3-03911-799-4
Last year I spoke at a conference in Aarhus, Denmark at the Centre for Lexicography at the Aarhus school of business. I was asked to write an article about what I had to say. I did, and with pleasure I received a complimentary copy in the post titled "Lexicography at a Crossroads".
Sadly the publication is not published as an Open Access work so you will have to find a copy when you want to read my essay "The Philosophy behind OmegaWiki and the Visions for the Future". There is always the presentation that I gave at the conference..
Thanks,
GerardM
Sadly the publication is not published as an Open Access work so you will have to find a copy when you want to read my essay "The Philosophy behind OmegaWiki and the Visions for the Future". There is always the presentation that I gave at the conference..
Thanks,
GerardM
Wednesday, May 06, 2009
WOTD publicatie
The word of the day is publicatie. The reason is, that a book with articles by the people who presented has been published. My article is about the philosphy behind OmegaWiki. Now I will have to figure out how to update my profile at wikiprofessional...Thanks,
GerardM
Tuesday, March 31, 2009
Ambaradan
OmegaWiki as a project works; concepts are added and we call them DefinedMeanings, we add Expressions to them and they are either synonyms or translations. We can add part of speech information, we can refer to Wikipedia articles or Commons pictures. All these things we can do in the language selected in your user interface.OmegaWiki does all these things but there is a problem; the software as it is, is convoluted. A programmer new to the code does not like it. Some hate it with a passion and some looked at it and walked away. This is not good. This is one reason why not much happened with OmegaWiki, the other reason is that we started the development of OmegaWiki mark II.
We created a proof of concept that demonstrated that we can provide multi lingual support to Commons and the next bit was that on the basis of the database backend we would create a new front end, it would be OmegaWiki mark II.
Bèrto went off the grid, it was not possible to contact him in a normal way and now, many months later he appears to have written something called Ambaradan. At this moment there is documentation what it is supposed to do. This documentation is very intriguing and I think that it may even work.
Obviously Ambaradan is welcome to the OmegaWiki content and once there is a user interface to the data, I will be interested to learn how it presents the data. What I will be looking for is how you can configure what information can be entered for a language and how you can relate information entered in one language to information in another.
There are several projects I am involved in that are anxious to learrn Ambaradan's potential. Many things have been on hold and for several projects alternatives are being looked at..
Ambaradan has surfaced, it may be great. At this stage there is not enough to go on.
Thanks,
GerardM
Wednesday, February 18, 2009
A new user ,,, many new languages
A new user came to OmegaWiki and he wanted to add content in the Mayan languages. All of them "except that funny one from veracruz which only recently got classified as Mayan". All of them because this new user comes with an existing dictionary and this has "all of them".
Now the Mayan languages are not easily classified; The way Ethnologue had it is not how the ISO-639-3 has it at this time. This means that we will have to carefully understand what the relation is between the languages in the dictionaries and the codes maintained by SIL.
Thanks,
GerardM
Now the Mayan languages are not easily classified; The way Ethnologue had it is not how the ISO-639-3 has it at this time. This means that we will have to carefully understand what the relation is between the languages in the dictionaries and the codes maintained by SIL.
Thanks,
GerardM
Monday, February 16, 2009
OmegaWiki has moved
The hosting of OmegaWiki is on a server by Knewco. Knewco has moved its servers from one location to another and as a consequence OmegaWiki would have been off line for a couple of days. This was not a good plan.A friend of mine, Tom Maaswinkel, provided us with temporary server space. Kim Bruning, another friend, has done the migration. At this moment, the DNS has been changed and we are testing the new server.
Everything should be working smoothly again.
Thanks,
GerardM
PS this is the 100th blog entry :)
Monday, January 26, 2009
What a difference a day makes
When you compare this screen shot with the one in my previous blog entry, you will agree that the right to left support has improved quite a lot.
One concern that has been raised is what the MediaWiki developers will say about this. It is obvious that there will be many people who get a different view and consequently the cache will be affected. My idea is that as this is about improved functionality, cache efficiency is secondary to a very large extend.
Thanks,
GerardM
One concern that has been raised is what the MediaWiki developers will say about this. It is obvious that there will be many people who get a different view and consequently the cache will be affected. My idea is that as this is about improved functionality, cache efficiency is secondary to a very large extend.
Thanks,
GerardM
Right, to the left
OmegaWiki is a multi lingual environment. We have been proud to support your language. When your language is not supported, you can ask. When your language is not properly supported, you can make the difference.
There was one thing though.. Languages like Arabic or Hebrew should go right to left. They did not.
Thanks,
GerardM
There was one thing though.. Languages like Arabic or Hebrew should go right to left. They did not.
Thanks,
GerardM
Sunday, January 25, 2009
artigo na Wikipédia
When you change the language in your user preferences, the presentation of the labels of the OmegaWiki data will change as well. It will tell you for instance where you can find the Wikipedia article in a given language.The phrase "artigo na Wikipédia" is todays "Word of the day". Yesterday this translation was added and consequently the user interface for Portuguese became more complete.
Typically phrases like this can be found here, and yes you can make a difference too.
Thanks,
GerardM
Friday, January 23, 2009
Malafaya had an itch
OmegaWiki suffered for quite some time from a missing functionality. Its annotations had gone away. This was really frustrating. Especially because nobody was looking after the maintenance of the OmegaWiki mark I database. This annoyance became too much and Malafaya started to look into the code. He found and fixed the problem.
Apparently it felt really good, and as there was this other annoyance, he had another look at the code. The trick will be to get the code in SVN. Once it is in SVN, updating OmegaWiki should not be a problem.
Thanks,
GerardM
Apparently it felt really good, and as there was this other annoyance, he had another look at the code. The trick will be to get the code in SVN. Once it is in SVN, updating OmegaWiki should not be a problem.
Thanks,
GerardM
Sunday, December 21, 2008
Words I do not know
I read and write a lot of English. Every now and again I come across a word, an idiom I do not know. When I have the time I add them to OmegaWiki. When I look at the ones I added today, fecklessness, valedictory, fricassee and enunciate, I wonder if these are the kind of words other people are looking for as well.
Thanks,
GerardM
Thanks,
GerardM
Friday, November 28, 2008
Usage
Some words that you look up in a dictionary, have a particular usage associated with it that is not always apparant. The word Balkans for instance is always preceded with "the" when it is used as a noun. The same is true for the Ivory Coast..
Thanks,GerardM
This new type of annotation was requested because people wanted to indicate the difference in usage between words that were otherwise synonymous.
Thanks,
Wednesday, November 05, 2008
Spanish at 30.000 expressions
Today I am proud to announce that Ascander added the thirty thousandth Spanish expression in OmegaWiki.Thanks,
GerardM
Sunday, November 02, 2008
The beauty of words
Last Friday [[nl:Afshin Afkari]], presented his Dutch Persian idiom dictionary. The publisher, Amsterdam University Press organised a dialogue with Ashin and Hugo Brandt Corstius in Spui27 in Amsterdam titled "the beauty of words". Going to presentations like this is something that I do rarely. It was however great fun particularly because of the people that you meet at such occasions.
When you read a book like this, you can enjoy a great print. It really looks good. In a book that contains both Latin script and the Perso-Arabic script, it takes more effort to achieve this. One of the things that Afshin included was the transscription of the Persia texts in Eurofarsi. This seems to me a smart move as it a great aid for people who learn Persian. Eurofarsi I am told is used quite a lot by people on the Internet and by people who SMS.
The exposure to Afshin's book made me think on how I would include idiom in OmegaWiki. When you have an idiom in a language, it has a meaning. The question is how do you deal with translations. When you think of it, there are two issues: there is idiom with the same meaning and, there is the need for a literal translation of the text.
It seems to me that you should treat idiom different from normal lexical content. Leaving it a DefinedMeaning without translations but with annotations to equivalent idiom, a defintion of the meaning and literal translations of the idiom itself. Idiom often has a few key words, eg in "blood is thicker then water" you would like to refer to both blood and water. Controry to what is usual in OmegaWiki, you want to refer to these concepts in the same language as the idiom.
The question is, how to support this in OmegaWiki mark II.
Thanks,
GerardM
When you read a book like this, you can enjoy a great print. It really looks good. In a book that contains both Latin script and the Perso-Arabic script, it takes more effort to achieve this. One of the things that Afshin included was the transscription of the Persia texts in Eurofarsi. This seems to me a smart move as it a great aid for people who learn Persian. Eurofarsi I am told is used quite a lot by people on the Internet and by people who SMS.
The exposure to Afshin's book made me think on how I would include idiom in OmegaWiki. When you have an idiom in a language, it has a meaning. The question is how do you deal with translations. When you think of it, there are two issues: there is idiom with the same meaning and, there is the need for a literal translation of the text.
It seems to me that you should treat idiom different from normal lexical content. Leaving it a DefinedMeaning without translations but with annotations to equivalent idiom, a defintion of the meaning and literal translations of the idiom itself. Idiom often has a few key words, eg in "blood is thicker then water" you would like to refer to both blood and water. Controry to what is usual in OmegaWiki, you want to refer to these concepts in the same language as the idiom.
The question is, how to support this in OmegaWiki mark II.
Thanks,
GerardM
Thursday, October 30, 2008
Fourtythousand DefinedMeanings

Having 40.000 DefinedMeanings is a nice milestone. The graphic shows how the number of expressions per language tails off.
Congratulations to all the people who contributed to this success. :)
Thanks,
GerardM
Tuesday, October 28, 2008
French has over 20K expressions
I expect that with 20k French starts to become relevant as a resource in OmegaWiki. The good news for our users is the amount of translations a concepts is connected to. Currently the ratio is 9,0402 Expressions per DefinedMeaning.A big thank you to everyone who contributed to the French success :)
Thanks,
GerardM
Tuesday, October 14, 2008
MySQL Workbench Beta program
SUN, is working on improving the MySQL workbench designer. We have been using workbench for OmegaWiki mark II and have been asking questions and making suggestions. It is with some satisfaction, that we find that we have been invited to the MySQL workbench beta program.
Currently our software uses some 30 business rules and these amount to some 6000 lines of code. They model a complex dynamic behaviour that is responsible for most on the "work behind the scenes". These triggers are currently not modelled by workbench, that can contain them, but not describe their behaviour from the analyst's point of view.
We are happy and satisfied being part of this program.
Thanks,
GerardM
Currently our software uses some 30 business rules and these amount to some 6000 lines of code. They model a complex dynamic behaviour that is responsible for most on the "work behind the scenes". These triggers are currently not modelled by workbench, that can contain them, but not describe their behaviour from the analyst's point of view.
We are happy and satisfied being part of this program.
Thanks,
GerardM
Monday, October 13, 2008
A Commons that supports multiple languages
Commons is great, it is the biggest resource of freely usable media files, some 3.363.734 of them. They are a great resource when you are looking for pictures and when you know English. If you do not know English, it might as well not exist.
Given that the Wikimedia Foundation cares about its educational value for everyone, it makes sense for Commons to be used by people who read and write other languages as well. This dream of supporting all the other people as well is an old one; I wrote about it as early as May 2005.
With support of the Digitale Pioniers, we have been enabled to create a proof of concept project. This project does demonstrate that we can do this. The languages that are currently supported are only a few but sufficient to demonstrate the principles.
The issue is that it is for the Wikimedia Foundation to show an interest. There is a growing group of people who have seen that it can do. When I give presentations about this people indicate that this is a "must have" feature. I will talk again at the Wikimedia Conferentie NL and I hope to discuss with the WMF soon what it takes to support multiple languages at Commons but more importantly if it wants to.
Thanks,
Gerard
Given that the Wikimedia Foundation cares about its educational value for everyone, it makes sense for Commons to be used by people who read and write other languages as well. This dream of supporting all the other people as well is an old one; I wrote about it as early as May 2005.
With support of the Digitale Pioniers, we have been enabled to create a proof of concept project. This project does demonstrate that we can do this. The languages that are currently supported are only a few but sufficient to demonstrate the principles.
The issue is that it is for the Wikimedia Foundation to show an interest. There is a growing group of people who have seen that it can do. When I give presentations about this people indicate that this is a "must have" feature. I will talk again at the Wikimedia Conferentie NL and I hope to discuss with the WMF soon what it takes to support multiple languages at Commons but more importantly if it wants to.
Thanks,
Gerard
Wednesday, September 17, 2008
Wikiprofessional is doing Portuguese
One function of Wikiprofessional that I really like is the Concept Web Linker. If you have not played with it, you should. What I am really happy with is that the Concept Web Linker is getting into other languages. I have an example of this..

This is the screen dump.
Thanks,
GerardM

This is the screen dump.
Thanks,
GerardM
Thursday, September 11, 2008
Commons but now multi lingual
This is the category "Felis silvestris catus" as it is on Commons with a twist. This screen makes best sense when you can read Dutch.. Obviously, the English text is still there but that does only helps when you understand that language.Thanks,
GerardMThursday, August 28, 2008
OmegaWiki goes Squeak
OmegaWiki was implemented in MediaWiki as an extension during its first iteration. At that time this was a great decision. It provided us with a lot of great functionality and it was the environment that we knew. MediaWiki is a wonderful collaborative application. In OmegaWiki mark II we are looking for other things; this has led to making OmegaWiki independent from MediaWiki.
OmegaWiki intended to do everything in one database. This meant that it was problematic to use our data for other purposes. We also had a situation where different applications wanted their specific data included in the database and needed control of the data involved. This could not be done in the first iteration of OmegaWiki so we had to rethink our ways.
As data may be connected and often will be shared, a peering model is required. Data will be used by many applications and consequently the data needs to be provided in a way that allows for many applications. For the data we will have one interface that uses a standard XML interface to provide the data. This will allow for many applications to use the same data and it will prevent the mixing of data and user interface elements as we saw in the past.
There will no longer be only one database; there will be many. The central or “global” database will provide basic information that will be CC-by licensed. This will allow anybody to have their own “regional” or “local” data and refer to shared concepts for instance for sharing or mapping purposes.
Regional and local databases can be licensed and maintained in a manner that makes sense to the people involved. They can include data that is not really acceptable from a pure linguistic point of view, for instance MALARIA as a synonym for malaria. They can include all kinds of relations between concepts, relations that may be really specialised or that require particular validations before they are published.
Another application that we have always wanted to give our data was was the OLPC and equivalent networks. This means that our database has to be able to function stand alone and be synchronised and share improved data when a connection becomes available.
As a consequence of all these considerations, we have been looking for the best technology that will serve our purpose; MySQL 5.1 will provide us with new functionality that makes an important difference. Squeak is a programming platform that we think will provide us with the tools to build the rich environment that we dream of having.
As the OLPC project also uses Squeak, it will allow us to bring our information to this great educational project, and in return we hope that people will find OmegaWiki an environment to contribute their Squeak work to and help us build dictionaries in the many languages spoken and written where the OLPC will become available.
Thanks,
GerardM
OmegaWiki intended to do everything in one database. This meant that it was problematic to use our data for other purposes. We also had a situation where different applications wanted their specific data included in the database and needed control of the data involved. This could not be done in the first iteration of OmegaWiki so we had to rethink our ways.
As data may be connected and often will be shared, a peering model is required. Data will be used by many applications and consequently the data needs to be provided in a way that allows for many applications. For the data we will have one interface that uses a standard XML interface to provide the data. This will allow for many applications to use the same data and it will prevent the mixing of data and user interface elements as we saw in the past.
There will no longer be only one database; there will be many. The central or “global” database will provide basic information that will be CC-by licensed. This will allow anybody to have their own “regional” or “local” data and refer to shared concepts for instance for sharing or mapping purposes.
Regional and local databases can be licensed and maintained in a manner that makes sense to the people involved. They can include data that is not really acceptable from a pure linguistic point of view, for instance MALARIA as a synonym for malaria. They can include all kinds of relations between concepts, relations that may be really specialised or that require particular validations before they are published.
Another application that we have always wanted to give our data was was the OLPC and equivalent networks. This means that our database has to be able to function stand alone and be synchronised and share improved data when a connection becomes available.
As a consequence of all these considerations, we have been looking for the best technology that will serve our purpose; MySQL 5.1 will provide us with new functionality that makes an important difference. Squeak is a programming platform that we think will provide us with the tools to build the rich environment that we dream of having.
As the OLPC project also uses Squeak, it will allow us to bring our information to this great educational project, and in return we hope that people will find OmegaWiki an environment to contribute their Squeak work to and help us build dictionaries in the many languages spoken and written where the OLPC will become available.
Thanks,
GerardM
Labels:
MediaWiki,
new functionality,
OLPC,
OmegaWiki mark II
Tuesday, June 17, 2008
Fast and furious
I just learned that OmegaWiki mark II has been updated; it now allows for multiple statements. Bèrto indicated that he will update the documentation to reflect this.. The example code I have does not get posted because the Blogger software wants to interpret it ...
I am sure that it will find its place in the documentation ... So, it is for you to RTFM :)
Thanks,"
GerardM
I am sure that it will find its place in the documentation ... So, it is for you to RTFM :)
GerardM
OmegaWiki mark II
It has been relatively silent around OmegaWiki, this silence however is very deceptive. Much hard work has gone into preparing for a new codebase for OmegaWiki. The new code has to deliver new functionality in order to justify the huge investment in time and money.
Why change..
At the bottom you find a document explaining about the API; it is written in the Open Office format..
Enjoy, have fun and tell us what you think of it. The data is experimental so we will replace it with more data at a later time.
Thanks,
Gerard
Why change..
- The existing user interface and the database routines are very much etwined, we want to separate them
- We want to provide services based on the OmegaWiki data, we will provide an XML interface into the data
- The underlying technology has changed a lot; we are using the bleeding edge of the MySQL database
- Our data can be used for applications; we will provide a way to separate data that is of general interest and data that is not
- For some applications it is really helpful when the data can be used in an off-line environment; we will provide a way of synchronising databases
- There is more ...
At the bottom you find a document explaining about the API; it is written in the Open Office format..
Enjoy, have fun and tell us what you think of it. The data is experimental so we will replace it with more data at a later time.
Thanks,
Gerard
Sunday, June 08, 2008
Internationalisation and localisation
When you write software, when the software is to be used by people who speak many languages, internationalisation is a key requirement. It is the precursor to localisation; the changes made by the localisers to support their language.
MediaWiki is really good at localisation and, it is being perfected all the time. When software is written and when internationalisation / localisation is not considered from the start, it is quite a job to get this right.
The OmegaWiki Vocabulary Trainer needed internationalisation / localisation badly. The first issue with the software; it was not German. The software was paid by the University of Bamberg to be used by its students. These students are expected to know German and this trainer is a tool to aid them to learn languages. OmegaWiki is very much dedicated to making information available in many languages and not considering internationalisation in its associated software is ... odd.
It is with relief that I can now announce that the Vocabulary trainer is supported in Betawiki. I am grateful to the kind developers at Betawiki who made this possible. You can localise for your language here ...
Thanks,
GerardM
MediaWiki is really good at localisation and, it is being perfected all the time. When software is written and when internationalisation / localisation is not considered from the start, it is quite a job to get this right.
The OmegaWiki Vocabulary Trainer needed internationalisation / localisation badly. The first issue with the software; it was not German. The software was paid by the University of Bamberg to be used by its students. These students are expected to know German and this trainer is a tool to aid them to learn languages. OmegaWiki is very much dedicated to making information available in many languages and not considering internationalisation in its associated software is ... odd.
It is with relief that I can now announce that the Vocabulary trainer is supported in Betawiki. I am grateful to the kind developers at Betawiki who made this possible. You can localise for your language here ...
Thanks,
GerardM
Tuesday, June 03, 2008
Getting a better look and feel ..
The vocabulary trainer was introduced... It looked awful; it now has been improved considerably. Like we wrote last time, it is open source and it is being worked on. One of the things on the "must have" details .... localisation :) 
Thanks,
GerardM
Thanks,
GerardM
Thursday, May 29, 2008
Publish early, publish often ...
For Open Source software, one of the mantras is to publish early and publish often. In this spirit I am happy to announce the first version of a vocabulary trainer on OmegaWiki.
This software uses the content of OmegaWiki; so the quality of the exercise is also determined by the amount of terminology that is contained in OmegaWiki.. The current functionality is not yet feature complete.. One of the things we want to do with this software is have a list of words and phrases that are good to know when you go abroad.. Wikimania 2008 anyone ??
NB the vocabulary trainer can be accessed from the OmegaWiki main page :)
Thanks,
GerardM
This software uses the content of OmegaWiki; so the quality of the exercise is also determined by the amount of terminology that is contained in OmegaWiki.. The current functionality is not yet feature complete.. One of the things we want to do with this software is have a list of words and phrases that are good to know when you go abroad.. Wikimania 2008 anyone ??
NB the vocabulary trainer can be accessed from the OmegaWiki main page :)
Thanks,
GerardM
Wednesday, April 16, 2008
Georgian
Georgian is a language that started with no content in OmegaWiki. It is spoken by some 4.4 million people mainly in Georgia, Turkey, Iran and Russia. It is written in the Georgian alphabet.
Statistics are in and off themselves not that relevant but they do allow you to tell a story. OmegaWiki started with the content of GEMET, the GEneral Multilingual Environmental Thesaurus, the languages that are part of this resource have a head start in size.
Today thanks to the hard work of Sopho, Georgian is the first language that grew bigger then one of the languages supported by Gemet. The OmegaWiki statistics show that Japanese might be the next language to grow bigger then Slovenian..
Because of the beautiful characters, Georgian is my favourite example of showing the value of the localisation in OmegaWiki. It is really special to see the same content optimised in such a way. :)
Thanks,
GerardM
Statistics are in and off themselves not that relevant but they do allow you to tell a story. OmegaWiki started with the content of GEMET, the GEneral Multilingual Environmental Thesaurus, the languages that are part of this resource have a head start in size.
Today thanks to the hard work of Sopho, Georgian is the first language that grew bigger then one of the languages supported by Gemet. The OmegaWiki statistics show that Japanese might be the next language to grow bigger then Slovenian..
Because of the beautiful characters, Georgian is my favourite example of showing the value of the localisation in OmegaWiki. It is really special to see the same content optimised in such a way. :)
Thanks,
GerardM
Saturday, April 05, 2008
Unicode 5.1
Today I learned that Unicode 5.1 has been released. The information that I received informs me that one major feature will be of particular relevance to Japanese, Chinese and Korean texts by enabling ideographic variation sequences. The linebreaking for Polish and Portuguese hyphenation has been improved. The Indic languages will be happy with improved text segmentation algorithm.
There are 1624 new encoded characters, this includes characters required for Malayam and Myanmar but there are also new characters for the Latin script. New is support for the Cham, Lepcha, Ol Chiki, Rejang, Saurashtra, Sundanese, and Vai scripts.
For the techies, the collation algorithms have been updated to include all the new characters. This has also an effect on contractions like the ch in the Slovak language.
Many of these things have an effect on languages supported in Wikimedia projects. My question is when will we have support for this. Is this a function of the MediaWiki / PHP code and is it also a function of the browser ??
Thanks,
GerardM
There are 1624 new encoded characters, this includes characters required for Malayam and Myanmar but there are also new characters for the Latin script. New is support for the Cham, Lepcha, Ol Chiki, Rejang, Saurashtra, Sundanese, and Vai scripts.
For the techies, the collation algorithms have been updated to include all the new characters. This has also an effect on contractions like the ch in the Slovak language.
Many of these things have an effect on languages supported in Wikimedia projects. My question is when will we have support for this. Is this a function of the MediaWiki / PHP code and is it also a function of the browser ??
Thanks,
GerardM
Tuesday, March 18, 2008
OpenStreetMap terminology
OpenStreetMap for those who do not know the project, is a free editable map of the whole world. Its data is freely licensed, it is build by volunteers and it is very much a work in progress.
OpenStreetMap has its own terminology, and given its origin it is British. The maps can be quite good providing you with sufficient information to plan your route.. There are routes for pubcrawls and other innovations :)
As OpenStreetMap intends to provide a map of the world, the people that make these maps have to literally put themselves on the map. In the Netherlands we have been blessed because maps have been made available to the project. A friend of mine is not yet on the map..
It is clear that the terminology used in Italy is not the same as in the UK or the Netherlands. It makes sense for an Italian to have Italian terminology available to him and Dutch would do me nicely.
At a meeting of the Digitale Pioniers, I met people of OpenStreetMap and we agreed to include their terminology in OmegaWiki. I have now entered the words that have the key "highway" and invite you all to come to OmegaWiki to add translations in your language.
Special in this data is that i have added the definitions provided by OpenStreetMap as alternate definitions as well. In this way they have been marked as definitions provided by OpenStreetMap.
All the OpenStreetMap terminology in OmegaWiki can be found here.
Thanks,
GerardM
As OpenStreetMap intends to provide a map of the world, the people that make these maps have to literally put themselves on the map. In the Netherlands we have been blessed because maps have been made available to the project. A friend of mine is not yet on the map..
At a meeting of the Digitale Pioniers, I met people of OpenStreetMap and we agreed to include their terminology in OmegaWiki. I have now entered the words that have the key "highway" and invite you all to come to OmegaWiki to add translations in your language.
Special in this data is that i have added the definitions provided by OpenStreetMap as alternate definitions as well. In this way they have been marked as definitions provided by OpenStreetMap.
All the OpenStreetMap terminology in OmegaWiki can be found here.
Thanks,
GerardM
Tuesday, March 11, 2008
Less is more
There are two types of content in OmegaWiki; there are the standard MediaWiki namespaces with the standard MediaWiki content and there is the OmegaWiki specific content. The specific content contain database records, they are the DefinedMeanings and Expressions.
As they are handled completely differently, it means that the functionality available to these types of data is different as well. Particularly the "move" functionality has given us problems in the past. This broken functionality has now been removed from the screens.
Thanks,
GerardM
As they are handled completely differently, it means that the functionality available to these types of data is different as well. Particularly the "move" functionality has given us problems in the past. This broken functionality has now been removed from the screens.
Thanks,
GerardM
Firefox 3 beta 4 is significantly faster
OmegaWiki is a web application that requires a lot from a system. I am really happy to report that the latest beta of Firefox gives me a significantly better performance on the same hardware. I reported in the past that the Firefox beta did a good job for me, but this time the performance is noticeably faster on a page like Nederland.
Firefox provides cutting edge technology and really makes a big difference to me. Now when you want OmegaWiki to perform even better, you can choose to move to this latest software. The other good news is that the spell checkers that are relevant to me are now available as well...
Thanks,
GerardM
Firefox provides cutting edge technology and really makes a big difference to me. Now when you want OmegaWiki to perform even better, you can choose to move to this latest software. The other good news is that the spell checkers that are relevant to me are now available as well...
Thanks,
GerardM
Saturday, March 01, 2008
Nagios
OmegaWiki has a technical problem; there are certain records that have problems and that crash the database. There are solutions to this problem and some are being tested at the moment.
This week I was in Bamberg at the Otto-Friedrich University, and discussed this with Martin Mai. He has now build a monitor that checks if OmegaWiki is still alive. This works fine. We now have permission for the Bamberg Nagios service to run a script when OmegaWiki is no longer alive.
This will solve one of my biggest worries which is the availability of the OmegaWiki service.
Thanks,
GerardM
This week I was in Bamberg at the Otto-Friedrich University, and discussed this with Martin Mai. He has now build a monitor that checks if OmegaWiki is still alive. This works fine. We now have permission for the Bamberg Nagios service to run a script when OmegaWiki is no longer alive.
This will solve one of my biggest worries which is the availability of the OmegaWiki service.
Thanks,
GerardM
Wednesday, January 16, 2008
Bounties for the localisation of MediaWiki
Hoi,
The "Stichting Open Progress" is happy to announce that it has received a grant from Hivos, to improve the localisation of MediaWiki. Open Progress is going to offer a bounty of up to 200 EURO for the full localisation for a language. Given the activity of Hivos, a Dutch NGO, a bounty will only be available for languages in Asia, Africa and Latin America that have a sizable number of speakers.
With this project we hope to achieve that MediaWiki can indeed claim to be one of the best Open Source projects that provides great localisation for many many languages out of the box. It will improve the usability not only for the WMF projects in that language but also for projects like Wikimedia Commons, Wikieducator, Wikihow, OmegaWiki .. the list goes on ..
The budget we have is substantial but limited. We will be sadly happy when we have to announce that we have ran out of money. Sad because we want to localise more languages, happy because so many languages will have been improved.
For the precise details of the project I refer to the details on Betawiki.
NB The amounts are inclusive of the money transfer costs and as these can be substantial, we offer some alternatives. In the past I have proposed a scheme called "Donations, putting your money where your mouth is". In this scheme you choose to get paid or donate the money for one of the projects that we advertise under this scheme. Another way of getting the money paid is when people in a country agree to work together and have us pay the money together as well. This could for instance work in the case of Wikimedia India...
Thanks,
GerardM
The "Stichting Open Progress" is happy to announce that it has received a grant from Hivos, to improve the localisation of MediaWiki. Open Progress is going to offer a bounty of up to 200 EURO for the full localisation for a language. Given the activity of Hivos, a Dutch NGO, a bounty will only be available for languages in Asia, Africa and Latin America that have a sizable number of speakers.
With this project we hope to achieve that MediaWiki can indeed claim to be one of the best Open Source projects that provides great localisation for many many languages out of the box. It will improve the usability not only for the WMF projects in that language but also for projects like Wikimedia Commons, Wikieducator, Wikihow, OmegaWiki .. the list goes on ..
The budget we have is substantial but limited. We will be sadly happy when we have to announce that we have ran out of money. Sad because we want to localise more languages, happy because so many languages will have been improved.
For the precise details of the project I refer to the details on Betawiki.
NB The amounts are inclusive of the money transfer costs and as these can be substantial, we offer some alternatives. In the past I have proposed a scheme called "Donations, putting your money where your mouth is". In this scheme you choose to get paid or donate the money for one of the projects that we advertise under this scheme. Another way of getting the money paid is when people in a country agree to work together and have us pay the money together as well. This could for instance work in the case of Wikimedia India...
Thanks,
GerardM
Tuesday, January 01, 2008
2008, International Year of Languages
The new year, 2008 has been designated the International Year of Languages by UNESCO. Countries and organisations are invited to participate, and I think that what we do in our projects qualifies us as participants. Our projects are relevant for many languages, we welcome new languages and we provide help and infrastructure to make it a success. Our communities are knowledge societies in which everyone can participate and benefit. We promote universal access to information and ensure in this way the use of an increasing number of languages.
Siebrand wrote an overview of the localisation of MediaWiki. Real support is provided to some 170 languages but only 47 have a minimal localisation. In a way it mirrors our projects; our projects do great in some languages while at the other end language versions are closed because there is not enough of a community that supports them.
The Wikimedia Foundation is becoming more mature; we aim to ensure that all our projects do well. Barriers to entry have been in place for new languages and it has led to improved localisation of MediaWiki. When new projects are finally approved, they are already of a size both in articles and participants that there is less need for anxiety for their future.
When we are to participate in the Year of Languages and continue to do what we do well and improve where we are weak, our projects will prove to be credible participants in this Year of Languages. Participating will give our Wiki way more credibility and it will give us access to people and organisations that can help us in fulfilling our goal; sharing the sum of all knowledge with every single human being.
Thanks,
GerardM
Siebrand wrote an overview of the localisation of MediaWiki. Real support is provided to some 170 languages but only 47 have a minimal localisation. In a way it mirrors our projects; our projects do great in some languages while at the other end language versions are closed because there is not enough of a community that supports them.
The Wikimedia Foundation is becoming more mature; we aim to ensure that all our projects do well. Barriers to entry have been in place for new languages and it has led to improved localisation of MediaWiki. When new projects are finally approved, they are already of a size both in articles and participants that there is less need for anxiety for their future.
When we are to participate in the Year of Languages and continue to do what we do well and improve where we are weak, our projects will prove to be credible participants in this Year of Languages. Participating will give our Wiki way more credibility and it will give us access to people and organisations that can help us in fulfilling our goal; sharing the sum of all knowledge with every single human being.
Thanks,
GerardM
Sunday, December 30, 2007
Numbers in a translation dictionary
In the Wikimedia Foundation, there are the inclusionists and the exclusionists. Some are of the opinion that certain topics should not be included in Wikipedia while others do. The most brilliant example is the inclusion of all the busstops in Japan. Someone took the effort to describe them and people find them actually useful.
Some people are enamoured by constructed languages and spend a lot of time making such languages their own. I personally have had dealings with at least three people that speak Volapük, and I know people that go to congresses because they meet people who speak Esperanto. Many constructed languages have more speakers than many natural languages (that are not yet extinct).
For exclusionists it is not palatable when constructed languages do well. There are always "good" reasons why those others need to be excluded. Hidden in the discussion about the "radical cleanup of the Volapük Wikipedia" is a discussion about the inclusion of numerals like 588 in the Limburgian Wiktionary and as you can imagine "it is not good".
In a translation dictionary there are reasons to include numbers. The point is that they are not written the same in all scripts. OmegaWiki has its fair share of numbers, and it did not address this issue.
By adding a new class, we now allow the representation of Arab and Roman numbers, and hundred has now as an annotation both 100 and C. In this way we do not have a separate page for each numerical representation.
Thanks,
GerardM
Some people are enamoured by constructed languages and spend a lot of time making such languages their own. I personally have had dealings with at least three people that speak Volapük, and I know people that go to congresses because they meet people who speak Esperanto. Many constructed languages have more speakers than many natural languages (that are not yet extinct).
For exclusionists it is not palatable when constructed languages do well. There are always "good" reasons why those others need to be excluded. Hidden in the discussion about the "radical cleanup of the Volapük Wikipedia" is a discussion about the inclusion of numerals like 588 in the Limburgian Wiktionary and as you can imagine "it is not good".
In a translation dictionary there are reasons to include numbers. The point is that they are not written the same in all scripts. OmegaWiki has its fair share of numbers, and it did not address this issue.
By adding a new class, we now allow the representation of Arab and Roman numbers, and hundred has now as an annotation both 100 and C. In this way we do not have a separate page for each numerical representation.
Thanks,
GerardM
Tuesday, December 18, 2007
Languages ...
Given a fixed point and you can move the world. In many ways, your language provides you with the tools to wrap your mind around the world, express its essence in a way that may be understood by the people you communicate with and help you to shape the world as you know it.
Language is both individual and shared. My English has been shaped by my schooling in the Netherlands, my stay in the United Kingdom and the many times I expressed myself in e-mail, articles, presentations and when skyping. Typically I get it right but a text can be understood by some and misunderstood by others. It is my English but to function it has to be expressed in a way that is shared with others.
For some subjects I prefer my native language, for others I prefer English. To communicate, the language used must be sufficiently shared by everyone involved. The language must be received; when I am surrounded by a vacuum, nobody will hear me talk. To read this blog, you either have access to a computer or someone must print it out for you.
A language lives when people use it, when it is part of a distinct community, a distinct culture. When the boundaries around such a community or culture disappear 0r change, the language either morphs or it dies. To understand history, you have to understand its artefacts and its language. Many languages die and died and with it we lose the history, the culture of the people that spoke that language. They may leave their literature, inscriptions and when enough is left, we may understand what is says. The trick will be to understand it as it was meant when the language, the culture was alive.
Thanks,
GerardM
Language is both individual and shared. My English has been shaped by my schooling in the Netherlands, my stay in the United Kingdom and the many times I expressed myself in e-mail, articles, presentations and when skyping. Typically I get it right but a text can be understood by some and misunderstood by others. It is my English but to function it has to be expressed in a way that is shared with others.
For some subjects I prefer my native language, for others I prefer English. To communicate, the language used must be sufficiently shared by everyone involved. The language must be received; when I am surrounded by a vacuum, nobody will hear me talk. To read this blog, you either have access to a computer or someone must print it out for you.
A language lives when people use it, when it is part of a distinct community, a distinct culture. When the boundaries around such a community or culture disappear 0r change, the language either morphs or it dies. To understand history, you have to understand its artefacts and its language. Many languages die and died and with it we lose the history, the culture of the people that spoke that language. They may leave their literature, inscriptions and when enough is left, we may understand what is says. The trick will be to understand it as it was meant when the language, the culture was alive.
Thanks,
GerardM
Saturday, December 15, 2007
Whe needs birthPlace
With some regularity I try to better understand Semantic Web and associated subjects. I find it hard going but also a compulsive subject. When you express the relation "Johan Cruijff" "birthPlace" "Amsterdam", it is understandable to you as a reader but for humans it should read like "Johan Cruijf was born in Amsterdam" or "Johan Cruijf werd geboren in Amsterdam" .. This magical statement "birthPlace" can be interpreted when you know your English otherwise it is truly for machines only.
OmegaWiki does express relations, you will find for instance that Amsterdam is the capital of the Netherlands. In essence it is expressed as a triple, but it is expressed in natural language and depending on the existence of a translation, you will read the relation in the language selected as your user preference.
How to combine what we do and what happens elsewhere, my latest idea is based in the RDF tag; "birthPlace". It is a construct that obviously needs a natural language equivalent and this is what OmegaWiki can provide. A method is needed to connect the two. In order to function, birthPlace has a very precise definition and this definition must be part of a collection of such definitions. These labels need to be linked to OmegaWiki DefinedMeanings as the identifier for an OmegaWiki collection.
To make this useful, an external application needs to call a function that provides the translation to a specified language. How to combine this with the notion of an URN I have not figured out yet.
Thanks,
GerardM
OmegaWiki does express relations, you will find for instance that Amsterdam is the capital of the Netherlands. In essence it is expressed as a triple, but it is expressed in natural language and depending on the existence of a translation, you will read the relation in the language selected as your user preference.
How to combine what we do and what happens elsewhere, my latest idea is based in the RDF tag; "birthPlace". It is a construct that obviously needs a natural language equivalent and this is what OmegaWiki can provide. A method is needed to connect the two. In order to function, birthPlace has a very precise definition and this definition must be part of a collection of such definitions. These labels need to be linked to OmegaWiki DefinedMeanings as the identifier for an OmegaWiki collection.
To make this useful, an external application needs to call a function that provides the translation to a specified language. How to combine this with the notion of an URN I have not figured out yet.
Thanks,
GerardM
Thursday, December 13, 2007
Eastern Yiddish
Eastern Yiddish, is one of the two varieties of Yiddish that have been recognised as languages in their own right in the ISO-639-3 (ydd). in OmegaWiki, we now have our first 213 Expressions in this language and I am impressed with the amount of work that has gone into it; most have annotations indicating hyphenation and the pronunciation using IPA notation.
Eastern Yiddish is not one of the languages supported by MediaWiki, and the mechanism for showing localised content is connected to the language selected in the "User Preferences". I have been given some help from Siebrand what files need to be changed and added. Kim helped me with doing it for the first time and now the first localisation is visible for Eastern Yiddish.
The MediaWiki localisation itself uses Yiddish as the fall back language so the experience is pretty good for now. What Siebrand indicated is that is is possible to include languages like Eastern Yiddish in the BetaWiki. This would create stubs that are of benefit to OmegaWiki. It would prepare for the moment when people start localising in earnest.
I think it would be a good thing, but I am interested to learn what other people think.
Thanks,
GerardM
Eastern Yiddish is not one of the languages supported by MediaWiki, and the mechanism for showing localised content is connected to the language selected in the "User Preferences". I have been given some help from Siebrand what files need to be changed and added. Kim helped me with doing it for the first time and now the first localisation is visible for Eastern Yiddish.
The MediaWiki localisation itself uses Yiddish as the fall back language so the experience is pretty good for now. What Siebrand indicated is that is is possible to include languages like Eastern Yiddish in the BetaWiki. This would create stubs that are of benefit to OmegaWiki. It would prepare for the moment when people start localising in earnest.
I think it would be a good thing, but I am interested to learn what other people think.
Thanks,
GerardM
Monday, December 03, 2007
Supporting American English
American- or British English are two variations of the English language. They have substantial differences. They are sufficiently the same and are unlikely be mistaken to be separate languages.
In OmegaWiki, it has been possible to add entries for English; this meant there is no difference between written the different versions of English or you had to specify both versions. Issues like this exist for other languages like Serbian and Mandarin as well.
In the OmegaWiki user interface, languages are considered ISO-639 entities. When a DefinedMeaning for a language is part of the appropriate collection, we use the translations in our user interface. The problem is that all these linguistic entities are needed now and that they are created to make OmegaWiki work.
For the ISO-639-6 there will be issues as the codes we make, using the RFC 4646 methodology, will be replaced. It will also be interesting to learn how in the end everything will be merged together.
In the mean time we now support localisation for these linguistic entities.
Thanks,
GerardM
In OmegaWiki, it has been possible to add entries for English; this meant there is no difference between written the different versions of English or you had to specify both versions. Issues like this exist for other languages like Serbian and Mandarin as well.
In the OmegaWiki user interface, languages are considered ISO-639 entities. When a DefinedMeaning for a language is part of the appropriate collection, we use the translations in our user interface. The problem is that all these linguistic entities are needed now and that they are created to make OmegaWiki work.
For the ISO-639-6 there will be issues as the codes we make, using the RFC 4646 methodology, will be replaced. It will also be interesting to learn how in the end everything will be merged together.
In the mean time we now support localisation for these linguistic entities.
Thanks,
GerardM
Sunday, December 02, 2007
Sinterklaas present
In the Netherlands we traditionally do not get presents with Christmas. For us Sinterklaas is celebrated on the fifth or sixth of December. In my family everyone no longer believes in Sinterklaas and consequently we can celebrate it on a more convenient moment like in a weekend.
I have had a wonderful Sinterklaas, and I do want to tell you about the present that Kipcool and Kim gave me. Kipcool wrote this wonderful functionality that show you what classes we have in OmegaWiki and, how many translations we have in your language.
With the new functionality you will see the concepts that are translated in your language. When you check out a concept, you will even find what attributes are available in your language. I have found it to be really addictive. :)
Thanks,
GerardM
I have had a wonderful Sinterklaas, and I do want to tell you about the present that Kipcool and Kim gave me. Kipcool wrote this wonderful functionality that show you what classes we have in OmegaWiki and, how many translations we have in your language.
With the new functionality you will see the concepts that are translated in your language. When you check out a concept, you will even find what attributes are available in your language. I have found it to be really addictive. :)
Thanks,
GerardM
Monday, November 19, 2007
New upload functionality
OmegaWiki is really happy to announce that we have, with thanks to the Otto-Friedrich University of Bamberg, for the first time used new upload functionality. The University of Bamberg has a need for a repository for its Destinazione Italia terminology and has found this in OmegaWiki.
We have uploaded translation in Persian and we have it nicely attributed to the person who did the work. We will upload for several more languages. What is of relevance is that we can make an export for any language, for any collection. So if you are interested to help on our OLPC collection.. just drop me a line..
Thanks,
GerardM
We have uploaded translation in Persian and we have it nicely attributed to the person who did the work. We will upload for several more languages. What is of relevance is that we can make an export for any language, for any collection. So if you are interested to help on our OLPC collection.. just drop me a line..
Thanks,
GerardM
Friday, October 19, 2007
Money
All projects need money to operate. OmegaWiki does need money to operate. We now have made it possible for you to support what we do in a practical way; we now have a link in our sidebar so that you can use Pay Pall to donate money. You can just give us money, you can help us fund a project that we want to do.. Check out our Donations, putting your money where your mouth is ..
Stiching Open Progress is a Dutch "not for profit" organisation and we can use all the money we can get to do all the cool development work we would like to do..
Thanks,
GerardM
Stiching Open Progress is a Dutch "not for profit" organisation and we can use all the money we can get to do all the cool development work we would like to do..
Thanks,
GerardM
Tuesday, October 16, 2007
Zimbabwe was formerly known as Rhodesia
Many countries have over time changed their nature. What they typically do is stay more or less in the same shape. As a consequence of war, the shapes do change. The change is often reflected in the name. The country that is now called Zimbabwe was once called Rhodesia. It is relevant information and can be expressed using relations.
In OmegaWiki the relation type "was formerly known as" has been introduced to express this relation for countries. It demonstrates that OmegaWiki is not strictly a dictionary, it also serves the functions of a dictionary. By including different types of attributes to classes, we provide more worthwhile information.
Most concepts are related to other concepts and when these relations become visible, a net develops of related information. This does not make OmegaWiki an encyclopaedia, it is what an ontology does. An encyclopaedia we are not; we refer to Wikipedia.. :)
Thanks,
GerardM
In OmegaWiki the relation type "was formerly known as" has been introduced to express this relation for countries. It demonstrates that OmegaWiki is not strictly a dictionary, it also serves the functions of a dictionary. By including different types of attributes to classes, we provide more worthwhile information.
Most concepts are related to other concepts and when these relations become visible, a net develops of related information. This does not make OmegaWiki an encyclopaedia, it is what an ontology does. An encyclopaedia we are not; we refer to Wikipedia.. :)
Thanks,
GerardM
Friday, October 12, 2007
Linking to Wikipedia
When people need information, OmegaWiki will not provide all information. What it can do is link to other sources of information and Wikipedia is the obvious and the only choice for encyclopaedic information. It is the only choice because it aims to be multi-lingual and it is an obvious choice because of the shared values.
At this stage, linking to the Wikipedia articles is done by hand so initially there will be few links. We hope to harvest these links from Wikipedia and insert them with a bot. In this way we will provide an encyclopaedic service without being encyclopaedic :)
Thanks,
GerardM
Friday, October 05, 2007
Antonym
An antonym is the complete opposite of something.. black and white are probably the best known examples. The great thing is that antonym is the first global relation type and in the way it is set up, the antonym is true on a concept level. This means that it does not allow for cultural differences in the appreciation of such a relation.
I wonder how many antonyms will prove to be problematic because of cultural differences. The good news is that we are now able to have global relation types in OmegaWiki. We will have to be REALLY careful what relation types we will include. The "is a" relation is not going to be part of it because that is what makes something a class member.
With the global and the class based relation types we only need the collection based relation types to get our full functionality :)
Thanks,
GerardM
I wonder how many antonyms will prove to be problematic because of cultural differences. The good news is that we are now able to have global relation types in OmegaWiki. We will have to be REALLY careful what relation types we will include. The "is a" relation is not going to be part of it because that is what makes something a class member.
With the global and the class based relation types we only need the collection based relation types to get our full functionality :)
Thanks,
GerardM
Tuesday, October 02, 2007
More localisation
OmegaWiki had a lot of new functionality go live, today I spend time on one aspect that is really dear to me; localisation. Much of the OmegaWiki content is localised by adding translations. With the latest software release much of the more programmatic parts are in the system messages.
I started to translate the Dutch messages and, Tosca caught on and started on the German messages, Malafaya did the Portuguese. We hope that in this way our data becomes even more accessible to our users :)
One issue remains, with our messages translated in OmegaWiki, how do we get them in MediaWiki proper ...
Thanks,
GerardM
I started to translate the Dutch messages and, Tosca caught on and started on the German messages, Malafaya did the Portuguese. We hope that in this way our data becomes even more accessible to our users :)
One issue remains, with our messages translated in OmegaWiki, how do we get them in MediaWiki proper ...
Thanks,
GerardM
Getting to grips with the new functionality
At OmegaWiki, a lot of new functionality has gone on line. This functionality is a mix of functionality that was needed for Wikiproteins and things we have been working towards for a long time. With the changes some functionality does not work as it used to. This is a good thing.
Our first content was the GEMET thesaurus, in this collection particular relation types were used. These relation types were available everywhere and consequently we have been reluctant to add more relation types. Now relation types are associated with "classes" and we can make a DefinedMeaning a member of a class. Nederland now has a capital, a motto, a nation anthem and entities bordering the country. For Nederlands it is now known what script it is written in, and in what countries it is spoken. And with the "incoming relations" we know where there is a reference to the DefinedMeaning.
Many of the existing relations will be changed from the GEMET relation types to the new relation types. The work that what is done in the past is a huge benefit as it helps a lot in identifying what needs doing. With the new functionality it makes sense to add the annotations straight away; we now know that they can be done properly.
Thanks,
GerardM
Our first content was the GEMET thesaurus, in this collection particular relation types were used. These relation types were available everywhere and consequently we have been reluctant to add more relation types. Now relation types are associated with "classes" and we can make a DefinedMeaning a member of a class. Nederland now has a capital, a motto, a nation anthem and entities bordering the country. For Nederlands it is now known what script it is written in, and in what countries it is spoken. And with the "incoming relations" we know where there is a reference to the DefinedMeaning.
Many of the existing relations will be changed from the GEMET relation types to the new relation types. The work that what is done in the past is a huge benefit as it helps a lot in identifying what needs doing. With the new functionality it makes sense to add the annotations straight away; we now know that they can be done properly.
Thanks,
GerardM
Friday, September 28, 2007
Major update for OmegaWiki
OmegaWiki has had some major update; the version of MySQL that is installed has been updated, several files have been changed to InnoDB and a lot of functionality has changed behind the scenes.
One of the effects is that the performance has improved noticeable; that was really needed. The difference in performance is a relief. It is fun again to work on the data.
One difference is that the way the relations work; relations are currently associated with a "class" and this class defines what relation types are possible. We have added a few classes so far; "linguistic entity" is one. The associated relation types allow us to indicate where the linguistic entity fits in and, where it is spoken. There will be many more classes and relation types, the quality of the classes and relation types will make a difference to the quality and the usefulness of our data.
Thanks,
GerardM
One of the effects is that the performance has improved noticeable; that was really needed. The difference in performance is a relief. It is fun again to work on the data.
One difference is that the way the relations work; relations are currently associated with a "class" and this class defines what relation types are possible. We have added a few classes so far; "linguistic entity" is one. The associated relation types allow us to indicate where the linguistic entity fits in and, where it is spoken. There will be many more classes and relation types, the quality of the classes and relation types will make a difference to the quality and the usefulness of our data.
Thanks,
GerardM
Friday, September 07, 2007
Demo Semantic Support on a new URL
At Wikimania I presented what we are doing to bring real time Semantic Support to Wikipedia. The URL in my presentation is no longer valid, the new location is at: wikipedia.wikitestsite.org.
You will find a dump of the English Wikipedia and you many of the expressions that we already now are in green. We are working towards a situation where new concepts defined in OmegaWiki will be recognised in the future data mining of the same article.
What we are discussing at the moment is adding functionality to the concepts found. Some are obvious like giving the definition, giving an option to go to OmegaWiki when the definition does not fit, showing translations for the expression in the language that is of interest to the reader.
We can imagine that there is more functionality that you would consider useful. Please let us know .. :)
Thanks,
GerardM
You will find a dump of the English Wikipedia and you many of the expressions that we already now are in green. We are working towards a situation where new concepts defined in OmegaWiki will be recognised in the future data mining of the same article.
What we are discussing at the moment is adding functionality to the concepts found. Some are obvious like giving the definition, giving an option to go to OmegaWiki when the definition does not fit, showing translations for the expression in the language that is of interest to the reader.
We can imagine that there is more functionality that you would consider useful. Please let us know .. :)
Thanks,
GerardM
Monday, September 03, 2007
Connecting data from different databases
In OmegaWiki there are different datasets. These represent different origins and have a different emphasis. What we are working on is to connecting the data in these different datasets. Currently over four percent of our Community data is connected to data of the UMLS.
These connections are not without problems. The UMLS does not have the same (lexical) outlook; it is quite happy to have a singular and a plural to be part of the same concept. In OmegaWiki we do not support the notion of plurals yet. For the UMLS it is not a problem to include Geologists as it is included as a subject heading. We have it connected to geologist.
Lyme disease has several synonyms that are problematic from a lexical point of view; only "Lyme borreliosis" is what I expect to find in a dictionary. This does not necessarily mean that "Borreliosis, Lyme" is not useful to have. The Community database knows some 15 translations and thereby adds value to the English only content for Lyme disease.
With four percent of the Community Database connected, in reality we haven't scratched the surface of the UMLS. The UMLS is a well explored resource and I am sure that there are many resources that have made connections already. I hope we will find the people, the organisations willing to share the work that they have already done.
Thanks,
GerardM
These connections are not without problems. The UMLS does not have the same (lexical) outlook; it is quite happy to have a singular and a plural to be part of the same concept. In OmegaWiki we do not support the notion of plurals yet. For the UMLS it is not a problem to include Geologists as it is included as a subject heading. We have it connected to geologist.
Lyme disease has several synonyms that are problematic from a lexical point of view; only "Lyme borreliosis" is what I expect to find in a dictionary. This does not necessarily mean that "Borreliosis, Lyme" is not useful to have. The Community database knows some 15 translations and thereby adds value to the English only content for Lyme disease.
With four percent of the Community Database connected, in reality we haven't scratched the surface of the UMLS. The UMLS is a well explored resource and I am sure that there are many resources that have made connections already. I hope we will find the people, the organisations willing to share the work that they have already done.
Thanks,
GerardM
Monday, August 27, 2007
Some more on Wolof
OmegaWiki wants to support all words of all languages and, it does not want to go into the issue of does this language exist or not. We make use of the ISO 639 standards and, when we feel like being adventurous, we look at what is recognised in the IANA language tags.
Deferring to standard organisations means that you take what they say as the "truth". It does not mean that we necessarily agree, but it saves us from a lot of mayhem. Yesterday I wrote about the first native Wolof speaker for OmegaWiki. Today Ibou changed the definition for Wolof and included Gambia as a country where Wolof is spoken. According to the description by Ethnologue of the Wolof language this is not the case. They do refer to another language, Gambian Wolof, this description makes it clear that Wolof is spoken in the Gambia as well.
The article on Wikipedia on Wolof is in my opinion wrong; it gives the impression that the ISO-639-1 and the ISO-639-2 codes are split into two. This is contrary to how standards work. When a language is split into two, the original meaning will stand as it is, it will get a new description to indicate that it has been split and two new codes will be created.
So Ethnologue is inconsistent. Ibou is probably right. I have send an e-mail to Ethnologue and I hope that they will amend their fine resource so that we will know for sure that he is right. :)
Thanks,
GerardM
Deferring to standard organisations means that you take what they say as the "truth". It does not mean that we necessarily agree, but it saves us from a lot of mayhem. Yesterday I wrote about the first native Wolof speaker for OmegaWiki. Today Ibou changed the definition for Wolof and included Gambia as a country where Wolof is spoken. According to the description by Ethnologue of the Wolof language this is not the case. They do refer to another language, Gambian Wolof, this description makes it clear that Wolof is spoken in the Gambia as well.
The article on Wikipedia on Wolof is in my opinion wrong; it gives the impression that the ISO-639-1 and the ISO-639-2 codes are split into two. This is contrary to how standards work. When a language is split into two, the original meaning will stand as it is, it will get a new description to indicate that it has been split and two new codes will be created.
So Ethnologue is inconsistent. Ibou is probably right. I have send an e-mail to Ethnologue and I hope that they will amend their fine resource so that we will know for sure that he is right. :)
Thanks,
GerardM
Sunday, August 26, 2007
One new user
Sometimes a new user is special. To me Ibou is special. He is the first Wolof native speaker on OmegaWiki. He is the first person where I have been told for whom communicating in English will be difficult.
I could not be more happy with what he has done so far; he created the Babel templates for Wolof. He has translated the first part of the main menu. Really, he makes the next Wolof speakers feel welcome..
Thanks,
GerardM
I could not be more happy with what he has done so far; he created the Babel templates for Wolof. He has translated the first part of the main menu. Really, he makes the next Wolof speakers feel welcome..
Thanks,
GerardM
Friday, August 24, 2007
Localisation of OmegaWiki
What makes OmegaWiki so special, is that the presentation of the data is shown in the language selected in the user preferences. The data is sorted properly. It is really nice.
In the next version of the software, it will be possible to localise the headers as well. These are currently in English only. In contrast to how the localisation is done, the headers will be system messages. In this way each Wikis for Professionals can choose the headers that provide the best fit for its application.
Thanks,
GerardM
In the next version of the software, it will be possible to localise the headers as well. These are currently in English only. In contrast to how the localisation is done, the headers will be system messages. In this way each Wikis for Professionals can choose the headers that provide the best fit for its application.
Thanks,
GerardM
Sunday, August 12, 2007
30.000 Expressions for English in the Community Database
The word "Southern Sierra Miwok" is a language particular to California. In 1994 there were still 7 people that spoke this language. According to the UMLS, the name of the language is Meewoc and as Ethnologue provides it as one of the alternate names, it was possible to link this DefinedMeaning in the Community Database with the concept in the UMLS.
With 30.000 Expressions, there are many that also exist in the UMLS Authoritative Database. The UMLS has some 1.93 million Expression at the moment, and the first 200+ DefinedMeanings have already been linked. They are mainly US-American languages and chemical elements with an occasional animal like guinea pig thrown in for good measure.
It is relevant to have the concept linked. It means that the information in one database can be seen as supplementary to what is available in another database. Currently we have three databases, but when you consider how they are structured, there are implicit connections known as many of the concepts known in the UMLS are also known in the Swiss-Prot database. The only thing left doing is making them explicit. :)
Thanks,
GerardM
With 30.000 Expressions, there are many that also exist in the UMLS Authoritative Database. The UMLS has some 1.93 million Expression at the moment, and the first 200+ DefinedMeanings have already been linked. They are mainly US-American languages and chemical elements with an occasional animal like guinea pig thrown in for good measure.
It is relevant to have the concept linked. It means that the information in one database can be seen as supplementary to what is available in another database. Currently we have three databases, but when you consider how they are structured, there are implicit connections known as many of the concepts known in the UMLS are also known in the Swiss-Prot database. The only thing left doing is making them explicit. :)
Thanks,
GerardM
Monday, July 30, 2007
UMLS
The UMLS or Unified Medical Language System is a collection of many resources it contains tools, a semantic network and a specialist lexicon. It is also a collection of many resources. These resources all have their own license and copyright. Effectively much of the UMLS can be used for many purposes because the particular license allows it. In the same way, there is much of the UMLS that can only be used when the copyright holder gives permission.
In OmegaWiki, we have our first Authoritative Database online. It is the UMLS and we are proud of it. Now as the UMLS is this collection of connected resources, we present it in the same way. There is one UMLS as an Authoritative Database and it has collections that are the parts that make up the UMLS as we have it. The important thing of the UMLS is that it did make the connection between the different databases and we do use their system to connect.
What makes the inclusion of the UMLS so special is that we have the cooperation of the NLM. It is what makes this such an exciting experiment.
Thanks,
GerardM
In OmegaWiki, we have our first Authoritative Database online. It is the UMLS and we are proud of it. Now as the UMLS is this collection of connected resources, we present it in the same way. There is one UMLS as an Authoritative Database and it has collections that are the parts that make up the UMLS as we have it. The important thing of the UMLS is that it did make the connection between the different databases and we do use their system to connect.
What makes the inclusion of the UMLS so special is that we have the cooperation of the NLM. It is what makes this such an exciting experiment.
Thanks,
GerardM
Sunday, July 29, 2007
IL7R-alpha and IL2R-alpha anyone ?
The OmegaWiki database is blocked for editing at the moment. The reason given is: "Importing new data". For me this is great news. It means that we are importing the data that we have been preparing for a long time. It means that we are closer to getting the first Wiki for Professionals life.
Today, on the BBC-news website there is an article where IL7R-alpha and IL2R-alpha play a major role. They are proteins, more specific they are genetic variants of proteins that play a role in the expression of multiple sclerosis.
OmegaWiki will contain terminology like IK7R-alpha, it is specialised terminology but as it can be found in sources like the BBC website, it is good to have it. Wikiproteins will be the first Wiki for Professionals and, it will allow for the further annotation of these proteins by people who know about these substances.
It is really thrilling to see all the development needed to come to a first public outing come to a close. I congratulate the members of the consortium that make Wikiproteins possible. I believe that this has the potential to become an important tool for scientists. Wikiproteins is possible because of the many people who believe that Open Access is essential to science.
By going life, we invite comments. These will help us to make sure that the functionality is just right. The best people that can help us identify what more needs to be done are the people who will become part of what will be the Wikiproteins community.
The "official" announcement of Wikiproteins going life is scheduled at Wikimania 2007 :)
Thanks,
GerardM
Today, on the BBC-news website there is an article where IL7R-alpha and IL2R-alpha play a major role. They are proteins, more specific they are genetic variants of proteins that play a role in the expression of multiple sclerosis.
OmegaWiki will contain terminology like IK7R-alpha, it is specialised terminology but as it can be found in sources like the BBC website, it is good to have it. Wikiproteins will be the first Wiki for Professionals and, it will allow for the further annotation of these proteins by people who know about these substances.
It is really thrilling to see all the development needed to come to a first public outing come to a close. I congratulate the members of the consortium that make Wikiproteins possible. I believe that this has the potential to become an important tool for scientists. Wikiproteins is possible because of the many people who believe that Open Access is essential to science.
By going life, we invite comments. These will help us to make sure that the functionality is just right. The best people that can help us identify what more needs to be done are the people who will become part of what will be the Wikiproteins community.
The "official" announcement of Wikiproteins going life is scheduled at Wikimania 2007 :)
Thanks,
GerardM
Saturday, July 28, 2007
DMM - Swahili content for OmegaWiki
Yesterday, I reintroduced the notion of "Donations, putting your money where your mouth is" on this blog. Today I want to tell you about one of the first such projects.
The Kamusi project is a really important project to create a dictionary for Swahili. The project was a project of Yale University and Martin Benjamin was its editor. The project is probably one of the most important resources for the Swahili language and it is therefore really sad that the activity of this project came to an end because of a lack of funding.
Martin is preparing a new project for African languages called PALDO or the Pan-African Living Online Dictionary. This project aims to create content in many of the important African languages. Martin has been given permission to use the content of the Kamusi project from Yale. This means that it becomes possible for him to collaborate with other projects as well.
It is with pride and gratitude that I can say that PALDO and OmegaWiki are going to work together. This means that we need to get the content of Kamusi analysed and imported. It also means that we have to analyse and build the functionality so that we can give back to the PALDO project. More information can be found here.
With a 70.000 word Swahili dictionary, we have sufficient data for the first two OLPC dictionaries that will amount to something. They will be Swahili and English.. the English content comes with Kamusi as well :)
So you can help; you can develop, you can edit and you can sponsor this project.
Thanks,
GerardM
The Kamusi project is a really important project to create a dictionary for Swahili. The project was a project of Yale University and Martin Benjamin was its editor. The project is probably one of the most important resources for the Swahili language and it is therefore really sad that the activity of this project came to an end because of a lack of funding.
Martin is preparing a new project for African languages called PALDO or the Pan-African Living Online Dictionary. This project aims to create content in many of the important African languages. Martin has been given permission to use the content of the Kamusi project from Yale. This means that it becomes possible for him to collaborate with other projects as well.
It is with pride and gratitude that I can say that PALDO and OmegaWiki are going to work together. This means that we need to get the content of Kamusi analysed and imported. It also means that we have to analyse and build the functionality so that we can give back to the PALDO project. More information can be found here.
With a 70.000 word Swahili dictionary, we have sufficient data for the first two OLPC dictionaries that will amount to something. They will be Swahili and English.. the English content comes with Kamusi as well :)
So you can help; you can develop, you can edit and you can sponsor this project.
Thanks,
GerardM
Friday, July 27, 2007
Donations, putting your money where your mouth is
When you want to get things done, you can do it yourself or you can get someone else to do it for you. Within many Open Source or Open Content projects you can donate your programming, your content and your money. When you are a programmer or an editor, you can choose what to develop, what to edit, because you are a volunteer. Nobody can tell you what to do. When you volunteer to give money, there is no such luck. You can give, you may be thanked, and that is it.
Unless of course you are a big time donor. When you give a sufficient amount of money and the purpose for this money fits within the aims of the organisation you give it to, you can determine what the money is spend on. This is not an option for small time donors.
Many small projects have been identified that need doing, projects that do not get done because they do not have priority or because nobody volunteers to do them. For such projects a specification can be made and a cost estimate can be given. These can be published and donations for these projects can be solicited. When enough people have contributed funding for a project, it can be executed.
For the complete policy read; Donations, putting your money where your mouth is.
Thanks,
GerardM
Unless of course you are a big time donor. When you give a sufficient amount of money and the purpose for this money fits within the aims of the organisation you give it to, you can determine what the money is spend on. This is not an option for small time donors.
Many small projects have been identified that need doing, projects that do not get done because they do not have priority or because nobody volunteers to do them. For such projects a specification can be made and a cost estimate can be given. These can be published and donations for these projects can be solicited. When enough people have contributed funding for a project, it can be executed.
For the complete policy read; Donations, putting your money where your mouth is.
Thanks,
GerardM
Thursday, July 26, 2007
What language is this text ?
When you write a text, you know your audience and you select a language accordingly. Given that English is the lingua franca of this day and age, and given that my public is international I do write in English. However, there is nothing that stops me or any of the other people who contribute to this blog from writing in a different language.
This is a bad thing. It would be so much better when I was able to actively indicate the language that I am writing. Obviously, Blogger can have its own routines to distinguish certain languages, but I am absolutely certain that they will not recognise the majority of languages.
While I am typing this blog, I have indicated to my spell checker that I am using UK English spelling. This means that many of the mistakes I make will not be seen by you. Having indicated that the languages IS UK English, it would have been great when it was picked up by the Blogger software.
Consider, when I inform Blogger that I am writing UK English, my Firefox spelling extension does not need to guess anymore. It would provide me with a much better functionality and it would make functionality possible in languages that are not well supported..
So please blogger.com, please allow me to tag the language of my texts.
Thanks,
GerardM
This is a bad thing. It would be so much better when I was able to actively indicate the language that I am writing. Obviously, Blogger can have its own routines to distinguish certain languages, but I am absolutely certain that they will not recognise the majority of languages.
While I am typing this blog, I have indicated to my spell checker that I am using UK English spelling. This means that many of the mistakes I make will not be seen by you. Having indicated that the languages IS UK English, it would have been great when it was picked up by the Blogger software.
Consider, when I inform Blogger that I am writing UK English, my Firefox spelling extension does not need to guess anymore. It would provide me with a much better functionality and it would make functionality possible in languages that are not well supported..
So please blogger.com, please allow me to tag the language of my texts.
Thanks,
GerardM
Wednesday, July 25, 2007
A new menu
In preparation for the presentations at Wikimania, we are making OmegaWiki extra nice. I am finishing the new main page where you will find information about the things we have been preparing that are not there yet.
There have been things in our wiki that do not have any functionality yet. A lot of work has been done in the last weeks in refactoring our code. We are making changes to the code in order to make it easier for new developers to get to grips with the code. These changes will make it possible to build some of the more complex functionality that we need.
While we are working hard at OmegaWiki, Knewco is working hard preparing their Desktop, with its Knowlet and Semantic Support. It is really cool that OmegaWiki will not only be useful in its own right, but that applications are going to be build on top of it.
I am really excited about going to Taipei. I will be happy to talk and demonstrate what we are on about.. It is less than a week I am thrilled to see all the everything coming together :)
Thanks,
GerardM
There have been things in our wiki that do not have any functionality yet. A lot of work has been done in the last weeks in refactoring our code. We are making changes to the code in order to make it easier for new developers to get to grips with the code. These changes will make it possible to build some of the more complex functionality that we need.
While we are working hard at OmegaWiki, Knewco is working hard preparing their Desktop, with its Knowlet and Semantic Support. It is really cool that OmegaWiki will not only be useful in its own right, but that applications are going to be build on top of it.
I am really excited about going to Taipei. I will be happy to talk and demonstrate what we are on about.. It is less than a week I am thrilled to see all the everything coming together :)
Thanks,
GerardM
Tuesday, July 10, 2007
Ch'orti', a language spoken in Guatemala and Honduras
Ch'orti' as a language has caa as its ISO 639-3 code, some 30.000 people speak the language and according to Reeck many more belong to the associated ethnic population.
I have been adding languages to the ISO 639-3 collection for some time now, I started with Ghotuo (aaa) and I have now progressed to Ch'orti' (caa). Many have few speakers, many are extinct, several are sign languages and almost all of them I have already forgotten.
So why do this, is there method to this madness.. OmegaWiki aims to include all words of all languages, but what languages are there ? Do we want to discuss the notion of yet another linguistic entity that we should support. Does something like Brithenig (bzt) deserve its place under the sun ?
I do not mind the discussion, but I do mind what the result will be of such a discussion. It needs to come to a conclusion and I do not want to be in the position that people look to me for a verdict. It is not a good idea either to have the OmegaWiki commission be in that position. It is for all these reasons that we decided on adopting standards and started with the creation of portals for the ISO 639-3 languages. We are now at the next phase, creating the DefinedMeanings for these languages and make them part of the ISO 639-3 collection.
This is only what is recognised by one standard, there are other standards that help indicate what the precise linguistic entity is that is to be documented in OmegaWiki. First we should finish this, there are currently 1365 entries in the ISO 639-3 collection .. there are many more thousands to go :)
Thanks,
GerardM
I have been adding languages to the ISO 639-3 collection for some time now, I started with Ghotuo (aaa) and I have now progressed to Ch'orti' (caa). Many have few speakers, many are extinct, several are sign languages and almost all of them I have already forgotten.
So why do this, is there method to this madness.. OmegaWiki aims to include all words of all languages, but what languages are there ? Do we want to discuss the notion of yet another linguistic entity that we should support. Does something like Brithenig (bzt) deserve its place under the sun ?
I do not mind the discussion, but I do mind what the result will be of such a discussion. It needs to come to a conclusion and I do not want to be in the position that people look to me for a verdict. It is not a good idea either to have the OmegaWiki commission be in that position. It is for all these reasons that we decided on adopting standards and started with the creation of portals for the ISO 639-3 languages. We are now at the next phase, creating the DefinedMeanings for these languages and make them part of the ISO 639-3 collection.
This is only what is recognised by one standard, there are other standards that help indicate what the precise linguistic entity is that is to be documented in OmegaWiki. First we should finish this, there are currently 1365 entries in the ISO 639-3 collection .. there are many more thousands to go :)
Thanks,
GerardM
Wednesday, July 04, 2007
Aklanon ...
Aklanon is a language spoken in the Philipines. The ISO-639-3 code is "akl" and according to a 1990 census some 394.545 people speak this language.
On OmegaWiki, Aklanon has its own portal and I was really thilled when Chief Mike indicated his interest in working on the Aklanon content. We do want Aklanon but we also have our own standards. One of these standards is that the Babel templates for a language are in that language. I really appreciate the notion that the Babel templates have to be understood however, the Babel templates are one of the first things that we hope to get in any language.
When we have the Aklanon Babel templates in Aklanon, it will be a privilege to have Aklanon as the next language that we support in OmegaWiki.
Thanks,
GerardM
On OmegaWiki, Aklanon has its own portal and I was really thilled when Chief Mike indicated his interest in working on the Aklanon content. We do want Aklanon but we also have our own standards. One of these standards is that the Babel templates for a language are in that language. I really appreciate the notion that the Babel templates have to be understood however, the Babel templates are one of the first things that we hope to get in any language.
When we have the Aklanon Babel templates in Aklanon, it will be a privilege to have Aklanon as the next language that we support in OmegaWiki.
Thanks,
GerardM
Saturday, June 23, 2007
Assorted statistics
OmegaWiki has reached 30.000 DefinedMeanings, we have some 258.000 Expressions. And as some people do not stop telling me there are about 28.000 Expressions in the largest language and this means that there is a close relation between the number of Expressions in a language and the number of concepts. This is said to indicate that OmegaWiki should be able to scale. :)
The Webaliser statistics have given me a surprise; there is now more info to be found. What is nice to see is that there is now a breakdown in where the traffic comes from. As we have a lot of traffic from crawlers, it would be good to exclude crawlders in order to see where interested PEOPLE come from. Erik told me that the new features are probably due to the upgrade of this week.
Malafaya is now the fourth person who has taken an interest in our statistics. He has worked on the reliability of the statistics of collections. His first effort improved the numbers, his second stab at it improved the performance of the queries a lot.
Finally the Alexa statistics have improved a lot for no apparent reason. We have had times when we were not ranked at all or we could be found above the 800.000 range.. Now we are for a few days hovering around the 368.500 mark. Still not impressive but it looks much better. When you compare the Alexa numbers with our Webaliser numbers, the only thing that can be said is that for Alexa the numbers are statistically not really valid.. This will improve as our community grows.
Thanks,
GerardM
The Webaliser statistics have given me a surprise; there is now more info to be found. What is nice to see is that there is now a breakdown in where the traffic comes from. As we have a lot of traffic from crawlers, it would be good to exclude crawlders in order to see where interested PEOPLE come from. Erik told me that the new features are probably due to the upgrade of this week.
Malafaya is now the fourth person who has taken an interest in our statistics. He has worked on the reliability of the statistics of collections. His first effort improved the numbers, his second stab at it improved the performance of the queries a lot.
Finally the Alexa statistics have improved a lot for no apparent reason. We have had times when we were not ranked at all or we could be found above the 800.000 range.. Now we are for a few days hovering around the 368.500 mark. Still not impressive but it looks much better. When you compare the Alexa numbers with our Webaliser numbers, the only thing that can be said is that for Alexa the numbers are statistically not really valid.. This will improve as our community grows.
Thanks,
GerardM
Monday, June 18, 2007
Server upgraded; dataset support online
The OmegaWiki.org server has been upgraded to Debian etch. This gives us PHP 5.2.0, which is needed to run the latest version of OmegaWiki. (In the process, we exchanged our hand-compiled PHP and Apache binaries with distribution packages.) OmegaWiki itself has also been upgraded. The current version of the code has support for so-called "data-sets".
A data-set is essentially an instance of OmegaWiki which can contain a completely separate set of DefinedMeanings and associated data. This is useful for importing authoritative sources which may either not yet be fully editable, or which are meant to be retained alongside an editable version. It also allows us to showcase imported databases, to convince organizations that own the data to release it freely and make it fully editable.
The current version already supports mapping DefinedMeanings across data-sets. So you can indicate that concept A in data-set 1 is the same as concept B in data-set 2. However, it does not yet support copying data from one data-set to another, which is what we are working on right now (some hints to it are already in the code).
Currently OmegaWiki has a single data-set only. We are considering to set up some example data-sets to let the user community play with this new functionality.
A data-set is essentially an instance of OmegaWiki which can contain a completely separate set of DefinedMeanings and associated data. This is useful for importing authoritative sources which may either not yet be fully editable, or which are meant to be retained alongside an editable version. It also allows us to showcase imported databases, to convince organizations that own the data to release it freely and make it fully editable.
The current version already supports mapping DefinedMeanings across data-sets. So you can indicate that concept A in data-set 1 is the same as concept B in data-set 2. However, it does not yet support copying data from one data-set to another, which is what we are working on right now (some hints to it are already in the code).
Currently OmegaWiki has a single data-set only. We are considering to set up some example data-sets to let the user community play with this new functionality.
A word of the day
Like so many other resources that are lexical in nature, OmegaWiki has a word of the day. Our word of the day is not prepared in advance and we leave it to the community to create one. I am always relieved when there is actually a word of the day when I wake up.
Today's word of the day is interesting for many reasons. The word is wheat. There are several issues to consider.
The issue here is that without being able to reference to both families that are grassy, it is hard to appreciate this definition. This word of the day clearly shows why there is a need for a dictionary of life, a dictionary that explains all these names and shows the relations between the different validly published taxonomical names.
Thanks,
GerardM
Today's word of the day is interesting for many reasons. The word is wheat. There are several issues to consider.
- It is marked as "English (United States)". There is however no "English (United Kingdom)" and as I cannot find this alternate, it should be just "English".
- The definition has not been translated into English. This is very much optional, but it makes it so much easier to translate the definition in yet another language
- In the definition, wheat is said to be part of the family ''Graminacee" of the genus "Triticum". According to Wikipedia the family should be "Poaceae".
The issue here is that without being able to reference to both families that are grassy, it is hard to appreciate this definition. This word of the day clearly shows why there is a need for a dictionary of life, a dictionary that explains all these names and shows the relations between the different validly published taxonomical names.
Thanks,
GerardM
Sunday, June 17, 2007
OmegaWiki only a translation dictionary ?
There is some misinformation about OmegaWiki, it is said for instance that OmegaWiki is only a translation dictionary. There are also people who do not consider OmegaWiki as relevant because it is not a Wikimedia Foundation project.
It is for the people that have not looked at OmegaWiki for a long time or have not really looked well that we want to state the obvious; OmegaWiki is not only but also a translation dictionary. When you look at the number of expressions per language, you will find that we have almost 30.000 DefinedMeanings, the reason why we have 11.000 more English Expressions then what we have for any other language is because we have collections that are at still mostly English. Collections like the ISO-DIS-639-6 are relevant because of the information that is included in the data.
OmegaWiki is becoming relevant because our data is starting to be used outside our project as well. Positano News uses OmegaWiki data for "assisted reading", this helps people to understand terminology that is in an Italian news article. It does give you definitions and translations.
It may be that the current possibilities at OmegaWiki are not immediately obvious; there are many DefinedMeanings that do not have any annotation. An annotation can identify the part of speech for a word, it can provide you with a sample sentence or how to hyphenate a word. We want to include links to other websites; we want to link to Wikipedia articles in order to make it convenient to our users to find good encyclopaedic information.
OmegaWiki is not feature complete. We want to add many more features, but our first priority is to make sure that it works well and that the features that matter most are included. We need to improve on our performance and, we need to make sure that we provide a framework that facilitates collaboration with other organisations.
The Wikimedia Foundation is one organisation that we really want to collaborate with. On a personal level we have been involved and we want to extend this by collaborating on an organisational level as well. This often repeated intention may be one reason why certain people are so apprehensive about OmegaWiki; we wanted it to be a WMF project, it is not a WMF project but we still see room for doing good together.
Thanks,
GerardM
It is for the people that have not looked at OmegaWiki for a long time or have not really looked well that we want to state the obvious; OmegaWiki is not only but also a translation dictionary. When you look at the number of expressions per language, you will find that we have almost 30.000 DefinedMeanings, the reason why we have 11.000 more English Expressions then what we have for any other language is because we have collections that are at still mostly English. Collections like the ISO-DIS-639-6 are relevant because of the information that is included in the data.
OmegaWiki is becoming relevant because our data is starting to be used outside our project as well. Positano News uses OmegaWiki data for "assisted reading", this helps people to understand terminology that is in an Italian news article. It does give you definitions and translations.
It may be that the current possibilities at OmegaWiki are not immediately obvious; there are many DefinedMeanings that do not have any annotation. An annotation can identify the part of speech for a word, it can provide you with a sample sentence or how to hyphenate a word. We want to include links to other websites; we want to link to Wikipedia articles in order to make it convenient to our users to find good encyclopaedic information.
OmegaWiki is not feature complete. We want to add many more features, but our first priority is to make sure that it works well and that the features that matter most are included. We need to improve on our performance and, we need to make sure that we provide a framework that facilitates collaboration with other organisations.
The Wikimedia Foundation is one organisation that we really want to collaborate with. On a personal level we have been involved and we want to extend this by collaborating on an organisational level as well. This often repeated intention may be one reason why certain people are so apprehensive about OmegaWiki; we wanted it to be a WMF project, it is not a WMF project but we still see room for doing good together.
Thanks,
GerardM
Saturday, June 16, 2007
If you love somebody set them free
On OmegaWiki we have many sysops. Giving people the abilities that comes with the sysop flag is what has prevented a lot of vandalism and spam. We are happy and grateful that this has worked out so well for us. As a consequence, we do not have the eternal admins versus the editors controversy, our admins do not have to do anything; they are kindly requested to do good and amazingly they do.
With some sadness, we learned that a Wiktionary admin is leaving Wiktionary; he was told to be more active or else. There is a silver lining in that this guy announced to become more active on OmegaWiki. Obviously every project makes his bed and lies in it. We have chosen to have as little bureaucracy as possible. The question is very much; how is it going to scale.
OmegaWiki will expand by including "Wikis for Professionals". Each will include the terminology for a specific domain extended with specific information and functionality. With more people signing up to such a community, it may acquire its own rules. These rules should fit in the larger community that is OmegaWiki. What I expect is that often the unwritten rules will be the more important ones. In a Wiki for Professionals, people will be interested when the project is relevant. When this proves to be demonstrably so, it may become important to be identifiable to gain the benefits of the association with the project. The flip side of the coin is that negative behaviour can damage a professional reputation.
In a year, the community of OmegaWiki will be different. We work hard to provide it with an environment that will enable it to do good. At this stage it is still very much basic functionality that we are building. There is much new functionality and data waiting to go live. When it has, we will love to hear what is good and what could be better. We will love it when people help us morph our functionality and make our environment more relevant.
The only thing that we will insist on is that things can coexist and people collaborate, in that way we set not only the data free but also the imagination free, we will love it and we will set them free.
Thanks,
GerardM
With some sadness, we learned that a Wiktionary admin is leaving Wiktionary; he was told to be more active or else. There is a silver lining in that this guy announced to become more active on OmegaWiki. Obviously every project makes his bed and lies in it. We have chosen to have as little bureaucracy as possible. The question is very much; how is it going to scale.
OmegaWiki will expand by including "Wikis for Professionals". Each will include the terminology for a specific domain extended with specific information and functionality. With more people signing up to such a community, it may acquire its own rules. These rules should fit in the larger community that is OmegaWiki. What I expect is that often the unwritten rules will be the more important ones. In a Wiki for Professionals, people will be interested when the project is relevant. When this proves to be demonstrably so, it may become important to be identifiable to gain the benefits of the association with the project. The flip side of the coin is that negative behaviour can damage a professional reputation.
In a year, the community of OmegaWiki will be different. We work hard to provide it with an environment that will enable it to do good. At this stage it is still very much basic functionality that we are building. There is much new functionality and data waiting to go live. When it has, we will love to hear what is good and what could be better. We will love it when people help us morph our functionality and make our environment more relevant.
The only thing that we will insist on is that things can coexist and people collaborate, in that way we set not only the data free but also the imagination free, we will love it and we will set them free.
Thanks,
GerardM
Sunday, June 03, 2007
250.000 expressions
Today we reached the milestone of 250.000 Expressions at OmegaWiki. It is special because most of this data has been entered by hand. We find that when people get enthused by the concept of OmegaWiki, they do make a difference for the language that they champion.
We have people who have a particular interest in Georgian, Khmer and Spanish, it shows in the statistics as these languages grow much faster than the others.
Aveyron is the 250.000th entry in OmegaWiki and, it is only fitting that Ascánder was the person adding it. Ascander is one of the most valuable contributors to OmegaWiki. Aveyron is part of a project to include information from the ISO-3166-2. In this standard it is detailed in what way countries are subdivided. It does not state that Italy has provinces, the USA has states or that Germany has Bundeslander. It does give the names of these entities.
So, OmegaWiki is evolving nicely. We hope that in line with how Wikis evolve, we will have an easier time to get 250.000 more Expressions.
Thanks,
GerardM
We have people who have a particular interest in Georgian, Khmer and Spanish, it shows in the statistics as these languages grow much faster than the others.
Aveyron is the 250.000th entry in OmegaWiki and, it is only fitting that Ascánder was the person adding it. Ascander is one of the most valuable contributors to OmegaWiki. Aveyron is part of a project to include information from the ISO-3166-2. In this standard it is detailed in what way countries are subdivided. It does not state that Italy has provinces, the USA has states or that Germany has Bundeslander. It does give the names of these entities.
So, OmegaWiki is evolving nicely. We hope that in line with how Wikis evolve, we will have an easier time to get 250.000 more Expressions.
Thanks,
GerardM
Subscribe to:
Posts (Atom)





