Sunday, May 02, 2010

Statistics ... a work in progress

At #OmegaWiki we have nice statistics, nice bar graphs that show the number of expressions for each language. The numbers are nice except.. except for the language of the labels.. they are in English. This is not really how we want to present them because we pride ourselves on our multilingual support.


Now that we are looking again at our numbers, we are considering what numbers to show. The DefinedMeanings, the Expressions or the Syntrans records. The first shows the number of concepts, the second the number of expressions that are used for a language and the last will include the homonyms as well. What do you think?

Obviously these numbers show the current status and you can help us improve these numbers.
Thanks,
       GerardM

Thursday, April 29, 2010

હાથકડી is the word of the day

Everyday at #OmegaWiki another word of the day. There are several people who divine what it should be. Sometimes the word is in the news, sometimes like today it is a word that was added the previous day.

OmegaWiki currently supports 247 languages and when you look at the distribution of the words over the languages you find that it tails off quite sharply. This does not mean that expressions in languages like Gujarati are not welcome. For from it and, to show our wish for collaborators for for instance Gujarati, હાથકડી is today's word of the day.
Thanks,
     GerardM

Monday, April 19, 2010

#Wikipedia article

At #OmegaWiki we have twenty translations of the phrase "Wikipedia article". This phrase is one of the Community class attributes, these are used to indicate a relation. 

The relation indicated by the "Wikipedia article" is an article in the Wikipedia of that language. Currently many articles are referred to in the English and the Spanish Wikipedia.

We would appreciate it when you help us with translations for this phrase in your language.
Thanks,
       GerardM

Sunday, April 18, 2010

Word not found

When an expression is not found at #OmegaWiki, a link will be proposed for you to create either a new article or a new expression..


Kipcool informed me about this on IRC, he also expressed his amazement at the speed in which localisations become available at translatewiki.net. There are already localisations for 10 languages for the new message.
Thanks,
      GerardM

What to do when a word is not part of your vocabulary

Every now and then, I find a word that I do not know. For a normal person, it is a trip to the dictionary, for me it is a trip to OmegaWiki and often I find that the word is not yet in there.

Today it was the word cohort that I did not know. I knew it as something to do with Roman military, but it has also something to do with statistics..

I like pictures in my blogs, people playing soldiers is more fun to look at then some bar charts.
Thanks,
      GerardM

Thursday, April 15, 2010

The ratio between DefinedMeanings and Expressions

#OmegaWiki as a resource provides information in many languages. Currently there are some 245 languages enabled for the inclusion of expressions. I often add words that feature as an Apertium concept and in order for it to be useful for the purpose of Apertium, it is important to add as many translations as possible.


As a consequence we currently have some 9.27 Expressions for every DefinedMeaning. Slowly but surely this ratio is moving upward and this provides some secret motivation :) This does not mean that we have something like 9.27 translations for every concept; an Expression is just a string of characters and when  this string of characters is used in multiple languages or is used for multiple concepts, it is still only one Expression.

I would love to know what the ratio is between the number of concepts and the number of translations; it will be a higher number ..
Thanks,
    GerardM

Tuesday, April 13, 2010

15,000 #Portuguese expressions in #OmegaWiki

We are happy to have a sufficient number of expressions in the Portuguese language at OmegaWiki. It shows how OmegaWiki provides the functionality that you will not find elsewhere.


The labels are shown in Portuguese when you have selected Portuguese as your default language. Better still, when you edit in OmegaWiki, the same labels can be selected in Portuguese.


We would dearly love to provide the same functionality to languages like Telugu, Kannada, Abchaz, Guarani.. We can help by pointing out how you can quickly provide a localised environment, we can help by adding the expressions for any languages when we find them. In the end we need people interested in supporting their language.
Thanks,
       GerardM

Friday, April 09, 2010

Working on the #Apertium M list


When you want to add content to #OmegaWiki, it helps when you have a goal. As you can see from the abundance of "red links", there is plenty left to do to add to the Apertium words that start with an "M".

Apertium as you may recall is a free/open-source machine translation platform. Working on the Apertium lists in OmegaWiki makes sense because its content is freely licensed as well and, the Apertium people are welcome to use it for their purposes.
Thanks,
     GerardM

Tuesday, March 23, 2010

OmegaWiki is back on line

When a technician does not know that a live application is running on a server, he will just turn the server off. That is what happened to OmegaWiki two weeks ago.

What happened next was an amazing group of people getting together, they found someone in the USA willing to go to the office where the server was, take it home and in a multi-continent operation get the data from that system.


OmegaWiki is now operational again; it runs on a server of Erik Moeller, Siebrand, Kim and Marc worked on the issue from the Netherlands and Cyde in the USA. Once the server was up and running, it was Kipcool who got everything working again.

I am really grateful for the hosting we received in the past from Knewco, I am really grateful and happy for the support of so many fine men who brought this project that has so much passion, love, work in it back from the brink.

I have added some translations for the concept Wikipedia, I enjoyed the fact that I could add all the different expressions all in one go. I needed them for some work at translatewiki.net
Thanks,
      GerardM

Saturday, March 20, 2010

Off line (progress report II)

The OmegaWiki server is now connected to the Internet. The data is being moved so happily a positive progress report :)
Thanks,
    GerardM

Friday, March 19, 2010

Off line (progress report)

Sadly things have not progressed as quickly as we would have hoped. I understand that today the server will be picked up to transfer the data from the server and prepare for a relaunch.

More info when I have it ..
Thanks,
    GerardM

Saturday, March 06, 2010

Off line

OmegaWiki is for the moment off line. Our host forgot that OmegaWiki was on their server. It will be available again soon I have been promissed.

It does however mean that we are looking for hosting elsewhere.
Thanks,
     GerardM

Thursday, February 18, 2010

400.000 expressions


Today at OmegaWiki we have reached the milestone of 400.000 expressions for more than 43.000 concepts in 242 languages.

It was longer to reach than expected, because the statistics where first too optimistic by showing 20.000 more expressions than we actually had. This is now fixed.

The statistics reveals that, among these 242 languages, 49 have more than 1.000 expressions, and 10 have more than 10.000, namely English, Castilian, German, Dutch, French, Italian, Portuguese, Swedish, Finnish and Polish. The first seven of these languages correspond to the languages in which the regular contributors are native (or fluent).

What is not known is how many definitions we have in these languages. An improvement of the statistics page is needed.

Thanks,
Kipcool.

Tuesday, February 16, 2010

Adding multiple translations at once

At Omegawiki, it used to be that if you want to add 10 translations or synonyms to a word, you have to add them one by one. This implies that the page is reloaded 10 times. It takes time for the contributor, and load for the server.

But now, with a bit of java, there is the possibility to add multiple translations without reloading the page.

For this, you just have to click on the green "+" and a new row appears.

I give translations as an example, but it works also for definitions, classes, annotations, etc.

I am glad, it is my first steps in the AJAX world, and it will allow us to reach faster the 400k milestone.

Thanks,
Kipcool.

Tuesday, January 26, 2010

Language specific annotations

In order to allow for transcriptions (see previous post), I implemented the possibility to have language specific annotations.

Up to now, it was already possible to have language specific annotations only when that annotation is a list of options. For example, gender shows "masculine" and "feminine" in French, and "masculine", "feminine" and "neutrum" in German (and nothing in English). However, the other annotations (texts and link) were always available for all languages.

Language specific annotations allow for example that the annotation "pinyin" shows up only for words in Mandarin. This is now possible, and configurable by adding the corresponding annotations to the language definedMeaning (e.g. Mandarin (simplified) , Japanese ).

At the moment, the following transcriptions are available:
- pinyin for simplified Mandarin and traditional Mandarin
- revised Hepburn romanization for Japanese
- Hiragana and Katakana for Japanese (as a way of reading a word in kanji)

More transcriptions can be easily added, as soon as a contributor shows interest in it.

It is possible to do more than just transcriptions with language specific annotations. For example, we could imagine to have links to some public domain (for example Webster) or authoritative English dictionaries (Oxford dictionary online), as a way of providing an attestation for a given syntrans (spelling + definition). Such a link would be available only for English words.

Other ideas and thoughts are welcome.

Kipcool.

Tuesday, January 19, 2010

Romanizations in Omegawiki

Romanizations are important to learn a language since it allows us, Latin alphabet readers, to read a word in a non-Latin script.

We are going to implement romanizations in OmegaWiki as text attributes. However, there are several concurring romanization systems for each non-Latin script, and it has to be decided which systems we are going to use.

First of all, there are several ISO norms for romanizations: (Wikipedia link). There are also many other romanization systems which are not ISO norms.

Among the languages and scripts listed in Wikipedia, I have knowledge and interest in Mandarin and Japanese. So, I'll discuss them below. For the rest, help is welcome.

For Mandarin Chinese, it is clear that the ISO norm, i.e. pinyin, is to be used. It is what is in the books when we learn the language, and what appears in almost all dictionaries, it is the most used system to write Chinese on a computer, and it is even taught to Chinese people at school.

For Japanese, the situation is not easy. The ISO norm is the Kunrei-shiki romanization . However, the most widely used seems to be the Hepburn romanization . It is the one that is used in my books at home. So we have the choice between several possibilities:
- should we use the Kunrei-shiki only, because it is ISO,
- should we use the Hepburn only, because it is the most widely used,
- or should we implement both?

For the other non-Latin scripts and languages , if you have knowledge or interest in Cyrillic, Arabic, Hebrew, Greek, Georgian, Armenian, Thai, Korean, Indic scripts or any other script that is not in the list, you are welcome to give your thoughts about the following question (here or at the International Beer Parlour):

"Should we use the ISO norm, or another system?"

Thanks,
Kipcool.

P.S.: there is also the cyrillization of Japanese. Is it a desired feature as well?

Wednesday, January 06, 2010

Etymologies

Thanks to Dh, it is now possible to add etymologies in Omegawiki (in fact, since about a month).

It has been implemented as a translatable text attribute. It means that the etymologies for any word can be explained in all languages.

This may seem like an overkill, and actually, the initial idea was to have only one text field, and to enter the etymology of a word in the language of that word, since we expected that only people who know a language would be interested in etymologies of words of that language.

However, while I might be interested in the etymology of a Latin word, I am not able to write it in Latin, and there is also no particular reason to write it only in English in that case.

What is missing now is the possibility to have etymons as links to the corresponding DefinedMeanings. This is also a desired feature for having the words of a definition linking to the corresponding DM to avoid ambiguity, and will need further development.

Have fun adding etymologies,
Kipcool.

Sunday, December 13, 2009

Linking to external ontologies

At Omegawiki, it is already possible to add links to Wikipedia and Commons.

In order to be part of the Web 2.0 , we are now considering adding links to external ontologies as well. There are many ontologies out there. For example, there is a small list on Wikipedia. From the discussions we had, there are several ways of linking to them:
  1. We only link to Wikipedia and Commons (this is the current situation).
  2. We link to as many ontologies as we like, as soon as there is a contributor willing to add links to it.
  3. We link to a few ontologies that we consider authoritative or relevant.
  4. We link to only one (or two?) super-ontologies, where we expect that this ontology will link back to other ontologies (and ideally to Omegaiki).
There are problems with each proposition.
  1. Why only Wikipedia and Commons?
  2. If we link to too many ontologies, we cannot keep track of what they are for, and then it is expected that we will be lost in too many links. We also have the risk that some user will come with their favorite ontology that do not bring any information to Omegawiki users, and are therefore useless links.
  3. What is an authoritative ontology? What is relevant? and for what purpose (OmegaWiki users, or automatic processing by programs)? In this case, each ontology of possible interested has to be discussed by the Omegawiki community, and it has to be clear what information it brings.
  4. What is this super-ontology? The name of Opencyc has been proposed. Opencyc links for example to Wikipedia, Umbel, Wordnet and Dbpedia. However, Opencyc does not link for example to geonames.
For more details, you can read and give your opinion on the discussions taking place in the omegawiki beer parlour: here and there.

Thanks,
Kipcool.

Saturday, December 05, 2009

Colors in the interface

In order to make OmegaWiki less austere, colors have been added in the interface:
- light red for each language entry
- light blue for each definition or concept
For an example, have a look at Expression:wild

Note that these colors have been chosen to match, more or less, the colors of the OmegaWiki logo ;-)

These changes involve only a modification of the Monobook.css, so each user can also re-configure it to his own taste.

We are looking for more ideas to make the interface more appealing. Please do not hesitate to make any suggestion, should it involve only a css change or even a change in the php.

Thanks,
Kipcool.

Thursday, October 08, 2009

Thank you Brion

Brion has added Kipcool to the illustrious roster of MediaWiki developers. His first contribution to OmegaWiki is a nice one; when you translate from one language to the the next, the definitions will be shown in the source language when available.



This is a much more friendly approach when translating from one language to the next ... Thank you Brion for making it possible to have cause to thank Kipcool :)
Thanks,
       Gerard

Thursday, September 03, 2009

Two improvements

Kipcool has written two nice improvements:
"patch-needstranslation" changes the SpecialNeedsTranslation in that :
instead of link to "Expression:", we have a link to the actual "DefinedMeaning:"
that needs to be translated.

"patch-classes" changes the way the classes are displayed when you want to
add a class to an element.
The current way displays in the first column the name of the classes
and in the second column, the collection to which the class belongs
(which is almost always "Omegawiki Community database")
I have changed it so that the second column now displays a definition of the class.

Note: "patch-classes" also includes additional documentation in the code,
that I am trying to add when I understand what a function does.
Thanks,
       GerardM

Wednesday, August 19, 2009

A new patch went life

Kipcool had written an improvement on the Expression needing translation special page. The page now shows the number of expressions that exist in one language and have no translation in the other language.

I have had fun playing with it, you find that there are words in Hindi that have no equivalent in Dutch.
Thanks,
       GerardM

Wednesday, July 22, 2009

Multi lingual support

When you support languages like OmegaWiki does, then there are two levels where this support is provided.
  • Localisation of the user interface
  • Localisation of the data
With the recent major upgrade of OmegaWiki we moved up from release 1.10a to 1.16a. As a consequence we gained the massive amount of localisations that had been provided to MediaWiki by the fine people at translatewiki.net. This is shown really well in for instance the Russian language..

The complete user interface is now in Russian. We will be able to do even better in the future. We will implement the "LocalisationUpdate" extension; this allow us to continuously adopt all relevant new localisations as committed to the code repository from translatewiki.net.

Making OmegaWiki a resource that is continuously updated for its multi language support is a dream come true. Really important in this has been the perseverence of RobertL. He got to grips with the complicated even convoluted code that makes OmegaWiki. His work included a lot of sanitising of the code base, this will make it not that hard for us to upgrade in the future.

One other upgrade is planned; this is the implementation of the Babel extension. We will be saying goodbye to all the current templates. They have served us well but they have been an absolute pain to maintain. Smaller project like OmegaWiki are best served by getting the localisations for this functionality from one central place. It makes thousands of templates redundant and, new localisations will now be provided on a daily basis through the LocalisationUpdate.

All in all, the latest major upgrade to OmegaWiki has been an important and most necessary boost to our code base. It adds a new sparkle to our project.
Thanks,
GerardM

Monday, July 20, 2009

Upgrading the system

OmegaWiki is expected to go completely offline on Monday 20th July at 14:00 UTC for a major upgrade lasting several hour

Friday, June 19, 2009

ISBN 978-3-03911-799-4

Last year I spoke at a conference in Aarhus, Denmark at the Centre for Lexicography at the Aarhus school of business. I was asked to write an article about what I had to say. I did, and with pleasure I received a complimentary copy in the post titled "Lexicography at a Crossroads".

Sadly the publication is not published as an Open Access work so you will have to find a copy when you want to read my essay "The Philosophy behind OmegaWiki and the Visions for the Future". There is always the presentation that I gave at the conference..
Thanks,
      GerardM

Wednesday, May 06, 2009

WOTD publicatie

The word of the day is publicatie. The reason is, that a book with articles by the people who presented has been published. My article is about the philosphy behind OmegaWiki. Now I will have to figure out how to update my profile at wikiprofessional...
Thanks,
       GerardM

Tuesday, March 31, 2009

Ambaradan

OmegaWiki as a project works; concepts are added and we call them DefinedMeanings, we add Expressions to them and they are either synonyms or translations. We can add part of speech information, we can refer to Wikipedia articles or Commons pictures. All these things we can do in the language selected in your user interface.

OmegaWiki does all these things but there is a problem; the software as it is, is convoluted. A programmer new to the code does not like it. Some hate it with a passion and some looked at it and walked away. This is not good. This is one reason why not much happened with OmegaWiki, the other reason is that we started the development of OmegaWiki mark II.

We created a proof of concept that demonstrated that we can provide multi lingual support to Commons and the next bit was that on the basis of the database backend we would create a new front end, it would be OmegaWiki mark II.

Bèrto went off the grid, it was not possible to contact him in a normal way and now, many months later he appears to have written something called Ambaradan. At this moment there is documentation what it is supposed to do. This documentation is very intriguing and I think that it may even work.

Obviously Ambaradan is welcome to the OmegaWiki content and once there is a user interface to the data, I will be interested to learn how it presents the data. What I will be looking for is how you can configure what information can be entered for a language and how you can relate information entered in one language to information in another.

There are several projects I am involved in that are anxious to learrn Ambaradan's potential. Many things have been on hold and for several projects alternatives are being looked at..

Ambaradan has surfaced, it may be great. At this stage there is not enough to go on.
Thanks,
     GerardM

Wednesday, February 18, 2009

A new user ,,, many new languages

A new user came to OmegaWiki and he wanted to add content in the Mayan languages. All of them "except that funny one from veracruz which only recently got classified as Mayan". All of them because this new user comes with an existing dictionary and this has "all of them".


Now the Mayan languages are not easily classified; The way Ethnologue had it is not how the ISO-639-3 has it at this time. This means that we will have to carefully understand what the relation is between the languages in the dictionaries and the codes maintained by SIL.
Thanks,
       GerardM

Monday, February 16, 2009

OmegaWiki has moved

The hosting of OmegaWiki is on a server by Knewco. Knewco has moved its servers from one location to another and as a consequence OmegaWiki would have been off line for a couple of days. This was not a good plan.

A friend of mine, Tom Maaswinkel, provided us with temporary server space. Kim Bruning, another friend, has done the migration. At this moment, the DNS has been changed and we are testing the new server.

Everything should be working smoothly again.
Thanks,
      GerardM

PS this is the 100th blog entry :)

Monday, January 26, 2009

What a difference a day makes

When you compare this screen shot with the one in my previous blog entry, you will agree that the right to left support has improved quite a lot.

One concern that has been raised is what the MediaWiki developers will say about this. It is obvious that there will be many people who get a different view and consequently the cache will be affected. My idea is that as this is about improved functionality, cache efficiency is secondary to a very large extend.
Thanks,
      GerardM

Right, to the left

OmegaWiki is a multi lingual environment. We have been proud to support your language. When your language is not supported, you can ask. When your language is not properly supported, you can make the difference.
There was one thing though.. Languages like Arabic or Hebrew should go right to left. They did not.
Thanks,
GerardM

Sunday, January 25, 2009

artigo na Wikipédia

When you change the language in your user preferences, the presentation of the labels of the OmegaWiki data will change as well. It will tell you for instance where you can find the Wikipedia article in a given language.

The phrase "artigo na Wikipédia" is todays "Word of the day". Yesterday this translation was added and consequently the user interface for Portuguese became more complete.

Typically phrases like this can be found here, and yes you can make a difference too.
Thanks,
     GerardM

Friday, January 23, 2009

Malafaya had an itch

OmegaWiki suffered for quite some time from a missing functionality. Its annotations had gone away. This was really frustrating. Especially because nobody was looking after the maintenance of the OmegaWiki mark I database. This annoyance became too much and Malafaya started to look into the code. He found and fixed the problem.

Apparently it felt really good, and as there was this other annoyance, he had another look at the code. The trick will be to get the code in SVN. Once it is in SVN, updating OmegaWiki should not be a problem.
Thanks,
GerardM

Sunday, December 21, 2008

Words I do not know

I read and write a lot of English. Every now and again I come across a word, an idiom I do not know. When I have the time I add them to OmegaWiki. When I look at the ones I added today, fecklessness, valedictory, fricassee and enunciate, I wonder if these are the kind of words other people are looking for as well.
Thanks,
     GerardM

Friday, November 28, 2008

Usage

Some words that you look up in a dictionary, have a particular usage associated with it that is not always apparant. The word Balkans for instance is always preceded with "the" when it is used as a noun. The same is true for the Ivory Coast..


This new type of annotation was requested because people wanted to indicate the difference in usage between words that were otherwise synonymous.

Thanks,
GerardM

Wednesday, November 05, 2008

Spanish at 30.000 expressions

Today I am proud to announce that Ascander added the thirty thousandth Spanish expression in OmegaWiki.
Thanks,
GerardM

Sunday, November 02, 2008

The beauty of words

Last Friday [[nl:Afshin Afkari]], presented his Dutch Persian idiom dictionary. The publisher, Amsterdam University Press organised a dialogue with Ashin and Hugo Brandt Corstius in Spui27 in Amsterdam titled "the beauty of words". Going to presentations like this is something that I do rarely. It was however great fun particularly because of the people that you meet at such occasions.

When you read a book like this, you can enjoy a great print. It really looks good. In a book that contains both Latin script and the Perso-Arabic script, it takes more effort to achieve this. One of the things that Afshin included was the transscription of the Persia texts in Eurofarsi. This seems to me a smart move as it a great aid for people who learn Persian. Eurofarsi I am told is used quite a lot by people on the Internet and by people who SMS.

The exposure to Afshin's book made me think on how I would include idiom in OmegaWiki. When you have an idiom in a language, it has a meaning. The question is how do you deal with translations. When you think of it, there are two issues: there is idiom with the same meaning and, there is the need for a literal translation of the text.

It seems to me that you should treat idiom different from normal lexical content. Leaving it a DefinedMeaning without translations but with annotations to equivalent idiom, a defintion of the meaning and literal translations of the idiom itself. Idiom often has a few key words, eg in "blood is thicker then water" you would like to refer to both blood and water. Controry to what is usual in OmegaWiki, you want to refer to these concepts in the same language as the idiom.

The question is, how to support this in OmegaWiki mark II.
Thanks,
GerardM

Thursday, October 30, 2008

Fourtythousand DefinedMeanings


Having 40.000 DefinedMeanings is a nice milestone. The graphic shows how the number of expressions per language tails off.

Congratulations to all the people who contributed to this success. :)
Thanks,
GerardM

Tuesday, October 28, 2008

French has over 20K expressions

I expect that with 20k French starts to become relevant as a resource in OmegaWiki. The good news for our users is the amount of translations a concepts is connected to. Currently the ratio is 9,0402 Expressions per DefinedMeaning.

A big thank you to everyone who contributed to the French success :)
Thanks,
GerardM

Tuesday, October 14, 2008

MySQL Workbench Beta program

SUN, is working on improving the MySQL workbench designer. We have been using workbench for OmegaWiki mark II and have been asking questions and making suggestions. It is with some satisfaction, that we find that we have been invited to the MySQL workbench beta program.

Currently our software uses some 30 business rules and these amount to some 6000 lines of code. They model a complex dynamic behaviour that is responsible for most on the "work behind the scenes". These triggers are currently not modelled by workbench, that can contain them, but not describe their behaviour from the analyst's point of view.

We are happy and satisfied being part of this program.
Thanks,
GerardM

Monday, October 13, 2008

A Commons that supports multiple languages

Commons is great, it is the biggest resource of freely usable media files, some 3.363.734 of them. They are a great resource when you are looking for pictures and when you know English. If you do not know English, it might as well not exist.

Given that the Wikimedia Foundation cares about its educational value for everyone, it makes sense for Commons to be used by people who read and write other languages as well. This dream of supporting all the other people as well is an old one; I wrote about it as early as May 2005.

With support of the Digitale Pioniers, we have been enabled to create a proof of concept project. This project does demonstrate that we can do this. The languages that are currently supported are only a few but sufficient to demonstrate the principles.

The issue is that it is for the Wikimedia Foundation to show an interest. There is a growing group of people who have seen that it can do. When I give presentations about this people indicate that this is a "must have" feature. I will talk again at the Wikimedia Conferentie NL and I hope to discuss with the WMF soon what it takes to support multiple languages at Commons but more importantly if it wants to.
Thanks,
Gerard

Wednesday, September 17, 2008

Wikiprofessional is doing Portuguese

One function of Wikiprofessional that I really like is the Concept Web Linker. If you have not played with it, you should. What I am really happy with is that the Concept Web Linker is getting into other languages. I have an example of this..

This is the screen dump.
Thanks,
GerardM

Thursday, September 11, 2008

Commons but now multi lingual

This is the category "Felis silvestris catus" as it is on Commons with a twist. This screen makes best sense when you can read Dutch.. Obviously, the English text is still there but that does only helps when you understand that language.

Thanks,

GerardM

Thursday, August 28, 2008

OmegaWiki goes Squeak

OmegaWiki was implemented in MediaWiki as an extension during its first iteration. At that time this was a great decision. It provided us with a lot of great functionality and it was the environment that we knew. MediaWiki is a wonderful collaborative application. In OmegaWiki mark II we are looking for other things; this has led to making OmegaWiki independent from MediaWiki.

OmegaWiki intended to do everything in one database. This meant that it was problematic to use our data for other purposes. We also had a situation where different applications wanted their specific data included in the database and needed control of the data involved. This could not be done in the first iteration of OmegaWiki so we had to rethink our ways.

As data may be connected and often will be shared, a peering model is required. Data will be used by many applications and consequently the data needs to be provided in a way that allows for many applications. For the data we will have one interface that uses a standard XML interface to provide the data. This will allow for many applications to use the same data and it will prevent the mixing of data and user interface elements as we saw in the past.

There will no longer be only one database; there will be many. The central or “global” database will provide basic information that will be CC-by licensed. This will allow anybody to have their own “regional” or “local” data and refer to shared concepts for instance for sharing or mapping purposes.

Regional and local databases can be licensed and maintained in a manner that makes sense to the people involved. They can include data that is not really acceptable from a pure linguistic point of view, for instance MALARIA as a synonym for malaria. They can include all kinds of relations between concepts, relations that may be really specialised or that require particular validations before they are published.

Another application that we have always wanted to give our data was was the OLPC and equivalent networks. This means that our database has to be able to function stand alone and be synchronised and share improved data when a connection becomes available.

As a consequence of all these considerations, we have been looking for the best technology that will serve our purpose; MySQL 5.1 will provide us with new functionality that makes an important difference. Squeak is a programming platform that we think will provide us with the tools to build the rich environment that we dream of having.

As the OLPC project also uses Squeak, it will allow us to bring our information to this great educational project, and in return we hope that people will find OmegaWiki an environment to contribute their Squeak work to and help us build dictionaries in the many languages spoken and written where the OLPC will become available.

Thanks,
GerardM

Tuesday, June 17, 2008

Fast and furious

I just learned that OmegaWiki mark II has been updated; it now allows for multiple statements. Bèrto indicated that he will update the documentation to reflect this.. The example code I have does not get posted because the Blogger software wants to interpret it ...

I am sure that it will find its place in the documentation ... So, it is for you to RTFM :)
Thanks,"
GerardM

OmegaWiki mark II

It has been relatively silent around OmegaWiki, this silence however is very deceptive. Much hard work has gone into preparing for a new codebase for OmegaWiki. The new code has to deliver new functionality in order to justify the huge investment in time and money.

Why change..
  • The existing user interface and the database routines are very much etwined, we want to separate them
  • We want to provide services based on the OmegaWiki data, we will provide an XML interface into the data
  • The underlying technology has changed a lot; we are using the bleeding edge of the MySQL database
  • Our data can be used for applications; we will provide a way to separate data that is of general interest and data that is not
  • For some applications it is really helpful when the data can be used in an off-line environment; we will provide a way of synchronising databases
  • There is more ...
The first data in the OmegaWiki Mark-II environment is now available. It is converted data from OmegaWiki; it is the DefinedMeanings with a SynTrans record in English.. You can find it here.

At the bottom you find a document explaining about the API; it is written in the Open Office format..

Enjoy, have fun and tell us what you think of it. The data is experimental so we will replace it with more data at a later time.
Thanks,
Gerard

Sunday, June 08, 2008

Internationalisation and localisation

When you write software, when the software is to be used by people who speak many languages, internationalisation is a key requirement. It is the precursor to localisation; the changes made by the localisers to support their language.

MediaWiki is really good at localisation and, it is being perfected all the time. When software is written and when internationalisation / localisation is not considered from the start, it is quite a job to get this right.

The OmegaWiki Vocabulary Trainer needed internationalisation / localisation badly. The first issue with the software; it was not German. The software was paid by the University of Bamberg to be used by its students. These students are expected to know German and this trainer is a tool to aid them to learn languages. OmegaWiki is very much dedicated to making information available in many languages and not considering internationalisation in its associated software is ... odd.

It is with relief that I can now announce that the Vocabulary trainer is supported in Betawiki. I am grateful to the kind developers at Betawiki who made this possible. You can localise for your language here ...
Thanks,
GerardM

Tuesday, June 03, 2008

Getting a better look and feel ..

The vocabulary trainer was introduced... It looked awful; it now has been improved considerably. Like we wrote last time, it is open source and it is being worked on. One of the things on the "must have" details .... localisation :)
Thanks,
GerardM

Thursday, May 29, 2008

Publish early, publish often ...

For Open Source software, one of the mantras is to publish early and publish often. In this spirit I am happy to announce the first version of a vocabulary trainer on OmegaWiki.

This software uses the content of OmegaWiki; so the quality of the exercise is also determined by the amount of terminology that is contained in OmegaWiki.. The current functionality is not yet feature complete.. One of the things we want to do with this software is have a list of words and phrases that are good to know when you go abroad.. Wikimania 2008 anyone ??

NB the vocabulary trainer can be accessed from the OmegaWiki main page :)

Thanks,
GerardM

Wednesday, April 16, 2008

Georgian

Georgian is a language that started with no content in OmegaWiki. It is spoken by some 4.4 million people mainly in Georgia, Turkey, Iran and Russia. It is written in the Georgian alphabet.

Statistics are in and off themselves not that relevant but they do allow you to tell a story. OmegaWiki started with the content of GEMET, the GEneral Multilingual Environmental Thesaurus, the languages that are part of this resource have a head start in size.

Today thanks to the hard work of Sopho, Georgian is the first language that grew bigger then one of the languages supported by Gemet. The OmegaWiki statistics show that Japanese might be the next language to grow bigger then Slovenian..

Because of the beautiful characters, Georgian is my favourite example of showing the value of the localisation in OmegaWiki. It is really special to see the same content optimised in such a way. :)
Thanks,
GerardM

Saturday, April 05, 2008

Unicode 5.1

Today I learned that Unicode 5.1 has been released. The information that I received informs me that one major feature will be of particular relevance to Japanese, Chinese and Korean texts by enabling ideographic variation sequences. The linebreaking for Polish and Portuguese hyphenation has been improved. The Indic languages will be happy with improved text segmentation algorithm.

There are 1624 new encoded characters, this includes characters required for Malayam and Myanmar but there are also new characters for the Latin script. New is support for the Cham, Lepcha, Ol Chiki, Rejang, Saurashtra, Sundanese, and Vai scripts.

For the techies, the collation algorithms have been updated to include all the new characters. This has also an effect on contractions like the ch in the Slovak language.

Many of these things have an effect on languages supported in Wikimedia projects. My question is when will we have support for this. Is this a function of the MediaWiki / PHP code and is it also a function of the browser ??
Thanks,
GerardM

Tuesday, March 18, 2008

OpenStreetMap terminology

OpenStreetMap for those who do not know the project, is a free editable map of the whole world. Its data is freely licensed, it is build by volunteers and it is very much a work in progress.
OpenStreetMap has its own terminology, and given its origin it is British. The maps can be quite good providing you with sufficient information to plan your route.. There are routes for pubcrawls and other innovations :)

As OpenStreetMap intends to provide a map of the world, the people that make these maps have to literally put themselves on the map. In the Netherlands we have been blessed because maps have been made available to the project. A friend of mine is not yet on the map..
It is clear that the terminology used in Italy is not the same as in the UK or the Netherlands. It makes sense for an Italian to have Italian terminology available to him and Dutch would do me nicely.

At a meeting of the Digitale Pioniers, I met people of OpenStreetMap and we agreed to include their terminology in OmegaWiki. I have now entered the words that have the key "highway" and invite you all to come to OmegaWiki to add translations in your language.

Special in this data is that i have added the definitions provided by OpenStreetMap as alternate definitions as well. In this way they have been marked as definitions provided by OpenStreetMap.

All the OpenStreetMap terminology in OmegaWiki can be found here.
Thanks,
GerardM

Tuesday, March 11, 2008

Less is more

There are two types of content in OmegaWiki; there are the standard MediaWiki namespaces with the standard MediaWiki content and there is the OmegaWiki specific content. The specific content contain database records, they are the DefinedMeanings and Expressions.

As they are handled completely differently, it means that the functionality available to these types of data is different as well. Particularly the "move" functionality has given us problems in the past. This broken functionality has now been removed from the screens.
Thanks,
GerardM

Firefox 3 beta 4 is significantly faster

OmegaWiki is a web application that requires a lot from a system. I am really happy to report that the latest beta of Firefox gives me a significantly better performance on the same hardware. I reported in the past that the Firefox beta did a good job for me, but this time the performance is noticeably faster on a page like Nederland.

Firefox provides cutting edge technology and really makes a big difference to me. Now when you want OmegaWiki to perform even better, you can choose to move to this latest software. The other good news is that the spell checkers that are relevant to me are now available as well...
Thanks,
GerardM

Saturday, March 01, 2008

Nagios

OmegaWiki has a technical problem; there are certain records that have problems and that crash the database. There are solutions to this problem and some are being tested at the moment.

This week I was in Bamberg at the Otto-Friedrich University, and discussed this with Martin Mai. He has now build a monitor that checks if OmegaWiki is still alive. This works fine. We now have permission for the Bamberg Nagios service to run a script when OmegaWiki is no longer alive.

This will solve one of my biggest worries which is the availability of the OmegaWiki service.
Thanks,
GerardM

Wednesday, January 16, 2008

Bounties for the localisation of MediaWiki

Hoi,
The "Stichting Open Progress" is happy to announce that it has received a grant from Hivos, to improve the localisation of MediaWiki. Open Progress is going to offer a bounty of up to 200 EURO for the full localisation for a language. Given the activity of Hivos, a Dutch NGO, a bounty will only be available for languages in Asia, Africa and Latin America that have a sizable number of speakers.

With this project we hope to achieve that MediaWiki can indeed claim to be one of the best Open Source projects that provides great localisation for many many languages out of the box. It will improve the usability not only for the WMF projects in that language but also for projects like Wikimedia Commons, Wikieducator, Wikihow, OmegaWiki .. the list goes on ..

The budget we have is substantial but limited. We will be sadly happy when we have to announce that we have ran out of money. Sad because we want to localise more languages, happy because so many languages will have been improved.

For the precise details of the project I refer to the details on Betawiki.

NB The amounts are inclusive of the money transfer costs and as these can be substantial, we offer some alternatives. In the past I have proposed a scheme called "Donations, putting your money where your mouth is". In this scheme you choose to get paid or donate the money for one of the projects that we advertise under this scheme. Another way of getting the money paid is when people in a country agree to work together and have us pay the money together as well. This could for instance work in the case of Wikimedia India...

Thanks,
GerardM

Tuesday, January 01, 2008

2008, International Year of Languages

The new year, 2008 has been designated the International Year of Languages by UNESCO. Countries and organisations are invited to participate, and I think that what we do in our projects qualifies us as participants. Our projects are relevant for many languages, we welcome new languages and we provide help and infrastructure to make it a success. Our communities are knowledge societies in which everyone can participate and benefit. We promote universal access to information and ensure in this way the use of an increasing number of languages.

Siebrand wrote an overview of the localisation of MediaWiki. Real support is provided to some 170 languages but only 47 have a minimal localisation. In a way it mirrors our projects; our projects do great in some languages while at the other end language versions are closed because there is not enough of a community that supports them.

The Wikimedia Foundation is becoming more mature; we aim to ensure that all our projects do well. Barriers to entry have been in place for new languages and it has led to improved localisation of MediaWiki. When new projects are finally approved, they are already of a size both in articles and participants that there is less need for anxiety for their future.

When we are to participate in the Year of Languages and continue to do what we do well and improve where we are weak, our projects will prove to be credible participants in this Year of Languages. Participating will give our Wiki way more credibility and it will give us access to people and organisations that can help us in fulfilling our goal; sharing the sum of all knowledge with every single human being.

Thanks,
GerardM

Sunday, December 30, 2007

Numbers in a translation dictionary

In the Wikimedia Foundation, there are the inclusionists and the exclusionists. Some are of the opinion that certain topics should not be included in Wikipedia while others do. The most brilliant example is the inclusion of all the busstops in Japan. Someone took the effort to describe them and people find them actually useful.

Some people are enamoured by constructed languages and spend a lot of time making such languages their own. I personally have had dealings with at least three people that speak Volapük, and I know people that go to congresses because they meet people who speak Esperanto. Many constructed languages have more speakers than many natural languages (that are not yet extinct).

For exclusionists it is not palatable when constructed languages do well. There are always "good" reasons why those others need to be excluded. Hidden in the discussion about the "radical cleanup of the Volapük Wikipedia" is a discussion about the inclusion of numerals like 588 in the Limburgian Wiktionary and as you can imagine "it is not good".

In a translation dictionary there are reasons to include numbers. The point is that they are not written the same in all scripts. OmegaWiki has its fair share of numbers, and it did not address this issue.

By adding a new class, we now allow the representation of Arab and Roman numbers, and hundred has now as an annotation both 100 and C. In this way we do not have a separate page for each numerical representation.

Thanks,
GerardM

Tuesday, December 18, 2007

Languages ...

Given a fixed point and you can move the world. In many ways, your language provides you with the tools to wrap your mind around the world, express its essence in a way that may be understood by the people you communicate with and help you to shape the world as you know it.

Language is both individual and shared. My English has been shaped by my schooling in the Netherlands, my stay in the United Kingdom and the many times I expressed myself in e-mail, articles, presentations and when skyping. Typically I get it right but a text can be understood by some and misunderstood by others. It is my English but to function it has to be expressed in a way that is shared with others.

For some subjects I prefer my native language, for others I prefer English. To communicate, the language used must be sufficiently shared by everyone involved. The language must be received; when I am surrounded by a vacuum, nobody will hear me talk. To read this blog, you either have access to a computer or someone must print it out for you.

A language lives when people use it, when it is part of a distinct community, a distinct culture. When the boundaries around such a community or culture disappear 0r change, the language either morphs or it dies. To understand history, you have to understand its artefacts and its language. Many languages die and died and with it we lose the history, the culture of the people that spoke that language. They may leave their literature, inscriptions and when enough is left, we may understand what is says. The trick will be to understand it as it was meant when the language, the culture was alive.
Thanks,
GerardM

Saturday, December 15, 2007

Whe needs birthPlace

With some regularity I try to better understand Semantic Web and associated subjects. I find it hard going but also a compulsive subject. When you express the relation "Johan Cruijff" "birthPlace" "Amsterdam", it is understandable to you as a reader but for humans it should read like "Johan Cruijf was born in Amsterdam" or "Johan Cruijf werd geboren in Amsterdam" .. This magical statement "birthPlace" can be interpreted when you know your English otherwise it is truly for machines only.

OmegaWiki does express relations, you will find for instance that Amsterdam is the capital of the Netherlands. In essence it is expressed as a triple, but it is expressed in natural language and depending on the existence of a translation, you will read the relation in the language selected as your user preference.

How to combine what we do and what happens elsewhere, my latest idea is based in the RDF tag; "birthPlace". It is a construct that obviously needs a natural language equivalent and this is what OmegaWiki can provide. A method is needed to connect the two. In order to function, birthPlace has a very precise definition and this definition must be part of a collection of such definitions. These labels need to be linked to OmegaWiki DefinedMeanings as the identifier for an OmegaWiki collection.

To make this useful, an external application needs to call a function that provides the translation to a specified language. How to combine this with the notion of an URN I have not figured out yet.

Thanks,
GerardM

Thursday, December 13, 2007

Eastern Yiddish

Eastern Yiddish, is one of the two varieties of Yiddish that have been recognised as languages in their own right in the ISO-639-3 (ydd). in OmegaWiki, we now have our first 213 Expressions in this language and I am impressed with the amount of work that has gone into it; most have annotations indicating hyphenation and the pronunciation using IPA notation.

Eastern Yiddish is not one of the languages supported by MediaWiki, and the mechanism for showing localised content is connected to the language selected in the "User Preferences". I have been given some help from Siebrand what files need to be changed and added. Kim helped me with doing it for the first time and now the first localisation is visible for Eastern Yiddish.

The MediaWiki localisation itself uses Yiddish as the fall back language so the experience is pretty good for now. What Siebrand indicated is that is is possible to include languages like Eastern Yiddish in the BetaWiki. This would create stubs that are of benefit to OmegaWiki. It would prepare for the moment when people start localising in earnest.

I think it would be a good thing, but I am interested to learn what other people think.

Thanks,
GerardM

Monday, December 03, 2007

Supporting American English

American- or British English are two variations of the English language. They have substantial differences. They are sufficiently the same and are unlikely be mistaken to be separate languages.

In OmegaWiki, it has been possible to add entries for English; this meant there is no difference between written the different versions of English or you had to specify both versions. Issues like this exist for other languages like Serbian and Mandarin as well.

In the OmegaWiki user interface, languages are considered ISO-639 entities. When a DefinedMeaning for a language is part of the appropriate collection, we use the translations in our user interface. The problem is that all these linguistic entities are needed now and that they are created to make OmegaWiki work.

For the ISO-639-6 there will be issues as the codes we make, using the RFC 4646 methodology, will be replaced. It will also be interesting to learn how in the end everything will be merged together.

In the mean time we now support localisation for these linguistic entities.

Thanks,
GerardM

Sunday, December 02, 2007

Sinterklaas present

In the Netherlands we traditionally do not get presents with Christmas. For us Sinterklaas is celebrated on the fifth or sixth of December. In my family everyone no longer believes in Sinterklaas and consequently we can celebrate it on a more convenient moment like in a weekend.

I have had a wonderful Sinterklaas, and I do want to tell you about the present that Kipcool and Kim gave me. Kipcool wrote this wonderful functionality that show you what classes we have in OmegaWiki and, how many translations we have in your language.

With the new functionality you will see the concepts that are translated in your language. When you check out a concept, you will even find what attributes are available in your language. I have found it to be really addictive. :)

Thanks,
GerardM

Monday, November 19, 2007

New upload functionality

OmegaWiki is really happy to announce that we have, with thanks to the Otto-Friedrich University of Bamberg, for the first time used new upload functionality. The University of Bamberg has a need for a repository for its Destinazione Italia terminology and has found this in OmegaWiki.

We have uploaded translation in Persian and we have it nicely attributed to the person who did the work. We will upload for several more languages. What is of relevance is that we can make an export for any language, for any collection. So if you are interested to help on our OLPC collection.. just drop me a line..

Thanks,
GerardM

Friday, October 19, 2007

Money

All projects need money to operate. OmegaWiki does need money to operate. We now have made it possible for you to support what we do in a practical way; we now have a link in our sidebar so that you can use Pay Pall to donate money. You can just give us money, you can help us fund a project that we want to do.. Check out our Donations, putting your money where your mouth is ..

Stiching Open Progress is a Dutch "not for profit" organisation and we can use all the money we can get to do all the cool development work we would like to do..

Thanks,
GerardM

Tuesday, October 16, 2007

Zimbabwe was formerly known as Rhodesia

Many countries have over time changed their nature. What they typically do is stay more or less in the same shape. As a consequence of war, the shapes do change. The change is often reflected in the name. The country that is now called Zimbabwe was once called Rhodesia. It is relevant information and can be expressed using relations.

In OmegaWiki the relation type "was formerly known as" has been introduced to express this relation for countries. It demonstrates that OmegaWiki is not strictly a dictionary, it also serves the functions of a dictionary. By including different types of attributes to classes, we provide more worthwhile information.

Most concepts are related to other concepts and when these relations become visible, a net develops of related information. This does not make OmegaWiki an encyclopaedia, it is what an ontology does. An encyclopaedia we are not; we refer to Wikipedia.. :)

Thanks,
GerardM

Friday, October 12, 2007

Linking to Wikipedia

OmegaWiki may be a lot, but it is not encyclopaedic. We do not want to be; Wikipedia does a great job at it and when it needs competition there are plenty of pretenders to its throne. So we do not compete.

When people need information, OmegaWiki will not provide all information. What it can do is link to other sources of information and Wikipedia is the obvious and the only choice for encyclopaedic information. It is the only choice because it aims to be multi-lingual and it is an obvious choice because of the shared values.

At this stage, linking to the Wikipedia articles is done by hand so initially there will be few links. We hope to harvest these links from Wikipedia and insert them with a bot. In this way we will provide an encyclopaedic service without being encyclopaedic :)

Thanks,
GerardM

Friday, October 05, 2007

Antonym

An antonym is the complete opposite of something.. black and white are probably the best known examples. The great thing is that antonym is the first global relation type and in the way it is set up, the antonym is true on a concept level. This means that it does not allow for cultural differences in the appreciation of such a relation.

I wonder how many antonyms will prove to be problematic because of cultural differences. The good news is that we are now able to have global relation types in OmegaWiki. We will have to be REALLY careful what relation types we will include. The "is a" relation is not going to be part of it because that is what makes something a class member.

With the global and the class based relation types we only need the collection based relation types to get our full functionality :)

Thanks,
GerardM

Tuesday, October 02, 2007

More localisation

OmegaWiki had a lot of new functionality go live, today I spend time on one aspect that is really dear to me; localisation. Much of the OmegaWiki content is localised by adding translations. With the latest software release much of the more programmatic parts are in the system messages.

I started to translate the Dutch messages and, Tosca caught on and started on the German messages, Malafaya did the Portuguese. We hope that in this way our data becomes even more accessible to our users :)

One issue remains, with our messages translated in OmegaWiki, how do we get them in MediaWiki proper ...

Thanks,
GerardM

Getting to grips with the new functionality

At OmegaWiki, a lot of new functionality has gone on line. This functionality is a mix of functionality that was needed for Wikiproteins and things we have been working towards for a long time. With the changes some functionality does not work as it used to. This is a good thing.

Our first content was the GEMET thesaurus, in this collection particular relation types were used. These relation types were available everywhere and consequently we have been reluctant to add more relation types. Now relation types are associated with "classes" and we can make a DefinedMeaning a member of a class. Nederland now has a capital, a motto, a nation anthem and entities bordering the country. For Nederlands it is now known what script it is written in, and in what countries it is spoken. And with the "incoming relations" we know where there is a reference to the DefinedMeaning.

Many of the existing relations will be changed from the GEMET relation types to the new relation types. The work that what is done in the past is a huge benefit as it helps a lot in identifying what needs doing. With the new functionality it makes sense to add the annotations straight away; we now know that they can be done properly.

Thanks,
GerardM

Friday, September 28, 2007

Major update for OmegaWiki

OmegaWiki has had some major update; the version of MySQL that is installed has been updated, several files have been changed to InnoDB and a lot of functionality has changed behind the scenes.

One of the effects is that the performance has improved noticeable; that was really needed. The difference in performance is a relief. It is fun again to work on the data.

One difference is that the way the relations work; relations are currently associated with a "class" and this class defines what relation types are possible. We have added a few classes so far; "linguistic entity" is one. The associated relation types allow us to indicate where the linguistic entity fits in and, where it is spoken. There will be many more classes and relation types, the quality of the classes and relation types will make a difference to the quality and the usefulness of our data.

Thanks,
GerardM

Friday, September 07, 2007

Demo Semantic Support on a new URL

At Wikimania I presented what we are doing to bring real time Semantic Support to Wikipedia. The URL in my presentation is no longer valid, the new location is at: wikipedia.wikitestsite.org.

You will find a dump of the English Wikipedia and you many of the expressions that we already now are in green. We are working towards a situation where new concepts defined in OmegaWiki will be recognised in the future data mining of the same article.

What we are discussing at the moment is adding functionality to the concepts found. Some are obvious like giving the definition, giving an option to go to OmegaWiki when the definition does not fit, showing translations for the expression in the language that is of interest to the reader.

We can imagine that there is more functionality that you would consider useful. Please let us know .. :)

Thanks,
GerardM

Monday, September 03, 2007

Connecting data from different databases

In OmegaWiki there are different datasets. These represent different origins and have a different emphasis. What we are working on is to connecting the data in these different datasets. Currently over four percent of our Community data is connected to data of the UMLS.

These connections are not without problems. The UMLS does not have the same (lexical) outlook; it is quite happy to have a singular and a plural to be part of the same concept. In OmegaWiki we do not support the notion of plurals yet. For the UMLS it is not a problem to include Geologists as it is included as a subject heading. We have it connected to geologist.

Lyme disease has several synonyms that are problematic from a lexical point of view; only "Lyme borreliosis" is what I expect to find in a dictionary. This does not necessarily mean that "Borreliosis, Lyme" is not useful to have. The Community database knows some 15 translations and thereby adds value to the English only content for Lyme disease.

With four percent of the Community Database connected, in reality we haven't scratched the surface of the UMLS. The UMLS is a well explored resource and I am sure that there are many resources that have made connections already. I hope we will find the people, the organisations willing to share the work that they have already done.

Thanks,
GerardM