Showing posts with label collaboration. Show all posts
Showing posts with label collaboration. Show all posts

Sunday, November 06, 2011

The Glosbe dictionary uses our data

Glosbe.com is a multilingual dictionary where many bilingual translation dictionaries are available. It started earlier this year, and already have an impressive number of translations.

Their data come from several freely available online resources, including OmegaWiki. For example, the many words I have added in Bavarian are there:
http://glosbe.com/bar/en/essn

It also shows definitions from OmegaWiki, and inflexion tables from Wiktionary.
They have managed somehow to merge all these sources, and present them in an organized manner. For an example of a page with many sources involved, see: http://glosbe.com/fr/en/manger

Furthermore, Glosbe is also a translation memory, and they claim to have gathered more than 1 billion sentences already. I think this is a great resource for translators. New features will come in the future, they told me.

It is always nice to learn that someone uses our data. It makes our work meaningful!

Sunday, January 25, 2009

artigo na Wikipédia

When you change the language in your user preferences, the presentation of the labels of the OmegaWiki data will change as well. It will tell you for instance where you can find the Wikipedia article in a given language.

The phrase "artigo na Wikipédia" is todays "Word of the day". Yesterday this translation was added and consequently the user interface for Portuguese became more complete.

Typically phrases like this can be found here, and yes you can make a difference too.
Thanks,
     GerardM

Tuesday, March 18, 2008

OpenStreetMap terminology

OpenStreetMap for those who do not know the project, is a free editable map of the whole world. Its data is freely licensed, it is build by volunteers and it is very much a work in progress.
OpenStreetMap has its own terminology, and given its origin it is British. The maps can be quite good providing you with sufficient information to plan your route.. There are routes for pubcrawls and other innovations :)

As OpenStreetMap intends to provide a map of the world, the people that make these maps have to literally put themselves on the map. In the Netherlands we have been blessed because maps have been made available to the project. A friend of mine is not yet on the map..
It is clear that the terminology used in Italy is not the same as in the UK or the Netherlands. It makes sense for an Italian to have Italian terminology available to him and Dutch would do me nicely.

At a meeting of the Digitale Pioniers, I met people of OpenStreetMap and we agreed to include their terminology in OmegaWiki. I have now entered the words that have the key "highway" and invite you all to come to OmegaWiki to add translations in your language.

Special in this data is that i have added the definitions provided by OpenStreetMap as alternate definitions as well. In this way they have been marked as definitions provided by OpenStreetMap.

All the OpenStreetMap terminology in OmegaWiki can be found here.
Thanks,
GerardM

Friday, September 07, 2007

Demo Semantic Support on a new URL

At Wikimania I presented what we are doing to bring real time Semantic Support to Wikipedia. The URL in my presentation is no longer valid, the new location is at: wikipedia.wikitestsite.org.

You will find a dump of the English Wikipedia and you many of the expressions that we already now are in green. We are working towards a situation where new concepts defined in OmegaWiki will be recognised in the future data mining of the same article.

What we are discussing at the moment is adding functionality to the concepts found. Some are obvious like giving the definition, giving an option to go to OmegaWiki when the definition does not fit, showing translations for the expression in the language that is of interest to the reader.

We can imagine that there is more functionality that you would consider useful. Please let us know .. :)

Thanks,
GerardM

Monday, September 03, 2007

Connecting data from different databases

In OmegaWiki there are different datasets. These represent different origins and have a different emphasis. What we are working on is to connecting the data in these different datasets. Currently over four percent of our Community data is connected to data of the UMLS.

These connections are not without problems. The UMLS does not have the same (lexical) outlook; it is quite happy to have a singular and a plural to be part of the same concept. In OmegaWiki we do not support the notion of plurals yet. For the UMLS it is not a problem to include Geologists as it is included as a subject heading. We have it connected to geologist.

Lyme disease has several synonyms that are problematic from a lexical point of view; only "Lyme borreliosis" is what I expect to find in a dictionary. This does not necessarily mean that "Borreliosis, Lyme" is not useful to have. The Community database knows some 15 translations and thereby adds value to the English only content for Lyme disease.

With four percent of the Community Database connected, in reality we haven't scratched the surface of the UMLS. The UMLS is a well explored resource and I am sure that there are many resources that have made connections already. I hope we will find the people, the organisations willing to share the work that they have already done.

Thanks,
GerardM

Monday, July 30, 2007

UMLS

The UMLS or Unified Medical Language System is a collection of many resources it contains tools, a semantic network and a specialist lexicon. It is also a collection of many resources. These resources all have their own license and copyright. Effectively much of the UMLS can be used for many purposes because the particular license allows it. In the same way, there is much of the UMLS that can only be used when the copyright holder gives permission.

In OmegaWiki, we have our first Authoritative Database online. It is the UMLS and we are proud of it. Now as the UMLS is this collection of connected resources, we present it in the same way. There is one UMLS as an Authoritative Database and it has collections that are the parts that make up the UMLS as we have it. The important thing of the UMLS is that it did make the connection between the different databases and we do use their system to connect.

What makes the inclusion of the UMLS so special is that we have the cooperation of the NLM. It is what makes this such an exciting experiment.

Thanks,
GerardM

Sunday, June 17, 2007

OmegaWiki only a translation dictionary ?

There is some misinformation about OmegaWiki, it is said for instance that OmegaWiki is only a translation dictionary. There are also people who do not consider OmegaWiki as relevant because it is not a Wikimedia Foundation project.

It is for the people that have not looked at OmegaWiki for a long time or have not really looked well that we want to state the obvious; OmegaWiki is not only but also a translation dictionary. When you look at the number of expressions per language, you will find that we have almost 30.000 DefinedMeanings, the reason why we have 11.000 more English Expressions then what we have for any other language is because we have collections that are at still mostly English. Collections like the ISO-DIS-639-6 are relevant because of the information that is included in the data.

OmegaWiki is becoming relevant because our data is starting to be used outside our project as well. Positano News uses OmegaWiki data for "assisted reading", this helps people to understand terminology that is in an Italian news article. It does give you definitions and translations.

It may be that the current possibilities at OmegaWiki are not immediately obvious; there are many DefinedMeanings that do not have any annotation. An annotation can identify the part of speech for a word, it can provide you with a sample sentence or how to hyphenate a word. We want to include links to other websites; we want to link to Wikipedia articles in order to make it convenient to our users to find good encyclopaedic information.

OmegaWiki is not feature complete. We want to add many more features, but our first priority is to make sure that it works well and that the features that matter most are included. We need to improve on our performance and, we need to make sure that we provide a framework that facilitates collaboration with other organisations.

The Wikimedia Foundation is one organisation that we really want to collaborate with. On a personal level we have been involved and we want to extend this by collaborating on an organisational level as well. This often repeated intention may be one reason why certain people are so apprehensive about OmegaWiki; we wanted it to be a WMF project, it is not a WMF project but we still see room for doing good together.

Thanks,
GerardM

Friday, May 11, 2007

OmegaWiki supports many linguistic entities

OmegaWiki aims to include all words in all languages and provide both lexical, terminological and ontological information. As the discussion of what makes a language is an endless one, those languages that are included in the ISO-639 codes are the ones that are supported.

Having chosen the ISO-639-3 to start of with has proven a great start. It did however not provide the granularity needed to categorize words to their linguistic entity. How to deal with languages that are written in several scripts, how to deal with regional differences? This is what this iteration of the ISO-639 standard does not deal with.

The implication is that this standard on its own does not suffice. By combining the data with other standards with other codes it is possible to provide more granularity, but how to deal with dialects like Westfries, that is spoken in the area where I grew up?

As the OmegaWiki project was evolving and getting traction, I got into contact with Debbie Garside. She is heading Geolang an organisation that has been preparing for a long time the next iteration of the ISO-639 standard, the ISO-639-6. The aim is to include at least 25.000 linguistic entities in a hierarchical structure. Adopting this data would allow OmegaWiki to better achieve its aim; include all words of all languages.

When a standard is published there is a prescribed period in which the public is invited to comment on a standard. So far this has been done using e-mail. Experience shows that when the amount of subject is too big, e-mail is not a tool to cope. Geolang had explored the option of using Wiki technology before, this sadly did not lead to the right synergy. In OmegaWiki however, there was both an active interest in language standards, it included not only the Wiki methodology, it even allows for the inclusion of the data in a true hierarchical way.

By publishing the data in a wiki, in essence everybody with an interest in orthographies and dialects is invited to comment, modify and add to the hierarchical data. To make this into a standard, there will be a need to assess the community generated data and assert the validity of the information provided. This is where the World Language Documentation Centre will play its role. As its name implies, it documents languages and it will do so in the broadest sense of the word. Obviously an organisation like this will only function well when it is as an organisation an inclusive organisation. The make-up of the current board reflects many specialities that make up linguistics and the language industry.

It is with a fair amount of satisfaction that I can announce that Sean Burke, one of the volunteers of OmegaWiki has imported the first batch of the ISO-DIS-639-6 data in time for the inaugural meeting of the World Language Documentation Centre. Both OmegaWiki and the WLDC will rely on collaboration, to get the necessary work done. Our challenge will be to provide the infrastructure and the minimal organisation to start and sustain our projects.

With the inaugural meeting, the WLDC it is proclaimed to the world that as an organisation the WLDC is ready for business. With the first data available in OmegaWiki, the first request to the world to collaborate on the languages that are spoken the orthographies that are written goes out. it is the start of acquiring the meta data that helps us understand the data that is already out there and consequently make from all this data information because we will become better able to parse the data.

Thanks,
GerardM

Monday, March 19, 2007

How do the four freedoms apply to one database?

The FSF defines four freedoms when it comes to software. What kind of freedoms applies to a database like OmegaWiki.

The OmegaWiki software is licensed, like MediaWiki what it is an integral part of, with a GPL license. This means that you can use the software as is. As OmegaWiki uses specific data structures that can be licensed separately for completeness sake, the database design is also available under a GPL license.

The data that is contained in OmegaWiki is licensed under a combined GFDL/CC-by license. Many people insist that these licenses are not compatible. At issue is that the data are just facts, it is only possible to copyright facts as a collection. We want people to make use of our collection. For us success is: "when people find a use for our data we did not think off".

We invite people to collaborate on our data, when they enter some Babel templates on their user page, we give them edit rights to the data. We invite organisations to collaborate on our data because there is so much data that organisations can share, there is so much labour invested in the type of data OmegaWiki can be a home for.

So the data and the software is Free. How about OmegaWiki itself ..

As there is only one OmegaWiki.org, the room to do whatever is limited. The data has to be useful to everyone and it has to fit in with the notion of the DefinedMeaning. When domain specific data is added, it needs to be domain specific and, there has to be agreement that this data provides a suitable extension for people involved in this domain.

When people find that this is not enough, they can have their own database. This does not mean that they cannot cooperate. Much of what they need in terms of extra functionality will be shared. It means that even when the OmegaWiki database is forked, there is still plenty of scope to improve all the things we do agree on and collaborate on those.

There is plenty of Freedom. All the Freedoms we provide. I think however we achieve the most success when we find that there is more that binds us than that drives us apart. It means that we have to work hard in understanding what our common needs are.

Thanks,
GerardM

Sunday, February 11, 2007

Why compete when you can collaborate ?

All words of all languages of the world.. that is what we eventually aim to include in OmegaWiki. This aim is of such a magnitude that you have to be certifiable to come up with such a project. The functional design for the project includes much more; everything including the kitchen sink..

When everything is to be included in one project, it is easy to suggest that people contribute to the project. When the project includes everything why have another?

In an Open Source / Open Content environment this is not necessarily how it works. Why should the others be seen as competitors? They do their own thing, sure. You may want to achieve the same thing, also true. It is however much possible to find the synergy between projects. This way you can build on each others accomplishments.

The Shtooka project is something I learned about the other day. The one thing it does really well is the way they make recording pronunciations easy. You can record a string of words and it will save them for you one at a time.

Wiktionarians saw this and they are working an upload facility so that it will also be saved automatically to Commons. I warned that the files should not only be saved as .ogg files. In order to make sure they are relevant for scientists there should also be a .wav file. The current thinking is that the flac file format will work as well and the benefit is that it provides a loss-less compression. To make sure that this is the case, the praat software, software that is also available under a GPL license, was analysed and it was considered that it is easy to incorporate this flac file format.

People from effectively five different communities are now working together. It will be even possible to include links to OmegaWiki in the Shtooka meta data. This will be possible even though both projects do their own thing. Both the data and the functionality can be shared.

I may be certifiable, but this kind of collaboration is awesome and, it is why there may be method to this madness.. :)

Thanks,
GerardM

Monday, January 15, 2007

Destinazione Italia

Destinazione Italia is a project of the University of Bamberg. It provides training for people learning an advanced level of Italian. Bamberg is a German University and many of its students are German. Many of the students do have a different mother tongue. Learning a third language based on the knowledge of a second language is less effective than learning based on the knowledge of the mother tongue.

I am really proud to announce that OmegaWiki has been selected by the University of Bamberg as the platform that will host the lexicological information for "Destinazione Italia". The initial phase of the project will create a lot of Italian based DefinedMeanings. In the second phase we will translate these words to English, German and Spanish. The third phase is to find translations in as many other languages as we can get.

Research done by Zdenek Broz learned, that when the combination of quality translations of German, English and Spanish is found, it will allow the inclusion of translations of other languages when these translations are shared in a different resource. According to Zdenek's figures this will get us an accuracy of around, probably better than 95%.

There is a budget to get us many translations in other languages. The sweet thing is, when we are able to provide quality translations, the budget can be used for other things. This can be to improve the OmegaWiki usability, it can also be to spend money on a language that is not part of the initial list of languages "Destinazione Italia" supports.

The challenge is therefore, how much can we do with a limited budget. What will be the added value of creating content in a Wiki environment. When will OmegaWiki reach the tipping point where collaboration in OmegaWiki is the obvious thing to do, "Destinazione Italia" will help us reach that point. :)

Thanks,
GerardM

Thursday, December 21, 2006

OmegaWiki now supports Cebuano

Cebuano is one of the languages spoken in the Philippines. Some 20 million people speak it as their first language and some 11 million speak it as a second language. It is good that Cebuano is now enabled for editing.

There is a Wikipedia in Cebuano, as far as I can tell it has been localised in MediaWiki. This means that when the names of languages are translated, OmegaWiki will start to look attractive for the people who speak Cebuano.

Out of interest I checked if Open Office supports Cebuano. Googling learns that people use OO in Cebuan, but there is no official localisation for Cebuano and there isn't one for Tagalog either. Open Office does not support many languages and it will be hard work to get the user interface localised in more languages.

This does however not mean that Open Office cannot create content in Cebuano. Of importance is that OO is able to indicate what language people use. It is not clear to me how to do this; it seems that OO only allows for the use of languages that it fully supports. This is in my opinion not the way to approach it.

If Open Office allows for people to select the languages that they edit in, and the languages are everything that ISO-639-3 supports that is written, it should be possible for people to select the user interface that suits them best and even allow for the use of spell checkers that are created for these languages. With the correct tagging that is implicit in using the language tags, it will become easier to support the documents that are produced because it is then possible to explicitly know what language a text is in.

The questions for me are:
  • Is my analysis correct .. please tell me it is not ..
  • How to convince the OO people to support the use of all recognised languages that can be written
  • Get support for spell checking in those languages as well.
Thanks,
GerardM