Sunday, June 03, 2007

250.000 expressions

Today we reached the milestone of 250.000 Expressions at OmegaWiki. It is special because most of this data has been entered by hand. We find that when people get enthused by the concept of OmegaWiki, they do make a difference for the language that they champion.

We have people who have a particular interest in Georgian, Khmer and Spanish, it shows in the statistics as these languages grow much faster than the others.

Aveyron is the 250.000th entry in OmegaWiki and, it is only fitting that Ascánder was the person adding it. Ascander is one of the most valuable contributors to OmegaWiki. Aveyron is part of a project to include information from the ISO-3166-2. In this standard it is detailed in what way countries are subdivided. It does not state that Italy has provinces, the USA has states or that Germany has Bundeslander. It does give the names of these entities.

So, OmegaWiki is evolving nicely. We hope that in line with how Wikis evolve, we will have an easier time to get 250.000 more Expressions.

Thanks,
GerardM

Friday, May 25, 2007

After a week of hacking, testing !!

A lot of work has been done on the OmegaWiki functionality. We have been working on functionality that is of importance to the organisations that we hope to collaborate with.

There were several issues that we have dealt with:
  • Support multiple "data-sets" within a single OmegaWiki installation. These sets can be used to store imported "authoritative databases," such as scientific databases.
  • Users can navigate within a data-set or choose a different one to look at. The default set can be configured globally, for a user group, or for an individual user.
  • Different data-sets can have different permission levels.
  • DefinedMeanings in different data-sets that are identical (describing the same concept) can be mapped to each other.
  • When data is imported, we can choose which data-set to import it into.
There are several parts of the puzzle that are still missing; we are however at a stage where we need to test our data. So we are going to make this functionality go live soon. The first thing is to know that after all the database changes and much refactored functionality everything still works.

The next thing will be to experiment with a first authoritative or additional database. The obvious first resources are the GEMET collection and the ISO-639-6 collection. This is all in preparation of more partners that will be collaborating in the OmegaWiki environment.

More functionality will be implemented in the coming weeks:
  • The possibility to add multiple values without having having to reload the editor each time
  • Allowing for annotations that are dependent on previously set values; this will for the first time provide us with terminological functionality
  • More functionality is in the pipe line, I think you will love it when we have it :)
Thanks,
GerardM

PS It was a fun week, we had a day with a negative number of lines added. We had to change functionality to enable the software to run under Windows. To relax, I have read several chapters of Accelerando. It was fun to watch Kim and Erik work together, my appreciation for both grew. It was gratifying to see my dream become more of a reality :)

Sunday, May 20, 2007

Annotations, hyphenations and IPA

On OmegaWiki we aannotate. In addition to the sample sentences, it is now possible to add hyphenations. A thank you to Sean Burke and Kim Bruning who made this possible.. :)

It is also possible to include the International Phonetic Alphabet or IPA. On the one hand we should feel confident that people will do good. On the other hand, a lot of the IPA notations out there are not useful because they assume that the persons using it have a specific background.

In OmegaWiki we have a public that is truly multi-lingual. This is best experienced when you change the user preferences to another language. Most of the language labels may be shown in the selected language. The consequence of a multi-lingual public is that only IPA notations without language specific shortcuts are useful.

I am sure that you have an opinion about this, we hope to learn your arguments ..

Thanks,
GerardM

Monday, May 14, 2007

Domains and OmegaWiki

Some days Lejocelyn added a feature request about Domains on OmegaWiki. Being one of those points that are also most relevant to me personally of course I answered. Why domains are so relevant? Well: let's say we have 1.000.000 expressions for English-German for a translator, but for us only a certain set of data is relevant when we do translations, so having all 1.000.000 Expressions to search, with all potential results in our glossary window is some kind of an overkill and instead of helping you to find the right term it would take you the triple of the time you need to look things up in a dictionary (let's say about physics or medicine).

Dictionaries are general, yes, but then the amount of specialistic terminology is limited to what is most often used, therefore each of us still has these very special dictionaries about just one topic and these are our most valuable tools besides Internet (well yes, there are terms that are not in our dictionaries, so we have to search for them in available texts about the topic we are translating).

What I would like to say with that: domains might not be relevant to somebody searching for just one word every now and then, but they are most relevant when you want to use a ressource in a professional way.

Thanks for considering to have Domains within OmegaWiki.

Friday, May 11, 2007

OmegaWiki supports many linguistic entities

OmegaWiki aims to include all words in all languages and provide both lexical, terminological and ontological information. As the discussion of what makes a language is an endless one, those languages that are included in the ISO-639 codes are the ones that are supported.

Having chosen the ISO-639-3 to start of with has proven a great start. It did however not provide the granularity needed to categorize words to their linguistic entity. How to deal with languages that are written in several scripts, how to deal with regional differences? This is what this iteration of the ISO-639 standard does not deal with.

The implication is that this standard on its own does not suffice. By combining the data with other standards with other codes it is possible to provide more granularity, but how to deal with dialects like Westfries, that is spoken in the area where I grew up?

As the OmegaWiki project was evolving and getting traction, I got into contact with Debbie Garside. She is heading Geolang an organisation that has been preparing for a long time the next iteration of the ISO-639 standard, the ISO-639-6. The aim is to include at least 25.000 linguistic entities in a hierarchical structure. Adopting this data would allow OmegaWiki to better achieve its aim; include all words of all languages.

When a standard is published there is a prescribed period in which the public is invited to comment on a standard. So far this has been done using e-mail. Experience shows that when the amount of subject is too big, e-mail is not a tool to cope. Geolang had explored the option of using Wiki technology before, this sadly did not lead to the right synergy. In OmegaWiki however, there was both an active interest in language standards, it included not only the Wiki methodology, it even allows for the inclusion of the data in a true hierarchical way.

By publishing the data in a wiki, in essence everybody with an interest in orthographies and dialects is invited to comment, modify and add to the hierarchical data. To make this into a standard, there will be a need to assess the community generated data and assert the validity of the information provided. This is where the World Language Documentation Centre will play its role. As its name implies, it documents languages and it will do so in the broadest sense of the word. Obviously an organisation like this will only function well when it is as an organisation an inclusive organisation. The make-up of the current board reflects many specialities that make up linguistics and the language industry.

It is with a fair amount of satisfaction that I can announce that Sean Burke, one of the volunteers of OmegaWiki has imported the first batch of the ISO-DIS-639-6 data in time for the inaugural meeting of the World Language Documentation Centre. Both OmegaWiki and the WLDC will rely on collaboration, to get the necessary work done. Our challenge will be to provide the infrastructure and the minimal organisation to start and sustain our projects.

With the inaugural meeting, the WLDC it is proclaimed to the world that as an organisation the WLDC is ready for business. With the first data available in OmegaWiki, the first request to the world to collaborate on the languages that are spoken the orthographies that are written goes out. it is the start of acquiring the meta data that helps us understand the data that is already out there and consequently make from all this data information because we will become better able to parse the data.

Thanks,
GerardM

Sunday, May 06, 2007

New functionality

OmegaWiki has collections. These collections serve to indicate that certain DefinedMeanings are related. Collections can serve a purpose; the GEMET collection for instance is a resource that was the data that started our project. The OLPC collection is a list of the first words that we want in all language to start a multilingual dictionary for the OLPC project.

In these statistics, we have a tool to tell people what projects we have within OmegaWiki. This allows people to work on things that are of interest to them. The really sweet thing is that it shows like a work in progress, it shows what needs doing and, what has already been done.

There are several projects that are dear to me and can use more attention:
When you do not see your language in a collection, just add one word to any of the DefinedMeanings that are part of the collection and the next time it will be there. When your language is not supported in OmegaWiki, let me know and I will see how to remedy this.

Thanks,
GerardM

Friday, April 27, 2007

A similar milestone of a different kind

The last post of the OmegaWiki blog was about statistics. Kipcool indicated that there was an issue with the numbers, in the query the deleted Expressions were not considered.

Today we have the opportunity to celebrate anew one of the milestones that went before. Today we have a cool 10.000 expressions in Italian. :)

I want to thank Kipcool, Kim Bruning and Zdenek Broz.
GerardM

Monday, April 16, 2007

Milestone

The nice thing of statistics is that when there is a milestone, it can be celebrated. It is with pleasure that I can announce that OmegaWiki now has one language with 15.000 words. It is German that has currently the most expressions. There are currently 213.550 Expressions in 133 languages.

For OmegaWiki it demonstrates that we have a nice autonomous growth. Of the languages that we started after the import of the GEMET data, Japanese currently has the highest number with almost 4.000 words. Esperanto is with some 1.500 the biggest artificial language.

You can help us with translations for the language names that we have translations in. This will improve the usability of our data when the user interface is selected in the language of people's mother tongue.

OmegaWiki is still in need of more functionality. This is something that we try to achieve in any way possible. It is however rewarding to notice what difference the existing functionality makes.

I am very happy with what we have achieved so far. It bodes well for the future :)

Thanks,
GerardM

Monday, April 09, 2007

Fall back languages

OmegaWiki has a relevant bug fix; it is now obvious what option you add a part of speech you will see "part of speech" in your own language and when we do not have it, you will get it English as the fall back language. This is a huge improvement from having the same text in some 24 other languages.

The next improvement will become possible when the Multilingual MediaWiki is finished; this will allow you to select the languages you are interested in. These languages are more personal than just falling back to English.

Another option would be to identify a fall back language for a language itself. This makes sense for languages that exist in a space where another language is well known and better supported in MediaWiki. Languages like French, German, Italian, Spanish, Portuguese and Mandarin come to mind.. I am sure there are more..

For the Incubator we want to define fall back languages to make it easier to localise the MediaWiki messages.. We hope to get answers what the best choice is.. otherwise we will have to guess..

Thanks,
GerardM

Friday, March 30, 2007

Protecting people from poisonous personalities

The BBC website has an article about the culture of abuse on line that allows people to be noxious, abusive, threatening to people and in this instance to women. The thing that sparked this off are death threats to a prominent blogger, Ms Kathy Sierra. I read a follow up article on oReillynet called "Open season on women".

We have been extremely lucky and fortunate on OmegaWiki so far with abuse and vandalism. We have also been extremely lucky with our community. I do not exaggerate when I say that we have some great females being extremely relevant to what we do. If both articles need an official reaction, it would be that we will not tolerate such things on OmegaWiki and, that I am still appalled by the way the Wikichix were driven away from the Wikimedia Foundation for all the "right" reasons.

At this stage in the life cycle of OmegaWiki, I can say that people who are deemed to be poisonous in what they do will find that there is little toleration for them. When we get to the stage where it becomes an issue if we should tolerate poisonous people, I can tell you know that I am prepared to fight tooth and nail to get such people out.

Thanks,
GerardM

Saturday, March 24, 2007

More new functionality & some stats

Lies, damned lies and statistics.. I always thought that OmegaWiki had more English words than German words. This turns out not to be the case, then I thought it may be because we make something English (United Kingdom) or English (United States) when the word is not shared between the two.. There are only 326 UK words; that does not fill the 441 gap.

These and more pondering are possible because of our new statistics functionality we have courtesy of Zdenek Broz. It shows you the actual numbers of Expressions in OmegaWiki.

Some more factoids, we now support 130 languages, 28 of these have less than 10 words.. When you compare us to Wiktionary, we would be the fourth in size when an article is considered the equivalent of an Expression or we would we the twenty-fifth in size when a DefinedMeaning is considered in this way. When you consider that all of our growth so far has been autonomously, I think we are doing well.

In the mean time, Alexa considers us the 480,601 website without statistics for the week.. I do not know what goes on there. I do like their T-shirts though.

To quote the Alexa T-shirts: "Will dance for better Alexa rank " :)

Thanks,
GerardM

Thursday, March 22, 2007

New functionality

OmegaWiki has new functionality, there are two bits of functionality. One bit is not obvious but really helpful; a function to (re)build the indexes of the database. What it does is ensure that the indexes are build in the right order. This improves performance considerably.

The other bit, is much more spectacular. It sorts the tables in the HTML based on the content of the first column. The order depends very much on the language selected in the user preferences. This leads me to my question; is it sorted well when your language is Persian, Arab or Hebrew. These are right to left languages and I have no clue if the sorting is done well.

Thanks,
GerardM

Monday, March 19, 2007

How do the four freedoms apply to one database?

The FSF defines four freedoms when it comes to software. What kind of freedoms applies to a database like OmegaWiki.

The OmegaWiki software is licensed, like MediaWiki what it is an integral part of, with a GPL license. This means that you can use the software as is. As OmegaWiki uses specific data structures that can be licensed separately for completeness sake, the database design is also available under a GPL license.

The data that is contained in OmegaWiki is licensed under a combined GFDL/CC-by license. Many people insist that these licenses are not compatible. At issue is that the data are just facts, it is only possible to copyright facts as a collection. We want people to make use of our collection. For us success is: "when people find a use for our data we did not think off".

We invite people to collaborate on our data, when they enter some Babel templates on their user page, we give them edit rights to the data. We invite organisations to collaborate on our data because there is so much data that organisations can share, there is so much labour invested in the type of data OmegaWiki can be a home for.

So the data and the software is Free. How about OmegaWiki itself ..

As there is only one OmegaWiki.org, the room to do whatever is limited. The data has to be useful to everyone and it has to fit in with the notion of the DefinedMeaning. When domain specific data is added, it needs to be domain specific and, there has to be agreement that this data provides a suitable extension for people involved in this domain.

When people find that this is not enough, they can have their own database. This does not mean that they cannot cooperate. Much of what they need in terms of extra functionality will be shared. It means that even when the OmegaWiki database is forked, there is still plenty of scope to improve all the things we do agree on and collaborate on those.

There is plenty of Freedom. All the Freedoms we provide. I think however we achieve the most success when we find that there is more that binds us than that drives us apart. It means that we have to work hard in understanding what our common needs are.

Thanks,
GerardM

Saturday, March 10, 2007

Managing data with some SQL

The great thing about OmegaWiki is that the data is in a database. You might say that this is not that special, every wiki uses a database. Today, we have as a first time done some curation on the data; everywhere where a word in en-US was written exactly the same as in English, we have deleted the English. One example is the word "competition", in the history you will find the deletions.

I am really grateful that Leftmost has started to use SQL to fix things for us. It saves us what is most valuable; the time of our editors.

There are other things that we can do, I have asked to have all Bulgarian words that are capitalised changed to lower case where the Russian words are lower case. This is to fix something that is done consistently this way in the GEMET database. With these improvements, the GEMET data becomes usable for other purposes; things like data mining .. :)

Thanks,
GerardM

Sunday, March 04, 2007

About reputation and education

Wikipedia has a new scandal. There are several issues here and, several interested parties. The scandal is about someone who is known by a nickname and who claimed that he had academic qualifications. He claimed to be a doctor of theology.

One of the interested parties who obviously relished this occasion was Dr Sanger of Citizendium fame. For Dr Sanger it is one of those occasions where he can sing the praises of his project. He did, I did not read anything new there. In the mean time, his project was announced in Nature and now, so many months later, there is still nothing to be seen. That does wonders for the credibility of that project ...

Given that scientific credentials will be very relevant on OmegaWiki, I have given it some thought. When people want to claim professional credentials, they would have to provide us at least with their real name and their e-mail address. These would be the requirements on the level of OmegaWiki.

For the Wikis for professionals, different rules might apply. When medical credentials are claimed, wrong information can kill. It is for such reasons that much more identifiable information is likely to be required. This does not mean that credentials are as important as Dr Sanger says. Relevant is the quality of the information provided. Relevant is the role a person plays in the community. This means that a person can become relevant in OmegaWiki by building a reputation. For this you do not need scientific credentials. Science has, like every part of society, its fair share of miserable people and I am sure Citizendium will learn that as well.

If this incident is to be a lesson, the lesson is that it does not pay to assume credentials and have a virtual reality meet real life. It is therefore sad that Essjay is now "retired". If this means that the person who did a lot of good work will do no more, than it is a sad outcome. If this is the only time that the issue of assumed credentials raises its ugly head, it has a silver lining. If this incident is sufficiently public that it will stop this phenomena, then I will be glad.

Thanks,
GerardM

Wednesday, February 28, 2007

What is a Word?

First, let me make it perfectly clear that this is a discussion that has raged for centuries. I know full well that everybody has his or her own opinion on this matter, and that I am not going to resolve this issue today. This is an overview and a bit of personal opinion, as relates to online dictionaries.

Intuitively, we all know the answer. A word is a unit of language conveying some meaning. But how do we decide what is a real word? We look in a dictionary, of course. What do we do if we're writing a dictionary?

We are caught between cataloging what is "right" (prescriptivism) and what is actually done (descriptivism). The pendulum has lately swung towards descriptivism, and I would say that there are some good reasons for that trend. The language that is spoken on the streets is not the same language that is written in academia. Somebody learning a language may genuinely need help sorting out the less proper terms in it.

Take, for instance, colloquialisms such as "irrespective" and "humongous", and all the phrases that have gotten squished together into amalgams like "gotcha" and "woulda". Most people would readily agree that these words do not belong in a college thesis paper.

What is the scrupulous lexicographer to do? Fortunately, it is not a strict either-or question, especially in a work not substantially limited by size. In an electronic resource, we can put them in, anyway. To satisfy the formal sorts, the perscriptivists, we can then place a prominent usage note in the entry, explaining just why a writer might wish to use caution with the term: ginormous is a colloquial term, regarded by many to be something less than a proper word. Thus, the reader is both informed and cautioned.

That's fine for most of the slang and jargon, but we have another problem. People keep making up new words. My sister in law coined the term "muskaroon" to mean generically any small, furry creature that scurries past too quickly to identify. Squirrels, chipmunks, gophers, and presumably rabbits would all qualify. So we have a unit of language with a symbol and a meaning. The trouble is, if you walked up to people on the street and inquired whether there were muskaroons in the area, nobody would be able to answer who hadn't talked lately to my sister in law, and that is a small minority of people, indeed.

The test here is usage. Can we demonstrate that the word is in common use? Now, depending on the character of the dictionary, we can define the rules various ways. Was it used by so many independent sources? Did anybody important (such as Shakespeare or a prominent academic journal) publish the word?

Generally, we also try to find and present examples of the term in what is called "running text". That means that it is in a paragraph, and isn't only used as somebody's nickname, say. The edge is still a fuzzy one. Are the citations in traditional print sources, such and books and journals, or are they sprinkled in a couple blogs and forums? Was the word used in only one limited context, or in a variety of sources and over a period of years? These sorts of tests can help to weed out many of the more questionable entries. At some point, though, it may yet come down to a judgment call, if not on whether a word is real, then on how to apply the rules. In these cases, I advise the users of a dictionary to bring a healthy dose of skepticism with them, to recall that even dictionaries are not infallible, and to trust at the very least that these decisions are made by real people who care for the project.

If, knowing all that, you find you don't like the way "they" are running the place, you are invited to do a better job.

Monday, February 19, 2007

Anyone may edit

Written in response to this project.

Somebody asked me today about what happens to a dictionary when anybody can edit it. As anybody who has ever edited a wiki knows, the openness is a mixed blessing.

It is a great thing, because many hands make light work. Dictionaries need to be every bit as large as the languages they catalog, so the process of gathering and maintaining the data is a huge one. As we start to add translations between languages, rather than simply defining a term, that task becomes orders of magnitude bigger. To capture all words in all languages is something that will take nothing less than a wiki and a worldwide community. It is a monumental task, but in a wiki, we can conceive of creating a resource on such a scale.

It is a great thing because so many regions and cultures can be represented. An American may understand most of the English spoken in South Africa or New Zealand, but both of those regions have slang all their own. Chile speaks Spanish far differently than Spain. All those variants can have their place.

It is a great thing because a wiki can evolve with a language. New terms come into use all the time, and a freely editable electronic resource is not limited in its capacity to store data or to accommodate a large, diverse set of editors.

The big trouble is this: if just anybody can edit, how on earth do we know it is right? I'd like to explore a few approaches here. Of course, any of these approaches could be considered as a barrier to entry, but these things are always trade-offs.
  1. Appoint trusted users to do the housekeeping. These are the sysops, administrators, bureaucrats, the librarians, or the janitors, depending on your point of view. These somebodies keep watch and undo the damage that some of the just-anybodies can do. If somebody writes an article containing typical vandalism, such as "asdfasdf" or "Dave is a dork!", an administrator can delete or undo it. Much vandalism is so predictable that even a bot can detect and remove it. Unfortunately, a select group of administrators, however well trusted or well-read, cannot be everywhere at once, and they cannot know everything. Things get missed, even with a checklist system such as patrolled edits. It is likewise impossible for an administrator or small group of administrators to know everything. Misinformation, intentional or otherwise, is not so easy to spot as out-and-out nonsense.
  2. Hold people accountable. Articles have histories, so you can see who did what. Even pseudonymous users develop reputations. Anonymous users tend to attract the most scrutiny. An active, healthy wiki often develops into a meritocracy, with leaders having sway (though not necessarily authority) based on reputation, seniority, and trust in the community. This effect generally works to improve content, but even a well-known, trusted user may make mistakes. If he or she is trusted well enough, there is a risk that an error or oversight may go unnoticed.
  3. Allow anybody and everybody to scrutinize and correct or flag the content. The process is not foolproof, especially in larger projects, but wikis have a remarkable capacity for self-cleaning. Of course, this approach can tend to result in a sort of groupthink effect: if enough people believe it, then it must be so.
  4. Demand credentials. Don't just let in any old riffraff. Wikipedia has clearly shown the power of amateurs and volunteers to create great content, but it is certainly possible to limit the users in a project, or part of the project, to a certain group. This approach is most appropriate to a wiki serving a closed community, such as a professional or academic group, especially one dedicated to a particularly narrow or specialized topic.
  5. Make the messes behind the scenes, and publish only the good stuff, with some review process. The German Wikipedia published a paper book containing selected articles. Online, there have been proposals for a "Stable Versions" system, where a mature article would be reviewed and locked, and any additional changes would go through a separate editing or discussion page.
  6. Demand references. There is a movement within Wikipedia to reference the articles and the claims made in them. In the context of a dictionary, references may be other dictionaries. Is the word recognized by RAE or OED (whom we trust to have done the requisite homework)? They may be other works about words. Or, they may be citations. Citations are quotations including the word in question. They show context and provide evidence that the word is or was in use. Of course, we must still question the validity of the evidence. Are the 400 Google hits because somebody prolific uses that nonsense word as a handle? Is a word more valid if it was used by a blogger or two, or by Thornton Wilder? Is an etymology known with reasonable certainty or is it apocryphal? Depending on the size and resources of the wiki, efforts to verify and reference articles may be systematic, or they may be requested when a given entry or fact is questioned.
A wiki is simply a website where anybody can post. With a bit of care and attention, its content can be as valid and accurate as any other reference, and certainly more complete and up-to-date.

Friday, February 16, 2007

An article in Nature ...

I am absolutely thrilled with the article that was published last Wednesday in Nature... It is a great article and it explains really well what we hope to achieve with relational data in MediaWiki. The only thing that is a bit sad is, that you have to pay $30 for the privilege of reading it.

The article is great, and what makes it special is the great presentation that Knewco has created to explain what we hope to achieve; this demo available at wikiprofessional.info. It presents some really impressive figures; it indicates the work done to integrate several important resources of the bio-medical domain, the numbers involved.

For me the most important point is that this is likely to be a very important stimulus to the Open Access movement. It indicates that it is possible to bring what was divided together. It allows people to work with the terminology of their field and also add data that is very specific. Information that goes much further than what was envisioned in what was once called the "Ultimate Wiktionary".

The whole notion of a resource that because of its roots already merged lexicology, terminology and ontology is really special. With the integration of such specialised data from different domains like the bio-medical, another really interested experiment will be under way when the data gets imported and merged. There is a nascent community for the bio-medical domain and, it will find that it will co-exist with the existing OmegaWiki community.

Both communities have everything to gain from collaboration; much of what the existing OmegaWiki community cares about will be seen as a fringe benefit. On the other hand, the translations that exist for concepts like malaria will prove to be of value when scientific articles are considered that were not published in English.

I am convinced that a bright future is ahead of us. We have this vision of what may come, I wish I could look into the future and see what it will be like. :)

Thanks,
GerardM

Sunday, February 11, 2007

Why compete when you can collaborate ?

All words of all languages of the world.. that is what we eventually aim to include in OmegaWiki. This aim is of such a magnitude that you have to be certifiable to come up with such a project. The functional design for the project includes much more; everything including the kitchen sink..

When everything is to be included in one project, it is easy to suggest that people contribute to the project. When the project includes everything why have another?

In an Open Source / Open Content environment this is not necessarily how it works. Why should the others be seen as competitors? They do their own thing, sure. You may want to achieve the same thing, also true. It is however much possible to find the synergy between projects. This way you can build on each others accomplishments.

The Shtooka project is something I learned about the other day. The one thing it does really well is the way they make recording pronunciations easy. You can record a string of words and it will save them for you one at a time.

Wiktionarians saw this and they are working an upload facility so that it will also be saved automatically to Commons. I warned that the files should not only be saved as .ogg files. In order to make sure they are relevant for scientists there should also be a .wav file. The current thinking is that the flac file format will work as well and the benefit is that it provides a loss-less compression. To make sure that this is the case, the praat software, software that is also available under a GPL license, was analysed and it was considered that it is easy to incorporate this flac file format.

People from effectively five different communities are now working together. It will be even possible to include links to OmegaWiki in the Shtooka meta data. This will be possible even though both projects do their own thing. Both the data and the functionality can be shared.

I may be certifiable, but this kind of collaboration is awesome and, it is why there may be method to this madness.. :)

Thanks,
GerardM

Thursday, February 08, 2007

Become an OmegaWiki developer

OmegaWiki is now running the latest version of the MediaWiki software used by Wikipedia. This is a major milestone, as it also makes it a lot easier for anyone to join in the fun of developing the open source OmegaWiki/Wikidata software. To give credit where credit is due, these are the people who have contributed to the code so far:
  • Peter-Jan Roes
  • Karsten Uil
  • Sean Burke
  • Rod A. Smith (sticky tree expansion via cookies)
  • Ævar Arnfjörð Bjarmason (namespace code installer)
  • Charles Pritchard (Multilingual MediaWiki development, ongoing)
  • Jelte Zeilstra (untranslated meaning script, under review)
  • Zdenek Broz (statistical scripts, under review)
  • Paa-Kwesi Imbeah (Wikimedia Commons support, under review)
  • Marc Carmen (TBX export, incomplete)
  • myself
There are probably others I forgot. These people, some of them volunteers, some paid developers, are helping to build the first truly multilingual, massively collaborative ontology. If you want to become a part of this history, there are now instructions that should help you get on your way. Please contact me under erik AT openprogress DOT org once you have read and followed these instructions. There are always plenty of things that need doing. And as the organization which runs OmegaWiki, Stichting Open Progress, develops more and more partnerships around the project, we will look to our team of existing developers to help us implement them.