Showing posts with label ISO 639. Show all posts
Showing posts with label ISO 639. Show all posts

Wednesday, February 18, 2009

A new user ,,, many new languages

A new user came to OmegaWiki and he wanted to add content in the Mayan languages. All of them "except that funny one from veracruz which only recently got classified as Mayan". All of them because this new user comes with an existing dictionary and this has "all of them".


Now the Mayan languages are not easily classified; The way Ethnologue had it is not how the ISO-639-3 has it at this time. This means that we will have to carefully understand what the relation is between the languages in the dictionaries and the codes maintained by SIL.
Thanks,
       GerardM

Thursday, December 13, 2007

Eastern Yiddish

Eastern Yiddish, is one of the two varieties of Yiddish that have been recognised as languages in their own right in the ISO-639-3 (ydd). in OmegaWiki, we now have our first 213 Expressions in this language and I am impressed with the amount of work that has gone into it; most have annotations indicating hyphenation and the pronunciation using IPA notation.

Eastern Yiddish is not one of the languages supported by MediaWiki, and the mechanism for showing localised content is connected to the language selected in the "User Preferences". I have been given some help from Siebrand what files need to be changed and added. Kim helped me with doing it for the first time and now the first localisation is visible for Eastern Yiddish.

The MediaWiki localisation itself uses Yiddish as the fall back language so the experience is pretty good for now. What Siebrand indicated is that is is possible to include languages like Eastern Yiddish in the BetaWiki. This would create stubs that are of benefit to OmegaWiki. It would prepare for the moment when people start localising in earnest.

I think it would be a good thing, but I am interested to learn what other people think.

Thanks,
GerardM

Monday, December 03, 2007

Supporting American English

American- or British English are two variations of the English language. They have substantial differences. They are sufficiently the same and are unlikely be mistaken to be separate languages.

In OmegaWiki, it has been possible to add entries for English; this meant there is no difference between written the different versions of English or you had to specify both versions. Issues like this exist for other languages like Serbian and Mandarin as well.

In the OmegaWiki user interface, languages are considered ISO-639 entities. When a DefinedMeaning for a language is part of the appropriate collection, we use the translations in our user interface. The problem is that all these linguistic entities are needed now and that they are created to make OmegaWiki work.

For the ISO-639-6 there will be issues as the codes we make, using the RFC 4646 methodology, will be replaced. It will also be interesting to learn how in the end everything will be merged together.

In the mean time we now support localisation for these linguistic entities.

Thanks,
GerardM

Monday, August 27, 2007

Some more on Wolof

OmegaWiki wants to support all words of all languages and, it does not want to go into the issue of does this language exist or not. We make use of the ISO 639 standards and, when we feel like being adventurous, we look at what is recognised in the IANA language tags.

Deferring to standard organisations means that you take what they say as the "truth". It does not mean that we necessarily agree, but it saves us from a lot of mayhem. Yesterday I wrote about the first native Wolof speaker for OmegaWiki. Today Ibou changed the definition for Wolof and included Gambia as a country where Wolof is spoken. According to the description by Ethnologue of the Wolof language this is not the case. They do refer to another language, Gambian Wolof, this description makes it clear that Wolof is spoken in the Gambia as well.

The article on Wikipedia on Wolof is in my opinion wrong; it gives the impression that the ISO-639-1 and the ISO-639-2 codes are split into two. This is contrary to how standards work. When a language is split into two, the original meaning will stand as it is, it will get a new description to indicate that it has been split and two new codes will be created.

So Ethnologue is inconsistent. Ibou is probably right. I have send an e-mail to Ethnologue and I hope that they will amend their fine resource so that we will know for sure that he is right. :)

Thanks,
GerardM

Tuesday, July 10, 2007

Ch'orti', a language spoken in Guatemala and Honduras

Ch'orti' as a language has caa as its ISO 639-3 code, some 30.000 people speak the language and according to Reeck many more belong to the associated ethnic population.

I have been adding languages to the ISO 639-3 collection for some time now, I started with Ghotuo (aaa) and I have now progressed to Ch'orti' (caa). Many have few speakers, many are extinct, several are sign languages and almost all of them I have already forgotten.

So why do this, is there method to this madness.. OmegaWiki aims to include all words of all languages, but what languages are there ? Do we want to discuss the notion of yet another linguistic entity that we should support. Does something like Brithenig (bzt) deserve its place under the sun ?

I do not mind the discussion, but I do mind what the result will be of such a discussion. It needs to come to a conclusion and I do not want to be in the position that people look to me for a verdict. It is not a good idea either to have the OmegaWiki commission be in that position. It is for all these reasons that we decided on adopting standards and started with the creation of portals for the ISO 639-3 languages. We are now at the next phase, creating the DefinedMeanings for these languages and make them part of the ISO 639-3 collection.

This is only what is recognised by one standard, there are other standards that help indicate what the precise linguistic entity is that is to be documented in OmegaWiki. First we should finish this, there are currently 1365 entries in the ISO 639-3 collection .. there are many more thousands to go :)

Thanks,
GerardM

Friday, May 11, 2007

OmegaWiki supports many linguistic entities

OmegaWiki aims to include all words in all languages and provide both lexical, terminological and ontological information. As the discussion of what makes a language is an endless one, those languages that are included in the ISO-639 codes are the ones that are supported.

Having chosen the ISO-639-3 to start of with has proven a great start. It did however not provide the granularity needed to categorize words to their linguistic entity. How to deal with languages that are written in several scripts, how to deal with regional differences? This is what this iteration of the ISO-639 standard does not deal with.

The implication is that this standard on its own does not suffice. By combining the data with other standards with other codes it is possible to provide more granularity, but how to deal with dialects like Westfries, that is spoken in the area where I grew up?

As the OmegaWiki project was evolving and getting traction, I got into contact with Debbie Garside. She is heading Geolang an organisation that has been preparing for a long time the next iteration of the ISO-639 standard, the ISO-639-6. The aim is to include at least 25.000 linguistic entities in a hierarchical structure. Adopting this data would allow OmegaWiki to better achieve its aim; include all words of all languages.

When a standard is published there is a prescribed period in which the public is invited to comment on a standard. So far this has been done using e-mail. Experience shows that when the amount of subject is too big, e-mail is not a tool to cope. Geolang had explored the option of using Wiki technology before, this sadly did not lead to the right synergy. In OmegaWiki however, there was both an active interest in language standards, it included not only the Wiki methodology, it even allows for the inclusion of the data in a true hierarchical way.

By publishing the data in a wiki, in essence everybody with an interest in orthographies and dialects is invited to comment, modify and add to the hierarchical data. To make this into a standard, there will be a need to assess the community generated data and assert the validity of the information provided. This is where the World Language Documentation Centre will play its role. As its name implies, it documents languages and it will do so in the broadest sense of the word. Obviously an organisation like this will only function well when it is as an organisation an inclusive organisation. The make-up of the current board reflects many specialities that make up linguistics and the language industry.

It is with a fair amount of satisfaction that I can announce that Sean Burke, one of the volunteers of OmegaWiki has imported the first batch of the ISO-DIS-639-6 data in time for the inaugural meeting of the World Language Documentation Centre. Both OmegaWiki and the WLDC will rely on collaboration, to get the necessary work done. Our challenge will be to provide the infrastructure and the minimal organisation to start and sustain our projects.

With the inaugural meeting, the WLDC it is proclaimed to the world that as an organisation the WLDC is ready for business. With the first data available in OmegaWiki, the first request to the world to collaborate on the languages that are spoken the orthographies that are written goes out. it is the start of acquiring the meta data that helps us understand the data that is already out there and consequently make from all this data information because we will become better able to parse the data.

Thanks,
GerardM

Sunday, January 28, 2007

Latin roots etc.

Well yesterday one thing came into mind - a dictionary a teacher of mine at the language school had. It was a dictionary that listed Latin words with many translations into other languages and one thing is obvious: all these words of course were similar in all languages. If you knew one of them and studied the other language it would have been easy to create the relative words following a set of rules for most of them.

So one thing should be obvious: to insert these words with their translations into OmegaWiki ... but well, there is one problem with Latin - the "normal" Latin language should not be mixed with the taxonomical Latin that is used in science ... so we need to create two languages: Latin and taxonomical Latin ... who knows if the relative language codes exist somewhere in the ISO 639 standards.

Technorati: