#Multilingual #MediaWiki is really helped by the Polyglot extension. What it does is show for instance the main page in your language if there is a translation. This helps the usability of OmegaWiki and any other installation a lot when multilinguality is important.
There are a few cotchas, the first is that it prefers the ISO 630-1 code when one is available. This means en in stead of eng, fr / fra, nl / nld .. This is the same for the Babel extension. However, this extension that the change to the short version for you.
Given that the ISO 639-6 will restore the primacy of existing two character codes, it makes sense for OmegaWiki to use this algorithm.
When Kipcool had a look at the code, he found that for OmegaWiki only a few lines were really needed... If you want to see the default or English text, you can always add "&redirect=no" to the URL.
Showing posts with label standards. Show all posts
Showing posts with label standards. Show all posts
Wednesday, May 05, 2010
Saturday, April 05, 2008
Unicode 5.1
Today I learned that Unicode 5.1 has been released. The information that I received informs me that one major feature will be of particular relevance to Japanese, Chinese and Korean texts by enabling ideographic variation sequences. The linebreaking for Polish and Portuguese hyphenation has been improved. The Indic languages will be happy with improved text segmentation algorithm.
There are 1624 new encoded characters, this includes characters required for Malayam and Myanmar but there are also new characters for the Latin script. New is support for the Cham, Lepcha, Ol Chiki, Rejang, Saurashtra, Sundanese, and Vai scripts.
For the techies, the collation algorithms have been updated to include all the new characters. This has also an effect on contractions like the ch in the Slovak language.
Many of these things have an effect on languages supported in Wikimedia projects. My question is when will we have support for this. Is this a function of the MediaWiki / PHP code and is it also a function of the browser ??
Thanks,
GerardM
There are 1624 new encoded characters, this includes characters required for Malayam and Myanmar but there are also new characters for the Latin script. New is support for the Cham, Lepcha, Ol Chiki, Rejang, Saurashtra, Sundanese, and Vai scripts.
For the techies, the collation algorithms have been updated to include all the new characters. This has also an effect on contractions like the ch in the Slovak language.
Many of these things have an effect on languages supported in Wikimedia projects. My question is when will we have support for this. Is this a function of the MediaWiki / PHP code and is it also a function of the browser ??
Thanks,
GerardM
Sunday, January 28, 2007
Latin roots etc.
Well yesterday one thing came into mind - a dictionary a teacher of mine at the language school had. It was a dictionary that listed Latin words with many translations into other languages and one thing is obvious: all these words of course were similar in all languages. If you knew one of them and studied the other language it would have been easy to create the relative words following a set of rules for most of them.
So one thing should be obvious: to insert these words with their translations into OmegaWiki ... but well, there is one problem with Latin - the "normal" Latin language should not be mixed with the taxonomical Latin that is used in science ... so we need to create two languages: Latin and taxonomical Latin ... who knows if the relative language codes exist somewhere in the ISO 639 standards.
Technorati: language, Latin, ISO 639, standards, translation, dictionary
So one thing should be obvious: to insert these words with their translations into OmegaWiki ... but well, there is one problem with Latin - the "normal" Latin language should not be mixed with the taxonomical Latin that is used in science ... so we need to create two languages: Latin and taxonomical Latin ... who knows if the relative language codes exist somewhere in the ISO 639 standards.
Technorati: language, Latin, ISO 639, standards, translation, dictionary
Labels:
dictionary,
ISO 639,
language,
Latin,
standards,
translation
Tuesday, December 19, 2006
The use of Standards to emancipate languages
More and more languages are supported by the localisation efforts of projects like Open Office. This has a major effect on the emancipation of these languages; much content is written and some of it ends up on the Internet. When it gets there, its information is often lost because it cannot be found by a user who is not sophisticated in the use of search engines. Sophistication is needed because Google currently only recognises 100 languages and only 15% of the content of the World Wide Web is tagged to indicate a language and much of it is tagged incorrectly. When the quality of the tagging is improved, it will be possible for search engines to provide information that is only in the requested language. This will have the added benefit that a growing corpus of content will become available and this in turn will stimulate the research in these languages.
OmegaWiki is a wiki based website that aims to provide information both of a lexical, terminological and ontological nature. It does this by extending the MediaWiki software with relational functionality. OmegaWiki aims to have all words in all languages.
In order to learn what languages exist, OmegaWiki adopted the ISO-639-3 standard. This leaves out many linguistic entities like dialects and orthographies. OmegaWiki has had the good fortune that it got into contact with the WLDC and GeoLang who are the organisations that deal with the ISO-639-6 standard that is under development. Together with the WLDC it has been proposed to the ISO task group to use the environment and the functionality that OmegaWiki provides to gather the data that is needed to learn about the different linguistic entities. This has been accepted by the task force at the LSGB conference in Vienna.
Much of the groundwork has now been done, the next step is to make this functional. To make it functional, we want to have software adopt the existing standards and the information provided in OmegaWiki. This is feasible for Open/Free software. We have already approached the OmegaT lead developer and, he will be happy to support this because it will make OmegaT more relevant. Because of what OmegaWiki aims to do, we will be able to build spell-checkers for linguistic entities on a regular basis, this in turn will have an impact on the standardisation and the emancipation of languages.
With the emancipation of languages, it will become increasingly an option to bring information to people in their native languages. Studies have shown that this leads to a much better understanding and appreciation of services provided. Particularly in what is called, the long tail of the language industry, there has been little support for any of the standards. Consequently the quality of service provided when translating for languages in the long tail are inconsistent. By developing the tools to include the support for any linguistic entities, it will become possible to use these languages for written communications, it will also become possible to raise the quality of these communications by applying the work that has been done in the language industry.
To make this project take off, it works best when many people and organisations collaborate. All of them will have their own reasons to want to be included. The challenge will be to coordinate things in such a way that all the necessary parts are realised.
There are sufficient reasons for many organisations to buy into this project. What is needed is not only to leverage this but also to explore how the data that is gathered, the functionality that is build can be extended to provide additional information that is of relevance to the project. Many organisations will have a need for information, we will want either their active participation and / or their support. We will find all types of technical complications, we need the buy in of the people that can resolve these issues.
All in all, it will take a lot of effort to provide the difference that we aim for.
Thanks,
GerardM
OmegaWiki is a wiki based website that aims to provide information both of a lexical, terminological and ontological nature. It does this by extending the MediaWiki software with relational functionality. OmegaWiki aims to have all words in all languages.
In order to learn what languages exist, OmegaWiki adopted the ISO-639-3 standard. This leaves out many linguistic entities like dialects and orthographies. OmegaWiki has had the good fortune that it got into contact with the WLDC and GeoLang who are the organisations that deal with the ISO-639-6 standard that is under development. Together with the WLDC it has been proposed to the ISO task group to use the environment and the functionality that OmegaWiki provides to gather the data that is needed to learn about the different linguistic entities. This has been accepted by the task force at the LSGB conference in Vienna.
Much of the groundwork has now been done, the next step is to make this functional. To make it functional, we want to have software adopt the existing standards and the information provided in OmegaWiki. This is feasible for Open/Free software. We have already approached the OmegaT lead developer and, he will be happy to support this because it will make OmegaT more relevant. Because of what OmegaWiki aims to do, we will be able to build spell-checkers for linguistic entities on a regular basis, this in turn will have an impact on the standardisation and the emancipation of languages.
With the emancipation of languages, it will become increasingly an option to bring information to people in their native languages. Studies have shown that this leads to a much better understanding and appreciation of services provided. Particularly in what is called, the long tail of the language industry, there has been little support for any of the standards. Consequently the quality of service provided when translating for languages in the long tail are inconsistent. By developing the tools to include the support for any linguistic entities, it will become possible to use these languages for written communications, it will also become possible to raise the quality of these communications by applying the work that has been done in the language industry.
To make this project take off, it works best when many people and organisations collaborate. All of them will have their own reasons to want to be included. The challenge will be to coordinate things in such a way that all the necessary parts are realised.
There are sufficient reasons for many organisations to buy into this project. What is needed is not only to leverage this but also to explore how the data that is gathered, the functionality that is build can be extended to provide additional information that is of relevance to the project. Many organisations will have a need for information, we will want either their active participation and / or their support. We will find all types of technical complications, we need the buy in of the people that can resolve these issues.
All in all, it will take a lot of effort to provide the difference that we aim for.
Thanks,
GerardM
Subscribe to:
Posts (Atom)
