Dhvani 0.94 Released

A new version of Dhvani -The Indian Language Text to Speech System is available now. The new version comes with the following improvements/features Support for 11 languages- Hindi, Panjabi, Gujarati, Marati, Bengali, Oriya, Telugu, Kannada, Tamil , Malayalam and Pashto(Afganistan) Pitch and Tempo modification for speech Direct ogg-vorbis speech output and optional wav output format C/C++ APIs for applications to use dhvani as a shared library. Generic driver for Speech-dispatcher and Integration to Orca through speech dispatcher Python binding through speech dispatcher Improved language detection algorithm Dhvani documentation is available here. ...

November 16, 2008 · 2 min · Santhosh Thottingal

Language Detection and Spellcheckers

A few weeks back there was a discussion on #indlinux IRC channel about automatic language detection. The idea is, spellcheckers or any language tools should not ask the users to select a language. Instead, they should detect the language automatically. The idea is not new. There is a KDE bug hereand Ubuntu has this as an brainstorm idea. It seems M$ word already have this. A sample use case can be this: “While preparing a document in Openoffice, I want to write in English as well as in Hindi. For doing spellcheck, I need to manually change the language rather than the application detect it automatically” ...

November 14, 2008 · 4 min · Santhosh Thottingal

Gedit plugin for showing unicode codepoints

While working with Unicode text, it is often required to get the Unicode code points of text for debugging. Using python, it is very easy to get the unicode codepoints of the text. Following examples illustrates it. ` “സന്തോഷ്”.decode(“utf-8”) u’\u0d38\u0d28\u0d4d\u0d24\u0d4b\u0d37\u0d4d’ ` or ` str=u"സന്തോഷ്" print repr(str) u’\u0d38\u0d28\u0d4d\u0d24\u0d4b\u0d37\u0d4d’ ` Well, But we need to take python console and type/paste the text etc..How can we make it more easy? What if pressing F12 key after selecting some text gives the codepoints? ...

November 12, 2008 · 1 min · Santhosh Thottingal

Screensavers in your language

I had written a blog post about hacking the glmatrix screensaver with the glyphs of our languages. Now I have those screensavers in the following languages: Hindi : Deb Package , RPM Gujarati : Deb Package , RPM Bengali : Deb Package , RPM Oriya: Deb Package , RPM Tamil : Deb Package , RPM Malayalam: Deb Package , RPM Try it and enjoy !! ps: I used the default fonts of Fedora 9 for these. If you have any specific font to be used please let me know. I used Dyuthi calligraphic font for Malayalam.

October 27, 2008 · 1 min · Santhosh Thottingal

Swanalekha M17N based Input Method for 11 Languages

Swanalekha is an Input method originally designed for Malayalam. It is works with scim. as well as m17n. The input method scheme is transliteration based and it has a unique feature of candidate list menu(which I will explain shortly). Now I have extended it to 10 other Indian languages. Before explaining how swanalekha is different from other phonetic/transliteration based input methods, let me explain some of the characteristics of transliteration. Transliteration based input methods were following a strict one to one mapping from english letters to another Indian language. For eg: The ka=क ,pa = प , ti = टि etc.. when you write bharath, you will easily transliterate it to hindi as भारत. But for a rule based transliteration system it is भरत unless the english is bhaarath. Some times it may be Bhaarat too.. See another example: Kartik. it should be transliterated to കാര്‍ത്തിക് in Malayalam. So some people write it as Karthik, and some others write it as karthick too. All these are based on personal preferences. But when it use transliteration based input methods, people find difficulty with using a strict rule based writing method. There they have to write kaa for കാ or કા or கா or কা. Users like to get what they mean without the difficulty following the strict rules of transliteration. In an Intelligent transliteration based system when somebody write linux they should be able to map it to लिनक्स . Some times a choice to select लैनक्स is also preferable. This is what google transliteration does. No rules, no learning.. just type in english… ...

October 27, 2008 · 5 min · Santhosh Thottingal

സോഫ്റ്റ്‌വേര്‍ സ്വാതന്ത്ര്യദിനാഘോഷം 2008: ഭാഷാ കമ്പ്യൂട്ടിങ്ങ് സെമിനാറും ഇന്‍സ്റ്റാള്‍ ഫെസ്റ്റും

സോഫ്റ്റ്‌വേര്‍ സ്വാതന്ത്ര്യദിനാഘോഷം 2008 ഭാഷാ കമ്പ്യൂട്ടിങ്ങ് സെമിനാറും ഇന്‍സ്റ്റാള്‍ ഫെസ്റ്റും മലബാര്‍ ക്രിസ്ത്യന്‍ കോളേജ്, കോഴിക്കോട് സപ്തംബര്‍ 20, രാവിലെ 10 മണി മുതല്‍ വൈകുന്നേരം 5 മണി വരെ സംഘാടനം: സ്വതന്ത്ര മലയാളം കമ്പ്യൂട്ടിംഗ്, മലബാര്‍ ക്രിസ്ത്യന്‍ കോളേജ്, കോഴിക്കോട്, ഫോസ്സ്‌സെല്‍ നാഷനല്‍ ഇന്‍സ്റ്റിറ്റ്യൂട്ട് ഓഫ് ടെക്ലനോളജി – കോഴിക്കോട് വിവരസാങ്കേതികവിദ്യയുടെ മാനുഷികവും ജനാധിപത്യപരവുമായ മുഖവും ധിഷണയുടെ പ്രതീകവുമാണു് സ്വതന്ത്രസോഫ്റ്റ്‌വേറുകള്‍. പരമ്പരകളായി നാം ആര്‍ജ്ജിച്ച കഴിവുകള്‍ വിജ്ഞാനത്തിന്റെ സ്വതന്ത്ര കൈ മാറ്റത്തിലൂടെ, ചങ്ങലകളും മതിലുകളും ഇല്ലാതെ, ഡിജിറ്റല്‍ യുഗത്തില്‍ ഏവര്‍ക്കും ലഭ്യമാക്കുന്നതിനും ലോകപുരോഗതിക്കു് ഉപയുക്തമാക്കുവാനുമാണു് സ്വതന്ത്ര സോഫ്റ്റ്‌വേറുകള്‍ നിലകൊള്ളുന്നതു്. സ്വതന്ത്ര സോഫ്റ്റ്‌വേറുകള്‍ വാഗ്ദാനം ചെയ്യുന്ന മനസ്സിലാക്കാനും പകര്‍ത്താനും നവീകരിക്കാനും പങ്കുവെക്കുവാനുമുള്ള സ്വാതന്ത്ര്യമാണു് സ്വതന്ത്ര വിവരവികസന സംസ്കാരത്തിന്റെ അടിത്തറ. ഈ സ്വാതന്ത്ര്യം പൊതുജനമദ്ധ്യത്തിലേക്കു് കൊണ്ടു വരുവാനും പ്രചരിപ്പിക്കാനുമായി ഓരോ വര്‍ഷവും സപ്തംബര്‍ മാസ ത്തിലെ മൂന്നാമത് ശനിയാഴ്ച ലോകമെമ്പാടും സോഫ്റ്റ്‌വേര്‍ സ്വാതന്ത്ര്യ ദിനമായി ആചരിക്കുന്നു. ...

September 19, 2008 · 3 min · Santhosh Thottingal

Geo-visualisation, the FOSS way

My friend Jaisen Nedumpala has been developing a Geo-visualisation system for Cheruvannoor Grama Panchayath(Page in ml_IN) of Kerala. The system, developed using FOSS tools is available here “Development of effective geo-visualisation based decision support system (DSS) involved primarily data compilation from collateral sources, setting up appropriate hardware configuration, design of database and design of a spatial DSS. ” Jaisen used softwares like GRASS, UMN MapServer and ka-Map. He has written a detailed documentation(English) on how he developed this and what are all the tools used.

September 5, 2008 · 1 min · Santhosh Thottingal

UTF8Decoder

zabeehkhan was trying to code a Pashto (ps_AF) module for dhvani. And he told me that “it is not saying anything” :). So I took the code and found the problem. Dhvani has a UTF-8 decoder and UTF-16 converter. It was written by Dr. Ramesh Hariharan and was tested only with the unicode range of the languages in India. It was buggy for most of the other languages and there by the language detection logic and text parsing logic was failing. So I did some googling, went through the code tables of gucharmap and got some helpful information from here and here ...

September 1, 2008 · 3 min · Santhosh Thottingal

Say NO to Software Patents

August 20, 2008 · 0 min · Santhosh Thottingal

കെ.ഡി.ഇ. 4.1 പുറത്തിറങ്ങി

മലയാളം കമ്പ്യൂട്ടിങ്ങിന്റെ ചരിത്രത്തിലെ ഒരു സുപ്രധാനനാഴികക്കല്ലായി KDE 4.1 പുറത്തിറങ്ങിയിരിക്കുന്നു…..! KDE യില്‍ ആദ്യമായി മലയാളത്തിനു് ഔദ്യോഗിക പിന്തുണയുമായി…..! SMC യുടെ ചരിത്രത്തിലെ നാഴികക്കല്ലുകളിലൊന്നാണിതു്. 10 ദിവസത്തിനുള്ളില്‍ രാത്രിയും പകലും 25 ല്‍ കൂടുതല്‍ കൂട്ടുകാരുടെ കഠിനപരിശ്രമത്തിന്റെ ഫലമായി 10000 ത്തില്‍ പരം വാചകങ്ങള്‍ തര്‍ജ്ജമ ചെയ്താണു് ഇതു സാധ്യമായതു്. മലയാളത്തില്‍ തന്നെയുള്ള പ്രസാധനക്കുറിപ്പു് വായിയ്ക്കൂ കൂടുതല്‍ വിവരങ്ങള്‍ : KDE 4.1 to Officially Support Malayalam- Praveen’s Blog KDE യെപ്പറ്റി. KDE 4.1 Malayalam Screenshots

July 29, 2008 · 1 min · Santhosh Thottingal