Grapholinguistics Proceedings 2024

In 2024, I presented a paper on “Parametric Type Design” based on the technology behind the Nupuram font at the Grapholinguistics conference in Italy. The proceedings of the conference—edited by Yannis Haralambous—are now published as a two-volume set. Print editions are available on Amazon, and free digital versions are accessible online. Grapholinguistics is an interdisciplinary field dedicated to the scientific study of written language, bringing together typographers, computer scientists, linguists, font designers, psychologists, archaeologists, and educators. ...

July 24, 2026 · 2 min · Santhosh Thottingal

പത്തുകോടിയുടെ മലയാളം കോർപ്പസ്

മലയാളത്തിൽ AI കോർപ്പസ് നിർമിക്കാനായി പത്തുകോടി രൂപ ഈയിടെ അവതരിപ്പിച്ച കേരളബഡ്ജറ്റിൽ നീക്കിയിരിത്തിയിട്ടുണ്ട്. ആർട്ടിഫിഷ്യൽ ഇന്റലിജൻസ് സമീപകാലത്ത് നേടിയിട്ടുള്ള വലിയ പുരോഗതിയുടെ പശ്ചാത്തലത്തിൽ മലയാളഭാഷ പുറംതള്ളാതെയിരിക്കാനായിട്ടാണ് ഈ “മലയാളം AI സംരംഭം” ബഡ്ജറ്റിൽ വിഭാവനം ചെയ്തിരിക്കുന്നത്. അത് സദുദ്ദേശപരവും കാലോചിതവുമാണ്. അതോടൊപ്പം തന്നെ ഈ പത്തുകോടികൊണ്ട് നമുക്കെന്തു ചെയ്യാനാകുമെന്നതിനെക്കുറിച്ചുള്ള കുറച്ചുചിന്തകൾ കൂടി പങ്കുവെക്കട്ടെ. എന്താണ് ഒരു കോർപ്പസ്? AI മോഡലുകളുടെ പരിശീലനത്തിന് അതിവിപുലമായ ഉള്ളടക്കം ആവശ്യമാണ്. ഇത് ടെക്സ്റ്റ് ആവാം, ചിത്രങ്ങളാവാം, സംഭാഷണങ്ങളുടെ റെക്കോർഡിങ്ങ് ആവാം. ഇതിനെയാണ് ട്രെയിനിങ്ങ് കോർപ്പസ്സ് എന്ന് വിളിക്കുന്നത്. എത്രത്തോളം വലുതാണ് കോർപ്പസ് അത്രത്തോളം ഈ മോഡലുകൾ മെച്ചമായിരിക്കുമെന്നാണ് നിലവിലെ ലാർജ് ലാംഗ്വേജ് മോഡലുകളുടെ സാങ്കേതികവിദ്യ. അതിനായി, ഇത്തരം മോഡലുകളുടെ നിർമാതാക്കളും ഗവേഷകരും ഇന്റർനെറ്റിൽ നിന്ന് കിട്ടാവുന്ന എല്ലാ ഡാറ്റയും ഉപയോഗിക്കാൻ ശ്രമിക്കുകയാണ്. ഈ മത്സരത്തിൽ പകർപ്പവകാശവും ഉള്ളടക്കത്തിന്റെ ഉടമസ്ഥാവകാശവുമൊക്കെ പലവട്ടം കോടതി കേറിയിറങ്ങിയെങ്കിലും വലിയവിജയമൊന്നും നേടിയിട്ടില്ല. ...

June 26, 2026 · 4 min · Santhosh Thottingal

ഡിറ്റക്ടീവ് ബ്യോംകേഷ് ബക്ഷി - മലയാള പരിഭാഷ

ശരദിന്ദു ബന്ദോപാധ്യായ രചിച്ച അനശ്വര ബംഗാളി കുറ്റാന്വേഷണപരമ്പരയായ ഡിറ്റക്ടീവ് ബ്യോംകേഷ് ബക്ഷിയുടെ മലയാളപരിഭാഷ മാതൃഭൂമി ബുക്ക്സ് പുറത്തിറക്കുന്ന വാർത്ത ഞാൻ വായിച്ചിരുന്നു. വാർത്തയോടൊപ്പം പുസ്തകത്തിന്റെ കവർപേജിന്റെ ചിത്രവും ഉണ്ടായിരുന്നു. കവർപേജിൽ മഞ്ജരി ഫോണ്ട് കണ്ടതുകൊണ്ടും ബ്യോംകേഷ് ബക്ഷിയെപ്പറ്റി പലയിടത്തും കേട്ടിട്ടുള്ളതുംകൊണ്ടാണ് പുസ്തകം വാങ്ങിയത്. മലയാളത്തിലേക്ക് പരിഭാഷപ്പെടുത്തിയത് പ്രശസ്ത വിവർത്തക ലീലാസർക്കാർ ആയതുകൊണ്ട് മോശമാവാനിടയില്ലെന്നും കരുതി. പുസ്തകം വായിച്ചു തുടങ്ങിയപ്പോൾ എന്നെ നിരാശപ്പെടുത്തിക്കൊണ്ട് വികലമായ പരിഭാഷകൾ കാണാൻ തുടങ്ങി. ലീലാ സർക്കാർ തന്നെയാണൊ ഇത് പരിഭാഷപ്പെടുത്തിയതെന്ന് സംശയിക്കാതെയുമിരുന്നില്ല. 1950 കൾക്ക് സമീപത്തെ കൊൽക്കത്തയുടെ സാംസ്കാരികപരിസരത്തെ മലയാളത്തിലേക്ക് കൊണ്ടുവരുന്നത് അത്ര എളുപ്പമല്ല. പക്ഷേ അതിന് ശ്രമിച്ചിട്ടുപോലുമില്ല. പകരം കണ്ടത് അക്ഷരത്തെറ്റുകൾ. നിഘണ്ടു ഉപയോഗിച്ചാൽ പോലും നന്നായി പരിഭാഷപ്പെടുത്താവുന്ന വാക്കുകൾ തെറ്റിയിരിക്കുന്നു. ഒരു പ്രാവശ്യം പോലും ആരും ഒന്നു മനസ്സിരുത്തി വായിക്കാത്രെ പ്രസിദ്ധീകരിച്ചോയെന്നു തോന്നിപ്പോയി. ...

June 5, 2026 · 2 min · Santhosh Thottingal

Home Maker: Declare Your Dev Tools in a Makefile

Your laptop has ripgrep, installed via cargo install. ruff is there too, via uv tool install. golangci-lint came from go install. bash-language-server was npm i -g. Neovim was a tarball download. Kitty was a curl script. Six months later you get a new machine, or you just want to upgrade or reinstall. What do you even have installed? How did you install each one? Which version? Good luck. This is a small system that answers those questions — a single Makefile that declares every tool you care about, grouped by purpose, with one command to install anything. ...

March 29, 2026 · 8 min · Santhosh Thottingal

Rendering complex scripts in terminal and OSC 66

As a programmer, I spend most of my time in a terminal application like Kitty. I use Neovim as my code editor. I use CLI based AI agents. But the biggest pain, even in 2026, is that there is no terminal that can render complex scripts like Indic languages or Arabic. This is a significant limitation for me, as most of my work involves language processing. In this article, I will give a brief overview of why this issue remains unsolved—covering the character-cell grid model, width measurement, and the distinction between text shaping and rendering—along with ongoing efforts and a small tool I built recently that illustrates a solution path. ...

March 22, 2026 · 11 min · Santhosh Thottingal

CLI for transforming Wikipedia articles to text, markdown, and JSON

We are witnessing a resurgence and evolution of Command Line Interfaces (CLIs), accelerated by AI agents. Text-based, scriptable CLI tools work very well with LLM-based workflows. Accessing Wikipedia articles during an agent session is common. Usually, a webfetch call is used to get the HTML for a page from a URL like https://en.wikipedia.org/wiki/2026_Winter_Olympics. That works, and LLMs are smart enough to read HTML. But there is a cost: HTML is for rendering, so the model must ignore a lot of non-content markup to get to the useful text. i That increases token usage and adds context noise. Can we improve this? ...

March 14, 2026 · 7 min · Santhosh Thottingal

Preparing sentence dataset from a wikipedia

I’m excited to announce two new resources for natural language processing researchers and developers: wikisentences - A Rust-based tool for extracting sentence datasets from Wikipedia dumps in any language ml-wiki-sentences - A dataset of 2.25 million Malayalam sentences extracted from Wikipedia, now available on HuggingFace, prepared using the above tool. The Wikisentences Tool The wikisentences project provides a complete pipeline for creating sentence datasets from Wikipedia content: Core Technology wiki-html-text-extractor (Rust) - Uses tree-sitter-html to parse article HTML and extract clean plain text sentencex (Rust) - Handles accurate sentence segmentation across languages. See my recent article about this library Four-Stage Pipeline Download enterprise HTML dumps from WikimediaThere is no recent html dumps for wikipedia, except this one year old dump Convert JSON dumps to Parquet format (id, name, url, language, html) Extract plain text from HTML (id, url, name, text) Segment text into sentences (id, url, name, sentence, sentence_index) Each stage is handled by a separate Python script, with the heavy lifting done by efficient Rust binaries. The pipeline is designed to be memory-efficient, streaming data between stages without writing intermediate files to disk. ...

March 14, 2026 · 3 min · Santhosh Thottingal

How to identify and annotate sentences in an HTML page

If you have ever needed to work with sentences inside an HTML page — highlight them, translate them, read them aloud — you quickly run into a deceptively awkward problem. The text is not plain text. It is interspersed with tags, attributes, inline elements, and markup that your sentence detector has no business reading. This post walks through an exploratory JavaScript project — html-sentence-segmenter — that I built to figure out how to do this properly. It is not a polished, reusable library. Think of it as a working proof-of-concept that demonstrates the approach, with a live demo using Wikipedia articles. ...

March 7, 2026 · 8 min · Santhosh Thottingal

Rewriting the multilingual sentence segmenter - sentencex - in Rust

In October 2023, we announced sentencex — a sentence segmentation library built for the wide language diversity of Wikipedia and Wikimedia projects. The original release comprised two separate libraries: one in Python (used by MinT, our machine translation service) and one in JavaScript (used by Content Translation). Both were rule-based, practical, and got the job done. But maintaining two implementations of the same logic was a constant source of drift. sentencex 1.0 is a ground-up rewrite in Rust, with the Python and JavaScript libraries replaced by thin bindings over a single, shared core. ...

March 7, 2026 · 6 min · Santhosh Thottingal

From Tokens to Text: A Trigram Markov Model for Malayalam

Ever wondered how a computer learns to generate text that actually looks like Malayalam? Not just random characters, but something with actual structure? I’m not talking about Large Language Models here. I’m talking about Small Language Models that are efficient and explainable. something you can build and run on your own laptop. In my previous “The Broken Token” article, I presented a Malayalam unigram tokenizer and analysed its strengths and weaknesses. I did fertility rate evaluation and then analysed the tokenization in the context of Malayalam language characteristics. A common evaluation method for tokenizers is using them in downstream tasks—so I decided to build a text generator. That’s where things got interesting. ...

February 28, 2026 · 11 min · Santhosh Thottingal