தமிழ்AI

Layer 4

பகிர்வு distribution

A grammar engine that nobody can reach is a hobby. This layer is about the distance between a working server and a Tamil teacher who has never heard of any of this.

An MCP server is not a thing people use

It is an adapter that lets an AI assistant call your logic. Nobody outside software will ever install one, and treating it as the product would strand this project among developers.

So the thing being distributed is the engine, and the MCP server is one adapter on it. A REST API is another, the browser page is another, and the command line is a fourth. Each reaches a different person, and only one of them requires the reader to be technical at all.

Assistants

Somebody asks Claude a Tamil grammar question, and it looks the answer up rather than recalling it.

App builders

A Tamil grammar helper for school students becomes an API call. The developer needs no grammar expertise, and the teacher can check the citation.

Everyone else

A web page where you type a word. No install, no account, no assistant. This is where most Tamil speakers are.

Researchers

Verified datasets, published where people who train models actually look.

The rungs, and where we actually are

Rung two is not done. Nothing has been released, the version is still 0.1.0, and today the only way to run this is to clone it. That is the honest position and it is the most useful thing on this page.

  1. Now

    Clone it and run it

    The code is public and Apache-2.0. Anyone can clone it, install the dependencies and run the server or the browser page today.

    Cost: Nothing to us. Somebody else’s machine.

  2. Next

    Install it in one command

    A published package, so an ordinary install works without cloning, and a container image for people who would rather not install a native binary at all.

    Cost: Nothing to us. Still their machine.

  3. Then

    Be findable

    The registries an MCP client actually reads, and the Tamil NLP catalogue that this project sources from in the first place. An unlisted server is invisible.

    Cost: Time, not money.

  4. Then

    A page anyone can use

    The browser page already exists and runs. Making it public for everyone is a hosting decision rather than a building one.

    Cost: Real, and small. This is the first rung that costs anything.

  5. Later

    A phone app

    A thin app over the hosted API. Running the analyser on the phone itself is not realistic: it needs a native binary and a lexicon, and that is a much larger, later project.

    Cost: Store fees, plus the backend it talks to.

Data is a channel too

Every analysis the server resolves is stored as verified, provenance-tagged data. Published properly, that becomes a resource other people can use without touching our software at all, and it is arguably the more durable contribution.

We surveyed what already exists for Tamil on the main model-sharing platform. There is a good deal of speech data, raw text and sentiment corpora. There is no morphological segmentation gold set, no loanword-to-equivalent dataset, and no origin-label dataset. Those three are precisely what this project produces as a by-product, which is either a happy accident or the reason the project is shaped the way it is.

Data published from a shared corpus needs a consent and licensing model before anyone else's usage goes into it. Our own hosted instance is a different question from pooling other people's installs, and the second one is deliberately not built.

What this costs, and why that shapes the plan

A server other people run costs us nothing. A service we host for everyone costs money every month, forever, and this is a nonprofit with volunteer time as its scarcest resource.

So the order is deliberate: everything that is free to us ships first, and the hosted service follows once there is reason to believe people want it. Adoption is the goal. Self-hostable and open-source is also the exit: if this project ever stops, a university or a school can keep running it without asking anyone.

And where it goes after that →