Starter Kit. Idea 308 of 500. Cluster 15
Sinhala and Tamil language technology: keyboards, speech, translation, sold to the platforms that need it
The mechanism, from the book
The proof exists in Malabe. What is missing is a thousand small versions of it.
The first step
Build a small, clean dataset: a few hundred recorded sentences in Sinhala or Tamil with accurate transcripts, made with the speakers' written consent. Then write to the language teams of speech and translation companies and offer more of it on contract, with the sample attached. A reply asking for a price or a larger sample is the proof that the demand exists.
- Who pays first
- A speech or translation company that needs accurate Sinhala or Tamil data and cannot find it, and will pay per hour of clean recording.
- What leaves today
- The platforms serving Sinhala and Tamil speakers build their language tools elsewhere, so the expertise here earns nothing from them.
Ask first
- Information and Communication Technology Agency. Ask whether any national language-technology work exists that a small firm could join.
- Data Protection Authority. Ask what consent you need to record people's voices for datasets.
Check before you spend
Signed consent from every speaker for commercial use, who owns the recordings and transcripts, and what the buying company's contract requires on quality.
Find out these three numbers
- What do language teams pay per hour of transcribed speech?
- How many speakers and dialects does a buyer want?
- How long does one hour of clean recording take to transcribe?
The test
Does it keep value that now leaves the island, or build the proof that lets somebody else do so?
From the appendix of Why Not Sri Lanka? by Dr Maheshika Halbeisen. The idea and the mechanism are the book's; this kit was written for the site.