Tulu is spoken across coastal Karnataka and northern Kerala, yet it remains poorly represented in mainstream speech technology. For builders, the challenge is not simply finding an audio folder labelled “Tulu”. A useful dataset needs clear transcripts, speaker consent, licensing terms, metadata, and enough variation to support the intended model.
This guide explains where to look, how to verify what you find, and how to create a responsible Tulu dataset when public resources are insufficient.
Start by defining the speech task
The right dataset depends on the product you are building:
- Automatic speech recognition (ASR): requires speech recordings paired with accurate transcripts. Read speech is easier to begin with; conversational audio is more representative of real use.
- Text-to-speech (TTS): needs clean, consistently recorded speech from one or more speakers, aligned with carefully edited text.
- Keyword spotting: needs many short examples of target words, plus background noise and negative examples.
- Voice agents: typically need ASR, text-to-speech, pronunciation handling, dialogue evaluation, and recordings of real accents and code-switching.
- Language documentation: may prioritise breadth, cultural context, and speaker metadata over model-ready formatting.
A narrow pilot—such as recognising appointment requests or local place names—usually needs less data than a general-purpose assistant. Review the practical capabilities and limitations described in what a voice agent is and how voice AI works in 2026 before setting a collection target.
Where to look for Tulu voice datasets
1. Mozilla Common Voice and other open speech repositories
Check Mozilla Common Voice for Tulu entries, but verify the language label, clip count, transcript quality, speaker diversity, and current licence before using any release. Contributions can change significantly between versions. Download the dataset card and preserve the exact version used in experiments.
Also search speech-resource catalogues such as OpenSLR, Hugging Face Datasets, Kaggle, Zenodo, and institutional research repositories. Search with several terms: Tulu, Tuluva, ತುಳು, Tulu speech, Tulu ASR, and South Indian language speech corpus. A dataset may be described in a paper while the files are hosted elsewhere.
Do not assume that a public download means unrestricted commercial use. Check whether the licence covers training, redistribution, derivatives, and commercial deployment.
2. Indian language technology programmes
Search resources produced through Indian language technology initiatives, university labs, and public research programmes. The Bhashini ecosystem is particularly relevant because it supports Indian-language datasets, translation, speech recognition, and language technology partnerships. Availability and access conditions vary, so contact the responsible programme or repository rather than relying on an old project page.
Useful searches include the names of institutions in Karnataka and Kerala, computational linguistics departments, speech-processing laboratories, and papers that mention Tulu as a low-resource or under-represented language. Research papers often reveal a contact author, corpus size, recording protocol, or follow-up project even when the original data is not openly downloadable.
3. Universities, linguists, and cultural organisations
Tulu documentation may sit with a university department, archive, language scholar, theatre group, or cultural organisation rather than a machine-learning platform. Contact researchers working in linguistics, Kannada and Tulu studies, oral history, or digital humanities. Ask specifically about:
- raw audio and edited audio;
- transcripts and translations;
- speaker consent forms;
- geographic and dialect metadata;
- restrictions on publication or redistribution;
- whether the data can be used to train commercial models.
A respectful partnership can produce better data than scraping public videos. It also gives speakers and local institutions a role in deciding how their language is represented.
4. Community-led collection
Tulu-speaking communities are the most important source for coverage, pronunciation, and cultural accuracy. Work through local colleges, community associations, radio stations, theatre groups, creators, and trusted NGOs. Offer clear information in accessible language about what will be recorded, where it will be stored, and whether participants can withdraw.
Public YouTube, Instagram, or podcast audio is not automatically available for model training. Recording a person in public does not remove privacy, copyright, or consent obligations. Obtain permission from both the speaker and the rights holder where applicable.
How to evaluate a dataset before using it
Create a dataset card before training. Record the source, version, licence, collection dates, number of speakers, hours of audio, dialect coverage, equipment, sample rate, transcript format, and known exclusions.
Then test a representative sample. Measure:
- Transcript accuracy: sample word error rate manually; names and local vocabulary need special attention.
- Speaker balance: check age, gender, geography, occupation, and speaking style without collecting unnecessary personal data.
- Audio quality: identify clipping, reverberation, background music, telephone compression, and inconsistent microphones.
- Language purity: quantify Kannada, Malayalam, English, and Hindi code-switching instead of silently removing it.
- Text normalisation: decide how numerals, abbreviations, borrowed words, punctuation, and variant spellings will be represented.
- Duplicate and leakage risk: ensure the same speaker or sentence does not appear across training, validation, and test sets.
For a serious benchmark, keep the test set private and speaker-disjoint. A model that performs well on familiar speakers can fail badly in villages, markets, homes, or phone calls.
Building a Tulu dataset when none is sufficient
A practical pilot can begin with 20–50 speakers and a carefully designed script. Include short and long sentences, questions, numbers, dates, names of local places, food, transport, health, and commerce. Add spontaneous speech only after the consent and transcription workflow is reliable.
Use a quiet but realistic recording environment, capture lossless audio where possible, and collect a short microphone and room sample. Store audio with stable identifiers rather than names. Keep consent records separate from model-training files, restrict access, and define a deletion process.
For ASR, double-check transcripts and adjudicate disagreements with a fluent reviewer. For TTS, prioritise consistent microphone placement, pacing, pronunciation, and text quality. Do not combine speakers casually: a multi-speaker TTS model requires speaker labels and enough balanced material per speaker.
Treat annotation as a local-language task. Transliteration may help engineers, but it should not replace a native-script or linguistically appropriate transcript. Build a pronunciation lexicon for names, places, and code-switched terms, and maintain a changelog for every normalisation rule.
Consent, licensing, and responsible release
A strong consent form should explain the purpose, intended users, storage period, publication plans, model training, commercial use, and withdrawal limits. Participants should know whether their voice can be used to generate synthetic speech that resembles them.
For an open release, choose a licence only after confirming that every contributor has granted compatible rights. Consider releasing metadata, scripts, evaluation protocols, or derived features separately if raw audio would expose participants. Never publish phone numbers, addresses, or identifiable background conversations.
If your end goal is a customer-facing system, plan evaluation with Tulu speakers before launch. Test recognition of accents, code-switching, names, noisy environments, and refusal or escalation flows. Teams considering a production voice system can also review multilingual voice agents for restaurants in India for a domain-specific deployment pattern.
A practical 30-day sourcing plan
- Days 1–5: define the task, target dialects, licence requirements, and success metrics.
- Days 6–10: search Common Voice, OpenSLR, Bhashini, Hugging Face, academic papers, and institutional contacts.
- Days 11–15: audit licences, sample quality, speaker coverage, transcript accuracy, and duplication.
- Days 16–22: run a small consented collection with native reviewers and document the pipeline.
- Days 23–27: prepare speaker-disjoint train, validation, and test splits.
- Days 28–30: benchmark a baseline model, publish a dataset card internally, and decide whether to expand.
The goal is not the largest possible corpus. It is a traceable, consented, representative dataset that can support a clearly defined Tulu use case. If you need outside engineering support, first understand the trade-offs in hiring voice agent developers and make dataset ownership part of the engagement.
FAQ
Is there one authoritative Tulu voice dataset?
Not generally. Availability changes, and many resources are research-led, small, private, or subject to access restrictions. Verify each source and version.
Can I train on Tulu videos found online?
Not without checking copyright, speaker consent, platform terms, and privacy. Public visibility is not the same as permission for AI training.
How much data is enough?
It depends on the task. A narrow keyword or intent pilot may need hours, while robust general ASR or high-quality TTS may require substantially more coverage and careful speaker balancing.
Should Tulu be transliterated into Latin script?
Use the representation that matches the task and users. Preserve the original script where possible, and document any transliteration so results remain reproducible.
AI Grants India supports Indian builders working on language technology and other high-impact applications. Explore AI Grants India for funding and support opportunities.