Bhashini is a practical entry point for adding Indian-language capabilities to products without training every speech or translation model yourself. Through the ULCA ecosystem, developers can discover available services for translation, automatic speech recognition (ASR), text-to-speech (TTS) and related language tasks across India’s scheduled languages.
The important distinction is that Bhashini is not a single fixed model endpoint. It is a pipeline-based model discovery and inference system. Your application first identifies a compatible pipeline and service, then sends a task-specific request to the inference endpoint. That extra step gives you access to multiple providers, but it also means that credentials, language codes, payload formats and service availability must be handled carefully.
This guide explains how to integrate the Bhashini Model API in a backend service, with examples for translation and practical guidance for ASR and TTS. Endpoint names and available services can change, so validate current values in the Bhashini API portal before deploying.
What you need before coding
Prepare the following items:
- A Bhashini or ULCA developer account.
- Your account-level
userIDandulcaApiKey. - A pipeline ID suitable for the task you want to run.
- Source and target language codes, such as
en,hi,taorbn. - A backend environment that can keep credentials private.
- Test inputs representing real users, including regional spellings, code-switching and noisy audio.
Do not put the account API key in browser JavaScript or a mobile application. Route requests through your server, authenticate your own users there, and apply rate limits before forwarding traffic to Bhashini.
If you plan to run an Indic model locally instead of using a hosted service, compare the operational trade-offs with how to deploy large language models locally. Local deployment can improve control and predictable latency, but requires model hosting, GPU capacity, monitoring and update processes.
How Bhashini’s two-stage flow works
A typical integration has two requests:
1. Pipeline discovery: send the task and language configuration to ULCA. The response identifies a service, inference endpoint and temporary or service-specific authorization details.
2. Inference: send the actual text or encoded audio to the returned endpoint using the service information from the first response.
The discovery request commonly uses an endpoint similar to:
POST https://meity-auth.ulcacontrib.org/ulca/gw/v1/get-models-pipelineUse the current endpoint and header names shown in the portal. A representative request looks like this:
{
"pipelineTasks": [
{
"taskType": "translation",
"config": {
"language": {
"sourceLanguage": "en",
"targetLanguage": "hi"
}
}
}
],
"pipelineId": "YOUR_PIPELINE_ID"
}Send the account credentials as headers, typically userID and ulcaApiKey, with Content-Type: application/json. Treat the response as configuration data, not as a permanent constant. Store it in a short-lived cache and refresh it when the service expires or becomes unavailable.
Translation integration example
The exact response shape can vary by pipeline. Extract the returned serviceId, inference URL and inference authorization value rather than hard-coding a model name from an old example. Then construct the task payload with the selected service:
import requests
def translate_text(text, source_language, target_language,
inference_url, inference_key, service_id):
payload = {
"pipelineTasks": [{
"taskType": "translation",
"config": {
"language": {
"sourceLanguage": source_language,
"targetLanguage": target_language
},
"serviceId": service_id
}
}],
"inputData": {
"input": [{"source": text}]
}
}
response = requests.post(
inference_url,
json=payload,
headers={
"Authorization": inference_key,
"Content-Type": "application/json"
},
timeout=30
)
response.raise_for_status()
return response.json()In production, validate the response before returning it to a user. Check that the expected output field exists, preserve the detected service ID in logs, and return a controlled fallback when a provider fails. Do not log API keys, raw personal data or sensitive audio.
Translation quality is highly domain-dependent. Test names, addresses, product terminology, government forms and mixed English-Indic sentences separately. If your product serves Hindi users, a comparison with open-source small language models for Hindi can help you decide whether Bhashini’s hosted pipeline or a specialised local model is the better fit.
Adding ASR and TTS
For ASR, the input is generally audio represented as Base64 inside the request payload. Capture audio in a predictable format, normalise sample rate and channels, and reject files that exceed your size or duration limit before encoding them. A simplified input structure may look like:
{
"pipelineTasks": [
{
"taskType": "asr",
"config": {
"language": {"sourceLanguage": "hi"},
"serviceId": "YOUR_ASR_SERVICE_ID"
}
}
],
"inputData": {
"audio": [{"audioContent": "BASE64_AUDIO"}]
}
}Confirm the accepted MIME type, sampling rate and field names for the selected ASR service. Do not assume that every provider supports the same audio encoding. For voice bots, separate recording, upload, transcription and response generation into measurable stages so you can identify where latency is introduced.
TTS reverses the flow: submit text and a language or voice configuration, then decode the returned audio content. Set limits on text length, cache repeated announcements, and use a queue for long-form synthesis. Browser playback also requires correct content types and a safe audio delivery path.
If your application combines speech with document or image inputs, review open-source vision-language models for Indian languages when designing the broader multilingual architecture.
Reliability, security and cost controls
A production integration should include:
- Timeouts and retries: retry transient
429,502and503responses with exponential backoff and jitter. Do not blindly retry malformed requests. - Circuit breaking: temporarily stop calling an unhealthy service and use a fallback message or provider.
- Caching: cache pipeline discovery results for a limited period, but refresh them after authorization or service errors.
- Chunking: split long text at sentence or paragraph boundaries, then preserve ordering when combining translations.
- Observability: track success rate, latency, input size, language pair, service ID and error category without storing sensitive content.
- Secret management: keep keys in environment variables or a managed secret store and rotate them when access changes.
- Input validation: enforce supported languages, maximum text length, audio duration and permitted file types.
For mobile or edge applications, measure payload size and network latency before selecting a hosted-only architecture. Guidance on AI model optimization for mobile devices is useful if you later move part of the workflow on-device.
Testing checklist for Indian-language products
A successful English-to-Hindi demo is not enough. Build a test set that includes:
- All target language pairs and scripts.
- Code-mixed sentences, abbreviations and numerals.
- Names, locations, units, currency and dates.
- Accents, background noise and overlapping speech for ASR.
- Short UI labels as well as long paragraphs.
- Low-bandwidth and repeated-request scenarios.
Measure task-specific quality: translation adequacy and terminology accuracy, ASR word error rate, TTS intelligibility and end-to-end response time. Have native speakers review high-impact flows such as payments, healthcare, education and public-service instructions.
Common integration failures
- 401 or invalid key: you may be using the account key where the inference authorization value is required, or sending the wrong header format.
- Unsupported language pair: the selected pipeline may not provide that combination; repeat discovery with a compatible pipeline.
- Empty output: inspect the complete response and confirm the input field matches the selected task.
- Audio rejection: check encoding, Base64 validity, sample rate, channel count and file duration.
- Intermittent 5xx errors: add bounded retries, record the service ID and test another available service rather than increasing request volume.
Bhashini works best when treated as an evolving service catalogue rather than a static API wrapper. Keep discovery logic isolated, make model selection configurable, test each language in realistic conditions and design a fallback before launch. That approach lets Indian startups ship multilingual features quickly while retaining the control needed for reliable production systems.