Choosing between multilingual models for Kannada sentiment analysis requires more than running two fine-tuning jobs and selecting the higher accuracy. Kannada text varies across formal writing, social media, code-mixed Kannada-English, transliterated Kannada, spelling variants, emojis, and domain-specific vocabulary. A useful AutoResearch experiment should make those differences visible and produce results another engineer can reproduce.
This guide presents a practical workflow for using Karpathy-style AutoResearch to compare MurIL and IndicBERT. Treat AutoResearch as an experiment loop and configuration pattern rather than assuming a particular package, command, or model alias. Verify the current repository, supported checkpoints, and launcher before running it; the original draft’s autorsearch installation command may not match the project you intend to use.
Define the comparison before training
Start with a precise question: which encoder gives the strongest and most reliable Kannada sentiment classifier under the same compute and data budget? Fix the following in a versioned configuration:
- Model checkpoints and tokenizer versions.
- Number of sentiment labels: binary, three-way, or a business-specific taxonomy.
- Maximum sequence length, batch size, learning-rate schedule, and number of epochs.
- Random seeds and number of repeated runs.
- Training, validation, and held-out test splits.
- Hardware, software versions, and evaluation script.
MurIL and IndicBERT may differ in architecture, vocabulary, pre-training data, and parameter count. Record those differences rather than presenting the experiment as a perfectly controlled scientific comparison. If one model has a substantially larger compute footprint, report both quality and cost.
For broader guidance on evaluation design, see how to evaluate Kannada small language models. The same discipline—fixed splits, explicit metrics, and qualitative error review—applies here.
Build a Kannada dataset that exposes real failures
A balanced dataset is useful, but balance alone does not make it representative. Include examples from the actual setting in which the classifier will run:
- Kannada news, product reviews, customer messages, or social posts.
- Formal Kannada and colloquial Kannada.
- Kannada written in Kannada script and transliterated into Latin script.
- Kannada-English code-mixing, abbreviations, emojis, hashtags, and punctuation.
- Positive, negative, neutral, mixed, and genuinely ambiguous examples.
Create annotation guidelines with examples of sarcasm, praise containing complaints, negation, comparisons, and statements that express an opinion without clear polarity. Have at least two annotators label a meaningful sample, then adjudicate disagreements. Report agreement and retain an uncertain or needs review category where appropriate instead of forcing every item into a noisy label.
Prevent leakage by splitting at the user, product, article, or conversation level when repeated content is possible. Near-duplicate reviews across train and test sets can make a weak model look excellent. Keep the test set untouched until the experiment is locked.
Configure a reproducible AutoResearch loop
Use a configuration file or equivalent experiment manifest. The exact field names depend on the AutoResearch implementation, but the manifest should express the following:
models:
- name: muril
checkpoint: google/muril-base-cased
- name: indicbert
checkpoint: ai4bharat/indic-bert
data:
train: data/kn_sentiment_train.jsonl
validation: data/kn_sentiment_valid.jsonl
test: data/kn_sentiment_test.jsonl
text_column: text
label_column: label
experiment:
seeds: [13, 42, 2026]
max_length: 256
epochs: 3
primary_metric: macro_f1
output_dir: runs/kannada-sentimentCheck each checkpoint’s actual model class and tokenizer requirements. Do not assume that a model name in a blog post is a valid Hugging Face identifier. Pin package versions, save the resolved configuration, and log the dataset hash. If AutoResearch modifies hyperparameters automatically, impose the same search budget for both models and preserve every trial result.
A sensible run structure is:
1. Confirm that both tokenizers process Kannada text correctly.
2. Run a small smoke test to catch label, padding, and device errors.
3. Launch the same training and tuning budget for each checkpoint.
4. Evaluate every trial on validation data only.
5. Select settings using the predefined primary metric.
6. Run the winning configuration once on the locked test set.
7. Repeat the comparison across several seeds.
Use metrics that reflect deployment risk
Accuracy is easy to understand but can hide poor performance on minority classes. Make macro-F1 the primary metric when positive, negative, and neutral classes matter equally. Also report:
- Per-class precision, recall, and F1.
- Weighted-F1 and accuracy for operational context.
- Confusion matrices, especially neutral-versus-negative errors.
- Confidence calibration, if predictions will trigger workflows.
- Mean and standard deviation across seeds.
- Inference latency, memory use, and model size.
For imbalanced data, include class counts and, if used, class weights or sampling strategy. Never compare one model’s best seed with another model’s average. Report mean performance and variation under the same protocol.
You can reuse practices from benchmarking Kannada generation quality with Evaluate, while adapting the metric set to classification. If the classifier will support a Kannada customer-service product, pair offline scores with a small review set from production-like messages; the workflow is related to fine-tuning a small language model for Kannada customer support.
Analyse errors, not just the leaderboard
After the automated runs, export predictions with text, gold label, predicted label, confidence, model, and seed. Inspect errors by slice:
- Script: Kannada script versus Latin transliteration.
- Language mix: Kannada-only versus Kannada-English.
- Length and tokenization rate.
- Domain, source, and topic.
- Negation, sarcasm, emojis, and spelling variation.
Compare tokenization on representative examples. Excessive fragmentation can increase sequence length and weaken performance, but it is only one diagnostic—not proof that a model is inferior. Review high-confidence mistakes first; these often reveal annotation problems, domain shift, or systematic blind spots.
Use bootstrap confidence intervals or paired significance tests where appropriate. A tiny macro-F1 difference may not justify switching models if it disappears across seeds or comes with much higher latency. Conversely, a model with similar average F1 but better recall for critical negative feedback may be the better product choice.
Improve the winning system responsibly
If both baselines underperform, improve the data before adding complexity. Correct inconsistent labels, add difficult examples, and expand underrepresented Kannada varieties. Then test targeted changes one at a time:
- Domain-adaptive pre-training on permitted Kannada text.
- Carefully reviewed transliteration and code-mixed examples.
- Class weighting or threshold tuning.
- Longer context only when the task requires it.
- Distillation or quantization for constrained deployment.
For a smaller production footprint, compare the selected encoder with a compact model using the same test slices. Related implementation paths include fine-tuning a Kannada model with Hugging Face AutoTrain and running a quantized Kannada model offline.
A decision rule for 2026 deployments
Choose MurIL or IndicBERT using a written rule agreed before reviewing the final test results. For example: select the model with the higher mean macro-F1, provided its lower confidence bound is competitive, minority-class recall clears the product threshold, and latency fits the service budget. Otherwise, prefer the operationally simpler model and document the trade-off.
Publish the model checkpoints, configuration, data card, label policy, slice metrics, and known limitations where licensing and privacy allow. Kannada sentiment systems can encode annotator and domain bias, so monitor drift after launch and route uncertain predictions for review. A reproducible comparison is not the end of the project—it is the baseline that makes later improvements credible.