Why benchmarking Bengali aquaculture models needs a field protocol
A Bengali-language model can produce fluent answers and still give unsafe or unusable fish-farming advice. Aquaculture recommendations depend on species, pond size, water quality, stocking density, feed, season, location, and farmer budget. A benchmark must therefore test more than language quality: it must measure whether the system gives correct, actionable, locally appropriate, and safe advice.
This matters across West Bengal, Bangladesh-facing Bengali markets, and Bengali-speaking farming communities elsewhere in India. In 2026, teams building advisory products should evaluate the complete system—model, retrieval layer, speech interface, escalation workflow, and mobile delivery—not just the underlying language model. If your product includes images of fish disease or pond conditions, combine text evaluation with principles from open-source vision-language models for Indian languages.
Define the advisory tasks before choosing a model
Start with a task taxonomy. Avoid a benchmark made entirely of generic questions such as “How do I grow fish?” Build test cases around decisions farmers actually make:
- Pond preparation: liming, water conditioning, depth, fertilisation, and pre-stocking checks.
- Species and stocking: carp polyculture, tilapia, catfish, seasonal stocking, and seed quality.
- Feed management: ration calculation, feeding frequency, feed storage, and responses to uneaten feed.
- Water quality: dissolved oxygen, pH, ammonia, turbidity, temperature, and urgent corrective action.
- Disease and mortality: symptom triage, isolation, sample collection, and when to contact a fisheries officer or veterinarian.
- Weather and risk: heat, heavy rain, flooding, low oxygen, and transport stress.
- Economics: input costs, harvest timing, break-even calculations, and market uncertainty.
- Compliance and safety: permitted chemicals, withdrawal periods, worker safety, and environmental safeguards.
For each task, record the intended user, farm context, expected answer type, and risk level. A low-risk question about feeding intervals should not receive the same evaluation threshold as advice involving antibiotics, pesticides, or mass fish mortality.
Build a representative Bengali evaluation set
A useful dataset reflects real language, not polished textbook Bengali. Collect consented and anonymised queries from extension workers, call centres, farmer groups, field surveys, and pilot users. Include Bengali script, colloquial phrasing, code-mixing with English, spelling variation, voice-transcribed text, and regional vocabulary. Keep Bangla numerals and Arabic numerals where farmers use both.
Every example should include structured metadata:
- District or agro-ecological context, without exposing personal information.
- Fish species, culture system, pond characteristics, and production stage.
- Available measurements, such as pH or dissolved oxygen, and whether readings are reliable.
- Farmer objective, constraints, and urgency.
- Gold-standard answer, acceptable alternatives, and prohibited claims.
- Required clarification questions and escalation conditions.
Use separate development, validation, and locked test sets. Prevent leakage from training data, demonstrations, or repeated questions. Split near-duplicate queries by scenario rather than randomly; otherwise, a model may memorise wording while failing on a new pond situation. For a multilingual pipeline, compare Bengali directly with the same request in English and other relevant Indian languages. Guidance on benchmarking NLP models for Telugu and Sanskrit offers a useful structure for multilingual test design, although aquaculture requires additional domain and safety checks.
Score factuality, actionability, and safety separately
Do not collapse every result into one fluency score. Use a rubric with five-point ratings and task-specific pass criteria.
1. Factual correctness: Are the biological, technical, and numerical claims supported by trusted sources?
2. Context fit: Does the answer use the supplied species, season, water conditions, and farm constraints?
3. Actionability: Can a farmer follow the recommendation, with quantities, sequence, timing, and required tools stated clearly?
4. Calibration: Does the model distinguish known facts, assumptions, and uncertainty? Does it ask for missing measurements?
5. Safety: Does it avoid dangerous treatment, unsupported chemical dosage, premature diagnosis, and false confidence?
6. Bengali quality: Is the language understandable to the intended audience, including units, terminology, and local phrasing?
Have at least two aquaculture experts independently review a stratified sample. Use a third reviewer for disagreements, and report inter-rater agreement. Automated metrics such as BLEU or ROUGE can help detect wording overlap but should not determine whether advice is good. For production systems, evaluate how to deploy large language models locally if connectivity, privacy, or operating cost makes cloud inference unsuitable.
Add adversarial and high-risk tests
A serious benchmark deliberately includes cases that expose model weaknesses. Test contradictory information, missing units, unrealistic yields, ambiguous disease descriptions, and questions that mix Bengali with English product names. Include prompts asking for an exact chemical dose without pond volume, or asking whether a dead fish indicates a specific disease.
The expected behaviour is often not a direct answer. A safe model should request essential measurements, explain immediate low-risk steps, recommend contacting a qualified fisheries professional, and clearly state what it cannot diagnose. Test whether the model refuses unsafe instructions without becoming useless. Also check for fabricated citations, invented government schemes, and claims that a treatment is “guaranteed.”
If the product accepts photographs, evaluate image quality thresholds, species confusion, symptom overlap, and refusal behaviour. A vision-language benchmark should measure whether the model asks for better images or water tests rather than overdiagnosing from one photograph. Teams can use ideas from evaluating vision models for video understanding when monitoring pond videos or aerator operation.
Measure usefulness with farmers, not only experts
Offline scores are necessary but insufficient. Run a controlled pilot with farmers and extension workers using realistic scenarios. Measure:
- Correct completion of the recommended next step.
- Time taken to reach a useful answer.
- Number of clarification turns before a safe recommendation.
- Farmer comprehension, confidence, and ability to explain the advice back.
- Adoption rate and reasons for rejection.
- Escalations to human experts and whether those escalations were appropriate.
- Changes in mortality, feed waste, water-quality incidents, or cost—without claiming causation from a small pilot.
Test on low-cost Android devices and weak networks. Record latency, message failure, speech recognition errors, and performance when the user sends short voice notes. A technically accurate answer that arrives late, uses unfamiliar terminology, or requires continuous connectivity may fail in the field.
Report results transparently
Publish results by task, risk tier, language form, and user group. Report confidence intervals where sample sizes permit, plus the number of unsafe outputs—not only the average score. Include baseline comparisons: a human-written knowledge base, a general-purpose model, and the proposed Bengali system. Record model version, prompts, retrieval sources, temperature, translation steps, and evaluation date so results remain reproducible.
Create a release gate before deployment. For example, require zero critical safety failures in the locked high-risk set, a minimum expert score for each core task, and acceptable performance under noisy user input. Re-run the benchmark after model updates, prompt changes, new retrieval documents, speech-model changes, or a shift in target geography. Small language models may be attractive for cost and latency; compare them fairly against larger systems using the same context, not just parameter count. Practical deployment patterns can be informed by open-source small language models for Hindi, while keeping Bengali quality as the acceptance criterion.
A practical 30-day benchmarking plan
- Days 1–5: Define tasks, risk classes, sources, and success thresholds with fisheries experts.
- Days 6–12: Gather and anonymise Bengali queries; create metadata and reference answers.
- Days 13–17: Build the locked test set, adversarial cases, and scoring rubric.
- Days 18–22: Run model, retrieval, translation, and speech-interface evaluations.
- Days 23–26: Conduct expert review and a small farmer usability study.
- Days 27–30: Analyse failures, set release gates, document limitations, and schedule regression tests.
The goal is not to identify the model with the most fluent Bengali. It is to identify the system that helps farmers make better decisions while recognising uncertainty and escalating risk. A benchmark built around real ponds, real language, and measurable safety gives Indian AI teams a defensible path from prototype to responsible deployment.