Short answer
For most new projects in 2026, Qwen2.5-3B-Instruct is a strong general-purpose starting point when you need affordable abstractive summaries, while BART-base remains a dependable choice for narrow, English-only summarization with predictable inference. For Indian-language workflows, test Qwen2.5-3B-Instruct alongside an Indic-focused model and validate performance on your own Hindi, Tamil, Bengali, or mixed-language data. If the model must run on a phone, edge device, or low-cost server, a quantised 1B–3B instruct model is usually more practical than a larger checkpoint.
There is no universally “best” small language model. The right choice depends on whether you need extractive or abstractive summaries, how long the source documents are, whether the output may contain new claims, and where inference will run.
What counts as a small language model?
For this guide, “small” means roughly under 7 billion parameters, with a practical focus on models that can run on a single modest GPU, CPU server, or edge device after quantisation. Parameter count is only one part of deployment cost. Context length, quantisation format, tokenizer efficiency, and generation speed can matter just as much.
A 3B model can be a better production choice than a 1B model if it avoids manual review. Conversely, a specialised encoder model may beat a generative model for sentence selection while using a fraction of the memory.
Best options by use case
Qwen2.5-3B-Instruct: best general starting point
Qwen2.5-3B-Instruct offers a useful balance of instruction following, multilingual capability, quality, and hardware requirements. It can produce concise summaries, follow formatting instructions, and handle common business documents. It is a sensible baseline for meeting notes, support tickets, research digests, and internal reports.
Use it when you need one model for several languages or summary formats. Test it carefully on code-mixed Indian text, because outputs may become verbose or silently omit details when the source contains Hindi-English or regional-language switching.
Phi-3.5-mini or Phi-4-mini: strong on structured, shorter documents
Microsoft’s Phi family is useful when summaries must follow a strict schema, such as key_points, risks, actions, and open_questions. These models can be effective for short to medium documents and are attractive for developers who want local inference without a large GPU.
They are not automatically reliable on very long documents. Use chunking and a second consolidation pass rather than placing an entire legal file or transcript into one prompt.
BART-base or DistilBART: predictable English summarization
BART-based checkpoints remain valuable for English abstractive summarization, especially when the task is stable and training data resembles the target content. DistilBART reduces latency and memory use, but quality may decline on specialised documents or unusual writing styles.
These encoder-decoder models are often easier to fine-tune than general chat models. Choose them for a fixed pipeline such as news, customer feedback, or support-ticket summaries—not for a multilingual assistant expected to handle arbitrary instructions.
T5-small or FLAN-T5-small: compact and adaptable
T5 and FLAN-T5 convert summarization into a text-to-text task. Their small variants are lightweight and straightforward to fine-tune, making them useful for experiments, offline prototypes, and narrow domains. They may need more task-specific tuning than newer instruction-tuned models to produce consistently natural summaries.
Extractive models: best when factuality is non-negotiable
If a summary must not introduce wording that is absent from the source, use an extractive approach. A sentence-ranking model based on BERT, DistilBERT, or a lightweight classifier can select the most important sentences, then apply a length and redundancy filter.
This approach is particularly appropriate for compliance, claims processing, government records, and early-stage healthcare tooling. It is less polished than abstractive generation, but reviewers can trace every sentence to the source. For Indian-language deployment, pair model testing with the principles in this low-resource Indic NLP builder’s guide.
A practical decision framework
Select the model using these questions:
- Need exact source wording? Start with extractive ranking or a hybrid pipeline.
- Need natural, short prose? Test Qwen2.5-3B-Instruct, Phi-mini, and BART-based models.
- Need Hindi or other Indian languages? Benchmark multilingual and Indic checkpoints on native text, not translated English samples. A dedicated open-source small language model for Hindi may outperform a general model for vocabulary and syntax.
- Need on-device inference? Prefer a 1B–3B model in 4-bit GGUF or another supported quantisation format, and measure time-to-first-token as well as tokens per second.
- Need long documents? Use hierarchical summarization: split the source, summarize each chunk, then summarize the summaries. Do not assume a large advertised context window guarantees accurate recall.
- Need sensitive data to stay in India or inside your network? Choose a self-hosted model and document retention, logging, and access controls before launching.
How to build a reliable summarization pipeline
A model alone is rarely enough. A production pipeline should:
1. Clean and segment the source. Preserve headings, speaker labels, page numbers, and dates. Avoid splitting in the middle of a sentence or table row.
2. Set a summary contract. Specify audience, target length, language, tone, required fields, and whether the model may infer anything.
3. Summarize in stages. For long documents, use map-reduce or hierarchical summarization with overlap between chunks.
4. Add factuality checks. Compare names, amounts, dates, and action items against the source. For high-risk use cases, require citations or highlighted source spans.
5. Retain uncertainty. Instruct the model to say “not stated” rather than fill gaps. This matters for legal, financial, and clinical content.
6. Keep a human review path. Route low-confidence, unusually short, or contradictory outputs to a reviewer.
If your product must work on low-cost Android devices or unreliable networks, review AI model optimization for mobile devices before selecting a checkpoint. Quantisation, batching, and prompt length often determine unit economics more than the model licence.
How to evaluate models properly
ROUGE-1, ROUGE-2, and ROUGE-L are useful for comparing systems against reference summaries, but they reward word overlap and do not reliably detect fabricated facts. Add BERTScore or semantic similarity, then conduct human review using a rubric covering:
- factual accuracy and unsupported claims;
- coverage of important points;
- concision and readability;
- faithfulness to numbers, names, dates, and negation;
- language quality for each target language;
- consistency across repeated runs.
Create a test set of at least 100–300 representative documents if possible. Include short and long inputs, noisy transcripts, tables, code-mixed text, spelling variation, and difficult edge cases. Report quality separately by language and document type; an aggregate score can hide poor Hindi or Marathi performance.
Measure operational metrics too: peak RAM or VRAM, latency, throughput, cost per document, failure rate, and review time. A slightly weaker model that is twice as fast and needs half the memory may be the better business decision.
Recommendation
Start with Qwen2.5-3B-Instruct for a multilingual generative baseline, BART-base or DistilBART for a stable English-only workflow, and an extractive DistilBERT-style pipeline where traceability matters most. Benchmark at least one Indic-focused alternative for Indian-language content, and do not deploy based on public leaderboard results alone.
For founders building summarization into a product, the strongest architecture is often hybrid: extract important source spans, generate a concise summary from those spans, and expose citations for review. That design improves trust while keeping inference affordable. If your application serves regional-language users, also study open-source vision-language models for Indian languages when documents contain scans, images, or tables.
FAQ
Is a smaller model always cheaper?
No. A smaller model generally uses less memory, but poor summaries can increase human-review and reprocessing costs. Compare total cost per accepted summary, not inference cost alone.
Can a small model summarize a two-hour meeting transcript?
Yes, but use diarization, cleanup, chunking, and hierarchical summarization. Feed the final model structured chunk summaries rather than the raw transcript alone.
Should I fine-tune the model?
Fine-tune only after establishing a strong prompt and evaluation baseline. Fine-tuning is worthwhile when you have consistent examples, a narrow domain, and a stable output format. For rapidly changing content, retrieval and better preprocessing may deliver more value.
What about privacy?
For confidential Indian business, legal, or health data, prefer self-hosting or a provider with appropriate contractual controls. Remove unnecessary personal data, restrict logs, encrypt stored documents, and define deletion policies before production.
Apply for AI Grants India
If you are building an Indian-language summarization product, an offline document assistant, or a sector-specific NLP system, explore AI Grants India for potential funding and support.