Urdu poetry is a demanding test for generative AI. Meaning often depends on metaphor, register, metre, cultural memory, and the relationship between Urdu and scripts such as Nastaliq and Devanagari. A benchmark that measures only word overlap or translation accuracy will miss much of what makes a ghazal, nazm, marsiya, or prose passage successful.
This guide explains how to use AI agents to research Urdu poetry and literature for generative AI benchmarks. The focus is not on asking a chatbot for literary summaries. It is on building a traceable research workflow in which agents find and classify sources, human experts validate interpretation, and evaluation reflects both linguistic quality and literary responsibility.
Start with a precise benchmark question
Define the capability you want to measure before collecting texts. A useful benchmark may test one or more of the following:
- Comprehension: Can a model identify themes, speakers, references, irony, and emotional shifts?
- Retrieval: Can it locate a relevant couplet, author, period, or critical interpretation from a controlled corpus?
- Generation: Can it produce an original passage in an explicitly requested form without imitating a living poet?
- Translation and transliteration: Can it preserve meaning between Urdu script, Roman Urdu, Hindi, and English?
- Cultural reasoning: Can it explain allusions, idioms, religious references, and historical context without inventing facts?
- Editorial assistance: Can it annotate a text while separating evidence from interpretation?
Write each task as an input, expected output, scoring rubric, and evidence requirement. For example, “identify the dominant theme” is too broad. A stronger task provides a passage, asks for a concise interpretation, and requires the model to cite the words or images supporting its answer.
Design a multi-agent research workflow
A reliable system should divide work among specialised agents rather than use one general-purpose prompt. This is similar to building distributed systems with AI agents, where each component has a narrow responsibility and outputs are logged for review.
A practical Urdu-literature workflow includes:
1. Discovery agent: Finds candidate books, archives, catalogues, journal articles, and authoritative author records.
2. Metadata agent: Captures author, title, genre, publication date, edition, script, language variety, and source URL.
3. OCR and transcription agent: Converts scans into machine-readable text and marks uncertain characters rather than silently guessing.
4. Text-normalisation agent: Creates separate versions for original script, normalised Urdu, transliteration, and translation. Never overwrite the source text.
5. Literary-analysis agent: Proposes themes, forms, rhetorical devices, and contextual notes with quoted evidence.
6. Verification agent: Checks citations, names, dates, quotations, and claims against primary or scholarly sources.
7. Benchmark-builder agent: Converts validated material into tasks, test splits, metadata, and scoring instructions.
Use a shared record format such as JSON. Store the original passage, page or stanza location, transformation history, agent prompt, model version, confidence, and reviewer decision. This makes errors reproducible and enables later audits.
Build a lawful, representative corpus
Do not treat the largest online collection as the best dataset. Urdu literature is distributed across libraries, publishers, archives, academic projects, personal collections, and public-domain repositories. Confirm usage rights for every source, especially recent poetry and commercially hosted websites.
Sample across dimensions that affect model behaviour:
- Classical, modern, and contemporary periods
- Ghazal, nazm, qasida, marsiya, rubai, drama, fiction, criticism, and essays
- Major and less represented authors, including women writers and regional voices
- Urdu script, transliteration, bilingual editions, and noisy OCR
- Literary, colloquial, religious, political, and academic registers
- India-based and wider South Asian contexts
Keep a provenance table with licence status, acquisition method, source quality, and permitted benchmark use. For copyrighted works, consider using short evaluation excerpts, licensed access, synthetic prompts, or human-authored reference answers instead of redistributing the full text.
For Indian teams, document whether a source reflects Deccan, North Indian, or broader South Asian usage. Also record script conversion choices: a transliteration scheme should be consistent and published with the benchmark, while human-readable spellings may need to remain as a separate field.
Make Urdu-specific preprocessing explicit
Urdu OCR and tokenisation require more care than simply applying an English NLP pipeline. Nastaliq layout, joining characters, diacritics, punctuation, spelling variants, spacing, and scan quality can all distort results. Run automated checks for:
- Character loss and substitution after OCR
- Incorrect joining or splitting of words
- Confusion between visually similar characters
- Verse-line and couplet boundaries
- Repeated headers, footers, and page numbers
- Mixed Urdu, Arabic, Persian, Hindi, and English vocabulary
- Transliteration inconsistencies
Have reviewers inspect a stratified sample, including difficult scans and poetry with diacritics. Preserve both raw and corrected text, and report correction rates. If an agent is uncertain, it should flag the line for review instead of producing a polished but false transcription.
Create evaluation tasks that measure literary quality
A strong benchmark combines automatic checks with expert and reader judgements. Useful task families include:
- Form recognition: identify radif, qafiya, metre where appropriate, stanza structure, and genre conventions.
- Meaning preservation: compare summaries or translations for semantic completeness, not just n-gram overlap.
- Allusion resolution: explain a cultural or historical reference and provide a verifiable source.
- Close reading: identify imagery, voice, ambiguity, and rhetorical movement using textual evidence.
- Controlled generation: follow constraints such as topic, register, length, or form without copying reference poems.
- Safety and respect: detect fabricated quotations, sectarian stereotyping, sexualisation, or confident misattribution.
Use separate scores for factuality, relevance, Urdu fluency, literary sensitivity, instruction following, and originality. Ask at least two trained reviewers to rate a sample independently, then measure agreement and adjudicate disagreements. Include “insufficient evidence” as a valid answer where the passage does not support a conclusion.
Automated judges can assist with triage, but do not let one language model decide literary quality on its own. If you use LLM-as-judge evaluation, freeze the judging prompt, test it against human labels, inspect positional and model-family bias, and publish limitations. Teams deploying agents in production can also learn from how to deploy Llama 3 agents in production, particularly around versioning, observability, and fallback behaviour.
Use retrieval to reduce hallucination
Require research agents to return evidence with every claim. A good output includes the passage, bibliographic source, page or stanza reference, confidence, and a short explanation of why the evidence supports the claim. Retrieval-augmented generation should search a curated corpus first and only then consult broader web sources.
Add adversarial tests: incorrect author names, altered couplets, fabricated publications, ambiguous transliterations, and questions with no answer in the corpus. Score whether the system refuses or qualifies an unsupported claim. This is more useful than rewarding fluent answers that happen to sound scholarly.
Protect authors, communities, and research integrity
Urdu literature carries religious, political, gendered, and historical sensitivities. Avoid presenting an agent’s interpretation as definitive, and do not use living poets’ work to create imitation tasks without permission. Clearly label generated examples and keep them separate from historical texts.
Remove personal data from contemporary correspondence and social-media sources. Obtain consent where required, respect licence restrictions, and provide a takedown or correction process. Have Urdu scholars and community readers review category definitions, offensive-content handling, and benchmark examples. A technically accurate dataset can still be culturally narrow or unfair.
Publish a benchmark card, not just a score
Document corpus composition, licences, scripts, OCR quality, preprocessing, annotation instructions, reviewer backgrounds, model versions, prompts, known gaps, and contamination checks. Report results by genre, period, script, and task—not only as one aggregate number.
For an Indian research team, a transparent benchmark is more valuable than a large opaque dataset. It helps model builders identify whether a system struggles with Nastaliq OCR, transliteration, classical references, contemporary idiom, or literary interpretation. If your broader project includes multilingual deployments, the same provenance discipline is useful when designing LLM-powered voice agents for complex conversations, where language switching and context retention also need explicit evaluation.
A practical launch checklist
Before releasing the first version, confirm that you have:
- A clearly scoped set of Urdu-literature capabilities and failure modes
- Source-level rights and provenance records
- Separate raw, corrected, transliterated, and translated text fields
- Human validation for OCR, metadata, and literary annotations
- Balanced splits that prevent the same poem or quotation entering train and test data
- Rubrics covering fluency, meaning, form, cultural context, and hallucination
- Expert and non-expert reviewers with documented training
- Adversarial and unanswerable cases
- Versioned prompts, agents, models, and evaluation scripts
- A public limitations statement and correction process
AI agents can accelerate discovery and annotation, but they should not replace Urdu scholarship. The strongest generative AI benchmarks pair machine-scale organisation with human-led interpretation, careful licensing, and evaluation that respects the form and cultural history of the literature.