PDF summarisation is not an LLM prompt wrapped around a file upload. A dependable system must recover document structure, handle scanned pages and tables, control context, cite source pages, and keep sensitive files secure. For Indian startups, research teams, and internal operations tools, these engineering decisions usually matter more than the choice between two similar language models.
This guide presents a practical architecture for building AI powered PDF summarizers with Python in 2026. It covers a local prototype, production safeguards, multilingual documents, cost controls, and a testing approach that catches confident but unsupported summaries.
Start with the output contract
Before choosing libraries, define what the summarizer must produce. “A summary” is too vague to test or operate. A useful contract might include:
- A 100-word executive overview.
- Five key findings with page references.
- Decisions, risks, deadlines, and named entities.
- A separate section for figures and tables.
- An uncertainty note when text extraction or OCR quality is poor.
- Structured JSON for downstream applications, plus readable Markdown for users.
For legal, financial, health, or government documents, require every material claim to carry a page number or source span. This turns the product from a black-box generator into an auditable document-intelligence workflow. Teams already building LLM APIs in Python web apps can expose this contract through a background job rather than blocking an HTTP request.
A production pipeline for PDFs
Treat each upload as a staged job with observable outputs:
1. Validate: Check file type, size, page count, encryption status, and malware scanning results.
2. Classify: Detect whether pages contain native text, images, tables, or mixed content.
3. Extract: Use a layout-aware parser and preserve page, block, heading, and table boundaries.
4. OCR selectively: Run OCR only on image-only pages, retaining confidence scores and the original page image.
5. Normalise: Remove repeated headers and footers, repair hyphenation, and preserve headings, lists, citations, and reading order.
6. Segment: Create semantically coherent chunks with metadata such as page number, section, language, and extraction method.
7. Summarise: Apply a strategy suited to document length and the required level of traceability.
8. Verify and deliver: Check citations, schema validity, unsupported claims, and sensitive-data handling before returning results.
Store intermediate artefacts. If a user reports an incorrect summary, you need to inspect the extracted text and chunk boundaries—not simply rerun the model.
Select an extraction stack by document type
No single Python library handles every PDF well. Build a fallback strategy and measure extraction quality on your own corpus.
- PyMuPDF: Fast native-text extraction with useful page and block metadata. A strong default for born-digital reports.
- pdfplumber: Helpful for character positions, lines, and table-oriented workflows.
- pypdf: Suitable for straightforward text and metadata operations, but less capable for complex layout reconstruction.
- OCRmyPDF with Tesseract: A practical open-source route for searchable scans; evaluate Indic-language accuracy separately.
- PaddleOCR or a managed document-AI service: Consider these for difficult scans, handwriting, or richer layout detection.
- Markdown conversion tools: Useful when headings, lists, tables, and figures need to be presented to an LLM in a consistent representation.
Keep extracted content tied to its page. Do not flatten a two-column paper into one string without testing reading order. Tables should become structured Markdown, CSV, or a typed object; sending visually scrambled table text to a model invites numerical errors.
Chunking and summarisation strategies
Chunk by meaning, not merely by character count. Prefer a heading-aware splitter, then enforce token limits as a safety boundary. Include modest overlap only when it prevents sentences or definitions from being separated. Every chunk should carry its source page and section metadata.
Use the following strategy decision tree:
- Short documents: A single pass can work if the model context is comfortably larger than the input and the output is tightly constrained.
- Long reports: Map-reduce summarises chunks in parallel, then combines the intermediate summaries. It is scalable but can lose cross-section relationships.
- Narrative or chronological documents: Refine processing preserves continuity but is slower and order-dependent.
- Question-specific synthesis: Retrieval-augmented generation fetches relevant sections before answering. It is better for targeted questions than for a complete overview.
- High-stakes output: Use hierarchical summarisation: section summaries, document-level synthesis, then a verification pass against source chunks.
RAG should not automatically replace summarisation. A vector index is valuable for follow-up questions and selective evidence retrieval, while a complete summary still requires coverage across the document. For a deeper architecture, see building high-performance AI applications with open-source tools.
A safer Python implementation pattern
The following sketch emphasises page-aware records rather than a one-line chain. Exact APIs vary by provider, but the separation of concerns is durable:
import fitz
def extract_pages(path: str) -> list[dict]:
doc = fitz.open(path)
pages = []
for number, page in enumerate(doc, start=1):
text = page.get_text("text").strip()
pages.append({"page": number, "text": text, "needs_ocr": len(text) < 40})
return pages
def make_chunks(pages: list[dict], max_chars: int = 7000) -> list[dict]:
chunks, buffer, start_page = [], [], None
size = 0
for item in pages:
if start_page is None:
start_page = item["page"]
if size + len(item["text"]) > max_chars and buffer:
chunks.append({"text": "\n".join(buffer), "start_page": start_page})
buffer, size, start_page = [], 0, item["page"]
buffer.append(item["text"])
size += len(item["text"])
if buffer:
chunks.append({"text": "\n".join(buffer), "start_page": start_page})
return chunksIn production, replace character limits with model-token counting, add overlap at section boundaries, and route needs_ocr pages through an OCR worker. The model prompt should specify the output schema, prohibit invented facts, request page citations, and instruct the model to return “not found” when evidence is absent. Validate the response with Pydantic or JSON Schema before storing it.
Multilingual and Indian document considerations
Indian business and public-sector PDFs frequently combine English with Hindi, Tamil, Bengali, Marathi, Telugu, or mixed-script names. Language detection should happen per page or block, not only once per file. Preserve the original text alongside translations so users can audit names, numbers, and legal wording.
Use multilingual embeddings when indexing mixed-language content, and test Indic OCR on representative scans rather than relying on English benchmarks. Do not silently translate a summary: label the output language and offer the source excerpt for verification. A multilingual product can borrow design principles from building multilingual chatbots for Indian startups, especially around fallback behaviour and language-aware evaluation.
Cost, latency, and deployment
Run extraction and OCR asynchronously. A queue such as Celery, Dramatiq, or a managed task system prevents large uploads from exhausting web workers. For elastic inference, serverless GPU platforms can help teams validate demand; the trade-offs are discussed in building serverless AI apps with Modal.
Control spend with a tiered pipeline:
- Cache extraction and intermediate summaries using a document hash.
- Use a smaller model for chunk summaries and a stronger model for synthesis or verification.
- Skip unchanged pages when users upload revised versions.
- Cap page counts and token budgets, with transparent upgrade paths.
- Batch independent chunk requests and retry only failed calls.
- Keep sensitive documents in an Indian region when contractual or sectoral requirements demand it.
Encrypt files in transit and at rest, set deletion windows, isolate tenants, and redact secrets before external inference where feasible. Log model version, prompt version, parser version, chunk IDs, and latency without logging raw document content by default.
Evaluation: measure faithfulness, not just fluency
Create a representative test set: scans, two-column papers, tables, contracts, multilingual reports, and PDFs with repeated headers. Have reviewers mark factual coverage, citation correctness, omission of critical details, numerical accuracy, and readability.
Automated checks can flag missing citations, invalid JSON, unsupported page numbers, excessive overlap with boilerplate, and disagreement between extracted figures and generated figures. Human review remains essential for high-stakes domains. Track extraction failures separately from model failures; otherwise, improving the prompt may appear to fix a parser problem.
A practical launch threshold is not “the summary sounds good.” It is: critical claims are traceable, known failure modes are disclosed, and the system degrades safely when extraction quality is low. For teams experimenting with local models, building open-source AI tools for Indian developers offers a useful direction for reducing vendor dependence while retaining control over data.
Recommended build sequence
Start with PyMuPDF, page-aware chunks, a constrained JSON schema, and citations. Add OCR for scanned pages, then table extraction and multilingual evaluation. Introduce embeddings and RAG only when follow-up retrieval is a real product requirement. Finally, add queues, caching, observability, access controls, and deletion policies before onboarding sensitive enterprise data.
This sequence keeps the first prototype small without locking you into an unsafe architecture. The winning PDF summarizer is not the one with the longest prompt; it is the one users can verify, operators can debug, and Indian organisations can trust with their documents.