0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to automate news summaries using python

How to Automate News Summaries Using Python

  1. aigi

    News moves faster than most teams can read it. A Python pipeline can collect articles from RSS feeds and licensed APIs, remove duplicate coverage, summarise the full text, and deliver a useful briefing to email, Slack, or a dashboard. The important work is not calling a language model once; it is building a workflow that handles unreliable feeds, paywalls, repeated wire stories, factual risk, and changing model costs.

    This guide explains how to automate news summaries using Python in a way that is practical for Indian builders, research teams, founders, and media products. It also shows where automation should stop and human review should begin.

    Define the output before writing code

    Start with a clear summary contract. Decide:

    • Audience: a founder, analyst, newsroom editor, sales team, or public reader.
    • Coverage: India, a state, an industry, or global developments affecting India.
    • Frequency: hourly alerts, a morning digest, or a weekly brief.
    • Length: one sentence, three bullets, or a structured briefing.
    • Required fields: headline, publisher, publication time, source link, summary, entities, and confidence notes.

    A useful news summary should answer what happened, who is affected, where it happened, and what may happen next. Avoid asking a model to “summarise this” without specifying these requirements. For developer and technical audiences, a personalized AI news feed for programmers can apply the same design principles with topic preferences and ranking.

    Collect articles responsibly

    Use RSS feeds, official publisher APIs, or licensed aggregators wherever possible. Scraping arbitrary pages can violate terms of service, break when layouts change, and produce incomplete text. Respect robots.txt, rate limits, copyright restrictions, and publisher attribution requirements. Store the canonical URL and publication timestamp for every item.

    A basic RSS collector can begin with feedparser:

    pip install feedparser requests beautifulsoup4 trafilatura transformers torch scikit-learn pandas python-dotenv
    import feedparser
    from datetime import datetime, timezone
    
    
    def collect_feed(feed_url, source_name):
        feed = feedparser.parse(feed_url)
        rows = []
        for entry in feed.entries:
            rows.append({
                "source": source_name,
                "title": entry.get("title", "").strip(),
                "url": entry.get("link"),
                "published": entry.get("published", ""),
                "collected_at": datetime.now(timezone.utc).isoformat(),
            })
        return rows

    For Indian coverage, keep sources separated by language and geography. English feeds alone will miss important reporting in Hindi, Tamil, Bengali, Marathi, and other languages. If you summarise translated material, retain the original URL and language so readers can verify the report.

    Extract and clean the article text

    Feed descriptions are often too short for a reliable summary. Fetch the article page only when permitted, then extract the main text while removing navigation, cookie banners, advertisements, and repeated captions. trafilatura is a useful starting point:

    import trafilatura
    import requests
    
    
    def fetch_article(url):
        response = requests.get(
            url,
            timeout=15,
            headers={"User-Agent": "NewsSummaryBot/1.0 (+contact@example.com)"},
        )
        response.raise_for_status()
        text = trafilatura.extract(response.text, include_comments=False)
        return text or ""

    Add safeguards for empty pages, paywalls, very short text, and unusually large documents. Do not send private or restricted content to an external model without checking your data-processing obligations. Keep the original text only as long as your retention policy allows.

    Remove duplicate and near-duplicate coverage

    The same press release or wire report may appear across dozens of websites. Summarising every copy wastes money and creates a misleading sense of corroboration. First normalise URLs and compare titles; then use similarity matching for near-duplicates.

    A simple baseline uses TF-IDF and cosine similarity:

    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.metrics.pairwise import cosine_similarity
    
    
    def duplicate_groups(titles, threshold=0.82):
        matrix = TfidfVectorizer(stop_words="english").fit_transform(titles)
        scores = cosine_similarity(matrix)
        return [(i, j) for i in range(len(titles)) for j in range(i + 1, len(titles))
                if scores[i, j] >= threshold]

    For production systems, combine title similarity, named entities, publication time, and article embeddings. Select a primary source rather than presenting duplicates as independent confirmation.

    Choose extractive or abstractive summarisation

    Extractive summarisation selects sentences from the article. It is easier to audit and less likely to invent details, making it appropriate for legal, policy, and high-risk alerts. Abstractive summarisation rewrites the story in new language and usually reads better, but it can introduce unsupported claims, wrong numbers, or confused attribution.

    Modern transformer models can handle abstractive summaries, but model choice depends on language, hardware, latency, and article length. A hosted API may be simplest for an early prototype; a self-hosted model can improve data control and predictable costs at scale. Test English and Indian-language coverage separately rather than assuming an English benchmark applies to all feeds.

    from transformers import pipeline
    
    summarizer = pipeline("summarization", model="facebook/bart-large-cnn")
    
    
    def summarise(text):
        text = text[:12000]
        result = summarizer(
            text,
            max_length=120,
            min_length=45,
            do_sample=False,
        )
        return result[0]["summary_text"]

    Long articles should be split into coherent chunks, summarised individually, and then combined in a second pass. Always preserve source links and label model-generated text. If the use case involves customer, employee, or regulated information, review the approach alongside your AI legal compliance automation plan in India.

    Use a structured prompt and validation layer

    A strong prompt asks for a fixed schema rather than free-form prose:

    Return valid JSON with these fields:
    summary: 2-3 sentences
    key_facts: 3 bullets containing only facts stated in the article
    entities: people, organisations, places
    why_it_matters: one cautious sentence
    uncertainties: missing or conflicting information

    Validate the JSON, enforce length limits, and reject outputs containing unsupported numbers or quotations. A second pass can compare each key fact with the source text, but it is not a substitute for editorial review. For sensitive topics such as elections, public health, markets, conflict, or allegations, route summaries to a human before publication.

    Automate delivery and operations

    Store articles and summaries in PostgreSQL, SQLite, or object storage with fields for source, URL, hash, language, model version, prompt version, and review status. Schedule collection with cron, GitHub Actions, Airflow, or a queue such as Celery. Add retries with exponential backoff, API-key management through environment variables, and alerts for feed failures.

    A practical daily workflow is:

    • Collect feeds every 15–60 minutes.
    • Canonicalise URLs and discard duplicates.
    • Extract permitted article text.
    • Summarise only new, sufficiently long articles.
    • Run validation and confidence checks.
    • Send a digest with source links and timestamps.
    • Log failures, latency, token usage, and reviewer corrections.

    For teams that already automate other information workflows, the same monitoring discipline used in automated lead generation for Indian B2B startups applies: measure quality and business outcomes, not just the number of items processed.

    Measure quality and control risk

    Create a small evaluation set of articles across business, politics, technology, regional language, and breaking news. Have reviewers score factuality, completeness, readability, attribution, and usefulness. Track recurring errors by source and topic.

    Do not present a generated summary as independent reporting. Keep the original headline, publisher, date, and link visible. Mark corrections, maintain an audit trail, and give readers a clear way to report an error. For an internal briefing, a cautious “the article reports” is often safer than stating an allegation as fact.

    A practical 2026 architecture

    For a small product, use RSS or licensed APIs, Python workers, a relational database, one summarisation model, and email or Slack delivery. Add a queue and observability when volume grows. Use model routing: a cheaper extractive method for routine items, a stronger model for important stories, and human review for high-risk categories.

    The best system is not the one that produces the shortest text. It is the one that helps readers understand verified developments quickly, shows where every claim came from, and fails safely when the source or model is uncertain.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.