0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build automated web scrapers with selenium

How to Build Automated Web Scrapers with Selenium

  1. aigi

    Selenium is useful when a website depends on JavaScript to render content, requires browser interactions, or exposes data only after a user flow. It is also expensive: every session runs a real browser, consumes memory, and can fail when selectors, network conditions, or page behaviour change. A production scraper therefore needs more than a working script. It needs clear scope, resilient waits, structured output, observability, and responsible access controls.

    This guide explains how to build automated web scrapers with Selenium in Python, with patterns that remain practical for Indian startups, research teams, and AI builders in 2026.

    Choose Selenium only when a browser is necessary

    Start by checking whether the target data is available through a documented API, a server-rendered HTML response, an export, or a public dataset. HTTP clients and parsers are faster and cheaper than browser automation. Selenium is a good fit when you need to:

    • Render content created by React, Vue, Angular, or client-side API calls.
    • Interact with filters, tabs, forms, pagination, or authenticated pages you are authorised to access.
    • Capture browser-visible state, downloads, screenshots, or accessibility information.
    • Reproduce a user journey for testing or controlled data collection.

    Do not treat Selenium as a way to defeat CAPTCHAs, access restricted systems, evade authentication, or circumvent technical controls. Review the site’s terms, robots guidance where relevant, privacy obligations, and applicable Indian law before collecting data. Minimise personal data, respect deletion requests, and document your lawful basis for sensitive use cases.

    For AI products, the extracted content is usually one component of a larger system. Teams building distributed systems with AI agents should define ownership, retries, queues, and data contracts before adding browser workers.

    Set up a reproducible Python environment

    Use a virtual environment and pin dependencies in a lockfile or requirements file. Recent Selenium releases can manage compatible browser drivers automatically, so webdriver-manager is often unnecessary.

    python -m venv .venv
    source .venv/bin/activate        # Windows: .venv\\Scripts\\activate
    pip install selenium pydantic python-dotenv

    A minimal browser factory keeps configuration in one place:

    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    
    def create_driver() -> webdriver.Chrome:
        options = Options()
        options.add_argument("--headless=new")
        options.add_argument("--window-size=1440,1200")
        options.add_argument("--disable-dev-shm-usage")
        return webdriver.Chrome(options=options)

    In a container, allocate adequate shared memory and run the browser with the permissions required by your image. Avoid blindly adding --no-sandbox; understand the security trade-off and use a hardened runtime instead.

    Build a scraper around explicit waits

    The most common Selenium failure is reading the page before the application has finished rendering. Fixed sleeps slow every run and still fail under variable network conditions. Use explicit waits for a meaningful state: an element becoming visible, a loading indicator disappearing, a URL changing, or a result count increasing.

    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    
    def scrape_title(url: str) -> str:
        driver = create_driver()
        try:
            driver.get(url)
            title = WebDriverWait(driver, 15).until(
                EC.visibility_of_element_located((By.CSS_SELECTOR, "h1"))
            )
            return title.text.strip()
        finally:
            driver.quit()

    Prefer stable selectors such as data-testid, accessible labels, semantic landmarks, or well-defined IDs. Avoid brittle selectors based on generated CSS classes and deep DOM paths. Keep locators in a separate module so a front-end redesign does not require editing business logic.

    Extract structured, validated records

    Return records rather than loose strings. Capture the source URL, retrieval time, stable identifier, and parser version so results can be audited and reprocessed.

    from pydantic import BaseModel, HttpUrl
    from datetime import datetime
    
    class Listing(BaseModel):
        title: str
        url: HttpUrl
        source: HttpUrl
        collected_at: datetime

    Use a separate parsing function for each page type. Normalise whitespace, dates, currency, and Indian numbering formats at the boundary. Reject incomplete records or send them to a review queue instead of silently storing bad data. For multilingual products, preserve the original text and language metadata; downstream systems may need Indic-language processing such as the methods covered in this guide to low-resource Indic NLP.

    Handle pagination and infinite scrolling safely

    For ordinary pagination, wait for the next page’s content to replace the previous page before extracting again. For infinite scroll, stop on a clear condition: no new record IDs, a known maximum, a repeated cursor, or a business-defined time limit. Do not rely only on document height.

    from selenium.webdriver.common.keys import Keys
    
    seen = set()
    for _ in range(20):
        cards = driver.find_elements(By.CSS_SELECTOR, "article[data-id]")
        before = len(seen)
        for card in cards:
            item_id = card.get_attribute("data-id")
            if item_id:
                seen.add(item_id)
    
        if len(seen) == before:
            break
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        WebDriverWait(driver, 10).until(
            lambda d: len(d.find_elements(By.CSS_SELECTOR, "article[data-id]")) > len(cards)
            or d.find_elements(By.CSS_SELECTOR, "[data-end-of-results]")
        )

    Add backoff for transient failures, but cap retries and record the reason. A retry policy should distinguish timeouts, stale elements, HTTP-level blocks, invalid data, and genuine end-of-results conditions.

    Treat login, privacy, and consent as product requirements

    Automate only accounts and data you are authorised to use. Store credentials in a secret manager or environment-backed vault, never in source code or scraped output. Do not collect passwords, session cookies, payment data, or unnecessary personal information. Keep session lifetime short and delete browser profiles after a run unless retention is explicitly required.

    For Indian deployments, map data flows before launch. Consider consent, purpose limitation, retention, access controls, and breach handling under the Digital Personal Data Protection Act, 2023, alongside contractual and sector-specific requirements. Mask personal fields in logs and encrypt data at rest and in transit.

    Make the scraper observable and testable

    A production worker should emit structured logs and metrics for:

    • Pages attempted, succeeded, retried, and failed.
    • Median and p95 page duration.
    • Records extracted, rejected, duplicated, and changed from the previous run.
    • Selector failures, browser crashes, and memory usage.
    • Source response or application states that require operator review.

    Use fixture pages or a staging copy for integration tests. Test empty results, slow rendering, missing optional fields, duplicate pages, stale elements, redirects, and partial network failure. Save failure screenshots and HTML snapshots only when they do not contain sensitive information. A schema version and parser version make historical corrections manageable.

    Scale with queues, not uncontrolled browser processes

    Selenium sessions are resource-intensive. Begin with one worker and measure CPU, RAM, browser startup time, and source limits. Then introduce a queue with bounded concurrency, one isolated browser session per job, and a dead-letter queue for repeated failures. Selenium Grid or a containerised browser service can distribute workloads, but scaling out does not justify sending traffic faster than a site can reasonably handle.

    Set per-domain rate limits, honour maintenance windows, and use caching or incremental collection. For larger pipelines, persist raw snapshots separately from cleaned records so parsing can be improved without repeatedly visiting the source. If your end product is a conversational workflow, connect the validated dataset to an agent architecture rather than letting an agent directly control unrestricted scraping; the principles in building generative AI agents are a useful starting point.

    Selenium alternatives and a practical decision rule

    Use an API when one exists and permits your use case. Use requests plus Beautiful Soup or lxml for static pages. Consider Playwright when you need modern browser automation, strong isolation, or multiple browser engines. Selenium remains a sensible choice when your organisation already operates WebDriver infrastructure or needs broad browser compatibility.

    A reliable decision rule is simple: use the least powerful collection method that can legally and accurately obtain the required fields. Browser automation should be the exception for dynamic interaction, not the default for every URL.

    Frequently asked questions

    Is Selenium better than Beautiful Soup?
    No tool is universally better. Beautiful Soup is faster for HTML already present in the response; Selenium is appropriate when browser execution is required.

    Can Selenium bypass CAPTCHAs or anti-bot systems?
    Do not design a system to evade access controls. If a legitimate workflow encounters a CAPTCHA, stop, request an approved integration, or obtain permission from the site owner.

    How should I run Selenium in production?
    Use pinned dependencies, isolated workers, explicit waits, bounded concurrency, secret management, structured logs, alerts, and a clear retention policy.

    Can scraped data be used to train an AI model?
    Possibly, but source terms, copyright, privacy, consent, and downstream licensing matter. Keep provenance and remove data that you cannot lawfully retain or use.

    If you are building an India-focused AI data product, AI Grants India supports founders working on practical infrastructure and applied research. Explore AI Grants India for grant and mentorship information.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.