What this workflow helps you do
A python script for network traffic analysis can turn packet captures into useful evidence: which hosts are communicating, which protocols dominate, when latency or volume changes, and where suspicious behaviour may warrant investigation. Python is especially useful for prototypes, internal observability tools, incident triage, and research projects because engineers can combine packet libraries with data processing and visualisation without maintaining a large platform.
This guide uses Scapy and standard Python patterns to build a small, extensible analyser. It is not a replacement for a network detection and response platform, packet broker, or lawful monitoring process. Capture only traffic you are authorised to inspect, document retention rules, and avoid collecting payloads when headers and flow metadata are sufficient.
Choose the right capture source
Before writing code, decide where the traffic will come from:
- Live interface: useful for short diagnostic sessions on a server, lab machine, or approved sensor.
- PCAP file: safer for repeatable analysis, testing, and sharing sanitised samples with a team.
- Flow records: NetFlow, IPFIX, or VPC flow logs scale better when full packets are unnecessary.
- Application telemetry: logs, traces, and metrics may answer performance questions without inspecting packets.
On Linux, an interface may be named eth0, ens5, or wlan0; on cloud platforms, visibility depends on the provider and network architecture. In India, teams should also align monitoring with organisational policies, contractual obligations, and applicable privacy requirements. Do not assume that access to a network automatically grants permission to inspect its contents.
Set up a reproducible Python environment
Use a virtual environment and pin versions for scripts that will run in production or during incident response:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install scapy pandas matplotlibScapy is convenient for packet capture and protocol dissection. For large PCAPs or formats that depend on Wireshark libraries, PyShark can be useful, but it introduces a tshark system dependency. Keep the first version focused: capture metadata, aggregate it, and write results to CSV or JSON. You can later connect the output to dashboards or a data pipeline, much like the structured stages used in Python scripts for automating data preprocessing.
Live capture often requires administrator privileges or specific capture capabilities. Run with the least privilege possible, and prefer reading a test PCAP when developing.
Capture packets without drowning in output
Printing every packet is useful for a five-minute learning exercise but quickly becomes unusable. Start with a BPF filter and collect only the fields needed for your question:
from collections import Counter
from datetime import datetime, timezone
from scapy.all import IP, TCP, UDP, sniff
protocols = Counter()
flows = Counter()
bytes_by_source = Counter()
def inspect_packet(packet):
if IP not in packet:
return
ip = packet[IP]
protocol = "TCP" if TCP in packet else "UDP" if UDP in packet else str(ip.proto)
src_port = packet[TCP].sport if TCP in packet else packet[UDP].sport if UDP in packet else "-"
dst_port = packet[TCP].dport if TCP in packet else packet[UDP].dport if UDP in packet else "-"
packet_size = len(packet)
protocols[protocol] += 1
flows[(ip.src, ip.dst, src_port, dst_port, protocol)] += 1
bytes_by_source[ip.src] += packet_size
print({
"time_utc": datetime.now(timezone.utc).isoformat(),
"src": ip.src,
"dst": ip.dst,
"protocol": protocol,
"src_port": src_port,
"dst_port": dst_port,
"bytes": packet_size,
})
sniff(iface="eth0", filter="ip", prn=inspect_packet, store=False, count=1000)
print("protocols:", protocols)
print("top_sources_by_bytes:", bytes_by_source.most_common(10))store=False prevents Scapy from retaining every packet in memory. In a real sensor, replace print with a structured logger and write periodic aggregates rather than raw records. If you are analysing a saved capture, use sniff(offline="sample.pcap", ...) and remove the interface argument.
Measure flows, not just individual packets
Packet counts alone can mislead. A single large file transfer and thousands of small requests have different operational implications. Useful flow-level fields include:
- source and destination IP or anonymised identifiers;
- source and destination port;
- protocol and packet count;
- total bytes and capture duration;
- TCP flags, retransmissions, and connection failures;
- DNS query volume and response errors;
- first-seen and last-seen timestamps.
For privacy-aware reporting, hash internal IP addresses with a protected key or aggregate by subnet. Avoid storing payloads, credentials, cookies, or full URLs unless there is a documented investigative need and suitable access control.
A simple next step is to export dictionaries to a Pandas DataFrame, group by destination port or source host, and calculate rates per minute. Plotting bytes and flows over time can reveal backups, software updates, capacity spikes, or unusual overnight activity. For broader Python data workflows, see Python data science automation for Indian startups.
Add practical anomaly checks
Rules should be explainable before they become alerts. Start with thresholds that an operator can validate:
- a host contacts an unusually high number of destinations;
- outbound bytes exceed a baseline for that service;
- a workstation generates repeated connections to uncommon ports;
- DNS failure rates rise sharply;
- TCP SYN packets increase without corresponding completed connections.
A baseline can be a rolling median by hour and day rather than a fixed global number. Label alerts with the evidence behind them: observed value, baseline, time window, and affected hosts. This makes the script useful during an incident instead of producing unexplained warnings. If you later add machine learning, treat it as prioritisation support, not proof of compromise; this is the same discipline needed when building customizable neural network architectures for beginners.
Production hardening checklist
Before deploying beyond a lab, address the operational details that simple examples omit:
- Permissions: grant only the capture capability required; do not run a permanent service as root.
- Performance: use BPF filters,
store=False, bounded queues, and batch writes. - Resilience: handle rotated PCAPs, interface changes, malformed packets, and interrupted captures.
- Observability: expose the script’s own packet rate, dropped packets, queue depth, and processing latency.
- Security: protect output files, rotate logs, encrypt transfers, and redact sensitive fields.
- Testing: replay representative PCAPs and verify results against Wireshark or known flow totals.
- Deployment: package dependencies, pin versions, and document interface names and retention settings.
Do not confuse packet capture with encryption inspection. HTTPS and modern application protocols intentionally limit payload visibility. Often, metadata plus application logs and distributed traces provide a safer and more complete diagnosis. If your team is building an AI-enabled operations tool, keep traffic summaries separate from model prompts and follow the same API hygiene described in integrating LLM APIs in Python web apps.
A sensible path from script to system
Use the script to answer one concrete question first: *Which hosts consumed the most outbound bandwidth during the last hour?* Once that works, add PCAP replay tests, flow aggregation, a small SQLite or Parquet output, and a dashboard. Only then consider streaming ingestion, alert routing, or a distributed collector.
For Indian startups, campuses, SaaS teams, and public-interest projects, this incremental approach keeps infrastructure costs and data exposure under control. It also creates a clear grant or product milestone: reproducible capture, validated metrics, privacy safeguards, and a measured operational outcome.
FAQs
Can this script inspect encrypted traffic?
It can analyse visible metadata such as IPs, ports, packet sizes, timing, and TLS handshake information. It cannot legitimately recover encrypted payloads without authorised keys or endpoint instrumentation.
Should I use Scapy or PyShark?
Use Scapy for a lightweight Python-native prototype and PyShark when you need Wireshark’s tshark dissectors or existing capture workflows. Benchmark both with your real PCAP sizes.
How do I analyse high-volume traffic?
Filter at capture time, aggregate into flows, sample where appropriate, and move from packets to NetFlow/IPFIX or platform flow logs when full packet visibility is unnecessary.
Is packet capture legal?
Permission and context matter. Capture only networks and devices your organisation is authorised to monitor, minimise personal data, restrict access, and obtain legal or compliance guidance for the intended deployment.