0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · Graph Neural Networks for Detecting UPI Mule Accounts and Fraud Rings

Graph Neural Networks for Detecting UPI Mule Accounts

  1. aigi

    UPI has made digital payments fast, interoperable, and widely accessible across India—but the same network effects can be exploited by fraud rings. Mule accounts receive, relay, split, or withdraw illicit funds, often making each individual transaction appear ordinary. Detecting them requires more than checking transaction amount, device reputation, or velocity in isolation. Graph neural networks (GNNs) model the relationships between accounts, UPI IDs, devices, merchants, bank accounts, phone numbers, and transactions to uncover coordinated behaviour.

    This article explains how to design a graph-based detection system for UPI mule accounts and fraud rings, including data modelling, GNN architectures, features, training strategies, evaluation, deployment, and India-specific compliance considerations.

    What Are UPI Mule Accounts and Fraud Rings?

    A mule account is a bank or payment account used to receive, move, conceal, or cash out funds originating from fraud or other illegal activity. The account holder may be knowingly involved, recruited through social media, or deceived into sharing credentials and access.

    A fraud ring is a coordinated group of accounts and entities that work together. Typical roles include:

    • Collection accounts: Receive funds from victims.
    • Layering accounts: Move money through multiple hops to obscure its source.
    • Aggregation accounts: Consolidate funds before withdrawal or transfer.
    • Cash-out accounts: Convert digital funds into cash or valuable goods.
    • Recruiter or controller entities: Coordinate multiple mule accounts.
    • Merchant fronts: Process apparently legitimate payments to disguise proceeds.

    The central detection challenge is that a mule account may not have extreme values on its own. Its risk becomes visible through connections: shared devices, repeated counterparties, synchronized transfers, common beneficiary patterns, or rapid movement through a network.

    Why Graph Neural Networks Fit UPI Fraud Detection

    Traditional fraud systems often represent a transaction as a row with columns such as amount, timestamp, sender, receiver, device, and location. This supports rules and tabular machine learning, but it loses much of the structure between records.

    A graph represents entities as nodes and interactions as edges. For UPI fraud detection, nodes may include:

    • Bank accounts and UPI IDs
    • Mobile numbers and email addresses
    • Devices, SIM cards, IP addresses, and app installations
    • Merchants, QR codes, and virtual payment addresses
    • PAN-linked or business entities where legally available
    • ATMs, cash-out points, and beneficiary accounts

    Edges can represent payments, logins, device associations, beneficiary additions, refunds, collect requests, or shared identifiers. Each edge can carry features such as amount, timestamp, channel, location, transaction status, and payment purpose.

    GNNs learn representations by aggregating information from neighbouring nodes. An account that looks normal in isolation may receive a high-risk embedding because it is connected to several confirmed fraud accounts, a suspicious device cluster, and a rapid multi-hop fund flow.

    Designing a UPI Fraud Graph

    A production system should avoid treating the entire payment ecosystem as one simplistic graph. A heterogeneous, temporal graph is usually more expressive.

    Heterogeneous graph schema

    Use separate node and edge types rather than flattening everything into account-to-account links:

    | Element | Examples | Useful attributes |
    |---|---|---|
    | Account node | Bank account, wallet account | Age, KYC status, segment, risk history |
    | Payment node | UPI transaction | Amount, timestamp, status, purpose, channel |
    | Device node | Phone, app installation | Device age, integrity signals, account count |
    | Identity node | Mobile, email, PAN-linked entity | Verification status, reuse count |
    | Merchant node | QR, store, online seller | Category, settlement pattern, tenure |
    | Edge | Sends, receives, logs in, shares | Frequency, recency, amount statistics |

    A transaction can be modelled as an edge with features or as its own node connected to sender, receiver, device, and merchant. The latter is useful when a payment has many attributes and participates in multiple analytical relationships.

    Temporal representation

    Fraud patterns change quickly. A static graph may incorrectly connect accounts that shared a device months apart or may leak future information into historical training data. Store event time and build time-aware neighbourhoods.

    Important temporal features include:

    • Number of new counterparties in the past 10 minutes, 1 hour, 24 hours, and 7 days
    • Time from account creation to first incoming payment
    • Time between receipt and onward transfer
    • Burstiness of incoming and outgoing transactions
    • Reuse of a device or beneficiary after an account is flagged
    • Change in transaction geography or merchant category

    For near-real-time decisions, neighbourhood aggregation should use only data available before the transaction under review.

    High-Value Features for Mule Account Detection

    GNNs learn from structure, but careful feature engineering improves accuracy, interpretability, and operational usefulness.

    Account and transaction behaviour

    • Incoming-to-outgoing amount ratio
    • Number of unique senders and receivers
    • Median holding time for received funds
    • Percentage of funds forwarded within a short window
    • Failed-to-successful transaction ratio
    • New beneficiary creation followed by immediate transfer
    • Unusual activity relative to the account’s historical baseline

    Network and centrality features

    • In-degree and out-degree
    • Weighted degree by transaction value
    • Two-hop exposure to confirmed fraud entities
    • PageRank or authority score within suspicious components
    • Community risk score
    • Number of weakly connected accounts using the same device
    • Shortest-path distance to known fraud clusters

    Centrality alone should not determine an adverse action. A legitimate payment aggregator, popular merchant, or business account can naturally have high degree. Context and entity type matter.

    Device and identity linkage

    Shared infrastructure is often a strong signal, particularly when multiple apparently unrelated accounts exhibit similar patterns. Useful indicators include:

    • Many accounts logged in from one device or device fingerprint
    • One phone number linked to multiple recently created accounts
    • Reused IP ranges combined with synchronized activity
    • Common beneficiary lists across accounts
    • Similar app, browser, language, or location signatures

    These signals should be governed by privacy, consent, access-control, and false-positive safeguards. A shared household device is not automatically evidence of collusion.

    GNN Architectures for UPI Fraud Rings

    Different models suit different graph structures and latency requirements.

    Graph Convolutional Networks

    GCNs aggregate normalized neighbour information. They are useful as a baseline for homogeneous graphs and can perform well when relationships are relatively consistent. They may be less suitable when UPI graphs contain many node and edge types.

    GraphSAGE

    GraphSAGE learns an aggregation function and can generate embeddings for previously unseen nodes. This is valuable for newly opened accounts, new devices, and streaming transactions. Sampling limits neighbourhood size, which helps control inference cost on large graphs.

    Graph Attention Networks

    GATs assign different weights to neighbours. Attention can help distinguish a high-value link to a confirmed fraud entity from a low-risk connection, although attention weights should not automatically be treated as explanations.

    Relational and heterogeneous GNNs

    Relational GCNs, heterogeneous graph transformers, and metapath-based models handle typed relationships such as account–device–account or account–merchant–account. They are often better aligned with UPI ecosystems than a single relation type.

    Temporal GNNs

    Temporal graph networks and event-based architectures model evolving interactions. They are particularly useful for detecting sudden bursts, fast fund movement, and coordinated campaigns that emerge over minutes or hours.

    Graph autoencoders and self-supervised learning

    Fraud labels are usually incomplete. Graph autoencoders can identify structural anomalies, while contrastive or masked-node pretraining can learn embeddings from large volumes of mostly unlabeled payment data. These representations can then support supervised fraud classification.

    Building the Training Dataset

    Fraud detection is a highly imbalanced learning problem. Confirmed mule accounts may represent a tiny fraction of all payment accounts, and labels often arrive weeks after the original event.

    Label sources

    Potential labels include:

    • Confirmed fraud investigations
    • Chargeback or dispute outcomes
    • Law-enforcement or bank-confirmed cases, subject to lawful sharing
    • Account restrictions that were later validated
    • Customer reports with investigation outcomes
    • Analyst-reviewed suspicious clusters

    Separate confirmed, suspected, cleared, and unknown states. Treating every alert as fraud creates noisy labels and can teach the model to reproduce analyst bias.

    Avoiding leakage

    Use chronological train, validation, and test splits. Do not allow post-investigation fields, future account restrictions, or future transaction activity into features for an earlier prediction. In graph settings, leakage can happen through neighbours even when row-level splitting appears correct.

    A practical design is:

    1. Choose an observation window.
    2. Build the graph using events available at that time.
    3. Predict a label over a future outcome window.
    4. Freeze features and relationships at prediction time.
    5. Evaluate on a later period containing new campaigns and entities.

    Evaluating Fraud-Ring Detection Models

    Accuracy is a poor metric when legitimate transactions vastly outnumber fraud. Use metrics that reflect investigation capacity and financial risk:

    • Precision at K: How many of the top K alerts are useful?
    • Recall at fixed false-positive rate: How much fraud is found within operational tolerance?
    • PR-AUC: More informative than ROC-AUC for rare positives.
    • Amount-weighted recall: Value of prevented or intercepted losses.
    • Time-to-detection: How quickly a ring is identified after its first suspicious activity.
    • Account-level and ring-level recall: Whether the model catches whole networks rather than isolated nodes.
    • Investigator workload: Alerts per analyst and average review time.
    • Stability: Performance across banks, regions, customer segments, and time periods.

    Evaluate separately on known rings, emerging rings, new accounts, and legitimate high-volume merchants. A model that performs well on familiar fraud but misses new structures may not be ready for production.

    Real-Time Architecture for UPI Risk Scoring

    A practical deployment separates online decisioning from deeper graph investigation.

    Streaming pipeline

    1. Ingest payment, authentication, device, and account events.
    2. Validate schema, deduplicate events, and enforce event-time ordering.
    3. Update online aggregates and temporal graph features.
    4. Retrieve relevant neighbourhoods or precomputed embeddings.
    5. Run GNN inference and combine it with rules and tabular models.
    6. Apply policy thresholds and step-up controls.
    7. Log features, model version, decision, and outcome for auditability.

    For latency-sensitive payment flows, use a lightweight online model with bounded neighbourhood sampling. Run more expensive community detection, subgraph matching, and investigator analytics asynchronously.

    Hybrid decisioning

    GNN output should usually be one component of a layered system:

    • Hard rules for known compromised devices or sanctioned entities
    • Tabular models for transaction-level risk
    • GNN scores for relational and ring-level risk
    • Velocity and behavioural baselines
    • Human review for ambiguous or high-impact cases

    This architecture improves resilience when one data source is unavailable and makes policy controls easier to govern.

    Explainability and Investigator Workflows

    A fraud score without actionable evidence is difficult to operate. Provide investigators with a compact subgraph showing:

    • The subject account and its recent counterparties
    • Shared devices, phone numbers, beneficiaries, and merchants
    • Time-ordered fund flows
    • Amounts and holding times
    • Links to known or previously reviewed entities
    • Model score, top contributing signals, and confidence

    Use counterfactual checks where possible: for example, estimate how the score changes if a suspicious shared-device link or rapid onward transfer is removed. Explanations should support review, not expose sensitive detection logic to fraudsters.

    India-Specific Governance and Privacy Considerations

    UPI fraud detection operates within a regulated Indian payments environment. Banks, payment service providers, fintechs, and technology vendors should involve compliance, legal, security, and model-risk teams from the design stage.

    Key considerations include:

    • Purpose limitation and data minimisation under applicable privacy requirements
    • Lawful processing, access controls, retention schedules, and audit trails
    • Secure handling of account, device, identity, and transaction data
    • Clear separation between fraud prevention, sanctions screening, and unrelated profiling
    • Appropriate customer communication and grievance or appeal processes
    • Coordination with regulated entities and authorised reporting channels
    • Controls for cross-border data access and vendor processing
    • Fairness testing across language, geography, income, device type, and customer segment

    Do not use sensitive attributes as shortcuts for risk. Network signals can indirectly encode socioeconomic or geographic patterns, so monitor disparate false positives and establish human review for consequential actions.

    Common Failure Modes

    Treating graph connectivity as guilt

    A shared device or common beneficiary can have legitimate explanations. Use multiple independent signals and investigate context.

    Building a static graph

    Static snapshots miss campaign onset, rapid layering, and changing infrastructure. Add event time and rolling windows.

    Overlooking concept drift

    Fraudsters adapt when rules or models change. Track performance by month, campaign, corridor, and entity type.

    Ignoring cold-start accounts

    New accounts have limited history but may connect to risky infrastructure. Use inductive models, identity and device signals, and conservative step-up controls.

    Optimising only ROC-AUC

    A high ROC-AUC can coexist with poor precision among the top operational alerts. Optimise for investigation capacity and loss prevention.

    Deploying without feedback loops

    Capture analyst decisions, customer outcomes, confirmed cases, and false-positive reasons. Retrain only after validating label quality and avoiding feedback bias.

    A Practical Implementation Roadmap

    Indian fintechs and payment teams can start incrementally:

    1. Map entities and relationships: Inventory transaction, account, device, identity, and merchant data.
    2. Create a graph data model: Define stable identifiers, event-time semantics, and retention rules.
    3. Build interpretable baselines: Start with rules, aggregates, connected components, and gradient-boosted models.
    4. Launch graph analytics: Add two-hop risk, community detection, shared-infrastructure counts, and flow analysis.
    5. Train an inductive GNN: Use a chronological split and compare against the baseline.
    6. Pilot in shadow mode: Score live traffic without automatically blocking payments.
    7. Measure operational outcomes: Review precision, analyst workload, prevented loss, and customer friction.
    8. Add governance controls: Document features, thresholds, model versions, approvals, and appeal processes.
    9. Deploy gradually: Use step-up authentication, temporary holds, or manual review according to risk and policy.
    10. Monitor and refresh: Detect drift, campaign changes, data-quality issues, and fairness regressions.

    FAQ: Graph Neural Networks for UPI Fraud Detection

    Can GNNs detect a mule account with no previous fraud label?

    Yes. Inductive GNNs can use the account’s current neighbours, device links, transaction behaviour, and temporal context. However, a model should express uncertainty and avoid treating weak connections as conclusive evidence.

    Are GNNs better than rule-based fraud systems?

    They solve different problems. Rules are transparent and effective for known patterns, while GNNs identify relational and evolving patterns. A layered system generally performs better than either alone.

    How much data is needed?

    A useful pilot can begin with historical transactions, confirmed cases, and entity-link data from a defined period. Scale depends on graph size, event volume, label quality, and latency requirements—not only the number of accounts.

    Can graph models run in real time for UPI payments?

    Yes, with bounded neighbourhood sampling, precomputed embeddings, online aggregates, and efficient inference. Deeper ring analysis can run asynchronously for investigator workflows.

    What is the biggest risk of using GNNs in payments?

    False positives and opaque decisions can harm legitimate customers and merchants. Strong validation, explainable subgraphs, human review, privacy controls, and continuous monitoring are essential.

    Apply for AI Grants India

    If you are an Indian AI founder building graph-based fraud detection, payment security, or trustworthy fintech infrastructure, apply through AI Grants India. Your project may be eligible for funding, ecosystem support, and guidance to move from research or prototype to responsible deployment.

    Last updated 26 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.