Transcriptomics data processing turns raw RNA sequencing reads into defensible biological conclusions. The challenge is not simply choosing a fast aligner or producing a volcano plot. It is designing a workflow that matches the experiment, preserves metadata, detects technical failures, and separates genuine biology from batch effects.
For Indian universities, hospitals, biotech companies, and AI-biology teams, the workflow must also account for uneven compute access, cloud costs, sensitive clinical data, and limited availability of experienced bioinformaticians. This guide covers bulk RNA-seq first, while noting where the choices change for single-cell and long-read data.
Start with the experimental design
Bioinformatics cannot rescue a poorly designed experiment. Before sequencing, define:
- The biological question and primary comparison.
- The unit of replication: use biological replicates, not repeated sequencing of the same library, to estimate biological variation.
- Sample metadata, including tissue, treatment, sex, age, collection site, extraction batch, library batch, and sequencing lane.
- The reference genome and annotation release you will use.
- Whether the study needs gene-level expression, transcript isoforms, fusion detection, allele-specific expression, or novel transcript discovery.
Aim for balanced batches. If all treated samples are sequenced in one lane and controls in another, statistical correction becomes risky because condition and batch are confounded. For clinical work, document consent, de-identification, access controls, and governance requirements before data leaves the institution. Teams handling medical datasets should also review ICMR-compliant medical AI data verification in India when transcriptomic outputs feed a diagnostic or clinical AI system.
1. Organise and verify raw data
Most workflows begin with paired-end FASTQ files. Create a sample sheet linking each file to a stable sample identifier and its metadata. Do not rely on filenames alone. Check that:
- Every expected sample has the correct number of FASTQ files.
- Read pairs are synchronised.
- File checksums match the transfer records.
- Read length, sequencing type, and library orientation are known.
- Human or patient-derived data is stored with appropriate permissions.
Keep raw FASTQ files immutable. Generate processed outputs in separate directories and record software versions, reference files, parameters, and run dates. A lightweight pipeline can be written in Python; reusable Python scripts for automating data preprocessing are particularly useful for validating sample sheets, renaming files, and checking expected outputs before expensive compute jobs begin.
2. Perform read-level quality control
Run FastQC on each FASTQ file and aggregate reports with MultiQC. Inspect, rather than blindly act on, the results. Important signals include:
- Per-base quality scores and the proportion of low-quality tails.
- Adapter contamination.
- Abnormal GC-content distributions.
- High duplication levels, which may reflect low complexity, targeted biology, or over-amplification.
- Unexpected overrepresented sequences.
- Read counts and read-length consistency across samples.
Trim adapters with Cutadapt or fastp when adapters or poor-quality bases are clearly affecting downstream analysis. Aggressive trimming can remove useful sequence and reduce alignment rates, so compare pre- and post-trimming reports. For degraded RNA, such as formalin-fixed samples, expect shorter fragments and interpret quality metrics in that context rather than applying a universal cutoff.
3. Align or quantify reads
Choose the processing route based on the biological objective.
Reference-guided alignment: STAR and HISAT2 map reads to a reference genome while accounting for splice junctions. This route is appropriate when you need genomic coordinates, novel junctions, fusion analysis, or detailed quality metrics. STAR is widely used for bulk RNA-seq but can require substantial memory, especially with mammalian genomes.
Alignment-free or lightweight quantification: Salmon and similar tools quantify transcripts using a transcriptome index. They are fast and resource-efficient, making them practical when compute is limited. Use a validated decoy-aware index and understand how transcript-level estimates are summarised to genes.
Use the same genome build and annotation release throughout the workflow. Mixing an older GTF with a newer genome can cause missing features, inconsistent counts, and misleading comparisons. Record whether strandedness is unstranded, forward, or reverse; an incorrect setting can make a good library appear biologically incoherent.
4. Quantify expression and assess mapping quality
For aligned reads, featureCounts or HTSeq can generate gene-level count matrices. Inspect mapping rate, uniquely mapped reads, splice-junction support, rRNA content, 3'-bias, insert size, and the fraction of reads assigned to annotated features. Low assignment can result from wrong strandedness, incompatible annotation, contamination, poor RNA integrity, or an unsuitable reference.
Avoid treating TPM, FPKM, or RPKM as interchangeable with raw counts. For differential expression, statistical packages such as DESeq2 and edgeR generally expect untransformed integer count data. TPM is useful for describing relative transcript abundance within a sample and for some visualisations, but it should not replace an appropriate count-based model for between-group inference.
5. Normalise and model the experiment
Normalisation corrects technical differences in library size and composition; it does not remove every batch effect. DESeq2 uses size-factor approaches, while edgeR commonly uses TMM-based normalisation. Choose the method together with the study design and inspect whether normalised samples behave plausibly.
Build a design formula that reflects the experiment, for example condition plus batch, rather than comparing groups after manually deleting inconvenient samples. Use independent filtering, pre-specified thresholds, and false-discovery-rate correction. Report both adjusted significance and effect size: a tiny change can be statistically significant in a large study, while a biologically important change may be uncertain in a small one.
For a typical bulk RNA-seq comparison, review:
- Library size and mapping metrics.
- Sample-sample correlation and principal component analysis.
- Count distributions before and after normalisation.
- Whether known biological groups separate as expected.
- Outliers and their technical explanations.
- The number of genes tested and the multiple-testing method.
Do not remove an outlier solely because it weakens a preferred result. Investigate its laboratory and sequencing records first, then document any exclusion.
6. Interpret genes, pathways, and cell context
Differentially expressed genes are a starting point, not the conclusion. Annotate gene identifiers carefully, especially when combining Indian cohorts, multiple species, or older clinical records. Use Gene Ontology, pathway databases, gene-set enrichment, and ranked-list methods to reduce dependence on arbitrary DEG cutoffs.
Bulk RNA-seq can hide changes in cell composition. A shift in immune-cell abundance may look like differential expression even when individual cell types have not changed their expression programs. Where the question demands cellular resolution, consider single-cell or spatial methods, but plan for larger matrices, stronger quality-control requirements, and different statistical assumptions.
Visualisation should support decisions: PCA for structure, sample correlation for consistency, MA plots for effect-size distribution, volcano plots for screening, and heatmaps for selected genes or pathways. For teams presenting results to collaborators, AI tools for data visualization design can help improve layout and accessibility, but generated charts still require scientific review and reproducible source data.
Reproducibility, compute, and data governance
Use a workflow manager such as Nextflow or Snakemake, package environments with Conda or containers, and retain a machine-readable configuration file. Pin reference genomes, annotations, and tool versions. MultiQC reports, logs, code, sample sheets, and result tables should travel together as an analysis package.
For large studies, estimate storage and compute before sequencing. Keep raw data in durable storage, use local scratch space for alignment, and delete only reproducible intermediates after approval. Sensitive human data should not be uploaded to public services or generative AI tools without institutional authorisation. If analysis will support an AI model, apply the same discipline expected for data veracity infrastructure for high-stakes AI: provenance, validation, access control, and audit trails.
A practical release checklist
Before sharing results, confirm that:
- The sample sheet, design formula, and reference versions are archived.
- FastQC and MultiQC findings have been reviewed.
- Counts, normalisation, and statistical assumptions are documented.
- Batch effects and outliers have been investigated.
- Gene identifiers and pathway databases are versioned.
- Figures can be regenerated from scripts.
- Raw and processed data access follows consent and institutional policy.
- Results distinguish exploratory findings from validated claims.
The strongest transcriptomics projects treat processing as an auditable measurement system, not a sequence of buttons in a software package. A clear design, conservative quality control, appropriate modelling, and reproducible reporting will produce results that collaborators, reviewers, and downstream AI systems can trust.