JWST data is not ready for a neural network when it first arrives. Detector effects, cosmic rays, varying exposure times, missing pixels, World Coordinate System (WCS) metadata, and instrument-specific artefacts can all become shortcuts that a model learns instead of the astrophysics. A robust workflow therefore treats preprocessing as part of the scientific method—not as cosmetic image enhancement.
This guide presents a practical pipeline for preprocessing JWST images for machine learning models in 2026, with an emphasis on reproducibility, physical validity, and dataset design.
Start with the right JWST data product
JWST observations are distributed through the Mikulski Archive for Space Telescopes (MAST) in stages. A typical progression is:
- **Uncalibrated data (
*_uncal.fits)**: detector-level exposures that still require the full calibration pipeline. - **Rate or rate-intensity data (
*_rate.fits,*_rateints.fits)**: products after detector-level corrections, useful for inspecting individual integrations. - **Calibrated exposures (
*_cal.fits)**: science-ready single exposures with physical units and associated data-quality information. - **Resampled products (
*_i2d.fits)**: distortion-corrected, combined images generated by Stage 3 processing.
For most image-classification or segmentation projects, begin with calibrated exposures or carefully generated mosaics rather than downloading rendered JPEG or PNG images. Keep the FITS headers, WCS, exposure metadata, filter, detector, and observing programme identifiers alongside every image.
The standard implementation is the open-source jwst Python pipeline, used with astropy, numpy, and asdf-compatible reference files. Record the pipeline version, CRDS context, parameters, and reference-data versions in a machine-readable manifest. These details can change pixel values and make apparently identical training runs irreproducible.
Calibrate before transforming pixels
Run the appropriate JWST pipeline stages for the instrument and observing mode. Calibration commonly includes bias or zero-point handling, reference-pixel correction, linearity correction, dark subtraction, saturation detection, flat-fielding, photometric calibration, and WCS assignment. The exact sequence differs between NIRCam, NIRISS, NIRSpec, and MIRI data, so do not copy parameters from another instrument without checking its documentation.
Treat the data-quality (DQ) array as first-class input. Saturated pixels, bad detector regions, cosmic-ray flags, unstable pixels, and missing data should not be silently replaced with zero. A zero can look like a real dark region and become a misleading feature. Instead:
- Preserve the DQ mask through every processing stage.
- Convert unusable pixels to NaN or a documented fill value.
- Add a binary validity mask as a second or third model channel.
- Report the fraction of masked pixels per image and reject pathological samples.
Do not use aggressive denoising before establishing a scientific baseline. Median, Gaussian, and wavelet filters can remove compact sources, weaken emission features, or create textures that are absent from the detector data. If denoising is scientifically justified, compare the model against an unfiltered pipeline and quantify what structures are lost.
Align, resample, and combine carefully
Alignment is necessary when combining exposures, filters, or epochs, but every resampling step changes noise correlations and can soften point sources. Use WCS-aware tools rather than image-only affine transformations when the observations cover different detectors or orientations. For mosaics, document the pixel scale, projection, kernel, and treatment of outliers.
A useful rule is to choose one representation for one task:
- Use single calibrated exposures when detector-level variation, transient events, or native resolution matters.
- Use drizzled or resampled mosaics for object-level morphology across a common field.
- Use multi-band stacks only when all channels have compatible astrometric coverage and a clearly defined registration process.
Keep the native science units where possible. Avoid mixing counts, count rates, surface brightness, and arbitrary display intensities in the same dataset. If mosaics have different exposure times or backgrounds, normalize using metadata and physically motivated estimates rather than per-image min-max scaling alone.
Normalize without erasing astrophysics
JWST images have high dynamic range. A practical input transform may use a signed or non-negative logarithmic mapping, such as asinh or log1p, after handling the background and invalid pixels. The transform must be fixed or fitted on the training split, then applied unchanged to validation and test data.
Recommended checks include:
- Estimate and subtract background using a method appropriate to the field, avoiding extended emission.
- Clip extreme values only with documented, training-derived thresholds.
- Scale channels consistently across filters and instruments.
- Preserve negative values when they carry information about background noise or calibrated flux.
- Store the inverse transform and units so predictions can be interpreted scientifically.
Display-oriented colour composites are useful for humans but should rarely be the primary model input. They combine filters through arbitrary stretches and can introduce correlations unrelated to the target. For computer vision experiments, start with calibrated single-band arrays or physically defined multi-channel tensors. Teams learning the broader computer-vision workflow can compare this approach with the practices in computer vision model projects on GitHub.
Build leakage-resistant datasets
The biggest machine-learning failure may occur before training. Randomly splitting individual cutouts can place nearly identical exposures, adjacent sky regions, or multiple observations of the same source in train and test sets. The resulting score will look strong while generalisation is poor.
Split by a meaningful observational group, such as:
- Target or astronomical source
- Proposal or observing programme
- Detector and visit
- Sky tile or field
- Observation epoch
Create labels and preprocessing manifests before augmentation. Store source identifiers, coordinates, filters, exposure dates, quality flags, and provenance. For rare classes, report class counts after grouping—not just after generating overlapping patches. A beginner-friendly project can use this workflow as a rigorous extension of machine learning portfolio projects for beginners in India, provided the scientific split is preserved.
Augmentation and model-ready packaging
Apply augmentation only when it reflects the invariances of the task. Rotations and flips may be reasonable for morphology classification, but not when orientation, detector artefacts, diffraction spikes, or instrument geometry are part of the label. Avoid arbitrary colour jitter on calibrated flux channels. Add realistic noise only when its distribution is understood and parameterised from the observations.
Package each sample with:
- Image tensor and validity mask
- Filter, instrument, pixel scale, and units
- WCS or a traceable sky-coordinate reference
- DQ summary and background statistics
- Label definition and source provenance
- Preprocessing configuration and software versions
Use FITS or Zarr for archival arrays and a compact index such as Parquet for metadata. Keep a small, untouched diagnostic set for visual and numerical checks.
Validate before training
Before fitting a model, inspect random samples, masked regions, histograms, background distributions, and per-filter statistics. Compare source counts and flux distributions across splits. Confirm that the model cannot predict the label from exposure ID, missingness patterns, border artefacts, or a pipeline version.
Useful baselines include a simple logistic or random-forest model on summary statistics, a CNN on minimally processed arrays, and a version with explicit masks. If performance collapses after grouping by source or programme, treat that as a dataset finding—not a reason to restore a leaky split. For more advanced vision work, document the experiment in the same disciplined way expected of machine learning projects for computer science students.
A reproducible minimum workflow
1. Download the relevant JWST products and metadata from MAST.
2. Run the instrument-appropriate jwst calibration pipeline with pinned CRDS settings.
3. Preserve DQ masks, WCS, units, and provenance.
4. Select calibrated exposures or documented mosaics for the task.
5. Estimate background and apply a fixed, training-derived normalization.
6. Align only when scientifically necessary; record resampling choices.
7. Split by source, field, programme, or epoch before augmentation.
8. Export tensors, masks, metadata, and configuration files.
9. Validate distributions and visual samples, then benchmark simple models.
The goal is not the cleanest-looking JWST image. It is a data product whose transformations are traceable, whose masks are respected, and whose test performance reflects real astronomical generalisation.