0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to prepare astronomical fits files for computer vision

How to Prepare Astronomical FITS Files for Computer Vision

  1. aigi

    Astronomical FITS files contain far more than pixels. A single file may include calibrated or raw detector values, multiple image extensions, exposure metadata, uncertainty maps, masks, and World Coordinate System (WCS) information. Treating it like an ordinary PNG can erase the scientific context your model needs.

    The right workflow is to preserve the original FITS data, define the computer-vision task, preprocess without leaking information, and export a model-ready representation with a reproducible record of every transformation. The same approach works for galaxy classification, transient detection, stellar segmentation, solar imaging, and remote-sensing workloads built by Indian research teams and student projects.

    1. Inspect the FITS file before changing it

    Start by identifying the HDUs (Header/Data Units), their dimensions, units, and role in the observation. Do not assume that hdul[0] contains the science image: many instruments place data in an extension such as SCI, while the primary HDU contains only a header.

    from astropy.io import fits
    
    with fits.open("observation.fits", memmap=True) as hdul:
        hdul.info()
        for index, hdu in enumerate(hdul):
            print(index, hdu.name, getattr(hdu.data, "shape", None))
            if hdu.data is not None:
                print(hdu.header.get("BUNIT"), hdu.header.get("EXPTIME"))

    Check at least the following:

    • Image dimensions and axes: distinguish a 2D image from a cube, mosaic, detector stack, or spectral array.
    • Data type and units: determine whether values are counts, counts per second, flux, magnitude, or already calibrated values.
    • Pixel scale and orientation: read WCS keywords before cropping, rotating, or resampling.
    • Missing-value conventions: look for NaN, BLANK, saturation markers, and detector masks.
    • Quality extensions: identify uncertainty, variance, exposure, and data-quality arrays.

    Keep an untouched copy and record the file checksum. If your project needs a broader computer-vision pipeline, the principles in this guide to building computer vision models on GitHub are useful for versioning code, configurations, and sample data.

    2. Calibrate and mask the science data

    Computer vision cannot compensate for uncorrected instrument effects. Use the observatory or instrument pipeline where available, then apply domain-specific corrections such as bias subtraction, dark correction, flat-fielding, cosmic-ray removal, background estimation, and astrometric calibration.

    For learning tasks, retain the calibrated science image and the quality mask as separate channels whenever possible. A saturated star, dead pixel, cosmic-ray hit, or chip gap should not be silently converted into an ordinary black pixel.

    import numpy as np
    from astropy.io import fits
    
    with fits.open("calibrated.fits", memmap=True) as hdul:
        image = np.asarray(hdul["SCI"].data, dtype=np.float32)
        mask = np.asarray(hdul["DQ"].data) if "DQ" in hdul else np.zeros(image.shape, dtype=np.uint16)
    
    valid = np.isfinite(image) & (mask == 0)
    image[~valid] = np.nan

    If uncertainty data exists, calculate a signal-to-noise feature or pass uncertainty as an additional model input. This is usually more defensible than aggressive denoising, which can remove faint structures that are scientifically meaningful.

    3. Define the target and split strategy

    Preprocessing depends on the task. A classifier may use fixed-size cutouts; a detector needs object-level coordinates; segmentation requires masks aligned pixel-for-pixel with the image. Write this specification before generating training data:

    • target label or annotation format;
    • input channels and pixel scale;
    • cutout size and overlap policy;
    • acceptable missing-pixel fraction;
    • train, validation, and test split rules;
    • evaluation metrics and expected deployment conditions.

    Avoid random splits when neighbouring cutouts come from the same sky field, exposure, object, or observing night. Such splits can place near-duplicates in training and test sets, producing misleading accuracy. Split by field, object, observation date, or telescope run, depending on the source of correlation. Also prevent augmented copies of one sample from crossing split boundaries.

    For a student or early-stage team, a small, well-documented baseline is more valuable than a large dataset with hidden duplication. Compare your design with practical guidance on building computer vision projects as a student before scaling collection and labelling.

    4. Normalize astronomical intensities safely

    Astronomical images often have extreme dynamic ranges. Direct min-max scaling can allow one cosmic ray or saturated source to determine the range for the entire image. A robust transformation should be fitted only on the training data and then applied unchanged to validation and test data.

    Common options include:

    • Percentile clipping: clip values to training-set percentiles, such as the 1st and 99th, while recording the chosen limits.
    • Asinh or sinh inverse scaling: preserve faint structure while compressing bright-source contrast.
    • Log scaling: useful for positive flux-like data, but requires careful handling of zero and negative background-subtracted values.
    • Per-image standardization: convenient, but potentially harmful if absolute brightness carries the label.

    Example using a robust asinh transform:

    import numpy as np
    
    def prepare_channel(image, low, high):
        clipped = np.clip(image, low, high)
        centre = (low + high) / 2.0
        scale = max((high - low) / 2.0, 1e-6)
        transformed = np.arcsinh((clipped - centre) / scale)
        return transformed.astype(np.float32)

    Document whether normalization is global, per band, per instrument, or per image. If combining observations from Indian and international facilities, harmonize units and calibration conventions before normalization rather than asking the model to learn instrument differences accidentally.

    5. Crop, align, and resize without losing meaning

    Generate cutouts using WCS or catalogue coordinates when the task is object-based. Pixel-coordinate crops can be wrong after mosaicking, rotation, or reprojection. Preserve the crop centre, sky coordinates, pixel scale, parent filename, and HDU in a sidecar table such as Parquet or CSV.

    Resize only when necessary. Interpolation changes point-spread functions and can blur compact sources. For segmentation masks, use nearest-neighbour interpolation; for science images, choose an interpolation method appropriate to the instrument and record it. If multiple bands are combined, align them to a common WCS and verify that their pixel grids match.

    Prefer lossless arrays such as NumPy .npy, compressed .npz, HDF5, or Zarr for training. PNG can be useful for visual inspection, but JPEG is generally unsuitable for quantitative astronomy because lossy compression creates artificial patterns and changes pixel values.

    6. Use augmentation that matches the telescope and task

    Astronomical augmentation is not a licence to apply every standard image transformation. Horizontal and vertical flips may be valid when orientation carries no physical meaning, but rotations can be inappropriate for fields with directional artefacts, trails, or instrument-specific structure. Brightness changes can also change the label if the task concerns flux, magnitude, or transient amplitude.

    Use augmentation only after deciding which properties should remain invariant. Suitable options may include small rotations, realistic noise injection based on the uncertainty map, PSF variation, masking of bad pixels, and modest translations. Store the random seed and augmentation configuration so that a result can be reproduced.

    7. Validate the exported dataset

    Create automated checks before training:

    • verify finite values, expected shapes, channel order, and units;
    • calculate the fraction of masked, saturated, and blank pixels;
    • compare class counts by field, instrument, and observing date;
    • inspect random samples before and after every major transform;
    • confirm that labels and segmentation masks remain aligned;
    • test that no source or field appears across data splits;
    • retain links to the original FITS file and processing version.

    Plot representative cutouts with identical colour scaling and inspect both bright and faint examples. A model can achieve strong validation scores while learning borders, detector defects, file provenance, or background gradients instead of astronomy. Dataset audits should therefore accompany model metrics. For implementation choices, review current open-source computer vision libraries in India and select tools that your team can inspect and maintain.

    A reproducible project structure

    A practical layout is:

    project/
      raw_fits/          # immutable originals
      calibrated_fits/   # pipeline outputs
      manifests/         # source paths, labels, WCS, checksums
      processed/         # model-ready arrays
      configs/            # preprocessing parameters
      notebooks/         # exploration only
      src/                # tested pipeline code
      reports/            # QA plots and split summaries

    Save the FITS header or a curated metadata record alongside each processed sample. Pin package versions, use a fixed configuration file, and make the preprocessing command rerunnable. In 2026, this level of provenance is especially important when datasets feed shared labs, grant-funded research, or foundation-model experiments.

    Final checklist

    Before training, confirm that you have:

    • preserved the original FITS files and checksums;
    • identified the correct science HDU and units;
    • applied calibration and retained quality information;
    • chosen normalization using training data only;
    • protected WCS, labels, and pixel-scale assumptions;
    • used lossless storage for scientific values;
    • split by independent fields, objects, or observing runs;
    • audited leakage, missing pixels, class balance, and visual quality;
    • recorded every preprocessing parameter and software version.

    The goal is not to make FITS files look like ordinary photographs. It is to convert calibrated measurements into stable, documented inputs while preserving the information needed to interpret a model's result. That discipline gives computer-vision projects a credible path from a local notebook to reproducible astronomical research.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.