0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build custom lora for stable diffusion

How to Build a Custom LoRA for Stable Diffusion

  1. aigi

    Stable Diffusion is useful out of the box, but a general model rarely understands a brand’s visual identity, a product category, a recurring character, or a person’s appearance consistently. A custom Low-Rank Adaptation (LoRA) lets you teach a base model a focused concept without retraining the entire checkpoint.

    A LoRA is a small adapter—often tens to a few hundred megabytes—that modifies selected model layers while keeping the original weights frozen. This makes it practical for product teams building private image workflows for fashion, gaming, advertising, architecture, and regional media. The goal is not to memorise every training image. The goal is to learn a reusable concept that performs reliably on new prompts, poses, settings, and compositions.

    Choose the right training target

    Start by defining what the LoRA must learn and what the base model should continue to control:

    • Subject LoRA: A person, character, mascot, vehicle, or product.
    • Style LoRA: A visual language such as editorial photography, illustration, textile rendering, or a brand art direction.
    • Object or product LoRA: A repeatable appearance for jewellery, furniture, packaging, or machinery.
    • Wardrobe or material LoRA: Sarees, embroidery, architectural finishes, or other specialised visual details.

    Use the newest model family your inference stack supports. SD 1.5 remains easier to train on 8GB GPUs and has a large ecosystem. SDXL generally produces better composition and detail but needs more VRAM and a larger, cleaner dataset. Do not mix SD 1.5 and SDXL training assumptions: the base checkpoint, text encoders, resolution, captions, and inference workflow must match.

    Teams already building broader custom-model pipelines may benefit from the same data discipline described in best practices for fine-tuning LLMs on custom data, even though image LoRA training uses different architectures.

    Prepare a dataset that teaches generalisation

    Dataset quality matters more than endlessly adjusting hyperparameters. For a subject LoRA, begin with roughly 15–30 strong images; complex styles or product ranges may require 40–100. Remove near-duplicates, screenshots, watermarks, compression artefacts, accidental text, and images where the subject is hidden.

    Include controlled variation:

    • Multiple angles, distances, poses, expressions, and lighting conditions.
    • Different backgrounds when the background is not part of the concept.
    • Several compositions so the model does not associate the subject with one crop.
    • Clear examples of details that matter, such as fabric weave, logos, jewellery, or facial structure.

    Respect consent, ownership, and licensing. Do not train on a person’s likeness, client assets, or copyrighted artwork without documented permission. For an Indian startup, maintain an asset register covering source, licence, consent status, intended use, and deletion requests. This is essential when a prototype becomes a customer-facing product.

    Use aspect-ratio bucketing rather than forcing every image into a square crop. For SD 1.5, 512 or 768 buckets are common; for SDXL, 1024-based buckets are typical. The output resolution should reflect how the LoRA will actually be used, while keeping memory limits realistic.

    Captioning and activation tokens

    Each image should have a matching text caption file. Automatic captioners such as WD14 can create a starting point, but review the results manually. Captions should describe elements the model must understand as variable—pose, clothing, camera angle, background, lighting, and composition—while the unique concept is represented by a consistent activation token.

    Choose an uncommon token such as bharatloom_subject rather than a normal word that already has strong meaning in the base model. For a person, captions might include the activation token, age range, pose, clothing, and setting. If clothing should remain changeable, caption it consistently instead of allowing it to become fused with the identity. If a particular visual feature is the concept, avoid over-describing it in every caption.

    Generate captions in batches, then inspect a representative sample. Bad captions can teach the model unwanted associations more effectively than a small learning-rate mistake.

    Set up Kohya_ss or an equivalent trainer

    Kohya_ss remains a practical interface for sd-scripts, especially for teams that need repeatable experiments. OneTrainer is another option. Use a clean Python environment, a compatible CUDA and PyTorch combination, and a fast SSD. Pin versions for a production workflow; a training run that cannot be reproduced is difficult to debug.

    Typical requirements are:

    • SD 1.5: 8GB VRAM can work with batch size 1, gradient checkpointing, and memory-efficient optimisers.
    • SDXL: 12–16GB is more comfortable; 24GB makes iteration substantially easier.
    • Storage: Reserve at least 30–50GB for checkpoints, caches, previews, and experiment outputs.
    • Cloud training: Use a GPU rental when local hardware is inadequate, but encrypt sensitive datasets and delete persistent volumes after verified cleanup.

    Create a fixed test prompt set before training. Keep the seed, sampler, steps, resolution, and negative prompt constant while comparing checkpoints. This turns subjective experimentation into a useful evaluation process.

    Starting hyperparameters

    There is no universal golden configuration, but these are sensible starting points:

    • Network dimension: 16–32 for simple styles; 32–64 for faces, products, and complex objects.
    • Network alpha: Usually half the dimension, such as alpha 16 for rank 32.
    • Batch size: 1 or 2, increased only when memory allows.
    • Optimizer: AdamW8bit is a dependable baseline. Test alternatives only after establishing a baseline.
    • Learning rate: Start around 1e-4 for the U-Net and lower for the text encoder, such as 5e-5. Disable text-encoder training initially if the concept is simple or the dataset is small.
    • Precision: fp16 is broadly compatible; bf16 is preferable where the GPU and software stack support it reliably.
    • Epochs: Save multiple checkpoints rather than trusting a single final epoch.
    • Regularisation: Use class or regularisation images when preserving the broader category matters, especially for identity training.

    Track steps, dataset repeats, resolution, rank, captions, base model hash, and software versions. “Ten epochs” means little without the number of images and repeats.

    Train, evaluate, and select the checkpoint

    Watch the loss curve, but do not use loss as the sole quality metric. A lower loss can coincide with a less useful adapter. At regular intervals, test each checkpoint with:

    • The activation token alone.
    • New poses, backgrounds, lighting, and camera distances.
    • Prompts that deliberately change clothing, environments, and composition.
    • Negative prompts and realistic production inputs.

    Use an XYZ grid that varies LoRA strength, for example 0.4, 0.6, 0.8, 1.0, and 1.2. The best weight is often below 1.0. Overfitting appears as copied poses, fixed backgrounds, harsh contrast, burnt details, or an inability to follow prompt changes. Underfitting produces weak identity or style transfer.

    For product use, evaluate more than aesthetics. Measure prompt adherence, failure rate, latency, memory use, and consistency across a fixed benchmark. Store preview grids and metadata with every release. If your team is also designing automated creative workflows, treat the evaluator like an agent component and document its inputs and outputs as you would when you build generative AI agents.

    Deploy safely in ComfyUI or Automatic1111

    Export the selected .safetensors file and place it in the correct LoRA directory for your inference tool. Keep the base checkpoint and LoRA versions paired in a model registry. A useful release record includes:

    • Base model name and hash.
    • Training dataset version and consent status.
    • Captioning method and reviewed-caption version.
    • Training configuration and random seed.
    • Recommended LoRA weight range.
    • Known failure cases and prohibited uses.

    For a product, expose only approved adapters, validate prompts and outputs, and log model versions. Do not assume that a LoRA preserves brand safety or prevents generated text, logos, or likeness misuse. Add moderation and human review where outputs can affect customers or public communications.

    Common mistakes and practical fixes

    • The subject is not recognisable: Improve image quality and captions before increasing rank or training time.
    • The model copies the dataset: Reduce repeats, lower the learning rate, stop earlier, or add more variation.
    • Clothing or background is locked: Caption those elements consistently and add contrasting examples.
    • Images look distorted: Check aspect-ratio buckets, resolution, latent settings, and corrupted files.
    • SDXL training runs out of memory: Use gradient checkpointing, memory-efficient attention, lower batch size, and a smaller rank.
    • Results change between machines: Pin dependencies, record seeds, and verify the base checkpoint hash.

    Building a useful LoRA business in India

    The strongest applications solve a repeatability problem, not merely an image-generation problem. Fashion sellers can create controlled catalog variations for regional garments; studios can maintain character continuity across episodes; architects can explore local materials and streetscapes; and consumer brands can generate campaign concepts without sending sensitive assets to an external API.

    Start with one narrow workflow, define acceptance criteria, and compare the LoRA against a prompt-only baseline. A small, well-evaluated adapter is more valuable than a large collection of untracked experiments. If the broader system includes customer-facing automation, connect the image workflow to your product architecture deliberately—similar principles apply when teams build computer vision models on GitHub and move from experiments to maintainable services.

    A custom LoRA is technically lightweight, but production quality comes from disciplined data collection, reproducible training, honest evaluation, and clear rights management. That combination lets Indian builders turn Stable Diffusion from a general tool into a dependable component of a local, specialised creative product.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.