Guides to building LoRA training datasets
Short, sourced answers to the questions people ask before training a character LoRA: how many images, how to caption them, how to stop identity drift, and how to hand the dataset to Flux, SDXL, Kohya, or Civitai. Each guide cites its sources and states when it was last checked.
Last updated 5 September 2026
How many images do you need to train a character LoRA?
Practical image counts for character LoRAs on Flux, SDXL, and SD1.5, why variety matters more than volume, and how repeats let small datasets train fully.
Updated 5 September 2026 · 4 min read
How to caption a LoRA dataset
Caption rules for character LoRAs: trigger words, what to describe and what to omit, natural-language captions for Flux versus tags for SDXL and SD1.5.
Updated 5 September 2026 · 3 min read
How to fix identity drift in a character LoRA
Why a character LoRA produces faces that look related but not identical, how to tell whether the dataset or the training is at fault, and the fixes for each.
Updated 5 September 2026 · 3 min read
IP-Adapter vs LoRA for consistent characters
IP-Adapter conditions on a reference image with no training; a LoRA bakes a character into trained weights. When to use each, and how to combine them.
Updated 5 September 2026 · 3 min read
FLUX.1 LoRA dataset requirements
What a FLUX.1 LoRA dataset needs: image count, resolution and aspect ratios, natural-language captions with a trigger word, folder layout, and licence notes.
Updated 5 September 2026 · 3 min read
Kohya vs Civitai trainer: how to format a LoRA dataset
How Kohya sd-scripts and the Civitai on-site trainer expect a LoRA dataset to be organised: folder names, repeats, TOML config, zip uploads, and caption files.
Updated 5 September 2026 · 3 min read
How to create a LoRA dataset from one image
Step by step: turn one reference photo into a captioned, identity-consistent LoRA training dataset with LoRA Dataset, free with your own Gemini API key.
Updated 5 September 2026 · 3 min read