- How-to
- Product LoRA
- Object LoRA
- Dataset creation
- Ecommerce
Product / Object LoRA from One Reference Photo (Angles, Materials, Captions)
Build a product or object LoRA dataset from one catalog photo: angle and material grid, white-bg vs lifestyle mix, captions, and shape-consistency curation.
Most LoRA advice assumes a face. Ecommerce sellers, prop artists, and brand marketers usually start with one catalog photo of a shoe, bottle, or weapon and need the product to stay consistent across scenes. This guide is the object sibling of creating a dataset from one image: when an object LoRA is the right tool, how to plan an angle / material / lighting grid, how captions differ from character work, and how to curate for shape consistency before you train.
What this guide covers (and what it does not)
Scope is building an object-consistent, captioned training set from one product or prop reference. Out of scope: Kohya folder trees, deep overfitting diagnosis, anime WD14 taxonomy, and a full base-model picker. You will leave with a decision table, a copyable product diversity grid, caption before/after examples, and a curation checklist.
Packaging after export: Kohya / Civitai packing. Baked white backgrounds or locked hero angles: overfitting vs undertraining. Base choice: Flux vs SDXL.
Object LoRA vs character LoRA
| Goal | Train as… | Dataset must show… | Trigger class examples |
|---|---|---|---|
| Face / OC identity across outfits | Character LoRA | Angles + expressions + wardrobe variety | ohwx woman / ohwx man |
| Fixed product design across scenes | Product / object LoRA | Sides, proportions, materials, markings you may reproduce | zvpack bottle, sksshoe sneaker, wpn01 sword |
| Wearable that drapes on bodies | Often garment (harder; more coverage) | Construction + drape on varied wearers if rights allow | Community garment starts often cite 25–50 — labeled starting range, not site law |
When a full LoRA may be overkill
This post deepens the object/product row on how many images a LoRA needs.
Image count for products (cite, do not invent)
- This site: object or product 15–30—rigid subjects need fewer samples; cover every side and two lighting setups.
- Community starting ranges (labeled as such): Everypixel Workroom cites 15–30 photos of the same product with angles, lighting, and clean/simple backgrounds; Offline Creator cites simple object/product 15–35.
Coverage beats count. Ten near-duplicate hero shots are not the same as 15–30 useful sides. Plan to the 15–30 site range and fill empty grid cells before adding duplicates.
One strong product reference
- Resolution: prefer ≥1024 px on the short side (same native class as Flux / SDXL training).
- Silhouette: product not tiny in frame; defining parts visible.
- Lighting: even light that reveals material (gloss, weave, wood grain, metal edge)—not crushed shadows that hide geometry.
- Angle: three-quarter or clear front that shows depth; a pure top-down crop alone is a weak sole reference.
- Avoid: heavy compositing, watermarks, mixed SKUs in one frame, occlusion of defining parts, extreme filters.
Lock an object spec before you expand: silhouette, proportions, materials, colors that are part of the SKU, and unique markings the LoRA should own.
Diversity grid for rigid objects
Axes differ from the character grid: no expression or wardrobe; add materials, scale, and state instead.
| Axis | Values to cover | Notes |
|---|---|---|
| View / angle | Front, rear, left, right, three-quarter L/R, top and/or underside when those surfaces matter | Multi-angle coverage is the core requirement |
| Framing / scale | Hero product shot, medium, detail/macro (logo, stitch, hinge), occasional wider in-scene | Avoid every frame filling the same % of canvas |
| Lighting | Soft studio, hard directional, natural daylight, warm indoor, one specular setup for glossy materials | Material response is the product analogue of expression |
| Background / context | Clean white or neutral, simple textured surface, 2–3 lifestyle settings | Mix on purpose — see next section |
| State (if applicable) | Closed/open, worn/unworn, filled/empty, sheathed/drawn | Caption state when visible |
| Material evidence | Matte vs gloss highlights, fabric folds, transparent parts | Same geometry, different light — proves material, not a new SKU |
Example cells across three SKUs
| Cell | Sneaker | Bottle | Sword |
|---|---|---|---|
| Front hero, soft studio, white | Generate | Generate | Generate |
| Three-quarter, hard light, concrete | Generate | Generate | Generate |
| Side / profile, daylight | Generate | Generate | Blade edge readable |
| Rear, warm indoor | Generate | Label side if needed | Hilt / scabbard |
| Macro detail | Stitch / logo | Cap / emboss | Fuller / guard |
| Lifestyle in-context | On foot / shelf | On desk | Held or sheathed in scene |
Object consistency
Filling every cell by hand is tedious. LoRA Dataset can expand one reference into object-consistent variations across angle, lighting, and background (free with your Gemini API key). Use it after you understand the coverage you need.
White background vs lifestyle (use both on purpose)
- Clean / white / neutral: teach silhouette, edges, labels, and proportions with less clutter.
- Lifestyle / in-context: teach scale, placement, and scene integration so gens do not only look like packshots.
- Only white-bg failure: hard to place in busy scenes; sticky studio correlation if uncaptioned; cutout look.
- Only lifestyle failure: clutter leaks into the concept; soft or occluded edges → deformed shape.
Bias toward clear product-first shots for geometry, then add enough lifestyle cells that backgrounds and support surfaces vary and get captioned. There is no binding 70/30 rule in the sources we cite. Caption every background you do not want fused to the trigger—same rule as character backgrounds. If bake still shows up, diagnose with overfitting vs undertraining.
Caption rules for products
Same philosophy as the captioning guide: trigger first, caption what should stay controllable. Specialize for objects.
What the trigger should own
- Fixed geometry / silhouette / unique design of this SKU.
- Permanent material identity when it defines the product (always brushed aluminum)—unless you want material promptable across a product line.
What captions should name (when visible)
- Rare trigger first + class noun (
zvpack bottle, not the real brand name as trigger). - View / orientation (front, three-quarter, rear, top-down).
- State (open, capped, laced, unfolded).
- Support surface / background / nearby props.
- Lighting / framing / scale.
- Material appearance as seen in that shot when lighting changes how it reads (
cream glaze,brushed steel,matte rubber outsole). - Readable brand / packaging text: only if you have rights. Prefer captioning that a logo is present and its placement; do not use a real brand string as the trigger token; review OCR-heavy auto-captions.
FLUX-style:
A three-quarter front view of zvpack ceramic bottle on a walnut desk, cream glaze, metal cap visible, daylight from the right.
Tag-style:
zvpack bottle, three-quarter view, cream glaze, metal cap, wooden table, daylight, product photo
Bad:
beautiful product photo, masterpiece, red sneaker
Good:
sksshoe sneaker, lateral view, white midsole, mesh upper, concrete floor, soft daylight, full product shotIllustrated anime props on Pony / Illustrious may prefer cleaned Danbooru tags—see the anime tagging guide.
Curation checklist: shape consistency (“object drift”)
Same curation mindset as identity drift, different tells. Object drift is shape / material / proportion inconsistency across the set or across seeds after training—not face morph.
- Proportions stay stable (bottle height/width, sole thickness, blade length).
- Materials do not flip (matte ↔ chrome, wood ↔ plastic) unless captioned as a lighting effect.
- Logo/label layout is stable when text fidelity matters (no gibberish mutation).
- No extra or missing parts (straps, buttons, screws).
- No near-duplicates (same angle + same background).
- Product is large enough in frame; no motion blur; silhouette not crushed by shadow.
- One SKU only—no mixed colorways or product generations.
Export toward Kohya and Civitai
Save paired image.png + image.txt (same stem) with trigger-first captions. Nest for Kohya or zip flat pairs for Civitai using packing a LoRA dataset for Kohya and Civitai.
Ecommerce, IP, and brand consent
Rights and trademarks
Doing this in LoRA Dataset
- 01
Upload one product reference
Use a sharp shot that already shows silhouette and materials clearly. - 02
Expand across the grid
Generate object-consistent variations that rotate angle, lighting, and background while holding SKU geometry fixed. - 03
Caption and download pairs
Export paired image +.txtfiles, curate for shape consistency, then pack for your trainer. Free with your own Gemini API key; Google bills model usage at their rates.
LoRA Dataset does not guarantee perfect logo OCR, ecommerce conversion, or a fixed magic output count. It turns one reference into a captioned, object-consistent backbone you still curate.
Frequently asked questions
How many images do I need for a product or object LoRA?
On this site we plan 15–30 for object/product LoRAs, with coverage of sides and lighting mattering more than raw count. Community product guides often cite similar bands (for example Everypixel 15–30; Offline Creator 15–35 as starting ranges). Fill empty angle cells before adding duplicates.
Can I train a product LoRA from a single catalog photo?
Not well if you train on that one file alone—the model memorizes one angle and background. Expand into a diverse, object-consistent set first (same workflow idea as the character one-image guide), then train.
Should every training image be on a white background?
No. Clean/neutral shots help teach shape; lifestyle/context shots help scene placement. Caption backgrounds you do not want stuck to the trigger. All-white uncaptioned sets are a common bake risk.
What should I put in product captions?
Trigger plus class noun first, then view, state, surface/background, lighting/framing, and visible material cues. Leave fixed geometry to the trigger. Do not use a real brand name as the trigger token.
Why does my product LoRA melt or warp the object?
Often weak shape evidence: cluttered frames, missing sides, or inconsistent SKUs. Add clean multi-angle shots, curate proportion drift, and remove near-duplicates.
Object LoRA vs character LoRA — which do I want for a mascot holding a bottle?
If the goal is the character’s face, that is a character dataset. If the goal is the bottle’s design across scenes, that is a product dataset. Mixing both concepts in one small LoRA makes failures hard to debug.
Can I include logos and packaging text?
Only if you have rights to reproduce them. Close-ups help legibility when text is in-scope; auto-captions need OCR review. This is not legal advice—skip third-party brands you do not own.
Sources and further reading
- LoRA Dataset: How many images to train a LoRA (object/product 15–30)
- Everypixel Workroom: Train a product LoRA
- Offline Creator: How many images for LoRA training
- Offline Creator: FLUX LoRA training dataset guide
- Offline Creator: LoRA captioning guide
- LoraAI.me: LoRA training dataset guide
- AI Wiki: LoRA training dataset preparation