- How-to
- Anime
- WD14
- Pony
- Illustrious
- Captioning
Anime Character LoRA Datasets: WD14/Danbooru Tags, Pony/Illustrious, What Not to Tag
Caption anime character LoRAs for Pony and Illustrious with cleaned WD14/Danbooru tags: trigger first, strip meta, and pick a deliberate score_ policy.
The general captioning guide teaches trigger first and “caption what changes.” Anime trainers on Pony Diffusion XL, Illustrious XL, and other Danbooru-fluent checkpoints still hit a last-mile question: which auto-tagger, which tag categories stay, and what do you delete so the trigger owns identity? This page is that anime specialization—WD14 / Danbooru cleanup, Pony score_ policy, and what not to leave in the line.
What this guide covers (and what it does not)
Scope is caption format and cleanup for anime character LoRAs aimed at Pony, Illustrious, and similar tag-native bases. Out of scope: learning rate, network dim, epoch counts, folder packing, and image counts. You will leave with:
- A clear tags-vs-natural-language decision by base model.
- A WD14 cleanup checklist and a concrete what-not-to-tag list.
- An explicit Pony score / source / rating training policy (pick one).
- Before/after caption examples you can copy as a pattern.
For packing after captions, see packing a LoRA dataset for Kohya and Civitai. For counts and steps, use how many images to train a LoRA—do not invent a magic N here.
Match the caption language to the base model
Train for the vocabulary the checkpoint already knows. Community workflows repeatedly push WD14 (or other booru taggers) for Pony/Illustrious and JoyCaption / VLMs when the base expects prose.
| Base / lineage | Caption language that usually fits | Why |
|---|---|---|
| Pony Diffusion V6 XL (+ many V6-derived) | Danbooru-style tags; optional NL; deliberate score_ / source_ / rating_ | Model card and community guides document tag + NL mix and that vocabulary |
| Illustrious XL lineage | Danbooru-style tags; later cards also describe NL | Built around annotated illustration / Danbooru-style data |
| Flux / photoreal SDXL | Natural-language sentences | Covered on Flux vs SDXL and the parent captioning post |
| Generic anime SD1.5 / SDXL anime merges | Usually tags; read the exact checkpoint card | Do not invent merge-specific rules |
Rule of thumb
WD14 / SmilingWolf in one screen
What the tagger outputs
SmilingWolf’s WD tagger family (for example wd-eva02-large-tagger-v3) emits ratings, character, and general tags. Models are trained on Danbooru images; v3 notes tags through early 2024 with filters such as dropping images that have too few general tags. Sibling checkpoints (wd-swinv2-tagger-v3, and others) exist—pick one and stay consistent rather than ranking them here.
Practical use: draft tags, then human-clean in TagGUI, BooruDatasetTagManager, or similar. Never ship raw dumps. Community tagging guides often start general thresholds around 0.65–0.75; treat that as a starting range and verify on your images—not as law.
Where WD14 sits vs JoyCaption / Florence / other VLMs
- WD14 family: fast, anime/illustration-oriented, emits Danbooru-like tags—default draft for Pony/Illustrious character sets.
- JoyCaption / Florence-2 / similar: natural-language descriptions—stronger fit when the base expects prose (Flux) or for hybrid experiments. These tools can mis-color, swap left/right, or invent details; always review.
- Hybrid (tags + short NL): a documented experiment in some workflows, not a required recipe. Consistency beats maximizing tools.
Danbooru tag taxonomy (only what trainers need)
Danbooru’s help docs define five categories. For character LoRAs, treat them differently:
| Category | Examples | Typical character LoRA handling |
|---|---|---|
| General | 1girl, blue_hair, sitting, school_uniform | Core of the caption — keep accurate visible tags |
| Character | hatsune_miku, … | Replace with your trigger for OCs; for canon, decide name vs trigger |
| Copyright | series / source work | Often strip for OCs; keep only if you want franchise conditioning |
| Artist | artist names | Almost always strip (style bleed) unless the LoRA is an artist style |
| Meta | absurdres, translated, scan | Strip from training captions |
Underscore convention: Danbooru uses blue_hair; some trainers normalize spaces. Pick one convention and stay consistent. Pony guides note both tags and natural language can work; consistency beats cleverness.
The anime version of “caption what changes”
Same philosophy as the parent guide, specialized for anime OCs. Still first on every line: trigger, then class/count (1girl / 1boy / solo as appropriate).
| Trait | Typical anime OC handling | If you tag it every time | If you omit it |
|---|---|---|---|
| Eye color | Often permanent identity | Reinforces look; may need the tag at inference | Trigger may absorb eyes (easier prompts; harder recolors) |
| Hair color / length | Modular for many OCs (wigs, forms) | Lets you prompt recolors/cuts | Trigger may bake one hairstyle |
| Signature accessory | Permanent if defining (ahoge, horns, halo) | Keeps the accessory reliable | May vanish unless the trigger absorbs it |
| Outfit / costume | Variable — tag per image | Outfit becomes promptable | Outfit bakes to the trigger |
| Pose / expression / framing / background | Always variable when the dataset varies | Controllable | Pose or background bake |
Tags included stay editable
What not to tag (identity leaks and noise)
- Meta / resolution / upload noise:
absurdres,highres,incredibly_absurdres,scan,comic, watermark-related meta—junk correlations. - Wrong or competing character/copyright names from WD14 when training an OC or a rename trigger—they fight the trigger.
- Artist tags—style contamination unless the LoRA is an artist style (out of scope here).
- Duplicate count tags (
1girltwice)—dedupe. - Negative / quality insults in the positive caption (
blurry,bad_anatomy,worst_quality)—describe what is present; do not train on “bad” words as content tags. - Pony score spam without a policy—see the next section.
- Outfit tags when every image shares one outfit and you want wardrobe freedom—add outfit diversity + tags, or accept bake-in. If likeness fails after clothing changes, see identity drift.
Before / after: raw WD14 to a cleaned caption
Fictional OC example—pattern only, not a franchise dump:
Before (raw dump):
1girl, absurdres, highres, hatsune_miku, twintails, blue_hair, school_uniform, outdoors, day, solo, looking_at_viewer
After (cleaned):
mychar, 1girl, solo, blue_hair, twintails, green_eyes, school_uniform, outdoors, day, looking_at_viewerStrip meta and the wrong character name, put the trigger first, keep accurate general appearance tags. Same character, second outfit—permanent traits and trigger stay; only variable tags change:
mychar, 1girl, solo, blue_hair, twintails, green_eyes, hoodie, jeans, indoors, night, sittingContrast with a Flux-style sentence for the same image (use this shape on NL-native bases, not as the default for Pony/Illustrious):
mychar, a girl with blue twintails and green eyes in a school uniform outdoors during the day, looking at the viewer, soloPony Diffusion XL caption quirks (score_, source_, rating_)
What score tags are
Pony V6 uses aesthetic-ranking-derived tags such as score_9, score_8_up, score_7_up, and so on—commonly used at inference to ask for “good” images. V6 labeling moved toward verbose chains (score_9, score_8_up, …); the score-tag writeup describes a Clever Hans-style effect where the long string correlated with good images. The same author does not give a strong personal recommendation for using score tags in LoRA training. Do not treat “always paste score_9 into every caption” as gospel.
Training policy options (pick one, document it)
- Omit score tags from training captions—learn identity/content only; add score tags at inference as usual.
- Include score tags that honestly match image quality—avoid labeling weak refs as
score_9. - Include a consistent score prefix on all images—may bind quality conditioning to the trigger; test with and without at inference.
Same image, three policy demos (not a ranking of “best”):
(a) no scores:
mychar, 1girl, solo, blue_hair, school_uniform, outdoors
(b) honest mid quality:
score_8_up, mychar, 1girl, solo, blue_hair, school_uniform, outdoors
(c) full chain prefix:
score_9, score_8_up, score_7_up, mychar, 1girl, solo, blue_hair, school_uniform, outdoorsAlso cover source_anime / source_cartoon / source_furry / source_pony and rating_safe / rating_questionable / rating_explicit: only when accurate; wrong source tags add noise. Pony V6 is SDXL lineage; Pony V7 / AuraFlow is a different architecture—do not assume V6 caption recipes transfer unchanged.
Illustrious: tags first, read the card
Illustrious is SDXL-based and developed around large annotated illustration / Danbooru-style data. Later versions describe stronger natural-language understanding. Practical advice: accurate Danbooru-style tags remain the safe default for character LoRAs; hybrid NL is optional if the exact release’s card supports it. Train against the exact Illustrious checkpoint you will use. For choosing the base family, see Flux vs SDXL LoRA training.
End-to-end anime caption workflow
- 01
Freeze trigger spelling + class token
Example:ohwxplus1girl/girl. Keep spelling identical across every file. - 02
Curate images for diversity
Outfit, pose, expression, and background should vary. Build that set with one-image expansion if you only have a reference; size it with the image-count guide. - 03
Auto-tag with WD14 (or equivalent)
Use a conservative threshold, export.txtsidecars beside each image. Kohya’s sd-scripts also ship a WD14 tagging utility if you already live in that toolchain. - 04
Bulk-clean and put the trigger first
Strip meta/artist/wrong names, fix synonyms, ensure trigger + class lead the line. In Kohya,keep_tokenscan protect the leading trigger when caption shuffle is on—set it to cover trigger (+ class) without turning this into a full trainer settings guide. - 05
Apply the what-changes pass
Image by image: tag variables; decide which fixed traits stay tagged vs absorb into the trigger. - 06
Decide Pony score/source/rating policy
Only if training on Pony. Pick omit / honest match / consistent prefix and stick to it. - 07
Spot-check, then pack
Read every caption beside its image. Then follow Kohya folder / Civitai zip packing.
Using LoRA Dataset exports for anime tag training
LoRA Dataset expands one reference into an identity-consistent set with captions (free with your own Gemini API key). Current product captions lean natural-language. For Pony / Illustrious, treat that export as the image backbone, then run a WD14 (or similar) tagger pass and/or manual tag edit so captions match the checkpoint. Do not expect Danbooru tags out of the box unless an explicit tag mode ships in-product.
Soft path
Troubleshooting
| Symptom | Likely caption cause | First fix |
|---|---|---|
| Trigger ignored / wrong franchise character appears | Competing character/copyright tags left in | Strip names; strengthen trigger-first |
| Always same outfit | Outfit never tagged (or never varied) | Tag outfits or add outfit diversity |
| Cannot recolor hair | Hair color omitted → absorbed by trigger | Add hair tags when you want control — or accept bake-in |
| Muddy “generic anime” look on Pony | Missing/wrong Pony vocab at inference or confused score policy | Test documented score/source prompts; do not invent quality phrases from other models |
| Great on Flux prompts, weak on Illustrious | NL captions on a tag-native base | Retag with WD14 + cleanup |
If tags look clean but likeness still wanders, packaging and captions are probably fine—see fixing identity drift.
Frequently asked questions
Do I need WD14 if I’m training on Flux?
Usually no. Flux workflows favor natural-language captions. WD14 shines when the base model’s strong vocabulary is Danbooru-style tags (Pony, Illustrious, many anime SD checkpoints).
Should every Pony training caption start with score_9, score_8_up, …?
Not automatically. Those tags come from Pony’s aesthetic labeling and are powerful at inference. For training, pick a deliberate policy (omit, match quality, or consistent prefix) and test—the score-tag author does not prescribe a LoRA training recipe.
Eye color: tag it or leave it to the trigger?
For most anime OCs, eye color is a stable identity cue. Tagging it every time keeps recolor control harder but likeness clearer; omitting it lets the trigger absorb eyes (easier prompts, harder recolors). Choose intentionally; don’t mix both strategies across the same dataset without a reason.
WD14 put the wrong character name on my OC images. Keep it?
No. Competing character tags fight your trigger. Replace with your trigger (and class/count tags) and keep accurate general appearance tags.
JoyCaption vs WD14 — which is better?
Neither universally. WD14 drafts booru tags for anime tag-native bases; JoyCaption drafts prose for NL-native bases. Match the tool to the checkpoint, and always human-review either output.
Can I mix natural language and Danbooru tags in one caption?
Some trainers do (especially on Pony, which was trained with both). If you mix, keep order and convention consistent across the set and don’t let prose contradict tags. Pure cleaned tags remain the simplest reliable path for Illustrious/Pony character LoRAs.
Will LoRA Dataset give me Danbooru tags automatically?
Not today. Product captions lean natural-language. Plan on a WD14 (or similar) tagger pass or manual conversion for Pony/Illustrious unless an explicit tag mode ships in-product.
Sources and further reading
- SmilingWolf wd-eva02-large-tagger-v3 model card
- SmilingWolf WD Tagger demo space
- Danbooru Help: Tags (five categories)
- Danbooru howto:tag
- AstraliteHeart: What is score_9 and how to use it in Pony Diffusion
- Offline Creator: Pony LoRA training guide
- Offline Creator: Illustrious LoRA training guide
- Civitai: Ultimate LoRA tagging guide
- kohya-ss/sd-scripts: WD14 tagger README