New ViT Anime Tagger Beats WD-tagger v3 by +0.053 mAP
darkth0ughts · reddit · 2026-08-12
A developer trained a new anime image tagger from scratch using a DINOv3 ViT backbone combined with a cross-attention tag query head.
- Architecture: Utilizes a query head design and was trained entirely on the Danbooru2025 dataset.
- Performance: On an intersection evaluation subset of 3,383 tags, the new L/16 model achieved 0.5352 mAP, outperforming the popular WD-eva02-large-tagger-v3 by +0.053 mAP while offering lower latency.
- Availability: Includes an interactive Hugging Face demo space and model card.
More from Multimodal
- AI Music Analyzer Showcased: Auto-Extracts Song Highlights — threepointone · 2026-08-12
- MiniMax Video Drift: Close-Ups Lose Identity After 3 Seconds — ItsMilaVoss · 2026-08-12
- LTX 2.5 Tested: Generates 15s 1080P Video in 8 Minutes on RTX 4080 — skyrimer3d · 2026-08-12
- MiniMax H3 local full-body shots suffer face distortion, likely due to VRAM limits — Loud-Guitar1920 · 2026-08-12
- How Cultural Differences Shape AI Video Model Preferences: H3 vs. LTX 2.5 — xbobos · 2026-08-12
- Pushing Generative AI Video to the Limits with Cinematic Lighting and Physics — eyishazyer · 2026-08-12