Machine Heart
Aug 12, 2026 · Artificial Intelligence
Why Longer Captions Don't Boost Text‑to‑Image Models: Insights from ByteDance Seed
The ByteDance Seed team shows that merely extending caption length adds little visual supervision for diffusion models; instead, the amount of image‑grounded information in captions predicts training loss, leading them to propose Structured Prompt, new metrics (GPG, ED), and a three‑stage LLM prompter to improve both Diffusability and Promptability.
Effective DetailnessGPGcaption scaling
0 likes · 15 min read
