Scaling Properties of Text Conditioning in Visual Generation

Scaling Properties of Text Conditioning in Visual Generation

We study empirical scaling properties for text conditioning in visual generation.

Hugging Face · Daily Papers ·Zilong Chen, Chaorui Deng · ·▲ 23 upvotes

Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.

Autores: Zilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, Haoqi Fan

  • 23 upvotes da comunidade

Resumo

Resumo original (em inglês), extraído do paper:

We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve diffusability by constructing structured prompts with semantic and geometric annotations derived from images, and improve promptability by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.

Onde ler

compartilhar: