Branda: Reasoning Bottlenecks in Ad Generators – A Compiler Engineer's Analysis
Data First: Assume one ad generation requires calling a 7B-parameter language model and a 1.5B-parameter image diffusion model. Typical inference latency: generating 128 tokens with the language model takes about 0.8 seconds (A100, FP16), while generating a 512x512 image with the image model takes about 2.5 seconds. Total: 3.3 seconds, where attention computation accounts for 60% of the language model's time, and UNet convolutions plus attention account for 75% of the image model's time. Regarding memory bandwidth, a single inference run for the 7B model needs to load approximately 14GB of parameters; if handling 10 concurrent requests, VRAM demand hits 140GB, far exceeding a single A100 card (80GB). These numbers come from actual deployment experience, not PR fluff.
Physix Frontier