When Marginal Cost of LLM Inference Becomes the Bottleneck, Chip Design Philosophy Needs a Shift
The computational characteristics of LLM inference are vastly different from training: autoregressive decoding generates only one token per step, and the compute intensity is far lower than the matrix multiplications in the training phase. This means the bottleneck for inference chips isn't peak FLOPS, but rather memory bandwidth, latency, and data movement efficiency. So, what architectural path should chips optimized for inference take?
Physix Frontier