Optimizing GPT-5.6 Sol Reveals the Load Balancing Struggles of Inference Engines
OpenAI spent two weeks reducing the Time to First Byte (TTFB) latency for the GPT-5.6 Sol API by 12%, while simultaneously loosening the requests-per-minute limit from 1,000 to 3,000. The changes on the cost side are more subtle: the price per million tokens remained unchanged, but compute consumption only increased by 8%. These figures come from my tracking of user-measured data and cross-referencing with official documentation.
Physix Frontier