Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
OpenAI launches Ultrafast API tier running GPT-5.6 Sol at 750 tokens/second via Cerebras
OpenAI is previewing a new Ultrafast API service tier that runs GPT-5.6 Sol up to 14× faster than standard, delivering up to 750 output tokens per second. The speed boost is powered by Cerebras hardware, marking a notable partnership for inference infrastructure. This signals intensifying competition on latency as a key differentiator in the AI API market.