ULTRAFAST//SOL
SPEED PTS 0

OPENAI · AUG 13 2026API LIMITED PREVIEW

GPT‑5.6 SOL, AT 14× THE SPEED

OpenAI’s most intelligent model now streams up to 750 output tokens per second on Cerebras wafer-scale silicon. Until now, real-time speed meant a smaller model. That trade-off is gone.

TOKEN RELAY one illustrative 1,500-token answer · three lanes · true relative rates
ULTRAFAST0 tok/s
STANDARD0 tok/s
YOU READ (≈300 wpm)0 tok/s
0.00s

HEADLINE NUMBERS

MORE USEFUL WORK
PER SECOND

Tap a die to flip it. Every card you absorb adds Speed Points — and every figure is straight from the announcement.

ANATOMY

THE LARGEST CHIP
EVER BUILT

Cerebras WSE‑3. One die, an entire 300 mm wafer. Keep scrolling — the surface opens up.

01 / 06

TO SCALE

57× THE SILICON
OF THE BIGGEST GPU

WSE‑3 spreads 46,225 mm² of compute across a whole wafer. The largest conventional GPU die — 814 mm² — is the faint rectangle inside it. Drag the loupe.

WSE‑3 · 46,225 mm² H100‑class die · 814 mm²
AREA
57×
AI CORES
52×
ON‑CHIP MEMORY
880×

PLAY THE FIVE

WHERE DOES
SPEED WIN?

Five real arenas from the announcement. Each one is a playable micro-challenge — beat the clock the way Ultrafast does, and bank the points.

ARCADE

TOKEN RUSH

Ride the interconnect fabric. Harvest SOL tokens and cache bursts; dodge thermal spikes. The fabric speeds up — like demand on a wafer-scale engine.

TOKEN RUSH

Harvest the token stream. Don’t thermal-throttle.

Works with mouse, touch, and keyboard. Every token you harvest is one the GPU never saw coming.

CERTIFICATION

PROVE YOU ABSORBED
THE SIGNAL

Eight questions. Streaks multiply. Finish certified — or finish Standard.

DEPLOYED WITH

ALREADY ON
THE WAFER

“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”

John Crepezzi, AI Assistants, Jane Street
JANE STREETPODIUMBASISROGO

INSIDE OPENAI

OpenAI runs Ultrafast on its own fires: incident response that reads logs, analyses traces and synthesises conversations while evidence is still changing — and research loops tightened from overnight batches into multiple same-day iterations.

THE PARTNERSHIP

Ultrafast mode is the next step in the OpenAI–Cerebras partnership to bring ultra-low-latency inference to OpenAI’s platform — Cerebras now supporting OpenAI’s most intelligent model.

AVAILABILITY

Limited preview today for a select group of API customers, expanding as capacity grows. OpenAI is collecting access requests through a sign-up form linked from the announcement.

READ THE ANNOUNCEMENT ↗
Get started with Devin CLI