Project Peak
Project Peak usable context

How much of the advertised context window actually holds up.

A model can advertise a million-token window and still start getting things wrong long before it fills. We hold a code task at a fixed difficulty and grow the context until the model stops being reliable. The result, per model, is the longest context it holds at each usage bar, next to the window the model claims.

173K
GPT-5 · reliable to (Critical, Frontier)
43%
of its 400K window
Leaderboard

Usable context vs advertised window

Pick a difficulty and a deployment. Each model shows the longest context it stays reliable to at three usage bars, the share of its advertised window that represents, and how much reliable context you get per dollar. Cheaper and self-hosted models often only bend in window at the lighter difficulties, so the difficulty filter is where they show their real reach.

Difficulty
Sort by
Deployment

One row per model at the selected difficulty. Top context ranks by usable context at the Critical bar; Value ranks by reliable tokens per dollar of input price, and can surface a different (cheaper) reasoning level than Top context does. A star marks a bar whose lower confidence bound is not resolved at this sampling depth.