Project Peak
The Premise

Advertised windows describe capacity. They do not describe capability.

A model that advertises a 400K-token window can read 400K tokens. Whether it can actually reason over them is a different question, and that is the one that matters when the output ships. Hard work means tracing dependencies across a long file and pulling together facts that sit far apart. In practice the range where a model holds a difficult task is much smaller than the range where it can simply read, the gap is different for every model, and nothing on a spec sheet tells you where it falls.

Project Peak measures this directly. We hold a task at a fixed difficulty and sweep the input from short to long, then report the longest context the model stays reliable to, at three usage bars: Critical (right 95% of the time, for work you cannot get wrong), Normal (85%), and Exploratory (75%). We put that usable context next to the advertised window, so the headline is the gap: how much of the window you can actually trust. Tasks are synthetic code problems that are scored exactly.

The goal is narrow and concrete. We want a defensible per-model number for how much context each model is reliable to, at a chosen reliability bar, reported with confidence intervals and a falloff shape, that you can route, budget, and compare against. A difficulty-axis score (how hard a task it sustains, 0 to 100) is kept as a secondary view. It is a measured number rather than a vibe or a one-off demo.