Project Peak
Studies

Controlled comparisons

Each study holds everything fixed but one variable and shows how it moves the result. The contenders are ranked by the hardest step each one holds, with a written analysis.

3 levels model analysis

GPT-5 family: usable context vs advertised window

How much of each GPT-5 model's advertised 400K window stays reliable, measured at three usage bars and the frontier difficulty tier. The smaller the model, the larger the gap between advertised and usable.

Read the study →
3 levels reasoning effort analysis

GPT-5-nano: does more reasoning buy more usable context?

GPT-5-nano measured at low, medium, and high reasoning effort, at a fixed difficulty tier. More reasoning shifts both usable context and the cost of a reliable result, so the cheapest level that clears your length can beat the deepest one.

Read the study →