Controlled comparisons
Each study holds everything fixed but one variable and shows how it moves the result. The contenders are ranked by the hardest step each one holds, with a written analysis.
3 levels model analysis
GPT-5 family: usable context vs advertised window
How much of each GPT-5 model's advertised 400K window stays reliable, measured at three usage bars and the frontier difficulty tier. The smaller the model, the larger the gap between advertised and usable.
Read the study → 3 levels reasoning effort analysis
GPT-5-nano: does more reasoning buy more usable context?
GPT-5-nano measured at low, medium, and high reasoning effort, at a fixed difficulty tier. More reasoning shifts both usable context and the cost of a reliable result, so the cheapest level that clears your length can beat the deepest one.
Read the study →