Project Peak
← Leaderboard
Hosted high reasoning window 400K $0.05/Mtok input

GPT-5 nano

At the Frontier difficulty, GPT-5 nano stays reliable to 4.5K of its 400K advertised window for critical use (1% of the window), at 806K /$ of reliable context.

Frontier difficulty

Hardest tasks, tuned to bend the strongest models

baseline 94% · value 806K /$
0 reliable to 4.5K window 400K
Critical ≥95%
4.5K *
CI 2.9K to 9.0K
Right 95% of the time. Use it where a wrong answer is costly.
Normal ≥85%
9.0K *
CI 2.9K to 17K
Right 85% of the time. A person reviews the output.
Exploratory ≥75%
13K
CI 5.2K to 20K
Right 75% of the time. Fine for drafts and ideation.
Reliability as context grows
2.9K
4.1K
8.2K
16K
33K
66K
131K
262K

Standard difficulty

A moderate task difficulty

baseline 100% · value 18M /$
0 reliable to 262K+ window 400K
Critical ≥95%
262K+ *
Right 95% of the time. Use it where a wrong answer is costly.
Normal ≥85%
262K+ *
Right 85% of the time. A person reviews the output.
Exploratory ≥75%
262K+
Right 75% of the time. Fine for drafts and ideation.
Reliability as context grows
2.0K
2.9K
4.1K
8.2K
16K
33K
66K
131K
262K

Light difficulty

Lighter tasks, where cheaper and self-hosted models still bend in window

baseline 100% · value 18M /$
0 reliable to 262K+ window 400K
Critical ≥95%
262K+ *
Right 95% of the time. Use it where a wrong answer is costly.
Normal ≥85%
262K+ *
Right 85% of the time. A person reviews the output.
Exploratory ≥75%
262K+
Right 75% of the time. Fine for drafts and ideation.
Reliability as context grows
2.0K
2.9K
4.1K
8.2K
16K
33K
66K
131K
262K

Usable context is the longest context a model stays at or above a usage bar, read from a fit that borrows strength across the difficulty tiers so the higher bars resolve without deep per-cell sampling. A star marks a bar whose lower confidence bound is not resolved at this depth. Value is reliable context per dollar of input price; self-hosted models are priced by amortized GPU-hours.

Reasoning levels measured
reasoning reliable (Critical) % of window reliable ctx / $
high reasoning 4.5K 1% 806K /$
medium reasoning 3.6K 1% 1.8M /$
low reasoning 3.3K 1% 4.3M /$

Higher reasoning usually buys more usable context but costs more. The leaderboard lists this model once and the Value sort surfaces its most cost-efficient level. Figures are at the Frontier difficulty tier.

Secondary metric: difficulty ceiling

Usable context above is the primary measure. As a secondary, difficulty-axis view (how hard a coupled reasoning task it sustains as context grows), GPT-5 nano scores 33.6/100.