Self-hosted no reasoning
window 33K
qwen2.5:7b
qwen2.5:7b does not bend in window at the Frontier difficulty. See the lighter tiers below for where it stays reliable.
Frontier difficulty
Hardest tasks, tuned to bend the strongest models
baseline 0% · value n/a
0 below floor window 33K
- Critical ≥95%
- below floor
- Right 95% of the time. Use it where a wrong answer is costly.
- Normal ≥85%
- below floor
- Right 85% of the time. A person reviews the output.
- Exploratory ≥75%
- below floor
- Right 75% of the time. Fine for drafts and ideation.
Reliability as context grows
2.9K
0% · 0/36
4.1K
0% · 0/16
8.2K
6% · 1/16
16K
0% · 0/16
33K
0% · 0/16
2.9K
4.1K
8.2K
16K
33K
Standard difficulty
A moderate task difficulty
baseline 0% · value n/a
0 below floor window 33K
- Critical ≥95%
- below floor
- Right 95% of the time. Use it where a wrong answer is costly.
- Normal ≥85%
- below floor
- Right 85% of the time. A person reviews the output.
- Exploratory ≥75%
- below floor
- Right 75% of the time. Fine for drafts and ideation.
Reliability as context grows
2.0K
0% · 0/16
2.9K
0% · 0/12
3.4K
0% · 0/8
4.1K
0% · 0/16
8.2K
0% · 0/16
16K
0% · 0/16
33K
0% · 0/16
2.0K
2.9K
3.4K
4.1K
8.2K
16K
33K
Light difficulty
Lighter tasks, where cheaper and self-hosted models still bend in window
baseline 0% · value n/a
0 below floor window 33K
- Critical ≥95%
- below floor
- Right 95% of the time. Use it where a wrong answer is costly.
- Normal ≥85%
- below floor
- Right 85% of the time. A person reviews the output.
- Exploratory ≥75%
- below floor
- Right 75% of the time. Fine for drafts and ideation.
Reliability as context grows
2.0K
0% · 0/16
2.9K
0% · 0/14
3.4K
0% · 0/6
4.1K
0% · 0/16
8.2K
0% · 0/16
16K
0% · 0/16
33K
0% · 0/16
2.0K
2.9K
3.4K
4.1K
8.2K
16K
33K
Usable context is the longest context a model stays at or above a usage bar, read from a fit that borrows strength across the difficulty tiers so the higher bars resolve without deep per-cell sampling. A star marks a bar whose lower confidence bound is not resolved at this depth. Value is reliable context per dollar of input price; self-hosted models are priced by amortized GPU-hours.