- Critical · 95%
- 173K *
- Normal · 85%
- 262K+
- Exploratory · 75%
- 262K+
- Value
- 652K /$
How much of the advertised context window actually holds up.
A model can advertise a million-token window and still start getting things wrong long before it fills. We hold a code task at a fixed difficulty and grow the context until the model stops being reliable. The result, per model, is the longest context it holds at each usage bar, next to the window the model claims.
Usable context vs advertised window
Pick a difficulty and a deployment. Each model shows the longest context it stays reliable to at three usage bars, the share of its advertised window that represents, and how much reliable context you get per dollar. Cheaper and self-hosted models often only bend in window at the lighter difficulties, so the difficulty filter is where they show their real reach.
- Critical · 95%
- 29K *
- Normal · 85%
- 63K
- Exploratory · 75%
- 165K
- Value
- 1.3M /$
- Critical · 95%
- 4.5K *
- Normal · 85%
- 9.0K *
- Exploratory · 75%
- 13K
- Value
- 806K /$
- Critical · 95%
- below floor
- Normal · 85%
- below floor
- Exploratory · 75%
- below floor
- Value
- n/a
- Critical · 95%
- 3.3K *
- Normal · 85%
- 4.4K
- Exploratory · 75%
- 5.8K
- Value
- 4.3M /$
- Critical · 95%
- 29K *
- Normal · 85%
- 63K
- Exploratory · 75%
- 165K
- Value
- 1.3M /$
- Critical · 95%
- 173K *
- Normal · 85%
- 262K+
- Exploratory · 75%
- 262K+
- Value
- 652K /$
- Critical · 95%
- below floor
- Normal · 85%
- below floor
- Exploratory · 75%
- below floor
- Value
- n/a
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 744K /$
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 3.7M /$
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 18M /$
- Critical · 95%
- below floor
- Normal · 85%
- below floor
- Exploratory · 75%
- below floor
- Value
- n/a
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 18M /$
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 3.7M /$
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 744K /$
- Critical · 95%
- below floor
- Normal · 85%
- below floor
- Exploratory · 75%
- below floor
- Value
- n/a
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 760K /$
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 3.8M /$
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 18M /$
- Critical · 95%
- below floor
- Normal · 85%
- below floor
- Exploratory · 75%
- below floor
- Value
- n/a
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 18M /$
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 3.8M /$
- Critical · 95%
- 262K+ *
- Normal · 85%
- 262K+ *
- Exploratory · 75%
- 262K+
- Value
- 760K /$
- Critical · 95%
- below floor
- Normal · 85%
- below floor
- Exploratory · 75%
- below floor
- Value
- n/a
One row per model at the selected difficulty. Top context ranks by usable context at the Critical bar; Value ranks by reliable tokens per dollar of input price, and can surface a different (cheaper) reasoning level than Top context does. A star marks a bar whose lower confidence bound is not resolved at this sampling depth.