What hardware do you need to run gpt-oss-120b?
gpt-oss-120b at MXFP4 is 63 GB of weights, with a context ceiling of 128k tokens. The best reasoning per gigabyte on this list. Needs about 64 GB and rewards it.
How good is it, really?
On the Artificial Analysis Intelligence Index v4.3 it scores 12 (high reasoning effort; 10 at low), which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.
- Summarising — good
- Translation — usable
- Everyday coding — good
- Reasoning & maths — good
- Agentic work — usable
What it costs either way
Renting the same model costs $0.03 per million input tokens and $0.17 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-03). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.
Machines that run it
| Machine | Price | Speed at 32k | Pay-back | |
|---|---|---|---|---|
| Strix Halo Framework Desktop, 128GB | $3,449 | 35 tok/s measured | Pays back in 687 years | Run the numbers |
| DGX Spark GB10 Grace Blackwell, 128GB | $4,699 | 41 tok/s measured | Pays back in 922 years | Run the numbers |
| Mac Studio M5 Max, 128GB | $5,099 | 46 tok/s estimated | Pays back in 946 years | Run the numbers |
| MacBook Pro M5 Max (16-inch), 128GB | $6,999 | 46 tok/s estimated | Pays back in 1,299 years | Run the numbers |
One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.
The specifics
- Parameters
- 116.8B, of which 5.1B are active per token
- Quantisation
- MXFP4
- Weights on disk
- 63 GB
- KV cache
- 1.2 GB at 32k context — Half the layers use a 128-token sliding window; KV cache is tiny.
- Maximum context
- 128k tokens (128k)
- Licence
- Apache 2.0
- Sources
- source 1, source 2