Sunk Cost sunkcost.ai Data checked 2026-09-03

gpt-oss-120b vs Llama 4 Scout 17B-16E

gpt-oss-120b scores higher on the intelligence index, 12 against 6. Both take the same machine to start: the cheapest here that runs either is the Framework Desktop, 128GB, at $3,449. On it, gpt-oss-120b is about 2.2× quicker: 35 tok/s against 16, both measured. At 500k tokens a day the Framework Desktop, 128GB pays for itself in 215 years running Llama 4 Scout 17B-16E, against 687 years running gpt-oss-120b.

gpt-oss-120bLlama 4 Scout 17B-16E
Intelligence index126
ClassBelow every hosted tierBelow every hosted tier
Weights63 GB65 GB
Needs at 32k65 GB68 GB
QuantisationMXFP4Q4_K_M
Parameters116.8B (5.1B active)108.6B (17B active)
Max context128k10240k
API price per 1M$0.03 in / $0.17 out$0.1 in / $0.3 out
LicenceApache 2.0Llama 4 Community License
Machines here that run it13 of 3713 of 37
Cheapest machine that runs itStrix Halo Framework Desktop, 128GB $3,449Strix Halo Framework Desktop, 128GB $3,449
Summarising good good
Translation usable good
Everyday coding good usable
Reasoning & maths good usable
Agentic work usable usable

Run gpt-oss-120b on the Framework Desktop, 128GB · or Llama 4 Scout 17B-16E

Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.

Side by side on the Framework Desktop, 128GB

The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Strix Halo Framework Desktop, 128GB, at $3,449.

gpt-oss-120bLlama 4 Scout 17B-16E
Speed at 32k35 tok/s measured16 tok/s measured
Pay-back on this machinePays back in 687 yearsPays back in 215 years
API cost per month$0.59$1.71

Run gpt-oss-120b on the Framework Desktop, 128GB · or Llama 4 Scout 17B-16E

How much use it takes to pay for the machine

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Framework Desktop, 128GB. Llama 4 Scout 17B-16E pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.

A day's usegpt-oss-120bLlama 4 Scout 17B-16E
50ka few chats a day6,875 years2,152 years
200klight assistant use1,719 years538 years
1Ma moderate coding-assistant day344 years108 years
4Mheavy coding with an agent86 years27 years
20Magents running most of the day17 years5.4 years

Run gpt-oss-120b at 20M tokens a day · or Llama 4 Scout 17B-16E

Memory is not what separates them

gpt-oss-120b needs 65 GB of memory at 32k of context and Llama 4 Scout 17B-16E needs 68 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.

gpt-oss-120b is also head to head with Qwen3.5 4B above it on the leaderboard and Ling 3.0 tiny below it. Llama 4 Scout 17B-16E is also head to head with Qwen3 14B above it on the leaderboard and Ministral 3 14B below it.

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.

More head to head: every machine that runs gpt-oss-120b · every machine that runs Llama 4 Scout 17B-16E · every other match-up · both against the frontier · the quickest pay-back at each level of use