Delphi Digital
Delphi Digital|Aug 19, 2026 16:01
A model can be better at the task and worse for the business. Model performance has to be judged against what the job requires. Frontier models earn their premium on work that cheaper models cannot complete reliably. Using them for routine tasks increases the cost without improving the result. In a task market, the buyer agrees on a price for a verified result, leaving the provider to decide how to deliver it profitably. OSWorld 2.0 evaluates agents on multi-step computer tasks. At the maximum step budget, Claude Opus 4.7 reached a 49% partial score across 108 tasks at a total cost of roughly $3,900. That works out to $79 per score point. MiniMax M3 reached 22% for about $260, or $12 per point. Opus delivered around 2.3 times the score at roughly six times the cost per score point. Its premium can be justified when the additional capability changes the outcome. Using it across an entire workflow would erode margins wherever a lower-cost model is sufficient. More compute keeps helping, but each increment buys less. The cheap models flatten below 25% regardless of spend. Routing allows providers to reserve frontier models for difficult work and send routine steps to lower-cost alternatives. They can also purchase inference or data through pay-per-call marketplaces and add human review when judgment is required. Buyers still see one price for the completed task, and providers keep the upside from delivering it more efficiently. Cost per task captures the combined effect of model choice, external services, retries, and verification. It determines how low a provider can price its work and how much margin remains after delivery.(Delphi Digital)
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads