AI Model Cost per Task 2026: Frontier vs Open-Weight Benchmarks
The short answer: in 2026 a single AI task costs anywhere from well under a cent to a few dollars, and the spread is almost entirely model choice. Published reference points (Artificial Analysis, Aug 2026) put Muse Spark 1.2 at ~$0.40 per task and Claude Opus 5 at ~$2.34 per task — a ~6x spread on the same class of work. This page benchmarks per-task cost across the frontier and open-weight models agencies actually route to, with methodology and dated sources.
Capacity context: Anthropic's reported $45 billion Nscale commitment and its SpaceX Colossus 1 compute deal signal real Claude capacity growth — but also pricing pressure ahead of the record IPO. For the full breakdown, see Anthropic's $45B Nscale compute deal: what it means for Claude capacity and pricing.
Published per-task reference points (Artificial Analysis, Aug 2026)
| Model | Reference cost per task | Source |
|---|---|---|
| Meta Muse Spark 1.2 | ~$0.40 | Artificial Analysis (Aug 2026) |
| Claude Opus 5 | ~$2.34 | Artificial Analysis (Aug 2026) |
| DeepSeek V4 Flash (pre-hike, historical) | ~$0.03 | Artificial Analysis (Aug 2026; pre-Aug 16, 2026 rates) |
These are vendor-independent benchmark figures, not our own measurements. The DeepSeek row is historical: it predates both the Aug 16, 2026 increase and the Sept 10, 2026 V4.1-Flash sheet that replaced it. On the current sheet the same 10K-in / 2K-out task costs about $0.0027 off-peak and about $0.0054 at peak (DeepSeek V4.1-Flash at $0.15/$0.60 off-peak) — read the $0.03 AA figure as a dated reference point, not the current rate.
Current per-1M rates (Aug 28, 2026; DeepSeek rows on DeepSeek's Sept 10, 2026 sheet) — the inputs for your own math
| Model | Input ($/1M) | Output ($/1M) | Notes |
|---|---|---|---|
| Qwen 3.8 Flash (Alibaba) | $0.15 | $0.47 | Open-weight 125B MoE, ~6B active (~95% sparsity), multimodal; qwen-community-1.0 license; cache read $0.016; 1M context (YaRN); verified Aug 27–28, 2026 |
| GPT-5.6 Luna | $0.20 | $1.20 | −80% Jul 30, 2026 (from $1/$6); 13.8x usage in the Jul 27 – Aug 14 discount window |
| DeepSeek V4.1-Flash (off-peak) | $0.15 | $0.60 | API name deepseek-flash (DeepSeek sheet, Sept 10, 2026); peak $0.30/$1.20; cache-hit input $0.003 off-peak / $0.006 peak. The retired deepseek-v4-flash / deepseek-v4-flash-vision-exp IDs are aliases billed at the Flash price |
| DeepSeek V4-Pro (off-peak) | $0.66 | $1.98 | Peak $1.32/$3.96; cache-hit input $0.022 off-peak / $0.044 peak |
| Gemini 3.8 Flash (intro) | $0.75 | $3.75 | Released Sept 2, 2026 — current Gemini Flash; intro through 2026-12-31, then $1.50/$7.50; 1M context |
| Gemini 3.7 Flash (intro — previous version) | $0.75 | $3.75 | Intro through 2026-12-31, then $1.50/$7.50; 1M context; superseded by 3.8 Flash Sept 2, 2026 |
| Meta Muse Glimmer (hosted, Together AI) | $0.35 | $1.50 | Local self-host = electricity after hardware amortized |
| GPT-5.6 Terra | $2.00 | $12.00 | −20% Jul 30, 2026 (from $2.50/$15); 5.6x usage in the discount window |
| Grok 4.6 (SpaceXAI) | $2.00 | $6.00 | 500K context; AA Intelligence Index 61; cache hit $0.50 |
| Qwen 3.8 Max (open weights) | $2.00 | $6.00 | Hosted API list price; open weights live Aug 12, 2026 |
| Claude Sonnet 5 | $2.00 | $10.00 | Permanent as of Aug 10, 2026; Sept 1 increase cancelled |
| Kimi K3 (Moonshot) | $3.00 | $15.00 | Open weights; $20M/12mo commercial-use threshold |
| GPT-5.6 Sol | $4.00 | $20.00 | Official OpenAI promo rate (Aug 21 – Nov 21, 2026); cached input $0.40; previously $5/$30 |
| GPT-6 Astra | $10.00 | $50.00 | NEW Sept 3, 2026 — OpenAI flagship; $1 cached input / $12.50 cache writes; prompts >272K input reprice the full request to $20/$75 per 1M; 1.05M context |
OpenAI's discount experiment (Jul 27 – Aug 14, 2026): when Terra and Luna were discounted 50% via OpenRouter, daily Terra token usage rose 5.6x and Luna 13.8x, while Sol at list price rose only 1.1x (control); ~1/3 of users who tried a discounted model kept using it after expiry. Terra/Luna went from 0.7% to 7.8% of OpenRouter tokens. For the agency-pricing breakdown, see OpenAI's discount experiment: what 13.8x usage growth means for agency pricing.
Cost per task on a typical 10K-in / 2K-out workload
Illustrative math (input tokens ÷ 1M × input price + output tokens ÷ 1M × output price), not a benchmark claim — your real token counts will differ:
| Model | Input cost | Output cost | Total per task |
|---|---|---|---|
| Qwen 3.8 Flash | $0.0015 | $0.0009 | ~$0.0024 |
| DeepSeek V4.1-Flash (off-peak) | $0.0015 | $0.0012 | ~$0.0027 |
| GPT-5.6 Luna | $0.002 | $0.0024 | ~$0.004 |
| Gemini 3.8 Flash (intro) | $0.0075 | $0.0075 | ~$0.015 |
| Gemini 3.7 Flash (intro — previous version) | $0.0075 | $0.0075 | ~$0.015 |
| Muse Glimmer (hosted) | $0.0035 | $0.0030 | ~$0.007 |
| Grok 4.6 | $0.020 | $0.012 | ~$0.032 |
| Qwen 3.8 Max | $0.020 | $0.012 | ~$0.032 |
| Claude Sonnet 5 | $0.020 | $0.020 | ~$0.040 |
| GPT-5.6 Terra | $0.020 | $0.024 | ~$0.044 |
| GPT-5.6 Sol | $0.040 | $0.040 | ~$0.080 |
| GPT-6 Astra | $0.100 | $0.100 | ~$0.200 — 2.5x Sol on this workload; over 272K input the full request reprices to $20/$75 per 1M |
At 10K tasks/month, that spread is $24/mo (Qwen 3.8 Flash) to $2,000/mo (GPT-6 Astra) on identical volume — GPT-6 Astra (Sept 3, 2026) at $10/$50 per 1M bills ~$0.20/task vs GPT-5.6 Sol's ~$0.08 (Sol's Aug 21 promo cut $5/$30 → $4/$20 dropped Sol from $1,100 to $800/mo on this volume). This is why model routing is a margin lever, not a footnote. The demand side of those prices is firming up too: Anthropic's reported $65B run rate signals a vendor with real pricing power behind Claude's per-token rates.
Coding-agent benchmark cost per task: SWE-2 (vendor-reported)
One benchmark row does not belong in the two tables above, because its unit is different. These are benchmark-run costs, not billable rates: the dollars spent running a full benchmark at list pricing including public discounts, measured by the model's own vendor.
| Model (effort) | FrontierCode 1.1 Main score | Benchmark-run cost per task |
|---|---|---|
| SWE-2 (medium) | 43.1% | $0.3712 |
| SWE-2 (high) | 47.2% | $0.7813 |
| SWE-2 (max) | 50.0% | $1.1761 |
| SWE-1.7 (Cognition's predecessor, no effort setting) | 42.0% | $1.9734 |
| Fable 5.1 (medium — the comparator Cognition names) | 50.9% | $3.2845 |
Basis: FrontierCode 1.1 Main, 100-task subset, three runs per task. Vendor-reported — Cognition ran it, and FrontierCode 1.1 is Cognition's own benchmark (methodology at cognition.com/frontiercode). The harnesses are not matched across models: SWE-2 ran in devin, SWE-1.7 in chisel, Fable 5.1 in claude-code. Every figure is read from Cognition's served dataset (cognition.com/data/swe-2/data.json, verified Sept 12, 2026); Cognition's two endpoints disagree for non-SWE-2 models (GPT-5.6 Sol is $2.10–$6.29 per task in this dataset against $1.89–$5.19 on Cognition's leaderboard, which its changelog attributes to promotional pricing), so the surface is declared here rather than assumed.
What these SWE-2 numbers do and do not say
- "Up to 70% lower cost" is Cognition's own wording, from Cognition's X post (Sept 10, 2026), and it is quoted here with "up to" intact. It is a bound across the top end of the benchmarks Cognition tested, not an exact per-task or per-benchmark figure — and the launch post itself contains no "70%" string at all.
- The 64% cheaper figure names its comparator: Fable 5.1. The effort pair behind it — SWE-2 max against Fable 5.1 medium — is our derivation, not vendor-stated: it is the only pair that reproduces 64% ($1.1761 against $3.2845 is 64.19% lower) while the two sit 0.91 points apart (50.0% vs 50.9%). Cognition does not publish the effort levels behind the claim.
- On Terminal-Bench 4.0 the gap to the frontier is wide. SWE-2 scores 27.3% against Fable 5.1's 55.8% — 28.5 points — and also trails GPT-6 Astra (57.9%) and GPT-5.6 Sol (37.3%). Any "parity with the frontier" reading is false on that benchmark. Terminal-Bench 4.0 is third-party (Stanford / Harbor / Laude); Cognition's own table labels it "Terminal-Bench 4". Source: Cognition's launch post table (verified Sept 12, 2026).
- The cleanest efficiency claim is against Cognition's own predecessor. On FrontierCode 1.1 Main, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average. Both are medium-effort figures, not an average over effort levels: SWE-2's mean across its three settings is $0.7762, which is 60.7% less than SWE-1.7 — so read the "81%" as the medium-effort row, not a blended average.
- Every SWE-2 figure on this page is vendor-reported. The 50.0% headline is on Cognition's own benchmark, and SWE-2's DeepSWE 1.1 and Terminal-Bench 4.0 rows exist only in Cognition's own reporting, not on those third-party benchmark owners' served pages.
- This row is not comparable to the tables above. It is a benchmark-run cost, not a pay-per-token API rate, so it must not be folded into the per-1M rate table or the 10K-in / 2K-out arithmetic — those price billed tokens, this prices a full benchmark run.
Picking the agent is a different question from costing its benchmark runs: which coding agent to run and how to budget for it — a calibration guide that deliberately carries no per-task dollar figure.
What the 2026 data says about routing
- Small structured tasks belong on cheap models. Published reference (AA): Muse Spark 1.2 ≈ $0.40/task vs Claude Opus 5 ≈ $2.34 — a 6x spread for the same class of work. Our 10K/2K math shows Qwen 3.8 Flash (~$0.0024/task), DeepSeek off-peak, Gemini 3.8 Flash intro (same rate as 3.7 Flash), and hosted Muse Glimmer under ~2 cents per task.
- Frontier costs cluster at the top. GPT-5.6 Sol at $4/$20 (official promo through Nov 21, 2026) is the highest per-task on the board; Grok 4.6 at $2/$6 matches Qwen 3.8 Max and undercuts GPT-5.6 Sol by ~2.5x on the illustrative task while matching its AA Intelligence Index of 61.
- DeepSeek's Sept 10, 2026 sheet resets the Flash tier. V4.1-Flash (API name
deepseek-flash) prices off-peak at $0.15 input / $0.60 output with $0.003 cache-hit input — below the Aug 16 sheet it superseded — and at exactly 2x those rates at peak ($0.30/$1.20, cache-hit $0.006). The retireddeepseek-v4-flashanddeepseek-v4-flash-vision-expIDs are still accepted and billed at the Flash price. DeepSeek V4 Pro keeps its own row at unchanged rates ($0.66/$1.98 off-peak), so the Flash tier and the Pro tier now price off different sheets. - Qwen 3.8 Flash resets the cheap-hosted tier. Released Aug 26, 2026 at $0.15/$0.47 per 1M on QwenCloud/OpenRouter ($0.016 cache reads), the open-weight 125B MoE activating ~6B per token (~95% sparsity) lands at ~$0.0024 on the 10K/2K task — below DeepSeek V4.1-Flash off-peak (~$0.0027) and GPT-5.6 Luna (~$0.004), and roughly a quarter of DeepSeek V4 Pro off-peak (~$0.0106).
- Sustained volume flips the answer to self-host. Open weights (Muse Glimmer 30B local, DeepSeek V4 MIT weights, Qwen 3.8 Max, Qwen 3.8 Flash-Next FP8 at 172.78 GiB, Kimi K3) run at electricity-cost marginal inference once hardware is amortized.
Routing decides which model reasons; it does not decide how the work is metered. Point the same routing table at a voice interface and the meter splits in two — duration billed by the second, reasoning billed by the token — so one call runs on two cost curves at once and a per-task figure stops being the whole answer. The two-meter voice agent cost model (voice seconds plus backend reasoning tokens) prices a cheap-reasoner preset and an expensive-reasoner preset against an identical $0.05-per-minute voice meter, so the pair is what you quote rather than the model alone.
How to compute your agency's real cost per task
- Instrument a pilot: log input/output tokens per task type for 100–500 real tasks.
- Multiply by your model's per-1M rate (use the table above; apply cache hits and off-peak where eligible).
- Sum per task type, then add retry/failure multipliers — real agent workflows rarely run clean the first time (levelsio reported ~$0.06 per request and "a few dollars per task" baselines, with failure modes turning into $500 loops and $2,000 overnight bills).
- Re-run quarterly — 2026 pricing moves weekly (Gemini 3.8 Flash Sept 2, GPT-5.6 Sol Aug 21, DeepSeek Aug 16 and its V4.1-Flash sheet Sept 10, Gemini 3.7 Flash Aug 13, Grok 4.6 Aug 12, Claude Sonnet 5 Aug 10).
And the model price is only one of three lines. Why cost per task is three lines, not one: model + tools + orchestration decomposes the same per-task unit into its model, tool and orchestration components.
Per-task math has one blind spot worth naming: it prices work, not time. A voice deployment bills two meters at once — realtime audio transport per connected minute, and the backend reasoning per turn — so a per-task figure understates the run. The voice agent two-meter cost model takes your calls per day and call length, bills the WebRTC session by the second with the 15-second init floor applied as a transport toggle, and returns cost per call, per day and per month with the month convention printed.
Model the blended cost of your exact task mix
Open the AI Agency Pricing Calculator →Setup fees, retainers, model strategy (DeepSeek, Gemini, Grok 4.6, Qwen, Sonnet 5, local), and margin — current 2026 rates.
Frequently asked questions
How much does an AI task cost in 2026?
It depends on the model and the task size. Published reference points (Artificial Analysis, Aug 2026): Muse Spark 1.2 ≈ $0.40 per task, Claude Opus 5 ≈ $2.34 per task, DeepSeek V4 Flash ≈ $0.03 per task at pre-hike rates (historical: the current V4.1-Flash sheet, effective Sept 10, 2026, puts the same 10K/2K task at ~$0.0027 off-peak and ~$0.0054 at peak). A 10K-token-in / 2K-token-out task on current 2026 rates ranges from ~$0.0024 (Qwen 3.8 Flash at its official $0.15/$0.47 rate) to ~$0.08 (GPT-5.6 Sol at its official $4/$20 promo rate), with frontier models like Claude Opus 5 and GPT-5.6 Sol at the high end for heavy tasks.
Which AI model is cheapest per task in 2026?
Among hosted APIs, Qwen 3.8 Flash ($0.15/$0.47 per 1M, ~$0.0024 per 10K/2K task) is now the cheapest per task, with DeepSeek V4.1-Flash off-peak ($0.15/$0.60 on DeepSeek's Sept 10, 2026 sheet, ~$0.0027 per 10K/2K task, cache-hit input $0.003) and Gemini 3.8 Flash intro pricing ($0.75/$3.75 through 2026-12-31, same rate as the 3.7 Flash it superseded) close behind for small work; Meta Muse Spark 1.2 at ~$0.40/task remains the published small-task reference. For sustained volume, self-hosted open weights (Meta Muse Glimmer local, DeepSeek V4 MIT weights) undercut every hosted API once hardware is amortized.
What is the cost per task for Grok 4.6 vs GPT-5.6 Sol?
On a 10K-token-in / 2K-token-out task, Grok 4.6 ($2/$6 per 1M) costs about $0.032 vs GPT-5.6 Sol (official $4/$20 per 1M promo through Nov 21, 2026) at about $0.08 — roughly 2.5x cheaper per task on that workload. Grok 4.6 also has a 500K context window and an AA Intelligence Index of 61, matching GPT-5.6 Sol max.
How much does GPT-6 Astra cost per task?
On a 10K-in / 2K-out task, GPT-6 Astra (launched Sept 3, 2026 at OpenAI's official $10/$50 per 1M, $1 cached input) costs about $0.20 — 2.5x GPT-5.6 Sol's ~$0.08 at Sol's promotional $4/$20. The cost cliff is long context: any prompt over 272K input tokens reprices the FULL request at $20 input / $75 output per 1M, so a 300K-in / 5K-out run bills roughly $6.375 vs about $3.25 at sub-272K rates (a ~96% jump). See GPT-6 Astra API Pricing for the full rate card.
How do I calculate my agency's cost per task?
Cost per task = (input tokens ÷ 1M × input price) + (output tokens ÷ 1M × output price). Measure real token counts per task type in a pilot, then multiply by the model's per-1M rate. Use cache hits and off-peak scheduling to cut input cost, and route small structured tasks to a cheap model while keeping frontier models for complex work.
Sources
- OpenAI developer docs — GPT-6 Astra model + pricing (verified Sept 3, 2026): developers.openai.com/api/docs/models/gpt-6-astra · developers.openai.com/api/docs/pricing
- Artificial Analysis, model pages + cost-per-task estimates (Aug 2026): artificialanalysis.ai/models
- DeepSeek official pricing page + V4.1-Flash launch note (live verified Sept 11, 2026 —
deepseek-flashsheet effective Sept 10, 2026: off-peak $0.15 input / $0.60 output / $0.003 cache-hit, peak exactly 2x;deepseek-v4-prounchanged): api-docs.deepseek.com/quick_start/pricing · deepseek.com/en/news/deepseek-v4-1-flash - Google, "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber" (Sept 2, 2026): blog.google · x.com/Google · Google, "Introducing Gemini 3.7 Flash" (Aug 13, 2026, previous version): blog.google
- SpaceXAI, "Grok 4.6" (Aug 12, 2026): x.ai/news/grok-4-6
- Anthropic, "Claude Sonnet 5" + pricing docs (Aug 10, 2026): anthropic.com
- Together AI, Muse Glimmer hosted pricing (verified Aug 12, 2026): together.ai/pricing
- OpenRouter Blog, "GPT 5.6 Discounts & Jevons Paradox" (Aug 25, 2026): openrouter.ai
- OpenAI API pricing (platform docs, verified Aug 28, 2026 — Terra $2/$12, Luna $0.20/$1.20, Sol $4/$20 promo): platform.openai.com/docs/pricing
- TLDR AI newsletter (Aug 28, 2026): tldr.tech
- QwenCloud, "Qwen3.8-Flash" official model & pricing page (verified Aug 27, 2026 — $0.15/$0.47/$0.016 per 1M): qwencloud.com/models/qwen3.8-flash
- OpenRouter, "Qwen3.8 Flash" API page + public models API (live listing, fetched Aug 28, 2026): openrouter.ai/qwen/qwen3.8-flash
- Hugging Face, "Qwen3.8-Flash-Next" model card (125B total / 6B active, multimodal MoE, qwen-community-1.0; released Aug 26, 2026): huggingface.co/Qwen/Qwen3.8-Flash-Next
- OrcaRouter, "Qwen 3.8 Flash release" (Alibaba’s Aug 28 announcement — ~95% sparsity, official QwenCloud rates): orcarouter.ai
- Hacker News / levelsio cost-per-task baseline (Aug 5, 2026): news.ycombinator.com/item?id=45931825
Accuracy note: GPT-6 Astra per-token pricing ($10/$50 per 1M, $1 cached input; >272K input reprices the full request to $20/$75) is OpenAI's official rate, verified Sept 3, 2026 from OpenAI's model + pricing docs; the ~$0.20 per 10K/2K task and ~$6.375 per 300K-in/5K-out run are our illustrative arithmetic on those official rates, not benchmark claims. The three per-task reference points (Muse Spark $0.40, Claude Opus 5 $2.34, DeepSeek V4 Flash $0.03) are Artificial Analysis figures from Aug 2026 and are attributed as such; the DeepSeek row is historical and predates both the Aug 16, 2026 increase and the Sept 10, 2026 V4.1-Flash sheet that replaced it. DeepSeek V4.1-Flash rates ($0.15/$0.60 off-peak, cache-hit input $0.003, peak exactly 2x at $0.30/$1.20/$0.006) are DeepSeek's own Sept 10, 2026 sheet, live-verified Sept 11, 2026; the retired deepseek-v4-flash and deepseek-v4-flash-vision-exp IDs are still accepted and billed at the Flash price, while deepseek-v4-pro keeps its own row at unchanged rates. The 10K/2K per-task column is our illustrative arithmetic on published per-1M rates — labeled as such, not a benchmark claim. GPT-5.6 Sol per-token pricing is OpenAI's official promotional rate ($4/$20 per 1M, cached input $0.40, valid Aug 21 – Nov 21, 2026; previously $5/$30). GPT-5.6 Terra ($2/$12) and Luna ($0.20/$1.20) are OpenAI's official Jul 30, 2026 list prices; the usage multiples (Luna 13.8x, Terra 5.6x, Sol 1.1x control) and ~1/3 retention are OpenRouter's measured discount-window data (Jul 27 – Aug 14, 2026), attributed as such (research brief t_5c8841ca). Per-1M rates are as of Aug 28, 2026 except the DeepSeek rows, which are on DeepSeek's Sept 10, 2026 sheet, and move frequently — re-verify before quoting. Qwen 3.8 Flash pricing is the official QwenCloud rate confirmed Aug 27, 2026, cross-checked on OpenRouter (fetched Aug 28); 125B total / 6B active and ~95% sparsity per the Hugging Face model card and Alibaba’s Aug 28 announcement (6/125 = 95.2% derived); Qwen benchmarks vendor-reported and unreproduced as of Aug 28. No hands-on model testing was performed for this page; all figures trace to named sources.