KEY TAKEAWAYS
- Use the price announcement to choose a candidate, not to forecast an equivalent reduction in total costs.
- Compare all spending and review effort across the same assigned tasks, including failures, then divide by accepted outputs.
- Keep the current model—or defer switching—when acceptance, access, risk or migration costs do not support the change.
Evaluate a workload, not a price ratio
For a small team maintaining a coding or document agent, the useful question is not whether a new model has cheaper tokens. It is whether the same queue of work produces enough acceptable outputs at lower total cost, without unacceptable failures or delay. Our recommendation is a bounded matched-task evaluation, not a wholesale switch.
OpenAI reports that GPT-6.1 Sol matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth of the cost. That is a vendor-reported result under evaluation conditions. OpenAI says its GPT evaluations used its research environment or API, and competitor results came from public reports; those comparisons are not one independently replicated production trial.[1] GitHub’s rollout announcement establishes another access route, not independent confirmation of the claimed cost-effectiveness.[2]
Four numbers that answer different questions
| Measure | What it tells you | What it cannot establish |
|---|---|---|
| API unit price | A charge per million tokens in a specified category and service tier.[9] | Tokens, retries, tools and review needed for your output. |
| Copilot usage charge | Model token consumption converted to AI credits; plan allowances affect additional billing.[3] | Your marginal invoice saving without your plan and usage history. |
| Benchmark cost per trial/task | Spend under the benchmark’s task set, harness, effort and accounting. DeepSWE v1 describes median cost per trial; v1.1 displays average cost.[4][7] | Cost per human-accepted production task, or a statistic interchangeable across benchmarks. |
| Accepted-task cost | Our proposed accounting: all measured workload cost divided by distinct accepted tasks in the same assigned cohort. | Safety, business value or a causal productivity gain on its own. |
The current OpenAI standard short-context table lists GPT-6.1 Sol at $2 input, $0.10 cached input, $2.50 cache writes and $10 output per million tokens; GPT-6 Astra is $10, $1, $12.50 and $50 respectively.[9] Calculation: Sol’s input and output rates are one-fifth of Astra’s, but its cached-input rate is one-tenth. There is no single discount to apply to every token mix. Both the API table and Copilot docs list long-context and cache-write pricing beyond the launch’s three-number summary.[3][9]
Copilot’s GPT-6.1 Sol default tier applies at up to 272K input tokens; the listed long-context tier above that threshold is $4 input, $0.20 cached input, $5 cache write and $15 output per million tokens.[3] A budget must use the actual billed token categories and applicable tier, not count all input as both uncached and cached. These tables are not a complete invoice estimate: tools, infrastructure and applicable service adjustments still matter.
GitHub defines one AI credit as $0.01 and bills usage beyond included allowances at the documented token rates. Existing annual Pro/Pro+ request-based plans have a separate legacy multiplier reference; Copilot code review also consumes Actions minutes.[3] The launch names Pro+, Max, Business and Enterprise, with gradual rollout and administrator controls.[2] Confirm your account’s access and billing mode before assigning an evaluation budget. Lower metered consumption may preserve an allowance without reducing this month’s cash bill.
What the benchmarks help you test
DeepSWE’s original method uses a shared mini-swe-agent harness with a bash tool and shared prompt. Its corpus covers 113 tasks across 91 open-source repositories and five languages; the authors warn about generalization to proprietary repositories and under-representation of bug localization and refactoring.[4] Version 1.1 keeps the tasks but grades committed patches in a separate clean container with named test reports.[7] A Copilot workflow with a different harness is therefore a different system to evaluate, even when its model name matches.
One inspected artifact makes the acceptance problem concrete. DeepSWE’s Cliffy config-parsing prompt requires CLI arguments to override environment variables, which override config values. It also explicitly requires false and zero to remain valid values and subcommands to inherit parent configuration.[5] Our inference: a patch that parses a file but drops zero-valued settings is not an accepted solution. Use similarly explicit edge cases in your own evaluation. We inspected the prompt; we did not execute its verifier or audit a model’s patch.
For document work, Surge describes GDP.pdf as 100 professional prompts and PDFs across ten domains.[6] Its harness distinguishes all_pass/mean—every rubric criterion satisfied—from mean_criteria/mean, the average fraction satisfied, and specifies five runs per task with Gemini 3.5 Flash as judge for its leaderboard configuration.[8] Partial rubric credit is not the probability that a whole document answer is acceptable. Nor should five attempts with at least one success be read as five consistently successful deliveries.
These methods suggest tests, not a universal winner. DeepSWE’s v1.1 report no longer presents wall-clock time because host performance and provider load affect it; that does not make latency irrelevant to your users.[7] OpenAI’s own launch cautions that its factuality prompts are deliberately difficult rather than typical usage, and flags omitted fallback costs in one AutomationBench comparison.[1] Our inference: preserve failure accounting and measure local latency rather than importing a favorable percentage or averaging incompatible benchmark scores.
Worked example: the cheaper attempt can cost more
The following is entirely hypothetical, not a measurement of GPT-6.1 Sol, Astra, Copilot or any customer. Two generic workflows receive the same 100 tasks during the same evaluation window. Acceptance is decided once per distinct task after the allowed retry process; a task still unresolved at cutoff does not count as accepted. All attempts and human time, including rejected work, stay in the numerator. A and B are assumed to satisfy the same minimum risk and deadline requirements.
| Same 100 assigned tasks | Current A | Candidate B |
|---|---|---|
| Total attempts, including retries | 120 | 150 |
| Assumed average model charge per attempt | $1.00 | $0.20 |
| Total model charges | $120 | $30 |
| Total review and repair minutes | 300 | 420 |
| Allocated human effort cost | $300 | $420 |
| Other tools/infrastructure cost | $30 | $30 |
| Total cost | $450 | $480 |
| Distinct accepted tasks | 90 | 80 |
| Cost per accepted task | $5.00 | $6.00 |
Calculation: A costs ($120 + $300 + $30) / 90 = $5 per accepted task. B costs ($30 + $420 + $30) / 80 = $6. Its assumed model charge per attempt is 80% lower, and its total model charge is 75% lower, but its accepted-task cost is 20% higher. The 80 accepted outputs are also less service delivered than A’s 90; the ratio must not hide the unresolved tasks.
Sensitivity: holding B’s 80 accepted tasks, $30 model charge and $30 other cost fixed, it ties A’s $5 rate at 340 total human minutes: ($30 + $340 + $30) / 80 = $5. It must fall below 340 minutes to beat that rate—more than 80 minutes below the illustrative 420. That is a conditional break-even threshold, not a forecast that review time can fall without changing acceptance.
For your own cohort, calculate cost per accepted task as (model charges for every attempt + tools/infrastructure + measured human minutes × hourly value / 60 + an explicitly allocated share of fixed and migration costs) / accepted task count. If no tasks are accepted, report the spend and zero accepted outputs, not a finite unit cost. Keep cash outlay and time-valued cost separate, do not double-count human repair, and report unresolved work and late-discovered defects alongside the ratio.
A small matched-task evaluation to run next
Proposal—not an executed experiment: begin with 20 recent, non-sensitive tasks from one recurring workflow, selected before seeing candidate results. For coding, use 10 ordinary changes, five multi-file changes and five known edge-case or recovery tasks; use a separate document cohort if document work is the decision. This deliberately stratified screening set is not an estimate of the production task mix or a statistically powered proof of equivalence.
- Freeze the comparison. Record task IDs, repository/document versions, model IDs, prompt, tools, harness version, reasoning effort, service tier and cache policy. Baseline A is the currently deployed workflow. If comparing two products rather than swapping a model in the same harness, call the result a workflow comparison; do not attribute all differences to the model.
- Define acceptance first. For code, require the requested behavior, edge cases, regression tests and a reviewer’s maintainability check. For documents, require every decision-critical answer and source/page check; partial rubric credit is diagnostic only. Name disqualifying privacy, security or material factual failures separately from the cost metric.
- Run each task from an isolated identical starting state in both arms. Randomize or alternate run order, hide model labels from reviewers where feasible, and prevent one arm’s answer from becoming the other’s context. Predeclare one initial attempt plus at most one repair attempt and a task time/spend ceiling. Keep timeouts, outages, refusals and unresolved tasks in the record.
- Record all billed categories and attempts, fallbacks, tool charges, reviewer and repair minutes without overlap, elapsed time, acceptance and failure reason. A fallback to the current model remains part of the candidate workflow’s cost. Never discard an expensive failed attempt because its successor passed.
- For each task pair, inspect A-only and B-only passes and review-time differences. Report accepted counts, total spend, cash-only and time-valued unit costs, latency and critical failures by stratum. Repeat borderline or unstable cases under a declared rule; do not keep rerunning only B until it wins.
- Set decision limits before execution: a workload-specific acceptance floor, maximum permitted regression, latency ceiling and minimum saving large enough to cover switching effort. Use the screen to reject a poor candidate or justify a larger representative trial—not to certify rare-failure safety from 20 tasks. No paid execution is part of this research article.
Choose a limited switch, stay, or defer
| Choice | Evidence that supports it |
|---|---|
| Test a limited switch | The candidate is accessible and a matched evaluation meets acceptance, critical-error and latency limits, with savings that survive measured review/retry costs and migration allocation. Start with the task class actually tested. |
| Stay with the current model | Candidate-only failures affect important tasks; extra review cancels model savings; or the current workflow already fits a paid allowance and switching adds cost without a useful capacity gain. |
| Defer the decision | Billing or acceptance cannot be measured, access is unavailable, the sample is unrepresentative, or results change materially across repeat runs. Repair that evidence gap before broad rollout. |
The countercase to caution is real: if tokens dominate spending and the candidate preserves acceptance and review time, cheaper rates can materially reduce cost. Conversely, a more capable model can be worth its higher rate if it avoids expensive repair. Neither case is established for your workflow by these sources. Evidence that would change our recommendation is a repeated, representative matched comparison showing lower total cost at the required acceptance and risk limits—not a second announcement quoting the first.
Sources & scope
Documentary research checked 1 October 2026 KST: OpenAI launch and pricing, GitHub Copilot billing, DeepSWE methods and one task prompt, and GDP.pdf methods and harness documentation. Sources were selected to explain cost definitions, not to rank the market. No models were run, customer results measured or independent replication of the launch established. The calculation is hypothetical; the evaluation is proposed. Publication/update dates not established by a source remain unknown.
- OpenAI — Introducing GPT-6.1 Sol ↗
Vendor launch and evaluation caveats; read 1 October 2026 KST. Publication/update date not established from the inspected page. Not independently replicated here.
- GitHub — GPT-6.1 Sol in GitHub Copilot ↗
Published 29 September 2026, 17:02:27 UTC; modified 17:34:06 UTC according to page metadata. Rollout/access report, not a controlled effectiveness study.
- GitHub Docs — Models and pricing for GitHub Copilot ↗
Live billing reference read 1 October 2026 KST; final URL includes /en/. Update date unknown. Account-specific allowances not measured.
- Datacurve — DeepSWE original methodology ↗
Byline dated 26 May 2026; historical v1 methods and limitations. Not a current Sol ranking or a production acceptance study.
- DeepSWE task — Add config file parsing to Cliffy commands ↗
Actual public v1 prompt artifact, identifier 132a437c40; read 1 October 2026 KST. Publication date unknown. No verifier executed.
- Surge AI — GDP.pdf benchmark introduction ↗
HTML datePublished is 14 April 2026; page also displays 16 September, whose precise update role is unconfirmed. Used for task scope, not its historical rankings.
- Datacurve — DeepSWE v1.1 ↗
Byline dated 14 June 2026; displayed leaderboard update 3 September 2026. Execution/grading revision and reference comparison inspected; exact measurement windows unknown.
- Surge AI — GDP.pdf harness README ↗
Public README and raw counterpart read 1 October 2026 KST. Metrics and leaderboard configuration inspected, not executed. Repository revision date unknown; sample pack is a synthetic placeholder.
- OpenAI API — Pricing ↗
Current standard short/long-context token categories read 1 October 2026 KST. Update date unknown; other modes and regional adjustments mean the table is not a full invoice estimate.
Publication history
- — Prepared initial research version. Version timestamp is not evidence of live publication.
Have a correction or a different perspective? Contact Palanthos.