Skip to content
Regolo Logo
Benchmarks & Cost Optimization

Best AI Coding Models in 2026: Claude Opus 5.5, GPT-6 Astra, GLM-5.3, and Qwen3.8-27B

Alex Genovese
9 min read
Share

Claude Code with Claude Opus 5.5 scores 66 points on the Artificial Analysis Coding Agent Index v1.5. Codex with GPT-6 Astra scores 62. During the same period, Qwen3.8-27B does not appear on the full coding-agent leaderboard, while its model card reports 61.7 on SWE-bench Pro and 84.3 on OSWorld-Verified. GLM-5.3 reports 66.9 on DeepSWE v1.1 and 88.2 on Terminal-Bench 2.1.

This guide compares the best AI coding models for agentic software work in 2026, including benchmark performance, cost per task, and European data controls. A leaderboard score combines the model, harness, reasoning settings, and provider. Start with the team’s real tasks, then measure reliability and total operating cost.

Key takeaways

  • Claude Opus 5.5 leads the Coding Agent Index v1.5 with 66 points, ahead of GPT-6 Astra at 62.
  • GLM-5.3 performs well on coding and long-running tasks. Its advantage still needs to be tested on the real workload and chosen infrastructure.
  • Qwen3.8-27B combines 27 billion parameters, visual capabilities, and the Apache License 2.0. Its OSWorld-Verified result makes it interesting for computer use.
  • API price is not total cost. Count retries, latency, supervision, infrastructure, and failures.
  • Zero Data Retention and European residency reduce specific risks, but they do not make an application compliant by law.
SOVEREIGN EUROPEAN INFERENCE

Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention

Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

Start 30-day free trial (no card required) → Free credits included · Live in 60s
✓ 100% EU Green Datacenters ✓ Certified ZDR

Numbers at a glance

The values below were checked on September 24, 2026. Model versions and settings affect the results.

MetricClaude Opus 5.5GPT-6 AstraGLM-5.3Qwen3.8-27B
Artificial Analysis Coding Agent Index v1.566 with Claude Code62 with CodexAbout 54 with OpenCode, based on the public chart snapshotNot listed on the full leaderboard
Artificial Analysis Intelligence Index v4.3.258534533.7, displayed as 34
Cost per Intelligence Index task$5.98$3.26$2.01$1.01
LicenseProprietaryProprietaryGLM-5.3 License, with conditionsApache 2.0
Strongest cited settingMax effort with fallbackMax effortMax effortXhigh

The Coding Agent Index compares complete systems. A score of 66 for Opus 5.5 is not a model-only score. It also includes Claude Code, the context given to the agent, the attempts, and provider routing. The OpenCode row with GLM-5.3 therefore does not automatically match the same model running in another harness.

The task cost in the table belongs to the Intelligence Index. It is not the average cost of a coding-agent session and does not include infrastructure, human intervention, or application retries.


How to read the AI coding agent benchmark

The current AI coding agent benchmark combines three tests with equal weight:

  1. DeepSWE v1.1: 113 software-engineering tasks.
  2. Terminal-Bench 4.0: 66 agentic terminal-use tasks.
  3. SWE-Atlas-QnA: 124 repository-understanding and technical Q&A tasks.

Artificial Analysis calculates average pass@1 across three attempts per task and gives the three components equal weight. The result provides a quick comparison between agents, but it does not replace an internal test.

The harness matters: the same model may receive different instructions, use different tools, and retry under different conditions


Where GLM-5.3 approaches the frontier

Z.ai’s GLM-5.3 report says the model uses the same base as GLM-5.2 and that the improvements come from post-training. The results published by the vendor show the jump:

IMAGE: source GLM report

KPI Highlights

BenchmarkGLM-5.2GLM-5.3Change
DeepSWE v1.146.266.9+20.7
Terminal-Bench 3.04.628.3+23.7
Agents’ Last Exam23.828.5+4.7
AutomationBench v1.0.626.248.2+22.0

The same chart gives GLM-5.3 an 88.2 score on Terminal-Bench 2.1, these benchmark versions must not be mixed with Terminal-Bench 4.0, which is part of the current Artificial Analysis Coding Agent Index. A score of 28.3 on Terminal-Bench 3.0 cannot be compared directly with a score on 4.0.

GLM-5.3 releases open weights under a custom license and It allows use, modification, distribution, and commercial use. However, the license requires a company that operates a Model as a Service and exceeds $10 billion in aggregate revenue over 12 months to undergo a Z.AI security review and legal teams should assess this condition before a Model as a Service deployment.

Qwen3.8-27B: a compact model for coding and computer use

Qwen’s model card presents Qwen3.8-27B as a dense 27-billion-parameter model with a visual encoder. It supports images and video, uses a native 262,144-token context window, and can be extended to one million tokens with RoPE scaling. The weights are released under the Apache License 2.0.

IMAGE: source Qwen model card

KPI Highlights

BenchmarkQwen3.8-27BWhat the result measures
Terminal-Bench 2.173.0Agentic terminal coding with the Terminus harness
SWE-bench Pro61.7Repository changes and software issue resolution
DeepSWE v1.142.2Software-engineering tasks on repositories
OSWorld-Verified84.3Computer use on desktop interfaces
IFBench79.5Instruction following
GPQA Diamond89.2Scientific reasoning
LiveCodeBench v690.3Competitive programming

Many tests used Claude Code and the settings listed in the model card, they help explain the stated capabilities but they are not an independent leaderboard identical to the Artificial Analysis Coding Agent Index.

The OSWorld-Verified result makes Qwen3.8-27B especially interesting for agents that need to use browsers, desktop applications, and visual tools. The Apache License 2.0 provides a simpler route to self-hosting and redistribution than a custom license.

Regolo’s Qwen3.8-27B benchmark analysis reviews the vendor table claim by claim, Its harness engineering guide shows how the same checkpoint can produce different results across coding agents.


Costs: from token price to task cost

Artificial Analysis measures cost with a consistent method across its own suite. On September 24, 2026, the Intelligence Index values were:

ModelIntelligence IndexCost per taskOutput tokens per task, when reported
Claude Opus 5.5 max58$5.98About 119,000
GPT-6 Astra max53$3.26About 27,000
GLM-5.3 max45$2.01About 71,000
Qwen3.8-27B xhigh34$1.01About 67,000

Regolo’s AI model comparison and pricing hub tracks provider prices, benchmark results, context limits, and serving assumptions.

In this suite, GLM-5.3 and Qwen3.8-27B have a lower Intelligence Index cost per task than Opus 5.5 and Astra. That does not prove that an open-weight model will always cost less.


US frontier APIs and EU inference: the compliance difference

US frontier APIs are not automatically outside GDPR. OpenAI says API data is not used for training by default, offers a DPA, and supports European processing plus Zero Data Retention for eligible projects and endpoints. Anthropic also provides contractual safeguards and ZDR for eligible features. Its privacy documentation says customer data is stored in the US by default, even when traffic can be routed to selected regions.

Regolo’s current model catalog lists Qwen3.8-27B and GLM-5.2, GLM-5.3 availability depends on a custom deployment review for now.

The comparison below describes the EU-hosted open-weight architecture

The architectural difference is the default path a European buyer can prove, a US provider can offer compliant processing, but the review may cover a US legal entity, subprocessors, system data, endpoint-specific retention, and transfer safeguards. EU-hosted inference can keep customer content, inference, and operational control in one jurisdiction.

That can simplify an EU-only policy and the evidence required under Articles 28 and 44 of the GDPR.

ControlUS frontier APIsEU-hosted open-weight inference on Regolo
Model qualityOften the strongest closed models, available through managed APIsDepends on the selected open model, quantization, and serving stack
EU processingOpenAI offers EU regional processing for eligible projects. Anthropic supports traffic routing, while documenting US storage by default.Regolo states that inference runs in European data centers and does not leave the EU region
Payload retentionOpenAI logs may be retained for up to 30 days by default; ZDR requires approval. Anthropic ZDR depends on the feature and model.Regolo states that prompts and responses are processed in volatile memory and are not retained after the response
Training and reuseOpenAI API data is not used for training by default. Anthropic retained data is not used without express permission.Regolo states that customer data is not reused for training or fine-tuning
Governance evidenceDPA, subprocessors, transfer mechanism, endpoint limits, and US jurisdiction must be reviewedEU DPA, certificate scope, infrastructure location, and operational metadata must still be reviewed

Regolo’s Zero Data Retention documentation and its independent certificate provide stronger evidence for an EU-only architecture than a generic “GDPR compliant” badge. They do not make the customer application compliant by themselves. The GDPR still requires a lawful basis, minimization, security, and appropriate contracts and the EU AI Act still depends on the system’s role and risk classification.

Latency: EU locality helps, measurement decides

Geography is one part of latency, a longer network path adds propagation delay and more routing hops – for a European application, an EU inference endpoint can avoid a transatlantic round trip and keep the request path inside one jurisdiction.

Model size, prompt length, reasoning effort, queueing, cache state, provider routing, and cold starts can dominate, the live Regolo OBS Qwen3.8-27B health page showed a 2.074-second average across its last 500 health pings, 95.8% success, and one recent slow ping of 12.432 seconds at 08:37 UTC on September 24, 2026.

When to choose each model family

This table separates measured model performance from provider posture. It is a decision aid, not a legal conclusion.

Legend: green marks an EU or Regolo advantage; orange marks an open-weight cost or capability advantage, or a close comparison; white marks an American-provider advantage; red marks a non-EU-only default path or a contractual transfer review requirement.

Decision criterionClaude Opus 5.5GPT-6 AstraGLM-5.3Qwen3.8-27B
Reasoning composite
Intelligence Index v4.3.2, same suite
US advantage
58 points
US advantage
53 points
Open-weight comparable
45 points
Open-weight comparable
33.7 points, displayed as 34
Agent composite
Coding Agent Index v1.5, same suite
US advantage
66 with Claude Code
US advantage
62 with Codex
Open-weight comparable
About 54 with OpenCode, public chart snapshot
Not listed
No full coding-agent leaderboard entry in checked sources
Long-horizon repo coding
DeepSWE v1.1, same version; harness differs
Not disclosed
No DeepSWE v1.1 figure in checked sources
US advantage
68 in Codex harness, per Artificial Analysis
Closest open signal
66.9, vendor-reported by Z.ai
Open-weight signal
42.2, vendor-reported by Qwen
Closest open coding signals
Vendor-reported; no frontier match in checked sources
Not disclosed
No matching SWE-bench Pro or OSWorld figure in checked sources
Not disclosed
No matching SWE-bench Pro or OSWorld figure in checked sources
Terminal signal
88.2 on Terminal-Bench 2.1; 28.3 on Terminal-Bench 3.0; versions differ from 4.0
Open-weight advantage
61.7 SWE-bench Pro; 73.0 Terminal-Bench 2.1; 84.3 OSWorld-Verified
Cost per Intelligence Index task
Artificial Analysis suite cost, not session cost
Higher cost
$5.98 per task
Higher cost
$3.26 per task
Open-weight advantage
$2.01 per task; score 45
Open-weight advantage
$1.01 per task; score 34
EU data path
Residency and jurisdiction
Contract review required
US provider; EU processing depends on eligible project and configuration
Contract review required
US provider; traffic routing and storage depend on product
EU/Regolo advantage
EU-hosted deployment can keep inference in Europe; confirm model availability
EU/Regolo advantage
Qwen3.8-27B is listed in the EU-hosted Regolo catalog
Retention and reuse
Payload controls
Eligibility dependent
Default abuse logs may retain content for up to 30 days; ZDR requires approval
Eligibility dependent
ZDR depends on feature and model; privacy documentation reports US storage by default
EU/Regolo advantage
Regolo states no payload retention and no training reuse for supported inference
EU/Regolo advantage
Qwen3.8-27B is available through the EU-hosted catalog with the same Regolo policy
License and control
Deployment freedom
Proprietary
Managed frontier model
Proprietary
Managed frontier model
Open-weight option
Custom GLM-5.3 License; review MaaS condition
Open-weight advantage
Apache 2.0; self-hosting and redistribution are simpler
Latency and EU locality
Network distance
US advantage
Strong managed performance; region and queue still matter
US advantage
Strong managed performance; region and queue still matter
EU locality can help
EU deployment can avoid a transatlantic hop; measure p95 on the workload
EU locality can help
EU deployment can avoid a transatlantic hop; measure p95 on the workload
Best fit
Decision guidance
Frontier quality
Hard reasoning and high-stakes tasks where budget is not the constraint
Frontier quality
Hard reasoning and high-stakes tasks where budget is not the constraint
Open-weight balance
EU-hosted coding, source code, and long-running agent workloads
Open-weight balance
EU-hosted coding, computer use, self-hosting, and cost-sensitive workloads

Red does not mean that an American provider is unlawful or automatically non-compliant, It means that an EU-only policy needs an explicit regional, contractual, retention, and subprocessor review.

Green means the EU path is easier to evidence under the stated architecture; it does not replace the customer’s legal assessment.


SOVEREIGN EUROPEAN INFERENCE

Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention

Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

Start 30-day free trial (no card required) → Free credits included · Live in 60s
✓ 100% EU Green Datacenters ✓ Certified ZDR

FAQ

Which model is best for coding agents in September 2026?

Claude Opus 5.5 has the highest score on the Coding Agent Index v1.5. The best operating choice still depends on task completion, p95 latency, data policy, and total cost. Qwen3.8-27B and GLM-5.3 can be stronger for EU-hosted or self-controlled workloads.

Are US AI models automatically non-compliant with GDPR?

No. OpenAI and Anthropic provide DPAs, retention controls, and regional options for eligible products. Compliance depends on the selected endpoint, contract, subprocessors, data flows, and customer configuration. EU-hosted inference offers a simpler default path when the policy requires EU-only processing.

What compliance advantage does EU inference provide?

The main advantage is verifiability. EU-only infrastructure can reduce transfer analysis, keep customer content under one jurisdiction, and simplify the provider assessment. Regolo additionally claims certified Zero Data Retention and no training reuse. Buyers should still inspect the certificate, DPA, metadata, and application controls.

Does EU-hosted inference always have lower latency?

No. It can remove a transatlantic network hop and reduce routing variability for European callers. Model size, queueing, cache state, reasoning effort, and cold starts can still make an EU endpoint slower. Compare p50 and p95 latency on the same workload and from the same region.

How close are GLM-5.3 and Qwen3.8-27B to frontier models?

GLM-5.3 scores 45 on the Intelligence Index, compared with 58 for Opus 5.5 and 53 for GPT-6 Astra. Qwen3.8-27B scores 33.7. These scores cover a broader suite than coding. In vendor-published benchmarks, Qwen reaches 61.7 on SWE-bench Pro and 84.3 on OSWorld-Verified; GLM reaches 66.9 on DeepSWE v1.1 and 88.2 on Terminal-Bench 2.1.

Can DeepSWE, Terminal-Bench 2.1, and Terminal-Bench 4.0 be compared directly?

No. Benchmark versions use different tasks and settings. Compare results only when they use the same version, harness, and similar evaluation conditions.

How much can an open-weight model save?

It depends on tokens per task, retries, latency, GPUs, operations, and completion rate. Measure total cost per completed task with the same protocol for every candidate.

Can you migrate from OpenAI or Anthropic through a compatible API?

The migration depends on the harness and provider. Qwen supports use through the OpenAI Python SDK, and Regolo states that it offers an OpenAI-compatible API. Its OpenCode configuration guide shows the provider configuration. Model IDs, multimodal features, caching, reasoning, and telemetry may require changes.