Claude Code with Claude Opus 5.5 scores 66 points on the Artificial Analysis Coding Agent Index v1.5. Codex with GPT-6 Astra scores 62. During the same period, Qwen3.8-27B does not appear on the full coding-agent leaderboard, while its model card reports 61.7 on SWE-bench Pro and 84.3 on OSWorld-Verified. GLM-5.3 reports 66.9 on DeepSWE v1.1 and 88.2 on Terminal-Bench 2.1.
This guide compares the best AI coding models for agentic software work in 2026, including benchmark performance, cost per task, and European data controls. A leaderboard score combines the model, harness, reasoning settings, and provider. Start with the team’s real tasks, then measure reliability and total operating cost.
Key takeaways
- Claude Opus 5.5 leads the Coding Agent Index v1.5 with 66 points, ahead of GPT-6 Astra at 62.
- GLM-5.3 performs well on coding and long-running tasks. Its advantage still needs to be tested on the real workload and chosen infrastructure.
- Qwen3.8-27B combines 27 billion parameters, visual capabilities, and the Apache License 2.0. Its OSWorld-Verified result makes it interesting for computer use.
- API price is not total cost. Count retries, latency, supervision, infrastructure, and failures.
- Zero Data Retention and European residency reduce specific risks, but they do not make an application compliant by law.
Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention
Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.
Numbers at a glance
The values below were checked on September 24, 2026. Model versions and settings affect the results.
| Metric | Claude Opus 5.5 | GPT-6 Astra | GLM-5.3 | Qwen3.8-27B |
|---|---|---|---|---|
| Artificial Analysis Coding Agent Index v1.5 | 66 with Claude Code | 62 with Codex | About 54 with OpenCode, based on the public chart snapshot | Not listed on the full leaderboard |
| Artificial Analysis Intelligence Index v4.3.2 | 58 | 53 | 45 | 33.7, displayed as 34 |
| Cost per Intelligence Index task | $5.98 | $3.26 | $2.01 | $1.01 |
| License | Proprietary | Proprietary | GLM-5.3 License, with conditions | Apache 2.0 |
| Strongest cited setting | Max effort with fallback | Max effort | Max effort | Xhigh |
The Coding Agent Index compares complete systems. A score of 66 for Opus 5.5 is not a model-only score. It also includes Claude Code, the context given to the agent, the attempts, and provider routing. The OpenCode row with GLM-5.3 therefore does not automatically match the same model running in another harness.
The task cost in the table belongs to the Intelligence Index. It is not the average cost of a coding-agent session and does not include infrastructure, human intervention, or application retries.
How to read the AI coding agent benchmark
The current AI coding agent benchmark combines three tests with equal weight:
- DeepSWE v1.1: 113 software-engineering tasks.
- Terminal-Bench 4.0: 66 agentic terminal-use tasks.
- SWE-Atlas-QnA: 124 repository-understanding and technical Q&A tasks.
Artificial Analysis calculates average pass@1 across three attempts per task and gives the three components equal weight. The result provides a quick comparison between agents, but it does not replace an internal test.
The harness matters: the same model may receive different instructions, use different tools, and retry under different conditions
Where GLM-5.3 approaches the frontier
Z.ai’s GLM-5.3 report says the model uses the same base as GLM-5.2 and that the improvements come from post-training. The results published by the vendor show the jump:

IMAGE: source GLM report
KPI Highlights
| Benchmark | GLM-5.2 | GLM-5.3 | Change |
|---|---|---|---|
| DeepSWE v1.1 | 46.2 | 66.9 | +20.7 |
| Terminal-Bench 3.0 | 4.6 | 28.3 | +23.7 |
| Agents’ Last Exam | 23.8 | 28.5 | +4.7 |
| AutomationBench v1.0.6 | 26.2 | 48.2 | +22.0 |
The same chart gives GLM-5.3 an 88.2 score on Terminal-Bench 2.1, these benchmark versions must not be mixed with Terminal-Bench 4.0, which is part of the current Artificial Analysis Coding Agent Index. A score of 28.3 on Terminal-Bench 3.0 cannot be compared directly with a score on 4.0.
GLM-5.3 releases open weights under a custom license and It allows use, modification, distribution, and commercial use. However, the license requires a company that operates a Model as a Service and exceeds $10 billion in aggregate revenue over 12 months to undergo a Z.AI security review and legal teams should assess this condition before a Model as a Service deployment.
Qwen3.8-27B: a compact model for coding and computer use
Qwen’s model card presents Qwen3.8-27B as a dense 27-billion-parameter model with a visual encoder. It supports images and video, uses a native 262,144-token context window, and can be extended to one million tokens with RoPE scaling. The weights are released under the Apache License 2.0.

IMAGE: source Qwen model card
KPI Highlights
| Benchmark | Qwen3.8-27B | What the result measures |
|---|---|---|
| Terminal-Bench 2.1 | 73.0 | Agentic terminal coding with the Terminus harness |
| SWE-bench Pro | 61.7 | Repository changes and software issue resolution |
| DeepSWE v1.1 | 42.2 | Software-engineering tasks on repositories |
| OSWorld-Verified | 84.3 | Computer use on desktop interfaces |
| IFBench | 79.5 | Instruction following |
| GPQA Diamond | 89.2 | Scientific reasoning |
| LiveCodeBench v6 | 90.3 | Competitive programming |
Many tests used Claude Code and the settings listed in the model card, they help explain the stated capabilities but they are not an independent leaderboard identical to the Artificial Analysis Coding Agent Index.
The OSWorld-Verified result makes Qwen3.8-27B especially interesting for agents that need to use browsers, desktop applications, and visual tools. The Apache License 2.0 provides a simpler route to self-hosting and redistribution than a custom license.
Regolo’s Qwen3.8-27B benchmark analysis reviews the vendor table claim by claim, Its harness engineering guide shows how the same checkpoint can produce different results across coding agents.
Costs: from token price to task cost
Artificial Analysis measures cost with a consistent method across its own suite. On September 24, 2026, the Intelligence Index values were:
| Model | Intelligence Index | Cost per task | Output tokens per task, when reported |
|---|---|---|---|
| Claude Opus 5.5 max | 58 | $5.98 | About 119,000 |
| GPT-6 Astra max | 53 | $3.26 | About 27,000 |
| GLM-5.3 max | 45 | $2.01 | About 71,000 |
| Qwen3.8-27B xhigh | 34 | $1.01 | About 67,000 |
Regolo’s AI model comparison and pricing hub tracks provider prices, benchmark results, context limits, and serving assumptions.
In this suite, GLM-5.3 and Qwen3.8-27B have a lower Intelligence Index cost per task than Opus 5.5 and Astra. That does not prove that an open-weight model will always cost less.
US frontier APIs and EU inference: the compliance difference
US frontier APIs are not automatically outside GDPR. OpenAI says API data is not used for training by default, offers a DPA, and supports European processing plus Zero Data Retention for eligible projects and endpoints. Anthropic also provides contractual safeguards and ZDR for eligible features. Its privacy documentation says customer data is stored in the US by default, even when traffic can be routed to selected regions.
Regolo’s current model catalog lists Qwen3.8-27B and GLM-5.2, GLM-5.3 availability depends on a custom deployment review for now.
The comparison below describes the EU-hosted open-weight architecture
The architectural difference is the default path a European buyer can prove, a US provider can offer compliant processing, but the review may cover a US legal entity, subprocessors, system data, endpoint-specific retention, and transfer safeguards. EU-hosted inference can keep customer content, inference, and operational control in one jurisdiction.
That can simplify an EU-only policy and the evidence required under Articles 28 and 44 of the GDPR.
| Control | US frontier APIs | EU-hosted open-weight inference on Regolo |
|---|---|---|
| Model quality | Often the strongest closed models, available through managed APIs | Depends on the selected open model, quantization, and serving stack |
| EU processing | OpenAI offers EU regional processing for eligible projects. Anthropic supports traffic routing, while documenting US storage by default. | Regolo states that inference runs in European data centers and does not leave the EU region |
| Payload retention | OpenAI logs may be retained for up to 30 days by default; ZDR requires approval. Anthropic ZDR depends on the feature and model. | Regolo states that prompts and responses are processed in volatile memory and are not retained after the response |
| Training and reuse | OpenAI API data is not used for training by default. Anthropic retained data is not used without express permission. | Regolo states that customer data is not reused for training or fine-tuning |
| Governance evidence | DPA, subprocessors, transfer mechanism, endpoint limits, and US jurisdiction must be reviewed | EU DPA, certificate scope, infrastructure location, and operational metadata must still be reviewed |
Regolo’s Zero Data Retention documentation and its independent certificate provide stronger evidence for an EU-only architecture than a generic “GDPR compliant” badge. They do not make the customer application compliant by themselves. The GDPR still requires a lawful basis, minimization, security, and appropriate contracts and the EU AI Act still depends on the system’s role and risk classification.
Latency: EU locality helps, measurement decides
Geography is one part of latency, a longer network path adds propagation delay and more routing hops – for a European application, an EU inference endpoint can avoid a transatlantic round trip and keep the request path inside one jurisdiction.
Model size, prompt length, reasoning effort, queueing, cache state, provider routing, and cold starts can dominate, the live Regolo OBS Qwen3.8-27B health page showed a 2.074-second average across its last 500 health pings, 95.8% success, and one recent slow ping of 12.432 seconds at 08:37 UTC on September 24, 2026.
When to choose each model family
This table separates measured model performance from provider posture. It is a decision aid, not a legal conclusion.
Legend: green marks an EU or Regolo advantage; orange marks an open-weight cost or capability advantage, or a close comparison; white marks an American-provider advantage; red marks a non-EU-only default path or a contractual transfer review requirement.
| Decision criterion | Claude Opus 5.5 | GPT-6 Astra | GLM-5.3 | Qwen3.8-27B |
|---|---|---|---|---|
| Reasoning composite Intelligence Index v4.3.2, same suite | US advantage 58 points | US advantage 53 points | Open-weight comparable 45 points | Open-weight comparable 33.7 points, displayed as 34 |
| Agent composite Coding Agent Index v1.5, same suite | US advantage 66 with Claude Code | US advantage 62 with Codex | Open-weight comparable About 54 with OpenCode, public chart snapshot | Not listed No full coding-agent leaderboard entry in checked sources |
| Long-horizon repo coding DeepSWE v1.1, same version; harness differs | Not disclosed No DeepSWE v1.1 figure in checked sources | US advantage 68 in Codex harness, per Artificial Analysis | Closest open signal 66.9, vendor-reported by Z.ai | Open-weight signal 42.2, vendor-reported by Qwen |
| Closest open coding signals Vendor-reported; no frontier match in checked sources | Not disclosed No matching SWE-bench Pro or OSWorld figure in checked sources | Not disclosed No matching SWE-bench Pro or OSWorld figure in checked sources | Terminal signal 88.2 on Terminal-Bench 2.1; 28.3 on Terminal-Bench 3.0; versions differ from 4.0 | Open-weight advantage 61.7 SWE-bench Pro; 73.0 Terminal-Bench 2.1; 84.3 OSWorld-Verified |
| Cost per Intelligence Index task Artificial Analysis suite cost, not session cost | Higher cost $5.98 per task | Higher cost $3.26 per task | Open-weight advantage $2.01 per task; score 45 | Open-weight advantage $1.01 per task; score 34 |
| EU data path Residency and jurisdiction | Contract review required US provider; EU processing depends on eligible project and configuration | Contract review required US provider; traffic routing and storage depend on product | EU/Regolo advantage EU-hosted deployment can keep inference in Europe; confirm model availability | EU/Regolo advantage Qwen3.8-27B is listed in the EU-hosted Regolo catalog |
| Retention and reuse Payload controls | Eligibility dependent Default abuse logs may retain content for up to 30 days; ZDR requires approval | Eligibility dependent ZDR depends on feature and model; privacy documentation reports US storage by default | EU/Regolo advantage Regolo states no payload retention and no training reuse for supported inference | EU/Regolo advantage Qwen3.8-27B is available through the EU-hosted catalog with the same Regolo policy |
| License and control Deployment freedom | Proprietary Managed frontier model | Proprietary Managed frontier model | Open-weight option Custom GLM-5.3 License; review MaaS condition | Open-weight advantage Apache 2.0; self-hosting and redistribution are simpler |
| Latency and EU locality Network distance | US advantage Strong managed performance; region and queue still matter | US advantage Strong managed performance; region and queue still matter | EU locality can help EU deployment can avoid a transatlantic hop; measure p95 on the workload | EU locality can help EU deployment can avoid a transatlantic hop; measure p95 on the workload |
| Best fit Decision guidance | Frontier quality Hard reasoning and high-stakes tasks where budget is not the constraint | Frontier quality Hard reasoning and high-stakes tasks where budget is not the constraint | Open-weight balance EU-hosted coding, source code, and long-running agent workloads | Open-weight balance EU-hosted coding, computer use, self-hosting, and cost-sensitive workloads |
Red does not mean that an American provider is unlawful or automatically non-compliant, It means that an EU-only policy needs an explicit regional, contractual, retention, and subprocessor review.
Green means the EU path is easier to evidence under the stated architecture; it does not replace the customer’s legal assessment.
Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention
Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.
FAQ
Which model is best for coding agents in September 2026?
Claude Opus 5.5 has the highest score on the Coding Agent Index v1.5. The best operating choice still depends on task completion, p95 latency, data policy, and total cost. Qwen3.8-27B and GLM-5.3 can be stronger for EU-hosted or self-controlled workloads.
Are US AI models automatically non-compliant with GDPR?
No. OpenAI and Anthropic provide DPAs, retention controls, and regional options for eligible products. Compliance depends on the selected endpoint, contract, subprocessors, data flows, and customer configuration. EU-hosted inference offers a simpler default path when the policy requires EU-only processing.
What compliance advantage does EU inference provide?
The main advantage is verifiability. EU-only infrastructure can reduce transfer analysis, keep customer content under one jurisdiction, and simplify the provider assessment. Regolo additionally claims certified Zero Data Retention and no training reuse. Buyers should still inspect the certificate, DPA, metadata, and application controls.
Does EU-hosted inference always have lower latency?
No. It can remove a transatlantic network hop and reduce routing variability for European callers. Model size, queueing, cache state, reasoning effort, and cold starts can still make an EU endpoint slower. Compare p50 and p95 latency on the same workload and from the same region.
How close are GLM-5.3 and Qwen3.8-27B to frontier models?
GLM-5.3 scores 45 on the Intelligence Index, compared with 58 for Opus 5.5 and 53 for GPT-6 Astra. Qwen3.8-27B scores 33.7. These scores cover a broader suite than coding. In vendor-published benchmarks, Qwen reaches 61.7 on SWE-bench Pro and 84.3 on OSWorld-Verified; GLM reaches 66.9 on DeepSWE v1.1 and 88.2 on Terminal-Bench 2.1.
Can DeepSWE, Terminal-Bench 2.1, and Terminal-Bench 4.0 be compared directly?
No. Benchmark versions use different tasks and settings. Compare results only when they use the same version, harness, and similar evaluation conditions.
How much can an open-weight model save?
It depends on tokens per task, retries, latency, GPUs, operations, and completion rate. Measure total cost per completed task with the same protocol for every candidate.
Can you migrate from OpenAI or Anthropic through a compatible API?
The migration depends on the harness and provider. Qwen supports use through the OpenAI Python SDK, and Regolo states that it offers an OpenAI-compatible API. Its OpenCode configuration guide shows the provider configuration. Model IDs, multimodal features, caching, reasoning, and telemetry may require changes.