Skip to content
Regolo Logo
Benchmarks & Cost Optimization

Bonsai 2 vs Qwen3.8-27B: enterprise benchmarks and deployment choices

Alex Genovese
8 min read
Share

Bonsai 2 and Qwen3.8-27B belong to the same approximately 27-billion-parameter family. Bonsai compresses the Qwen parent into ternary weights: its developer reports a capabilities average of 84.78/100, versus 86.32/100 for the original Qwen3.8-27B FP16 reference, across the same 14 thinking-mode benchmarks. On Regolo, we offer Qwen3.8-27B through our managed API at €0.50 per million input tokens and €2.10 per million output tokens, excluding VAT.

The practical decision is whether your company benefits more from Bonsai’s smaller deployable artifact or from using our managed Qwen inference. The published capability gap is modest in aggregate but varies by task. It does not establish equivalent code-review accuracy, batch-extraction quality, latency or cost per successful task.

SOVEREIGN EUROPEAN INFERENCE

Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention

Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

Start 30-day free trial (no card required) → Free credits included · Live in 60s
✓ 100% EU Green Datacenters ✓ Certified ZDR

Important terminology: 27B means billions of parameters, not billions of tokens. Parameters are the model’s learned weights; tokens are units of input, context, reasoning and output. Bonsai remains a 27B-class model after compression—it is not a 5.95-billion-parameter model because its file occupies 5.95 GB.

What this comparison covers

This article separates three layers:

  • Model capability: a controlled publisher-run comparison of Bonsai against the original full-precision Qwen parent.
  • Deployment specification: context, modalities, file size and runtime requirements.
  • Service economics: our actual Qwen list price on Regolo, versus costs that must be measured for a self-hosted Bonsai deployment.

No alternative Qwen quantization is used in the comparisons below; the FP16 reference is not asserted to be our exact serving precision – our public platform documentation identifies the available model, but does not establish its serving precision, effective context limit or maximum completion limit.

Requested metricBonsai 2Qwen3.8-27BInterpretation
Parameter classApproximately 27B; derived from Qwen3.8-27B27B language modelSame underlying model family, different weight representation.
Capabilities score84.78/10086.32/100, original FP16 referenceprismml’s 14-test average, not an independent index or our platform benchmark.
Pricing per 1M tokensNo verified standardized Bonsai API rate in the acquired sourcesRegolo: €0.50 input / €2.10 outputLocal or dedicated compute costs must not be represented as zero token cost.
Native context window262,144 tokens, documented in runtime setup262,144 tokensCombined working context, not a maximum final-answer length.
Extended contextNo independently verified 1M deployment in this articleUp to 1M with appropriate scaling and serving configurationA Qwen model capability, not our default production endpoint setting.
Maximum output capacityNo separately verified hard model/API output maximumNo verified hard maximum on our endpoint; Qwen recommends up to 262,144 reasoning and 131,072 final-answer tokens in supported 1M-context configurationsRecommendations are not universal hard limits; do not copy them into an API comparison as guarantees.
Release dateSeptember 17, 2026August 14, 202615 versus 49 days old on October 2; newer compression does not mean fresher knowledge.
Documented versatilityText and image input; optional vision packText, image and video inputConfirm actual endpoint support; model modality support does not prove every provider exposes it.
Commercial licenseApache 2.0Apache 2.0Check the license and dependencies for your intended deployment.
Production composite rankingNot determinable from current evidenceNot determinable from current evidenceMatched costs, output limits and operational evaluations are missing.

Evidence and independence

prismml reports these capability results using evalscope and vllm on NVIDIA H100, with the same decoding and scoring in thinking mode. That is a valuable within-suite comparison, but the evaluator also develops Bonsai.

Artificial Analysis independently reports 34 on its Intelligence Index for Qwen3.8-27B at xhigh effort. We did not acquire an equivalent Bonsai score on that index. 34 and 84.78 are different scales and must not be placed on the same ranking. AA’s reference API costs and speed also must not be attributed to our infrastructure.

The prismml launch post uses a different aggregate, 83.9 versus 85.4, while the current model card uses the 14-test aggregate above. We use only the current model-card suite throughout the numerical comparison instead of mixing evaluation sets.

Capabilities and business workloads

Capability categories

Capability categoryBonsai 2Original Qwen FP16How a business should read it
Knowledge and reasoning79.8685.55Qwen leads this category; validate analytical answers against evidence.
Mathematics96.5797.06Very close on the selected math tests.
Coding89.4289.07Near parity in code-generation tests; not proof of superior autonomous coding.
Instruction following82.6681.25Useful signal for following constraints; not an extraction accuracy guarantee.
Tool calling74.9276.74Function-call correctness, not complete workflow success.
Vision66.1971.36Qwen leads the selected vision tests; use real scans and images for validation.
Overall, 14 tests84.7886.32Publisher-reported macro-average, not a business success rate.

All scores come from the same prismml model-card evaluation. The published overall is the average of 14 tests, not the unweighted average of the six category rows, which contain different numbers of tests.

Coding and code review

TestBonsai 2Original Qwen FP16
humaneval+95.1293.29
MBPP+83.0783.86
livecodebench90.0790.05
BFCL v3, function calling74.9276.74

These publisher results suggest that Bonsai preserves much of the parent’s code-generation performance. They do not evaluate bug detection, security-review recall, false-positive review comments or an agent resolving a real issue.

For your pilot, measure tests passed, seeded bugs found, false positives, regressions introduced, human minutes spent reviewing, wall-clock time per verified fix and cost per successful fix. Use the same repository snapshot, tool permissions and acceptance criteria for both deployments. Do not substitute token throughput for code-review quality.

Long document analysis

Both models document a 262K-class native context, but the model limit and the serving configuration are separate. A server must still allocate working memory and reserve room for reasoning and output.

prismml’s musr scores—70.63 for Bonsai and 79.63 for Qwen FP16—measure multi-step reasoning, not retrieval accuracy at 262K context. A larger context window alone does not prove that a model can locate a clause or preserve its meaning across an entire contract collection.

Measure grounded-answer accuracy, citation correctness, unsupported claims and fact recall at increasing input lengths, such as 8K, 32K, 64K and the largest limit supported by your target endpoint. A representative question is: “which clauses allow termination, and what notice periods apply?” Require exact evidence spans in the answer.

Creative writing

The acquired evidence does not contain a comparable creative-writing benchmark for both models. Neither the math average nor instruction-following scores justify declaring a writing winner.

Run a blinded evaluation of the same briefs: product pages, emails or campaign copy. Reviewers should score brand voice, clarity, originality, factual consistency and constraint adherence. Track first-draft acceptance, editing time and cost per accepted draft. Keep reasoning settings and length requirements visible, since these affect cost and responsiveness.

Image understanding and OCR

Visual testBonsai 2Original Qwen FP16
MMMU-Pro75.4981.73
OCR Bench v256.8860.99

Qwen leads both published tests. Bonsai requires its optional vision component for image input; the compact language-model file alone is not an OCR deployment.

For document operations, benchmark character or word error rate, exact matching of critical fields, table reconstruction and end-to-end extraction correctness. Include skew, small print, handwriting where relevant and the languages your business uses. Confirm image inputs are supported on our target endpoint before designing your pipeline around them.

Pricing efficiency and scoring

When workloads become steady, our subscription plans reduce effective token costs significantly. For example, our Core Plan costs €39.00 per month (with an introductory 70% discount for the first 3 months at €11.70 per month) and includes up to 20 million tokens per day, representing a monthly capacity of up to 600 million tokens.

If a team utilizes that full 600M token allocation under the Core Plan:

€39.00600 M tokens=€0.065 per million tokens. \frac{€39.00}{600\text{ M tokens}} = €0.065\text{ per million tokens}.

Even with our standard list price of €39.00/month, the effective rate drops from €0.82 blended pay-as-you-go down to €0.065 per million tokens (over a 92% reduction). During the introductory 70% discounted period (€11.70/month), the effective rate falls further to €0.0195 per million tokens. Even at a conservative 50% utilization (300M tokens consumed in a month), the cost is €0.13 per million tokens, still more than 6× cheaper than on-demand pay-as-you-go rates.

SOVEREIGN EUROPEAN INFERENCE

Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention

Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

Start 30-day free trial (no card required) → Free credits included · Live in 60s
✓ 100% EU Green Datacenters ✓ Certified ZDR

An auditable composite score

A global composite is a decision model, not an established benchmark. We propose the following weights for a general enterprise shortlist; change them before seeing the results when your use case requires different priorities.

DimensionWeightProposed scoring rule
Benchmark performance35%Current prismml 14-test score, 0–100; replace with matched business-task results for production.
Pricing efficiency25%100 × the lowest measured cost per success / the candidate’s measured cost per success.
Native context15%100 × min(verified native context / 262,144, 1). Confirm endpoint limits before production scoring.
Maximum output10%100 × min(verified completion cap / required completion cap, 1). Unknown caps remain unknown.
Versatility10%Documented coverage of text, image and video input, one-third each; only a modality-coverage proxy.
Recency5%max(0, 100 × (1 − days since release / 180)). Release age is not a quality or knowledge-cutoff score.

Here the price-efficiency and maximum-output dimensions, together worth 35%, are missing for a matched comparison, the other dimensions are model-level descriptors—not validation of our production endpoint.

Using the explicit proxy rules above, the possible reference-composite ranges are:

CandidateContributions from documented dimensionsUnknown contributionPossible total / 100
Bonsai 255.920–3555.92–90.92
Qwen3.8-27B58.850–3558.85–93.85

These are calculated uncertainty ranges, not confidence intervals or achieved production scores. They use 15 versus 49 days of age, identical native context and a two-versus-three documented-modality proxy. The ranges overlap substantially, so a defensible global ranking cannot yet be assigned. The known contributions must not be used as a winner’s podium.

Freshness also needs interpretation: bonsai was released later but derives from Qwen. Its release date does not prove a later training-data cutoff or better knowledge of recent events.

Use open-weight models with Regolo

Start with our ready-to-use Qwen API

We list qwen3.8-27b as a core model on Regolo. Our public integration example uses an openai-compatible endpoint, so you can keep your application logic while replacing the API base URL and key.

import os
from openai import openai

client = openai(
    api_key=os.environ["REGOLO_API_KEY"],
    base_url="https://api.regolo.ai/v1",
)

response = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[
        {"role": "system", "content": "extract only facts supported by the source. Mark missing fields as null."},
        {"role": "user", "content": "extract supplier, invoice number and total from this text:\n..."},
],
)

print(response.choices[0].message.content)
print(response.usage)Code language: Python (python)

Choose by workload, not ranking

Business requirementSensible starting pointAcceptance evidence
Local/offline inference or tight memoryBonsai 2Actual memory at target context, supported runtime and task success.
Managed inference without owning servingQwen3.8-27B on RegoloEndpoint limits, measured latency, error rate and real cost per accepted task.
Coding and code reviewPilot both on the same repositoryTests passed, seeded defects found, false positives and regressions.
Long document analysisVerify endpoint capacity, then test bothGrounded answers and evidence spans at increasing input lengths.
Batch extractionCompare accepted records, not JSON validity aloneField-level F1, rejection/retry rates, throughput and cost per 1,000 accepted records.
Creative writingBlinded human evaluationBrand fit, factuality, editing minutes and first-draft acceptance.
Image understanding / OCRQwen as the first quality baselineReal scans, critical-field exact match and confirmed multimodal endpoint support.

The architecture and hosting options above follow the documented model/runtime specifications and our platform integration paths; the workload tests are our proposed evaluation protocol, not results already achieved.

Run the same labeled task set on both candidates. Record the model and runtime revisions, input/output budgets, prompts, tool permissions, concurrency and evaluation rules. Track end-to-end P50/P95 latency and verified successes, not only generation speed. Promote a deployment when it meets explicit business thresholds, then calculate the composite with the missing dimensions measured.

SOVEREIGN EUROPEAN INFERENCE

Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention

Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

Start 30-day free trial (no card required) → Free credits included · Live in 60s
✓ 100% EU Green Datacenters ✓ Certified ZDR

FAQ

Are these genuinely comparable model sizes?

Yes: they are in the same approximately 27B parameter family; Bonsai is derived from Qwen3.8-27B, compression changes the weight representation and deployment footprint, not the comparison into a different trillion-parameter class.

Is the reported Qwen capability score measured on Regolo?

No, 86.32 is prismml’s FP16 reference score, a matched pilot on our infrastructure is needed before claiming the same score or production performance on our endpoint.

Is Bonsai automatically supported by Regolo Custom Models?

Not established. Our documented custom route requires serving compatibility, while Bonsai’s ternary packs require specialized runtime kernels. Talk to our team first to confirm support.

Which has the best global score?

No verified winner is available. The evidence supports a task-dependent capability comparison, but not a complete matched ranking on cost, output limits and production suitability.


Ship Private AI. Not Infrastructure.

You have the private AI App architecture, bow give it an inference layer built for production.

Regolo gives European teams fast, OpenAI-compatible access to Mistral, Llama, Qwen, DeepSeek, GLM, and more — with zero data retention, EU data residency, and no new SDK to learn.

Change your base_url. Keep your LangChain code. Start shipping.

🚀 Start your 30-day free trial →

Build, test, and deploy with no infrastructure to maintain.
No credit card. No migration project. No compromise on data control.

💬 Join the Regolo Discord →

Meet builders working on private RAG, local LLMs, LangChain, Ollama, and production AI systems. Share your setup, get feedback from the community, and speak directly with the Regolo team.

🤝 Talk to an AI Infrastructure Engineer →

Running a sensitive workload, scaling beyond a proof of concept, or assessing a managed EU inference provider? Get a tailored architecture and commercial proposal for your team.

📂 Clone the GitHub repository →

Get the full implementation from this guide: ingestion scripts, ChromaDB setup, hybrid retrieval, the 30-Question RAG Floor, evaluation examples, and deployment configuration.

Private AI should not require a private data center.
Regolo gives your team an EU-native path from local experimentation to production-grade inference.


Build with Regolo


Built with ❤️ by the Regolo team. Questions? regolo.ai/contact or chat with us on Discord