Bonsai 2 and Qwen3.8-27B belong to the same approximately 27-billion-parameter family. Bonsai compresses the Qwen parent into ternary weights: its developer reports a capabilities average of 84.78/100, versus 86.32/100 for the original Qwen3.8-27B FP16 reference, across the same 14 thinking-mode benchmarks. On Regolo, we offer Qwen3.8-27B through our managed API at €0.50 per million input tokens and €2.10 per million output tokens, excluding VAT.
The practical decision is whether your company benefits more from Bonsai’s smaller deployable artifact or from using our managed Qwen inference. The published capability gap is modest in aggregate but varies by task. It does not establish equivalent code-review accuracy, batch-extraction quality, latency or cost per successful task.
Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention
Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.
Important terminology: 27B means billions of parameters, not billions of tokens. Parameters are the model’s learned weights; tokens are units of input, context, reasoning and output. Bonsai remains a 27B-class model after compression—it is not a 5.95-billion-parameter model because its file occupies 5.95 GB.
What this comparison covers
This article separates three layers:
- Model capability: a controlled publisher-run comparison of Bonsai against the original full-precision Qwen parent.
- Deployment specification: context, modalities, file size and runtime requirements.
- Service economics: our actual Qwen list price on Regolo, versus costs that must be measured for a self-hosted Bonsai deployment.
No alternative Qwen quantization is used in the comparisons below; the FP16 reference is not asserted to be our exact serving precision – our public platform documentation identifies the available model, but does not establish its serving precision, effective context limit or maximum completion limit.
| Requested metric | Bonsai 2 | Qwen3.8-27B | Interpretation |
|---|---|---|---|
| Parameter class | Approximately 27B; derived from Qwen3.8-27B | 27B language model | Same underlying model family, different weight representation. |
| Capabilities score | 84.78/100 | 86.32/100, original FP16 reference | prismml’s 14-test average, not an independent index or our platform benchmark. |
| Pricing per 1M tokens | No verified standardized Bonsai API rate in the acquired sources | Regolo: €0.50 input / €2.10 output | Local or dedicated compute costs must not be represented as zero token cost. |
| Native context window | 262,144 tokens, documented in runtime setup | 262,144 tokens | Combined working context, not a maximum final-answer length. |
| Extended context | No independently verified 1M deployment in this article | Up to 1M with appropriate scaling and serving configuration | A Qwen model capability, not our default production endpoint setting. |
| Maximum output capacity | No separately verified hard model/API output maximum | No verified hard maximum on our endpoint; Qwen recommends up to 262,144 reasoning and 131,072 final-answer tokens in supported 1M-context configurations | Recommendations are not universal hard limits; do not copy them into an API comparison as guarantees. |
| Release date | September 17, 2026 | August 14, 2026 | 15 versus 49 days old on October 2; newer compression does not mean fresher knowledge. |
| Documented versatility | Text and image input; optional vision pack | Text, image and video input | Confirm actual endpoint support; model modality support does not prove every provider exposes it. |
| Commercial license | Apache 2.0 | Apache 2.0 | Check the license and dependencies for your intended deployment. |
| Production composite ranking | Not determinable from current evidence | Not determinable from current evidence | Matched costs, output limits and operational evaluations are missing. |
Evidence and independence
prismml reports these capability results using evalscope and vllm on NVIDIA H100, with the same decoding and scoring in thinking mode. That is a valuable within-suite comparison, but the evaluator also develops Bonsai.
Artificial Analysis independently reports 34 on its Intelligence Index for Qwen3.8-27B at xhigh effort. We did not acquire an equivalent Bonsai score on that index. 34 and 84.78 are different scales and must not be placed on the same ranking. AA’s reference API costs and speed also must not be attributed to our infrastructure.
The prismml launch post uses a different aggregate, 83.9 versus 85.4, while the current model card uses the 14-test aggregate above. We use only the current model-card suite throughout the numerical comparison instead of mixing evaluation sets.
Capabilities and business workloads
Capability categories
| Capability category | Bonsai 2 | Original Qwen FP16 | How a business should read it |
|---|---|---|---|
| Knowledge and reasoning | 79.86 | 85.55 | Qwen leads this category; validate analytical answers against evidence. |
| Mathematics | 96.57 | 97.06 | Very close on the selected math tests. |
| Coding | 89.42 | 89.07 | Near parity in code-generation tests; not proof of superior autonomous coding. |
| Instruction following | 82.66 | 81.25 | Useful signal for following constraints; not an extraction accuracy guarantee. |
| Tool calling | 74.92 | 76.74 | Function-call correctness, not complete workflow success. |
| Vision | 66.19 | 71.36 | Qwen leads the selected vision tests; use real scans and images for validation. |
| Overall, 14 tests | 84.78 | 86.32 | Publisher-reported macro-average, not a business success rate. |
All scores come from the same prismml model-card evaluation. The published overall is the average of 14 tests, not the unweighted average of the six category rows, which contain different numbers of tests.

Coding and code review
| Test | Bonsai 2 | Original Qwen FP16 |
|---|---|---|
| humaneval+ | 95.12 | 93.29 |
| MBPP+ | 83.07 | 83.86 |
| livecodebench | 90.07 | 90.05 |
| BFCL v3, function calling | 74.92 | 76.74 |
These publisher results suggest that Bonsai preserves much of the parent’s code-generation performance. They do not evaluate bug detection, security-review recall, false-positive review comments or an agent resolving a real issue.
For your pilot, measure tests passed, seeded bugs found, false positives, regressions introduced, human minutes spent reviewing, wall-clock time per verified fix and cost per successful fix. Use the same repository snapshot, tool permissions and acceptance criteria for both deployments. Do not substitute token throughput for code-review quality.

Long document analysis
Both models document a 262K-class native context, but the model limit and the serving configuration are separate. A server must still allocate working memory and reserve room for reasoning and output.
prismml’s musr scores—70.63 for Bonsai and 79.63 for Qwen FP16—measure multi-step reasoning, not retrieval accuracy at 262K context. A larger context window alone does not prove that a model can locate a clause or preserve its meaning across an entire contract collection.
Measure grounded-answer accuracy, citation correctness, unsupported claims and fact recall at increasing input lengths, such as 8K, 32K, 64K and the largest limit supported by your target endpoint. A representative question is: “which clauses allow termination, and what notice periods apply?” Require exact evidence spans in the answer.
Creative writing
The acquired evidence does not contain a comparable creative-writing benchmark for both models. Neither the math average nor instruction-following scores justify declaring a writing winner.
Run a blinded evaluation of the same briefs: product pages, emails or campaign copy. Reviewers should score brand voice, clarity, originality, factual consistency and constraint adherence. Track first-draft acceptance, editing time and cost per accepted draft. Keep reasoning settings and length requirements visible, since these affect cost and responsiveness.
Image understanding and OCR
| Visual test | Bonsai 2 | Original Qwen FP16 |
|---|---|---|
| MMMU-Pro | 75.49 | 81.73 |
| OCR Bench v2 | 56.88 | 60.99 |
Qwen leads both published tests. Bonsai requires its optional vision component for image input; the compact language-model file alone is not an OCR deployment.

For document operations, benchmark character or word error rate, exact matching of critical fields, table reconstruction and end-to-end extraction correctness. Include skew, small print, handwriting where relevant and the languages your business uses. Confirm image inputs are supported on our target endpoint before designing your pipeline around them.
Pricing efficiency and scoring
When workloads become steady, our subscription plans reduce effective token costs significantly. For example, our Core Plan costs €39.00 per month (with an introductory 70% discount for the first 3 months at €11.70 per month) and includes up to 20 million tokens per day, representing a monthly capacity of up to 600 million tokens.
If a team utilizes that full 600M token allocation under the Core Plan:
Even with our standard list price of €39.00/month, the effective rate drops from €0.82 blended pay-as-you-go down to €0.065 per million tokens (over a 92% reduction). During the introductory 70% discounted period (€11.70/month), the effective rate falls further to €0.0195 per million tokens. Even at a conservative 50% utilization (300M tokens consumed in a month), the cost is €0.13 per million tokens, still more than 6× cheaper than on-demand pay-as-you-go rates.

Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention
Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

An auditable composite score
A global composite is a decision model, not an established benchmark. We propose the following weights for a general enterprise shortlist; change them before seeing the results when your use case requires different priorities.
| Dimension | Weight | Proposed scoring rule |
|---|---|---|
| Benchmark performance | 35% | Current prismml 14-test score, 0–100; replace with matched business-task results for production. |
| Pricing efficiency | 25% | 100 × the lowest measured cost per success / the candidate’s measured cost per success. |
| Native context | 15% | 100 × min(verified native context / 262,144, 1). Confirm endpoint limits before production scoring. |
| Maximum output | 10% | 100 × min(verified completion cap / required completion cap, 1). Unknown caps remain unknown. |
| Versatility | 10% | Documented coverage of text, image and video input, one-third each; only a modality-coverage proxy. |
| Recency | 5% | max(0, 100 × (1 − days since release / 180)). Release age is not a quality or knowledge-cutoff score. |
Here the price-efficiency and maximum-output dimensions, together worth 35%, are missing for a matched comparison, the other dimensions are model-level descriptors—not validation of our production endpoint.
Using the explicit proxy rules above, the possible reference-composite ranges are:
| Candidate | Contributions from documented dimensions | Unknown contribution | Possible total / 100 |
|---|---|---|---|
| Bonsai 2 | 55.92 | 0–35 | 55.92–90.92 |
| Qwen3.8-27B | 58.85 | 0–35 | 58.85–93.85 |
These are calculated uncertainty ranges, not confidence intervals or achieved production scores. They use 15 versus 49 days of age, identical native context and a two-versus-three documented-modality proxy. The ranges overlap substantially, so a defensible global ranking cannot yet be assigned. The known contributions must not be used as a winner’s podium.
Freshness also needs interpretation: bonsai was released later but derives from Qwen. Its release date does not prove a later training-data cutoff or better knowledge of recent events.
Use open-weight models with Regolo
Start with our ready-to-use Qwen API
We list qwen3.8-27b as a core model on Regolo. Our public integration example uses an openai-compatible endpoint, so you can keep your application logic while replacing the API base URL and key.
import os
from openai import openai
client = openai(
api_key=os.environ["REGOLO_API_KEY"],
base_url="https://api.regolo.ai/v1",
)
response = client.chat.completions.create(
model="qwen3.8-27b",
messages=[
{"role": "system", "content": "extract only facts supported by the source. Mark missing fields as null."},
{"role": "user", "content": "extract supplier, invoice number and total from this text:\n..."},
],
)
print(response.choices[0].message.content)
print(response.usage)Code language: Python (python)
Choose by workload, not ranking
| Business requirement | Sensible starting point | Acceptance evidence |
|---|---|---|
| Local/offline inference or tight memory | Bonsai 2 | Actual memory at target context, supported runtime and task success. |
| Managed inference without owning serving | Qwen3.8-27B on Regolo | Endpoint limits, measured latency, error rate and real cost per accepted task. |
| Coding and code review | Pilot both on the same repository | Tests passed, seeded defects found, false positives and regressions. |
| Long document analysis | Verify endpoint capacity, then test both | Grounded answers and evidence spans at increasing input lengths. |
| Batch extraction | Compare accepted records, not JSON validity alone | Field-level F1, rejection/retry rates, throughput and cost per 1,000 accepted records. |
| Creative writing | Blinded human evaluation | Brand fit, factuality, editing minutes and first-draft acceptance. |
| Image understanding / OCR | Qwen as the first quality baseline | Real scans, critical-field exact match and confirmed multimodal endpoint support. |
The architecture and hosting options above follow the documented model/runtime specifications and our platform integration paths; the workload tests are our proposed evaluation protocol, not results already achieved.
Run the same labeled task set on both candidates. Record the model and runtime revisions, input/output budgets, prompts, tool permissions, concurrency and evaluation rules. Track end-to-end P50/P95 latency and verified successes, not only generation speed. Promote a deployment when it meets explicit business thresholds, then calculate the composite with the missing dimensions measured.
Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention
Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.
FAQ
Are these genuinely comparable model sizes?
Yes: they are in the same approximately 27B parameter family; Bonsai is derived from Qwen3.8-27B, compression changes the weight representation and deployment footprint, not the comparison into a different trillion-parameter class.
Is the reported Qwen capability score measured on Regolo?
No, 86.32 is prismml’s FP16 reference score, a matched pilot on our infrastructure is needed before claiming the same score or production performance on our endpoint.
Is Bonsai automatically supported by Regolo Custom Models?
Not established. Our documented custom route requires serving compatibility, while Bonsai’s ternary packs require specialized runtime kernels. Talk to our team first to confirm support.
Which has the best global score?
No verified winner is available. The evidence supports a task-dependent capability comparison, but not a complete matched ranking on cost, output limits and production suitability.
Ship Private AI. Not Infrastructure.
You have the private AI App architecture, bow give it an inference layer built for production.
Regolo gives European teams fast, OpenAI-compatible access to Mistral, Llama, Qwen, DeepSeek, GLM, and more — with zero data retention, EU data residency, and no new SDK to learn.
Change your base_url. Keep your LangChain code. Start shipping.
🚀 Start your 30-day free trial →
Build, test, and deploy with no infrastructure to maintain.
No credit card. No migration project. No compromise on data control.
💬 Join the Regolo Discord →
Meet builders working on private RAG, local LLMs, LangChain, Ollama, and production AI systems. Share your setup, get feedback from the community, and speak directly with the Regolo team.
🤝 Talk to an AI Infrastructure Engineer →
Running a sensitive workload, scaling beyond a proof of concept, or assessing a managed EU inference provider? Get a tailored architecture and commercial proposal for your team.
📂 Clone the GitHub repository →
Get the full implementation from this guide: ingestion scripts, ChromaDB setup, hybrid retrieval, the 30-Question RAG Floor, evaluation examples, and deployment configuration.
Private AI should not require a private data center.
Regolo gives your team an EU-native path from local experimentation to production-grade inference.
Build with Regolo
- Discord: Join the community →
- GitHub: Explore open-source workflows →
- X / Twitter: Follow @regolo_ai →
- Reddit: Join the community →
- Documentation: Read the API docs →
- Contact: Talk to the team →
Built with ❤️ by the Regolo team. Questions? regolo.ai/contact or chat with us on Discord