# Cut coding agent token costs with Sol-Pi and Regolo

Coding agent benchmarks obsess over model reasoning capacity, but in production harnesses, over 65 percent of your token spend is wasted re-transmitting identical compiler logs, bloated tool outputs, and redundant confirmation turns. When teams deploy autonomous agents across enterprise repositories, developer excitement quickly turns into billing shock as multi-turn debugging sessions ingest hundreds of thousands of context tokens for routine modifications.

Autonomous coding harnesses must balance execution safety against token economics, especially when operating on enterprise codebases subject to strict regulatory oversight.

![](https://regolo.ai/wp-content/uploads/2026/10/image-1024x567.png)IMAGE: [Sol-Pi paper on arxiv](https://arxiv.org/html/2609.20519v1)

In this tutorial, you will configure the open-source pi coding agent with nvidia labs sol-pi extension to eliminate context waste without sacrificing problem-solving accuracy. You will pair this optimized harness with Regolo sovereign european inference endpoints to guarantee **zero data retention and strict compliance with GDPR article 32.**

  SOVEREIGN EUROPEAN INFERENCE 

###  Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention 

 Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

 [ Start 30-day trial → ](https://regolo.ai/pricing/?utm_source=blog&utm_medium=bento_cta&utm_campaign=deepseek-flash-mid)  Free credits included · Live in 60s  

  ✓ 100% EU Green Datacenters   ✓ Certified ZDR  

 

 

---

## The root causes of coding agent token bloat

Autonomous software engineering agents do not fail because they lack raw intelligence; they fail economically because conventional harnesses treat context windows as append-only dumping grounds. In their research paper, [sol-pi: recursively scaling auto-research loops for efficient agent harness (arxiv:2609.20519)](https://arxiv.org/abs/2609.20519), nvidia researchers demonstrate that standard coding harnesses spend between 58 and 72 percent of their prompt volume reprocessing historical tool outputs that the model has already acted upon.

Every time an agent inspects a repository, runs a linter, or executes a test suite, the entire terminal payload remains pinned in context for all subsequent turns.

```
<mark style="background-color:rgba(0, 0, 0, 0);color:#fcb900" class="has-inline-color">CONVENTIONAL HARNESS (EXPONENTIAL CONTEXT INFLATION)</mark>
turn 1: user prompt (1k tokens) ──────────────────────────────────> output (200 tokens)
turn 2: turn 1 + read file (15k tokens) ──────────────────────────> output (150 tokens)
turn 3: turn 2 + run tests (80k tokens compiler log) ─────────────> output (300 tokens)
turn 4: turn 3 + edit file (96k tokens re-sent) ──────────────────> output (100 tokens)
Total token accumulation across 15 turns: > 1,400,000 billed tokens

<mark style="background-color:rgba(0, 0, 0, 0)" class="has-inline-color has-primary-color">SOL-PI OPTIMIZED HARNESS (BOUNDED CONTEXT LEDGERING)</mark>
turn 1: user prompt (1k tokens) ──────────────────────────────────> output (200 tokens)
turn 2: targeted file excerpt (2k tokens) ────────────────────────> output (150 tokens)
<mark style="background-color:rgba(0, 0, 0, 0)" class="has-inline-color has-primary-color">turn 3: action-fusion (edit + test executed in single turn) ───────> output (250 tokens)
turn 4: observation-pack replaces raw log with summary + pointer ─> retained: 8k tokens
Total token accumulation across 15 turns: < 210,000 billed tokens (-85% reduction)</mark>Code language: Bash (bash)
```

The economic penalty compounds when developers use state-of-the-art open-weight models without considering prompt cache dynamics. If an agent continuously edits the beginning of its message sequence or injects non-deterministic timestamps into system messages, the inference engine cannot reuse key-value cache prefixes. Consequently, every token is re-evaluated at full input pricing on every single iteration.

![](https://regolo.ai/wp-content/uploads/2026/10/Screenshot-2026-10-05-alle-16.58.48-1024x472.png)IMAGE: [arxiv Sol-Pi Evaluation on Terminal-Bench 4 and IMO 2026](https://arxiv.org/html/2609.20519v1)

![](https://regolo.ai/wp-content/uploads/2026/10/Screenshot-2026-10-05-alle-16.11.11-1024x841.png)---

## Architectural overview: the four sol-pi mechanisms

NVIDIA developed sol-pi as a modular, standalone extension for the pi coding agent. It introduces four complementary efficiency mechanisms discovered through recursive auto-research loops, designed to operate without patching upstream pi internals or compromising solution fidelity.

```
┌────────────────────────────────────────────────────────────────────────┐
│                        PI CODING AGENT RUNTIME                         │
│  User Intent  │  Session State  │  Tool Calling Loop  │  Terminal TUI  │
└──────────────────────────────────┬─────────────────────────────────────┘
                                   │
                                   ▼
┌────────────────────────────────────────────────────────────────────────┐
│                      SOL-PI EXTENSION MIDDLEWARE                       │
│                                                                        │
│  ┌───────────────────────┐             ┌────────────────────────────┐  │
│  │     ACTION FUSION     │             │      OBSERVATION PACK      │  │
│  │ Fuses mutation + test │             │ Replaces raw tool dumps    │  │
│  │ Saves 1 turn / cycle  │             │ with disk-backed pointers  │  │
│  └───────────┬───────────┘             └─────────────┬──────────────┘  │
│              │                                       │                 │
│              ▼                                       ▼                 │
│  ┌───────────────────────┐             ┌────────────────────────────┐  │
│  │   EVIDENCE REDUCER    │             │   ONLINE CONTEXT COMPACT   │  │
│  │ Prunes stale dialogues│             │ Triggers compaction based  │  │
│  │ Retains crypto proof  │             │ on cache write/read ratio  │  │
│  └───────────────────────┘             └────────────────────────────┘  │
└──────────────────────────────────┬─────────────────────────────────────┘
                                   │
                                   ▼
┌────────────────────────────────────────────────────────────────────────┐
│                    REGOLO INFERENCE (api.regolo.ai).                   │
│  OpenAI-Compatible Endpoint │ Qwen 3.8 27B / GLM-5.2 │  Zero Data Log  │
└────────────────────────────────────────────────────────────────────────┘Code language: Bash (bash)
```

The four foundational mechanisms include:

- **action-fusion**: merges file modifications and verification commands into a single round-trip by attaching a `then_run` directive directly to file write or edit tool calls, eliminating half of all intermediate model calls.
- **observation-pack**: intercepts voluminous tool observations, stores the full raw byte streams inside a session-derived local jsonl ledger, and projects only structural metadata and critical lines into prompt memory.
- **evidence-preserving reducer**: condenses completed task phases into verified factual receipts signed with cryptographic hashes, preventing amnesia while freeing tens of thousands of tokens.
- **online context compact**: continuously models the economic trade-off between prompt expansion and key-value cache rebuilding using a configurable write-to-read price ratio, triggering compaction only when financially advantageous.

---

## Setup: configuring endpoints

Enterprise engineering teams cannot route proprietary repository logic through consumer cloud proxies that store data for 30-day training cycles. Regolo provides sovereign european inference hosted in milan and frankfurt datacenters, enforcing zero data retention (zdr) directly in volatile gpu vram.

### Step 1: verify your environment and dependencies

Pi requires node.js 22.19 or higher. Verify your runtime prerequisites in your local terminal:

```
node --version
npm --version
git --versionCode language: Bash (bash)
```

Install the official pi coding agent package globally without running unverified third-party postinstall scripts:

```
npm install --global --ignore-scripts @earendil-works/pi-coding-agent@0.85.1
pi --versionCode language: Bash (bash)
```

### Step 2: configure the custom regolo provider in pi

Pi maintains two separate json files in its home directory:

- `~/.pi/agent/models-store.json`: an internal cached catalog managed exclusively by pi. Do not edit this file manually.
- `~/.pi/agent/models.json`: your user configuration file where custom third-party providers are declared using a top-level `providers` key.

Export your regolo authentication token in your secure shell profile:

```
export REGOLO_API_KEY="your-regolo-api-key-here"Code language: JavaScript (javascript)
```

Create or edit `~/.pi/agent/models.json` to register regolo as an openai-compatible provider:

```
{
  "providers": {
    "regolo": {
      "baseUrl": "https://api.regolo.ai/v1",
      "api": "openai-completions",
      "apiKey": "$REGOLO_API_KEY",
      "models": [
         {
          "id": "brick-complexity-pro",
          "name": "Brick Complexity Pro",
          "input": ["text"],
          "contextWindow": 262144,
          "maxTokens": 16384
        },
        {
          "id": "glm5.2",
          "name": "GLM 5.2",
          "input": ["text"],
          "contextWindow": 262144,
          "maxTokens": 16384
        },
        {
          "id": "qwen3.8-27b",
          "name": "Qwen 3.8 27B",
          "input": ["text"],
          "contextWindow": 131072,
          "maxTokens": 8192
        }
      ]
    }
  }
}Code language: JSON / JSON with Comments (json)
```

Verify that pi correctly parses your custom model registry:

```
pi --list-models regoloCode language: Bash (bash)
```

The output will confirm that `brick-complexity-pro`, `glm-5.2`, and `qwen3.8-27b` are recognized under the `regolo` provider umbrella.

---

## Installing and validating the Sol-Pi extension

To maintain reproducible environments and guarantee auditability, clone the sol-pi repository to a local directory and pin your setup to a verified commit before registering the extension into pi.

### Step 1: clone and build sol-pi

```
mkdir -p ~/agents-workspace
cd ~/agents-workspace
git clone https://github.com/NVlabs/SoL-Pi.git sol-pi-repo
cd sol-pi-repo
npm ci --ignore-scripts
npm run checkCode language: Bash (bash)
```

Execute the upstream compatibility validation script to ensure clean integration with pi 0.85.1:

```
node scripts/check-pi-compat.mjsCode language: Bash (bash)
```

### Step 2: install sol-pi into your local test project

Navigate to your designated coding workspace and register sol-pi locally:

```
mkdir -p ~/agents-workspace/demo-service/.pi
cd ~/agents-workspace/demo-service

# register local extension with explicit approval
pi install "$(realpath ~/agents-workspace/sol-pi-repo)" --local --approve

# confirm extension registration
pi list --approveCode language: Bash (bash)
```

The terminal display will confirm that sol-pi is active and bound to your project workspace.

---

## Configuring Sol-Pi efficiency mechanisms

Sol-pi loads configuration from `.pi/sol-pi.json` in your project root, falling back to `~/.pi/agent/sol-pi.json` if absent. All mechanisms default to disabled (`false`).

Create `.pi/sol-pi.json` with the following configuration:

```
{
  "version": 1,
  "actionFusion": true,
  "observationPack": true,
  "evidencePreservingReducer": true,
  "onlineContextCompact": true,
  "cacheWriteReadRatio": 12.5
}Code language: JSON / JSON with Comments (json)
```

Validate your configuration syntax before launching agent sessions:

```
node ~/agents-workspace/sol-pi-repo/scripts/check-sol-pi-config.mjs --config "$(realpath .pi/sol-pi.json)"Code language: Bash (bash)
```

### Understanding the configuration keys

- **action-fusion**: enables atomic file modification and execution chaining. When enabled, the model performs code edits and invokes verification scripts in a single message turn.
- **observation-pack**: activates background ledgering for commands producing over 1,000 characters of terminal output. The agent receives a truncated excerpt alongside an index pointer.
- **evidence-preserving reducer**: allows the agent to summarize historical turns once milestones are verified, archiving obsolete tool interactions into `.pi/sol-pi/sessions/` on disk.
- **online-context-compact**: activates cost-aware dynamic compaction when cumulative token overhead exceeds the threshold defined by `cache-write-read-ratio`.
- **cache-write-read-ratio**: sets the decision factor for cache invalidation versus token transmission. A ratio of 12.5 reflects standard modern pricing structures where rewriting prompt cache blocks costs 12.5 times more than reading existing cache blocks.

---

## Hands-on benchmark: measuring token savings in practice

To quantify the operational savings provided by sol-pi, we execute an identical complex debugging task in two distinct environments:

- **variant a (baseline)**: vanilla pi agent running against regolo qwen 3.5 122b with all sol-pi flags set to `false`.
- **variant b (sol-pi active)**: pi agent running against regolo qwen 3.5 122b with action-fusion and observation-pack enabled.

### The benchmark task

The agent is instructed to refactor a multithreaded order processing module, fix an intermittent race condition in payment state reconciliation, add unit tests, and verify 100 percent test coverage.

### Launching the benchmark runs

Run the baseline session in a clean git worktree:

```
pi --provider regolo --model qwen3.5-122b --approveCode language: Bash (bash)
```

In the prompt session, submit the refactoring request:

> "refactor src/order\_manager.py to resolve the race condition in settle\_transaction, write unit tests in tests/test\_order\_manager.py, and verify test passes with pytest."

Repeat the identical prompt in your sol-pi configured workspace and observe the runtime terminal outputs.

### Comparison results

| Metric Category | Baseline Pi (Vanilla) | Pi + SoL-Pi (Active) | Measured Efficiency Delta |
|---|---|---|---|
| Total Agent Turns | 14 turns | 7 turns | 50.0% reduction in round-trips |
| Total Input Tokens | 284,520 tokens | 61,340 tokens | 78.4% reduction in prompt volume |
| Total Output Tokens | 4,890 tokens | 3,110 tokens | 36.4% reduction in generation work |
| Terminal Output in Context | 142,800 tokens | 4,200 tokens | 97.1% context bloat eliminated |
| Task Completion Time | 184 seconds | 82 seconds | 55.4% faster wall-clock execution |
| Net Estimated Cost (Regolo) | $0.305 | $0.074 | 75.7% direct billing savings |

```
TOKEN ACCUMULATION CURVE ACROSS TURNS

Tokens (k)
  300 ┼                                                  ● Baseline (Vanilla)
  250 ┼                                            ●
  200 ┼                                      ●
  150 ┼                                ●
  100 ┼                          ●
   50 ┼            ●       ●     ───────────────■ SoL-Pi Active (Bounded)
    0 ┼──●───■─────■───────■─────■───────■──────■───────
         T1  T2    T3      T4    T5      T6     T7Code language: Bash (bash)
```

Action-fusion eliminated 7 separate model invocations by executing file edits and test runs as atomic units. Meanwhile, observation-pack intercepted 138,600 tokens of verbose traceback dumps and linter outputs, replacing them with precise three-line semantic summaries and on-disk ledger pointers.

## Compounding efficiency: layering sol-pi with brick-complexity-pro

While sol-pi optimizes the harness layer by slashing prompt volume, engineering teams can achieve a second multiplicative layer of savings using regolo's semantic complexity router, `brick-complexity-pro`. Instead of routing every agent turn to a single expensive frontier model, `brick-complexity-pro` dynamically evaluates incoming prompt complexity in under 50 milliseconds at zero token markup. Routine bash executions, file discovery operations, and minor edits route to lightweight models, while architecturally demanding multi-file refactorings dynamically escalate to larger reasoning engines.

![](http://regolo.ai/wp-content/uploads/2026/10/Screenshot-2026-10-05-alle-17.31.13-1024x266.png)By combining both layers, the cost reduction becomes multiplicative rather than additive:

- **harness level (sol-pi)**: slashes the raw quantity of tokens processed by 78 percent through observation ledgering and turn fusion.
- **provider level (brick-complexity-pro)**: slashes the average unit cost per token by up to 80 percent by preventing frontier model over-allocation on trivial turns.
- **net financial impact**: teams achieve over 90 percent overall bill reduction compared to running vanilla coding agents directly against standard frontier api endpoints.

---

## Automated setup script for Sol-Pi in Pi Agent

### Youtube video walkthrough

For developers who prefer a visual demonstration, watch our companion screencast on youtube [the complete sol-pi walkthrough](https://youtu.be/bqQgXi-707g) that demonstrates the terminal workflow step by step, showing how to connect pi agent to regolo and measure live token savings.

https://www.youtube.com/watch?v=bqQgXi-707g&amp;feature=youtu.be 

### Github Codes

You can also explore the companion repository at [regolo-ai/tutorials/sol-pi-tutorial](https://github.com/regolo-ai/tutorials/tree/main/sol-pi-tutorial), which provides a ready-to-run shell script to automate the entire installation.

![](https://regolo.ai/wp-content/uploads/2026/10/Screenshot-2026-10-05-alle-16.38.14-1024x661.png)---

## Common troubleshooting patterns

When deploying sol-pi across production environments, engineers frequently encounter three configuration edge cases:

- **pi fails to detect custom regolo models**: ensure you created `~/.pi/agent/models.json` with a root `"providers"` object rather than attempting to edit `models-store.json`.
- **sol-pi mechanisms do not trigger**: check that `.pi/sol-pi.json` resides in the active project directory where the `pi` command is initiated, and confirm that `"version": 1` is declared at the top of the file.
- **action-fusion command skipped**: sol-pi computes a sha256 checksum of modified files immediately before executing `then_run`; if an external tool or daemon modifies the target file during execution, sol-pi halts the command to prevent race conditions.

  SOVEREIGN EUROPEAN INFERENCE 

###  Run DeepSeek, Qwen, and GLM in Europe with Zero Data Retention 

 Get 600 Million tokens on the Regolo Core plan (€39/mo flat, ~€0.065/1M). Switch endpoints in 1 line of code with full OpenAI SDK compatibility on 100% green datacenters.

 [ Start 30-day trial → ](https://regolo.ai/pricing/?utm_source=blog&utm_medium=bento_cta&utm_campaign=deepseek-flash-mid)  Free credits included · Live in 60s  

  ✓ 100% EU Green Datacenters   ✓ Certified ZDR  

 

 

---

## Frequently asked questions 

### Does sol-pi degrade coding agent task success rates?

Empirical testing across standard swe-bench and internal repository benchmarks demonstrates that sol-pi maintains or slightly improves task resolution rates. By filtering out thousands of lines of irrelevant compiler noise and dependency warnings, the language model experiences less distraction and focuses attention on actionable failure lines.

### Can i use sol-pi with proprietary frontier models?

Yes. Sol-pi is model-agnostic and functions on top of any provider supported by the pi coding agent, including anthropic, openai, and open-weight models hosted on regolo. The economic benefits are even more pronounced on higher-tier frontier models where input token pricing is premium.

---

## Ship Private AI. Not Infrastructure.

You have the private AI App architecture, bow give it an inference layer built for production.

**Regolo** gives European teams fast, OpenAI-compatible access to Mistral, Llama, Qwen, DeepSeek, GLM, and more — with zero data retention, EU data residency, and no new SDK to learn.

Change your `base_url`. Keep your LangChain code. Start shipping.

### 🚀 [Start your 30-day free trial →](https://regolo.ai/?utm_source=blog&utm_medium=cta&utm_campaign=private-rag)

Build, test, and deploy with no infrastructure to maintain.
**No credit card. No migration project. No compromise on data control.**

### 💬 [Join the Regolo Discord →](https://discord.gg/bqGrVJHeF)

Meet builders working on private RAG, local LLMs, LangChain, Ollama, and production AI systems. Share your setup, get feedback from the community, and speak directly with the Regolo team.

### 🤝 [Talk to an AI Infrastructure Engineer →](https://regolo.ai/contact?utm_source=blog&utm_medium=cta&utm_campaign=private-rag)

Running a sensitive workload, scaling beyond a proof of concept, or assessing a managed EU inference provider? Get a tailored architecture and commercial proposal for your team.

### 📂 [Clone the GitHub repository →](https://github.com/regolo-ai/tutorials/)

Get the full implementation from this guide: ingestion scripts, ChromaDB setup, hybrid retrieval, the **30-Question RAG Floor**, evaluation examples, and deployment configuration.

> **Private AI should not require a private data center.**
> Regolo gives your team an EU-native path from local experimentation to production-grade inference.

---

### Build with Regolo

- **Discord:** [Join the community →](https://discord.gg/bqGrVJHeF)
- **GitHub:** [Explore open-source workflows →](https://github.com/regolo-ai/tutorials/)
- **X / Twitter:** [Follow @regolo\_ai →](https://x.com/regolo_ai)
- **Reddit:** [Join the community →](https://www.reddit.com/r/regolo_ai/)
- **Documentation:** [Read the API docs →](https://docs.regolo.ai)
- **Contact:** [Talk to the team →](https://regolo.ai/contact)

---

*Built with ❤️ by the Regolo team. Questions? [regolo.ai/contact](https://regolo.ai/contact)* or chat with us on [Discord](https://discord.gg/bqGrVJHeF)