AI RACE— The AI Race
Research

Nvidia's SoL-Pi System Halves AI Agent Token Usage by Optimizing Control Logic

Nvidia researchers have developed SoL-Pi, an automated system that rewrites the control harness for AI coding agents to reduce token consumption by nearly half without sacrificing task performance.

09/26/2026, 17:30
Hệ thống SoL‑Pi của Nvidia giảm tới nửa lượng token tiêu thụ của các agent lập trình nhờ tối ưu “harness”

Automated Harness Optimization Cuts Agent Token Costs

Nvidia researchers have introduced SoL-Pi, a system that automatically optimizes the control layer—known as the harness—used by autonomous AI coding agents. Detailed in a paper published on September 26, 2026, the framework slashes agent token usage by 44.7 to 49 percent while keeping task accuracy roughly equal to baseline setups.

As AI agents run unsupervised across extended workflows, costs expand dramatically as simple model queries compound into long chains of reasoning, repeated tool calls, and execution feedback loops. While standard efficiency techniques focus on model-level optimizations—such as quantization, faster attention kernels, or swapping in smaller models—SoL-Pi targets the underlying harness that manages how agents track state, execute tools, and compress context. Building on principles of recursive self-improvement, the framework uses a research agent to observe execution traces from tools like Codex, Claude Code, and OpenClaw, propose leaner control logic, and automatically evaluate candidate changes in sandboxed environments.

Four Key Mechanisms and Benchmark Performance

To search for leaner control code, Nvidia ran more than 3,000 runs comprising over 60,000 agent-environment interactions across 535 executable environments, which included 495 GitHub issue-pull-request pairs and 40 synthetic test cases. To prevent the automated system from overfitting to its training tasks, the researchers strictly isolated search feedback from final evaluation by holding out the EdgeBench benchmark. Of EdgeBench's 51 public tasks, 11 were used for a one-time candidate validation, while the remaining 40 were reserved strictly for un-blinded final testing.

The automated search yielded four distinct operational mechanisms:

  • Action Fusion: Merges two sequential steps, such as code editing and test execution, into a single action to eliminate an entire language model call.
  • Online Context Compact: Trims non-essential accumulated context directly after planning steps.
  • ObservationPack: Archives lengthy tool outputs and replaces them with concise summaries in subsequent processing steps.
  • Evidence-Preserving Reducer: Reroutes large error logs and test outputs through a cheaper secondary model to pull key findings, backed by an automated verification step to recover critical details if missed.

On EdgeBench, SoL-Pi's most efficient setup—combining all four mechanisms—reduced token consumption by 49 percent while achieving 93.7 percent of the baseline Pi harness's performance, lowering overall testing run costs from $1,339 to $894. A performance-focused variant selecting only the single strongest mechanism beat the original Pi harness score by 5.3 percent while still reducing token usage. At current API rates, the authors estimate hourly operational savings of $8.75 to $13.50 compared to native Codex and Claude Code harnesses, and $4.36 to $5.71 per hour compared to standard Pi setups.

The researchers initially developed the system using GPT-5.6 Sol and subsequently transferred it to Anthropic's Opus 5 without structural modifications, maintaining 94.3 percent of baseline Pi performance. Beyond EdgeBench, SoL-Pi solved 15 out of 63 CPU tasks on Terminal-Bench 4—compared to 18 solved by standard Codex and Pi—at a 25 percent cost reduction. On formal mathematical verification tasks in Lean 4 from the IMO 2026 benchmark, SoL-Pi solved three of six problems at the lowest cost per solved task. Additionally, a swarm experiment involving 20 SoL-Pi workers achieved top results in kernel optimization while cutting costs by 26.8 percent compared to a baseline Pi swarm.

Industry Demand Surges as Agentic Overhead Grows

The development of SoL-Pi comes as software teams face ballooning costs driven by autonomous agent workflows. OpenRouter analyst Peter Walker noted that agentic token consumption has increased 14-fold since February 2026, with nearly 70 percent of that volume stemming from cached prompts. The structural impact of harnesses was highlighted in an August evaluation by tooling firm Composio, which ran DeepSeek V4 Flash across four framework setups—including Claude Code and Oh My Pi—and found that the cost per solved task varied by nearly three times despite using the exact same underlying model.

However, aggressive harness trimming presents technical trade-offs. Shortening context windows can disrupt prompt cache reuse, occasionally offsetting theoretical token savings. Furthermore, separate studies on context compression show that condensed prompts preserve an average of only 17 percent of initial user instructions. Multi-agent designs also encounter scaling limits; Codex developer Eric Provencher recently cautioned that deploying more than two sub-agents almost always burns tokens without improving output quality, as agents spend excessive compute validating each other's work. Looking forward, Nvidia researchers propose pretraining harnesses across broad task libraries to establish recursive efficiency gains in future systems.

◗ Sources

The Decoder09/26

Related stories