Research Analysis
What NVIDIA’s SoL-Pi Research Tells Us About Agent Costs
NVIDIA’s SoL-Pi paper shows that the orchestration layer between model and environment is a primary lever for token efficiency. Key findings and enterprise implications.
Published on


What NVIDIA’s SoL-Pi Research Tells Us About Long-Running Agent Costs
A new paper from NVIDIA’s Efficient AI team shows that the orchestration layer between model and environment is a primary lever for token efficiency. What the research found, what its limits are, and what it means for enterprise teams governing long-running agents.
By Tahir Mahmood, Co-founder and CTO, OpenBox | 20 September 2026
The Cost Problem the Paper Actually Addresses
Coding agents no longer stop at a file edit and a test run. As they take on longer tasks across multiple files and repositories, they accumulate something engineers do not always price carefully: token overhead that compounds across every turn of the conversation.
A paper from researchers at NVIDIA, NTU, and MIT, submitted to arXiv on 17 September 2026, puts a number on this. The researchers studied what they call the harness layer: the orchestration code that presents state to the model, exposes tools, and processes feedback. Their argument is that the harness is where substantial token overhead accumulates, and where it can be reduced without retraining the model or switching to a cheaper one.
Their finding is specific and defensible: the same frontier model, running the same benchmark tasks, can be made substantially less expensive by changing the orchestration layer, though the most cost-intensive mechanism (Evidence-Preserving Reducer) does delegate log compression to a lower-cost auxiliary model, so the intervention is not purely deterministic wrapper code. That the harness is a primary lever for cost is the claim the paper tests, and the one that matters for engineering and governance teams planning for agent deployments at scale.
How the Research Was Conducted
The team took what they call an RSI-inspired (recursive self-improvement) approach to harness discovery. Rather than manually inspecting execution traces and writing fixes by hand, they built an automated research pipeline in which an AI system proposes harness changes, tests them in prepared environments, and filters candidates through capability and efficiency gates.
The outer search considered 152 proposed directions across six families: context, progress, tools, delegation, prompt and policy, and improvement and evaluation. The inner search developed each candidate independently, following an iterative implementation and review loop. Across the full search, the team report more than 3,000 runs and more than 60,000 agent-environment interactions across 535 executable environments.
Four mechanisms survived capability-constrained selection and were integrated into SoL-Pi. The paper keeps EdgeBench separate from the iterative search: none of the EdgeBench tasks are used to guide or revise candidate development. Of the 51 public EdgeBench tasks, 11 are used for one-way acceptance of already-frozen candidates after search concludes, and the remaining 40 are reserved for the final generalization evaluation. Held-out results do not feed back into the search loop.
The Four Discovered Mechanisms
The mechanisms target different parts of the agent workflow. The paper evaluates them in combination to verify their joint effect, and the evidence is consistent with complementarity, though the paper notes that comparisons on each mechanism’s own triggered-task subset do not isolate interaction effects between them.
Action Fusion
When an agent edits a file, the next step is usually a predictable validation command. Action Fusion combines the edit and the follow-up into one tool call, returning both outcomes in a single observation and eliminating the intermediate model round trip. Commands that require inspecting the result of the mutation stay separate.
Online Context Compact
Context compaction carries a cost because it breaks prompt-cache reuse. Online Context Compact uses plan-step completion as the trigger point and applies compaction only when projected input savings exceed the estimated cost of rewriting the cache. Later compactions require a larger savings margin to account for previously unrecovered rewrite costs.
ObservationPack
Large tool outputs are replayed in full on every subsequent request even when little of their content remains relevant. ObservationPack archives results exceeding 10 KiB locally, sends them in full for the first two provider requests, then substitutes a stable handle plus a 1 KiB excerpt from the third request onward. The agent retrieves exact pages on demand.
Evidence-Preserving Reducer
Build and test logs of at least 4 KiB are compressed by a lower-cost auxiliary model (the paper specifies GPT-5.6 Luna) into compact receipts. A deterministic verifier checks the receipt against the original log for schema, source hash, exit status, exact quotations, and size. The harness falls back to the original if verification fails, if credentials are suspected, or if the receipt provides no size reduction. File reads and search results bypass the reducer entirely.
What the Results Actually Show
The paper evaluates SoL-Pi on EdgeBench (51 public tasks), Terminal-Bench 4 (63 CPU-only tasks), IMO 2026 (six formal-proof problems), and a multi-agent kernel-optimisation experiment. The EdgeBench results are the primary evidence; the others serve as generalisability checks.
All token cost figures in the paper use API prices as of 17 August 2026. Costs will differ at different price points.
EdgeBench
The paper reports two operating points. SoL-Pi [Efficiency], the complete four-mechanism stack, reduces recorded token traffic by 49.0% and API cost by 33.2% compared with Pi (from $1,339 to $894), while retaining 93.7% of Pi’s average score (42.0 versus 44.8 on a 51-task set). Against the native Codex harness, API cost falls 50.0%; against Claude Code on Opus 5, it falls 54.3%.
SoL-Pi [Performance], which uses ObservationPack alone under GPT-5.6 Sol, raises the average score from 44.8 to 47.2 (a 5.3% gain) while reducing token traffic by 6.1% and improving token efficiency by 9.8%.
The paper estimates hourly savings of $8.75 to $13.50 relative to native Codex and Claude Code harnesses, and $4.36 to $5.71 relative to Pi. These are the paper’s own estimates based on fixed benchmark conditions and should be treated as indicative rather than guaranteed for any given workload.
Terminal-Bench 4 and IMO 2026
On 63 CPU-only Terminal-Bench 4 tasks, SoL-Pi solves 15 tasks compared with 18 for both Codex and Pi. The paper reports that SoL-Pi’s total model cost of $211.12 represents a 26.3% reduction from Pi’s $286.45, and a lower cost per solved task ($14.07 versus $15.91). On the IMO 2026 evaluation, SoL-Pi and Pi each pass three of the six problems while Codex passes five; SoL-Pi achieves the lowest cost per passed problem at $20.90, compared with $22.89 for Codex and $25.32 for Pi.
These results show that efficiency gains generalise somewhat beyond EdgeBench, though the Terminal-Bench 4 task-count comparison is unfavourable to SoL-Pi: it solves fewer tasks than its baselines in that setting.
The Multi-Agent Dimension
The paper includes a two-hour multi-agent kernel-optimisation experiment relevant to teams running agent swarms. Three configurations are compared: a single Codex agent, a Codex coordinator with 20 Pi-baseline workers, and a Codex coordinator with 20 SoL-Pi workers.
The SoL-Pi swarm achieves 1,127 optimisation cycles at a total API cost of $60.11. The Pi-baseline swarm achieves 1,366 cycles at $82.12. The single agent achieves 1,333 cycles at $39.20 and remains the cheapest configuration overall. The SoL-Pi swarm and Pi-baseline swarm use the same number of workers; the difference between them is the harness. Under the same budget, the SoL-Pi swarm reduces API cost by 26.8% relative to the Pi-baseline swarm and achieves a better optimisation result (lower cycle count is better in this benchmark).
The paper also notes that the SoL-Pi swarm and the single agent both pass all eight speed thresholds in the benchmark, while the Pi-baseline swarm passes seven, missing the final threshold of fewer than 1,363 cycles. For enterprise teams evaluating agent swarms, the experiment shows that a more efficient harness can produce meaningfully better results than a less efficient one at lower cost for the same worker count, and that the single-agent configuration remains cheaper in absolute terms for this particular optimisation task.
Why the Harness Layer Is Also Where Governance Sits
In OpenBox’s architecture, governance controls for long-running agents operate at the harness boundary: policy is evaluated before each action executes, evidence is collected and linked across turns, and multi-step behavioural patterns are detected within a session. The orchestration code that moves tokens is the same layer at which governance enforcement decisions are made.
This has a practical implication. An agent that runs for two hours benefits from sustained runtime behavioural monitoring, not just a log review after the fact. A more efficient harness does not automatically produce a more governable agent, but it does mean that a fixed governance budget covers more turns of agent activity for the same cost.
The SoL-Pi paper’s Evidence-Preserving Reducer reflects a design philosophy that is relevant to audit practice. The mechanism will not compress a log if the compressed version cannot be verified against the original. This mirrors the principle behind tamper-evident audit trails: the compact record should be provably derivable from the source. Efficiency and verifiability are not in tension; they are compatible requirements on the same artefact.
Online Context Compact reduces the active context at plan-step boundaries, which changes what the model can see. Behavioral Rules, which detect multi-step patterns within a session, need to be aware of compaction events. A compaction changes what is active in the model’s context; the governance audit record should be kept separately so that compaction does not reduce what is available for post-hoc verification.
The multi-agent swarm experiment is a reminder that governance at the agent-to-agent boundary matters as much as governance within a single agent. In SoL-Pi’s swarm architecture, workers exchange notes and the coordinator relays findings across groups. OpenBox’s per-agent Trust Scores and multi-agent session instrumentation are designed to give the coordinator-level view that each worker cannot generate for itself, though this is a governance architecture observation, not a conclusion the SoL-Pi paper itself draws.
What the Research Does Not Show
The paper is careful about its own claims, and any summary should be equally careful. The core EdgeBench evidence covers 51 of 134 total tasks in the benchmark. The paper acknowledges that SoL-Pi was developed with GPT-5.6 Sol as the primary backend.
On Terminal-Bench 4, SoL-Pi solves three fewer tasks than either Codex or Pi (15 versus 18). The paper reports this without elision. Efficiency gains transfer to task cost, but not to task completion rate in that setting.
The mechanisms trigger less frequently on Opus 5 than on GPT-5.6 Sol, which the authors attribute to the harness being developed exclusively on GPT-5.6 Sol trajectories. The paper describes the Opus 5 results as consistent with cross-model transfer, and its conclusion characterises them as strong cross-model generalisation within the evaluated setting.
The hourly savings estimates are derived from benchmark conditions at August 2026 API prices. Real-world savings will vary with workload composition, task duration, and model pricing, which changes over time.
What Enterprise Teams Should Take From This
The SoL-Pi paper is evidence that the model is not the only lever. Orchestration design, taken seriously, can produce substantial cost reductions while preserving the majority of task performance. That matters for enterprise AI budgeting, where the cost of a long-horizon agent run is often treated as fixed by the choice of model.
It also matters for governance programme design. Governance controls at the harness layer add latency and token overhead to each turn. A harness that is already efficient makes governance additions easier to justify. The argument that governance is too expensive for production agents is weakest when the underlying harness cost has already been reduced.
The evidence-preserving design philosophy in SoL-Pi translates directly into audit practice. If you are running a long-horizon coding agent and need to reconstruct what it saw and why it acted, the pattern of archiving the source and verifying any compressed version against it is the same principle behind session replay and post-hoc verification. Research and governance requirement converge on the same architecture: keep the source, compress for efficiency, verify the compression.
For teams evaluating agentic design patterns, SoL-Pi adds empirical weight to a pattern that is otherwise argued from first principles: the orchestration layer is not neutral infrastructure. It is a design choice with measurable consequences for cost, capability, and the surface area on which governance controls operate.
Key Sources
1. SoL-Pi arXiv:2609.20519 (submitted 17 Sep 2026): https://arxiv.org/abs/2609.20519 (accessed 20 Sep 2026) |
2. NVlabs/SoL-Pi GitHub repository (MIT licence): https://github.com/NVlabs/SoL-Pi (accessed 20 Sep 2026) |
3. SoL-Pi project blog: https://nvlabs.github.io/SoL-Pi/ (accessed 20 Sep 2026) |
4. OpenBox: Trust Lifecycle: https://docs.openbox.ai/trust-lifecycle.md (accessed 20 Sep 2026) |
5. OpenBox: Governance Decisions: https://docs.openbox.ai/core-concepts/governance-decisions.md (accessed 20 Sep 2026) |
6. OpenBox: Trust Scores: https://docs.openbox.ai/core-concepts/trust-scores.md (accessed 20 Sep 2026) |
7. OpenBox: Behavioral Rules: https://docs.openbox.ai/trust-lifecycle/authorize/behaviors.md (accessed 20 Sep 2026) |
8. OpenBox: Verify (post-hoc verification): https://docs.openbox.ai/trust-lifecycle/verify.md (accessed 20 Sep 2026) |
9. OpenBox: Session Replay: https://docs.openbox.ai/trust-lifecycle/session-replay.md (accessed 20 Sep 2026) |
10. OpenBox: Attestation and Cryptographic Proof: https://docs.openbox.ai/administration/attestation-and-cryptographic-proof.md (accessed 20 Sep 2026) |
11. OpenBox: Multi-Agent Sessions: https://docs.openbox.ai/administration/organization/teams/multi-agent-sessions.md (accessed 20 Sep 2026) |
12. OpenBox blog: Why AI Agents Need Runtime Monitoring: https://www.openbox.ai/blog/why-ai-agents-need-runtime-monitoring (accessed 20 Sep 2026) |
13. OpenBox blog: What Is AI Agent Governance?: https://www.openbox.ai/blog/what-is-ai-agent-governance (accessed 20 Sep 2026) |
14. OpenBox blog: AI Agent Governance Cannot Be Optional: https://www.openbox.ai/blog/ai-agent-governance-cannot-be-optional (accessed 20 Sep 2026) |
15. OpenBox blog: AI Governance Needs Proof, Not Logs: https://www.openbox.ai/blog/ai-governance-needs-proof-not-logs (accessed 20 Sep 2026) |
16. OpenBox blog: Your agent ran two hours. Who was watching?: https://www.openbox.ai/blog/your-agent-ran-two-hours.-who-was-watching (accessed 20 Sep 2026) |
17. OpenBox blog: Governance Starts Where Monitoring Ends: https://www.openbox.ai/blog/governance-starts-where-monitoring-ends (accessed 20 Sep 2026) |
18. OpenBox blog: Runtime Enforcement Comes Before Compliance: https://www.openbox.ai/blog/runtime-enforcement-comes-before-compliance (accessed 20 Sep 2026) |
19. OpenBox blog: Govern the Handoffs Between Your AI Agents: https://www.openbox.ai/blog/multi-agent-ai-governance (accessed 20 Sep 2026) |
20. OpenBox blog: Agentic Design Patterns: https://www.openbox.ai/blog/agentic-design-patterns (accessed 20 Sep 2026) |

