The model is no longer the whole agent article image

The model is no longer the whole agent

An AI agent can fail while using a perfectly capable model. It can stop reading a file too early, call the wrong tool, carry the wrong context forward, or take an action outside the boundary a company intended.

That is why the most interesting part of a new LangChain and NVIDIA release is not the model at its center. It is the argument that the surrounding harness and runtime are becoming products companies need to own.

On July 8, the companies released the NemoClaw for LangChain Deep Agents blueprint. It combines NVIDIA's open Nemotron 3 Ultra model, LangChain's Deep Agents Code harness, and NVIDIA OpenShell, a sandboxed runtime that applies policies to tools, systems, and data.

The three layers do different jobs. The model reasons. The harness plans tasks, manages context and memory, and decides how tools are used. The runtime controls where the agent runs and what it is allowed to touch. That separation gives a business more places to improve a system without retraining the model or replacing the entire stack.

LangChain reported an aggregate score of 0.86 at a cost of $4.48 for Nemotron 3 Ultra with its tuned harness. The next closest model in its evaluation cost $43.48. That is a large gap, but it needs the right label. This was LangChain's own agent evaluation suite, not an independent measure of every enterprise workload.

The useful result is not simply “10 times cheaper.” It is evidence that agent economics can change when a team tunes the system around a model. LangChain says the main lesson is that “agent performance improves when the model, harness, evals, and runtime are tuned together.”

The model is no longer the whole agent
The harness connects model reasoning to context, memory, tools, and task execution.

NVIDIA's technical walkthrough makes the point concrete. In one evaluation, the agent was asked for the last non-empty line in a large file. The first read returned a full page, but the answer was farther down. The model answered early instead of continuing with another offset.

The fix did not require new model weights. Developers added middleware that tells the agent when a file result probably continues and instructs it to keep paging. The failing read test moved from zero passes in three runs to three passes. The broader benchmark improved from 94 to 96 correct results out of 127.

That example is small enough to understand and important enough to generalize. Many agent failures are not pure intelligence failures. They are interface failures. The model does not know that a tool response was truncated, that an approval is required before a write, or that a customer record came from a stale source unless the surrounding system makes those facts visible.

Companies already encode operating knowledge in procedures, permissions, quality checks, and exceptions. An agent harness turns some of that knowledge into software. Its prompts, tool descriptions, traces, evaluation cases, routing rules, and memory design become a record of how the company wants work done.

Free HubSpot workshopBring one HubSpot problem to a free 30-minute callA screen-share walkthrough of your portal with me, not a salesperson, and a short roadmap at the end. No contract or credit card.Book the free workshop

LangChain calls those assets valuable intellectual property. That is a fair description. A generic model can be rented by every competitor. A tested harness that knows when to ask for clarification, which system owns a field, how to recover from a partial response, and when to stop is much harder to copy.

OpenShell adds the other half of the argument. An agent that can act needs boundaries outside its own instructions. NVIDIA describes NemoClaw as a path from prototype to governed deployment, with model routing, skill execution, state, observability, and runtime controls in one setup. The runtime can sandbox execution and apply policy to data access and sensitive actions.

That matters because a prompt is not a security boundary. Telling an agent not to access a directory or send data to an unapproved service is weaker than preventing that connection at runtime. Businesses need both behavioral instructions and enforced limits, especially for long-running agents working with code, financial records, customer data, or production systems.

The open label also deserves scrutiny. The blueprint gives teams more control over model, harness, and runtime choices, and NVIDIA says it can run anywhere. But the stack still creates dependencies. Teams must understand the licenses, hosting requirements, NVIDIA infrastructure options, LangChain interfaces, and the work required to operate the components themselves.

Ownership is not the same as simplicity. A company can avoid one closed platform and still end up maintaining a complicated collection of open parts. The deciding question is whether that control produces better reliability, lower total cost, or a real compliance advantage for the workload.

Agent design is turning into process design. The business value sits in making the hidden operating rules explicit: what the agent may read, what it may change, how success is checked, and what evidence triggers a human review. The model is one component inside that design.

NemoClaw will be tested less by its launch benchmark than by the failure records companies build around it. Every truncated file, rejected tool call, policy block, and corrected answer can become a reusable test. Over time, that collection may be more valuable than the original agent prompt because it captures what the business learned the hard way. The companies that keep that learning in their own harness will own more than an AI agent. They will own the instructions for making it dependable.

The model is no longer the whole agent supporting image
Runtime policy creates a boundary outside the prompt for sensitive tools, systems, and data.
Sam C BarthAI sticks when the CRM underneath it is cleanI help teams get HubSpot, data, and handoffs in shape so new tools have something solid to run on.Visit samcbarth.com