

Deep Agents v0.7 Cuts Per-Turn Tokens by 65%
See how Deep Agents v0.7 reduces per-turn input tokens by 65% through a configurable harness, leaner tool descriptions, opt-in todos, and middleware.
Deep Agents v0.7 cuts base input tokens per turn by 65%. That means less fixed prompt overhead on every call, lower API spend, and more context room for the parts of the task that matter.
If I had to boil the update down, it’s this:
- The default base system prompt is gone
- Built-in tool descriptions are shorter
- Todos are no longer attached by default
- Middleware is picked on purpose, not bundled in
- The savings come from the harness, not the model
In plain English: if an agent makes 10 calls, the same fixed wrapper used to get sent 10 times. In v0.7, that wrapper is much smaller by default. So I get a leaner request on each turn, especially for simple tasks like reading or writing files.
A few points stand out:
- Cost scales with
llm_calls × input_tokens_per_call - The old setup sent planning, file-system, and sub-agent text even when a task didn’t need it
- The new setup shifts control to me, so I choose the prompt, tools, and middleware for the job
- The biggest cuts come from removing the base prompt and trimming tool text
- Long, multi-agent workflows gain the most because fixed overhead repeats on every step
Here’s the simple before-and-after view:
| Area | Before v0.7 | v0.7 |
|---|---|---|
| Base prompt | Sent every turn | Removed |
| Tool descriptions | Long | Shorter |
| Todos | On by default | Opt-in |
| Middleware | Bundled | Selected per task |
| Base input tokens | 100% | ~35% |
My takeaway: v0.7 is less about model changes and more about prompt discipline. If I keep the harness lean, I keep the full 65% gain. If I pile middleware and tools back in, I give part of that gain away.
That’s the core of the update, and it sets up the rest of the article.

What changed in the configurable harness
The 65% token cut came from changing what the harness injects on each turn, not from using a smarter model. Put simply, v0.7 strips out a lot of default text that used to ride along with every request.
Removed base system prompt and trimmed tool descriptions
Before v0.7, the harness packed a default system prompt, long tool descriptions, planning middleware, and sub-agent logic into every request, even when the task didn't call for any of it. v0.7 changes that setup by making the harness configurable, so developers decide what gets injected on each turn.
The biggest sources of overhead were the default base system prompt and the long built-in tool descriptions. The base prompt included instructions for planning tools, file system tools, and sub-agents, and it was sent on every turn. In v0.7, that prompt was removed. Developers can now supply prompt text that fits the task instead.
Built-in tool descriptions for utilities like ls, read_file, and write_file were shortened too. The tools behave the same way. There's just less fixed text wrapped around each request. Neither change affects the underlying model. They simply cut the token load that every turn used to carry.
Todos are opt-in and middleware is now explicitly selectable
Before v0.7, todoListMiddleware was attached by default, so planning text went out on every turn. In v0.7, todos are opt-in. That means you add them only when a task gets better with multi-step planning.
The same shift applies to the rest of the middleware stack. FilesystemMiddleware and SubAgentMiddleware are no longer bundled in by default. Developers can now put together only the middleware they need. A file-reading task can skip sub-agent logic. A verification workflow can add checklist middleware only when it helps.
How orchestration becomes more explicit
The practical change is simple: the setup moves from implicit defaults to explicit configuration. Instead of the harness deciding which tools are visible, which middleware runs, and what the system prompt says, developers make those calls themselves. They now control prompt assembly, tool visibility, and middleware on a per-task basis.
These changes show up in the default-agent stack below. [2]
| Feature | Pre-v0.7 | v0.7 |
|---|---|---|
| Base system prompt | Included by default | Removed |
| Tool descriptions | Verbose, built-in | Trimmed and configurable |
| Todo list middleware | Auto-attached every turn | Opt-in only |
| Middleware stack | Implicitly bundled | Explicitly composed |
Before-and-after token usage
Those harness changes show up right away in the payload sent on each turn. The savings come from trimming the fixed request envelope, not from changing the user prompt or model.
Default-agent turn before v0.7
Before v0.7, each turn included planning, file-system, and sub-agent scaffolding even when none of it was used. Todo text and middleware prompts were also included by default. That meant a big fixed payload repeated on every turn.
Default-agent turn after v0.7
After v0.7, a simple file-reading task sends only the tools and middleware it needs. So the extra overhead stops repeating across turns. The base system prompt is removed, tool descriptions are shorter, and a file-reading task no longer carries unused planning or sub-agent text.
As Aaron Jewitt notes, agent cost scales with llm_calls × input_tokens_per_call.[1]
Where the token savings come from
Here’s where the 65% reduction comes from.
| Harness Component | Before v0.7 | After v0.7 | Estimated Token Impact |
|---|---|---|---|
| Base system prompt | Sent every turn | Removed | High |
| Tool descriptions | Full built-in descriptions | Shortened | Moderate |
| Todo management | Bundled by default | Opt-in only | Depends on task |
| Middleware stack | Bundled by default | Explicitly composed per task | Depends on task |
| Total per-turn input | 100% (baseline) | ~35% | 65% reduction |
The biggest savings come from removing the base prompt and cutting down tool descriptions. That’s why selective middleware and shorter tool visibility make such a clear difference in day-to-day workflows.
The next section shows how developers keep these savings by choosing only the harness pieces a task needs.
Harness configuration patterns for developers
Once you've locked in the token savings, the next move is picking a lean harness profile for each task. Those savings come from smaller prompts and tighter harness defaults. In Deep Agents v0.7, the harness - not the model - accounts for most per-turn token overhead. That means optimization in Deep Agents v0.7 is less about model choice and more about how you put the harness together.
Choose only the middleware your task needs
Use middleware only when a task calls for gating, planning, or state management. On short tasks, those extra layers just add overhead.
The rule is simple: Start with the smallest middleware stack the task needs. The same goes for tools. Only expose the tools the current task actually needs.
Limit tool visibility and shorten descriptions
Showing every tool on every turn is one of the fastest ways to bloat per-turn input. Keeping tool visibility narrow helps keep the prompt small.
Shorter tool descriptions help trim prompt size even more. Tight, precise descriptions reduce prompt size and make it easier for the model to pick the right tool without a bunch of extra context.
Use opt-in todos and profile-based defaults
Todo handling helps on long-horizon tasks where the agent needs to track progress across many steps. For short, single-step work, todos add overhead and give you nothing back.
Make todos opt-in for long-horizon tasks, and set lean or rich defaults by agent class. Lean agent-class defaults preserve the 65% gain on simple agents, while richer profiles are kept for planning-heavy workflows.
Those choices decide whether the 65% cut shows up as lower cost and faster iteration in day-to-day use.
What the v0.7 update means for cost, speed, and scale
Lower inference costs and faster iteration loops
That smaller request envelope adds up on every turn in a long workflow. Token costs stack across multi-turn runs, so a leaner harness doesn't just save money on turn one. It keeps the starting point smaller on every turn after that, and the gap gets bigger as the conversation grows.
Here’s where that cut shows up in practice:
| Cost Factor | Impact of 65% Reduction | Business Value |
|---|---|---|
| Base Input | Lower starting point for every turn | Direct reduction in per-run spend |
| Context Accumulation | Slower growth of conversation history | Supports longer-running, more complex tasks |
| Payload Size | Smaller request payloads | Shorter iteration loops |
Smaller payloads also tighten iteration loops. If you're testing a prompt change or trying a new tool setup, lighter requests come back faster. That removes a lot of the drag from daily development work.
Better scalability for long-running and multi-agent workflows
Those same token savings matter even more when several agents share the same workflow budget. Fixed overhead is the quiet budget drain in multi-agent systems. If each agent carries an oversized harness, that overhead multiplies across every coordinated turn.
A leaner harness package keeps each agent’s footprint smaller. That improves throughput and cuts the odds of running into context caps or rate limits in the middle of a workflow.
As context windows fill up, performance can slip. A leaner harness gives each agent more usable context room for actual task data. In plain English: longer-running workflows can stay accurate longer, without needing compaction or summarization logic to step in too early.
Lower per-call overhead matters most in workflows with lots of turns, lots of agents, or both. The 65% token cut reduces cost per call. But the bigger scale win comes from the explicit orchestration model in v0.7. With explicit stopping criteria instead of open-ended loops, agents make fewer total calls to finish a task.
Key takeaways from the Deep Agents v0.7 release

Deep Agents v0.7 improves cost and quality by making the harness configurable. Middleware selection, tool visibility, and profile-based defaults now decide whether a workflow keeps the full savings or gives much of it back through extra overhead.
Configuration is the main optimization lever. And it matters most in workflows with many turns, many agents, or both.
FAQs
How do I keep the full 65% token savings in real-world workflows?
Treat your harness configuration like a living system, not a set-it-and-forget-it task. Start by measuring token usage and LLM call counts. Then use the harness to lock in tight behavior rules.
Use PreCompletionChecklistMiddleware to stop redundant reasoning loops. Use LocalContextMiddleware to pass in only the context and tools the model needs. Add prompt caching so system instructions stay stable across runs.
When token usage jumps, watch for drift and tighten the harness rules.
Which tasks should still use todos or extra middleware?
Use todoListMiddleware when the agent is working through a complex, multi-part problem. It helps the agent keep track of what’s finished and what still needs attention as the work moves forward.
When you prompt use of the write_todos tool, the agent can update progress as new details come in. That makes long, tough workflows easier to follow and helps the agent stay on track.
Does the smaller harness affect agent quality or reliability?
Not by default. A smaller, tuned harness can keep reliability the same - or even improve it - by cutting tokens per turn with better context handling, tool offloading, and structured prompt packaging.
Quality holds up when those savings are paired with deterministic guardrails, like verification loops and pre-completion checklists. That helps the agent stay locked on the task data that matters and double-check its work before it responds.
Choose the model you want in the model marketplace
Try chat, image and video models in the APIMart model marketplace, and experience model capabilities quickly with one unified API.
