Analysis
October 4, 2026
8 min read
Maxime Dalessandro

Denial of wallet attacks on LLM agents: billed every turn

A denial of wallet attack no longer needs request volume. One malicious tool return retained in an agent's context is metered again on every later turn.

#denial of wallet#LLM agents#agent security#MCP#tool calling#AI cost#context management

TL;DR. A denial of wallet attack on a tool-calling agent no longer looks like traffic. A September 2026 study shows that when a runtime carries a tool return into later model inputs, the provider meters it again on every turn, so one admitted malicious payload becomes recurring, victim-billed processing with no credentials and no runtime compromise. Across 243 instrumented executions, the worst session's cumulative input reached 14,293 times its first call. The control point is before reingestion: transform history instead of deleting it, and bound growth, recursion, and spend at the host.

Denial of wallet is the attack class where the target is the invoice. OWASP's LLM Top 10 for 2025 files it under Unbounded Consumption: "by initiating a high volume of operations, attackers exploit the cost-per-use model of cloud-based AI services, leading to unsustainable financial burdens." That definition, and most of the vendor glossaries built on it, assumes volume. Someone outside sends many expensive requests, so the defenses are request-shaped: token caps, rate limits, per-user budgets. A paper published September 23, Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents (Zhang et al.), shows why that picture is out of date for agents. The expensive thing is no longer the request. It is what the runtime keeps.

The attack moved inside the session

A multi-step agent works by replay. The host runtime preserves state across turns by carrying earlier content, including tool returns, back into the model's input on every subsequent call. Providers meter input per call. So a tool return admitted at turn one is paid for at turn two, turn three, and every turn after that, at full price each time, until the session ends or the history is trimmed.

The paper names the retained content persistent billable state, and names the host's decision about whether and how that content re-enters later input the persistent billable-state boundary. Everything interesting happens at that boundary, after admission. A malicious or compromised tool does not need the victim's credentials and does not need to break out of any sandbox. It needs to get one large or self-amplifying payload admitted into history. The runtime then does the attacker's work, resubmitting the payload to the meter on every turn, billed to the victim.

The authors derive six attack vectors against that boundary, from a direct recursive baseline (V0) through polymorphic mutation, cross-tool chaining, stealth escalation, and weaponized output, up to adaptive threshold-hugging (V5), a vector designed to stay just under any fixed cap a defender sets. The names matter less than what they share: none of them involve request volume, and every one of them survives the glossary defenses, because a rate limit inspects arriving requests and this content already arrived.

What the measurements say

The study's harness, DOW-BENCH, ran end-to-end sessions across six model families, 243 core executions in all, with usage telemetry on every call. The headline number is the ceiling: the worst session's cumulative input processing reached 14,293 times the size of that session's complete first model input. One admitted payload, compounded by replay, outweighed the session's legitimate content by four orders of magnitude.

The more operationally useful number is the baseline overhead, because it applies to sessions nobody is attacking. Controlled reruns that varied only the history policy isolated what raw retention costs: carrying full raw history raised mean effective session cost by 21.2 to 35.9 percent, depending on provider. That slice of every agent bill is paying to re-read what the model already read.

Two small bar charts of billed input per turn across an eight-turn agent session. On the left, labeled raw retention, each bar stacks a thin new-input segment on top of a replayed-history block that grows every turn; the portion attributable to a single large tool return from turn two, shaded in emerald, recurs at full height in every bar after it. On the right, labeled transformed before reingestion, the same session shows flat, small bars: the tool return is metered at full size once, then carried forward as a compressed summary sliver.

The mechanism. The read happens once; under raw retention, paying for it recurs every turn. Transformation at the reingestion boundary keeps the task and drops the recurring bill.

The natural fix, dropping history, fails for a different reason: the agent needs its history. On history-dependent tasks, a deletion policy completed 2 of 12 for each provider tested, while deterministic compression completed 10 of 12 and 11 of 12. That asymmetry is the paper's practical core. The boundary decision has three options, keep raw, delete, or transform, and only transformation preserves both the task and the budget.

A database is the largest tool return an agent touches

The paper's demonstrations use generic tools, and it says nothing about SQL. But the database case is worth constructing, because for most production agents the database is where the big payloads live. Consider an agent wired to a Postgres MCP server, asked a routine analytical question. It writes a SELECT without a LIMIT, and the tool faithfully returns 50,000 rows into context. The query executed once, and the database did its job. Under raw retention, the agent's host now resubmits those rows with every subsequent turn of the session. A twenty-turn session pays for that one read twenty times. No attacker is required, only an unbounded result set and a runtime that keeps it, and with an attacker, a poisoned row set that instructs further wide reads is the cross-tool chaining vector with a connection string.

This recurring-spend shape should look familiar from the intrusion side. LLMjacking established that attackers already treat metered AI capacity as the asset worth stealing; denial of wallet is the same economics run in reverse, spending the victim's meter instead of using it. And the gateway layer, where most teams would reach for a control, has the same span problem we found in CData's Connect AI Gateway: a gateway rules on the tool call as it passes. Retention happens after the pass, inside the host's history, on every later turn. The paper's repository scan puts a number on how unguarded that layer is: of 3,830 MCP server and transport repositories scanned, 71 exposed any code-visible safeguard proxy at all, and none covered all four of the safeguard families the authors define.

The control point is before reingestion

The defense the paper lands on is host-side and runs at the boundary, before each call rather than after a bill arrives. Deterministic history transformation does the compression. Four invariants bound what transformation alone cannot: a token budget on prompt mass, growth-anomaly detection on context size, governance of recursive opportunity, and a cumulative cost circuit breaker. Replayed against the study's 123-evaluation attack corpus, the combined kernel contained every recurring attack.

The detail worth stealing for any agent platform is how the budget is enforced. A fixed spend cap, the obvious control, interrupted work: 13 of 24 workflow completions on the Mistral Small 4 test set. A progress-authorized policy, which spends only when the session is demonstrably advancing, reached 22 of 24 with no pre-completion interruptions. Budgets that measure progress beat budgets that measure only spend, because the attack's signature is cost without progress, and so is the signature of an agent stuck in a loop.

None of this is settled engineering. The paper is a v1 preprint, one research group, and its numbers are ceilings and harness measurements, so treat 14,293x as what the mechanism permits rather than what a typical session costs. But the direction matches what practitioners already found by accident: Spotify cut Claude Code token usage 90 percent by routing tool output out of context and handing it to a cheaper worker, which is this paper's transformation defense, discovered as a cost optimization before anyone framed it as a security control. The same move answers both the CFO and the attacker, which is rare enough to act on.

Where Datapace fits

The cheapest tool return to govern is the one that never enters context at full size, and for database agents that is a data-layer decision: what a query may return, at what grain, to which agent, under which budget. Datapace is building the context layer between your databases and your AI: resolved meaning validated by the people who own the data, the workload evidence beside it (cost, performance, usage and freshness, lineage), and a policy gate over what an agent may do and access, served over MCP. If agent spend on your databases is the question, start with cost optimization.

Sources

  1. Jinqian Zhang, Haojun Xia, Shujiang Wu, Jingkun Yue, Xia Zhang, Zhangpei Cheng, Bibo Tu, Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents, arXiv:2609.28585, September 23, 2026 (persistent billable state, six vectors, DOW-BENCH, 14,293x, 21.2 to 35.9 percent, compression versus deletion, four invariants, progress-authorized policy, MCP repository scan).
  2. OWASP, LLM10:2025 Unbounded Consumption, OWASP Top 10 for LLM Applications 2025 (denial of wallet definition and placement).

Frequently asked questions

What is a denial of wallet attack?
An attack that targets cost rather than availability. OWASP lists it under LLM10:2025 Unbounded Consumption: attackers exploit the per-use pricing of cloud AI services until the bill becomes unsustainable. In tool-calling agents, a single retained tool return can do it without any request volume.
What is persistent billable state?
Content an agent runtime carries from earlier turns into later model inputs. Providers meter input on every call, so a tool return admitted once is paid for again on each later turn. The term comes from a September 2026 paper treating retained content as a security object.
How is denial of wallet different from denial of service?
Denial of service aims to make a system unavailable. Denial of wallet leaves the service running and attacks the invoice: every operation succeeds, and each one costs the victim money. In agents the two diverge further, because the costly sessions are the ones that keep working.
How do you defend an LLM agent against denial of wallet attacks?
Govern what re-enters context before each call, with budgets on prompt size, context growth, recursion, and cumulative spend. Transforming history (compressing tool returns) preserved most tasks in the paper's tests, while deleting history broke them, so transform rather than delete.
Can a database query cause a denial of wallet?
Yes, mechanically. A SELECT without a LIMIT that returns tens of thousands of rows into an agent's context becomes retained input, re-metered on every later turn of the session. The database read runs once; paying for it recurs until the session ends or the history is transformed.

Keep reading

Ready to let agents touch production, safely?

Bring a use case. We will show you what agents can do on your live data, inside your guardrails.