Skip to main content

What caps do

A grant says which tools an agent may call. Caps say how often. They bound the damage of a looping agent, protect a paid provider account and make spend predictable. Two places declare caps. The manifest declares tenant-wide caps per tool, and the cost of a call in cost units:
The grant declares caps for that grant only, in its urn:grantex:tools:v1 authorization details: per tool, and a cost-unit budget for the connector under the reserved key cost_units:
Every declared cap is a separate counter, and a call must fit all of them. Manifest caps are shared by every grant in the tenant; grant caps count only that grant’s calls. A cap of 0 disables the tool.

Metering a call

Configure a meter on the client. enforce() then reserves the call’s units as its last step, after the token, scope, purpose and decision checks have passed, so a denied call never uses up a cap.
The tenant is the grant’s developer (dev claim) unless you pass caps_tenant_id / capsTenantId, which applies to every counter of that call. A call’s cost is the sum of the manifest cost_units for the components it incurs. Metering attaches to the call that incurs the cost, in the code that makes that call, and never to parsing its result afterwards.

Case and cost components come from the gateway

case_id and cost_components decide which counters a call is charged to, so the tool gateway sets them from its own context: the case being worked, and the provider request it is about to send. Never take them from the agent’s or model’s tool arguments, or an agent could name a fresh case to escape a per-case cap or claim a cheaper component. The same applies to wrap_tool(case_id=..., cost_components=...) / wrapTool({ caseId, costComponents }) and to enforceMiddleware({ extractCaseId, extractCostComponents }), which should read trusted request context. An empty cost_components list for a tool that declares cost units is denied (invalid_cost_component); omit the argument to charge every declared unit.

Check early, reserve once

An agent platform often checks a call twice, once when the plan is validated and again at the tool gateway. Only the second check should consume a unit:
reserve=False (reserve: false) compares current usage with the caps and returns cap_limits / capLimits and caps_tenant_id / capsTenantId. The check is point in time: another call can take the last unit before you reserve, so the reserving enforce() (or meter.reserve(result.caps_tenant_id, result.cap_limits)) is the decision that counts.

Rolling caps out: caps_mode

Set it on the client (Grantex(caps_mode="warn"), new Grantex({ capsMode: 'warn' })) or per call. Malformed grant caps are a token problem and are denied in every mode. Log would_deny_all / wouldDenyAll while in warn, review it, then switch to enforce. would_deny / wouldDeny is the first denial warn mode let through on the call. When several steps would deny the same call (for example a missing decision grant under decisions_mode="warn", then a missing amount, then an exhausted call cap), would_deny_all / wouldDenyAll lists every one of them, in the order the steps run: decision, amount cap, call caps, decision consumption. Read the full list: a later warning never replaces an earlier one in would_deny, so reading only that field hides it. The client’s separate enforce_mode="permissive" (development only) turns every denial into an allow, including cap_exceeded and meter_unavailable; the result keeps its reason_code, but nothing is reserved for such a call.

When a cap is exceeded

enforce() denies with reason_code cap_exceeded, sub_reason limit_reached and details that include error code E1008, the limit and the window:
Other cap_exceeded sub-reasons: Malformed grant caps deny as token_invalid / malformed_authorization_details.

Amount caps

A scope such as tool:merchant:write:*:capped:50 caps the amount of every call on that connector (the tightest capped:N on the connector wins), whatever permission the scope names: with tool:merchant:read:* and tool:merchant:write:*:capped:50, a read tool such as get_order is capped too. The call’s amount is passed as amount; the SDK does not read it from the tool’s arguments on its own. From version TypeScript 0.8.0 / Python 0.7.0 (breaking): a call with no amount under a capped:N scope is denied with cap_exceeded / amount_missing, and a malformed cap is denied whether or not an amount is given. In the current release such a call is allowed and the cap is never checked. Because the cap is connector-wide, this includes read-only tools with no monetary amount on a connector that carries a capped scope: give them an amount (an extractor may return 0). To keep the old behaviour while you add amounts, set caps_mode="warn" / capsMode: 'warn' and log would_deny_all / wouldDenyAll; it covers both amount_missing and a malformed cap on a call without an amount, including calls that also report a decision warning. From version TypeScript 0.8.0 / Python 0.7.0, the wrappers take an amount extractor, a function from the call to its amount:
Without an extractor, or when it returns None / undefined / null, a capped scope denies the call with amount_missing. An extractor that raises refuses the call before enforce() runs, whatever the grant says: wrap_tool raises PermissionError, wrapTool throws, and enforceMiddleware answers 403 with cap_exceeded / invalid_amount. A value that is not a finite number (a string, a boolean, NaN) is passed on and denied with invalid_amount. Read the amount the tool will actually spend, from the same arguments the tool receives, so the cap bounds what the call does.

Failed calls are not refunded

The reservation is made before the provider call and stays counted if the call then fails or times out. A timeout does not prove the provider did no work or did not bill for it, and refunding on error would let a flaky or hostile upstream reset the cap. Refund only when you know the request never left your process, for example when it failed validation locally or a connection could not be opened:

Backends

Both SDKs derive the same counter keys and run the same Lua script and SQL, so Python and TypeScript workers can share one Redis or Postgres. Fifty parallel calls against a cap of ten reserve exactly ten, on both backends, in CI. Redis (6.0 or later; the scripts use SET ... KEEPTTL). Keys look like grantex:caps:{<tenant hash>}:<counter hash>:z|s. Everything a reservation touches shares one hash tag, so it works on Redis Cluster. Time comes from the Redis server. Per-hour and per-day keys expire one window plus 60 seconds after the last reservation on the counter (each reservation resets the TTL); entries older than the window are dropped whenever the counter is used. Per-case keys never expire unless you set case_ttl_seconds / caseTtlSeconds, because an expired per-case counter would reset the cap. Run Redis with maxmemory-policy noeviction: an evicted counter forgets reservations. Postgres. Create the tables with grantex.caps.SCHEMA_SQL or CAPS_SCHEMA_SQL in your migrations, or call ensure_schema() / ensureSchema(). The Python backend takes a factory for DB-API connections with %s placeholders, for example pg8000. The TypeScript backend takes a pg Pool. Reservations run in READ COMMITTED transactions and time comes from the database. Rows carry a hash of the tenant id rather than the id, so row-level security policies keyed on tenant ids do not apply to these tables. Expired rows are removed when their counter is next used; run prune() periodically (for example hourly, from a scheduled job) to delete expired reservations and empty counters from all tenants. Per-case reservations are kept unless the backend was given case_ttl_seconds / caseTtlSeconds. prune() skips counters that are being reserved; if a race makes it fail, run it again.

No automatic failover

The meter does not fall back from Redis to Postgres when Redis fails. Two stores would hold two sets of counters, and calls spread across them could exceed every cap. Choose one backend per deployment. If it is unavailable, enforce() denies with meter_unavailable until it recovers.

Remaining budget

meter.usage(tenant_id, limits) returns what each counter holds and what remains. Build the limits for a tool with grantex.caps.build_cap_limits (buildCapLimits), or use result.cap_limits from enforce().

Not yet covered

This is the metering library. Showing caps and remaining budget on the consent page and in a per-case view, and a per-tenant caps.enforce rollout flag in the platform that drives caps_mode, are separate work.
Last modified on September 28, 2026