Plan-aware authenticated throughput is implemented in current repository source. A managed deployment must be verified separately; source completion is not evidence that a hosted rollout has occurred.
Default Limits
A route-specific Fastify
config.rateLimit replaces the 5,000 requests/minute Fastify default for that route; those two Fastify policies are not stacked. The Redis-backed standard developer plan budget remains an additional policy when standard API-key authentication applies.
Containment and Revocation-Status Budgets
Stopping an agent must not wait out a quota that ordinary traffic used up, and an SDK checking for revocations must not be starved by the tenant’s other calls. These routes therefore draw on budgets of their own instead of the plan budget:
Calls to these routes do not reduce the plan budget, and the plan budget being exhausted does not refuse them. Each route’s Fastify per-IP policy still applies first: 20/min for
POST /v1/emergency-stop and POST /v1/consent-bundles/:id/revoke, 600/min for GET /v1/revocations, 6,000/min for GET /v1/revocations/status, 120/min for GET /v1/revocations/stream, 100/min for a consent bundle’s revocation-status, and the 5,000/min default for the others. Resuming a suspended grant, POST /v1/tokens/verify and POST /v1/grants/verify stay in the plan budget. So do DPDP consent withdrawal (POST /v1/dpdp/consent-records/:recordId/withdraw, even with revokeGrant: true) and erasure (POST /v1/dpdp/data-principals/:principalId/erasure), which can mark grants revoked: they are compliance operations, not the incident path. They revoke only the grants their records name, without the cascade to delegated grants that DELETE /v1/grants/:id performs, and erasure also rewrites the principal’s audit entries, so they keep the plan budget and answer 503 while the limiter is unavailable. To contain an incident, revoke the grant or use the emergency stop.
GET /v1/revocations/status is called once per enforce() by an SDK client checking revocation online, which is the SDK default. Its per-IP limit is therefore the same 6,000/min as the developer’s revocation-status budget, so a server running many tools behind one address is held to the developer’s budget rather than refused below it; the per-IP limit remains as an abuse ceiling and counts unauthenticated calls too. That budget is 100 checked calls a second across all of a developer’s instances. A client that needs more, or checks on a hot path, should use revocationCheck: 'feed' (revocation_check="feed"), which reads GET /v1/revocations and /stream once per process rather than once per call; those routes keep their lower per-IP limits because they are not called per enforce().
The operator route POST /v1/admin/emergency-stop uses the admin key, not a developer key, so no per-developer budget applies to it; its 20/min per-IP policy does.
Commerce, the SCIM Bearer data-plane routes under /scim/v2/*, admin, and other custom-auth routes remain outside the standard developer plan bucket. Their active Fastify policy still applies. The standard API-key-authenticated /v1/scim/tokens management routes do consume the plan bucket.
Response Headers
Protected responses include rate-limit headers so your application can track the policy represented by that response. Successful standard developer API-key responses report the per-developer budget the route draws on — the plan budget, or the containment or revocation-status budget; a request rejected earlier by the active Fastify default or route policy reports that policy instead. Treat the active Fastify policy and Redis per-developer policy as separate applicable controls, and always honor a429 plus Retry-After.
429 Error Response
When you exceed a rate limit, the API returns a429 Too Many Requests status with the following body:
Retry-After header tells you the minimum number of seconds to wait. Generic IP/route-limit responses are produced by @fastify/rate-limit and can use a different error code/message; clients should branch on HTTP status 429 and honor the header rather than matching message text.
Authenticated Limiter Availability
The standard developer API-key plan and revocation-status limiters fail closed when their Redis transaction cannot be completed:503 Service Unavailable, not 429. Retry it as a transient service failure; do not treat it as proof that the request reached the protected handler.
Containment routes fail open instead: a revocation is written to Postgres, which is authoritative, so an outage of the limiter’s cache does not refuse it. When the Redis counter fails, or does not answer within 500 ms, each auth-service instance counts containment calls in its own memory against the same 2,000/min ceiling, so the ceiling holds per instance until Redis returns. The response carries the usual X-RateLimit-* headers from that count. Each instance holds at most 10,000 of these counters; past that, the least recently used counter is dropped, so the call is still served and counted and memory stays bounded.
Reading Rate Limits from SDKs
All three SDKs automatically parse rate limit headers from every response. You can read them viaclient.lastRateLimit (TypeScript/Python) or client.LastRateLimit() (Go).
After a Successful Call
Handling 429 Errors
When a429 is returned, the error object includes rate limit info with the retryAfter value:
Retry Strategy
Use exponential backoff with jitter to avoid thundering-herd problems when multiple clients hit the limit simultaneously.Best Practices
The JWKS endpoint (
/.well-known/jwks.json) is exempt from rate limits. Local
JWKS-based verification avoids the POST /v1/tokens/verify rate limit when a
revocation lookup is not required, but current standalone helpers still make a
JWKS network request per call.- Cache tokens — Grant tokens are valid JWTs. Store and reuse them until they expire instead of requesting new ones per operation.
- Use local verification deliberately —
verifyGrantToken()validates with JWKS and avoids the online verification rate limit, but the standalone helper fetches JWKS on each call. - Use webhooks instead of polling — Subscribe to webhook events like
grant.createdandgrant.revokedrather than polling grant or audit endpoints. - Honor
Retry-After— When you receive a429, always use theRetry-Afterheader value as your minimum wait time. - Spread requests — If your system makes burst requests (e.g., batch token exchanges), add short delays between calls.
Self-Hosted Deployments
If you’re running the Grantex auth service yourself, rate limits are configurable. The default Fastify IP policy is set inapps/auth-service/src/server.ts; a route-specific Fastify policy lives beside its route and replaces that default on the route. Standard developer API-key budgets, the containment and revocation-status budgets, and their window are defined in apps/auth-service/src/plugins/dynamicRateLimit.ts and remain additional; a route joins the containment or revocation-status budget with config.rateLimitClass. All instances must share Redis for standard developer budgets to remain global across the deployment. Fastify IP counters are process-local unless you configure a shared store or equivalent ingress enforcement.
RATE_LIMIT_ROUTE_CLASSES_ENABLED=false (default true) puts the containment and revocation-status routes back in the plan budget, failing closed when Redis is unavailable, as releases before these budgets did. grantex_rate_limit_decisions_total{bucket, outcome} counts every per-developer decision: bucket is plan, containment or status; outcome is allowed, limited, unavailable (refused with 503), or local_allowed / local_limited for a containment call counted in-process while Redis was unreachable. A rising local_* count means the limiter’s Redis is down.
See the Self-Hosting guide for deployment instructions.