Skip to main content

Why an event bridge

A grant is issued against the facts at the time: the business was active, the case was open, the principal was employed. Facts change after issuance. A provider that monitors a business learns it was dissolved; an identity provider learns a session was revoked. The event bridge lets those systems tell the auth service, so the grants that relied on the old facts stop working. The bridge has three rules:
  1. Only verified events count. A delivery that does not verify is refused with 401, counted, and never acted on.
  2. An event is acted on at most once. Every event id is claimed in a replay store before anything happens.
  3. Nothing is inferred. A verified event that no mapping rule matches is logged, counted and ignored. It never widens or narrows a grant by default.
The bridge is off unless EVENT_BRIDGE_ENABLED=true. With it off every event bridge route answers 404 (ingestion) or 403 FEATURE_DISABLED (registration), and nothing else in the auth service changes.

Registering a source

Sources belong to one developer and are managed with the developer API key.

SSF/CAEP transmitter

A transmitter pushes Security Event Tokens (RFC 8417) to the receiver, as in RFC 8935 push delivery. Register its issuer and its public keys, inline or by URL:
The response carries ingestUrl (/v1/event-bridge/ssf/<id>). audience defaults to that URL; set it to what the transmitter puts in aud if it differs. jwksUri is fetched through the same outbound guard as webhook URLs (no private hosts in production), cached for five minutes and refetched at most every 30 seconds when a token names an unknown kid. Inline jwks must contain public keys only. A SET is accepted only when all of these hold, and each failure has its own reason code: The subject is read from the SSF sub_id claim, or from an event’s subject member for transmitters on earlier CAEP drafts.

Generic signed webhook

For systems that do not speak SSF, register a webhook source:
The response includes secret once. It is stored encrypted with VAULT_ENCRYPTION_KEY, bound to the source id. The sender signs the timestamp and the exact body bytes:
This is the scheme the auth service uses for its own outbound webhooks (X-Grantex-Signature-V2), so one signer serves both directions. The body is JSON:
The signature is checked before the body is parsed; then the timestamp must be within toleranceSeconds of the receiver’s clock, in either direction (timestamp_out_of_window). Several sha256= values may be sent, comma separated, while the sender changes secrets. Rotation. POST /v1/event-sources/<id>/rotate-secret returns a new secret. The previous one keeps verifying for previousSecretTtlSeconds (default one day, at most seven). Pass 0 for a leaked secret so it stops immediately.

Mapping an event to an action

A verified event does nothing until a rule says what it means. A rule is data:
Paths start at type, subject or data and address members and array elements: subject.business_ref, data.filing.parties.0. Inherited JavaScript members are not event data and never match.

Targets

The path may hold one identifier or a list of up to 50. Anything else — a number, an object, an absent path — records target_invalid: the rule acts on nothing. Every resolution is scoped to the rule’s developer. A rule that names a grant, principal or agent of another developer resolves to nothing and records no_target; no grant of theirs is ever touched, and a rule may only name an event source of its own developer.

Binding a subject to a grant

Provider events talk about businesses and cases, not grant ids. Bind the identifiers a grant was issued for:
kind is the developer’s own vocabulary (lower case, up to 64 characters); values are opaque, up to 50 per grant. Bindings on a parent grant are enough: actions cascade to everything delegated beneath it.

What the actions do

revoke — cascade revocation

The grant and every grant delegated beneath it, to any depth are revoked in one transaction: their status changes, wallet reservations are released, credentials issued for them are revoked, the developer’s audit hash chain gets one entry per grant, and a grant.revoked event is emitted. The Redis revocation key is written after the commit as an accelerator; the database stays authoritative, so a cache outage cannot undo a committed revocation. Revocation is irreversible and reaches suspended grants too, so a suspended subtree can never be resumed under a revoked ancestor. Cascade revocation takes the same per-developer lock as delegation, so a child being delegated while its parent is revoked either loses the race (the parent is gone when it commits) or is included in the cascade. No active grant is ever left under a revoked one.

suspend — reversible revocation

The grant and its subtree move to suspended. Every authorisation check requires status = 'active', so a suspended grant authorises nothing, and its tokens stop verifying. Undo it with:
This restores exactly the grants that suspension suspended. It is refused with 409 ANCESTOR_INACTIVE while any grant above the root is revoked or suspended, and it keeps working when the event bridge is turned off, so a suspension can always be undone. A grant is suspended once, under the first root that reached it. So suspending an ancestor of an already-suspended subtree reports zero affected: everything below it is already suspended, and each grant keeps the root it was suspended under, so resuming that original root still restores exactly what it suspended. Nothing is lost — the second call simply has nothing left to do — but do not read “0 affected” as “the suspension did not work”.

re_evaluate — hand the decision back

Nothing about the grant changes. One audit entry per grant is written and a grant.re_evaluation_requested event is emitted to the developer’s webhooks and event stream, carrying the grant ids, the event id and type, the rule id and the event’s subject. The relying platform decides what to do. and a bounded copy of the event’s subject (scalar members, short values, and subject_truncated when anything was dropped — the subject is provider-supplied). Revoking and suspending are idempotent, so a retried delivery costs nothing. Asking the platform to look again is not, so it is claimed per source, event and rule and happens at most once however often the delivery is retried.

Audit records

Every action writes an entry to the developer’s audit hash chain: trigger uses the evidence-package vocabulary (api, event, admin, cascade), so an evidence export names why a grant stopped. The grantex. prefix and the grantex:platform marker are reserved: POST /v1/audit/log refuses them, so a tenant cannot forge a revocation record. These entries are written whatever the plan’s audit limit — a security record a full plan could suppress would be worthless.

Delivery outcomes

The receipt for each delivery records what the rules did, and the response carries the same status:

Replay protection

Each verified delivery claims (source, event id) — the SET jti or the webhook id — before it is processed: A replay outside the webhook window, or of a SET older than maxAgeSeconds, is refused before it reaches the replay store. Receipts are pruned hourly, but only once they are older than the window in which their own source would still accept the delivery, and never sooner than EVENT_BRIDGE_RECEIPT_RETENTION_HOURS. Removing one earlier would make an old delivery acceptable again. That window is twice the tolerance, measured from when the delivery arrived. A webhook’s timestamp check is two-sided, so a delivery may arrive timestamped up to toleranceSeconds in the future and stays acceptable until received_at + 2 × toleranceSeconds. A SET is bounded by maxAgeSeconds plus twice the 60 s clock skew, because its iat may also be ahead of the receiver’s clock. Raising a source’s tolerance widens this for future deliveries only. Receipts already pruned under the old, narrower window cannot come back, so for the length of the new window there are old deliveries that would verify again and have no receipt to refuse them. If you raise a tolerance materially, rotate the source’s secret at the same time: that invalidates every old signature and closes the gap immediately.

Responses

  • 202 {"status": "unmapped" | "applied" | "observed" | "duplicate"}
  • 401 {"err": "unverifiable", "code": "EVENT_UNVERIFIABLE"} for every verification failure: an unknown source, a disabled one, a source whose developer is outside EVENT_BRIDGE_DEVELOPER_IDS, a bad signature, a stale timestamp, a reused event id. One code for all of them, on purpose: a sender that could tell them apart could probe for source ids and for how far a guess had got. The precise reason is in the structured log (alert: event_bridge_verification_failure) and in the reason label of grantex_event_bridge_verification_failures_total. The response body is what is indistinguishable, not the work behind it. A real source id proceeds to secret decryption and HMAC comparison, so it takes measurably longer than an unknown one — around 1.4 ms in the reviewer’s measurement, over 400 samples each. For an SSF source with a jwksUri the difference can be far larger, because a real id can trigger an outbound key fetch. Treat the endpoint as resistant to reading which ids exist, not to a patient attacker timing it
  • 415 for the wrong media type
  • 404 when the bridge is off entirely
  • 5xx when processing failed; the receipt is left failed so the sender’s retry is processed again

How an agent finds out: the revocation feed

enforce() verifies a grant token offline against the issuer’s JWK Set. That is what makes it fast and what makes revocation invisible to it: a revoked grant’s token stays cryptographically valid until it expires. Revoking a grant stops the auth service issuing anything new; it does not, on its own, stop an SDK that already holds a token. The revocation feed closes that gap. Three modes, chosen per client, and tightened (never loosened) per call: From the next SDK release a client created without revocationCheck / revocation_check checks online, and denies with grant_revoked / status_unavailable when the auth service cannot answer, including when a deployment has turned the feed endpoints off. To keep the previous behaviour, pass offline explicitly:
online costs one round trip to the auth service per enforce() call, and those calls draw on the developer’s revocation-status budget: 6,000 a minute (100 a second) across all of the developer’s instances, whatever the plan (see Rate limits). The status endpoint’s per-address limit is the same 6,000 a minute, so a server running many tools behind one egress address can use the whole budget; past it the SDK is answered 429, retries, and then denies with status_unavailable. A client making more checked calls than that, or one on a hot path, should use feed, which costs one connection per process however many calls it makes.

Per-call overrides only tighten

The modes are ordered by how soon a revocation is seen: offline never, feed within its staleness bound, online on the next call (offline < feed < online; exported as REVOCATION_CHECK_STRENGTH). The per-call revocationCheck / revocation_check option of enforce() may choose the client’s mode or a stricter one. A weaker one is refused before anything is checked: enforce() rejects (TypeScript) or raises ValueError (Python), so code that reaches enforce() cannot switch off what the deployment configured. A client configured offline can still ask for feed or online on a sensitive call.
A denial carries grant_revoked with a sub-reason: revoked, suspended, parent_revoked, feed_stale, feed_unavailable or status_unavailable.

Failing closed

The last three matter most. An agent whose feed has gone quiet does not know what has been revoked, so it stops authorising calls:
  • the feed records when it last heard from the auth service;
  • the stream sends a heartbeat every second, and only while the server has read the database successfully;
  • if the client has heard nothing for longer than its staleness bound (default 5 seconds), enforce() denies with feed_stale;
  • if the deployment does not serve the feed, or the feed is not ready, every call is denied with feed_unavailable;
  • in online mode, a check that cannot be completed denies with status_unavailable, and a grant the auth service does not recognise is refused rather than assumed live.
This is the opposite of a cache: it is a claim about freshness that expires.

The endpoints

A client starting cold reads the cursor and then pages the snapshot — every grant currently revoked or suspended and not yet expired, and every individually revoked token — then streams from the cursor. Because the cursor is read first, nothing that happens while it pages can fall between the two. Entries are {seq, action, grantId, jti, expiresAt, at} with action one of revoked, suspended, resumed or token_revoked. They are a set of identifiers, so a duplicate delivery changes nothing. The cursor never advances past an entry that could still be overtaken. A transaction that took its sequence number before another but committed after it would otherwise be skipped, so entries younger than the settle window (REVOCATION_FEED_SETTLE_SECONDS, default 15 s) are delivered but do not move the cursor. That costs a few repeated entries and removes a way to miss one. And never past the page it returned. A page is bounded by limit (1000 by default) while the settled maximum is not; a cursor taken from the larger number would skip everything in between while the client believed itself up to date. The cursor a response carries is therefore the lowest of the two — which matters exactly when there is a lot to deliver: a large cascade, an emergency stop, a sweep. Delivered entries are kept for REVOCATION_FEED_RETENTION_HOURS past the expiry of the credential they are about, and an hourly worker prunes the rest, so the table the snapshot reads does not grow without bound. The triggers fill the table even while the feed is off, so the first prune after it turns on can face a large backlog: the worker deletes in batches of 1000 rows, at most 50 batches or 60 seconds a run, and leaves the rest to the next run. Only one instance prunes at a time (it takes a Postgres advisory lock and the others skip that run), and each instance’s first run waits a random delay of up to REVOCATION_FEED_PRUNE_JITTER_SECONDS (default 300), so the instances of one deploy do not start together.

Where the entries come from

Database triggers on grants and grant_tokens, not from each revocation path. Every way a grant stops — DELETE /v1/grants/:id, a cascade from a provider event, an emergency stop, a consent withdrawal, an anomaly, a DPDP erasure, OAuth revocation, and the hard delete behind DELETE /v1/agents/:id — writes a feed entry in the same transaction as the revocation itself. Four triggers cover it: two on status changes and two on deletion, since a deleted row can appear in no snapshot. pg_notify wakes the receivers on commit; each instance also polls (REVOCATION_FEED_POLL_MS, default 500 ms), so a lost notification costs latency and never correctness. If those triggers are missing (the migration could not take the lock at startup), the feed endpoints answer 503 FEED_UNAVAILABLE rather than an empty feed, and clients fail closed.

Measured

scripts/revocation-release-test.sh starts the auth service against real Postgres and Redis, builds a delegation tree, revokes each parent and measures how long the child kept being authorised, through both SDKs. G-6 requires two seconds at the ninety-fifth percentile. Two things make that number mean something. The TypeScript measurement runs in a plain Node process loading the build from the checkout, and the Python one refuses to start unless grantex was imported from the checkout — an ambient install would otherwise “prove” the criterion against code nobody reviewed. And the clock starts when the revocation is committed (when the API call returns), not when the call was made: the revoke call is rate limited, and the SDK waiting out a Retry-After is not propagation. That wait is reported separately as revoke_call_max_ms — and it has a budget of its own (10 s, REVOCATION_REVOKE_CALL_BUDGET_MS), because a release that prints a minute-long wait on the containment path and passes anyway is not telling you the truth. Revoking used to share the plan’s rate-limit bucket with ordinary traffic (FINDINGS G-23); it now has a containment bucket of its own (see Rate limits below), so the wait no longer depends on the tenant’s plan or its other traffic. Any figure quoted from a run is environment-specific. The numbers depend on the machine, the container runtime, whether Postgres and Redis are local, and what else is running: an independent reviewer measured p95 100–506 ms and max 515–673 ms where this checkout’s machine measured tens of milliseconds. What the release test asserts is the requirement — p95 within two seconds, no failed trial — not a particular number.

The emergency stop

Cascade revocation is the documented emergency stop for the whole platform. One authenticated call halts every agent under a grant, an agent, a principal or a whole developer:
  • confirm must be exactly stop <type>:<id>; anything else is refused with 412 CONFIRMATION_REQUIRED, and the expected phrase is not echoed back. Nothing is revoked before that check passes.
  • dryRun: true reports how many grants the scope covers and revokes nothing.
  • A developer API key can only stop its own grants. The platform operator uses POST /v1/admin/emergency-stop with ADMIN_API_KEY and a developerId.
  • GET /v1/emergency-stops lists what has been stopped, when, by whom and why, and the lockouts still in force, a page at a time (page, pageSize) with freezesTotal giving how many there are in all.
  • Underneath it is an ordinary cascade revocation per matched grant, so the stop appears in the audit hash chain (one grantex.grant.revoked per grant plus one grantex.emergency_stop summary) and on the revocation feed, and agents following the feed are denied within seconds.
  • Off unless EMERGENCY_STOP_ENABLED=true. The revocations are irreversible: principals have to authorise again.
  • A sweep, and a lockout only when asked for. Without lockout, it revokes what exists, re-reading the scope until it comes back empty so a grant delegated mid-stop is caught, and then it is done: the same API key can mint a new grant immediately afterwards, and the response says "lockout": false. status is completed, incomplete (grants kept appearing) or failed (a batch did not finish — the record says what was revoked, and the call can be repeated).
  • lockout: true freezes issuance under the scope until the freeze is lifted. The freeze is recorded before the first sweep, in the same transaction as the stop’s record. Every issuance path then refuses with 403 ISSUANCE_FROZEN (access_denied on the OAuth endpoints): authorize, code exchange, refresh, delegation, the OAuth profile’s pushed request and token endpoint, consent bundles and agent passports (POST /v1/passport/issue). Commerce passports and decision grants are not issued from grants and are outside any lockout; the runbook says how to contain commerce passports. A freeze covers what a stop over the same scope would revoke, so a refresh, delegation or passport is checked against every grant above it. If the freeze state cannot be read, issuance fails closed with 503 FREEZE_STATE_UNAVAILABLE. The response says "lockout": true with a freezeId, and grantex.issuance_frozen goes on the audit chain.
  • POST /v1/emergency-stop/unfreeze lifts a lockout; confirm must be exactly unfreeze <type>:<id>. The operator lifts any lockout with POST /v1/admin/emergency-stop/unfreeze. A lockout the operator placed can only be lifted by the operator, because the tenant’s key may be the leaked credential. Lifting it writes grantex.issuance_unfrozen on the audit chain.
The runbook — rehearsing it, working out the blast radius, what to do when an agent keeps running, and what to do if the API itself is unreachable — is section 11 of docs/self-hosting.md.

Rate limits

Revoking and stopping are rate limited in a containment bucket of their own, and the feed in a status bucket of its own, rather than in the developer’s plan budget. A tenant that has spent its plan quota on ordinary calls can still revoke, and its SDKs can still learn about revocations. The per-address limits on each route still apply first (20 a minute for the emergency stop; 600, 6,000 and 120 for the feed, status and stream routes). The status route’s per-address limit is the developer’s status budget, because an SDK checking online calls it once per enforce(): a lower one would refuse a server running many tools behind one address before the developer reached its budget. The feed and stream are called once per SDK process (a long poll or a reconnect), so their per-address limits scale with processes, not calls, and stay lower. A revocation fails open when the limiter cannot count it because it is written to Postgres, which is authoritative, and refusing it would prolong an incident over an outage of a cache. RATE_LIMIT_ROUTE_CLASSES_ENABLED=false puts these routes back in the plan budget. See docs/guides/rate-limits.mdx.

Observability

Every refused delivery also logs alert: "event_bridge_verification_failure" with the source id and reason, never the payload or signature. Alert rules are in deploy/prometheus/event-bridge-alerts.yml.

Settings

What this does not defend against

  • A transmitter whose signing key is stolen can send events that verify. Scope what its events can do with mapping rules, and disable the source (PATCH /v1/event-sources/<id> with {"status": "disabled"}) if its key leaks.
  • Events are only as timely as the sender. The bridge bounds how old an accepted event may be, not how late the sender is.
  • A disabled source’s events are refused, not queued: re-enable it and have the sender retransmit.
  • A rule is only as good as the bindings behind it. A grant with no subject_ref binding is not reached by a rule that targets one; the delivery records no_target and says so in the metrics.
  • Revocation stops new authorisation decisions. An SDK holding a verified token still needs to learn about it; see the revocation feed for how quickly, and what happens when it cannot.
  • The feed tells an SDK what the auth service knows. An agent that opts out (revocationCheck: 'offline', and any client of @grantex/sdk 0.7.0 or grantex 0.6.0 and earlier, which do not check revocation at all) keeps calling until its token expires, which is why short grant lifetimes still matter.
  • A revocation is bounded by the client’s staleness bound, not by zero: an agent can make calls in the window between the revocation and the entry arriving. Lower staleAfterMs and the poll interval to narrow it; the measured propagation on a local stack is well under 100 ms.
Last modified on September 27, 2026