> ## Documentation Index
> Fetch the complete documentation index at: https://docs.grantex.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Self hosting

# Self-Hosting Grantex

This guide covers running your own Grantex auth service — from a quick local spin-up to a
production-grade Kubernetes deployment.

***

## 1. Quick Start (Dev)

The fastest way to run the full stack locally:

```bash theme={null}
git clone https://github.com/mishrasanjeev/grantex.git
cd grantex
docker compose up --build
```

This starts PostgreSQL, Redis, and the auth service. Two developer accounts are seeded
automatically:

| Account | API key | Mode |
| - | - | - |
| Live | `dev-api-key-local` | Normal consent flow |
| Sandbox | `sandbox-api-key-local` | Auto-approves grants, returns `code` immediately |

Verify it's running:

```bash theme={null}
curl http://localhost:3001/health
# { "status": "ok" }

curl http://localhost:3001/.well-known/jwks.json
# { "keys": [{ "kty": "RSA", "alg": "RS256", ... }] }
```

> **Note:** The dev compose exposes database and Redis ports and uses hardcoded credentials.
> Never use it in production.

***

## 2. Generating a Production Signing Key

Grantex signs grant tokens with RS256 by default, or with ES256 when `JWT_SIGNING_ALG=ES256`.
Generate the private key once, in PKCS#8 form, and store it securely:

```bash theme={null}
# RS256 (default): RSA_PRIVATE_KEY
openssl genpkey -algorithm RSA -pkeyopt rsa_keygen_bits:2048 -out private.pem

# ES256: EC_PRIVATE_KEY
openssl genpkey -algorithm EC -pkeyopt ec_paramgen_curve:P-256 -out private-ec.pem
```

For use in environment variables or Kubernetes secrets, collapse it to a single line with
literal `\n` between each PEM line:

```bash theme={null}
awk 'NF {sub(/\r/, ""); printf "%s\\n", $0}' private.pem
```

Copy the output (starting with `-----BEGIN PRIVATE KEY-----\n...`) and use it as
`RSA_PRIVATE_KEY` (or `EC_PRIVATE_KEY` for the EC key).

Instead of supplying keys, `SIGNING_KEY_STORE=postgres` lets the service generate the key on
first start and store it in `platform_signing_keys`, encrypted with `VAULT_ENCRYPTION_KEY`
(see Section 7).

> Keep `private.pem` out of source control. The JWKS endpoint (`GET /.well-known/jwks.json`)
> exposes only the public key, so tokens remain verifiable after key rotation.

***

## 3. Production Docker Compose

### Prerequisites

* Docker 24+ with Compose v2
* A domain name with DNS pointing to your server
* TLS certificate (self-signed for testing; Let's Encrypt for production)

### Step 1 — Copy and fill in the env file

```bash theme={null}
cp .env.prod.example .env.prod
```

Edit `.env.prod` and replace every `change-me-*` placeholder with strong randomly generated
values. Set `RSA_PRIVATE_KEY` to the collapsed PEM from Section 2, and `JWT_ISSUER` to your
public base URL (e.g. `https://auth.example.com`).

### Step 2 — Provide TLS certificates

Place your certificate and private key at:

```
deploy/nginx/certs/server.crt
deploy/nginx/certs/server.key
```

**Self-signed (testing only):**

```bash theme={null}
mkdir -p deploy/nginx/certs
openssl req -x509 -nodes -newkey rsa:2048 -days 365 \
  -keyout deploy/nginx/certs/server.key \
  -out deploy/nginx/certs/server.crt \
  -subj "/CN=localhost"
```

**Let's Encrypt (production):**

```bash theme={null}
certbot certonly --standalone -d auth.example.com
cp /etc/letsencrypt/live/auth.example.com/fullchain.pem deploy/nginx/certs/server.crt
cp /etc/letsencrypt/live/auth.example.com/privkey.pem   deploy/nginx/certs/server.key
```

### Step 3 — Start the stack

```bash theme={null}
docker compose -f docker-compose.prod.yml --env-file .env.prod up -d
```

Verify:

```bash theme={null}
curl https://your-domain.example.com/health
# { "status": "ok" }
```

### Architecture

```
Internet → nginx (:443) → auth-service:3001
                ↓
           postgres + redis  (internal network only, ports not exposed)
```

***

## 4. Kubernetes / Helm

### Prerequisites

* Kubernetes 1.26+
* Helm 3.x
* A managed PostgreSQL instance (RDS, Cloud SQL, Neon, etc.)
* A managed Redis instance (ElastiCache, Upstash, etc.)
* An RSA private key (see Section 2)

### Install

```bash theme={null}
helm install grantex deploy/helm/grantex/ \
  --namespace grantex --create-namespace \
  --set externalDatabase.url="postgres://user:pass@host:5432/grantex" \
  --set externalRedis.url="redis://:pass@host:6379" \
  --set rsaPrivateKey="$(awk 'NF {sub(/\r/, ""); printf "%s\\n", $0}' private.pem)" \
  --set config.jwtIssuer="https://auth.example.com"
```

### Enable Ingress

```bash theme={null}
helm upgrade grantex deploy/helm/grantex/ \
  --reuse-values \
  --set ingress.enabled=true \
  --set ingress.className=nginx \
  --set "ingress.hosts[0].host=auth.example.com" \
  --set "ingress.hosts[0].paths[0].path=/" \
  --set "ingress.hosts[0].paths[0].pathType=Prefix" \
  --set "ingress.tls[0].secretName=grantex-tls" \
  --set "ingress.tls[0].hosts[0]=auth.example.com"
```

### Use an existing Secret

If you manage secrets externally (Vault, Sealed Secrets, External Secrets Operator):

```bash theme={null}
kubectl create secret generic grantex-secrets \
  --namespace grantex \
  --from-literal=RSA_PRIVATE_KEY="$(cat private.pem)"

helm install grantex deploy/helm/grantex/ \
  --namespace grantex \
  --set existingSecret=grantex-secrets \
  --set externalDatabase.url="..." \
  --set externalRedis.url="..."
```

### Upgrading

```bash theme={null}
docker build -t grantex/auth-service:0.2.0 ./apps/auth-service
docker push grantex/auth-service:0.2.0

helm upgrade grantex deploy/helm/grantex/ \
  --reuse-values \
  --set image.tag=0.2.0
```

### Rollback

```bash theme={null}
helm rollback grantex 1   # roll back to revision 1
```

***

## 5. Environment Variable Reference

This table is a quick-start subset, not an exhaustive schema. Consult `apps/auth-service/src/config.ts` and `.env.example` from the exact release you deploy for all feature-specific settings and validation rules.

| Variable | Required | Default | Description |
| - | - | - | - |
| `DATABASE_URL` | Yes | — | PostgreSQL connection string |
| `DATABASE_POOL_MAX` | No | `3` | Maximum PostgreSQL connections per auth-service replica (1-20); multiply by maximum replicas and leave headroom for other clients and rolling deploys |
| `REDIS_URL` | Yes | — | Redis connection string (include password if set) |
| `JWT_SIGNING_ALG` | No | `RS256` | Signing algorithm: `RS256` or `ES256` |
| `RSA_PRIVATE_KEY` | Yes\* | — | PKCS#8 PEM RSA private key (RS256). \*Required for `JWT_SIGNING_ALG=RS256` with the env key store, unless `AUTO_GENERATE_KEYS=true` (dev only). When `JWT_SIGNING_ALG=ES256`, a configured RSA key is published for verification only |
| `EC_PRIVATE_KEY` | Yes\* | — | PKCS#8 PEM EC P-256 private key (ES256). \*Required for `JWT_SIGNING_ALG=ES256` with the env key store. When `JWT_SIGNING_ALG=RS256`, a configured EC key is published for verification only |
| `JWT_VERIFICATION_PUBLIC_KEYS` | No | — | JWK Set (JSON) of public keys published for verification only: a key about to sign, or one that no longer signs; each key needs its thumbprint `kid` (as shown in the JWK Set) and `alg` |
| `JWT_LEGACY_KID_KEY` | No | the `RSA_PRIVATE_KEY` key | Thumbprint `kid` of the RSA key that signed tokens carrying a pre-0.6 `grantex-YYYY-MM` kid; set it when that key is no longer `RSA_PRIVATE_KEY` |
| `JWT_LEGACY_KID_MONTHS` | No | `13` | Months of `grantex-YYYY-MM` kid aliases published for the legacy key (current month and earlier); raise it if pre-0.6 grants live longer than a year; `0` publishes none. From the releases after `grantex` 0.6.0 and `@grantex/sdk` 0.7.0, SDK verifiers that set `boundedJwksFetch: true` / `bounded_jwks_fetch=True` (off by default until a later major release turns it on) refuse a JWK Set of more than 128 keys or 64 KiB, so keep these aliases and the other published keys within that |
| `SIGNING_KEY_STORE` | No | `env` | `env` (keys from the settings above) or `postgres` (stored encrypted; needs `VAULT_ENCRYPTION_KEY`) |
| `SIGNING_KEY_ACTIVATION_DELAY_SECONDS` | No | `900` | How long a new key is published before it signs (postgres rotations, and the switch from the legacy kid after start); at least 90 |
| `SIGNING_KEY_RETIRED_GRACE_SECONDS` | No | `2592000` | How long a retired stored key stays in the JWK Set; must cover your longest grant lifetime |
| `MAX_GRANT_LIFETIME_SECONDS` | No | — | Longest grant `expiresIn` accepted by authorization and delegation; with the postgres store, start-up refuses a grace shorter than this |
| `AGENT_KEY_ROTATION_OVERLAP_SECONDS` | No | `604800` | Default overlap of an agent key rotation: how long the replaced key stays usable (0 to 2592000); see `docs/providers/registering-agents.md` |
| `AGENT_KEY_HISTORY_MIRROR_ENABLED` | No | `false` | Mirror the key `POST` and `PATCH /v1/agents` write into the agent key history, and refuse there a key held in another agent's history, a compromised key, or a non-P-256 key under a payments rail. Off, those routes behave as before the history existed; only exactly `true` turns it on. See `spec/agent-keys.md` §7 |
| `SSO_STATE_SECRET` | No | derived | HMAC key for SSO state; derived from `RSA_PRIVATE_KEY`, `EC_PRIVATE_KEY` or `VAULT_ENCRYPTION_KEY` when unset, so every instance agrees |
| `AUTO_GENERATE_KEYS` | No | `false` | Auto-generate the signing key at startup (dev only — invalidated on restart) |
| `GRANT_TOKEN_LEGACY_CLAIMS` | No | `true` | Issue the pre-0.6 claim aliases (`agt`, `dev`, `grnt`, `scp`, `parentAgt`, `parentGrnt`, `delegationDepth`, `bdg`) next to the standard claims. Defaults to `false` in 0.7; see `docs/migration-0.6.md` |
| `JWT_ISSUER` | Yes | `https://grantex.dev` | `iss` claim in every JWT; your public base URL |
| `PORT` | No | `3001` | Port the auth service listens on |
| `HOST` | No | `0.0.0.0` | Bind address |
| `SEED_API_KEY` | No | — | Pre-seed a live developer API key (dev only — omit in prod) |
| `SEED_SANDBOX_KEY` | No | — | Pre-seed a sandbox API key (dev only — omit in prod) |
| `STRIPE_SECRET_KEY` | No | — | Enable Stripe billing integration |
| `STRIPE_WEBHOOK_SECRET` | No | — | Stripe webhook signature validation |
| `STRIPE_PRICE_PRO` | No | — | Stripe price ID for Pro tier |
| `STRIPE_PRICE_ENTERPRISE` | No | — | Stripe price ID for Enterprise tier |
| `MIGRATION_LOCK_TIMEOUT` | No | `2s` | How long a migration statement waits for a lock before the boot fails loudly (section 6) |
| `EVENT_BRIDGE_ENABLED` | No | `false` | Accept provider events (SSF/CAEP SETs, signed webhooks); see `docs/concepts/event-bridge-and-revocation.md` |
| `EVENT_BRIDGE_DEVELOPER_IDS` | No | — | Limit the event bridge to these developers (comma separated) |
| `EVENT_BRIDGE_RATE_LIMIT_PER_MINUTE` | No | `30000` | Event ingestion requests per client address (read per request) |
| `EVENT_BRIDGE_RECEIPT_RETENTION_HOURS` | No | `48` | Floor for how long delivery receipts are kept; never shorter than the source's own replay window (twice its tolerance) |
| `REVOCATION_FEED_ENABLED` | No | `true` | Serve the revocation feed and status endpoints SDKs use to see revocations (`docs/concepts/event-bridge-and-revocation.md`). On by default from the next release (it was `false`); `false` is the opt-out, and any other value leaves it on. SDK clients from the next release check revocation online by default and deny every call while it is off, so opt out only when every client sets `revocationCheck: 'offline'` |
| `REVOCATION_FEED_DEVELOPER_IDS` | No | — | Limit the feed to these developers (comma separated); every other developer's default SDK clients deny every call |
| `REVOCATION_FEED_POLL_MS` | No | `500` | How often an instance looks for new revocations when no notification arrives |
| `REVOCATION_FEED_SETTLE_SECONDS` | No | `15` | How long a feed entry may still be uncommitted; the cursor never advances past younger entries |
| `REVOCATION_FEED_HEARTBEAT_MS` | No | `1000` | How often a live stream confirms it is up to date; must stay well below a client's staleness bound |
| `REVOCATION_FEED_MAX_CONNECTIONS` | No | `200` | Revocation streams one developer may hold on one instance |
| `REVOCATION_FEED_RETENTION_HOURS` | No | `48` | How long delivered feed entries are kept after the credential expires |
| `REVOCATION_FEED_PRUNE_JITTER_SECONDS` | No | `300` | Upper bound, 0 to 3600, of the random delay before an instance first prunes the feed table; it then prunes hourly from that point. Only one instance prunes at a time (a Postgres advisory lock; the others skip that run), in batches of 1000 rows, at most 50 batches or 60 seconds a run, so a large backlog is removed over several hours rather than in one statement. `0` prunes at start |
| `EMERGENCY_STOP_ENABLED` | No | `false` | Serve the emergency stop and its lockout (section 11); revocations are irreversible, and issuance reads lockouts only while this is on |
| `REGISTRY_OPERATOR_API_KEYS` | No | — | Keys for the registry operator routes, `POST` and `PATCH /v1/registry/issuers` (comma separated, each at least 32 characters, separate from `ADMIN_API_KEY`); unset, those routes answer `503`, and a shorter key stops the service from starting. See `docs/issuers/becoming-an-accredited-issuer.md` |
| `REGISTRY_ATTESTATION_EDDSA_ENABLED` | No | `false` | `true` accepts EdDSA (Ed25519) signatures on attestations posted to `POST /v1/registry/attestations`, on issuers' withdrawal and refresh requests and on issuers' status lists; ES256 is always accepted. See `spec/attestation-1.0.md` |
| `REGISTRY_DEV_ISSUER_ORIGIN_MAP` | No | — | **Development and tests only.** Comma-separated `https-origin=loopback-origin` pairs, such as `https://mock-issuer.example=http://127.0.0.1:56901`: the registry fetches that issuer's status lists from the loopback server instead (targets must be `127.0.0.1`, `localhost` or `[::1]`). The service refuses to start with it set unless `NODE_ENV` is `development` or `test`, and never uses it in production. See `spec/attestation-1.0.md` section 10 |
| `PASSPORT_BOUND_GRANTS_ENABLED` | No | `false` | `true` (exactly) lets `POST /v1/authorize` take an Agent Passport in `passport`: the registry checks it (accredited issuer, signature with the recorded key, a registered and accepted attestation, the registry's and the issuer's status, the agent's proven key, the declared limits) before recording the request, and the grant token carries an `authorization_details` entry of type `urn:grantex:commerce:v1` and `cnf.jkt` for the passport's key. It also takes the merchants of a bound grant in `authorization_details`, and an RFC 8693 token exchange on `POST /v1/token` that issues a per-merchant child grant (section 8). Off, `passport` and `authorization_details` are ignored and a token exchange is answered as before. See `spec/passport-binding.md` |
| `REGISTRY_STATUS_RECONCILIATION_ENABLED` | No | `false` | `true` (exactly) runs status-list reconciliation: one instance at a time (a Postgres advisory lock) reads each accredited issuer's status lists at their `ttl`, one fetch per list, keeps the registry's acceptance entries in line (INVALID for a revoked passport, SUSPENDED while the passport or its issuer is suspended), and revokes, suspends or resumes the grants bound to those passports; `PATCH /v1/registry/issuers/{id}` cascades a suspension, reinstatement or revoked kid at once. Requires `DATABASE_POOL_MAX` of at least 2 (a run holds one connection for its advisory lock and works through another); with `true` and a pool of 1 the service refuses to start. Off, the per-attestation recheck worker runs as before and nothing is cascaded. See `spec/registry-federation.md` (Status reconciliation) and `docs/runbooks/status-list-incident.md` |
| `REGISTRY_STATUS_POLL_MIN_INTERVAL_MS` | No | `30000` | The minimum interval between two reads of one issuer status list by reconciliation; each list is otherwise read at its own `ttl`. At least `30000`, or at least `1000` when `NODE_ENV` is `development` or `test` (for the mock issuer and CI only); at most `86400000`. Any other value stops the service from starting |
| `RATE_LIMIT_ROUTE_CLASSES_ENABLED` | No | `true` | Revocation and emergency-stop routes, and the revocation feed, draw on per-developer budgets of their own instead of the plan (`docs/guides/rate-limits.mdx`); `false` puts them back in the plan budget, failing closed when Redis is unavailable |
| `TRUST_REGISTRY_ADMIN_LISTING_ENFORCED` | No | `false` | `true` makes `GET /v1/trust-registry`, which lists every developer's registry records, take `ADMIN_API_KEY` instead of a developer API key, answering `503` while that is unset; recommended, since the listing crosses tenants. Off (and any value other than exactly `true`) keeps the existing developer-API-key access. Read at startup |
| `RATE_LIMIT_ROUTE_CLASSES_ENABLED` | No | `true` | Revocation and emergency-stop routes, and the revocation feed, draw on per-developer budgets of their own instead of the plan (`docs/guides/rate-limits.mdx`); `false` puts them back in the plan budget, failing closed when Redis is unavailable. The revocation-status budget, 6,000 requests a minute per developer, is also the ceiling on SDK calls checked `online` (the default), one status request per `enforce()`; `GET /v1/revocations/status` allows the same 6,000 a minute per client address. With `false`, online checks share the plan budget (100 to 2,000 a minute), so opt out only if clients use `feed` or `offline` |
| `ED25519_PRIVATE_KEY` | Recommended in production | — | PKCS#8 PEM Ed25519 key. It signs DPDP consent proofs and Ed25519 credentials, and its public key is published in `/.well-known/jwks.json`. Unset, each process generates its own key at boot, so a proof signed before a restart or by another instance cannot be verified; with `NODE_ENV=production` the service logs a warning at startup (it does not refuse to start). Each DPDP consent proof says which key signed it (`keyPersistence`: `persistent` or `ephemeral`); `DPDP_REQUIRE_PERSISTENT_PROOF_KEY=true` refuses consent records while the key is ephemeral |
| `ED25519_STABLE_KID` | Recommended with `ED25519_PRIVATE_KEY` | `false` | `true` (exactly) gives a configured `ED25519_PRIVATE_KEY` the key id `grantex-ed25519-<RFC 7638 thumbprint>`, the same across restarts and months, and also publishes the key (in `/.well-known/jwks.json` and the DID document) under the `grantex-ed25519-YYYY-MM` ids of this month and the previous `JWT_LEGACY_KID_MONTHS` months, so proofs it signed before still verify. Off, the key id is the month the process started in, and a proof signed in an earlier month stops resolving once the service restarts in a later one. A generated (ephemeral) key keeps the month id either way |
| `DPDP_WITHDRAWAL_REVOKES_GRANT` | No | `false` | `true` (exactly) makes a DPDP consent withdrawal (`POST /v1/dpdp/consent-records/{id}/withdraw`) that omits `revokeGrant` revoke the record's grant (with `DPDP_REVOCATION_CASCADE=true` also the grants delegated from it), as if `revokeGrant: true` had been sent: after a withdrawal the Data Fiduciary must cease processing (DPDP Act s.6(6)). An explicit `revokeGrant: false` still leaves the grant alone. Off, an omitted `revokeGrant` means `false`, as before. Read at request time |
| `DPDP_ENFORCE_GRANT_PRINCIPAL` | No | `false` | `true` (exactly) refuses a DPDP consent record whose `dataPrincipalId` is not the grant's principal (`400 PRINCIPAL_MISMATCH`), the documented model. Off, the two are not compared, for integrations that key data principals differently. Read at request time |
| `DPDP_CONSENT_EXPIRY_ENABLED` | No | `false` | `true` (exactly) runs the DPDP consent expiry worker every five minutes: active consent records past `processingExpiresAt` become `expired`, each with a `grantex.dpdp.consent_expired` audit entry and a `dpdp.consent.expired` event. Rows are claimed with `FOR UPDATE SKIP LOCKED`, so several instances can run it. Read at startup |
| `DPDP_CONSENT_EXPIRY_REVOKES_GRANT` | No | `false` | With the expiry worker on, `true` (exactly) also revokes an expired record's grant, with `grant.revoked` (and its delegated grants with `DPDP_REVOCATION_CASCADE=true`). Off, the grant is left as it is. A run takes up to 100 developers with due records in `developer_id` order from a cursor that resumes after the previous run and wraps, so none is starved |
| `DPDP_REQUIRE_PERSISTENT_PROOF_KEY` | No | `false` | `true` (exactly) refuses DPDP consent record creation with `503 CONSENT_PROOF_KEY_NOT_PERSISTENT`, storing nothing, while the Ed25519 key is ephemeral (generated in-process because `ED25519_PRIVATE_KEY` is unset). Off, records are created and the proof carries `keyPersistence: 'ephemeral'`: such a proof is not verifiable on another instance or after a restart. Set `ED25519_PRIVATE_KEY` before turning this on. Read at request time |
| `DPDP_REVOCATION_CASCADE` | No | `false` | `true` (exactly) makes DPDP revocations (a withdrawal with `revokeGrant`, an erasure, and the expiry worker under `DPDP_CONSENT_EXPIRY_REVOKES_GRANT`) go through the grant cascade: the grants delegated from the record's grant, their credentials and wallet reservations are revoked too. Off, only the record's own grant is revoked (if still active), as before. Read at request time |
| `DPDP_ERASURE_EXPANDED` | No | `false` | `true` (exactly) makes a DPDP erasure (`POST /v1/dpdp/data-principals/{id}/erasure`) also replace the principal's grievance description and evidence with a fixed marker and delete stored exports about the principal. Off, both are kept, and the response's `retained` lists them with the reason. Read at request time |
| `DPDP_NOTICE_REQUIRE_RULE3` | No | `false` | `true` (exactly) refuses a DPDP consent notice (`POST /v1/dpdp/consent-notices`) that is missing an element of DPDP Rules 2025 r.3 (itemised personal data, specific purposes, the means to withdraw consent, to exercise rights and to complain to the Board) or whose `language` is not English or an Eighth Schedule language by ISO 639 code (`400 NOTICE_INCOMPLETE`). Off, the response's `validation` block only reports what is missing. See `docs/api-reference/dpdp/consent-notice-content.mdx`. Read at request time |
| `DPDP_REQUIRE_NOTICE_LANGUAGE` | No | `false` | `true` (exactly) refuses a DPDP consent record (`POST /v1/dpdp/consent-records`) without `consentNoticeLanguage` when the notice version it would bind exists in more than one language (`400 NOTICE_LANGUAGE_REQUIRED`). Off, such a request binds, as before, the pinned version (or the version of the newest notice row) through its newest row and records that row's language. With `consentNoticeLanguage` the flag makes no difference: the pinned version, or the newest notice row, in that language is bound. Read at request time |
| `DPDP_EXPORT_GDPR_REQUIRES_PRINCIPAL` | No | `false` | `true` (exactly) refuses a `gdpr-article-15` export (`POST /v1/dpdp/exports`) without `dataPrincipalId` (`400`): GDPR Art. 15 is one data subject's access right. Off, such an export is produced as before, without the per-person `article15` block, which is added whenever `dataPrincipalId` is given. Read at request time |
| `DPDP_BREACH_DEADLINE_ALERTS_ENABLED` | No | `false` | `true` (exactly) runs the DPDP breach deadline worker every five minutes: for each recorded breach whose detailed report to the Board is not recorded as sent, it emits `dpdp.breach.board_report_due` once when the 72-hour deadline (DPDP Rules 2025 r.7(2)(b)), or a granted extension, is within `DPDP_BREACH_ALERT_LEAD_MINUTES` (`stage: "approaching"`) and once when it has passed (`stage: "overdue"`), each with a `grantex.dpdp.breach_deadline_alerted` audit entry. Rows are claimed with `FOR UPDATE SKIP LOCKED`, so several instances can run it. Grantex files nothing with the Board. Read at startup |
| `DPDP_BREACH_ALERT_LEAD_MINUTES` | No | `720` | With the breach deadline worker on, how many minutes before the deadline the `approaching` alert goes out: an integer from 1 to 4320. Any other value stops the worker from starting. Read at startup |
| `REGISTRY_PUBLIC_ENDPOINTS_ENABLED` | No | `false` | Serve the registry's unauthenticated reads: the accredited issuer list, `GET /v1/registry/issuers` (paged with `page` and `pageSize`, 60 requests a minute per address), and the attestation-acceptance status lists, `GET /status/attestations/{list}`, `.../bitstring` and `.../bitstring/suspension` (`spec/registry-federation.md`), plus the minimised agent lookup (`GET /v1/registry/agents/{did}` and `GET /v1/registry/agents?...`) without an API key and the signed manifest `/.well-known/agent-registry.json`. On only for exactly `true`, read at startup. Off, the paths are not routes: they answer like any unknown path (`404` to an authenticated caller) and no CORS preflight is granted for them; the agent lookup still answers a request with a developer API key. Allocating and setting entries inside the service does not depend on it |

***

## 6. Database Migrations

Migrations run **automatically on every startup**, and each file is applied **once per database**.
The built-in runner (`src/db/migrate.ts`) reads all `*.sql` files from the `migrations/` directory in
alphabetical order, applies the ones this database has not seen, and records them in the
`schema_migrations` ledger (filename, checksum, applied-at). A start that has nothing to apply
touches no table at all.

That matters during a rolling deploy. Postgres takes an `ACCESS EXCLUSIVE` lock **before** it
evaluates `ADD COLUMN IF NOT EXISTS`, so a no-op `ALTER TABLE grants …` still queues behind
whatever transaction is touching `grants` — and every reader arriving after it waits behind that
queued request, including `/v1/authorize`, token exchange and delegation on the instance that is
still serving traffic. With the ledger a repeat start issues no DDL, so it cannot stall anything.

While applying, the runner sets `lock_timeout` (`MIGRATION_LOCK_TIMEOUT`, default 2 s) and retries
a few times, so a migration that cannot take its lock fails the boot loudly instead of stalling
the table. The setting is reset before the connection returns to the pool, so no application
statement inherits it. A file whose content changed after it was applied is reported as a warning
and **never re-applied** — ship a new migration instead. That is a warning and not a failure on
purpose: the edit has already had no effect on this database, and refusing to boot over it would
take the service down for nothing. The same applies to a ledger row whose file is no longer on
disk: it is warned about, because a renamed migration counts as a new pending file and its
statements run again.

### Adopting a database that is already at head

A database migrated by a release **before** the ledger existed has the full schema and no
`schema_migrations` table, so the first start after the upgrade treats all files as pending and
re-executes them. That is safe — every file is idempotent — and against ordinary traffic it takes
a few seconds. But if a single transaction is holding a row in `grants` for longer than
`MIGRATION_LOCK_TIMEOUT`, the `ALTER TABLE grants` files cannot take their lock and **the boot
fails**. Nothing is corrupted and no traffic is affected (migrations run before the server
listens, so the new instance never becomes ready and the old one keeps serving), but the deploy
is broken and has to be retried.

Baselining removes that risk. It records every file as applied **without executing any of them**,
so the upgrade's first start is a no-op like every start after it:

The command runs from the built service, so run it inside the image you are about to deploy
rather than from a source checkout (`dist/` does not exist until `npm run build`). It needs the
same `DATABASE_URL` as the service and nothing else:

```bash theme={null}
# Against the running container (Docker Compose)
docker compose -f docker-compose.prod.yml exec auth-service \
  node dist/cli/migrate-baseline.js --dry-run

# Or a one-off container on the release you are deploying
docker run --rm -e DATABASE_URL="$DATABASE_URL" ghcr.io/<org>/grantex-auth-service:<tag> \
  node dist/cli/migrate-baseline.js --dry-run

# Kubernetes
kubectl exec -n grantex deploy/grantex -- node dist/cli/migrate-baseline.js --dry-run

# From a built checkout
cd apps/auth-service && npm run build && node dist/cli/migrate-baseline.js --dry-run
```

`--dry-run` writes nothing at all — not even the ledger table — and prints the verdict:

```
this database is at head: all 183 tables and columns the migration files build are present
{"baselined":false,"dryRun":true,"atHead":true,"objectsChecked":183,...,"recorded":103}
```

Drop `--dry-run` to record. Then deploy: the new instance logs `applied 0` and takes no lock on
any table.

**It checks the precondition itself.** Before recording anything it reads every
`CREATE TABLE IF NOT EXISTS` and `ALTER TABLE … ADD COLUMN IF NOT EXISTS` out of the migration
files and confirms each object exists in the database. A database that is behind is refused, with
the missing objects named:

```
this database is NOT at head: 44 of 183 objects are missing (table evidence_records, …)
```

That matters because recording a file as applied means **no later start will ever apply it**. A
partly-migrated database that was baselined would run on an incomplete schema indefinitely, and
the only way back is editing `schema_migrations` by hand. `--dry-run` exits non-zero on such a
database, so it can be used as a pre-deploy check in a script.

Rules:

* Run it **only** against a database whose schema is already at head — and let the command
  confirm that rather than taking it on trust.
* It is not needed for a new database. Start the service and it applies everything itself.
* It is safe to repeat: files already in the ledger are left alone.
* If you skip it, the upgrade still works — retry the deploy at a quieter moment, or during a
  short maintenance window.

The `migrations/` directory of the release you are deploying is the only authoritative list of what will be applied; a number copied into this page goes stale on the next merge, so there is none here. The files cover core authorization, webhooks, policy, enterprise identity, credentials, budgets, offline operation, trust registry, DPDP, commerce, MCP certification-state integrity, query-performance indexes, agent prepaid wallets, layered wallet spend controls, the event bridge and the revocation feed. Index builds use `CREATE INDEX CONCURRENTLY`, and the runner serializes migrations across service instances with a PostgreSQL advisory lock. An index a cancelled concurrent build left `INVALID` is dropped before the file that creates it is retried, because `CREATE INDEX CONCURRENTLY IF NOT EXISTS` matches such an index by name and would otherwise skip it forever.

**Upgrade procedure** — just restart the service:

```bash theme={null}
# Docker Compose
docker compose -f docker-compose.prod.yml pull auth-service
docker compose -f docker-compose.prod.yml up -d auth-service

# Kubernetes
kubectl rollout restart deployment/grantex -n grantex
```

New migration files are applied automatically on startup. No manual SQL execution required.

***

## 7. Key Rotation

`GET /.well-known/jwks.json` publishes every platform signing key with `kid`, `alg` and
`use: "sig"`. Verifiers select the key by `kid` and refuse a key whose type does not match the
token's algorithm. No step below invalidates an outstanding token: a key leaves the JWK Set only
after the tokens it signed have expired.

### Key ids

A key's `kid` is its RFC 7638 thumbprint, `grantex-rs256-…` or `grantex-es256-…`, so every
instance publishes the same `kid` for the same key whenever it started.

Before 0.6 the RS256 `kid` was `grantex-YYYY-MM` of the month the process started. Tokens carrying
such a kid keep verifying:

* the auth service verifies an RS256 token whose `kid` is `grantex-YYYY-MM`, or that has no
  `kid`, with the *legacy key* — `RSA_PRIVATE_KEY`, or the key named by `JWT_LEGACY_KID_KEY`;
* the JWK Set also publishes the legacy key under `grantex-YYYY-MM` for the current month and the
  previous `JWT_LEGACY_KID_MONTHS - 1` months, so SDK verifiers find it;
* for `SIGNING_KEY_ACTIVATION_DELAY_SECONDS` after start, an instance still signs with the legacy
  kid, so resource servers holding a JWK Set fetched from a pre-0.6 instance keep accepting new
  tokens until they refresh it.

Upgrading needs no action. Do not remove the RSA key, and set `JWT_LEGACY_KID_KEY` if you replace
it, until pre-0.6 tokens have expired.

### Postgres key store

**Switching from the env store.** Set `SIGNING_KEY_STORE=postgres` and keep the existing key
settings for the first start. Every instance imports them: the env signing key becomes the stored
active key (same `kid`, so nothing changes for verifiers), and the other configured keys are stored
as retired public keys, keeping the legacy kid marker. Once the table holds them, the private key
settings can be removed.

**Rotation** is publish-then-sign:

```bash theme={null}
node dist/cli/rotate-signing-key.js            # new key for JWT_SIGNING_ALG
node dist/cli/rotate-signing-key.js --alg ES256 # switch algorithm
```

The command stores a new pending key and prints when it activates. Every instance publishes it
within a minute. After `SIGNING_KEY_ACTIVATION_DELAY_SECONDS` the next reload makes it the signing
key and retires the previous key, erasing its stored private key. The retired public key stays in
the JWK Set for `SIGNING_KEY_RETIRED_GRACE_SECONDS`, and the legacy kid key for the legacy alias
window. A second rotation is refused while one is pending. The stored active key is authoritative:
instances with a different `JWT_SIGNING_ALG` keep using it.

Set `MAX_GRANT_LIFETIME_SECONDS`; start-up refuses a grace window shorter than it, and without it a
warning says grants may outlive their key.

**Erasure limits.** Retiring a key sets its encrypted private key to `NULL`. The ciphertext can
remain in dead tuples until vacuum, in WAL and replicas, and in backups for their retention. It is
encrypted with `VAULT_ENCRYPTION_KEY` and bound to its `kid`, so it is useless without that key.
If a private key may have been exposed, rotate at once, and rotate `VAULT_ENCRYPTION_KEY` as part
of the response.

### Env key store

To replace a key (RSA to RSA, EC to EC, or a change of algorithm):

1. **Publish the new key.** Add its public JWK, with its thumbprint `kid` and `alg`, to
   `JWT_VERIFICATION_PUBLIC_KEYS` (for a change of algorithm you can instead set the other private
   key setting, for example `EC_PRIVATE_KEY` while `JWT_SIGNING_ALG=RS256`). Restart and wait at
   least `SIGNING_KEY_ACTIVATION_DELAY_SECONDS` so verifiers see it.
2. **Sign with it.** Set the new private key (and `JWT_SIGNING_ALG` if it changes). Keep the old key
   verifiable: add the old public JWK to `JWT_VERIFICATION_PUBLIC_KEYS` (the entry for the new key
   may stay; the same key listed twice is one key). If the old key is an RSA key that signed
   pre-0.6 tokens, set `JWT_LEGACY_KID_KEY` to its thumbprint `kid`. Restart.
3. **Clean up** only after every token the old key signed has expired: remove its entry from
   `JWT_VERIFICATION_PUBLIC_KEYS`, and unset `JWT_LEGACY_KID_KEY` once pre-0.6 tokens have expired.

```bash theme={null}
# Docker Compose
docker compose -f docker-compose.prod.yml up -d auth-service

# Kubernetes
kubectl rollout restart deployment/grantex -n grantex
```

***

## 8. Health Checks & Monitoring

### Health endpoint

```
GET /health
→ 200 { "status": "ok" }
```

Returns `200` when the service is up and connected. The Docker Compose healthcheck and
Kubernetes liveness/readiness probes both use this endpoint.

### Structured logging

All logs are emitted as JSON to stdout, compatible with Datadog, Loki, and CloudWatch Logs.
No configuration needed — just forward stdout from your container runtime.

### Prometheus metrics

When `METRICS_ENABLED=true` (the default), the auth service exposes Prometheus text at
`GET /metrics`. The endpoint is unauthenticated and limited to 10 requests per minute per IP;
restrict it at your network boundary if metrics must remain private.

***

## 9. Backup & Recovery

### PostgreSQL

Back up with `pg_dump`:

```bash theme={null}
docker compose -f docker-compose.prod.yml exec postgres \
  pg_dump -U "$POSTGRES_USER" grantex | gzip > "grantex-$(date +%Y%m%d).sql.gz"
```

Restore:

```bash theme={null}
gunzip < grantex-20260101.sql.gz | \
  docker compose -f docker-compose.prod.yml exec -T postgres \
  psql -U "$POSTGRES_USER" grantex
```

Schedule daily backups with cron or your cloud provider's managed snapshot feature.

### Redis

Redis holds ephemeral token metadata and rate-limiting state — not primary data. For
durability enable AOF persistence:

```
appendonly yes
appendfsync everysec
```

If Redis data is lost, in-flight auth requests will fail temporarily, but no permanent data
is lost. PostgreSQL is the source of truth for all grants, audit entries, and agent records.

While Redis is unreachable, standard API-key routes answer `503 RATE_LIMIT_UNAVAILABLE`
because their per-developer rate limit cannot be counted — once the Redis client gives up on
the command, which against a stopped Redis took more than a minute. Revoking a grant, token,
passport or consent bundle and the emergency stop are the exception: after at most 500 ms
they are counted in each instance's memory against the containment ceiling instead, and the
revocation is committed, so an incident can be contained during a Redis outage. The response
can still wait on the best-effort cache write that follows the commit.
`grantex_rate_limit_decisions_total{bucket="containment",outcome=~"local_.*"}` shows it
happening.

***

## 10. Production Readiness Checklist

Before going live, verify each item:

* [ ] `RSA_PRIVATE_KEY` is a real 2048-bit (minimum) RSA key — **not** `AUTO_GENERATE_KEYS=true`
* [ ] `POSTGRES_PASSWORD` and `REDIS_PASSWORD` are strong, randomly generated values (e.g. `openssl rand -hex 32`)
* [ ] `SEED_API_KEY` and `SEED_SANDBOX_KEY` are **not** set in production
* [ ] TLS is enabled end-to-end — nginx terminates HTTPS; internal services are on a private network with no exposed ports
* [ ] Database and Redis ports are **not** exposed to the public internet
* [ ] `JWT_ISSUER` matches your public base URL exactly — clients validate this claim during token verification
* [ ] Automated database backups are scheduled and have been tested with a restore
* [ ] Health checks are wired into your load balancer or uptime monitor
* [ ] CPU and memory limits are set to prevent runaway containers
* [ ] Log forwarding is configured (stdout → your observability stack)

## 11. Emergency Stop (Runbook)

One call halts every agent under a grant, an agent, a principal or a whole
developer. Use it when an agent is doing damage, a provider credential has
leaked, or a tenant must be stopped now and questions asked afterwards. With
`lockout: true` the same call also freezes issuance under that scope until the
freeze is lifted.

It is off unless `EMERGENCY_STOP_ENABLED=true`. Grants stopped this way are
**revoked, not paused**: there is no undo, and the principals involved have to
authorise again.

### A sweep, and a lockout only when you ask for one

**Without `lockout`, it is a sweep, not a lockout.** It revokes what exists —
repeatedly, until the scope comes back empty, so a grant delegated while it
runs is caught by a later sweep — and then it is finished. It does **not**
prevent new grants from being issued a second later. Anyone still holding the
developer's API key can call `POST /v1/authorize` and mint another one, and
`POST /v1/agents` to register another agent. The response says
`"lockout": false`.

**With `"lockout": true`, it also freezes issuance.** The stop records a freeze
over its scope *before* it sweeps, in the same transaction as its own record,
and until the freeze is lifted nothing is issued under that scope. These are
refused with `403 ISSUANCE_FROZEN` (`403 access_denied` on the OAuth
endpoints):

* `POST /v1/authorize`, and `POST /v1/token` for a code approved before the
  stop (the code is not consumed, so it works again once the freeze is lifted);
* `POST /v1/token/refresh` and `POST /v1/grants/delegate`;
* the OAuth profile's `POST /oauth/par` and `POST /oauth/token`
  (authorization code, refresh token and token exchange);
* `POST /v1/consent-bundles`, `POST /v1/consent-bundles/:id/refresh` and
  `POST /v1/passport/issue`.

A freeze covers what a stop over the same scope would revoke, and anything
new that would be issued under it. That means a new grant for a frozen agent
or principal. It also means a refresh, delegation, exchange or passport from
any grant with a frozen grant, agent or principal anywhere above it. An
`agent` or `principal` lockout does not stop the key registering a *new*
agent and asking for grants for it. A `developer` lockout covers every path
listed above for every agent and principal of the tenant, including agents
registered after it. Commerce passports and decision grants are outside any
lockout ("What a lockout does not refuse", below). The response says
`"lockout": true` and gives the `freezeId`.

So an incident that starts with a leaked credential:

1. **Stop with a lockout**, as the platform operator if you can
   (`POST /v1/admin/emergency-stop`). Only the operator can lift a lockout
   the operator placed. A lockout the tenant places with its own key can be
   lifted by any key of that tenant, *including the leaked one*.
2. **Rotate or disable the leaked credential**: `POST /v1/keys/rotate` for a
   developer API key, or remove the agent (`DELETE /v1/agents/:id`). The
   platform operator can also disable the developer.
3. **Lift the lockout** once the credential is safe ("Lifting a lockout",
   below).

Without a lockout the order is the other way round: rotate first, *then* stop.
Run a sweep-only stop first and it will be clean while the attacker mints a
fresh grant behind it. `status` tells you how the sweep ended:

* `completed`;
* `incomplete`: grants were still appearing after five sweeps. Something is
  still issuing them; add a lockout or go back to step 2;
* `failed`: a batch did not finish. The row records what was revoked before
  it stopped, and the call is safe to repeat.

A lockout recorded before a sweep fails **stays in place**. The grants the
sweep had not reached yet are still live, but nothing new is issued under
them, refresh and delegation included, until you repeat the stop or lift the
freeze.

The lockout is part of the emergency stop and is off with it: while
`EMERGENCY_STOP_ENABLED` is not `true`, no issuance path reads the freeze
state. Turning the flag off while a freeze is in force therefore stops
enforcing it. The freeze stays recorded and is enforced again when the flag
comes back. **Lift freezes before turning the stop off.** If the freeze state
cannot be read, issuance fails closed: every path answers
`503 FREEZE_STATE_UNAVAILABLE` (`503 temporarily_unavailable` on the OAuth
endpoints) and logs `alert: "issuance_freeze_unavailable"`.

### Before the incident

* Keep the revocation feed on (`REVOCATION_FEED_ENABLED` unset or `true`; it
  is on by default from the next release) and make sure the agents you need to
  stop check revocation: `revocationCheck: 'online'` (the default from the
  next SDK release) or `'feed'`. An agent checking neither (`'offline'`, or a
  client of `@grantex/sdk` 0.7.0 or `grantex` 0.6.0 and earlier, which do not
  check revocation at all) keeps working with the token it already holds until that
  token expires — the stop revokes the grant, but nothing tells that agent.
* Keep grant lifetimes short enough that the tokens of an agent you cannot
  reach expire in a time you can live with.
* Rehearse it: `scripts/revocation-release-test.sh` runs agents under a grant
  tree, stops them and measures how long each kept working. Every production
  release should have rehearsed it (PRD section 10).

### Stopping

Work out the blast radius first — same call, `dryRun: true`. Nothing is
revoked; the rehearsal itself **is** recorded, as a row with `dryRun: true`,
so `GET /v1/emergency-stops` shows who has been measuring the blast radius of
a tenant and when:

```bash theme={null}
curl -sS -X POST "$BASE_URL/v1/emergency-stop" \
  -H "Authorization: Bearer $DEVELOPER_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"scope": {"type": "agent", "id": "ag_01..."},
       "reason": "incident 4102: provider credentials leaked",
       "confirm": "stop agent:ag_01...",
       "dryRun": true}'
```

Then run it for real by dropping `dryRun`. `confirm` must be exactly
`stop <type>:<id>` — `stop agent:ag_01...` for the call above. Anything else
is refused with `412 CONFIRMATION_REQUIRED`. The expected phrase is **not**
echoed back: the point of the confirmation is that the caller knows what they
are stopping, which is lost if the endpoint hands them the answer to paste.

| `scope.type` | Stops |
| - | - |
| `grant` | That grant and everything delegated beneath it |
| `agent` | Every live grant of that agent, and their subtrees |
| `principal` | Every live grant that principal authorised, and their subtrees |
| `developer` | Every live grant of the developer |

To freeze issuance as well, add `"lockout": true` to the real call. It must be
a boolean; anything else is refused with `400` before anything is recorded. A
dry run never freezes anything, and says `lockout: false`.

The response names the stop (`stopId`), its `status`, how many sweeps it took,
how many grants matched and were revoked, which agents were stopped, and
`lockout`: `true` with the `freezeId` when the stop placed a lockout, or
reaffirmed one already in force over the same scope, and `false` otherwise.
`agentsStopped` lists at most 100 ids; when more were stopped,
`agentsStoppedTruncated` is true and `agentsStoppedTotal` gives the real
number.

**Suspended grants are swept up too.** A suspension is reversible; a stop is
not. Every grant the scope covers with status `active` *or* `suspended` is
revoked permanently, and the suspension bookkeeping that `POST
/v1/grants/:id/resume` needs goes with it. If a subtree is suspended pending
an investigation and you stop its scope, that investigation's subject cannot
be resumed afterwards — the principals must authorise again.

As the platform operator, use `POST /v1/admin/emergency-stop` with
`ADMIN_API_KEY` and the same body plus `developerId` (not needed for a
`developer` scope, where the scope names it). A developer API key can only ever
stop, or freeze, its own grants.

### What happens

With `lockout: true`, the freeze goes in first, in one transaction with the
stop's `emergency_stops` row and a `grantex.issuance_frozen` entry on the
developer's audit hash chain. From then on, nothing new is issued under the
scope. An issuance that had already passed its check when the freeze arrived
is waited for, and the sweep finds what it wrote: a new grant, which it
revokes, or a passport's credential, whose status bit it sets. The same holds
for the verifiable credential a code exchange or a delegation issues after its
grant is committed (`credentialFormat: "vc-jwt"` or `"both"`): it is written in
a transaction of its own that reads the grant and the freeze again under the
same lock. If a lockout committed in between, the call is refused with
`403 ISSUANCE_FROZEN` and no credential is written; the grant it had just
created is revoked by the stop's sweep. If the grant was revoked in between
without a lockout, the call returns the grant token without the credential, as
a failed best-effort issuance always has. Then, with or without a lockout:

1. Every matched grant and everything delegated beneath it is revoked in one
   transaction per batch, with wallet reservations released and credential
   revocation started. The scope is then read again and swept until it comes
   back empty, so a grant delegated mid-stop is caught.
2. One audit entry per grant (`grantex.grant.revoked`, cause
   `emergency_stop`) plus a summary entry (`grantex.emergency_stop`) go on the
   developer's audit hash chain, and a row goes into `emergency_stops`.
3. Each revocation reaches the revocation feed in the same transaction, so
   SDKs in feed mode deny the agents' next calls — measured in well under a
   second on a local stack, with two seconds as the requirement.
4. The auth service logs `alert: "emergency_stop"`, and
   `grantex_emergency_stops_total{scope,outcome}` and
   `grantex_grant_revocations_total{cause="emergency_stop"}` move. A lockout
   also moves `grantex_issuance_freeze_changes_total{action,scope}`. Every
   refusal it causes moves `grantex_issuance_refusals_total{path,reason}` and
   logs `alert: "issuance_frozen"`.

### Lifting a lockout

A freeze stays in force until someone lifts it. There is no expiry, and
repeating the stop does not lift it. Lift it once the credential it was
containing is rotated or disabled:

```bash theme={null}
curl -sS -X POST "$BASE_URL/v1/emergency-stop/unfreeze" \
  -H "Authorization: Bearer $DEVELOPER_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"scope": {"type": "agent", "id": "ag_01..."},
       "reason": "credential rotated; incident 4102 closed",
       "confirm": "unfreeze agent:ag_01..."}'
```

* `confirm` must be exactly `unfreeze <type>:<id>`. It is deliberately not
  the stop's phrase, so pasting the stop's confirmation lifts nothing. It is
  not echoed back either. A mismatch is `412 CONFIRMATION_REQUIRED`.
* `404 NOT_FROZEN`: no freeze is in force for exactly that scope. A freeze is
  lifted by the scope it was placed with; a `developer` freeze is not lifted
  by unfreezing one agent.
* `403 FREEZE_HELD_BY_OPERATOR`: the platform operator placed it, or
  reaffirmed it with a stop of its own, so only the operator can lift it. The
  operator uses `POST /v1/admin/emergency-stop/unfreeze` with `ADMIN_API_KEY`
  and the same body, plus `developerId` for a `grant`, `agent` or `principal`
  scope. The operator can lift any freeze.
* The response is the lifted freeze: `freezeId`, `scope`, the `stopId` that
  placed it, `placedBy`, `frozenAt`, `clearedAt`, `clearedBy` and
  `clearReason`. The row is kept, so the table is also the history of every
  lockout. A `grantex.issuance_unfrozen` entry goes on the audit chain in the
  same transaction.
* `GET /v1/emergency-stops` lists the freezes still in force under `freezes`,
  beside the stops, oldest first, 50 to a page. `freezesTotal` is how many are
  in force in all; when it is larger than the page, ask for the next one with
  `?page=2`, or for up to 200 at a time with `?pageSize=200`, as on the other
  paged lists. A stop's own `lockout` field says whether it asked for one, not
  whether that freeze is still in force.
* A second stop with a lockout over the same scope reaffirms the freeze in
  force rather than stacking another. If the operator does it, the freeze
  becomes the operator's.

### Afterwards

```bash theme={null}
curl -sS "$BASE_URL/v1/emergency-stops" -H "Authorization: Bearer $DEVELOPER_API_KEY"
```

* Check `status` in `GET /v1/emergency-stops`: anything other than
  `completed` means the sweep did not finish cleanly, and the row says how far
  it got.
* Check that agents stopped: `grantex_revocation_feed_entries_total` and the
  agents' own denial logs (`grant_revoked`).
* Any agent still running is one that is not watching the feed. Rotate or
  block its credentials, or wait out the token lifetime.
* Check `freezes` in the same response, and every page of it when
  `freezesTotal` is larger than the page. A lockout you placed is still
  refusing issuance until you lift it.
* To restore service, lift any lockout, and then the principals authorise
  again; the revoked grants cannot come back.
* Keep the `stopId`: the audit entries, the `emergency_stops` row, the freeze
  and the feed entries all carry it.

### What the stop cannot see

It revokes grants. Anything already handed out and cached elsewhere is
outside its reach:

* **Decision grants and passports already issued** keep verifying until they
  expire; they are signed artefacts. Whatever consumes them has to check
  revocation itself. For an agent passport (`POST /v1/passport/issue`) that
  check works: the sweep sets its status-list bit when it revokes the grant
  behind it. Decision grants and commerce passports are not touched by the
  sweep.
* **Work already in flight** — a tool call the agent has already made, a
  payment already authorised downstream — is not recalled. The stop denies
  the *next* call.
* **An agent that checks neither the feed nor the status endpoint** keeps
  using the token it holds until that token expires.
* **Anything below the depth your rehearsal covered.** The release test
  exercises three levels of delegation; deeper chains are handled by the same
  recursive query, but they are not measured.
* **Agents beyond the hundredth**: the response and the audit summary name at
  most 100. The summary also carries `agents_stopped_total`, so the true
  number is written down, but the list is not.
* **What a lockout does not refuse.** It refuses new issuance only; tokens
  already held keep working until the sweep revokes their grants. It does not
  stop the key registering new agents: under an `agent` or `principal`
  lockout, a new agent can still be given grants. It does not refuse decision
  grants, which are issued when an approver signs in and approves, not by the
  tenant's key. It does not refuse resuming a suspended grant that a sweep
  which did not finish left behind; repeat the stop, which revokes it. It does
  not refuse commerce passports (`POST /v1/commerce/passports/exchange`),
  which a commerce tenant's agent mints from a consent the shopper approved
  rather than from a grant, and the stop does not sweep them either. Contain
  those with the commerce tenant's own control: disable the tenant
  (`PATCH /v1/commerce/tenants/:tenant_id` with `{"status": "disabled"}`, as
  the platform operator or the tenant's owner), which refuses new ones with
  `403 tenant_disabled`, and revoke any already issued with
  `POST /v1/commerce/passports/revoke`.

### If the stop itself fails

* `403 FEATURE_DISABLED` / `404`: `EMERGENCY_STOP_ENABLED` is not `true` on
  the instance you reached.
* `412 CONFIRMATION_REQUIRED`: the `confirm` phrase does not match.
* `429`: the stop's own limits — 20 calls a minute from one address, and the
  developer's containment budget of 2,000 revocation calls a minute, which
  ordinary traffic does not use up. Wait out `Retry-After`; the stop is
  idempotent. A Redis outage does not refuse the stop.
* A 5xx: the stop is idempotent — run it again. Grants already revoked are
  left alone, and a partly finished stop finishes on the retry. A lockout the
  failed call had already placed is still in force; the retry reaffirms it
  rather than adding another.
* If the API cannot be reached at all, revoke at the database
  (`UPDATE grants SET status = 'revoked', revoked_at = NOW() WHERE …`): the
  feed triggers fire on that too, so agents still find out. The audit chain
  will not record it, so write it up. A row inserted straight into
  `issuance_freezes` is enforced as a lockout too. It bypasses the audit chain
  and the lock that keeps a concurrent issuance from slipping past a freeze,
  so repeat the stop with `lockout: true` as soon as the API is back.

## Ownership

Grantex is owned by Orchestrum Technologies LLP. Inventor and owner: Sanjeev Kumar. Ownership contact: [sanjeev@orchestrum.in](mailto:sanjeev@orchestrum.in) or [mishra.sanjeev@gmail.com](mailto:mishra.sanjeev@gmail.com).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.