The two limits
A sliding window, keyed on your API key for authenticated routes — so a shared egress address does not collapse tenants into one bucket. 60/min by default; enterprise keys carry their own rate. The one unauthenticated route, /demo/filter, is keyed on client IP instead, at a fixed 10/min.
The plan allowance, counted on POST /filter and /filter/batch, reset at the start of each UTC calendar month. Enterprise is unlimited and publishes no monthly policy.
What every response carries
Structured fields from draft-ietf-httpapi-ratelimit-headers-09, alongside the earlier draft spelling and the pre-draft X- spelling. Three generations of client parse three different things, and the cost of emitting all of them is a handful of bytes.
| RateLimit-Policy "requests-per-minute";q=60;w=60, "requests-per-month";q=5000;w=2592000 | The static contract. One member per policy: q is the quota, w the window in seconds. Present on every response. |
| RateLimit "requests-per-minute";r=59;t=41 | The live reading. r is what is left, t is seconds until that policy resets. Present only when the request actually consumed the window. |
| RateLimit-Limit 60 | The earlier draft spelling of the same quota, for clients that parse it. |
| RateLimit-Remaining 59 | The earlier draft spelling of r. |
| RateLimit-Reset 41 | The earlier draft spelling of t — delta-seconds, not a timestamp. |
| X-RateLimit-* X-RateLimit-Reset: 1788426730 | The pre-draft spelling, kept for compatibility. Note that X-RateLimit-Reset is a Unix epoch, not a delta. |
| Retry-After 41 | Delta-seconds. Sent on 429 only. When both are present it names the same instant as the reset, and it takes precedence. |
See it
$ curl -sS -D- -o/dev/null https://api.shadow-warden-ai.com/health RateLimit-Policy: "requests-per-minute";q=60;w=60 RateLimit-Limit: 60 X-RateLimit-Limit: 60
A route that consumes no quota — such as /health — publishes the policy and no counter. There is no bucket reading that belongs to that request, and inventing one would be a number with nothing behind it.
How to pace against it
- Read RateLimit on each response and slow down as r approaches zero, rather than waiting for a 429.
- On 429, honour Retry-After — it takes precedence over the reset parameter. Do not retry sooner.
- A monthly 429 carries a long Retry-After and an upgrade URL in the body. Retrying will not clear it.
- Both limits are per API key on authenticated routes. Running two workers on one key halves each worker's share.
The same conventions are described in the OpenAPI document so a generated client sees them too.