This is the full developer documentation for Limitry
# Introduction
> Limitry decides, at request time, whether a user, workspace, API key or agent may do something at a given cost.
Limitry is the runtime that decides whether a request is allowed. When your product is about to do something that costs — call a model, run a job, spend credits, open another connection — it asks Limitry one question: **may this subject do this, at this cost?** Limitry answers in milliseconds, and keeps everything behind the answer: counters and resets, rolling windows, reservations, shared credit pools, budgets, concurrency and entitlements.
* **Subjects** are whoever acts in your product: a user, a workspace, a service, an API key — or an AI agent.
* **A check** is the question; the answer is *allowed* or *not allowed*, with what remains and why.
* **Built for agents as well as people.** Every operation is available through the [API](/docs/quickstart/), the [command line](/docs/cli/) and the [MCP server](/docs/mcp/), so an agent can set Limitry up in your codebase and ask before doing anything that needs a person ([Approvals](/docs/approvals/)).
Start with [limits](/docs/limits/) — the rules — and [checks](/docs/checks/) — the question your code asks.
# Use with AI agents
> Let an AI agent work with your workspace — through the MCP server, the command line, or the API — and teach it how with the agent skill.
AI agents can do in your workspace what you could do yourself — within the access you give them, as **you**, and recorded in the audit log. Three ways, from the least setup:
| Your agent | Use |
| ------------------------------------------ | ------------------------------------------------- |
| Claude, ChatGPT or another MCP-capable app | the [MCP server](/docs/mcp/) — nothing to install |
| A coding agent with a terminal | the [command line](/docs/cli/), with `--json` |
| Your own software | the [API](/docs/quickstart/) or [SDK](/docs/sdk/) |
## MCP server
[Section titled “MCP server”](#mcp-server)
Add `https://mcp.limitry.com/mcp` as a remote MCP server in your agent. You sign in once in the browser, choose **one workspace** and what the agent may do; its tools are exactly that. Details and safeguards: [MCP server](/docs/mcp/).
## Command line
[Section titled “Command line”](#command-line)
Coding agents (Claude Code, Cursor, Codex …) work well with the [command line](/docs/cli/):
```bash
npm install -g @limitry/cli
limitry login # you approve it in the browser
limitry workspace list --json
limitry workspace webhook-endpoints list --workspace acme --json
```
`--json` gives the agent the API’s exact JSON, errors as JSON on stderr, and a 0/1 exit code. In CI or a sandbox without a browser, set `LIMITRY_TOKEN` to a [personal access token](/docs/authentication/) limited to what the agent needs.
## Teach your agent: the skill
[Section titled “Teach your agent: the skill”](#teach-your-agent-the-skill)
The agent skill is one file that teaches an agent everything above — how to sign in, choose the workspace, every command and tool, the error codes, and what to leave to a person: [`/docs/skill/SKILL.md`](https://limitry.com/docs/skill/SKILL.md). It is generated from the API, so it always matches.
For Claude Code, install it for every project:
```bash
mkdir -p ~/.claude/skills/limitry
curl -fsSL https://limitry.com/docs/skill/SKILL.md \
-o ~/.claude/skills/limitry/SKILL.md
```
(or into a project’s `.claude/skills/limitry/`). Other skills-aware agents take the same file in their own skills folder.
## Docs for agents
[Section titled “Docs for agents”](#docs-for-agents)
The documentation is available as plain text for agents: [`/docs/llms.txt`](/docs/llms.txt) (an index of every page) and [`/docs/llms-full.txt`](/docs/llms-full.txt) (everything in one file). The site’s [`/llms.txt`](https://limitry.com/llms.txt) indexes these, the API, the MCP server and the skill.
## What agents cannot do
[Section titled “What agents cannot do”](#what-agents-cannot-do)
Some operations stay with people even when an agent has full access — anything that sends your workspace’s data somewhere new or hands out access (for example, creating API keys or adding webhook endpoints). An agent asks instead, and you approve or deny — see [Approvals](/docs/approvals/); a secret it creates never reaches the agent.
# Approvals
> When an agent needs something it may not do alone — an API key, a webhook — it asks, and a person decides.
Some actions are never an agent’s to take alone, because they hand out access or send your data somewhere new: creating an API key, adding, removing or re-enabling a webhook endpoint. An agent that needs one **asks**, and a person who could do it themselves decides.
## How it works
[Section titled “How it works”](#how-it-works)
1. **The agent asks** — with the action, its details and a reason (“I’m connecting acme-web and need an API key that can read the workspace”). Everyone who could approve it gets a notification.
2. **A person decides** at the link — the notification, or the one the agent shows you. You see exactly what will happen and why. **Approve** or **Deny**.
3. **It happens as you** — with your permissions, recorded in the audit log as yours, noting the agent that asked.
Requests expire after 24 hours, and each can be decided only once.
## Secrets stay out of the agent’s hands
[Section titled “Secrets stay out of the agent’s hands”](#secrets-stay-out-of-the-agents-hands)
When the action creates a secret (an API key, a webhook signing secret), the agent never sees it:
* **Asked from the command line**: after you approve, the command line writes the secret straight into a file in your project — it is never printed, so an agent driving the terminal cannot read it:
```bash
limitry workspace approval-requests create --action workspace.apiKey.create \
--input '{"name":"acme-web","scopes":["workspace:read"]}' \
--reason "Connect acme-web" --json
limitry workspace approval-requests redeem --wait --write-env .env
# Done: … Wrote LIMITRY_API_KEY to .env (the value is not shown).
```
* **Asked by an agent (MCP)**: the action runs when you approve, and the secret is shown to **you**, once, on the approval page.
## For developers
[Section titled “For developers”](#for-developers)
`GET /v1/workspace/approval-actions` lists what can be asked (each with its input schema and who approves it); `POST /v1/workspace/approval-requests` asks; `GET /v1/workspace/approval-requests/{id}` reports the decision. Asking needs the `workspace.approvals:write` scope and a person’s credential — a personal access token or an app a person connected, not a workspace API key.
# Authentication
> Authenticate public API requests with Workspace API keys.
The public API authenticates requests with **Workspace API keys** using the standard Bearer scheme:
```http
Authorization: Bearer YOUR_API_KEY
```
```bash
curl \
-H "Authorization: Bearer YOUR_API_KEY" \
https://api.limitry.com/v1/workspace
```
A successful response identifies the workspace the key belongs to:
```json
{ "id": "…", "name": "Example Workspace" }
```
## API keys belong to a workspace
[Section titled “API keys belong to a workspace”](#api-keys-belong-to-a-workspace)
Keys authenticate a **workspace**, not a person. A key keeps working even if the person who created it leaves the workspace, and it never carries a user session — it cannot be used to sign in.
## Creating a key
[Section titled “Creating a key”](#creating-a-key)
Workspace owners can create keys in the application under **Settings → API Keys**:
1. Choose a descriptive name (for example `production-backend`).
2. Choose what the key may do (see [Scopes](#scopes)) — **Read only** by default.
3. Copy the secret when it is shown — **it is displayed exactly once** and cannot be viewed again. Store it in your secret manager.
4. The key list shows only safe metadata (name, first characters, access, creation and last-used times).
If a secret is lost, create a replacement key and revoke the old one.
## Revoking and rotating
[Section titled “Revoking and rotating”](#revoking-and-rotating)
Owners can revoke a key at any time from **Settings → API Keys**. Revocation is immediate and permanent: the same secret will fail authentication on the next request.
To rotate a key without downtime:
1. Create a replacement key.
2. Update your integration to use it.
3. Revoke the old key.
## Scopes
[Section titled “Scopes”](#scopes)
Every key and token is granted **scopes** when it is created, and can do only what it was granted. A scope names a group of endpoints and whether it may read or change them — for example `workspace.audit:read` (read the audit trail), `workspace.members:read` (read members and pending invitations) or `workspace:read` (read the workspace profile). Each endpoint in the [API reference](https://api.limitry.com/v1/docs) states the scope it needs.
When you create a credential you choose one of:
* **Read only** (the default) — every read scope available at creation. A scope added to the API later is not granted automatically.
* **Full access** — everything, including scopes added later.
* **Custom** — exactly the scopes you tick.
Grant only what an integration needs. A request for an endpoint the credential was not granted is a `403` with code `forbidden`:
```json
{
"error": {
"code": "forbidden",
"message": "This credential is not granted the workspace.audit:read scope",
"requestId": "…"
}
}
```
To change what an integration may do, create a new credential with the access it needs and revoke the old one.
## Personal access tokens
[Section titled “Personal access tokens”](#personal-access-tokens)
A second credential type exists for acting as **yourself** rather than as a workspace — for scripts, CI jobs and the command line: personal access tokens (`pat_…`), created in the application under **Account → API tokens** — or by the command line itself: `login` opens your browser, you approve it, and the token for that computer is created and stored for you. A personal access token:
* acts as you, in the workspaces you choose at creation (all of yours, or specific ones), with the [scopes](#scopes) you grant it;
* is subject to your workspace role as well: `GET /v1/workspace/audit-events` is owner-only, like the audit page in the application (a workspace API key is the workspace and has no role);
* expires on the schedule you pick (30 days, 90 days, a year, or never);
* answers the identity check on `/v1/user/me`:
```bash
curl \
-H "Authorization: Bearer YOUR_PERSONAL_ACCESS_TOKEN" \
https://api.limitry.com/v1/user/me
```
```json
{
"id": "…",
"name": "Casey Customer",
"email": "casey@example.com",
"workspaces": [
{ "id": "…", "name": "Example", "slug": "example", "role": "owner" }
]
}
```
On every other endpoint a personal access token must say which workspace it acts in, with the `X-Workspace` header (the workspace slug); the call then runs with your membership in that workspace and the token’s scopes:
```bash
curl \
-H "Authorization: Bearer YOUR_PERSONAL_ACCESS_TOKEN" \
-H "X-Workspace: example" \
https://api.limitry.com/v1/workspace
```
A missing `X-Workspace` is a `400`; a workspace you are not a member of, or one the token was not scoped to, is a `403` — as is a scope the token was not granted. Workspace API keys ignore the header: a key is its workspace. A workspace API key is never accepted on `/v1/user/*`. Like API keys, tokens are shown once at creation, revocable from the same page, and stop working immediately if your account is suspended or the token expires.
## Apps and AI agents (OAuth)
[Section titled “Apps and AI agents (OAuth)”](#apps-and-ai-agents-oauth)
Apps — including AI agents’ tool connections — can act for a person without ever seeing a key: the product is an OAuth 2.1 authorization server. The app sends the person to sign in, they choose a workspace and approve what the app may do, and the app receives a short-lived access token for the API.
* Discovery: `https://api.limitry.com/.well-known/oauth-authorization-server/auth` lists every endpoint. Apps may register themselves (dynamic client registration); a registration grants nothing until a person approves.
* An app asks for the same [scopes](#scopes) (`workspace.audit:read`, …), plus `offline_access` for a refresh token. The person can narrow them.
* A token acts for one person in one workspace — no `X-Workspace` header — with that person’s role. Request it for the resource `https://api.limitry.com/v1`.
* The person can disconnect the app at any time under **Account → Connected apps**; its tokens stop working on the next request.
## Authentication failures
[Section titled “Authentication failures”](#authentication-failures)
Requests with a missing, malformed, invalid, or revoked credential receive a `401` with the standard [error envelope](/docs/errors/):
```json
{
"error": {
"code": "unauthorized",
"message": "Unauthorized",
"requestId": "…"
}
}
```
The response is deliberately identical across failure causes (no enumeration help). Browser session cookies are never accepted on the public API.
# Checks
> Ask at request time whether a subject may do an action at a cost — and use it.
A **check** is the question your product asks Limitry before doing something that costs: **may this subject do this action, at this cost?** Limitry evaluates every enabled [limit](/docs/limits/) that applies and answers in milliseconds, from Cloudflare’s edge.
```bash
curl -X POST https://api.limitry.com/v1/checks \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: 7f9c…" \
-d '{"subject":{"kind":"user","id":"usr_123"},"action":"generate-image","cost":1}'
```
```json
{
"id": "chk_9d2e…",
"allowed": false,
"reason": { "code": "limit_exceeded", "limit": "lim_3f0c…" },
"remaining": [
{
"limit": "lim_3f0c…",
"name": "Daily images",
"unit": "images",
"window": "day",
"remaining": 0,
"resetsAt": "2026-10-05T00:00:00.000Z"
}
],
"retryAfterSeconds": 41200
}
```
* **A “no” is an answer, not an error.** Every decision is HTTP 200 — branch on `allowed`. Error statuses mean the request itself failed (a bad key, an invalid body, or calling Limitry too fast).
* **All or none.** When several limits apply, the check is allowed only if the cost fits every one of them, and then uses all of them.
* **Shared pools.** Add `"shared": [{ "kind": "team", "id": "t_42" }]` (up to 4) and the check also draws from those subjects’ limits — a team’s, org’s or project’s pool. Every limit on every subject must fit; all are charged, or none, and a pool is never overspent however many users draw from it at once.
* **`mode`:** `consume` (the default) decides and uses; `preview` only decides — “would this be allowed?”; `reserve` decides and holds an estimate you settle later — see [Reservations](/docs/reservations/).
* **Retries are safe** with an `Idempotency-Key` header: the same key returns the same decision for 24 hours and uses nothing again.
* **`remaining`** lists every limit that applied, so your app can show “3 left today” without another call.
* **When Limitry can’t be reached**, follow each limit’s `failMode`: `open` (the default) allows, `closed` denies — so an outage of Limitry need not become an outage of your product.
The key needs the `checks:write` scope. Limit changes reach checks within a minute. A revoked key can keep working here for up to 60 seconds (its check is cached for speed). Checks count toward your plan (**100,000 per month** on the free plan; see Settings → Billing).
The same operation is `limitry checks create` on the [command line](/docs/cli/) and a tool on the [MCP server](/docs/mcp/).
# Command line
> Install the CLI, sign in, and call every API operation from a terminal or script.
The command line talks to the same public API as your code, as **you** (a personal access token), across every workspace you allow it.
## Install
[Section titled “Install”](#install)
```bash
npm install -g @limitry/cli
```
## Sign in
[Section titled “Sign in”](#sign-in)
```bash
limitry login
```
Your browser opens: sign in, choose which workspaces the command line may use and what it may do, and click **Allow**. A personal access token for this computer is created (valid for 90 days, listed under **Account → API tokens**) and stored in `~/.config/limitry/`.
On a server or in CI there is no browser — pass a token instead, or set it in the environment (nothing is written to disk):
```bash
limitry login --token "$TOKEN" # or: echo "$TOKEN" | limitry login
LIMITRY_TOKEN=pat_… limitry whoami
```
`limitry logout` revokes the token and forgets it.
## Choose a workspace
[Section titled “Choose a workspace”](#choose-a-workspace)
```bash
limitry workspace list
limitry workspace use acme # the default for later commands
limitry workspace audit-events list --workspace other-co # one call elsewhere
```
## Every API operation is a command
[Section titled “Every API operation is a command”](#every-api-operation-is-a-command)
Each operation in the [API reference](https://api.limitry.com/v1/docs) is a command, named after its resource and action:
```bash
limitry workspace webhook-endpoints list
limitry workspace webhook-endpoints create --url https://example.com/hooks
limitry workspace webhook-deliveries list we_123 --limit 10
limitry workspace webhook-endpoints delete we_123 --yes
```
* IDs in the path are arguments; other inputs are flags (`limitry --help` lists them, with descriptions).
* `--data '{…}'` sends a whole JSON body.
* Lists print one page; the last line names the `--cursor` for the next.
* Writes send an [idempotency key](/docs/idempotency/) for you.
* An operation that cannot be undone asks for `--yes`.
* `limitry api ` calls any endpoint directly.
## Scripts and agents
[Section titled “Scripts and agents”](#scripts-and-agents)
Add `--json` to any command for the API’s exact JSON on stdout. Errors then go to stderr as JSON too — the API’s [error envelope](/docs/errors/) with the HTTP status:
```bash
limitry workspace webhook-endpoints list --json | jq '.items[].url'
```
```json
{
"error": {
"code": "forbidden",
"message": "…",
"status": 403,
"requestId": "…"
}
}
```
Exit codes: `0` success, `1` failure (any error, including a refused API call or a missing flag).
# Errors & API rate limits
> The error envelope, stable error codes, and how fast each key may call the API.
Every non-2xx response from the API uses one envelope:
```json
{
"error": {
"code": "invalid_request",
"message": "Invalid request",
"requestId": "9f0c1c2e-…",
"docs": "https://limitry.com/docs/errors/#invalid_request",
"details": [{ "path": "name", "message": "Required" }]
}
}
```
* **`code`** is a small, stable set your integration may branch on.
* **`message`** is human-readable and may change — never parse it.
* **`requestId`** matches the `X-Request-Id` response header; include it when contacting support so a request can be found in the logs.
* **`details`** appears on validation failures, with one entry per field.
* **`docs`** links to the code’s entry below — how to fix it.
## Error codes
[Section titled “Error codes”](#error-codes)
Branch on `code`, never on `message`. Every error from a public code also carries `docs`, a link to its entry below. New codes may be added over time, so treat an unknown code like `internal_error`.
### `unauthorized`
[Section titled “unauthorized”](#unauthorized)
`401` — the credential is missing, malformed, invalid, revoked, or expired. **Fix:** send `Authorization: Bearer ` with a live credential; create a new one if it was revoked or has expired.
### `forbidden`
[Section titled “forbidden”](#forbidden)
`403` — the credential is valid but not allowed to do this: it was not granted the endpoint’s [scope](/docs/authentication/#scopes), your workspace role does not allow it, or a personal access token cannot act in the workspace named by `X-Workspace`. **Fix:** use a credential with the needed access (the `message` names the missing scope).
### `not_found`
[Section titled “not\_found”](#not_found)
`404` — the resource does not exist, or is not yours to see. **Fix:** check the path and the id; ids from another workspace are not found.
### `method_not_allowed`
[Section titled “method\_not\_allowed”](#method_not_allowed)
`405` — the path exists, but not with this method. **Fix:** use one of the methods in the `Allow` response header.
### `invalid_request`
[Section titled “invalid\_request”](#invalid_request)
`400` — the request failed validation. **Fix:** read `details`, one entry per invalid field (`path` and `message`), and correct the request.
### `rate_limited`
[Section titled “rate\_limited”](#rate_limited)
`429` — the credential’s quota is exhausted ([API rate limits](#api-rate-limits)). **Fix:** wait for the `Retry-After` seconds, then retry with backoff.
### `idempotency_key_reused`
[Section titled “idempotency\_key\_reused”](#idempotency_key_reused)
`409` — this `Idempotency-Key` was already used for a different request. **Fix:** use a new key for a new operation; reuse a key only to retry the same request ([Idempotency](/docs/idempotency/)).
### `idempotency_in_progress`
[Section titled “idempotency\_in\_progress”](#idempotency_in_progress)
`409` — a request with the same `Idempotency-Key` is still running. **Fix:** wait for the `Retry-After` seconds, then retry with the same key.
### `internal_error`
[Section titled “internal\_error”](#internal_error)
`500` — something failed on our side. **Fix:** retry with backoff; if it persists, contact support with the `requestId`.
## API rate limits
[Section titled “API rate limits”](#api-rate-limits)
Every authenticated request counts against a per-credential quota (per API key on the workspace endpoints; per token on `/v1/user/*`). Exhausting it returns:
```http
HTTP/1.1 429 Too Many Requests
Retry-After: 60
```
with `code: "rate_limited"`. Honor `Retry-After` and retry with backoff; limits are per credential, so one busy integration does not starve another key’s traffic.
# Idempotency
> Retry writes safely with the Idempotency-Key header.
Networks fail: a request can time out after the API has already acted on it. To retry a write **without doing it twice**, send an `Idempotency-Key` header — any unique string up to 255 printable ASCII characters, such as a UUID you generate per operation:
```bash
curl -X POST \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Idempotency-Key: 6f1c2e9a-8b3d-4c47-9a51-0d2f7e4b8c13" \
-H "Content-Type: application/json" \
-d '{ … }' \
https://api.limitry.com/v1/…
```
The first request with a key runs normally. Any retry with the **same key and the same request** returns the first response again — same status, same body — without running the operation a second time, and carries the header `Idempotency-Replayed: true`.
## The rules
[Section titled “The rules”](#the-rules)
* **Writes only.** The header applies to `POST`, `PUT`, `PATCH` and `DELETE`; reads are naturally safe to repeat and ignore it.
* **One key, one request.** Reusing a key for a different request (a different endpoint or body) is a `409` with code `idempotency_key_reused`. Generate a new key per operation, and reuse it only for retries of that operation.
* **Concurrent retries wait.** A retry that arrives while the first request is still running gets a `409` with code `idempotency_in_progress` and a `Retry-After` header; retry after it.
* **Failures you can retry.** Responses with a `5xx` status are not stored, so retrying the same key runs the request again. Any other response — success or a `4xx` — is what every retry receives.
* **Per credential, for 24 hours.** Keys are scoped to the API key or token that sent them (two integrations never collide), and are remembered for 24 hours.
Sending a key is optional, and recommended on every write.
# Limits
> Rules for how much of an action each subject may use per window — what checks enforce.
A **limit** says how much of an **action** a **subject** may use per **window** — for example, every `user` may `generate-image` 10 times per day. Subjects are whoever acts in your product: your users, teams, API keys or AI agents. You choose the labels; Limitry attaches no meaning to them beyond matching.
| Field | Meaning |
| ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `name` | Your label for the limit; unique in the workspace. |
| `subjectKind` | Which subjects it applies to (`user`, `apikey`, `agent` …). |
| `subjectId` | One subject, or empty for every subject of that kind. |
| `action` | The action it limits, or `*` for every action. |
| `amount` | How many units per window. |
| `unit` | What the amount counts (`images`, `credits` …) — for display. |
| `window` | `minute`, `hour`, `day` or `month` — fixed windows that reset at UTC boundaries; `total`: a balance that never resets ([Pools, balances and budgets](/docs/pools-and-budgets/)); `concurrent`: at most `amount` at once ([Rolling windows and concurrency](/docs/rolling-and-concurrency/)). |
| `rolling` | With `minute`, `hour` or `day`: count the last window at the moment of the check instead of the calendar one. Default `false`. |
| `failMode` | What your code should do when Limitry can’t be reached: `open` (allow, the default) or `closed` (deny). |
When several limits match a subject and action, all of them apply.
## Managing limits
[Section titled “Managing limits”](#managing-limits)
In the app: **Settings → Limits** (owners and admins change them; every member can see them). Through the API, with a key or token holding the `limits:read` / `limits:write` scopes:
```bash
curl -X POST https://api.limitry.com/v1/limits \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name":"Daily images","subjectKind":"user","action":"generate-image","amount":10,"unit":"images","window":"day"}'
```
The same operations are in the [command line](/docs/cli/) (`limitry limits create …`) and available to agents through the [MCP server](/docs/mcp/). Every change is in the workspace’s audit log.
Limits are enforced by [checks](/docs/checks/).
# MCP server
> Connect Claude or any MCP client to your workspace — what it can do, and how you stay in control.
AI agents that speak the Model Context Protocol (MCP) — Claude, and many other assistants and coding agents — can work in your workspace through our MCP server:
```text
https://mcp.limitry.com/mcp
```
## Connect
[Section titled “Connect”](#connect)
Add the address above as a remote MCP server (in Claude: **Settings → Connectors → Add custom connector**). The first time, your browser opens:
1. Sign in, if you are not already.
2. Choose **one workspace** the agent will work in.
3. Choose what it may do. Everything that can change data starts unticked — tick only what the agent needs.
4. Click **Allow**.
There is nothing to copy and no key to store: the agent receives its own sign-in, valid only for that workspace and that access.
## What the agent can do
[Section titled “What the agent can do”](#what-the-agent-can-do)
The agent’s tools are the [API](https://api.limitry.com/v1/docs) operations your choice allows — reading the workspace, listing webhook deliveries, and so on — always as **you**, with your role in the workspace: if you could not do something in the app, neither can the agent. Everything it changes appears in the workspace’s audit log as you, via the app.
Some things are never available to agents, even with full access — anything that sends your workspace’s data somewhere new or hands out access (for example, creating webhook endpoints or credentials). The agent can **ask** for them: you approve or deny in the app, and any secret is shown to you, never to the agent — see [Approvals](/docs/approvals/).
## Stay in control
[Section titled “Stay in control”](#stay-in-control)
* **Account → Connected apps** lists every agent you connected; disconnect one and it stops working immediately.
* Your workspace’s **audit log** records what each agent did.
* Removing you from the workspace cuts off your agents too.
## For developers
[Section titled “For developers”](#for-developers)
The server follows the MCP authorization specification: an unauthenticated request gets a `401` pointing at `/.well-known/oauth-protected-resource/mcp`, which names the authorization server (`https://api.limitry.com/auth`, with dynamic client registration and PKCE). Tokens are issued for the MCP server alone. Prefer the [command line](/docs/cli/) or the [SDK](/docs/sdk/) for your own scripts, and see [Use with AI agents](/docs/agents/) for the agent skill.
# Pagination
> Cursor pagination on list endpoints.
List endpoints use **cursor pagination**: results come newest-first in pages, with an opaque cursor pointing at the next page.
```bash
curl \
-H "Authorization: Bearer YOUR_API_KEY" \
"https://api.limitry.com/v1/workspace/audit-events?limit=25"
```
```json
{
"items": [{ "id": "…", "action": "member.invited", "createdAt": "…" }],
"nextCursor": "eyJ0IjoiMjAyNi0wOC0uLi4ifQ"
}
```
To fetch the next page, pass the cursor back unchanged:
```bash
curl \
-H "Authorization: Bearer YOUR_API_KEY" \
"https://api.limitry.com/v1/workspace/audit-events?limit=25&cursor=eyJ0IjoiMjAyNi0wOC0uLi4ifQ"
```
Iterate until `nextCursor` is `null` — that is the last page.
Rules worth knowing:
* `limit` accepts 1–100 (default 25).
* Cursors are **opaque**: never construct or modify one; a malformed cursor returns `invalid_request`.
* Pages are stable under concurrent writes — new rows created while you paginate never cause skips or duplicates within your walk.
* Treat the response shape as additive: new fields may appear on items, existing ones will not be renamed.
# Pools, balances and budgets
> Share an allowance across a team, sell prepaid credits, and cap what an agent can spend — all with limits.
Everything in Limitry is a **limit**. Three common needs are limits used a particular way.
## Pools — an allowance shared by a team
[Section titled “Pools — an allowance shared by a team”](#pools--an-allowance-shared-by-a-team)
Define a limit on the shared subject — say 10,000 AI calls a month for each `team` — and name the team on each user’s check:
```json
{
"subject": { "kind": "user", "id": "u_7" },
"shared": [{ "kind": "team", "id": "t_42" }],
"action": "ai-call"
}
```
The check must fit **every** limit on the user and on the team, and charges all of them — or none. However many of the team’s users check at once, the pool is never overspent. Name up to 4 shared subjects (a team, its org, a project…).
## Balances — prepaid credits that never reset
[Section titled “Balances — prepaid credits that never reset”](#balances--prepaid-credits-that-never-reset)
Choose the window **total (a balance)**: the amount is the starting balance and it never resets. Top a subject’s balance up when they buy credits:
```bash
curl -X POST https://api.limitry.com/v1/limits/$LIMIT_ID/grants \
-H "Authorization: Bearer $LIMITRY_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "subject": { "kind": "team", "id": "t_42" }, "amount": 5000, "note": "October top-up" }'
```
or in the app: **Settings → Limits → Credits** on the balance. Every grant is kept (`GET /v1/limits/{id}/grants`). When a balance is used up, the check’s `retryAfterSeconds` is `null` — waiting won’t help; credits will. Owners and admins add credits; an agent can ask, and a person approves (`limit.grant`).
## Budgets — a spend cap for an agent
[Section titled “Budgets — a spend cap for an agent”](#budgets--a-spend-cap-for-an-agent)
A budget is a limit in money. Count in cents, on the agent, for every action:
| Field | Value |
| ------- | --------------------------- |
| Subject | `agent` (every, or one id) |
| Action | `*` |
| Unit | `cents` |
| Amount | `5000` — $50 |
| Window | `day` (or `total` for once) |
Each check passes the call’s price as `cost`. When the cost isn’t known up front, [reserve](/docs/reservations/) an estimate and commit what it actually cost. The agent stops at its budget — every check after that is a “no” until the window resets (or, for a balance, until you add more).
# Quickstart
> From zero to your first authenticated API call.
## 1. Sign up and create a workspace
[Section titled “1. Sign up and create a workspace”](#1-sign-up-and-create-a-workspace)
Create your account in the application, verify your email, and create your first workspace during onboarding. Everything in the API belongs to a workspace.
## 2. Create an API key
[Section titled “2. Create an API key”](#2-create-an-api-key)
As a workspace owner, go to **Settings → API Keys**, create a key, and copy the secret — it is shown exactly once. Store it in your secret manager, never in code.
## 3. Make your first call
[Section titled “3. Make your first call”](#3-make-your-first-call)
```bash
curl \
-H "Authorization: Bearer YOUR_API_KEY" \
https://api.limitry.com/v1/workspace
```
```json
{ "id": "…", "name": "Your Workspace" }
```
That response is the API’s “who am I” — a healthy integration smoke test.
## Next steps
[Section titled “Next steps”](#next-steps)
* [Authentication](/docs/authentication/) — how keys and personal access tokens work, and how to rotate them.
* [Webhooks](/docs/webhooks/) — push events to your systems.
* The [TypeScript SDK](/docs/sdk/) and the [command line](/docs/cli/).
* [Errors & API rate limits](/docs/errors/) and [Pagination](/docs/pagination/).
* The **interactive API reference** — every endpoint, schema, and a try-it console — lives at [`api.limitry.com/v1/docs`](https://api.limitry.com/v1/docs), generated from the same definitions that validate requests, so it is always current.
# Reservations
> Hold an estimate now, settle the actual amount when the work is done.
Some work only knows its cost afterwards — an AI call billed by tokens, a job billed by minutes. A **reservation** holds an estimate while the work runs, then settles at the actual amount.
## 1. Reserve
[Section titled “1. Reserve”](#1-reserve)
Check with `mode: "reserve"` and your estimate as the `cost`:
```bash
curl -X POST https://api.limitry.com/v1/checks \
-H "Authorization: Bearer $LIMITRY_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "subject": { "kind": "user", "id": "u_42" }, "action": "ai-call",
"cost": 2000, "mode": "reserve" }'
```
It decides exactly like a normal check. When allowed, the estimate is **held** — every other check for that subject sees it as used — and the answer carries the reservation:
```json
{
"allowed": true,
"remaining": [{ "limit": "lim_…", "remaining": 8000, "…": "…" }],
"reservation": { "id": "rsv_…", "expiresAt": "2026-10-04T12:05:00Z" }
}
```
A “no” holds nothing and has no `reservation`.
## 2. Settle
[Section titled “2. Settle”](#2-settle)
When the work is done, **commit** what it actually used:
```bash
curl -X POST https://api.limitry.com/v1/reservations/rsv_…/commit \
-H "Authorization: Bearer $LIMITRY_API_KEY" \
-H "Content-Type: application/json" -d '{ "cost": 1310 }'
```
Less than you held comes back at once. More is charged in full — the work happened — even past the limit, so the next check sees the true total.
If the work did not happen, **release** it: `POST /v1/reservations/{id}/release` gives the whole hold back.
Settling is safe to retry: the same commit (or release) answers the same. Releasing a committed reservation, or committing a released one, is refused (`400`).
## Expiry
[Section titled “Expiry”](#expiry)
A hold lasts `ttlSeconds` (default 300, at most 3,600 — set it on the check). One that nobody settles **gives itself back** when it expires, so a crashed worker never keeps a subject blocked. If your work finished after all, a late commit (within 24 hours) still charges what it used.
For longer work, **extend** a held reservation as a heartbeat — `POST /v1/reservations/{id}/extend` with `{ "ttlSeconds": 600 }` — it then expires that long from now. A reservation can also hold a slot of a [concurrent limit](/docs/rolling-and-concurrency/).
## Good to know
[Section titled “Good to know”](#good-to-know)
* A subject may hold at most 1,000 open reservations; past that a reserve is a normal “no” with `reason.code: "too_many_reservations"`.
* A reserve counts as one check on your plan; settling is free.
* Reservations need the same `checks:write` scope as checks.
# Rolling windows and concurrency
> Limits like 100 in any 24 hours, and at most 5 at once — the two that fixed windows cannot express.
## Rolling windows
[Section titled “Rolling windows”](#rolling-windows)
A `minute`, `hour` or `day` limit resets at the UTC boundary by default — so 100 calls at 23:59 and 100 more at 00:01 are both allowed. Set `rolling: true` and the limit counts **the last 60 seconds, 60 minutes or 24 hours** at the moment of each check instead: no burst at the boundary.
```bash
curl -X POST https://api.limitry.com/v1/limits \
-H "Authorization: Bearer $LIMITRY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name":"Calls in any 24 hours","subjectKind":"user","action":"call",
"amount":100,"window":"day","rolling":true}'
```
Use is counted in 60 slices of the window (a second for a minute, a minute for an hour, 24 minutes for a day), and the oldest slice counts whole — so a rolling limit **never allows more than its amount**, and may refuse at most one slice early. In a check’s answer, `resetsAt` (and `retryAfterSeconds` on a denial) is when enough use leaves the window.
`month`, `total` and `concurrent` limits cannot be rolling. In the app: **Window → In any 60 seconds / 60 minutes / 24 hours**.
## Concurrency
[Section titled “Concurrency”](#concurrency)
A limit with window `concurrent` caps work **in flight**: `amount` is how many at once. Each [reservation](/docs/reservations/) takes one slot — whatever its `cost`, which still counts against the subject’s other limits — and gives it back when it is committed, released or expires.
```bash
# At most 3 exports at once per user
curl -X POST https://api.limitry.com/v1/limits ... \
-d '{"name":"Exports at once","subjectKind":"user","action":"export",
"amount":3,"window":"concurrent"}'
# Start one: a reservation holds the slot
curl -X POST https://api.limitry.com/v1/checks ... \
-d '{"subject":{"kind":"user","id":"u_42"},"action":"export","mode":"reserve"}'
# Finished (or failed): give it back
curl -X POST https://api.limitry.com/v1/reservations/$ID/release ...
```
* A `consume` or `preview` check against a concurrent limit is allowed while a slot is free but takes none — instant work holds nothing.
* When every slot is held, `retryAfterSeconds` is when the earliest hold expires (a release can free one sooner).
* Always settle in a `finally`: an unsettled hold keeps its slot until it expires.
### Long-running work
[Section titled “Long-running work”](#long-running-work)
A hold lasts `ttlSeconds` (default 300, at most 3,600). For work that runs longer, extend it as a heartbeat:
```bash
curl -X POST https://api.limitry.com/v1/reservations/$ID/extend ... \
-d '{"ttlSeconds":600}'
```
If your worker dies, it stops extending and the slot frees itself.
A limit cannot change to or from `concurrent` (one counts use, the other open holds) — create a new limit instead.
## In reports
[Section titled “In reports”](#in-reports)
Rolling limits report like others; the Usage page’s “Right now” shows the last window, to the hour. For a concurrent limit the chart shows the work it let start each day, and “Right now” the slots each subject holds — among subjects checked in the last hour. `limit.exceeded` fires at most once an hour per subject for both.
# TypeScript SDK
> Call the API from TypeScript with types generated from the API itself.
The SDK is a small typed client: every path, parameter and response is typed from the same definitions that validate requests, so it matches the [API reference](https://api.limitry.com/v1/docs) exactly.
## Install
[Section titled “Install”](#install)
```bash
npm install @limitry/sdk
```
## Make a call
[Section titled “Make a call”](#make-a-call)
```ts
import { createClient, unwrap } from "@limitry/sdk";
const api = createClient({ token: process.env.API_KEY! });
const endpoints = unwrap(await api.GET("/workspace/webhook-endpoints"));
for (const endpoint of endpoints.items) console.log(endpoint.url);
```
* With an **API key**, every call acts in the key’s workspace.
* With a **personal access token**, name the workspace: `createClient({ token, workspace: "acme" })`.
* `unwrap` returns the data or throws an `ApiError` carrying the [error envelope](/docs/errors/)’s `code`, `message`, `status` and `requestId`.
* Writes may send an [idempotency key](/docs/idempotency/): `api.POST("/workspace/webhook-endpoints", { body, headers: { "idempotency-key": id } })`.
# Usage and reports
> See how each limit is used over time, who is nearest it now, and get told when a limit is hit.
Every check is counted per limit, per subject, per hour — and reported in the app and the API, at most about a minute behind live checks.
## In the app
[Section titled “In the app”](#in-the-app)
**Usage** shows, for the limit you pick:
* **Last 30 days** — units used per day (UTC), by every subject together.
* **Right now** — this window’s subjects, most used first: what each used, what is left, and **Over — n denied** for anyone turned away.
## Through the API
[Section titled “Through the API”](#through-the-api)
With a key or token holding `limits:read`:
```bash
# A limit's use over time: by day (default) or hour, all subjects or one
curl "https://api.limitry.com/v1/usage?limitId=$LIMIT_ID&granularity=day" \
-H "Authorization: Bearer $LIMITRY_API_KEY"
# Who is using it in the current window
curl "https://api.limitry.com/v1/usage/subjects?limitId=$LIMIT_ID" \
-H "Authorization: Bearer $LIMITRY_API_KEY"
```
`/v1/usage` answers every bucket in the range (zeros included) with `used`, `allowed` and `denied`; add `subjectKind` and `subjectId` for one subject, and `from` / `to` (ISO times) for another range — the default is the last 30 days. `used` is net: a reservation’s unused part, or a released one, is given back.
For a balance (window `total`), “what is left” includes the credits you’ve added.
History is kept for 13 months.
## Webhook: `limit.exceeded`
[Section titled “Webhook: limit.exceeded”](#webhook-limitexceeded)
The first time a subject is denied by a limit in a window, your webhook endpoints get `limit.exceeded`:
```json
{
"type": "limit.exceeded",
"data": {
"limit": "lim_…",
"subject": { "kind": "user", "id": "u_42" },
"windowStart": "2026-10-05T00:00:00.000Z",
"resetsAt": "2026-10-06T00:00:00.000Z"
}
}
```
Once per subject, limit and window — not on every denied check. `resetsAt` is `null` for a balance. A budget is a limit, so an agent running out of budget is this event too. Subscribe in **Settings → Webhooks**.
# Versioning
> What can change in the API, what never changes, and how deprecations are announced.
The API is versioned in its path: `/v1`. Within a version we only make changes that do not break a correct integration — and the same promise covers the command line and agent tools built from the API, whose names come from the API’s operations.
## Changes we may make at any time
[Section titled “Changes we may make at any time”](#changes-we-may-make-at-any-time)
* New endpoints.
* New optional request parameters and fields.
* New fields in responses.
* New error codes, new webhook event types, new scopes.
* New commands and tools for the command line and agents.
Build your integration to tolerate these:
* **Ignore response fields you do not recognize.**
* **Treat an unknown error `code` like `internal_error`**, and branch on `code`, never on `message`.
* **Treat cursors and ids as opaque** strings.
## Changes we never make within `/v1`
[Section titled “Changes we never make within /v1”](#changes-we-never-make-within-v1)
* Removing or renaming an endpoint, a field, or an operation (and so a command or an agent tool).
* Changing the type or meaning of an existing field.
* Making an optional parameter required.
* Changing how requests authenticate, or which error a given failure returns.
A change like these ships only in a new version (`/v2`), served alongside `/v1` while integrations move.
## Deprecations
[Section titled “Deprecations”](#deprecations)
Before anything in `/v1` is retired, it is announced in the changelog and marked **deprecated** in the [API reference](https://api.limitry.com/v1/docs), with the replacement to use. Deprecated endpoints keep working for the life of `/v1`.
# Webhooks
> Receive workspace events at your endpoint, verify signatures, and handle retries.
Webhooks push workspace events to an HTTPS endpoint you host, as they happen. Workspace owners manage endpoints in the application under **Settings → Webhooks** — add a URL, copy the signing secret (shown exactly once), and events start flowing.
## Managing endpoints with the API
[Section titled “Managing endpoints with the API”](#managing-endpoints-with-the-api)
Software can manage endpoints too, with a credential granted the `workspace.webhooks:write` scope (reading needs `workspace.webhooks:read`; a personal access token must also belong to a workspace owner):
```bash
curl -X POST \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Idempotency-Key: $(uuidgen)" \
-H "Content-Type: application/json" \
-d '{ "url": "https://example.com/hooks" }' \
https://api.limitry.com/v1/workspace/webhook-endpoints
```
The response carries the endpoint’s signing secret — once. The API also lists endpoints, re-enables one after repeated failures, sends a test event, and lists and redelivers deliveries; see the [API reference](https://api.limitry.com/v1/docs).
AI agents connected to the workspace can read endpoints and delivery history, send test events and redeliver — but not add, remove or re-enable endpoints: those change where your workspace’s data is sent, so a person configures them.
## The delivery
[Section titled “The delivery”](#the-delivery)
Each delivery is an HTTP `POST` with a JSON body:
```json
{
"id": "evt_5f0c…",
"type": "workspace.member.joined",
"createdAt": "2026-08-20T12:00:00.000Z",
"data": { "userId": "…", "email": "casey@example.com" }
}
```
and three signature headers:
```http
webhook-id: del_9a1b…
webhook-timestamp: 1755691200
webhook-signature: v1,MEQCIB…
```
* **`webhook-id`** identifies this delivery. It stays the same across retries — use it to deduplicate.
* **`id`** inside the body identifies the *event*. If you registered multiple endpoints, each receives its own delivery of the same event — deduplicate across endpoints by event `id` if you need to.
## Verifying signatures
[Section titled “Verifying signatures”](#verifying-signatures)
Deliveries are signed with your endpoint’s secret (`whsec_…`) using the same scheme as [Svix](https://docs.svix.com/receiving/verifying-payloads/how), so any standard Svix library verifies them:
```ts
import { Webhook } from "svix";
const wh = new Webhook(process.env.WEBHOOK_SECRET);
// Express-style handler; `payload` must be the RAW request body string.
app.post("/webhooks", (req, res) => {
let event;
try {
event = wh.verify(req.body, req.headers);
} catch {
return res.status(400).send("bad signature");
}
// handle event…
res.status(200).send("ok");
});
```
Verifying by hand: the signature is `v1,` followed by a Base64 HMAC-SHA256 of `` `${webhookId}.${timestamp}.${body}` `` keyed with the secret after its `whsec_` prefix (Base64-decoded). Always verify against the **raw** body — re-serializing JSON breaks the signature — and reject timestamps older than a few minutes to prevent replays.
## Respond fast, process later
[Section titled “Respond fast, process later”](#respond-fast-process-later)
Return a `2xx` within 10 seconds. If your processing is slow, acknowledge first and process asynchronously — a timeout counts as a failed delivery.
## Retries and failures
[Section titled “Retries and failures”](#retries-and-failures)
Failed deliveries (non-2xx, or timeout) retry automatically with backoff, up to 6 attempts, with the same `webhook-id`. An endpoint that fails **20 consecutive** deliveries is disabled automatically — delete and re-add it (new secret) once your endpoint is healthy.
Under **Settings → Webhooks → Deliveries** you can see each delivery’s status and attempts, send a test event (`type: "ping"`), and **redeliver** any recorded delivery — a redelivery arrives with a fresh `webhook-id` but the same event `id`, so event-level deduplication still applies.
## Events
[Section titled “Events”](#events)
| Type | Fires when | `data` |
| ------------------------- | ------------------------------ | ----------------- |
| `workspace.member.joined` | An invitation is accepted. | `userId`, `email` |
| `ping` | You send a test from Settings. | a test message |
Event payloads only ever gain fields — build tolerant parsers. `GET /v1/workspace/event-types` returns the full list, with what each means.
## Choosing events
[Section titled “Choosing events”](#choosing-events)
An endpoint receives **every** event unless you choose: in **Settings → Webhooks** pick “Only these” when adding it, or pass `eventTypes` when creating it through the API — only listed types are accepted:
```bash
curl -X POST https://api.limitry.com/v1/workspace/webhook-endpoints \
-H "Authorization: Bearer YOUR_API_KEY" -H "content-type: application/json" \
-d '{"url":"https://example.com/hooks","eventTypes":["workspace.member.joined"]}'
```
A test event (`ping`) always reaches the endpoint you test.
# Workspace usage
> How much of your plan the workspace used this month, and the limits.
Some things are counted per calendar month (UTC) — each meter, with your plan’s limit for it. See them under **Settings → Billing**, or through the API:
```bash
curl -H "Authorization: Bearer YOUR_API_KEY" \
https://api.limitry.com/v1/workspace/usage
```
```json
{
"periodStart": "2026-10-01T00:00:00.000Z",
"periodEnd": "2026-11-01T00:00:00.000Z",
"meters": [
{
"id": "webhook_deliveries",
"description": "Webhook deliveries this month (each event, each endpoint)",
"used": 1520,
"limit": null
}
]
}
```
`limit` is `null` when the plan has no limit for that meter. Counters reset at the start of each month. The key needs the `workspace:read` scope.