Rate limits
The API throttles requests to protect the instance. When you blow a limit the response is 429 Too Many Requests — wait and try again.
Overall limit
3,000 requests per minute, per IP. It applies to everything, including what is not the API.
Per-endpoint limits
Some expensive endpoints have their own limits, counted per account:
| Endpoint | Limit |
|---|---|
GET /api/v1/accounts/{id}/contacts/search | 100 / minute |
POST /api/v1/accounts/{id}/upload | 60 / hour |
GET /api/v1/accounts/{id}/conversations/{id}/transcript | 1,000 / hour |
GET /api/v2/accounts/{id}/reports | 1,000 / minute per account, and 100 / minute per user |
GET /api/v1/accounts/{id}/conversations/meta | 30 / minute per user |
On the "per user" limits, the API key itself is the identity — each key has its own quota.
Authentication endpoints
These are stricter, to hold back brute force: login accepts 5 attempts per IP every 5 minutes and 10 per e-mail every 15 minutes; password recovery, 5 per IP every 30 minutes.
Tuning it on your instance
If you host AzChat yourself, the limits are environment variables:
| Variable | Default |
|---|---|
RACK_ATTACK_LIMIT | 3000 |
RATE_LIMIT_CONTACT_SEARCH | 100 |
RATE_LIMIT_REPORTS_API_ACCOUNT_LEVEL | 1000 |
RATE_LIMIT_REPORTS_API_USER_LEVEL | 100 |
RATE_LIMIT_CONVERSATION_TRANSCRIPT | 1000 |
RATE_LIMIT_CONVERSATIONS_META | 30 |
RACK_ATTACK_ALLOWED_IPS | — list of IPs that bypass the limits |
ENABLE_RACK_ATTACK | true |
:::info No limits in development
Rate limiting is only applied in production. A script that works on your local
machine can start taking 429s the moment it ships — test your handling
beforehand.
:::
Handling the 429
On a 429, wait before retrying, widening the interval on each attempt:
async function withRetry(request, attempts = 5) {
for (let attempt = 0; attempt < attempts; attempt += 1) {
const response = await request();
if (response.status !== 429) return response;
// 1s, 2s, 4s, 8s…
await new Promise((r) => setTimeout(r, 1000 * 2 ** attempt));
}
throw new Error('Rate limit exceeded');
}
Good practices that avoid the problem at the source:
- Use webhooks instead of polling the API in a loop.
- Raise
perin pagination to make fewer calls. - Spread large imports over time instead of firing everything at once.