Skip to main content

Rate limits

The API throttles requests to protect the instance. When you blow a limit the response is 429 Too Many Requests — wait and try again.

Overall limit​

3,000 requests per minute, per IP. It applies to everything, including what is not the API.

Per-endpoint limits​

Some expensive endpoints have their own limits, counted per account:

EndpointLimit
GET /api/v1/accounts/{id}/contacts/search100 / minute
POST /api/v1/accounts/{id}/upload60 / hour
GET /api/v1/accounts/{id}/conversations/{id}/transcript1,000 / hour
GET /api/v2/accounts/{id}/reports1,000 / minute per account, and 100 / minute per user
GET /api/v1/accounts/{id}/conversations/meta30 / minute per user

On the "per user" limits, the API key itself is the identity — each key has its own quota.

Authentication endpoints​

These are stricter, to hold back brute force: login accepts 5 attempts per IP every 5 minutes and 10 per e-mail every 15 minutes; password recovery, 5 per IP every 30 minutes.

Tuning it on your instance​

If you host AzChat yourself, the limits are environment variables:

VariableDefault
RACK_ATTACK_LIMIT3000
RATE_LIMIT_CONTACT_SEARCH100
RATE_LIMIT_REPORTS_API_ACCOUNT_LEVEL1000
RATE_LIMIT_REPORTS_API_USER_LEVEL100
RATE_LIMIT_CONVERSATION_TRANSCRIPT1000
RATE_LIMIT_CONVERSATIONS_META30
RACK_ATTACK_ALLOWED_IPS— list of IPs that bypass the limits
ENABLE_RACK_ATTACKtrue

:::info No limits in development Rate limiting is only applied in production. A script that works on your local machine can start taking 429s the moment it ships — test your handling beforehand. :::

Handling the 429​

On a 429, wait before retrying, widening the interval on each attempt:

async function withRetry(request, attempts = 5) {
for (let attempt = 0; attempt < attempts; attempt += 1) {
const response = await request();
if (response.status !== 429) return response;

// 1s, 2s, 4s, 8s…
await new Promise((r) => setTimeout(r, 1000 * 2 ** attempt));
}
throw new Error('Rate limit exceeded');
}

Good practices that avoid the problem at the source:

  • Use webhooks instead of polling the API in a loop.
  • Raise per in pagination to make fewer calls.
  • Spread large imports over time instead of firing everything at once.