Quick fixes
Retry 429 and 5xx. Fix everything else.
429: wait for the Retry-After header, then back off (rejected requests aren’t billed). 401: check your key and the Authorization header. 403 agent_api_migration_required: move from Sonar to the Agent API. 400: fix the request body.
Reference
Perplexity API error codes
What each error means, whether to retry, and how to fix it.
| Code | Meaning | Retry? | How to fix |
|---|---|---|---|
| 400 | Bad request Malformed JSON, a wrong parameter name or type, or a value out of range (e.g. max_results above 20, wrong date format). | Don’t retry | Fix the request body. Check parameter names against the docs and validate inputs before sending. |
| 401 | Unauthorized The API key is missing, mistyped, revoked, or not sent in the Authorization header. | Don’t retry | Send Authorization: Bearer YOUR_KEY, confirm the environment variable is loaded, and create a new key if it was revoked. |
| 403 | Forbidden / migration required Calling the retired Sonar Chat Completions endpoint, or using a feature your account can’t access.agent_api_migration_required | Don’t retry | If the error is agent_api_migration_required, move to the Agent API (migration guide). Otherwise check your account and key permissions. |
| 404 | Not found Wrong URL or endpoint path, or an unknown resource ID (for example when retrieving a background response). | Don’t retry | Check the base URL and path: /v1/agent, /search. Confirm the response ID is correct. |
| Credit | Out of credit / billing issue Your API credit is used up or the payment method failed. The exact status code can vary. | Don’t retry | Add credit or update your payment method in the API console, and consider auto top-up. |
| 429 | Too many requests You exceeded your rate limit. The response includes a Retry-After header. Rejected 429 requests are not billed.Retry-After header | Retry | Wait for Retry-After, then retry with exponential backoff and jitter. Throttle on your side and spread out bursts. |
| 500 | Internal server error A temporary problem on Perplexity’s side. | Retry | Retry with backoff. If it continues, check Perplexity’s status page and contact api@perplexity.ai. |
| 502 / 503 / 504 | Gateway, unavailable or timeout Temporary overload, maintenance or a slow upstream. | Retry | Retry with backoff. For long jobs, use background=True and poll instead of holding a connection open. |
| Timeout | Client timeout Your HTTP client gave up before the answer arrived, common with high or xhigh presets. | Adjust | Raise your client timeout, stream the response, or use background runs for long research. |
Status codes follow standard HTTP meanings. Read the error message in the response body for specifics, and log it with the request ID.
Decide
Retry or fix? A quick decision guide
Rate limits
Usage tiers and rate limits
Your tier is set by your cumulative API purchases over the account’s lifetime, not your current balance. Higher tiers get higher Agent API limits.
| Tier | Cumulative purchases | Agent API | Per minute (max) |
|---|---|---|---|
| Tier 0 | $0 (new account) | 1 QPS | 60 / min |
| Tier 1 | $50+ | 3 QPS | 180 / min |
| Tier 2 | $250+ | 8 QPS | 480 / min |
| Tier 3 | $500+ | 17 QPS | 1,020 / min |
| Tier 4 | $1,000+ | 33 QPS | 1,980 / min |
| Tier 5 | $5,000+ | 33 QPS | 1,980 / min |
Agent API queries per second by tier
| API | Limit | Notes |
|---|---|---|
| Agent API | 1–33 QPS by tier | Rolling one-second window, shared across your organization |
| Search API | 50 query units / second | Same on all tiers; burst of 50. Each query in a multi-query request uses 1 unit |
| Embeddings | ~85–335 QPS by tier | Contextualized embeddings are limited by chunks per second (about 415–1,670) |
Source: Perplexity rate limits and usage tiers, checked 6 October 2026.
Mechanics
How Perplexity’s rate limits work
Each accepted Agent API request counts against your limit for one second. Ten requests at once on a 3 QPS tier means seven get a 429.
The Search API refills steadily at 50 units per second, with room for a short burst of 50.
Every key and app on the account draws from the same Agent API limit, so plan for your total traffic.
Rate limit planner
Check whether your expected peak traffic fits your tier.
Code
Retry with backoff: copy-ready code
Retries only 429 and 5xx, honors Retry-After when present, and otherwise uses exponential backoff with jitter.
import random
import time
from perplexity import Perplexity
client = Perplexity()
RETRYABLE = {429, 500, 502, 503, 504}
def create_with_retry(max_attempts: int = 6, **kwargs):
for attempt in range(max_attempts):
try:
return client.responses.create(**kwargs)
except Exception as error:
status = getattr(error, "status_code", None)
if status not in RETRYABLE or attempt == max_attempts - 1:
raise # 400/401/403 etc.: fix the request, don't retry
# Prefer the server's Retry-After header when present
response = getattr(error, "response", None)
retry_after = response.headers.get("retry-after") if response is not None else None
if retry_after and retry_after.replace(".", "", 1).isdigit():
wait = float(retry_after)
else:
wait = min(60, 2 ** attempt) + random.uniform(0, 1) # backoff + jitter
print(f"{status}: retrying in {wait:.1f}s (attempt {attempt + 1})")
time.sleep(wait)
answer = create_with_retry(preset="fast", input="What is the capital of Canada?")
print(answer.output_text)// Node.js 18+ (ESM). Server-side only.
const RETRYABLE = new Set([429, 500, 502, 503, 504]);
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
export async function agentRequest(body, maxAttempts = 6) {
for (let attempt = 0; attempt < maxAttempts; attempt++) {
const res = await fetch("https://api.perplexity.ai/v1/agent", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.PERPLEXITY_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
});
if (res.ok) return res.json();
const text = await res.text();
if (!RETRYABLE.has(res.status) || attempt === maxAttempts - 1) {
throw new Error(`Perplexity API ${res.status}: ${text}`);
}
const retryAfter = Number(res.headers.get("retry-after"));
const wait = Number.isFinite(retryAfter) && retryAfter > 0
? retryAfter * 1000
: Math.min(60000, 2 ** attempt * 1000) + Math.random() * 1000;
await sleep(wait);
}
}
const data = await agentRequest({ preset: "fast", input: "What is the capital of Canada?" });
console.log(data.output_text);Prevention
Stay under the limit
Retries handle the odd 429; a client-side limiter stops them happening. This async example spaces requests to stay just under a 3 QPS tier.
import asyncio
import time
from perplexity import AsyncPerplexity
client = AsyncPerplexity()
class RateLimiter:
"""Allow at most `qps` request starts per second."""
def __init__(self, qps: float):
self.interval = 1.0 / qps
self.lock = asyncio.Lock()
self.next_time = 0.0
async def wait(self):
async with self.lock:
now = time.monotonic()
if now < self.next_time:
await asyncio.sleep(self.next_time - now)
self.next_time = max(now, self.next_time) + self.interval
limiter = RateLimiter(qps=2.5) # stay a little under your tier's limit (e.g. 3 QPS)
async def ask(question: str) -> str:
await limiter.wait()
response = await client.responses.create(preset="fast", input=question)
return response.output_text
async def main():
questions = [f"Give me one fact about the number {n}" for n in range(1, 11)]
answers = await asyncio.gather(*(ask(q) for q in questions))
for a in answers:
print(a)
asyncio.run(main())Target about 80–90% of your QPS to leave headroom for other apps on the account.
Push batch jobs through a queue with a fixed worker count instead of firing them all at once.
Every cached answer is one fewer request against your limit and your bill.
Lighter presets finish sooner; the Search API has its own, higher limit.
Record codes, timings and IDs so you can spot patterns and share them with support.
Sustained high traffic? Higher tiers unlock with cumulative purchases.
FAQ
Errors and rate limit questions
What are the Perplexity API rate limits?
For the Agent API, limits depend on your usage tier: 1 query per second at Tier 0, 3 at Tier 1, 8 at Tier 2, 17 at Tier 3 and 33 at Tiers 4–5. The Search API allows 50 query units per second on every tier. Embeddings range from about 85 to 335 QPS by tier.
How do I move up a usage tier?
Tiers are based on your cumulative API purchases over the account’s lifetime, not your current balance: $50 for Tier 1, $250 for Tier 2, $500 for Tier 3, $1,000 for Tier 4 and $5,000 for Tier 5.
Am I charged for requests that get a 429?
No. Perplexity says requests rejected with a 429 are not billed.
What does agent_api_migration_required mean?
You’re calling the retired Sonar Chat Completions API. Sonar was retired on 27 September 2026; switch to the Agent API at /v1/agent. Our Sonar migration guide shows how.
Should I retry every error?
No. Retry 429 and 5xx errors with backoff. Don’t retry 400, 401, 403 or 404: those mean the request, key or endpoint needs fixing, and retrying just wastes time.
Are rate limits per API key or per account?
Agent API limits are shared across your organization, so several keys or apps on one account draw from the same limit.
Why am I getting 429s below my limit?
Bursts can exceed the per-second limit even if your per-minute average is fine. Spread requests evenly with a client-side limiter, and remember multi-query Search API calls use one unit per query.
More API guides
Still stuck?
Check your usage and contact support
Look at usage, tier and credit in the API console. For persistent 5xx errors, email api@perplexity.ai with the request ID and timestamp.