Independent guide. perplexitiai.com is not affiliated with or endorsed by Perplexity AI, Inc. The official site is perplexity.ai.

Perplexity API Errors and Rate Limits

Got a 429, 401 or 403? This page explains every common Perplexity API error and what to do about it, shows the official usage tiers and rate limits, and gives you copy-ready retry code and a planner to check your traffic fits your tier.

Quick fixes

Retry 429 and 5xx. Fix everything else.

429: wait for the Retry-After header, then back off (rejected requests aren’t billed). 401: check your key and the Authorization header. 403 agent_api_migration_required: move from Sonar to the Agent API. 400: fix the request body.

Reference

Perplexity API error codes

What each error means, whether to retry, and how to fix it.

CodeMeaningRetry?How to fix
400Bad request
Malformed JSON, a wrong parameter name or type, or a value out of range (e.g. max_results above 20, wrong date format).
Don’t retryFix the request body. Check parameter names against the docs and validate inputs before sending.
401Unauthorized
The API key is missing, mistyped, revoked, or not sent in the Authorization header.
Don’t retrySend Authorization: Bearer YOUR_KEY, confirm the environment variable is loaded, and create a new key if it was revoked.
403Forbidden / migration required
Calling the retired Sonar Chat Completions endpoint, or using a feature your account can’t access.agent_api_migration_required
Don’t retryIf the error is agent_api_migration_required, move to the Agent API (migration guide). Otherwise check your account and key permissions.
404Not found
Wrong URL or endpoint path, or an unknown resource ID (for example when retrieving a background response).
Don’t retryCheck the base URL and path: /v1/agent, /search. Confirm the response ID is correct.
CreditOut of credit / billing issue
Your API credit is used up or the payment method failed. The exact status code can vary.
Don’t retryAdd credit or update your payment method in the API console, and consider auto top-up.
429Too many requests
You exceeded your rate limit. The response includes a Retry-After header. Rejected 429 requests are not billed.Retry-After header
RetryWait for Retry-After, then retry with exponential backoff and jitter. Throttle on your side and spread out bursts.
500Internal server error
A temporary problem on Perplexity’s side.
RetryRetry with backoff. If it continues, check Perplexity’s status page and contact api@perplexity.ai.
502 / 503 / 504Gateway, unavailable or timeout
Temporary overload, maintenance or a slow upstream.
RetryRetry with backoff. For long jobs, use background=True and poll instead of holding a connection open.
TimeoutClient timeout
Your HTTP client gave up before the answer arrived, common with high or xhigh presets.
AdjustRaise your client timeout, stream the response, or use background runs for long research.

Status codes follow standard HTTP meanings. Read the error message in the response body for specifics, and log it with the request ID.

Decide

Retry or fix? A quick decision guide

Request failed What’s the status code? 429 · 5xx 400 · 401 · 403 · 404 Wait, then retry Retry-After or backoff + jitter Fix, don’t retry Check body, key or endpoint Timeout? Stream or use background runs

Rate limits

Usage tiers and rate limits

Your tier is set by your cumulative API purchases over the account’s lifetime, not your current balance. Higher tiers get higher Agent API limits.

TierCumulative purchasesAgent APIPer minute (max)
Tier 0$0 (new account)1 QPS60 / min
Tier 1$50+3 QPS180 / min
Tier 2$250+8 QPS480 / min
Tier 3$500+17 QPS1,020 / min
Tier 4$1,000+33 QPS1,980 / min
Tier 5$5,000+33 QPS1,980 / min
APILimitNotes
Agent API1–33 QPS by tierRolling one-second window, shared across your organization
Search API50 query units / secondSame on all tiers; burst of 50. Each query in a multi-query request uses 1 unit
Embeddings~85–335 QPS by tierContextualized embeddings are limited by chunks per second (about 415–1,670)

Source: Perplexity rate limits and usage tiers, checked 6 October 2026.

Mechanics

How Perplexity’s rate limits work

Rolling one-second window

Each accepted Agent API request counts against your limit for one second. Ten requests at once on a 3 QPS tier means seven get a 429.

Leaky bucket for search

The Search API refills steadily at 50 units per second, with room for a short burst of 50.

Shared across your org

Every key and app on the account draws from the same Agent API limit, so plan for your total traffic.

Rate limit planner

Check whether your expected peak traffic fits your tier.

Code

Retry with backoff: copy-ready code

Retries only 429 and 5xx, honors Retry-After when present, and otherwise uses exponential backoff with jitter.

Python
import random
import time

from perplexity import Perplexity

client = Perplexity()
RETRYABLE = {429, 500, 502, 503, 504}

def create_with_retry(max_attempts: int = 6, **kwargs):
    for attempt in range(max_attempts):
        try:
            return client.responses.create(**kwargs)
        except Exception as error:
            status = getattr(error, "status_code", None)
            if status not in RETRYABLE or attempt == max_attempts - 1:
                raise  # 400/401/403 etc.: fix the request, don't retry

            # Prefer the server's Retry-After header when present
            response = getattr(error, "response", None)
            retry_after = response.headers.get("retry-after") if response is not None else None
            if retry_after and retry_after.replace(".", "", 1).isdigit():
                wait = float(retry_after)
            else:
                wait = min(60, 2 ** attempt) + random.uniform(0, 1)  # backoff + jitter

            print(f"{status}: retrying in {wait:.1f}s (attempt {attempt + 1})")
            time.sleep(wait)

answer = create_with_retry(preset="fast", input="What is the capital of Canada?")
print(answer.output_text)

Prevention

Stay under the limit

Retries handle the odd 429; a client-side limiter stops them happening. This async example spaces requests to stay just under a 3 QPS tier.

Python
import asyncio
import time

from perplexity import AsyncPerplexity

client = AsyncPerplexity()

class RateLimiter:
    """Allow at most `qps` request starts per second."""
    def __init__(self, qps: float):
        self.interval = 1.0 / qps
        self.lock = asyncio.Lock()
        self.next_time = 0.0

    async def wait(self):
        async with self.lock:
            now = time.monotonic()
            if now < self.next_time:
                await asyncio.sleep(self.next_time - now)
            self.next_time = max(now, self.next_time) + self.interval

limiter = RateLimiter(qps=2.5)   # stay a little under your tier's limit (e.g. 3 QPS)

async def ask(question: str) -> str:
    await limiter.wait()
    response = await client.responses.create(preset="fast", input=question)
    return response.output_text

async def main():
    questions = [f"Give me one fact about the number {n}" for n in range(1, 11)]
    answers = await asyncio.gather(*(ask(q) for q in questions))
    for a in answers:
        print(a)

asyncio.run(main())
Throttle below your limit

Target about 80–90% of your QPS to leave headroom for other apps on the account.

Queue background work

Push batch jobs through a queue with a fixed worker count instead of firing them all at once.

Cache repeat questions

Every cached answer is one fewer request against your limit and your bill.

Use fast and Search where possible

Lighter presets finish sooner; the Search API has its own, higher limit.

Log status and request IDs

Record codes, timings and IDs so you can spot patterns and share them with support.

Grow your tier

Sustained high traffic? Higher tiers unlock with cumulative purchases.

FAQ

Errors and rate limit questions

What are the Perplexity API rate limits?

For the Agent API, limits depend on your usage tier: 1 query per second at Tier 0, 3 at Tier 1, 8 at Tier 2, 17 at Tier 3 and 33 at Tiers 4–5. The Search API allows 50 query units per second on every tier. Embeddings range from about 85 to 335 QPS by tier.

How do I move up a usage tier?

Tiers are based on your cumulative API purchases over the account’s lifetime, not your current balance: $50 for Tier 1, $250 for Tier 2, $500 for Tier 3, $1,000 for Tier 4 and $5,000 for Tier 5.

Am I charged for requests that get a 429?

No. Perplexity says requests rejected with a 429 are not billed.

What does agent_api_migration_required mean?

You’re calling the retired Sonar Chat Completions API. Sonar was retired on 27 September 2026; switch to the Agent API at /v1/agent. Our Sonar migration guide shows how.

Should I retry every error?

No. Retry 429 and 5xx errors with backoff. Don’t retry 400, 401, 403 or 404: those mean the request, key or endpoint needs fixing, and retrying just wastes time.

Are rate limits per API key or per account?

Agent API limits are shared across your organization, so several keys or apps on one account draw from the same limit.

Why am I getting 429s below my limit?

Bursts can exceed the per-second limit even if your per-minute average is fine. Spread requests evenly with a client-side limiter, and remember multi-query Search API calls use one unit per query.

More API guides

Still stuck?

Check your usage and contact support

Look at usage, tier and credit in the API console. For persistent 5xx errors, email api@perplexity.ai with the request ID and timestamp.