Independent guide. perplexitiai.com is not affiliated with or endorsed by Perplexity AI, Inc. The official site is perplexity.ai.

Perplexity API Python Tutorial

Go from pip install to a working research app in Python. Twelve short lessons cover the Agent API and Search API: your first call, presets, streaming, sources, filters, JSON output, async and error handling, plus a complete project and a code generator.

This tutorial uses the Agent API. Perplexity retired Sonar Chat Completions (client.chat.completions with model="sonar") on 27 September 2026. Updating old code? Read the Sonar migration guide.

Tutorial

Learn the Perplexity API in Python, step by step

You’ll need a recent version of Python 3 and a Perplexity API key with some credit.

Set up your project

Create a virtual environment, install the official perplexityai package, and keep your key in a .env file. Get a key from the API console first; our API key guide shows how to keep it safe.

Terminal
# 1. Create and activate a virtual environment
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate

# 2. Install the official SDK (and dotenv for local keys)
pip install perplexityai python-dotenv
.env
# .env  (add this file to .gitignore)
PERPLEXITY_API_KEY=your_api_key_here

Make your first request

The Agent API uses client.responses.create. Give it a preset and your question, then print output_text.

Python
from dotenv import load_dotenv
from perplexity import Perplexity

load_dotenv()           # loads PERPLEXITY_API_KEY from .env
client = Perplexity()   # reads the key from the environment

response = client.responses.create(
    preset="fast",
    input="What are the latest developments in AI agents?",
)

print(response.output_text)
Recent developments in AI agents include … [1][2]

Choose a preset or a model

Presets bundle a model, tools and limits, and Perplexity keeps them up to date. You can also pass a specific model instead.

Python
# Use a preset (recommended): Perplexity picks the model and tools
response = client.responses.create(preset="low", input="Compare solar and wind costs in 2026")

# Or choose a specific model yourself
response = client.responses.create(
    model="openai/gpt-5.6-sol",
    input="Explain supervised learning in simple terms",
)
print(response.output_text)
PresetDefault modelToolsBest for
fastopenai/gpt-6-lunaweb_searchSingle facts, definitions, quick summaries
lowopenai/gpt-6-lunaweb_search, fetch_urlEveryday research, light multi-step lookups
mediumopenai/gpt-6-lunaweb_search, fetch_urlMulti-hop browsing across many sources
highopenai/gpt-6-solweb_search, fetch_urlExpert-level, exhaustive research
xhighanthropic/claude-opus-5-5web_search, finance_search, sandboxOpen-ended agentic work with code execution

Presets update automatically when Perplexity improves them. Source: Agent API presets.

Add instructions and hold a conversation

Use instructions for a system prompt. For multi-turn chat, pass a list of messages as input and append each answer.

Python
response = client.responses.create(
    preset="fast",
    instructions="You are a friendly tutor. Answer in under 120 words, in simple English.",
    input="How do vaccines train the immune system?",
)
print(response.output_text)
chat.py
from perplexity import Perplexity

client = Perplexity()
history = []

while True:
    question = input("You: ").strip()
    if question.lower() in {"quit", "exit"}:
        break
    history.append({"role": "user", "content": question})

    response = client.responses.create(preset="fast", input=history)
    answer = response.output_text
    print(f"Perplexity: {answer}\n")

    history.append({"role": "assistant", "content": answer})

Stream answers as they’re written

Set stream=True and print each response.output_text.delta event. Great for chat UIs and long answers.

Python
stream = client.responses.create(
    preset="fast",
    input="Summarize today's top technology news in 5 bullets",
    stream=True,
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print(f"\n\nUsage: {event.response.usage}")

Read the sources and citations

Numbers like [1] in output_text match the id of results in the search_results output items. Show them so users can check the answer.

Python
response = client.responses.create(
    preset="fast",
    input="How much of the Amazon rainforest has been lost?",
)

print(response.output_text, "\n")

# Sources arrive as output items of type "search_results"
for item in response.output:
    if item.type == "search_results":
        for result in item.results:
            print(f"[{result.id}] {result.title}")
            print(f"    {result.url}  ({result.date})")

Filter the web search

Pass a web_search tool with filters to limit domains (prefix - to exclude) and recency.

Python
response = client.responses.create(
    preset="fast",
    input="Latest renewable energy policy updates",
    tools=[
        {
            "type": "web_search",
            "filters": {
                "search_domain_filter": ["iea.org", "energy.gov"],
                "search_recency_filter": "month",
            },
        }
    ],
)
print(response.output_text)

Get structured JSON output

Use response_format with a JSON schema, then parse output_text with json.loads.

Python
import json

response = client.responses.create(
    preset="low",
    input="List the 3 largest countries by area with their area in km2",
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "countries",
            "schema": {
                "type": "object",
                "properties": {
                    "countries": {
                        "type": "array",
                        "items": {
                            "type": "object",
                            "properties": {
                                "name": {"type": "string"},
                                "area_km2": {"type": "number"},
                            },
                            "required": ["name", "area_km2"],
                        },
                    }
                },
                "required": ["countries"],
            },
        },
    },
)

data = json.loads(response.output_text)
for c in data["countries"]:
    print(c["name"], c["area_km2"])

Run long research in the background

For the high preset or big reports, set background=True and poll with client.responses.retrieve. The job keeps running even if your script disconnects.

Python
import time

response = client.responses.create(
    preset="high",
    input="Write a detailed report on battery recycling technologies, with sources",
    background=True,
)

while response.status in ("queued", "in_progress"):
    time.sleep(2)
    response = client.responses.retrieve(response.id)

print(response.status)
print(response.output_text)

Get raw results with the Search API

Need links instead of an answer? The Search API returns ranked results for $5 per 1,000 requests ($1 in fast mode). Full guide: Search API tutorial.

Python
search = client.search.create(
    query="EV battery recycling 2026",
    max_results=5,
    search_recency_filter="month",
)

for result in search.results:
    print(f"{result.title}: {result.url}")

Run many questions at once (async)

Use AsyncPerplexity with asyncio.gather to answer several questions concurrently. Mind your rate limits.

Python
import asyncio
from perplexity import AsyncPerplexity

client = AsyncPerplexity()

async def ask(question: str) -> str:
    response = await client.responses.create(preset="fast", input=question)
    return response.output_text

async def main():
    questions = [
        "What is the capital of Australia?",
        "Who discovered penicillin?",
        "How tall is Mount Everest?",
    ]
    answers = await asyncio.gather(*(ask(q) for q in questions))
    for q, a in zip(questions, answers):
        print(f"Q: {q}\nA: {a}\n")

asyncio.run(main())

Handle errors and retry

Retry rate limits (429) and server errors (5xx) with exponential backoff and jitter. Don’t retry 400 or 401; fix the request or key instead.

Python
import random
import time

def ask_with_retry(question: str, attempts: int = 5) -> str:
    for attempt in range(attempts):
        try:
            response = client.responses.create(preset="fast", input=question)
            return response.output_text
        except Exception as error:
            status = getattr(error, "status_code", None)
            retryable = status == 429 or (status is not None and status >= 500)
            if not retryable or attempt == attempts - 1:
                raise
            wait = (2 ** attempt) + random.random()
            print(f"Got {status}, retrying in {wait:.1f}s...")
            time.sleep(wait)

Project

Build a cited research CLI

One script that puts the lessons together: it streams an answer to your question, applies optional domain and recency filters, then lists the sources.

research.py
"""research.py: ask a question, stream a cited answer, list the sources.

Usage:  python research.py "your question" [--domains iea.org,nature.com] [--recent week]
"""
import argparse

from dotenv import load_dotenv
from perplexity import Perplexity


def main() -> None:
    parser = argparse.ArgumentParser(description="Cited research from the terminal")
    parser.add_argument("question")
    parser.add_argument("--preset", default="low", choices=["fast", "low", "medium", "high"])
    parser.add_argument("--domains", default="", help="comma-separated, prefix - to exclude")
    parser.add_argument("--recent", default="", choices=["", "day", "week", "month", "year"])
    args = parser.parse_args()

    load_dotenv()
    client = Perplexity()

    filters = {}
    if args.domains:
        filters["search_domain_filter"] = [d.strip() for d in args.domains.split(",") if d.strip()]
    if args.recent:
        filters["search_recency_filter"] = args.recent
    tool = {"type": "web_search"}
    if filters:
        tool["filters"] = filters

    stream = client.responses.create(
        preset=args.preset,
        input=args.question,
        tools=[tool],
        stream=True,
    )

    final = None
    for event in stream:
        if event.type == "response.output_text.delta":
            print(event.delta, end="", flush=True)
        elif event.type == "response.completed":
            final = event.response

    print("\n\nSources:")
    for item in (final.output if final else []):
        if item.type == "search_results":
            for r in item.results:
                print(f"  [{r.id}] {r.title} - {r.url}")


if __name__ == "__main__":
    main()
$ python research.py "EV battery recycling breakthroughs" --recent month --domains iea.org,nature.com Recent advances include … [1][2] Sources: [1] Battery recycling report - https://www.iea.org/… [2] Direct recycling of cathodes - https://www.nature.com/…

Python code generator

Pick your options and copy ready-to-run Python.

Python

Reference

Python cheat sheet

TaskCode
Installpip install perplexityai
Create clientclient = Perplexity()
Ask a questionclient.responses.create(preset="fast", input="…")
Read answerresponse.output_text
System promptinstructions="…"
Streamstream=True → event.delta
Sourcesitem.type == "search_results" → item.results
Filterstools=[{"type": "web_search", "filters": {…}}]
JSON outputresponse_format={"type": "json_schema", …}
Backgroundbackground=True + responses.retrieve(id)
Raw resultsclient.search.create(query="…")
AsyncAsyncPerplexity() + await

FAQ

Perplexity Python questions

How do I install the Perplexity Python SDK?

Run pip install perplexityai, then import it with from perplexity import Perplexity. The client reads your key from the PERPLEXITY_API_KEY environment variable.

Should I still use client.chat.completions with model="sonar"?

No. Sonar Chat Completions was retired on 27 September 2026. Use client.responses.create with a preset such as fast instead. See our Sonar migration guide.

Which preset should I use in Python?

Start with fast for quick lookups. Use low for everyday research, medium for multi-source browsing, high for expert-level research, and xhigh for agentic work with code execution.

How do I stream responses in Python?

Pass stream=True and loop over the events, printing event.delta when event.type is response.output_text.delta.

How do I get the sources?

Loop over response.output and look for items with type search_results. Each result has an id, title, url, snippet and date.

Is there an async client?

Yes. Import AsyncPerplexity and await client.responses.create(...), for example inside asyncio.gather to run several questions concurrently.

How much does it cost?

You pay for model tokens plus tool calls. Standard web search costs $2.50 per 1,000 invocations and fast search $1.00 per 1,000. See our API pricing guide and the official pricing page for current rates.

Official sources

  1. Agent API quickstart
  2. Agent API presets
  3. Streaming, structured output and background runs
  4. Web search tool

More API guides

Build

Run your first Python call in two minutes

Install the SDK, set your key, and paste lesson 2. Then grow it into the research CLI.