This tutorial uses the Agent API. Perplexity retired Sonar Chat Completions (client.chat.completions with model="sonar") on 27 September 2026. Updating old code? Read the Sonar migration guide.
Tutorial
Learn the Perplexity API in Python, step by step
You’ll need a recent version of Python 3 and a Perplexity API key with some credit.
Set up your project
Create a virtual environment, install the official perplexityai package, and keep your key in a .env file. Get a key from the API console first; our API key guide shows how to keep it safe.
# 1. Create and activate a virtual environment python -m venv .venv source .venv/bin/activate # Windows: .venv\Scripts\activate # 2. Install the official SDK (and dotenv for local keys) pip install perplexityai python-dotenv
# .env (add this file to .gitignore) PERPLEXITY_API_KEY=your_api_key_here
Make your first request
The Agent API uses client.responses.create. Give it a preset and your question, then print output_text.
from dotenv import load_dotenv
from perplexity import Perplexity
load_dotenv() # loads PERPLEXITY_API_KEY from .env
client = Perplexity() # reads the key from the environment
response = client.responses.create(
preset="fast",
input="What are the latest developments in AI agents?",
)
print(response.output_text)Choose a preset or a model
Presets bundle a model, tools and limits, and Perplexity keeps them up to date. You can also pass a specific model instead.
# Use a preset (recommended): Perplexity picks the model and tools
response = client.responses.create(preset="low", input="Compare solar and wind costs in 2026")
# Or choose a specific model yourself
response = client.responses.create(
model="openai/gpt-5.6-sol",
input="Explain supervised learning in simple terms",
)
print(response.output_text)| Preset | Default model | Tools | Best for |
|---|---|---|---|
| fast | openai/gpt-6-luna | web_search | Single facts, definitions, quick summaries |
| low | openai/gpt-6-luna | web_search, fetch_url | Everyday research, light multi-step lookups |
| medium | openai/gpt-6-luna | web_search, fetch_url | Multi-hop browsing across many sources |
| high | openai/gpt-6-sol | web_search, fetch_url | Expert-level, exhaustive research |
| xhigh | anthropic/claude-opus-5-5 | web_search, finance_search, sandbox | Open-ended agentic work with code execution |
Presets update automatically when Perplexity improves them. Source: Agent API presets.
Add instructions and hold a conversation
Use instructions for a system prompt. For multi-turn chat, pass a list of messages as input and append each answer.
response = client.responses.create(
preset="fast",
instructions="You are a friendly tutor. Answer in under 120 words, in simple English.",
input="How do vaccines train the immune system?",
)
print(response.output_text)from perplexity import Perplexity
client = Perplexity()
history = []
while True:
question = input("You: ").strip()
if question.lower() in {"quit", "exit"}:
break
history.append({"role": "user", "content": question})
response = client.responses.create(preset="fast", input=history)
answer = response.output_text
print(f"Perplexity: {answer}\n")
history.append({"role": "assistant", "content": answer})Stream answers as they’re written
Set stream=True and print each response.output_text.delta event. Great for chat UIs and long answers.
stream = client.responses.create(
preset="fast",
input="Summarize today's top technology news in 5 bullets",
stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.completed":
print(f"\n\nUsage: {event.response.usage}")Read the sources and citations
Numbers like [1] in output_text match the id of results in the search_results output items. Show them so users can check the answer.
response = client.responses.create(
preset="fast",
input="How much of the Amazon rainforest has been lost?",
)
print(response.output_text, "\n")
# Sources arrive as output items of type "search_results"
for item in response.output:
if item.type == "search_results":
for result in item.results:
print(f"[{result.id}] {result.title}")
print(f" {result.url} ({result.date})")Filter the web search
Pass a web_search tool with filters to limit domains (prefix - to exclude) and recency.
response = client.responses.create(
preset="fast",
input="Latest renewable energy policy updates",
tools=[
{
"type": "web_search",
"filters": {
"search_domain_filter": ["iea.org", "energy.gov"],
"search_recency_filter": "month",
},
}
],
)
print(response.output_text)Get structured JSON output
Use response_format with a JSON schema, then parse output_text with json.loads.
import json
response = client.responses.create(
preset="low",
input="List the 3 largest countries by area with their area in km2",
response_format={
"type": "json_schema",
"json_schema": {
"name": "countries",
"schema": {
"type": "object",
"properties": {
"countries": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"area_km2": {"type": "number"},
},
"required": ["name", "area_km2"],
},
}
},
"required": ["countries"],
},
},
},
)
data = json.loads(response.output_text)
for c in data["countries"]:
print(c["name"], c["area_km2"])Run long research in the background
For the high preset or big reports, set background=True and poll with client.responses.retrieve. The job keeps running even if your script disconnects.
import time
response = client.responses.create(
preset="high",
input="Write a detailed report on battery recycling technologies, with sources",
background=True,
)
while response.status in ("queued", "in_progress"):
time.sleep(2)
response = client.responses.retrieve(response.id)
print(response.status)
print(response.output_text)Get raw results with the Search API
Need links instead of an answer? The Search API returns ranked results for $5 per 1,000 requests ($1 in fast mode). Full guide: Search API tutorial.
search = client.search.create(
query="EV battery recycling 2026",
max_results=5,
search_recency_filter="month",
)
for result in search.results:
print(f"{result.title}: {result.url}")Run many questions at once (async)
Use AsyncPerplexity with asyncio.gather to answer several questions concurrently. Mind your rate limits.
import asyncio
from perplexity import AsyncPerplexity
client = AsyncPerplexity()
async def ask(question: str) -> str:
response = await client.responses.create(preset="fast", input=question)
return response.output_text
async def main():
questions = [
"What is the capital of Australia?",
"Who discovered penicillin?",
"How tall is Mount Everest?",
]
answers = await asyncio.gather(*(ask(q) for q in questions))
for q, a in zip(questions, answers):
print(f"Q: {q}\nA: {a}\n")
asyncio.run(main())Handle errors and retry
Retry rate limits (429) and server errors (5xx) with exponential backoff and jitter. Don’t retry 400 or 401; fix the request or key instead.
import random
import time
def ask_with_retry(question: str, attempts: int = 5) -> str:
for attempt in range(attempts):
try:
response = client.responses.create(preset="fast", input=question)
return response.output_text
except Exception as error:
status = getattr(error, "status_code", None)
retryable = status == 429 or (status is not None and status >= 500)
if not retryable or attempt == attempts - 1:
raise
wait = (2 ** attempt) + random.random()
print(f"Got {status}, retrying in {wait:.1f}s...")
time.sleep(wait)Project
Build a cited research CLI
One script that puts the lessons together: it streams an answer to your question, applies optional domain and recency filters, then lists the sources.
"""research.py: ask a question, stream a cited answer, list the sources.
Usage: python research.py "your question" [--domains iea.org,nature.com] [--recent week]
"""
import argparse
from dotenv import load_dotenv
from perplexity import Perplexity
def main() -> None:
parser = argparse.ArgumentParser(description="Cited research from the terminal")
parser.add_argument("question")
parser.add_argument("--preset", default="low", choices=["fast", "low", "medium", "high"])
parser.add_argument("--domains", default="", help="comma-separated, prefix - to exclude")
parser.add_argument("--recent", default="", choices=["", "day", "week", "month", "year"])
args = parser.parse_args()
load_dotenv()
client = Perplexity()
filters = {}
if args.domains:
filters["search_domain_filter"] = [d.strip() for d in args.domains.split(",") if d.strip()]
if args.recent:
filters["search_recency_filter"] = args.recent
tool = {"type": "web_search"}
if filters:
tool["filters"] = filters
stream = client.responses.create(
preset=args.preset,
input=args.question,
tools=[tool],
stream=True,
)
final = None
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.completed":
final = event.response
print("\n\nSources:")
for item in (final.output if final else []):
if item.type == "search_results":
for r in item.results:
print(f" [{r.id}] {r.title} - {r.url}")
if __name__ == "__main__":
main()Python code generator
Pick your options and copy ready-to-run Python.
Reference
Python cheat sheet
| Task | Code |
|---|---|
| Install | pip install perplexityai |
| Create client | client = Perplexity() |
| Ask a question | client.responses.create(preset="fast", input="…") |
| Read answer | response.output_text |
| System prompt | instructions="…" |
| Stream | stream=True → event.delta |
| Sources | item.type == "search_results" → item.results |
| Filters | tools=[{"type": "web_search", "filters": {…}}] |
| JSON output | response_format={"type": "json_schema", …} |
| Background | background=True + responses.retrieve(id) |
| Raw results | client.search.create(query="…") |
| Async | AsyncPerplexity() + await |
FAQ
Perplexity Python questions
How do I install the Perplexity Python SDK?
Run pip install perplexityai, then import it with from perplexity import Perplexity. The client reads your key from the PERPLEXITY_API_KEY environment variable.
Should I still use client.chat.completions with model="sonar"?
No. Sonar Chat Completions was retired on 27 September 2026. Use client.responses.create with a preset such as fast instead. See our Sonar migration guide.
Which preset should I use in Python?
Start with fast for quick lookups. Use low for everyday research, medium for multi-source browsing, high for expert-level research, and xhigh for agentic work with code execution.
How do I stream responses in Python?
Pass stream=True and loop over the events, printing event.delta when event.type is response.output_text.delta.
How do I get the sources?
Loop over response.output and look for items with type search_results. Each result has an id, title, url, snippet and date.
Is there an async client?
Yes. Import AsyncPerplexity and await client.responses.create(...), for example inside asyncio.gather to run several questions concurrently.
How much does it cost?
You pay for model tokens plus tool calls. Standard web search costs $2.50 per 1,000 invocations and fast search $1.00 per 1,000. See our API pricing guide and the official pricing page for current rates.
Official sources
- Agent API quickstart
- Agent API presets
- Streaming, structured output and background runs
- Web search tool
More API guides
Build
Run your first Python call in two minutes
Install the SDK, set your key, and paste lesson 2. Then grow it into the research CLI.