What it is
/contents returns clean, structured content from any URL, handling JavaScript-rendered pages, PDFs, and complex layouts automatically. You pass in URLs and get back full page text, targeted highlights, LLM-generated summaries, or all three. It can also be used to crawl linked subpages to pull content from entire site sections in a single request. All contents features are also available in /search for returned URLs, at no extra charge up to 10 results per search ($1/1000 pages afterwards). We recommend using /search in this way instead of /contents for web search tool use cases.
Key Capabilities
Content Modes
Choose how you receive content, or combine them in a single request:
| Mode | What You Get | Best For |
|---|---|---|
| Text | Full page content as clean markdown | Deep analysis, full context research |
| Highlights | Key excerpts relevant to your query | Agent workflows, factual lookups (10x fewer tokens) |
| Summary | LLM-generated abstract | Quick overviews, structured extraction with JSON schema |
Subpage crawling
Automatically discover and extract content from linked pages within a site. Pass subpages: 10 and optionally subpageTarget: ["docs", "about"] to focus on relevant sections.
Content freshness
Control whether results come from cache or are freshly crawled with maxAgeHours:
| Setting | Behavior |
|---|---|
| Omit (default) | Livecrawl only when no cache exists |
24 |
Use cache if < 24 hours old, otherwise livecrawl |
0 |
Always livecrawl (slowest, freshest) |
-1 |
Cache only (fastest, may be stale) |
Common use cases
Token-efficient info from an article
Get the most relevant excerpts without needing the full page.
result = exa.get_contents(
["https://example.com/research-paper"],
highlights={"query": "methodology and results"}
)
Structured outputs using summaries
Extract specific fields from any page using a JSON schema.
result = exa.get_contents(
["https://example.com/company-page"],
summary={
"query": "Extract company information",
"schema": {
"type": "object",
"properties": {
"name": {"type": "string"},
"industry": {"type": "string"},
"founded": {"type": "number"}
},
"required": ["name", "industry"]
}
}
)
Crawl many pages from a website
Pull content from a docs site, targeting specific sections.
result = exa.get_contents(
["https://docs.example.com"],
subpages=15,
subpage_target=["api", "models", "embeddings"],
max_age_hours=24,
text={"max_characters": 5000}
)
Human Quickstart
Get your API key from the Exa Dashboard.
- Python
pip install exa-py
from exa_py import Exa
exa = Exa(api_key="your-api-key")
result = exa.get_contents(
["https://example.com/article"],
highlights=True
)
- JavaScript
npm install exa-js
import Exa from "exa-js";
const exa = new Exa("your-api-key");
const result = await exa.getContents(
["https://example.com/article"],
{
highlights: true
}
);
- cURL
curl -X POST "https://api.exa.ai/contents" \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_API_KEY" \
-d '{
"urls": ["https://example.com/article"],
"highlights": true
}'