Scrape and crawl
Send a URL and get the page back as clean markdown, with a record of how it was fetched.
Endpoints
- POST
/v1/scrapeOne page as markdown, raw HTML, links, JSON, branding or a screenshot. - POST
/v1/crawlA whole site as an async job. Poll it, then read the pages. - POST
/v1/mapEvery URL on a site from its sitemap and links, without page bodies. - POST
/v1/batchMany URLs in one request, sync or async, up to your plan cap. - POST
/v1/screenshotA viewport or full-page capture as PNG or JPEG.
Try it
Example outputhttps://acme-plumbing.example
- Completeness
- 0.92
- Fetched by
- plain HTTP
- Read by
- JSON-LD
- Cost
- 1 credit
Example values. One completeness score covers the whole result.
Shows example output until you press Run. Demo runs are free, up to 10 every 5 minutes.
What it returns
POST /v1/scrape
| Field | What it is | Present |
|---|---|---|
| markdown | The page as markdown. When you set onlyMainContent, includeTags or excludeTags, this is the trimmed version. | Always |
| formats | The formats you asked for: markdown, rawHtml, links, json, branding, summary or screenshot. | When you send formats |
| metadata.fetchMethod | How the page was fetched: plain (HTTP), playwright (a real browser) or stealth (a proxy route). | Every HTML page scrape |
| metadata.cached | True when the result came from cache. metadata.cacheAgeMs says how old it is. | Every HTML page scrape |
| cascade | The fetch steps that ran, each with its level, result, latency and error class. level is null when no step ran (a cache hit or a PDF). | Always |
| listing._completeness | A 0 to 1 completeness score for the business data found on the page. One score for the whole listing. | When structured data is found |
| listing._extraction_method | How that data was read: json-ld, or regex-cascade for page text. | When structured data is found |
| partial | True when the content served is incomplete, with partial_reason (for example auth_wall). | Only when it happens |
| formats.screenshot.images | One base64 PNG per screenshot action, in order. | With screenshot actions and screenshot in formats |
| isPdf, text | For a PDF URL: isPdf is true, text and markdown hold the document, and there is no metadata object. The PDF page covers page counts and parse status. | When the URL is a PDF |
The X-Credits-Consumed response header says what the call cost.
Browser actions
Send actions to click, type, scroll or wait in a real browser before the page is captured. They run in order. Requests run fresh by default; if you set maxAge, a cached result for the same URL and actions can be returned instead. Requests with screenshot actions always run fresh.
| type | Fields | What it does |
|---|---|---|
| click | selector | Clicks the first element that matches. |
| write | selector, text | Fills a field with text. |
| press | key, selector (optional) | Presses a key, such as Enter, on the page or on an element. |
| scroll | selector (optional), ms (optional) | Scrolls to the bottom of the page, or brings an element into view, then waits ms. |
| wait | ms or selector | Waits for a number of milliseconds (default 1,000), or for an element to appear. |
| screenshot | selector (optional) | Captures the viewport or one element. Returned in formats.screenshot.images when formats includes screenshot. |
| executeJavascript | script | Coming soon. Not callable yet: a request that includes it is refused with 400 action_not_available before anything runs or is charged. |
pdf is an accepted action type but adds nothing to the /v1/scrape response today. minAge with screenshot actions returns 400 cache_not_applicable. waitFor is capped at 30,000 ms and shares one 30-second budget with the actions. An action that fails is skipped and the rest still run.
One request
Send your key as a Bearer token. Every parameter in these samples is one the endpoint reads today. The Example tab shows a trimmed response for a fictional business.
Every parameter in the docscurl -X POST https://api.superscraper.dev/v1/scrape \
-H "Authorization: Bearer $SUPERSCRAPER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"formats": [
"markdown",
"links"
],
"onlyMainContent": true
}'Pricing
Compare the plans- Scrape
- 1credit per page
- Crawl
- 1credit per page crawled
- Map
- 1credit per call
- Batch
- 1credit per URL
- Screenshot
- 1credit per call
- A scrape that returns no data costs 0 credits.
- A cache hit costs 0 credits, unless the same request asks for new work: a schema, a summary or a screenshot.
Pages per crawl
- free
- 50
- hobby
- 500
- pro
- 10,000
- scale
- 50,000
Per request, by plan.
Questions
Does it respect robots.txt?
Yes, by default. Set ignoreRobotsTxt to true only on a site you have permission to fetch.
Are results cached?
Only when you allow it. maxAge defaults to 0, so every call fetches the live page. Set maxAge in milliseconds to accept a cached copy, or zeroDataRetention to store nothing.
What happens when the URL is a PDF?
/v1/scrape returns the document text and markdown with isPdf set to true, and no metadata object. For page counts, chunks and a parse status, use /v1/parse (see the PDFs and documents page).
How do I crawl a site?
POST /v1/crawl returns a jobId. Poll GET /v1/crawl/:id for status, then read the pages from GET /v1/crawl/:id/results. includePaths, excludePaths, maxDepth and limit shape the crawl.
Coming soon
- MonitorsComing soon
- Interactive browser sessionsComing soon
Not callable yet. Beta items open as beta first.