Skip to main content
Live

Scrape and crawl

Send a URL and get the page back as clean markdown, with a record of how it was fetched.

Endpoints

  • POST/v1/scrapeOne page as markdown, raw HTML, links, JSON, branding or a screenshot.
  • POST/v1/crawlA whole site as an async job. Poll it, then read the pages.
  • POST/v1/mapEvery URL on a site from its sitemap and links, without page bodies.
  • POST/v1/batchMany URLs in one request, sync or async, up to your plan cap.
  • POST/v1/screenshotA viewport or full-page capture as PNG or JPEG.

Try it

TryAny page as clean markdown a model can read.

Example outputhttps://acme-plumbing.example

# Acme Plumbing
24/7 emergency plumbing in Austin, TX.
## Services
- Water heater repair and install
- Drain cleaning and leak detection
[Call now](tel:+15125550148)
Completeness
0.92
Fetched by
plain HTTP
Read by
JSON-LD
Cost
1 credit

Example values. One completeness score covers the whole result.

Shows example output until you press Run. Demo runs are free, up to 10 every 5 minutes.

What it returns

POST /v1/scrape

Response fields for POST /v1/scrape
markdownThe page as markdown. When you set onlyMainContent, includeTags or excludeTags, this is the trimmed version.Always
formatsThe formats you asked for: markdown, rawHtml, links, json, branding, summary or screenshot.When you send formats
metadata.fetchMethodHow the page was fetched: plain (HTTP), playwright (a real browser) or stealth (a proxy route).Every HTML page scrape
metadata.cachedTrue when the result came from cache. metadata.cacheAgeMs says how old it is.Every HTML page scrape
cascadeThe fetch steps that ran, each with its level, result, latency and error class. level is null when no step ran (a cache hit or a PDF).Always
listing._completenessA 0 to 1 completeness score for the business data found on the page. One score for the whole listing.When structured data is found
listing._extraction_methodHow that data was read: json-ld, or regex-cascade for page text.When structured data is found
partialTrue when the content served is incomplete, with partial_reason (for example auth_wall).Only when it happens
formats.screenshot.imagesOne base64 PNG per screenshot action, in order.With screenshot actions and screenshot in formats
isPdf, textFor a PDF URL: isPdf is true, text and markdown hold the document, and there is no metadata object. The PDF page covers page counts and parse status.When the URL is a PDF

The X-Credits-Consumed response header says what the call cost.

Browser actions

Send actions to click, type, scroll or wait in a real browser before the page is captured. They run in order. Requests run fresh by default; if you set maxAge, a cached result for the same URL and actions can be returned instead. Requests with screenshot actions always run fresh.

Browser actions
clickselectorClicks the first element that matches.
writeselector, textFills a field with text.
presskey, selector (optional)Presses a key, such as Enter, on the page or on an element.
scrollselector (optional), ms (optional)Scrolls to the bottom of the page, or brings an element into view, then waits ms.
waitms or selectorWaits for a number of milliseconds (default 1,000), or for an element to appear.
screenshotselector (optional)Captures the viewport or one element. Returned in formats.screenshot.images when formats includes screenshot.
executeJavascriptscriptComing soon. Not callable yet: a request that includes it is refused with 400 action_not_available before anything runs or is charged.

pdf is an accepted action type but adds nothing to the /v1/scrape response today. minAge with screenshot actions returns 400 cache_not_applicable. waitFor is capped at 30,000 ms and shares one 30-second budget with the actions. An action that fails is skipped and the rest still run.

One request

Send your key as a Bearer token. Every parameter in these samples is one the endpoint reads today. The Example tab shows a trimmed response for a fictional business.

Every parameter in the docs
curl -X POST https://api.superscraper.dev/v1/scrape \
  -H "Authorization: Bearer $SUPERSCRAPER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "formats": [
      "markdown",
      "links"
    ],
    "onlyMainContent": true
  }'
Scrape
1credit per page
Crawl
1credit per page crawled
Map
1credit per call
Batch
1credit per URL
Screenshot
1credit per call
  • A scrape that returns no data costs 0 credits.
  • A cache hit costs 0 credits, unless the same request asks for new work: a schema, a summary or a screenshot.

Pages per crawl

free
50
hobby
500
pro
10,000
scale
50,000

Per request, by plan.

Questions

Does it respect robots.txt?

Yes, by default. Set ignoreRobotsTxt to true only on a site you have permission to fetch.

Are results cached?

Only when you allow it. maxAge defaults to 0, so every call fetches the live page. Set maxAge in milliseconds to accept a cached copy, or zeroDataRetention to store nothing.

What happens when the URL is a PDF?

/v1/scrape returns the document text and markdown with isPdf set to true, and no metadata object. For page counts, chunks and a parse status, use /v1/parse (see the PDFs and documents page).

How do I crawl a site?

POST /v1/crawl returns a jobId. Poll GET /v1/crawl/:id for status, then read the pages from GET /v1/crawl/:id/results. includePaths, excludePaths, maxDepth and limit shape the crawl.

Coming soon

  • MonitorsComing soon
  • Interactive browser sessionsComing soon

Not callable yet. Beta items open as beta first.

Built with it