Skip to main content
Live

PDFs and documents

Turn a PDF or DOCX into markdown, or into heading-aware chunks ready for your vector store.

Endpoints

  • POST/v1/parseA PDF or DOCX, uploaded as a file or linked by URL, as markdown, text or chunks.

Try it

TryA PDF or DOCX link to markdown. Scanned PDFs with no text layer are not read yet.

Example outputhttps://acme-plumbing.example/price-list.pdf

status ok · 4 of 4 pages parsed
# Acme Plumbing price list
## Water heaters
- Tank install, parts and labor
- Tankless install, parts and labor
## Drains
Completeness
not scored
Fetched by
document URL
Read by
pdfjs
Cost
1 credit

Example values. This call returns no completeness score.

Shows example output until you press Run. Demo runs are free, up to 10 every 5 minutes.

What it returns

POST /v1/parse

Response fields for POST /v1/parse
statusok, partial or failed.Always
parserWhich parser read the file: pdfjs, docx or none.Always
markdown, textThe document as markdown and as plain text.On ok and partial
pages_total, pages_parsedHow many pages the document has and how many were read.Always (null for DOCX)
truncatedTrue when the parser stopped early, with truncated_reason.Always
pages_without_textPages with no text layer, with partial_reason.When some pages have no text
chunks[]Each chunk has text, heading, headingPath, headingLevel, charStart, charEnd and position. totalChunks counts them.When format is rag
source, cachedupload or url, and whether the result came from cache.Always

Parse responses carry no completeness score. status and the page counts say what was read.

Limits today

  • No OCR today. A scanned or image-only PDF fails with 422 no_text_layer.
  • Up to 300 pages per document by default. More pages returns 413 page_limit, with max_pages in the body, before any page is parsed.
  • Up to 25 MB per document by default. A larger upload returns 413 file_too_large; a larger download from a URL returns 413 download_too_large.
  • PDF and DOCX only.

One request

Send your key as a Bearer token. Every parameter in these samples is one the endpoint reads today. The Example tab shows a trimmed response for a fictional business.

Every parameter in the docs
curl -X POST https://api.superscraper.dev/v1/parse \
  -H "Authorization: Bearer $SUPERSCRAPER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
    "format": "rag",
    "chunkSize": 1000
  }'

# Or upload a file (multipart)
curl -X POST https://api.superscraper.dev/v1/parse \
  -H "Authorization: Bearer $SUPERSCRAPER_API_KEY" \
  -F "file=@report.pdf" \
  -F "format=markdown"
Parse
1credit per call
  • One document per call, whatever the output format.

Questions

Can it read scanned PDFs?

Not today. A document with no text layer fails with 422 no_text_layer. In a mixed document, pages without text are listed in pages_without_text and the status is partial.

How are chunks split?

On headings first, then on paragraphs, sentences and spaces when a section runs past chunkSize. The default is 1,000 characters with 100 characters of overlap, and chunkSize accepts 100 to 20,000.

Can I upload a file instead of a URL?

Yes. Send multipart/form-data with the file in the file field. format, chunkSize and chunkOverlap work as form fields too.

What does a truncated result mean?

The parser stopped before the end, for example on a time or text budget. truncated_reason says which, and pages_parsed says how far it got.

Coming soon

  • Large PDF jobsComing soon
  • Tables and chartsBetaComing soon

Not callable yet. Beta items open as beta first.

Built with it