PDFs and documents
Turn a PDF or DOCX into markdown, or into heading-aware chunks ready for your vector store.
Endpoints
- POST
/v1/parseA PDF or DOCX, uploaded as a file or linked by URL, as markdown, text or chunks.
Try it
Example outputhttps://acme-plumbing.example/price-list.pdf
- Completeness
- not scored
- Fetched by
- document URL
- Read by
- pdfjs
- Cost
- 1 credit
Example values. This call returns no completeness score.
Shows example output until you press Run. Demo runs are free, up to 10 every 5 minutes.
What it returns
POST /v1/parse
| Field | What it is | Present |
|---|---|---|
| status | ok, partial or failed. | Always |
| parser | Which parser read the file: pdfjs, docx or none. | Always |
| markdown, text | The document as markdown and as plain text. | On ok and partial |
| pages_total, pages_parsed | How many pages the document has and how many were read. | Always (null for DOCX) |
| truncated | True when the parser stopped early, with truncated_reason. | Always |
| pages_without_text | Pages with no text layer, with partial_reason. | When some pages have no text |
| chunks[] | Each chunk has text, heading, headingPath, headingLevel, charStart, charEnd and position. totalChunks counts them. | When format is rag |
| source, cached | upload or url, and whether the result came from cache. | Always |
Parse responses carry no completeness score. status and the page counts say what was read.
Limits today
- No OCR today. A scanned or image-only PDF fails with 422 no_text_layer.
- Up to 300 pages per document by default. More pages returns 413 page_limit, with max_pages in the body, before any page is parsed.
- Up to 25 MB per document by default. A larger upload returns 413 file_too_large; a larger download from a URL returns 413 download_too_large.
- PDF and DOCX only.
One request
Send your key as a Bearer token. Every parameter in these samples is one the endpoint reads today. The Example tab shows a trimmed response for a fictional business.
Every parameter in the docscurl -X POST https://api.superscraper.dev/v1/parse \
-H "Authorization: Bearer $SUPERSCRAPER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
"format": "rag",
"chunkSize": 1000
}'
# Or upload a file (multipart)
curl -X POST https://api.superscraper.dev/v1/parse \
-H "Authorization: Bearer $SUPERSCRAPER_API_KEY" \
-F "file=@report.pdf" \
-F "format=markdown"Pricing
Compare the plans- Parse
- 1credit per call
- One document per call, whatever the output format.
Questions
Can it read scanned PDFs?
Not today. A document with no text layer fails with 422 no_text_layer. In a mixed document, pages without text are listed in pages_without_text and the status is partial.
How are chunks split?
On headings first, then on paragraphs, sentences and spaces when a section runs past chunkSize. The default is 1,000 characters with 100 characters of overlap, and chunkSize accepts 100 to 20,000.
Can I upload a file instead of a URL?
Yes. Send multipart/form-data with the file in the file field. format, chunkSize and chunkOverlap work as form fields too.
What does a truncated result mean?
The parser stopped before the end, for example on a time or text budget. truncated_reason says which, and pages_parsed says how far it got.
Coming soon
- Large PDF jobsComing soon
- Tables and chartsBetaComing soon
Not callable yet. Beta items open as beta first.