x402AgentTools

๐Ÿ“„PDF to Text

Fetch a PDF (by URL, SSRF-guarded) or accept it as base64 and extract all readable text with page count. Deterministic parsing โ€” no language model, so nothing is invented. Text-based PDFs only; scanned images need OCR (not included).

๐Ÿ”Œ Server-side tool

This tool runs on the server (it fetches third-party URLs directly and cannot run in your browser). Use the paid JSON API below โ€” the worked examples show real server-side results and refresh automatically.

Worked examples

Live examples refresh automatically โ€” call the API for guaranteed-fresh results.

Machine API (x402)

$0.003 / call

This tool is also a JSON API for AI agents. Requests without payment receive 402 Payment Required plus instructions; agents pay USDC on Base via the x402 protocol โ€” no accounts, no API keys.

GET /api/v1/dev/pdf-to-text?url=https%3A%2F%2Fwww.w3.org%2FWAI%2FER%2Ftests%2Fxhtml%2Ftestfiles%2Fresources%2Fpdf%2Fdummy.pdf HTTP/1.1
Host: agenttools-hub.vercel.app

โ†’ 402 (payment required, instructions in headers)
โ†’ 200 (after X-PAYMENT header; JSON body below)

{
  "tool": "dev/pdf-to-text",
  "input": {"url":"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"},
  "result": { "value": null, "answer": "" }
}

Agent docs: /llms.txt ยท OpenAPI spec ยท integration guide

About this tool

Language models read text, not PDFs. This tool extracts the full text layer from any PDF โ€” by URL or base64 โ€” with page counts and hard size caps, server-side and SSRF-guarded. Deterministic extraction: the output is exactly what the document contains, nothing invented.

Frequently asked questions

Does it work on scanned documents?

No โ€” scanned PDFs are images and need OCR. The tool detects this and returns a clear error instead of empty text.

What are the limits?

10MB per file, 150 pages, and a configurable text cap (default 50k characters, extendable to 200k). Fetches are SSRF-guarded.

Related tools