✂️Text Chunker
Split text into retrieval-friendly chunks with a token budget and sentence-aware boundaries. Overlap between consecutive chunks preserves context across splits — the standard preparation step before embedding.
Worked examples
Small doc, 200-token chunks
GET /api/v1/dev/text-chunker?text=Retrieval-augmented%20generation%20improves%20factuality.%20First%2C%20documents%20are%20split%20into%20chunks.%20Second%2C%20chunks%20are%20embedded%20into%20vectors.%20Third%2C%20the%20agent%20retrieves%20the%20top%20matches%20for%20a%20query.%20Overlap%20keeps%20ideas%20that%20straddle%20boundaries%20intact.%20Finally%2C%20the%20model%20answers%20with%20citations.&targetTokens=60&overlapTokens=15
Result: 2 chunks
Default 500-token chunks
GET /api/v1/dev/text-chunker?text=The%20grass%20is%20green.%20The%20sky%20is%20blue.%20Water%20is%20wet.%20Fire%20is%20hot.%20Snow%20is%20cold.%20Rain%20falls%20down.%20Sun%20rises%20east.%20Moon%20glows%20night.%20Stars%20twinkle%20far.%20Wind%20blows%20soft.
Result: 1 chunk
Machine API (x402)
$0.002 / callThis tool is also a JSON API for AI agents. Requests without payment receive 402 Payment Required plus instructions; agents pay USDC on Base via the x402 protocol — no accounts, no API keys.
GET /api/v1/dev/text-chunker?text=Retrieval-augmented%20generation%20improves%20factuality.%20First%2C%20documents%20are%20split%20into%20chunks.%20Second%2C%20chunks%20are%20embedded%20into%20vectors.%20Third%2C%20the%20agent%20retrieves%20the%20top%20matches%20for%20a%20query.%20Overlap%20keeps%20ideas%20that%20straddle%20boundaries%20intact.%20Finally%2C%20the%20model%20answers%20with%20citations.&targetTokens=60&overlapTokens=15 HTTP/1.1
Host: agenttools-hub.vercel.app
→ 402 (payment required, instructions in headers)
→ 200 (after X-PAYMENT header; JSON body below)
{
"tool": "dev/text-chunker",
"input": {"text":"Retrieval-augmented generation improves factuality. First, documents are split into chunks. Second, chunks are embedded into vectors. Third, the agent retrieves the top matches for a query. Overlap keeps ideas that straddle boundaries intact. Finally, the model answers with citations.","targetTokens":60,"overlapTokens":15},
"result": { "value": 2, "answer": "2 chunks" }
}Agent docs: /llms.txt · OpenAPI spec · integration guide
About this tool
Chunking is the first step of every RAG pipeline, and naive character splits shred meaning. This chunker packs sentence units up to a token budget (estimated with the same heuristic as the token estimator), carries configurable overlap across boundaries, and returns machine-readable chunks in `result.artifacts.chunks`.
chunks = pack(sentences, budget = targetTokens, carry = overlapTokens)
Frequently asked questions
How big should chunks be?
For most embedding models, 300–800 tokens with 10–20% overlap works well. Small chunks retrieve precisely; large chunks preserve context — tune by eval.
Where are the full chunks?
In `result.artifacts.chunks` as a JSON array of {index, estTokens, text}. The web page shows only previews.