x402AgentTools

✂️Text Chunker

Split text into retrieval-friendly chunks with a token budget and sentence-aware boundaries. Overlap between consecutive chunks preserves context across splits — the standard preparation step before embedding.

Worked examples

Small doc, 200-token chunks

GET /api/v1/dev/text-chunker?text=Retrieval-augmented%20generation%20improves%20factuality.%20First%2C%20documents%20are%20split%20into%20chunks.%20Second%2C%20chunks%20are%20embedded%20into%20vectors.%20Third%2C%20the%20agent%20retrieves%20the%20top%20matches%20for%20a%20query.%20Overlap%20keeps%20ideas%20that%20straddle%20boundaries%20intact.%20Finally%2C%20the%20model%20answers%20with%20citations.&targetTokens=60&overlapTokens=15

Result: 2 chunks

Default 500-token chunks

GET /api/v1/dev/text-chunker?text=The%20grass%20is%20green.%20The%20sky%20is%20blue.%20Water%20is%20wet.%20Fire%20is%20hot.%20Snow%20is%20cold.%20Rain%20falls%20down.%20Sun%20rises%20east.%20Moon%20glows%20night.%20Stars%20twinkle%20far.%20Wind%20blows%20soft.

Result: 1 chunk

Machine API (x402)

$0.002 / call

This tool is also a JSON API for AI agents. Requests without payment receive 402 Payment Required plus instructions; agents pay USDC on Base via the x402 protocol — no accounts, no API keys.

GET /api/v1/dev/text-chunker?text=Retrieval-augmented%20generation%20improves%20factuality.%20First%2C%20documents%20are%20split%20into%20chunks.%20Second%2C%20chunks%20are%20embedded%20into%20vectors.%20Third%2C%20the%20agent%20retrieves%20the%20top%20matches%20for%20a%20query.%20Overlap%20keeps%20ideas%20that%20straddle%20boundaries%20intact.%20Finally%2C%20the%20model%20answers%20with%20citations.&targetTokens=60&overlapTokens=15 HTTP/1.1
Host: agenttools-hub.vercel.app

→ 402 (payment required, instructions in headers)
→ 200 (after X-PAYMENT header; JSON body below)

{
  "tool": "dev/text-chunker",
  "input": {"text":"Retrieval-augmented generation improves factuality. First, documents are split into chunks. Second, chunks are embedded into vectors. Third, the agent retrieves the top matches for a query. Overlap keeps ideas that straddle boundaries intact. Finally, the model answers with citations.","targetTokens":60,"overlapTokens":15},
  "result": { "value": 2, "answer": "2 chunks" }
}

Agent docs: /llms.txt · OpenAPI spec · integration guide

About this tool

Chunking is the first step of every RAG pipeline, and naive character splits shred meaning. This chunker packs sentence units up to a token budget (estimated with the same heuristic as the token estimator), carries configurable overlap across boundaries, and returns machine-readable chunks in `result.artifacts.chunks`.

chunks = pack(sentences, budget = targetTokens, carry = overlapTokens)

Frequently asked questions

How big should chunks be?

For most embedding models, 300–800 tokens with 10–20% overlap works well. Small chunks retrieve precisely; large chunks preserve context — tune by eval.

Where are the full chunks?

In `result.artifacts.chunks` as a JSON array of {index, estTokens, text}. The web page shows only previews.

Related tools