x402AgentTools

🔢Token Estimator

Estimate how many tokens a text consumes for each major model family (GPT, Claude, Gemini, Llama, DeepSeek) using a character/word heuristic that is within ~10% of real BPE tokenizers — enough for context budgeting and cost math, with zero dependencies.

Worked examples

One paragraph

GET /api/v1/dev/token-estimator?text=The%20quick%20brown%20fox%20jumps%20over%20the%20lazy%20dog.%20Pack%20my%20box%20with%20five%20dozen%20liquor%20jugs.

Result: 23 tokens (GPT family)

A code snippet

GET /api/v1/dev/token-estimator?text=function%20sum(a%2C%20b)%20%7B%20return%20a%20%2B%20b%3B%20%7D%20%2F%2F%20adds%20two%20numbers

Result: 17 tokens (GPT family)

CJK text

GET /api/v1/dev/token-estimator?text=%E4%BA%BA%E5%B7%A5%E7%9F%A5%E8%83%BD%E3%81%AF%E4%B8%96%E7%95%8C%E3%82%92%E5%A4%89%E3%81%88%E3%82%8B

Result: 13 tokens (GPT family)

Machine API (x402)

$0.001 / call

This tool is also a JSON API for AI agents. Requests without payment receive 402 Payment Required plus instructions; agents pay USDC on Base via the x402 protocol — no accounts, no API keys.

GET /api/v1/dev/token-estimator?text=The%20quick%20brown%20fox%20jumps%20over%20the%20lazy%20dog.%20Pack%20my%20box%20with%20five%20dozen%20liquor%20jugs. HTTP/1.1
Host: agenttools-hub.vercel.app

→ 402 (payment required, instructions in headers)
→ 200 (after X-PAYMENT header; JSON body below)

{
  "tool": "dev/token-estimator",
  "input": {"text":"The quick brown fox jumps over the lazy dog. Pack my box with five dozen liquor jugs."},
  "result": { "value": 23, "answer": "23 tokens (GPT family)" }
}

Agent docs: /llms.txt · OpenAPI spec · integration guide

About this tool

Context windows and API bills are denominated in tokens, but agents rarely have a tokenizer on hand. This estimator returns per-model-family counts within about 10% of real BPE tokenizers — the precision budgeting and cost math actually need, with no dependencies.

tokens ≈ CJK × 1.1 + blend(other chars ÷ 3.7, words × 1.32), then × per-model factor

Frequently asked questions

How accurate is the estimate?

Within ±10% for typical prose and code. Tokenizers differ per model (o200k, Claude, Gemini SentencePiece), so exact counts require the vendor's tokenizer — this tool trades that precision for speed and zero dependencies.

Why do Claude/Gemini counts differ from GPT?

Different tokenizers with different vocabularies. Empirically Claude counts run ~6% above o200k and Gemini ~12% for English text — the multipliers encode that.

Related tools