Skip to content

Runs entirely in your browser. Nothing you paste leaves this page.

Free / No sign-up

LLM token counter and estimator.

Paste a prompt or document to get an approximate token count and see how much of a context window it would use. A quick, free heuristic that runs in your browser, never an exact count.

Estimated tokens

~81

Likely range 79 to 84 tokens

Characters
334
Words
59
Lines
3
Context window to budget against

Up to <0.1% of a 128K window. Leave room for the system prompt, history and the reply, which all count too.

Heuristic estimate, not an exact count: about 4 characters or 0.75 words per token for English text. Each model's tokenizer differs, and code usually needs more tokens. Learn more about tokens and context windows.

How to use it.

  1. 01

    Paste a prompt, a document or a chat transcript.

  2. 02

    Read the estimated token count and its likely range.

  3. 03

    Pick a context window size to see how much room is left for history and the reply.

What it does.

Everything this tool handles, all of it inside your browser tab.

  • Instant token estimate as you type or paste
  • Likely range from two rules of thumb, with the midpoint shown
  • Character, word and line counts
  • Context window presets: 8K, 32K, 128K, 200K and 1M
  • Custom window size in tokens
  • Meter showing the share of the window used, with warning colours
  • Warning when much of the text is in a non-Latin script
  • Runs locally: nothing is sent to any AI model

Worked examples.

  • How many tokens is a sentence?

    The quick brown fox jumps over the lazy dog.
    
    44 characters / 4    = 11 tokens
    9 words / 0.75       = 12 tokens
    > Estimated ~12 tokens (likely range 11 to 12)

    Short, common English words tokenise efficiently, so the two rules agree closely here.

  • How many tokens is 1,000 words?

    1,000 words of English prose
    
    by words: 1,000 / 0.75    = about 1,333 tokens
    by chars: 5,500 / 4       = about 1,375 tokens
    > Roughly 1,300 to 1,400 tokens

    The character figure assumes about 5.5 characters per word including spaces. Technical writing with long terms or code runs higher.

  • Will this document fit in a 128K context window?

    Window             128,000
    - system prompt     -1,500
    - tool definitions  -2,000
    - chat history      -6,000
    - reserved reply    -4,000
    = room for documents 114,500 tokens
      (about 86,000 words of English)

    Budget for everything that shares the window, not only the document. The word figure uses the 0.75 words per token rule.

  • Estimate the cost of an API call

    input  20,000 tokens x $3.00 per million  = $0.060
    output  1,000 tokens x $15.00 per million = $0.015
    > about $0.075 per request
    
    (example prices only: use your provider's current rates)

    Input and output are priced separately, and output usually costs more per token, so long replies add up quickly.

What is a token in an LLM?

Large language models do not read characters or words. A tokenizer first splits text into tokens: pieces from a fixed vocabulary that the model learned to work with. Most current tokenizers use byte-pair encoding (BPE) or a close relative, with vocabularies of tens of thousands to a few hundred thousand entries.

Common English words are usually one token, often including the space before them. Rare words, names and typos are split into several pieces. Numbers are cut into short chunks, punctuation and indentation take tokens of their own, and text in scripts that were rarer in the tokenizer's training data is broken into many small pieces.

Tokens matter because everything is measured in them: the context window a model can see, the maximum length of its reply, rate limits and the price of each request.

How this token estimator works

Each model family has its own tokenizer, so only that tokenizer gives an exact count. The estimator uses two widely quoted rules of thumb for English instead: about 4 characters per token, and about 0.75 words per token. It applies both, shows the range between them, and uses the midpoint as the headline figure.

Alongside the estimate you get character, word and line counts. If more than about a fifth of the text is in a non-Latin script, the tool warns that the real count is likely to be higher than the estimate. Everything is simple arithmetic in your browser, so it is instant even for long documents, and nothing is sent to any model or server.

Use the estimate for planning questions: will this document fit, how much room is left for the conversation, roughly what will a batch of requests cost. Do not use it to enforce hard limits in production code, where an error of ten percent can mean a rejected request; count with the real tokenizer there instead.

Why token counts differ between models

The same paragraph can produce noticeably different token counts in different model families, because their vocabularies differ. A newer tokenizer with a larger vocabulary often needs fewer tokens for the same text, especially for code and non-English languages.

Content type matters as much as model choice. Source code, JSON, tables, URLs and long numbers need more tokens per character than prose. Hindi, Arabic, Japanese, Thai and many other languages typically need more tokens than English for the same meaning, which means less room in the window and a higher bill.

Requests also carry hidden tokens: chat formatting adds a few per message, tool and function definitions are sent with every call, and images, audio and files are converted into tokens of their own. This estimator counts only the text you paste.

How to count tokens exactly

For an exact figure, use the tokenizer that belongs to your model. OpenAI publishes the open-source tiktoken library, which can look up the right encoding from a model name. Anthropic and Google offer token-counting endpoints in their APIs for their models, and open-weight models ship their tokenizer with the weights, usually loadable through Hugging Face's tokenizers or transformers libraries.

The simplest source of truth is the response itself: most LLM APIs return a usage object with the input and output token counts for each request. Log it, and you can compare estimates with reality over time.

Budgeting a context window

The context window holds everything a model sees in one request: the system prompt, tool definitions, retrieved documents, the conversation so far and the reply it writes. If your input already fills most of it, the model has little room to answer, and many chat applications quietly drop older messages to make space.

Models also have a separate limit on output tokens, often much smaller than the window, and reasoning models may spend part of the budget thinking before they answer. Reserve room for the reply before you decide how much context to send.

The presets on this page (8K, 32K, 128K, 200K and 1M tokens) are common window sizes, not claims about any particular model. Check your provider's documentation for the exact limit, or type a custom size.

Tokens, words and pages: quick conversions

For English prose, these rough conversions are handy for planning. 100 tokens is about 75 words. 1,000 words is about 1,300 to 1,400 tokens. A dense page of around 500 words is about 650 to 700 tokens. A 128K-token window therefore holds very roughly 96,000 words of plain English, before you leave any room for instructions or the reply.

Treat these as planning figures for English only. Code, tables, JSON and most non-English text use more tokens for the same length, and the only reliable count comes from the model's own tokenizer. When the difference matters, for example near a limit or when estimating cost at scale, measure a representative sample with the real tokenizer and use the ratio you observe.

How to reduce token usage and cost

Trim the system prompt to the instructions that change behaviour. Summarise or drop old conversation turns instead of resending the whole history. Ask for concise, structured output when you do not need prose, since output tokens are usually priced higher than input tokens.

When documents do not fit, retrieval-augmented generation sends only the passages relevant to the question instead of everything. Several providers also offer prompt caching, which makes a repeated prefix such as a long system prompt cheaper and faster on later calls. Our comparison of RAG vs fine-tuning covers when each approach fits.

Questions, answered

Something else on your mind? Ask a consultant and get a reply within one business day.

Is this an exact token count?

No. It is a heuristic based on characters and words. For an exact figure, use the tokenizer or token-counting API of the model you are calling.

How many tokens is one word?

In English, about 1.3 tokens per word on average, or roughly 0.75 words per token. Common short words are often a single token; long, rare or technical words take several.

How many characters is a token?

About 4 characters of English text, including spaces. Code, numbers and non-Latin scripts usually have fewer characters per token, so they cost more.

Why do non-English languages use more tokens?

Tokenizers learn their vocabulary mostly from English-heavy text, so other scripts are split into more, smaller pieces. The same meaning can cost noticeably more tokens, which also raises API cost.

Do spaces and punctuation count as tokens?

Yes, they are part of the text the model reads. A space is usually merged into the following word's token, while punctuation, line breaks and indentation often take tokens of their own.

Does the system prompt count toward the token limit?

Yes. The system prompt, tool definitions, conversation history and retrieved documents all count as input tokens and share the context window with the reply.

Do output tokens count against the context window?

Yes. The window covers the input and the model's reply together, and reasoning models may also spend tokens on internal reasoning before they answer.

What happens if my prompt is longer than the context window?

Most APIs reject the request with an error. Chat apps usually drop or summarise older messages instead, which is why long conversations can forget early details.

How do I get an exact count for my model?

Use the model's own tokenizer, such as OpenAI's tiktoken library, or the token-counting endpoint your provider offers. Every API response also reports the tokens actually used.

Is my text sent to an AI model?

No. The estimate is calculated in this tab with simple arithmetic. Nothing you paste leaves your browser.

More free tools.

All tools

Need tooling like this inside your product?

We build internal tools, developer platforms and APIs. Tell us what your team keeps doing by hand.