Tokens are the small chunks of text, often whole words, pieces of words, punctuation or spaces, that a large language model actually reads and writes, and the context window is the fixed number of tokens a model can handle in one request, counting your input and its output together. Almost everything practical about using an LLM comes back to tokens: how much text fits, how long a response takes, and what you pay. This guide explains the mechanics in plain language, with approximate examples and an illustrative budget for a RAG app. Exact context sizes and prices vary by model and change often, so check your provider's current documentation for real figures.
What is a token in an LLM?
A model does not see letters or words the way you do. Before your text reaches the network, a component called a tokenizer chops it into tokens and maps each one to an ID number from a fixed vocabulary. The model works only with those IDs, predicts the next token ID, and the tokenizer turns the IDs back into text. For what happens next, see our explainer on what an LLM is and how it works walks through embeddings, attention and next-token prediction.
Why not whole words? A vocabulary covering every word, name, typo and code identifier would be impossibly large. Single characters would make sequences very long and slow. Tokens sit in the middle.
Subword tokenization, simply
Most modern tokenizers use subword tokenization. An algorithm scans a large body of text and learns which character sequences appear together often. Frequent words become one token; rarer words are built from reusable pieces; and any new word can still be spelled out from smaller pieces.
Here is an approximate, illustrative split. Real tokenizers differ, and the same sentence will split differently from one model family to another:
| Text | Possible tokens (approximate) | Why |
|---|---|---|
the | the | Very common word, one token |
unbelievable | un | believ | able | Built from common pieces |
Ameerpet | Am | eer | pet | Proper noun, rare in training text |
getUserName() | get | User | Name | () | Code identifiers split on patterns |
Leading spaces are usually part of a token, and numbers, dates and IDs often break into several pieces, one reason LLMs are clumsy at character-level tasks such as counting letters.
Tokens vs words
In ordinary English prose, a passage usually has somewhat more tokens than words. Treat any words-to-tokens ratio as a rough rule for English prose only; it breaks down for code, JSON, URLs, numbers and non-English text. When the count matters, measure it.
Why token counts differ across languages
A tokenizer learns its vocabulary from its training text. Scripts and languages that were heavily represented get efficient, longer tokens. Less represented scripts get split into shorter fragments, sometimes single characters or raw bytes.
In practice, this means text in Indian languages such as Telugu, Hindi, Tamil or Kannada often needs noticeably more tokens than the same meaning written in English. The gap depends heavily on the tokenizer, and newer tokenizers are generally more efficient for many languages, so measure rather than assume.
Code-mixed text, the Hinglish or Tenglish that real users type ("mera leave balance kitna hai?", or Telugu written in Roman script), is a special case. It uses Latin letters, but the word patterns are unusual compared with standard English, so it frequently splits into more pieces than English text of similar length.
Consider an insurer's customer support assistant serving users across Andhra Pradesh and Telangana (an illustrative example). The team budgets using English test queries. In production, many questions arrive in Telugu script and code-mixed Telugu-English, with replies in the same language. Requests use more tokens than planned, so costs rise, responses slow and long chats hit the limit sooner. The fix: build your test set from real user language and budget from measured counts.
The context window explained
The context window is the maximum number of tokens a model can work with in a single request. The crucial beginner point is that input and output share the same window. Everything you send counts, and so does everything the model writes back.
+------------- context window -------------+
| system prompt | tools | history | |
| retrieved docs | user question | |
|------------------------------------------|
| model output (generated token by token) |
+------------------------------------------+
input tokens + output tokens <= window
"Input" is much more than the user's message: system prompt, tool definitions, earlier turns, retrieved documents, files converted to tokens and formatting the API adds. Deciding what fills each slot is context engineering, covered in our context engineering guide; this article stays with counting, limits and cost.
One more thing to understand: the model has no memory between API calls. A chat app "remembers" because your code re-sends the earlier conversation every turn. So long chats get steadily more expensive and eventually run out of room. Persistent memory is built on top, as explained in our article on memory in AI agents.
What happens when you exceed the token limit
Behaviour depends on the provider and the framework, but you will usually see one of these:
- The request is rejected. Many APIs return an error if input plus the requested maximum output exceeds the window.
- The output is cut off. If the input fits but the model runs out of room or hits your maximum output setting, the response stops mid-sentence. The API usually reports a length-limit finish reason, and truncated JSON will not parse.
- Something is silently dropped. Some interfaces and frameworks trim old messages or truncate documents to fit. The request succeeds, but the model no longer sees what it needs.
The last case is the most dangerous because nothing looks broken: if a document is truncated before the answering clause, the model may fill the gap with something plausible but wrong. That failure pattern is covered in LLM hallucinations explained.
Why long context isn't free
Context windows have grown a lot, and it is tempting to paste in everything. Long context helps some tasks, such as reviewing one large contract in full, but it has three real costs.
- Money. Hosted models charge per token, so every token you send on every call costs something. A large prompt repeated on every request adds up quickly.
- Latency. All input is processed before the first output token appears, and output is generated one token at a time, so more of either means a longer wait. Techniques for reducing this are in our guide to LLM latency optimization.
- Attention dilution. More text is not more understanding. A relevant fact buried in loosely related material can be overlooked, or the model distracted by similar but wrong passages. Long-input quality varies by model and task, so test on your own data.
A well-chosen small context usually beats a huge unfiltered one on cost, speed and accuracy.
Input vs output tokens: how LLM pricing works
Most hosted LLM APIs bill per token, quoted per large block of tokens. The numbers vary by model, provider and region and change often, but the pattern is consistent:
- Input tokens are everything you send. They are usually cheaper per token.
- Output tokens are everything the model generates. They are usually more expensive per token, because generation happens one token at a time.
- Some models generate internal reasoning tokens before answering. Depending on the provider, these may be billed as output even if you never see them in full. Our sibling article on reasoning models explains why.
So the basic formula for one request is:
cost = input_tokens x price_per_input_token
+ output_tokens x price_per_output_token
Multiply by requests per day and you have a first estimate. Self-hosted models swap per-token bills for GPU and operations costs, but tokens still drive how much hardware you need. For the wider picture of keeping AI workloads affordable, see cloud cost optimization for AI.
Cached tokens
Many applications send the same large prefix on every call: a long system prompt, tool definitions, a policy handbook. Several providers support prompt caching, where the model reuses its processing of a repeated prefix instead of recomputing it. Cached input tokens are typically billed at a reduced rate and can also cut latency.
Details differ by provider: some cache automatically, others need explicit markers; caches expire after inactivity; and there may be a minimum prefix length. The general rule holds everywhere: caching only works on an identical prefix, so put stable content first and anything that changes per request (the user's question, retrieved chunks, timestamps) after it.
If you want to practise these mechanics hands-on, counting tokens, comparing models and building real RAG and agent apps on Amazon Bedrock, Azure OpenAI and Gemini, Cloudsoft's AI, GenAI and Agentic AI course is built around labs rather than slides.
How to count tokens in code
Never guess token counts in production code. There are three reliable ways to get them:
- Tokenizer libraries. Open-weight models ship their tokenizer, which you can load with libraries such as Hugging Face
transformersortokenizers. Some API providers publish tokenizer libraries too. Use the tokenizer matching the model you call; another one gives only an estimate. - Token-counting endpoints. Some providers offer an API call that returns a request's token count without generating anything, including the formatting the API adds.
- Usage fields in responses. API responses report the input, output and (where supported) cached tokens actually used. Log them on every call; they are the ground truth for cost tracking.
A minimal Python sketch for an open-weight model looks like this (the model ID is a placeholder):
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("your-model-id")
text = "mera leave balance kitna hai?"
ids = tok.encode(text)
print(len(ids), tok.convert_ids_to_tokens(ids))
Printing the tokens shows the language effects described earlier. If you are new to Python for this kind of work, start with Python for AI engineers.
The max output tokens setting
Most APIs let you set a maximum output tokens value (named differently by different providers) for each request. It is a ceiling, not a target. The model will stop earlier if it finishes, but it will never write more than this.
- Set it deliberately. A classification needs a handful of tokens, a long report far more. A sensible cap controls cost and stops runaway responses.
- Do not set it too low. JSON gets cut off and fails to parse, and reasoning models need room to think before answering.
- Remember the shared window. A large output allowance leaves less room for input.
- Check the finish reason. A length-limited response is incomplete: retry with a higher limit or ask for a shorter format.
Practical token budgeting for a RAG app
Consider a GCC IT team in Hyderabad building an internal HR policy assistant with retrieval-augmented generation (an illustrative example). They draw up a token budget first. Every number below is an illustrative placeholder, not a real model limit or price. Substitute measured values from your own tokenizer and your provider's current pricing.
| Slot | Illustrative tokens | Notes |
|---|---|---|
| Context window (assumed) | 16,000 | Placeholder; real windows vary by model |
| System prompt and rules | 800 | Stable, so a good caching candidate |
| Tool / output schema definitions | 600 | Stable |
| Conversation history | 1,500 | Recent turns plus a summary of older ones |
| Retrieved chunks | 5 x 500 = 2,500 | Depends on chunk size and top-k |
| User question | 150 | Measure in real user languages |
| Total input | 5,550 | |
| Max output tokens | 700 | Room for an answer plus citations |
| Safety margin | ~1,500 | For code-mixed queries and long turns |
| Planned peak use | 7,750 | Well inside the assumed window |
The team then estimates cost with placeholder prices P_in and P_out per token:
per request ~= 5,550 x P_in + 700 x P_out
per day ~= per request x expected requests
(cached prefix of 1,400 tokens may bill lower)
Three lessons usually emerge. First, retrieved chunks are often the largest controllable slot, so chunk size and count drive cost and quality; our guide to RAG chunking strategies covers the trade-offs. Second, history grows every turn, so cap or summarise it. Third, the margin is not waste: real traffic is messier than test data.
The budget is a hypothesis: compare it with the usage you log once the app runs, and adjust. Taking an assistant like this from a working prototype to a measured, cost-controlled system inside a customer's environment is a large part of what Forward Deployed Engineers do, and it is the focus of Cloudsoft's FDE PRO program.
Common mistakes with tokens and context windows
- Treating words and tokens as the same. Budgets built on word counts are wrong for code, numbers and most non-English text.
- Forgetting hidden input. System prompts, tools and history all count.
- Ignoring output in the window. A full window leaves no room for the answer.
- Not checking finish reasons. Truncated responses mean broken JSON and half answers.
- Assuming a big window means good recall. Test long-context accuracy on your own documents.
- Testing only in English. Telugu, Hindi and code-mixed users change your token profile.
- Hardcoding limits and prices. Keep them in configuration; models are updated and retired. Our guide on choosing an LLM for enterprise use covers how to compare options.
Frequently asked questions
What are tokens in an LLM?
Tokens are the units of text an LLM reads and generates: whole words, parts of words, punctuation and spaces, each mapped to an ID number. Context limits and API pricing are measured in tokens, not words.
How many tokens are in a word?
There is no fixed number. In English prose a passage usually has somewhat more tokens than words, while code, numbers and many non-English languages use more, so measure with the model's own tokenizer.
What is a context window in simple terms?
It is the maximum number of tokens a model can handle in one request, counting everything you send plus the response it generates. Window sizes vary by model and change as new models are released.
What happens if my prompt is longer than the context window?
Usually the API returns an error or the response is cut off. Some tools silently drop older messages or truncate documents instead, which quietly reduces answer quality.
Why do Indian languages use more tokens than English?
Tokenizers learn their vocabulary from training text, and less represented scripts get split into smaller pieces. Telugu, Hindi and code-mixed text such as Hinglish often need more tokens than equivalent English, depending on the tokenizer.
Why are output tokens more expensive than input tokens?
Input tokens are processed together in parallel, while output tokens are generated one at a time. Generation uses more resources, so most providers price output higher, though exact prices vary and change often.
What are cached tokens?
They are a repeated prompt prefix, such as a long system prompt, that the provider has already processed and can reuse. Where supported, cached input is usually billed at a reduced rate, but only an identical prefix qualifies.
Is a bigger context window always better?
No. Every extra token adds cost and latency, and models can miss facts buried in long text. Sending a small amount of highly relevant context usually works better.
Tokens and context windows are the first practical layer of building with LLMs; RAG, agents, evaluation and deployment sit on top of them. To learn that full stack with guided labs, classroom in Ameerpet or live online, explore Cloudsoft's AI, GenAI and Agentic AI training in Hyderabad, or call +91 96660 19191 to book a free demo.



