TPS stands for tokens per second, and it is the number AI labs quote when they say a model is fast. If a model runs at 50 TPS, it writes about 50 tokens every second, which is roughly 37 words. That is the whole definition, but the number is used loosely, so the details below are worth knowing before you compare models or judge a speed claim.
This matters more now than it did a year ago. OpenAI just said it made GPT-6 Astra and GPT-6.1 Sol about 50 percent faster by default, from roughly 30 to roughly 50 TPS, and its paid Ultrafast tier has been described at up to 750 tokens per second. To judge whether those numbers change your day, you need to know what they measure and what they leave out.
TL;DR: TPS in one table
| Question | Short answer |
|---|---|
| What does TPS mean? | Tokens per second: how fast a model generates output |
| What is a token? | A piece of text, about 4 characters or three quarters of an English word |
| How do I turn TPS into a wait? | Output tokens divided by TPS |
| Is 50 TPS fast? | Faster than you can read; fine for chat, helpful for coding, modest for agents |
| Is higher always better? | No. Quality, cost and time to first token matter too |
| Does TPS include thinking time? | Usually not. Reasoning models spend extra time before the visible answer |
| Is it the same on every run? | No. Load, prompt length and settings move it |
What is a token?
A language model does not read letters or whole words. It reads tokens, small chunks of text produced by a tokenizer. A common English word is often one token, a long or rare word may be two or three, and punctuation and spaces count too. As a rule of thumb, one token is about four characters, so 100 tokens is about 75 words.
Different models use different tokenizers, so the same paragraph can be 800 tokens on one and 950 on another. That is why a more efficient tokenizer lowers both cost and wait time: fewer tokens are needed for the same text. We covered how far tokenizer speed can go in the GigaToken tokenizer write-up. For the formal term, see the Tokens Per Second dictionary entry.
How does a model produce tokens one at a time?
Language models are autoregressive. They generate one token, append it to the text so far, and run again to produce the next one. Every token needs a pass through the network, which is why long answers take time. The architecture behind this loop is explained in our guide to the transformer.
The process has two phases:
- Prefill: the model reads your whole prompt in one parallel pass. This is fast per token, but a very long prompt still takes noticeable time.
- Decode: the model writes the answer token by token. This is the part TPS measures.
Decode is slower per token than prefill because each step depends on the one before it. It is also limited mostly by memory bandwidth, not raw compute: for every token the hardware must read the model's weights from memory. That is why chips with very fast memory can reach hundreds of TPS, as laid out in our AI chip architectures guide.
How do you calculate the wait from TPS?
The formula is short:
generation_time_seconds = output_tokens / tps
A 1,000-token answer, about 750 words, takes 20 seconds at 50 TPS and about 33 seconds at 30 TPS. A 4,000-token code file takes 80 seconds at 50 TPS and 133 at 30.
Notice what the formula leaves out. Real tasks also include the delay before the first token, time spent reasoning, reading files, running tools and waiting for tests. TPS speeds up only the writing part, so a faster number does not cut total task time by the same percentage.
The lab below lets you move both sliders and watch the wait change.
What counts as a good TPS?
There is no universal threshold, but there are useful anchors.
| TPS | How it feels | Typical use |
|---|---|---|
| Under 15 | Visibly slow; you wait and watch | Heavily loaded or very large models |
| 15 to 30 | Readable, but long answers drag | Older default speeds on frontier models |
| 30 to 60 | Comfortable for chat and coding | Current default for many frontier models |
| 100 to 300 | Near instant for most answers | Small models, optimized serving |
| 500 and up | Whole files appear almost at once | Specialized hardware and premium tiers |
Reading speed is the baseline. An adult reads around 240 words per minute, which is about 5 to 6 tokens per second. Anything above that stays ahead of you while you read a streaming answer. Higher speeds help when you skim, when output is long, and above all when a program, not a person, is consuming the output.
TPS is not the only speed number
If you only look at TPS, you can pick a model that feels slow. Four other measurements matter.
Time to first token
Time to first token is how long you wait before anything appears. It includes queueing at the provider, network delay and prefill. For short questions it dominates how fast a model feels. A model that starts in 300 milliseconds and runs at 40 TPS can feel quicker than one that starts after 4 seconds and runs at 100.
Reasoning time
Reasoning models spend tokens thinking before the visible answer. Those tokens may be hidden or shown as a summary, and they are billed as output. A reply that shows 300 tokens may have involved several thousand tokens of thinking. When someone quotes TPS for a reasoning model, check whether the figure counts those hidden tokens, because the wait you feel includes them.
Tokens needed per task
A faster model that needs twice as many tokens to solve a problem is not faster in practice. Efficiency means fewer tokens to reach a correct result. Artificial Analysis found that GPT-6.1 Sol costs about 78 percent less per task than Astra, a reminder that cost and time per completed task are the fairer comparisons. Total time is roughly tokens needed divided by TPS, so a model can win by being smarter, faster, or both.
Throughput versus per-user speed
Providers often quote total throughput across many users at once, such as tokens per second across a whole GPU. That is not what one person experiences. Your speed is the per-request figure. When you see a dramatic number, ask whether it is per user or per system.
What makes TPS higher?
Speed comes from three places, and understanding them helps you read launch announcements critically.
Hardware. Memory bandwidth sets the ceiling for decoding. Chips built around large on-chip memory, such as those from Cerebras and Groq, reach much higher TPS than a standard GPU serving the same model.
Model design. Smaller models, mixture-of-experts layouts that activate only part of the network per token, and quantization to lower precision all cut the data moved for each token.
Serving techniques. Speculative decoding lets a small draft model propose several tokens that the big model verifies in one pass, which can raise TPS without changing the output. Our DeepSeek speculative decoding guide walks through a real implementation. A cached key-value store, described in the KV cache entry, avoids recomputing earlier tokens, and continuous batching keeps hardware busy.
When a provider announces a speedup with no new model, one of these is usually responsible. The October 2026 default speed change for GPT-6 models, for instance, was described as an optimization of existing models, which points to serving improvements and not a new network.
Fast tiers and what they cost
Speed is increasingly sold as a product. OpenAI has offered Ultrafast as a paid tier for Astra, tied to its highest plan, as covered in our Ultrafast and Pro 500 breakdown. The trade is straightforward: the provider dedicates faster hardware or more of it to your requests, and you pay for the privilege through a higher plan or per-token premium.
When you evaluate such a tier, ask three questions. What is the measured TPS, not the marketing maximum? Does it apply to the model and effort level you use? And will the savings in waiting actually matter for your task? For an overnight batch job, faster generation is worth little. For a live coding session or a voice assistant, it can be the main feature.
How to measure TPS yourself
You can verify a speed claim in a few minutes.
- Choose a prompt that produces a long answer, such as "explain how a B-tree works in about 1,200 words."
- Send it and note the time the first token appears and the time the last one arrives.
- Count the output tokens. Most APIs return a usage field, and many harnesses show it at the end of a turn.
- Compute
output_tokens / (end_time - first_token_time). - Repeat at least three times, at different hours, and keep the range.
import time
start = time.time()
first = None
tokens = 0
for chunk in stream: # your streaming API iterator
if first is None:
first = time.time() # time to first token = first - start
tokens += 1 # replace with real token count from usage
end = time.time()
print("TTFT:", round(first - start, 2), "s")
print("TPS:", round(tokens / (end - first), 1))
Counting chunks is only an approximation, since a streamed chunk can hold more than one token. Prefer the exact count returned by the API when it is available.
Common mistakes when reading TPS
- Treating TPS as intelligence. Speed and quality are separate axes.
- Comparing across different tokenizers. A model with a less efficient tokenizer needs more tokens for the same text, so equal TPS does not mean equal words per second.
- Ignoring reasoning tokens. Hidden thinking can multiply the wait.
- Trusting a single run. Load varies, so a single measurement is noise.
- Quoting the peak. Vendors cite best-case numbers; plan around the typical figure.
- Confusing percent and ratio. Going from 30 to 50 TPS is a 67 percent throughput gain but only a 40 percent drop in time per token, so "50 percent faster" is an approximation depending on which way you measure.
What this means for what you build
If you are building a chat interface, aim for fast time to first token and a comfortable 30 to 50 TPS; users care more about the start than about the rate. If you are building an agent that chains dozens of calls, TPS and token efficiency compound, and each saved second is multiplied by the number of steps. If you pay by the token, remember that speed and price are independent: a faster tier may or may not cost more per token.
For choosing between models, the Astra versus Sol comparison shows the kind of trade you will face, and why AI companies want you using agents explains why token volume is central to the economics. Sampling settings also affect outputs, as described in our guide to temperature, top-p and top-k, though they do not change speed meaningfully.
Related reading on explainx.ai
- Astra 28 days of updates: day-by-day log
- GPT-5.6 Sol Ultrafast mode and 750 tokens per second
- Ultrafast and Pro 500 at DevDay 2026
- AI chip architectures: GPU, TPU, Trainium, Cerebras, Groq
- GigaToken: a Rust tokenizer 1000x faster
- DeepSeek speculative decoding guide
- GPT-6.1 Sol cost efficiency vs Astra
- What is the transformer architecture?
Figures, product tiers and model speeds are accurate as of October 5, 2026. Token-to-word ratios are rules of thumb and vary by tokenizer and language.
