explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: TPS in one table
  • What is a token?
  • How does a model produce tokens one at a time?
  • How do you calculate the wait from TPS?
  • What counts as a good TPS?
  • TPS is not the only speed number
  • What makes TPS higher?
  • Fast tiers and what they cost
  • How to measure TPS yourself
  • Common mistakes when reading TPS
  • What this means for what you build
  • Related reading on explainx.ai
← Back to blog

explainx / blog

What Is TPS? Tokens Per Second Explained for AI Models

LLM Basics, Inference, Tokens, AI Performance, Guides

TPS means tokens per second: how fast an AI model writes its answer. Learn how to read it, what a good number is, and why 30 vs 50 TPS changes your wait.

Oct 5, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
What Is TPS? Tokens Per Second Explained for AI Models

TPS stands for tokens per second, and it is the number AI labs quote when they say a model is fast. If a model runs at 50 TPS, it writes about 50 tokens every second, which is roughly 37 words. That is the whole definition, but the number is used loosely, so the details below are worth knowing before you compare models or judge a speed claim.

This matters more now than it did a year ago. OpenAI just said it made GPT-6 Astra and GPT-6.1 Sol about 50 percent faster by default, from roughly 30 to roughly 50 TPS, and its paid Ultrafast tier has been described at up to 750 tokens per second. To judge whether those numbers change your day, you need to know what they measure and what they leave out.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: TPS in one table

table · 2 cols
QuestionShort answer
What does TPS mean?Tokens per second: how fast a model generates output
What is a token?A piece of text, about 4 characters or three quarters of an English word
How do I turn TPS into a wait?Output tokens divided by TPS
Is 50 TPS fast?Faster than you can read; fine for chat, helpful for coding, modest for agents
Is higher always better?No. Quality, cost and time to first token matter too
Does TPS include thinking time?Usually not. Reasoning models spend extra time before the visible answer
Is it the same on every run?No. Load, prompt length and settings move it

What is a token?

A language model does not read letters or whole words. It reads tokens, small chunks of text produced by a tokenizer. A common English word is often one token, a long or rare word may be two or three, and punctuation and spaces count too. As a rule of thumb, one token is about four characters, so 100 tokens is about 75 words.

Different models use different tokenizers, so the same paragraph can be 800 tokens on one and 950 on another. That is why a more efficient tokenizer lowers both cost and wait time: fewer tokens are needed for the same text. We covered how far tokenizer speed can go in the GigaToken tokenizer write-up. For the formal term, see the Tokens Per Second dictionary entry.

How does a model produce tokens one at a time?

Language models are autoregressive. They generate one token, append it to the text so far, and run again to produce the next one. Every token needs a pass through the network, which is why long answers take time. The architecture behind this loop is explained in our guide to the transformer.

The process has two phases:

  1. Prefill: the model reads your whole prompt in one parallel pass. This is fast per token, but a very long prompt still takes noticeable time.
  2. Decode: the model writes the answer token by token. This is the part TPS measures.

Decode is slower per token than prefill because each step depends on the one before it. It is also limited mostly by memory bandwidth, not raw compute: for every token the hardware must read the model's weights from memory. That is why chips with very fast memory can reach hundreds of TPS, as laid out in our AI chip architectures guide.

How do you calculate the wait from TPS?

The formula is short:

text
generation_time_seconds = output_tokens / tps

A 1,000-token answer, about 750 words, takes 20 seconds at 50 TPS and about 33 seconds at 30 TPS. A 4,000-token code file takes 80 seconds at 50 TPS and 133 at 30.

Notice what the formula leaves out. Real tasks also include the delay before the first token, time spent reasoning, reading files, running tools and waiting for tests. TPS speeds up only the writing part, so a faster number does not cut total task time by the same percentage.

The lab below lets you move both sliders and watch the wait change.

Lab · tokens per second stopwatch

What counts as a good TPS?

There is no universal threshold, but there are useful anchors.

table · 3 cols
TPSHow it feelsTypical use
Under 15Visibly slow; you wait and watchHeavily loaded or very large models
15 to 30Readable, but long answers dragOlder default speeds on frontier models
30 to 60Comfortable for chat and codingCurrent default for many frontier models
100 to 300Near instant for most answersSmall models, optimized serving
500 and upWhole files appear almost at onceSpecialized hardware and premium tiers

Reading speed is the baseline. An adult reads around 240 words per minute, which is about 5 to 6 tokens per second. Anything above that stays ahead of you while you read a streaming answer. Higher speeds help when you skim, when output is long, and above all when a program, not a person, is consuming the output.

TPS is not the only speed number

If you only look at TPS, you can pick a model that feels slow. Four other measurements matter.

Time to first token

Time to first token is how long you wait before anything appears. It includes queueing at the provider, network delay and prefill. For short questions it dominates how fast a model feels. A model that starts in 300 milliseconds and runs at 40 TPS can feel quicker than one that starts after 4 seconds and runs at 100.

Reasoning time

Reasoning models spend tokens thinking before the visible answer. Those tokens may be hidden or shown as a summary, and they are billed as output. A reply that shows 300 tokens may have involved several thousand tokens of thinking. When someone quotes TPS for a reasoning model, check whether the figure counts those hidden tokens, because the wait you feel includes them.

Tokens needed per task

A faster model that needs twice as many tokens to solve a problem is not faster in practice. Efficiency means fewer tokens to reach a correct result. Artificial Analysis found that GPT-6.1 Sol costs about 78 percent less per task than Astra, a reminder that cost and time per completed task are the fairer comparisons. Total time is roughly tokens needed divided by TPS, so a model can win by being smarter, faster, or both.

Throughput versus per-user speed

Providers often quote total throughput across many users at once, such as tokens per second across a whole GPU. That is not what one person experiences. Your speed is the per-request figure. When you see a dramatic number, ask whether it is per user or per system.

What makes TPS higher?

Speed comes from three places, and understanding them helps you read launch announcements critically.

Hardware. Memory bandwidth sets the ceiling for decoding. Chips built around large on-chip memory, such as those from Cerebras and Groq, reach much higher TPS than a standard GPU serving the same model.

Model design. Smaller models, mixture-of-experts layouts that activate only part of the network per token, and quantization to lower precision all cut the data moved for each token.

Serving techniques. Speculative decoding lets a small draft model propose several tokens that the big model verifies in one pass, which can raise TPS without changing the output. Our DeepSeek speculative decoding guide walks through a real implementation. A cached key-value store, described in the KV cache entry, avoids recomputing earlier tokens, and continuous batching keeps hardware busy.

When a provider announces a speedup with no new model, one of these is usually responsible. The October 2026 default speed change for GPT-6 models, for instance, was described as an optimization of existing models, which points to serving improvements and not a new network.

Fast tiers and what they cost

Speed is increasingly sold as a product. OpenAI has offered Ultrafast as a paid tier for Astra, tied to its highest plan, as covered in our Ultrafast and Pro 500 breakdown. The trade is straightforward: the provider dedicates faster hardware or more of it to your requests, and you pay for the privilege through a higher plan or per-token premium.

When you evaluate such a tier, ask three questions. What is the measured TPS, not the marketing maximum? Does it apply to the model and effort level you use? And will the savings in waiting actually matter for your task? For an overnight batch job, faster generation is worth little. For a live coding session or a voice assistant, it can be the main feature.

How to measure TPS yourself

You can verify a speed claim in a few minutes.

  1. Choose a prompt that produces a long answer, such as "explain how a B-tree works in about 1,200 words."
  2. Send it and note the time the first token appears and the time the last one arrives.
  3. Count the output tokens. Most APIs return a usage field, and many harnesses show it at the end of a turn.
  4. Compute output_tokens / (end_time - first_token_time).
  5. Repeat at least three times, at different hours, and keep the range.
python
import time

start = time.time()
first = None
tokens = 0
for chunk in stream:          # your streaming API iterator
    if first is None:
        first = time.time()   # time to first token = first - start
    tokens += 1               # replace with real token count from usage
end = time.time()

print("TTFT:", round(first - start, 2), "s")
print("TPS:", round(tokens / (end - first), 1))

Counting chunks is only an approximation, since a streamed chunk can hold more than one token. Prefer the exact count returned by the API when it is available.

Common mistakes when reading TPS

  • Treating TPS as intelligence. Speed and quality are separate axes.
  • Comparing across different tokenizers. A model with a less efficient tokenizer needs more tokens for the same text, so equal TPS does not mean equal words per second.
  • Ignoring reasoning tokens. Hidden thinking can multiply the wait.
  • Trusting a single run. Load varies, so a single measurement is noise.
  • Quoting the peak. Vendors cite best-case numbers; plan around the typical figure.
  • Confusing percent and ratio. Going from 30 to 50 TPS is a 67 percent throughput gain but only a 40 percent drop in time per token, so "50 percent faster" is an approximation depending on which way you measure.

What this means for what you build

If you are building a chat interface, aim for fast time to first token and a comfortable 30 to 50 TPS; users care more about the start than about the rate. If you are building an agent that chains dozens of calls, TPS and token efficiency compound, and each saved second is multiplied by the number of steps. If you pay by the token, remember that speed and price are independent: a faster tier may or may not cost more per token.

For choosing between models, the Astra versus Sol comparison shows the kind of trade you will face, and why AI companies want you using agents explains why token volume is central to the economics. Sampling settings also affect outputs, as described in our guide to temperature, top-p and top-k, though they do not change speed meaningfully.

Related reading on explainx.ai

  • Astra 28 days of updates: day-by-day log
  • GPT-5.6 Sol Ultrafast mode and 750 tokens per second
  • Ultrafast and Pro 500 at DevDay 2026
  • AI chip architectures: GPU, TPU, Trainium, Cerebras, Groq
  • GigaToken: a Rust tokenizer 1000x faster
  • DeepSeek speculative decoding guide
  • GPT-6.1 Sol cost efficiency vs Astra
  • What is the transformer architecture?

Figures, product tiers and model speeds are accurate as of October 5, 2026. Token-to-word ratios are rules of thumb and vary by tokenizer and language.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 5, 2026

Strata Runs Qwen3.8-Flash-Next 125B on a 12 GB Gaming GPU: Speed vs Quality

Strata, an MIT-licensed open-source engine, claims to run the 125B-parameter Qwen3.8-Flash-Next on a normal gaming PC with 12 GB of VRAM and 32-64 GB of RAM, at 44-124 tokens per second. A 617-point Hacker News thread tested the claim. Here is the setup, the speed table, the quant guide, and the honest quality caveats.

Oct 6, 2026

Codex Auto-Review Is Now Free: How the Reviewer Agent Cuts Approval Prompts

Codex Auto-review swaps the human approval click for a separate reviewer agent, and OpenAI now reports it is free for anyone signed in with a ChatGPT account. The docs claim roughly 200 times fewer stops for approval. This guide covers how it works, what it does not change, how to configure it, and why it is not a security boundary.

Oct 5, 2026

Impeccable 4.5: The AI Design Skill Now Has 24 Commands and 61 Detector Rules

Impeccable, Paul Bakaus's design skill for AI coding agents, now has 76.5K GitHub stars, 24 commands, 61 deterministic detector rules, a Rust engine and edit-time hooks for Claude Code, Codex, Cursor and more. Here is what changed since June, how to install it, how the detector and live mode work, and the security trade-offs to read first.