explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • What “parameters” means in one paragraph
  • Why people still talk in “billions”
  • Total vs “active” parameters: MoE in plain terms
  • Frontier API models: strong specs, often no public parameter line
  • How to use parameter counts in practice
  • Can parameter count tell you whether a model fits in memory?
  • What does a larger model know that a smaller one does not?
  • How should you discuss undisclosed model sizes?
  • Read next
← Back to blog

explainx / blog

What are parameters in a large language model? Billions, MoE, and what 2026 model cards really say

LLM basics, Model parameters, MoE, Meta Llama, AI research, Anthropic, OpenAI

Part of Open-Weight Models

Model parameters are the learned numbers inside a neural net—roughly, how big the model is. Here is a clear picture of total vs active parameters, why frontier APIs often hide counts, and a table of top models with public figures (Meta Llama 4) next to the undisclosed front tier.

Apr 22, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
What are parameters in a large language model? Billions, MoE, and what 2026 model cards really say

Update — July 20, 2026: A viral X thread claimed Claude Sonnet ~1T, Opus ~5T, Fable/Mythos ~10T from an unnamed compute partner — UNVERIFIED. Musk's Grok 4.20 reply independently implies the same Sonnet/Opus ratios. Full debate breakdown →

Update — July 31, 2026: For a full explainer plus the top 10 disclosed model sizes (Kimi K3 2.8T → ~400B class), see What Are LLM Parameters? Top 10 Model Sizes (July 2026).

Parameters (often billions of parameters, or B / bn) are the usual shorthand for how big a neural language model is in terms of learned weights. They are not the same as tokens (text units) and not the same as context length (how much text fits in one request)—but all three get compared when people discuss GPT-class, Claude-class, Gemini-class, and open weights like Llama.

This article hews to what vendors publish in 2026 and to one open line with full public tables: Meta’s Llama 4 on the official model card. Frontier APIs often list behavior and limits without a single headline parameter count; we cover why below.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


What “parameters” means in one paragraph

LLM parameters explained — a dense particle cube with one glowing active subset slice representing MoE routingLLM parameters explained — a dense particle cube with one glowing active subset slice representing MoE routing

A transformer-style LLM is a stack of layers that transform vectors representing tokens. Pretraining and fine-tuning adjust the entries of large weight matrices (and related biases) so the model improves at next-token prediction, tool use, or multimodal tasks—depending on the architecture.

Parameter count is how many of those scalar weights sit in the shipped checkpoint. Some model cards also break out separate totals for a tokenizer, vision tower, or audio encoder—read the specific card for the SKU you run.


Why people still talk in “billions”

  1. Capacity — With similar data and training, a larger weight budget can represent richer patterns; in practice, data and recipe still dominate outcomes.
  2. Serving cost — More weights (especially active per forward pass) tend to mean more FLOPs and memory at inference, though quantization and hardware matter.
  3. MoE (mixture of experts) — A model can have a huge total while routing each token through only a subset of “expert” blocks, so active width is the better first-order handle on per-step compute.

Scaling laws in research usually relate loss to compute, data, and size together; a headline “B” count is one line in a larger system.


Total vs “active” parameters: MoE in plain terms

In a dense model, “70B parameters” generally means on the order of 70B weights on the main path of each token (implementation details aside).

In an MoE design, many parallel feedforward experts exist, but a router sends each token to one or a few of them. Cards often list total parameters (all experts) and activated parameters (roughly what runs for a typical forward pass).

Neural network diagram with only a small cluster of nodes lit up, illustrating how a mixture-of-experts router activates a sparse subset of parameters per token instead of the full networkNeural network diagram with only a small cluster of nodes lit up, illustrating how a mixture-of-experts router activates a sparse subset of parameters per token instead of the full network

Meta Llama 4 (from the model card table, April 2025 release; confirm on GitHub for updates):

table · 4 cols
ModelActivated (per card)Total (per card)Context length (per card)
Llama 4 Scout (17B × 16E)17B109B10M tokens
Llama 4 Maverick (17B × 128E)17B400B1M tokens

E denotes expert count in Meta’s notation. Scout and Maverick share the same activated width in this table but differ in total size and in context length by design. Always re-read the model card for the exact checkpoint you deploy.


Frontier API models: strong specs, often no public parameter line

  • OpenAI documents GPT-5.4 and lists model behavior, context, and API model pages—without a public total parameter count in the same way open releases do.
  • Anthropic publishes Claude Opus 4.7 and a models overview with context, pricing, and features—not a “N billion parameters” headline.
  • Google DeepMind lists Gemini 3.1 Pro capabilities, modalities, and context—again, typically without a full parameter count in the consumer-facing card.

If you see a billion-scale number for a closed model in a third-party post, treat it as analysis or speculation unless the vendor or a vetted system card states it.


How to use parameter counts in practice

  • Open-weight models (Llama, others): the model card, license, and memory notes tell you if a run fits your GPUs—active size and quantization usually matter more than a huge MoE total for download size vs runtime.
  • APIs: use vendor docs for latency, context window, tools, and $/M tokens (tokens explainer, Caveman economics).
  • Benchmarks: treat headline size as weak evidence without measurements on your task and data.

Can parameter count tell you whether a model fits in memory?

It gives you a starting estimate for the weights, rather than a complete memory budget. Each stored parameter occupies space according to its numerical representation. As a simple illustration, eight billion weights represented with two bytes each require roughly sixteen billion bytes for the weights alone. This arithmetic excludes additional runtime allocations and is not a promise that the model runs on a machine with that amount of RAM.

Quantization reduces the storage used per weight, usually with a trade-off in numerical precision and sometimes quality. A four-bit representation starts from half a byte per weight, but real formats also need metadata and implementation overhead. Compare actual checkpoint file sizes and the inference engine's requirements rather than treating the theoretical calculation as a buying recommendation.

The conversation consumes memory too. In transformer serving, cached information from earlier tokens can grow with sequence length and the number of simultaneous requests. A model that fits with one short prompt may run out of memory with a long document or several users. Architecture and serving settings determine how much that extra memory costs.

Why active parameters do not equal download size

An MoE model may activate only some expert weights for a token, while still needing the full checkpoint available to the serving system. That distinction is easy to miss when a model name advertises both total and active size. Small active size can reduce computation; it does not automatically turn a large checkpoint into a small download.

Think of a reference library: answering one question may require only a few shelves, but keeping the whole library available still takes space. This analogy is about storage versus use, not a literal description of routing. The model card and serving documentation tell you whether experts are resident, distributed, or moved between storage and memory.

The official Llama 4 model card is useful because it reports both categories. When reading another model's card, look for the same distinction and note what is included in the total. Different authors may count embedding matrices or multimodal components differently, which makes loose comparisons misleading.

What does a larger model know that a smaller one does not?

Parameter count describes the number of learned numerical values. It is not a table of facts, a count of documents memorized, or a guarantee that particular knowledge is present. Training data, optimization, instruction tuning, and the way a model is evaluated all affect what it can do with those values.

This matters when a team wants to answer questions about its own documents. A larger model may reason better about supplied context, but more parameters do not give it access to a private handbook it has never received. Retrieval and grounding address the missing-information problem. Model size and access to relevant evidence are separate decisions.

Similarly, a smaller specialized model can be a sensible choice for a narrow extraction or classification task. The important question is whether its errors are acceptable on your input distribution. A broad benchmark result cannot tell you whether it handles your scanned invoices, internal terminology, or unusually long support messages.

A practical comparison exercise

Build a small set of representative questions before choosing a model. Include easy cases, ambiguous cases, and inputs where the correct behavior is to abstain. For extraction, include documents with missing fields. For coding, include repository conventions and a relevant test. State what counts as success before looking at the outputs.

Run candidate models with comparable instructions and the same evidence. Record correctness, latency, output length, and failure behavior. If a model produces a confident answer where the evidence is missing, count that separately from a formatting mistake. A smaller model that asks for clarification may be safer for that workflow than a larger one that invents details.

Then assess cost under your expected traffic. For a local deployment, consider hardware, peak concurrency, and operational work. For an API, consider billed tokens, retries, and provider constraints. These comparisons depend on your workload; parameter count is contextual information rather than a final score.

How should you discuss undisclosed model sizes?

Use an explicit unknown when the publisher has not disclosed a parameter count. Do not infer a precise value from a product name, subscription price, or response quality. A rumor can be described as a rumor, but it should not become the basis for memory calculations or a table that looks authoritative.

Keep three fields separate in your notes: disclosed architecture facts, your own measured behavior, and third-party estimates. That separation prevents a plausible estimate from acquiring the status of a specification through repetition. It also makes later updates straightforward when an official card supplies information that was previously unavailable.

Read next

  • What Are LLM Parameters? Top 10 Model Sizes (July 2026)
  • What are tokens?
  • Context window: what it is, with 2026 model snapshot
  • Claude Opus 4.7: limits and pricing in Anthropic’s own table

Parameter and architecture details change with each release. Prefer the model card and API docs of the specific checkpoint you use.

Spotted something out of date? Let us know.

People in this article

  • Elon Musk →Tesla CEO and technology entrepreneur
Explore people in AI →
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Apr 22, 2026

What is a context window? LLM 'working memory' and a 2026 snapshot of top models

Context length is the cap on 'how much the model can read at once,' not the same as how many parameters it has. This guide defines the window, input vs max output, long-context tradeoffs, and what Anthropic, OpenAI, Google, and Meta publish today.

Apr 22, 2026

What are tokens? A plain guide to how LLMs count (and charge for) text

If you have ever seen

Oct 9, 2026

OpenAI Says $50B Run Rate, Not $70B: What Was Claimed and What Is Verified

A $20 billion gap between two OpenAI revenue numbers moved AI stocks on October 8. Reuters, CNBC and the Financial Times report the company told investors it was near $50 billion annualized at the end of September, not the $68 to 70 billion that had been widely repeated. Here is what is confirmed, what is not, and what it changes for people who build on these models.