Related: Which Surface Laptop Ultra configuration to buy for local AI.
On October 7, 2026, Microsoft said a coding model it has been selling through the cloud will now run on your own PC. In a post on X summarizing the Windows and Surface event, Satya Nadella listed MAI-Code-1.1 Flash first: a "137B parameter coding model w/ 256K context window, which is now optimized to run on your PC." He added that GitHub Copilot "now hands off work to local models like MAI-Code-1.1 Flash, helping projects cost a lot less without sacrificing quality."
That is a bigger claim than it first sounds. A 137-billion-parameter model on a laptop is not a typical developer setup. This article lays out what Microsoft and reviewers have reported about the model, the memory it needs, how Copilot decides what runs where, and which claims you should hold until documentation lands. It belongs next to our coverage of the Windows hybrid intelligence announcement, which explains the sandbox and routing layer this model plugs into.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | Microsoft's small-tier coding model, 137B total and 6.8B active parameters, 256K context |
| Is it new? | The cloud model shipped to GitHub Copilot on August 11, 2026; the local build was announced October 7 |
| What hardware? | Reference numbers come from a Surface Laptop Ultra with Nvidia RTX Spark and up to 128GB unified memory |
| How much memory? | About 75.5GB peak at full 256K context; Microsoft reportedly recommends over 120GB of RAM |
| Is it free locally? | Nadella says no cloud token spend; Microsoft has not published local pricing |
| When can I try it? | Limited experimental rollout reported for the end of October 2026 |
| Can my current laptop run it? | Probably not unless it has well over 64GB of fast unified memory |
| Does it replace the cloud? | No. Copilot routes between local and cloud |
What was actually announced
Nadella's post bundled six points. The first two matter most for developers: the local model, and Copilot's ability to hand work to it. The others were Hybrid Intelligence acting directly on the PC, "Code in Copilot" for building software on the desktop "without any cloud token spend," the pairing of Windows with Agent 365 and the MXC sandbox, and new hardware such as the Surface Laptop Ultra powered by Nvidia RTX Spark.
The phrase doing the heavy lifting is "infinite software factory." Nadella argued that you are "no longer limited to what's in an app store" because your PC can generate the software you need. Treat that as a pitch. What is checkable today is the model card and the hardware numbers.
The numbers: size, memory and speed
Here is what has been reported, with the source of each figure noted.
| Item | Reported value | Source type |
|---|---|---|
| Total parameters | 137B | Microsoft figures as reported |
| Active parameters | 6.8B (mixture of experts) | Microsoft figures as reported |
| Context window | 256,000 tokens | Microsoft announcement coverage |
| Local precision | 3-bit, reducing model size by nearly 80% | Press coverage |
| Quantized file size | About 53GB | Press coverage |
| Peak memory at 256K context | About 75.5GB | Reference measurement on Surface Laptop Ultra |
| Decode speed at 64K context | 923.5 tokens per second | Reference measurement |
| Decode speed at 128K context | 769.8 tokens per second | Reference measurement |
| SWE-Bench Verified, local | 70.8 percent | Microsoft claim, reported |
| SWE-Bench Verified, cloud | 72.6 percent | Microsoft claim, reported |
Two cautions. First, the benchmark pair is Microsoft's own, and one outlet noted it found no independent testing. A gap of 1.8 points between local and cloud would be a small price for running offline, but it needs a third-party reproduction. Second, sources disagree on the exact size: one model-card summary lists 138B total and 5B active, while Microsoft's figures in other coverage are 137B and 6.8B. We used the Microsoft numbers and flag the discrepancy.
The speed figures are fast for a model this size, which makes sense for a mixture-of-experts design. Only about 6.8 billion parameters fire per token, so memory capacity is the constraint, not raw compute. If you want the broader reasoning on why capacity beats peak FLOPS for local inference, our DGX Spark local LLM setup guide and the MacBook versus dedicated GPU comparison cover it.
Can your PC run it?
Almost certainly not today, unless you bought for it. Microsoft's reference box was the Surface Laptop Ultra with Nvidia RTX Spark, which has been reported at up to 128GB of unified memory. The model alone is about 53GB, and the key-value cache for a long context takes the rest, which is how you reach roughly 75.5GB at 256K tokens.
Reviewers note that the recommendation of more than 120GB of RAM is a "best performance" figure, not a published hard minimum. Microsoft has not given a minimum spec or a supported-device list. One write-up put it plainly: local inference "is not equivalent to running a lightweight autocomplete tool on any laptop."
A practical rule from the reference numbers:
- Under 64GB of memory: do not expect to run it. Use the cloud model through Copilot.
- 64GB to 96GB: possible at reduced context, but unproven. Wait for community tests.
- 128GB unified memory: the class Microsoft measured. Expect near-full context to fit.
Our write-up of the Surface RTX Spark Dev Box and Laptop Ultra has the hardware prices and ship dates.
How Copilot hands work to a local model
According to reporting that describes the architecture, there are two paths.
- Automatic routing. Copilot decides whether a request runs locally or in the cloud. The reporting ties this to Microsoft's HydraFusion orchestration approach, which we covered in GitHub Copilot HydraFusion multi-model orchestration.
- Explicit selection. You pick MAI-Code-1.1 Flash yourself, either through a Windows ML provider or by pointing Copilot at a local endpoint.
Tool execution is meant to run inside Microsoft Execution Containers (MXC), the Windows sandbox for agents. The surfaces named are the GitHub Copilot app, Copilot CLI and Visual Studio Code. Microsoft also said it is bringing local versions of other models to Windows, including DeepSeek V4 Flash and Nvidia's upcoming Nemotron, per coverage of the event.
The routing rule has not been published. That is the part that decides your bill and your privacy: what goes to the cloud, what stays on the device, and whether you can force one or the other. Until Microsoft documents it, assume the auto mode may send some requests off the machine.
What it costs
There are three separate prices to keep apart.
| Path | Price or status |
|---|---|
| Cloud MAI-Code-1.1 Flash | List price reported at $0.20 per million input, $0.02 cached, $1.20 per million output tokens |
| GitHub Copilot premium requests | 0.25 times multiplier for annual subscribers, per GitHub's August 11 changelog |
| Local inference | Nadella: no cloud token spend. Microsoft has not published local pricing or limits |
GitHub's August changelog also gives the plan picture for the cloud model. It is available to Copilot Free and Student through auto model selection, and Pro, Pro+, Max, Business and Enterprise plans can pick it manually. Business and Enterprise administrators must turn on the policy, which is "off by default." That matters if you manage a team: nothing will route to this model until an admin flips the setting.
"Cost a lot less" is plausible for tasks a smaller local model can finish. It is not a guarantee. If a local attempt fails and Copilot retries in the cloud, you pay twice in time and possibly in tokens. Measure your own mix.
What is still unconfirmed
- Local pricing and limits. No official figure. The "zero inference charge" claim comes from a social post and Nadella's wording.
- The routing policy. No published rule for local versus cloud, and no stated guarantee that sensitive work stays local in auto mode.
- Independent benchmarks. The 70.8 and 72.6 percent SWE-Bench Verified figures are Microsoft's.
- Minimum hardware. No supported-device list.
- Rollout. A limited experimental rollout by the end of October is the only date reported.
- Parameter count. 137B with 6.8B active versus 138B with 5B active in one summary.
We have not run this model. Everything above is from Microsoft's announcement as relayed by Nadella and by reviewers, plus GitHub's own changelog.
What this means for what you build or pay
- If you pay for Copilot, check the policy toggle. On Business and Enterprise it is off by default for the cloud model, and local routing will likely need its own switch.
- Plan memory before you plan models. If you want to run this class of model, the buying decision is 96 to 128GB of fast memory, not a faster GPU.
- Keep the cloud path. A hybrid router means failures fall back. Design agent workflows that tolerate either path and log which one ran.
- Do not tell your security team "it stays local" yet. Wait for the routing documentation. Pair any file-access agent with containment, as in our agent sandbox isolation coverage.
- Test the end-of-October preview. That is the first chance to measure real speed, quality and cost on your own repo.
Why Microsoft is doing this
Microsoft has been building its own model line for coding, speech and images, as our look at Nadella's MAI strategy and the MAI hill-climbing story show. A model Microsoft owns and can ship to the device removes a per-token bill to a third party, ties Windows PCs and RTX Spark hardware to Copilot, and gives it a cheaper tier for the volume of small coding tasks. It also competes with the open-weight local options developers already use, such as the DeepSeek V4 Flash vision model.
The risk for Microsoft is the same as for any local model: if quality drops on hard tasks, developers will manually switch back to the cloud and the savings story weakens.
Related reading on explainx.ai
- Microsoft Windows hybrid intelligence: local and cloud agents
- GitHub Copilot HydraFusion multi-model orchestration
- Surface RTX Spark Dev Box and Laptop Ultra
- NVIDIA DGX Spark: best local LLM setup
- MacBook vs dedicated GPU for local LLMs
- Satya Nadella and the MAI model line
- GitHub Copilot SDK for multi-platform agents
Specs and availability are as reported on October 8, 2026 from Microsoft's announcement, Satya Nadella's post and GitHub's changelog, and may change as documentation is published. We have not tested the model.
