Copilot on Windows will hand some tasks to local models
A coding model with 137 billion parameters, only 6.8 billion of them active at a time, is coming to Windows PCs at 3-bit precision. In Microsoft's announcement, that model, MAI Code 1.1 Flash, anchors a plan called "hybrid intelligence": Windows picks local or cloud for each task, and Copilot, with your permission, reads your files, acts for you and calls on-device models, "helping your tokens go further."
At a glance
- With your permission, Copilot on Copilot+ PCs will use local files and recent activity, take actions inside Windows and call on-device models, with rollout across Home, Code and Autopilot starting in the coming months.
- Microsoft compresses MAI Code 1.1 Flash to 3-bit precision, cutting its size by nearly 80% while keeping a 256K context window, and GitHub HydraFusion will route work between cloud and local models.
- Microsoft gives no exact date for the Copilot features, limits them to Copilot+ PCs, and the main announcement does not say how much memory MAI Code 1.1 Flash needs on a typical machine.
According to the Windows Experience Blog, Microsoft introduced Copilot+ PCs in May 2024 and unveiled Phi Silica alongside them, an on-device model based on a derivative of Phi-3.5-mini with a 4k context length. Windows Forum reports that Copilot+ PCs are defined around NPUs capable of 40+ TOPS. Microsoft now says those machines run over 2 trillion local inferences per month, and that over 40% of laptops being built for business are Copilot+ PCs.
A 137-billion-parameter coding model and the PCs meant to run it
Microsoft introduced MAI Code 1.1 Flash at Build: 137 billion total and 6.8 billion active parameters, designed for real-world coding workloads. The local version uses 3-bit precision, which reduces the model size by nearly 80% while, Microsoft says, preserving coding quality and supporting a 256K context window on the device.
It is not the only large model headed to PCs. Microsoft lists an upcoming Nemotron model from NVIDIA with over 70 billion parameters, quantized to 2-bit so it uses just over 20GB of memory, and DeepSeek V4 Flash, a 284B parameter model, among others running locally on RTX Spark.
The hardware ranges from always-on mini desktops, where OpenClaw users get a native Windows gateway and a few-click setup, to DGX Station for Windows systems, which Testing Catalog says support models up to one trillion parameters. Surface Laptop Ultra, powered by NVIDIA RTX Spark, is on pre-order now; Testing Catalog reports up to 128 GB of unified memory and support for models above 120 billion parameters.
HydraFusion, which launched as a cloud-only router, gets local models in October
GitHub launched HydraFusion earlier this year to route each task to the right model, but only among models in the cloud. Microsoft is extending it to Windows so it can also tap models running on the device. That version arrives in experimental preview later in October in the GitHub Copilot app, GitHub Copilot CLI and Visual Studio Code.
Underneath sits Windows ML, Microsoft's runtime for deploying models across GPU, NPU and CPU. It now gets llama.cpp support, which Microsoft frames as more open-source model choices and a quick way to try emerging models as they land. According to Testing Catalog, a separate technical deep dive covers Copilot's automatic routing and explicit local model selection.
Copilot on Copilot+ PCs gets local context, local actions and local models
Microsoft recently rebuilt Copilot around three experiences: Home, Code and Autopilot. On Copilot+ PCs it gains three local capabilities. With your permission, it can read relevant content on your PC, including files and recent activity. It can act across Windows: organizing files, assessing device diagnostics, troubleshooting, coding. And it can call models on the PC, combining them with cloud intelligence when needed.
Each experience uses this differently. In Home, Copilot pulls in files you have been working on and helps turn them into collaboration-ready artifacts. In Code, it builds native Windows apps from a single prompt, runs code under MXC and uses local models to manage token costs. Autopilot, described as a persistent, proactive and personal agent, keeps working on your behalf while you do something else.
Microsoft Execution Containers reach general availability on Windows 11
Microsoft Execution Containers, or MXC, let organizations define which files and networks an agent can access, with those policies enforced at runtime. Containment options range from process and session isolation to WSLc, virtual machines and Windows 365 for Agents. Identity and management tie in through Agent 365 and Intune, so IT can tell an agent's actions apart from those of the person using the device.
Codex from OpenAI, GitHub Copilot, OpenClaw, Replit, LM Studio, OpenShell from NVIDIA and Unsloth AI already support MXC. Claude Code, Box, Egnyte, Heidi Health, Hermes Agent, Manus, Perplexity, Raycast and Simular are among those due to follow, and Meta's Muse for Windows is coming as a native app with MXC integration.
What does a 3-bit model actually give up?
Every parameter in a model is a number, and precision is how many bits store each one. Dropping to 3 bits means far fewer bits per number, which is where the nearly 80% size cut for MAI Code 1.1 Flash comes from. Think of saving a photo with a smaller color palette: the file shrinks, and the open question is whether you notice the difference.
The total-versus-active split matters as well. Of 137 billion parameters, only 6.8 billion do the work at any given step, so the model needs memory for all of them but compute for a fraction. That combination is what makes a model this size plausible on a desk rather than in a data center.
Routing is the other half. A router such as HydraFusion looks at each task and sends it to a model that can handle it, now including models on your own machine. Tasks that a local model handles well stay on the PC, and only the ones that need a frontier cloud model spend cloud tokens.
Where the plan stays vague
The Copilot features are tied to Copilot+ PCs and carry no date beyond "the coming months", and Microsoft offers no benchmark behind its claim that 3-bit MAI Code keeps its coding quality. Oddly, in our view, a pitch built on making tokens "go further" puts no number on the savings, and MAI Code's memory needs are left to a separate deep dive, so you cannot yet tell whether an ordinary Copilot+ laptop can run it.
HydraFusion in October, Copilot later
The first checkpoint is October 16, when Surface Laptop Ultra is due to arrive, according to Testing Catalog. HydraFusion's experimental preview follows later in October, and Testing Catalog says Windows Search actions for toggling settings, managing windows and sending messages from the taskbar go to Insiders first. For the hybrid Copilot features themselves, no release date has been given.
Related stories
- Microsoft Copilot gets a Code mode on GitHub Copilot tech
- OpenAI's Decisions API beats its own engine by up to 10x
- Claude Code's /diff is now a mod you can delete
- DeepSeek Harness v0.2 lands on macOS and Windows, not Linux
- OpenAI's Decisions API picks an answer in 150 milliseconds
- Cursor's new bot follows each pull request into production
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
