2026-10-07

Microsoft Brings 137B-Parameter MAI-Code-1.1 Flash to Windows PCs, Hands Copilot Work to Local Models and Makes Its Agent Sandbox MXC Generally Available

AIModelsInfrastructure🌍 North America

Microsoft announced on October 7, 2026 that Windows will be the home for "hybrid intelligence", where agents run locally when it makes sense and reach the cloud when needed. The headline items are MAI-Code-1.1 Flash, a 137B-parameter coding model now optimized to run on the PC with a 256K context window, GitHub Copilot handing work to local models, general availability of the MXC agent sandbox on Windows 11, and the Surface Laptop Ultra on NVIDIA RTX Spark. Satya Nadella's announcement thread framed it as bringing "unmetered intelligence to every desk and every home". The blog post is by Pavan Davuluri, Microsoft's EVP for Windows and Devices.

Local models: MAI-Code-1.1 Flash and others

MAI-Code-1.1 Flash, introduced at Build, has 137B total and 6.8B active parameters and is designed for real-world coding. Microsoft runs it on the device at 3-bit precision, which it says cuts model size by nearly 80% while preserving coding quality. The arithmetic roughly agrees: 137B parameters at 16 bits is about 274 GB, at 3 bits about 51 GB, an 81% reduction before overhead. That fits within the 128 GB of unified memory on the new laptops.

More models are coming to RTX Spark machines: an upcoming NVIDIA Nemotron model of over 70B parameters, quantized to 2-bit and needing just over 20 GB (70B at 2 bits is about 17.5 GB, consistent), and DeepSeek V4 Flash at 284B parameters. See earlier coverage of DeepSeek V4.1 Flash and Nemotron 3.5 Lightning. Windows ML, the runtime across GPU, NPU and CPU, adds llama.cpp support for more open-source model choices.

HydraFusion and Copilot

GitHub HydraFusion, launched earlier this year, routes each task to the right cloud model. It now extends to Windows to tap models running locally, which Nadella says helps projects "cost a lot less without sacrificing quality". The feature reaches the GitHub Copilot app, GitHub Copilot CLI and Visual Studio Code in experimental preview later in October.

The new Copilot is built around Home, Code and Autopilot, with local context, local actions and local models on Copilot+ PCs. Home turns recent files into collaboration-ready artifacts. Code builds native Windows apps from a single prompt, runs code more securely with MXC and uses local models to manage token costs; with no cloud token spend, Nadella says, "your PC becomes an infinite software factory". Autopilot is a persistent, proactive personal agent. All three roll out on Copilot+ PCs in the coming months.

MXC and agent security

Microsoft's security case rests on containment (what an agent can access), identity (which agent acted, so IT can tell it from the person) and manageability through Agent 365 and Intune. Microsoft Execution Containers (MXC) are generally available on Windows 11: organizations define which files and networks agents can reach, enforced at runtime, while on personal PCs the safeguards are built into the agent experience. Options range from process and session isolation to WSLc, virtual machines and Windows 365 for Agents.

Codex, GitHub Copilot, OpenClaw, Replit, LM Studio, NVIDIA's OpenShell and Unsloth AI already support MXC. Claude Code, Box, Egnyte, Heidi Health, Hermes Agent, Manus, Perplexity, Raycast, Simular and others are coming. Meta's Muse for Windows, a personal agent whose security design this blog covered, arrives soon as a native app with MXC integration. OpenClaw users on mini desktop PCs get a native Windows gateway.

The hardware

Microsoft says over 2 trillion inferences run locally each month on Copilot+ PCs, and over 40% of laptops built for business are Copilot+ PCs.

The Surface Laptop Ultra, the most powerful Surface yet, has a 15-inch touchscreen and up to 128 GB of unified memory, and Microsoft says it runs models above 120B parameters locally. It ships October 16, as do RTX Spark laptops from ASUS (ProArt P16 and P14), Dell (XPS 16 Creator Edition), HP (OmniBook Ultra 16), Lenovo (Yoga 9n) and MSI (Prestige N16 Flip AI+). A Surface RTX Spark Dev Box ships in the U.S. in November. Microsoft and NVIDIA claim up to 2.1x faster time to first token, 4.3x faster image generation and 6.2x faster video generation than a 16-inch MacBook Pro with M5 Pro; these are vendor figures.

DGX Station for Windows, with the NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip, arrives later this year and runs models above a trillion parameters locally, such as Llama 4 Maverick, Kimi K2.6 and DeepSeek V4 Pro. As a "Scalable Windows Token Factory" it can serve 32 or more simultaneous agents for a team. Dell Pro Precision with GB300 and HP ZGX Fury AI Station follow. Perplexity's local agent already targets RTX Windows PCs.

Search actions and gaming

Windows Search on the taskbar gains thousands of actions, such as "switch to dark mode" or "arrange my windows", reaching Windows Insiders today, with an opt-in Copilot integration in select markets later this year. For gaming, RTX Spark on Windows on Arm runs Gears of War: E-Day, the first title to expand Advanced Shader Delivery, which cuts first-launch shader compilation from minutes to seconds. Call of Duty arrives on RTX Spark in 2027.

Read next