Coming by the end of the month, GitHub Copilot will determine when a task is best handled by on-device intelligence and when it should leverage cloud-scale models. Microsoft is introducing local models and sandboxed tools across Windows and GitHub Copilot, deploying Microsoft Execution Containers to secure developer sessions without requiring separate virtual machines.
Developers working with intelligent agents need both choice and clear boundaries. According to Microsoft, they require technologies that make it straightforward to select the right model speed, performance, and cost profile for individual tasks. To address these needs, Microsoft and GitHub are rolling out local model selection alongside isolated execution controls across Windows and development tools.
Local Model Performance on Surface Laptop Ultra
For coding agents, smaller models must still complete complex tasks reliably. Microsoft notes that a smaller footprint matters only if changes in code quality and tool use remain well-understood. The initial shipping version of MAI Code 1.1 Flash on Surface Laptop Ultra reaches a peak memory usage of 75.5GB at a 256k context length. At 64k and 128k context lengths, prompt-processing throughput hits 923.5 and 769.8 tokens per second, respectively.
The quantized version of MAI Code 1.1 Flash used on devices achieves an 80% reduction in size down to 53GB, while retaining capability compared to the Bfloat16 cloud variant. In benchmark comparisons tested on October 5, 2026—utilizing mixed-precision quantization at approximately 3.3 bits per weight with DFlash2 sliding-window speculative decoding and a Windows ARM64 llama.cpp CUDA runtime—the quantized local model scored 70.80% on SWE-Bench Verified and 66.29% on Terminal-Bench 2.1. By comparison, the non-quantized MAI Code 1.1 Flash scored 72.6% on SWE-Bench Verified and 62.9% on Terminal-Bench 2.1, while the comparison model, Unsloth’s GPT-OSS-120B GGUF, scored 32.0% and 23.6% on those same benchmarks.
Deployment in GitHub Copilot and VS Code
GitHub Copilot is introducing two ways to use local models across the GitHub Copilot CLI, Copilot app, and VS Code. Explicit local-model selection supports workflows requiring a specific provider, model, or endpoint. Developers can select MAI Code 1.1 Flash via the Windows ML provider or connect GitHub Copilot to OpenAI-compatible local endpoints to choose from whatever models those endpoints expose.
Additionally, GitHub Copilot will determine automatically whether a given task fits best on device or on cloud-scale infrastructure. GitHub also offers frontier models alongside orchestrators such as Project HydraFusion, which balances performance, cost, and latency across one or multiple models.
Sandboxing and Microsoft Execution Containers
Moving inference directly onto a device does not change the access permissions of an agent’s shell commands, which normally inherit the account running them. To secure interactive and non-interactive agentic coding sessions, Windows developed Microsoft Execution Containers (MXC). MXC is an open-source library from the Windows team that translates policy into native operating-system controls, controlling access to files, networks, credentials, system capabilities, and execution paths regardless of which model requests the work.
- On Windows, it uses the BaseContainer tier of the ProcessContainer backend.
- On macOS, it uses Seatbelt.
- On Linux, it uses bubblewrap.
When sandboxing is enabled, shell commands and local Model Context Protocol servers and language servers run inside the process boundary by default. Built-in file tools execute within GitHub Copilot itself, where the agent harness checks requests against effective policy rather than relying on OS-enforced child-process isolation. Remote Model Context Protocol servers remain outside the local process sandbox, meaning connection policies are checked in process.