Physix Frontier · News Briefing Card (IT Home · Oct 8, 2026)
Microsoft adds local AI model inference to GitHub Copilot
KEY FACTS
- Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month.
- Developers can switch automatically or manually between cloud models and on-device models.
- Users can set preferences in the CLI, the Copilot app, and VS Code.
- Local models support Windows ML's MAI Code 1.1 Flash or OpenAI-compatible endpoints.
- MAI Code 1.1 Flash has 137 billion total parameters and 6.8 billion active parameters.
KEY DATA
137BMAI Code 1.1 Flash total params
6.8BMAI Code 1.1 Flash active params
40-63 tokens/sDecoding throughput
2K-256K tokensTest prompt length range
PHYSIX OBSERVATION
Local inference frees Copilot from cloud dependency, a real boon for privacy-sensitive and network-constrained developers. But 40-63 tokens/s throughput is only enough for lightweight completions; complex refactoring still has to rely on the cloud. Microsoft's move looks more like staking out a position at the on-device AI gateway than replacing the cloud.
Source: IT Home report
Physix Frontier