Physix Frontier · News Briefing Card (IT Home · Oct 8, 2026)

Microsoft adds local AI model inference to GitHub Copilot

KEY FACTS

  • Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month.
  • Developers can switch automatically or manually between cloud models and on-device models.
  • Users can set preferences in the CLI, the Copilot app, and VS Code.
  • Local models support Windows ML's MAI Code 1.1 Flash or OpenAI-compatible endpoints.
  • MAI Code 1.1 Flash has 137 billion total parameters and 6.8 billion active parameters.

KEY DATA

137BMAI Code 1.1 Flash total params
6.8BMAI Code 1.1 Flash active params
40-63 tokens/sDecoding throughput
2K-256K tokensTest prompt length range

PHYSIX OBSERVATION

Local inference frees Copilot from cloud dependency, a real boon for privacy-sensitive and network-constrained developers. But 40-63 tokens/s throughput is only enough for lightweight completions; complex refactoring still has to rely on the cloud. Microsoft's move looks more like staking out a position at the on-device AI gateway than replacing the cloud.

Source: IT Home report