Physix Frontier · News Briefing Card (enbrief · Oct 8, 2026)

Microsoft adds local AI model inference to GitHub Copilot

KEY FACTS

  • Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month.
  • Developers can switch automatically or manually between cloud models and on-device models.
  • Users can set preferences in the CLI, the Copilot app, and VS Code.
  • Local models support Windows ML's MAI Code 1.1 Flash or OpenAI-compatible endpoints.
  • MAI Code 1.1 Flash has 137 billion total parameters and 6.8 billion active parameters.

KEY DATA

137 billionMAI Code 1.1 Flash total parameters
6.8 billionMAI Code 1.1 Flash active parameters
40-63 tokens per secondDecoding throughput
2K-256K tokensTest prompt length range

PHYSIX OBSERVATION

Local inference frees Copilot from cloud dependency, a real win for privacy-sensitive and network-constrained developers. But 40-63 tokens per second is only enough for lightweight completions; complex refactoring still has to rely on the cloud. Microsoft's move looks more like staking out a position at the on-device AI entrance than replacing the cloud.

Source: enbrief original report ↗

Microsoft adds local AI model inference to GitHub Copilot | Physix Frontier Briefing