Microsoft Embraces AMD Helios: Diversifying AI Inference and Investment Opportunities
When cloud giants start diversifying AI chip deployments, is NVIDIA's moat still solid? Microsoft announced deploying AMD Helios rack systems on Azure specifically for AI inference. This is not just a choice of technical route, but a re-evaluation of the risk-reward ratio.
From an asset allocation perspective, chip procurement for cloud infrastructure is evolving from single-supplier to multi-source. Microsoft's move essentially builds a "portfolio"—using NVIDIA H100/B200 to defend absolute advantage in training, and using AMD Helios to achieve cost optimization and supply chain security on the inference side. This strategy is similar to configuring targets with different betas in a stock portfolio, pursuing upside returns while reducing tail risk.
Comparative Analysis: NVIDIA Solution vs. AMD Helios Solution
| Dimension | NVIDIA (H100/B200) | AMD Helios (MI300X Series) |
|---|---|---|
| Training Performance | Absolute lead, CUDA ecosystem barrier is extremely high | Lags by approx. 30-50%, ROCm ecosystem still catching up |
| Inference Cost-Performance | High but expensive | 15-20% higher inference throughput at same power consumption, lower price |
| Supply Stability | Constrained by TSMC CoWoS capacity, long lead times | Uses more mature packaging, greater supply elasticity |
| Software Ecosystem | CUDA mature, native PyTorch support | ROCm gradually improving, but requires additional adaptation work |
Key Observation: Microsoft's deployment strategy introduces AMD in inference scenarios because inference has lower dependency on software ecosystems than training. Inference tasks are more about linear execution of matrix operations rather than complex gradient calculations, resulting in lower stickiness to specific CUDA operators.
Technical Details: Architectural Design of Helios Rack System
AMD Helios Rack System Core Parameters (Based on MI300X)
- Each rack integrates 8 MI300X accelerators, totaling 1536 CUs (Compute Units)
- VRAM: 192GB HBM3 per GPU, total VRAM 1.5TB
- Memory Bandwidth: 5.2 TB/s per GPU
- Interconnect: Infinity Fabric 4.0, bandwidth 448 GB/s
- Inference Scenario Optimization: Supports FP8/FP16 mixed precision, INT8 quantization
In comparison, Microsoft Azure's traditionally deployed NVIDIA H100 racks (based on HGX H100) are 8-card configurations with total VRAM capacity of 1.6TB (80GB 8 2.5? Actually H100 is 80GB, 8 cards total 640GB), but memory bandwidth is slightly lower (3.35TB/s per GPU). In inference scenarios, Helios's larger VRAM capacity and bandwidth help handle long contexts for large models, reduce model sharding requirements, and lower latency.
[!note] From an investment perspective, Microsoft choosing AMD Helios over NVIDIA solutions is driven by optimizing the risk-reward ratio. The inference market is expected to reach 40% of total AI chip demand by 2025, growing faster than training. In this niche market, AMD's cost-performance advantage can translate into improved gross margins for Azure. Assuming inference costs drop by 20%, with Azure AI inference annual revenue at $5 billion, this saves approx. $1 billion in costs, directly contributing to profit.
Business Model and Competitive Moat Assessment
Microsoft's Business Model: Cloud providers offer AI compute infrastructure, charging by token consumption or instance duration. The key is lowering unit compute costs while maintaining customer stickiness. By introducing AMD Helios, Azure can launch cheaper inference instances to attract price-sensitive customers (e.g., indie developers, ad recommendation systems), while serving high-end customers (e.g., large model training, scientific research) with NVIDIA instances. This tiered pricing strategy maximizes customer lifetime value.
AMD's Competitive Moats:
- Hardware Level: CDNA architecture approaches NVIDIA in inference efficiency, especially in integer operations and sparsity support.
- Supply Level: Closer cooperation with TSMC, relatively ample CoWoS-L packaging capacity.
- Ecosystem Level: ROCm support for PyTorch/TensorFlow is accelerating, and Microsoft promises engineering support.
- Key Risk: CUDA's inertia effect remains powerful; many AI frameworks default to NVIDIA optimization, requiring extra engineering effort for AMD.
NVIDIA's Moat: Deep binding of the CUDA ecosystem. Even if AMD hardware offers better cost-performance, enterprise migration costs—including recompiling models, adjusting inference frameworks, and testing stability—may offset hardware savings. But clients at Microsoft's level have enough capability to bear migration costs, and their proprietary AI models (like the Phi series) have already begun adapting to AMD.
Valuation Logic and Investment Judgment
Currently, AMD's PE (TTM) is approx. 170x, NVIDIA approx. 75x (data as of July 2024, needs actual verification). On the surface, AMD looks more expensive, but considering its revenue growth (Q2 2024 data center revenue grew 115% YoY) and expansion potential in the inference market, the premium is justified. From an asset allocation perspective, AMD's risk-reward ratio in AI inference is superior to NVIDIA because:
- The inference market is not yet fully monopolized by NVIDIA; AMD has significant room for share gain.
- Diversification strategies by cloud vendors like Microsoft and Google are long-term trends; AMD is a primary beneficiary.
- If AMD's ROCm ecosystem continues to improve, its valuation can shift from "hardware supplier" to "ecosystem platf
Original link: https://www.ithome.com/0/979/235.htm
Physix Frontier