NVIDIA CorporationKimi K3's GPU capacity crunch drives demand for Nvidia's accelerators, interconnects, and software stack.
Moonshot AI paused new consumer subscriptions for its Kimi K3 model after demand pushed its GPU clusters close to capacity, echoing a similar crunch that hit DeepSeek V3 in early 2025. The episode underscores a split in AI economics where competition compresses model-level margins while compute remains scarce, benefiting infrastructure providers like Nvidia and Microsoft. Kimi K3’s API costs $3 per million input tokens and $15 per million output tokens, but its token inefficiency—generating about 67% more output tokens than GPT-5.6 Sol on a standardized workload—narrows its per-token discount. SemiAnalysis notes that the 2.8-trillion-parameter model requires more than 1.5 terabytes of HBM capacity, driving heavy GPU, storage, and networking demand that directly benefits Nvidia across its accelerators, interconnects, and software stack. Microsoft, through Azure and Foundry, can route workloads among models and monetize compute and distribution even as model prices fall, with 282 hedge funds holding its shares at the end of Q1 2026.
NVIDIA CorporationKimi K3's GPU capacity crunch drives demand for Nvidia's accelerators, interconnects, and software stack.
Microsoft CorporationMicrosoft's Azure and Foundry benefit from increased demand for compute and model routing as AI models like Kimi K3 strain GPU capacity.
Moonshot AI paused new subscriptions due to GPU capacity constraints, limiting growth.