Moonshot AI Suspends Kimi Subscriptions Due to Kimi K3 Computing Strain

The company is currently prioritizing existing users while expanding computing capacity to manage the surge in request volume. The company stated that it is dedicated to ensuring that the rights and interests of existing subscribers remain unaffected by the current infrastructure constraints.

Kimi K3 Launch Triggers Unforeseen Computing Infrastructure Strain

Kimi K3 Performance and Infrastructure Strain

The decision to halt new subscriptions comes just three days after the July 16 launch of Kimi K3. The model, which features 2.8 trillion parameters, is currently the largest open-source model globally. According to independent testing by Artificial Analysis, Kimi K3 has achieved a comprehensive intelligence score of 57, placing it in the top tier of global models alongside Claude Opus 4.8 and GPT-5.5, while trailing only Claude Fable 5 and GPT-5.6 Sol. The model’s release also drove a significant financial impact, with 月之暗面 recording its largest single-day increase in annualized revenue on July 17.

The rapid adoption rate pushed the company’s existing computing clusters to their limits. To manage the current strain, 月之暗面 announced that it will separate Kimi’s primary features—including KimiWeb, Kimi APP, and Kimi Work—from Kimi Code capabilities for future subscribers. This move is intended to match hardware resources more precisely to specific user needs. The company is actively working on infrastructure expansion and plans to reopen subscriptions in stages as new computing capacity becomes available.

月之暗面 Adopts Capability-Driven Pricing for Kimi K3

Market Positioning and Pricing Strategy

The release of Kimi K3 marks a strategic shift for the company, moving away from low-price competition toward a model based on capability-driven pricing. K3’s pricing structure is set at $0.30 for cache-hit inputs, $3.00 for cache-miss inputs, and $15.00 per million tokens for output. These prices are notably higher than those of its predecessor, Kimi K2.6, as well as competitors like GLM-5.2 and MiniMax-M3. While the output pricing is now comparable to GPT-5.6 Terra, it remains lower than that of GPT-5.6 Sol and Claude Fable 5. According to Huatai Securities, the increase in intelligence directly supports higher token pricing, thereby strengthening the commercialization capabilities of domestic models.

China’s Moonshot AI Releases Kimi K3 | Nvidia Expands In Japan | SpaceX Stock Falls | CNBC TV18

Technical optimizations are central to this commercial strategy. 月之暗面 has implemented a Mooncake decoupled inference architecture, which the company claims achieves a cache hit rate of over 90% for coding workloads. By utilizing techniques such as Prefill and Decode separation, KDA prefix caching, static expert parallelism, key path host-free synchronization, and expert load balancing, the platform aims to maximize cluster efficiency during long-context and high-sparsity MoE (Mixture-of-Experts) scenarios. As noted by analysts, the industry is shifting from simple token price competition toward a comprehensive struggle over cache reuse, memory management, computing power scheduling, and cluster throughput efficiency.

Global Investors and Analysts Evaluate China’s AI Progress

Industry and Investor Reactions

The Kimi K3 launch has sparked significant discussion among global investors, with many drawing parallels to the DeepSeek moment of early 2025. Dean W. Ball, OpenAI’s head of strategic future and a former White House AI policy advisor, remarked on the model’s performance in coding scenarios, noting that it stands on par with the top-tier public models available as of the first quarter of 2026. However, Ball also noted that the model is very 'token hungry' and consumes significant resources during inference.

Financial analysts view the development as a broader indicator of China’s progress in the AI sector. Gary Yu of Morgan Stanley suggested that the positive feedback for K3 demonstrates that Chinese large language models are effectively closing the gap with U.S. leaders in scale, performance, and commercial viability. Similarly, Bernstein analyst Robin Zhu described the release as a “home run,” emphasizing that the pace of innovation within the Chinese AI market is increasingly difficult for global investors to overlook. Zhu noted that the launch confirms that China’s top AI labs have the capacity to maintain pace with the global state-of-the-art.

The current scarcity of high-quality, low-cost token production systems remains the primary challenge for the industry as it transitions toward an agent-based model of computing. Kevin Kelly, the futurist and "father of the Silicon Valley spirit," also commented at the WAIC that the development of open-source models in China is a positive attempt that provides a competitive advantage in managing token costs.

Leave a Comment