Why Kimi K3’s GPU Capacity Is Exhausted, New Subscriptions Paused, and IPO Looms
Less than a week after Kimi K3’s launch, soaring demand has maxed out Moonlight Dark’s GPU resources, forcing a pause on new and upgrade subscriptions, prompting a split of membership plans, driving six‑fold revenue growth, and accelerating a Hong‑Kong IPO within six months.
Kimi K3 was released less than a week ago, and the surge in usage quickly pushed Moonlight Dark’s GPU capacity to its limits. The company announced that it will pause new member subscriptions to protect the experience of existing users.
According to the official statement, the pause also affects active members who cannot upgrade their compute quota, and all membership packages are shown as sold out.
Some observers questioned whether the shortage is genuine or a marketing stunt, but the mismatch between the rapid increase in model calls after Kimi K3’s launch and the current compute supply provides a direct explanation.
In addition to halting subscriptions, Kimi plans to split its membership offering into two tiers: Kimi Membership for general products such as Kimi Web, App, and Work, and Kimi Code Membership for programming‑focused workflows. This separation allows independent quota, concurrency limits, pricing, and routing of tasks to the most suitable inference resources, addressing the question of how much GPU consumption differs between a casual chat user and a long‑running coding agent paying the same monthly fee.
Financially, Kimi K3’s release drove daily sales up at least six‑fold. The company’s annual recurring revenue rose from about $200 million in April to $300 million in June. Concurrently, Moonlight Dark is advancing a new financing round and preparing a Hong‑Kong IPO that could occur within the next six months, with a potential valuation exceeding $30 billion.
The article notes that the compute shortage is not solely a chip‑count issue; the AI industry must also secure data‑center land, construction permits, power supply, grid connections, server delivery, networking, and cooling. While model demand can spike within days, planning and building new data‑centers typically takes months to years, making supply‑side expansion inherently slower than user‑growth rates. Consequently, AI competition now hinges as much on capital and infrastructure as on model architecture and data.
In summary, the rapid sell‑out of Kimi K3 highlights a broader challenge: delivering enough GPU capacity for popular models is a far more expensive and difficult problem than creating the models themselves.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
