Discussion about this post

User's avatar
Todd Royer's avatar

This is an unusually valuable analysis because it shows why a cheaper open-weight model does not necessarily reduce hardware demand. K3 economizes on active compute by routing each token through only a small group of experts, but the full 2.8 trillion parameter set still has to remain rapidly accessible. The savings occur in FLOPs; the pressure shifts toward HBM capacity, bandwidth, and interconnect.

The widening gap between total and active parameters may be the most important signal here. If model size continues expanding much faster than the number of parameters activated per token, then architecture is not eliminating the infrastructure requirement. It is redistributing that requirement across memory, racks, networking, foundries, packaging, and test. K3 reaching Moonshot’s capacity limit within days seems like real-world confirmation of that direction.

The unresolved question may therefore sit less with the hardware vendors than with the hyperscalers financing the buildout. They either see credible paths toward profitable AI revenue that remain difficult for outside investors to measure, or competition and network effects are compelling them to continue investing before the final revenue model is clear.

Given how early this development remains—and how AI may eventually combine with adjacent technologies such as optimization, robotics, and quantum computing—I wonder whether hyperscalers view the present uncertainty as a reason to slow down, or as the very reason they cannot afford to fall behind.

The market may continue producing corrections whenever investors question the eventual returns. But your analysis suggests those corrections are discounting pressure on the model layer and demand growth in the hardware layer as though they were the same thing, when they may point in opposite directions.

No posts

Ready for more?