TLDR: Chinese AI chip startup Houmo.AI anticipates a significant paradigm shift in AI computing, with its leadership predicting that a vast majority, specifically 90%, of generative AI inference will eventually be processed directly on edge devices rather than relying solely on cloud infrastructure. The company is developing advanced chips with a novel storage-computing integration architecture to facilitate this transition, with a new ‘Tianxuan’ architecture-based chip slated for release in 2025 to accelerate large model deployment at the edge.
Chinese AI chip innovator, Houmo.AI, is at the forefront of a predicted revolution in artificial intelligence computing, asserting that the future of generative AI inference lies predominantly at the edge. Ni Xiaolin, Vice President of Houmo Intelligence, has articulated a vision where an estimated 90% of generative AI inference operations will transition from centralized cloud environments to direct processing on edge devices. This strategic shift is poised to redefine how AI applications are deployed and utilized across various sectors.
The company, which specializes in AI chips built on an innovative storage-computing integration architecture, is actively developing solutions to enable this transition. This architecture is designed to achieve a high degree of integration between storage and computing units, allowing computations to be performed directly within storage units. The result is a significant reduction in power consumption and a substantial increase in bandwidth, critical factors for efficient edge AI processing.
In line with this forward-looking strategy, Houmo.AI plans to launch its latest chip in 2025. This new chip, based on the next-generation ‘Tianxuan’ architecture, is expected to deliver significantly improved performance, specifically engineered to accelerate the deployment of large AI models on edge devices. The company is not only focusing on chip development but also providing a range of standardized product forms, including the Limou®️ LM30 intelligent accelerator card (PCIe) and the Limou®️ SM30 computing module (SoM), to facilitate rapid deployment for AI device solution providers and manufacturers.
Also Read:
- China’s AI Chip Strategy Shifts Following US Commerce Secretary’s Remarks on Nvidia Exports
- Specialized Storage Solutions Propel Enterprise Generative AI into Production
This anticipated shift comes amidst the rapid evolution of the AI landscape, often referred to as the AI 2.0 era, marked by the proliferation of advanced large models. While cloud models continue to scale in size and parameters, exploring the boundaries of general intelligence, Houmo.AI’s focus underscores a parallel and equally vital development path: bringing powerful AI capabilities closer to the data source, enhancing real-time processing, privacy, and efficiency.


