SenseTime’s Galaxy Project targets domestic AI chip scale-up
SenseTime launched the "Galaxy Project," a collaborative initiative with nearly 20 domestic partners to scale China's indigenous AI chip infrastructure. The company claims its heterogeneous hybrid inference technology significantly boosts Model FLOPs Utilization and cost-effectiveness compared to domestic homogeneous setups and Nvidia H-series parts. Strategic partnerships extend beyond traditional chips to include space computing, optical computing, and quantum computing applications, with a sa
Analysis
TL;DR
- SenseTime launched the "Galaxy Project," a collaborative initiative with nearly 20 domestic partners to scale China's indigenous AI chip infrastructure.
- The company claims its heterogeneous hybrid inference technology significantly boosts Model FLOPs Utilization and cost-effectiveness compared to domestic homogeneous setups and Nvidia H-series parts.
- Strategic partnerships extend beyond traditional chips to include space computing, optical computing, and quantum computing applications, with a satellite constellation planned for launch in 2026.
- SenseTime introduced a new "Tokens Per Watt" metric and an automated resource scheduling agent to optimize energy efficiency and electricity costs in AI data centers.
- The ecosystem includes major domestic chipmakers like Huawei Ascend, Cambricon, and Biren Technology, aiming to create a closed-loop from chip-level technology to commercial deployment.
Why It Matters
This initiative highlights the accelerating consolidation of China's domestic AI hardware supply chain, moving from isolated chip development to integrated ecosystem solutions. For global observers, it signals a maturing alternative infrastructure that could reduce reliance on Western semiconductor technology for large-scale enterprise AI deployments within China.
Technical Details
- Heterogeneous Hybrid Inference: SenseTime reports an 85–152% increase in Model FLOPs Utilization on mainstream domestic chips and claims inference cost-effectiveness 1.25x that of Nvidia’s H-series, though these figures lack independent verification.
- Full-Stack Adaptation Layer: A software layer designed to span models, frameworks, operators, and hardware, enabling workload migration across different domestic chip vendors without extensive code rewrites.
- Performance Claims: Specific optimizations include a threefold reduction in prediction time for AI4S protein workloads and a 93% multi-card parallel acceleration ratio for AIGC video generation using DiT models.
- Energy Efficiency Metrics: Introduction of "Tokens Per Watt" as a benchmark, coupled with a Computing-Power Collaboration Agent that achieves claimed 80% increases in token output per electricity cost unit and 96% accuracy in load prediction.
- Future Infrastructure: Plans for a "token factory," five "10,000-calorie" computing clusters, and a space computing constellation involving thousands of satellites by 2030.
Industry Insight
- Supply Chain Resilience: The breadth of partners suggests that domestic AI infrastructure in China is becoming more standardized and interoperable, potentially lowering integration barriers for enterprises adopting local hardware.
- Verification Gap: Stakeholders should approach performance and efficiency claims with caution, as the significant disparity between controlled test environments and production realities remains unaddressed by independent benchmarks.
- Diversification of Compute: The move into space and optical computing indicates a long-term strategy to bypass terrestrial infrastructure limitations, suggesting that future AI scaling may involve non-traditional distributed computing paradigms.
Disclaimer: The above content is generated by AI and is for reference only.