The next-generation flagship general-purpose GPU, the Tiankai 300, was formally launched by ILUVATAR COREX (ASX: ILU) during the 2026 World Artificial Intelligence Conference. As the inaugural product based on the company's new self-developed architecture, the Tiankai 300 is designed to align with the evolution of AI from "information intelligence" and "action intelligence" to "exploration intelligence." It delivers comprehensive upgrades in chip architecture, key operators, system scalability, and software ecosystem, aiming to match or surpass the performance efficiency of the mainstream international Hopper architecture solution, thereby providing high-quality computing power to support the large-scale application of AI.
Built on a SIMT general-purpose computing architecture, the Tiankai 300 efficiently supports various computation types including scalar, vector, and tensor operations. It incorporates optimizations for critical technologies such as Attention, Mixture of Experts (MoE), AF separation, PD separation, and large-scale system expansion. These enhancements further improve the conversion efficiency of underlying computing power into actual token output and business value. Under complex workloads, the product balances computational efficiency, task latency, and Total Cost of Ownership, meeting the comprehensive demands for generality, performance, and cost in large-scale AI applications.
Core Design Principles for Evolving AI
The head of AI and accelerated computing at ILUVATAR COREX stated that artificial intelligence is progressing from "information intelligence," which understands and generates information, to "action intelligence" with planning and execution capabilities, and further extending into "exploration intelligence" capable of modeling, predicting, and exploring the unknown. These three intelligence types impose different computational requirements, forming the central design philosophy for the Tiankai 300.
Enhancing Large Model Training and Inference for Information Intelligence
As model parameter sizes expand, context lengths increase, and inference requests grow, computing platforms must enhance training throughput while also considering inference response times and usage costs. Focusing on core aspects of large model computation, the Tiankai 300 places significant emphasis on optimizing Attention and MoE. For Attention computation, by improving the parallel efficiency of matrix and vector calculations and reducing data read/write waits and redundant computations, the product achieves over 90% Attention efficiency across different precisions, outperforming mainstream Hopper architecture solutions. In long-sequence tasks, it utilizes computational resources more fully, boosting training and inference efficiency. In scenarios with a 64k context length, the Tiankai 300's Attention efficiency is 10% higher than the Hopper solution.
Addressing expert scheduling and data routing losses in MoE architectures, the Tiankai 300 employs co-optimization of algorithms and hardware to improve expert and compute unit utilization. In DeepSeek V4 MoE scenarios, computational efficiency surpasses 70%, with overall training performance and cost-effectiveness exceeding the Hopper solution. For inference scenarios, it can handle more concurrent requests while maintaining the same response experience, completing more inference tasks with equivalent computational investment, delivering an average efficiency improvement of 10% over the Hopper solution.
Boosting Agent Responsiveness and Execution for Action Intelligence
Agents must not only understand goals and plan tasks but also continuously invoke tools and execute actions, demanding higher performance for first-token latency, generation speed, and multi-card communication. To address this, the Tiankai 300 strengthens PD separation capabilities, applying differentiated optimizations for the context-processing Prefill stage and the token-by-token generation Decode stage, enabling Agents to both "think fast" and "act fast."
In tests with mainstream models like Qwen, GLM, Kimi, and DeepSeek, the Tiankai 300 reduces first-token latency by approximately 20% compared to mainstream Hopper architecture solutions. For the high-frequency, multi-card, small-data-volume communication typical of the Decode stage, the product reduces average communication latency by about 13% by lowering protocol overhead and optimizing link paths and data flow mechanisms. In Decode tests with models like GLM 5.2, efficiency is 10% higher than the Hopper solution, providing more timely and fluid interactive experiences for Agent applications such as intelligent programming, office assistants, and business automation. Both Prefill and Decode efficiencies within PD separation surpass the Hopper solution, representing a comprehensive breakthrough.
Supporting Diverse Algorithm Evolution for Exploration Intelligence
In fields like new materials discovery, weather prediction, digital twins, and world models, AI participates in complex system modeling, experimental validation, and future forecasting. The computational paradigm thus expands from relatively concentrated large model workloads to diverse tasks including matrix solving, dynamic programming, sparse computation, and simulation.
Adhering to the principle of "returning to the essence of computation," the Tiankai 300 is not limited to the Transformer architecture. It efficiently supports scalar, vector, and tensor computations, covering various precisions like FP4 and FP8, and instructions such as MMA, DPX, and FMA. Beyond Attention, FFN, and MoE, it can also adapt to tasks like matrix solving, dynamic programming, sparse computation, and digital twins. This provides room for optimizing existing algorithms and evolving future computational paradigms, helping clients reduce real-world experimental costs and trial-and-error risks through virtual simulation. In the Cosmos3 world model, the Tiankai 300's inference efficiency is also 10% higher than the Hopper solution.
Beyond the Chip: Unleashing Value Through Foundational Innovation and Full-Stack Ecosystem
Supporting these three intelligence types are continuous innovations in the Tiankai 300's underlying architecture. The ixSMEX feature increases data reuse to reduce redundant memory access; ixDPX compresses computations that originally required six instructions in dynamic programming into a single instruction; ixTrans utilizes lossless matrix transposition to reduce storage conflicts and video memory overhead, further unleashing the chip's computational efficiency.
ILUVATAR COREX believes the long-term value of a general-purpose GPU stems not only from the performance of a single chip but also from building a universal, functional, and easy-to-use ecosystem. Since 2018, the company has continuously refined its software stack, developing nearly a hundred acceleration libraries and establishing a tool system covering communication, compilation, drivers, quantization, and performance analysis. The Tiankai 300 concurrently supports globally mainstream large model frameworks, models, and key acceleration components, reducing clients' redundant investment in migration, adaptation, and performance tuning.
The Tiankai 300 is now ready for large-scale application and is undergoing deep adaptation with domestic cloud providers, server manufacturers, interconnect ecosystems, and super-node systems. It supports scalable deployment from single machines to large-scale clusters. The Tiankai 300 is not an isolated chip product; it is supported by an ecosystem comprising hardware, software, systems, and partners. Moving forward, ILUVATAR COREX will collaborate with more industry chain partners to enhance its general-purpose GPU products and ecosystem, empowering an intelligent society with high-quality computing power and promoting the transformation of industries by artificial intelligence.