Google TPU Runs Kimi 57% Faster Than Nvidia GPU! Using DeepSeek Inference Framework
Published on · Sep 26 · Sat Source · 量子位 (CN)

Google TPU Runs Kimi 57% Faster Than Nvidia GPU! Using DeepSeek Inference Framework

A startup founded by core members of the vLLM team ran Moonshot AI's Kimi large model on Google TPU, achieving 57% faster inference speed than on Nvidia GPU. The key was using DeepSeek's open-source inference framework for low-level adaptation and optimization. This breakthrough reveals TPU's architectural advantages in matrix operations and marks a key step in the engineering feasibility of cross-hardware migration for inference frameworks, posing a substantial challenge to Nvidia's CUDA ecosystem moat.

Key Takeaways

  • Key Highlight:A startup founded by core members of the vLLM team ran Moonshot AI's Kimi large model on Google TPU, achieving 57% faster inference speed than on Nvidia GPU. The key was using DeepSeek's open-source inference framework for low-level adaptation and optimization. This breakthrough reveals TPU's architectural advantages in matrix operations and marks a key step in the engineering feasibility of cross-hardware migration for inference frameworks, posing a substantial challenge to Nvidia's CUDA ecosystem moat.
  • Innovation & Tech:Highlights advancements in Google, DeepSeek, TPU, demonstrating rapid progress in model capabilities.
  • Industry Impact:Reported via 量子位 (CN), offering actionable signals for developers and technology leaders.
KeywordsGoogleDeepSeekTPURunsKimiFasterThanNvidia

[Core Event and Technical Overview]

This breakthrough comes from a startup founded by core members of the vLLM development team. Its technical team successfully deployed Moonshot AI's Kimi large model on Google TPU (Tensor Processing Unit), achieving a 57% throughput improvement during inference compared to Nvidia GPU. The core of this optimization does not rely on stacking raw hardware computing power, but rather on deep adaptation of DeepSeek's previously open-sourced inference framework. The framework's native design accommodates cross-hardware abstraction layers and operator fusion capabilities, allowing the TPU's systolic array architecture to fully exert its structural advantages in large-scale matrix multiplication. The entire technical verification process covered a comprehensive reconstruction from weight format conversion and attention mechanism restructuring to KV Cache memory management strategies, proving that the ceiling of large model inference performance depends not only on chip peak FLOPS but also on the depth of software stack excavation of hardware features.

From a technical parameter perspective, Kimi, as a large language model known for its ultra-long context window, puts far more pressure on memory bandwidth and KV Cache management during inference than conventional models. The team completed full-link verification on TPU v5e and v5p instances. The DeepSeek inference framework version used supports dynamic batching and PagedAttention variants, demonstrating better memory access locality under TPU's unified high-bandwidth memory (HBM) architecture than GPU's distributed memory. Regarding open-source licensing, the DeepSeek inference framework uses the MIT license, allowing free modification and distribution for commercial scenarios, providing a compliant and flexible technical foundation for startups to build inference services on non-Nvidia hardware. The model parameter scale verified covered multiple version configurations of Kimi, with stable performance advantages observed across context lengths from 2048 to 128K tokens, indicating that the optimization solution has strong generalization capabilities rather than being an overfitting result for specific scenarios.

[Technical Principles and Core Breakthroughs]

There is a fundamental divergence in underlying architectural design between TPU and GPU. Nvidia GPUs adopt the SIMT (Single Instruction, Multiple Threads) architecture, relying on massive thread parallelism and complex cache hierarchies to hide memory access latency. Its versatility makes it excel at handling irregular computations and branch-intensive tasks, but it has a ceiling in computing power utilization for dense matrix multiplication, the core bottleneck of Transformer inference. TPUs, on the other hand, adopt a Systolic Array architecture, implementing a data-flow-driven computing model through MXUs (Matrix Compute Units). Weight data is passed level-by-level within the array, drastically reducing read/write operations to register files and shared memory. This architecture can run at near-theoretical peak efficiency when processing GEMM operations in Transformers, with significantly better power efficiency than GPUs. When adapting the TPU, the DeepSeek inference framework restructured the Attention computation graph, replacing the traditional CUDA-based FlashAttention implementation with fused operators optimized for the TPU XLA compiler, reducing the materialized storage of intermediate results.

The core technical contribution of the DeepSeek inference framework lies in its hardware abstraction layer design. The framework introduces hardware-agnostic descriptions at the Intermediate Representation (IR) level, allowing upper-level inference logic to be mapped to different backends without major modifications. In TPU adaptation, key engineering work focused on three aspects: first, converting PyTorch's dynamic graph semantics into the static graph compilation flow supported by TPU, eliminating runtime overhead via XLA's Just-in-Time compilation; second, restructuring the memory layout of KV Cache, leveraging TPU's unified HBM address space feature for more efficient page table management, avoiding the extra overhead introduced by PagedAttention on GPUs due to memory fragmentation; third, optimizing the tensor parallelism strategy for TPU's interconnect topology. The 3D Torus interconnect network of TPU v5p showed lower latency fluctuations in all-reduce communication compared to NVLink. Benchmark tests showed that in generation tasks with 2048 tokens input and 512 tokens output, the per-token latency of TPU was reduced to 63.7% of GPU, and the advantage further expanded in long-context scenarios.

In terms of quantized inference, the framework enabled a mixed-precision mode of INT8 matrix computation and BF16 accumulation on TPU, utilizing TPU's native bfloat16 data format support to avoid the precision loss and dequantization overhead common in GPU quantized inference. Compared with Nvidia TensorRT-LLM's W8A8 quantization scheme, the TPU solution controlled accuracy degradation within 0.3% on HumanEval and GSM8K benchmarks, while inference throughput increased by 57.3%. Notably, the team also verified the feasibility of speculative decoding on TPU. By deploying both the draft model and the verification model within the same TPU Pod, leveraging the low-latency characteristics of inter-chip interconnects, an additional 1.8x speedup was achieved. This result is typically difficult to reach theoretical speedup ratios on GPUs due to SM resource contention.

[Industry Background and Competitive Landscape]

From the perspective of global computing power competition, this breakthrough constitutes a substantial technical impact on Nvidia's CUDA ecosystem moat. For a long time, Nvidia relied on the deep moat of its CUDA ecosystem, binding large model training and inference almost exclusively to GPUs, with developers facing huge engineering costs to migrate to other hardware. However, the combination of the DeepSeek inference framework's cross-hardware abstraction capabilities and the vLLM team's deep accumulation in inference optimization proves that in the specific scenario of inference, the marginal cost of hardware migration is rapidly decreasing. Compared with OpenAI's Triton inference server, Anthropic's reliance on TensorRT-LLM, and Google's own Pathways framework, the open-source nature of the DeepSeek framework makes it an important variable in breaking hardware lock-in. Google TPUs previously served mainly internal models (like Gemini, PaLM). This efficient inference of a third-party large model on TPU marks TPU's transition from a closed ecosystem to open competition.

Comparing current mainstream inference frameworks horizontally, vLLM dominates the GPU ecosystem with PagedAttention, SGLang excels in structured generation optimization, while the DeepSeek framework chose the differentiated route of cross-hardware compatibility. In the domestic market, DeepSeek's own open-source models (like DeepSeek-V2/V3 series) are deeply integrated with this framework. The successful deployment of Kimi on TPU further validates the framework's model-agnostic nature. Compared with open-source model ecosystems like Qwen and Llama, Kimi, as a closed-source API service model, having its underlying inference framework publicly adapted means that inference infrastructure is moving towards layered decoupling: the model layer, framework layer, and hardware layer each evolve independently. This trend forms a tripartite win-win situation for Moonshot AI to reduce inference costs, for Google to expand TPU's cloud market share, and for the DeepSeek framework to expand its industry influence.

From a broader industry perspective, the diversified supply of inference computing power is reshaping the investment logic of AI infrastructure. Nvidia GPU's dominant position in training is hard to shake in the short term, but the diversity of inference scenarios—from cloud batch inference to edge real-time inference—provides room for differentiated competition for TPUs, Groq LPUs, AMD MI300s, and even domestic Ascend chips. This result of TPU outperforming GPU by 57% is essentially a local optimal solution under a specific model architecture and specific inference framework optimization, and does not mean that TPU comprehensively surpasses GPU in all scenarios. But it does prove one point: when the optimization depth of the inference framework is sufficient, the cost-performance advantage of specialized architecture chips will be fully released, which will exert long-term pressure on Nvidia's pricing power in the inference market.

[Developer and Industry Implementation Insights]

For developers and enterprise users, the engineering implementation insight of this technical verification is that the selection logic for inference infrastructure is changing. In the past, choosing Nvidia GPUs was almost the default option; the maturity and community support of the CUDA ecosystem made migration risks manageable. Now, the hardware abstraction layer provided by the DeepSeek inference framework allows developers to switch the inference backend from GPU to TPU without modifying upper-level application code. The actual integration complexity mainly focuses on two aspects: first, model weight format conversion, requiring conversion from Safetensors format to TF Record or equivalent formats supported by TPU, for which the DeepSeek framework provides an automated conversion toolchain; second, environment configuration, where the TPU runtime environment relies on Google Cloud TPU VM and JAX/XLA compilation stack, differing from the PyTorch CUDA environment, but the framework has encapsulated a unified inference API interface, keeping upper-level calling code unchanged.

Regarding hardware memory overhead, TPU's HBM capacity is comparable to GPU (v5p equipped with 80GB HBM), but the unified memory architecture makes KV Cache management more efficient, improving memory utilization by about 22% in 128K context scenarios. In terms of commercial implementation value, comparing Google Cloud TPU's on-demand billing price with Nvidia A100/H100 cloud instance prices, factoring in the 57% throughput improvement, the per-token inference cost can be reduced to about 40% of the GPU solution. Migration costs mainly come from the operations team's learning curve for the TPU environment and the offline conversion process of model weights. It is estimated that for teams with existing vLLM deployment experience, the complete migration cycle is about 2-3 weeks. For online services with large volumes of inference requests and sensitivity to latency (such as Kimi's own conversational API), this input-output ratio is significantly attractive.

Evaluating from the toolchain support perspective, the DeepSeek inference framework has provided an OpenAI-compatible REST API interface, supporting mainstream features like streaming output, function calling, and JSON mode, making business-side changes almost zero when migrating from existing GPU inference services to TPU backends. The framework also integrates Prometheus monitoring metric export and distributed tracing support, facilitating operations teams to reuse existing observability infrastructure in TPU environments. Notably, the current solution's TPU adaptation for multimodal inference (such as vision-language models) is still being promoted; the optimization results for pure text scenarios cannot yet be directly transferred to multimodal pipelines, which is a key focus for subsequent engineering implementation.

[Comprehensive Review and Key Points]

The core insight of TPU inference of Kimi outperforming GPU in speed is that the competition in large model inference performance has evolved from a simple comparison of hardware peak computing power to a systems engineering of software-hardware collaborative optimization. The combination of the DeepSeek inference framework's cross-hardware abstraction capabilities and the vLLM team's engineering experience in inference engine optimization successfully unleashed the structural advantages of TPU's systolic array architecture in matrix computation-intensive tasks. This indicates that in the next 1-2 years, the strategic importance of inference frameworks as the middleware connecting models and hardware will continue to rise, and teams mastering framework optimization capabilities will take the initiative in computing power bargaining. Hardware diversification in inference scenarios will accelerate. Nvidia's monopoly in training and its dominant position in inference will diverge, and the market share of specialized inference chips is expected to increase from the current less than 15% to over 30%.

Judging from the industry impact, this breakthrough has a profound impact on three types of entities. For large model vendors, inference costs account for 60-80% of operating expenses. The diversification of hardware selection will significantly improve gross margins. Companies like Moonshot AI are expected to leverage this to reduce API call prices to expand market share. For chips.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.

Industry Insights & Analysis

As artificial intelligence rapidly evolves, breakthroughs surrounding Google, DeepSeek, TPU, Runs are shifting toward scalable, robust real-world implementations.

Driven by both open-source ecosystems and proprietary model architectures, the integration between compute optimization, data engineering, and agentic workflows is accelerating. This development provides a strategic benchmark for upcoming AI tooling and developer workflows.