
OpenAI Launches First Self-Developed Inference Acceleration Chip Jalapeño
OpenAI released its first self-developed inference chip, Jalapeño, built in collaboration with Broadcom and Celestica, tailored for models like ChatGPT and Codex. The project was led by former Google TPU core member Richard Ho, completing design to tape-out in just 270 days, setting the fastest record for high-performance ASICs. The chip improves actual utilization by optimizing data movement and resource balance, planned for deployment by the end of 2026, with performance per watt significantly better than existing solutions.
1 Late last night, OpenAI released a "chili pepper"
June 24, 2026, OpenAI jointly launched Jalapeño with Broadcom (Broadcom) — a custom AI chip optimized specifically for large language model inference. No press conference, no CEO appearance, just a brief announcement on the official website, but everyone in the AI industry knows: OpenAI has revealed its trump card.
Jalapeño means "chili pepper" in Spanish. The choice of name is interesting — it implies this chip is not for the "main course" (training), but for "seasoning" (inference). But don't be fooled by the name. Jalapeño's significance goes far beyond a single chip; it is a landmark move for OpenAI's transformation from a "model company" to a "full-stack AI company". From this moment on, OpenAI is no longer just a customer of NVIDIA, but a potential competitor.
"Jalapeño is not a chip; it is OpenAI's manifesto for transforming from a 'model company' to a 'full-stack AI company'."
2 Why an inference chip, not a training chip?
OpenAI's choice of an inference chip as the entry point for self-developed hardware is a strategically calculated decision.
Training chips (such as NVIDIA H100/B200) are "nuclear weapons" for large-scale parallel computing, requiring the processing of massive matrix operations, with extremely high demands on bandwidth, interconnects, and power consumption. NVIDIA took twenty years to build technical barriers in this field, which OpenAI cannot surpass in the short term. But the logic for inference chips is completely different: inference is a fine job of "a needle", pursuing low latency, low power consumption, and low cost, rather than stacking computing power.
More importantly, OpenAI possesses the world's largest AI inference traffic pool. ChatGPT serves 230 million users weekly, and the daily average inference call volume for the GPT-5.5 series is an astronomical number. This means OpenAI can validate Jalapeño's performance on its own inference load without relying on external customer feedback. This "self-production and self-use" model allows the chip's iteration speed to far exceed that of traditional chip companies.
🌶️ Key Information on Jalapeño
Positioning: Custom chip optimized specifically for large language model inference
Partner: Broadcom, a global leading custom chip design company
Significance: OpenAI's first release of self-developed hardware, transforming from a "model company" to a "full-stack AI company"
Strategic Logic: Reduce dependence on NVIDIA GPUs, use own inference traffic to feed back into chip iteration
Industry Impact: May reshape the cost structure of the AI inference market, impacting NVIDIA's inference chip market share
3 Broadcom's Role: OpenAI's "Chip Arms Dealer"
Another key detail of Jalapeño is the partner — Broadcom. Broadcom is the world's largest custom chip (ASIC) design company; behind Google's TPU and Meta's inference chips, there is Broadcom's shadow. Choosing Broadcom instead of building a chip team in-house indicates that OpenAI has adopted a "light asset, heavy design" strategy: OpenAI defines the chip architecture and requirements, while Broadcom is responsible for engineering implementation.
The cleverness of this strategy lies in avoiding the "novice trap" of the chip industry. Chip design is not software engineering; a single tape-out can cost tens of millions of dollars, and one bug could lead to a six-month project delay. By collaborating with Broadcom, OpenAI gains the performance advantages of a custom chip while avoiding the high trial-and-error costs of building an in-house chip team. This is also why Jalapeño could move from concept to product in a relatively short time — Broadcom provides a mature engineering process already validated by Google's TPU.
4 What Does This Mean for NVIDIA?
The release of Jalapeño is not good news for NVIDIA, but it is not a "catastrophe" either.
In the short term, Jalapeño only affects OpenAI's own inference load and will not directly impact NVIDIA's revenue. NVIDIA's inference chip business was not as profitable as its training chips anyway — H100/B200 training orders are the foundation. But in the long term, Jalapeño's demonstration effect cannot be ignored: if OpenAI can significantly reduce costs through self-developed inference chips, Google, Meta, Microsoft, and Amazon will accelerate their pace of self-developed chips. NVIDIA's market share in the inference market may be gradually eroded over the next 3-5 years.
More critically, Jalapeño exposes a structural weakness in NVIDIA's business model: when your largest customer is also your competitor, your pricing power is unreliable. NVIDIA's GPU gross margin is as high as 70%+, which is unsustainable in any industry. Jalapeño is a clear signal from OpenAI to NVIDIA: we can build it ourselves.
"When your largest customer is also your competitor, your pricing power is unreliable."
5 Implications for Chinese AI Chips
The release of Jalapeño has three direct implications for the Chinese AI chip industry:
First, inference chips are a more cost-effective breakthrough point
Chinese AI chip companies have long struggled in the narrative of "benchmarking against NVIDIA" — Cambricon, Hygon, Moore Threads are all trying to create training chips that can replace A100/H100. But Jalapeño proves a different path: do not fight NVIDIA head-on in the training track, but enter from inference. Inference chip design is less difficult, customer needs are clearer, and the commercialization path is clearer. For Chinese companies with limited computing power, this is a more pragmatic road.
Second, the "vertical integration" of AI companies is accelerating
OpenAI making chips, Google making TPU, Meta making MTIA, ByteDance investing in chip companies — this wave of "AI companies developing their own chips" indicates that the ultimate form of AI competition is vertical integration. Whoever masters the full-stack capability from chip to model to application will have pricing power and cost advantages. If Chinese AI companies only make models and not chips, they may face the dilemma of "gaining reputation but losing money" in the future.
Third, the Broadcom model is worth attention
Broadcom's role as a "chip arms dealer" provides a reference for Chinese chip design service companies. AI companies do not need to build their own chip teams; they only need to find reliable chip design partners. This model of "AI companies defining requirements, chip companies responsible for implementation" may become the industry standard.
6 In Conclusion: A Chili Pepper, A Grand Game
The release of Jalapeño appears on the surface to be OpenAI launching an inference chip, but in reality, it is a reshuffling of the power structure in the AI industry. It marks that AI competition has moved from the "battle of models" to the "battle of full-stack" — chips, models, and applications are all indispensable.
For OpenAI, Jalapeño is the first step toward achieving "AI inference costs approaching zero". For NVIDIA, this is a warning bell that customers are becoming competitors. For the Chinese AI industry, this is a clear signal: companies that make models will sooner or later make chips. Model companies that do not make chips will eventually become "laborers" for model companies that do make chips.
A chili pepper, a grand game. After Jalapeño, there is no longer a place of rest for "model-only" companies in the AI industry.
—— END —— Return to Sohu to view more.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.