Benchmark Highlights
The Hot Chips conference unveiled numbers that an independent firm verified with the InferenceX suite. OpenAI calls Jalapeño a general-purpose inference accelerator—it runs models but doesn’t train them. In the tests the chip:
- Delivered 1.5 × to 1.9 × more work per watt at peak throughput than today’s leaders.
- Cut latency 1.7 × to 3.6 ×, with interactive gains of 2.1 × to 4.1 ×.
Model-specific results:
- GPT-OSS 120B: ~1,400 tokens / second / user.
- Deepseek R1 670B: ~700 tokens / second on a single concurrent request.
- Kimi K2.5 1T: tested in the high-performance suite (exact numbers withheld).
OpenAI stressed that these figures came without speculative decoding or multi-token prediction—tricks many rivals use to boost headline numbers.
How Jalapeño Was Built
OpenAI finished the chip’s design in nine months, teaming with Broadcom. The firm fed its own models into the design loop, speeding blueprinting, programming and optimization—a method dubbed “AI-assisted silicon design.” The sprint highlights a shift from the multi-year cycles that once defined high-performance silicon.
Implications for Nvidia’s Lead
Nvidia leans on raw performance and the CUDA ecosystem. The new data makes the performance edge contestable. Analysts warn that if more AI firms roll out custom inference silicon with similar efficiency, developers may drift away from CUDA, opening space for alternative stacks and a fragmented market.
OpenAI’s Mixed-Signal Strategy
Chief financial officer Kelly Huang framed Jalapeño as one piece of a broader compute plan, not a wholesale replacement for existing partners. OpenAI still works with Nvidia, AMD, AWS and Cerebras across workloads. Jalapeño remains an engineering sample; it has yet to face the biggest upcoming models like Deepseek V4 Pro. Huang’s comments suggest OpenAI will mix its own silicon with third-party accelerators, chasing the best cost-performance mix for each task.
Counterpoints
Early-stage samples don’t guarantee volume production, yields or long-term reliability, so excitement should be tempered.
What to Watch Next
- Production rollout: Can OpenAI turn Jalapeño into a mass-produced part, and at what price?
- Software stack: How mature will the programming model and toolchain be for developers?
- Competitor response: How will Nvidia and other vendors tweak roadmaps or pricing to protect market share?
- Model scaling: Will Jalapeño keep its edge on the next wave of trillion-parameter models?
Takeaway: Custom inference silicon is moving from a niche experiment to a credible challenge to Nvidia’s hardware dominance.
