Seven hundred watts per compute die
SemiAnalysis puts numbers on OpenAI's Jalapeno: 13.4 PFLOPs of MXFP4, 15.4TB/s of HBM4, and a volume ramp that only arrives late in 2027.
3 minDePIN Compute & Bandwidth
SemiAnalysis published its analysis of Jalapeno on 25 August. It is OpenAI's inference chip, not a training part, designed with Broadcom, integrated by Celestica and fabricated by TSMC on N3P. The piece pushes back on the framing that it only runs OpenAI's own models, describing it as a general inference chip capable of diverse workloads.
The specification
- 13.4 PFLOPs of MXFP4 on the B0 stepping.
- 15.4TB/s of HBM4 bandwidth per package.
- 700W TDP per compute die.
- Out-of-order scalar cores alongside FP32 and INT32 vector units, with systolic arrays that take variable matrix dimensions.
- Scale-up to 2,048 ASICs across sixteen racks, over a hybrid copper and optical interconnect.
Where it wins and where it does not
On throughput per megawatt the analysis has Jalapeno ahead of every other chip without using speculative decoding, while the parts it is measured against are using it. On single-token prediction it is reported above 700 tokens per second per user on DeepSeek R1 at low concurrency. Performance per dollar, though, comes out comparable to Nvidia's Vera Rubin, and the two are optimised differently enough that the comparison needs qualifying.
The number that governs everything else
Engineering samples exist now. Production ramps gradually through 2027 with volume concentrated at the end of the year. For anyone modelling inference capacity, that timeline matters more than the FLOPs: a chip that wins on throughput per megawatt but does not arrive in quantity until late 2027 does not change a 2026 or 2027 buildout. It changes what the buildout after that one looks like.
Retold from SemiAnalysis. This is a summary in our own words; follow the link for the original reporting.