OpenAI has released the first measured performance figures for Jalapeño, the inference chip it designed with Broadcom, and they put the part ahead of Nvidia's top-scoring machines on two counts at once.
Across the three public models tested — GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T — the chip got 1.5 to 1.9 times as much AI work done per watt at peak throughput, and its end-to-end latency came in 1.7 to 3.6 times lower than the Nvidia GB200 and GB300 systems used as the baseline, which held the strongest scores on record when the testing took place.
Jalapeño is the first piece of silicon the company has designed in-house, an application-specific integrated circuit built purely to execute models that have already been trained rather than to carry out the training itself.
Those figures come from InferenceX, an openly available benchmark from SemiAnalysis that gauges everything involved in answering an AI inference request from start to finish. On workloads that demand heavy interactivity, that benchmark scored Jalapeño 2.1 to 4.1 times above the machines it was measured against, a margin that holds for those workloads alone.
The margins are not even. On Kimi K2.5, the biggest public model put through the tests, peak performance per watt was roughly 1.5 times higher — the bottom of the range — while end-to-end latency was 3.4 times lower than the machine it was compared with. To keep the comparison level, the per-watt figures were scaled against the chip power rating each accelerator's maker publishes; Jalapeño's rating is 700 watts, though its sustained draw across the workloads run was measured at 550 watts or under.
The chip is not carrying traffic yet. OpenAI intends to put it into service in limited quantities before the close of 2026 and to scale those quantities through 2027, without disclosing how many it plans to field that year. Richard Ho, OpenAI's vice president for hardware, described the chip as one element of a compute approach that continues to lean on partners such as Nvidia and Cerebras; the GPU fleet is not being retired.
The baseline it beat is Nvidia's Blackwell generation, which SemiAnalysis calls an unfair fight: the chip Jalapeño really competes with is Rubin, already shipping to customers while OpenAI has only engineering samples. Against the Rubin figures Nvidia and CoreWeave published in July it still comes out ahead on throughput per megawatt — and does it without the speculative decoding Rubin's runs use.
The numbers themselves all originated with OpenAI. SemiAnalysis watched the InferenceX runs on site, but it did not execute the complete benchmark suite on its own, and it was not shown any results from AgentX, a separate benchmark.
Asked whether OpenAI would make the chip available to outside users, Ho replied that the company does not have compute enough even to cover its own requirements. Jalapeño cannot be rented, and nobody can buy one.
Cover image: “green chili pepper” by Brunxs, Unsplash, cropped.