Procurement evidence is changing
Stop Buying the Peak Score. Buy the Performance Curve.
MLCommons says MLPerf Endpoints will replace datacenter MLPerf Inference. The practical change is a benchmark that compares a deployed endpoint across the load and latency trade-offs buyers actually operate.

The familiar AI infrastructure scorecard begins with a flattering number: maximum throughput, best latency, or the fastest result on a selected workload. It is useful evidence, but it is incomplete evidence for a buyer. Production systems are not bought to operate at a vendor’s preferred point. They are bought to serve a particular number of concurrent users while keeping interactive delays within an acceptable range.
MLCommons is moving its datacenter benchmark family toward a more operational unit of comparison. In results published on September 16, 2026, the organization said that MLPerf Endpoints will replace MLPerf Inference for datacenter benchmarking. More than half of submitters to MLPerf Inference v6.1 already used the API-centric client/server harness that forms Endpoints’ foundation.
The important development is not this round’s fastest accelerator. It is the change in what a credible comparison can look like: a working AI endpoint, measured across a range of concurrent load, rather than a single score detached from the operating point a customer needs.
The endpoint becomes the unit that buyers can compare
An endpoint is the deployed combination of model, serving software and hardware. That combination is what users experience and what an infrastructure team must run. A model name alone does not establish its production behaviour; neither does an accelerator name. Batch handling, request scheduling, serving configuration and the rest of the system affect the result.
MLPerf Endpoints measures a working model-serving endpoint under varying concurrency. It presents the relationship among system throughput, per-user interactivity and P95 time to first token. That creates a performance curve rather than a single winner’s podium.
This is a material distinction. Aggregate throughput answers how much work a system can process. Per-user interactivity addresses what an individual user experiences as load rises. P95 time to first token exposes the experience nearer the slow end of the distribution, rather than only an average. A configuration can look excellent on one of these measures while failing the operating condition that matters to a buyer.
The useful question is no longer “Which system is fastest?” It is “Which submitted endpoint meets our load and responsiveness requirement, and what evidence supports that claim?”
A curve does not eliminate judgment. It makes the judgment explicit. A service supporting a small group of analysts may accept lower aggregate throughput in return for stronger responsiveness. A high-volume workflow may choose a different point, provided its interaction requirements are still met. These are operating decisions. A fixed peak score cannot make them for the buyer.
What changed—and what did not
MLCommons published MLPerf Inference v6.1 with submissions from a record 30 organizations, spanning cloud providers, chipmakers, infrastructure vendors and system builders. The scale of participation matters because a comparison is more useful when buyers can evaluate multiple classes of supplier within a common methodology.
More than 50% of those submitters used the API-centric client/server harness. MLCommons also stated its direction plainly: Endpoints will replace Inference in the datacenter benchmark family. This is adoption evidence and a declared transition, not confirmation that the replacement is already complete. MLPerf Inference v6.1 remains the published result round.
The release also added an end-to-end RAG test covering corpus ingestion and multi-hop question answering, alongside an edge-agentic workload with latency and accuracy requirements. Those additions are supporting evidence that benchmark workloads are becoming more representative of deployed systems. They should not be mistaken for a claim that the benchmark measures business value, factual reliability in every application, or completed outcomes.
Nor does a performance curve settle commercial questions. MLCommons says submissions receive peer review, require reproducibility artifacts and use availability labels to distinguish systems available now from preview systems. That is valuable procurement evidence. It does not mean every result represents a configuration a buyer can order today, nor does it establish cost-performance without pricing and a comparison at the same operating point.
The headline 5.7× year-over-year improvement in the release has an equally important limit: it applies only to the best per-accelerator DeepSeek R1 server result. It is not a measure of inference progress generally, system-wide performance, or customer economics.
Turn the curve into an RFP requirement
Procurement teams should respond by changing the evidence they request, not by treating a new chart as another marketing attachment. An RFP for inference capacity can require a verified endpoint curve for the model, serving configuration and availability window being purchased. It can define the relevant concurrency range and set the service condition to be evaluated—for example, a required level of per-user interactivity and a P95 time-to-first-token threshold.
The request should also require the supplier to identify whether the submitted system is available now or in preview, and to provide the reproducibility and verification artifacts associated with the result. This separates two questions that are frequently merged: whether a configuration performed well in a benchmark, and whether it is the configuration a buyer can obtain and operate.
- Specify the actual model-serving endpoint to be evaluated, rather than requesting an accelerator score in isolation.
- Define the concurrent-load range that reflects the expected service, including planned growth where relevant.
- State responsiveness conditions in operational terms, including P95 time to first token and the required level of per-user interactivity.
- Ask for the full verified curve at those conditions, not only the peak throughput point.
- Require the result’s availability label and the associated peer-review and reproducibility artifacts.
- Compare commercial proposals only after aligning configurations at the same operating point.
This approach gives technical teams a cleaner basis for capacity planning and gives commercial teams a firmer basis for comparing offers. It also makes trade-offs visible early. If a supplier can meet the required responsiveness only at a lower concurrency than expected, that is not a benchmark footnote. It is a design and budget issue.
A better benchmark asks a harder question
Single benchmark results will remain useful. They can expose capability advances and establish reference points. But datacenter AI is increasingly delivered as a live service, and live services operate under changing concurrent demand. The benchmark should therefore reveal the shape of performance under that demand.
That is the significance of MLCommons’ declared move to Endpoints. It shifts the centre of gravity from the component or selected scenario toward the deployed serving system. Buyers gain a way to ask for evidence at their required operating point, rather than inheriting a vendor’s preferred comparison point.
The discipline is straightforward: do not buy an impressive peak and infer the rest. Define the service condition, obtain the verified endpoint curve, check availability, then make the commercial comparison. The curve will not choose the system. It will make the real choice visible.

