Products & Models · Oct 4, 2026
MLPerf's recommendation benchmark moves to HSTU, which reads user history as a sequence; DLRMv4's 560GB of embeddings won't fit on one accelerator
The benchmark drops the older feature-crossing DLRM and will measure generative recommendation workloads instead. AMD, Meta, NVIDIA and ByteDance co-wrote the announcement
Rie Suzuki · Technology Editor

Key points
- MLPerf's recommendation training test is moving from the older DLRM, which aggregates features and crosses them, to HSTU, which reads interaction history as a sequence with one token per event
- The embeddings alone total 560GB, which is too large for a single accelerator, so the test assumes they are split across several
- The authors give two reasons: production recommendation systems have moved to sequential models, and these models keep improving as compute grows. AMD, Meta, NVIDIA and ByteDance co-wrote the announcement
MLCommons said on October 1 that it is replacing the recommendation test in MLPerf Training, the industry-standard benchmark, with DLRMv4 (MLCommons blog, primary source). The earlier DLRM combined user and item features into aggregates and then crossed them to produce recommendations. DLRMv4 replaces that part with HSTU, which reads a user's behavior history directly as a sequence. This also changes what the benchmark measures. The new workload is "generative recommendation": the embeddings alone total 560GB and are split across multiple accelerators.
The post was written by the MLCommons DLRMv4 task force, including Chris Cai, Linjian Ma, Nan Zhang and Haolei Pei, and was co-written by AMD, Meta, NVIDIA and ByteDance. Companies that run recommendation at massive scale and chipmakers that sell them the compute have rebuilt the benchmark together.
What changes: from crossing features to reading history as a sequence
The task force first says that production systems have already changed.
production recommendation systems at hyperscale operators have shifted from compressing user behavior into a small set of aggregated dense features to representing each user's recent interaction history as a per-event token sequence.
The older approach takes what a user recently did, such as which videos they watched and which products they bought, and reduces it to a small set of numbers. Those aggregated features are crossed to estimate the probability of a click. The newer approach treats each interaction as a single token and reads the history in order, the way a language model reads text. Information such as order and timing is no longer thrown away during aggregation.
HSTU is at the center of the new approach. The task force describes it this way:
DLRMv4 fills that gap by replacing the feature-interaction stack with HSTU, a transducer that reads a user's interaction history directly as a sequence
"That gap" is the difference between the recommendation systems running in production and the ones the benchmark measures. If MLPerf's recommendation test stays on the older feature-interaction design, a good score does not show how fast a system runs the workloads operators actually use. DLRMv4 is a revision meant to bring the benchmark back in line with production. HSTU is an architecture that Meta researchers introduced in 2024 as a foundation for "generative recommendation."
Why now: more compute keeps paying off
The task force gives a second reason for the switch: how these models behave as they grow.
Models built this way keep improving as compute and parameters grow, while earlier feature-interaction designs flatten out.
The claim is that the scaling behavior familiar from LLMs also appears in recommendation once models become sequential: performance keeps rising as compute and parameters increase. Feature-interaction designs, by contrast, hit a ceiling at some point. If operators plan to spend more compute on recommendation, it makes sense to benchmark the approach that keeps improving. The revision is also a joint industry acknowledgment that recommendation has become a field where throwing more compute at the problem pays off.
What 560GB means: a workload too large for one accelerator
For hardware readers, the headline figure matters most. DLRMv4's embeddings alone total 560GB, which exceeds the memory of any single current accelerator. The test therefore assumes the embedding tables are split across multiple accelerators.
In recommendation training, the costliest step is often looking up the needed rows in huge embedding tables, not the computation itself. Once the tables are split, the system has to decide which accelerator holds which rows, and accelerators have to fetch rows from one another. Raw compute speed alone will not decide results. Memory capacity, memory bandwidth, the bandwidth of the links between accelerators, and how well the software decides where to place each part of the table will all show up in the score.
HSTU's sequence computation comes on top of that. Longer histories mean more attention-like computation. DLRMv4 therefore tests embedding lookups and transformer-like sequence computation in a single benchmark. It gives a named benchmark to a third type of workload, generative recommendation, which is neither LLM training nor LLM inference.
Who the revision is for
The list of co-authors is telling. Meta and ByteDance run sequential recommendation in production at large scale. AMD and NVIDIA sell the accelerators that run it. In future MLPerf results, the recommendation numbers will show whose systems can train hyperscale recommendation models fastest. A key point of comparison will be how each vendor's system splits the 560GB of embeddings and how many accelerators it uses.
The post confirms only a few things: the change of approach, the adoption of HSTU, the size of the embeddings, and the requirement to split them across multiple accelerators. Which round of MLPerf Training will first use DLRMv4 officially, the quality target, the dataset contents, and each vendor's first submissions have not yet been announced. We will watch MLCommons' Insights page for follow-up posts.
Editorial cartoon
