Skip to content
Apixo
Blog
news· 4 min read· via Towards AI

How Generative AI is Rewriting the Rules of Recommendation Systems

A hands-on breakdown of generative recommendation engines, exploring how YouTube, Netflix, and Meta use Semantic IDs and Transformers to suggest the next item.

How Generative AI is Rewriting the Rules of Recommendation Systems

Large-scale recommendation systems are undergoing a fundamental paradigm shift. Traditionally, platforms relied on collaborative filtering to determine what users might like next by matching similar user profiles and item interactions. However, industry leaders like YouTube, Netflix, and Meta are increasingly moving toward generative recommendation systems. Instead of treating recommendation as a search or matrix factorization problem, these modern architectures treat a user's interaction history like a sentence and generate the next item's identity, much like a large language model predicts the next word in a sequence.

To demonstrate how this works in practice, developer Juan Miguel Gutierrez built a generative recommendation engine from scratch using the Amazon Reviews 2023 dataset (specifically the "5-core" Video Games category from UCSD's McAuley Lab). The project breaks down the transition from classic baselines to a fully generative transformer-based retriever paired with a gradient-boosted tree ranker.

Constructing Semantic IDs and Generative Retrieval

The foundation of a generative recommender is giving items an identity that carries actual meaning. In classic systems, products are represented by arbitrary database IDs. In a generative system, items are assigned "Semantic IDs." Gutierrez achieved this using a multi-step pipeline. First, a pretrained Sentence-T5-base model converts product titles into a 768-dimensional vector space, grouping similar titles close together. Next, a small encoder compresses these vectors down to 32 dimensions. Finally, a Residual Quantization Variational Autoencoder (RQ-VAE) snaps these coordinates to three successive 256-vector codebooks, generating a three-digit ID representing coarse, finer, and finest categories. A fourth digit is appended to resolve ties among highly similar products.

With every product represented by a four-digit Semantic ID, a user's purchase history is converted into a stream of tokens. Gutierrez trained a custom, decoder-only Transformer with 4 attention layers, 4 heads, and 1.1 million parameters to read this stream. At inference, the model uses beam search constrained by a lookup tree of real catalog items to generate the most probable next four-digit ID, ensuring that only actual products are recommended.

What it means for developers

For developers, this shift from database-lookup recommendations to generative retrieval introduces both massive opportunities and new engineering challenges. Traditionally, cold-start problems plagued collaborative filtering: if an item was brand new with zero purchases, the algorithm could not recommend it. Because Semantic IDs are built from content (titles and descriptions), a generative model can recommend completely new products based on their semantic characteristics alone.

While implementing custom transformers and autoencoders from scratch offers deep architectural control, developers looking to integrate state-of-the-art language models into their applications can also try top AI models cheaply through one API at https://apixoai.online. This simplifies the experimentation process when working with text embeddings or prompt-based classification.

Furthermore, developers must pay close attention to how they train their ranking layers. In a two-stage recommendation pipeline, a cheap retriever generates a shortlist of candidates, and a ranker (like LightGBM using LambdaMART) orders them using richer features like price, popularity, and cosine similarity. Gutierrez warns against an "injection bug" during ranker training: inserting true positive items that the retriever missed with placeholder low scores teaches the ranker to erroneously promote poorly scored, unseen candidates. Rerankers must be trained strictly on the actual candidates and feature values that the retrieval stage outputs.

Evaluating the Architecture

The project evaluated multiple tiers—popularity baselines, Alternating Least Squares (ALS), the generative Transformer, and the reranked pipeline—using Recall@K and Normalized Discounted Cumulative Gain (NDCG@K).

The evaluation revealed how heavily the choice of testing protocol influences results. Under a realistic time-based split (where training data ends in 2019 and testing spans 2021–2023), the gap between interactions averaged nearly four years, and 72% of test items were completely unseen in training. In this harsh scenario, all retrieval models collapsed to the floor of recommending popular items, though the LightGBM reranker emerged as the clear winner, boosting performance by 35% to 50% over retrieval-only models.

However, when evaluated under the standard "leave-one-out" next-item prediction protocol used in academic papers (such as SASRec), the generative model proved highly effective. It achieved a Recall@10 of 0.069 to 0.077, matching benchmarks established by models like Google's TIGER. While it trailed the purely ID-based SASRec model by roughly 15% on this specific dataset, the experiment confirmed that generating content-derived digits is a viable, powerful alternative to traditional lookup systems.


Source: Generative AI for Recommendations: what YouTube, Netflix and Meta are Moving to, Built From Scratch — Towards AI. Written by the Apixo team from that report.

#ai-news#machine-learning#recommender-systems#transformers#ai-engineering#data-science
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading