Don't Waste the Last Token: A Quantizer-Agnostic Quality Boost from Semantic ID Collision Resolution
Code for the paper "Don't Waste the Last Token: A Quantizer-Agnostic Quality Boost from Semantic ID Collision Resolution" (under review; authors and citation will be added after the review period).
TL;DR. Generative recommenders represent each item as a Semantic ID (SID): a short sequence of discrete tokens. Popular quantization algorithms, such as RQ-VAE and residual K-means, can map several items to the same token sequence, producing collisions. Collisions must be resolved before training a single-stage generative recommender, since the model cannot distinguish items that share an identifier. The widely adopted disambiguation technique appends one extra codebook and uses the last token as an arbitrary item counter within each collision group. We argue that this last codebook can be used more effectively as a source of additional content or collaborative signal. Keeping the quantizer, the generative model, and the SID length unchanged, we replace the ordinal counter with a token that both resolves collisions and carries a useful signal. Across several datasets, the collision resolution choice alone yields an easy-to-implement, quantizer-agnostic gain in recommendation quality (up to +12.7% in NDCG@10).
After quantization, each item has a Semantic ID prefix $\mathbf{p}i = (p{i,1}, \ldots, p_{i,L})$, where
The appending approach is attractive because it is quantizer-agnostic: it operates on obtained semantic codes after quantization, so it applies to any quantizer, leaving the SID structure and the training objective unchanged. However, the usual way is to fill the extra token with an arbitrary counter that carries no interpretable signal. We argue that this collision codebook can carry a useful content-based or collaborative signal at essentially no cost.
Appending strategies add one token quantization.collision_solver=<name>:
-
Sequential (
sequential, baseline, TIGER-style). Items in a collision group are numbered$z_i = 0, 1, 2, \ldots$ in arbitrary order. The last token carries no additional information; every non-colliding item, as well as the first item of each group, receives token 0, so the last-token distribution is heavily skewed. -
In-group popularity (
ingroup_pop, baseline, MMGRec-style). Items in a collision group are sorted by training-set popularity, and the rank is used as the last token:$z_i = 0$ denotes the most popular item within that group. The signal is local: the same token means different popularity levels in different groups. -
Random (
random, ours). For each collision group (including groups of size one), we build a random permutation of${0, \ldots, C_c - 1}$ and assign the first$|G_\mathbf{p}|$ values, one per item, as the last token. The token carries no interpretable signal, but it spreads items almost uniformly over the codebook, acting closely to a randomly-hashed item identifier. Each distinct token value is shared by far fewer items than in the Sequential baseline, which we hypothesize gives the model more capacity to learn collaborative patterns. -
Global Popularity (
global_pop, ours). We rank all items by training-set popularity and split them into$C_c$ equal-count bins. The bin index is the item's preferred token; smaller values mean more popular items. Non-colliding items keep their bin index exactly. Inside a group, items are processed from the most popular bin down, and each takes the still-free token nearest to its bin. The token therefore approximates the global popularity quantile with the same meaning in every prefix, which could allow the model to explicitly see and propagate popularity preferences, e.g., favoring long-tail items for users with niche tastes. -
ZCR (
zcr, baseline from the rebalancing family). Rewrites the last semantic code by reassigning colliding items to last-level centroids with minimal distortion: the SID stays fully semantic and no token is appended, but the quantizer's centroids and residuals are required. In our setup, its last codebook matches the collision codebook size of the appending strategies but carries a signal of a different nature (content-based and collision-free versus popularity or random), which isolates the effect of the collision token signal type.
The code additionally ships global_pop_opt, a variant of Global Popularity that repairs each group with a cost-optimal Hungarian assignment instead of the greedy nearest-free-code rule.
Whether the strategy is quantizer-agnostic; whether a token value has a consistent meaning across different collision groups (— means not applicable, since the token carries no interpretable signal); whether the collision codebook size
| Strategy | Signal | Quantizer-agnostic | Globally consistent | Varying |
Side info | |
|---|---|---|---|---|---|---|
| Sequential | counter | ✓ | — | ✗ | none | |
| In-group pop. | local rank | ✓ | ✗ | ✗ | popularity | |
| ZCR | content | ✗ | ✓ | ✗ | quantizer | |
| Random (ours) | spread (hash) | ✓ | — | ✓ | none | |
| Global pop. (ours) | pop. quantile | ✓ | ✓ | ✓ | popularity |
-
Datasets: three widely used Amazon Reviews 2014 datasets — Beauty, Sports and Outdoors, and Toys and Games — 5-core filtered, with consecutive repeats removed:
Dataset #Users #Items #Inter. Density Avg. length Beauty 22,363 12,101 198,502 0.073% 8.876 Sports 35,598 18,357 296,337 0.045% 8.324 Toys 19,412 11,924 167,597 0.072% 8.633 -
Split: a global temporal split at the 0.9 timestamp quantile with the last-item target (
gts_q09_val_by_time). The standard leave-one-out protocol may introduce temporal leakage, which is particularly problematic when evaluating popularity-based collision resolution strategies (see below). -
Embeddings:
sentence-transformers/all-MiniLM-L6-v2over item text metadata. -
Quantizer: residual K-means (FAISS) into 3 codebooks (4 for the ZCR baseline) of size
$C_q = 256$ , a configuration widely used in prior work. Appending strategies add one disambiguation token with collision codebook size$C_c = C_q$ ; ZCR instead rewrites a fourth codebook of the same size, so all methods use 4-token SIDs. -
Model: a decoder-only GPT-2-style Transformer with 8 layers, 8 attention heads, an embedding dimension of 512, and beam search for autoregressive decoding.
-
Protocol: NDCG@10 and HitRate@10 as the main ranking metrics, averaged over 5 random seeds.
Generative recommender quality for different collision resolution strategies: mean ± std over 5 seeds, with relative change vs. the Sequential baseline. The best results are shown in bold.
Beauty
| Strategy | NDCG@10 | Δ | HR@10 | Δ |
|---|---|---|---|---|
| Sequential | 0.0182 ± .0015 | — | 0.0344 ± .0023 | — |
| In-group pop. | 0.0181 ± .0017 | −0.8% | 0.0348 ± .0031 | +1.1% |
| ZCR | 0.0174 ± .0009 | −4.6% | 0.0345 ± .0012 | +0.2% |
| Random (ours) | 0.0182 ± .0014 | −0.3% | 0.0348 ± .0032 | +1.2% |
| Global pop. (ours) | 0.0191 ± .0008 | +5.1% | 0.0364 ± .0012 | +6.0% |
Sports
| Strategy | NDCG@10 | Δ | HR@10 | Δ |
|---|---|---|---|---|
| Sequential | 0.0131 ± .0004 | — | 0.0252 ± .0006 | — |
| In-group pop. | 0.0129 ± .0007 | −1.7% | 0.0245 ± .0010 | −2.6% |
| ZCR | 0.0138 ± .0006 | +5.6% | 0.0254 ± .0008 | +0.8% |
| Random (ours) | 0.0134 ± .0004 | +2.1% | 0.0250 ± .0005 | −0.8% |
| Global pop. (ours) | 0.0134 ± .0007 | +2.7% | 0.0253 ± .0011 | +0.5% |
Toys
| Strategy | NDCG@10 | Δ | HR@10 | Δ |
|---|---|---|---|---|
| Sequential | 0.0243 ± .0010 | — | 0.0450 ± .0023 | — |
| In-group pop. | 0.0233 ± .0005 | −3.9% | 0.0441 ± .0009 | −2.0% |
| ZCR | 0.0242 ± .0013 | −0.3% | 0.0453 ± .0029 | +0.5% |
| Random (ours) | 0.0247 ± .0009 | +1.8% | 0.0464 ± .0014 | +3.1% |
| Global pop. (ours) | 0.0274 ± .0010 | +12.7% | 0.0514 ± .0025 | +14.1% |
Findings. The Sequential strategy is outperformed by at least one alternative approach on every dataset; most notably, Global Popularity improves metrics on all three datasets, by up to +12.7% in NDCG@10 (Toys). The ZCR strategy outperforms the others on the Sports dataset. Global Popularity clearly outperforms Random, indicating that an additional popularity signal is beneficial compared to the simple expansion of the codebook with random assignment. There is no universal winner, highlighting the need for deeper analysis of the connection between data properties and SID/quantization properties, such as embedding quality, SID sparsity and distribution, and the severity of the collision problem. The takeaway: collision resolution is a design choice, not an implementation detail — the strategy alone can, on its own, increase recommendation quality with no change to the quantizer, the model, or the SID length.
We additionally evaluate under the conventional leave-one-out protocol to verify that our pipeline demonstrates performance comparable to that reported in prior work. The leave-one-out numbers are higher than the global-temporal-split numbers above; however, we do not report them as our main results, because the leave-one-out protocol introduces popularity (temporal) leakage that popularity-based collision resolution strategies can exploit, making the comparison unreliable.
Beauty
| Strategy | NDCG@10 | Δ | HR@10 | Δ |
|---|---|---|---|---|
| Sequential | 0.0335 ± 0.0010 | — | 0.0620 ± 0.0018 | — |
| In-group pop. | 0.0321 ± 0.0018 | -3.9% | 0.0569 | -3.9% |
| ZCR | 0.0357 ± 0.0009 | +6.5% | 0.0647 ± 0.0015 | +4.3% |
| Random (ours) | 0.0356 ± 0.0002 | +6.5% | 0.0647 ± 0.0010 | +4.4% |
| Global pop. (ours) | 0.0367 ± 0.0009 | +9.6% | 0.0669 ± 0.0008 | +7.9% |
Sports
| Strategy | NDCG@10 | Δ | HR@10 | Δ |
|---|---|---|---|---|
| Sequential | 0.0170 ± 0.0004 | — | 0.0323 ± 0.0008 | — |
| In-group pop. | 0.0169 ± 0.0001 | -0.3% | 0.0321 ± 0.0002 | -0.6% |
| ZCR | 0.0184 ± 0.0007 | +8.3% | 0.0344 ± 0.0012 | +6.3% |
| Random (ours) | 0.0176 ± 0.0005 | +3.4% | 0.0329 ± 0.0011 | +1.7% |
| Global pop. (ours) | 0.0183 ± 0.0001 | +7.7% | 0.0345 ± 0.0015 | +6.8% |
Toys
| Strategy | NDCG@10 | Δ | HR@10 | Δ |
|---|---|---|---|---|
| Sequential | 0.0319 ± 0.0007 | — | 0.0611 ± 0.0011 | — |
| In-group pop. | 0.0310 ± 0.0005 | -2.6% | 0.0602 ± 0.0008 | -1.5% |
| ZCR | 0.0328 ± 0.0010 | +3.0% | 0.0628 ± 0.0018 | +2.7% |
| Random (ours) | 0.0321 ± 0.0006 | +0.8% | 0.0614 ± 0.0014 | +0.5% |
| Global pop. (ours) | 0.0362 ± 0.0004 | +13.5% | 0.0694 ± 0.0009 | +13.5% |
Python 3.11 is recommended.
git clone <repository-url>
cd sid-collisions
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txtAll commands run from the repository root with PYTHONPATH="./" (set inside the provided scripts) and SEQ_REC_DATA_PATH pointing to the directory where datasets should be stored (exported by you).
export SEQ_REC_DATA_PATH=/path/to/data
bash sh/prepare_datasets.shFor Beauty2014, Sports2014, and Toys2014 this performs: (0) download of the Amazon 2014 ratings and metadata from SNAP; (1) 5-core preprocessing with consecutive-duplicate removal; (2) the global temporal split (gts_q09_val_by_time) and the leave_one_out split; (3) all-MiniLM-L6-v2 embedding generation (set EMBED_GPU=<id> to choose the GPU, default 0).
bash sh/run_exp.sh all gts_q09_val_by_time # main (global temporal split)
bash sh/run_exp.sh all leave_one_out # supplementary (leave-one-out)A single run can also be launched directly:
python runs/quantize_and_train.py \
recipe=GPTRec_SID \
dataset=Beauty2014 \
quantization.collision_solver=global_pop \
random_state=17 \
dataset_params.train_dataset.shift=0 \
dataset_params.predict_dataset.shift=0 \
split_name=gts_q09_val_by_timeRuns log with DiskTracker (local write on disk) by default (experiment_tracker in runs/config/quantize_and_train.yaml).
Also we have ClearmlTracker for log with ClearML. If you have credentials, enable tracking with:
python runs/quantize_and_train.py ... experiment_tracker._target_=src.trackers.ClearmlTrackerIf HuggingFace model downloads hang behind a proxy, set HF_HUB_DISABLE_XET=1.
runs/ entry points (Hydra)
preprocess.py raw data -> filtered interactions + metadata
split.py global temporal / leave-one-out splits
build_semantic_embeddings.py
quantize_and_train.py quantize items, resolve collisions, train, evaluate
get_data/ raw data download scripts
config/ Hydra configs (dataset, quantizer, recipe, ...)
src/
semantic_ids/ quantizers and collision solvers (collision_solvers.py)
models/, modules/ GPT-2-style decoder and Lightning module
metrics.py, callbacks.py, trackers.py
sh/
prepare_datasets.sh
run_exp.sh
