Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Don't Waste the Last Token: A Quantizer-Agnostic Quality Boost from Semantic ID Collision Resolution

Code for the paper "Don't Waste the Last Token: A Quantizer-Agnostic Quality Boost from Semantic ID Collision Resolution" (under review; authors and citation will be added after the review period).

TL;DR. Generative recommenders represent each item as a Semantic ID (SID): a short sequence of discrete tokens. Popular quantization algorithms, such as RQ-VAE and residual K-means, can map several items to the same token sequence, producing collisions. Collisions must be resolved before training a single-stage generative recommender, since the model cannot distinguish items that share an identifier. The widely adopted disambiguation technique appends one extra codebook and uses the last token as an arbitrary item counter within each collision group. We argue that this last codebook can be used more effectively as a source of additional content or collaborative signal. Keeping the quantizer, the generative model, and the SID length unchanged, we replace the ordinal counter with a token that both resolves collisions and carries a useful signal. Across several datasets, the collision resolution choice alone yields an easy-to-implement, quantizer-agnostic gain in recommendation quality (up to +12.7% in NDCG@10).

Why SIDs collide, and what signal the appended token can carry

The problem

After quantization, each item has a Semantic ID prefix $\mathbf{p}i = (p{i,1}, \ldots, p_{i,L})$, where $L$ is the number of quantization codebooks. Items with similar content can receive an identical prefix, producing a collision: without resolution, decoding cannot recover a unique item. Existing quantizer-agnostic responses either number the colliding items with an arbitrary counter (TIGER) or with an in-group popularity rank (MMGRec), while rebalancing methods such as ZCR rewrite the last semantic code but require access to the quantizer's internals (centroids and residuals).

The appending approach is attractive because it is quantizer-agnostic: it operates on obtained semantic codes after quantization, so it applies to any quantizer, leaving the SID structure and the training objective unchanged. However, the usual way is to fill the extra token with an arbitrary counter that carries no interpretable signal. We argue that this collision codebook can carry a useful content-based or collaborative signal at essentially no cost.

Collision resolution strategies

Appending strategies add one token $z_i \in {0, \ldots, C_c - 1}$ from a collision codebook of size $C_c$ to form the final SID $\mathbf{s}_i = (\mathbf{p}_i, z_i)$. All solvers live in src/semantic_ids/collision_solvers.py and are selected via quantization.collision_solver=<name>:

  • Sequential (sequential, baseline, TIGER-style). Items in a collision group are numbered $z_i = 0, 1, 2, \ldots$ in arbitrary order. The last token carries no additional information; every non-colliding item, as well as the first item of each group, receives token 0, so the last-token distribution is heavily skewed.
  • In-group popularity (ingroup_pop, baseline, MMGRec-style). Items in a collision group are sorted by training-set popularity, and the rank is used as the last token: $z_i = 0$ denotes the most popular item within that group. The signal is local: the same token means different popularity levels in different groups.
  • Random (random, ours). For each collision group (including groups of size one), we build a random permutation of ${0, \ldots, C_c - 1}$ and assign the first $|G_\mathbf{p}|$ values, one per item, as the last token. The token carries no interpretable signal, but it spreads items almost uniformly over the codebook, acting closely to a randomly-hashed item identifier. Each distinct token value is shared by far fewer items than in the Sequential baseline, which we hypothesize gives the model more capacity to learn collaborative patterns.
  • Global Popularity (global_pop, ours). We rank all items by training-set popularity and split them into $C_c$ equal-count bins. The bin index is the item's preferred token; smaller values mean more popular items. Non-colliding items keep their bin index exactly. Inside a group, items are processed from the most popular bin down, and each takes the still-free token nearest to its bin. The token therefore approximates the global popularity quantile with the same meaning in every prefix, which could allow the model to explicitly see and propagate popularity preferences, e.g., favoring long-tail items for users with niche tastes.
  • ZCR (zcr, baseline from the rebalancing family). Rewrites the last semantic code by reassigning colliding items to last-level centroids with minimal distortion: the SID stays fully semantic and no token is appended, but the quantizer's centroids and residuals are required. In our setup, its last codebook matches the collision codebook size of the appending strategies but carries a signal of a different nature (content-based and collision-free versus popularity or random), which isolates the effect of the collision token signal type.

The code additionally ships global_pop_opt, a variant of Global Popularity that repairs each group with a cost-optimal Hungarian assignment instead of the greedy nearest-free-code rule.

Strategy comparison

Whether the strategy is quantizer-agnostic; whether a token value has a consistent meaning across different collision groups (— means not applicable, since the token carries no interpretable signal); whether the collision codebook size $C_c$ can vary after quantization; the number of distinct disambiguation tokens ($g_{\max}$ is the largest collision group size); and the required side information:

Strategy Signal Quantizer-agnostic Globally consistent Varying $C_c$ $C_c$ Side info
Sequential counter ✓ — ✗ $g_{\max}$ none
In-group pop. local rank ✓ ✗ ✗ $g_{\max}$ popularity
ZCR content ✗ ✓ ✗ $C_q \geq g_{\max}$ quantizer
Random (ours) spread (hash) ✓ — ✓ $C_c \geq g_{\max}$ none
Global pop. (ours) pop. quantile ✓ ✓ ✓ $C_c \geq g_{\max}$ popularity

Experimental setup

  • Datasets: three widely used Amazon Reviews 2014 datasets — Beauty, Sports and Outdoors, and Toys and Games — 5-core filtered, with consecutive repeats removed:

    Dataset #Users #Items #Inter. Density Avg. length
    Beauty 22,363 12,101 198,502 0.073% 8.876
    Sports 35,598 18,357 296,337 0.045% 8.324
    Toys 19,412 11,924 167,597 0.072% 8.633
  • Split: a global temporal split at the 0.9 timestamp quantile with the last-item target (gts_q09_val_by_time). The standard leave-one-out protocol may introduce temporal leakage, which is particularly problematic when evaluating popularity-based collision resolution strategies (see below).

  • Embeddings: sentence-transformers/all-MiniLM-L6-v2 over item text metadata.

  • Quantizer: residual K-means (FAISS) into 3 codebooks (4 for the ZCR baseline) of size $C_q = 256$, a configuration widely used in prior work. Appending strategies add one disambiguation token with collision codebook size $C_c = C_q$; ZCR instead rewrites a fourth codebook of the same size, so all methods use 4-token SIDs.

  • Model: a decoder-only GPT-2-style Transformer with 8 layers, 8 attention heads, an embedding dimension of 512, and beam search for autoregressive decoding.

  • Protocol: NDCG@10 and HitRate@10 as the main ranking metrics, averaged over 5 random seeds.

Results

Generative recommender quality for different collision resolution strategies: mean ± std over 5 seeds, with relative change vs. the Sequential baseline. The best results are shown in bold.

Beauty

Strategy NDCG@10 Δ HR@10 Δ
Sequential 0.0182 ± .0015 — 0.0344 ± .0023 —
In-group pop. 0.0181 ± .0017 −0.8% 0.0348 ± .0031 +1.1%
ZCR 0.0174 ± .0009 −4.6% 0.0345 ± .0012 +0.2%
Random (ours) 0.0182 ± .0014 −0.3% 0.0348 ± .0032 +1.2%
Global pop. (ours) 0.0191 ± .0008 +5.1% 0.0364 ± .0012 +6.0%

Sports

Strategy NDCG@10 Δ HR@10 Δ
Sequential 0.0131 ± .0004 — 0.0252 ± .0006 —
In-group pop. 0.0129 ± .0007 −1.7% 0.0245 ± .0010 −2.6%
ZCR 0.0138 ± .0006 +5.6% 0.0254 ± .0008 +0.8%
Random (ours) 0.0134 ± .0004 +2.1% 0.0250 ± .0005 −0.8%
Global pop. (ours) 0.0134 ± .0007 +2.7% 0.0253 ± .0011 +0.5%

Toys

Strategy NDCG@10 Δ HR@10 Δ
Sequential 0.0243 ± .0010 — 0.0450 ± .0023 —
In-group pop. 0.0233 ± .0005 −3.9% 0.0441 ± .0009 −2.0%
ZCR 0.0242 ± .0013 −0.3% 0.0453 ± .0029 +0.5%
Random (ours) 0.0247 ± .0009 +1.8% 0.0464 ± .0014 +3.1%
Global pop. (ours) 0.0274 ± .0010 +12.7% 0.0514 ± .0025 +14.1%

Findings. The Sequential strategy is outperformed by at least one alternative approach on every dataset; most notably, Global Popularity improves metrics on all three datasets, by up to +12.7% in NDCG@10 (Toys). The ZCR strategy outperforms the others on the Sports dataset. Global Popularity clearly outperforms Random, indicating that an additional popularity signal is beneficial compared to the simple expansion of the codebook with random assignment. There is no universal winner, highlighting the need for deeper analysis of the connection between data properties and SID/quantization properties, such as embedding quality, SID sparsity and distribution, and the severity of the collision problem. The takeaway: collision resolution is a design choice, not an implementation detail — the strategy alone can, on its own, increase recommendation quality with no change to the quantizer, the model, or the SID length.

Leave-one-out results

We additionally evaluate under the conventional leave-one-out protocol to verify that our pipeline demonstrates performance comparable to that reported in prior work. The leave-one-out numbers are higher than the global-temporal-split numbers above; however, we do not report them as our main results, because the leave-one-out protocol introduces popularity (temporal) leakage that popularity-based collision resolution strategies can exploit, making the comparison unreliable.

Beauty

Strategy NDCG@10 Δ HR@10 Δ
Sequential 0.0335 ± 0.0010 — 0.0620 ± 0.0018 —
In-group pop. 0.0321 ± 0.0018 -3.9% 0.0569 -3.9%
ZCR 0.0357 ± 0.0009 +6.5% 0.0647 ± 0.0015 +4.3%
Random (ours) 0.0356 ± 0.0002 +6.5% 0.0647 ± 0.0010 +4.4%
Global pop. (ours) 0.0367 ± 0.0009 +9.6% 0.0669 ± 0.0008 +7.9%

Sports

Strategy NDCG@10 Δ HR@10 Δ
Sequential 0.0170 ± 0.0004 — 0.0323 ± 0.0008 —
In-group pop. 0.0169 ± 0.0001 -0.3% 0.0321 ± 0.0002 -0.6%
ZCR 0.0184 ± 0.0007 +8.3% 0.0344 ± 0.0012 +6.3%
Random (ours) 0.0176 ± 0.0005 +3.4% 0.0329 ± 0.0011 +1.7%
Global pop. (ours) 0.0183 ± 0.0001 +7.7% 0.0345 ± 0.0015 +6.8%

Toys

Strategy NDCG@10 Δ HR@10 Δ
Sequential 0.0319 ± 0.0007 — 0.0611 ± 0.0011 —
In-group pop. 0.0310 ± 0.0005 -2.6% 0.0602 ± 0.0008 -1.5%
ZCR 0.0328 ± 0.0010 +3.0% 0.0628 ± 0.0018 +2.7%
Random (ours) 0.0321 ± 0.0006 +0.8% 0.0614 ± 0.0014 +0.5%
Global pop. (ours) 0.0362 ± 0.0004 +13.5% 0.0694 ± 0.0009 +13.5%

Installation

Python 3.11 is recommended.

git clone <repository-url>
cd sid-collisions
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

Reproducing the experiments

All commands run from the repository root with PYTHONPATH="./" (set inside the provided scripts) and SEQ_REC_DATA_PATH pointing to the directory where datasets should be stored (exported by you).

1. Prepare the data

export SEQ_REC_DATA_PATH=/path/to/data
bash sh/prepare_datasets.sh

For Beauty2014, Sports2014, and Toys2014 this performs: (0) download of the Amazon 2014 ratings and metadata from SNAP; (1) 5-core preprocessing with consecutive-duplicate removal; (2) the global temporal split (gts_q09_val_by_time) and the leave_one_out split; (3) all-MiniLM-L6-v2 embedding generation (set EMBED_GPU=<id> to choose the GPU, default 0).

2. Run the collision-solver comparison

bash sh/run_exp.sh all gts_q09_val_by_time     # main (global temporal split)
bash sh/run_exp.sh all leave_one_out   # supplementary (leave-one-out)

A single run can also be launched directly:

python runs/quantize_and_train.py \
    recipe=GPTRec_SID \
    dataset=Beauty2014 \
    quantization.collision_solver=global_pop \
    random_state=17 \
    dataset_params.train_dataset.shift=0 \
    dataset_params.predict_dataset.shift=0 \
    split_name=gts_q09_val_by_time

Experiment tracking

Runs log with DiskTracker (local write on disk) by default (experiment_tracker in runs/config/quantize_and_train.yaml).

Also we have ClearmlTracker for log with ClearML. If you have credentials, enable tracking with:

python runs/quantize_and_train.py ... experiment_tracker._target_=src.trackers.ClearmlTracker

If HuggingFace model downloads hang behind a proxy, set HF_HUB_DISABLE_XET=1.

Repository layout

runs/                     entry points (Hydra)
  preprocess.py           raw data -> filtered interactions + metadata
  split.py                global temporal / leave-one-out splits
  build_semantic_embeddings.py
  quantize_and_train.py   quantize items, resolve collisions, train, evaluate
  get_data/               raw data download scripts
  config/                 Hydra configs (dataset, quantizer, recipe, ...)
src/
  semantic_ids/           quantizers and collision solvers (collision_solvers.py)
  models/, modules/       GPT-2-style decoder and Lightning module
  metrics.py, callbacks.py, trackers.py
sh/
  prepare_datasets.sh
  run_exp.sh

About

Code for ACM RecSys 2026 paper "Don’t Waste the Last Token: A Quantizer-Agnostic Quality Boost from Semantic ID Collision Resolution"

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages