UTTSI: Test-Time Selective Inference Boosts CTR Prediction Without Retraining

Alibaba researchers propose UTTSI, a training-free, model-agnostic framework that estimates per-sample uncertainty at inference time using frequency priors and gradient-based attribution, then adaptively allocates compute — filtering noisy features and exploring multiple inference paths only for uncertain samples — achieving a 5.3% CTR lift in a 7-day online A/B test with just 2.8× average model calls.

Alibaba International Intelligent Technology
Alibaba International Intelligent Technology
Alibaba International Intelligent Technology
UTTSI: Test-Time Selective Inference Boosts CTR Prediction Without Retraining

Background

Click-through rate (CTR) prediction is the core component of recommender systems. Recent years have seen extensive architecture innovations — DeepFM, DCN, PEPNet, HSTU — all focusing on training-time optimization. However, a fundamental mismatch exists: models predict reliably on frequent feature combinations seen during training, but their embeddings for rare or unseen combinations collapse toward random initialization. Adaptive gating (e.g., PEPNet) learns a deterministic selection function that remains unreliable for sparsely trained features. Moreover, training and inference have asymmetric objectives: training needs diverse patterns for generalization, while inference should rely only on well-learned feature interactions. This motivates test-time optimization as a complementary paradigm, yet CTR models lack per-sample confidence estimation and industrial latency budgets are extremely tight.

UTTSI: Uncertainty-Triggered Test-Time Selective Inference

UTTSI is a training-free, model-agnostic framework that plugs into any trained CTR model without modifying parameters, retraining, or changing architecture. Its core idea: at inference, estimate each sample's uncertainty, then allocate compute on demand — deep exploration for uncertain samples, zero extra cost for confident ones.

2.1 Frequency Prior Estimation

Motivation: Feature values appearing more often in training have better-trained embeddings; rare values yield unreliable embeddings.

Method: Because industrial feature spaces are huge, UTTSI uses a Count-Min Sketch probabilistic hash structure — maintaining d independent hash tables per feature field, taking the minimum across tables to estimate frequency with an upper bound while controlling storage. The normalized frequency score lies in [0, 1] for downstream use.

Benefit: Fully precomputed offline; online query overhead is near zero. Provides a data-level reliability signal for uncertainty estimation.

2.2 Dual-Signal Uncertainty Estimation

Motivation: A single signal is insufficient. Low model confidence may stem from epistemic uncertainty (sparse feature coverage) or aleatoric uncertainty (inherent label noise). Frequency priors flag rare features but cannot distinguish reliable from unreliable predictions. The two signals complement each other to precisely identify samples needing extra exploration.

Method: Perform one standard forward pass, then backpropagate to obtain gradient norms (attribution strength) for each feature embedding. Compute two confidence scores:

Model confidence: Normalized absolute logit to [0, 1]; near 0 means the model sits on the decision boundary.

Frequency confidence: Attribution-weighted average of per-feature frequencies — measuring not just how common features are, but how common the features actually driving the prediction are.

The final uncertainty score is a weighted complement of the two confidences; the number of exploration paths scales proportionally — higher uncertainty allocates more paths .

Benefit: Continuous allocation enables fine-grained budget control; average ~2.8 model calls cover all samples.

2.3 Adaptive Feature Filtering

Motivation: Each sample may contain noisy features — undertrained embeddings that disturb prediction — which should be filtered before entering the model.

Method: For each feature, compute a composite score (frequency reliability + gradient attribution strength) and compare against a per-field adaptive threshold . Thresholds are computed separately per field from training data because score distributions differ vastly across user-ID, item-attribute, and context fields; a single global threshold would systematically over-filter or under-filter.

Benefit: Features are removed only when both frequency and attribution are low, a conservative strategy that avoids discarding informative rare features. For confident samples (uncertainty = 0), the filtered features go straight to output — their combinations are already well learned, so a single refined prediction suffices.

2.4 Selective Multi-Path Exploration & Consistency Aggregation

Motivation: Filtering addresses feature-level noise, but uncertain samples still face combinatorial-level uncertainty — even if each retained feature is reliable individually, their joint interaction pattern may be sparse in training.

Method: For uncertain samples (uncertainty > 0), UTTSI runs iterative Bernoulli sampling on the filtered feature subset to generate multiple inference paths. At each step, a candidate feature is selected with probability proportional to its reliability-attribution composite score — informative rare features are preferentially kept while noisy features are more likely excluded. All paths (plus the refined path, K +1 total) are aggregated via consistency-weighted voting : predictions agreeing with the majority receive higher weight; outlier paths are automatically down-weighted.

Benefit: All exploration paths are fully independent and parallelizable. In a horizontally scaled serving system, worst-case latency equals a single forward pass.

Experiments

Evaluated on four large-scale datasets: Criteo (45M), Avazu (40M), KDD12 (60M), and Industrial (513M), against 10+ SOTA models including FM, DeepFM, DCN, AutoInt, GDCN, MaskNet, PEPNet, HSTU.

3.1 Main Results

OptFu+UTTSI achieves the best AUC on all four datasets, with statistically significant improvements over the strongest baseline OptFu (p < 0.01). Gains are most pronounced on KDD12 and Industrial — the datasets with higher feature sparsity — directly validating UTTSI's targeted improvement for sparse features.

3.2 Ablation Study

Three ablations confirm each component's necessity:

Remove dual-signal estimation (logit confidence only, RD): consistent drop across all datasets — single signal cannot separate epistemic from aleatoric uncertainty.

Remove attribution-guided sampling (replace with uniform random sampling, RA): largest drop on high-sparsity datasets — per-path quality matters more when data is sparse.

Single-path instead of multi-path (WS): even with correct signals and sampling, single path underperforms multi-path ensemble — robustness comes from exploration diversity.

3.3 Online A/B Test

Deployed on a large-scale e-commerce platform for a 7-day A/B test (Apr 15–21, 2026) against a PEPNet baseline. CTR improved +5.3% relative (p < 0.01) .

Compute overhead: ~62% of samples require only feature filtering (no multi-path exploration); average ~2.8 model calls per sample (2.8× base cost). All paths run in parallel, so actual latency matches a single forward pass.

3.4 Uncertainty Calibration

Samples split into low/medium/high uncertainty tiers (Industrial dataset). High-uncertainty samples gain the largest AUC improvement (+0.0088). The uncertainty score achieves a Spearman rank correlation of 0.91 with prediction error (vs. 0.76 for logit confidence alone), confirming the dual-signal design reliably directs extra budget to the model's weakest spots.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

CTR predictionuncertainty estimationrecommendation systemsCount-Min SketchCIKM 2026selective inferencetest-time inference
Alibaba International Intelligent Technology
Written by

Alibaba International Intelligent Technology

Alibaba International Tech – Official channel of the Intelligent Technology team, sharing cutting‑edge AI applications and innovations in Alibaba's global e‑commerce business.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.