ME-Decoding: Reframing LLM Decoding as Ensemble Pruning with Semantic Redundancy Control
Renmin University researchers propose ME-Decoding, a method that selects candidate tokens by jointly optimizing probability confidence and semantic diversity via Mahalanobis distance, outperforming probability-only truncation on reasoning and open-ended generation across multiple models and temperatures with only 3% latency overhead.
Large language model decoding traditionally selects the next token by sampling from a probability distribution truncated by methods such as Top-p or Min-p. These approaches retain tokens based solely on individual probability, ignoring that multiple high-probability tokens may be semantically similar — essentially redundant expressions of the same generation path — while a slightly lower-probability token offering a distinct semantic direction may be discarded.
Researchers from Renmin University (first author Xue Dunyao, corresponding authors Dai Wenlin and Meng Cheng) introduce Mahalanobis-Ensemble Decoding (ME-Decoding) , accepted at EMNLP 2026 Main Conference. The paper, titled "Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning" (arXiv:2609.18723, code: https://github.com/sapphirexdy/ME_decoding), reframes token selection as a dynamic subset optimization problem: choose a sampling support that maximizes joint confidence while minimizing internal semantic redundancy.
Method: Three-Step Ensemble Pruning
Step 1: Mahalanobis-Ensemble Score (MES) for Joint Probability-Redundancy Evaluation
Given a candidate subset S, probability vector p, and token similarity matrix K, the paper defines Mahalanobis-Ensemble Energy (MEE) : MEE(S) = p_S^T K_S^{-1} p_S where K_S is the submatrix of K for tokens in S. MEE captures the joint contribution of the subset: high probabilities increase the score, while high similarity (redundancy) increases K_S 's eigenvalues, reducing the quadratic form. To automatically determine subset size, a scale penalty based on the current distribution's uncertainty is added, yielding MES = MEE(S) / |S|^α (α adapts to entropy). MES rises only when a new token's marginal gain outweighs the size penalty.
Step 2: Adaptive Gaussian Kernel for Local Semantic Structure
Similarity between tokens i and j is computed from normalized embeddings e_i, e_j via a Gaussian kernel: K_ij = exp(-||e_i - e_j||^2 / (2σ^2)) The bandwidth σ is set adaptively per decoding step using the probability-weighted semantic dispersion of the current candidate pool. When the distribution is concentrated, σ shrinks, making the kernel focus on very local neighborhoods; when dispersed, σ expands to capture broader semantic relations. This ensures the redundancy measure matches the current step's semantic scale.
Step 3: Greedy Expansion with Automatic Stopping
Exhaustive subset search is infeasible. ME-Decoding starts with the highest-probability token, then iteratively adds the candidate yielding the largest marginal increase in MEE. If MES increases, the set expands; otherwise the process stops immediately. The final subset's original probabilities are renormalized and sampled from. The resulting support sets are typically small; with early stopping, computation scales near-linearly with candidate pool size.
Experiments
Evaluated on Qwen3-4B, Phi-4-mini, and Mistral-7B across temperatures T=1.0, 1.5, 2.0. Baselines: Min-p, p-less, Top-H, Top-W.
Reasoning: Robustness Across Temperatures
On GSM8K and GPQA, ME-Decoding achieves the highest average accuracy among compared methods: 72.66% on GSM8K and 32.96% on GPQA . At high temperatures (T=2.0), probability-truncation baselines degrade sharply, while ME-Decoding maintains stable performance, demonstrating cross-model, cross-temperature reasoning robustness.
Open-Ended Generation: Instruction Following & Dialogue
Using DeepSeek-V4-Pro as judge on AlpacaEval and MT-Bench, ME-Decoding attains the highest average win rate and average score respectively, indicating improved instruction following and multi-turn dialogue quality.
Ablation: Is Improvement Just from Support Size or Entropy Changes?
The paper matches probability-based baselines to ME-Decoding on three statistics: average support size, post-pruning entropy, and average pairwise cosine similarity within the support. Even when these are matched, ME-Decoding still achieves the highest GSM8K/GPQA accuracy, highest MES, and lowest semantic similarity. This confirms the gain stems from MES's joint probability-redundancy evaluation, not merely from smaller supports, lower entropy, or lower average cosine similarity.
Theoretical Guarantees & Efficiency
The paper provides sufficient conditions for the greedy MES trajectory to be unimodal, with corresponding approximation guarantees. Empirically, with early stopping disabled, all reasoning trajectories on GSM8K and GPQA exhibit unimodal MES curves, validating the stopping rule.
On synthetic CPU benchmarks, ME-Decoding runs 3-7× faster than Top-W . End-to-end GPU tests show total latency increases only ~3% over the fastest probability-truncation baseline — a negligible fraction of model forward-pass time.
Summary of Contributions
Redefines token selection : treats candidates as ensemble members to be pruned, extending pointwise probability selection to subset-level joint optimization.
Jointly leverages confidence and semantic geometry : MES retains high-probability tokens while discounting directions already covered by the current set, yielding a more compact, informative sampling support.
Efficiently adapts to the current distribution : kernel bandwidth adjusts per step; greedy search stops automatically when MES plateaus, controlling support size and compute.
ME-Decoding demonstrates that decoding can be understood as dynamic ensemble pruning rather than mere probability thresholding, improving average performance on tested tasks while remaining stable across temperatures and models.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
