SMELT: Looped Transformers Outperform Baselines Under Fair Budget Matching
The SMELT framework from Tsinghua and ByteDance Seed fairly compares Looped Transformers against baselines by matching compute, parameters, and KV cache, finding that looping the middle 50% of layers twice with a larger depth-width ratio consistently reduces validation loss across scales, saves 6.8–18% training compute, and yields downstream gains beyond loss reduction.
