MIT Researchers Embed Generalization in Harness: Short-Task Training Unlocks 32× Length Extrapolation
MIT CSAIL’s study shows that training a Recursive Language Model with a harness that keeps each model call locally in-distribution allows the system to extrapolate up to 32-fold longer sequences, achieve superior cross-domain transfer, and outperform transformer baselines despite higher training cost.
