Tagged articles

Skill²‑Bench

1 articles · Page 1 of 1
PaperAgent
PaperAgent
Aug 11, 2026 · Artificial Intelligence

Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM

Researchers introduce Skill‑Entropy, a metric quantifying the difficulty of switching between reasoning skills in long‑horizon tasks, build the 558‑skill Skill²‑Bench, and show that Skill‑Entropy‑RL training dramatically improves cross‑skill performance of LLMs such as Qwen3, closing the gap observed in standard benchmarks.

Cross‑Skill ReasoningLLM BenchmarkingLong‑Horizon Tasks
0 likes · 12 min read
Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM