Why Large Models Excel at Table Lookup Yet Fail at Future Prediction – Insights from TopBench
TopBench, a new benchmark for implicit predictive reasoning in table question answering, shows that current large language models can retrieve tabular facts but often miss the hidden prediction intent, leading to low accuracy across four task types and revealing two key bottlenecks: intent alignment and robust modeling.
