Choosing the Right Agent Harness: NTU Study Shows Model-Harness-Task Fit Matters Most

NTU researchers evaluated four configurable harnesses across five models and three benchmarks, finding harness performance depends on specific model-harness-task combinations rather than fixed rankings, with native harnesses not always optimal and cost not correlating with performance.

Machine Heart
Machine Heart
Machine Heart
Choosing the Right Agent Harness: NTU Study Shows Model-Harness-Task Fit Matters Most

Experimental Setup

The study from Nanyang Technological University (NTU) by Professor An Bo's team, titled "Finding the Right Fit: Model–Harness Interactions across Agent Tasks" (arXiv:2610.00917), systematically evaluates how harness choice affects agent performance. Four configurable harnesses were tested: OpenHands (openhands-tools 1.44.1), DeepSeek Harness (DSH) (0.1.1-rc.2), PI (v0.84.4), and openJiuwen (0.1.18). Each harness was paired with five models: Claude Opus 5 , GPT-6 Astra , GLM-5.3 , Kimi K3 , and DeepSeek V4 Pro , yielding a 4×5×3 cross design across three task suites: TUA-Bench (120 general terminal tasks), ALE-CLI (99 professional workflow tasks), and Terminal-Bench 4 (63 challenging command-line tasks). Two native pairings — Claude Code–Claude (v2.1.251) and Codex–GPT (v0.150.1) — served as references but did not participate in the full cross evaluation. All harnesses called models via OpenRouter with high reasoning intensity and default context/retry settings. Scoring used fixed denominators per benchmark; partial credit allowed on TUA-Bench and ALE-CLI, pass/fail on Terminal-Bench 4. The final matrix contains 66 records (60 cross + 6 native).

Key Results

Harness Significantly Alters Model Rankings

On Terminal-Bench 4, Claude Opus 5 with OpenHands solved 36 tasks (60.32%), while GPT-6 Astra solved 31 (49.21%) — a 7.94-point lead for Claude. Switching to PI reversed the order: GPT-6 Astra solved 38 tasks (60.32%), Claude only 19 (30.16%), giving GPT a 30.16-point advantage. The harness change hurt Claude (−17 tasks) but helped GPT (+7 tasks), proving that model rankings are not intrinsic but reflect model-harness adaptation.

Best Harness Varies by Task

Four of five models changed their top-performing harness across the three benchmarks. GLM-5.3 and DeepSeek V4 Pro each had a different best harness on every benchmark (openJiuwen on TUA-Bench, OpenHands on ALE-CLI, DSH on Terminal-Bench 4). GPT-6 Astra preferred openJiuwen on TUA-Bench but PI on the other two. Claude Opus 5 led with Claude Code on TUA-Bench and ALE-CLI, yet OpenHands outperformed on Terminal-Bench 4. Only Kimi K3 consistently achieved its highest score with openJiuwen across all three suites, leading the next-best configuration by 5.61, 6.91, and 11.11 points respectively.

Native Harnesses Are Not Always Optimal

Claude Code gave Claude its top score on TUA-Bench and ALE-CLI but fell behind OpenHands on Terminal-Bench 4. Codex–GPT scored lower than both openJiuwen–GPT and PI–GPT on every benchmark. Thus, vendor-provided harnesses do not guarantee the best fit for their own models.

Higher Cost Does Not Guarantee Better Performance

On Terminal-Bench 4 with GPT-6 Astra, PI cost $293.30 and achieved 60.32% score, while DSH cost $1,256.48 (4.3× more) but scored only 52.38%. Cost differences stem from how harnesses organize model calls and context. openJiuwen reused 94–99% of input tokens from cache, versus 49–77% for other harnesses, indicating superior context management.

Trajectory Analysis: Why Interactions Differ

The authors examined 192 failure events from 10 OpenHands–PI pairs on Terminal-Bench 4 (five models) plus 6 Kimi K3 pairs (openJiuwen vs PI). In 180 events the model initiated recovery; diagnosis or targeted fix was the most common action (133 occurrences, 116 successes). The critical factor was whether the harness turned failures into usable feedback.

Case Study: TUA-Bench Task 056-move-textbox-left

Kimi K3 made the same error under both harnesses: a GIMP script run via pipe missed an exit command, causing the command to wait indefinitely. Under PI , the shell had no default timeout and the model did not set one; the command hung until the 39.8-minute deadline, yielding zero score. Under openJiuwen , the first two failures returned error messages; the third hung but the shell returned a timeout after 300 seconds. The model used that signal to diagnose the missing exit, verify the generated image, and check canvas size, background, and text position — completing the task in 11 minutes with a score of 1. OpenHands and DSH also enforce command timeouts; the point is that returning a failure signal enables recovery .

Model-Specific Habits Determine Harness Fit

GPT-6 Astra proactively set timeouts in 56–67% of its shell calls under PI, so PI's lack of a default timeout mattered little. GPT-6 Astra achieved its best Terminal-Bench 4 and ALE-CLI scores with PI. Harness intervention should match model tendencies: reduce scaffolding where the model already handles a step, add support where the model tends to omit it.

Why openJiuwen–Kimi K3 Is Stably Strong

Per-task comparison shows openJiuwen led on 22/120 (TUA-Bench), 25/99 (ALE-CLI), and 8/63 (Terminal-Bench 4) tasks. Even removing the three largest-margin tasks, average leads remained 3.11, 3.88, and 6.35 points. Trajectory analysis identifies three complementary mechanisms:

Tool interface: OpenHands merges "write file" and "edit file"; Kimi often omitted required parameters, triggering 55 terminations (48 from Kimi). openJiuwen splits them into two tools, aligning with Kimi's calling style.

Command timeout: openJiuwen returns a timeout signal when a command stalls, converting a hang into actionable feedback.

Truncation continuation: When output is cut by length limits, openJiuwen preserves existing reasoning and prompts the model to continue. In 6 Kimi paired trajectories, this mechanism was decisive in 2 cases.

Conclusion

The paper argues that the evaluation unit for agents should be the full configuration of model, harness, and task , not model or harness alone. If the model is fixed, compare its performance across harnesses on the target task; if the harness is fixed, do not transfer model rankings from other harnesses. When moving an agent to a new business domain, re-validate the combination. For multi-task systems, preserving configuration flexibility is valuable, but "allowing choice" differs from "automatically choosing correctly" — the reported best configurations were identified post-hoc. Validating automatic routing requires a dev-set selection phase followed by held-out testing, with latency, API cost, and tool overhead weighed alongside task completion.

Agent evaluation should treat the model, harness, and task as a single integrated configuration.

References: Paper (https://arxiv.org/abs/2610.00917), Code (https://github.com/liyix/finding-the-right-fit), Dataset (https://huggingface.co/datasets/yixuanli97/finding-the-right-fit), TUA-Bench (https://arxiv.org/abs/2606.28480), ALE (https://arxiv.org/abs/2606.05405), Terminal-Bench (https://arxiv.org/abs/2601.11868), OpenHands (https://github.com/OpenHands/OpenHands), DeepSeek Harness (https://github.com/deepseek-ai/deepseek-harness), PI (https://github.com/earendil-works/pi), openJiuwen (https://github.com/openJiuwen-ai/), Codex (https://github.com/openai/codex), Claude Code (https://github.com/anthropics/claude-code).

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsAgent Evaluationagent harnessTerminal-BenchopenJiuwenOpenHandsModel-Harness FitNTU Research
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.