Cut Agent Token Costs 75% by Eliminating Wasted Turns, Not Just Compressing Prompts
AgentSight, an observability component of Alibaba Cloud's Agentic OS, analyzes agent execution traces to identify wasted turns and reusable experience, demonstrating a 75% token reduction (36K to 8.9K) and 65% time savings in a frontend development case study by capturing project-specific verification rules in AGENTS.md.
Background: Token Waste from Repeated Agent Mistakes
Agent token costs include significant waste: the same project pits are re-encountered across tasks, and previously resolved information is re-investigated. Each unnecessary turn burns tokens.
Existing Approaches: Per-Call Compression
Industry-standard cost reduction focuses on compressing single calls: tools like Headroom (headroomlabs-ai/headroom) compress redundant structures, rtk (rtk-ai/rtk) filters shell output, and context window management like Claude Code's compaction and LangChain's trim_messages use sliding windows and summarization. These reduce per-call volume but cannot address whether a turn should have occurred at all.
AgentSight: Observability for Turn-Level Optimization
AgentSight is the observability component of the Agentic OS (ANOLISA) intelligent agent operating system. It runs on the user's machine, parsing trace data locally by default. The current version requires no agent code changes: on macOS it scans local JSONL session files; on Linux it uses a local session collection mode by default, with full eBPF pipeline available (requires root, Linux 5.8+, BTF). It supports Kubernetes/Docker container identification and integrates with Claude Code, Codex CLI, and Qwen Code.
Trace Collection and Reconstruction
Raw data is fully reconstructed. In full eBPF mode, kernel-side traffic capture combines with application-layer log parsing. Disparate HTTP/1.x, HTTP/2, and SSE requests and responses are reassembled into individual calls, session turns, and complete sessions. Built-in parsers cover OpenAI, Anthropic, Alibaba Cloud Bailian, and OpenAI-compatible endpoints, outputting OpenTelemetry GenAI semantic conventions and ATIF v1.7 standards. Traces include prompts, model outputs, tool call parameters and results, linked to process trees, file writes, and network actions, enabling correlation of an LLM call with triggered system actions and the information driving the next step.
Multi-Dimensional Analysis: Cost, Performance, Accuracy
After reconstruction, AgentSight analyzes execution from three angles:
Cost analysis decomposes context windows per turn, using token flame graphs to show amplification from repeated history replay, and identifies duplicate calls, compressible prompts, and invalid turns.
Performance analysis breaks latency into model wait, tool execution, and idle gaps to pinpoint true bottlenecks.
Accuracy analysis uses semantic recognition to locate defects, attributing them to Skill, tool, or context, and detects multi-turn spinning. The system includes 18 session interruption detectors and ready-to-use dashboards filterable by time, agent, and session.
Actionable Reports with Before/After Comparison
Reports pinpoint recommendations to specific calls, suggesting adjustments to Skill definitions, context organization, and prompt structure. Results are retained for re-running similar tasks to compare before/after data. All suggestions are read-only; the user decides adoption.
Workflow: Select Session, Analyze, Review Report
The dashboard lists collected sessions with source/agent filters and semantic search (e.g., "fix build error"). Clicking analyze launches a dedicated optimization agent that inspects the session end-to-end.
The report page shows a trajectory summary, basic stats (issue count, tool calls, total latency, events), and three tabs:
Accuracy : Defects table with phenomenon, type (Workflow, Tool Error, Skill Gap), attribution, fix location, confidence. Expanding reveals full root-cause analysis and copyable optimization prompts for Skill definitions or AGENTS.md.
Performance : Latency breakdown (model inference, tool execution, user idle) with proportion chart and slowest-calls table listing tools and commands.
Cost : Total tokens and peak context size. Stacked bar chart replays context window per step, split into static region, user prompt, assistant history, tool returns. Clicking a step shows exact composition. Red line tracks output tokens per step. Sudden context growth without task progress flags waste. Bottom waste analysis quantifies duplicate calls, compressible prompts, invalid turns, and estimable token savings.
Case Study: Frontend Verification Port Confusion
Scenario from AgentSight's own frontend development . The hybrid architecture has two frontend access paths: :3004 (standalone Vite dev server) and :7396 (frontend embedded in Rust binary). The agent modified AgentSessionsPage.tsx to add a sub-agent count badge, then verified on :7396 where changes weren't visible without build:embed and cargo build. It spent multiple turns debugging cache, build artifacts, and packaging before realizing the correct verification path was :3004.
Initial run metrics : 36K tokens, 18 LLM calls, 21 tool calls, 370 seconds. Context window stacked bars show sharp growth from step 4 (start of misguided debugging). Peak context 3.4K; by step 17 tool output occupied 63% of context, largely from prior screenshots and command returns.
Waste analysis flagged a high-confidence "predictable pit" and produced an executable optimization prompt: "Prefer :3004 dev server for frontend verification. Do not verify on embedded frontend :7396 without running build:embed and cargo build."
The rule was added to AGENTS.md . The article emphasizes that experience rules should be concrete actions: "Prefer :3004 for frontend verification; do not check :7396 without rebuild" is more effective than vague reminders.
Re-run after adding rule : Agent chose correct verification path immediately. Trajectory summary collapsed to five key stages with no repeated debugging.
Total tokens: 36K → 8.9K ( ~75% reduction )
LLM calls: 18 → 8
Latency: 370s → 130s ( ~65% reduction )
Peak context: 3.4K → 1.6K
Context bars show stable growth; waste analysis found no further issues.
Conclusion: Continuous Turn Reduction Loop
Compression tools reduce per-turn tokens; AgentSight reduces the number of turns by surfacing where tokens are spent and codifying reusable experience. The loop—run task, review report, write rule, re-test—compounds: each effective rule gives subsequent tasks a clearer action basis, reducing repeated mistakes and lowering token spend over time. As agent capabilities grow, managing cost separately remains essential; eliminating repeated errors reserves compute for novel problems.
Source code and product links: https://github.com/agentic-os-org/ANOLISA (open source), https://help.aliyun.com/zh/alinux/agentic-os-getting-started (Alibaba Cloud product).
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
