Highlights and Insights from the Shenzhen Stop of the Agent Observation & Optimization Tour
The Shenzhen session of the Agent Observation & Optimization tour gathered nearly a hundred technologists to discuss evaluation paradigms, showcase AgentScope 2.0’s enterprise‑grade features, demonstrate a Java e‑commerce chatbot assessment with AgentLoop, and offer a hands‑on workshop, while previewing the upcoming Shanghai event.
Event Overview
The Shenzhen stop of the "Agent Observation & Optimization – Close the Loop" tour concluded successfully, co‑hosted with AgentScope and attended by close to one hundred technical professionals. The theme centered on "Evaluate / Assessment and Experimentation," featuring deep dives into agent evaluation paradigms, the AgentScope 2.0 enterprise foundation, and a practical e‑commerce customer‑service assistant case built with AgentLoop.
Topic 1 – Evaluation Paradigm
Speaker Xie Jingjie explained the gap between feeling an agent is better and having measurable proof. He introduced a trajectory‑based evaluation paradigm that expands a single dialogue into a full reasoning trace, using an "Agent‑as‑a‑Judge" to automatically score production traces while maintaining over 90% agreement with human experts. Evaluation is embedded in a "run‑evaluate‑experiment‑gate" loop, with built‑in evaluators such as task completion, groundedness, and tool‑selection rationality quantifying A/B differences and preventing negative iterations.
Topic 2 – AgentScope 2.0 Enhancements
Speaker Liu Jun described how enterprise‑grade agents differ from single‑call LLM apps, requiring long‑running tasks, state persistence, and multi‑agent collaboration. AgentScope 2.0 introduces the Harness runtime kernel and a Workspace as the factual source. Features include an abstract file system, four‑layer context compression, and a dual‑layer long‑term memory that preserve context and facts in extended conversations. The platform also supports sub‑agent orchestration, delegation, parallelism, asynchronous notifications, and integrates the E2B Sandbox for file and shell isolation with snapshot recovery.
Topic 3 – AgentLoop Java E‑commerce Assistant Evaluation
Using a Java‑based e‑commerce customer‑service assistant built on AgentScope, the demo ran six fixed service cases, generating A/B traces for two prompt versions (12 traces total). The AliyunJavaAgent 5.1.2 automatically reported these traces to AgentLoop, where three built‑in evaluators—task completion, groundedness (evidence support), and tool‑selection rationality—performed batch assessment.
Results showed that the B‑version prompt retained task completion and tool‑selection quality while significantly improving groundedness. In high‑risk cases, B avoided fabricating platform details, phone numbers, or refund timelines, and prioritized official sources over conflicting forum content. The workflow serves as a pre‑release gate: only when key metrics do not decline and no new severe errors appear does the next version proceed.
Hands‑On Workshop
Attendees followed a step‑by‑step lab that connected a Java AgentScope application to AgentLoop, generated A/B traces, created batch evaluations, and interpreted the results, demonstrating the full path from intuition to evidence‑based assessment.
Next Event Preview
The final stop of the tour will be in Shanghai on September 4, 14:00‑17:00, co‑hosted with LangChain and focusing on "Optimize / Evolution." Registration is by approval, with successful applicants receiving SMS notifications.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Alibaba Cloud Native
We publish cloud-native tech news, curate in-depth content, host regular events and live streams, and share Alibaba product and user case studies. Join us to explore and share the cloud-native insights you need.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
