Tagged articles

Agent Safety

9 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 17, 2026 · Artificial Intelligence

Beyond Prompts: Harness-Policy Co-Evolution for Agent Safety by Shanghai AI Lab

Shanghai AI Lab and university collaborators propose SHE and SafeEvolve, two frameworks that evolve agent safety by learning from execution trajectories: SHE updates a modular safety harness via trajectory-driven evolution, while SafeEvolve distills verified harness experience into the policy model through SFT and RL, reducing attack success rates on benchmarks.

Agent SafetyHarness EvolutionLLM Agents
0 likes · 12 min read
Beyond Prompts: Harness-Policy Co-Evolution for Agent Safety by Shanghai AI Lab
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 12, 2026 · Artificial Intelligence

SafeEvolve: Co-Evolving Harness and Policy for Self-Improving Agent Safety

SafeEvolve introduces a co-evolution framework where Agent Harness and Policy jointly learn from execution trajectories, reducing attack success rates to 0.79% on AgentDojo and 2.42% on Qwen3-4B while improving task utility, enabling continuous safety improvement from real-world experience.

AI safetyAgent SafetyHarness-Policy Co-Evolution
0 likes · 9 min read
SafeEvolve: Co-Evolving Harness and Policy for Self-Improving Agent Safety
DataFunTalk
DataFunTalk
Jul 25, 2026 · Artificial Intelligence

When Errors Spread Among AI Agents, Who Pulls the Brakes? Safe Collaborative Growth

The article analyses how self‑evolving AI agents shift from simple tool use to autonomous planning, proposes a fast‑slow thinking architecture, curriculum learning, organizational structures, continuous evaluation, and a four‑layer safety framework to ensure they grow responsibly while collaborating with humans.

AI AgentsAgent SafetySelf-Evolution
0 likes · 17 min read
When Errors Spread Among AI Agents, Who Pulls the Brakes? Safe Collaborative Growth
ThinkingAgent
ThinkingAgent
Jul 25, 2026 · Artificial Intelligence

From Next Token to Deployable AI: A Comprehensive Overview of Large Model Technology

This article maps the entire large‑model production chain—from data collection, token prediction, and architecture design through training, alignment, inference, multimodal perception, agentic action, deployment, evaluation, and safety—highlighting key engineering decisions, trade‑offs, and concrete examples.

Agent SafetyLarge Language ModelsRetrieval-Augmented Generation
0 likes · 47 min read
From Next Token to Deployable AI: A Comprehensive Overview of Large Model Technology
Architect Practice
Architect Practice
Jun 1, 2026 · Artificial Intelligence

When AI Rate Limiting Goes Wrong: A Four‑Dimension Framework and Three‑Layer Gateway in Practice

A midnight alarm at a fintech AI platform revealed that traditional QPS throttling missed a runaway Agent that consumed hundreds of times more tokens, prompting a detailed analysis of four token‑based limiting dimensions, three‑layer gateway design, agent‑specific controls, semantic caching, and tool selection to prevent similar “ghost avalanche” failures.

AI rate limitingAgent SafetyLLM Operations
0 likes · 20 min read
When AI Rate Limiting Goes Wrong: A Four‑Dimension Framework and Three‑Layer Gateway in Practice
ArcThink
ArcThink
May 23, 2026 · Artificial Intelligence

Why a 65‑Line CLAUDE.md Can Put Brakes on AI Coding Assistants

The article dissects the Karpathy‑inspired 65‑line CLAUDE.md file, showing how its four concise constraints—think before coding, simplicity first, surgical changes, and goal‑driven execution—prevent common AI coding agent failures, why the approach works, and its limits.

AI coding assistantAgent SafetyCLAUDE.md
0 likes · 14 min read
Why a 65‑Line CLAUDE.md Can Put Brakes on AI Coding Assistants
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Apr 14, 2026 · Artificial Intelligence

Balancing Usability, Fun, and Safety: How Fudan’s Post‑00 Team Built XSafeClaw for Controllable AI Agents

Amid soaring hype for autonomous agents, a Meta incident exposed how hidden execution steps can cause real‑world damage, prompting Fudan’s XSafeClaw project to deliver a visual, layer‑by‑layer security framework that makes agent behavior observable, auditable, and safely interceptable.

Agent SafetyObservabilityRuntime monitoring
0 likes · 10 min read
Balancing Usability, Fun, and Safety: How Fudan’s Post‑00 Team Built XSafeClaw for Controllable AI Agents
AI Large Model Application Practice
AI Large Model Application Practice
Apr 13, 2026 · Artificial Intelligence

How Hermes-Agent Enables Self‑Learning Skills for Autonomous AI Agents

Hermes‑Agent introduces a novel self‑learning Skill system that lets AI agents automatically capture, refine, and patch reusable knowledge from complex tasks, using a dual front‑end awareness and back‑end inspection loop, reinforced by safety guards and a reinforcement‑learning training pipeline.

AI AgentsAgent SafetySkill Management
0 likes · 18 min read
How Hermes-Agent Enables Self‑Learning Skills for Autonomous AI Agents
AntTech
AntTech
Apr 2, 2026 · Information Security

How ClawAegis Secures OpenClaw AI Agents with a Native Immunity System

Ant Group’s AI Security Lab and Tsinghua University have open‑sourced ClawAegis, a native security‑immune framework for OpenClaw agents that protects the entire lifecycle—from initialization to execution—by detecting malicious skill injections, memory poisoning, permission abuse, and providing dynamic auditing, configurable policies, and resource‑level safeguards.

AI securityAgent SafetyOpenClaw
0 likes · 5 min read
How ClawAegis Secures OpenClaw AI Agents with a Native Immunity System