Tagged articles

context caching

3 articles · Page 1 of 1
Top Architecture Tech Stack
Top Architecture Tech Stack
Sep 9, 2026 · Artificial Intelligence

Xiaomi's MiMo Desktop: Model-Harness Integration Signals Agent System Engineering Shift

Xiaomi launches MiMo Desktop, a desktop AI agent integrating MiMo-X-Pro and MiMo-X-Flash models with a model-native harness that automates model routing, achieves 99% cache hit rates, and demonstrates that agent performance depends on system-level context management rather than model capability alone.

AI AgentsMiMo DesktopMiMo-X-Flash
0 likes · 10 min read
Xiaomi's MiMo Desktop: Model-Harness Integration Signals Agent System Engineering Shift
Architecture Development Notes
Architecture Development Notes
Sep 8, 2026 · Artificial Intelligence

Rethinking Agent Composition: Single-Loop Skills vs. Sub-Agent Handoffs

This article analyzes why default multi-agent architectures leak state in long conversations, advocating for single-loop agents with dynamically loaded skills based on usage frequency, using Anthropic's commerce-agents reference implementation to illustrate caching-aware design, handoff vs. delegation distinctions, and evaluation strategies.

Agent ArchitectureAnthropicLLM applications
0 likes · 10 min read
Rethinking Agent Composition: Single-Loop Skills vs. Sub-Agent Handoffs
AI Insight Log
AI Insight Log
Feb 14, 2026 · Artificial Intelligence

Why Claude Code’s Context Caching Suddenly Fails and Costs Skyrocket

After Claude Code was updated to version 2.1.37, developers observed a sharp drop in context‑caching hit rates, causing unexpected cost spikes and slower responses, and a community investigation revealed that random headers and spaces injected by the tool break the model’s cache matching.

APIAnthropicBug Analysis
0 likes · 6 min read
Why Claude Code’s Context Caching Suddenly Fails and Costs Skyrocket