OpenSquilla: Local Routing Slashes AI Agent Costs 9x Without Quality Loss

OpenSquilla, a 7K-star open-source AI agent, uses on-device routing to classify each conversation turn by complexity and dispatch it to the cheapest suitable model, achieving 9x cost reduction on 25 benchmark tasks while maintaining near-identical scores, plus adaptive reasoning, dynamic prompts, and pluggable providers.

Geek Labs
Geek Labs
Geek Labs
OpenSquilla: Local Routing Slashes AI Agent Costs 9x Without Quality Loss

OpenSquilla is a microkernel personal AI agent (GitHub: TokenRhythm/opensquilla, 7k+ stars, Apache-2.0, Python, v0.5.4) that reduces model invocation costs by routing each conversation turn to the cheapest adequate model. In the project's own PinchBench benchmark of 25 tasks, a baseline agent cost $6.23 while OpenSquilla cost $0.69 — a 9x reduction — with scores of 0.9255 vs 0.9251, essentially identical.

Core Idea: Smarter Scheduling, Not Stronger Models

The agent includes a local router called SquillaRouter that runs entirely on the user's machine (LightGBM + ONNX). For every turn it examines conversation length, language type, code density, keywords, and a local semantic vector, then assigns a complexity tier:

C0: casual chat, simple QA → fast economy model

C1: routine tasks → mid-tier model

C2: complex → flagship model with deep reasoning

C3: extremely hard → multi-model ensemble

Because the prompt never leaves the device for routing, privacy is preserved.

Three-Layer Token Savings

Tiered routing as described above.

Adaptive reasoning : only turns routed to C2/C3 request the model's extended thinking; simple turns skip "thinking" tokens.

Dynamic system prompts : lightweight prompts for simple turns, full instruction set only for complex ones. Additionally, 15 built-in skills (code, GitHub, PPT/Excel/PDF, cron, etc.) are loaded on-demand — unused skills consume zero context tokens.

Microkernel + TurnRunner: Consistent Behavior Across Channels

Web UI, CLI, Feishu, Telegram, Slack, DingTalk, QQ, and WeCom all plug into a single TurnRunner loop. Tool dispatch, retry logic, and decision logs behave identically regardless of entry point. Model providers are pluggable: OpenAI, Anthropic, Gemini, DeepSeek, Qwen, Moonshot, Zhipu, SiliconFlow, plus OpenRouter aggregation — 20+ services configurable without structural changes. Chinese model support is notably complete; installation packages include Aliyun OSS mirrors.

Memory and Sandbox: Two Baselines for a Personal Agent

Memory : MEMORY.md plus date-organized Markdown notes for long-term storage. Retrieval combines SQLite full-text search with sqlite-vec vector recall; embeddings run locally via ONNX. Optional features: exponential memory decay (unused memories down-weighted) and a "dream" consolidation mode that merges memories during idle periods.

Sandbox : three policies (Standard/Strict/Locked). Linux uses Bubblewrap, macOS uses Seatbelt. A refusal ledger pauses autonomous execution after repeated permission denials. Tool results are XML-escaped to mitigate prompt injection.

Quick Start and Migration

Recommended desktop installers (macOS ARM DMG, Windows x64 EXE) or a one-line uv command:

uv tool install "opensquilla[recommended]@<release-wheel-url>" opensquilla onboard opensquilla gateway run

A one-click migration from OpenClaw and Hermes Agent is provided (dry-run first, then apply). The README openly acknowledges OpenClaw as inspiration; OpenSquilla re-implements the personal agent along a different value axis: cost.

Caveats

The cost comparison is self-reported from the project's README using PinchBench; treat the exact 9x figure as directional, not a guarantee.

Local dependencies: macOS needs libomp, Windows needs VC++ runtime. Missing deps trigger automatic fallback to single-model direct connection; restart after installing libraries restores routing.

Early stage: repo created May 2024, version 0.x, 300+ open issues — expect rough edges.

Routing gains correlate with usage pattern: if every task is high-complexity, savings diminish; mixed workloads benefit most.

The Second Track in the Model Arms Race

Instead of chasing "which model is strongest," OpenSquilla treats model selection as an optimizable engineering problem. The accompanying arXiv paper Agentic Routing: The Harness-Native Data Flywheel coins the term harness-native data flywheel : the router lives inside the agent's skeleton, and daily usage traffic becomes training data for the router. More usage → better routing → a deepening moat. Model prices will fall, but "allocate on demand" remains a first-principle cost saver. This project turns the intuition that "most conversations don't need a flagship model" into a runnable, open-source system.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LLMCost OptimizationOpen SourceAI AgentBenchmarkOn-Device Inferencemodel routingOpenSquilla
Geek Labs
Written by

Geek Labs

Daily shares of interesting GitHub open-source projects. AI tools, automation gems, technical tutorials, open-source inspiration.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.