Why 80% of Agent Decisions Fail: Deep Dive into ADPS Perception Patterns P1–P4
The article reveals that most agent failures stem from perception flaws, not model limits, and details ADPS's four-layer perception funnel, four ingestion modes, four core design patterns (Context Triage, Semantic Compaction, Progressive Discovery, Multi-Modal Fusion), three anti-patterns, and a security warning — all grounded in production case studies from Tencent, Weibo, and game teams.
The article opens with a striking claim from Tencent expert engineer Zhang Dong: "80% of reasoning decisions fail in production not because the model is weak, but because the perception layer is broken." His team's data shows today's LLMs already solve 60–80% of real tasks; only 20–40% need heavy RAG, SFT, or RL. The bottleneck is what the model sees — missing input, lost context, noisy signals — which caps the model's innate ability.
Perception Is Not an Input Layer
ADPS defines perception as: turning heterogeneous, noisy, ultra-long raw input into high signal-to-noise representations the model can use. It governs four questions: what to look at, how much to compress, how deep to drill, how to fuse. Crucially, tool outputs are also perception — traditional tools were built for human eyes (tables, pagination, color), but agents need machine-readable schemas, statuses, receipts. If tool calls are flaky, check whether the tool still returns human-oriented formats.
Four-Layer Perception Funnel (Enterprise Reference Structure)
Layer 1 – Signal Input: Protocol adaptation, auth, fault tolerance, idempotency. Raw signals enter untouched.
Layer 2 – Data Preprocessing: Cleaning, deduplication, normalization, filtering. Standardized input stabilizes model output (which is hypersensitive to input variance).
Layer 3 – Semantic Aggregation: Context stitching, entity linking, semantic completion, cross-source verification. This is the divide between agent perception and traditional alerting . Example: a lone SQL injection alert becomes "business context + historical handling + team risk level" structured cognition.
Layer 4 – Event Adjudication: Outputs confidence scores, classifications, trigger recommendations — "advisor, not general." Perception never executes business actions; decoupling enables independent calibration and iteration.
Four Ingestion Modes (Often Combined)
Event-driven (passive): Low latency, low resource, simple. E.g., code push, vulnerability ticket.
Periodic polling (active): Full coverage, controllable, simple. High latency (daily/weekly), heavy resource at scale.
Streaming perception: Parallel scalable, high throughput, near-real-time. Complex ops/debugging.
Multi-source fusion: Cross-source input, context completion. Complex scenarios: high-severity threats, intelligent fault localization, full-chain audit. "Multi-source" ≠ "multi-modal" — sources determine permissions/timeliness/cross-check; modalities determine parsers/representations.
Production proof — Code Audit Chain: Combines event trigger + daily full scan + multi-source evidence aggregation. Three layers: ingestion (4 signal sources), aggregation (code context completion, historical vuln linking, team habit profiling → structured cognition), output (critical block / regular alert / dev suggestion). Measured results: ~40% false-positive drop, 90% accuracy, perception latency from hours to minutes, full-code coverage retained.
ADPS Four Perception Patterns (P1–P4)
P1 Context Triage (Perception × Routing)
Problem: Candidate info exceeds context window. Solution: Explicitly classify every piece:
Must enter context (P0): current goal, hard constraints, recent tool results, key evidence, pending actions.
Compressed entry.
Handle (lazy fetch) — handle must be resolvable.
Discard.
Engineering rule: Triage rules must be explicit, not left to model "intuition." Common failure: Equating relevance with priority. In long-horizon tasks, original goal, non-goals, and key evidence may have low relevance but must be protected. Three constraints added post-workshop: goal-backward constraint, missing-item logging, input scope.
P2 Semantic Compaction (Perception × Chain)
Problem: Context grows unbounded. Risk: Compression discards decision-critical detail. Must preserve: numbers, paths, error codes, decision rationales, rejected alternatives, evidence citations. Post-workshop, "architectural decisions" added to protected list for coding agents. Textbook counter-example: Compressing ConnectionError line 47 max_connections=20 to "database error" loses line number and parameter — next repair loses grounding. Higher compression ⇒ stricter traceability: original must be retrievable on demand.
P3 Progressive Discovery (Perception × Loop)
Problem: Unknown information space. Metaphor: Animal foraging — wide scan → zoom in → deep dive. Each round's findings drive the next. Engineering must-haves per round: keywords, candidates, selection rationale, stop condition, token budget. Hard rule: Three rounds no hit → redefine problem/keywords, don't keep burning tokens. Prerequisite: A "map" (AST, code graph) for directed exploration, not random walk. Post-workshop added "local reverse discovery & architectural constraint retrieval in brownfield codebases." Common failure: Dumping entire codebase/docs into context — hurts visibility (lost-in-the-middle).
P4 Multi-Modal Fusion (Perception × Parallel)
Problem: PDFs, tables, charts, screenshots, logs, structured data mixed. Approach: Separate parse paths → unified evidence representation. Preserve structure per modality: table row/col relations, chart units/axes, PDF page/section, screenshot verifiable references. Before reasoning, align four fields: entity, time, scope, source. Case: Unreal Engine Blueprint — team chose direct image reading over script conversion (higher efficiency, fewer tokens). Fatal failure: Blind OCR everything to plain text — flattens charts/tables/layout, reasoning looks smooth but evidence is corrupted. Key clarification: Multi-source (where data comes from) ≠ Multi-modal (how data is represented). Sources govern permissions, timeliness, cross-verification; modalities govern parsers and representation.
Three Anti-Patterns + One Security Reef
Anti-pattern 1: Perception Omnipotence. "More sources = smarter agent" — actually drowns signal, raises latency, degrades quality. Fix: Ford-assembly-line decomposition — separate agents for attack point, entry point, path; then a fusion agent. Principle: Define decision goal first, reverse-design perception scope, keep subtracting.
Anti-pattern 2: Perception-Decision Coupling. Hard-coding business logic in perception layer kills reusability; rule changes cascade. Fix: Strict layering — perception outputs structured facts + labels; decision layer owns judgment. Essence: classify & layer.
Anti-pattern 3: Static Perception. "Design once, run forever." Business & architecture drift → blind spots grow, accuracy decays. In probabilistic world, no set-and-forget: continuous iteration on bad cases (misses, false positives, latency) across sources, rules, algorithms.
Security Reef: Every integrated tool is a master key. External tools, docs, pages, images, logs may carry malicious instructions (prompt injection, image poisoning). Perception must record source, scope, data-vs-instruction boundary at ingress. ADPS elevates security (cross-cutting X3, extension G5) to a first-class unit.
When Perception Engineering Peaks, Software Engineering Rewrites Itself
The seminar naturally expanded to AI-driven Software Engineering because coding agents' primary objects — code, requirements, architecture, tests — are perception targets.
Coding is solved; pre/post-coding is the frontier. Huang Jia outlines four coding paradigms' "judges": Vibe Coding (no judge), Prompt-driven (existing CI/CD), TDD (pre-defined tests), SDD (specs as judge). Core shift: "What humans judged, judges now judge." But judges are compressed info → lossy → must monitor judge completeness/effectiveness.
Weibo's Li Qingfeng: Perception = Goal + Constraints. Goal = what to do (with background, no overload); Constraints = how to know it's done (model self-verifies via automated tests). TDD resurges because model must judge against auto tests. Real-world: new projects run smoothly; hundred-million-user legacy systems — only modular, edge-to-core rollout. Front-end UI verification remains weak spot.
OpenLogos (Huang Xianglong): Give the "foreman" blueprints. Vibe Coding = oral instructions for a 100-story tower → information loss → black box → endless rework. OpenLogos supplies: Why-What-How docs, vertical scenario slices (each with req/design/sequence/test), TDD-first, docs-as-context (full linked doc set fed to AI). RunLogos productizes: proposal → plan → docs → slices → code → verify → multi-env deploy → smoke → archive, with two loops (test-fail→fix, smoke-fail→fix) and convergent AI peer review (Claude writes, Codex reviews; 5 issues/round, must cite line & fix, max 5 rounds). Six-dim slice scoring: impact scope, behavior complexity, spec count, new test volume, risk, uncertainty. Low score → AI; high score → vertical slice (each independently verifiable).
DeerFlow maintainer Jiang Ning: Brownfield is "mud-piling." Greenfield toys work; production systems need guardrails. Global contributors use diverse agents, only AGENTS.md constrains architecture → PR quality varies. Typical rot: over-design (middleware bloat hard to catch in review). "Mud-piling": incremental patches without global control or periodic cleanup/reflect. DeerFlow 2.0 code bloat visible in 1.5 months. Answer to Huang Jia's question — "When AI code volume exceeds human review capacity, what assets remain?" → architecture readability, ADRs, tests, task logs, periodic architecture guardianship: make boundaries discoverable, checkable, executable by agents.
Game team Huang Cheng: 3 devs → 6-7x throughput. Named agents (Builder-1, Handheld-2…). Context organization = perception design masterclass: ADR (long-term ambiguous choices), TDD (per-change, no semantic compression, keeps window focused), Task/daily logs (short-term memory/scratchpad), bidirectional flows (forward: GDD→dev; reverse: hidden engineering like save/map-switch written back to task/TDD/ADR). Validates ADPS tenet: perception and memory are interlocked — layered retention (Memory M1) determines perception focus scale.
Big-tech Team Lead "Monk": Solve brownfield first, then full automation. 70-80 person backend team. Three concrete observations: (a) Core contradiction = brownfield = context problem; humans still in loop for req/design/split/integration test. (b) AI outputs too fast for human review → codify "good human habits" into workflow: Claude Code scripts for CR, differentiated workflows, cross-model review (Claude writes, Codex reviews). (c) Measure AI impact via observability: track human-AI round-trips in "tech design" and "coding" phases — 40+ demands avg 10+ rounds in design; OKR targets reducing rounds, raising first-pass accuracy. This yields a management lens : when AI coding is norm, "human rounds per demand, tokens burned, bottleneck phase" become new R&D efficiency metrics.
Context Contract: From Trick to Engineering
All practices converge on the whitepaper artifact — Context Contract . Every production agent must explicitly declare:
What enters current context? What gets a handle? What is deferred? What is discarded?
Plus goal, must-read materials, timeliness, source, conflict rules, missing items — all recorded.
This elevates P1–P4 from techniques to engineering: not ad-hoc improvisation, but a reviewable, traceable, retrospectable contract.
Self-Check: Eight Questions Your Perception Layer Must Answer
Where are the raw goal, non-goals, hard constraints stored? Will they survive compression?
Context/handle/deferred/discard — are triage rules explicit or model-improvised?
Are all handles guaranteed resolvable?
During compression, are numbers, paths, error codes, rejected alternatives, architectural decisions, evidence citations preserved?
Facing unknown codebase/docs: build map first, then progressive explore? At three-round miss, pivot or burn tokens?
Multi-modal: table row/col, chart axes, PDF page/section kept? Any blind OCR-to-text?
Multi-source vs multi-modal modeled separately? Source permissions/timeliness/cross-check vs modality parsers/lineage clear?
Every external tool/data source: at ingress, is data-vs-instruction boundary drawn?
If you can't answer these, your next investment isn't a bigger model — it's fixing perception. Because that 80% decision-failure root cause lives here.
Closing consensus: Model sets capability ceiling; perception determines how much reaches business. Goal-backward perception design, layered signal processing, strict "advisor not general" boundary, vigilance against three anti-patterns and input security — do these solidly and current models jump a tier. Extending perception engineering to software engineering reveals the same logic at scale: Making AI write code is easy; making it see complete requirements, constraints, architecture, and acceptance is hard. OpenLogos blueprints, Weibo's goal+constraints, DeerFlow's architecture guardianship, game team's ADR/TDD layering — all are building high-quality perception systems for coding agents.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Smart Era Software Development
Committed to openness and connectivity, we build frontline engineering capabilities in software, requirements, and platform engineering. By integrating digitalization, cloud computing, blockchain, new media and other hot tech topics, we create an efficient, cutting‑edge tech exchange platform and a diversified engineering ecosystem. Provides frontline news, summit updates, and practical sharing.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
