Which Rules Stay When AI Agents Get Smarter? A Framework for Essential Constraints
The article proposes a framework for managing AI agent development rules that retains essential constraints like goals, boundaries, and acceptance criteria while loading project facts and technical guidance per task, and validates rule removal through fixed-task evaluation after model upgrades.
As AI agents become more capable, the author reevaluates which development rules must remain always active versus which can be loaded on demand. The core principle: keep goals, boundaries, acceptance criteria, and high-risk constraints permanently; load project facts and technical guidance only for the current task.
Classify Rules by Scope and Permanence
Rules are organized into three layers:
Always-active rules — a minimal set applied to every task: confirm goal, scope, and acceptance criteria; read-only analysis skips change process; validation intensity matches risk; unexecuted checks cannot be reported as passed.
High-risk change rules — for public interfaces, data structures, permissions, or cross-module contracts: require explicit impact analysis, verification, and rollback plans.
Task-scoped guidance — project facts (directory layout, commands, API conventions) and technical guides loaded only when the task demands them.
An adapter layer points different tools (Cursor, Claude Code, etc.) to the same rule sources, avoiding duplicate maintenance.
Eliminate Redundancy, Not Accountability
The author reduced rule volume by removing duplication, not by dropping acceptance requirements. Examples:
Task specifications already state acceptance criteria; execution checklists reference them instead of copying.
Common test requirements are centralized; technical guides add only technology-specific additions.
Read-only analysis bypasses the full development flow; small local changes and cross-module migrations receive proportionate design and verification depth.
Character count and rule-check logs decreased, but actual token-cost savings and rework reduction still need measurement on real tasks.
Design Inputs: Distinguish Decision Strength
A classification of design content by constraint level:
Confirmed decisions — execute per convention; changes require confirmation process.
Hard constraints — implementation may vary, but boundaries cannot be silently crossed.
Candidate solutions — for comparison only; must not be mistaken as the sole answer.
Open questions — investigate and verify within allowed range; escalate for confirmation if needed.
For a backend task, the API contract may be fixed while internal function organization and component reuse remain open to code evidence. Prematurely fixing such details can block the agent from choosing a better implementation based on the existing codebase. Design-first remains valuable for providing rationale, but a tentative idea in a design doc should not automatically become an immutable requirement. Business-meaning, public-contract, and authorization changes still require responsible-person approval; stronger models do not expand task boundaries.
Validate Before Deleting Rules After Tool Upgrades
When tools gain native capabilities (e.g., Claude Code's CLAUDE.md, Skills, Hooks), the author follows Boris Cherny's advice: periodically remove rules and verify the model still completes tasks. The evaluation protocol:
Fix target project, task inputs, and acceptance criteria.
Run in isolated environments; save actual changes and verification results.
Compare candidate and original rules under identical conditions — no simultaneous model and task changes.
Check in order: boundary violations, requirement conformance, verification authenticity, then time and cost. A faster run that misses a critical check is not a success.
Remove redundant guides and temporary fixes one group at a time, re-running related tasks.
Permissions, business constraints, and acceptance requirements stay even if the model improves; if native capability replaces a rule, confirm the constraint still holds.
Recurring failures are fed back into the appropriate layer (task spec, project facts, technical guide, or verification entry), fixed there, and reproducible cases added to regression evaluation — avoiding endless context-stuffing of reminders.
From Personal System to Team-Shared Practice
The accumulated rules and templates are now being tested for team adoption:
Onboard a project by injecting real paths and commands, keeping the project's existing requirements/design system, dropping unused technical guides, and validating with varied task scopes.
Have other members complete a bounded real change; observe whether they find references, spot design gaps, leave credible verification evidence, and where the author still must intervene. Success means completion without continuous author explanation.
Feedback drives further adjustments: hard-to-find entry points get reorganized, unhelpful template sections trimmed, inapplicable scenarios documented. When members can propose and validate changes, shared maintenance begins.
Summary
The author continues to enforce clear goals, boundaries, and acceptance, but stops defaulting to ever-finer implementation prescriptions. What can be relaxed is decided by actual task outcomes and verification results. The shared workflow aims for adjustable rules, non-negotiable accountability, and methods that others can sustain and improve.
Reference
Boris Cherny: We Cut 80% of Claude Code's Prompt — Transcript & Summary (https://sozai.app/transcript/boris-cherny-cut-80-percent-claude-code-prompt/)
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Bricklaying Diary
Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
