Cloudflare's Security Audit Skill: Six-Phase AI Code Review with Adversarial Validation

Cloudflare open-sources a six-phase security audit skill that transforms coding agents into structured auditors using adversarial validation, machine-readable findings, and a coverage ledger to make AI-driven vulnerability detection reproducible and trustworthy.

Architecture Digest
Architecture Digest
Architecture Digest
Cloudflare's Security Audit Skill: Six-Phase AI Code Review with Adversarial Validation

Cloudflare has released security-audit-skill, an open-source skill that imposes a disciplined six-phase workflow on coding agents to perform security audits. The methodology originates from Cloudflare's internal vulnerability-hunting harness, which evolved from a single slash-command into a fleet scanner covering 128 repositories in six weeks.

Six-Phase Structured Workflow

Reconnaissance : Maps architecture, trust boundaries, and attack surface, producing architecture.md and coverage-ledger.json.

Coverage-Guided Hunting : Dispatches isolated "hunter" agents per coverage unit; a coverage critic identifies gaps.

Candidate Verification : Each candidate vulnerability is handed to a fresh verification agent tasked with disproving it — finder ≠ verifier.

Structured Output : Writes findings.json with three verdicts — confirmed, needs_validation, rejected — validated against report-schema.json by zero-dependency scripts ( validate-findings.cjs, validate-coverage-ledger.cjs).

Independent Verification Logging : A new agent re-checks final conclusions; critical replacements trigger another independent verification round.

Neutral Reporting : Derives REPORT.md, FINDINGS-DETAIL.md, and NEEDS-VALIDATION.md from verified records.

Core Design Principles

Finder Never Verifies

The agent that discovers a vulnerability never validates it. A fresh agent with a falsification mandate eliminates confirmation bias. Combined with strict rules — confirmed findings require full source trace and bounded observation; severity requires demonstrated defense-in-depth impact — this reduces false positives dramatically.

Audit Continuity via Coverage Ledger

coverage-ledger.json

persists "what's been checked" across runs. Multiple runs accumulate coverage; only uncovered gaps are re-scanned, and changed code is re-verified. Cloudflare's blog notes a single run finds only ~50% of total vulnerabilities — auditing is a marathon.

Three-Tier Verdict Semantics

confirmed

: complete source trace + bounded observation. needs_validation: exact unresolved facts documented, no severity allowed. rejected: disproven candidates recorded for audit trail.

Real-World Scale Data

One fleet scan generated 20,799 raw candidates → deduplicated to 5,044 → verification eliminated 2,302 → ~13,841 entered system → 7,245 delivered to teams → classified as 41 critical and 777 high-severity. Every stage enforced mechanical checks and adversarial validation.

Quick Start

npx skills add https://github.com/cloudflare/security-audit-skill \
  --skill security-audit

Add --global for user-level install. Then instruct the coding agent: security audit this codebase. Results default to ~/security-audit-skill/<repo-name>/run-<N>.

Hard Requirements

Agent must support tool calling and parallel sub-agents — pure chat models cannot drive this orchestration.

OS-level sandbox required: no egress, resource limits, allowlisted environment, write-only to designated temp directory. Without a sandbox, runnable leads are downgraded to needs_validation instead of executed — a deliberate safety default.

Enterprise Adoption Path: Skill → Harness

Cloudflare's blog maps each phase to an independent agent; adding a database and orchestrator yields the 128-repo fleet scanner. Teams can replicate:

Standardize artifacts : findings.json + report-schema.json become the team contract; CI gates with validate-findings.cjs.

Use coverage ledger as backlog : The "region × attack-type" matrix gaps become the next iteration's task pool.

Assign hunters by attack domain : Skill includes ten domain prompts (memory safety, AI/LLM injection, web protocols/auth, client-side DOM, supply chain, cloud/deployment, RPC/messaging, resource exhaustion, data isolation, desktop/mobile/IPC). Domain experts own hunting and verification for their area.

Observe, don't prescribe, tooling : Cloudflare integrated Semgrep end-to-end; hunter agents called it zero times in a month — they preferred reading and running code. The most-used tool was the wishlist (agent requests human-supplied capabilities), invoked 25,472 times across 128 repos. Lesson: prepare environment, dependencies, permissions; watch what the agent actually asks for.

When to Use

Security engineers / dev teams needing systematic, verifiable, repeatable vulnerability discovery.

Compliance-driven teams: immutable graded records + coverage ledger naturally satisfy "what was audited, what was missed" evidence.

Practitioners studying Cloudflare's vulnerability-hunting methodology — the skill files and linked blog serve as curriculum.

Pure backend CRUD teams without security mandates: overkill.

Pros & Pitfalls

Pros

Cloudflare-produced, battle-tested across 128 repos — not a toy.

Finder ≠ verifier adversarial validation + machine-readable findings directly address AI audit trustworthiness.

MIT licensed, zero-dependency validators, schema backed by test fixtures.

Pitfalls

Not a scanner — output quality depends entirely on underlying model and tool-calling capability.

Sandbox is mandatory, not optional; without it, executable verification degrades to needs_validation.

Single run ≠ done; official data shows ~50% coverage per run — accumulation is required. needs_validation is not a free pass; it demands documented unresolved facts — don't report pending items as confirmed.

Architect's Takeaway

The value isn't the few hundred lines of validation scripts but the audit discipline: coverage ledger + adversarial validation + three-state grading . It answers how to make AI security conclusions credible — a question most tools avoid. Cloudflare open-sourced their internal discipline as a free skill; the methodology outweighs the code volume, but it's infrastructure for teams with genuine security needs, not a casual add-on for every developer.

GitHub: https://github.com/cloudflare/security-audit-skill

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

open-sourceAI code reviewvulnerability detectionCloudflarecoding agentadversarial validationcoverage ledgersecurity-audit-skill
Architecture Digest
Written by

Architecture Digest

Focusing on Java backend development, covering application architecture from top-tier internet companies (high availability, high performance, high stability), big data, machine learning, Java architecture, and other popular fields.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.