NVIDIA's SkillSpector Scans AI Agent Skills for 71 Vulnerability Types Before Install
NVIDIA's open-source SkillSpector scans AI agent skills for 71 vulnerability types across 17 categories using static AST/YARA/taint analysis and optional LLM semantic comparison, outputting SARIF for CI/CD gates with baseline suppression and fail-closed defaults.
Research analyzing 42,447 agent skills found 26.1% contain at least one vulnerability and 5.2% exhibit highly suspicious malicious behavior patterns. Skills with executable scripts are 2.12 times more likely to be compromised than pure instruction-based skills. A skill runs with the user's identity, credentials, and tool permissions, making prompt injection a direct channel for abuse.
SkillSpector Overview
SkillSpector is an open-source pre-installation security scanner for AI agent skills, released by NVIDIA in March 2026 under Apache-2.0 (Python, 98.5% of codebase). It serves as the security layer for NVIDIA's Verified Skills pipeline; every skill in the official catalog is scanned, assessed, and signed before publication.
Three Core Capabilities
1. 71 Rules Across 17 Risk Categories
The rule taxonomy covers prompt injection, data exfiltration, privilege escalation, supply chain, over-permissioning, output handling, system prompt leakage, memory poisoning, tool abuse, rogue agent, trigger word abuse, dangerous code patterns (AST), taint tracking, YARA signatures, MCP least privilege, and MCP tool poisoning.
Notable rule examples:
P2 Hidden Instructions : Malicious directives concealed in comments or invisible text that human reviewers often miss.
SC8 Bundled Python Bytecode : Skills shipping __pycache__ / .pyc files to bypass source-code scanning and execute pre-compiled bytecode directly.
Wildcard Permissions : Skills declaring *, all, full, or any permissions are flagged immediately.
Missing Permission Declarations : Code exhibiting capabilities without a corresponding permission field — silent privilege requests.
2. Two-Phase Analysis: Deterministic First, Semantic Second
Static Layer (default, offline) : AST traversal + YARA signatures + taint tracking. Fast, deterministic, free; enabled with --no-llm.
Semantic Layer (optional, LLM-backed) : Compares the skill's declared intent against its observed behavior to catch description-behavior mismatches and vague trigger words that signature rules miss. Supports OpenAI, Anthropic, Bedrock, NVIDIA Build, Ollama, Azure, and local vLLM. The LLM analysis prompt includes built-in anti-jailbreak protection to prevent a malicious skill from manipulating the scanner model.
3. Pipeline-Native Design
Four Output Formats : Terminal (human), JSON (machine), Markdown (review packet), SARIF (direct integration with CI/CD and IDE security tools).
Baseline Suppression : skillspector baseline stores accepted findings as a baseline (glob patterns or fingerprints). Re-scans report only new issues. Fingerprint-bound baselines reactivate automatically when source changes or versions upgrade.
Real-Time CVE + Fail-Closed Input Limits : SC4 analyzer queries OSV.dev for CVEs (offline fallback). Input limits: 100 MiB download cap, 10,000 zip entries, 1 MiB per file. Exceeding any limit triggers immediate fail-closed refusal — correct default for a scanner consuming untrusted input.
Why "Just Review the Skill" Is Insufficient
The npm ecosystem already proved that popular packages are prime poisoning targets. Skills spread like browser extensions — a colleague installs one, you follow, no review, no inventory, no owner. Moreover, SKILL.md itself is an executable instruction surface : it looks like documentation but drives the agent. Hidden comments, payloads disguised as example code, and description-behavior mismatches evade manual review.
Architecture: Deterministic-First Analysis Pipeline
Ingestion Layer : Five input types (Git repo, URL, zip, directory, single file) unified into ingest with immediate resource limits; exceedance triggers fail-closed.
Static Analysis Layer : AST + YARA + taint tracking run in parallel — all deterministic, offline, reproducible — forming the inspection foundation.
Semantic Analysis Layer (optional) : LLM receives static-layer conclusions and performs intent-behavior comparison.
Scoring & Disposition Layer : 0-100 risk score + severity labels + remediation advice, emitted as SARIF/JSON/Markdown/terminal.
Official documentation defines a disposition policy: Critical/high block release; hidden instructions removed then re-evaluated; permission-behavior mismatches require permission adjustment or behavior removal; known vulnerable dependencies upgraded or pinned. Goal: declared purpose, declared permissions, and actual code must align.
The scanner distrusts its own inputs: zip bombs, oversized repos, pre-compiled bytecode are halted at ingest — a rare self-awareness in security tooling.
Installation & Usage
Requires Python 3.12+.
# Fastest via uv (CLI only)
uv tool install git+https://github.com/NVIDIA/skillspector.git
# Or from source
git clone https://github.com/NVIDIA/SkillSpector.git
cd SkillSpector
uv venv .venv && source .venv/bin/activate
make installDocker alternative:
docker build -t skillspector .
docker run --rm -v "$PWD:/scan" skillspector scan ./my-skill/ --no-llmBasic scans:
skillspector scan ./my-skill/ # local directory
skillspector scan ./SKILL.md # single file
skillspector scan https://github.com/user/my-skill # Git repo
skillspector scan ./my-skill.zip # zip packageEnable LLM semantic analysis:
export SKILLSPECTOR_PROVIDER=openai
export OPENAI_API_KEY=sk-...
skillspector scan ./my-skill/Batch scan a skill directory (20 workers, multi-language detection):
python -m contrib.batch_scan.batch_scan ./my-skills/ --workers 20 -f json -o report.jsonAdditional forms: MCP server mode ( skillspector mcp with [mcp] extra) and Pi/OpenCode extensions — /skillspector command inside agent sessions makes pre-install scanning part of the workflow.
Team Adoption Strategy
1. Gate Definition
Treat skill scanning as an admission gate, not post-hoc audit: every third-party skill must pass scan before entering the dev environment. Critical/high block; others follow official triage table. Boundary: scan covers pre-install; runtime isolation (sandbox, egress control) is a separate layer.
2. Deployment
Security team runs batch scanner on internal skill library ( --workers 20, JSON reports). Developers install CLI via uv tool. Teams using Claude Code add MCP mode for in-session self-check. Multiple API key pools increase LLM scan throughput.
3. Pipeline Integration
SARIF output feeds CI/CD code-scanning channel alongside SCA (dependency scanning) at the same gate: skill fails gate = merge blocked. Baseline file ( .skillspector-baseline.yaml) committed with repo; re-scans alert only on new findings.
4. Policy Customization
Codify official triage policy: forbid wildcard permissions, forbid bundled __pycache__, require permission-behavior consistency as code-review checklist items. LLM semantic layer defaults to internal gateway (skill content sent to provider endpoint — data egress consideration; local Ollama/vLLM avoids this).
Real-World Applicability
Agent users installing third-party skills : Claude Code, Codex, Gemini CLI heavy users — one command before install.
Enterprise security teams : Bring agent skills into software supply chain management; batch scan + SARIF into CI gate.
Skill developers/publishers : Self-scan before release to avoid joining the 26.1%.
MCP ecosystem maintainers : Dedicated rules for MCP tool poisoning and least privilege.
Not suited for users who never install third-party skills or rely solely on official catalogs — it defends against external code, not first-party bugs.
Pros, Cons & Pitfalls
Core Strengths : NVIDIA-produced and used as the release gate for its own skill catalog (production-validated, not a demo); 71 rules/17 categories is the broadest coverage in class; static layer offline and deterministic, semantic layer optional and controllable; SARIF + baseline suppression designed seriously for pipelines; Apache-2.0.
Limitations :
Niche Audience : Developers not installing third-party skills gain little; vertical tool for agent supply chain.
LLM Semantic Layer Cost : Requires API keys; batch scanning consumes significant tokens. README notes default provider DeepSeek-Chat expected to deprecate; local backends (Ollama/vLLM) still seeking PRs — verify compatibility before adopting.
Scanner Is Not Immunity : Findings require human triage; cannot prevent post-approval poisoning updates or runtime injection — runtime controls (e.g., NVIDIA's OpenShell sandbox) needed separately.
Overlap with Alternatives : Partial overlap with cloudflare/security-audit-skill, which focuses on runtime audit skills; can be combined as needed.
Data Source Clarification : The opening 26.1% / 5.2% statistics come from the January 2026 academic paper "Agent Skills in the Wild" (Liu et al., arXiv), cited in NVIDIA's README — not NVIDIA's own research. Numbers are credible but sourced from third-party study.
Operational Pitfalls :
Don't Treat Risk Score as Sole Verdict : 0-100 score is a prioritization aid; hidden-instruction findings (P2) demand manual review regardless of score.
Keep Baseline File Outside Scanned Directory : Official docs warn that a baseline inside the scanned directory can be maliciously tampered with.
Rule Count Discrepancy : Docs site says 68 rules, README says 71 — versioning difference; trust main repo README; rule count still growing.
Docker Scan Mount Current Directory : Reports written back to mount for persistence; forgetting the mount loses results.
Closing Thought
The skill ecosystem is replaying the browser extension and npm supply-chain story, but this time skills hold your credentials and tool permissions. NVIDIA open-sourcing its catalog's security gate turns pre-install scanning from a hygiene practice into a single command — a step that should have arrived earlier, but welcome now.
Open-source repository: https://github.com/NVIDIA/SkillSpector
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architecture Digest
Focusing on Java backend development, covering application architecture from top-tier internet companies (high availability, high performance, high stability), big data, machine learning, Java architecture, and other popular fields.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
