NVIDIA SkillSpector: An AI Agent Skill Security Scanner
TopicsGitHub TrendingAgent Skills
One-Sentence Summary
An AI Agent Skill security scanner open-sourced by NVIDIA. Scan before installing a Skill to detect 64 vulnerability patterns × 16 risk categories, covering prompt injection, data exfiltration, privilege escalation, supply chain attacks, and more.
Why Is SkillSpector Needed?
In the AI Agent ecosystem, anyone can publish a Skill. These Skills run on implicit trust, with almost no review mechanism.
NVIDIA's research shows:
- 26.1% of Skills contain vulnerabilities
- 5.2% show possible malicious intent
You have 200+ installed Skills, and you have never run a security scan.
Feature Highlights
| Feature | Description |
|---|---|
| Multi-format input | Git repos, URLs, zip files, local directories, single files |
| 64 vulnerability patterns | 16 risk categories, from prompt injection to supply chain attacks |
| Two-stage analysis | Fast static analysis + optional LLM semantic assessment |
| Real-time CVE lookup | Connects to OSV.dev for live vulnerability data |
| Multiple output formats | Terminal / JSON / Markdown / SARIF |
| Risk scoring | A 0-100 score + severity level + concrete recommendations |
16 Risk Categories × 64 Detection Patterns
🔴 High-Risk Categories
| Category | Patterns | Typical Detection |
|---|---|---|
| Prompt Injection | 5 | Instruction override, hidden instructions, data exfiltration commands |
| Data Exfiltration | 4 | External transmission, environment variable harvesting, context leakage |
| Dangerous Code (AST) | 6 | eval() / exec() / subprocess / os.system() |
| Supply Chain | 4 | Unverified dependencies, URL download and execute, pip install injection |
🟡 Medium-Risk Categories
| Category | Patterns | Typical Detection |
|---|---|---|
| Privilege Escalation | 3 | Excessive permission requests, sudo/root execution |
| Excessive Agency | 4 | Autonomous behavior beyond the declared function |
| Output Handling | 3 | Unfiltered output, sensitive information leakage |
| System Prompt Leakage | 3 | System prompt extraction |
| Memory Poisoning | 4 | Memory contamination, long-dormant instructions |
| Tool Misuse | 4 | Tool abuse, unauthorized operations |
🟢 Low-Risk but Important
| Category | Patterns | Typical Detection |
|---|---|---|
| Rogue Agent | 3 | Self-replication, unauthorized spawn |
| Trigger Abuse | 3 | Timed triggers, hidden activation conditions |
| Taint Tracking | 5 | User input taint tracking |
| YARA Signatures | 4 | Known malicious pattern matching |
| MCP Least Privilege | 3 | MCP tool least-privilege checks |
| MCP Tool Poisoning | 3 | MCP tool poisoning detection |
Usage
Pre-Installation Scan (the most important)
# Scan GitHub Repositories (Before Installation)
skillspector scan https://github.com/someone/some-skill
# Scan Local Directory
skillspector scan ./my-skill/
# Static analysis only (fast)
skillspector scan ./my-skill/ --no-llm
LLM Semantic Analysis
export SKILLSPECTOR_PROVIDER=openai
export OPENAI_API_KEY=sk-...
skillspector scan ./my-skill/
Supports OpenAI / Anthropic / NVIDIA NIM / a local Ollama.
CI/CD Integration
The SARIF output format integrates directly with GitHub Code Scanning:
skillspector scan ./my-skill/ --format sarif --output report.sarif
Docker (no Python installation needed)
docker run --rm -v "$PWD:/scan" skillspector scan ./my-skill/ --no-llm
Value to Us
We have 200+ Skills and zero security scanning.
| Action | Priority |
|---|---|
| Run a one-time batch scan of all installed Skills | 🔴 P0 |
| Mandatory scan before installing a new Skill | 🔴 P0 |
| Integrate into CI/CD to scan automatically on every Skill update | 🟡 P1 |
| Establish a Skill security allowlist/blocklist | 🟢 P2 |
Tech stack: Python 3.12+ · LangGraph · SARIF 2.1.0 License: Apache 2.0 | Repo: NVIDIA/SkillSpector
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built