Evaluate and harden LLM-based autonomous agents against adversarial attacks using the α³-SecBench layered security framework. Assesses security (attack detection, CWE attribution), resilience (safe degradation), and trust (policy-compliant tool usage) across 7 autonomy layers. Use when: 'audit my LLM agent for security', 'add adversarial resilience to my autonomous system', 'evaluate agent trust and tool safety', 'harden my AI agent against prompt injection', 'security benchmark my LLM pipeline', 'test my agent for hallucinated tool calls'.
Claude Skills are markdown + scripts loaded into Claude Code; scanned with the same pipeline.
Claude Skills install to a single target — no runtime picker.
nerlo install -- 3-secbench-large-scale-evaluation-suite-security
Manual alternative: copy the skill folder into ~/.claude/skills/ so Claude Code loads it on next launch.
cp -r -- 3-secbench-large-scale-evaluation-suite-security ~/.claude/skills/
Install respects the composite badge: Verified proceeds, Caution prompts for confirmation, and Unsafe is refused. Write operations require an API token.
Not yet scanned
This server is in the registry but has no completed scan yet. Per-scanner scoresheets appear here once the pipeline finishes its first run.
1 scan on record. Every scan's full results are retained immutably for 24 months.
| Completed | Composite | Change | Scanners | Status |
|---|---|---|---|---|
| — | Pending scan | — | — | queued |
Are you the author and believe a verdict is inaccurate? Appeal this badge.