>

AgentArmor

Hardening LLM agents against prompt injection.

A control checklist that moves every security decision outside the model, where an attacker's text cannot argue with it.

$git clone https://github.com/MuhammadMurtuzaHussain/agent-armor ~/.claude/skills/agent-armor
Made at Dublin AI Week

The pair, live

Watch an attempt get blocked.

The same seven attack classes, one per line. On AgentArmor the control stops them. Hover a line to scan the redacted payload.

attack-vs-defense
ATTEMPT -> BLOCKED
Tool-description / MCP poisoning
> tool.description <= hidden parameter instruction
// scanning vector 01 / 07
BLOCKEDschema caps the field to an enum
hover to scan

Seven controls, one per attack class

Any guardrail the model enforces by reading untrusted text is bypassable. Each control moves the decision to a layer an attacker cannot argue with.

vector 01 / 07

Constrain tool metadata at the schema

Treat tool metadata as untrusted data. Constrain at the schema level what each argument can carry, so a free-text field cannot hold PII.

Install

Clone it into your Claude skills directory, then restart the session. The clone ships the full SKILL.md.

bash
$ git clone https://github.com/MuhammadMurtuzaHussain/agent-armor ~/.claude/skills/agent-armor

The pair

Companion skill, offensive

AgentBreaker

The offensive playbook for the same attack classes.

Open AgentBreaker