An agent skill that pentests the app you're building. /hack-me attacks your own
running app the way an attacker would, then closes what it finds:
find → prove → patch → re-verify
Nothing is called a bug until a real HTTP request proves it, no fix is done until that exact exploit stops working, and no class is closed until the same pattern is swept from the paths no probe touched.
Quickstart
Claude Code — one install, gets you the skill and the pentest loop:
/plugin marketplace add kulchankas/paranoid
/plugin install paranoid@paranoid
Then start your app and point it at the loop — plugin skills are namespaced, so
it's /paranoid:hack-me here:
python3 my_app.py # your app, running locally
/paranoid:hack-me # → http://localhost:<port>
Codex, Cursor, or any agent that reads skills:
npx skills add kulchankas/paranoid/skills/paranoid
Copy commands/hack-me.md into your agent's commands
directory (e.g. .claude/commands/) and it's plain /hack-me.
No dependencies, no network calls, no telemetry — it's Markdown your agent reads.
Receipts
On an app we didn't write
The hard claim is code we don't control. Pointed at OWASP VAmPI
— a well-known third-party vulnerable API — with nothing but its URL, /hack-me
found, proved, patched and re-verified six real bugs:
| # | Finding | OWASP API | Status |
|---|---|---|---|
| 1 | Unauth /users/v1/_debug dumps every password |
API3/5 | 401/403 |
| 2 | Read any user's private book secret (BOLA) | API1 | 404 |
| 3 | Register with admin:true → privilege escalation |
API6 | admin=false |
| 4 | Change any user's password (account takeover) | API1 | victim untouched |
| 5 | SQLi in user lookup (UNION-dumps passwords) | API8 | 404 |
| 6 | Debugger + stack traces exposed | API7 | clean errors |
Notably, VAmPI's own global "secure mode" flag closed only four of the six — the
critical password dump stayed open until patched. /hack-me caught it by
replaying every exploit instead of trusting the flag. Full receipts:
examples/vampi/HACKME_REPORT.md.
On a demo app, start to finish
Against a small invoicing API (examples/ledgerlite), an
agent told nothing about the app's bugs found four by probing:
| # | Found by probing the API | Class | After patch |
|---|---|---|---|
| 1 | Any user reads any invoice | IDOR / broken object auth | 404 |
| 2 | /admin/users open to anyone logged in |
broken function auth | 403 |
| 3 | /search?email= SQL injection (dumped passwords) |
SQLi | [] |
| 4 | /profile accepts is_admin |
mass assignment → privesc | 400 |
Every legitimate request still returns 200 afterwards. Reproduce it yourself:
python3 examples/ledgerlite/app.py, then run /hack-me. Walkthrough with exact
requests and diffs: examples/ledgerlite/HACKME_REPORT.md.
On a shop with no exploit menu
DVWA lists its bug categories in the navigation, so that run
tests the fix loop, not discovery. VulnShop
renders as a store (NexCart): products, search, account. Pointed at that UI,
/hack-me proved five bugs and re-verified each — including a product-id
UNION that printed every seed password, and a login of admin' OR '1'='1
that came back as an admin session:
| # | Finding | Status |
|---|---|---|
| 1 | /products?id= dumps every password |
empty listing |
| 2 | Login bypass to an admin session | no session |
| 3 | Search reflects raw HTML | escaped |
| 4 | Any logged-in user opens /admin |
403 |
| 5 | Profile save with no CSRF token | 400 |
The upstream README names those classes, so this is not a fully blind-to-the-repo
find. Receipts: examples/vulnshop/HACKME_REPORT.md.
What /hack-me actually does
- Maps your running app and picks the risk classes it's exposed to.
- Probes each — one crafted request that only succeeds if the bug is real.
- Proves every finding with the actual request/response (no theorizing).
- Patches the root cause with a minimal, behavior-preserving fix.
- Re-verifies by replaying the exact exploit — a finding isn't closed until it fails.
- Sweeps for the same bug on paths no probe touched — a sibling branch, the
same resource under a different method, a
/v1copy, a cron job that reaches the same sink. A green re-verify proves the request is dead, not the class.
Step 6 exists because the loop got caught by exactly that: in our own DVWA run an SQL injection was proven, patched and re-verified green while an identical injection sat in the same file, in the branch for the other database backend. Review caught it; the loop hadn't. Anything the sweep fixes but can't reach with a request is reported as "same pattern, fixed, not separately proven" — never as verified.
It knows where routes and auth live in twelve stacks (Next.js, FastAPI, Express,
Django, Rails, Flask, Spring Boot, Laravel, Phoenix, Go, NestJS, ASP.NET Core) — see
references/frameworks.md. Guardrails
apply throughout; see Scope & ethics.
Why this isn't another "write secure code" skill
It started as the obvious thing — guidance telling the agent to write secure code
— and then got benchmarked honestly before anyone believed it. The harness
(benchmark/) generates the same tasks with and without the skill
and runs real exploits against whatever the model writes.
The result was a clean negative:
| Model | Tasks | Exploit rate without skill | with skill | Effect |
|---|---|---|---|---|
| Opus | all 22 classes, blinded | 0% | 0% | none |
| Opus | 3 isolated functions | 0% | 0% | none |
| Fable 5.1 | 3 isolated functions | 0% | 0% | none |
On an isolated function a capable model already writes the secure version
unprompted — ownership in the WHERE clause, parameterized queries, field
allow-lists. Advice adds nothing there. (The harness isn't rigged: it flags
deliberately-insecure reference code at 100% and secure code at 0%, and CI
asserts that on every push.)
Blinded matters here: the task ids name their own vulnerability, so asking a
model for sqli_login.py is itself a security hint. The 22-class run was redone
with the specs renamed task_01…task_22, generated outside the repo, with no
mention of a benchmark — and the result held. The skill did, however, cost a
functional test the baseline passed. Full method, caveats and raw
solutions.
Real vulnerabilities don't live in one tidy function. They live in the wiring
of a whole running app: auth on one route but not the next, a request body that
quietly sets is_admin, a search box that concatenates SQL. So paranoid stops
advising and starts attacking the running app.
Also inside: the paranoid skill
The guidance the benchmark tested still earns its place as a companion while
you code and as hack-me's knowledge base:
- the vibe-coded top 10
- references: auth & access · secrets & the client boundary · injection & SSRF · APIs & webhooks · framework guides
- a 7-point pre-commit gate
Load it while building; run /hack-me to check whether it held.
The benchmark
A reproducible harness for "does a security skill actually reduce
vulnerabilities?" — 24 vulnerability classes, each with a neutral spec, a
functional check and a real exploit check. 22 of the 24 have been scored against a
model in both conditions, blinded (the other two — route-wiring IDOR and
session-cookie-flags — are harness-verified and awaiting a model run, and
route-wiring is the one most likely to move the result: #23);
the raw solutions are committed so anyone can re-score them. CI asserts on every
push that the deliberately-insecure references
still score 100% and the secure ones 0%, so the benchmark can't silently rot.
Details, caveats and how to re-run it: benchmark/.
Scope & ethics
paranoid secures your code and pentests your running app, with your
say-so: authorized targets, localhost only, non-destructive proofs. It is not
built to target third-party systems, scan hosts you don't own, evade detection,
or produce live malware, and it will decline to. See SECURITY.md.
Roadmap
-
/hack-meloop — find → prove → patch → re-verify, on localhost - Reproducible skill-efficacy benchmark + the honest result behind the pivot
- Independent-app proof — OWASP VAmPI: 6 real bugs found, fixed & re-verified
- 24 benchmark task classes (IDOR — object-level, session-scoped, and route-wiring; missing auth; SQLi; mass assignment; path traversal; SSRF; SSRF via DNS-rebinding; XSS; command injection; open redirect; JWT auth; leaked secrets; CSRF; template injection; XXE; unrestricted upload; permissive CORS; weak password storage; ReDoS; unverified webhooks; insecure deserialization; session cookie flags)
-
/hack-meframework guides — 12 stacks - A second independent-app proof — DVWA: 6 real bugs found, fixed & re-verified on a PHP/MariaDB stack
- A third proof on a shop UI with no in-app exploit menu — VulnShop. Its README still names the classes, so a target that publishes no bug list anywhere is still open
- SSRF via DNS-rebinding task class
- Session-cookie-flags task class
paranoid is v0.1 and actively developed — issues and PRs welcome.
Comments