first-pass
first-pass makes a coding agent check the code around a change, not only the lines it writes, and prove its work before it calls anything done. It's a Claude Code plugin; Cursor, Codex and other agents get the same rules and skills (the hooks are Claude Code only). You set it up once, in the folder that holds your repos.
What it does
- Before code, the agent answers ten questions about the change against the real code:
file:line, a test, or a gap you decide to accept.
- While building, anything the task didn't ask for (a helper, a cap, a retry, an option)
- Before "done", a second agent that didn't write the code reviews it in a fresh
- On a bug, it fixes the whole class: it reproduces the bug, finds the same pattern
- On a pull request,
/first-pass:reviewchecks your own branch before you ask for
- In every session, hooks send back a "done" that has no evidence behind it (edits made
- In replies (optional, you choose at setup): plain, short answers. The first line says
What it helps with
- Bugs next to the change: the other code path that writes the same field, the webhook
- "Fixed" and "verified" that only meant "it compiles and the mocks pass".
- An agent reviewing its own work in the same context that wrote it.
- Many repos opened from one folder, where each repo's own rules and hooks don't load.
- Prompts like "be 100% sure", which change how sure the answer sounds, not what gets
- Code nobody asked for: the extra cache layer, guard or option that becomes one more thing
- Replies too long, or too full of the agent's own names for things, to follow once you've
/plugin marketplace add joetawil7/first-pass
/plugin install first-pass@first-pass
Then, in the folder where you start your sessions: /first-pass:setup-first-pass.
Why I made this
I built a product feature by feature with Claude Code. Every feature request ended with some version of "make sure the code is correct and bug free, cover all gaps, be 100% sure." Then I ran a full audit. It found 318 issues, 25 of them high severity.
When I sorted them, only 45 were mistakes in the lines being written. The rest were in the code around those lines:
| What went wrong | Issues | | ----------------------------------------------------------------- | -----: | | Another code path using the same data was not updated | 54 | | Runs twice, runs at the same time, or stops halfway | 54 | | An outside service fails or is slow, and the error is hidden | 45 | | A plain mistake in the code itself | 45 | | Time, units, rounding | 26 | | Scale: no limit, no index, lists that stop at one page | 25 | | Endings: cancel, delete, expire, reconnect, downgrade | 19 | | UI, help, legal or pricing text that says what the code doesn't do | 18 | | Pipeline: CI red, no tests against a real database | 17 | | Hostile user or uncapped cost | 15 |
(One private codebase, sorted by hand with one cause per issue. Your mix will differ.)
Several of these had been found by earlier audits and fixed. Each fix patched one spot, and nothing carried the lesson into the next session.
"Be 100% sure" changed how sure the answers sounded. It never made the agent open the other file that writes the same field, or ask what happens when the webhook arrives twice. So first-pass names those checks, and asks for proof before anything is called done:
- Ten questions before code (the pre-mortem): twice, halfway, outside call, failure is
file:line, a test, or "Not handled, because ___" for you to accept.
- A reviewer that didn't write the code. Models are poor at catching their own mistakes
breaker agent reviews each change in a fresh context, starting from the other
code paths that touch the same data.
- "Done" means a test that failed before the change, the breaker's review, and CI's own
- Bugs get fixed as a class: reproduce, find the same pattern elsewhere, and add the
What's in it
| Piece | What it does | When |
| --- | --- | --- |
| premortem | The ten questions, answered against the code | Before code |
| breaker (agent) | Fresh-context review of the diff and of every other path touching the same data; concrete findings only | Before done |
| ship-check | The definition of done, ending in a report where every "Verified" line says what was run and its result | Before done |
| fix-the-class | Reproduce, name the class, search for it everywhere, run the ten questions on the fix, fix or record each hit, make it hard to repeat | On any bug |
| setup-first-pass | Writes the rules once, a map of your repos, and a section per repo with its real commands, test limits and a drafted INVARIANTS.md | Once, then to update |
| habit-words | Reads what you typed in your recent sessions and maps words like "be 100% sure" to the checks they should mean | At setup, then when due |
| sharpen | Rewrites the prompt you type after it: numbered asks, habit words turned into checks, names and numbers kept exactly. Shows you the rewrite, then works from it | Only when you type it |
| review | Reviews your own branch before you ask for review (type nothing after it), or a teammate's PRs, several repos at once. Checks the change against its ticket, judges the failed checks and every Bugbot comment, runs the breaker, traces the other code that uses what changed, proves each finding or marks it unproven, and says what the merge needs and how to check the deploy. Reads GitHub and never posts; pushes a fix only on your yes | Only when you type it |
| jev | Sets up the optional Jev judge (TypeSafe's decision model), which ship-check asks whether a review finding is real harm, which small ones to fix now, and what proof a small fix needs | Only when you type it |
| Hooks | Run each repo's own hooks from the main folder, send back a "done" with no evidence, say what drifted at session start, and hold sharpen's work until its rewrite is shown | Every session |
If you keep all your repos in one folder
A lot of us open one folder with every repo in it and start each session there. Claude Code
then finds agents, skills and hooks only in that folder and above it: a repo's own
.claude/agents, its .claude/settings.json hooks and its .cursor/rules/*.mdc imports
never load, and its CLAUDE.md loads only once a file in it is read. first-pass is built for
that:
- One set of rules at the root, in
AGENTS.md(imported byCLAUDE.md), loaded from
- A section per repo, written from what setup finds in it: CI's exact commands, how to
- Each repo's own hooks still run. Setup lists them and you approve each one. An
- Drift is reported at session start: a repo's CI changed since its section was written,
- One repo on its own works too: setup writes the rules and the section into that repo.
Your habit words
Most of us have words we type out of habit: "be 100% sure", "don't assume", "full review", "all fine, right?". They name no place to look, so the answer sounds more certain without anything more being checked.
habit-words reads what you typed in your last 20 Claude Code sessions, shows how often you
use each phrase, what went wrong after it, and what to say instead. Then it writes a short
block that maps each phrase to the checks it should trigger. You never have to type them
again, and if you do, they mean the checks, not a more confident tone.
What it reads and keeps:
- Only your own transcripts in
~/.claude/projects, and only what you typed, plus the end
- Anything that looks like a key, token, password, email or phone number is replaced before
- The block holds only your phrases and the checks, and never goes into a file your
What it writes to your machine
AGENTS.mdandCLAUDE.mdat the root (the rules and, if you want them, the reply style
CLAUDE.md or
AGENTS.md, all between first-pass markers. Outside the markers it adds only import
lines, and it lists each one it adds.
INVARIANTS.mdin each repo (a draft for you to review),.first-pass/workspace.jsonat
.claude/cursor-rules/*.md copies where a repo imports .mdc files.
- A small state folder in
~/.claude/plugins/data/for the hooks, and a temp file while
habit-words runs, deleted when it's done.
~/.claude/first-pass/jev.json, only if you set up the Jev judge: which variable holds
Setup never commits or pushes, and sends nothing anywhere except one test request per key
to TypeSafe if you set up the Jev judge: you review the files and commit them. The plugin's own
scripts make no network calls and read no keys or tokens from your environment, with one
exception you have to switch on: the Jev judge reads the one key you named and sends review
findings (with secrets, emails and phone numbers blanked) to TypeSafe's API, from
ship-check and jev. Besides that, the only skill that reaches a server is review: it
reads the PRs, checks and comments through your own gh login, and it pushes a fix only
when you picked that fix and said yes to the push.
Install
Claude Code
/plugin marketplace add joetawil7/first-pass
/plugin install first-pass@first-pass
From a terminal: claude plugin marketplace add joetawil7/first-pass, then
claude plugin install first-pass@first-pass. To update later:
claude plugin marketplace update first-pass, then claude plugin update first-pass@first-pass.
The hooks need Node.js 18 or later on your PATH. The tests run on Windows and Linux.
Cursor, Codex and other agents
npx skills add joetawil7/first-pass -a cursor --copy # this project
npx skills add joetawil7/first-pass -a cursor --copy -g # every project
This uses the skills CLI. Swap -a cursor for your
agent (-a codex, -a windsurf, ...). --copy writes real files instead of symlinks,
which Windows checkouts turn into plain text. Don't add -a claude-code if you use the
plugin, or you'll get every skill twice. Other tools get the rules and skills; the hooks are
Claude Code's.
Then run setup once
In the folder where your sessions start: /first-pass:setup-first-pass (or
/setup-first-pass outside Claude Code). It looks through your repos, asks only what the
code can't tell it (which repos have a UI, which depend on which, which repo hooks to run),
and writes the files above. Re-run it after an update: your own text, each repo's section and
INVARIANTS.md are kept.
Teammates
Once a repo's files are committed, every teammate's agent loads its section. Tell setup
which repos people open on their own: those also get the rules and a copy of the breaker.
Each person installs the skills: the plugin for Claude Code, the npx skills line for
Cursor.
Day to day
- Plan. Ask for the change. The agent runs the pre-mortem before it edits. Read the
- Build, with the tests the pre-mortem named, each one failing on the old code first.
- Done.
ship-checkruns the breaker and CI's checks in a clean checkout, then reports
Verified: → and what wasn't verified.
- Bug.
fix-the-classfixes the one you found, the others like it, and adds the check
- Review. On your own branch,
/first-pass:reviewwith nothing after it reviews your
Instead of "make sure it's bug free", try: *"Run the pre-mortem, show me the tests that fail without the change, and list what you didn't verify."*
Or write the prompt the way you would anyway and put /first-pass:sharpen in front of it
(/sharpen in Cursor). It splits a message that mixes three asks into a numbered list,
turns the habit words into the checks they stand for, shows you the rewrite, and works from
that. It stops to ask only when it would have to guess, and it never adds work you didn't
ask for. In Claude Code a hook holds back edits, shell commands, subagents, MCP tools,
publishing, scheduling and other skills until the rewrite is on screen (reading files is
never held back): with the
instruction alone, the rewrite was skipped in 7 of my 11 test runs.
If you have a TypeSafe API key, /first-pass:jev turns on a
second opinion for ship-check: Jev, a model that only picks from a list and says how sure
it is, judges whether each review finding is real harm, which small ones are worth fixing
now, and what proof a small fix needs (the checks alone, a unit test, a real-database test
or a browser test). It is cheap and fast, but when I had it sort 269 audit issues it gave
more weight to how many people meet an issue than to how bad the harm is. So it can make a
finding serious, never clear one the reviewer called serious, and when it is unsure or
fails, the rules decide as if it weren't there.
How the rules were tested
I test a rule the way I'd test code: fresh Claude Code sessions and agents (Opus, high effort) that load nothing else, hidden checks they never see, and judges who read the output blind.
- Less code nobody asked for. 52 runs on two tasks (add caching to a product page, add
- Plain replies. 10 real replies were rewritten under the old and the new reply rules.
- The reviewer still reports everything. In 3 runs on a small change with 8 planted
These are small samples on small tasks. They show the rules do what they say, not how much they help on your code.
What it won't do
- It's slower per change. A test that fails first, a second agent's review and CI in a
fix-the-class and then running ship-check used about $29 at API list prices
(three reviewer passes). The trade is fewer rounds after "done". The pre-mortem scales
with the change: a copy tweak answers it in one line. To keep the cost down, a prompt with
several items builds them all first, then reviews them and runs CI's full checks once at
the end. During the work, only serious findings (money, data, something done twice or sent
wrong, security, legal, a crash, stuck work, the change not doing its job) are fixed and
checked again. The smaller ones
come to you once, as one list at the end, and the ones you pick get one review together.
Before this, a reviewer finding a small case, a question to you, a fix and another review
of that fix could chain for hours: in one two-day stretch, about a third of 84 reviews
found only small points, and most of them still started another round.
- It won't make code bug free. The aim is fewer and smaller escapes: no high-severity
- Rules alone fade. The fixes that last are the ones
fix-the-classpushes toward: a
- Shell edits are seen late, and not all of them. A file a shell command changes
sed -i, a heredoc, a copy, a formatter) is found at the end of the turn, from
git status and the file's time, so the done check and end-of-turn hooks see it; hooks
that run after each edit don't. Not seen: files git ignores; on Windows, a copy
(Copy-Item, copy) over a file git showed no change in, which keeps the source's
times; writes from a command still running in the background after it returned; from a
main folder, a repo the command neither runs in, cds into nor names a path in; and, in
a single repo, any other repo. Counted anyway: a file rewritten with the same content
(git stash, then git stash pop); another program's write or delete while a shell
command runs; and after you refuse a command, writes until your next prompt. A delete is
dated by its folder, so a later change in the same folder can count or hide it.
Related
- superpowers: a full development method for coding
- anthropics/skills: Anthropic's examples of skills.
- ponytail: makes coding agents write less
- i-have-adhd: replies shaped so you can act on
- Huang et al., Large Language Models Cannot Self-Correct Reasoning Yet (ICLR 2024)
- Tambon et al., Bugs in Large Language Models Generated Code
- Anthropic, Best practices for Claude Code
Feedback
If you try it and it gets in your way, open an issue and tell me where. Reports of a bug that got past all of this are the most useful kind.