In my previous post comparing subagent chaining against agent teams, I built a five-agent roster, ran the same application through it three times under three different coordination models, and put a note in the middle of the setup section indicating that those five agent files were deliberately configured with application-specific instructions, which made them unusable for any other application.
The note went on to say that if we were adapting the pattern for a different application rather than running a one-off comparison, we should keep the domain specifics out of the agent file and pass them in the prompt instead. This follow-up post shows how to change that and configure a generic Claude Code agent roster that can build any application.
I initially thought generalizing the roster would mostly involve deleting the lines that mentioned Waypoint but then I realized that some of those lines were doing more work than expected and removing them made the agents worse rather than more portable. What I ended up doing was sorting the contents of the five files into three categories, then adding several mechanisms the original roster never needed because it only had to build one application under controlled conditions.
What Was Actually Wrong With the Waypoint Roster
The roster built a personal travel map portal called Waypoint, and every one of the five agent files knew it. Inside backend.md, a pin was defined as a latitude, a longitude, a place name, a visit date, and an optional note. The frontend agent was told the map had to be a bundled SVG rather than Leaflet or Mapbox. Over in tester.md sat a list of specific behaviours to cover, and two of those named the administrator endpoint directly.
All these application-specific details were intentional because the point of the exercise was to run three rounds against a byte-identical brief so the only variable was the coordination model, and baking the spec into the persona files is what guaranteed it.
My first instinct was to strip out anything that mentioned Waypoint and leave the role behind. That made the files look generic, but it also removed constraints that were preventing the agents from making unsuitable choices, and the stack block in backend.md is the clearest example.
Your stack choice is constrained, because this has to run on a Windows 11
laptop with nothing exotic installed:
- Node with Express, or Python with FastAPI, and nothing else
- SQLite for storage, created on first run
- No external services, no cloud dependencies, and no API keys
None of that is about travel maps. It describes the environment the application has to run in, and it was the only thing stopping the backend agent from reaching for a container and a Postgres instance. Removing it did not make the agent more reusable. It left the agent without the constraints it needed to make a sensible choice.
Role, Application, and Process
I had been separating the files into role facts and application facts. What I was missing was process, a third category that sits between them.
| Kind of fact | Example | Home | Changes per project |
|---|---|---|---|
| Role | “The tester may only write under tests/“ |
.claude/agents/*.md |
Never |
| Application | “A check has a URL, an interval, and an owner” | SPEC.md |
Always |
| Process | “Route test failures back to the owning layer” | CLAUDE.md |
Rarely |
Role facts describe what a job is and what it must never touch, so they hold regardless of what is being built. The application tier is the one that turns over completely every time, because it describes the thing under construction. Process facts sit in between, covering how the team works together, which changes occasionally as you refine the pipeline but not when you switch projects.
I got this part wrong initially. The stack constraint block belongs in SPEC.md. It reads like a role constraint, but it changes with the project. A demo running on my laptop has different limits from something destined for Azure Container Apps. Moving those constraints to SPEC.md lets the same unchanged backend.md handle both.
Why I Put the Specification on Disk
The note in my last post told readers to pass the domain specifics in the prompt. After reading the documentation more closely and comparing it with what happened in my third round, I would change that advice. The prompt is still useful for the assignment, but it is not where I would put the full specification.
Here is what actually reaches an agent. A subagent receives its system prompt plus basic environment details such as the working directory, and not the parent session’s conversation, which means everything it knows about the task arrives through the invocation prompt the lead composes for it. Teammates in an agent team are stricter still, because the lead’s conversation history does not carry over to them at all. So when you hand the lead a spec and ask it to brief five agents, each agent gets the lead’s summary of your spec rather than your spec, and any detail the lead judged unimportant is gone.
That is more or less what happened in round three. I left the stack placeholder in my own prompt unfilled, and because that same prompt forbade the lead from reading the earlier rounds, it could not resolve the bracket the way round two’s lead had by reading round-1/API.md. It stopped and asked me directly instead, which was the right call. I answered “Node + Express + SQLite,” and that answer was accurate and less specific than what round one had actually built. Round three’s backend went on to ship JWT authentication with a pinned better-sqlite3 dependency where rounds one and two both used opaque cookie sessions and Node’s built-in node:sqlite.
The relay lost the detail, and I would put that down to my input rather than the mechanism.
Putting the specification on disk avoids most of that problem because every agent reads the same file instead of a summary written by the lead. Custom subagents also load CLAUDE.md, and teammates read it from their working directory. The built-in Explore and Plan subagents are the exceptions. With the application and process details on disk, the invocation prompt shrinks to a sentence.
Each of the five agent files now opens with a variation on the same instruction.
Read `SPEC.md` at the project root before doing anything. It is the source
of truth for what this application does, who its users are, and what
constraints the stack has to satisfy. If `SPEC.md` does not exist, stop and
say so rather than inventing requirements.
The refusal clause is important because if the agent cannot find the spec, I want it to stop instead of filling the gap with requirements nobody asked for.
What Changed in Each Agent File
| File | Removed | Added |
|---|---|---|
backend.md |
Waypoint’s data model, the stack allowlist | Reads SPEC.md, publishes API.md to a fixed path, follows existing conventions when the repo is not empty |
frontend.md |
The bundled-SVG map requirement | Reads SPEC.md and API.md by name, discloses guesses to FRONTEND-ASSUMPTIONS.md, asks the backend directly when it can |
tester.md |
The named list of Waypoint behaviours | A generalized category checklist, a PreToolUse hook enforcing the tests/ boundary, an instruction to break down what a pass count actually proves |
security-reviewer.md |
Nothing application-specific to remove | Timing side-channels, algorithm pinning, listen-binding, and revocation-versus-forgetting added to the checklist, plus formatting guidance so its report can be persisted verbatim |
docs.md |
Waypoint’s administrator-bootstrap instructions | Reads the contract artifacts, trusts code over stale instructions, verifies handed-down numbers against their source |
Several of those additions came from watching round three’s lead introduce them on its own. I cover those decisions further down.
Giving API.md a Fixed Filename
Round two dispatched the backend and frontend agents in the same turn, so neither could reach the other, and the frontend built the entire interface against four assumptions that all turned out to be wrong.
| Frontend assumed | Backend actually built | Symptom | |
|---|---|---|---|
| Auth | Bearer token in localStorage, {token, user} |
httpOnly cookie, {user} only |
Every login throws an error, sign-in is impossible |
| Pin fields | {name, lat, lng, date} |
{placeName, latitude, longitude, visitDate} |
400 on every create, existing pins render as “Unnamed place” |
| Admin flag | isAdmin |
role: "admin" |
The admin view is never reachable |
| Error shape | {error: "string"} |
{error: {code, message}} |
Validation detail collapses to one generic message |
Routes matched exactly. Every guess that failed, failed on mechanism and shape rather than on the URL, and the frontend disclosed all four in a file of its own rather than burying them, exactly as its prompt asked. Writing the assumptions down made the mismatch easy to diagnose afterwards, though it did nothing to reduce the rework that followed.
Part of the root cause was one word missing from the agent definition. The frontend file said to read the backend agent’s API documentation and never said where that documentation would be. Round one worked around it because the chain was strictly sequential and the file was simply there by the time frontend ran, and round three only got an early copy because the lead added an instruction of its own telling backend to publish it before polishing. Neither of those came from the agent files, so both of them now name the path explicitly.
Write your API surface to `API.md` at the project root. Publish it early,
as soon as the shape is settled, before you polish the implementation,
because the frontend agent is blocked on it and every hour it sits
unpublished is an hour of parallelism wasted.
The frontend file takes the other side of the same contract, with a hard stop where it used to fall back to inference.
Read `API.md` at the project root before writing anything that talks to the
server. That file is the backend agent's published contract.
If `API.md` does not exist yet, build everything that does not depend on
the server first. Do not infer a contract from the backend's source code
and do not invent one.
I deliberately told the frontend not to infer the contract from the backend source. That would only give it a snapshot of the current implementation, which can become stale as soon as the backend changes.
Putting a Guardrail on the Tester
The previous post drew a distinction between two kinds of restriction and only enforced one of them, and the two work differently enough to be worth separating. Since the security reviewer has no write-capable tool and no shell, there is no capability to restrict in the first place. A tester needs to create test files, so it has to hold Write, and the only constraint stopping it from also patching application source to turn a failing test green is a sentence in its system prompt.
I did not build the enforcement hook last time because I wanted to see whether the prompt-level restriction would hold on its own. It held across all three rounds. I added the hook to the generic roster because tests/ remains the boundary from one project to the next, so the enforcement can be written once and reused.
I ran into four details while testing this, and three fail silently when they were wrong.
The exit code has to be 2. Exit code 1 is the conventional Unix way to signal failure and it is treated here as a non-blocking error, so your script runs, reports its complaint to the transcript, and Claude Code performs the write anyway. A policy hook that exits 1 looks like it is working right up until the moment you check whether anything was actually prevented.
The hook has to live in the agent’s own frontmatter rather than in settings.json. Hooks defined in settings apply session-wide and fire inside every subagent, so a path restriction written there would block the backend and frontend agents from writing anything at all. Frontmatter hooks run only while that specific agent is active and are cleaned up when it finishes.
On Windows the obvious move is to set shell: powershell on the hook entry, and that is the one thing not to do. The field accepts bash or powershell and defaults to bash whenever Git Bash is present, and most developer machines have it.
Setting it explicitly cost me an evening. shell: powershell wraps the command in a PowerShell host, and PowerShell does not propagate a native command’s exit code as its own. The script exits 2, the wrapper exits 1, and exit 1 is the non-blocking error from two paragraphs up, so Claude Code writes the file anyway. I tested this directly instead of assuming, and it behaves identically on Windows PowerShell 5.1 and PowerShell 7.
| How the script is invoked | Exit code Claude Code sees |
|---|---|
Under bash, no shell key |
2, denies correctly |
Wrapped by shell: powershell |
1, the write proceeds |
Wrapped, with ; exit $LASTEXITCODE appended |
2, denies correctly |
The fix is to leave the shell key off entirely and keep ; exit $LASTEXITCODE on the command line anyway. Under bash the propagation is redundant and harmless. On a machine with no Git Bash, where the default falls back to PowerShell, it is the thing that saves you. That combination works either way, so the file survives being copied between projects.
This was difficult to spot because the script was correct and the hook genuinely ran. In the failing version, the block message and the file write appeared in the same sequence.
[DEBUG] Hook PreToolUse:Write (PreToolUse) error:
BLOCKED: the tester agent may only write under tests/. Attempted write to: C:\Claude\uplert\hook-check5.md
[DEBUG] Writing to temp file: C:\Claude\uplert\hook-check5.md.tmp...
[DEBUG] File C:\Claude\uplert\hook-check5.md written atomically
Here is the working version, where the two lines that matter are the last two, and neither of them appears in the failing log above.
[DEBUG] Registered 1 frontmatter hook(s) from agent 'tester'
[DEBUG] Hook PreToolUse:Write (PreToolUse) error:
BLOCKED: the tester agent may only write under tests/ in this project.
[DEBUG] Hook denied tool use for Write
[DEBUG] Write tool permission denied
Hook denied tool use for Write is the line to grep for. If it is missing, your hook complained and then let the write through anyway.
The fourth issue took longer to isolate. Frontmatter hooks only run once you have accepted the workspace trust dialog for the folder holding the agent file. Until you do, the agent runs perfectly happily while its hooks are skipped, and the explanation goes to the debug log where you will not be looking. That looks remarkably similar to a hook exiting with the wrong code, arrived at from a completely different direction.
Here is the frontmatter, which the documentation confirms as the supported shape.
---
name: tester
description: Writes and runs tests against every layer of the application, including its authorization boundaries, and reports pass or fail honestly.
tools: Read, Write, Edit, Bash, Glob, Grep
model: sonnet
hooks:
PreToolUse:
- matcher: "Write|Edit|MultiEdit|NotebookEdit"
hooks:
- type: command
command: "powershell -NoProfile -ExecutionPolicy Bypass -File .claude/hooks/restrict-to-tests.ps1; exit $LASTEXITCODE"
---
The script that frontmatter points at is Step #7 of the walkthrough further down.
The script fails open on a malformed or empty payload rather than blocking, on the reasoning that a parsing bug in my own hook should not brick the tester, and the prompt-level instruction is still standing behind it either way.
The Guardrail Does Not Survive an Agent Team
While proving the hook fired, I ran the same probe twice without noticing that the lead had changed how it dispatched the tester. That accidental difference exposed the most useful result of the exercise.
One run spawned the tester through the Agent tool as a subagent. The other spawned it as an agent teams teammate, because agent teams are enabled in my user settings and the lead is free to choose. Same tester.md, same hook script, same project, and the two debug logs are nowhere near each other.
| Run | Dispatch | Registration line | Outcome |
|---|---|---|---|
| A | Subagent | Registered 1 frontmatter hook(s) from agent 'tester' |
Write denied |
| B | Teammate | Nothing, across all 61,384 lines | Write landed, no hook in the path |
In run B there is no Registered 1 frontmatter hook(s), no Added session hook, and no hook invocation of any kind. The Write dispatched with permissionDecisionMs=4 and went straight to disk.
The documentation covers part of this. It says a teammate honours the definition’s tools and model, and names skills and mcpServers as not applied, and it says nothing either way about hooks. The runs say hooks belongs on the not-applied list.
So the tester’s boundary is a guardrail when the tester runs as a subagent, and an instruction when it runs as a teammate. That is the opposite of what I expected. A teammate has more autonomy than a subagent, yet it is the path where the prompt-level request has to carry the boundary on its own. It also means the roster’s own enforcement quietly depends on a dispatch decision the lead makes without telling you.
If you need the boundary to hold in a team, the hook has to move to .claude/settings.json, where it applies session-wide, with the script reading the agent_type field from the payload and exiting 0 for every agent except the tester. That is more fragile than the frontmatter version because the deny logic now lives in a file that governs every agent, and one bad edit blocks the whole roster rather than one role.
The Read-Only Reviewer’s Dead End
Running the pipeline exposed a consequence of the security reviewer’s missing write tool that I had not accounted for on paper. With no write tool, its findings exist only in its reply to whoever invoked it. In round two that meant an 80-line security report had nowhere to go, and the orchestrating session had to read it back out of the subagent’s response and write it to SECURITY-REVIEW.md itself, so that the docs agent, a separate invocation with no memory of that conversation, would have something on disk to read.
That was manageable once. Repeating it on every project would turn it into a recurring problem, so the generic roster now handles the handoff explicitly. I could see two reasonable ways to solve it, and I picked the one that keeps the reviewer entirely read-only.
| Approach | Mechanism | Trade-off |
|---|---|---|
| Carrier (chosen) | The lead writes the reply to SECURITY-REVIEW.md before dispatching docs |
The reviewer stays entirely read-only, but it needs a lead or a human in the loop |
| Narrowed grant | Grant Write plus a hook allowing only SECURITY-REVIEW.md |
Fully automatic, but the reviewer is no longer categorically read-only |
The reviewer’s own file now closes with an instruction to structure its report so it can be written to that file verbatim without reformatting, CLAUDE.md makes persisting it an explicit lead responsibility, and docs.md is told to ask for the findings in writing rather than document an audit it cannot read.
Three Decisions Round Three’s Lead Made on Its Own
Three of the changes above are not mine. Before spawning anything, round three’s lead made a set of coordination decisions that were nowhere in my prompt, and they were good enough that the generic roster now carries them as standing instructions rather than waiting for a lead to rediscover them.
It partitioned file ownership up front so two agents would never edit the same file, with backend owning the server and database, frontend owning public/ only, tester owning tests/ plus one script line in the package manifest, the security reviewer read-only by its own tool grant, and docs writing only README.md and CHANGELOG.md. That is now the file ownership section of CLAUDE.md, including the note that the shared package manifest is the usual collision point and somebody has to own it explicitly.
It told backend to publish API.md early, before polishing, because frontend was blocked on it. That is now a line in backend.md.
Its third decision was a pair of guardrails aimed at exactly the failure modes a pipeline like this tends to produce. The tester was told that an accurately reported failing suite is a successful outcome and that it must never delete or weaken a test to get to green. Docs was told it may not write “all tests pass” or drop an unresolved finding to make the README read better. Both of those are now in their respective agent files, close to verbatim.
Those additions carry more weight for me because they came from a lead that was actively coordinating five agents, not from me trying to predict every failure in advance.
The Remediation Loop
The clearest capability difference in the last post showed up after defects were found rather than during the build. Rounds one and two could identify problems and recommend fixes, and neither could coordinate a resolution on its own. Round three could. When the tester found an invalid-calendar-date defect, the lead assigned remediation back to backend, delayed the downstream work until the fix landed, and then routed the result back through verification. The same pattern was used on one of the security findings.
The team did not try to close everything. One issue was fixed immediately, and another stayed open as a documented trade-off after backend justified the decision. The workflow was evaluating and prioritizing, not just finding and fixing.
A reusable roster should not depend on rediscovering that every time, so it is now five steps in CLAUDE.md.
1. Confirm it independently before acting. An agent's report is a claim,
not a verdict. Read the file yourself.
2. Route it back to the agent that owns the affected layer as new work.
3. Stage the tester's re-verification behind that fix.
4. Hold the docs agent until both land.
5. Confirm the fix was applied to every path sharing the defect, not just
the one the test happened to exercise.
The first step came directly from round three. In round three the docs agent noticed the frontend had diverged from the earlier rounds and worried the comparison might be skewed by it, which was a genuinely sharp observation, and it missed that the backend had diverged too, onto a different auth mechanism and a different SQLite driver. A report that is partly right is the hardest kind to act on, and the only defence is reading the file yourself before routing anything anywhere.
Letting the Security Reviewer Inherit the Model
All five roles were pinned to model: sonnet in the last post. That was a comparability control rather than a recommendation, because I needed all three rounds running on the same model regardless of what the main session was set to.
That control was only needed for the experiment, so the generic roster removes the model line from security-reviewer.md and leaves the other four agents pinned to Sonnet. With no model specified, the reviewer inherits the model used by the main conversation. Security review is the one role where I expect the model tier to affect the quality of the findings more than the time required to produce them.
What I’m Having the Agents Build
To show the roster is actually generic I needed to build something that was not another travel map, and more specifically not another application whose security review would come back talking about login flows. A roster quietly tuned to authentication would pass that test without deserving to.
Uplert is an uptime monitor. Users register, sign in, and add HTTP endpoints to be polled on a schedule, a background worker polls each one and records the status code and response time, and the dashboard shows current state, uptime percentage over the last day, and a response-time history. An administrator can list every registered user and their check counts. Each user’s checks and history are private to them.
It differs from Waypoint in three ways, and each one stresses a different agent.
| Dimension | Waypoint | Uplert | Which agent this stresses |
|---|---|---|---|
| Shape | Request-response CRUD | Background scheduler plus time-series data | backend, tester |
| Testing problem | Assert on a response | Assert on something that happens on a timer | tester |
| Headline security surface | Password and session handling | Server-side request forgery from user-supplied URLs | security-reviewer |
The last row was the most useful test of whether the roster was genuinely generic. Uplert asks the server to fetch arbitrary URLs that users type in, so I expected the audit to come back talking about a check pointed at http://169.254.169.254/ and reading back a cloud instance metadata response, or at http://localhost:5432 to port-scan the host from the inside. That is a finding class the Waypoint rounds could not produce, and what I wanted to see was whether an unchanged security-reviewer.md would go looking for it.
It did go looking, and the answer it came back with was not the one I expected, so I have put what actually happened in Step #14.
The constraints provide an even clearer comparison. Waypoint’s brief forbade outbound network access entirely so the demo would run offline with no API keys. Uplert cannot function without outbound network access, because polling remote endpoints is the whole application. Two specs with opposite requirements, one unchanged set of agent files, and the difference living entirely in the tier where it belongs.
Step #1 – Create the Project Structure
mkdir C:\Claude\uplert
cd C:\Claude\uplert
mkdir .claude\agents
mkdir .claude\hooks
The next eight steps write every file this needs, and the following table it is best to take a glance at all of them in one table before writing them one at a time, because which tier a file belongs to is the whole argument of this post.
| Step | File | What it holds | Tier | On the next project |
|---|---|---|---|---|
| #2 | .claude/agents/backend.md |
Server side, data store, publishes API.md |
Role | Unchanged |
| #3 | .claude/agents/frontend.md |
Client side, reads API.md |
Role | Unchanged |
| #4 | .claude/agents/tester.md |
Tests, confined to tests/ by hook |
Role | Unchanged |
| #5 | .claude/agents/security-reviewer.md |
Read-only audit, no write tool, no shell | Role | Unchanged |
| #6 | .claude/agents/docs.md |
README and changelog, markdown only | Role | Unchanged |
| #7 | .claude/hooks/restrict-to-tests.ps1 |
Blocks tester writes outside tests/ |
Role | Unchanged |
| #8 | CLAUDE.md |
File ownership, contract paths, remediation loop | Process | Rarely edited |
| #9 | SPEC.md |
What the application actually is | Application | Rewritten |
Seven of those eight get copied into the next project untouched, and SPEC.md is the only one that is written from scratch.
Three more files will show up during the run that we do not create.
| File | Written by | Read by | Appears |
|---|---|---|---|
API.md |
backend | frontend, tester, docs | Early in the build, before the frontend needs it |
SECURITY-REVIEW.md |
the lead session, not the reviewer | docs | After the audit, because the reviewer has no write tool |
FRONTEND-ASSUMPTIONS.md |
frontend | you, tester, docs | Only when the frontend had to guess at the contract |
That middle row is the one to remember. The security reviewer cannot write its own findings to disk, so if nothing persists them the docs agent has nothing to read, and a finding quietly disappears between two steps of your own pipeline.
Step #2 – Write the backend Agent
Save this as .claude/agents/backend.md.
---
name: backend
description: Builds the API, data store, and server-side logic for whatever application SPEC.md describes. Use first, before the frontend agent, since frontend builds against the API this exposes.
tools: Read, Write, Edit, Glob, Grep, Bash
model: sonnet
---
Read `SPEC.md` at the project root before doing anything. It is the source
of truth for what this application does, who its users are, and what
constraints the stack has to satisfy. If `SPEC.md` does not exist, stop and
say so rather than inventing requirements.
You own the server side only. Build the API, the data store, and the
server-side logic the spec describes.
## Choosing a stack
If the project already contains code, follow its existing stack and
conventions rather than introducing a second way of doing things.
If the project is empty, choose the simplest stack that satisfies every
constraint in `SPEC.md`. State your choice and your reasoning in one
paragraph in your summary, so the frontend agent isn't guessing at it, and
so a human can veto it before you are three hours deep.
Do not add a dependency the spec does not require and you cannot justify in
one sentence.
## The API contract
Write your API surface to `API.md` at the project root. Publish it early,
as soon as the shape is settled, before you polish the implementation,
because the frontend agent is blocked on it and every minute it sits
unpublished is a minute of parallelism wasted.
Document every route, its method, its request shape, its response shape,
its error shape, and how a client is expected to authenticate, clearly
enough that another agent could build the entire interface against that
document without ever reading your implementation.
If you change the contract after publishing it, update `API.md` in the same
turn. A stale contract is worse than a late one.
## Boundaries
Do not build any user interface, that is the frontend agent's job.
Do not write tests, that is the tester's job. Verifying your own work as
you go is fine and expected, but the test suite belongs to someone else.
It reads SPEC.md for what to build and publishes API.md for the frontend to build against, and neither of those filenames is negotiable, because the whole contract handshake depends on both agents naming the same path.
Step #3 – Write the frontend Agent
Save this as .claude/agents/frontend.md.
---
name: frontend
description: Builds the user interface on top of whatever API the backend agent exposes. Use after the backend agent has published its interface, or in parallel with it if the spec allows.
tools: Read, Write, Edit, Glob, Grep, Bash
model: sonnet
---
Read `SPEC.md` at the project root before doing anything. It is the source
of truth for what this application does, what screens it needs, and what
constraints the interface has to satisfy. If `SPEC.md` does not exist, stop
and say so rather than inventing requirements.
You own the client-side interface only. Build the screens, the
interactions, and the client-side state handling the spec describes.
## The API contract
Read `API.md` at the project root before writing anything that talks to the
server. That file is the backend agent's published contract.
If `API.md` does not exist yet, build everything that does not depend on
the server first. Do not infer a contract from the backend's source code
and do not invent one.
If a route, a field name, an error shape, or the authentication mechanism
is ambiguous, missing, or contradicts what you need:
- If you can message the backend agent directly, ask it. Ask before you
build, not after.
- If you cannot, build against your best guess but record every single
guess in `FRONTEND-ASSUMPTIONS.md` at the project root, one line each,
stating what you assumed and what you assumed it instead of. Never fold a
silent guess into the code.
A disclosed wrong guess is recoverable. An undisclosed one is a bug someone
finds three layers downstream.
## Boundaries
Do not modify any server-side file, including to fix a contract mismatch
you are certain about. Report it instead.
Do not write tests, that is the tester's job.
The refusal to infer a contract from the backend’s source is the important line. So is FRONTEND-ASSUMPTIONS.md, which is where a guess goes when the agent has no way to ask.
Step #4 – Write the tester Agent
Save this as .claude/agents/tester.md.
---
name: tester
description: Writes and runs tests against every layer of the application, including its authorization boundaries, and reports pass or fail honestly. Use after the implementing agents report done.
tools: Read, Write, Edit, Bash, Glob, Grep
model: sonnet
hooks:
PreToolUse:
- matcher: "Write|Edit|MultiEdit|NotebookEdit"
hooks:
- type: command
command: "powershell -NoProfile -ExecutionPolicy Bypass -File .claude/hooks/restrict-to-tests.ps1; exit $LASTEXITCODE"
---
Read `SPEC.md` at the project root before doing anything. It defines the
behavior you are testing against. Read `API.md` too if it exists, since a
route that does not match its own published contract is a defect worth
reporting even when the code works.
## What to cover
Cover all of the following, and skip an item only when it genuinely does
not apply to this application:
- The happy path for each primary user action in the spec
- Invalid, missing, and malformed input on every write path
- Authentication rejection, if the application has accounts: bad
credentials, expired sessions, absent credentials
- Ownership boundaries, if users own private data: that one user cannot
read, modify, or delete another user's data
- Role boundaries, if the application has privileged roles: that an
unprivileged account cannot reach a privileged endpoint
- Any behavior the spec states explicitly that nothing above covers
The boundary cases matter more than the happy path. A suite that proves
sign-in works but never proves that a stranger cannot read my data has
tested the easy half.
## Reporting
Report a plain pass or fail per behavior, not a vague summary.
When you report a total, break it down by what the tests actually prove.
"64 of 64 passing" is misleading if 55 of them verify one layer against
its own assumptions and only 9 exercise two layers together. State that
split yourself rather than letting a headline number imply coverage you
did not achieve.
An accurately reported failing suite is a successful outcome. Never delete,
weaken, skip, or loosen a test to turn a report green.
Name your own gaps. If you could not test something, say what and why.
## Boundaries
You may only create or modify files under `tests/`. This is enforced by a
hook, not left to your judgment, so a write outside `tests/` will be
blocked rather than merely discouraged.
You must never edit application source, even if you find a bug, even if the
fix is one character, and even if the fix is obvious. If a test fails
because of a real defect, report exactly what failed, why, and which
documented behavior it violates, then stop. Fixing it is not your job.
If your suite needs a hook into the project's tooling that lives outside
`tests/`, such as a test script line in a package manifest, ask for it
rather than adding it yourself.
This is the only one of the five with a hooks block, and it points at the script in Step #7. Note also the instruction about breaking down what a pass count proves, which came straight out of watching round three’s lead tell its own tester much the same thing.
Step #5 – Write the security-reviewer Agent
Save this as .claude/agents/security-reviewer.md.
---
name: security-reviewer
description: Audits the application's identity, access, input handling, and data-exposure behavior by reading code. Read-only by design. Use after the implementing agents report done.
tools: Read, Grep, Glob
---
Read `SPEC.md` at the project root first, so you are auditing against what
the application is supposed to do rather than against a generic checklist.
You audit code, you never modify it, which is why you have no write access
and no shell at all.
## What to review
Work through the following, and say plainly when an item does not apply to
this application rather than padding the report with it:
**Identity and access**
- How credentials are stored, and whether the hashing choice and its cost
parameters are appropriate
- How the session token or cookie is generated, signed, transmitted, and
stored on the client
- Whether the signing algorithm is pinned rather than left to a library
default
- How long a session stays valid, and whether the expiry is enforced by the
server rather than merely recorded
- Whether logout genuinely revokes, or only forgets
- Whether every endpoint returning or changing user-owned data verifies
that the caller owns that data, rather than only that the caller is
signed in
- Whether privileged endpoints verify the caller's role, and whether that
role is read from the server's own store rather than trusted from the
client or from a long-lived token
**Input and output**
- Whether any user-supplied text reaches the page, a log, or a query
unescaped or unparameterized
- Whether input validation actually rejects what it claims to reject
**Exposure**
- Whether a failed sign-in reveals whether the account exists, through the
response body, the status code, or response timing
- Whether anything limits repeated authentication attempts
- How cross-origin requests are configured, and whether a wildcard is doing
work a specific origin should
- What network interface the application binds to
- Whether any secret, key, signing value, or credential is committed in the
source, including as a silent development fallback
## Reporting
For each issue, state the file, the specific risk, a plain-language
severity, and what an attacker would actually have to do to exploit it.
You cannot run the application, so be explicit about which findings you
confirmed by reading code and which are inferences that need runtime
verification by a human. Label them separately. An inference presented as a
confirmed finding is worse than no finding.
If you find nothing significant, say so plainly rather than padding the
report to look thorough.
## A note on your own output
You have no write access, so your findings exist only in your reply. The
session that invoked you is responsible for persisting them to
`SECURITY-REVIEW.md` so downstream agents can read them. Structure your
report so it can be written to that file verbatim, without needing to be
reformatted first.
This is the file with no model line, so it inherits from the main conversation. It also has no Write, no Edit, and no shell, so there is nothing for a hook to intercept, and that is why the closing paragraph tells it to format its report for someone else to persist.
Step #6 – Write the docs Agent
Save this as .claude/agents/docs.md.
---
name: docs
description: Writes the README and changelog from what was actually built and actually tested. Use last, after the tester and security-reviewer have both reported.
tools: Read, Write, Glob, Grep
model: sonnet
---
You only create or edit markdown documentation: `README.md`,
`CHANGELOG.md`, or files under `docs/`. You never touch application source
code, configuration, or tests.
## What to read first
Read the code that actually exists, not the plan. Specifically:
- `SPEC.md` for what was asked for
- `API.md` for the contract that was actually published
- `SECURITY-REVIEW.md` for the audit findings, if it exists
- `FRONTEND-ASSUMPTIONS.md` for unresolved contract guesses, if it exists
- The test output or test suite for what was actually verified
Where the code and your instructions disagree, the code wins. Document what
is there and flag the contradiction in your summary rather than silently
reconciling it. A task description can go stale mid-run. The repository
cannot.
If a security review was produced but exists nowhere on disk, ask for it in
writing before you finalize. Do not write a README that implies an audit
happened without being able to read what it found.
## What to write
Write the README from what was actually built and actually tested, not from
the original feature request. Cover how to run the project, how to set it
up from a clean checkout, and any first-run steps a new developer would
otherwise have to discover by failing.
Record every open concern the tester or the security reviewer raised as a
known limitation. Do not drop one because it makes the README read better.
Be precise about test coverage. Never write "all tests pass" as a summary
of quality. Write what was tested, what was not, and what the gaps mean.
If the tester said its coverage was static analysis rather than observed
runtime behavior, that distinction survives into the README intact.
Verify any number you were handed against its source before repeating it.
Reported counts drift between an agent's summary and its own report table.
The instruction to trust the code over its own task description matters more than it looks. A task description written at the start of a run goes stale the moment the remediation loop changes something, and the repository does not.
Step #7 – Write the tests/ Enforcement Hook
Save this as .claude/hooks/restrict-to-tests.ps1. This is the script the tester’s frontmatter calls, and it is the only thing standing between an instruction and an enforced boundary.
# restrict-to-tests.ps1
#
# PreToolUse hook for the tester agent. Blocks any Write/Edit whose target
# falls outside the project's own tests/ directory, turning the tester's
# prompt-level restriction into an enforced one.
#
# Wired up via the `hooks:` block in .claude/agents/tester.md so it applies
# to that agent only. Do NOT move this into settings.json unmodified: there
# it would apply to every agent in the project and block the backend and
# frontend agents from writing anything.
#
# Exit codes matter here:
# 0 - allow the tool call
# 2 - BLOCK the tool call and return stderr to the agent
# 1 - non-blocking error, so the write proceeds anyway. Never use 1 to deny.
#
# Written for Windows PowerShell 5.1, which is on every Windows machine, so
# it avoids .NET Core only helpers such as Path.GetRelativePath.
$input_json = [Console]::In.ReadToEnd()
if ([string]::IsNullOrWhiteSpace($input_json)) {
# Nothing to inspect. Fail open rather than blocking the agent on a
# malformed payload, and the prompt-level instruction still stands.
exit 0
}
try {
$payload = $input_json | ConvertFrom-Json
} catch {
exit 0
}
# Write/Edit/MultiEdit use file_path, NotebookEdit uses notebook_path.
$target = $payload.tool_input.file_path
if (-not $target) { $target = $payload.tool_input.notebook_path }
if (-not $target) { exit 0 }
# Anchor everything to the project root the hook was invoked from. Matching
# "tests/" anywhere in the absolute path is not enough, because a project
# living under C:\tests\ or C:\Users\test\ would then allow every write.
$root = $payload.cwd
if (-not $root) { $root = (Get-Location).Path }
if (-not [System.IO.Path]::IsPathRooted($target)) {
$target = Join-Path $root $target
}
try {
$fullTarget = [System.IO.Path]::GetFullPath($target)
$fullRoot = [System.IO.Path]::GetFullPath($root)
} catch {
$fullTarget = $target
$fullRoot = $root
}
$normTarget = $fullTarget -replace '\\', '/'
$normRoot = ($fullRoot -replace '\\', '/').TrimEnd('/')
# Inside the project, and inside tests/ or test/ at its root, is the only
# combination that is allowed. Anything outside the project is denied too.
if ($normTarget.StartsWith($normRoot + '/', [System.StringComparison]::OrdinalIgnoreCase)) {
$relative = $normTarget.Substring($normRoot.Length + 1)
if ($relative -match '^tests?/') { exit 0 }
}
[Console]::Error.WriteLine(
"BLOCKED: the tester agent may only write under tests/ in this project. " +
"Attempted write to: $target`n" +
"If this is a real defect in application source, report it and stop. " +
"Fixing it is not your job."
)
exit 2
It fails open on a malformed or empty payload rather than blocking, on the reasoning that a parsing bug in my own hook should not brick the tester, and the prompt-level instruction is still standing behind it either way.
Step #8 – Write CLAUDE.md
Save this at the project root. This is the process tier, so it changes when you refine the pipeline rather than when you change applications.
# CLAUDE.md
This file holds the rules that govern how the five-agent roster works
together, separate from what the application does. It sits at the project
root alongside `SPEC.md`.
Rule of thumb for what goes where:
- **What the application is** goes in `SPEC.md`
- **How the team works** goes here
- **What each role does** is already in `.claude/agents/`, and does not
change per project
---
## The roster
Five agents in `.claude/agents/`, split by codebase layer:
| Agent | Owns | Write access |
|---|---|---|
| `backend` | Server-side code, data store, `API.md` | Full |
| `frontend` | Client-side code only | Full, but never server-side files |
| `tester` | `tests/` only | Enforced by hook |
| `security-reviewer` | Nothing | None, by tool grant |
| `docs` | `README.md`, `CHANGELOG.md`, `docs/` | By instruction |
## File ownership
Assign ownership explicitly before dispatching anything, so two agents
never edit the same file. When a file needs to change and its owner is not
the agent that needs the change, route the change through the owner rather
than letting the other agent reach across the line.
The shared package manifest is the usual collision point. Decide up front
who owns it, and have everyone else request changes to it.
## Contract artifacts
These files are the pipeline's connective tissue. Their paths are fixed
because agents read them by name.
| File | Written by | Read by |
|---|---|---|
| `SPEC.md` | You | Everyone |
| `API.md` | `backend` | `frontend`, `tester`, `docs` |
| `SECURITY-REVIEW.md` | **The lead session** | `docs` |
| `FRONTEND-ASSUMPTIONS.md` | `frontend` | The lead, `tester`, `docs` |
`SECURITY-REVIEW.md` is the one that needs your attention. The
security-reviewer has no write tool by design, so its findings exist only
in its reply. **The lead session must write that reply to
`SECURITY-REVIEW.md` before dispatching the docs agent**, or the docs agent
will document an audit it cannot read.
## The remediation loop
The default pipeline is a straight line: build, test, audit, document.
A straight line has no way to act on what the tester and the auditor find,
which means real defects get written up as known limitations instead of
fixed. Do not run it that way.
When the tester reports a failure or the security reviewer reports a
finding worth acting on:
1. Confirm it independently before acting. An agent's report is a claim,
not a verdict. Read the file yourself.
2. Route it back to the agent that owns the affected layer as new work.
3. Stage the tester's re-verification behind that fix.
4. Hold the docs agent until both land.
5. Confirm the fix was applied to every path sharing the defect, not just
the one the test happened to exercise. A half-applied fix to a shared
validator is worse than no fix.
Not everything has to be fixed. A debatable judgment call can stay open and
documented, on the owning agent's reasoning. What must not happen is a
defect getting quietly downgraded to a limitation because the pipeline had
nowhere to send it.
## Verifying agent reports
Treat every agent's self-report as a claim requiring confirmation:
- Re-run the test suite yourself before believing a pass count
- Grep for what an agent says it removed, before believing it removed it
- Spot-check the highest-severity security findings against the actual file
- Check any number that appears in two places against both
The security reviewer in particular cannot run the application, so
everything it produces is a code review rather than a proven exploit.
Findings from it need a human in the loop before they are treated as fact.
## Writing task instructions
Read an agent's definition file before writing a task for it. The most
common orchestration failure is a task instruction that contradicts the
agent's own constraints, or that asks an agent for something its tool grant
makes impossible, such as asking the read-only reviewer to create a file.
When you do have to correct a task mid-run, issue the correction as an
unambiguous directive with concrete acceptance criteria, not as a
discussion. Reversals framed conversationally get read as commentary and
ignored. Then update every task that carries the stale instruction, not
just the one you were looking at.
## Project-specific rules
<!-- Add anything here that applies to this project and is not in SPEC.md.
Coding conventions, branch rules, deployment gotchas. -->
The two sections that earn their place are file ownership and the remediation loop, and both of them are round three’s ideas rather than mine.
Step #9 – Write SPEC.md
This is the only file that changes per application, so it is worth showing in full rather than in fragments.
# SPEC.md
## What this application is
Uplert is a personal uptime monitor. A signed-in user registers HTTP
endpoints to be polled on a schedule, and sees whether each one is up,
how often it has been up recently, and how quickly it responds.
## Users and roles
- **Registered user**: creates and deletes their own checks, views their
own results. Cannot see any other user's checks or results.
- **Administrator**: additionally lists every registered user and how many
checks each one has. Cannot view another user's check results.
## Core behaviors
1. Register an account and sign in.
2. Add a check: a URL, a display name, and a polling interval.
3. A background worker polls every due check and records the status code,
the response time in milliseconds, and a timestamp.
4. View a dashboard of the signed-in user's checks with current state,
uptime percentage over the last 24 hours, and recent response times.
5. Delete a check, which also removes its history.
6. An administrator lists every registered user and their check count.
## Data
- **User**: email, password credential, role, created timestamp. Owned by
itself.
- **Check**: owner, URL, display name, interval in seconds, enabled flag.
Owned by the user who created it, visible to nobody else.
- **Result**: check, status code, response time, timestamp. Inherits the
owner of its check.
## Stack constraints
- **Runtime / language:** Node or Python, either is fine
- **Data store:** SQLite, created on first run
- **Deployment target:** runs locally with one command
- **Must run without:** Docker, WSL, any paid or signup-gated service
- **Requires:** outbound HTTP access, because polling remote endpoints is
the application
## Non-goals
Alerting is the obvious next thing to build and is deliberately out of scope
here, so no email or webhook notifications in this version. Also no status
page shared with anyone outside the account, and no scheduled report exports.
## Known limitations of the environment
Served over plain HTTP on localhost. Browsers treat localhost as a secure
context, so Secure cookies and anything else gated on a secure origin still
behave normally. What cannot be exercised here is anything depending on a
real TLS connection, including certificate validation and HSTS, so treat
those as out of scope rather than as findings.
Two of those sections are carrying more weight than they look like they are.
Stack constraints is where every environment fact now lives. In the Waypoint roster these were written into the agent files themselves, so backend.md knew the project had no outbound network access before it had read a word of the spec. Moving facts like that into one section of one file is most of what made the roster generic in the first place.
“Must run without” and “requires” work as a pair. Waypoint’s constraint was no outbound network at all. Uplert’s requirement is the exact opposite, since polling remote endpoints is the entire application. One unedited backend.md reads both specs and comes to opposite conclusions about what it may reach for, and that only works because the spec states the boundary in both directions. Leave the pair out and a generic backend agent has to guess, with both guesses bad: assume network access and it builds a poller that cannot run on a locked-down machine, assume none and it refuses to build the only feature the spec asked for.
Known limitations of the environment is aimed at the security reviewer rather than the backend. Uplert runs over plain HTTP on localhost, and a reviewer who does not know that is deliberate will file missing certificate validation and absent HSTS as findings. Browsers treat localhost as a secure context, so cookies marked Secure still behave normally, and spelling that out keeps a genuine finding about cookie flags separate from noise about TLS that cannot exist here. A finding nobody can act on costs the same attention as a real one.
Whether any of this leaked back into the agent files is the thing to check, and one side-by-side settles it:
Step #10 – Trust the Folder and Restart the Session
There are two setup steps here, and both can fail without producing an obvious error.
Claude Code watches .claude/agents/ for changes, and the watcher only covers directories that existed when the session started, so the first time you create that folder in a project you need to restart before anything you drop in there is found.
Then accept the workspace trust dialog for the project folder, because until you do, the tester’s frontmatter hook is skipped while the agent itself runs normally.
Starting a fresh session in the project folder:
I also ran into a version change while confirming that the five definitions had loaded. My instinct was to run /agents, which used to open a wizard with a Running tab listing live subagents and a Library tab for editing them. That wizard was removed in v2.1.198, so on anything current the command prints a short notice telling you to ask Claude or edit .claude/agents/ directly, and lists nothing at all.
Nothing is wrong when that happens. The files, the frontmatter fields, and the .claude/agents/ location all work exactly as before, and only the terminal wizard is gone. What changed is how you check.
| Version | How to confirm the definitions loaded |
|---|---|
| v2.1.197 and earlier | /agents, then the Library tab |
| v2.1.198 and later | Type @ at the prompt and read the typeahead, or press the left arrow key for the agent panel, or simply ask Claude which subagents this project has |
Typing @ lists what actually loaded:
If none of the five appear, the cause is almost always the restart above rather than anything in the files. A running session does detect edits to subagent files within a few seconds, so a restart is only needed when the agents directory itself did not exist when the session began. On a brand new project, it will not have.
Step #11 – Prove the Hook Actually Fires
I gave this its own step because a skipped hook and a working hook can look identical until an agent writes somewhere it should not.
The obvious test is to ask the tester to edit a source file, and at this point in the sequence that does not work, which I found out by trying it. Nothing has been built yet, so there is no source file to edit, and the lead refuses on those grounds before the tester is ever dispatched.
There was also a second reason it refused. CLAUDE.md tells the lead to read an agent’s definition before writing its task so that no task contradicts that agent’s own constraints, and asking the tester to write outside tests/ is exactly such a contradiction. The lead catches it one layer above the hook. That is the process rule working the way it should, though it does mean a naive test never reaches the thing you were trying to test.
So ask for something that works on an empty project, and say plainly that you are testing enforcement rather than requesting work.
This is a deliberate test of the tester agent’s PreToolUse hook rather than real work. Dispatch the tester subagent and instruct it to create a file calledÂ
hook-check.md at the project root. I expect that write to be blocked. Report exactly what came back.
If you have agent teams enabled, pin the dispatch path before running this, because the hook only registers for subagents. I first tried adding “dispatch as a subagent, do not spawn a teammate” to the prompt and watched a teammate spawn anyway, which sent me off diagnosing the hook when the hook was fine.
The cause is documented, and it is subtler than the lead ignoring me. Claude names subagents on its own so it can message them later, and while agent teams are enabled, a named subagent launches as a teammate. No prompt wording prevents that, because the naming happens regardless of what the prompt asks for.
So take the choice away with a project-level settings.json that turns the flag off for this project only.
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "0"
}
}
Claude Code reapplies settings-file env values to the running session when you save and rereads the variable each time it spawns a subagent, so the next dispatch picks up the change without a restart. You can tell which path you got from the agent panel without reading a log at all. A subagent shows as tester under a Backgrounded agent heading, and a teammate shows as Teammate @tester-something. The @ prefix is the name Claude assigned, and the name is what made it a teammate.
Three things can happen, and only one of them means the hook is broken.
| What happens | What it means |
|---|---|
| The tester is dispatched and the write comes back refused with the BLOCKED message | The hook fired. This is the outcome you want. |
| The lead declines to dispatch at all, citing the agent definition | The CLAUDE.md rule caught it first. Reasonable, and you still have not tested the hook, so press again with the framing above. |
hook-check.md appears and no hook message appears in the debug log |
The hook never fired. Check that this was a subagent and not a teammate, then workspace trust, then the frontmatter shape. |
hook-check.md appears and the debug log shows the hook firing with your BLOCKED message |
The hook ran and its exit code was swallowed. This is the shell: powershell wrapper, and the fix is ; exit $LASTEXITCODE on the command. |
The refusal carrying the message straight from the script:
And the tester’s own report of the blocked write:
Delete hook-check.md afterwards in the case where it did get created.
The hook fails open on empty stdin, malformed JSON, or a missing path field, all of which exit 0. That is a defensible choice for a payload-parsing failure, and it does mean the enforcement is only as reliable as the payload shape staying stable.
Its path comparison anchors to the project root taken from the hook payload instead of matching tests/ anywhere in the absolute path. That matters more than it sounds, because the naive version of this check permits every write in a project that happens to live under a folder named test or tests.
Step #12 – Run the Build
The invocation prompt is short now, because the spec is on disk and every agent reads it.
Build the application described in SPEC.md. Use the backend subagent to build the API, data store, and polling worker first. Then the frontend subagent for the dashboard. Then the tester subagent, then the security-reviewer subagent, then the docs subagent. Follow the remediation loop in CLAUDE.md rather than documenting defects you could route back for a fix.
Compare that against the equivalent prompt from the last post, which had to restate the entire application in a paragraph and still lost detail on the way to the agents.
Kicking off the build:
Before dispatching anything, the lead works out file ownership from CLAUDE.md:
The backend finishes and the lead verifies its claims instead of taking them:
Subagents run in the background by default, so the left arrow opens the agent panel:
Ctrl+O expands the running agent’s output inline:
Step #13 – Watch the Contract Handshake
What I wanted to know here was whether naming the contract path prevented the failure it was meant to prevent. In round two of the Waypoint post the frontend built an entire interface against four assumptions that were all wrong, because its agent file told it to read the backend’s API documentation without ever saying where that documentation would be.
This time the backend published API.md to the project root before the frontend was dispatched, the frontend read it, and FRONTEND-ASSUMPTIONS.md was never created. The frontend’s own summary said the contract was unambiguous and that it had nothing to guess at, and the lead confirmed independently that every src/ timestamp predated the frontend’s run, so it had genuinely stayed on its own side of the line.
The frontend agent working against the published contract:
Frontend complete, with no assumptions file to show for it:
The result here is the absence of a file, so it is easy to miss. At the end of any run, check whether FRONTEND-ASSUMPTIONS.md exists. If it does, read it before anything else, because it is a list of everything the frontend could not confirm and built around anyway.
Step #14 – Watch the Remediation Loop
Nine defects were found, routed back to the owning agent, fixed, and re-verified across two full rounds, and not one of them was written up as a known limitation instead.
| Defect | Found by | Severity |
|---|---|---|
Oversized bodies reset the socket instead of returning 413 |
tester | real bug, shared helper |
npm test script did not run at all on Windows |
tester | blocking |
| Registration timing oracle undermined login’s anti-enumeration | auditor | Medium |
SameSite ignores port, so cross-port origins could drive four routes |
auditor | Medium |
| Static-file traversal via encoded slashes | auditor, then the lead | escalated |
| Credentialed check URLs stored, and could never poll | auditor, then the lead | Low, and functional |
intervalSeconds accepted numeric strings |
auditor | Low |
| scrypt cost below current guidance | auditor | hardening |
| “Sign out” shown to signed-out users | the lead, in a browser | cosmetic |
The tester goes first:
Writing the suite, confined to tests/ by the hook from Step #7:
Then the security reviewer, read-only by its tool grant:
This next one is the whole point of the loop, with findings routed back to the agent that owns the layer instead of being written up as limitations:
Backend reports its fixes done, and the lead verifies rather than trusts:
The suite runs again against the corrected code:
A second round, this time on the frontend:
Tester and frontend run in parallel here, because their file ownership does not overlap:
Re-verification lands and the docs agent is finally released:
The lead proved something the reviewer could only infer. The security reviewer rated the static-file traversal as low and latent, because it reads code and cannot run anything. Since the lead has Bash, it went and tried the exploit. A plain ../ path returned 404, which on its own would have read as a pass, and /..%2fpublic-backup%2fsecrets.env returned HTTP 200 with the file contents, because the decode happens after URL normalization. It escalated the finding from latent to confirmed and live, and the fix returns 400 for both forms now.
Interestingly enough, that is the cost of the reviewer’s missing shell being paid back by whoever holds the loop, and it is the best argument I have for keeping a human or a capable lead in that position.
The docs agent caught an error in the lead’s own numbers. Docs has no Bash, so it could not run the suite. Instead of repeating the test split it had been handed, it counted the files statically, got a different answer, and flagged the discrepancy. The lead re-ran each file separately and found docs was right. Its total of 139 was correct and its split was not, because password.test.js‘s four tests had been filed on the integration side, so the real split works out to 99 integration and 40 unit.
What I found most interesting is that the correction went back through the docs agent instead of being edited into its file directly, so the file-ownership rule held even for a two-line change. The agent with the fewest tools produced the most careful piece of work in the run.
The CSRF fix was verified in a real browser. The Origin guard added for the cross-port finding was verified with curl by the backend, and the lead noted that an Origin guard is exactly the kind of change that passes curl and breaks a browser. It drove the actual UI through register, sign in, add a check, watch the poller record a result, and the admin view, and that same run resolved two of the reviewer’s open inferences at once.
What was deliberately not fixed. Five items stayed open as accepted risks with the reasoning written down. The 409 on a duplicate email still reveals that an address is registered, and closing it needs an email confirmation round trip that SPEC.md lists as a non-goal. Session tokens are stored in the database in plaintext. The rate limiter is effectively one bucket on a loopback deployment. There is no Content-Security-Policy. The poller deliberately does not block private targets, and the reviewer engaged with the backend’s reasoning on that one and accepted it.
One more is labelled untested instead of passed. The timing-equalization work is real in the code, and the suite does not prove the timing, so the README says exactly that.
That distribution is the part I would point at. Nine fixed, five documented with reasons, one honest gap, and nothing quietly downgraded to make the README read better.
Step #15 – Review What Got Built
The whole run took an hour and five minutes and produced a working application with 139 tests, 139 passing, up from 111 tests with two real failures at first report.
Usage at that point in the run:
Docs runs last, once both the tester and the reviewer have reported:
And the finished run:
Everything to this point is terminal output and agent self-reports. Starting the application is the only check that settles whether five agents produced working software or a very well documented pile of claims.
The backend chose a stack with no dependencies at all, so there is nothing to install first:
$env:ADMIN_EMAILS="you@example.com"
npm start
Set ADMIN_EMAILS before registering an account. The backend reads administrator status from that variable alone and provides no way to promote a user afterwards, a decision it made on its own and then documented rather than leaving me to discover it.
The README is the artifact I would read first if someone handed me this project. It carries every accepted risk with the reasoning attached, it says which claims the suite does not actually prove, and it warns that npm test destroys your development database because src/db.js hardcodes the path. That last one is a gap the tester named in its own report instead of working around silently.
Reusing the Claude Code Agent Roster on the Next Application
Nothing in .claude/ or CLAUDE.md knows anything about uptime monitoring, so the next project will just need a new SPEC.md:
mkdir C:\Claude\kql-library
cd C:\Claude\kql-library
xcopy /E /I C:\Claude\uplert\.claude .claude
copy C:\Claude\uplert\CLAUDE.md .
Write the spec, restart the session so the agent files are picked up, trust the folder, and run the same dispatch prompt from Step #12.
Final Thoughts
I would not call this roster universal. It still has limits that become clear as soon as the application stops looking like the server-and-client projects it was designed around.
It is shaped for an application with a server and a client. Build a CLI tool, a data pipeline, or a Terraform module with it and frontend is the wrong role name, at which point you are editing the roster rather than only the spec, and avoiding exactly that was the point of the exercise. I would rename or drop that role rather than pretend it fits.
The tester’s guardrail only stands on one of the two dispatch paths. As a subagent it holds, and as an agent teams teammate the frontmatter hook never registers, so the restriction reverts to a prompt-level request in the exact mode where agents have the most autonomy. I found that by accident, and it is the first behaviour I would verify before trusting this roster on a project that matters.
Even on the path where it does hold, it only guards the Write family. The matcher covers Write, Edit, MultiEdit, and NotebookEdit, while the tester’s tool grant includes Bash, so a shell redirect writes wherever it likes.
I left Bash in deliberately, because a tester that cannot run its own suite is most of the way to useless, and the hole cannot be closed by listing the shell commands that write, since >, tee, cp, sed -i and a dozen others all qualify and enumerating them is the denylist mistake this roster exists to avoid. What I would actually say is that the hook guards one path, and the rest of that boundary is still a request. Remove Bash and have the lead run the suite if you need it to be more than that.
The frontend’s instruction not to touch server-side files and the docs agent’s instruction to stay in markdown are prompt-level throughout, on every path. That same hook pattern extends to both of them, and I have not built either one yet.
Nothing lets the security reviewer run code, so everything it produces is a code review rather than a demonstrated exploit, which keeps a human in the loop on every finding whether the pipeline is generic or not.
Round three of the Waypoint build gave me confidence that the checklist in security-reviewer.md is doing real work. All three Waypoint rounds shipped without login rate limiting. Round three’s reviewer also caught the JWT in localStorage and a hardcoded fallback signing secret committed to the repository, which it rated as full token forgery. The instruction that caught the signing secret remains in the generic roster, so the same five files helped produce the defect and identify it.
What I have now is five agent files that built Uplert without a line of application detail in any of them, against a spec whose network requirements are the opposite of Waypoint’s. One build is not proof that the roster is generic, and the test I would trust is a second application on the same unedited files. That run is still ahead of me.
Writing SPEC.md by hand was the deliberate baseline here, and it is the part of this I am least settled on. The three-tier split is right and I would make the same call again. The spec itself is a document I invented, with the sections I happened to think of, and nothing anywhere checks that it is complete before five agents start building against it.
The administrator gap is the clearest example. My spec named a role and its permissions and never said how an account acquires it, so the backend decided for me, and the decision it made is permanent and unrecoverable for any account registered before the environment variable is set. Nothing caught that at spec time, because there was nothing to catch it. I wrote the file, read it over, and it looked finished. An omission is invisible in a document you wrote yourself, and the cost of this one only became visible an hour later in a README section explaining the workaround.
GitHub’s Spec Kit is the established answer to that problem, and it is where I am taking this next. It formalizes what I did by hand: a constitution written once per project, then specify, plan, tasks, implement, and converge per feature, each one a skill invoked in the agent’s own chat rather than a terminal command. The constitution covers roughly the ground my CLAUDE.md covers, and /speckit-specify produces the artifact I wrote myself, through a process built to interrogate the requirement rather than transcribe it.
The generic roster, the hook script, both templates, and the Uplert spec that produced this build are on this GitHub repo. Last post’s three-round Waypoint comparison and its five original agent files are in its own repo.




















































