AI Coding
Prompt Engineering for a Code Review That Finds Real Bugs
Ask an AI reviewer for specific failure modes, with the diff and the invariant, instead of a generic look at this code.
- Prompt Engineering
- Code Review
- AI Coding
- Quality
On this page
"Review this code" is a request for a summary and a compliment. The model will tell you the function is clear, suggest a rename, and miss the missing auth check. A review prompt has to name the failure you care about, paste the diff rather than the whole repository, and say what must remain true after the change. You are not asking for a vibe. You are asking a careful reader to try to break one claim.
Use the assistant as a second pass after you already understand the change. It is good at spotting a branch you did not test, a null you assumed, and a secret that slipped into a log. It is bad at knowing your product's politics, your on-call reality, and whether a TODO is load-bearing. Keep those decisions with a human. Put the mechanical claims in the prompt.
State the claim the diff makes
Before you ask for bugs, write one sentence the author believes. "A signed-out user cannot read another account's invoice." "This action rejects a budget above one million." The review then has a job: find an input that makes the sentence false. Without that sentence, every comment is optional style.
Review this diff against one claim: a signed-out caller cannot create a project.
Look only for:
- missing authentication or authorization
- validation that trusts the client
- error paths that still write
- secrets or raw provider errors returned to the client
Ignore naming and formatting. Quote the line you do not trust and say the input that breaks it. If you cannot break the claim, say so.
Give the diff and the boundary
Paste the changed function and the type of its input. Paste the caller if the bug is likely at the boundary. Do not paste node_modules, lockfiles, or a generated client. Tell the model which framework assumptions are real: params are promises, this route is a Server Action, the anon key is public. A review that "fixes" a Server Component by adding useEffect has not seen the constraint.
- Include the invariant in the first lines. Models weight the start of the prompt.
- List the bug classes you want. Security, data loss, and broken error paths beat style.
- Forbid drive-by refactors. A review that rewrites the file hides the one real note.
- Ask for a concrete input. "Might be null" is weaker than "formData.get returns null when the field is omitted."
Ask in rounds, not in a pile
One round for security and authorization. A second round for failure handling. A third, only if you need it, for tests that would lock the claim. When you ask for everything at once you get a list of ten nits and the auth bug is item eight. Stop when the comments are about taste. Taste is a human conversation.
| Round | Ask | Reject |
|---|---|---|
| 1 | Who can call this, and what proves it? | Rename suggestions |
| 2 | What happens when the dependency fails? | A new architecture |
| 3 | Which test would fail if the claim broke? | Tests for getters that only return a field |
Check the finding before you change code
A confident paragraph is not a bug. Reproduce it. If the model says the action is unauthenticated, find the function that was supposed to check the session and see if it is actually called. If you cannot reproduce the issue, say so in the review thread and move on. Applying every suggestion is how a safe function grows a useless abstraction and a new dependency.
Do not paste production data, customer rows, or live keys into the prompt to "make the review realistic." A synthetic FormData payload is enough. If the bug only appears with a real token, you have a logging problem as well as a review problem. Strip secrets before you ask.
Make the good prompt the default
Save the three-round prompt in the repo's contributor notes. Change the claim per pull request. The rest can stay. Teams that rewrite the prompt from scratch each time slide back to "review this." The value is the constraint: one claim, a diff, a bug class, and a request to quote the line. The model is a reviewer with a checklist. You still merge.
A finding worth keeping
Useful output looks like this: the line that returns the signed URL is reachable without requireUser when the id is passed as a query string, and the input that proves it is a request with no cookie. You can write a test or a curl for that. Useless output looks like this: consider extracting a helper for readability. Apply the first, decline the second, and tell the model which class of comment you will not take so the next round stays on the claim. If three rounds produce only style, the diff may actually be fine. Say that out loud. An endless review is not higher quality. It is a way to avoid merging. The prompt worked if it either produced a bug you could reproduce or gave you confidence to stop.
