How to verify an AI agent's fix before you trust it
An AI agent's passing test is not proof of a security fix. Check the original failure, legitimate access and deployed code before accepting the repair.
Sep 28, 2026 · 8 min read
Field Notes from Thunkle, a studio that takes AI-built apps from prototype to secure, production-ready software.
By Thunkle · Verification workflow, checked 28 September 2026
An AI coding agent changes the code, adds a test and reports that everything passes. Before you accept that, run the original failing request again. It is the most useful thing to run after any repair, and a new green test proves very little if it would also pass on the broken version. This is how we review AI-generated code changes on client projects, whether the agent is Claude Code, Cursor, Codex, Lovable or Replit.
The example throughout is an illustrative missing ownership check on a document. Adapt the workflow to the specific failure and its safe test conditions.
If an agent has made security changes you cannot independently verify, do not treat its confidence as sign-off. Thunkle's paid AI-generated code review checks the implementation and the evidence behind the claimed repair. We can also scope the hardening work when the fix is incomplete.
The three results to ask for
For a missing ownership check, we want to see three cases run on both the broken and the fixed version:
- Owner reads their own private document. Broken: allowed. Fixed: allowed.
- Unrelated customer requests that document. Broken: incorrectly allowed. Fixed: refused.
- Signed-out visitor requests it. Broken: should be refused. Fixed: refused.
The failed customer boundary should change. The legitimate workflow should keep working. If the agent can only show you the second line, you do not yet know whether the fix broke the product.
Save the failing case before you ask for the repair
Keep the request, the account relationship involved, the expected result and the actual result. Use synthetic data in an environment you own or are authorised to test.
For the document example, the useful fact is that account B can request a document belonging to account A. That specific permission failure matters more than a generic instruction to "improve security". The two-account test is the quickest way to capture it.
Do not put real credentials or private documents in a public issue or a prompt. Reproduce the relationship safely and keep only what the developer needs. Now the repair has a result to aim for, and something independent of the agent's explanation that can be checked afterwards.
Make the new test fail for the right reason
Run the same test against the pre-fix revision in an isolated environment, then against the repaired revision.
The first run should fail because the wrong customer received access. A missing dependency, a broken fixture or a disconnected database does not establish that the test catches the vulnerability. Keep the assertion and account setup identical between runs. A test rewritten to match each implementation cannot demonstrate the change.
An instruction you can give an agent:
Run the authorization test against the version before your fix and the fixed version, using the same synthetic records and assertions. Show the relevant response and why the first run fails. Confirm that the owner still has access. If you cannot execute either version, state what is unverified.
If both versions pass, check whether the test exercised the real access decision. Mocking that decision away removes the very behaviour you meant to test.
Check whether the fix belongs in more than one route
The reported issue may concern previewing a document. The app may also extract its text, generate a download link or delete it. Trace whether those operations share the corrected authorisation logic or implement separate checks. Do not assume the preview repair protects every file operation.
For a workspace product, include an authorised colleague and a removed member if those relationships affect access. The expected result comes from the product rules. Refusing everybody is not a successful security fix.
Keep this expansion proportionate to the finding. The goal is to test related paths, not to turn every small repair into a full audit.
Look at what the test is allowed to assume
Suppose the test replaces the authorisation function with a mock that returns "permitted" for the owner and "refused" for everyone else. It may accurately test how the interface handles those two results. It cannot establish that the real authorisation function makes the right decision.
That is not an argument against mocks. It is a reason to name the boundary of each test. A unit test isolates a calculation. A route test checks how a request is handled. A test against the database policies examines a different link in the permission chain. For the original finding, choose the test that includes the decision that failed.
Ask the agent to explain what remains real and what is substituted. That explanation often reveals whether a reassuring test name corresponds to a meaningful check.
Watch for a repair that changes the question
There are several ways to make a failing example disappear without correcting its cause. The endpoint could stop returning any documents. The test account could quietly gain permission. The fixture could stop creating the unrelated record.
Inspect those changes before accepting the result. Keep the account relationship and the intended operation stable. If the product rule really has changed, document the new rule and test it explicitly rather than calling it the same security repair.
The positive case matters here. The owner should still be able to process the document. An authorised colleague should still succeed if sharing is part of the product. Otherwise the fix traded an access control issue for an unusable feature.
Read the patch for unrelated changes too. A small permission repair should not alter subscription limits or remove error handling. If more work was necessary, the agent should be able to explain why it belongs to the same defect.
When the old version cannot be run safely
A dependency may be gone, the test environment may be incomplete, or the original behaviour may depend on configuration that was not preserved. Do not pretend the before-and-after experiment happened if it did not.
Keep the original evidence and state the limitation. You may still inspect the code path, reproduce the issue in a safe fixture or verify the repaired behaviour in an appropriate environment. Those are useful checks that support a narrower conclusion.
Never restore a known vulnerable version to production to get a cleaner test result. Use an isolated setup with synthetic data, or record why that part of the verification is unavailable.
Ask for an evidence summary you can challenge
A completion note worth trusting names:
- the rule that changed, in one sentence;
- the tests that actually ran, distinguished from tests the agent merely wrote;
- the original failing request, run again, now refused;
- the owner's normal access, still working;
- any configuration or deployment step still required;
- anything that could not be tested, named plainly.
If the note says "all tests pass", ask which tests and on which revision. If the answer is a build command, you have evidence that the project builds, not that the customer boundary is repaired. That distinction lets you use an agent productively without treating its confidence as a security control.
Then test what is actually deployed
A passing local test does not establish that the published app contains the fix. Confirm the intended revision and any database policies, migrations or configuration it needs.
Repeat the safe original case against the deployed environment. Check both the refused operation and normal use. If a scanner originally reported the issue, rescan as another source of evidence within its coverage. When the result disagrees with local testing, investigate the deployment, the environment and the exact request. Do not declare the scanner wrong or edit the test to make the status green. Our note on automated scanners versus manual review explains what each can and cannot see.
Close the finding with a small evidence record
Keep the original behaviour, the affected revision, the repair, the test results and the deployed verification together. If something could not be tested, say so.
That record tells the next person what "fixed" means, and it makes a future regression easy to recognise. If you would like an independent review of the fixes an agent has made, Thunkle's paid security audit includes evidence-backed findings, remediation guidance and a free re-review of agreed fixes. Implementation is scoped separately. Request a review quote with the app URL and the changes you need checked. Tool-specific guidance is in our Claude Code and Cursor security guides.
Common questions
How do I review AI-generated code for security?
For a claimed security fix, start from the original failing case. Run it on the version before and after the change, keep the assertions identical, confirm the legitimate path still works, then inspect the patch and deployed behaviour. A broader code review also covers paths where no finding has been reported yet.
Is "all tests pass" enough to accept an agent's fix?
No. Ask which tests ran, on which revision, and whether any of them would have failed on the broken version. A test that passes on both versions has not tested the fix.
Should I let the agent write the verification test?
Yes, but pin the scenario yourself: the accounts, the request and the expected results. The agent can implement it. It should not be free to redefine what counts as passing.
What if the agent says it cannot run the old version?
Accept the narrower conclusion and record the gap. Inspect the code path and reproduce the issue in a safe fixture. Do not restore a vulnerable build to production to get a cleaner result.
Does this apply to Lovable and Replit as well as Claude Code and Cursor?
Yes. The platform changes how you run the tests, not what needs proving. The vibe coding security launch kit has the checklist version for builder platforms.
Working on something like this?
If you're migrating, securing or building an AI-made app, Thunkle can help. Get a fixed quote in 24 hours.
