Implementation worksheet · 6 min read
A Requirements Traceability Matrix for an AI-Built App
A traceability matrix has one row per requirement in the brief, linking it to where it was built, the acceptance test that checks it, the evidence that the test passed, and the build that evidence came from. For an AI-built app the last column matters most: a later change can alter code that satisfied an earlier requirement, so 'verified' only means something next to a build identifier. Keep requirement IDs stable, record evidence as artifacts rather than ticks, re-test the rows a change could touch, and treat any row without evidence on the current build as unverified, whatever its status once said.
Scope: the actor is a founder, product lead or agency delivering an AI-built app to a client or stakeholder who asks 'is everything in the brief done, and how do you know?'. The starting state is a versioned brief with numbered requirements and a deployed build. The boundary is the functional, permission and data requirements in the brief; performance, security and accessibility have their own tests, which can appear as rows that point to those results. The outcome is a matrix a reviewer can audit row by row without asking the builder.
Put it into practice
1. Number the requirements and never renumber
R1 to Rn, each a single testable statement. Split compound requirements ('members can book and cancel rooms') into separate IDs. A retired requirement keeps its ID and is marked retired, so older evidence still points at something.
2. Record where each requirement was built
The change request, prompt or commit that implemented it. In an AI builder that may be a chat turn or a version in the project's history; record whatever identifier lets someone find it again. Later, this is what shows which requirements a new change could have affected.
3. Write one acceptance test per requirement
State the exact input and the expected result: 'Book room A 10:00-11:00 as member 1, then 10:30-11:30 as member 2: the second booking is refused with a message naming the conflict.' A requirement without a test can only ever be marked 'believed done'.
4. Attach evidence, not ticks
A screenshot with the URL and time visible, a short screen recording, a test-run log or an exported record. A tick in a status column is a claim; an artifact is evidence someone else can check without taking your word for it.
5. Stamp each result with the build it ran on
Record the build or version identifier. A pass on yesterday's preview doesn't cover today's production build, and the matrix has to say which one was tested.
6. Re-test the rows a change could touch
Before accepting a change, note which rows share its area: the same screen, table or workflow. Re-run those tests once it lands. A change can reach further than its request, as the guide on scoping changes describes, and the matrix tells you what to re-check instead of re-testing everything or nothing.
7. Report by status on the current build
Three counts: verified on the current build, verified only on an earlier build, and never verified. Watch the middle one. It grows with every change and only shrinks when someone re-tests, so it is the honest measure of how much of the brief is currently proven.
8. Worked example (illustrative, synthetic data)
A room-booking app with twelve requirements, one of them retired. After build 14: nine verified on build 14, one verified only on build 11 (R7, maintenance block), one never verified (R12, month-end export, because the test data never crossed a month end). Change 15 edits the booking form, and the rows sharing that area are R1 to R4. After build 15, R1 to R3 pass again and R4 fails: the late-cancellation fee is now charged 30 hours ahead instead of only inside 24 hours. The report after build 15 reads: 3 verified on build 15, 1 failed on build 15, 6 verified only on earlier builds, 1 never verified. Without the matrix, R4's last recorded status would still read 'done'.
Requirements traceability matrix (worked example)
Copy this structure into your review document and record your observed result for each row.
| Requirement | Built in | Acceptance test | Evidence | Verified on build |
|---|---|---|---|---|
| R1 Member books a free slot | Change 3 | Book room A 10:00-11:00 as member 1 | Recording and exported booking row | Build 15 |
| R2 Overlapping booking refused | Change 3 | Book 10:30-11:30 as member 2 after R1 | Screenshot of the refusal message | Build 15 |
| R3 Member cancels own booking | Change 5 | Cancel R1's booking as member 1 | Recording | Build 15 |
| R4 Fee only for cancellations inside 24 hours | Change 7 | Cancel 3 hours ahead, then 30 hours ahead | Payment test-mode log | Failed on build 15: fee charged at 30 hours |
| R5 Guest sees only the shared booking | Change 8 | Open the share link signed out; try another booking's ID | Two screenshots with URLs visible | Build 14 |
| R6 Admin sees the day's schedule | Change 2 | Sign in as admin and open today | Screenshot | Build 14 |
| R7 Maintenance block cancels and notifies | Change 9 | Block room B next week with 3 bookings | Test inbox showing 3 emails | Build 11 only |
| R8 Members can't read others' booking details | Change 8 | Two-account test | Test log | Build 14 |
| R9 Times shown in the space's local time | Change 10 | Book 09:00 with the browser in another time zone | Screenshot | Build 14 |
| R10 Audit event on every cancellation | Change 5 | Cancel, then read the audit table | Exported rows | Build 14 |
| R11 Room photos | Not built | None | None | Retired in brief v1.2 |
| R12 Month-end export of billable hours | Change 12 | Export a month with seeded bookings | None yet | Never verified |
A failure worth checking
The matrix that is really a to-do list. Every row says Done; there is no evidence column and no build number. It was accurate on the day each row was ticked. Twelve changes later nobody knows which ticks still hold, and the matrix gives a stakeholder more confidence than the app has earned. The counterexample: for a throwaway prototype the matrix is overhead. If the app will be rebuilt rather than extended, a short document of the prototype's limits is the better artifact, and the matrix can start with the rebuild.
Common questions
Isn't this heavy for a small app?
For ten to twenty requirements it is one table, and most of the work (writing acceptance tests) needs doing anyway. The additions are two columns, evidence and build, and they are what let you answer 'how do you know?' without re-running everything.
How is this different from a release evidence checklist?
The release checklist asks whether a release is ready, area by area. The matrix asks whether each promised requirement is met, and keeps asking after every change. The release checklist can point to the matrix as its evidence for the core journey.
What counts as the build?
Whatever identifier the platform gives a deployed version: a version number, a commit hash or a deployment ID. If there isn't one, use the deployment date and time and note that it is a weaker reference.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.