Greta.sh

Implementation worksheet · 6 min read

A Requirements Traceability Matrix for an AI-Built App

A traceability matrix has one row per requirement in the brief, linking it to where it was built, the acceptance test that checks it, the evidence that the test passed, and the build that evidence came from. For an AI-built app the last column matters most: a later change can alter code that satisfied an earlier requirement, so 'verified' only means something next to a build identifier. Keep requirement IDs stable, record evidence as artifacts rather than ticks, re-test the rows a change could touch, and treat any row without evidence on the current build as unverified, whatever its status once said.

Scope: the actor is a founder, product lead or agency delivering an AI-built app to a client or stakeholder who asks 'is everything in the brief done, and how do you know?'. The starting state is a versioned brief with numbered requirements and a deployed build. The boundary is the functional, permission and data requirements in the brief; performance, security and accessibility have their own tests, which can appear as rows that point to those results. The outcome is a matrix a reviewer can audit row by row without asking the builder.

Put it into practice

1. Number the requirements and never renumber

R1 to Rn, each a single testable statement. Split compound requirements ('members can book and cancel rooms') into separate IDs. A retired requirement keeps its ID and is marked retired, so older evidence still points at something.

2. Record where each requirement was built

The change request, prompt or commit that implemented it. In an AI builder that may be a chat turn or a version in the project's history; record whatever identifier lets someone find it again. Later, this is what shows which requirements a new change could have affected.

3. Write one acceptance test per requirement

State the exact input and the expected result: 'Book room A 10:00-11:00 as member 1, then 10:30-11:30 as member 2: the second booking is refused with a message naming the conflict.' A requirement without a test can only ever be marked 'believed done'.

4. Attach evidence, not ticks

A screenshot with the URL and time visible, a short screen recording, a test-run log or an exported record. A tick in a status column is a claim; an artifact is evidence someone else can check without taking your word for it.

5. Stamp each result with the build it ran on

Record the build or version identifier. A pass on yesterday's preview doesn't cover today's production build, and the matrix has to say which one was tested.

6. Re-test the rows a change could touch

Before accepting a change, note which rows share its area: the same screen, table or workflow. Re-run those tests once it lands. A change can reach further than its request, as the guide on scoping changes describes, and the matrix tells you what to re-check instead of re-testing everything or nothing.

7. Report by status on the current build

Three counts: verified on the current build, verified only on an earlier build, and never verified. Watch the middle one. It grows with every change and only shrinks when someone re-tests, so it is the honest measure of how much of the brief is currently proven.

8. Worked example (illustrative, synthetic data)

A room-booking app with twelve requirements, one of them retired. After build 14: nine verified on build 14, one verified only on build 11 (R7, maintenance block), one never verified (R12, month-end export, because the test data never crossed a month end). Change 15 edits the booking form, and the rows sharing that area are R1 to R4. After build 15, R1 to R3 pass again and R4 fails: the late-cancellation fee is now charged 30 hours ahead instead of only inside 24 hours. The report after build 15 reads: 3 verified on build 15, 1 failed on build 15, 6 verified only on earlier builds, 1 never verified. Without the matrix, R4's last recorded status would still read 'done'.

Requirements traceability matrix (worked example)

Copy this structure into your review document and record your observed result for each row.

Requirements traceability matrix (worked example)
RequirementBuilt inAcceptance testEvidenceVerified on build
R1 Member books a free slotChange 3Book room A 10:00-11:00 as member 1Recording and exported booking rowBuild 15
R2 Overlapping booking refusedChange 3Book 10:30-11:30 as member 2 after R1Screenshot of the refusal messageBuild 15
R3 Member cancels own bookingChange 5Cancel R1's booking as member 1RecordingBuild 15
R4 Fee only for cancellations inside 24 hoursChange 7Cancel 3 hours ahead, then 30 hours aheadPayment test-mode logFailed on build 15: fee charged at 30 hours
R5 Guest sees only the shared bookingChange 8Open the share link signed out; try another booking's IDTwo screenshots with URLs visibleBuild 14
R6 Admin sees the day's scheduleChange 2Sign in as admin and open todayScreenshotBuild 14
R7 Maintenance block cancels and notifiesChange 9Block room B next week with 3 bookingsTest inbox showing 3 emailsBuild 11 only
R8 Members can't read others' booking detailsChange 8Two-account testTest logBuild 14
R9 Times shown in the space's local timeChange 10Book 09:00 with the browser in another time zoneScreenshotBuild 14
R10 Audit event on every cancellationChange 5Cancel, then read the audit tableExported rowsBuild 14
R11 Room photosNot builtNoneNoneRetired in brief v1.2
R12 Month-end export of billable hoursChange 12Export a month with seeded bookingsNone yetNever verified

A failure worth checking

The matrix that is really a to-do list. Every row says Done; there is no evidence column and no build number. It was accurate on the day each row was ticked. Twelve changes later nobody knows which ticks still hold, and the matrix gives a stakeholder more confidence than the app has earned. The counterexample: for a throwaway prototype the matrix is overhead. If the app will be rebuilt rather than extended, a short document of the prototype's limits is the better artifact, and the matrix can start with the rebuild.

Common questions

Isn't this heavy for a small app?

For ten to twenty requirements it is one table, and most of the work (writing acceptance tests) needs doing anyway. The additions are two columns, evidence and build, and they are what let you answer 'how do you know?' without re-running everything.

How is this different from a release evidence checklist?

The release checklist asks whether a release is ready, area by area. The matrix asks whether each promised requirement is met, and keeps asking after every change. The release checklist can point to the matrix as its evidence for the core journey.

What counts as the build?

Whatever identifier the platform gives a deployed version: a version number, a commit hash or a deployment ID. If there isn't one, use the deployment date and time and note that it is a weaker reference.

Basis and scope

This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.

Continue with Greta.sh

Explore Greta →