Greta.sh

Implementation worksheet · 6 min read

A Monthly Release Review Template for Small Product Teams

Once a month, review every release the app shipped as a list, not a feeling: what was requested, what was delivered, what broke afterwards, what was reverted or hotfixed, and which requirements were re-verified. Then decide four things: which kinds of change caused rework and what rule or test prevents it, which invariants broke more than once, what carries into next month with an owner, and how many changes next month allows. The output is decisions with owners, not a changelog. For a team changing an app through an AI builder, the most useful line is requested versus delivered, because a change that did more than it was asked is the one whose side effects nobody tested.

Scope: the actor is a small product team of one to five people shipping changes to an AI-built app through a builder, plus whoever handles support. The starting state is a month of releases, a change history, error tracking and a support inbox. The boundary is releases and their consequences; roadmap planning and experiment reviews are separate meetings. The outcome is a one-page record of decisions and owners, kept alongside earlier months so patterns show up.

Put it into practice

1. Build the release list before the meeting

Every release in the month with its identifier, its date and the request that produced it, taken from the builder's version history or your deployment log. If the list can't be built, that is the first finding: releases you can't list are releases you can't confidently roll back to.

2. Compare what was requested with what was delivered

For each release, one line on what was asked and one on what changed, with the size of the change where you can see it. Mark every release that changed more than it was asked to.

3. Trace every regression to its release

For each bug this month that broke something that used to work, find the release that introduced it and the invariant it broke. Two regressions against the same invariant is a pattern that deserves a named test run on every related change.

4. Count reverts and hotfixes, then read their causes

A revert means a release reached users and had to be undone; a hotfix means it had to be patched in a hurry. The counts aren't a score. The causes are the material for the month's decisions.

5. Group support messages and errors by release

Put support messages and new error types against the release after which they started. A release followed by a burst of one kind of question deserves a close read even if nobody filed a bug.

6. Update the requirement and limits records

Which requirements were re-verified this month, and which are now verified only on an older build? Which known limits were lifted, and did any new ones appear? The review is the natural moment to bring the traceability matrix and the limits register up to date.

7. Set next month's change budget and owners

Decide how many changes the team will ship, what must be re-tested with each, and whether any area stays frozen until its regressions stop. Every open issue leaves the room with an owner and a date.

8. Worked example (illustrative, synthetic data)

A fictional three-person team running a client-portal app reviews September. 14 releases. 3 changed more than requested; one also reformatted the invoice table and broke its sorting. 4 regressions, 2 of them against 'clients only see their own projects', both from releases that touched shared query code. 1 revert and 2 hotfixes. 11 support messages, 6 of them after release 9. 5 requirements now verified only on an August build. Decisions: run the two-account permission test on every release that touches queries (owner: developer); freeze the invoice area for two weeks; allow 10 releases next month; re-verify the 5 stale requirements by mid-month (owner: product lead).

Monthly release review template

Copy this structure into your review document and record your observed result for each row.

Monthly release review template
SectionQuestionEvidence sourceThis month (synthetic example)Decision and owner
Release listWhat shipped, when, from which request?Version history or deployment log14 releasesNo change needed
Requested vs deliveredWhich releases changed more than they were asked to?Change requests and diffs3, including an invoice table reformatName the flows to keep unchanged in every request; product lead
RegressionsWhat broke that used to work, and from which release?Bug reports traced to releases4; 2 against client isolationTwo-account test on every query change; developer
Reverts and hotfixesWhat had to be undone or patched in a hurry?Deployment log1 revert, 2 hotfixesFreeze the invoice area for two weeks; team
Support by releaseDid a release change what users asked about?Support inbox grouped by date6 of 11 messages after release 9Rewrite release 9's new screen copy; designer
Errors by releaseDid a release introduce new error types?Error tracking2 new error types after release 12Fix both before new features; developer
RequirementsWhich requirements are verified only on old builds?Traceability matrix5 verified only on an August buildRe-verify by mid-month; product lead
Known limitsWhich limits were lifted or added?Limits register1 lifted (CSV export), 1 added (upload size)Update the register; product lead
Open issuesWhat carries into next month?Issue list6 openEach given an owner and a date
Next month's budgetHow many changes, and what must each re-test?This review10 releases; query changes re-run the permission testAgreed by the team

A failure worth checking

The review that counts releases as output. The team reports '14 releases shipped' as the month's headline and nobody asks what those releases broke. Next month they ship 20, regressions rise with them, and the review celebrates again. A release count without regressions, reverts and requested-versus-delivered beside it measures activity, not progress. The counterexample: a review that turns into blame over individual requests or people. The question is which kind of change causes rework and which test or rule prevents it, not who wrote the request.

Common questions

How long should the review take?

About an hour for a small team, if the release list and evidence are gathered beforehand. If building the list takes longer than the meeting, fix the release log first.

What if we only shipped two releases this month?

Hold it anyway, briefly. Two releases can still carry a regression, and keeping requirements and limits current doesn't depend on volume.

Is this the same as a sprint retrospective?

It overlaps, but it starts from different material. A retrospective starts from how the work went for the team; this review starts from the list of what reached users and what happened next. Some teams run both.

Basis and scope

This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.

Continue with Greta.sh

Explore Greta →