Implementation worksheet · 6 min read
A Project Complexity Estimator Based on Workflows and Roles
Estimate complexity by counting what multiplies work, not screens: roles, workflows, the states each workflow's main record passes through, the role-and-workflow pairs where permissions differ, each direction of each integration, workflows that move money, time-dependent rules, and jobs that run without a user. Give each driver a weight, total them, and compare the total with projects you have already built rather than with a universal scale. The weights here are an illustrative starting point, not a benchmark. The value is in the counting, which makes a 'simple app' visibly not simple before anyone commits to a date or a plan.
Scope: the actor is a founder, product lead or agency lead scoping an app before choosing a builder, a plan or a timeline. The starting state is a written brief with roles, entities and workflows. The boundary is the complexity of the app's behaviour; the estimator does not convert to hours, cost or credits, which depend on the tool and the people. The outcome is a points total with the counts written beside it, used to compare candidate scopes and to see where complexity concentrates.
Put it into practice
1. Count roles, then the pairs where permissions differ
A role is a group of users who see or can do different things. Count them. Then, for each workflow, count the roles whose permissions differ within it. Two roles that can do everything everywhere are less work than two roles whose rights change at each step, and the pair count is what captures that difference.
2. Count workflows and the states in each
A workflow is a start, action and end sequence from the brief. For each, count the states its main record moves through: draft, submitted, approved, rejected and archived is five. Every state adds transitions to build and test, and the failure transitions (rejected, expired, cancelled) are the ones a short prompt is most likely to leave unstated.
3. Count integrations by direction
An integration is not one unit of work. Count each direction separately: calls the app makes to the service, and webhooks the service sends back. Each direction has its own failure: the call can time out, and the webhook can arrive twice, late or not at all.
4. Flag money, time and background work
Workflows that charge, refund or pay out; rules that depend on time (expiries, reminders, recurring billing, time zones); and jobs that run with nobody watching (scheduled tasks, imports). Each carries its own testing burden that screens don't reveal, so count each one.
5. Weight, total and compare with your own history
Multiply each count by its weight and add the subtotals. Then put the total beside two or three past projects whose outcome you know. A points total means nothing on its own; it becomes useful when you can say 'this scores about the same as the portal we built in spring, and that took longer than planned'.
6. Look at where the points concentrate
Sort the drivers by subtotal. A project whose points sit mostly in one workflow with money and an integration carries a different risk from one spread thinly across many simple workflows. The first suggests building and testing that workflow first; the second suggests trimming scope.
7. Recount every time the scope changes
When a feature is requested mid-build, count what it adds before accepting it. 'Let clients approve invoices' adds a role-and-workflow pair, two states and a notification. The recount makes that visible; the one-line request doesn't.
8. Worked example (illustrative, synthetic data)
Two candidate scopes for a fictional equipment-rental business, using the illustrative weights in the table. Scope A: 2 roles, 4 workflows, 11 states, 3 differing pairs, no integration, no money, 1 time rule (overdue reminders) and 1 background job: 6 + 16 + 11 + 6 + 0 + 0 + 3 + 4 = 46. Scope B adds online deposits and a staff role: 3 roles, 5 workflows, 16 states, 7 differing pairs, 2 integration directions (charge call and payment webhook), 1 money workflow, 2 time rules (overdue reminders, deposit refund deadline) and 2 background jobs: 9 + 20 + 16 + 14 + 10 + 6 + 6 + 8 = 89. Of the 43 extra points, the payment integration and money handling account for 16, and the third role with its differing permissions for 11. That says where review and testing effort will go; it doesn't say how many days B takes.
Complexity estimator worksheet (worked example: scope B)
Copy this structure into your review document and record your observed result for each row.
| Driver | How to count | Weight (illustrative) | Count | Subtotal |
|---|---|---|---|---|
| Roles | Groups of users who see or can do different things | 3 | 3 | 9 |
| Workflows | Start, action, end sequences in the brief | 4 | 5 | 20 |
| Record states | States each workflow's main record passes through, summed | 1 | 16 | 16 |
| Differing role-and-workflow pairs | Pairs where a role's permissions differ within a workflow | 2 | 7 | 14 |
| Integration directions | Each outbound call type and each inbound webhook, per service | 5 | 2 | 10 |
| Money-moving workflows | Workflows that charge, refund or pay out | 6 | 1 | 6 |
| Time-dependent rules | Expiries, reminders, recurring schedules, time-zone rules | 3 | 2 | 6 |
| Background jobs | Anything that runs without a user present | 4 | 2 | 8 |
| Total | Sum of subtotals, compared with your past projects | Not weighted | 8 drivers | 89 |
A failure worth checking
The estimate that counted screens. A brief with eight screens gets called small because eight screens sounds small. Three of those screens belong to one workflow with a payment, a webhook, a refund deadline and a role that can approve only below a threshold, and counting screens hides every one of them. The limit of the method is just as real: points compare scopes, they don't convert to days. Two projects with the same total can take very different effort if one reuses an integration you have built before and the other doesn't, and the worksheet has no row for familiarity. Keep the counts next to the total so anyone checking can see what was counted.
Common questions
Where do the weights come from?
They are illustrative, set so that drivers with their own failure modes (integrations, money) count for more than a single state. They are not derived from measured data. Adjust them after two or three of your own projects: if integrations keep taking more effort than their share of the points suggests, raise their weight.
Can this predict what building with an AI builder will cost?
Not directly. Cost depends on the tool's pricing model and on how many build-and-revise cycles you need, which the fixed-brief cost comparison covers. The complexity total helps you check that two scopes you are pricing are actually comparable.
What counts as a separate workflow?
A sequence with its own trigger and its own end state. 'Create booking' and 'cancel booking' are two workflows. 'Create booking from the calendar' and 'create booking from the list' are one workflow with two entry points.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.