A software architecture review checks whether the system can support the workflows, failure conditions and changes the business actually needs. A useful review ends with evidence, decisions and owners. A diagram alone cannot establish that a backup restores, a queued field action arrives, or a release can be rolled back.
This checklist is for a team preparing a review and a buyer deciding what to commission. Use it on one important workflow first. Expand across the system once the team agrees what “working” means. It is a planning aid, not a security certification or a substitute for testing a live environment.
Download the architecture review checklist (CSV). It opens in a spreadsheet and includes columns for evidence, consequence, owner and next action. No email address is required.
Collect the inputs before reviewing the design
Start with the current system, not the intended system. Ask for a component map, repositories, deployment instructions, integration contracts and a sample of recent incidents. Identify who can explain each external dependency and who has authority to change it. Record missing access as a finding rather than filling the gap with assumptions.
Then select a business journey. For a field service platform this could be a technician receiving a job, completing a survey and making the result available in the client portal. For a safety app it could be a person recording an alarm without connectivity. Following the action from start to finish exposes boundaries that a list of technologies hides.
- What action starts the journey, and who is allowed to perform it?
- What counts as accepted, completed and visible to the next person?
- What evidence can demonstrate the current behavior?
- What failure would cause the greatest operational harm?
- Which decision does the team need the review to resolve?
Keep requirements specific. “Highly scalable” is difficult to review. “Support the expected number of simultaneous field submissions without losing accepted work” leads to a workload, a test and a recovery question. A review can help refine an unknown requirement; it should make the uncertainty explicit.
1. Follow workflow and component boundaries
Trace the action through the mobile or web client, API, storage, background processing and external services. At each boundary, ask which component owns the state and how another component learns that it changed. If two services both believe they own the same business decision, document how disagreement is resolved.
The UAE field-service ERP case joins field mobile workflows, an enterprise ERP and a client portal for a network of more than 7,000 ATMs. The published scope establishes why those boundaries matter; it does not prove a particular message bus or database design. A review of a similar system should inspect the actual route from field evidence to client reporting.
Check where business rules live. If the mobile client and the server each calculate approval eligibility, compare their behavior when the rules change. If a reporting job reads directly from operational tables, ask how schema changes reach its owner. The issue is the consequence of a change, not whether the architecture uses a fashionable pattern.
2. Establish data and integration ownership
List the important records and the authority for each. A customer identifier, appointment, payment or field incident may be represented in multiple systems. That can be necessary, but the review must explain which copy is authoritative, how updates propagate and how a failed synchronization is discovered.
For each integration, capture the contract, credentials owner, timeout behavior and reconciliation method. A successful HTTP response may mean “received for processing,” not “business action completed.” Make that distinction visible to both product designers and operators. Test what happens when the external system accepts a request but its response never reaches your server.
Include schema evolution and deletion. An added optional field is different from changing the meaning of an existing field. Removing a record from the main database is different from removing copies from logs, exports and devices. These questions deserve a named owner and a demonstrable procedure.
3. Review failure, retry and recovery
Inspect the moments when the system can become uncertain: after a local save, during a network request, between a database update and a notification, and during a deployment. Ask the team to explain the user-visible state at each point. “Try again” is incomplete if a retry can duplicate an important action.
AWS's discussion of idempotent APIs describes why retries need a way to identify the caller's original intent. For a review, the practical question is whether repeating a request produces another business action or returns the status of the existing one. Verify this against the implementation.
The duty-of-care platform includes an offline panic queue, retry when connectivity returns and escalation mail after delivery. Its relevance to a review is the need to distinguish a locally captured action from a remotely delivered one. Offline capture cannot guarantee remote receipt while the device has no connection.
Ask for a restore exercise, not just a backup setting. What data was restored, how long did it take, and could the application operate afterward? If the team has never performed the exercise, the finding is “recovery unverified.” The next step is an appropriately isolated test, not an assertion that the current backup is useless.
4. Check deployment and operational ownership
A system is maintainable when someone other than its original builder can deploy it, observe it and recover it. Inspect the release path from source to production. Identify environment configuration, secrets, database migrations and third-party callbacks that cannot be understood from the repository alone.
Discuss rollback precisely. Returning to an earlier application version may be possible while reversing a data migration is not. The review should identify compatible release sequences, recovery options and the point beyond which a rollback needs a separate data decision.
Choose one recent incident and follow how it was detected. Which signal changed? Who saw it? Could that person tell whether users were affected? Logs that exist only on an engineer's machine cannot support a distributed operating team. Alerts should point toward an action and an owner, with enough context to investigate without exposing unnecessary personal data.
5. Test performance and cost assumptions
Write down the workload before comparing architecture options. Include peak activity, background jobs, data growth and external service limits. Averages can hide the exact period that matters, such as everyone checking in after an alert or submitting reports at the end of a shift.
Record which measurements are from production, which come from a test and which are estimates. A small proof of concept can establish feasibility; it cannot establish every behavior at full deployment scale. The video analytics engagement covered architecture for 1,400 cameras and a six-camera live proof of concept. Those are different evidence scopes.
Cost review should include the operating model. Compute, storage and model calls are visible charges; device replacement, support, manual review and deployment effort also matter. Use the video analytics TCO calculator when comparing that particular workload, replacing its assumptions with your own quotes.
A compact review table
| Area | Evidence to request | Decision it supports |
|---|---|---|
| Workflow | A trace of one accepted action through all components | Where completion is established |
| Data | Record ownership and synchronization behavior | Which copy wins a disagreement |
| Retry | A repeated request after a lost response | Whether duplication is controlled |
| Recovery | Results of an isolated restore exercise | What can be recovered and how |
| Release | Deployment and migration sequence | How changes are introduced safely |
| Operations | Incident signals and responsible owners | Who can act when behavior changes |
| Capacity | A defined workload and representative measurements | Which limit matters first |
Turn findings into a short decision record
A finding should state the observation, its consequence, the evidence and the proposed next step. Separate a verified defect from an untested assumption. “The integration may duplicate jobs” is a hypothesis. “Repeating this accepted request created two jobs in the test environment” is an observation with a reproducible check.
Prioritize according to the consequence and the decision blocked. A critical unanswered recovery question can deserve attention before a verified cosmetic defect. Include the owner and what would demonstrate completion. Avoid scores that make uncertain evidence look precise.
A useful closing document contains the workflow map, prioritized findings, unresolved questions and a sequence of bounded changes. If the evidence points toward replacing part of the system, use the refactor versus rewrite guide to compare the transition cost as well as the destination.
What a paid architecture review can resolve
A paid scoping or diagnostic week can establish the review boundary, inspect available evidence and identify the next decisions. It cannot automatically certify every subsystem or resolve every unknown dependency. The scope should specify the inputs, the access available and the deliverables before work begins.
For teams working remotely across the US, UK or other markets, bring a named technical contact and the people who understand the business workflow. Shared written evidence helps the review continue across time zones. Describe the system, the failure or planned change, and the decision you need to make through the software architecture consulting page.