Project analysis

Test the whole workflow before adding capacity

GOV.UK’s email migration shows why a dependency’s capacity and a service’s end-to-end behaviour need separate investigation.

An analysis of a publicly documented project. Sources are linked in the article.

The published account

During a move to GOV.UK Notify, the GOV.UK team examined an email workflow from detecting a content change through identifying subscribers, generating messages, and delivering them. Load testing exposed bottlenecks in message generation and subscriber matching, as well as an assumption in rate limiting. The team used profiling and operational measurements to choose improvements. GOV.UK’s May 2018 engineering account.

This article analyses the published work. BE A VIKING was not involved. The exercise below is our own suggested application.

Our interpretation

A dependency can meet its advertised responsibility while the complete task remains too slow. A customer experiences the time from asking for something to receiving a useful result. They do not get much comfort from an efficient component in the middle of an unfinished workflow.

Before proposing more workers or a faster server, sketch one actual task from its starting event to its observable completion. Include time spent waiting for another process, permission, response, or retry. Mark where the team currently has evidence and where it is guessing.

Use that sketch to choose the next measurement. An elaborate dashboard is unnecessary if a temporary trace can answer the immediate question responsibly. Conversely, a single average may hide the waiting that affects the people you are trying to help. Choose evidence to support the decision, and give any temporary instrumentation a removal or maintenance plan.

Keep the acceptance condition human

Consider a hypothetical report-generation service. A submission response tells the user that their request was accepted. It does not mean the report is ready, correct, or accessible to the intended recipient.

Describe a complete result before tuning the implementation: the authorised person can retrieve the right report, understands its status while waiting, and receives a useful explanation if it cannot be produced. These conditions give design, quality, security, and engineering colleagues something shared to evaluate.

Then investigate a representative mix of tasks. Include work that takes a different route through the system, not just several copies of the easiest case. Keep the test isolated from real recipients or otherwise establish explicit controls against unintended delivery. Record what the test environment does not represent, so the result is not mistaken for broader evidence.

Change one decision with the findings

A load test should leave the crew able to choose something. The choice might be to change a query, adjust scheduling, limit the first release, investigate another constraint, or make no infrastructure change at all.

State the competing explanations before running the test. If the concern is time spent creating work, measure that separately from time spent delivering it. If the concern is contention, make the competing work visible. This helps the team distinguish an improvement in the observed symptom from a change in the condition producing it.

Keep the results with the decision they informed. Record enough context for someone else to repeat the investigation, including workload shape, relevant limits, and known gaps. Avoid turning one successful test run into a promise that every future workload will behave the same way.

For your next performance task, write down the user’s start and finish before opening the infrastructure configuration. Then identify one missing observation that could change the proposed investment. Investigate that first.

Use the expedition brief Back to The Logbook

Copy the text manually

Your browser could not copy this automatically. Select the text below and use your device’s copy command.