A small compressed violet interval interrupts much broader blue optical depth.

Software Development

Find the Bottleneck Before Scaling Your Application

Measure what is slowing the service before adding capacity.

MT BYTES6 min read
Read the perspective

Describe the failure as a customer task

“The application is slow” is too broad to guide a scaling decision. Establish what users are trying to do, when the difficulty appears and what happens next.

An online order may load quickly until checkout. A report may become slow only when it covers a large date range. An internal system may work normally until several scheduled jobs begin at once. These are different demand patterns.

Define a successful transaction from the user’s perspective. It should include the final business result, such as an accepted order or a completed export, rather than only a response from the first server.

Microsoft’s reliability-target guidance connects service objectives with business expectations. Agree which tasks matter and what acceptable performance means before choosing infrastructure changes.

Collect timing and failure evidence around those tasks. Compare ordinary periods with the periods where users struggle. Record changes in volume, request type, data size and dependencies.

This gives the technical team a question it can investigate: which part of this transaction becomes constrained under this pattern of demand? It is much more useful than a general instruction to make the system scalable.

A scaling design is incomplete until the team understands how the service behaves at its limit.

Follow the transaction through its dependencies

A customer-facing request may pass through an application, database, cache, queue and external service. The slowest or most constrained part can limit the whole transaction.

Measure where time is spent. High processor use suggests a different investigation from a database waiting on locks or an application waiting for a third-party response. Low average processor use does not prove spare end-to-end capacity.

Look for repeated work. The application may request the same data many times, process more records than the task requires or repeat an expensive calculation for every user. Reducing that work may be more effective than increasing the number of servers performing it.

Check concurrency limits too. A database connection pool, external API allowance or shared resource can become saturated even when individual requests are small. Adding more application instances may increase pressure on that constrained dependency.

Queues deserve separate attention. A queue can absorb a temporary burst, but a steadily growing backlog means the service is receiving work faster than it completes it. Determine whether the backlog can clear within the business’s acceptable delay.

Map the relevant boundaries and identify which team or supplier controls each one. A scaling plan that assumes unlimited capacity from an external dependency may fail even if the application changes work exactly as intended.

Understand what happens when capacity runs out

Failure behaviour can turn a local constraint into a wider outage. A slow dependency may cause requests to remain open. Users retry. Automated clients retry too. The additional work increases pressure on a service already struggling.

Google’s SRE chapter on cascading failures describes mechanisms involving overload, reduced capacity and retries. The relevant lesson is to examine what happens beyond the first delay.

Set appropriate time limits and retry behaviour for the task. A retry should not repeat a consequential action without a way to recognise that it has already happened. Creating two orders is not a successful recovery from an uncertain response.

Decide which work can wait and which should be rejected or deferred visibly. An internal report may tolerate a queue. A checkout needs a clear result and a safe way to establish whether the order was accepted.

Protect the service from unbounded work. Limits on request size, concurrency or expensive operations may preserve useful capacity, provided users receive an understandable response.

Also consider partial degradation. The business may prefer to keep essential transactions available while temporarily limiting a secondary feature. That choice should be designed and tested, rather than improvised during overload.

A scaling design is incomplete until the team understands how the service behaves at its limit.

Address the constraint you have measured

Changes to application design and cloud capacity address different constraints. The evidence should determine which one is worth testing.

ConstraintCandidate interventionWhat still needs checking
Repeated expensive computationReuse a valid result or simplify the workFreshness and correctness of reused information
Inefficient data accessImprove queries, indexes or access patternsWrite cost and behaviour with realistic data
Insufficient application capacityAdjust resources or distribute workDownstream capacity and operating cost
Slow external dependencyChange request flow, caching or asynchronous handlingThe customer’s required consistency and timing
Unbounded background workSchedule, limit or prioritise jobsWhether the backlog clears within the needed period

Do not assume that a more distributed architecture is the natural next step. More components can introduce deployment, monitoring and data-consistency work. A simpler application with a corrected query may serve the requirement well.

Likewise, do not reject additional capacity when the workload genuinely needs it. The point of diagnosis is to choose a proportionate intervention, not to avoid infrastructure spending at all costs.

Consider who will operate the changed design. A technically capable architecture that the team cannot monitor or recover can create a new constraint elsewhere in the business.

Test the demand the business actually expects

A useful performance test represents important tasks, data sizes and demand patterns. It should exercise the path that previously became constrained.

Microsoft’s performance-testing guidance emphasises measurable targets and realistic conditions. Write a brief that identifies the workload, the environment, the target and the observations needed.

Include the mix of activity. Ten thousand simple page views do not reproduce a smaller number of expensive searches or report exports. Test realistic records and relevant dependencies, using safe arrangements for payments or other consequential external actions.

Examine sustained demand and bursts where the business expects both. Watch response times, completed transactions, failures, backlog and resource use. An average can hide a group of users experiencing much longer delays, so inspect the spread as well.

Increase load deliberately and observe the point where the service stops meeting its target. The purpose is to understand behaviour and capacity, with safeguards against unintended disruption.

Repeat the relevant test after the proposed change. Confirm that the original constraint improved and that another component did not become the limiting factor. Record the configuration so the result can inform later planning.

Connect capacity planning to cost and change

The resulting evidence should state what the service handled, under which conditions and with what operating cost. It should also identify the next constraint likely to matter as demand changes.

Set monitoring around those conditions. A business may need to watch database saturation, queue delay or external-service failures rather than only server utilisation. Assign someone to review signals and act before customer impact becomes severe.

Revisit the assessment when the product introduces a more expensive task, the data grows materially or a dependency changes. Capacity established for one request profile does not automatically apply to another.

Connect the engineering result with business plans. A campaign, new customer group or additional location may alter the timing and shape of demand. Share those expectations early enough to test the relevant scenario.

The team should be able to state the load the service can handle, the warning that it is approaching a limit and the next change available. That turns the growth plan into an operating decision the business can revisit.

MT
MT BYTES

Perspectives on technology and business.

Explore perspectives

Find the constraint behind a slow or overloaded application

MT BYTES can assess a critical transaction, identify its bottleneck and scope a realistic performance test. Bring the task that slows down, the conditions that trigger it and the demand the business expects.

Discuss your project