Small unresolved intervals contrast with a larger coherent ivory optical region.

AI and Automation

Is AI Saving Time Your Business Can Actually Use?

Count time saved only when the team can put it to useful work.

MT BYTES6 min read
Read the perspective

Separate staff effort from customer waiting time

One clock measures active effort: the time a person spends reading, preparing, checking and completing a task. The other measures elapsed time: how long the customer or colleague waits for the result. AI can improve the first while making little difference to the second.

Take a hypothetical business preparing service quotations. An assistant reduces the time spent drafting each quotation, but a manager still approves them in one batch at the end of the day. The team has gained working capacity. Customers may receive their quotations at exactly the same time as before. Both observations can be true.

The distinction matters because the appropriate next step changes. If the objective is to free staff for other work, the pilot may already be helping. If the objective is a quicker customer response, the approval arrangement needs attention. Treating both objectives as “productivity” obscures the real result.

Time saved inside a task becomes business value only when the surrounding process can use it. That use might be shorter waiting time, more capacity, better preparation or less overtime. Name it before the trial begins so that the project has a commercial purpose beyond a favourable demonstration.

Time saved inside a task becomes business value only when the surrounding process can use it.

Choose an endpoint the recipient would recognise

A generated answer is an output. A resolved enquiry is an outcome. A completed form is an output. An application ready for assessment may be the outcome the next team needs. The measurement boundary should end where the recipient can move forward.

For customer support, distinguish an acknowledgement from a substantive response and a substantive response from resolution. An automated message can make the first-response figure look excellent while adding no information. A technically complete answer may still require correction when it overlooks the customer's particular circumstances.

For internal work, define what makes an output usable. A summary may need links to its source material. A draft proposal may need verified scope and pricing. An extracted record may need validation against required fields before another system accepts it.

Write the endpoint in operational language: “the customer has received a checked quotation that they can accept” is easier to evaluate than “quotation generation improved”. Then identify the events needed to measure it. Arrival time, draft readiness, review completion and customer delivery can reveal where work still waits.

Preserve these distinctions in reports. Combining all response events into a single average hides whether the system accelerates substantive progress or simply produces an earlier message.

Include untidy work in the baseline

A trial built from tidy examples says little about the workload a business receives. Include short requests, complicated requests, missing information and cases handled by staff with different experience. Document the mix before comparing results.

A field study of generative AI in customer support found differing effects across worker experience and skill levels. The relevant lesson for measurement is to inspect variation. An average can conceal substantial help for one group and extra checking for another.

Where possible, compare similar work over the same period or introduce the tool in a way that preserves a credible comparison. If the business is too small for a formal experiment, keep a structured case log and be candid about uncertainty. Avoid comparing a busy seasonal month with a quiet week and attributing the entire difference to AI.

Record changes that occur alongside the trial. A new policy, an experienced hire, a revised form or a reduced backlog may affect the result. These observations do not invalidate a pilot; they prevent the organisation from crediting the tool for everything that improved.

Use medians and slower-case measures where they reveal more than an average. A handful of requests stuck for days can matter more to customers than a modest improvement across requests that were already handled promptly.

Quality belongs inside the speed measure

A response that arrives quickly and needs correction can create more work than a slower, accurate one. Count rework through to completion, including the colleague who repairs the answer and the customer who supplies information twice.

Define a small set of quality failures that matter to the task. For a quotation, these might include unsupported pricing, omitted requirements or an incorrect commitment. For document extraction, they may be missing mandatory fields or values assigned to the wrong customer. Separate material errors from cosmetic edits so the review produces a meaningful account.

Review effort also needs measurement. Staff may describe a draft as helpful while spending substantial time verifying every factual statement. That effort may still be worthwhile, but it belongs in the calculation.

Human control is part of the design. Microsoft's human-AI interaction guidance addresses correction and user control. In a pilot, reviewers should be able to reject or amend a result without fighting the interface.

Track reopened work and escalation as well as first-pass acceptance. A high acceptance rate can be misleading when staff approve quickly and problems emerge later. The measurement window should be long enough to see whether completed work stays complete.

Account separately for capacity, cash and experience

Freed time is potential capacity. It becomes a financial saving only under particular operating conditions, such as reduced paid overtime or an avoided expense the business can substantiate. It may instead allow staff to follow up enquiries, improve quality or manage a growing workload.

Do not convert every saved minute into salary savings and present the result as cash returned. A team member who completes a recurring task sooner still has a working day. Explain what happens to the available time and whether the organisation can put it to productive use.

The cost side should include licences, usage, integrations, review, support and maintenance. Reference material changes. Workflows need adjustments. Someone investigates unusual failures. A narrow pilot can reveal these costs before they become embedded in an operating budget.

There is a technical trade-off too. Anthropic's guidance on building effective agents discusses starting with simpler solutions and considering latency and cost. More elaborate processing must earn its place through a better task outcome.

A business case can therefore report three separate results: the change in staff effort, the change in elapsed time for the recipient and the change in quality. Attach costs and commercial consequences to each only where the evidence supports them.

For an enquiry-driven business, link operational measures to later commercial outcomes cautiously. Faster quotations may coincide with more accepted work, but price, demand and lead quality also matter. Retain the individual case history so the team can examine those influences instead of assuming that a shorter response caused every additional sale.

Use the result to change the process

Imagine the pilot produces faster drafts, stable quality and no improvement in customer waiting time. That is a clear finding. Investigate the review queue, responsibility for approval and the availability of information needed to finish the request. Replacing the model may have little effect on the remaining delay.

A different result may show faster responses but rising corrections. Tighten scope, improve reference material or restore review at the point where errors enter. Another may reveal that experienced staff gain little while newer colleagues benefit substantially. Deployment and training can reflect that pattern.

Agree in advance who will read the results and what action each plausible outcome would support. Continue, narrow, redesign and stop should all remain available. Measurement is valuable when it influences how the system operates.

An AI automation engagement should leave the organisation able to explain the change in ordinary business terms. The strongest account is specific: which work improved, for whom, by what measure and with what remaining cost. That account is more durable than a headline about how quickly a tool can generate text.

MT
MT BYTES

Perspectives on technology and business.

Explore perspectives

Build a measurement plan into the pilot

MT BYTES can help define the baseline, quality checks and process measures for an AI automation project. We can connect the technical trial to the outcome your team is responsible for delivering.

Discuss your project