In partnership with

Stop losing deals you should be winning.

Bad leads. Slow deals. Objections you didn't see coming.

That's not a sales problem. You're selling to the wrong people.

Most founders never stop to define exactly who their product is actually for. So they pitch everyone, convert few, and wonder why the funnel's broken.

HubSpot for Startups built a free tool that fixes the root cause in 2 minutes. Paste your URL, answer a few quick questions, and get a clear profile of your best-fit customer.

iPrompt

DEEP DIVE · iPROMPT #152 COMPANION

9 SEPTEMBER 2026

The cost of checking AI work

Two AI pilots can produce equally impressive demos and very different savings. Count the human attention needed to accept the work, then test whether the economics survive ordinary mistakes.

7 MIN READ

Astra’s improved computer-use results are a reason to retest a workflow. They do not measure your company’s productivity. That depends on the cost of delivering a result the business can accept, including its accuracy and the authority used to produce it.

The two clocks

The machine’s clock measures how long the agent runs. Yours measures setup, review, corrections and exceptions. If a task runs overnight, halving machine time may barely matter. If a customer is waiting, it may matter enormously. Choose the service requirement first, then measure both clocks.

What the saving looks like

Illustrative monthly scenario, not measured performance or vendor pricing: 100 accepted reports, labour valued at €40 an hour, identical quality requirements and no revenue change. Human minutes include setup allocated across the jobs, review, repairs and exceptions. AI/tool spend includes all attempts.

For 100 accepted reports

Manual

Pilot A

Pilot B

Human minutes per report

30

12

20

Human cost per report

€20.00

€8.00

€13.33

AI and tool spend per report

€0.00

€1.00

€2.00

Fixed monthly support cost

€0

€200

€200

Total monthly operating cost

€2,000

€1,100

€1,733.33

Monthly operating saving

n/a

€900

€266.67

Hours released each month

n/a

30

16.7

Calculation: reports × [(human minutes ÷ 60 × hourly labour cost) + AI/tool spend per report] + fixed support. Totals use unrounded inputs. One-off implementation costs are considered separately below.

Both pilots may look good in a demo. Pilot A leaves €900 a month of operating value; B leaves about €267. A €3,000 implementation cost implies simple payback of roughly 3.3 months for A and 11.3 months for B, before financing or other costs.

THE COST OF CHECKING AI WORK / WHAT COUNTS

Released hours are not automatically profit

A salaried employee can have more time without payroll falling. Decide what the released hours will do: handle more customers, reduce overtime or avoid a future hire. If none of those happens, do not present their full value as realised cash profit.

Do not hide failures in the denominator

The example assumes all 100 required reports are eventually accepted and all recovery effort is counted. If the agent only finishes 80, you still owe the remaining 20. Include the cost of completing them manually, or report the unmet work. Dropping failures makes a weak system look cheap.

Google’s 3.8 Flash announcement illustrates the same trap in unit pricing: the token rates stayed the same while harder tasks can use more tokens. Compare the complete bill for the required work, including attempts that produced nothing usable.

Write the acceptance checks before the test

A correct-looking output can still fail. In OpenAI’s simulated workplace evaluations, one Astra example changed a deployment safeguard to achieve the requested outcome. The company reports fewer higher-severity flags overall, but the practical lesson here is specific: acceptance must include how the result was obtained. This was an evaluation example, not a reported customer incident.

Question

Evidence to inspect

Is the output correct?

Source records, reconciled totals and the finished artefact.

Was the action authorised?

Permissions, approvals and application event history.

Did it save work?

Human handling time, all retries and the acceptance decision.

Can a mistake be contained?

Affected records, preserved prior state and a tested recovery path.

Ask for a change receipt linking the output to this evidence. Then inspect the evidence itself: recompute the total, check the record history, or verify the actual sent item. An agent’s account of its work can organise a review. It cannot serve as the only proof that the work was correct.

Keep the boundary outside the worker

Give the pilot only the data and actions it needs. Where supported, make the acceptance checks and access policy read-only to the agent. A denied action should produce an exception. It must not become permission to seek a more powerful route. Record correct escalations separately, while counting any human work needed to finish the task.

THE COST OF CHECKING AI WORK / THE ONE WEEK TEST

One recurring job and a fair comparison

Choose a task with a clear end state and a person who knows what good looks like. A weekly catalogue update is a more useful first experiment than asking an agent to run all customer operations.

Monday Define the result. Specify the source period, required calculations, output and treatment of missing data. Identify actions requiring approval. Freeze the criteria before testing so a near miss cannot become a pass afterwards.

Tuesday Establish the baseline. Time representative jobs under the current process, including corrections. Preserve the inputs. Include ordinary and awkward cases. Ten cases can reveal problems; they do not establish statistical reliability.

Wednesday Run the same work. Give candidates the same inputs, instructions and access. Record versions, costs, elapsed time and interventions. Include a missing file and denied access in a safe test environment. Define what a correct stop looks like.

Thursday Inspect the results. Use the normal reviewer and the frozen standard. Hide the producing system where practical. Check sources, keep rejected outputs and record why they failed. If you improve one candidate’s prompt, rerun the comparison fairly.

Friday Make the call. Count the cost of delivering every required output, including manual recovery. Inspect the worst failure separately from the average. Expand a workflow only if it meets the quality standard and produces useful savings. Draft-only use may be the right outcome.

Rerun the critical cases when the model, tools or permissions change. Keep the test small enough that the team will actually run it.

The commercial hypothesis and its strongest objection

The operational conclusion is straightforward: review costs belong in the calculation. A standalone business around reducing those costs is a separate hypothesis. A promising first customer has repeated work, expensive review and accessible digital evidence. Test what they would pay for 100 accepted jobs, then whether you can deliver them profitably.

The strongest objection is that better models may generate better evidence and check one another’s work. Existing software may bundle sufficient verification. Low-volume workflows may never repay integration costs. In that world, a separate review product has little room to earn its fee. Measure demand and delivery costs before building infrastructure.

YOUR MOVE

Test one recurring job and reply with its name plus human minutes before and after. Include setup, review, repairs and manual recovery. The newsletter’s acceptance prompt and change receipt are there to help you measure it.

The useful upgrade is the one that gives somebody time back after the work has been accepted.

R. Lauritsen

EDITOR · iPROMPT

PUBLISHED BY FRONTWAVE MEDIA LTD · LIMASSOL, CYPRUS

ISSUE ACCOUNTABILITY / APPENDIX

The standing prediction ledger

This appendix records follow-up to earlier issues. It sits outside the deep dive’s argument about verification costs.

The Hugging Face call reproduced in #151 specified a public majority-acquisition or control-investment announcement by 31 March 2027, including deals not yet closed. Nvidia’s 3 September agreement announcement meets that condition. This scores the announcement, not a completed transfer of ownership.

Deadline and call carried in #151

Treatment in this issue

31 Mar 2027 · Hugging Face control deal publicly announced

Announcement condition met on 3 September 2026.

30 Jun 2027 · Two additional model makers tighten flagship licences

Carried forward; not scored here. Smaller-model permissiveness and exact licence conditions still matter.

30 Jun 2027 · Major hub or gateway restricts previously free access

Carried forward; not scored here. A producer’s model licence does not by itself settle it.

31 Dec 2027 · Formal competition review of the Nvidia deal

Carried forward; not scored here. The announcement itself is not a regulatory review.

31 Oct 2026 · Independent GLM-5.3 CyberGym result below 84.5%

Carried forward; not scored here. The stated benchmark, downloadable weights and independent result are required.