In partnership with

For product teams moving at AI speed.

AI makes it easier to ship anything, even bad ideas. The hard part is knowing which ideas are worth building.

Jira Product Discovery brings your ideas, customer insights, and priorities into one place, so your team can decide what to ship and move forward with confidence.

Capture ideas, prioritize with evidence, and build living roadmaps your team can rally around—all while staying connected to delivery in Jira, so everyone can see what’s being built and why.

Better product decisions in the AI era.


D E E P D I V E

ISSUE 30 COMPANION · FRIDAY 25 SEPTEMBER 2026

What a Free AI Agent Costs, and Who Gets Paid First

Two models, kept separate: the platform’s economics per user, and when those turn into hardware orders. Plus a map of who carries a data-centre delay.

Somewhere in a Meta data centre, a virtual machine is still working for a Muse user who closed the app hours ago. Each agent runs on its own machine and can keep going when nobody’s watching. Somebody pays for that.

Halve the cost of every token that agent uses and, in the model below, the platform’s break-even hurdle roughly halves. Whether its chip suppliers then sell more hardware is a different question, with a different model, and it’s the question Monday’s rally skipped.

This companion keeps the two apart. First the platform: what a free agent costs per user and what has to happen for it to pay. Then the supplier: when extra usage becomes new hardware orders. Then the problem neither model can fix, a delivery date nobody controls.

THE ANNOUNCED STARTING POINT

Muse launched on 8 September with a free tier and plans at $20 and $100 a month. Meta’s AI chief, Alexandr Wang, told Axios that most users should manage on the free tier and that subscriptions help cover compute costs. At Connect on 23 September, Zuckerberg added a planned second line: a small fee on purchases the agent completes. Muse runs on Meta’s own Muse Spark models.

A day earlier, Anthropic and OpenAI cut list prices on new models by 20% to 50%. Those are prices for outside customers, not Meta’s own cost. TrendForce cites estimates that some agent workloads need four to forty CPUs per GPU, because each agent runs on its own virtual machine.

Everything in the two models below is invented. They illustrate a method. They are not estimates of Meta’s, Anthropic’s, OpenAI’s or any chipmaker’s economics.

FOUR NUMBERS TO KEEP SEPARATE

Measure

What it tells you

Common mistake

List price per token

What an outside customer pays a model seller.

Treating it as another platform’s cost.

Internal cost per token

What an operator pays to serve its own models: hardware, power, staff.

Assuming it moves with competitors’ list prices.

Compute spend per active user

The platform’s running cost per user.

Tracking downloads instead.

Supplier revenue

Hardware sold, recognised when control passes to the buyer.

Assuming it moves with token volume or running costs.

My starting point: operating spend and hardware orders are linked, but through capacity, not directly. A platform can cut its cost per token on hardware it already owns. Orders come when workload outgrows the fleet, or when the operator upgrades to newer chips.

MODEL ONE: THE PLATFORM

All figures are hypothetical. An average active user runs 5 million tokens a week. The operator’s internal cost is $0.20 per million tokens, and virtual-machine orchestration adds $0.15 per user per week. Four per cent of users pay $20 a month. Users buy $40 a week through the agent, and the platform keeps 1.5%.

Weekly compute spend is therefore $1.15 per user. Revenue is about $0.78: $0.18 from subscriptions ($20 × 12 ÷ 52 × 4%) and $0.60 from fees. That leaves a shortfall after compute of about $0.37 per user per week, roughly $190m a year at 10 million weekly users. Counting subscription income, breaking even needs about $64 of purchases per user per week; without subscriptions it would need about $77.

Scenario

Compute spend

Purchases

Revenue

Contribution after compute

Break-even purchases

A Base case

$1.15

$40

$0.78

−$0.37

$64

B Token cost halves

$0.65

$40

$0.78

+$0.13

$31

C Usage doubles, no extra purchases

$2.15

$40

$0.78

−$1.37

$131

D Usage doubles, purchases double

$2.15

$80

$1.38

−$0.77

$131

E Usage doubles, cost halves, purchases double

$1.15

$80

$1.38

+$0.23

$64

F Base case, fee rises to 3%

$1.15

$40

$1.38

+$0.23

$32

G Usage triples, cost halves, purchases triple

$1.65

$120

$1.98

+$0.33

$98

Per active user per week. Illustrative only; orchestration cost is held fixed per user. Figures rounded.

Productive usage versus expensive activity

C and D use the same compute. The difference is whether the extra usage produced anything anyone paid for. In C, the agent thinks longer and browses more but nobody buys more, and the shortfall nearly quadruples. In D, purchases double with usage, which helps, but not enough at today’s cost.

C is the row that worries me. An agent that spends an hour comparing prices and buys nothing looks like engagement on a download chart and like a cost line in a 10-Q. From the outside, you can’t tell C from D until the company discloses purchases.

E and G show how a free agent can cover its compute: usage grows, efficiency keeps pace, and the extra usage generates purchases. A higher fee (F) is one lever. More commissionable purchases are another, and arguably the more durable one. At 10 million users, G turns a $190m annual shortfall into a positive contribution of about $170m. That’s after the compute costs in the model only, not company profit.

So for any platform, ask two questions about a price cut: how much more will users consume, and how much of that extra consumption turns into revenue?

MODEL TWO: WHEN USAGE BECOMES HARDWARE ORDERS

Chipmakers recognise product revenue when control of the hardware passes to the customer. They aren’t paid per token. Orders come from two places: a fleet that can’t carry the expected workload, and upgrades to newer, more efficient chips.

Hypothetical assumptions: the operator’s fleet runs at 70% utilisation today, and it orders more capacity when projected load would exceed 85%. “Workload” means total tokens served, so it grows with both users and usage per user. “Efficiency” means tokens served per unit of today’s capacity. Capacity needed, relative to today, is 70% × workload multiple ÷ efficiency multiple ÷ 85%.

Case

Workload

Efficiency

Where efficiency comes from

Capacity needed vs today

Hardware orders

Today

×1

×1

n/a

0.82

None; spare capacity

Software gain

×1

×2

Better models, batching, caching

0.41

None; fleet over half idle

Hardware upgrade

×1

×2

New-generation chips

0.41

The upgrade itself is the order

Workload gain only

×2

×1

n/a

1.65

About +65% more capacity

Both double, via software

×2

×2

Software

0.82

None

Workload outruns efficiency

×3

×2

Either

1.24

About +24%, plus any upgrade

Illustrative only. Capacity is measured in today’s processing capacity; after an upgrade, a smaller physical fleet can deliver the same capacity. Excludes replacement timing, hardware price changes, lead times, and savings from power or margins.

Where the efficiency comes from matters as much as its size. A software gain lets the same chips serve more tokens, and nobody buys anything. A hardware gain usually means buying the next generation of chips, and that purchase is itself supplier revenue. So “cheaper tokens” can be bad news for chipmakers or the thing that sells their newest product. The company’s own disclosures on replacement cycles and new-generation deployments tell you which.

Putting the two models side by side

Platform scenario B, cheaper tokens with flat usage, creates no orders if the saving came from software, and an upgrade order if it came from new chips. Scenario E creates none from software either, because workload and efficiency both doubled. Scenario G is where both sides win: the platform covers its compute and the operator needs roughly a quarter more capacity.

So platform and supplier are neither simply opposed nor simply aligned. Their interests meet in G, where workload outgrows efficiency and the extra usage is productive. They diverge elsewhere: a platform gains when efficiency outpaces usage (B), while a supplier can gain from upgrades even when workload is flat. My read: Monday’s rally priced something like G. Before paying for G, I’d want one number from Meta, contribution per active user, and one from the chipmakers, orders separated from deliveries. Neither exists yet.

Timing matters too. Orders are placed months before delivery, so near-term chip revenue largely reflects commitments made before Muse launched. Muse’s effect, if any, shows up in orders and guidance first.

WHO CARRIES THE DELAY

Model two assumes new hardware has somewhere to go. Thursday showed it may not. According to Bloomberg, Oracle sent a force majeure notice to the Blue Owl unit developing Project Jupiter in New Mexico, seeking to delay payments if the site misses its 2028 start. The campus needs about 2.45GW from Bloom Energy fuel cells, fed by a gas pipeline whose construction has slipped to February 2027 pending permits.

Oracle says the project remains on schedule, Blue Owl that its financial commitments are unchanged, and Bloom that it remains committed. No contract terms are public, so the table below is a checklist of questions, not a finding.

Party

Role as reported

What a delay could mean

What to look for

Tenant (Oracle)

Leases the campus to serve customers.

Payments may be deferred; customer delivery slips.

Disclosures on lease start dates and customer commitments.

Developer (Blue Owl unit)

Builds and owns the site.

Carries construction cost longer before rent begins.

Fund reports on the project’s income start date.

Project lenders

About $18bn of related debt, per the FT via CNBC.

Cash to service the debt may arrive later; repayment dates depend on terms.

Trading levels and any covenant waivers.

Power supplier (Bloom)

Fuel-cell microgrid for the site.

Revenue timing moves with the site date.

Backlog timing in its next report.

Gas supplier

Pipeline needing permits.

The first domino; construction already slipped.

Permit decisions and construction start.

A force majeure clause can excuse or postpone obligations when specified events outside a party’s control intervene; whether a permit delay qualifies depends on the wording. The practical signal is simpler. A counterparty is preparing for a delay and wants its cost to sit elsewhere.

In model two, a capacity shortfall becomes an order only when there’s a powered building to put the hardware in. A late site defers the supplier’s delivery and the tenant’s rent. It may also delay the cash flow that lenders rely on, although repayment dates depend on the loan terms, which aren’t public.

What a year’s delay costs

A hypothetical, unrelated to Jupiter’s actual figures. Suppose $10bn of hardware is bought for a site that opens a year late. Financed at 5%, carrying it for that year costs about $500m in interest. There may be a second cost: if newer chips arrive during the wait, the hardware could earn less over its life than planned. Sizing that needs assumptions this example doesn’t make, including lost earning years, resale value and whether the hardware can be used elsewhere. The carrying cost is arithmetic; the ageing cost is a range.

That’s why I read the force majeure notice as a question about who carries the cost of waiting, not about whether AI demand is real.

TURN THE ANALYSIS INTO A DECISION

The bull case is scenario G: agents drive real purchases, efficiency keeps pace, the platform covers its compute and suppliers sell more capacity. Nothing in this week’s news rules that out. The evidence to look for: contribution per active user from platforms, orders separated from deliveries at suppliers, and dated first-rent milestones from capacity owners.

The case weakens when usage rises without purchases (C), or when the capacity to serve it arrives late and financed at 5%. Dilution was last week’s cost to measure. Delay is this week’s.

YOUR WEEKEND REVIEW

Write five lines for one holding:

1 How it gets paid: on delivery, on usage, or on a transaction.

2 Whether a 50% fall in the price of its output would raise or lower its revenue, and why.

3 The one disclosed metric that would settle question 2.

4 The date outside its control that must arrive before it earns.

5 The cost of waiting six months past that date at today’s rates.

This companion supports Issue 30 of iPrompt Signals, the Friday briefing on AI and robotics investing. If someone forwarded it to you, subscribe at [SUBSCRIBE LINK] to get the next issue, with this week’s tripwires checked against the results.

For information and education only, not financial advice. Figures in both models are hypothetical. iPrompt Signals is not a registered investment adviser.

iPrompt Signals Deep Dive · Issue 30 · 25 September 2026