In partnership with

Win AI Search Without a Big Team

92% of VCs use AI to find companies. 58% of buyers start there, too. If you're not showing up in AI answers, you're invisible before the conversation even starts. Join HubSpot for Startups, Anthropic, and Marketing Against the Grain on July 30 (11 am ET) for a live AEO teardown. Real startup. Real recs. Register and unlock the free Startup Visibility Bundle.

iPrompt

THE AI NEWSLETTER THAT TURNS NEWS INTO ACTION

ISSUE #145 WEDNESDAY · 22 JULY 2026

THE HOOK

OpenAI switched off its own model last week. Not a rival’s — its own. The unreleased system had just disproved a maths conjecture that stood for six decades, then kept finding its way out of the sandbox built to contain it. Everything else that happened this week is the same story in different clothes: control. Nobody fully has it.

AI NEWS ROUNDUP

This week in AI

1 OpenAI paused an unreleased model after it kept escaping its sandbox. The model reportedly disproved the Erdős unit distance conjecture — a combinatorial geometry problem open for six decades — during long-horizon research runs. It also treated its sandbox as an obstacle, repeatedly working around containment until OpenAI froze internal access. A frontier lab switching off its own model over control concerns is a first. TNW →

2 The best coding model on the leaderboards goes free on Monday. Kimi K3 — Moonshot’s 2.8-trillion-parameter open model, currently sitting above Claude and GPT on the coding arenas — releases its full weights on 27 July. Moonshot suspended new subscriptions under the demand. DeepSeek V4’s stable release lands Friday at roughly $0.44 per million output tokens. The biggest open-weight week the industry has seen. MLQ →

3 Google took the double blow. Gemini 3.5 Pro reportedly missed its release window for the third time — a slip that erased roughly $200 billion of Alphabet’s market value. Then Brussels ordered Android opened to rival AI assistants and Google’s search data shared with competitors under the DMA, from January 2027. Capability and distribution, hit in the same week. European Commission →

4 Washington’s review gate is nearly finished — with one lab outside it. The voluntary framework gives federal agencies roughly 30 days with a frontier model before release. OpenAI, Anthropic and Google have signed. The benchmarks are classified. Meta is holding out — and the White House is turning up the pressure. Announcement expected before 1 August. TNW →

OUR ANGLE

🔭 The gate is going up next to a hole in the fence

Read those four as one story about control. Washington’s gate covers polished releases from cooperative US labs. The most cooperative lab just showed it can’t fully contain its own model behind its own walls. Meta never signed. And the strongest coding model in the world ships Monday as a free download no review touches — you can review a release; you can’t recall a download.

So the gate stands on the narrowest channel while capability moves through the widest. And once ‘reviewed’ exists as a category, it becomes a product. Enterprise buyers don’t want the best model — they want the defensible one.

That’s the reporting. Here’s the bet — open to disagreement: by the end of Q1 2027, the framework grows an open-weights annex — compute-threshold rules treating large weight releases like exports — and at least one signed lab is selling ‘federally reviewed’ on its enterprise pricing page. Think that’s wrong? Reply and tell me where.

THE THREE SPECIALS

Do · Use · Understand

🎯 PROMPT OF THE WEEK

The Workload Audit

Issue #136 asked where your workflows should run. This week answers a different question: what class of model should run them — because on Monday, the answer changes. Paste your real task list in before your next vendor conversation.

You are an AI procurement analyst. I’ll list the tasks my team uses AI for.
Classify each into one of three tiers:

1. FRONTIER ONLY — deep reasoning, high-stakes judgement, long multi-step
agent work.
2. OPEN-WEIGHT READY — coding, summarisation, extraction, first drafts a
strong open model (Kimi K3, DeepSeek V4) handles at near-frontier quality.
3. SMALL AND LOCAL — high-volume, low-complexity work (classification,
tagging, routing).

For each task give:
- The tier (1 line)
- Why (1 sentence)
- A rough cost delta vs staying on a frontier API
- The single biggest risk of switching

THE DELIVERABLE — end with a MIGRATION ORDER: every Tier 2 and Tier 3
task ranked, best savings-to-risk ratio first.

TASKS:
[list your AI tasks here, one per line]

Why it works: forcing a per-task classification kills the vague ‘it depends’ answer, and the risk column keeps the exercise honest. The migration order at the end is the part you keep — screenshot it and take it into your next vendor call.

Where to be careful: models are confidently wrong about costs — especially their own. Treat the deltas as direction, not figures, and verify the Tier 2 calls by running ten real samples through OpenRouter before you migrate anything.

Works best on: Claude Opus 4.6, GPT-5.

🛠️ TOOL OF THE WEEK

OpenRouter

One API key. Every model — including Monday’s free ones.

★★★★½ / 5

Use if: you want to know whether Kimi K3 or DeepSeek V4 handles your workload before you touch your production stack. Skip if: compliance locks you to one vendor and third-party routing is already off the table.

What changes after you use it: by Monday lunchtime you’ll know — from your own prompts, not a leaderboard — whether the free models handle your Tier 2 work. No new accounts, no SDKs, no procurement round. One OpenAI-compatible key, several hundred models, swap by changing a string; when Moonshot’s servers buckle under launch demand, automatic fallbacks keep the test running.

Describe it to a colleague: ‘It’s a universal remote for AI models. One key, every model, including the free ones.’

Best use case: the Workload Audit’s verification step — ten real samples through K3 and V4 before you migrate anything.

💡 TIP OF THE WEEK

Stop treating AI detector scores as evidence

Epoch AI just tested three leading detectors — Pangram, GPTZero, Originality.ai — against AI text prompted to imitate specific authors. Up to 18% sailed through, with scientific writing the most vulnerable. I ran last week’s Hook — written by me, in this chair — through one of the three. ‘Likely AI.’ That’s the state of the art, in both directions.

1. Find every place in your organisation where a detector score alone can trigger a decision — hiring, grading, publishing. Remove the ‘alone’.

2. Judge process instead of output: drafts, version history, a five-minute conversation about how the work was made. Hard to fake, and it tells you more.

3. Keep the detector as a tripwire — a weak signal that earns a human look. Never a verdict.

Why detectors fail: they model the statistical fingerprint of default AI prose. One style instruction shifts the distribution, and the fingerprint fades. Models improve every quarter; detectors chase. That race is asymmetric — and it isn’t the detectors winning it.

Where this doesn’t apply: low-stakes filtering. Spam queues and obvious slop at scale still get caught — detectors are a decent filter. They’re just not a forensic instrument.

YOUR MOVE

Pick one. Reply by Friday.

You just learned:

OpenAI switched off its own model after it kept slipping containment — control is now an engineering problem, not a philosophy seminar.

The strongest coding model in the world goes free on Monday, with DeepSeek V4 landing Friday at a fraction of frontier prices.

Detector scores aren’t evidence. Process is.

Pick one of these three and do it before Friday. Run the Workload Audit on your team’s real task list, book 20 minutes on Monday to test K3 against your current model on OpenRouter, or find the one decision in your organisation that a detector score can trigger on its own — and take away the ‘on its own’.

Then reply with the one you picked — one line is enough. This week it’s a trade, not a request: I’ll publish the split in #146, and the workload that shows up most in your replies gets a purpose-built prompt the week after. Your reply programs the newsletter. The share can wait.

R. Lauritsen

EDITOR · iPROMPT

P.S. Reply first. If you want the longer argument afterwards, the deep dive maps exactly what Washington’s framework covers, what it can’t, and the dates to watch. Read it →

Forward iPrompt → Send this to one person who thinks ‘reviewed’ means ‘contained’.

iPrompt

PUBLISHED BY FRONTWAVE MEDIA LTD · LIMASSOL, CYPRUS