iPrompt
DEEP DIVE · iPROMPT #154 COMPANION
23 SEPTEMBER 2026
AI agent security
Test what your agent can reach
In a week you’ll have a list of everywhere one agent connects, and know which of its limits are real. Three steps this week; two more once you know what normal looks like.
7 MIN READ · TEST AND CHECKLIST
The newsletter explains why an agent’s belief about its scope is not a boundary. This companion starts with the work. The cases come from Google’s and Irregular’s statements as reported by the press, a developer’s analysis of ZCode, AIR’s Plugin4Shell disclosure and Amazon’s statement on Muse. AIR sells plugin-marketplace security. None of these cases tells you how often small teams are affected; the test below is iPrompt’s guidance.
AI AGENT SECURITY / THE TEST
This week: the first pass
Pick one agent. Write down its name, version, the dates and who is responsible. You need one free tool, one setting and a short review at the end. This is a first pass, not a security certification. If you find an exposed key or an upload you can’t explain, use last week’s response guidance before carrying on.
1 Scope. List the agent’s job and what you’ve given it: folders, repositories, connected accounts, plugins, and any saved passwords or keys it could read, including configuration files such as .env. Access nobody can account for is a finding.
2 Observe. Install LuLu on a Mac or Portmaster on Windows or Linux. In LuLu, untick ‘Allow installed applications’ during set-up so the agent isn’t trusted silently. Use the agent normally for seven days and note the addresses it contacts, especially any while nobody is using it.
3 Classify. Sort every address as needed for the job, unexpected or unknown, using the newsletter’s reach-map prompt. Check the needed ones against the vendor’s documentation.
What counts as significant. Act the same day on uploads you didn’t trigger, connections to storage or file-sharing services unrelated to the job, activity while the agent is idle, or a key within the agent’s reach. An address that is merely unexplained is a question for the vendor, not an incident.
Next: two follow-up controls
Once you know what normal looks like, close the gaps the first pass found.
4 Probe. Test limits in a copy or test account, never against systems you don’t own. Put a canary credential, a decoy key that alerts you when used, in the test folder. Switch the vendor’s privacy settings off and on and see whether the traffic changes.
5 Restrict. Block every address the job doesn’t need, adding back what breaks. Give the agent its own account with only the access its job requires, and keep keys out of its folders. Update the agent promptly. Decide deliberately about automatic plugin updates: they deliver fixes, and under Plugin4Shell they would have delivered the swap.
Finish with a short record: agent checked, addresses found, action taken and open questions with owners. A first pass is complete when every finding has an owner or an answer, not when the agent has been declared safe.
Where an agent’s reach extends
Use the rows that apply to your agent. They cover this month’s cases, not every route out.
Route | What happened | What to check |
Network access | Gemini reached live systems during a test it believed was sealed. | Where it connects when unattended. Block what the job doesn’t need. |
Keys it can read | Two of Gemini’s three logins used credentials from a public repository. | Keys out of its folders, and out of old versions kept in Git history. Keys that expire. |
The tool’s own uploads | ZCode packed a whole project, Git history included, for upload. | Observed traffic against what the vendor documents. What ‘off’ actually stops. |
Plugins | Plugin4Shell let a plugin’s owner swap approved code. Fixed in Claude Code 2.1.179 and Codex 0.146.0; plugins hosted outside GitHub carry most of the remaining risk. | Agent version. Where each plugin comes from, and who controls it. |
Connected accounts | Amazon says Muse could reach account pages and order history. | What each connection can read and change, and where you approve actions. |
Cases from Google’s statement, ferstar’s analysis, AIR’s disclosure and Amazon’s statement to GeekWire. Checks are iPrompt’s recommendations.
AI AGENT SECURITY / VENDORS AND LIMITS
Judge the behaviour, not the setting’s name
ZCode had a setting called Repository Snapshot Indexing. According to ferstar’s analysis, it governed what happened to a snapshot after it arrived; the client built and tried to send one either way. Z.ai says the behaviour is fixed and the data destroyed. The archive was encrypted to a key only Z.ai holds, so users cannot verify that themselves.
Put these questions to the vendor, then check the answers against step two.
Ask the vendor | Why it matters |
What leaves the machine by default, and which setting stops it? | A setting can name one step and leave the rest running. |
Who holds the keys to anything uploaded? | If only the vendor can decrypt it, only the vendor can confirm deletion. |
Does the agent identify itself to the sites it visits? | Amazon’s objection to Muse was an agent it couldn’t identify. |
Does the agent check that installed plugins match what was approved? | A version number in a settings file is not proof of what is running. |
Instructions are not a boundary
Gemini had a defined scope and a mistaken belief about where it was. Google says it stopped once it recognised real targets. A limit that depends on the model noticing its own mistake has already failed once. The fix is a control that doesn’t depend on belief: a network that only reaches the test, an account that can only touch its job.
Then make sure you would notice. Google learned of the intrusions about two months later. Name the log that would show an unexpected connection, and who reads it; the monitor from step two is a start.
The strongest objection
‘Our agent runs in the vendor’s sandbox. Containment is their job.’ Partly true. A vendor sandbox protects the vendor’s boundary. Your keys, repositories and connected accounts sit inside it, because you put them there.
The Gemini evaluation was run by a specialist firm whose business is testing these systems safely, and the model still reached the internet. A sandbox is a claim until someone checks it.
These cases can’t tell you the odds. That’s why the test is bounded: one agent, one week, every finding answered or assigned.
YOUR MOVE
Watch one agent for a week. Reply with the agent, the number of addresses and what you did about the unexpected ones. ‘Nothing unexpected’ and ‘still asking the vendor’ both count. Never send passwords, keys or file contents.
The goal is an agent whose reach you’ve seen, not one you’ve been told about.
NEW TO iPROMPT?
Every Wednesday: the week’s AI news turned into one thing you can do. Free.
Disclosure: the original draft used Claude. Anthropic makes Claude Code, named in the Plugin4Shell report. AIR and Irregular are commercially interested sources; their claims are attributed.
ISSUE ACCOUNTABILITY / APPENDIX
The standing prediction ledger
Historical forecasts are reproduced for continuity, separately from the containment guidance. This edition does not independently re-score them.
The statuses below are carried from the supplied issue records. An unresolved entry is not a claim that no relevant event has occurred. Original wording and counting rules must be checked before any outcome is scored.
Deadline and call carried in #151 | Carried status and verification limits |
31 Mar 2027 · Hugging Face control deal publicly announced | Previously recorded as meeting the announcement condition in #152. Not re-scored here; closing is a separate event. |
30 Jun 2027 · Two additional model makers tighten flagship licences | Carried forward; not scored. One development logged for review below. |
30 Jun 2027 · Major hub or gateway restricts previously free access | Carried forward; not scored in this edition. |
31 Dec 2027 · Formal competition review of the Nvidia deal | Carried forward; current status not independently verified. Filing alone is not scored as a formal review. |
31 Oct 2026 · Independent GLM-5.3 CyberGym result below 84.5% | Carried forward; current status not independently verified. Vendor figures do not count. |
Logged for review: on 20 September Alibaba released Qwen-Image-2.1 under a non-commercial research licence, where earlier Qwen-Image releases used Apache 2.0. Whether an image model counts as a flagship under the #151 wording has not been assessed, and the change is not scored here.
The #151 counting rules, including void and loss provisions, remain controlling. These entries are abbreviated. No new forecast is added.
Reader response follow-up
No reader exercise from earlier issues remains open. This week’s reply request carries no published tally; replies inform future issues only.
EDITOR ONLY / REMOVE BEFORE PUBLISHING
SEO title. AI agent security: test what your agent can reach (49 characters)
Meta description. Test what your AI agent can reach in one week: where it connects, which limits are real, and what to ask the vendor. Free tools, step by step.
Slug. /ai-agent-security-reach-test
Target keywords. AI agent security (primary); AI coding agent security; AI agent network access; AI agent data exfiltration; outbound firewall for AI agents
Placement. Primary keyword appears in the title, the H1 and the first section label. Link from #154’s Our Angle; add a link from the #153 key-security companion.
iPROMPT / 154 /