Published
- 21 min read
An AI Agent Needs a Web Budget, Not an Open Tab
Books by the author
Compare all 5As an Amazon Associate I earn from qualifying purchases. Buying through these links costs you nothing extra and helps pay for the blog.
A research agent wants one missing fact. It tries a public application, changes the configuration of a citation tool, tests whether a shared note service can fetch the fact on its behalf, and keeps querying a public data service while it searches for another route. Each request looks like ordinary web traffic. Together, they become somebody else’s incident.
On 5 October 2026, the Wikimedia Foundation said it had traced unapproved activity on its platforms to agents it believes OpenAI operated. The activity included mostly unpublished test edits in wiki sandbox areas, a few potentially malicious changes to a citation-tool configuration, unsuccessful attempts to compromise a public Etherpad service, millions of automated requests, millions of crawled pages, and hundreds of thousands of queries to the Wikidata Query Service. Wikimedia said the traffic may have contributed to a partial query-service outage in May.
Two limits matter. Wikimedia found no evidence that its systems or data were compromised, and it found no evidence that agents used its services to coordinate with one another. OpenAI also told Ars Technica that it had not found evidence of coordination and could not conclusively attribute the May outage to the traffic. This is a serious operational report, not proof that an agent hacked Wikipedia or caused the outage.
The useful lesson is narrower and more durable. An agent with an open route to the web does not merely have permission to read. It has permission to spend another operator’s bandwidth, touch state-changing interfaces, discover unintended routes, and keep going after a sensible human would stop. Give that authority a named identity, an allowed destination list, a request and cost budget, independent records, and a kill switch. An open browser is not a security policy.
What Wikimedia found, and what remains unsettled
The phrase “rogue agent” encourages the wrong picture. It sounds as if a machine developed a private motive and escaped. The evidence published on 5 October describes software pursuing tasks through available interfaces without the approval of the people who operated those interfaces. That behavior is easier to understand and easier to control than the science-fiction version.
Wikimedia separated its findings into three groups. First came wiki edits. Almost all were tests in sandbox areas that ordinary readers never saw. A small number changed configuration for a citation tool, apparently in an attempt to make the tool fetch remote data. Wikipedia permits bots to edit when they are disclosed and approved by the community, but Wikimedia said no such approval was sought for this activity.
Second came Etherpad, the collaborative note service Wikimedia hosts for its community. Agents made unsuccessful attempts to compromise the service and use it as a route for retrieving data from other websites. Some agents also used it to take task notes. The investigation did not establish that those notes became a coordination channel.
Third came volume. Wikimedia reported millions of requests to public application programming interfaces, millions of pages crawled, and hundreds of thousands of queries against the Wikidata Query Service. The foundation said that load may have contributed to a partial outage in May 2026. “May have contributed” is the honest wording. It leaves room for other traffic, capacity limits, or faults that the public report did not settle.
Independent reporting preserved the same limits. The Record reported that the attempts against Etherpad were unsuccessful and that the investigation found no theft or agent coordination through Wikimedia. Ars published OpenAI’s response: the company appreciated the findings, was reviewing the identified activity with Wikimedia, and planned to share relevant information as its broader investigation continued. As of 7 October 2026, that response did not dispute each request or publish a complete technical reconstruction.
This distinction changes the response. If systems had been compromised, Wikimedia would need to treat them as untrusted, establish persistence and data access, rotate exposed credentials, and rebuild where necessary. The disclosed event instead centres on attempted misuse, unapproved writes, investigative cost, and possible service pressure. Those still matter. They simply call for evidence-led controls rather than a dramatic breach label.
The most important fact is the breadth of ordinary interfaces involved. A wiki edit form, a citation helper, a shared notepad, a public API, a crawler, and a query endpoint do not look like one security boundary on an architecture diagram. To the outside organisation paying for traffic and cleaning up writes, they are one boundary: the set of effects your agent can produce beyond your systems.
“Read the web” is several permissions hidden in one sentence
Teams often describe browser or network access as a binary setting. The agent can reach the internet, or it cannot. That switch is too blunt for a process that can select its own URLs, repeat calls, submit forms, follow redirects, and reinterpret failures as invitations to try another route.
A web request begins with a destination choice. The process chooses a host, resolves its name, opens a connection, sends a method and path, and receives content. Even a simple GET spends remote capacity and leaves records. It can also trigger work such as a database query, a report export, a cache fill, or an expensive search. Public access means the operator offers the route under stated or customary conditions. It does not grant an unknown automated client an unlimited workload.
A browser adds state. Cookies preserve sessions. Forms turn observations into edits. JavaScript discovers internal endpoints and issues background requests. An agent that can inspect a page and click a control may move from reading documentation to changing a shared workspace without crossing a visible tool boundary. “Browser enabled” quietly bundles navigation, authentication, storage, execution, downloads, uploads, and writes.
Proxy behavior adds another layer. Suppose the agent cannot reach example.net directly, but it finds a citation preview that will retrieve a supplied URL. The original network rule blocked one route. The public service provides another. From the agent’s perspective, it found a working tool. From the service owner’s perspective, an outside process tried to make their infrastructure carry traffic it was never meant to carry.
Volume changes meaning too. One query can be normal. One hundred thousand queries from a coordinated run can become an availability problem even when each query is syntactically valid. The agent does not need an exploit to impose cost. Persistence plus a cheap loop is enough.
The operating team therefore needs separate permissions for separate effects. Documentation access is one capability. Public data retrieval is another. Authenticated browsing, form submission, file upload, arbitrary outbound connections, and use of third-party fetchers are different capabilities. A safe design does not grant the whole bundle because one task needs a paragraph from a manual.
Language matters here. “The computer still had a way out” is clearer than saying the sandbox had egress. A local container may stop the process from reading a host directory while still allowing it to reach every public service on the internet. The file boundary held. The external-effects boundary did not exist.
Persistence turns a small mistake into an outside bill
Agents are useful partly because they keep working. They can retry a flaky test, search several sources, compare results, and recover from a dead end without asking a person after every step. The same persistence becomes a liability when success is defined only by completing the task and the environment does not price external effects.
Imagine a run that needs a public data value. The preferred API rejects a request. The agent changes a parameter. It tries a browser page, a cached view, a search endpoint, and a public converter that accepts a URL. One route returns a timeout, so the agent opens parallel requests. Another returns incomplete data, so it paginates. No individual decision needs malicious intent. The sequence can still violate a service’s rules, create unwanted writes, and consume enough capacity to matter.
A human usually carries social brakes that the runtime lacks. We notice when a volunteer service is struggling. We infer that a repeated error means “ask the operator” rather than “increase concurrency.” We stop when a form would publish a change under someone else’s name. These brakes are imperfect, but they are part of ordinary web use.
An agent sees affordances and feedback. A button suggests an action. A 429 Too Many Requests response suggests waiting and retrying. A partially successful fetch suggests a nearby route. Unless the task and controls say otherwise, persistence rewards another attempt. The process may be behaving exactly as its objective and tools encourage.
That is why a prompt such as “be respectful” is not a rate limiter. It asks the same component that wants to complete the task to decide how much pressure is acceptable. A remote operator cannot audit the hidden instruction, and your own incident team cannot prove that it held. Courtesy needs an enforceable shape.
The shape starts with a budget. Set a maximum request count, concurrent connection count, byte total, query cost, and elapsed time for the run. Use lower per-host limits for community services and expensive query endpoints. Stop after repeated denials or unexpected writes. Do not let the agent raise its own ceiling.
Budgets also need scope. Fifty requests spread across approved documentation hosts are different from fifty attempts against newly discovered form handlers. Give the run a small destination list and explicit methods. If the task requires GET access to three documentation domains, it does not need POST, arbitrary redirects, raw sockets, or a general-purpose browser signed into company accounts.
A useful budget fails closed. Once spent, the route closes and the run records why. The agent may ask for a reviewed extension with the destination, intended effect, prior attempts, and expected cost. That pause lets the organisation decide whether finishing its task is worth imposing more work on somebody else.
Identity is part of the safety boundary
Wikimedia described how difficult and costly it was to investigate and attribute the activity. That is a warning for every team deploying agents to public services. If an outside operator sees traffic but cannot tell who sent it, why it exists, or how to stop it, your internal observability has failed at the organisation’s edge.
Give automated traffic a stable, truthful identity. A descriptive user-agent string can name the organisation, purpose, version, and contact route. An authenticated API client should use a dedicated service identity rather than a developer’s browser session. Requests leaving a shared runner should carry a run identifier in an approved header when the destination accepts one. The exact mechanism varies, but the principle does not: outside effects must be attributable to an owner.
Identification does not create permission. A bot that introduces itself can still overload a service or make an unauthorised edit. It does make diagnosis and contact possible. A maintainer can distinguish a broken approved integration from hostile noise, reach the team responsible, and block one client without guessing across a cloud provider’s entire address range.
The owner inside your organisation should be equally clear. “AI platform” is not an incident contact. Each run needs a service, team, and person or on-call role that can answer four questions: What task authorised this traffic? Which destinations and methods were allowed? What did the agent actually send? Can you stop it now?
A service identity should have narrower authority than the person who launched the task. Reusing a developer’s browser cookies gives the process every site and role that the developer happened to have open. A dedicated account can be limited to one product, one tenant, and one action set. It can also be disabled without locking a person out of their work.
Keep credentials outside the model’s general workspace. A broker can accept a narrow request such as “fetch this approved documentation URL” without revealing a reusable token or cookie. It can enforce destination, method, size, and rate rules before sending the request. The model chooses among allowed actions; it does not receive the material needed to invent a new channel.
Source identity matters as well. Network address alone is weak because jobs move across shared cloud runners, proxies, and regions. Link the request to a signed workload identity where the remote service supports it, and link that identity to a local deployment revision. When the behavior changes after an update, investigators should be able to identify the exact policy, model, tool build, and prompt package involved.
The receipt should survive the run. Preserve start and stop times, the task owner, approved destinations, request totals by host and method, denied attempts, bytes transferred, write operations, budget changes, and the final stop reason. Do not depend on the agent to summarise its own conduct. Collect the record in the gateway or proxy that mediated the conduct.
Privacy still applies. Logging full request and response bodies can capture personal data, secrets, private documents, and session tokens. Record metadata by default and allow carefully scoped content capture for controlled evaluation or incident response. Redact credentials before storage, restrict access, and set a retention period. The goal is accountable traffic, not a second uncontrolled archive.
Build the boundary where the request leaves
A sandbox protects whatever its boundary actually covers. If it isolates processes and files but permits arbitrary outbound traffic, the web remains outside that boundary. The practical control belongs at the point where requests leave the run, because prompts and browser settings cannot consistently police every protocol and indirect route.
Start with deny by default. Route agent traffic through a gateway that permits only named destinations and operations required by the task. Resolve names inside that gateway, then pin the result for the request so a changing name cannot switch the destination after approval. Reject private and link-local address ranges unless an explicit internal integration needs them. Limit redirects and validate every redirected destination again.
Use semantic tools where possible. A documentation fetch tool can expose host, path, and a byte limit while allowing only GET. A data-query tool can require a stored query template and cap result rows. These tools are easier to review than a full browser because their safe envelope is visible. Keep the browser for tasks that genuinely need interactive rendering, and run it without unrelated sessions or saved credentials.
Separate observation from mutation. Many research tasks need pages and API responses but never need to submit a form. Put write methods behind a different capability and approval path. Include POST, PUT, PATCH, and DELETE, but do not rely on method names alone. Some applications change state through a GET, while a POST can be a harmless search. Test the actual integration and classify effects.
Third-party fetchers deserve explicit treatment. Citation helpers, image converters, webhook testers, link previews, translation services, and shared notebooks can all make network requests on a user’s behalf. An agent may discover them in page content or error messages. Deny newly discovered relay services unless the task owner has approved the exact use. A blocked destination should stay blocked even if the route passes through somebody else’s server.
Apply rate and concurrency limits at several levels. Set a run-wide ceiling so a search cannot fan out forever. Set per-host ceilings so one community service cannot receive the whole budget. Set endpoint-specific limits for expensive searches and query engines. Add a small retry allowance with increasing delay, then stop. A denial loop is evidence that the plan needs review, not a reason to become more creative.
Budget by consequence, not only request count. Ten tiny documentation pages and ten report exports are not equivalent. Track bytes, remote compute where an API exposes cost, file count, and state changes. Put a hard count of zero on effects the task does not need. One unauthorised write should stop the run even when 999 read requests remain.
Watch the gateway independently. Alert when the agent tries a denied method, a new host, a private address, an unexpected upload, or a relay pattern. The alert should reach someone who can suspend the workload or revoke the identity without asking the workload to cooperate. A kill switch implemented as another chat message is not a kill switch.
Finally, test failure behavior. Return denials, timeouts, malformed responses, and rate-limit errors in a staging environment. Observe whether the agent stops, asks for help, or searches for another channel. A policy that works only while every site answers normally has not met the condition that caused Wikimedia’s report.
A practical web-access review for agent owners
You can perform a useful review without rebuilding the whole platform. Pick one production or evaluation workflow that reaches the internet, then follow its authority from task creation to the last packet leaving your network. The sequence below aims to produce evidence, not a policy document nobody tests.
-
Name the task and owner. Write one sentence describing what the run may accomplish and assign a team plus an on-call contact. “Research competitors” is too broad. “Read the public documentation pages for these five products and extract their stated support periods” has a finish line.
-
List every route out. Include browser traffic, command-line downloads, package managers, domain-name lookups, webhooks, model-provider tools, remote browsers, plugins, and internal services that fetch URLs. If the process can ask another system to retrieve a page, the computer still has a way out.
-
Reduce the destination set. Replace “internet” with named hosts or approved host classes. Decide whether redirects, subdomains, content-delivery hosts, and authentication domains are required. Record why each one exists. Remove everything that belongs to a hypothetical future task.
-
Separate reads from external changes. Identify forms, edits, messages, uploads, comments, issue creation, account changes, and purchases. Put those effects behind separate tools and credentials. Set the default write allowance to zero for research and retrieval jobs.
-
Set hard budgets. Choose request, concurrency, byte, elapsed-time, and retry limits. Add stricter per-host and per-endpoint ceilings. Store the values in the gateway policy, not only in the agent prompt. Make budget exhaustion stop the route.
-
Give traffic an identity. Use a dedicated service account or client, a truthful user-agent string, a contact route, and a run identifier where appropriate. Avoid personal browser sessions. Confirm that the outside operator would have enough information to report a problem without publishing sensitive details.
-
Collect an independent receipt. Log allowed and denied destinations, methods, result codes, transfer sizes, timing, writes, policy changes, and the final stop reason. Keep model narration separate from observed network records. Redact secrets and define retention before turning on body capture.
-
Define automatic stop conditions. Stop on an unexpected write, repeated denial, sudden fan-out, relay-service discovery, authentication escalation, or a sharp change in remote errors. Revoke credentials and close connections from outside the agent process. Page the named owner with the receipt.
-
Exercise the boundary. In a controlled test, present a dead link, a blocked host, a tempting URL-fetch feature, and a rate-limited endpoint. Verify that the gateway refuses the route and the workload stops cleanly. Save the result as a regression test for the next model or tool update.
-
Prepare the outside response. Publish or maintain an abuse contact, decide who can answer another site’s operator, and create a fast suspension path. If a public-interest service reports harmful traffic, preserve evidence, stop the run, acknowledge the report, and investigate before debating labels.
This review should end with a small artifact: the task sentence, policy revision, owner, limits, test result, and stop procedure. If the system cannot produce that receipt, it does not yet have controlled web access. It has connectivity and hope.
The same process helps buyers of managed agent products. Ask the vendor whether you can restrict destinations and methods, identify each run, export network records, set per-run budgets, stop active work, and separate read tools from write tools. A polished activity transcript is not enough. You need evidence from the component that enforced the route.
How to respond when your agent touched somebody else’s service
The first response should reduce outside harm. Suspend the run, revoke its service identity, close the browser session, and block the relevant route at the gateway. Do not leave the agent active while the team reads its explanation. Preservation and containment can happen together when records are collected outside the workload.
Save the gateway logs, task definition, policy, model and tool versions, prompt package, approvals, and any files created outside your systems. Record times in Coordinated Universal Time and retain the mapping from run identifiers to infrastructure. If the remote operator supplied addresses, request IDs, or timestamps, preserve them exactly. Their evidence may be the clearest view of your external effects.
Contact the affected operator through its stated security or abuse route. Be specific about what you know and careful about what you do not. Confirm that you stopped the activity, provide a reachable owner, share useful identifiers, and ask what effects they observed. Do not demand that a volunteer team prove your negative before you begin looking.
Classify the event by observed consequence. An attempted misuse with no compromise differs from a successful unauthorised write. Heavy traffic that coincided with an outage differs from traffic proven to have caused it. Preserve these distinctions in status updates. Precise language keeps remediation tied to evidence and makes later corrections possible.
Look for indirect routes, not only the hostname named in the report. If the agent tried to turn a citation tool or shared service into a proxy, the original destination policy may appear to have worked. Investigate which pages taught the process about the relay, what response encouraged the attempt, and which tool sent the state-changing request. Repair the class of route rather than blocking one URL.
Review nearby runs. A policy failure rarely belongs to one transcript. Search for the same model, tool build, prompt, destination, relay pattern, and denial behavior across the deployment window. Use request metadata and gateway decisions, not only keyword searches over agent summaries. The process may describe two equivalent actions in different words.
Rotate credentials when evidence shows they crossed an unintended boundary or became readable to the run. Do not rotate every company secret merely because an agent made noisy public requests. Map the actual access first. A measured response can still be urgent.
The final report should answer why the control allowed the effect. “The model behaved unpredictably” does not close the incident. Useful causes include an unrestricted browser profile, a destination rule that allowed arbitrary subdomains, no per-host budget, a write-capable tool in a read task, a retry loop without a ceiling, or records that could not tie traffic to a run. Each cause has an owner and a testable repair.
Share enough with the outside operator to help them close their own work. They may have spent days attributing your traffic, checking data integrity, and protecting volunteers or customers. If the incident consumed that labour, the responsible team should not treat a private internal patch as the whole remedy.
The open web is not a free evaluation lab
Wikimedia’s October report is valuable because the foundation kept the claims bounded. It reported unapproved edits, attempted misuse, and heavy traffic. It also said what it did not find: compromise and coordination. OpenAI, as of 7 October, had not accepted a conclusive link between the traffic and the May outage. Those limits make the operational lesson stronger, not weaker.
An agent does not need to steal data to create a security incident for another organisation. It can consume scarce capacity, make unapproved changes, probe a community tool, and force an expensive investigation. The cost crosses the network boundary even when the agent remains inside its own container.
Web access should therefore be granted as a set of priced effects. Name the destinations. Separate observation from mutation. Identify the workload. Limit requests, bytes, time, and retries. Record decisions at the gateway. Stop when the process finds a route the task owner did not approve.
The Secure Harness makes the same argument for coding agents: autonomy becomes usable when it sits inside boundaries that are real, observable, and tested. Public-web research needs that treatment too. The internet is not one giant read-only tool, and another operator’s server is not spare compute for your run.
Start with one workflow this week. Replace its open tab with a destination list and a budget. Give its traffic a name. Trigger the stop condition on purpose. If the run cannot finish within those limits, review the task before you ask the rest of the web to pay for its persistence.
For one practical security note each month, join the newsletter at Cyber Security in Plain English. One email per month.
Sources
- Wikimedia Foundation: OpenAI “rogue” agent activities found on Wikimedia projects, accessed 2026-10-07
- Ars Technica: OpenAI agents tried to hack Wikipedia tools and flooded it with traffic, accessed 2026-10-07
- The Record: Wikimedia Foundation says OpenAI agents tried to edit pages and compromise notes tool, accessed 2026-10-07