CSIPE

Published

- 21 min read

Google's Bug Bounty Pause Shows the Cost of Unproven AI Findings


Books by the author

Compare all 5

As an Amazon Associate I earn from qualifying purchases. Buying through these links costs you nothing extra and helps pay for the blog.

A coding agent spots a suspicious bounds check in a library. Within minutes it can name the weakness, draft a severity score, write a polished report, and send that report to a maintainer. Only one task in that sequence is expensive: proving the bug exists.

On 1 October 2026, Google stopped accepting new product-vulnerability reports through its Open Source Software Vulnerability Reward Program. The company said automated submissions had risen significantly and that the vast majority were invalid. TechCrunch reported the pause on 4 October and quoted Google’s promise of an update in the first quarter of 2027.

The broad headline, “AI broke a bug bounty,” misses the operational lesson. AI made a report cheap for the sender while leaving verification expensive for the receiver. A fluent claim entered a scarce human queue before it had met a useful standard of proof. Google can pause a program. A small open-source project may lose an evening, burn out a maintainer, or stop accepting private reports altogether.

My position is simple: nobody should send an AI-assisted vulnerability report that they cannot reproduce against the named revision. A model may find the lead and help investigate it. The sender still owns the claim. Put an evidence gate between generated suspicion and another person’s inbox.

What Google paused, and what it did not

Google’s decision was narrower than several headlines made it sound. The pause covers new product-vulnerability submissions to the Open Source Software Vulnerability Reward Program, usually shortened to OSS VRP. Google’s current program rules state that product-vulnerability reports have not been accepted since 1 October 2026. Reports submitted before that date continue through the existing process.

Supply-chain submissions remain in scope. Those concern routes into source, build systems, package credentials, release artifacts, or signing material rather than a defect in the runtime behavior of one product. SecurityWeek independently confirmed on 5 October that only product vulnerabilities were covered by the pause. Researchers with eligible Google Cloud findings may also have a route through the separate Cloud Vulnerability Reward Program.

That distinction matters because “Google closed its open-source bounty” suggests a retreat from all outside research. The verified facts support a more precise statement. One intake lane was paused after the volume and validity of automated reports made that lane hard to operate. As of 6 October 2026, Google had promised an update in the first quarter of 2027, not a reopening date or a final replacement policy.

The pause also followed an earlier attempt to raise the bar without closing the lane. In March 2026, Google published new OSS VRP rules that asked for stronger proof for certain reward tiers. The examples included a reproduction through OSS-Fuzz, Google’s continuous fuzz-testing service for open-source projects, or a patch already accepted by the affected project. The point was practical: direct review time toward claims with demonstrated effect.

By October, that filter had not solved the intake problem. Google said the rise in automated submissions was significant and the vast majority were invalid. It did not publish a submission count, a percentage, or a breakdown by model. We should resist filling those gaps with invented precision. What we can say is that the program operator judged the queue unhealthy enough to stop new product reports for at least several months.

A product bug may still exist even when a reward program will not accept it. Researchers should follow the affected project’s security policy and private reporting channel. The pause changes one route and its financial incentive. It does not turn public issue trackers into a substitute disclosure desk, and it does not make an unverified report more responsible because the usual form is unavailable.

The bottleneck moved from writing to verification

Before coding agents, a detailed vulnerability report imposed some friction on its author. The researcher had to locate the code, understand the path, build the project, trigger the behavior, and explain why the result crossed a security boundary. Weak reports still existed, but producing a plausible ten-page narrative took time.

A model removes much of the writing cost. It can scan many repositories, identify patterns associated with known weakness classes, produce proof-looking code, assign a Common Weakness Enumeration label, estimate a severity score, and fill a disclosure template. The finished report may look more complete than the investigation behind it. Good formatting becomes dangerously easy to confuse with good evidence.

The maintainer cannot verify the report at model speed. Someone must check out the exact revision, reconstruct the build, determine whether the named code ships, inspect callers, trace input control, test the proposed path, and look for a guard the report missed. If the result fails, that person must decide whether the reproduction was wrong, the environment differed, or the claim was false. A polished invalid report can cost more time than a short honest hunch because every confident detail invites checking.

This is a queueing problem before it is an intelligence problem. Suppose a project receives two reports a week and each takes 90 minutes to settle. Three hours of triage may be manageable. If automation raises the intake to twenty reports while the confirmation rate falls, the project now owes thirty hours before it writes a fix. The model created more candidate text, not more maintainer time.

The queue then changes behavior. Real reports wait behind generated ones. Maintainers start skimming. Severity words lose force. Outside researchers receive slower replies. A volunteer who expected to improve code spends a weekend disproving claims about code that the reporter never ran. Eventually the project adds barriers, closes rewards, or narrows who may report.

That last step hurts careful researchers too. A noisy sender does not carry the full cost of a bad report. The receiver pays first, and later reporters pay through stricter access and slower response. The externality is the central fact of AI-assisted disclosure: generation is private and cheap; verification is shared and scarce.

The correct metric is therefore not findings per hour. It is maintainer minutes per confirmed, actionable vulnerability. A scanner that produces one real defect and nine quick rejections may help. A coding agent that produces one real defect and ninety persuasive investigations can make the same project less secure by burying the useful signal.

A convincing report can still describe an impossible bug

Models are good at completing familiar security stories. They have seen paths where an unchecked length reaches a copy, a user-controlled URL reaches a server request, or an access check happens after an action. When code resembles that pattern, the model can explain the expected impact in clear language. The missing question is whether this program, in this build, under these conditions, actually follows the story.

A suspected memory error offers a concrete example. The model sees arithmetic on an input length and a write into a fixed buffer. It generates an input that should exceed the boundary. Yet the public function may reject that input first, the build may compile the feature out, or the allocator may receive a larger dynamic size than the report assumed. The pattern is worth investigating. It is not yet a vulnerability.

Authentication claims fail in similar ways. A model finds one handler that appears to skip a role check and calls it an authorization bypass. Runtime routing may make the handler internal. Middleware may enforce the check before dispatch. The supposedly low-privilege account may require an administrative capability to create the object used in the demonstration. One omitted precondition can reverse the severity.

Dependency reports have another trap: reachable code versus present code. A vulnerable function can exist in the source tree without being used by the affected product configuration. Conversely, a wrapper can make an apparently safe function reachable with attacker-controlled values. Text search finds presence. A security verdict needs the path through the shipped system.

Generated proof-of-concept code does not settle the matter if nobody ran it. Code can import the wrong module, assume a nonexistent endpoint, use a patched interface, or print “exploited” after receiving an ordinary error. The decisive evidence is observed behavior against a named revision in a controlled environment. Save the command, input, output, logs, and environmental details that let another person see the same result.

Severity deserves a separate check after reproduction. A crash in a test process does not automatically become remote code execution. Reading the reporter’s own temporary file does not prove cross-user disclosure. A request that works only with an administrator’s token does not show privilege escalation. The model may propose the worst plausible consequence; the researcher must establish the consequence that occurred.

This is why confidence language cannot replace a test. “Highly likely exploitable” still leaves the maintainer with the whole expensive job. A useful report states what was observed, under which conditions, and which part remains inference. That separation lets the receiver trust the evidence without inheriting the author’s optimism.

Open source pays the bill in human attention

The Google pause is new, but the cost pattern is not. On 31 January 2026, the curl project ended its six-year HackerOne bounty. Curl is the networking library and command-line tool embedded in an enormous range of systems, which makes its security queue valuable and demanding.

Lead maintainer Daniel Stenberg wrote that the program had produced 87 confirmed vulnerabilities and paid more than $100,000 in rewards. He also reported a sharp deterioration in the queue. In earlier years, more than 15 percent of submissions became confirmed vulnerabilities. Starting in 2025, the confirmation rate fell below 5 percent, alongside what he described as an explosion of low-quality AI reports.

Those figures belong to curl, not every bounty. They should not be turned into a universal failure rate for AI-assisted research. They do show the receiving side of the economics. Even a successful program can become unsustainable when report production grows faster than qualified review and each rejection still consumes care.

Curl did not stop accepting security concerns. The project moved suspected security problems to private vulnerability reporting on GitHub. A month later, Stenberg wrote that the first routing change had created its own problems and adjusted the process again. Intake design is operational work. Closing a payment mechanism does not remove the need to receive, protect, investigate, and answer credible reports.

Small projects have less room to absorb mistakes. The person reading the disclosure may also be the release engineer, documentation writer, support contact, and only reviewer for a delicate subsystem. An invalid critical report can interrupt planned work because ignoring it would be irresponsible. The reporter gets to move on; the maintainer must close the uncertainty.

There is also a confidentiality cost. A private report often includes the affected code path, a proposed trigger, and speculation about impact. Maintainers have to protect that material while deciding whether it is real. Sending bulk-generated reports to private channels creates a sensitive archive full of unverified claims. Publishing them as issues is worse because a mistaken report may still advertise a useful attack idea before the project can check it.

Responsible disclosure therefore begins before submission. The sender should reduce uncertainty, not merely transfer it. That means reproducing the behavior, removing irrelevant model output, checking the project’s scope, and stating limits. A report is a request for another person to interrupt their work. Evidence is how the sender earns that interruption.

The lesson for engineering teams is close to home. If your company points an agent at its dependency tree and tells it to report everything upstream, your company has become a security-program operator. It owns the quality of those messages. “The agent sent it” is no defence to the volunteer who spent Saturday disproving it.

Put a proof gate before the send button

The safest design separates discovery from disclosure. Let a coding agent search, classify, and draft inside a private workspace. Do not give that discovery process the credential or tool that contacts a maintainer. The send path should require an evidence packet and a named human decision.

This boundary is useful even when the model is excellent. Discovery rewards recall: surface every path that might matter. Disclosure needs precision: interrupt another team only when the claim is supported. One threshold cannot serve both jobs. If you tune discovery to avoid every false positive, you miss leads. If you publish every lead, you export the triage bill.

A practical pipeline uses four states: candidate, reproduced, validated, and ready to disclose. The agent may create a candidate. A controlled test moves it to reproduced. Someone who understands the affected technology checks security impact and moves it to validated. A disclosure owner confirms scope, wording, confidentiality, and contact details before the report leaves the organisation.

Each transition needs evidence, not a more emphatic paragraph. Candidate to reproduced requires the exact source revision, build instructions, trigger, and observed result. Reproduced to validated requires a demonstrated security boundary, realistic attacker prerequisites, and a check for mitigating controls. Validated to ready requires a minimal report that another engineer can repeat without access to the sender’s chat history.

Keep failed candidates too. If the proposed path stopped at an input check, save that result beside the relevant code and model version. Otherwise the next scan may rediscover the same pattern and charge another reviewer to reject it. Negative evidence turns a false positive from recurring waste into a regression case.

The agent’s network and tools should reflect the stage. Discovery may need source access, a compiler, tests, and a small set of documentation sites. It does not need email, issue creation, social posting, or a bounty submission token. A separate disclosure process can hold those capabilities behind approval. The safest send button is one the discovery agent cannot press.

Rate limits belong at the human boundary as well. Cap how many new candidates one run can place into review. Stop a scan when its early sample produces too much noise. Limit external reports per project until the receiver has answered the first few. These controls sound slow only if report count is the goal. They are fast when useful fixes are the goal.

The Secure Harness makes this broader point about coding agents: useful autonomy needs permissions, evidence, and stop points around the action. Vulnerability research is no exception. A security purpose does not grant a process permission to consume somebody else’s attention without proof.

What a report should prove before it leaves

A strong report lets a maintainer challenge the claim quickly. It names the artifact, shows the behavior, separates fact from inference, and avoids making the receiver reconstruct the researcher’s environment. The following sequence is deliberately demanding because the report may trigger confidential work and an urgent release.

  1. Name the exact affected revision. Record the repository URL, commit hash, release tag, build options, operating system, architecture, and any feature flags. “Latest” expires. A hash lets the project determine whether the code changed between your test and its review.

  2. Confirm that the code ships and the path is reachable. Show how attacker-controlled input enters the affected component. Include the caller, route, parser, or file format that reaches it. If the path exists only in a test utility or disabled feature, say so before assigning impact.

  3. Build a minimal controlled reproduction. Remove unrelated setup and use harmless data. The reproduction should show the claimed boundary crossing, not merely execute suspicious code. A crash log, unauthorized read of a canary value, rejected-versus-accepted request pair, or failing security test is more useful than generated exploit prose.

  4. Save the observed output. Capture the command, input, exit status, relevant logs, stack trace, and result. Mark which lines came from the system and which are your interpretation. Never present model-generated output as a run result.

  5. Test the nearest safe control case. Change one condition that should prevent the behavior: use the fixed revision, remove the permission, shorten the input, switch to an ordinary account, or enable the documented guard. A good control helps prove that the observed effect comes from the suspected defect rather than the test harness.

  6. Check prerequisites and mitigating paths. State what access the attacker already needs, whether user interaction is required, which configurations are affected, and which checks run before the suspect line. If an administrator must perform the final action, the report must not describe an unauthenticated attack.

  7. Separate demonstrated impact from possible impact. “The test read this canary file across the intended boundary” is demonstrated. “The same path may expose other files” is a hypothesis until tested or established by the code path. Both can be useful when labelled honestly.

  8. Read the project’s security policy. Use its private channel, scope, encryption instructions, response expectations, and embargo rules. Check whether the project accepts AI-assisted reports and whether a bounty lane is open. Google’s October pause is a good example of why yesterday’s submission form may not be today’s route.

  9. Remove filler and unsupported severity language. The maintainer needs the path, evidence, impact, and conditions. They do not need a generic explanation of why memory safety matters or a generated score whose assumptions are hidden. If you provide a severity estimate, list each assumption that drives it.

  10. Have a qualified person sign the claim. The reviewer should be able to answer questions and rerun the test. Their name means the organisation accepts responsibility for the report’s accuracy and conduct. Human approval is weak when it means glancing at generated prose; it is meaningful when the approver owns the evidence.

This packet does not guarantee the project will agree. Different builds, threat models, and product boundaries can produce a reasonable dispute. It does make the disagreement concrete. The maintainer can challenge a revision, condition, observation, or impact instead of trying to disprove a fog of confident text.

A merged patch can add useful evidence, but it is not a substitute for coordination. Some projects do not want public fixes before private review because a patch can reveal the defect. Follow the project’s policy. Google’s earlier use of an accepted patch as one example of stronger proof worked in its program context; researchers should not push a sensitive fix publicly merely to satisfy a self-imposed checklist.

Build an internal review lane that rewards restraint

Teams adopting AI security review often celebrate total candidates because that number rises first. It is also the least useful measure. A healthy lane measures how many candidates survive reproduction, how much reviewer time each one costs, how often severity changes, and how many reports the receiver accepts as actionable.

Start with code your team owns. Seed a test set with resolved vulnerabilities, harmless suspicious patterns, unreachable code, and cases where a mitigating control wins. Run the agent against a fixed revision. The goal is to learn where it produces evidence and where it produces stories. Do not learn that by experimenting on maintainers who never agreed to join your evaluation.

Give every candidate a reason code at closure. Useful categories include not reachable, attacker lacks control, guard executes first, test does not reproduce, impact overstated, duplicate, and confirmed. These labels expose repeated failure modes. If half the queue closes because a framework middleware runs first, the next prompt or analysis stage should check that middleware before drafting a report.

Measure precision at the disclosure threshold, not only across raw discovery. A broad internal scan may tolerate many candidates if cheap automation rejects most of them before a person sees them. The report lane should be far stricter. One accepted report with a ten-minute maintainer reproduction is better than ten elaborate messages that each need an hour to dismiss.

Human review time must appear on the dashboard beside model cost. A run that costs $20 in inference and twelve engineer-hours did not cost $20. Track minutes spent reproducing, resolving build problems, checking reachability, correcting severity, and preparing disclosure. Those numbers reveal whether the agent reallocates expert attention or merely generates work for experts.

Limit who can promote a finding. A model should not mark its own report validated. The engineer who built the affected feature may understand reachability but miss the security consequence. The security reviewer may understand impact but need the maintainer’s build knowledge. For consequential findings, two perspectives are worth more than another model pass.

Keep disclosure credentials outside the scan environment. If reports go through a web form, private repository advisory, or email account, use a separate service identity or manual step. Log who approved the send, which evidence packet they reviewed, what text left the organisation, and when. That receipt protects both the receiver and the reporting team when details later change.

Set a receiver budget. One open-source project should not receive ten reports from one automated campaign on the same day. Group related findings when the policy permits, wait for initial feedback, and adapt to the project’s requested format. The maintainer’s response is evidence about your process. If the first three reports fail for the same reason, stop sending and repair the lane.

Reward researchers for closed false positives that improve the evaluation set. Otherwise incentives push every candidate toward an external ticket. A careful “no vulnerability because this guard runs first” can save dozens of future reviews. The organisation should value that work even though it creates no bounty headline.

What maintainers can do without building a second company

The burden should not fall entirely on maintainers, but a few intake rules can reduce avoidable work. Publish a SECURITY.md file with the private contact route, supported versions, expected evidence, response window, and out-of-scope categories. Clear rules give responsible researchers a target and make it easier to close submissions that ignore the basics.

Ask for structured fields before free-form narrative. Require the affected revision, build configuration, entry point, prerequisites, minimal reproduction, observed result, control case, and reporter contact. A model can still fill every box badly, but missing evidence becomes visible before a long essay occupies the queue.

Use an acknowledgement that distinguishes receipt from validation. “We received your report” should never imply “we confirmed a vulnerability.” Assign a private identifier and state the next step. That wording protects researchers from silence while preventing a generated claim from acquiring authority merely because a project answered it.

Triage by evidence, not claimed severity. A reproducible low-impact boundary failure deserves attention before an unreproduced critical claim. Severity can change after the mechanism is established. Evidence quality tells you whether there is a mechanism to score at all.

Rate limits and staged access can be reasonable. A new reporter might submit one active case until the project resolves it. Researchers with a history of accurate, well-formed reports can receive more capacity. The policy should be transparent enough that good-faith newcomers understand how to earn trust.

Templates can also ask whether AI assisted the report, but disclosure of tool use is not the main control. A human can submit nonsense without AI, and a model can assist excellent research. The decisive questions remain: Did the sender test the claim? Can the project repeat it? Does the observed behavior cross a security boundary?

Projects should preserve closed reports and their reasons. Repeated invalid patterns can become automated prechecks or documentation. Confirmed findings can become regression tests. A queue is painful when each case disappears after closure; it becomes useful when settled cases improve the next decision.

Downstream companies can help by funding maintenance, lending qualified triage time, and fixing issues in the projects they depend on. Sending a machine-generated report is easy. Supporting the human process that turns a verified defect into a safe release is the contribution open source needs.

The useful unit is a proved claim

Google’s October 2026 pause is a warning about throughput without responsibility. The program did not run out of candidate text. It ran into a verification bottleneck after automated submissions rose and most were invalid. Earlier evidence requirements had already signalled where the bottleneck lived.

Curl’s experience shows why this matters beyond one large company. Its bounty delivered 87 confirmed vulnerabilities and more than $100,000 in rewards, then ended after report quality deteriorated and the confirmation rate fell below 5 percent. A security channel can create real value and still become too expensive to operate.

Coding agents should remain in the research loop. They can map attack surfaces, trace values across files, draft tests, compare revisions, and help a researcher examine more code. The mistake is promoting their first coherent story directly into someone else’s urgent queue.

The durable boundary is an evidence gate. Discovery stays broad and private. Reproduction names the revision and records behavior. Validation checks the security boundary and realistic prerequisites. Disclosure happens through the project’s chosen route after a person accepts responsibility for the claim.

Count proved claims, useful fixes, and receiver time saved. Treat report volume as load, not success. The model may write the first draft, but the sender owes the proof.

For one practical security note each month, join the newsletter at Cyber Security in Plain English. One email per month.

Sources