# GitHub's 24 Android Bugs Show What an AI Security Review Needs

> GitHub Security Lab used targeted AI taskflows to find and report 24 Android vulnerabilities. The useful lesson is how to make an AI security review repeatable, testable, and subject to a human verdict.

- **Author:** Kubilay Tunca
- **Published:** 2026-10-01
- **Category:** For Developers
- **Tags:** AI Agents, Application Security, Android, Secure Development
- **Canonical URL:** https://cyber-security-in-plain-english.com/post/developers/news/github-android-taskflow-needs-human-verdict

---

A general code-review prompt can miss the one Android screen that another app is allowed to open. GitHub Security Lab changed the job. It first made an AI agent list those externally reachable screens and links, then asked narrower questions about each one. The resulting workflow found and reported 24 Android vulnerabilities, according to research GitHub published on 28 September 2026.

Two disclosed examples make the number tangible. One chain could let an unprivileged app quietly alter settings in the OsmAnd navigation app and expose map locations. Another combined weak link checking and cookie handling in the Wikipedia Android app into an account-takeover path. [Help Net Security independently described both findings on 29 September](https://www.helpnetsecurity.com/2026/09/29/github-ai-android-app-vulnerabilities/) and reported the same central limit: the agent was better at finding suspicious code than judging its real severity.

That limit is the useful part of the story. The result does not prove that a model can replace a mobile security researcher. It shows that a researcher can package part of a review method, run it repeatedly, and keep the final verdict with someone who understands the platform. The achievement lies beyond a clever prompt, in a review process with stages, records, and a place for doubt.

My position is that teams should adopt this pattern before they adopt the tool. Turn expert questions into a repeatable test package. Give the coding agent a narrow workspace and a fixed budget. Require evidence for every finding. Then make a qualified person decide whether the evidence describes a reachable defect or an attractive story about one.

## What GitHub reported on 28 September

GitHub Security Lab researcher Kevin Stubbings built Android-specific workflows on top of the lab's open-source Taskflow Agent. GitHub calls these workflows taskflows: YAML files that divide a larger investigation into ordered jobs for one or more agents. [The GitHub report says the Android taskflows had found and reported 24 vulnerabilities as of 28 September 2026](https://github.blog/security/how-we-found-24-android-vulnerabilities-using-our-open-source-ai-security-agent/). It presents two disclosed cases in detail rather than publishing all 24 at once.

The first stage maps the application. Android code may sit beside server code, desktop code, libraries, and build scripts in the same repository. A broad instruction to “find security bugs” leaves the model to decide what matters. The targeted workflow instead gathers mobile entry points, the places where data or control can cross into the app. These include exported activities, services, broadcast receivers, deep links, and web views that connect JavaScript to native functions.

A second stage classifies those entry points against mobile-specific failure modes. If the first stage finds an exported activity receiving an Android intent, the next stage asks about unsafe extras, confused-deputy behavior, and unprotected broadcasts. If it finds a web view, the questions change. Hostname validation, script execution, bridges into native code, and cookie scope become more important than the generic warnings that dominate many automated scans.

This separation sounds administrative, but it changes the search. The agent no longer has to remember every Android concern while wandering through a mixed repository. One task produces an inventory. Later tasks inspect each item with the right set of questions. The intermediate record also gives a human somewhere to intervene: an entry point can be added, reclassified, or removed before the expensive reasoning begins.

GitHub says the strict prompts run more than once, while a broader prompt gives the model room to connect behavior across components. Repetition covers the model's variability. The broader pass looks for logic chains that a fixed checklist may not name. The two modes have different jobs, and neither is treated as a verdict by itself.

The public route is accessible but not free in practice. GitHub's instructions use a Codespace and require a GitHub Copilot licence. The company warns that a medium-sized repository can take one or two hours, make many tool calls, and consume a substantial number of premium model requests. [The companion taskflows repository gives the same cost warning](https://github.com/GitHubSecurityLab/seclab-taskflows) and writes the candidate results to an SQLite database for review.

Those details matter more than a demo that finishes in thirty seconds. A real review has an input scope, a runtime, a bill, an output format, and a person expected to read the output. If any one of those is missing, “run the agent on the repository” is a wish rather than an engineering process.

## The model was given a map before it was asked to hunt

Security reviewers rarely begin by reading files in alphabetical order. They build a threat model. They ask what an outsider can reach, which data can cross a boundary, where identity changes, and which component makes the final decision. GitHub's Android taskflow turns part of that habit into an explicit pipeline.

Consider an app with 300,000 lines of Kotlin, Java, XML, and build configuration. Perhaps twelve components are exported to other apps. Three custom links open internal screens. Two web views can run scripts. Those seventeen surfaces deserve different attention from a colour theme, a local database migration, or a unit-test helper. A model that receives the whole repository and one broad prompt must discover that hierarchy while spending context and tool calls on the discovery.

The taskflow makes the hierarchy an artifact. The public repository says its mobile collection records exported Android components, deep links, URL schemes, iOS Universal Links, app extensions, web-view bridges, permissions, export status, and input filters. [Its README also says the mobile taskflows use container-backed source access and require Docker](https://github.com/GitHubSecurityLab/seclab-taskflows). The container helps package the job, although the agent framework itself warns that its Docker image is a deployment convenience rather than a security boundary.

That last distinction deserves a red line under it. Packaging and isolation are different properties. A container can make dependencies repeatable while still mounting a repository, credentials, and network paths that give the process broad authority. [The Taskflow Agent README states plainly that its Docker image is not intended as a security boundary](https://github.com/GitHubSecurityLab/seclab-taskflow-agent). A team running the workflow on private code still has to decide what the agent may read, where it may connect, and which credentials it receives.

The map also improves the prompts. “Review this activity for intent attacks” is a better question after another stage has proved that the activity is exported and recorded its declared permission. “Can this external value reach a sensitive operation?” becomes testable when the workflow carries the source and sink into the same task. The model spends less effort guessing the shape of the app and more effort tracing a named path.

There is a wider engineering lesson here. Prompts are often discussed as if they were the product. In this system, the product is the sequence around the prompts: collect, classify, inspect, repeat, record, and review. Change the model next month and that sequence can survive. Change a single giant prompt and the team may not know which part of the old behavior disappeared.

A good taskflow is closer to a test suite than a conversation. It has named inputs. It produces intermediate state. It can be checked before it runs. It should carry a version. Its failures should be visible rather than softened into a plausible paragraph. Most importantly, somebody can compare two runs and ask why the candidate set changed.

This does not make the output deterministic. GitHub used repeated runs precisely because models vary. It makes the surrounding process inspectable. The distinction is small in wording and large in operations: you cannot force a model to reason identically every time, but you can force every run to leave the same kinds of receipts.

## Two bugs show why component boundaries matter

The disclosed OsmAnd case began at a boundary between apps. Android uses intents as messages that ask a component to do something, such as open a screen or handle a link. An exported activity is available to callers outside its own application. That can be intentional. A navigation app needs to receive map links. The security question is whether an outside caller can also set values that were designed for a trusted internal path.

GitHub's report says OsmAnd's exported `MapActivity` handled settings files and accepted extras including `silent_import`, `replace`, and a list of settings types. Those values were expected to arrive through an internal service path, but Android does not reserve an intent extra for one trusted sender. Another app could set the same fields. According to GitHub, that allowed an app with no special permission to import replacement settings without a visible warning.

The practical effect came from what those settings controlled. The researcher changed the source used to fetch map tiles. Requests that normally supported the map could then carry tile coordinates to an outside server, while that server returned the expected image so the map continued to work. Route requests could reveal origins and destinations as well. [Cyber Security News reported the same chain on 29 September](https://cybersecuritynews.com/github-ai-finds-24-android-vulnerabilities/), including the role of the exported activity and the silent replacement settings.

Notice the shape of the defect. No single line says “share the user's location.” One component accepts a message. Several flags suppress the normal friction. A setting changes a network destination. Ordinary map requests then disclose where the user is looking or travelling. The security impact appears only when a reviewer follows the values across those steps.

The Wikipedia Android case crossed a different boundary. The app registered a `wikipedia://` link so a tap could open content inside the application. GitHub says the link handler checked whether a host name ended with the trusted domain string. A lookalike host whose name merely ended with that string could therefore pass. The app could load attacker-controlled content inside a web view that looked like part of Wikipedia.

A second suffix check affected cookie handling. GitHub's account says the combination could send long-lived Wikimedia session material to the lookalike page, creating an account-takeover path after a user tapped the crafted link. Help Net Security and Cyber Security News both reported the same two-part chain. The finding is useful because it joins two individually narrow checks into one outcome.

This is the kind of connection models can help search for. A reviewer may find the link parser during one pass and the cookie manager during another. A workflow can keep both candidates available and ask whether they compose. Models know many common application-programming interfaces and bug patterns, so they can propose links across files quickly. GitHub says that capability was strong enough to surprise its researcher when the agent reasoned about security-relevant APIs without receiving their source.

The same ability creates a trap. A coherent chain can still be impossible at runtime. A value may be overwritten later. An operating-system rule may block the caller. A release build may disable the component. A cookie may have a scope that the initial reading missed. The model's explanation can become more convincing as it adds steps, even while one untested assumption invalidates the whole path.

The disclosed examples survived researcher review and reporting. They should not be read as proof that every checked row in the output database is a vulnerability. They show what a useful candidate looks like: named entry point, controllable input, traced effect, realistic conditions, and evidence that the parts work together.

## Finding suspicious code is different from issuing a verdict

GitHub is unusually direct about the weakness of its result. The report says the models repeatedly returned low-severity issues even when prompted not to. They also estimated severity incorrectly when the real application had a mitigating behavior that was easy to miss. In one example, attacker-controlled data in external storage looked dangerous until internal storage took priority, leaving no working vulnerability.

That failure is not a cosmetic scoring error. Severity determines who gets interrupted, whether a release pauses, how a maintainer communicates with users, and how quickly a fix has to move. A tool that finds an interesting path but inflates its impact can consume more security time than it saves. After several false emergencies, maintainers stop reading.

The report recommends expert review and, where possible, a proof of concept that exercises the original code. In plain terms, make the claim collide with the program. Can the supposedly external caller reach the component? Does the proposed value survive parsing? Does the sensitive operation occur under the stated conditions? Does a mitigating branch win first? A debugger, test harness, emulator, or controlled reproduction can answer questions that a fluent explanation cannot.

Evidence has levels. A model pointing to a suspicious suffix check is a lead. A trace showing attacker-controlled input reaching that check is stronger. A test that fails on the vulnerable revision and passes after the proposed fix is stronger again. A human who understands Android then decides whether the conditions and impact match the report. Teams should record the level rather than flatten every output into “AI found a vulnerability.”

A useful internal status model is `candidate`, `reproduced`, `validated`, and `disclosed or fixed`. The first label belongs to the agent. The later labels require evidence and named ownership. A candidate may be worth investigation without being suitable for a ticket marked critical. That wording protects both urgency and credibility.

Human review does not mean asking a person to read 500 polished essays. The workflow should help the reviewer reject weak candidates quickly. Each record needs the affected revision, entry point, source, sink, assumptions, relevant code locations, attempted reproduction, observed result, and known mitigating checks. If the agent cannot identify those fields, it has produced a hunch, not a review package.

The person also needs the negative evidence. Suppose a reproduction failed because the activity was not exported in the release manifest. That result should remain attached to the candidate. Otherwise a later run may rediscover the same pattern, write a fresh explanation, and charge the team to investigate it again. Closed findings are training data for the process even when no model is trained on them.

Here the review method starts paying compound interest. A rejected path becomes a regression case for the taskflow. A confirmed bug becomes a fixture that future model or prompt versions should still find. A severity correction becomes an explicit question added before triage. Over time, the team builds an evaluation set from its own code and decisions rather than judging the tool by a vendor's headline number.

The model can assist with discovery. The test can establish behavior. The researcher owns the verdict. Moving those responsibilities into one box makes the output look simpler and the system less trustworthy.

## Treat the taskflow like code that touches valuable code

An audit agent is not a passive document reader. It may clone repositories, query hosting APIs, build databases, run analysis tools, and write artifacts. The public Taskflow Agent supports Model Context Protocol servers, which are tool connections the agent can call. Its configuration can pass environment variables into those server processes. This is useful machinery, and it deserves the same review as any build tool that handles source and credentials.

Start with the source revision. Pin the agent framework and taskflow repository to reviewed commits instead of pulling a moving default into a sensitive environment. Pin container images by digest when the runner supports it. Record the model name and configuration used for the run. GitHub's agent can produce a machine-readable manifest with task status, model choice, timing, and named outputs; [that manifest behavior is documented in the framework README](https://github.com/GitHubSecurityLab/seclab-taskflow-agent). Keep it with the finding set.

Credentials should be narrow and temporary. A token that can read one audit repository should not also administer the organisation or publish releases. If a task only needs public code, do not provide a private token out of convenience. The framework documents an environment-variable denylist for tool processes, but a denylist is a backstop rather than a reason to launch the parent process with every secret from a developer's shell.

Network access deserves an explicit answer. The computer still has a way out if the container can contact arbitrary hosts through the host network, a browser tool, a package manager, or an attached server. An audit of private code may need the model endpoint and a small set of repository services. It does not automatically need the whole internet. Route those destinations through a controlled path and keep the network record outside the agent's writable workspace.

Repository mounts should default to read-only. Write findings into a separate output directory. Do not give the audit process a deployment key, package-publishing token, production configuration, or live customer dataset. If reproduction requires building the app, use a disposable worker and fake accounts. A security review should not create a second incident while looking for the first one.

Tool permissions need the same separation. Source navigation, syntax queries, and local builds fit the ordinary path. Creating issues, opening pull requests, contacting maintainers, or publishing an advisory are external effects. Keep those tools out of the discovery run, or place them behind a separate approval and identity. A model that can find a plausible bug should not be able to announce it before a researcher has checked the claim and coordinated disclosure.

Budget is also a boundary. GitHub warns that the mobile audit can use many premium requests. Set a ceiling for model calls, wall-clock time, repository size, and concurrent jobs. A failed stage should stop or checkpoint rather than retrying indefinitely. Cost records belong beside security records because an uneconomic scan will be disabled, quietly skipped, or restricted to the smallest repositories.

Finally, inspect what the workflow itself changed between runs. A new prompt may improve Android coverage while flooding the queue with path traversal reports. A model upgrade may alter tool behavior. A renamed database field may break the triage view. Review taskflow changes through pull requests, run them against the fixed evaluation set, and require an owner to accept the new precision, recall, runtime, and cost profile.

The Secure Harness makes the same argument for coding agents: autonomy becomes dependable when permissions, evidence, and stop points are designed around the action. An audit agent belongs inside that boundary too. Security purpose does not grant security properties.

## Build a review lane before scanning the whole organisation

The fastest way to waste this capability is to point it at every repository and send the output to a shared inbox. Volume arrives before trust. Teams then face a queue whose cost and quality they have not measured, while maintainers learn that “AI security finding” means extra work with uncertain value.

Begin with one application whose owners can explain its architecture. Pick a revision with known resolved vulnerabilities and a clean set of ordinary code around them. The known cases tell you whether the workflow can find relevant defects. The clean surface tells you how much noise reaches a person. Neither measure is perfect, but both are better than counting generated reports.

Define the review lane before the first broad run. One person owns execution. A mobile security reviewer owns technical validation. The application team owns fixes. A disclosure owner handles contact with an outside project when needed. The tool does not own any of those decisions, and a database checkmark does not move a candidate between stages by itself.

Set a service level based on evidence. A reproduced path with a clear sensitive effect can interrupt planned work. A pattern match without a working path enters ordinary triage. A speculative concern missing its caller, input, or sink goes back to the workflow or closes. This keeps severity language tied to what the team knows rather than how alarming the generated prose sounds.

Measure reviewer minutes per confirmed finding. Model cost is visible on an invoice; human triage cost hides in calendars. If ten hours of scanning creates forty candidates and one real issue, measure model spend plus reviewer time per useful fix instead of cost per scan. Record false positives by reason so the next taskflow version can remove repeated waste.

Run the first lane on pull requests only after the offline review is stable. A pull-request gate has a short patience budget. It should comment only when evidence is strong enough to help the author and the suggested action is specific. Lower-confidence candidates can run nightly or enter a private security queue where a specialist has room to investigate.

Do not let the new lane replace existing controls. Dependency scanning, compiler warnings, static analysis, tests, code review, and platform hardening catch different classes of failure. The taskflow adds structured exploration, especially for logic paths spread across components. Removing a deterministic check because an agent sometimes notices the same pattern trades a reliable guard for an expensive opinion.

Aim for a better allocation of attention rather than full automation. Let machines enumerate surfaces, trace repetitive paths, and propose cross-file connections. Let tests settle behavior. People can spend their judgement on exploitability, impact, fix quality, and responsible communication.

## What to do this week

A team can test this pattern without turning its security process upside down. The first pass should be small enough to stop, inspect, and repeat. Use a repository you are authorised to assess, and keep any discovery private until the project owner has validated it.

1. **Choose one owned application and one fixed revision.** Record the commit hash before the run. Prefer an Android application with an engaged maintainer and a staging build. Do not begin with the company's largest monorepo or with somebody else's project merely because it is public.

2. **Write the expected attack-surface inventory first.** Ask the application's developer to name exported components, deep links, web views, bridges, and sensitive data paths. Then compare that list with the agent's collection stage. A missing entry point is a pipeline failure even if the later prose looks excellent.

3. **Run in a disposable, narrow environment.** Mount source read-only, use a separate output directory, provide only the token scope the run needs, and restrict network destinations. Pin the workflow revision and record the model configuration. Treat GitHub's container as packaging, not as the final wall.

4. **Set a budget and a stop condition.** Cap requests, runtime, retries, and concurrent tasks before execution. Decide what happens when collection is incomplete or one tool fails. An agent that silently continues with half an inventory can produce a clean report about the wrong application surface.

5. **Require a finding packet.** For every candidate, collect the source revision, entry point, attacker-controlled value, path to the sensitive operation, assumptions, code locations, attempted reproduction, and observed result. Reject prose-only findings from the urgent queue.

6. **Reproduce in a controlled build.** Use an emulator, test account, and non-sensitive data. Confirm each required condition. Save the test, trace, or debugger evidence that supports the verdict. If a mitigating branch blocks the path, record that result so the same candidate stays closed on the next run.

7. **Have a mobile specialist assign impact.** The model may suggest possibilities, but the reviewer decides reachability and severity. Separate “interesting bug pattern” from “security vulnerability” and from “critical incident.” Each label should carry an owner and evidence threshold.

8. **Turn outcomes into an evaluation set.** Keep confirmed cases, false positives, missed entry points, and severity corrections. Run the next workflow or model version against that set before promotion. Compare useful findings, reviewer minutes, cost, and runtime rather than comparing the fluency of reports.

9. **Keep external actions in a separate lane.** Issue creation, pull requests, maintainer contact, and public disclosure happen only after validation. Use the project's security policy and coordinated disclosure channel. Discovery speed does not shorten the maintainer's right to verify and repair.

10. **Decide where the workflow earns a permanent place.** If it reliably maps attack surfaces but struggles with impact, use it for mapping. If one narrow class produces strong evidence, gate that class and leave the rest in research mode. Adoption should follow measured value, not the broadest claim the tool can make.

After two or three controlled runs, the team should be able to answer plain questions. What does one scan cost? How many candidates need a person? Which bug classes survive reproduction? Which surfaces does collection miss? What changed when the model or taskflow changed? If those answers are unavailable, expansion is premature.

## The durable result is the review package

Twenty-four reported Android vulnerabilities are a serious research result. The OsmAnd and Wikipedia examples also show why logic defects remain hard: impact emerges across component boundaries, trusted assumptions, and ordinary features that behave badly together. A model can search those combinations at a scale that helps a skilled reviewer.

The number alone is a poor procurement plan. GitHub's own account says the system still returned low-impact candidates, misread mitigating behavior, and needed a mobile security researcher to judge the findings. The public workflow can also take hours and consume many paid requests. Those are not footnotes. They define the work required to use the capability responsibly.

The part worth copying is the shape. GitHub turned expert attention into staged, shareable taskflows. The run leaves structured output instead of ending with a chat transcript. Repeated passes cover variability. Concrete reproduction and human review stand between a model's claim and a security verdict. The workflow is open for inspection and change.

That shape will outlast this month's model. A team can replace one model, add a platform-specific question, tighten a tool permission, or improve a reproduction step without discarding the whole method. It can test the change against findings that its own reviewers have already settled. Security knowledge becomes a maintained artifact rather than advice that disappears when an expert closes a terminal.

Use the agent to widen the search, not to lower the standard of proof. Give it a map, a boundary, a budget, and a required evidence packet. Keep the decision with the person who can make the program demonstrate the claim.

For one practical security note each month, join the newsletter at Cyber Security in Plain English. One email per month.

## Sources

- [GitHub Blog: How we found 24 Android vulnerabilities using our open source AI security agent](https://github.blog/security/how-we-found-24-android-vulnerabilities-using-our-open-source-ai-security-agent/), accessed 2026-10-01
- [GitHub Security Lab: seclab-taskflows repository](https://github.com/GitHubSecurityLab/seclab-taskflows), accessed 2026-10-01
- [GitHub Security Lab: Taskflow Agent repository](https://github.com/GitHubSecurityLab/seclab-taskflow-agent), accessed 2026-10-01
- [Help Net Security: GitHub's AI agent found 24 Android app vulnerabilities](https://www.helpnetsecurity.com/2026/09/29/github-ai-android-app-vulnerabilities/), accessed 2026-10-01
- [Cyber Security News: GitHub AI Security Agent Finds 24 Android Vulnerabilities Including Account Takeover Flaws](https://cybersecuritynews.com/github-ai-finds-24-android-vulnerabilities/), accessed 2026-10-01

---

## About the author

Kubilay Tunca — Senior Full Stack Developer and Author. Founded Cyber Security in Plain English to translate complex security concepts into clear, practical advice, and writes the accompanying books on security, privacy, secure development, and AI systems.

## Books by this author

- **The Digital Fortress** — Your Everyday Guide to a Safer Digital Life. A warm, plain-English guide for people with real lives and finite patience. Learn the handful of habits that genuinely protect your money, accounts, and family, and get honest permission to ignore the rest. [Amazon](https://buy.cyber-security-in-plain-english.com/digital-fortress) · [Details](https://cyber-security-in-plain-english.com/books/the-digital-fortress)
- **The Anonymity Playbook** — Digital Survival for Whistleblowers, Journalists, Activists, and Everyone Else. A practitioner’s field manual for journalists protecting sources, whistleblowers, and activists. It explains how the surveillance actually works, what each technique costs you, and exactly where it fails. [Amazon](https://buy.cyber-security-in-plain-english.com/anonymity-playbook) · [Details](https://cyber-security-in-plain-english.com/books/the-anonymity-playbook)
- **Secure Software Development** — Practical patterns for building secure software. A hands-on security guide for developers and IT professionals who ship real software. Build, deploy, and maintain secure systems without slowing down or drowning in theory. [Amazon](https://buy.cyber-security-in-plain-english.com/secure-software-development) · [Details](https://cyber-security-in-plain-english.com/books/secure-software-development)
- **The Secure Harness** — Shipping Production Code with AI Coding Agents. A calm, practical guide to letting agents do useful work inside boundaries you set, enforce, and audit. Ships with 15 copy-pasteable artifacts: hook scripts, permission configs, release gates, and MCP templates. [Amazon](https://buy.cyber-security-in-plain-english.com/secure-harness) · [Details](https://cyber-security-in-plain-english.com/books/the-secure-harness)
- **The AI Native Engineer** — Build, Evaluate, and Ship AI Systems That Work in Production. Sixteen hands-on chapters, one real product. Grow it from a single model call into a retrieved, tool-using, observable, production-grade system, with evaluation treated as a habit from the first feature. [Amazon](https://buy.cyber-security-in-plain-english.com/ai-native-engineer) · [Details](https://cyber-security-in-plain-english.com/books/the-ai-native-engineer)

Full catalogue with contents and intended audience: https://cyber-security-in-plain-english.com/books

_As an Amazon Associate I earn from qualifying purchases. Buying through these links costs you nothing extra and helps pay for the blog._
