GitHub Security Lab researcher Kevin Stubbings built custom AI-driven audit workflows, called taskflows, on top of the lab's open source Taskflow Agent, and used them to find and report more than 20 vulnerabilities in Android apps. The work shows how agentic AI can scale mobile security auditing, but also where it still falls short.
Two disclosed bugs illustrate the stakes. In OsmAnd, a navigation app with over 10 million Play Store downloads, an exported activity accepted intent extras that should have stayed restricted to an internal channel. Any app on the phone, with no permissions needed, could use those extras to silently import malicious settings, including swapping the map tile source for an attacker-controlled server. That would let the attacker log the coordinates of every tile a victim loaded and reconstruct their routes without the user noticing. The Wikipedia Android app had a different problem: a hostname check in its deeplink handler used endsWith() instead of matching the full domain, allowing a wikipedia:// link to point at a lookalike domain and load it inside the app's WebView. A second flawed check in the cookie manager could then hand over long-lived session cookies valid across every Wikimedia project.
Stubbings tailored the taskflows for mobile apps by separating mobile entry points from web or desktop ones and prompting the model to check for intent-based bugs, such as confused deputy issues and insecure broadcasts, that generic security prompts tend to miss. The AI proved better at finding bugs than judging how bad they are: it kept flagging low-severity issues even after being told not to, and misjudged real-world impact when mitigating factors quietly canceled out what looked like a working exploit. Every finding still needs a human reviewer who understands mobile apps. The taskflows are open source and free to run against any repository, though they require a GitHub Copilot license and can consume a large number of premium model requests even on a single medium-sized codebase.