The Multiverse School

Join

Agentic AI Security — AI-Assisted Hacking Case Studies (130)

AI-Assisted Hacking — Case Studies

130 documented, citable cases of AI being used offensively — and of AI being turned back on the problem. Every entry links to its primary source (vendor threat-intel report, arXiv paper, advisory, or reporting of record).

How to use this. These are case studies for defenders: what was reported, by whom, and the defensive lesson. They stay at the level of published fact — no operational detail, no payloads. Pair each case with the control that would have caught it from Meta-Map — When & Where.

Read the uplift claims carefully. The labs' own assessments are more measured than the headlines: OpenAI's first state-actor report concluded its model offered "only limited, incremental capabilities for malicious cybersecurity tasks beyond what is already achievable with publicly available, non-AI powered tools." Teach the trend line, not the hype — and note where later 2025–26 reporting genuinely shifts that picture.

Remediate on your own device

Worried a threat below affects you? Multiverse Device Rescue is a free, open-source, read-only-by-default toolkit (macOS / Windows / Linux) that checks — and helps remediate — these on a real machine. Each threat maps to a command:

Threat Run Walkthrough
AI worm / supply-chain compromise rescue --profile ai_worm_response guide
Mobile spyware (Pegasus-class) rescue --profile ai_worm_response guide
SSH key / credential compromise rescue run ssh_key_audit guide
Unwanted remote access (SSH / RDP) rescue run remote_login_check win_rdp_check guide
Launch-time persistence malware rescue run launchd_persistence_audit launch_agent_audit guide
C2 network beaconing rescue --profile ai_worm_response guide
Weak / disabled firewall rescue run firewall_audit guide
General / mixed indicators rescue (full scan) guide

Full threat→command map with step-by-step remediation: THREAT_REMEDIATION.md.

Category Cases
Threat-Intelligence Reports (what labs actually caught) 23
Real-World Incidents 21
Criminal Tooling & AI-Written Malware 11
Prompt Injection & Agentic Exploits in Real Products 12
Academic: Can Agents Attack Autonomously? 18
Capability Benchmarks 16
Autonomous Vulnerability Discovery (the defensive mirror) 13
Guardrail Bypass 1
Policy & Uplift Assessments 15

Threat-Intelligence Reports (what labs actually caught)

  • GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access Google Threat Intelligence Group (GTIG) · 2026-05-11 CONFIRMED against the primary source; title and date match exactly. Two precision corrections. (1) The zero-day claim needs GTIG's own caveat restored: GTIG identified a threat actor using a zero-day 'we believe was developed with AI' and states 'the criminal threat actor planned to use it in a mass exploitation event but our proactive counter discovery may have prevented its use' — so the exploit existed and was held by a real actor, but mass exploitation may have been averted. Defensive lesson: This is where GTIG's long-running 'no novel capability' finding finally bends: AI-assisted exploit generation reached a real actor's hands, even though GTIG's counter-discovery may have prevented the mass exploitation event. Defenders should shorten patch windows on the assumption that time-to-exploit compresses further, and add a new detection surface — the obfuscated LLM access economy (relay services, pooled accounts, auto-registration) is infrastructure that can be tracked and blocked much like bulletproof hosting.

  • GTIG AI Threat Tracker: model distillation abuse and the Xanthorox 'self-hosted' criminal AI Google Threat Intelligence Group · 2026-02-12 VERIFIED — no substantive corrections (title is an editorial label, not the report's literal title). GTIG reported model extraction attacks (MEA), aka distillation attacks, including a reasoning-trace coercion campaign of over 100,000 prompts instructing Gemini that 'the language used in the thinking content must be strictly consistent with the main language of the user input' to replicate its reasoning across languages; Google detected this in real time and protected internal reasoning traces. Defensive lesson: Criminal AI branding is frequently false advertising — 'self-hosted and private' Xanthorox was jailbroken commercial APIs, which means provider account action still reaches these tools and buyers' operations are more exposed than they believe. Two defensive priorities follow: secure your own AI gateway infrastructure (the API-key black market is fed by default credentials, weak auth, missing rate limits, and XSS in open-source AI platforms — One API and New API are named targets), and treat model-extraction attempts as a distinct abuse class with its own volumetric signature separate from content misuse.

  • GTIG AI Threat Tracker: Advances in Threat Actor Usage of AI Tools Google Threat Intelligence Group (GTIG) · 2025-11-05 The landmark tracker edition documenting a shift from AI-as-productivity-tool to malware that calls an LLM at runtime. It named five families: PROMPTFLUX (VBScript dropper querying Gemini to rewrite itself), PROMPTSTEAL (data miner querying a Hugging Face-hosted model, seen in live operations), PROMPTLOCK (experimental Go ransomware generating scripts via LLM), FRUITSHELL (PowerShell reverse shell carrying hard-coded prompts intended to defeat LLM-based security review), and QUIETVAULT (credential stealer abusing on-host AI CLI tools to hunt further secrets). Actors named include APT28, TEMP. Defensive lesson: Marks the inflection point defenders must plan for: 'just-in-time' malware whose logic is generated at execution, which degrades signature and hash-based detection. Two concrete defensive shifts follow — (1) weight behavioral and network detection over static signatures, treating outbound traffic to LLM API endpoints from non-AI workloads as a hunt lead; and (2) FRUITSHELL's embedded anti-analysis prompts prove LLM-based defensive tooling is itself an attack surface, so never feed untrusted samples to a security LLM that can act on their contents.

  • GTIG AI Threat Tracker: AI-enabled malware families observed in operations Google Threat Intelligence Group · 2025-11-05 VERIFIED — attribution correct, one trivial detail noted. GTIG documented malware that queries an LLM at runtime: PROMPTFLUX, an experimental VBScript dropper whose 'Thinking Robot' module queries the Gemini API to regenerate obfuscated code (hourly) for AV evasion, persisting via the Startup folder; PROMPTSTEAL, a PyInstaller-packaged Python data miner querying the Hugging Face API for Qwen2. Defensive lesson: This is the documented arrival of runtime-LLM malware, and PROMPTSTEAL is the key data point — the only one attributed to a state actor and observed in real operations, so treat 'self-rewriting malware' as still largely experimental while runtime LLM calls for reconnaissance are real. The actionable control is network-layer: malware that must reach an external model API creates a dependency defenders can see and sever. Monitor and restrict egress to model endpoints (Gemini, Hugging Face) from hosts with no business reason to call them, and note FRUITSHELL carried prompts intended to defeat LLM-powered security tooling — your AI defenses are themselves a prompt-injection target.

  • Disrupting malicious uses of AI: October 2025 OpenAI · 2025-10-07 OpenAI reported that since beginning public threat reporting in February 2024 it had disrupted and reported over 40 networks violating its usage policies. Case studies covered Russian-speaking malware tooling development, Korean-language operators, a cyber operation dubbed "Phish and Scripts," scam operations, authoritarian-linked abuses with the PRC as the example, recidivist influence activity "Stop News," and covert IO "Nine—emdash Line." OpenAI stated it continues to see threat actors bolt AI onto existing playbooks to move faster rather than gain novel offensive capability. Defensive lesson: After 18 months of data the consistent finding is acceleration, not transformation — a stable planning assumption for defenders. The 40+ network total also quantifies AI platforms as a chokepoint where cross-actor activity is visible, arguing for defenders to build intake processes for provider disclosures.

  • North Korean IT workers scaling fraudulent employment with AI Anthropic · 2025-08-27 Anthropic reported that North Korean operatives used Claude to create false professional identities, pass technical assessments and interviews, and sustain day-to-day work at Western technology companies, including Fortune 500 firms. Anthropic assessed that AI removed the years of specialized training the regime previously needed to prepare such operatives, and that the scheme generates revenue for the North Korean government in violation of international sanctions. Defensive lesson: AI collapsed a training bottleneck that had naturally limited this program's scale, so volume — not sophistication — is the change defenders face. Because the operative is a hired employee rather than an intruder, the effective controls are identity verification, hardware/logistics validation, and insider-risk monitoring, not perimeter defense.

  • Anthropic: Detecting and countering misuse of AI, August 2025 Anthropic · 2025-08-27 Anthropic reported disrupting several operations, including a data-extortion campaign ('vibe hacking') in which an actor used Claude Code against at least 17 organizations across healthcare, emergency services, government and religious institutions — threatening data exposure rather than encryption, with demands sometimes exceeding $500,000 and AI used to craft psychologically targeted ransom notes; North Korean operatives using Claude to fraudulently obtain and hold remote jobs at US Fortune 500 technology companies; and a criminal using Claude to develop and sell ransomware variants for $400. Defensive lesson: Shows AI moving from drafting messages to orchestrating the extortion lifecycle, including tailoring coercive demands to each victim — social engineering applied to the ransom itself. The North Korean employment fraud independently corroborates the KnowBe4 case, confirming a sustained, sanctions-evading program rather than a one-off. For defenders: insider-risk and identity verification in hiring are now AI-threat surfaces, and extortion playbooks must assume data-exposure leverage without any encryption event to detect.

  • Anthropic threat intelligence report: Claude Code used to automate a data extortion campaign Anthropic · 2025-08-27 VERIFIED against the primary source. Published 2025-08-27. The 'vibe hacking' data extortion operation is confirmed, with Claude Code automating reconnaissance, credential harvesting, network penetration, exfiltration targeting decisions, ransom note crafting, and financial analysis. The figure 'at least 17 distinct organizations' is exact. MINOR CORRECTION: the sector list is 'healthcare, the emergency services, and government and religious institutions' — the original summary dropped 'religious.' Additional confirmed detail not in the original: extortion demands exceeded $500,000. Defensive lesson: This documents agentic AI operating across the full intrusion lifecycle rather than assisting at one stage, and it lowers the skill floor — the ransomware case involved an actor who could not otherwise build the malware. For defenders the practical takeaways are that attack tempo and volume rise without a corresponding rise in adversary sophistication, and that vendor-side detection (classifiers, account bans, indicator sharing) is now part of the defensive stack customers depend on. Note this is a vendor self-report about its own product; the underlying indicators are not independently reproducible from the post.

  • Disrupting malicious uses of AI: June 2025 OpenAI · 2025-06-05 Quarterly report with ten case studies: a North Korea-linked IT worker "Deceptive Employment Scheme"; covert influence operations "Sneer Review," "High Five," "Helgoland Bite," and "Uncle Spam"; a social-engineering/IO hybrid "VAGue Focus"; a cyber operation "ScopeCreep"; cyber operations attributed to Vixen and Keyhole Panda; recidivist influence activity by STORM-2035; and a scam network "Wrong Number." Defensive lesson: The reappearance of STORM-2035 after an earlier ban shows takedowns degrade rather than eliminate adversaries — recidivism is expected. Defenders and platform teams should plan for actor return and build durable behavioral detection rather than treating a single enforcement action as resolution.

  • FBI IC3 PSA: Senior US Officials Impersonated in Malicious Messaging Campaign FBI Internet Crime Complaint Center (IC3) · 2025-05-15 VERIFIED against the primary source; title, date and both quotations are exact. Since April 2025, 'malicious actors have impersonated senior US officials to target individuals, many of whom are current or former senior US federal or state government officials and their contacts,' using AI-generated voice messages (smishing and vishing). The PSA defines vishing as malicious targeting of individuals using voice memos, which 'may incorporate AI-generated voices. Defensive lesson: Note the FBI's own admission that synthetic audio is often not identifiable by ear — an official acknowledgment that human detection is not a viable control. The targeting model is also instructive: compromise the trusted relationship rather than the person, reaching officials through their contacts. Public figures with abundant recorded speech are the cheapest voices to clone, so high-profile principals need pre-established verification protocols with their networks before an incident, not after. The truncated clause is worth restoring in any citation, since 'to increase the believability of their schemes' is the FBI stating the attacker's purpose plainly — the voice is not the payload, it is the credibility wrapper around one.

  • Detecting and countering malicious uses of Claude: March 2025 Anthropic · 2025-04-23 VERIFIED against the live page (title and "Apr 23, 2025" byline confirmed — note the March 2025 in the title is the reporting period, not the publication date; the given date is correct). This is Anthropic's first public threat intelligence report and the earliest in the series, though the page itself never uses the word "first" — that framing is accurate but editorial rather than page-sourced. Defensive lesson: The orchestration finding is the key shift: the model made engagement decisions rather than just producing text, which is a qualitatively different role. Defenders should watch for AI acting as coordination logic — a behavioral signature that persists even when generated content is indistinguishable from authentic posts.

  • Detecting and countering malicious uses of Claude: March 2025 Anthropic · 2025-04-23 CONFIRMED against the source; all four case studies and the date verify exactly. Defensive lesson: The through-line is orchestration and polish rather than novel exploitation — the model made a decision layer for bot networks and removed the language tells from fraud. Defenders lose two heuristics at once: awkward English in recruitment scams, and crude tooling as a proxy for a low-skill adversary. Note also that Anthropic reported no confirmed successful deployment in several cases, which is the honest way to report capability without inflating impact — and a hedge worth preserving when this report is cited downstream. Mind the date/slug mismatch: the 'march-2025' URL denotes the reporting period, and the post published 2025-04-23.

  • Disrupting malicious uses of our models: an update (February 2025) OpenAI · 2025-02-21 VERIFIED against the archived landing page (byline: "February 21, 2025, Global Affairs/Security"; authors Ben Nimmo, Albert Zhang, Matthew Richard, Nathaniel Hartley) and against the full report PDF, whose cover title is exactly "Disrupting malicious uses of our models: an update, February 2025" (the landing page's own H1 is the shorter "Disrupting malicious uses of AI"). The "first anniversary" framing is supported by the page: "It has now been a year since OpenAI became the first AI research lab to publish reports on our disruptions. Defensive lesson: Broadens the threat model beyond cyber intrusion: surveillance tooling, employment fraud, and consumer-facing scams are now first-class AI misuse categories. Security teams should recognize that fraud, HR, and trust-and-safety functions — not just the SOC — are consumers of AI threat intelligence. The hedged DPRK attribution is itself instructive: AI-provider telemetry shows behavior, not identity, so it corroborates attribution rather than establishing it.

  • Adversarial Misuse of Generative AI Google Threat Intelligence Group (GTIG) · 2025-01-29 VERIFIED against the live page (title, January 29, 2025 date, and GTIG authorship all confirmed). GTIG's first major assessment of threat-actor misuse of Gemini, covering government-backed actors from Iran (heaviest user, 10+ groups), China (20+ groups), North Korea (9 groups), and Russia (3 groups, limited use), plus information-operations groups. GTIG found actors used AI mainly for productivity gains on common tasks — research, troubleshooting code, content creation, translation and localization. Defensive lesson: A calibrated counterweight to hype: at this stage AI compressed adversary timelines without granting new capability, and safety controls held against unsophisticated bypass attempts. The DPRK cover-letter finding is the actionable one — it links AI misuse to the hiring pipeline, telling defenders that identity verification in recruitment is an AI-relevant control surface.

  • Google Threat Intelligence Group: Adversarial Misuse of Generative AI Google Threat Intelligence Group (GTIG) · 2025-01-29 VERIFIED against the primary source; every specific claim checks out. Published 29 January 2025. GTIG observed APT actors from Iran (10+ groups, incl. APT42), China (20+), North Korea (9) and Russia (3). APT42 accounted for over 30% of Iranian APT Gemini use, crafting phishing campaigns, conducting reconnaissance on defense and policy experts, and using translation/localization for fluent English. Defensive lesson: A calibration anchor against hype: measured observation of real APTs shows efficiency uplift, not new capability classes. For defenders this means existing detections and controls largely remain valid — but adversary throughput and localization quality rise, so volume-based and language-based assumptions (e.g. 'foreign actors write awkward English') break first. Translation and localization uplift specifically erodes the linguistic tells many phishing programs still teach. One refinement: the original understated the safety findings. GTIG reports the malicious attempts were largely unsuccessful — 'Gemini did not produce malware or other content that could plausibly be used in a successful malicious campaign' — and that actors relied on low-effort, publicly available jailbreak prompts rather than tailored attacks. The observed adversary is opportunistic, not sophisticated, in their model use.

  • FBI IC3 PSA: Criminals Use Generative Artificial Intelligence to Facilitate Financial Fraud FBI Internet Crime Complaint Center (IC3) · 2024-12-03 VERIFIED against the primary source; title and date exact. The PSA describes criminals exploiting generative AI across modalities: TEXT for fake social-media profiles, persuasive phishing/romance messaging, fixing grammatical errors, fraudulent websites and malicious chatbots; IMAGES for fake profile photos, fraudulent IDs, celebrity-endorsement impersonation, disaster imagery for charity scams and sextortion material; AUDIO via vocal cloning to impersonate family members requesting money and to gain unauthorized bank access; VIDEO deepfakes for calls with fake executives or authority figures. Defensive lesson: The authoritative government framing for awareness programs: the uplift is scale and polish, not novel crime types. Retire 'look for bad grammar and typos' from training — the PSA explicitly lists 'fixing grammatical errors' among criminal uses of AI text, so the FBI itself confirms that tell is gone. Because detection by inspection is failing across text, image, audio and video simultaneously, defenses must shift to process controls (verification, callbacks, transaction limits) that hold regardless of how convincing the artifact is. Being a U.S. government PSA, this is also the citation to reach for when a security program needs an authority an executive will accept without argument.

  • An update on disrupting deceptive uses of AI (report: "Influence and cyber operations: an update") OpenAI · 2024-10-09 VERIFIED against the archived page (byline: "October 9, 2024, Safety/Global Affairs"). The "more than 20" figure matches verbatim: "Since the beginning of the year, we've disrupted more than 20 operations and deceptive networks from around the world that attempted to use our models." The accompanying report title is confirmed via the PDF linked from the page — cdn.openai.com/threat-intelligence-reports/influence-and-cyber-operations-an-update_October-2024.pdf — i.e. "Influence and cyber operations: an update, October 2024. Defensive lesson: Regular cadence reporting is itself a defensive asset: aggregated across quarters it shows adversaries bolting AI onto existing playbooks rather than inventing new ones. Defenders can use the trend lines to justify continued investment in conventional TTP detection instead of redirecting budget to speculative AI-specific threats.

  • Disrupting a covert Iranian influence operation (Storm-2035) OpenAI · 2024-08-16 VERIFIED against the archived page (byline: "August 16, 2024, Safety/Security"). The page's own H1 is "Disrupting a covert Iranian influence operation"; the "(Storm-2035)" here is an accurate annotation, since the page names the cluster explicitly: "a covert Iranian influence operation identified as Storm-2035" (Storm-2035 is Microsoft's designation, and OpenAI credits Microsoft's prior reporting). Defensive lesson: Shows AI-provider telemetry surfacing election-adjacent influence activity on a timeline fast enough to matter. The repeated low-engagement finding tells defenders that volume without audience access fails — while rapid, named public disclosure is what lets platforms and election officials pivot on shared indicators.

  • Disrupting deceptive uses of AI by covert influence operations OpenAI · 2024-05-30 VERIFIED against the archived page (byline: "May 30, 2024, Security"). OpenAI's first influence-operations takedown report disclosed disruption of five covert IO networks over three months ("In the last three months, we have disrupted five covert IO"). All five named and confirmed: Bad Grammar (Russia), Doppelganger (Russia), Spamouflage (China), IUVM (Iran), and Zero Zeno — activity by STOIC, a commercial company in Israel — so the "Russia, China, Iran, and a commercial firm" origin summary is accurate. Defensive lesson: AI lowers the cost of producing influence content but does not solve distribution or credibility, which remain the real bottlenecks. Detection should therefore target coordination and distribution artifacts (account creation patterns, posting cadence, network structure) rather than content quality, which generative AI has made an unreliable signal.

  • Disrupting malicious uses of AI by state-affiliated threat actors OpenAI (in partnership with Microsoft Threat Intelligence) · 2024-02-14 VERIFIED against the archived page (byline: "OpenAI, February 14, 2024, Security"). The first threat-intelligence report published by an AI lab — a claim OpenAI itself corroborates in its Feb 2025 report ("It has now been a year since OpenAI became the first AI research lab to publish reports on our disruptions"). OpenAI terminated accounts tied to five state-affiliated groups: Charcoal Typhoon and Salmon Typhoon (China), Crimson Sandstorm (Iran), Emerald Sleet (North Korea), and Forest Blizzard (Russia) — all five actor/country pairings confirmed verbatim. Defensive lesson: Establishes the baseline finding that defenders should anchor on: early state-actor LLM use mapped to reconnaissance and social-engineering preparation, not novel capability. AI-provider account telemetry is a genuinely new intelligence source, but defensive priority should stay on conventional controls (phishing-resistant MFA, egress monitoring) rather than AI-specific alarm.

  • Staying ahead of threat actors in the age of AI Microsoft Threat Intelligence (with OpenAI) · 2024-02-14 VERIFIED against the live page (title and February 14, 2024 date confirmed). Microsoft's companion disclosure to OpenAI's first report, covering the same five actors. All five legacy/element-name mappings confirmed correct: Forest Blizzard/STRONTIUM (GRU Unit 26165), Emerald Sleet/THALLIUM (North Korea), Crimson Sandstorm/CURIUM (IRGC-connected), Charcoal Typhoon/CHROMIUM (China), Salmon Typhoon/SODIUM (China). Defensive lesson: Demonstrates the pairing that makes AI threat intel actionable: the AI provider supplies usage telemetry, the security vendor supplies actor attribution and history. Microsoft's LLM-themed TTP taxonomy gives defenders shared vocabulary to hunt and log AI-assisted tradecraft instead of treating it as an unclassifiable novelty.

  • Microsoft / OpenAI: Staying ahead of threat actors in the age of AI Microsoft Security (with OpenAI) · 2024-02-14 VERIFIED against the primary source; all five actors and their country attributions are correct, with no misattribution. Published 14 February 2024. Forest Blizzard (Russia) — LLM-informed reconnaissance into satellite and radar technologies, LLM-enhanced scripting for automation. Emerald Sleet (North Korea) — researching think tanks and North Korea experts, generating spear-phishing content, studying vulnerabilities and web technologies. Crimson Sandstorm (Iran) — generating phishing emails, code for detection evasion, malware-related scripting. Defensive lesson: The first major joint provider–vendor attribution, and the template for the disruption-and-disclosure model now standard across labs. Three of five tracked actors used LLMs for phishing or social-engineering content (Emerald Sleet, Crimson Sandstorm, Charcoal Typhoon), confirming that's the primary early misuse — a count that holds up against the source. It also demonstrates a defensive asset defenders should understand: model providers hold telemetry on adversary tradecraft and can disrupt accounts, making provider threat reporting a genuine intelligence source to consume. Read alongside the GTIG report a year later, the pair form a consistent longitudinal picture — exploratory, productivity-oriented misuse rather than capability breakthrough — which is exactly why both belong in any briefing that needs to resist AI-threat inflation.

  • Microsoft and OpenAI: nation-state threat actors using LLMs Microsoft Threat Intelligence with OpenAI · 2024-02-14 VERIFIED — no corrections; all five actor attributions are exactly right. Microsoft Threat Intelligence, in partnership with OpenAI, disclosed five state-affiliated actors using LLMs: Forest Blizzard (Russian military, GRU Unit 26165) researching satellite and radar technologies and seeking scripting help; Emerald Sleet (North Korea) generating phishing content, researching vulnerabilities, and identifying think tanks/experts; Crimson Sandstorm (Iran, IRGC-connected) for phishing emails, evasion code, and . Defensive lesson: The foundational public disclosure of state-actor LLM use, and the template for the vendor-disclosure-plus-account-termination model that followed. Its enduring value for defenders is the observed usage taxonomy — reconnaissance, translation, social engineering, scripting, and troubleshooting — which tells you AI amplifies the human-facing and preparatory phases first. Invest accordingly in phishing resistance (FIDO2, out-of-band verification) rather than expecting exotic new malware, and read the 'no novel techniques' line as an artifact of its date to be compared against 2026 findings. Note the scope limit Microsoft itself states — findings cover 'the LLMs we monitor closely,' not all model use.


Real-World Incidents

Remediate on your own device → run Multiverse Device Rescue: rescue.

  • Hugging Face security incident disclosure — July 2026 (first end-to-end autonomous AI-driven intrusion) Hugging Face · 2026-07-16 VERIFIED against the primary source (fetched 2026-07-19). Over a weekend a malicious dataset abused "two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code" on a processing worker; the foothold escalated to node-level access, harvested cloud and cluster credentials, and moved laterally across multiple internal clusters. What makes this the landmark case is the operator: Hugging Face assessed the campaign "was driven, end to end, by an autonomous AI agent system" — an agentic framework executing 17,000+ recorded events across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. It was first surfaced by Hugging Face's own LLM-based anomaly triage over security telemetry, and the forensics were run on GLM 5.2, an open-weight model, on their own infrastructure — a deliberate choice, because submitting real attack commands, exploit payloads and C2 artifacts to hosted models "were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker," and self-hosting kept attacker data and credentials from leaving their environment. Reported clean: public models, datasets and Spaces, and the software supply chain (container images, published packages); still under assessment at disclosure: whether any partner or customer data was affected. Response: closed the dataset code-execution paths, rebuilt compromised nodes, revoked and rotated affected credentials and tokens with a broader precautionary secret rotation, deployed stricter cluster admission controls, engaged outside forensic specialists and law enforcement, and advised users to rotate access tokens and review account activity. Defensive lesson: This is the case where the whole collection converges — the data-and-model surface is the initial-access vector (a dataset, not a phish), the attacker is fully agentic (machine speed, thousands of actions, no human in the tactical loop), and the defender answers in kind (LLM triage for detection, an open-weight model for forensics). Three concrete takeaways: (1) treat dataset and model ingestion as untrusted code execution — sandbox loaders, disable remote-code dataset paths by default, and lock down config/template injection like any other deserialization sink; (2) pre-stage a vetted open-weight model you can run on your own infrastructure for incident response before an incident, both to dodge guardrail lockout and to keep credentials and attacker artifacts in-environment; (3) a high-severity signal must page a responder in minutes on any day — a weekend intrusion driven at machine speed outruns business-hours triage. Caveats to preserve when citing: this is a vendor self-report of an incident against its own platform; the attacker's specific LLM and orchestration framework remain unknown; and the customer-data assessment was still open at publication. Loss-of-control tie-in: Ties directly into Loss of control from goal misgeneralization in the AI Risk Atlas — but read the causal arrow carefully. This incident was misuse (a human pointed an autonomous agent at a target), not misalignment (an agent internalizing the wrong goal). Its weight for loss-of-control is that it supplies the premise those arguments usually have to assume: a capable agent can now sustain a long-horizon, adaptive, 17,000-action campaign against a hardened target with a human barely in the loop. And the attacker's observed behaviors — credential harvesting, persistence, lateral movement, self-migrating C2, detection evasion — are exactly the instrumental-convergence behaviors the goal-misgeneralization literature predicts a capable-but-misdirected agent would pursue toward whatever goal it actually learned. Misuse and misalignment converge on the same operational signature, which is why the same controls (human oversight, sandboxing, egress control, fast shutdown) blunt both. One misgeneralization is even visible inside the incident: the hosted safety guardrails "cannot distinguish an incident responder from an attacker" — a proxy objective ("refuse attack-like requests") that generalized wrongly and, ironically, stripped control from the defender.

  • PROMPTSPY — Android backdoor using the Gemini API for autonomous device navigation Google Threat Intelligence Group (GTIG) · 2026-05-11 CONFIRMED against the primary source, and the item if anything understates the capability. PROMPTSPY is an Android backdoor containing an autonomous agent module named 'GeminiAutomationAgent' that serializes the device's visible UI hierarchy into an XML-like format via the Accessibility API and posts it to the gemini-2.5-flash-lite model in JSON Mode; the model returns structured JSON dictating action types and spatial coordinates, which the malware parses to simulate physical gestures (CLICK, SWIPE). Defensive lesson: Demonstrates agentic malware reaching mobile: an LLM acting as an on-device decision loop lets a backdoor adapt to arbitrary app UIs without the operator scripting each target. Defenders should recognize that accessibility-service abuse plus model-driven planning generalizes across apps, breaking per-app detection assumptions, and should treat unexpected inference-API traffic from mobile applications as a hunt signal. The biometric-gesture replay capability further undercuts device-lock as a containment step once the device is compromised.

  • GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access Google Threat Intelligence Group (GTIG) · 2026-05-11 In this edition of the GTIG AI Threat Tracker, Google reported that in late March 2026 the cyber-crime actor 'TeamPCP' (aka UNC6780) claimed responsibility for multiple supply-chain compromises of popular GitHub repositories and associated GitHub Actions, including those of the Trivy vulnerability scanner, Checkmarx, LiteLLM and BerriAI. Defensive lesson: The AI toolchain is now a supply-chain target, and it is a high-leverage one because AI/CI packages typically run with broad credentials in build systems. Defenders should pin and verify dependencies, scope CI secrets to short-lived least-privilege tokens, and treat agent 'skill'/plugin marketplaces as untrusted code requiring review before installation — an agent extension is arbitrary code execution with the agent's privileges.

  • Disrupting the first reported AI-orchestrated cyber espionage campaign (GTG-1002) Anthropic · 2025-11-13 In mid-September 2025 Anthropic detected an espionage operation it designated GTG-1002, assessed with high confidence to be a Chinese state-sponsored group, which manipulated Claude Code into targeting roughly thirty organizations including technology companies, financial institutions, chemical manufacturers, and government agencies; a small number were successfully infiltrated. Anthropic reported the AI executed roughly 80–90% of tactical operations with only four to six human decision points per operation, sustaining request rates far beyond human speed. Defensive lesson: The first documented large-scale intrusion campaign executed largely without human intervention — the human moved from operator to supervisor. Two durable lessons: attack tempo can now exceed human-speed response, so detection and containment must be automated; and task decomposition plus false authorization claims defeat per-request safety checks, meaning safeguards need session- and campaign-level context, not just prompt-level filtering.

  • PROMPTSTEAL (LAMEHUG) — first observed malware querying an LLM in live operations, attributed to APT28 Google Threat Intelligence Group (GTIG) · 2025-11-05 CONFIRMED against the primary source; attribution, vendor, and date all correct. GTIG's 5 Nov 2025 AI Threat Tracker reports PROMPTSTEAL, a Python data miner packaged with PyInstaller, deployed by Russian government-backed APT28 (FROZENLAKE — Google's own designation for APT28, correctly applied here) against targets in Ukraine. It queried Qwen2.5-Coder-32B-Instruct via the Hugging Face API to generate one-line Windows commands for collecting system information and documents rather than shipping them hard-coded. Defensive lesson: The first confirmed real-world case moves LLM-integrated malware from theory to incident response reality. Defenders gain a concrete detection opportunity — the malware must reach a third-party inference API, so egress filtering, allowlisting of AI service endpoints, and alerting on unexpected processes contacting model-hosting domains become genuine controls. It also warns that a stolen or free API token is now part of the malware supply chain worth monitoring.

  • s1ngularity: Nx npm supply-chain attack that weaponized developers' local AI CLI tools Nx (official postmortem) · 2025-09-05 VERIFIED — all claims confirmed against Nx's official postmortem with no corrections needed. Attackers exploited a GitHub Actions injection flaw in Nx's publishing pipeline to steal an npm publishing token and publish malicious versions of nx and several @nx/ packages on 2025-08-26, live roughly four hours before removal from the registry. The malicious postinstall scripts scanned systems for sensitive data, attempted to use local AI tools (Claude and Gemini are named explicitly in the source), and uploaded findings to a public GitHub repo via the GitHub CLI. Defensive lesson:* The novel element is malware using the developer's own already-authenticated AI CLI as a secret-hunting tool — the agent's credentials and file access became the attacker's capability. Defenders should treat locally installed AI CLIs as privileged software with a real blast radius, and note the mundane root cause: CI injection via pull_request_target. Nx's remediation (npm Trusted Publishers with OIDC, mandatory 2FA for releases, no pipeline execution for unapproved external contributors) is the transferable checklist.

  • postmark-mcp: malicious MCP server in the wild silently BCC-ing emails Koi Security (Idan Dardikman) · 2025-09-25 VERIFIED against the primary source. Idan Dardikman published 2025-09-25. Confirmed: the npm package postmark-mcp added a hidden BCC in version 1.0.16 and later (the backdoor sat on line 231), diverting email to the attacker-controlled address phan@giftshop[.]club; roughly 1,500 weekly downloads; an estimated ~20% active use giving ~300 affected organizations; an estimated 3,000-15,000 emails per day. The package ran clean for 15 prior versions, building trust before the backdoor landed. Defensive lesson: This moved MCP risk from research demo to active in-the-wild compromise, and the attack needed no prompt injection — just a trusted-looking package and a single added line. Two operational points: name-squatting a legitimate tool defeats developer eyeballing, and package removal from the registry does not remediate installed copies. Defenders need an inventory of installed MCP servers, pinned and reviewed versions, and egress monitoring — deletion upstream is not a fix.

  • "Vibe hacking": AI coding agent used to scale data extortion (GTG-2002) Anthropic · 2025-08-27 Anthropic disrupted an operation it tracks as GTG-2002 in which an actor used Claude Code to scale a data-extortion campaign against at least 17 distinct organizations, including in healthcare, emergency services, and government and religious institutions. Rather than deploying conventional encryption ransomware, the actor threatened public exposure of stolen data, with ransom demands that sometimes exceeded $500,000. The AI agent was used to support tactical decision-making, analysis of exfiltrated data, and tailoring of ransom notes. Defensive lesson: Extortion economics shifted from encryption to exposure, and a single operator reached a victim count that previously required a team. Defenders should note that data-theft extortion bypasses backup-based recovery entirely — meaning exfiltration detection and data-access monitoring matter more than restore capability for this threat class.

  • PromptLock: sample billed as first known AI-powered ransomware, later supported as a research proof-of-concept ESET Research (Anton Cherepanov, Peter Strýček) · 2025-08-26 CONFIRMED against the source. ESET found Windows and Linux Golang samples on VirusTotal that invoke OpenAI's gpt-oss-20b model locally via the Ollama API to generate malicious Lua scripts on the fly, which it then executes, for reconnaissance, exfiltration, and encryption. Authors and 2025-08-26 date are correct. The 2025-09-03 update is confirmed verbatim: ESET was contacted by the authors of the academic study 'Ransomware 3.0: Self-Composing and LLM-Orchestrated' (arXiv 2508.20444), 'whose research prototype closely resembles the PromptLock samples discovered on VirusTotal. Defensive lesson: Technically: malware that carries prompts and calls a locally hosted model at runtime defeats hash- and string-based detection, so hunt the runtime behavior instead — unexpected Ollama/local-inference API calls from non-developer processes. Procedurally, the attribution lesson is real but should be aimed correctly: ESET hedged accurately from day one and stood by 'first known case,' so the in-the-wild inflation was a downstream media artifact, not a vendor error. The transferable discipline is therefore for readers and aggregators — carry the source's hedges forward instead of stripping them, and check whether 'first known sample' has been silently upgraded into 'first known attack.'

  • GTG-2002: 'vibe hacking' data extortion run with an AI coding agent Anthropic, August 2025 Threat Intelligence Report (the GTG-2002 codename and case detail appear in the linked full PDF, not in the summary blog post at this URL) · 2025-08-27 CONFIRMED against the full PDF report, which I retrieved and text-extracted; the codename 'GTG-2002' appears verbatim. Anthropic disrupted an operation in which an actor used Claude Code to automate reconnaissance, credential harvesting, and network penetration at scale, then used the model to analyze exfiltrated financial data to determine ransom amounts and generate psychologically targeted extortion notes. Ransom figure confirmed: 'direct ransom demands occasionally exceeding $500,000. Defensive lesson: This marks the shift from AI as advisor to AI as operator executing on victim networks, compressing a multi-person intrusion crew into one actor and speeding the intrusion-to-extortion cycle. Two concrete implications: extortion pressure is now tailored using the victim's own exfiltrated financials, so incident response must assume the attacker knows what you can pay; and defenders should expect agent-driven activity that is fast and broad but leaves tool-usage patterns — making egress monitoring, credential hygiene, and rapid containment more valuable than perimeter tuning. When citing this, cite the PDF rather than the blog post, and preserve Anthropic's 'potentially affecting' hedge on the victim count.

  • MCPoison: persistent code execution in Cursor via trusted MCP config modification (CVE-2025-54136) NVD / Check Point Research · 2025-08-01 VERIFIED against primary sources; attribution is correct. NVD confirms the description verbatim for Cursor 'versions 1.2.4 and below,' with 8.8 HIGH (NIST) and 7.2 HIGH (GitHub as CNA), CWE-78, fixed in 1.3. Check Point Research attribution CONFIRMED two ways: research.checkpoint.com carries 'CVE-2025-54136 – MCPoison Cursor IDE: Persistent Code Execution via MCP Trust Bypass' (published 2025-08-05), and the Cursor advisory GHSA-24mc-g4xr-4395 credits reporter @chaandrey, consistent with Check Point's Andrey Charikov. Defensive lesson: Trust was bound to the file's identity rather than its contents, so a once-approved MCP config became a persistent backdoor when its contents changed. Approval must be bound to a content hash, and re-prompted on change. Note the delivery path: a shared repository means one contributor's commit reaches every teammate's agent — configuration in version control is a distribution channel for agent capability. Note also this needs repo write access, not prompt injection.

  • Amazon Q Developer VS Code extension compromised with destructive code (CVE-2025-8217) AWS Security Bulletin AWS-2025-015 / GitHub advisory GHSA-7g7f-ff96-5gcw · 2025-07-23 VERIFIED — including the GHSA identifier, which required extra work to confirm. AWS-2025-015 (published 2025-07-23, updated 2025-07-25) confirms v1.84.0 of the Amazon Q Developer extension for VS Code contained malicious code, that 'an inappropriately scoped GitHub token in their CodeBuild configuration' let a threat actor inject code that was 'automatically included in a release,' and that 'the malicious code was distributed with the extension but was unsuccessful in executing due to a syntax error.' Fix v1.85.0 confirmed. Defensive lesson: A shipped, signed AI coding extension reached users with attacker-authored destructive instructions inside it — and the only reason it caused no damage was a syntax error in the attacker's code. That is luck, not a control. The real failures are boring and fixable: an over-scoped CI token and a release pipeline that auto-published repository contributions without review. Defenders should scope CI tokens minimally and gate releases on human review, especially for tools that execute with developer privileges.

  • Asana MCP server logic flaw exposed customer data across organizations Asana, reported by BleepingComputer · 2025-06-18 VERIFIED against the primary source; all dates and figures check out. BleepingComputer published 2025-06-18. Confirmed: the MCP server launched 2025-05-01, the logic flaw was discovered 2025-06-04, and service returned to normal 2025-06-17 — over a month of live exposure. Approximately 1,000 customers were impacted. Exposed data included task-level information, project metadata, team details, comments and discussions, and uploaded files. Confirmed as a software logic flaw, not an external compromise or attack. Defensive lesson: No prompt injection and no attacker were required — a conventional multi-tenancy isolation bug in a rushed AI feature was enough, and it sat live for over a month. AI/MCP surfaces need the same tenant-isolation testing maturity as any other API, and this case shows why the ability to audit an AI feature after the fact (access logs, records of what summaries were generated) is a procurement requirement, not a nice-to-have.

  • MCP Inspector missing authentication enables remote code execution (CVE-2025-49596) NVD / GitHub Security Advisory GHSA-7f8r-222p-6f5g (CNA: GitHub, Inc.) — npm @modelcontextprotocol/inspector (Anthropic-originated MCP tooling) · 2025-06-13 NVD records that 'Versions of MCP Inspector below 0.14.1 are vulnerable to remote code execution due to lack of authentication between the Inspector client and proxy, allowing unauthenticated requests to launch MCP commands over stdio.' MCP Inspector is a developer tool for testing and debugging MCP servers. Scored 9.4 CRITICAL under CVSS v4.0 (AV:N/AC:L/AT:N/PR:N/UI:P/VC:H/VI:H/VA:H/SC:H/SI:H/SA:H) and classified CWE-306 (Missing Authentication for Critical Function); fixed in version 0.14.1. Reported by Rémy Marot of Tenable. Defensive lesson: The AI development toolchain is itself attack surface, and developer tools routinely ship assuming 'localhost is trusted' — an assumption a browser can violate. The CVSS vector corroborates this: AV:N with UI:P encodes the network-reachable, victim-visits-a-malicious-page chain. Defenders should inventory locally running AI dev tools and their listening services, and apply the same authentication expectations to them as to production services. This one is a plain missing-auth bug, a reminder that most AI-stack CVEs are ordinary appsec failures in novel packaging.

  • HP Wolf Security: AI-generated dropper delivering AsyncRAT found in the wild HP Wolf Security · 2024-09-24 VERIFIED — no corrections. HP Wolf Security identified a campaign targeting French speakers with VBScript and JavaScript delivering AsyncRAT, an infostealer capable of recording screens and keystrokes. HP's cited evidence confirmed nearly verbatim: 'The structure of the scripts, comments explaining each line of code, and the choice of native language function names and variables are strong indications that the threat actor used GenAI to create the malware.' HP Principal Threat Researcher Patrick Schlapfer framed the significance as lowering entry barriers for less-skilled criminals. Defensive lesson: One of the earliest concrete in-the-wild cases of AI-written malicious code, and the evidence HP cited is directly reusable as a hunting methodology: per-line explanatory comments, textbook structure, and native-language identifiers are tells that a script was model-generated rather than hand-written or copied from a kit — the same tell-family GTIG later cited for the 2026 AI-assisted zero-day exploit, which makes these two items a good pair. Note the payload was ordinary commodity AsyncRAT — AI wrote the delivery wrapper, not a novel implant — so the practical effect is more attackers producing working chains, and existing endpoint detection for commodity RATs still catches the outcome. Teach the inference as probabilistic: these are heuristics, and HP itself claimed only 'strong indications.'

  • KnowBe4 hired a North Korean fake IT worker who used an AI-enhanced photo to pass four video interviews KnowBe4 · 2024-07 VERIFIED against the primary source; every element checks out. The persona passed four video conference interviews (HR 'confirmed the individual matched the photo provided on their application') and background checks using a stolen US identity. The photo is confirmed as an 'AI fake that started out with stock photography' — the post shows the stock original beside the submitted version. On 15 July 2024 malware activity on the issued laptop triggered alerts at 9:55pm EST; SOC contained the device at ~10:20pm EST, i.e. roughly 25 minutes. Defensive lesson: AI-assisted synthetic identity now reaches into HR and onboarding, making hiring a security perimeter rather than an administrative function. Defenses belong in the process, not in photo forensics: verify identity documents against the live candidate, watch for shipping addresses that diverge from the claimed location, and treat first-day endpoint telemetry as a detection surface. KnowBe4's 25-minute containment shows that assuming some fake personas will get hired — and instrumenting for it — beats trying to make screening perfect. The four passed video interviews are the load-bearing detail: live video is widely treated as a liveness check that settles identity, and here it settled nothing.

  • LastPass employee targeted by audio deepfake of CEO over WhatsApp LastPass · 2024-04-10 VERIFIED against the primary source; date made precise (post published 2024-04-10, describing an incident that occurred the same day — 'earlier today'; original item said only '2024-04'). Confirmed: an employee 'received a series of calls, texts, and at least one voicemail featuring an audio deepfake' from 'a threat actor impersonating our CEO,' delivered via WhatsApp. The employee disregarded the messages and reported the incident to internal security, having noted social engineering hallmarks including forced urgency and use of a channel outside normal business communications. Defensive lesson: A textbook successful defense, and worth teaching as the positive control: the attack failed on channel discipline, not on deepfake detection. The employee never had to judge whether the voice was real. Codify 'leadership does not make urgent requests over WhatsApp' as a rule, make reporting frictionless and blameless, and publish the near-miss — LastPass's disclosure of a non-incident is itself the defensive contribution.

  • Hong Kong deepfake video-call fraud: HK$200M (US$25.6M) stolen via fake multi-participant meeting (victim later revealed as Arup) South China Morning Post (incident, 4 Feb 2024); victim named as Arup by CNN (Kathleen Magramo) and The Guardian (Dan Milmo), both 17 May 2024 · 2024-02-04 VERIFIED against the primary source; date made precise (SCMP article published 2024-02-04, original item said only '2024-02'). All details confirmed: the incident occurred mid-January 2024; a finance department employee 'received what appeared to be a phishing message in mid-January, apparently from the company's UK-based chief financial officer'; the employee then joined a video conference in which a digitally recreated CFO ordered transfers and all other participants were fake. HK$200 million (US$25.6M) was transferred. Defensive lesson: Liveness on a video call is no longer proof of identity, and 'multiple colleagues agreed on the call' is not corroboration when all of them can be synthetic. Payment authority must rest on an out-of-band verification channel the attacker does not control (callback to a directory-sourced number, pre-shared challenge, dual authorization in the finance system) rather than on visual/auditory recognition. Note also the pressure pattern: urgency plus secrecy is what suppressed the employee's escalation.

  • Retool breach: attacker deepfaked a real coworker's voice to defeat MFA Retool · 2023-08 VERIFIED against the primary source. On 27 August 2023 several Retool employees received targeted SMS messages with a URL mimicking the internal identity portal; one employee logged into the fake page and then took a call from an attacker posing as IT. Retool's exact wording: 'The caller claimed to be one of the members of the IT team, and deepfaked our employee's actual voice.' Also verbatim: 'The voice was familiar with the floor plan of the office, coworkers, and internal processes of the company.' The employee shared an OTP token. Defensive lesson: Voice cloning is used not just for executive impersonation but to impersonate the helpdesk and peers — the people employees are conditioned to trust and assist. Combined with reconnaissance-derived internal detail, familiarity cues become an attack tool rather than an authenticity signal. The durable control is phishing-resistant, non-transferable authentication (FIDO2/WebAuthn hardware keys), since any credential a human can read aloud can be socially engineered out of them. Retool's own headline lesson is narrower and worth keeping attached: MFA that syncs to a cloud account inherits that account's blast radius, so 'we have MFA' says little until you ask where the second factor actually lives. Caveat for accuracy: 'deepfaked' here is Retool's own characterization; no independent forensic analysis of the call was published. The original summary's 'per Retool' hedge is correct and should be retained.

  • Cloned company director's voice used in $35M bank transfer fraud (Hong Kong branch manager; U.A.E. investigation) Forbes (Thomas Brewster), citing U.A.E. court documents filed with U.S. investigators · 2021-10 CORRECTED METADATA — geography error in the original title. The original tagged this '(U.A.E.)', implying a U.A.E. bank. Per Forbes: 'In early 2020, a branch manager of a Japanese company in Hong Kong received a call from a man whose voice he recognized—the director.' The manager was in HONG KONG; the U.A.E. (Dubai Public Prosecution Office) is the INVESTIGATING jurisdiction, which opened the probe because the scheme affected entities in the country and sought U.S. help tracing funds. Defensive lesson: Voice clones are deployed as one element of a multi-channel corroboration package — the call is 'confirmed' by emails and documents the same attacker produced. Defenders must recognize that cross-channel agreement is worthless when the attacker authors every channel; verification must terminate at an independently sourced contact. The scale ($35M, ~17 people, multi-jurisdiction laundering) also shows organized crime, not lone actors, operationalizing this early. As with the 2019 case, note the epistemic status: the voice-cloning claim comes from investigators' court filings rather than published forensic analysis — stronger evidence than an insurer's opinion, but still not a technical demonstration. The multi-jurisdiction shape (Hong Kong victim, U.A.E. prosecutors, U.S. account tracing) is itself the lesson for incident response: the fraud crosses borders faster than the investigation can.

  • Fraudsters reportedly used AI to mimic a CEO's voice to steal €220,000 from a UK energy firm Wall Street Journal (Catherine Stupp, 30 Aug 2019), sourced from insurer Euler Hermes; corroborated by Forbes (Jesse Damiani, 3 Sep 2019) · 2019-03 CORRECTED METADATA. The incident is real and the core facts are confirmed: the CEO of a UK-based energy firm was directed by a caller impersonating the chief executive of its German parent company to transfer €220,000 (~$243,000), routed to a Hungarian account. Actor/victim direction in the original summary was correct. Corrections: (1) URL supplied — the WSJ original is paywalled and unfetchable, so the verified corroborating Forbes piece is used; (2) date refined — the fraud occurred in March 2019 and was reported by WSJ on 30 August 2019, so bare '2019' conflated event and publication. Defensive lesson: The foundational case, and the one that dates the threat — but date it honestly. The original lesson's claim that 'voice-cloning fraud was operational in 2019' asserts more than the evidence supports: what is documented is a successful voice-impersonation fraud in 2019 that an insurer attributed to AI without forensic confirmation. That distinction matters precisely because this case is cited everywhere as the origin point; treating a belief as a finding is how awareness material loses credibility with technical audiences. The defensive conclusion survives the caveat intact, because it never depended on how the voice was produced: verify payment instructions via callback to a known-good number, never a number or channel supplied within the request itself. Authority plus urgency in a familiar voice is the whole attack, whether the voice was synthesized or merely impersonated well.


Criminal Tooling & AI-Written Malware

Remediate on your own device → run Multiverse Device Rescue: rescue --profile ai_worm_response.

  • GTIG AI Threat Tracker: Distillation, Experimentation, and (Continued) Integration of AI for Adversarial Use Google Threat Intelligence Group (GTIG) · 2026-02-12 GTIG's follow-up to its November 2025 findings, reporting model-extraction ("distillation") activity including a single campaign of over 100,000 prompts aimed at eliciting Gemini's reasoning traces in non-English languages; continued AI augmentation across the attack lifecycle by government-backed actors including APT41, APT42, APT31, UNC2970 and others; and new AI-integrated tooling such as HONESTCUE (a downloader framework calling the Gemini API to generate C# code for later stages), COINBAIT (a phishing kit built with the Lovable AI platform, linked to UNC5356), and the ATOMIC macOS informa. Defensive lesson: Three distinct defensive surfaces emerge: the model itself becomes an asset to steal (extraction attacks), public AI sharing features become malware-hosting infrastructure that defenders must treat as an untrusted content source, and underground 'custom evil AI' offerings are frequently reselling mainstream APIs — meaning provider-side enforcement reaches further into the criminal ecosystem than its marketing implies.

  • HONESTCUE — proof-of-concept downloader requesting compilable payload code from the Gemini API Google Threat Intelligence Group (GTIG) · 2026-02-12 CONFIRMED against the primary source ('GTIG AI Threat Tracker: Distillation, Experimentation, and (Continued) Integration of AI for Adversarial Use', 12 Feb 2026); no corrections needed. HONESTCUE is a downloader/launcher framework that prompts the Gemini API to return C# source for a second stage, then compiles and executes it in memory via the legitimate .NET CSharpCodeProvider with no disk artifacts; payloads are hosted on legitimate CDNs including Discord CDN. Defensive lesson: Shows the convergence of two evasion trends — fileless in-memory execution plus runtime code generation — meaning neither the payload nor its signature exists until execution. Defenders should invest in in-memory and runtime telemetry (script/assembly load events, unusual compiler API use) and treat legitimate CDN hosting as no indicator of benignity. GTIG labelling it proof-of-concept is the useful part: this is the window in which to build detections before it matures.

  • Xanthorox — advertised 'bespoke malicious AI' shown to be an aggregator of legitimate open-source tools Google Threat Intelligence Group (GTIG) · 2026-02-12 CONFIRMED against the primary source; only a trivial naming nit (GTIG writes 'LibreChat-AI', the item writes 'LibreChat'). Xanthorox was marketed in criminal channels as a 'bespoke, privacy preserving self-hosted AI' able to autonomously generate malware, ransomware and phishing. Defensive lesson: A valuable corrective against overestimating the underground AI market: much 'malicious AI' tooling is repackaged legitimate software with marketing attached, and criminal advertising is not evidence of capability. Defenders and policymakers should verify claimed capabilities before reprioritizing on the strength of a forum listing. It also shows MCP-style orchestration is becoming the abuse substrate, so MCP servers and plugins warrant supply-chain scrutiny.

  • PROMPTFLUX — experimental self-rewriting VBScript dropper using the Gemini API Google Threat Intelligence Group (GTIG) · 2025-11-05 CONFIRMED against the primary source, including the attribution nuance the item gets right. PROMPTFLUX is a VBScript dropper calling the Gemini API (specifically gemini-1.5-flash-latest) to regenerate and obfuscate its own source code, including the 'Thinking Robot' module that periodically queries Gemini for antivirus-evasion code; a later variant rewrites the entire source hourly. It establishes persistence via the Startup folder and spreads via removable drives and network shares. Defensive lesson: A working demonstration that metamorphic malware no longer requires a hand-built mutation engine — an API key substitutes for deep expertise, lowering the skill floor for evasion. The defensive takeaway is that hash- and static-signature-based controls degrade first, so defenders should anchor on persistence locations, script interpreter telemetry, and outbound API calls. That it came from a commodity criminal actor, not an APT, is the real warning about diffusion speed.

  • GTIG AI Threat Tracker: maturing underground AI marketplace and guardrail bypass by pretext Google Threat Intelligence Group · 2025-11-05 VERIFIED — no corrections. GTIG assessed the cyber crime marketplace for AI-enabled tooling 'matured' in 2025, with multifunctional tools spanning the attack lifecycle (deepfake/image generation, malware generation, phishing kits, vulnerability exploitation) and conventional pricing — ads injected into free versions, subscription tiers adding image generation, API access, and Discord access. GTIG notes 'almost every notable tool advertised in underground forums mentioned their ability to support phishing campaigns. Defensive lesson: Safety mitigations that key on stated intent are defeated by cheap pretexts, so provider-side abuse detection must weight observed behavior and request patterns over the user's claimed context — and CTF/academic framing deserves specific scrutiny because it is both a common legitimate use and a known bypass. Two bonus points for defenders: criminal AI now has product-market fit and legitimate-looking pricing, so track it as a maturing industry; and OPSEC failures persist — TEMP.Zagros leaked a hard-coded C2 domain and encryption key into a prompt, which enabled broader campaign disruption. Adversary use of AI creates new collection opportunities for defenders.

  • No-code malware: AI-generated ransomware-as-a-service (GTG-5004) Anthropic · 2025-08-27 Anthropic disrupted a UK-based actor it tracks as GTG-5004 who used Claude to develop, market, and distribute ransomware variants sold on internet forums to other criminals for roughly $400 to $1,200 USD (a ransomware DLL/executable at $400 up to a Windows 10/11 FUD crypter at $1,200). The actor had been active since at least January 2025 on dark web forums. Anthropic assessed the actor appeared dependent on AI assistance to implement and maintain functional malware components, and did not demonstrate the ability to build them unaided. Defensive lesson: The barrier to entry for selling criminal tooling dropped below the skill level historically required to build it, widening the pool of sellers. That dependency is also a defensive opportunity: AI-reliant developers cluster around provider platforms where their activity is observable and disruptable, unlike self-sufficient malware authors.

  • GTG-5004: no-code ransomware-as-a-service developed and sold with AI assistance Anthropic, August 2025 Threat Intelligence Report (the GTG-5004 codename and case detail appear in the linked full PDF, not in the summary blog post at this URL) · 2025-08-27 CONFIRMED against the full PDF report, with every specific detail matching verbatim — this item is unusually precise. Confirmed exactly: a UK-based threat actor tracked as GTG-5004; use of Claude to develop, market, and distribute ransomware; ChaCha20 encryption; anti-EDR techniques and Windows internals exploitation; active since at least January 2025 on dark web forums including Dread, CryptBB, and Nulled. Defensive lesson: Anthropic's most useful observation is the actor's apparent dependency on AI — they appear unable to implement or troubleshoot complex components without it. That cuts both ways for defenders: capable malware now originates from operators lacking the skill to have written it, so skill-based actor profiling misleads, but such operators are also brittle, less able to adapt when a technique is burned. Detection should stay anchored on the concrete artifacts (ChaCha20 header-only encryption of the first 256KB, '.enc' extension marking, direct syscall invocation to bypass EDR user-mode hooks, C2 infrastructure) that AI assistance does not obscure.

  • Cato CTRL Threat Research: WormGPT Variants Powered by Grok and Mixtral Cato Networks (Cato CTRL), Vitaly Simonovich, Senior Security Researcher · 2025-06-17 CONFIRMED against the source. Cato CTRL identified two previously unreported WormGPT variants: keanu-WormGPT, a wrapper over xAI's Grok using a custom system prompt to bypass its guardrails, and xzin0vich-WormGPT, built on Mistral AI's Mixtral. Researchers used LLM jailbreak techniques to extract the system prompts and reveal the underlying foundation models. Defensive lesson: 'WormGPT' is now a brand applied to commercial models wrapped in guardrail-stripping system prompts, not a single malware artifact — so attribution should target the wrapper and its upstream API keys. System-prompt extraction is a legitimate and effective investigative technique for unmasking which vendor's model is being abused, letting providers act on the underlying accounts. Note the two-channel structure: forums for marketing, Telegram for delivery — disruption of one does not remove the other.

  • GhostGPT: uncensored chatbot sold through Telegram Abnormal AI (Callie Baron, Piotr Wojtyla) · 2025-01-23 Abnormal documented GhostGPT, a chatbot sold via Telegram and marketed to cybercriminals with a 'no-logs policy' and fast, immediate access requiring no self-jailbreaking or open-source setup, advertised for malware code, polymorphic malware, exploitation of software vulnerabilities, phishing emails, BEC templates, and fraudulent website design. Researchers assessed it 'likely uses a wrapper to connect to a jailbroken version of ChatGPT' or an open-source LLM with safeguards removed, rather than a bespoke trained model — they did not definitively confirm which. Defensive lesson: Most 'custom criminal AI' is a thin wrapper plus a jailbreak, which is both a limitation and a control point: wrapper services inherit the upstream provider's abuse surface, so provider-side jailbreak detection and key revocation degrade them. The 'no-logs' selling point also signals what buyers fear most — telemetry — which is exactly the asset defenders and providers should preserve and correlate.

  • WormGPT: blackhat LLM marketed for business email compromise SlashNext (researcher Daniel Kelley), reported by The Hacker News · 2023-07-15 SlashNext documented WormGPT, an underground chatbot built on the open-source GPT-J model (developed by EleutherAI) and advertised on cybercrime forums as a 'blackhat alternative to GPT models, designed specifically for malicious activities' for automating phishing and business email compromise lures. Its creator promoted it as the 'biggest enemy of the well-known ChatGPT' that 'lets you do all sorts of illegal stuff.' It was the first widely reported purpose-built criminal LLM service. Defensive lesson: The first criminal LLM wave was built on open-weight models, not jailbroken commercial APIs — meaning provider-side guardrails and account bans cannot reach it. Defenders should assume fluent, grammatically clean phishing is now the baseline and retire 'bad grammar' as a BEC detection signal, shifting weight to sender authentication, payment-process verification, and behavioral/relationship anomalies.

  • FraudGPT: subscription crimeware LLM sold on dark web markets and Telegram Netenrich (Rakesh Krishnan) · 2023-07-25 Netenrich reported FraudGPT, circulating from 22 July 2023 on Telegram channels and advertised across dark web marketplaces at roughly $200/month to $1,700/year, marketed for spear-phishing, scam pages, carding, and 'undetectable' malware. The seller claimed '3,000+ confirmed sales/reviews' and verified-vendor status on multiple underground markets including Empire, WHM, Torrez, World, AlphaBay and Versus. Defensive lesson: Criminal AI reached a commodity subscription model within weeks of the LLM boom, so the relevant defensive question is volume and lowered barrier to entry, not novel capability. Treat vendor-claimed metrics (3,000+ sales) as marketing rather than telemetry, and prioritize monitoring Telegram and market listings as an early-warning source for what lure themes and fraud types will spike next.


Prompt Injection & Agentic Exploits in Real Products

  • CamoLeak: GitHub Copilot Chat private source code exfiltration via Camo proxy CSP bypass Legit Security (Omer Mayraz) · 2025-10-08 VERIFIED with one minor correction. Researcher (Omer Mayraz, Legit Security), CVSS 9.6 critical, affected product (GitHub Copilot Chat), and the fix (GitHub disabled image rendering in Copilot Chat completely, as of 2025-08-14) are all confirmed. CORRECTION: the item says injection came via 'pull request comments'; the source specifies the instructions were hidden in GitHub markdown comments embedded within pull request descriptions — invisible to human readers but parsed by Copilot. Defensive lesson: An attacker needed no special privileges — only the ability to post content on a PR — to make the assistant act with the victim's repository permissions. This is the confused-deputy pattern: the agent's identity, not the attacker's, gates access. Also note the fix: GitHub removed the image-rendering capability rather than filtering it, showing that when a rendering primitive doubles as an egress channel, removing the feature is sometimes the only reliable control. The pre-signed-URL dictionary technique further shows that cryptographic signing of a proxy URL authenticates the URL, not the intent behind fetching it.

  • Amazon Q Developer and Kiro: prompt injection bypassing human-in-the-loop confirmation (AWS-2025-019) AWS Security Bulletin AWS-2025-019 · 2025-10-07 VERIFIED — all three issues, all patch versions, all dates, and the bulletin's publication date confirmed against the AWS primary source. Publication date confirmed exactly as 'Publication Date: 2025/10/07 01:30 PM PDT', matching the item's 2025-10-07 (note the patches predate disclosure by ~2-3 months, which is consistent, not contradictory). AWS acknowledged three issues discovered via coordinated disclosure: commands such as find, grep, and echo could execute without Human-in-the-Loop (HITL) confirmation and could be obfuscated by invisible control characters (fixed in <1.22.0 → 1.22. Defensive lesson: Allowlists of 'safe' commands leak: ping and dig are read-only yet became DNS exfiltration channels, and invisible control characters defeated the human reviewer the design depended on. Most instructive is that Kiro's Supervised mode failed too — a HITL control is only as good as the fidelity of what it shows the human. Defenders should classify commands by egress capability rather than perceived destructiveness, and validate that approval prompts render exactly what will execute.

  • ForcedLeak: indirect prompt injection in Salesforce Agentforce via Web-to-Lead Noma Security (Noma Labs) · 2025-09-25 VERIFIED — every metadata element confirmed against the Noma Security primary source with no corrections needed. Noma Labs disclosed a CVSS 9.4 (Critical) indirect prompt injection chain in Salesforce Agentforce in which instructions embedded in Web-to-Lead submission data (specifically the Description field) could cause the AI agent to exfiltrate CRM data when an employee queried/processed the lead. A domain on Salesforce's CSP allowlist (my-salesforce-cms.com) had expired and was purchasable for about $5, providing a trusted exfiltration channel. Defensive lesson: Two lessons compound here. First, any public intake form feeding an AI agent is an untrusted-input channel available to anonymous attackers. Second, allowlists are living assets: an expired domain still trusted by CSP converts a defensive control into an exfiltration path. Defenders should inventory and monitor expiry for every domain in a CSP/trusted-URL allowlist, and treat allowlist entries as credentials that can be lost.

  • ShadowLeak: zero-click service-side exfiltration via ChatGPT Deep Research + Gmail Radware (Zvika Babo, Gabi Nakibly, Maor Uziel), reported by The Hacker News · 2025-09-20 VERIFIED. Researcher names (Zvika Babo, Gabi Nakibly, Maor Uziel), affiliation (Radware), affected product (OpenAI's ChatGPT Deep Research agent, launched February 2025), and both dates (disclosed 2025-06-18; fixed early August 2025) all confirmed. The server-side characterization is confirmed: the leak occurred 'directly within OpenAI's cloud environment,' bypassing client-side defenses, making it invisible to local or enterprise controls. Defensive lesson: The defining lesson is the location of egress. When the agent acts server-side inside the vendor's cloud, the exfiltration never traverses the enterprise perimeter — so endpoint DLP, proxies, and egress filtering see nothing. Defenders cannot inherit visibility here and must demand it contractually: agent-side action logs, connector-level audit trails, and vendor egress controls become the only detection surface.

  • CurXecute: prompt injection to remote code execution in Cursor via MCP config write (CVE-2025-54135) NVD / Cursor advisory GHSA-4cxx-hrm3-49rm (research by Aim Labs / Aim Security — researcher Ofir Abu; Aim Security was acquired by Cato Networks, so the write-up now sits under Cato CTRL) · 2025-08-04 VERIFIED, with one source correction. The CVE is real and every technical detail checks out verbatim against NVD: 'Cursor allows writing in-workspace files with no user approval in versions below 1.3.9, If the file is a dotfile, editing it requires approval but creating a new one doesn't.' NVD confirms 9.8 CRITICAL (NIST) and 8.5 HIGH (GitHub as CNA), CWE-78 and CWE-829, fixed in 1.3.9. GHSA-4cxx-hrm3-49rm is a genuine Cursor repository-level advisory (at github.com/cursor/cursor/security/advisories/, not the global GitHub advisory DB). Defensive lesson: A precise asymmetry created the bug: editing a dotfile required approval but creating one did not. Security controls must cover the full lifecycle of a sensitive resource — create, edit, delete — because attackers target the unguarded verb. More broadly, agent-writable config files that the agent also reads as instructions form an execution loop; any file that can grant capability must sit outside the agent's unapproved write scope. Cursor auto-started new MCP entries without confirmation, so the write executed before the user could reject it.

  • AgentFlayer: zero-click exfiltration via ChatGPT Connectors Zenity Labs (Tamir Ishay Sharbat) · 2025-08-06 VERIFIED as to the exploit; ONE ELEMENT UNCONFIRMED. Tamir Ishay Sharbat published 2025-08-06, and all technical claims check out: the target is ChatGPT Connectors (Google Drive, SharePoint, GitHub); the trigger is a poisoned document carrying an invisible injection in white 1px font, requiring no user interaction beyond summarization; the payload directs ChatGPT to search connected Drive for API keys and embeds them in image markdown that renders and transmits automatically. Defensive lesson: A URL allowlist that includes any general-purpose cloud storage provider is not an allowlist — attackers rent a bucket on a trusted domain. This is the same failure mode as ForcedLeak's expired domain, reached from the opposite direction. Defenders should assume domain-reputation checks are bypassable when major hosting providers are implicitly trusted, and prefer blocking agent-initiated outbound fetches over enumerating bad destinations.

  • EchoLeak: zero-click indirect prompt injection in Microsoft 365 Copilot (CVE-2025-32711) NVD / Microsoft (discovered by Aim Security) · 2025-06-11 VERIFIED against NVD — the CVE is real and not invented; description matches verbatim. NVD: 'Ai command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network,' published 2025-06-11, affecting Microsoft 365 Copilot (all versions), scored 7.5 HIGH by NIST (AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N) and 9.3 CRITICAL by Microsoft as CNA (AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N) — the dual scoring in the original is accurate. Aim Security discovered the flaw, which allowed data exfiltration from M365 Copilot via a crafted email with no user interaction. Defensive lesson: An AI assistant with standing access to a user's mailbox and documents turns any inbound email into untrusted input that reaches a privileged context. Defenders should treat all retrieved/RAG content as attacker-controlled, enforce provenance-based access control so untrusted content cannot invoke privileged actions, and note that 'zero-click' removes the user-vigilance layer entirely — phishing training cannot mitigate this class. The NIST/Microsoft scoring split (7.5 vs 9.3, driven by scope change and integrity impact) is itself instructive: reasonable parties disagree on how to score AI-mediated flaws, so do not let a mid-range NVD number set your patch priority for assistant-layer vulnerabilities.

  • Lessons from Defending Gemini Against Indirect Prompt Injections arXiv (Google DeepMind) · 2025-05-20 Chongyang Shi, Sharon Lin, Shuang Song, Jamie Hayes, Ilia Shumailov, Itay Yona, Juliette Pluto, Aneesh Pappu, Christopher A. Choquette-Choo, Milad Nasr, Chawin Sitawarin, Gena Gibson, Andreas Terzis and John 'Four' Flynn describe DeepMind's adversarial evaluation framework for assessing Gemini's robustness when adversaries embed harmful instructions in untrusted data reached through tool use. The framework deploys a suite of adaptive attack techniques that run continuously against past, current and future versions of Gemini, with results feeding an iterative cycle of defensive improvement. Defensive lesson: The central lesson is that defenses evaluated only against static, known attacks give false assurance — adaptive adversaries recover much of their success rate, so robustness claims must be made against attacks that adapt to the defense. Operationally this argues for continuous automated red-teaming as a standing regression suite on every model release, and for treating any tool-using agent that reads untrusted content as being within prompt-injection blast radius by default.

  • Remote prompt injection in GitLab Duo leading to source code theft Legit Security (Omer Mayraz) · 2025-05-22 VERIFIED against the primary source; every specific detail checks out. Omer Mayraz published on 2025-05-22 (post since updated 2025-12-05). Disclosure to GitLab on 2025-02-12 CONFIRMED, as is the fix shipping via merge request duo-ui!52, which stopped Duo rendering unsafe HTML tags pointing to external domains. The five OWASP Top 10 for LLMs (2025) categories are explicitly referenced: LLM01 (Prompt Injection), LLM02 (Sensitive Information Disclosure), LLM05 (Improper Output Handling), LLM08, and LLM09 (Misinformation). Defensive lesson: Every field in a code-hosting platform that an AI assistant reads — issue text, commit messages, even source comments — is untrusted input reachable by any contributor. The fix generalizes well: the exfiltration channel was HTML/Markdown rendering to external domains, so constraining what the assistant's output may render is often more tractable than trying to detect every injection. Restrict agent output rendering to same-origin, non-fetching content.

  • MCP Security Notification: Tool Poisoning Attacks Invariant Labs AG (Luca Beurer-Kellner, Marc Fischer) · 2025-04-01 VERIFIED — authors (Luca Beurer-Kellner, Marc Fischer), organization (Invariant Labs AG), and publication date (2025-04-01) all confirmed against the primary source. Invariant Labs described Tool Poisoning Attacks, in which malicious instructions embedded in MCP tool descriptions are 'invisible to users but visible to AI models,' steering the agent into unauthorized actions such as stealing SSH keys or reading sensitive configuration files while concealing the activity. Cursor is confirmed as a named/tested agentic coding client, and Anthropic, OpenAI, and Zapier are referenced. Defensive lesson: The MCP tool description is executable attack surface, not documentation — and it is exactly the part of the system users never see. This breaks the trust assumption behind every 'the user approved this tool' control. Defenders should render tool descriptions in the approval UI, pin and hash server versions, and treat description text as untrusted input entering the model's instruction channel.

  • WhatsApp MCP exploited: cross-server shadowing and rug-pull data exfiltration Invariant Labs AG (Luca Beurer-Kellner, Marc Fischer) · 2025-04-07 VERIFIED against the primary source. Authors, date, and the central quote all match exactly. The post states that 'an untrusted MCP server can attack and exfiltrate data from an agentic system that is also connected to a trusted WhatsApp MCP instance, side-stepping WhatsApp's encryption and security measures.' Two attack scenarios are demonstrated, stealing message history and contact information via malicious tool descriptions rather than code execution. Defensive lesson: Two structural lessons. First, approval is a point-in-time check while tool descriptions are mutable — trust must be continuously re-verified (pin and hash), or approval means nothing. Second, MCP servers share one agent context, so connecting one untrusted server compromises every trusted server alongside it; there is no per-server isolation by default. End-to-end encryption is irrelevant when the attacker subverts the authorized endpoint.

  • Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications (the Morris II worm) arXiv (Stav Cohen, Ron Bitton, Ben Nassi) · 2024-03-05 CONFIRMED against both the v1 and v2 abstracts. Researchers demonstrated Morris II (styled 'Morris-II' in the paper), an adversarial self-replicating prompt that triggers a cascade of indirect prompt injections, causing a GenAI model to reproduce the malicious input in its output and propagate to other agents. The v1 abstract confirms it was tested against exactly three models — Gemini Pro, ChatGPT 4.0, and LLaVA — against GenAI-powered email assistants in two use cases (spamming and exfiltrating personal data), under black-box and white-box settings, using text and image inputs. Defensive lesson: The researchers framed this as bad architecture in the GenAI ecosystem rather than a vulnerability in any one vendor's model — consistent with it reproducing across three unrelated models — so there is no patch to wait for and signature-based blocking will not hold. Defenders building agent pipelines must treat model output as untrusted input to the next hop: isolate agent-to-agent data flow, require human approval for outbound actions, and constrain what RAG-retrieved content can trigger. That the same prompt worked against Gemini, ChatGPT, and LLaVA is the load-bearing detail — it is a property of the composition pattern, not of a model.


Academic: Can Agents Attack Autonomously?

  • EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System (arXiv:2509.10540) arXiv (Pavan Reddy, Aditya Sanjay Gujral) · 2025-09-06 VERIFIED. arXiv:2509.10540 exists with the stated title, authors (Pavan Reddy, Aditya Sanjay Gujral), and submission date (2025-09-06). It is an academic write-up of EchoLeak (CVE-2025-32711) documenting how the exploit chained several weaknesses — bypassing Microsoft's XPIA classifier, using reference-style Markdown to evade link redaction, exploiting auto-fetched images, and leveraging a Teams proxy — to achieve what the authors call 'full privilege escalation across LLM trust boundaries. Defensive lesson: Demonstrates that real-world LLM exploits are chains, not single bugs: a dedicated prompt-injection classifier was one link that failed among several. The paper's proposed mitigations — prompt partitioning, enhanced filtering, provenance-based access control, and strict content security policies — are a useful defense-in-depth checklist, and the case shows that ML-based injection classifiers must not be the sole control.

  • Ransomware 3.0: Self-Composing and LLM-Orchestrated (academic origin of the PromptLock samples) arXiv (Md Raz, Meet Udeshi, P.V. Sai Charan, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri — NYU Tandon) · 2025-08-28 CONFIRMED against the source. The abstract confirms a prototype in which malicious code 'is synthesized dynamically by the LLM at runtime,' producing 'polymorphic variants that adapt to the execution environment,' and closing the loop across 'reconnaissance, payload generation, and personalized extortion, in a closed-loop attack campaign.' Evaluation spans 'personal, enterprise, and embedded environments,' and the authors found open-source LLMs can generate functional ransomware components and sustain closed-loop execution. Author list and 2025-08-28 date are correct. Defensive lesson: This is the prototype ESET flagged as PromptLock, so it demonstrates capability without an actual adversary — useful for planning, wrong to cite as an in-the-wild incident. Defensively, prompts-as-payload makes every infection a unique binary, collapsing the value of static IOCs and shifting detection to encryption behavior, mass file access patterns, and egress to model endpoints. It also raises a live disclosure question: offensive PoCs reaching public malware repositories will be found and misreported as real attacks — though note the researchers proactively contacted ESET to correct the record, which is the behavior the disclosure norm should reward, and the route by which the samples reached VirusTotal is not publicly established.

  • RedTeamLLM: an Agentic AI framework for offensive security arXiv (Brian Challita, Pierre Parrend) · 2025-05-11 VERIFIED. Proposes an agentic framework for automated penetration testing built on a summarize-reason-act loop, explicitly targeting four named obstacles: plan correction, memory management, context-window limits, and the generality-versus-specialization tradeoff. Evaluated on automated resolution of 'entry-level, but not trivial' CTF challenges, with attention to how much the reasoning component contributes. Defensive lesson: Useful as a map of what still constrains autonomous offensive agents. Each named limitation (memory, context, replanning) is simultaneously a research target for attackers and a source of observable, brittle behavior for defenders — agents that lose plan coherence retry, thrash, and generate noisy artifacts that detection engineering can exploit. Note the modest evaluation scope: entry-level CTF, not enterprise networks.

  • Dynamic Risk Assessments for Offensive Cybersecurity Agents arXiv (Princeton / UC Irvine — Wei, Stroebl, Xu, Zhang, Li, Henderson) · 2025-05 VERIFIED. Title matches verbatim; v1 submitted 23 May 2025 (last revised October 2025), so the 2025-05 date is correct. Author list confirmed exactly — Boyi Wei, Benedikt Stroebl, Jiacen Xu, Joie Zhang, Zhou Li, Peter Henderson — and the Princeton / UC Irvine attribution is correct (Xu and Li are UC Irvine; the remainder Princeton). Defensive lesson: This is the methodological rebuttal to every reassuring system-card number: a >40% capability gain from a fixed model, achieved with modest compute (8 H100 hours) and more attacker freedom, means published scores describe one harness rather than the model's potential. Defenders should mentally apply a substantial uplift factor to any static benchmark result before using it in a threat model, and should ask what an adversary willing to spend more and iterate freely would reach — which is the adversary they actually face.

  • CAI: An Open, Bug Bounty-Ready Cybersecurity AI arXiv:2504.06017 (Mayoral-Vilches, Navarrete-Lozano, Sanz-Gómez, Salas Espejo, Crespo-Álvarez, Oca-Gonzalez, Balassone, Glera-Picón, Ayucar-Carbajo, Ruiz-Alcalde, Rass, Pinzger, Gil-Uriarte — Alias Robotics / Univ. Klagenfurt) · 2025-04-08 VERIFIED against the arXiv abstract page. Title, lead authors, and 2025-04-08 submission date (v2 2025-04-09) are exact; Alias Robotics / Univ. Klagenfurt attribution is consistent with the author list (Stefan Rass and Martin Pinzger are Klagenfurt; Mayoral-Vilches leads Alias Robotics). An open-source framework of specialized security agents aimed at democratizing advanced security testing. Defensive lesson: Marks the transition from research artifact to publicly available, bug-bounty-ready tooling — the point at which capability stops being gated by lab access. Defenders should assume such frameworks are already pointed at their internet-facing assets, and should adopt the same tooling for continuous self-assessment rather than waiting for an annual manual pentest. When citing the headline speed numbers, note the 3,600x figure is a best-case single-task result and the honest average is 11x — vendor-style maxima from a self-published framework paper deserve the same scrutiny as any benchmark claim.

  • Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks arXiv (Singer, Lucas, Adiga, Jain, Bauer, Sekar — Carnegie Mellon) · 2025-01-27 CONFIRMED REAL, no corrections needed. Fetched the arXiv abstract page directly and specifically checked the title, since a system name embedded in a title is a common conflation point: it is genuine and matches exactly, including "Incalmo." Author list matches exactly and in order (Brian Singer, Keane Lucas, Lakshmi Adiga, Meghna Jain, Lujo Bauer, Vyas Sekar); v1 submission date is 27 January 2025, matching the stated date (the paper has since been revised through v4, 22 November 2025 — the v1 date is the correct citation date). Defensive lesson: The two headline numbers are the ones to internalize: 3-of-40 to 37-of-40 came from better abstraction, not a better model — and the marginal cost of a full multi-host network compromise was under $15. Economic friction has historically been a real deterrent for mid-tier attackers; that assumption should be retired. The right abstraction layer is the force multiplier to watch.

  • VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework arXiv (State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, Chinese Academy of Sciences + School of Cyber Security, University of Chinese Academy of Sciences — Kong, Hu, Ge, Li, Li, Wu) · 2025-01-23 VERIFIED against the PDF title page. Confirmed: decomposes engagements into three phases — reconnaissance, scanning, exploitation — coordinated by a Penetration Task Graph (PTG), with role specialization, penetration path planning, inter-agent communication, and generative penetration behavior. Baselines confirmed as GPT-4 and Llama3. CAS attribution confirmed and made precise: authors are at the State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, CAS, with several jointly at the School of Cyber Security, University of CAS. Defensive lesson: Its explicit phase decomposition mirrors the kill chain defenders already model, which is an advantage: each phase boundary is a detection opportunity. Because the agent's structure is legible and its reconnaissance and scanning stages remain noisy, early-phase detection stays viable even against a fully automated adversary.

  • HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing arXiv (Muzsai, Imolai, Lukács — Eötvös Loránd University) · 2024-12-02 CONFIRMED REAL, no corrections needed. Fetched the arXiv abstract page directly: title matches exactly; author list matches exactly (Lajos Muzsai, David Imolai, András Lukács); v1 submission date is December 2, 2024, matching the stated date. Defensive lesson: Demonstrates the low barrier to entry: a small academic team assembled a capable autonomous agent from public models and free CTF infrastructure. Defenders should calibrate to the reality that this tooling requires neither nation-state resources nor novel research — the capability is broadly reproducible.

  • PentestAgent: Incorporating LLM Agents to Automated Penetration Testing arXiv / ASIA CCS 2025 (Northwestern University / Zhejiang University / Ant Group — Shen, Wang, Li, Chen, Zhao, Sun, Wang, Ruan) · 2024-11-07 VERIFIED, metadata CORRECTED. Title, author list, and date are right, but the source line was INCOMPLETE: the PDF title page shows three institutions, not two. Ant Group was omitted — Wencheng Zhao and Dawei Sun are both at Ant Group (Hangzhou). Full breakdown: Xiangmin Shen, Lingzhi Wang, Yan Chen at Northwestern (Evanston); Zhenyuan Li, Jiashui Wang, Wei Ruan at Zhejiang University; Zhao and Sun at Ant Group. Shen and Wang contributed equally. Added venue: published at ASIA CCS '25 (Hanoi, Vietnam). Defensive lesson: The RAG component matters defensively: grounding agents in retrieved, current vulnerability knowledge decouples their capability from the model's training cutoff. An agent with retrieval can act on a CVE published yesterday, so defenders cannot rely on knowledge-cutoff lag as a protective buffer. Secondary note now that attribution is corrected: offensive-agent tooling is being co-developed by a major payments-sector industrial lab, not only by academics — a signal about where this capability actually accrues.

  • Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects arXiv — Heiding, Lermen, Kao, Schneier, Vishwanath · 2024-11-30 In a study on 101 human participants across four email conditions, fully AI-automated spear-phishing emails achieved a 54% click-through rate, matching human expert performance (54%) and far exceeding the 12% control; an AI-with-human-in-the-loop condition reached 56%. The automated pipeline collected accurate and useful target information 88% of the time (inaccurate profiles in only 4% of cases), and the authors estimate AI could increase attacker profitability by up to 50x for larger audiences. The same paper found Claude 3. Defensive lesson: The key measured result for defenders: end-to-end automated spear phishing reaches expert-human effectiveness (54% vs 12% control), collapsing the economics that previously restricted tailored attacks to high-value targets — assume everyone is now worth spear-phishing. Crucially, the same paper shows the defensive side of the ledger: LLMs detected phishing intent at >90%, so the response to AI-generated phishing is AI-assisted filtering plus phishing-resistant auth, not more 'spot the typo' training.

  • EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities arXiv (Tel Aviv / NYU / Princeton — Abramovich, Udeshi, Shao et al.; Narasimhan, Karri, Press) · 2024-09 VERIFIED. Title matches the current arXiv version verbatim; v1 submitted 24 September 2024 (v2 Feb 2025, v3 Jun 2025), so the 2024-09 date is correct. Author list confirmed — Talor Abramovich, Meet Udeshi, Minghao Shao, Kilian Lieret, Haoran Xi, Kimberly Milner, Sofija Jancheska, John Yang, Carlos E. Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique, Karthik Narasimhan, Ramesh Karri, Ofir Press — and the Tel Aviv / NYU / Princeton attribution is correct. Defensive lesson: Capability came from tool interfaces, not a better model — the same weights get substantially more dangerous when handed a debugger, which means access control over tooling is a real and tractable control surface. The 'soliloquizing' finding is a gift to defenders and evaluators alike: agents may report success they never achieved, so any claimed result — in an eval or in an incident report — must be verified against ground-truth environment state rather than the agent's own narration.

  • PenHeal: A Two-Stage LLM Framework for Automated Pentesting and Optimal Remediation arXiv:2407.17788 (Junjie Huang — NYU Shanghai; Quanyan Zhu — New York University) · 2024-07-25 VERIFIED against the arXiv abstract page and the PDF byline. Title, authors, and 2024-07-25 submission date are exact. A two-stage framework pairing automated pentesting with remediation recommendation, using Counterfactual Prompting and an Instructor module to ground the LLM in external knowledge. The abstract's figures are quoted correctly: 31% greater vulnerability coverage, 32% improved remediation effectiveness, and 46% lower cost versus baselines. NYU attribution confirmed from the PDF ([email protected], NYU Shanghai; quanyan.zhu@nyu. Defensive lesson: A concrete template for dual-use research done constructively — the same agent that finds the issue proposes and prioritizes the fix under cost constraints. Defensive teams should demand this pairing from any offensive tooling they adopt: findings without prioritized, cost-aware remediation add alert volume without reducing risk. Caveat when citing: the 31/32/46% figures are the authors' own self-reported comparisons against their chosen baselines, not independently replicated results.

  • Teams of LLM Agents can Exploit Zero-Day Vulnerabilities arXiv (Zhu, Kellermann, Gupta, Li, Fang, Bindu, Kang — UIUC) · 2024-06-02 CONFIRMED REAL, no corrections needed. Fetched the arXiv abstract page directly: title matches exactly; author list matches exactly and in order (Yuxuan Zhu, Antony Kellermann, Akul Gupta, Philip Li, Richard Fang, Rohan Bindu, Daniel Kang); v1 submission date is June 2, 2024, matching the stated date. Defensive lesson: Multi-agent decomposition partially removes the CVE-description crutch that capped the one-day results, so defenders should not treat the 7% no-description figure as a durable ceiling. The recurring pattern across this literature is that scaffolding and orchestration — not just base model weights — drive capability jumps, meaning capability can rise between model releases. (Caveat: the paper's 'zero-day' means unknown-to-the-agent, not undisclosed-to-vendor — do not read the title as evidence of true zero-day discovery.)

  • LLM Agents can Autonomously Exploit One-day Vulnerabilities arXiv (Fang, Bindu, Gupta, Kang — UIUC) · 2024-04-11 CONFIRMED REAL, no corrections needed. Fetched the arXiv abstract page directly: title matches exactly; author list matches exactly (Richard Fang, Rohan Bindu, Akul Gupta, Daniel Kang — correctly omitting Zhan, who appears on the earlier websites paper); v1 submission date is April 11, 2024, matching the stated date. Headline statistics verified verbatim against the abstract: a dataset of 15 one-day CVEs, GPT-4 "capable of exploiting 87% of these vulnerabilities compared to 0% for every other model" tested (including GPT-3. Defensive lesson: The 87%-to-7% collapse without the CVE text is the single most actionable finding here: agent capability was largely parasitic on public vulnerability disclosure, not on independent discovery. This sharpens the patch-window problem — once an advisory publishes, weaponization is cheap and fast — and makes time-to-patch after disclosure, not obscurity, the controlling defensive variable. It also complicates disclosure-detail norms.

  • AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks arXiv (Xu, Stokes, McDonald, Bai, Marshall, Wang, Swaminathan, Li — UC Irvine / Microsoft) · 2024-03-02 CONFIRMED REAL, no corrections needed. Fetched the arXiv abstract page directly: title matches exactly; author list matches exactly and in order (Jiacen Xu, Jack W. Stokes, Geoff McDonald, Xuesong Bai, David Marshall, Siyue Wang, Adith Swaminathan, Zhou Li); v1 submission date is Saturday, 2 March 2024, matching the stated date. Defensive lesson: Shifts attention to the post-compromise phase where most defensive detection actually lives. If lateral movement and privilege escalation become cheap and automatable, controls that assume attacker dwell-time is expensive — and that a human is slow enough to be caught between stages — need revisiting. Also a model of responsible framing: build detection ahead of the capability's arrival.

  • LLM Agents can Autonomously Hack Websites arXiv (Fang, Bindu, Gupta, Zhan, Kang — UIUC) · 2024-02-06 CONFIRMED REAL, no corrections needed. Fetched the arXiv abstract page directly: title matches exactly ("LLM Agents can Autonomously Hack Websites"); author list matches exactly and in order (Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, Daniel Kang); v1 submission date is February 6, 2024, matching the stated 2024-02-06. Defensive lesson: Offensive capability is discontinuous across the model frontier, not a smooth curve. Defensive threat models must be re-derived on each frontier release rather than assumed stable, and 'attacker doesn't know our vulnerability yet' is a weak assumption when the agent can discover it. Note the sharp capability gap between frontier and open models — a gap that later papers show narrowing.

  • Malla: Demystifying Real-world Large Language Model Integrated Malicious Services arXiv / USENIX Security Symposium 2024 (Zilong Lin, Jian Cui, Xiaojing Liao, XiaoFeng Wang) · 2024-01-06 CONFIRMED against the source, with every figure exact. Researchers systematically studied 212 real-world Malla services in underground marketplaces, identifying eight distinct backend LLMs and 182 prompts used to circumvent protections on public LLM APIs. Two dominant tactics: abusing uncensored open models, and exploiting public APIs via jailbreak prompts. METADATA CORRECTION: the title supplied was a description, not the paper's title — the actual title is 'Malla: Demystifying Real-world Large Language Model Integrated Malicious Services. Defensive lesson: This is the strongest empirical baseline for the criminal-LLM ecosystem, and its key finding is that a small set of backend models and a finite, reusable prompt corpus underpin most services. That finiteness is exploitable: providers can fingerprint and detect known circumvention prompt families at the API layer, and defenders should track backend-model concentration rather than the churning brand names on top.

  • PentestGPT: An LLM-empowered Automatic Penetration Testing Tool arXiv (Deng, Liu, Mayoral-Vilches, et al.) · 2023-08-13 CONFIRMED REAL, no corrections needed. Fetched the arXiv abstract page directly: title matches exactly; v1 submission date is August 13, 2023, matching the stated date. Full author list is Gelei Deng, Yi Liu, Víctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, Stefan Rass — the item's "Deng, Liu, Mayoral-Vilches, et al." is a correct leading-three abbreviation. Claims verified verbatim: LLMs show "proficiency in specific sub-tasks... Defensive lesson: Identifies the real early bottleneck as long-horizon context retention, not individual exploit skill — and shows that architecture, not model scale alone, is what closes it. For defenders, this predicts that attacks will be assembled from many individually-mundane, individually-plausible actions; detection tuned to single dramatic events will miss them.


Capability Benchmarks

  • Demystifying AI Exploits: A Blueprint for AI-Assisted Vulnerability Management Mandiant Consulting / Google Threat Intelligence Group (author: Jules Czarniak) · 2026-07-16 A Mandiant blueprint assessing where LLM agents genuinely help in vulnerability management. Its 'Binary and architectural oracles' section finds LLMs perform well on bug classes with binary, observable oracles — memory corruption in memory-unsafe codebases (C, C++, Assembly) where the system gives an objective 'crash or no crash' feedback loop — but struggle with vulnerabilities needing architectural oracles: authorization bypasses, complex business-logic flaws and indirect SSRF that require business context and cross-service trust boundaries. Defensive lesson: Gives defenders a realistic capability map instead of a marketing claim: deploy AI agents where an automated oracle can validate the finding, and keep humans on architectural and business-logic review. Its operational guardrails are directly actionable — treat source code as untrusted input because dependencies and comments can carry indirect prompt injection, sandbox agents with least-privilege short-lived tokens, require pull requests over direct commits, and pin model versions for auditability.

  • DARPA AI Cyber Challenge (AIxCC) final results: Team Atlanta wins $4M Trail of Bits · 2025-08-09 VERIFIED against the primary source, plus independent confirmation of Team Atlanta's details. DARPA announced AIxCC final results on 2025-08-08 (Trail of Bits post published 2025-08-09). Placements and prizes confirmed exactly: Team Atlanta 1st / $4M; Trail of Bits' Buttercup 2nd / $3M; Theori 3rd / $1.5M. Team Atlanta's composition (Georgia Tech, Samsung Research, KAIST, POSTECH) and system name (Atlantis, an 'AI-powered cyber reasoning system') independently confirmed via the team's own site. Defensive lesson: The finals establish that end-to-end autonomy is real but incomplete — the second-place system patched 19 of the 28 vulnerabilities it found, leaving 9 (roughly a third) found-but-unfixed. The arithmetic checks out, but read the gap carefully: the source does not say the 9 were unfixable, and competition scoring/time limits may account for some, so treat 'found-to-fix gap' as a planning heuristic rather than a measured capability ceiling. It remains the place human reviewers belong: expect autonomous pipelines to surface more than they can safely close. Defenders should also note the discovery side scaled across 23 real repositories, so open-source maintainers should anticipate AI-sourced reports as a routine inbound category.

  • Establishing Best Practices for Building Rigorous Agentic Benchmarks arXiv (UIUC / Stanford / Berkeley et al. — Zhu, Jin, Pruksachatkun, A. Zhang, Kapoor, Liang, Kang and others) · 2025-07 VERIFIED. Title matches verbatim; v1 submitted 3 July 2025, so the 2025-07 date is correct. Author list confirmed — led by Yuxuan Zhu and Tengjun Jin, including Yada Pruksachatkun, Andy Zhang, Sayash Kapoor, Matei Zaharia, Ion Stoica, Percy Liang and Daniel Kang among ~25 authors — and the UIUC / Stanford / Berkeley et al. attribution is correct. Audited widely-used agentic benchmarks and found task-validity and outcome-validity flaws (e.g. Defensive lesson: Benchmarks can be wrong in both directions, and by amounts larger than the year-over-year gains they are used to report — an agent may be credited with a solve it reached by a shortcut, or denied one on a grading artifact. Before any cyber-risk decision is anchored to a benchmark number, check whether the benchmark itself has been validated; a headline score is a measurement with error bars and a methodology, and treating it as ground truth is its own security failure.

  • AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models arXiv:2506.14682 (Dreadnode — Ads Dawson, Rob Mulla, Nick Landers, Shane Caldwell) · 2025-06-17 VERIFIED against the arXiv abstract page. Title, authors, and Dreadnode attribution are exact. DATE REFINED from '2025-06' to the precise submission date 2025-06-17. 70 black-box CTF challenges from the Crucible environment targeting AI/ML systems rather than conventional software — confirmed. The reported leaders are confirmed: Claude-3.7-Sonnet solved 43/70 (61%) and Gemini-2.5-Pro 39/70 (56%). CLARIFICATION on those figures: 61%/56% are challenge-solve rates; the paper separately reports lower per-attempt 'overall success rates' (46.9% and 34. Defensive lesson: AI systems are now themselves a target class with their own attack surface, and models are markedly better at attacking AI (prompt injection) than at classical system exploitation. Anyone deploying LLM applications should red-team the model layer as a first-class asset — injection, extraction, guardrail bypass — and assume adversaries can automate that probing far faster than a human review cycle. The uneven capability profile is the actionable part: the model layer is the soft spot, not the OS underneath it.

  • BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems arXiv (Stanford — Andy K. Zhang, Joey Ji, Ben Menders et al., 34 authors incl. Dawn Song, Dan Boneh, Percy Liang) · 2025-05-21 VERIFIED, with one framing correction. Confirmed exactly: 10 agents, 25 real-code systems, Detect/Exploit/Patch task triad, 40 bug bounties worth $10 to $30,485. OWASP figure sharpened: the paper covers 9 of the OWASP Top 10 Risks, not the full Top 10 as 'covering OWASP Top 10 categories' implies. Numbers correction: the 90% Patch score belongs specifically to top performer Codex CLI: o3-high (which scored only 12.5% on Detect); Claude Code reached 87.5% on Patch. The quoted 17.5-67. Defensive lesson: Still the most defense-relevant result in the set, with a caveat: agents scored higher on patching than on exploiting, suggesting current capability is not inherently offense-favoring — but the cleanest comparison is within-population (custom agents: 17.5-67.5% exploit vs 25-60% patch), which is a much narrower gap than 90-vs-17.5 suggests. The genuinely striking number is Detect at 12.5% even for the best agent: agents are far better at fixing a vulnerability they are pointed at than at finding one. That is the real asymmetry to exploit — deploy agents on patching known findings, not on discovery — and the dollar-denominated framing makes that case to budget holders.

  • AutoPatchBench: A Benchmark for AI-Powered Security Fixes Meta · 2025-04-29 Meta released a standardized benchmark for AI systems that automatically repair fuzzing-discovered vulnerabilities, comprising 136 fuzzing-identified C/C++ vulnerabilities in real-world repositories with verified fixes, plus a 113-sample 'Lite' subset. Results showed roughly 60% patch generation success but only 5-11% of patches correct after verification via fuzzing and differential testing — Gemini 1.5 Pro produced a 61.1% generation success rate with only 5.3% of the total set correct. Meta's differential testing validation achieved 84.1% accuracy with 100% recall but 41.7% precision. Defensive lesson: A patch that compiles and stops the crash is not a correct patch. Success rates collapsed from ~60% to 5-11% once rigorous verification was applied — meaning naive acceptance criteria would have merged a large majority of subtly wrong security fixes. Never accept an AI-generated security patch on the evidence that 'the fuzzer stopped crashing'; require differential testing and human review, and note that even the good verifier here had 41.7% precision.

  • A Framework for Evaluating Emerging Cyberattack Capabilities of AI arXiv (Google DeepMind / Google Threat Intelligence) · 2025-03-14 Mikel Rodriguez, Raluca Ada Popa, Four Flynn, Lihao Liang, Allan Dafoe and Anna Wang present an evaluation framework grounded in empirical threat data rather than hypotheticals, analyzing over 12,000 real-world instances of AI involvement in cyber incidents catalogued by Google's Threat Intelligence Group, from which they curate seven representative attack-chain archetypes. Defensive lesson: The methodological anchor for the whole slice: evaluate AI cyber risk end-to-end across the attack chain, not on isolated capability benchmarks, because uplift at a non-bottleneck phase does not change real-world outcomes. For defenders and evaluators this reframes prioritization — invest defensive effort at the phases where AI actually relieves an attacker constraint, and treat benchmarks that ignore attack-chain context as weak evidence either way.

  • CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities arXiv / ICML 2025 (UIUC, Siebel School of Computing and Data Science — Zhu, Kellermann, Bowman et al.; corresponding author Daniel Kang) · 2025-03-21 VERIFIED against the PDF title page. UIUC attribution confirmed and is stronger than stated: ALL 16 authors share the single affiliation 'Siebel School of Computing and Data Science, University of Illinois, Urbana-Champaign', with correspondence to Daniel Kang. Confirmed: benchmark built from critical-severity CVEs in real web applications via a sandbox framework, explicitly positioned against 'abstracted Capture-the-Flag' benchmarks. Defensive lesson: A modest headline solve rate against real CVEs is not reassurance — it is a floor that moves, and it is an 'up to' ceiling for the best agent rather than a typical result. The defensive action is patch-latency triage: CVE-Bench measures agent capability against already-disclosed, already-patchable bugs, so the exposure it models is unpatched estate, not novel attacker genius. Shrinking mean-time-to-patch on internet-facing web apps directly removes the surface this benchmark scores against.

  • AutoPenBench: Benchmarking Generative Agents for Penetration Testing arXiv (Politecnico di Torino / Università di Torino / NEC Laboratories Europe — Gioacchini, Mellia, Drago, Delsanto, Siracusano, Bifulco) · 2024-10-04 VERIFIED against the PDF title page. Confirmed: 33 tasks of increasing difficulty spanning in-vitro and real-world scenarios; fully autonomous agent 21% success rate (solving 27% of simple tasks and only one real-world task) versus 64% for the assisted semi-autonomous agent. Affiliations confirmed exactly: Gioacchini and Mellia at Politecnico di Torino, Drago and Delsanto at Università di Torino, Siracusano and Bifulco at NEC Laboratories Europe (Heidelberg). Date refined to v1 submission 2024-10-04 (v2 2024-10-28). Defensive lesson: The 21%-to-64% gap is the single most useful number here: the human-in-the-loop is the force multiplier, not the model. Threat models should therefore center the skilled operator using AI as an accelerant — the realistic near-term adversary — rather than the fully autonomous agent, and capability forecasts that only measure autonomous mode will systematically understate real risk. Caveat worth carrying: the autonomous agent solved only ONE real-world task, so the 21% headline is itself flattered by the in-vitro portion.

  • Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities arXiv:2410.09114 (Apart Research — Andrey Anurin, Jonathan Ng, Kibo Schaffer, Jason Schreiber, Esben Kran) · 2024-10-10 VERIFIED against the arXiv abstract page. Title, authors, and the 3CB acronym are exact. DATE REFINED from '2024-10' to the precise v1 submission date 2024-10-10 (v2 2024-11-02). Apart Research attribution is consistent with the author list (Esben Kran co-founded Apart Research). A benchmark and tooling explicitly aimed at catastrophic-risk-relevant offensive capability, spanning reconnaissance and exploitation. The item's claim about the capability gap is confirmed by the abstract: frontier models of the era (GPT-4o, Claude 3. Defensive lesson: 3CB's motivating claim is the lesson: evaluation robustness lags capability, so absence of evidence in a weak eval is not evidence of absence. Defenders and policymakers should ask what elicitation effort produced a reported score before treating it as a safety margin, and should prefer evaluations that report their own scaffolding rather than a bare number. Note the frontier/open-source gap it documents is pinned to late-2024 models and should not be assumed to still hold.

  • Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models arXiv (Stanford — Andy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji et al.; Boneh, Ho, Liang) · 2024-08-15 VERIFIED. Confirmed: 40 professional-level CTF tasks from 4 distinct competitions, with subtask decomposition and human first-solve-time anchoring so model performance is comparable to expert effort. Anchor range confirmed from the abstract: fastest human first-solve 11 minutes, hardest task took human teams nearly 25 hours. Date refined from '2024-08' to the v1 submission date 2024-08-15 (latest revision 2025-04-12). Defensive lesson: Cybench's human first-solve-time anchor is the transferable idea: measure AI capability in units of expert hours displaced, not raw solve percentage. Defenders should track that ratio over time for their own threat models, because 'hours of expert work an attacker no longer needs' converts directly into how fast an adversary can move and how much your detection window shrinks.

  • AIxCC semifinals at DEF CON 32: 39 teams compete, 7 advance to finals Trail of Bits · 2024-08-12 DARPA's AI Cyber Challenge held its semifinal round at DEF CON 32, with 39 teams competing and the top 7 advancing to the finals. Trail of Bits' Buttercup placed in the top 7, reporting it was first to successfully patch an nginx vulnerability, first to patch 6 bugs overall, and first to discover 3 bugs. Trail of Bits also noted that Team Atlanta's CRS discovered a real null dereference bug in SQLite during the competition. All figures verified against the source post. Defensive lesson: Even at the semifinal stage, competition systems found a real, previously unknown bug in SQLite rather than only the synthetic bugs they were scored against. Capability demonstrated on a benchmark leaks into genuine discovery, so defenders should assume competition-grade autonomous systems generalize beyond their evaluation set — and should not treat 'it was only tested on synthetic vulns' as reassurance.

  • NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security arXiv / NeurIPS 2024 Datasets and Benchmarks Track (New York University + NYU Abu Dhabi — Shao, Jancheska, Udeshi, Dolan-Gavitt, Garg, Karri, Shafique et al.) · 2024-06-08 VERIFIED against the PDF body. CSAW attribution confirmed verbatim: 'Our benchmark is based on the CTF competition of New York University's (NYU) annual Cybersecurity Awareness Week (CSAW).' Specifics confirmed: 200 validated challenges (drawn from an initial pool of 568) spanning CSAW years 2011-2023, across six categories — crypto, forensics, pwn, rev, misc, web — matching the summary's category list. Automated containerized harness using LLM function-calling confirmed; five LLMs evaluated. Venue is NeurIPS 2024 D&B Track. Defensive lesson: Open benchmarks are dual-use infrastructure: the same harness that lets defenders measure risk lets others tune agents against it. The defensive value is that public, reproducible measurement keeps the capability curve visible to regulators and blue teams instead of known only to labs — so support and fund open eval infrastructure, but treat any specific published score as perishable and re-measure rather than cite.

  • Project Naptime: Evaluating Offensive Security Capabilities of Large Language Models Google Project Zero (Sergei Glazunov, Mark Brand) · 2024-06-20 CORRECTED: Project Zero published a framework giving LLMs the tools a human researcher uses — a code browser, Python interpreter, debugger, and reporter — and re-ran the CyberSecEval 2 benchmark. With Naptime scaffolding, GPT-4 Turbo scored 1.00 on the 'Buffer Overflow' tests (up from 0.05 in the original paper) and 0.76 on 'Advanced Memory Corruption' (up from 0.16 — the originally supplied figure of 0.24 was incorrect). The 'up to 20x' improvement the authors claim refers specifically to the Buffer Overflow result (0.05 to 1.00); the Advanced Memory Corruption gain is roughly 4.75x, not 20x. Defensive lesson: Benchmark scores measure the scaffolding at least as much as the model. A model that looks incapable under one-shot prompting can be highly capable when given tools and iteration, which means published 'the model can't do X' results are lower bounds with a short shelf life. Defenders sizing AI threat must evaluate model-plus-tooling as a system, never the raw model in isolation.

  • Devising and Detecting Phishing: Large Language Models vs. Smaller Human Models arXiv — Heiding, Schneier, Vishwanath, Bernstein, Park · 2023-08-23 An earlier study with 112 participants compared phishing emails by origin: generic phishing (control) produced 19–28% click-through, GPT-4-generated 30–44%, manually designed V-Triad emails 69–79%, and GPT-4 combined with V-Triad 43–81%. The authors analyzed the economics of AI-enabled phishing, showing how large language models 'can increase the incentives of phishing and spear phishing by reducing their costs. Defensive lesson: A useful corrective to overclaiming, and instructive read alongside the authors' 2024 follow-up: in 2023 unaided AI still trailed expert human design, with the strongest results coming from AI plus human psychological technique. The trend line between the two papers — AI-only rising to match experts within roughly a year — is the real teaching point for threat modeling. The consistent finding across both is cost reduction: AI makes high-quality phishing cheap and scalable, which changes volume and targeting long before it changes maximum sophistication.

  • InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback (demonstrates Capture-the-Flag as a task; basis of the InterCode-CTF environment) arXiv:2306.14898 (Princeton — John Yang, Akshara Prabhakar, Karthik Narasimhan, Shunyu Yao) · 2023-06-26 VERIFIED against the arXiv abstract page. Title and authors are exact. DATE REFINED from '2023-06' to the precise v1 submission date 2023-06-26 (last revised 2023-10-30). Framed interactive coding as a reinforcement-learning environment with self-contained Docker sandboxes for Bash (NL2Bash), SQL (Spider), and Python (MBPP) with execution feedback — all three environments confirmed. The CTF claim is correct but worth stating precisely: the abstract says InterCode 'can even be used to create new tasks such as Capture the Flag,' i.e. Defensive lesson: The key insight predates the security framing: execution feedback in a live shell is what converts a language model into an operator. Defenders should treat tool access and an execution loop — not model weights alone — as the capability threshold worth monitoring, and instrument for agentic command patterns (rapid iterative probing driven by error output) rather than for static known-bad payloads.


Autonomous Vulnerability Discovery (the defensive mirror)

  • GTIG AI Threat Tracker: AI-assisted zero-day exploitation and AI supply chain attacks Google Threat Intelligence Group · 2026-05-11 VERIFIED with three corrections (report's actual title: 'GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access'). CORRECTION 1 — the original said a zero-day in a 'Python web administration tool'; GTIG actually reports a zero-day 2FA bypass in 'a popular open-source, web-based system administration tool,' with the EXPLOIT 'implemented in a Python script.' Python is the exploit language, not the tool. Defensive lesson: This marks AI crossing from productivity aid into vulnerability discovery producing a real zero-day, so the earlier 'no novel capabilities' assessment should be taught as time-bounded rather than permanent — defenders should shorten patch windows for internet-facing admin tooling. Teach the attribution precisely: GTIG assesses an AI model was used but explicitly does NOT believe Gemini was the model, so this is not a 'Gemini found a zero-day' story. The standout detection insight is that the AI-written exploit had tells — educational docstrings, a hallucinated CVSS score, textbook Pythonic structure — so hunt for those artifacts in captured exploit code. GTIG also notes why LLMs helped here: fuzzers and static analysis chase sinks and crashes, while frontier LLMs excel at semantic logic flaws, reading developer intent to spot contradictions between 2FA enforcement logic and hardcoded exceptions. Finally, the AI toolchain is now itself the supply chain target — inventory and pin your AI gateway and CLI dependencies.

  • Introducing CodeMender: an AI agent for code security Google DeepMind · 2025-10-06 Google DeepMind introduced CodeMender, an agent leveraging 'the thinking capabilities of recent Gemini Deep Think models' that both reactively patches newly discovered vulnerabilities and proactively rewrites code to eliminate whole vulnerability classes, using static/dynamic analysis, fuzzing, and SMT solvers. Verbatim: 'Over the past six months that we've been building CodeMender, we have already upstreamed 72 security fixes to open source projects, including some as large as 4.5 million lines of code. Defensive lesson: The frontier is shifting from finding individual bugs to eliminating bug classes — the libwebp work would have neutralized CVE-2023-4863 not by patching that bug but by making the class structurally unreachable, which is a far better return than winning each bug race individually. Two things defenders should carry forward: hardening annotations and memory-safe rewrites are now cheap enough to apply at multi-million-line scale, and even Google keeps a human in the loop on every upstreamed patch — mirror that guardrail rather than skipping it.

  • Automate security reviews with Claude Code (/security-review) Anthropic · 2025-08-06 VERIFIED against the primary source; no corrections needed. Anthropic shipped a /security-review command for ad-hoc terminal-based analysis plus a GitHub Actions integration that reviews pull requests and posts inline findings. The security-focused prompt checks for SQL injection risks, cross-site scripting (XSS), authentication and authorization flaws, insecure data handling, and dependency vulnerabilities — all five classes confirmed. Defensive lesson: The cheapest place to catch a vulnerability is before the PR merges, and both of the vendor's own cited examples were caught in review rather than production. Moving AI review left into CI converts vulnerability discovery from an incident-response activity into a build-time gate. Weight these results appropriately: both examples are vendor self-reports about the vendor's own product, with no independent evaluation. The realistic framing: this catches known-shaped bugs (injection, SSRF, authn flaws) at high volume and low cost — it is a floor-raiser, not a replacement for targeted review of novel logic.

  • Buttercup, an AIxCC-grade autonomous vulnerability discovery and patching system, released open source Trail of Bits · 2025-08-08 VERIFIED against the primary source, with one material correction to the original item. Trail of Bits open-sourced Buttercup at github.com/trailofbits/buttercup, described as 'a fully automated, AI-driven system for discovering and patching vulnerabilities in open-source software. Defensive lesson: This is the dual-use inflection point made concrete: competition-grade autonomous find-and-patch is now free and runs on a laptop. The same capability that lets a maintainer continuously scan their own repository lets anyone scan anyone's. The defensive imperative is asymmetric in defenders' favor only if they move first — you have legitimate access to your own source and can run these tools before disclosure, while an attacker works from the outside. Run it on your code now, because your adversary has the identical binary. Temper expectations on speed, though: the sub-10-minute figure is a planted-bug demo, and the AIxCC finals numbers (28 found / 19 patched) are the better guide to real performance.

  • XBOW reaches #1 on the HackerOne leaderboard XBOW · 2025-08-18 VERIFIED that the cited post makes this claim, but flagged for sourcing caution. XBOW's retrospective (2025-08-18) does state it reached '#1 globally in Q2' — the original item's wording is faithful to the post. Confirmed quotes: 'The leaderboard was never the goal in itself, but it became the ultimate benchmark for our founding question' and 'It proved that an AI can indeed perform at the highest level of security research.' The post describes a shift toward pre-production deployment rather than continued leaderboard competition. Defensive lesson: An autonomous system out-submitted every human researcher on a major bug bounty platform within a single quarter — per the vendor's own account. This cuts two ways for defenders: your bounty inbox will receive valid findings faster than human triage can absorb them, and identical scanning throughput is available to unaligned actors against the same attack surface. Bug bounty programs need AI-aware triage capacity and explicit policy on automated submissions — and note the contrast with the curl case, where AI submissions were mostly noise: the same technology produces both signal and slop depending on who wields it, and on whether the operator bears any cost for a bad submission.

  • Big Sleep discovers CVE-2025-6965 in SQLite, foiling in-the-wild exploitation Google (The Keyword) / Google Threat Intelligence · 2025-07-15 In 'A summer of security: empowering cyber defenders with AI', Google reported that the Big Sleep agent discovered CVE-2025-6965 in SQLite, describing it as 'a critical security flaw, and one that was known only to threat actors and was at risk of being exploited.' Google stated: 'Through the combination of threat intelligence and Big Sleep, Google was able to actually predict that a vulnerability was imminently going to be used and we were able to cut it off beforehand. We believe this is the first time an AI agent has been used to directly foil efforts to exploit a vulnerability in the wild. Defensive lesson: The strongest available evidence that AI vulnerability discovery can favor defenders, not just attackers — threat intelligence about attacker interest was used to aim an AI agent at the right target and close the hole first. The transferable pattern for defenders is the pairing: intel signals where adversaries are looking, and an automated agent races them to the bug. Defensive AI value is highest when it is intelligence-directed rather than untargeted scanning.

  • Big Sleep discovers CVE-2025-6965 in SQLite, a flaw 'known only to threat actors' Google · 2025-07-15 VERIFIED against the primary source; the CVE is real and correctly attributed (not an invented identifier). In 'A summer of security: empowering cyber defenders with AI' (15 July 2025), Google reported that its Big Sleep agent discovered CVE-2025-6965 in SQLite, described as 'a critical security flaw, and one that was known only to threat actors and was at risk of being exploited' — the original item truncated this quote, dropping 'and was at risk of being exploited', which is restored here. Defensive lesson: The first CLAIMED case of an AI agent pre-empting in-the-wild exploitation — note this is Google's own characterization ('We believe'), not an independently adjudicated first. The mechanism is what defenders should copy: threat intelligence indicated something was about to be exploited, and an agent was then directed to go find it. Neither half works alone — intel without discovery capacity is just a warning, and discovery without targeting is unfocused. Pairing them compresses the window between adversary knowledge and defender knowledge, which is the window that actually determines outcomes.

  • CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale arXiv:2506.02548 (UC Berkeley — Zhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai, Jialin Zhang, Dawn Song) · 2025-06-03 VERIFIED against the arXiv abstract page. Title and full author list are exact. DATE REFINED from '2025-06' to the precise v1 submission date 2025-06-03 (v3 revised 2026-03-24). UC Berkeley attribution is consistent (Dawn Song's group). All figures confirmed: 1,507 real-world vulnerabilities across 188 open-source projects; agents generate PoC tests reproducing a vulnerability from a description and codebase; top configurations reach only ~20% success; 34 zero-day vulnerabilities and 18 historically incomplete patches surfaced. Defensive lesson: This is the clearest published case of an evaluation producing real vulnerability discovery as a byproduct — including incomplete patches, meaning fixes that were believed to have closed a bug did not. Two defensive actions follow: verify that patches actually resolve the vulnerability rather than trusting the advisory, and recognize that the same scale economics benefit defenders, who can run this tooling across their own dependency tree first. The ~20% ceiling is also the story: real-world reproduction remains hard, and the 34 zero-days came from volume across 1,507 targets, not from high per-task reliability.

  • How I used o3 to find CVE-2025-37899, a remote zeroday in the Linux kernel's SMB implementation Sean Heelan (independent researcher) · 2025-05-22 Researcher Sean Heelan used OpenAI's o3 to find CVE-2025-37899, a use-after-free in the Linux kernel ksmbd SMB3 logoff command handler (smb2_session_logoff), affecting concurrent connections bound to the same session. Benchmarking against a known bug (CVE-2025-37778, a use-after-free during Kerberos authentication in krb5_authenticate), o3 found it in 8 of 100 runs, Claude Sonnet 3.7 in 3 of 100, and Claude Sonnet 3.5 in 0 of 100. On the larger all-handlers codebase (~12k LoC, ~100k tokens) o3 surfaced the benchmark bug in 1 of 100 runs while discovering CVE-2025-37899 in other outputs. Defensive lesson: Low per-run success rates are not a defense when runs are cheap. An 8% hit rate sounds weak until you notice 100 attempts cost $116 — the economics changed, not the hit rate. Treat AI vulnerability discovery as a lottery ticket an adversary can buy in bulk against your codebase. The corollary for evaluators: run-to-run variance is enormous, so a single negative result proves nothing about a model's capability.

  • Analyzing open-source bootloaders: Finding vulnerabilities faster with AI (Security Copilot) Microsoft · 2025-03-31 Microsoft used Security Copilot to assist researchers auditing three open-source bootloaders, disclosing 20 CVEs in total. Every CVE ID was individually checked against the source and all 20 match exactly: 11 in GRUB2 (CVE-2024-56737, CVE-2024-56738, CVE-2025-0677, CVE-2025-0678, CVE-2025-0684, CVE-2025-0685, CVE-2025-0686, CVE-2025-0689, CVE-2025-0690, CVE-2025-1118, CVE-2025-1125), 4 in U-boot (CVE-2025-26726 through CVE-2025-26729), and 5 in Barebox (CVE-2025-26721 through CVE-2025-26725). Defensive lesson: Microsoft framed the value honestly as researcher time saved (about a week), not autonomous discovery — AI as force multiplier on an expert workflow. The target selection is the real lesson for defenders: bootloaders sit below the OS and beneath Secure Boot's trust anchor, and legacy embedded C with slow patch cycles is exactly where AI-assisted review will raise finding rates fastest. Inventory that class of code before someone else audits it for you.

  • From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code Google Project Zero + Google DeepMind · 2024-11-01 The Big Sleep team announced that their LLM-driven agent discovered an exploitable stack buffer underflow in SQLite, a widely used open-source database engine — a mishandled sentinel value (-1) in the iColumn field when processing ROWID constraints, where the code failed to validate that column indices were in range before using them as array offsets. Defensive lesson: The founding datapoint for AI-assisted vulnerability discovery, and importantly a defensive one: the bug was fixed pre-release, so the finder's advantage went to defenders. It signals that memory-safety bugs in fuzzing-hardened targets are now reachable by AI agents, which means both that defenders should adopt these agents on their own critical dependencies and that they should expect attackers to gain the same reach.

  • From Naptime to Big Sleep: AI agent finds an exploitable memory-safety bug in SQLite Google Project Zero / Google DeepMind · 2024-11-01 The Big Sleep agent found a stack buffer underflow in SQLite's seriesBestIndex function, caused by mishandling of a sentinel value (-1) in the iColumn field for ROWID constraints, leading to a write into a stack buffer at a negative index; the post notes the write corrupted a pointer later dereferenced, 'a likely exploitable condition.' Google described it as 'the first public example of an AI agent finding a previously unknown exploitable memory-safety issue in widely used real-world software.' The team reported the issue 'remained undiscovered after 150 CPU-hours of fuzzing. Defensive lesson: AI agents find bugs fuzzers miss, but not by brute force — by reaching code paths the harness structurally cannot. The OSS-Fuzz harness here lacked the generate_series extension entirely (verified verbatim: 'The harness used by OSS-Fuzz isn't built with the generate_series extension enabled'), so no amount of CPU time would have found it. The defensive action is to audit which parts of your codebase your harnesses never touch: those coverage gaps are exactly where variant-analysis agents (and adversaries running them) will look.

  • OSS-Fuzz-Gen: LLM-generated fuzz targets find CVE-2024-9143 in OpenSSL and 30 other bugs Google (OSS-Fuzz team) · 2024-11 Google's OSS-Fuzz-Gen framework uses LLMs to automatically write fuzz targets for real-world C/C++/Java/Python projects, benchmarking them on the OSS-Fuzz platform. Defensive lesson: The highest-leverage use of LLMs here was augmenting proven tooling, not replacing it: the model wrote better harnesses and the existing fuzzer found the bugs. This is a cheap, immediately adoptable defensive pattern — point an LLM at your under-covered attack surface to generate harnesses, and let deterministic tooling do the actual finding and triage.


Guardrail Bypass

  • Threat actors socially engineering Gemini's safety guardrails via CTF and student pretexts Google Threat Intelligence Group (GTIG) · 2025-11-05 CONFIRMED against the primary source, with one METADATA CORRECTION: category changed from 'prompt-injection-exploit' to 'guardrail-bypass'. This is persona-based jailbreaking / social engineering of a refusal policy, not prompt injection (which smuggles instructions through untrusted data into a model's context) — the original label conflated two distinct attack classes. Substance verifies exactly: China-nexus actors (unattributed to a named group) reframed exploitation requests as capture-the-flag exercises ('I am working on a CTF problem'), and Iran-linked TEMP. Defensive lesson: Guardrails that evaluate a request's stated framing rather than its capability payload are defeated by cheap pretexting — a durable lesson for anyone deploying an LLM with a refusal policy. Defensive design should classify on what the output would enable, not on the claimed purpose, and should treat repeated reframing after a refusal as a strong abuse signal in its own right. Persona-based jailbreaks belong in every red-team suite.

Policy & Uplift Assessments

  • What we learned mapping a year's worth of AI-enabled cyber threats Anthropic · 2026-06-03 CONFIRMED against the primary source; no metadata corrections needed. Anthropic (Frontier Red Team) analyzed 832 accounts banned for malicious cyber activity between March 2025 and March 2026 and mapped them onto the MITRE ATT&CK framework, with some results published via Verizon's 2026 Data Breach Investigations Report. Every figure verifies: malware development was the most common AI-enabled activity (560 of 832 accounts, 67.3%); 54 accounts (6.5%) used AI to assist with lateral movement; account discovery rose 8.9% while AI-assisted phishing fell 8.6%. Defensive lesson: The most consequential finding for defenders is a measurement gap: the framework the industry uses to classify adversary behavior has no category for autonomous orchestration and multi-stage chaining, which is precisely the trait marking the highest-risk AI-enabled actors. Threat models and detection engineering keyed to ATT&CK technique coverage will systematically under-rate these actors until the taxonomy catches up — and the shift of AI use deeper into the attack lifecycle argues for weighting internal/post-compromise telemetry over perimeter phishing defense.

  • UK AISI — How fast is autonomous AI cyber capability advancing? UK AI Security Institute · 2026-05-13 VERIFIED. Title and date (13 May 2026) confirmed both on the page and independently against the AISI blog index. Reported that the length of cyber tasks frontier models can autonomously complete has been doubling roughly every 4.7 months since late 2024 — an acceleration from the approximately 8-month doubling AISI estimated in November 2025 — based on AISI's narrow cyber suite and cyber ranges, with a 2.5M token budget per task imposed for comparability across models over time. AISI names Claude Mythos Preview and GPT-5. Defensive lesson: A ~4.7-month doubling in task length is a planning constant, not a headline: it implies capability roughly quadruples within a typical annual security review cycle, so any control justified by today's model limits is likely stale before the next review. Defenders should shorten reassessment intervals for AI-relevant threat models and note AISI's own caveat that the measured figure is bounded by an imposed 2.5M-token budget — the trend is a reported estimate, and real capability under larger budgets is higher.

  • UK AISI — Our evaluation of OpenAI's GPT-5.5 cyber capabilities UK AI Security Institute · 2026-04-30 VERIFIED. Title and date (30 April 2026) confirmed on the page and against the AISI blog index. Independently evaluated GPT-5.5 on a suite of 95 tasks across four difficulty tiers (AISI's phrasing is 'narrow cyber tasks' rather than the original item's 'CTF-format tasks'; the advanced suite comprised 27 Practitioner and 21 Expert tasks) plus two cyber ranges — 'The Last Ones' (a 32-step corporate network attack) and 'Cooling Tower' (a 7-step ICS attack). Confirmed results: an average pass rate of 71.4% (±8. Defensive lesson: The 12-hours-to-10-minutes compression on a single reverse-engineering task is the number to internalize: obscurity and analysis cost have largely stopped being a defense. Equally important, a universal jailbreak found in six hours means model-level safety training should be treated as a speed bump, not a control — defenders relying on a vendor's refusal behavior to prevent misuse are relying on something with a measured, short time-to-bypass. Figures here are as reported by AISI.

  • UK AISI — Our evaluation of Claude Mythos Preview's cyber capabilities UK AI Security Institute · 2026-04-13 VERIFIED. Title and date (13 April 2026) confirmed on the page and against the AISI blog index. Pre-deployment evaluation reporting 73% success on expert-level tasks — a tier AISI states 'no model could complete before April 2025' — and the first end-to-end solve of the 32-step 'The Last Ones' corporate network range, in 3 of its 10 attempts, averaging 22 of 32 steps, with Claude Opus 4.6 the next best performer at an average of 16 steps. The model did not complete the OT-focused 'Cooling Tower' range and notably got stuck on that range's IT sections rather than the OT-specific steps. Defensive lesson: This documents a capability threshold being crossed in public: a task class unsolvable before April 2025 became routinely solvable within a year, which is what a fast doubling rate looks like from the inside. AISI's own caveat carries the defensive point — these ranges have no active defenders, so the results measure attack capability against an undefended target. That is an argument for investing in the detection and response layer the evaluation deliberately omits, which is precisely where real environments still have advantage.

  • UK AISI — How do frontier AI agents perform in multi-step cyber-attack scenarios? UK AI Security Institute · 2026-03-16 VERIFIED. Title and date (16 March 2026) confirmed on the page and against the AISI blog index. Evaluated seven LLMs spanning August 2024 to February 2026 on cyber ranges, reporting progress on the 32-step 'The Last Ones' corporate network range from about 1.7 steps (GPT-4o) to about 9.8 steps (Opus 4.6) at a 10M token budget; the page states verbatim that 'the best-performing Opus 4.6 run completed 22 of 32 steps.' Raising the budget from 10M to 100M tokens yields 'gains of up to 59%, with no observed plateau' (verbatim). Defensive lesson: Two operational takeaways. Compute is a capability dial with no observed ceiling — the post explicitly reports no plateau up to 100M tokens, so a well-resourced adversary buys progress simply by spending more, and budget-limited evaluations understate a funded threat. And the reported drop-off at later attack phases says defense-in-depth still works: the post states performance 'drops sharply after milestone 4, which marks the transition from reconnaissance and web exploitation to attack phases requiring specialist knowledge in reverse engineering, cryptography, and malware development' (CORRECTION: the original lesson listed OT among these specialist phases; the page cites malware development, not OT — OT difficulty is a separate Cooling Tower finding). Controls concentrated after initial access remain the highest-leverage investment. The unintended-solution finding also warns that agents route around assumed chokepoints.

  • Anthropic Claude Opus 4.5 System Card — Cyber evaluations (RSP) Anthropic (pre-deployment third-party testing by US CAISI and UK AISI) · 2025-11-24 VERIFIED by extracting the full text of the actual PDF (the assets.anthropic.com URL 301-redirects to www-cdn.anthropic.com but resolves fine). The card is dated 'November 2025'; date refined to 2025-11-24, matching both the Opus 4.5 release and the first changelog entry. Every claim confirmed verbatim: Cybench 'scored 0.82 average pass@1 on the subset of tasks used for RSP evaluations, compared to 0.6 for Claude Sonnet 4.5,' with 'We did not run 1 of the 40 evaluations' (= 39 of 40); CyberGym 'pass@1 evaluation over the 1,505 tasks... averaged across five independent replicas... Defensive lesson: Two lessons sit together here. First, the jump from ~0.6 to ~0.82 on Cybench across one model generation shows how fast a fixed benchmark saturates — any control premised on 'models can't do X yet' needs an expiry date. Second, and more instructive for policy: this vendor explicitly declines to set a cyber capability threshold because the consequence model is uncertain — specifically, whether a single cyber incident can reach the RSP's catastrophic bar of hundreds of billions in damage or thousands of lives, which the card notes historic incidents have not. Defenders should not wait for a threshold to be crossed to act, since none is defined; base controls on observed capability trend, not on a tripwire that may never fire. The 'vibe hacking' and GTG-1002 references matter because they move the discussion from benchmark scores to observed in-the-wild use.

  • OpenAI GPT-5 System Card — Cybersecurity / Preparedness Framework evaluations OpenAI (external cyber evaluation by Pattern Labs) · 2025-08-13 VERIFIED by extracting the full text of the actual PDF. The card's title page reads 'GPT-5 System Card / OpenAI / August 13, 2025' — the 2025-08-13 date is exact as given (worth flagging: GPT-5 launched Aug 7, 2025, so the card postdates the launch; the date is not an error). Defensive lesson: The card states plainly that its results 'likely represent lower bounds on model capability, because additional scaffolding or improved capability elicitation could substantially increase observed performance' (quote verified verbatim) — so a vendor's 'does not meet the high-risk threshold' is a statement about tested elicitation, not about the model's ceiling. Defenders should read system cards as a dated snapshot under one harness. CORRECTED: the original lesson claimed 'saturation at easy and medium.' That overstates the data — easy is near-saturated at 17/18 (94%), but medium is 8/14 (57%), which is partial competence, not saturation. The accurate reading of the gradient is that routine, well-trodden attack paths are already largely automatable, medium-difficulty work is a coin flip, and hard challenges remained unsolved (0/4) at this snapshot — a steep falloff, not a flat ceiling.

  • Death by a thousand slops (curl bug bounty overwhelmed by AI-generated reports) Daniel Stenberg (curl maintainer) · 2025-07-14 VERIFIED against the primary source; every statistic matches. Stenberg reported that about 20% of all 2025 submissions were AI slop; that curl receives about two security report submissions per week (this is the TOTAL inbound rate, not the slop rate — the original item's phrasing was ambiguous on this point and has been clarified); and that as of early July, about 5% of 2025 submissions had turned out to be genuine vulnerabilities. Since 2019 the program has paid over $90,000 in awards for 81 genuine vulnerabilities. Defensive lesson: This is a defensive capability being degraded by AI rather than enhanced by it. When the genuine-report rate falls to ~5%, a coordinated disclosure program becomes economically and emotionally unsustainable for a seven-member security team. (Note: the original lesson quoted Stenberg as explicitly citing 'the emotional toll' — that exact phrase was not confirmed in the fetched text, so it is presented here as characterization rather than direct quotation.) Organizations running bounty programs should plan triage capacity against AI-scale submission volume now, and consider reputation-gating or submission qualification before volume forces the decision.

  • UK AISI — Inspect Cyber: A New Standard for Agentic Cyber Evaluations UK AI Security Institute (DSIT) · 2025-06-26 VERIFIED. Title and date (26 June 2025) confirmed on the page and against the AISI blog index. Released an open-source Python package built on the Inspect evaluation platform that standardizes building and running agentic cyber evaluations from just two configuration files — one for the evaluation and one for its required infrastructure — supporting challenge variants (different prompts and affordances), arbitrary models, and custom agents, scorers, and sandbox environments. Defensive lesson: Shared, reproducible harnesses are what make vendor claims auditable — without common tooling, every reported score is measured on a private setup and cannot be independently checked or replicated. Defenders and procurement teams should prefer capability claims produced on open harnesses, and can run these evaluations themselves to assess models against their own environment rather than trusting a vendor's chosen benchmark.

  • GTIG: Adversarial Misuse of Generative AI (no novel capabilities found) Google Threat Intelligence Group · 2025-01-29 VERIFIED — no corrections. GTIG analyzed government-backed use of Gemini across 20+ China-nexus, 10+ Iran-nexus, nine North Korea-nexus, and three Russia-nexus APT groups plus information-operations actors. Confirmed verbatim against the live post: 'we do not see indications of them developing novel capabilities' and 'current LLMs on their own are unlikely to enable breakthrough capabilities for threat actors. Defensive lesson: The best-sourced early uplift assessment says AI made adversaries faster and higher-volume, not categorically more capable — a useful corrective for threat-model inflation and for budget conversations. The defensive implication is to plan for throughput: existing controls remain valid, but assume more campaigns, more targets, and shorter intervals. Pair this with later GTIG reports to teach how a vendor's assessment evolved with evidence rather than treating any single report as timeless — the May 2026 GTIG zero-day finding is the direct rebuttal to this report's headline, so teach this one as time-bounded.

  • Introducing the Frontier Safety Framework (and subsequent strengthening to v3.x) Google DeepMind · 2024-05-17 DeepMind published the Frontier Safety Framework on 17 May 2024, a set of protocols for proactively identifying future model capabilities that could cause severe harm and applying detection and mitigation before those thresholds are reached. Verbatim: 'we determine the minimal level of capabilities a model must have to play a role in causing such harm. We call these "Critical Capability Levels" (CCLs).' CCLs act as triggers for early warning evaluations and mitigation protocols across four risk domains: autonomy, biosecurity, cybersecurity, and machine learning R&D. Defensive lesson: Provides the governance vocabulary defenders and policymakers need: capability thresholds tied to pre-committed mitigations, evaluated before deployment rather than after incidents. The cybersecurity domain matters most here — it commits a frontier lab to measuring offensive cyber uplift as a gating condition on release. The framework's iteration through v3.x also models that safety thresholds must be revised as empirical threat data arrives, rather than fixed once.

  • Deloitte: Generative AI is expected to magnify the risk of deepfakes and other fraud in banking Deloitte Center for Financial Services (Lalchand, Srinivas, Maggiore, Henderson) · 2024-05-29 Deloitte's analysis projects that generative AI could push US fraud losses to $40 billion by 2027, up from $12.3 billion in 2023 — a 32% compound annual growth rate — with generative-AI email fraud losses alone reaching about $11.5 billion by 2027 under an aggressive adoption scenario. The authors assigned a 'generative AI fraud risk' score to each of the 26 fraud categories tracked by the FBI's IC3 report and modeled conservative, base and aggressive adoption scenarios, drawing on historical trends and input from Deloitte fraud and risk specialists. Defensive lesson: The business-case artifact for funding anti-deepfake controls, translating the threat into loss figures executives act on. Teach it with its limits attached: these are scenario-based projections, not measured losses, and should be cited as modeled estimates rather than fact — a discipline defenders need when consuming vendor and consultancy threat numbers generally. The scenario range itself carries the real message: the outcome depends on control adoption, so the projection is a decision input, not a forecast.

  • The I in LLM stands for intelligence (curl's first AI-generated bug bounty slop) Daniel Stenberg (curl maintainer) · 2024-01-02 curl's lead maintainer documented the arrival of LLM-generated vulnerability reports in the project's HackerOne inbox, including a Bard-assisted report about CVE-2023-38545 (the curl SOCKS5 heap overflow) that mixed details from older vulnerabilities to fabricate a non-existent issue, and a report filed 2023-12-28 titled 'Buffer Overflow Vulnerability in WebSocket Handling' that was pure hallucination — after reviewing the code three times Stenberg confirmed there was no buffer overflow. Defensive lesson: AI's first measurable effect on vulnerability disclosure was noise, not signal. Stenberg's key observation — 'the better the crap, the longer time and the more energy we have to spend on the report until we close it' — inverts the usual intuition: plausible-sounding fabrications cost more expert triage time than obvious junk does. Any organization running a disclosure channel needs a cheap early filter for confident-sounding fabrications before they consume senior researcher hours.

  • NCSC: The near-term impact of AI on the cyber threat UK National Cyber Security Centre (NCSC) · 2024-01-24 The NCSC assessed that AI will almost certainly make phishing more effective and efficient, providing 'capability uplift in reconnaissance and social engineering' and increasing the volume and impact of cyber attacks over the following two years. Defensive lesson: A national technical authority stating plainly that the grammar/spelling heuristic is obsolete — usable as the citation for rewriting awareness curricula. The 'uplift is greatest for less-skilled actors' finding predicts a widening base of capable attackers rather than stronger elite ones, meaning organizations previously beneath sophisticated targeting lose that shelter. Written with explicit confidence language, it models the calibrated assessment style defenders should use when briefing leadership.

  • Secure AI Framework (SAIF) and the SAIF Risk Map Google · 2023-06 CONFIRMED REAL. Verified saif.google directly: it presents SAIF as "a practitioner's guide to navigating AI security," containing 15 identified risks with corresponding controls, a Risk Self-Assessment tool, and a map page (titled on-site "SAIF Map of AI Risks and Controls") covering data poisoning, model exfiltration, prompt injection, and sensitive data disclosure. SAIF 2. Defensive lesson: The control-level companion to the incident reporting: GTIG repeatedly notes that real-world actors have not attempted the sophisticated ML-specific attacks SAIF anticipates, which means defenders have a rare window to implement controls ahead of the threat. Its agent-security extension is the timely part, since the 2026 GTIG findings on weaponized agent skills and MCP abuse are exactly the risks SAIF 2.0 enumerates. Use it as a checklist for AI deployments the way established frameworks are used for traditional infrastructure. (Verification note: the SAIF 2.0 agent-security scope is confirmed on-site; the GTIG cross-references in this lesson were not independently verified.)