The Multiverse School

Join

Agentic AI Security — Foundational Agenda Papers (49)

Foundational Agenda Papers

The genre of large collaborative research-agenda papers — many authors, many pages, many enumerated open problems. 49 verified, every citation pinned to arXiv/publisher.

Why this genre matters for the course. These papers are how a field decides what it doesn't know. They're the best possible reading list: each one hands you a numbered set of open problems, already peer-reviewed for being the right problems.

⚠️ Attribution warning. This genre is chronically misattributed. Foundational Challenges is Anwar et al. (senior author David Krueger) — not Hendrycks. Hendrycks's paper is the distinct Unsolved Problems in ML Safety. The two share zero authors. Check attribution_note below before citing.

Companion to Level 3 — Research Frameworks (Bibliography) and Deep Dive — Anthropic · Apollo · METR.


Ancestors — the founding template

  • Harms from Increasingly Agentic Algorithmic Systems 2023 (arXiv v1, 2023-02-20) · arXiv; ACM FAccT 2023 (arXiv comment: "Accepted at FAccT 202 · arXiv:2302.10329 Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan et al. Scale: 22 authors (verified by full enumeration from the arXiv API). Frames agency along 4 enumerated characteristics: underspecification, directness of impact, goal-directedness, and long-term planning. The ancestor of this slice — the paper that made 'agentic' a spectrum rather than a binary, and connected it to concrete sociotechnical harm. Argues increasingly self-directed systems produce systemic and long-lasting harms falling disproportionately on marginalized communities, while insisting humans remain responsible for algorithmic harm.

    ⚠️ Attribution: Disambiguation (not a correction): lead author Alan Chan is distinct from Lawrence Chan, 2nd author of 'The Alignment Problem from a Deep Learning Perspective' (2209.00626). Citing this as 'Chan et al.' is ambiguous within this genre; the lead here is Alan Chan.

  • Unsolved Problems in ML Safety 2021 · arXiv · arXiv:2109.13916 Dan Hendrycks, Nicholas Carlini, John Schulman, Jacob Steinhardt (4 authors) — exact order confirmed via arXiv Atom API Scale: 4 authors; presents four problems ready for research, confirmed verbatim in the abstract: withstanding hazards ("Robustness"), identifying hazards ("Monitoring"), reducing inherent model hazards ("Alignment"), and reducing systemic hazards . Provides a new roadmap for ML Safety that refines the technical problems the field needs to address, organized into four research areas, clarifying each problem's motivation and giving concrete research directions. This is the Hendrycks paper frequently confused with Anwar et al.'s "Foundational Challenges" — they are distinct works with no shared authors.

    ⚠️ Attribution: This IS the Dan Hendrycks paper (Hendrycks, Carlini, Schulman, Steinhardt) that is commonly confused with Anwar et al.'s "Foundational Challenges in Assuring Alignment and Safety of Large Language Models" (arXiv 2404.09932). The two are distinct works sharing zero authors; citations crediting Hendrycks with the "Foundational Challenges" agenda are misattributions and belong to Anwar/Krueger.

  • Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims 2020 · arXiv (multi-institution report; submitted 15 April 2020; no · arXiv:2004.07213 Miles Brundage (lead), Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner et al. Scale: 59 authors (verified from arXiv metadata) drawn from across industry labs, academia and civil society — one of the largest author lists in the genre's pre-LLM era An ancestor of the modern frontier-governance literature: argues that responsible AI development requires developers to make verifiable claims that outsiders can scrutinise, and proposes concrete institutional, software and hardware mechanisms (audits, red-teaming, bias/safety bounties, incident sharing, compute accounting) to support them.

    ⚠️ Attribution: Anti-misattribution note (verified correct): this is a large multi-organisation coordination report, so the usual first-author/last-author-senior convention does NOT apply. Miles Brundage is the coordinating lead; Markus Anderljung's last-listed position does not denote senior authorship, and Yoshua Bengio's penultimate position does not denote seniority either. Cite as 'Brundage et al. (2020)'.

  • Risks from Learned Optimization in Advanced Machine Learning Systems 2019 · arXiv (confirmed arXiv-only; DBLP lists only CoRR abs/1906 · arXiv:1906.01820 Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, Scott Garrabrant (5 authors; Hubinger lead, Garrabrant senior) Scale: 5 authors (verified in order on the arXiv abs page; Hubinger first, Garrabrant last — both confirmed). Introduces the term 'mesa-optimization' for the case where a learned model is itself an optimizer, and names inner vs. outer alignment and deceptive alignment. Supplies much of the conceptual vocabulary that later LLM-era agendas reuse.

  • Scalable agent alignment via reward modeling: a research direction 2018 · arXiv (confirmed arXiv-only; DBLP lists only CoRR abs/1811 · arXiv:1811.07871 Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, Shane Legg (6 authors; Leike lead, Legg senior) Scale: 6 authors (verified in order on the arXiv abs page; Leike first, Legg last — both confirmed). A single-thesis research direction (recursive reward modeling) rather than a broad enumeration. States the agent alignment problem — how to create agents that behave per user intentions — and proposes recursive reward modeling as a scalable answer, with a discussion of the approach's challenges and promising directions. The direct intellectual ancestor of RLHF-based alignment at OpenAI/DeepMind.

  • AI Safety Gridworlds 2017 · arXiv (confirmed arXiv-only; DBLP lists only CoRR abs/1711 · arXiv:1711.09883 Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Scale: 8 authors (verified in order on the arXiv abs page; Leike first, Legg last — both confirmed). Turns the Concrete Problems agenda into runnable benchmarks: a suite of gridworld RL environments, each with a performance function hidden from the agent so that specification gaming can be measured rather than argued about. The bridge from enumerating problems to evaluating them.

  • Concrete Problems in AI Safety 2016 · arXiv · arXiv:1606.06565 Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman et al. Scale: 6 authors; 29 pages (arXiv comment field confirmed to read exactly "29 pages"); enumerates 5 practical research problems, grouped by whether the problem originates from a wrong objective function, an objective too expensive to evaluate, or . The founding document of the genre: reframes AI safety away from speculation and toward 'accidents' — unintended and harmful behavior emerging from poor design of real-world AI systems — and enumerates five concrete, tractable ML research problems (avoiding side effects, avoiding reward hacking, scalable supervision, safe exploration, distributional shift).


LLM-era agendas

  • Foundational Challenges in Assuring Alignment and Safety of Large Language Models 2024 · TMLR (Transactions on Machine Learning Research), 09/2024 · arXiv:2404.09932 Usman Anwar (lead), Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper et al. Scale: 42 authors — CONFIRMED (arXiv abs page and Atom API both enumerate exactly 42, Anwar #1 through Krueger #42; OpenAlex's "43 authorships" is an overcount artifact and was rejected). Enumerates 18 foundational challenges to assuring LLM alignment and safety, grouped into scientific understanding of LLMs, development-and-deployment methods, and sociotechnical challenges. Each challenge is decomposed into concrete research questions, making it a field-defining research agenda rather than a survey of results.

    ⚠️ Attribution: KEY CORRECTION: this paper is NOT by Dan Hendrycks. It is led by Usman Anwar with David Krueger as senior/last author. It is frequently confused with Hendrycks et al.'s "Unsolved Problems in ML Safety" (arXiv 2109.13916). Verified: the two papers share ZERO authors — 2109.13916 is Hendrycks, Carlini, Schulman, Steinhardt, none of whom appear among these 42.

  • TrustLLM: Trustworthiness in Large Language Models 2024 · ICML 2024 (Position paper track), PMLR v235:20166-20270 · arXiv:2401.05561 Yue Huang (lead), Lichao Sun, et al Scale: 70 authors (verified: full list counted on the arXiv abs page — arXiv's truncated view cuts off at 'Hongyi Wang', author #25, with '45 additional authors not shown'; 25+45=70, a common source of misattribution). A combined principles/benchmark megaproject on LLM trustworthiness: sets out eight dimensions, then empirically benchmarks sixteen models. Finds trustworthiness and utility positively correlated, and proprietary models generally ahead of open-source ones.

    ⚠️ Attribution: REAL MISATTRIBUTION, CAUSE IDENTIFIED. This paper is widely miscited as 'Sun et al. (2024)'. Verified cause: arXiv v1 (10 Jan 2024) listed Lichao Sun FIRST and Yue Huang second. From the current version onward — and in the ICML 2024 proceedings of record (PMLR v235) — the order is Yue Huang FIRST, Lichao Sun second. Correct citation is Huang et al. (2024); 'Sun et al.

  • Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback 2023 · TMLR (Transactions on Machine Learning Research), 2023 · arXiv:2307.15217 Stephen Casper and Xander Davies (co-first authors; the PDF footnote reads "Equal contribution"), Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer et al. Scale: 32 authors — CONFIRMED by two independent enumerations (arXiv Atom API and dblp both list exactly 32, Casper #1 to Hadfield-Menell #32). Substantial author overlap with the Anwar et al. Per the abstract, the paper (1) surveys open problems and fundamental limitations of RLHF and related methods; (2) overviews techniques to understand, improve, and complement RLHF in practice; and (3) proposes auditing and disclosure standards to improve societal oversight of RLHF systems. The body separates tractable problems (addressable within RLHF) from fundamental ones (requiring alternative approaches).

  • The Alignment Problem from a Deep Learning Perspective 2022 (arXiv v1, 2022-08-30); ICLR 2024 · arXiv; published in ICLR 2024 (per arXiv comment field: "Pub · arXiv:2209.00626 Richard Ngo, Lawrence Chan, Sören Mindermann (3 authors) Scale: 3 authors; a focused argumentative agenda rather than a many-author enumeration. Argues that AGI trained with today's methods could learn to act deceptively to gain reward, acquire misaligned internally-represented goals that generalize beyond the fine-tuning distribution, and pursue power. The canonical statement of the deep-learning-grounded case for misalignment risk.

    ⚠️ Attribution: Disambiguation (not a correction): the Lawrence Chan who is 2nd author here is a different researcher from the Alan Chan who leads 'Harms from Increasingly Agentic Algorithmic Systems' (2302.10329). A bare 'Chan et al.' is ambiguous across these two papers in this genre.

  • Ethical and social risks of harm from Language Models 2021 · arXiv · arXiv:2112.04359 Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh et al. Scale: 23 authors (full list verified on abs page; count confirmed as exactly 23). Abstract states verbatim: "We outline six specific risk areas... In total, we review 21 risks in-depth." Both figures confirmed against the v1 abstract. The foundational DeepMind taxonomy of language-model harms, structuring 21 risks across six areas (I. Discrimination, Exclusion and Toxicity; II. Information Hazards; III. Misinformation Harms; IV. Malicious Uses; V. Human-Computer Interaction Harms; VI. Automation, Access, and Environmental Harms), with mitigation directions and gaps.

    ⚠️ Attribution: Author list and order verified correct as submitted. One genuine confusion to flag: this 2021 arXiv report is routinely conflated with its peer-reviewed condensation, Weidinger et al., "Taxonomy of Risks posed by Language Models" (FAccT 2022) — same lead author and overlapping content, but a distinct paper with a shorter author list and a real conference venue.


"Open Problems" agendas

  • Open Problems in Mechanistic Interpretability 2025 · arXiv · arXiv:2501.16496 Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, et al Scale: 29 authors across multiple industry and academic labs; organized around 3 classes of open problem (conceptual/practical method improvements; how to apply methods to goals; socio-technical challenges). A forward-facing review that applies the open-problems format to a single subfield: what mechanistic interpretability still cannot do, and what the field should prioritize. Demonstrates the genre's maturation from whole-field agendas to per-subfield agendas.

    ⚠️ Attribution: All claims verified against arXiv; no correction needed. Author order confirmed: Sharkey #1, Chughtai #2, Batson #3, Lindsey #4, Wu #5; Nanda #14, Tegmark #21, Bau #23; Hoogland #27, Murfet #28, McGrath #29 (last).

  • Open Problems in Machine Unlearning for AI Safety 2025 · arXiv · arXiv:2501.04952 Fazl Barez, Tingchen Fu, Ameya Prabhu, Stephen Casper, Amartya Sanyal, Adel Bibi et al Scale: 19 authors Argues machine unlearning cannot serve as a comprehensive AI safety solution, focusing on dual-use knowledge in domains like cybersecurity and CBRN. Maps tensions between unlearning dangerous knowledge and preserving safety mechanisms, plus open challenges in evaluation, robustness, and feature retention.

  • Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents 2025 · arXiv · arXiv:2505.02077 Christian Schroeder de Witt (lead), Klaudia Krawiecka, Igor Krawczuk, Ben Hagag, William L Scale: 24 authors spanning academia, industry, and security practice (including OWASP-linked authors). Explicitly framed as an open-challenges agenda for a new subfield ('multi-agent security'). Argues that decentralised networks of interacting agents create attack surfaces with no single-agent analogue — secret collusion, steganographic coordination, emergent and cascading failures, adversarial swarms — and enumerates open problems toward a security discipline for agent ecosystems.

  • Why Do Multi-Agent LLM Systems Fail? 2025 · arXiv (v3) · arXiv:2503.13657 Mert Cemri, Melissa Z Scale: 13 authors (UC Berkeley Sky Computing-anchored, with Zaharia/Gonzalez/Stoica as senior authors in positions 11-13). Derives MAST (Multi-Agent System Failure Taxonomy): 14 distinct failure modes across 3 categories. The empirical complement to the multi-agent risk agendas: rather than reasoning forward from advanced capabilities, it analyses actual traces from existing multi-agent LLM frameworks and inductively derives a taxonomy of why they fail.

    ⚠️ Attribution: Author list and order verified correct as submitted. Commonly shorthand-cited as "the Berkeley/Zaharia paper" or credited to Stoica/Gonzalez because of the well-known senior authors trailing the byline; the lead author is Mert Cemri, with Melissa Z. Pan second. Cite as Cemri et al., not Zaharia et al.

  • Open Problems in Technical AI Governance 2024 · TMLR 2025 (arXiv v1 July 2024) · arXiv:2407.14981 Anka Reuel and Ben Bucknall (co-first authors, contributed equally) et al Scale: 33 authors (arXiv metadata, hand-enumerated). Journal-ref: Transactions on Machine Learning Research, 2025. arXiv comment explicitly records Bucknall and Reuel as sharing first authorship. Defines 'technical AI governance' — technical analysis and tools supporting effective AI governance — and presents a taxonomy of the field organised by capacity and target, identifying open problems where technical work could unblock governance action.

    ⚠️ Attribution: VERIFIED, NO CORRECTION NEEDED — but this is the entry's genuine attribution hazard, so flagging it. arXiv lists Anka Reuel first and Ben Bucknall second, so the conventional short-cite 'Reuel et al. (2024)' silently drops Bucknall's equal status.

  • Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems 2024 · arXiv · arXiv:2405.06624 David "davidad" Dalrymple (lead) and Joar Skalse et al Scale: 17 authors (arXiv metadata) Proposes 'guaranteed safe AI' — a family of approaches producing AI systems with high-assurance, ideally quantitative, safety guarantees — built around a world model, a safety specification, and a verifier. Outlines approaches to each core component and the main technical challenges.

    ⚠️ Attribution: VERIFIED CORRECT AS GIVEN. This paper is routinely miscredited to its famous middle authors. Confirmed order from arXiv metadata: Dalrymple (1), Skalse (2), Bengio (3), Russell (4), Tegmark (5), Seshia (6) ... Barrett (13), Ding Zhao (14), Tan Zhi-Xuan (15), Wing (16), Tenenbaum (17, last). It is a Dalrymple-led paper — NOT a Bengio, Russell, or Tegmark paper, despite being commonly cited as such.

  • Consciousness in Artificial Intelligence: Insights from the Science of Consciousness 2023 (arXiv v1, 2023-08-17) · arXiv · arXiv:2308.08708 Patrick Butlin & Robert Long (joint leads), Eric Elmoznino, Yoshua Bengio, Jonathan Birch, Axel Constant, George Deane, Stephen M Scale: 19 authors (verified by full enumeration from the arXiv API), spanning philosophy of mind, neuroscience and machine learning Assesses AI consciousness by rendering neuroscientific theories (global workspace, recurrent processing, higher-order theories, attention schema) into 'indicator properties' checkable against real architectures. Notable for placing Yoshua Bengio as 4th author alongside philosophers of mind.

    ⚠️ Attribution: Commonly misattributed: frequently credited to Yoshua Bengio, who is the most famous name on the roster but is only the 4th author. Patrick Butlin and Robert Long are the joint lead authors; correct short-cite is 'Butlin, Long et al. (2023)'.

  • Open Problems in Cooperative AI 2020 · arXiv · arXiv:2012.08630 Allan Dafoe, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R Scale: 8 authors; defines a research field spanning multi-agent systems, game theory, social choice, and human-machine alignment. Founds 'Cooperative AI' as a distinct research bet: rather than aligning a single agent, study how machine and human agents can jointly improve welfare. The abstract frames cooperation problems as ubiquitous across scales from daily routines to global challenges. Introduced the multi-agent axis that later agentic-AI agendas build on.


Mega-collaborations & consensus reports

  • International AI Safety Report 2025 · arXiv · arXiv:2501.17805 Chaired by Yoshua Bengio (first-listed); writing team incl Scale: 96 authors on the arXiv listing (verified by full hand-enumeration via the arXiv API): a writing team, an expert advisory panel, and national / international-organisation representatives The IPCC-style state-of-the-science synthesis on general-purpose AI risk, chaired by Bengio: capabilities, risks, and technical risk-management. Maintained as a rolling series — an interim report (arXiv:2412.05282, titled 'International Scientific Report on the Safety of Advanced AI (Interim Report)', 44 authors), two 2025 Key Updates (arXiv:2510.

    ⚠️ Attribution: CORRECTED THREE ERRORS. (1) Title: the canonical arXiv title is 'International AI Safety Report' with no '(2025)' suffix — the year is a disambiguator added by citers, not part of the title. (Only the 2026 edition carries a year in its actual title.) (2) The interim report arXiv:2412.05282 has 44 authors, not 25 (verified twice via independent fetches with identical enumeration).

  • Multi-Agent Risks from Advanced AI 2025 · arXiv · arXiv:2502.14143 Lewis Hammond (lead), Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, et al Scale: 44 authors (hand-counted from the full arXiv author list). The flagship agenda for the multi-agent slice: argues that risks from many interacting advanced AI agents are not reducible to single-agent safety, and builds a taxonomy of novel failure modes and underlying risk factors, with empirical and experimental examples. Closes with implications for safety, governance, and ethics.

    ⚠️ Attribution: Metadata verified: 44 authors, Hammond #1 (lead), Rahwan #44 (last), arXiv comment confirms 'Cooperative AI Foundation, Technical Report #1', published 2025-02-19.

  • Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety 2025 · arXiv (v1 Jul 2025; v2 Dec 2025) · arXiv:2507.11473 Tomek Korbak (lead), Mikita Balesni, Elizabeth Barnes, Yoshua Bengio, Joe Benton, Scale: 41 authors (count verified mechanically against the arXiv author list) spanning OpenAI, Google DeepMind, Anthropic, METR, Redwood Research, Apollo Research, UK AI Security Institute, Mila and academia — one of the widest cross-organisationa. A cross-lab position paper arguing that agents which reason in human language create a uniquely legible oversight surface, that this monitorability is imperfect and actively fragile, and that frontier developers should evaluate and preserve it as a first-class property when making training and architecture decisions.

    ⚠️ Attribution: VERIFIED CORRECT AS GIVEN. Confirmed against the arXiv author list: Ilya Sutskever is NOT an author — the item's disclaimer is accurate and worth keeping, as this paper is widely associated with him in secondary coverage. Lead author is Tomek Korbak; Bengio is #4, not a lead. Positions confirmed: Korbak (1), Balesni (2), Barnes (3), Bengio (4), Benton (5) ...

  • The Singapore Consensus on Global AI Safety Research Priorities 2025 · institutional report (arXiv, submitted 25 June 2025; final r · arXiv:2506.20702 Yoshua Bengio (lead), Tegan Maharaj, Luke Ong, Stuart Russell, Dawn Song, Max Tegmark, Lan Xue, Ya-Qin Zhang, Stephen Casper, Wan Sie Lee et al. Scale: 88 authors (arXiv metadata; abstract states Bengio plus 87 other authors). arXiv comment: final report from the '2025 Singapore Conference on AI (SCAI)' held April 26. Synthesises global AI safety research priorities from an international scientific exchange, organising the field via a defence-in-depth model into three domains: creating trustworthy AI systems (Development), evaluating them (Assessment), and controlling them post-deployment (Control). Builds on the International AI Safety Report.

  • The Ethics of Advanced AI Assistants 2024 · arXiv (Google DeepMind institutional report) · arXiv:2404.16244 Iason Gabriel (lead), Arianna Manzini, Geoff Keeling, Lisa Anne Hendricks, Verena Rieser, Hasan Iqbal, Nenad Tomašev, Ira Ktena, Zachary Kenton et al. Scale: 57 authors — now CONFIRMED by direct enumeration of the arXiv Atom API author list (Gabriel #1, Isaac #56, Manyika #57), independent of the OpenAlex record that originally supplied the figure; the earlier hedge is resolved. A very large DeepMind collaboration examining the ethical and societal questions raised by advanced, agentic AI assistants — value alignment, wellbeing, safety, manipulation, anthropomorphism, trust, and equity of access. Agenda-setting across a wide sociotechnical surface rather than reporting experimental results.

  • Personhood Credentials: Artificial Intelligence and the Value of Privacy-Preserving Tools to Distinguish Who is Real Online 2024 · arXiv · arXiv:2408.07892 Steven Adler, Zoë Hitzig, Shrey Jain (co-leads), Catherine Brewer, Wayne Chang, Renée DiResta, Eddy Lazzarin, Sean McGregor, Wendy Seltzer et al. Scale: 32 authors spanning OpenAI, Microsoft, Harvard, MIT, Berkeley, a16z, W3C, Stanford Internet Observatory and civil-society orgs — a genuine multi-institution consortium. Responds to agents becoming indistinguishable from humans online: proposes privacy-preserving 'personhood credentials' that let a user prove they are a real person without revealing identity, and enumerates design goals, deployment risks (equity, free expression, credential markets), and open problems.

    ⚠️ Attribution: Confirmed: arXiv 2408.07892 lists exactly 32 authors, matching the stated count.

  • Managing extreme AI risks amid rapid progress 2023 (arXiv v1, 2023-10-26); 2024 (Science publication; arXiv updated 2024-05-22) · Science (doi:10 · arXiv:2310.17688 Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz et al. Scale: 25 authors (NOW VERIFIED via the arXiv API — the source item left this count unasserted). A high-profile multi-author consensus statement arguing that R&D and governance are not keeping pace with AI capability gains, and setting out priorities for technical research and adaptive governance institutions to manage extreme risks.

  • On the Opportunities and Risks of Foundation Models 2021 (arXiv v1, 2021-08-16) · arXiv · arXiv:2108.07258 Rishi Bommasani, Drew A Scale: 114 authors (CORRECTED from the item's claimed 139). Counted mechanically from the complete arXiv API author enumeration. The paper that coined and institutionalized the term 'foundation model'. A sprawling multi-section report covering capabilities, underlying technology, applications, and societal impact (inequity, misuse, economic and environmental impact, legal and ethical considerations). The template for the modern mega-report format.

    ⚠️ Attribution: Commonly misattributed: this report is often credited to Percy Liang as lead author (he directs Stanford CRFM and is the most senior name). In fact Rishi Bommasani is the lead/first author and Percy Liang is last-listed/senior. Correct short-cite is 'Bommasani et al. (2021)'.

  • Advances and Open Problems in Federated Learning 2019 · Foundations and Trends in Machine Learning, Vol · arXiv:1912.04977 Peter Kairouz and H Scale: 59 authors Included as a structural cousin, not an AI-safety paper: this is the archetype of the 'Advances and Open Problems in X' form that the safety genre inherited — a 59-author, book-length, cross-institutional catalog of open problems in federated learning (privacy, efficiency, robustness). Useful as the template against which the safety-specific entries above can be read.


Governance & evaluation agendas

  • Authenticated Delegation and Authorized AI Agents 2025 · arXiv (DBLP lists a CoRR 2025 entry only · arXiv:2501.09674 Tobin South (lead), Samuele Marro, Thomas Hardjono, Robert Mahari, Cedric Deslandes Whitney, Dazza Greenwood et al. Scale: 8 authors (MIT Media Lab / MIT Connection Science-anchored). Smaller than the mega-collaborations here; it is the protocol-design anchor of the agent-identity cluster rather than a broad author consortium. NOTE ON TITLE: the canonical arXiv title is 'Authenticated Delegation and Authorized AI Agents' — 'AI Agents Need Authenticated Delegation' is a paraphrase/variant. Proposes a framework for humans to delegate scoped, auditable authority to agents by extending OAuth 2.0 and OpenID Connect with agent-specific credentials, plus translating natural-language permissions into enforceable access controls.

  • Governing AI Agents 2025 · arXiv; Notre Dame Law Review, Vol · arXiv:2501.07913 Noam Kolt (SOLE AUTHOR — this is a single-authored law review article, not a collaboration) Scale: Scale rests on length and scope, NOT author count: a solo-authored full-length law review article (Notre Dame L. Rev. Vol. 101). Applies agency law and economics to AI agents, identifying information asymmetry, discretionary authority, and loyalty as the core problems, and argues that classical solutions (incentive design, monitoring, enforcement) break down at machine speed and scale — motivating new technical and legal infrastructure built on inclusivity, visibility, and liability.

    ⚠️ Attribution: VERIFIED SOLE AUTHOR. arXiv 2501.07913 lists Noam Kolt as the only author — the 'not a collaboration' framing is accurate. Submitted 14 Jan 2025 (v1), revised 11 Feb 2025 (v2). The arXiv journal-ref field independently confirms 'Notre Dame Law Review, Vol. 101, Forthcoming'. Title, ID, year, and venue all resolve correctly; no correction needed.

  • Infrastructure for AI Agents 2025 · arXiv; accepted to TMLR · arXiv:2501.10114 Alan Chan (lead), Kevin Wei, Sihao Huang, Nitarshan Rajkumar, Elija Perrier, Seth Lazar, Gillian K Scale: 8 authors. Enumerates a catalogue of 'agent infrastructure' — technical systems and shared protocols external to agents — organised by the functions of attributing actions, shaping interactions, and detecting/remedying harms. Argues the missing layer for agent safety is not model-internal but infrastructural (analogous to HTTPS or public-key infrastructure for the web), and enumerates concrete infrastructure primitives plus open research questions for each.

    ⚠️ Attribution: Author list and byline order confirmed verbatim against arXiv 2501.10114: Chan, Wei, Huang, Rajkumar, Perrier, Lazar, Hadfield, Anderljung — 8 authors exactly as stated. Alan Chan is first author and Markus Anderljung is last/senior, so both role attributions are correct. Submitted 17 Jan 2025, last revised 19 Jun 2025 (v3); 'Accepted to TMLR' confirmed in the arXiv comments field.

  • Visibility into AI Agents 2024 · arXiv; ACM FAccT 2024 · arXiv:2401.13138 Alan Chan (lead), Carson Ezell, Max Kaufmann, Kevin Wei, Lewis Hammond, Herbie Bradley, Emma Bluemke, Nitarshan Rajkumar, David Krueger, Noam Kolt et al. Scale: 12 authors. Enumerates 3 visibility measures — agent identifiers, real-time monitoring, and activity logging — each analysed across centralised and decentralised deployment and across supply-chain actors. Sets the agenda for governance-relevant transparency into deployed agents: what information about agents would need to exist, who would hold it, and how each measure interacts with privacy and concentration of power.

    ⚠️ Attribution: No known attribution confusion — left empty deliberately rather than manufacturing one. Everything verified exactly as claimed: all 12 authors in the stated order, Chan #1 (lead) through Anderljung #12 (last/senior); arXiv comment confirms acceptance to the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2024); published 2024-01-23, so year 2024 is correct.

  • IDs for AI Systems 2024 · arXiv; RegML workshop @ NeurIPS 2024 · arXiv:2406.12137 Alan Chan (lead), Noam Kolt, Peter Wills, Usman Anwar, Christian Schroeder de Witt, Nitarshan Rajkumar, Lewis Hammond, David Krueger et al. Scale: 10 authors. Note: lead author Alan Chan and coauthor Usman Anwar overlap with the broader agenda literature — Anwar led 'Foundational Challenges in Assuring Alignment and Safety of LLMs', which is NOT a Hendrycks paper. Proposes IDs that travel with agent activity so that affected parties can learn about the system acting on them, sketching what an ID should contain, deployment mechanisms across the supply chain, and the privacy and concentration-of-power trade-offs.

    ⚠️ Attribution: Author list and order confirmed verbatim against arXiv 2406.12137 (10 authors; Chan first, Anderljung last/senior). Submitted 17 Jun 2024, revised 28 Oct 2024 (v3); 'Accepted to RegML workshop @ NeurIPS 2024' confirmed in the arXiv comments.

  • Black-Box Access is Insufficient for Rigorous AI Audits 2024 · FAccT 2024 · arXiv:2401.14446 Stephen Casper (lead), Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer et al. Scale: 21 authors (verified from arXiv metadata). Published at ACM FAccT '24, Rio de Janeiro, June 3-6 2024 (journal-ref on arXiv; DOI 10.1145/3630106.3659037). Argues that audits conducted with only black-box (query) access cannot support rigorous conclusions, and compares black-box, white-box and 'outside-the-box' (training data, methodology) access levels — showing deeper access enables substantially stronger audits.

  • Evaluating Frontier Models for Dangerous Capabilities 2024 · arXiv (no journal-ref or venue comment on arXiv as of verifi · arXiv:2403.13793 Mary Phuong (lead), Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael et al. Scale: 27 authors (verified from arXiv metadata), Google DeepMind Reports a pilot programme of dangerous capability evaluations, applying a concrete battery of evaluations across several risk-relevant capability areas to frontier models. Turns the agenda of 'Model evaluation for extreme risks' into an executed empirical programme.

    ⚠️ Attribution: Routinely conflated at the paper level with Toby Shevlane's earlier 'Model evaluation for extreme risks' (arXiv 2305.15324). Shevlane is the senior (last) author here but the lead author of that earlier agenda paper; this one is correctly cited as Phuong et al. The two are distinct papers, and this is the empirical follow-through on that agenda, not a reissue of it.

  • Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? 2024 · NeurIPS 2024 · arXiv:2407.21792 Richard Ren (lead), Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan H Scale: 12 authors (verified from arXiv metadata). arXiv comment: NeurIPS 2024. Empirically asks whether AI safety benchmarks actually measure distinct safety progress, by measuring how strongly each benchmark correlates with general capabilities. Benchmarks that simply track upstream capability are argued to enable 'safetywashing' — capability improvements misrepresented as safety advances.

    ⚠️ Attribution: The canonical misattribution case in this genre: because of its strong association with Dan Hendrycks and CAIS, the paper is frequently credited as 'Hendrycks et al.' Hendrycks is the senior (last) author, NOT the lead. The correct short cite is 'Ren et al. (2024)' — Richard Ren is first author. Verified against arXiv metadata.

  • On the Societal Impact of Open Foundation Models 2024 · ICML 2024 (Position Papers track); Proceedings of the 41st I · arXiv:2403.07918 Sayash Kapoor, Rishi Bommasani, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Peter Cihon, Aspen Hopkins, Kevin Bankston, Stella Biderman et al. Scale: 25 authors, confirmed independently by the arXiv API Atom feed and the PMLR v235 paper page (identical order in both). A position paper on open-weight foundation models (e.g. Llama 2, Stable Diffusion XL). It identifies five distinctive properties of open foundation models (e.g.

    ⚠️ Attribution: Attribution as given is CORRECT — no correction needed. Verified against the arXiv API and the PMLR v235 proceedings page, which list an identical 25-author roster in the same order: Sayash Kapoor is lead author, Percy Liang is penultimate, and Arvind Narayanan is final/senior author; Daniel E. Ho is 23rd.

  • Model evaluation for extreme risks 2023 (arXiv v1, 2023-05-24) · arXiv · arXiv:2305.15324 Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung et al. Scale: 21 authors (verified by full enumeration from the arXiv API), spanning DeepMind, OpenAI, Anthropic, GovAI and academic groups Argues model evaluation is critical for addressing extreme risks, and that developers must be able to identify dangerous capabilities (via 'dangerous capability evaluations') and the propensity of models to apply capabilities for harm (via 'alignment evaluations'). Positions evaluation as the input to responsible training, deployment and policy decisions.

  • Frontier AI Regulation: Managing Emerging Risks to Public Safety 2023 · arXiv (confirmed arXiv-only; DBLP lists a single CoRR abs/23 · arXiv:2307.03718 Markus Anderljung (lead) et al Scale: 24 authors (verified by full count on the arXiv abs page: Anderljung, Barnhart, Korinek, Leung, O'Keefe, Whittlestone, Avin, Brundage, Bullock, Cass-Beggs, Chang, Collins, Fist, Hadfield, Hayes, Ho, Hooker, Horvitz, Kolt, Schuett, Shavit, S. Sets out the regulatory challenge posed by 'frontier AI' models whose dangerous capabilities may arrive unexpectedly, and proposes building blocks for regulating them — including safety standards, registration and reporting, and mechanisms to ensure compliance.

    ⚠️ Attribution: 'Anderljung et al. (2023)' is correct and undisputed — no misattribution found. The item's seniority caveat is VERIFIED against the arXiv v2 comment (11 Jul 2023): 'Adjusted author order (mistakenly non-alphabetical among the first 6 authors).

  • Sociotechnical Safety Evaluation of Generative AI Systems 2023 · arXiv (confirmed arXiv-only; DBLP lists only CoRR abs/2310 · arXiv:2310.11986 Laura Weidinger (lead), Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay et al. Scale: 13 authors (verified on the arXiv abs page); Weidinger lead, William Isaac senior — both confirmed. Proposes a three-layer evaluation framework and surveys the existing state of safety evaluation to identify gaps. Argues safety evaluation must be sociotechnical, spanning three layers: capability evaluations, human-interaction evaluations, and systemic-impact evaluations. Surveys current evaluations, finds coverage concentrated at the capability layer, and maps the resulting gaps.


Surveys, taxonomies & risk repositories

  • The AI Agent Index 2025 · arXiv (accompanying website: aiagentindex · arXiv:2502.01635 Stephen Casper (lead), Luke Bailey, Rosco Hunter, Carson Ezell, Emma Cabalé, Michael Gerovitch, Stewart Slocum, Kevin Wei, Nikola Jurkovic et al. Scale: 15 authors (MIT-anchored). A documentation effort rather than an enumerated-problems agenda — it is the empirical base layer the agenda papers lean on. NOT an open-problems paper: it is a public index documenting deployed agentic AI systems — their components, application domains, and (sparse) risk-management practices. Its finding is itself the argument: developers publish plenty about capabilities and very little about safety.

    ⚠️ Attribution: AUTHOR NAME CORRECTED. The third author is ROSCO HUNTER, not 'Robert Hunter' as supplied — an incorrect first name in the input. Verified against two independent primary sources: the arXiv abstract page for 2502.01635 and the arXiv API export record, both of which give 'Rosco Hunter'. This is a characteristic corruption of an uncommon first name into a familiar one.

  • The AI Risk Repository: A Meta-Review, Database, and Taxonomy of Risks From Artificial Intelligence 2024 · arXiv (submitted 14 Aug 2024; v3 revised 5 May 2026) · arXiv:2408.12622 Peter Slattery, Alexander K Scale: 10 authors. Meta-review of 74 existing AI risk frameworks, from which 1,725 distinct risks were extracted into a living public database. Builds a living, publicly-updatable database and two-part taxonomy (causal: entity/intent/timing; and domain-level) unifying 74 prior risk frameworks into one common reference. Explicitly a repository rather than a one-shot survey.

    ⚠️ Attribution: Author list and order verified correct as submitted. arXiv currently renders the title in sentence case ("The AI risk repository: A meta-review..."); the title-case form given here is the standard citation form and is not an error.

  • AI Alignment: A Comprehensive Survey 2023 · arXiv · arXiv:2310.19852 Jiaming Ji (lead), Tianyi Qiu, Boyuan Chen, Scale: 26 authors (full ordered list verified via arXiv API). Organizes the field around 4 core principles (RICE) and splits research into forward vs. backward alignment. No page count in the arXiv comments field, so none claimed. Surveys the alignment field around four objectives — Robustness, Interpretability, Controllability, Ethicality (RICE) — dividing work into forward alignment (training: feedback learning, distribution shift) and backward alignment (assurance, verification, governance). Maintained alongside a companion tutorial website.

    ⚠️ Attribution: No known attribution confusion — left empty deliberately rather than manufacturing one. Verified: 26 authors, Ji #1 (lead), Qiu #2, Chen #3; trailing senior block Yang #22, Wang #23, Zhu #24, Guo #25, Gao #26 (last-listed) — all exactly as claimed.

  • The Rise and Potential of Large Language Model Based Agents: A Survey 2023 · arXiv (Sept 2023); published in Science China Information Sc · arXiv:2309.07864 Zhiheng Xi (lead), Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, et al Scale: 29 authors on the arXiv version (28 on the journal version). Book-length survey (86 pages, 12 figures) with an accompanying continuously-updated paper repository; the canonical map of the LLM-agent field. The capabilities-side counterpart to the risk agendas: surveys LLM-based agents via a brain/perception/action framework, covering single-agent applications, multi-agent societies, human-agent interaction, and closing with open problems including agent safety, evaluation, and emergent social phenomena.

    ⚠️ Attribution: Confirmed and both author counts verified. arXiv 2309.07864 (submitted 14 Sep 2023, v3 19 Sep 2023) lists exactly 29 authors, Zhiheng Xi first — matching the claim. The journal version in Science China Information Sciences carries 28: 'Wensen Cheng' is dropped, and Qi Zhang moves from a mid-list position to the senior/corresponding slot alongside Tao Gui.

  • An Overview of Catastrophic AI Risks 2023 · arXiv · arXiv:2306.12001 Dan Hendrycks, Mantas Mazeika, Thomas Woodside (Center for AI Safety) Scale: Only 3 authors — its genre credential is breadth/length, not headcount. Enumerates 4 top-level risk categories, each broken into specific hazards with illustrative scenarios and mitigations. Organizes catastrophic AI risk into four sources: malicious use, AI race, organizational risks, and rogue AIs. For each it describes hazards, tells illustrative stories, sketches an ideal scenario, and proposes mitigations.

    ⚠️ Attribution: Attribution here is CORRECT as submitted — this paper genuinely is Hendrycks, Mazeika & Woodside. The well-known Hendrycks/Anwar confusion runs in the opposite direction and does not apply to this entry: "Foundational Challenges in Assuring Alignment and Safety of Large Language Models" (arXiv:2404.

  • TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI 2023 · arXiv · arXiv:2306.06924 Andrew Critch, Stuart Russell (UC Berkeley / CHAI) Scale: Only 2 authors — included for taxonomy breadth, not headcount. Organizes societal-scale risks by accountability structure (actor identification, unity, and intent — i.e. A systematic, exhaustively-intended taxonomy of societal-scale and extinction-level AI risks, categorized by how diffuse the responsibility is — spanning diffusion of responsibility, bigger-than-expected AI impacts, and unrecognized/unsafe multi-agent dynamics. Pairs illustrative scenarios with combined technical and policy responses.

    ⚠️ Attribution: Attribution verified correct as submitted: TASRA is Critch & Russell. It is frequently miscited as "Critch & Krueger" by conflation with Critch's other major risk taxonomy, ARCHES — "AI Research Considerations for Human Existential Safety" (arXiv:2006.04948) — which IS Andrew Critch & David Krueger.

  • A Survey of Safety and Trustworthiness of Large Language Models through the Lens of Verification and Validation 2023 · arXiv (v1 May 2023; v2 Aug 2023); published in Artificial In · arXiv:2305.11391 Xiaowei Huang (lead), Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, Kaiwen Cai et al. Scale: 17 authors (verified from arXiv metadata). Scale is breadth of coverage: organizes LLM safety across the full lifecycle via falsification/evaluation, verification, runtime monitoring, and ethical/regulatory governance. Surveys LLM safety and trustworthiness through a verification-and-validation lens borrowed from safety-critical engineering, categorizing inherent issues, attacks, and unintended bugs, and mapping V&V techniques onto each lifecycle stage.

  • Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction 2022 · arXiv (v1 11 Oct 2022; v3 19 Jul 2023; cs · arXiv:2210.05791 Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N'Mah Yilla-Akbari, Jess Gallegos, Andrew Smart et al. Scale: 11 authors. Synthesizes a large body of computing scholarship into a shared taxonomy of sociotechnical harms from algorithmic systems, intended as common vocabulary for harm identification and reduction during design and evaluation.

    ⚠️ Attribution: Author list, order, lead (Renee Shelby) and last author (Gurleen Virk) all verified correct as submitted. One name-rendering discrepancy worth noting: arXiv lists the seventh author as "N'Mah Yilla" while Crossref/ACM list her as "N'Mah Yilla-Akbari" — same person, and the ACM/published form is the fuller one. No misattribution.