From Toronto to Seoul: SRI researchers at ICML 2026
SRI researchers are heading to Seoul this July for ICML 2026, and they're arriving with a lot to share. Across 24 papers, seven position papers, and a co-organized workshop, 21 members of our community are presenting work spanning machine learning theory, AI safety, clinical AI, legal compliance, and more. (Image credit: Milad Fakurian/Unsplash.)
The International Conference on Machine Learning (ICML) draws researchers from around the world, and at this year’s gathering in Seoul, South Korea, researchers affiliated with the Schwartz Reisman Institute for Technology and Society (SRI) will be well represented. Now in its 43rd year, ICML is one of the most selective and influential venues in machine learning, drawing thousands of researchers from academia and industry worldwide. Across 24 accepted papers, seven position papers, and a co-organized workshop, 21 members of the SRI community are presenting work that reflects the range, rigor, and social relevance of research happening across the Institute.
Research spanning machine learning, safety, and society
Several papers examine how increasingly capable AI systems behave in real-world settings, and how they can be evaluated, monitored, and made safer. David Duvenaud contributes work on disempowerment patterns in real-world LLM usage, examining how users can lose control or agency when interacting with large language models. Zhijing Jin contributes multiple papers on LLM behaviour and evaluation, including work on training models with “honeypots” to reshape failure modes, benchmarking cooperation-sustaining mechanisms in social dilemmas, and evaluating causal inference for scientific research. Xujie Si contributes work on auditing malicious scam endpoints in production LLMs and on evaluating conversational agents in dual-control environments, including an oral presentation on τ²-Bench.
This theme extends into social and multi-agent settings as well. Gillian Hadfield contributes work on gossip-driven indirect reciprocity among self-interested LLM agents, studying how decentralized reputation systems can sustain cooperation without centralized enforcement—a question with direct relevance to how autonomous AI agents might coordinate, or fail to, at scale.
SRI researchers are also advancing work on machine learning foundations, optimization, and evaluation. Igor Gilitschenski contributes research on test-time graph search for goal-conditioned reinforcement learning, exploring new approaches for decision-making in complex environments. Chris Maddison presents work on predicting large model test losses through a noisy quadratic system, contributing to the broader challenge of understanding and forecasting model scaling behaviour. Vardan Papyan contributes work on gradient smoothing, proposing ways to couple layer-wise updates for improved optimization. Tim G. J. Rudner presents research on semantic isotropy as a signal of nonfactuality in long-form text generation, offering a new way to understand when generated text may be unreliable. Rudner also presents two additional papers this year, on temporal straightening for latent planning in world models and on sparse, maximum-entropy representations in joint-embedding predictive architectures—extending his contributions beyond long-form text reliability into representation learning more broadly.
Other contributions focus on privacy, verification, and the governance of AI systems. Nicolas Papernot presents work on efficient public verification of private machine learning through regularization, as well as research on the principled design of privacy canaries. Noam Kolt contributes two papers on AI compliance in legal contexts, including work on contextual illegality in corporate law scenarios and copyright compliance in agentic AI systems. Together, these projects point to the growing importance of technical tools that can evaluate whether AI systems comply with privacy expectations, legal obligations, and institutional rules.
Several SRI contributions address how AI systems should be evaluated before deployment, especially in high-stakes domains. Marzyeh Ghassemi contributes papers on multi-answer reinforcement learning in language models and automated diagnosis of LLM failures, as well as a position paper arguing that benchmarks alone do not measure deployment readiness in clinical AI. This work highlights a central challenge for the field: systems that perform well on benchmarks may still fail in ways that matter for patients, institutions, and real-world decision-making.
This year's delegation also includes a strong contingent working at the intersection of machine learning and genomics. Bo Wang contributes three papers: dnaHNet, a scalable hierarchical foundation model for genomic sequence learning (Spotlight); research showing that predicting evolutionary rate as a pretraining task improves genome language model representations; and work addressing signal dilution in modeling and evaluating single-cell perturbation response data. Together, these projects reflect a growing effort to bring foundation-model methodology to bear on long-standing problems in molecular biology.
The position papers presented by SRI researchers expand these themes across alignment, scientific research, trustworthy AI, sociopolitical risk, and formal reasoning. Ashton Anderson and Karina Vold co-author a new work arguing LLMs should be designed and evaluated against human flourishing, not just task performance. Zhijing Jin contributes work arguing that LLMs for physics research require domain-specialized training and tooling; that trustworthy AI is shaped by conflicts between different forms of invariance; and that safe models do not, by themselves, guarantee safe societies. Xujie Si contributes a position paper on theory-level autoformalization, making the case for more rigorous connections between formal methods and machine learning. Mamatha Bhat contributes a position paper arguing that multi-agentic AI systems in transplant medicine risk amplifying disparities without targeted explainability and deployment strategies, extending the piece's throughline on deployment readiness in clinical AI into a specific, high-stakes clinical domain.
David Lie, Zhijing Jin, Kexin Li and Wenjun Qiu are among the co-organizers of the Trustworthy AI for Good Workshop, which bridges technical advances in trustworthy AI with real-world societal impact—from protecting democratic institutions and civic discourse to accountable use of AI in government settings. The workshop features five keynote speakers: Yoshua Bengio (Université de Montréal, Mila, and LawZero), Joel Leibo (Google DeepMind), Oana Ignat (Santa Clara University), Maksym Andriushchenko (ELLIS Tübingen and MPI-IS), and Naman Goyal and Jenny Ni (Google DeepMind and Google). A panel discussion will be moderated by Milind Tambe (Harvard University), with panelists including Lewis Hammond (Cooperative AI Foundation and University of Oxford), Gopal Sarma (RAND Corporation), and SRI’s own David Duvenaud.
The Institute congratulates all SRI researchers and collaborators whose work will be presented at ICML 2026.
Accepted papers
Mrinank Sharma, Miles McCain, Raymond Douglas and David Duvenaud, “Who’s in Charge? Disempowerment Patterns in Real-World LLM Usage”
Isha Puri, Mehul Damani, Idan Shenfeld, Marzyeh Ghassemi, Jacob Andreas and Yoon Kim, “Escaping the Mode: Multi-Answer Reinforcement Learning in LMs”
Yue Huang, Zhengzhe Jiang, Yuchen Ma, Yu Jiang, Xiangqi Wang, Yujun Zhou, Yuexing Hao, Kehan Guo, Pin-Yu Chen, Marzyeh Ghassemi, Stefan Feuerriegel and Xiangliang Zhang, “ProbeLLM: Automating Principled Diagnosis of LLM Failures”
Evgenii Opryshko, Junwei Quan, Claas Voelcker, Yilun Du and Igor Gilitschenski, “Test-Time Graph Search for Goal-Conditioned Reinforcement Learning”
Shuhui Zhu, Yue Lin, Shriya Kaistha, Wenhao Li, Baoxiang Wang, Hongyuan Zha, Gillian Hadfield, Pascal Poupart, “Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents”
Samuel Simko, Punya Syon Pandey, Zhijing Jin and Bernhard Schölkopf, “Training with Honeypots: Reshaping How LLMs Fail”
Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer and Zhijing Jin, “CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas”
Sawal Acharya, Terry Jingchen Zhang, Andrew Kim, Anahita Haghighat, Xianlin Sun, Rahul Babu Shrestha, Maximilian Mordig, Furkan Danisman, Clijo Jose, Yahang Qi, Pepijn Cobben, Bernhard Schölkopf, Mrinmaya Sachan and Zhijing Jin, “CauSciBench: Evaluating LLM Causal Inference for Scientific Research”
Hilal Aka, Joe Kwon and Noam Kolt, “Evaluating Contextual Illegality: AI Compliance in Corporate Law Scenarios”
Zheng Hui, Doni Bloomfield and Noam Kolt, “Copyright-Bench: Agentic Evaluation of Copyright Law Compliance” — Spotlight
Valentyn Melnychuk, Vahid Balazadeh, Stefan Feuerriegel, Rahul G. Krishnan, “Frequentist Consistency of Prior-Data Fitted Networks for Causal Estimation”
Adnan Mohammed, Rohan Jain, Tom Jacobs, Ekansh Sharma, Rahul G. Krishnan, Rebekka Burkholz, Yani Ioannou, “SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training”
Chuning Li and Chris J. Maddison, “Predicting Large Model Test Losses with a Noisy Quadratic System”
Zoë Ruha Bell, Anvith Thudi, Olive Franzese-McLaughlin, Nicolas Papernot and Shafi Goldwasser, “Efficient Public Verification of Private ML via Regularization”
Mohammad Yaghini, Michael Aerni, Junrui Zhang, Nicolas Papernot and Florian Tramèr, “OptiFluence: Principled Design of Privacy Canaries” — Spotlight
Haoming Meng, Anton Sugolov and Vardan Papyan, “Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization”
Dhrupad Bhardwaj, Julia Kempe and Tim G. J. Rudner, “Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation”
Ying Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero, Tim G. J. Rudner, Yann LeCun, Mengye Ren, “Temporal Straightening for Latent Planning”
Yilun Kuang, Yash Dagade, Tim G. J. Rudner, Randall Balestriero, Yann LeCun, “Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations”
Zhiyang Chen, Tara Saba, Xun Deng, Xujie Si and Fan Long, “Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs”
Victor Barres, Honghua Dong, Soham Ray, Xujie Si and Karthik Narasimhan, “τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment” — Oral
Mica Consens, Kevin Yang, James Hall, Ashley Conard, Bo Wang, Lorin Crawford, Alan Moses, Alex Lu, “Predicting evolutionary rate as a pretraining task improves genome language model representations”
Arnav Shah, Junzhe Li, Parsa Idehpour, Adibvafa Fallahpour, Brandon Wang, Sukjun Hwang, Bo Wang, Patrick Hsu, Hani Goodarzi, Albert Gu “dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning” – Spotlight
Gabriel Mejia, Henry Miller, Francis Leblanc, Bo Wang, Brendan Swain, Lucas Paulo de Lima Camillo, “Needles in the Haystack: Addressing Signal Dilution Improves scRNA-seq Perturbation Response Modeling and Evaluation”
Position papers
Ashton Anderson, Harsh Kumar, Louis Tay, and Karina Vold, “Position: We Need Large Language Models Optimized For Our Well-Being”
Divya Sharma, Ghazal Azarfar, Bima Hasjim and Mamatha Bhat, “Position: When AI Decides Who Gets an Organ: Multi-Agentic AI Systems in Transplant Medicine Risk Amplifying Disparities Without Targeted Explainability and Deployment Strategies”
Haoran Zhang, Hyewon Jeong, Olawale Salaudeen, Walter Gerych, Nigam Shah and Marzyeh Ghassemi, “Position: Benchmarks Do Not Measure Deployment Readiness in Clinical AI”
David Guzman Piedrahita, Dave Banerjee, Changling Li, Terry Zhang, Zhijing Jin and collaborators, “Position: Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk” — Spotlight paper
Sirui Lu, Zhijing Jin, Terry Jingchen Zhang, Pavel Kos, J. Ignacio Cirac and Bernhard Schölkopf, “Position: LLM for Physics Research Requires Domain-Specialized Training and Tooling”
Ruta Binkyte, Ivaxi Sheth, Zhijing Jin, Mohammad Havaei, Bernhard Schölkopf and Mario Fritz, “Position: Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution”
Marcus J. Min, Deyuan Mike He, Zhaoyu Li, Zixuan Yi, Sharad Malik, Aarti Gupta, Xujie Si and Osbert Bastani, “Position: The Case for Theory-Level Autoformalization” — Spotlight paper
AI4GOOD Workshop
Trustworthy AI for Good (AI4GOOD) Workshop, co-organized by 16 researchers including SRI’s David Lie, Zhijing Jin, Kexin Li and Wenjun Qiu, and Vector’s Terry Jingchen Zhang. Featuring keynotes by Yoshua Bengio (Université de Montréal), Joel Leibo (Google DeepMind), Oana Ignat (Santa Clara University), Maksym Andriushchenko (ELLIS Tübingen), and Naman Goyal and Jenny Ni (Google DeepMind), and a panel moderated by Milind Tambe (Harvard University) with Lewis Hammond (Cooperative AI Foundation), Gopal Sarma (RAND Corporation), and SRI Chair David Duvenaud (University of Toronto).
About ICML
The International Conference on Machine Learning (ICML) is one of the leading global conferences for research in machine learning and artificial intelligence. ICML brings together researchers, practitioners, and students working on the foundations, methods, applications, and impacts of machine learning. ICML 2026 will take place from July 6 to 11 at the COEX Convention & Exhibition Center in Seoul, South Korea, with tutorials, main conference sessions, workshops, and an expo.
Want to learn more?
Browse stories by tag:
- AI
- AI safety
- AI trust
- Computer science
- Computer security
- Copyright
- Cybersecurity
- Data
- Democracy
- Economics
- Education
- Engineering
- Ethics
- GPO-AI
- Governance
- Health
- Human Rights
- In the Media
- Jobs
- LLMs
- Law
- Normativity
- Philosophy
- Political Science
- Privacy
- Privacy Series
- Psychology
- Public Policy
- Recommenders
- Regulation
- Religion
- Reports
- Trust
- Workshops
