The big picture of dangerous capability evaluations: David Duvenaud at the Seminar Series

 

“There will be systems much smarter than us that may not be aligned with us,” — Schwartz Reisman Chair in Technology and Society David Duvenaud.


What happens when AI systems grow too powerful for human oversight? That question anchored an urgent and thought-provoking session led by SRI Chair in Technology and Society David Duvenaud, professor of computer science at the University of Toronto. The seminar was jointly presented with the Department of Computer Science’s C.C. “Kelly” Gotlieb Distinguished Lecture Series and supported by the Webster Family Charitable Giving Foundation.

Watch the full session here:

Duvenaud opened with a challenge: to be honest about what we don’t yet know about AI safety. “At some point, there will be systems much smarter than us that may not be aligned with us,” he said. The discussion ranged from safeguards to philosophical reflections, all aimed at understanding when (and how) we might lose control of our technological creations. 

Among the key ideas explored were responsible scaling policies (RSPs) and AI safety plans, frameworks for gradually testing increasingly capable systems to detect catastrophic risks before deployment. Duvenaud explained how such measures might include restricting access to dangerous capabilities or enforcing nested oversight measures.

But the deeper question was one of trust. “It’s easy to show ability,” he noted. “It’s much harder to show inability.” Can we rely on models to be honest about what they know—or on other models to fairly evaluate them? Duvenaud discussed his work on the first suite of “sabotage evaluations” that probe whether systems can deceive or mislead humans, and “control protocols” that use simpler models to monitor more complex ones. 

As the conversation turned to governance, Duvenaud invited reflection on what it means for humans to remain at the helm of our species’ agency. “In what sense do humans currently control civilization?” he asked, pointing to bureaucracies, states, corporations, and ideologies that already regularly and predictably fail to serve the interests of those they’re meant to help.

Duvenaud explained that, in addition to the technical problem of alignment, we also will need to learn to keep our institutions aligned to human interests, even when competitive pressures will be pushing them to marginalize us and they no longer need to support humans for labor or involve them in decision-making.

Watch the recording to explore how researchers are rethinking control, oversight, and alignment in the age of accelerating AI capability.

Want to learn more?


Browse stories by tag:

Related Posts

 
Previous
Previous

Therapy bots: Regulating the future of AI-enabled mental health support

Next
Next

SRI appoints Bruce Schneier as visiting senior policy fellow