Beyond algorithms: Travis LaCroix on AI and the value alignment problem

 

SRI Faculty Affiliate Travis LaCroix discusses his new book Artificial Intelligence and the Value Alignment Problem, exploring how alignment is as much about power and justice as it is about technology.


Artificial intelligence raises profound questions about values, power, and responsibility. In his new book, Artificial Intelligence and the Value Alignment Problem (Broadview Press, 2025), SRI Faculty Affiliate Travis LaCroix examines these questions through both a technical and social lens. Drawing on his experiences teaching AI ethics to students in philosophy and computer science, LaCroix traces how value alignment is not only a technical challenge of encoding human goals into machines, but also a deeply political question about whose values are represented, who benefits, and who bears the risks.

An assistant professor in the Department of Philosophy at Durham University, LaCroix has also previously taught at Dalhousie University and held postdoctoral fellowships at Mila – Québec Artificial Intelligence Institute, the Université de Montréal, and the University of Toronto. His research brings together philosophy of science, ethics, and AI to explore how societies can grapple with the risks and possibilities of emerging technologies.

In this interview, LaCroix reflects on the development of his book, including why alignment should be understood as a matter of community and justice, and considers how educators, researchers, and policymakers can respond to the social impacts posed by AI systems.

The following conversation has been lightly edited for length and clarity.

Schwartz Reisman Institute: In the book’s introduction, you describe that it grew out of challenges in finding the right teaching material. How did your experience as an educator shape the book’s development and structure?

Travis LaCroix: I first taught an undergraduate philosophy course on AI and value alignment at the University of Toronto in 2021, which explored how many AI ethics issues such as algorithmic racial bias could be seen as specific cases of value misalignment. Later that year, I joined Dalhousie, where I taught a required first-year computer science course that adapted material from my philosophy seminar. Around the same time, Stephen Latta at Broadview Press approached me about writing a book on the philosophy of AI. I had already been compiling weekly lecture notes for my students, which eventually became the foundation for the book—the book’s structure mirrors the twelve lectures of the course.

Artificial Intelligence and the Value Alignment Problem: A Philosophical Introduction (2025) is published by Broadview Press.

Teaching the class over multiple years showed how quickly interest in AI was changing. In my first course, before ChatGPT’s public release, many students were somewhat indifferent to AI. The following year, just six weeks after ChatGPT launched, students arrived with much more curiosity, signalling that AI had suddenly appeared more urgent and widely relevant.

Student feedback also played a key role in shaping the book. In the second iteration of the course, I used their comments to refine the sequencing and presentation of topics. In the final iteration, students read draft chapters of the book itself, offering insights on what was clear, engaging, or confusing. Interestingly, many found the first chapter on the history of AI most compelling—several students said they hadn’t realised that AI existed beyond the last few years. This iterative teaching experience directly influenced the book’s development and helped ensure it was accessible, relevant, and engaging.

SRI: You frame value alignment using the economic model of the principal-agent problem. How can this lens help clarify our understanding of AI alignment?

LaCroix: The “value alignment problem” is often described as the challenge of making sure AI systems act in line with human values, but this quickly runs into complications. What exactly are “human values”? Are they individual preferences, cultural norms, or universal principles? Whose values count, and how do we reconcile disagreements across people, societies, or even within ourselves? Even if we could fully articulate our values, we would still face the difficult task of formally encoding them into AI systems and ensuring those systems follow them as intended.

My approach starts by stepping back from the question of what the right values are, and instead asking how misalignment happens in the first place. By exploring the contexts under which misalignment can occur, we shift focus from the normative (what the correct values are) or technical aspects (how we encode values in an AI system) to exploring the structure of the value alignment problem.

The model of the principal-agent problem describes the structure of value alignment: we defer authority to an artificial system to act on our behalf—as a principal does with an agent in the economic context—and that system may or may not satisfy our objectives as intended. One insight that emerges from the economic analysis is that it is not just misspecified objectives that drive value misalignment, but information asymmetries.

Applying this economic model to AI clarifies several things. First, there isn’t a single “alignment problem” with a single solution. Instead, there’s a whole family of problems that must be addressed on a case-by-case basis. Second, it highlights a limitation of current AI research: as models grow larger and more general-purpose, the potential for misalignment also grows. Larger systems require more data and more complex objectives, which increases the risk of divergence. This model helps us understand why misalignment arises, what forms it can take, and why aiming for “general” AI, as many leading tech companies claim to be, will necessarily exacerbate the problem.

“Misalignment is not just a technical problem but fundamentally a problem of power. Alignment, by contrast, should be understood as a matter of community.”

SRI: You argue that alignment is not just a technical challenge but a question of power. How do we ensure that aligning AI with “human values” doesn’t simply mean aligning it with the values of the most powerful actors—governments, corporations, or dominant cultures—at the expense of democratic and pluralistic values?

LaCroix: The idea of “aligning AI with human values” sounds simple, but it quickly becomes problematic if we don’t ask whose values are being aligned. Without that scrutiny, we risk building systems that reinforce those with power while sidelining marginalised groups.I think it’s more accurate to say that misalignment is not just a technical problem but fundamentally a problem of power. Alignment, by contrast, should be understood as a matter of community. It’s a social and political project: deciding together what kinds of technologies we want, who they should serve, and what trade-offs we are willing to accept.

If alignment is framed only as a technical problem, the solutions will remain in the hands of technical elites. But if we recognise it as a question of power and justice, then alignment becomes a collective responsibility—one that demands democratic debate, pluralistic values, and the empowerment of those who have historically been excluded from shaping technological futures.

SRI: How can we address the deeper social and political structures that produce misaligned AI?

LaCroix: Part of the emphasis on structural misalignment is to shift the conversation to focus on the complex socio-political systems and power structures within which AI technologies are embedded: Whose interests does this system serve? Who benefits, and who bears the risks? Should this system even exist?

Moving forward requires shifting power. That means broadening the scope of who gets a say in AI development. This could take many forms: community-led governance structures, participatory design processes, citizen assemblies on technology policy, or even cooperatives where communities directly shape how AI is deployed in their lives. It also means challenging business models that drive harmful applications—such as predictive policing or exploitative labour platforms—and pushing for structural regulation, including data protections, labour rights, antitrust measures, and outright bans where necessary.

Ultimately, AI systems do not create new injustices so much as they automate old ones at scale. The task, then, is not just to make algorithms “less biased” but to confront the broader social and political systems in which these technologies are embedded. Sometimes, the most ethical choice may be to decide that certain AI systems should not be built at all.

SRI: Could you elaborate on how this contrasts with the current prevailing “more data, more power” narrative in AI research? How should policymakers and researchers respond to the tension between scaling and alignment?

LaCroix: The dominant narrative in AI research today is the scaling hypothesis: the bigger the model, the more data and compute you throw at it, the better it performs. But as these systems scale, so do the risks of misalignment.

One way to see this is through information asymmetry. Smaller models can be trained on carefully curated datasets, so we have some idea of what goes in. But once models reach massive scale, hand-curation becomes impossible, so we rely on indiscriminate web crawling. That means no one really knows what data the model has been trained on, which makes it harder to predict or control its behaviour. In other words, the very process of scaling creates new points of opacity that amplify misalignment.

For policymakers and researchers, the implication is clear: scaling should not be treated as a universal path to progress. Bigger is not always better. Narrower, more targeted systems—trained on curated, transparent datasets—may offer safer and more accountable paths forward. Sometimes the responsible choice isn’t to scale up, but to deliberately constrain.

“It should be obvious that interdisciplinary collaboration isn’t optional: it is essential for meaningful, responsible AI alignment.”

SRI: Which case study do you find most illustrative of the alignment problem?

LaCroix: Although it is somewhat low-hanging fruit, from an ethics perspective, one of the clearest examples is predictive policing, because it demonstrates why some AI applications can never be truly aligned, regardless of the amount of data or technical refinement applied.

Predictive policing tries to forecast future crime, but this objective is a poor proxy for what communities really want—safety and harm reduction. Because it’s based on predicting contingencies that can never be observed directly, the system’s outputs cannot ever align with those broader goals. In other words, misalignment is baked into these systems. This example makes clear that alignment isn’t just about data or accuracy—it’s about power, values, and who gets to decide what “good” means.

SRI: What advice would you give to students or researchers entering this space who want to tackle alignment from an interdisciplinary perspective?

LaCroix: As Emily Bender often says: Resist the urge to be impressed! It’s essential to engage directly with the systems themselves—learn how they work, what their limitations are, and speak about them in precise, literal terms. Avoid anthropomorphising AI systems; these are tools and processes, not sentient beings.

Equally important is listening to those most affected by these technologies. Marginalised communities, historically excluded voices, and individuals living at the intersections of social, economic, and environmental vulnerability often possess insights that are invisible from within technical labs or policy circles. Their experiences should guide the questions we ask and the solutions we design.

Interdisciplinary engagement is crucial. Beyond philosophy and computer science, fields in the social sciences are already contributing valuable perspectives on power, inequality, and institutional dynamics. Critical disciplines like gender studies, critical race theory, and disability studies offer essential frameworks for understanding how values are embedded in technology. Economics, law, and environmental studies can shed light on the incentives, regulatory structures, and systemic consequences of AI deployment.

It should be obvious that interdisciplinary collaboration isn’t optional: it is essential for meaningful, responsible AI alignment.

Want to learn more?


Browse stories by tag:

Related Posts

 
Previous
Previous

Rethinking knowledge in the age of AI

Next
Next

AI and digital innovation needs science too