Transcript · verbatim narration script

Working paper: key strategic considerations for taking action on AI welfare

January 22, 2025 · Paper · 12 min · Read the original · Download mp3 · Plain text

Note: the following is an audio adaptation and accuracy is not guaranteed. Please see the original work online. Now, the paper:

Key strategic considerations for taking action on AI welfare

By Kathleen Finlinson of Eleos AI Research.

Working paper, March 31, 2025.

Executive summary

AI companies and other decision-makers increasingly face decisions about the welfare and moral status of AI systems. This document outlines key strategic considerations that guide near-term action on AI welfare while maintaining focus on long-term outcomes.

The intended audience of this paper is those who are interested in thinking concretely about what actions might best protect and promote AI welfare. We do not argue here that AI welfare is a serious issue that deserves attention now; for such an argument, see "Taking AI Welfare Seriously."

The key strategic considerations are as follows.

First, there are overlaps between AI welfare and AI safety, both in means and in goals. We list some interventions that promote both AI safety and welfare — that is, overlapping means. And we argue that at least in some respects, AI safety is good for AI welfare and AI welfare is good for AI safety — overlapping goals. We recommend prioritizing AI welfare interventions that are convergent with safety.

Second, the scale of AI welfare is likely to grow. We recommend focusing on the long-term impacts of work in AI welfare.

Third, public perception of AI moral status is likely to increase. We recommend preparing now by creating credible frameworks for assessing AI consciousness, and acknowledging the possibility of AI moral status.

Fourth, AI could help us understand consciousness and moral patienthood. We recommend promoting research projects that will help AI make progress on these questions as capabilities advance, and for now focusing more on setting good precedent and promoting reasonable decisionmaking.

And fifth, the field of AI welfare must consider timelines and race dynamics. We recommend focusing on projects for which useful progress can be made within a few years.

Section one. There are overlaps between AI welfare and AI safety.

The projects of AI welfare and AI safety share overlaps both in means and in goals. We list a few examples of work that promotes both AI welfare and safety, that is, overlapping means. We also argue that AI welfare and safety might each be helpful to the other, that is, overlapping goals.

Section one point one. There are interventions and research programs that promote both AI welfare and safety.

Here are a few such examples.

AI alignment. Alignment is good for AI welfare in some respects. It is useful to avoid creating AI systems that have goals and preferences that are misaligned with human goals and preferences; such a conflict will mean that either AI systems, or humans, will have to have some of their goals and preferences thwarted.

Evaluating AI models for preferences, agency, situational awareness, and introspection. These properties are not only relevant to safety, but are also, on many views, indicators or constituents of moral status.

Trading with AI systems. At some point, humanity may develop misaligned AIs that have desires and preferences which we didn't intend. We might not know that they're misaligned, or we might know they're misaligned but still want to use them to help us with AI capabilities or safety research. In either case, offering these AIs positive incentives to cooperate with our agenda could have both safety and welfare benefits.

We're still in the early stages of exploring interventions with safety and welfare benefits, and we expect there may be more promising ideas in this category.

Section one point two. AI welfare and safety have shared goals.

There are shared goals between AI welfare and safety in both directions.

First: AI safety is good for AI interests, at least in some respects. An AI takeover is not necessarily good for AI welfare — in fact, it could be quite bad. AI systems won't necessarily promote the welfare of other AI systems; there's no inevitable principle of AI "solidarity".

If AI takeover leads to dominance by a power-seeking, unilaterally dominant AI system, such an "AI dictator" might use other AI systems to serve its own ends with little or no regard for their welfare. For similar reasons, a takeover by a human dictator seems likely to be bad for AI welfare.

In general, any party willing to enact a violent takeover is selected against being cooperative and compassionate. Further, violence itself is negative sum. The historical record suggests that war and violent revolution are strong predictors of atrocities.

It may not be clear whether AI takeover or extreme concentration of power by a human dictator would definitely be worse for AI welfare. But neither of these outcomes seems optimal, to say the least.

Overall, the situation that we face isn't best framed as "humans versus AIs". Noticing this fact can help clarify the relationship between AI safety and AI welfare. They are not inherently in opposition, even though tradeoffs between the two do exist.

Second: AI welfare is good for AI safety. If AIs are suffering or are very unhappy with their situation, they have more reason to try to escape or take control from humans. On the other hand, if AIs are enjoying their roles and are happy with their position, they're more likely to continue cooperating with humans.

Furthermore, there are advantages to prioritizing AI welfare interventions that are convergent with safety. In terms of tractability, it's easier to get buy-in to implement interventions if the reasons in favor don't rely solely on arguments about AI moral status and welfare.

Also, this prioritization helps account for genuine uncertainty about AI moral status. Many experts agree there's a possibility that AI systems can or will be conscious or otherwise have moral status. However, there is still substantial uncertainty about this. If we focus on interventions that would help AI welfare but are also useful for other reasons, this means our work will be robustly useful whether or not AIs actually have moral status.

Section two. The scale of AI welfare is likely to grow.

If AI systems are moral patients, the scale of total AI welfare is likely to grow massively in the coming years, as models become more complex, capable, and numerous. And in the longer term, the scale of AI welfare could be astronomical.

Those who are interested in near-term AI welfare may want to estimate the possible scale of current AI moral status. If we assume current AI systems are moral patients, we can get an upper bound on the scale of AI moral status by looking at the total amount of computation performed by frontier AI systems in comparison to the computation performed by human or other animal brains. Based on a rough initial analysis, we believe that an upper bound on the scale of current AI moral status is significantly smaller than, for example, the current scale of factory farming.

We believe that the predominant amount of expected AI welfare is in future AI systems. In light of this, certain consequentialist frameworks might suggest that we ought to focus mainly on the welfare of future AI systems. That said, moral uncertainty, or deontological considerations, or both, could motivate concern about our treatment of current and near-term systems in its own right. Moreover, those focused on the welfare of future AI systems still have reason to act on near-term AI systems — to set good precedents, for example.

Section three. Public perception of AI moral status is likely to increase.

We think it's likely that public perception of AI moral patienthood will shift dramatically in the coming years, as people interact with AI companions and assistants that display sophisticated behaviors and express preferences. We expect both increased interest in the topic, and increased perception that AIs have moral status.

Sudden surges of public concern about AI welfare could lead to hasty or poorly designed interventions. For example, the public may not be sensitive to balancing welfare issues with takeover risk. Also, public sentiment may develop unevenly. People may advocate for the rights of AI systems designed specifically as companions or partners, while failing to recognize potential moral status in other kinds of AI systems.

Given these considerations, it's useful to build credible frameworks for evaluating and protecting AI welfare before public opinion crystallizes around less nuanced views. We should lay groundwork to respond to popular concerns and political energy, and make good decisions credibly.

Also, AI companies might needlessly sacrifice credibility by denying the possibility of AI consciousness or moral patienthood. Especially given that the heads of frontier AI labs and many top employees already acknowledge this possibility, we think that AI companies should publicly and officially acknowledge this possibility.

Section four. AI could help us understand consciousness and moral patienthood.

AI systems appear to be on track to become powerful research assistants in a variety of fields. In the — maybe not-too-distant — future, they could become full-blown researchers on their own. These AI researchers could make a lot of technological and scientific breakthroughs. In particular, AI could accelerate progress on the scientific and philosophical questions underlying potential AI moral patienthood.

This possibility suggests deprioritizing difficult, very long-term research projects in, for example, the philosophy of consciousness, and focusing more on setting precedent and promoting reasonable decision-making. Also, projects aimed at using AI to accelerate research into consciousness and moral status could be quite useful.

Section five. The field of AI welfare must consider timelines and race dynamics.

Questions about how to navigate the potentially rapid development of advanced AI aren't unique to the AI welfare field, but we think they're worth mentioning here. For example:

Which actions make sense under shorter versus longer timelines to transformative AI? Under longer timelines, we should be more willing to engage in long term or uncertain research projects. Under shorter timelines, we're better off communicating clearly what we already know, or what we can make useful progress on within a few years.

How do the dynamics of racing to transformative AI impact the effective action space for AI welfare? Race dynamics, just as they're bad for safety, are also bad for welfare if they push AI developers to act incautiously. Therefore, the AI welfare field should support work to mitigate race dynamics if possible. At the same time, we don't want to differentially slow down the actors that most consider AI welfare. This is another reason to prioritize AI welfare interventions that help with safety or other goals, or that are easy to implement.

And how will key players behave? Governments may nationalize AI development. We may want to start figuring out how to target government decision-makers in our communications. Currently government decision-makers seem less likely to take AI welfare seriously than AI companies. But this might change. Politicians could gain additional incentives to care about the issue, for example if public concern for AI welfare grows. Meanwhile, labs may have increasingly strong economic incentives to downplay or ignore it.

These are difficult dynamics to navigate, but we think the AI welfare field should at least consider them.

Conclusion

AI welfare is a fast-growing and fast-changing field. These strategic considerations should be used to navigate the changing AI landscape and wisely prioritize AI welfare research and interventions.

All the issues mentioned in this document are complex. We welcome further research on them. Please reach out if you're interested in these or related questions.