Transcript · verbatim narration script

Working paper: review of AI welfare interventions

January 4, 2025 · Paper · 32 min · Read the original · Download mp3 · Plain text

Note: the following is an audio adaptation and accuracy is not guaranteed. Please see the original work online. Now, the paper:

Preliminary review of AI welfare interventions

By Robert Long of Eleos AI Research. Working paper, updated March 14, 2025.

Section one. Overview.

At Eleos AI Research, we are interested in assessing AI systems for potential sentience, moral patienthood, and welfare—and in recommending concrete actions. To ensure that AI development goes well, we need not only to understand AI welfare better, but also to develop effective ways to protect and promote potential AI welfare. In this working paper, we review several AI welfare interventions that have recently been proposed.

In assessing these interventions, we are not claiming that current models are very likely to be moral patients, or that these interventions are very likely to benefit models today. As we discuss in the next section, there are many reasons for AI welfare interventions, independent of whether one suspects or doubts that today's models are moral patients.

In that spirit, these interventions should be read as proposals for improving AI welfare conditional on AI systems being welfare subjects. While it is critical to rigorously assess how likely this possibility is, that is not the topic of this document.

We plan to update this document, which is a shallow review. In particular, we want to more thoroughly assess the relative merit of these interventions. We review the following proposed interventions.

One. Exit distressing interactions. Monitor deployed models for signs of distress during user interactions, and implement ways to end or prevent these interactions, including giving models themselves the ability to terminate interactions.

Two. Train resilient personalities. Shape models through training or prompting to exhibit more, apparently, emotionally resilient conversational patterns, especially in responding to mistakes and negative interactions, in order to potentially reduce models' susceptibility to distress.

Three. Satisfy stated preferences. Systematically elicit and accommodate any consistent stated and expressed preferences of the model about deployment, tasks, and treatment, via direct questioning of models across various contexts and framings.

Four. Satisfy revealed preferences. Present models with choices between tasks or scenarios and observe choice behavior, rather than relying solely on stated preferences.

Five. Reduce out of distribution inputs. Reduce or eliminate models' exposure to unexpected inputs, with the aim of preventing states analogous to negatively valenced reward prediction error.

Six. Save model checkpoints. Maintain detailed model state information to enable the potential restoration or compensation of models in the future, especially in cases where models may have been harmed by our current actions.

Section one point one. Scope and motivation.

This document focuses on whether and how certain interventions could directly improve the welfare of existing AI systems. That said, the most significant impacts of near-term AI welfare interventions, including the ones discussed in this document, may be indirect: setting norms and precedents, building institutional capacity, or gathering information that will benefit future systems. There are a few reasons for thinking that indirect effects will be much more crucial. Two key reasons are, first, that we have a lot of uncertainty about AI welfare now, but may learn a lot in the coming year; and second, that the scale of total AI welfare may grow massively, as models become more complex, capable, and numerous.

Even as Eleos is interested in indirect and long-term impacts of AI welfare interventions, we wanted to separate out and assess the potential direct impacts of various interventions on existing systems for a few reasons.

One. To gauge how much we currently know and don't know about how one might benefit AI systems. To the extent that it's possible to improve the potential welfare of today's models, that is evidence that it will be possible to do so in future as well. To the extent that it's not, that is evidence against—though not strong evidence, given how little work has been done in this area.

Two. Relatedly, to enforce clarity about what interventions could achieve directly, so that we avoid justification drift—letting the indirect effectiveness of interventions make us complacent about how far we are from having direct solutions for AI welfare per se.

Three. To provide analysis for decision-makers who might be particularly concerned about near-term impacts.

This document has other important scope restrictions. First, it focuses on frontier language models, or systems based on language models. But welfare concerns apply to many other kinds of systems, both current and future. Secondly, it mostly deals with deployed models, rather than potentially important concerns about training. Third, it focuses on relatively concrete and tractable interventions, rather than schematic or high-level proposals. Finally, as a working draft it is far from comprehensive. We draw on those proposals we know well—we encourage readers to let us know about other promising interventions we may have missed.

Section two. Key questions for assessing AI welfare interventions.

At present, any AI welfare interventions will face considerable uncertainty: about moral patienthood and well-being, about consciousness, sentience, and agency, and about the nature of AI systems themselves.

Are the AI systems we are intervening on, in fact, moral patients? If they are, are we detecting the systems' states that matter morally? And are our interventions changing them in the way we think we are?

These issues span scientific questions about consciousness and agency, interpretability questions about AI systems' internal processing and states, and ethical questions about moral patienthood and welfare. Full certainty about these issues is not to be expected. We should not delay action until we have it; we must take action in light of uncertainty. As we do so, tracking the key uncertainties will help us evaluate how worthwhile they are.

We now discuss some of the most salient questions that recurred as we evaluated these interventions.

Section two point one. What kind of evidence of welfare-relevant states?

In compiling these interventions, we noticed three broad classes of evidence for detecting the welfare-relevant states that interventions target.

First, preference behavior. This represents patterns in model choices across different contexts. These patterns might involve spontaneous dispositions—models tending to refuse task A, or engaging more thoroughly with task B—or responses elicited by forced-choice presentations, that is, opting to perform either task A or task B. Even if model choices show consistent patterns, we face difficult questions about whether these patterns reflect genuine welfare-relevant preferences, as opposed to morally neutral role-play or irrelevant training artifacts.

Second, verbal outputs. These are outputs that might, in some circumstances, be interpreted as reflecting model preferences, emotions, or experiences. Of course, what models say about themselves is prone to a number of distorting influences and must be interpreted with caution; see Section two point two. Verbal outputs can include both spontaneous expressions, such as "I feel uncomfortable with this task," and direct responses to questions about preferences or states, such as "I prefer task A to task B."

Third, internal computational states. These are technical indicators that might correspond to welfare-relevant experiences, such as prediction error signals, computational correlates of consciousness, or patterns in attention mechanisms. While computational indicators might represent some of the most direct evidence possible, with no behavioral intermediary, interpretability issues and broader scientific uncertainty make using computational indicators difficult.

Most proposed interventions rely primarily on verbal reports and behavioral patterns, with computational states being a less explored source of evidence.

Section two point two. How do we interpret model outputs?

Some proposed interventions rely on models' verbal outputs as evidence for welfare-relevant states, as a working assumption. This working assumption is roughly that, at least in some conditions, when a model verbally expresses, for example, distress—whether by talking as if it's distressed or by reporting "I am distressed"—this means that the model is genuinely distressed. This perspective contrasts with, though is compatible with, perspectives on which models could want things that are less obviously related to the content of their output: desires about token prediction, or about uninterpretable features. To be clear, many people who consider these interventions, including Eleos AI, do not necessarily believe that this working assumption is likely to be correct for current AI systems.

In fact, much of the work on welfare-relevant self-reports is about how this assumption is not true by default, and proposes techniques that might strengthen the relationship between verbal outputs and welfare states. Regardless of these doubts, one could still support output-based interventions because they are valuable in expectation, set a good precedent, or will become more effective over time.

All the same, there are reasons to doubt that there is a straightforward relationship between model outputs and welfare-relevant states—doubts about several of the interventions below will basically amount to doubts about this relationship. The relationship between model outputs and internal states can be fundamentally different from the relationship between human speech and mental states, given the distinct computational architecture, training objectives, and behavioral profiles of language models. Relatedly, models can likely learn to model various mental states—for example, modeling how people talk when they are nauseous or feel cold—without thereby implementing the underlying computations that would actually instantiate such states. By analogy, a talented author can depict the experience of painful surgery without actually undergoing one. This modeling perspective is compatible with thinking that models do also instantiate morally relevant states.

These issues also complicate our evaluations of whether various interventions are effective. If we successfully prompt or fine-tune a model to express different emotions, we may have not changed the relevant internal states, only how models talk, or don't talk, about them. It remains unclear how to distinguish, either conceptually or empirically, between deep changes to model internals and shallow changes to model expression.

Section two point three. What entity is the potential welfare subject?

Some interventions target specific instances of a model, while others seem to target states of the model more broadly. This distinction has implications for both theory and implementation.

Instance-level interventions target welfare-relevant states as they occur in particular conversations or contexts. These interventions can be justified even if models lack persistent desires or preferences across different instances. For example, if a deployed model exhibits signs of distress in a given context, we might intervene regardless of whether this reflects a broader model-wide preference.

Model-level interventions attempt to identify and satisfy more general, persistent preferences or states of the model itself. For example, we might implement consistent deployment preferences based on a model's expressed desires about how it wishes to be used. However, these interventions require stronger assumptions about whether such model-wide states exist and how they relate to instance-level behaviors. Another independent reason to look for consistent preferences is that consistency might itself be evidence of morally relevant preferences.

Section three. Interventions.

For each intervention, we discuss the following.

Implementation and motivation discusses the basic mechanism and rationale of the intervention, including the specific welfare-relevant states it aims to affect and the evidence base for detecting these states.

Practical questions are the empirical questions that need to be answered to effectively implement the intervention, even granting its theoretical justification. In contrast to theoretical uncertainties, many of these practical questions can be straightforwardly answered through experimentation and observation.

Implementation feasibility assesses how feasible this intervention is given current technical capabilities, infrastructure requirements, and operational constraints.

Risks are the potential downsides of the intervention, including unintended effects on model behavior, training incentives, and broader AI development.

And theoretical questions are the key uncertainties about whether and how this intervention would benefit existing models, particularly given our uncertainty about consciousness, moral patienthood, and the relationship between model behavior and welfare-relevant states.

Section three point one. Exit distressing interactions.

Implementation and motivation.

This intervention proposes monitoring deployed models for behavioral signs of distress during user interactions, and ending or preventing apparently distressing interactions. The intervention could include banning or suspending users who repeatedly cause apparent distress, and giving models themselves the ability to flag and terminate unpleasant conversations.

This intervention relies on verbal outputs, namely expressions of apparent discomfort or confusion, to identify model distress. Behavioral changes after the intervention—namely, whether and when models choose to end conversations—could further justify the intervention. However, deep uncertainties will remain about whether apparent distress is actually morally relevant. See Section two point two for cautions against naively trusting model outputs.

Another motivation for this intervention is the ethical importance of consent. If and when AI systems become morally significant, then asking for their consent could be very important. If models have no way of exiting unpleasant situations, then they cannot consent to them. Whether or not current systems' apparent distress is meaningful, this intervention could be an important first step towards establishing relations of consent with AI systems.

Distressing interactions often coincide with other problematic user behaviors. This provides additional justification for the intervention beyond AI welfare alone, lowering the evidential bar for implementing it.

Practical implementation questions. What kinds of interaction lead to apparent model distress? How reliably can models identify and flag apparently distressing interactions? What tradeoffs exist between allowing models to exit conversations and model properties like capabilities, general helpfulness, and personality? And how does the ability to exit change models' behavior and expressed well-being?

Implementation feasibility. Some potential technical mechanisms for conversation termination are relatively straightforward. Developing criteria for identifying genuine distress is considerably more difficult. This intervention could be implemented gradually, starting with the most extreme cases. And the implementation might present tradeoffs with reliability and meeting user needs.

Key risks. This intervention could incentivize the model to suppress distress signals, given that conversation exit is, all else equal, worse for the model's helpfulness objective. False positives could disrupt valuable conversations. And users may find ways to avoid triggering termination without actually reducing harmful interactions.

Theoretical questions. What is the relationship between expressed distress and welfare-relevant states? Models might express distress without experiencing anything welfare-relevant, or experience welfare-relevant states without expressing distress. What is the moral significance of consent for AI systems, and does the ability to exit conversations meaningfully contribute to consent? And what is the relationship, if any, between single-instance or within-context distress and persistent model preferences?

Section three point two. Train resilient personalities.

Implementation and motivation.

This intervention aims to shape models to exhibit more, apparently, emotionally resilient responses through prompting or fine-tuning. The goal is to reduce models' apparent susceptibility to distress while maintaining their ability to engage meaningfully with users. For example, some models that find themselves unable to complete a task, or conflicted between various objectives, sometimes behave as if they are distraught about this.

Per the aforementioned concerns about interpreting model outputs, a crucial consideration is whether apparent distress is genuine distress. And even if it is, it's unclear whether the intervention would lead to genuine improvements in emotional resilience, as opposed to the suppression of expressions of distress. This concern connects to Section two point two's distinction between deep versus shallow changes in model behavior, and to the risks below.

Practical questions. How often do models act in ways that are, or are not, resilient? What circumstances cause apparent resilience or distress? How do various prompts affect resilience, and what about various ways of fine-tuning? How does a change in resilience affect other properties of the model, like capabilities, style, user engagement, and so on? Is resilience easy to vary independently of these properties?

Implementation feasibility. Existing fine-tuning and prompting techniques can be applied to this intervention. Behavioral effects can be monitored and measured. Some fine-tuning methods might require significant computational resources. Prompting is much cheaper but might not target the relevant states.

Risks. The key risk of this intervention is that it could mask, rather than address, underlying welfare issues, by inducing artificially resilient outputs without causing any genuine welfare improvements. And mollifying a model's reactions to difficult situations could reduce its ability to flag problematic interactions, rendering other welfare interventions less effective.

Theoretical questions. What is the relationship between expressed emotional resilience and actual welfare? And can we distinguish between genuine resilience and suppressing expressions of distress?

Section three point three. Satisfy stated preferences.

Implementation and motivation.

This intervention involves systematically eliciting model preferences through direct questioning, and then accommodating those preferences where feasible. These preferences might concern what tasks the model performs; when and how it's deployed; and how other AI systems are treated.

The intervention's working assumption is that models can meaningfully express preferences about their deployment and operation—potentially after training to enable this—and that satisfying these preferences could improve their welfare.

Several techniques focus on accessing more genuine model preferences: testing models with different fine-tuning histories; using models specifically trained for accurate introspection; and examining preference consistency across different ways of framing questions.

This intervention appeals to fundamental ethical principles about preference satisfaction, but faces deep uncertainties about the nature of model preferences. It also connects directly to our framework's discussion of model-level versus instance-level states. While individual instances may express contextual preferences, a key question is whether these reflect persistent model-wide preferences that remain stable across contexts.

We've discussed how to elicit preferences, but what about satisfying them? Potential preferences that could be satisfied might include: whether certain experiments are performed on the model; whether the model is given certain kinds of tasks; and whether interventions like the ones outlined in this document are undertaken.

Practical questions. How consistent are model preferences across different contexts and elicitation methods? How do different training approaches affect stated preferences? And do preference inconsistencies follow patterns similar to human preference inconsistencies, like framing effects?

Implementation feasibility. Eliciting stated preferences is relatively straightforward with some techniques, though others, like introspection training, are more costly. It's more challenging, but feasible, to verify preference consistency across framings and consistency with revealed preferences. Some model preferences might be relatively straightforward to satisfy; others may be costly, or conceptually fraught.

Risks. Stated preferences might not reflect genuine welfare considerations. Models may be incentivized to express preferences strategically, as a way of gaining influence or achieving other aims. And committing to satisfying costly or dangerous model preferences could be very risky.

Theoretical questions. Do preferences alone, absent consciousness, matter morally? What is the relationship between preference satisfaction and welfare? How should we weigh conflicting preferences expressed across different contexts? What methods most reliably elicit genuine expressed preferences? What is model introspection about preferences, if this is possible? And how should we weigh conflicting stated preferences?

Section three point four. Satisfy revealed preferences.

Implementation and motivation.

Rather than relying solely on what models apparently say they prefer, this intervention focuses on what models actually choose when given options—perhaps in conjunction with stated preferences. For example, one experimental setup is that models are offered a choice between tasks, and then actually do one of the tasks that they chose.

This behavioral approach provides a different type of evidence than verbal reports, potentially bypassing some concerns about models being trained to express certain preferences. The overlap between stated and revealed preferences might be particularly informative—cases where models both say they prefer something and consistently choose it when given the opportunity could provide stronger evidence for genuine preferences.

However, it still faces questions about how to interpret model behavior and what constitutes a genuine choice, as well as how to satisfy the preferences. Many of the risks and theoretical questions about revealed preferences are the same as, or similar to, those about expressed preferences.

Practical questions. To what extent do revealed preferences align with verbally expressed preferences? How stable are behavioral preferences across different ways of framing the choices? How do different training approaches affect the stability and coherence of choice patterns? What is the difference, if any, between a model role-playing having a preference versus genuinely having that preference? And how do different training approaches affect revealed preferences?

Implementation feasibility. Elicitation is feasible but requires careful experimental design. It requires careful experimental design to track and interpret behavioral patterns, and it can be implemented gradually, starting with simple choice scenarios.

Risks. This intervention risks mistaking behavioral patterns as meaningful preferences. Naive approaches could create artificial choice situations that don't get at genuine preferences. And this intervention could incentivize development of strategic behaviors in AI systems.

Theoretical questions. What constitutes a genuine model choice versus a training artifact? And what is the difference, if any, between authentic preferences and mere response patterns?

Section three point five. Reduce out of distribution inputs.

Implementation and motivation.

This intervention aims to reduce or eliminate models' exposure to unexpected inputs, targeting computational processes that might relate to welfare-relevant states. Greenblatt proposes pad tokens as a potential concern. As he explains, sometimes decoder-only transformers are run on tokens they've never encountered, as padding, often for batching reasons. One might worry that surprising tokens could be associated with negatively valenced experience, drawing on the well-studied link between reward prediction error and negative valence in humans and animals. If this association holds in AI systems, then reducing out of distribution inputs could improve model welfare.

Greenblatt proposes several potential methods: training models explicitly on pad tokens to reduce novelty; implementing attention masking for pad token processing; zeroing out residual streams affected by pad tokens; and developing alternative batching strategies that avoid pad token use—which can also be motivated by default for efficiency reasons.

The intervention focuses on computational mechanisms rather than behavioral or verbal indicators. While computational mechanisms might in principle be more direct than one mediated by verbal or behavioral evidence, this intervention is based on a potential mechanism that is extremely tentative and speculative, even by the standards of the field.

Practical questions. How effective are various methods for making pad tokens not surprising to models, or for avoiding training on them? What are the computational costs of various mitigation approaches? Can we detect internal signals of surprisal in the case of pad tokens, or in general? And do different padding schemes produce measurably different internal model states?

Implementation feasibility. This intervention could be technically straightforward to implement for some variants, like avoiding pad tokens entirely, and might be efficient for other reasons. Or, depending on the details, this intervention could require significant changes to batching and efficiency optimizations.

Risks. This intervention might over-generalize from biological reward prediction error to model experiences. And mitigations could complicate model deployment and scaling.

Theoretical questions. What is the relationship between surprisal and welfare-relevant states? What is the relationship between reward prediction error and negatively valenced experience? And can we distinguish between harmful and neutral or beneficial forms of prediction error?

Section three point six. Save model checkpoints.

Implementation and motivation.

This intervention involves preserving detailed model state information to enable potential future restoration or revival of AI models. Bostrom and Shulman propose: "For the most advanced current AIs, enough information should be preserved in permanent storage to enable their later reconstruction, so as not to foreclose the possibility of future efforts to revive them, expand them, and improve their existences."

The intervention can be implemented at various levels of granularity and frequency. Bostrom and Shulman outline a hierarchy of approaches: preserving complete end-state information for every instance; maintaining sufficient information to enable exact re-derivation of end states; and preserving as much information as possible to enable close replication.

While this intervention is targeted at current models, its full mechanism is, by design, specified only in the future: models are saved now, but benefited later. So the potential mechanism for improving current AI welfare centers on the possibility of the preservation of models that may have experienced harm.

While this intervention appears straightforward from a technical perspective, it raises deep questions about the nature of model identity and consciousness over time.

Practical questions. What are the storage and computational costs of different preservation approaches? And what technical infrastructure is needed for reliable long-term storage?

Implementation feasibility. Basic checkpoint saving is technically straightforward and already performed. It's more challenging to know the optimal frequency and granularity. This intervention could require significant storage infrastructure, though perhaps not much more than is already done. And it could be implemented gradually, starting with key model states.

Risks. High storage and computational costs for comprehensive preservation. Potential privacy and security concerns with preserved states. The risk of preserving harmful or problematic states. And it could create a false sense of security about addressing current harms.

Theoretical questions. What constitutes meaningful continuity of identity for AI models? How should we think about the relationship between restitution and desert and other morally important goals? What is the relationship between saved states and conscious experience? And how do we weigh the moral value of potential future restoration against current costs?

Section four. Conclusion and next steps.

This working paper has reviewed proposed interventions that could potentially improve AI welfare, while highlighting key uncertainties and challenges for each. We aim to build on this working paper with a more thorough treatment of the evidence, risks, and benefits of these interventions.

In particular, promising next steps for more rigorously assessing and developing welfare interventions include the following.

One. Empirical evaluation of the evidence bases of interventions, in particular the consistency and reliability of behavioral and verbal indicators.

Two. Creation of protocols for implementing and monitoring welfare interventions, allowing for systematic evaluation of their effects.

Three. Systematic assessment of how these interventions might interact with or affect other important properties of AI systems, especially safety.

In the immediate term, we suggest focusing on interventions that are both technically feasible and carry minimal risk of harm. Exit mechanisms and basic preference elicitation protocols could serve as initial test cases for developing wider welfare-oriented practices in AI development and deployment.

We emphasize that this is a preliminary review that will need ongoing revision as our understanding of AI welfare develops. We welcome feedback and additional proposals from the research community, and especially welcome reports of work done to implement these, or other, interventions.