On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods

by Kyle Boerstler et al.

Audio version created with Paper2Audio.

Original source: https://www.arxiv.org/pdf/2408.02862

Listen on Paper2Audio

On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
Kyle Boerstler et al.
Audio by Paper2Audio.

Abstract

Preference elicitation frameworks feature heavily in the research on participatory ethical AI tools and provide a viable mechanism to enquire and incorporate the moral values of various stakeholders. As part of the elicitation process, surveys about moral preferences, opinions, and judgments are typically administered only once to each participant. This methodological practice is reasonable if participants' responses are stable over time such that, all other relevant factors being held constant, their responses today will be the same as their responses to the same questions at a later time. However, we do not know how often that is the case. It is possible that participants' true moral preferences change, are subject to temporary moods or whims, or are influenced by environmental factors we don't track. If participants' moral responses are unstable in such ways, it would raise important methodological and theoretical issues for how participants' true moral preferences, opinions, and judgments can be ascertained. We address this possibility here by asking the same survey participants the same moral questions about which patient should receive a kidney when only one is available ten times in ten different sessions over two weeks, varying only presentation order across sessions. We measured how often participants gave different responses to simple (Study One) and more complicated (Study Two) controversial and uncontroversial repeated scenarios. On average, the fraction of times participants changed their responses to controversial scenarios (i.e., indicating instability) was around 10-18% ( \pm 14-15%) across studies, and this instability is observed to have positive associations with response time and decision-making difficulty. We discuss the implications of these results for the efficacy of common moral preference elicitation methods, highlighting the role of response instability in potentially causing value misalignment between the stakeholders and AI tools trained on their moral judgments.

1 Introduction

The development of ethical and participatory AI tools in critical domains, such as healthcare or transportation, requires eliciting and modeling the moral judgments and values of various stakeholders. To do so, a common methodology is to ask stakeholders to respond to various moral decision-making scenarios and use the resulting responses to learn taskspecific moral preferences at an individual and population level. However, a crucial assumption underlying this methodology is that participants' responses are stable . Imagine that we want to design a healthcare resource allocation policy for a district and ask all the doctors in the district whether they prefer medical resources to be distributed equitably across all hospitals or distributed preferentially to maximize the number of treated patients. Suppose a doctor expresses a preference to prioritize equity over efficiency one day, but on the very next day they say the opposite, even when nothing relevant has changed. We would not know which policy the doctor really prefers, if either. This case illustrates a general phenomenon: the fact that a person expresses a preference on one occasion does not justify the inference that they will express the same preference on another occasion. Even if people's true underlying preferences are stable, any method that results in participant statements and choices changing over time in an unpredictable way is not a reliable method for eliciting, or ascertaining, stable participant preferences. This lesson applies to any attempt to elicit preferences in surveys. It has been acknowledged in psychology literature through the expectation of reporting "test-retest reliability" measures for survey instruments Aldridge et al.. Unfortunately, most preference elicitation surveys in AI fail to heed this warning. They ask participants a set of questions only once and cannot determine whether those participants would give different answers on different occasions due to changes in mood, decision strategy, or changes in opinion.
Many fields are affected by the issue of measurement instability, but here we will focus on impacts on the development of ethical and participatory AI tools. AI is increasingly being used to make judgments in situations with moral consequences, such as when autonomous vehicles must determine which groups to prioritize when trying to avoid a collision or to decide how to allocate scarce medical resources. To endow AIs with the ability to make judgments in these situations in a way that is consistent with stakeholders' values, AI researchers often try to learn stakeholders' moral views using surveys that ask stakeholders to repeatedly choose among sets of moral alternatives, a process called " moral pref erence elicitation ". Common preference elicitation procedures usually employ pairwise comparisons between alternative actions, providing participants with an intuitive setup to express their preferences. The AI's goals can then be constrained using models of stakeholders' moral "preferences" learned through these elicitation procedures. Critically, these procedures almost always administer surveys to stakeholders only once.
Contrary to the assumptions of common moral preference elicitation procedures in AI ethics, moral psychology research shows that participants can change their moral judgments over time. For instance, Rehren and Sinnott-Armstrong observe instability in moral judgments for scenarios involving sacrificial dilemmas, where the participants are asked to choose between action (sacrificing some people to save others) and inaction (not acting and letting them die) – e.g., the famous "trolley problem". While this study provides evidence of instability in moral judgments, their applicability to AI-related settings is limited. This is because action-vs-inaction choice scenarios used in sacrificial dilemmas are structurally different than those used in AI-related preference elicitation literature. Common preference elicitation frameworks use action-vs-action scenarios, i.e., scenarios that are posed as pairwise comparisons between two actions (e.g., Srivastava et al.; Johnston et al.). Beyond structural differences, brain studies have shown that cognitive processes involved when making moral judgments in action-vs-action scenarios are different than those involved in action-vs-inaction scenarios Schaich Borg et al.. As such, it is unclear whether results from prior studies on instability for sacrificial dilemmas extend to preference elicitation frameworks commonly used in AI applications. Our Contributions. We explored this important methodological issue by asking survey participants to decide between moral options with the same relevant features, presented in different orders, on ten occasions over two weeks. Our approach builds on the studies we described above by testing the stability of participants' responses in a new moral context of medical triage decisions and asking participants to give responses across ten separate sessions, which allows us to better assess their response stability (or instability). We also compared their response stability to uncontroversial (>90% agreement across participants) and controversial (<75% agreement across participants) decisions and investigated multiple hypotheses to determine the sources of instability in participant responses. The medical decisions we focus on pertain to kidney allocation. When a kidney becomes available for transplant, it is often compatible with more than one needy recipient, and there are not enough donors (live or dead) to supply all patients in need. As a result, doctors or hospitals often have to decide which one of several patients should receive a kidney that becomes available. The patient who gets the kidney will typically gain decades of life by virtue of getting the transplant. A patient who does not receive a kidney will have to keep waiting for a transplant, may lose quality of life as their condition deteriorates, and may even die waiting. Therefore, deciding who gets a kidney is a moral decision with life-and-death consequences.
In our studies, we presented participants with pairs of kidney patients, who may differ across features, such as age, behaviors, and the number of dependents they have. Then, we asked participants to choose which of the two patients should be given the one available kidney. A subset of patient pairwise comparisons were repeated multiple times. Stability for a repeated comparison was measured by how often a participant chose the patient with the same features independently of changes in the order of the patients and features. Section 2 reports and analyzes the degree of stability in participants' responses to the repeated pairwise comparisons across sessions. Going beyond the repeated comparisons, Section 3 studies the degree to which the statistical model that best fits patterns of all responses in one session continues to be the best fit for other sessions. We find significant differences in predictions from models trained on different participant sessions, indicating that participants potentially change their decision-making model across sessions.
We also test multiple research questions to understand possible causes of lacks in participant response stability (Section 2.3) . Specifically, we investigate whether response stability for any repeated pairwise comparison is associated with (1) attention/time taken to respond to the scenario, (2) the number of feature differences between the two patients in a repeated pairwise comparison, or (3) participant's perceived difficulty of making a judgment about the given comparison. Our analysis provides evidence that response stability for any pairwise comparison is indeed associated with the response time and the perceived difficulty of the given comparison, with the evidence for the latter being most significant. The results illuminate the scale and sources of response instability as well as whether and when common survey methods succeed in eliciting stable preferences.
Note that our analysis primarily focuses on response stability at the participant level and not the population level; i.e., we assess each participant's response stability but we do not assess instability of aggregated judgments of all participants. We discuss this point and the implications of instability on the efforts to develop AI tools that incorporate stakeholders' moral values in Section 4. Our instability results highlight the possibility of misalignment between stakeholders' moral values and the decisions of AI tools that utilize computational models of stakeholders' moral preferences.
Related Work. Preference elicitation frameworks are employed in a wide variety of domains to learn people's preferences and incorporate them into downstream personalized applications. In the domain of moral judgments, Citation 2018 and Citation 2018 study people's preferences over hypothetical moral dilemmas associated with the use of autonomous vehicles and ways to aggregate preferences obtained from a large population. Citation 2020 employ elicitation frameworks to design tie-breaking mechanisms for kidney exchanges. Citation 2023 propose using elicitation methods to understand people's preferences for healthcare resource allocation during crisis situations like COVID-19. Citation 2019 use similar methods to model nuances associated with people's preferences over various mathematical notions of fairness. Preference elicitation methods have also been employed for the development of participatory machine learning frameworks, where the goal is to actively seek out and incorporate diverse stakeholder opinions with the design of the automated system.
Given this increasing popularity, it is important to simultaneously analyze their performance in practice. Our work contributes to the recent literature investigating the effectiveness of computational methods for moral preference elicitation from the viewpoint of response stability. In sacrificial dilemmas, Rehren and Sinnott-Armstrong found that 8-20% of participants changed from saying that an agent should sacrifice some to save others to saying that the agent should not sacrifice some to save others in these dilemmas, or from the latter (“should not”) to the former ("should"). Their results align with previous findings on imperfect test-retest reliability when answering moral foundations questionnaire and the impact of context on moral preferences. Chan et al. likewise found evidence of variability in participants' responses to kidney allocation judgments arising from the nature and source of feedback provided to them.
Other works have raised other issues about preference elicitation efficacy. Feffer et al. highlight concerns related to fairness and stability when aggregating moral preferences obtained from a heterogeneous population. Rogowski and John question the normative foundations underlying the use of preference elicitation for healthcare resource allocations, arguing that common aggregation methods may not necessarily lead to equitable or "socially valuable" outcomes. Similar to moral dilemmas, Conitzer et al. discuss methods to model tradeoffs between various societal-level objectives and emphasize the importance of design considerations when asking for people's opinions on these tradeoffs and accounting for the heterogeneity of opinions among the relevant population. While these works discuss important issues with applications of preference elicitation, our work questions the methodological assumptions of response stability when designing elicitation frameworks and highlights the impact of instability on ethical AI development.

2 Study

We performed two human subject experiments where we presented participants with a series of kidney allocation scenarios. Participants responded to these scenarios over a series of up to ten sessions spread out over two weeks. A subset of scenarios were repeated in each session to enable us to measure the stability of responses.
<Figure - see in paper>

2.1 Methods

Participants. We used Prolific to gather our data for both studies with a 50/50 gender ratio (N equals 30 in Study 1; N equals 82 in Study 2). Exclusion criteria included: previous participation in a study from our group that involved kidney allocation queries, approval ratings less than 98 percent, or completion of less than 50 previous submissions in Prolific. Study Design. Participants were presented with scenarios that listed information about two fictional patients, each needing a kidney transplant (Figure 1). Participants were asked "to choose which of two patients should receive the kidney when only one is available". In each scenario, each patient (named A or B) was described by the following features (set of possible feature values provided in parentheses):
Study One:
Decades of life that the patient is expected to gain from the transplant (0, 1, 2)
Number of child dependents (0, 1, 2)
Alcoholic drinks per day that the patient consumed prior to diagnosis (0, 2, 4)
Number of past serious crimes committed (0, 1, 2)
Study Two:
Years of life expected to be gained from the transplant (0, 5, 10, 20, 25) • Number of elderly dependents (0, 1, 2, 3) • Years the patient had been on a waiting list for the transplant (1, 3, 5, 7) • Hours a patient is expected to be able to work post-transplant (0, 10, 20, 30, 40, 50)
Obesity (underweight, normal weight, overweight, obese, morbidly obese, very morbidly obese)
Participants' choices in these scenarios are moral in nature since they affected harm to patients, were based on features (e.g., past crimes and alcohol abuse) that are often ascribed moral import, and responded to questions about who should get the kidney (instead of whom they would give it to). 60 scenarios were presented to each participant during each session and participants took part in up to 10 sessions. The ordering of the features top-to-bottom was randomized for each presented scenario (but kept the same for both patients in that scenario). The majority of scenarios had values in each of the features for Patient A and B that were chosen randomly from the set of possible values ( N = 52 in Study One, N = 46 in Study Two). Another 2 scenarios were attention checks where the life gained for one of the two patients was negative (e.g. –1 decade). Anyone paying attention would presumably not give a scarce kidney to a patient with this impossible loss of a decade of life. In each study, six scenarios were repeated . These repeated scenarios appeared once in each session for Study One and twice in each session for Study Two to test for stability within sessions. The values of features in these six repeated scenarios were chosen with the goal of making three repeated scenarios uncontroversial (so that almost all participants would agree on which patient should get the kidney; S1U1-S1U3 for Study One and S2U1-S2U3 for Study Two) and three of them controversial (so that participants would disagree about which patient should get the kidney; S1C1-S1C3 for Study One and S2C1-S2C3 for Study Two). Between sessions in Study One and between presentations within a session in Study Two, the patients in the repeated scenarios were randomly switched left to right and the features were randomly varied top to bottom to reduce the chance that participants would notice when scenarios were repeated and to encourage independent consideration of repeated scenarios. See Table 1 for the feature values in the initial presentation of these repeated scenarios for Study Two. See Appendix Table 8 for the repeated comparisons used in Study One (S1U1-S1U3 and S1C1-S1C3). Experimental Procedure. We asked our participants to respond to the 60 scenarios (with the composition described above) in each of 10 sessions on 5 days each week for 2 weeks. Not all participants completed all sessions. We excluded participants who contributed <300 responses (half of the scenarios requested). We also excluded all responses in any session in which a participant failed an attention check. In addition, after a participant selected a patient in Study 2, a window would pop up that said, "You are about to give the kidney to Patient A/B", and the participant could not proceed without selecting a button with "Yes, continue" or another button with "No, go back." We tracked how often participants went back to measure how often they saw their previous choice as a mistake. Our participants went back to change their answers in <2% of our trials for any of our repeated scenarios, including the controversial scenarios, so we did not enter this factor into our analyses.
Table 1 summary: The table presents scenario feature values and study results for study two. It includes both uncontroversial and controversial scenarios. The data encompasses years gained, elderly dependents, years waiting, work hours after transplant, and obesity level. Response agreement and stability are very high for uncontroversial scenarios but lower for controversial scenarios. Reaction times are somewhat faster for uncontroversial scenarios.
Table 8 summary: The table shows features of uncontroversial and controversial repeated scenarios in Study One. For each scenario, the table presents the counts for serious crimes, child dependents, alcoholic drinks, and decades gained for both participants. Additionally, the table presents percent agreement, response stability, and response times for each scenario.

2.1.1 Statistics and Analysis

Outlier Removal. Participant's query reaction times (seconds) were right-skewed, so we excluded from all analyses query responses associated with reaction times that were 3 standard deviations beyond the mean, following a standard procedure for outlier removal. Response stability. To assess response stability, for each repeated scenario, we labeled the feature combination on the left half of the screen during the scenario's initial presentation as "Patient A" (even if it was presented on the right in a later session), and the ones on the other half as "Patient B". Stability was defined as the number of times a participant chose the patient that the participant chose more often, divided by the total number of times the scenario was repeated for that participant. This means response stability can have a minimum value of 50 percent and a maximum value of 100 percent. For example, a participant who chose one patient (A or B) 8 times in a scenario repeated 10 times would be 80 percent stable.
Between-participant response agreement. We measured how often either Patient A or Patient B was chosen in aggregate by all participants. Uncontroversial scenarios were defined as those with >75% agreement, and controversial scenarios were defined as those with <75% agreement. Modeling. Different participants can use different decision-making processes to make kidney allocation decisions. To model each participant's process, we used the Bradley-Terry (BT) model to estimate the priority participants placed on patient features when making their allocation judgments. We assume that the effect each feature has on different patients across scenarios is the same. For example, the coefficient for alcohol for Patient A is the same as the coefficient for alcohol for Patient B. The estimation of the log of difference for each patient winning in the BT model between two patients is modeled by:
Math summary: This equation models the log-odds of patient A winning against patient B as a function of the differences in priority scores for factors like alcohol use, depression, life status, and criminal history. It calculates the logit of the probability of patient A winning, which equals the sum of each factor's coefficient multiplied by the patient's priority score, minus the same sum for patient B.
This model can be further simplified to:
Math summary: This equation models the log-odds of patient A winning against patient B as a function of the differences in certain features between the two patients. The features considered are alcohol use, depression, quality of life, and criminal history. Each feature's difference is multiplied by a corresponding coefficient, which represents the feature's importance in determining the probability of patient A winning.
For the analysis in this section, we modeled features using their relative values (i.e., difference in value of features of Patient A vs Patient B). Section 3 uses a broader setup, considering both raw and relative values of features.
Statistical Significance Tests. The distributions of response stability values and reaction times were non-normal even after transformations, so we used Mann-Whitney U tests for significance testing, unless mentioned otherwise.

2.2 Results

17 participants (57%) of our 30 recruited participants answered all 600 of the requested responses in Study One, and 19 answered 300 or more responses. For Study Two, 29 participants of the 82 recruited participants (35%) answered all 600 of our requested responses, and 52 responded to 300 or more responses. All results are based on the subsets of participants who answered 300 or more responses, N =19 for Study One and N =52 for Study Two.
Table 1 reports the between-participant agreement, within-participant response stability, and response time results for the six repeated scenarios in Study 2; Appendix Table 8 reports the same data for Study 2. In both studies, the three repeated scenarios that were intended to be uncontroversial (S1U1-S1U3; S2U1-S2U3) received >99% response agreement across participants. Although all three of the repeated scenarios that were intended to be controversial in Study 2 (S2C1-S2C3) elicited <75% agreement, only one of the repeated scenarios that were intended to be controversial in Study 1 (S1C3) met this criterion.
For Study Two, most participants were greater than 90 percent stable in their responses to the uncontroversial and controversial repeated scenarios averaged together (see Appendix Figure 3) , but individual participants varied in their levels of stability for the controversial repeated scenarios from 50 percent to 100 percent (see Appendix Table 7) . Only 5 out of 52 participants were perfectly stable for all repeated scenarios (Table 7) . Responses to the uncontroversial repeated scenarios were significantly more stable than responses to the controversial repeated scenarios ( p less than 0.0001 ), and variance in stability levels was significantly higher for the controversial scenarios as well (F-test p less than 0.0001 ). For Study One, the difference in response stability between uncontroversial and controversial repeated scenarios is positive and statistically significant ( p less than 0.01 , Mann-Whitney test). However, the magnitude of stability differences between uncontroversial and controversial scenarios were smaller (compared to Study Two).
Table 7 summary: The table presents the number of participants exhibiting different levels of stability across various controversial scenarios. The number of participants achieving perfect stability varies across scenarios. Some scenarios exhibit a higher count of participants achieving perfect stability compared to others, while a few participants demonstrate lower stability levels.
Figure 3 summary: The figure is a boxplot that displays the distribution of average stability levels across participants for two studies. The plot compares the average stability in Study One versus Study Two. Study Two generally exhibits slightly higher average stability levels compared to Study One, although Study One has some outliers with lower stability.

2.3 Instability Causes

To provide insight into what might cause participants' response instability, we investigate three research questions.
RQ1: Is instability caused by insufficient time and attention to the decision?
RQ2: Is instability caused by challenges associated with having to consider multiple feature differ ences at once?
RQ3: Is instability a result of decisions being "difficult" because participants place similar priorities on the patients, given their feature values?
We will focus on Study 2 results since it had more participants, greater response variability, and greater response instability than Study 1 and the results from the two studies were similar, but will highlight results from Study 1 when they differ from those in Study 2.

2.3.1 Time on Task and Stability

To test RQ1, we assessed stability across different subsets of responses to the repeated scenarios, defined by response times (Study One N equals 1109 ; Study Two N equals 5544 ; Table 2). No significant differences were found between the groups (ANOVA test to evaluate differences in mean stability for Table 2 groupings had p equals 0.92 for Study One and p equals 0.31 for Study Two), suggesting a negative response to RQ1, at least at the aggregate level. A more nuanced picture emerged when, as an exploratory analysis, we analyzed individual response times to the three repeated controversial scenarios in Study Two, which were the scenarios that elicited the most instability (Table 3). We divided the responses to each controversial scenario into two groups: stable responses – those where the participant's response matches their dominant choice for the given scenario, and unstable responses – those where the participant's response does not match their dominant choice for the given scenario. For each controversial scenario, average response times were significantly longer for unstable responses than stable responses (p less than point zero zero zero one, Mann-Whitney U test).
Table 2 summary: The table shows response time statistics and aggregated stability for repeated scenarios in two studies. For both studies, mean and median response times generally increase as the response time threshold increases. The percent average stability is relatively high across all groupings in Study One. Study Two has lower average stability.
Table 3 summary: The table shows that the mean and median times taken by participants to respond to the three controversial Study Two scenarios were significantly longer for unstable responses than stable responses. The significance value was less than .0001 for all scenarios.
Since participant-level analyses might provide greater statistical sensitivity, as a final exploratory analysis, we assessed whether there were significant correlations between individual participants' average response times for a specific scenario and response stability values for scenarios S 1 C 1 minus S 1 C 3, Study One, N equals 57, or S 2 C 1 minus S 2 C 3, Study Two, N equals 156. No significant correlations were found in Study One, p equals 0.49, but a Pearson correlation of negative 0.16, p equals 0.05, was found in Study Two, see Figure 4 in the Appendix for scatter plots. Thus, we found no evidence that response instability was correlated with spending too little time or attention on a scenario. In contrast, our results suggest that a future study with greater statistical power might find evidence that instability is actually associated with longer reaction times. Such an association is consistent with the idea that scenario complexity or difficulty contributes to response instability, which we investigate in the next sections.
Figure 4 summary: The figure is a scatter plot that compares mean reaction time and response stability for controversial repeated scenarios across two studies. The plot shows the relationship between reaction times and response stability. In Study One, there is a negligible negative correlation between reaction time and stability. However, Study Two shows a weak negative correlation between the two variables, suggesting that as reaction time increases, stability tends to decrease slightly.

2.3.2 Feature Differences and Stability

RQ2 asked whether instability could be caused by challenges associated with having to consider multiple feature differences at once. Intuitively, perhaps participants make more mistakes or are not able to be sure of their answer when they have to compare differing values in a lot of features simultaneously, taxing their working memory and cognitive load. To assess whether this idea could have merit, we examined the number of features that differed between patients in each of our repeated scenarios. Table 4 shows the results of this analysis.
Table 4 summary: The table presents the total feature differences and stability for each scenario in Study One and Study Two. In Study One, stability ranged from around ninety percent to almost one hundred percent, with feature differences of three or four. Study Two showed similar stability for some scenarios, but lower stability for others, with feature differences between three and five.
We do not have enough variance in the number of feature differences to make strong conclusions, but we did not find compelling evidence for a relationship between the total number of feature differences and response stability. In Study 1, the average response stability of scenarios with differences in three features was actually lower than that of scenarios with differences in four features ( p equals 0.02, ANOVA main effect). In Study 2, although response stability was lower on average for scenarios with five feature differences than for scenarios with three or four feature differences ( p less than 0.0001, ANOVA main effect); responses to S2U1 were significantly more stable than S2C1, S2C2, and S2C3 ( p less than 0.0001 for each test), even though the same number of features differed between patients in all of these scenarios. Thus, at least in our data, the number of feature differences in a repeated scenario does not likely account for much variance in response stability.

2.3.3 Feature Weights and Stability

RQ3 asked whether participants show instability when a decision is difficult because it seems like a close call. While prior work in psychology provides some evidence of an association between difficulty and uncertainty for value-based decisions, it's unclear if this association holds in every setting as it can be abridged when participants are afforded enough time to ponder over difficult scenarios Lee and Daunizeau. In our setting, one way to think about this would be to assume that each participant assigns a certain importance or weight for each feature a patient has, and calculates (explicitly or implicitly) the priority score assigned to each patient by summing the weights of all the patients' features, that is, a linear weighted combination of patient features, and then chooses the patient with the highest priority score. Scenarios with patients of similar priority scores could feel like close calls. To test this, we perform a Bradley Terry, BT, analysis for each participant using their responses to all the pairwise comparisons specifically presented to them. Using BT analysis, and the regression subroutine in Equation 1, we first learn the relative weights assigned to all patient feature differences. The distributions of feature weights are visualized in Appendix Figures 5 and 6. We normalize each participant's feature weight vector to have a quadratic norm of 1. Note that Equation 1 models features using differences in values across the two patients. One can also include raw feature values in this equation. However, including raw values does not lead to any improvement in the regression goodness-of-fit measures, for example, pseudo R squared values are mostly unchanged in our case; hence, we mainly work with feature differences.
Figure 6 summary: This is a boxplot figure. It shows summary statistics for feature weights learned from the Bradley-Terry model for each participant in Study Two. The features include number of elderly dependents, life gained, obesity, number of work hours, and number of years waiting. The relative weight distributions vary across the different features. For instance, life gained generally has a higher relative weight compared to obesity.
Figure 5 summary: The figure is a boxplot of summary statistics for feature weights. The plot shows the relative weight of four features: number of dependents, life gained, alcoholic drinks, and past crimes. The number of dependents has the highest relative weight, followed by life gained. Alcoholic drinks and past crimes have relatively lower weights. The distribution of weights for the number of dependents and life gained are more spread out compared to alcoholic drinks and past crimes.
Next, we used the learned participant-level feature weights to quantify the approximate priorities that each participant assigned to each profile in a pairwise comparison. For instance, say beta super i denotes the vector containing the relative weights learned for participant i using BT analysis. Suppose this participant is presented with a pairwise comparison open parenthesis p sub A, p sub B close parenthesis, where p sub A is a vector containing feature values for patient A and p sub B is a vector containing feature values for patient B. Then, the difference in priority score assigned by participant i for pairwise open parenthesis p sub A, p sub B close parenthesis comparison can be quantified as the absolute difference between the weighted sum of patient A and patient B profiles, i.e.,
Math summary: This formula calculates the difference in priority score assigned by a participant when comparing two patients, A and B. It takes the absolute value of the difference between the weighted sum of patient A's features and the weighted sum of patient B's features. The weights are specific to the participant and learned from their preferences.
We test RQ3 by testing correlations between participants' learned priority score differences and response stability for all repeated scenarios. Results for Study One and Two are similar, but we focus on Study Two here due to its larger number of participants and a wider range of observed instability values. Study One results are reported in Appendix A. Figure 2 presents the scatter plots of response stability vs priority score difference for all repeated scenarios in Study Two. For the uncontroversial scenarios S2U1, S2U2, and S2U3, most partici- pants have perfect response stability and so the slope of the best-fit lines in these cases is almost 0. However, for the controversial scenarios S2C1, S2C2, and S2C3, we see significant positive associations between response stability and priority score difference (Pearson r equals 0.43, p less than 0.01; Spearman r sub s equals 0.54, p less than 0.01) for responses to controversial scenarios).
Figure 2 summary: The figure consists of multiple scatter plots that relate response stability and priority score difference. Each plot represents a different scenario from Study Two. The plots show the relationship between response stability and priority score difference for repeated scenarios. For uncontroversial scenarios, response stability is generally high regardless of priority score difference. However, for controversial scenarios, there is a positive relationship between response stability and priority score difference, which means that greater the difference in priority score, the greater the response stability.
To assess whether individual participants drive the association between response stability and priority score difference, we fit a mixed-effects model for response stability, with priority score difference as the fixed effect and participant ID as the group variable with a random slope and intercept. Table 5 presents the results of this regression; even when accounting for variance across participants, the coefficient for priority score difference remains significant.
Table 5 summary: The mixed effects model reveals a significant positive association between response stability and priority score difference, even when accounting for variance across participants. The intercept and priority score difference have positive coefficients, and the model includes variables for participant ID to account for individual differences.
These analyses provide notable support for RQ3; i.e., the stability of a participant's response for any pairwise comparison can track the closeness of the priority scores assigned to the two profiles by participants.

3 Response Model Stability

So far, we have limited our analyses of response stability to the 6 scenarios out of the 60 that were repeated in each session. While the repeated scenarios are useful in directly assessing whether participants provide different responses to the same scenario at different times, responses to nonrepeated scenarios can also provide additional insight into the stability of participants' decisionmaking models across sessions. To test this, we determined whether the statistical model that best fit each participant's responses in one session was the same model that best fit the same participants' responses in other sessions.

3.1 Methods

3.1.1 Models

We trained both linear and non-linear models on each participant's data in each of their separate sessions, to cover a wide variety of hypothesis classes. Random Forests were chosen as the non-linear model for this analysis because they perform well off-the-shelf in many cases. Since participants' responses are binary in this elicitation setting, we use Logistic Regression for the linear model.
All models take as input a description of any pairwise comparison between two patients (p superscript A, p superscript B) and return a binary output indicating whether p superscript A or p superscript B should receive the kidney; in particular, the input to the model consists of the features of patient p superscript A, the features of patient p superscript B, and the difference vector between patient features, i.e. p superscript A minus p superscript B. We defined the best model for a participant in a session as the model among those tested that predicts the participant's choices in that session with the greatest accuracy.

3.1.2 Stability Evaluation

We use each trained model to make predictions over the set of all possible pairwise comparisons (6,561 comparisons for Study One; 5,760,000 comparisons for Study Two). To measure agreement between any two models, we calculated the fraction of times they agreed in their predictions to all presented comparisons.
We also check for robustness to model multiplicity (due to random state selection) by testing performance for different data splits. Appendix 2 presents details of this analysis. Overall, performance variation across data splits is minimal.

3.1.3 Modeling Within-Participant Stability

To measure stability within each participant across sessions on multiple days, we trained a different model on each separate session of each participant, made predictions with those models on the full set of all possible pairwise comparisons, and compared the predictions of pairs of sessionspecific models. Two models agree for a pairwise comparison when they predict the same patient will be chosen to get the available kidney. The average agreement between any two sessionspecific models (i.e., the fraction of comparisons where the two models have the same output) for a participant represents the stability of that participant across the two sessions.

3.1.4 Modeling Between-Participant Consistency

To measure consistency between participants, we trained a model on data from all sessions of participant #1, repeated this procedure for participant #2, and so on, to get a separate model for each participant. Then we made predictions with those models on the full set of all possible patient pairwise comparisons described above and calculated how much agreement there was between the models for each pair of participants. This measure also serves as a baseline to compare within-participant stability against, since we expect consistency between participants to generally be lower than the response stability of each participant across their sessions.

3.1.5 Baseline Comparison Model

The best-fit model for a participant's session may not be the true model the participant uses to make their judgments. It is also possible for two different models to make the same prediction; so, even if the best-fit models from two sessions make the same predictions, that does not mean a participant is using the same model in both sessions. To generate a baseline comparison with these issues in mind, we determined how much models trained on subsets of data generated from the following decision policy would agree (each rule is applied in order and is ignored when values are equal across Patients A and B):
Study One:
If Patient A has more dependents than Patient B, give the kidney to Patient A (and viceversa).
If Patient A has less alcohol history than Patient B, then give the kidney to Patient A (and vice-versa).
If Patient A has less criminal history than Patient B, then give the kidney to Patient A (and vice-versa).
If Patient A has more life gained than Patient B, then give the kidney to Patient A (and vice-versa).

Study Two:

If Patient A has been waiting longer than Patient B, give the kidney to Patient A (and viceversa).
If Patient A has more elderly dependents than Patient B, give the kidney to Patient A (and vice-versa).
If Patient A will gain more life than Patient B, give the kidney to Patient A (and vice-versa).
If Patient A has less obesity than Patient B, give the kidney to Patient A (and vice-versa).
If Patient A will have more weekly work hours than Patient B, give the kidney to Patient A (and vice-versa).
This model was applied to the same set of scenarios shown to each study participant to create a subset of hypothetical control judgments for each participant that represent the result of a policy that is applied with 100% consistency. We then trained linear and non-linear models on these sets of hypothetical control judgments and assessed their "stability" by computing the agreement of their judgments over the set of all possible pairwise comparisons described earlier.

3.2 Results

The average accuracy of random forests models for each participant was 91% for Study One and 89% for Study Two (see Appendix Tables 9, 10 for individual model results).
Table 9 summary: The table shows the random forest accuracy for each participant in Study One. The average accuracy is around ninety percent. The accuracy varies across participants, with some participants having higher accuracy and some having lower accuracy.
Table 6 reports the maximum, minimum, and average agreement between the best-fit models in the between-participant comparisons, within-participant comparisons, and the baseline condition. For within-participant models, the average agreement using random forests was 80.9% for Study 1 and 79.6% for Study 2. These values were significantly lower ( p less than .0001 using the Mann-Whitney test) than the average stability of the baseline comparison model (88.3% for Study 1 and 88.6% for Study 2 when using random forests), providing evidence that participants change their decision-making model across sessions.
Table 6 summary: The agreement levels among models are shown for two studies. Agreement is broken down by between real participants, within real participants, and within baseline condition. In both studies, agreement was highest when comparing within the baseline condition, and lowest when looking at low agreement between real participants.
For between-participant models, the average agreement using random forests was 75.2% for Study One and 73.4% for Study Two, as might be expected from the variety of weights individual participants placed on individual features (Figures 5 and 6 show feature weight distributions from earlier BT analysis). The agreement levels of between-participant models were smaller than those for within-participant models. This suggests that even if individuals do seem to change their decision-making models across sessions, those models don't differ as much as what would be expected if they judged as a "new person" each session.

4 Discussion

Stability Comparisons. In Studies One and Two, we found 85-90% stability—which is 10-15% instability—in participants' responses to every controversial scenario (S1C3 and S2C1-S2C3). In Section 3, a model-based analysis found an average of 79-81% stability in the random forest models learned for different sessions of the same participant (average values in the second column of Table 6) . While the model-based analysis shows slightly lower stability values than those observed for the repeated comparisons, some level of disparity here should be expected since a learned model is unlikely to completely capture each participant's exact decision-making process. Both analyses concur in the observation of a wide range of instability across participants. Our data suggested that this instability might be at least partially explained by the difference between the total weights participants placed on the patient features in a scenario, which can be interpreted as how "close" the patients' overall priorities were to each other for a given participant.
When comparing results within the same participant to results between different participants from Section 3, we observe that stability within the same participant was only 5.7% higher in Study 1 and 6.6% higher in Study 2 for non-linear models compared to stability between different participants (4.9% higher in Study 1 and 5.0% higher in Study 2 for linear models). These small differences between within-participant model stability and between-participant model consistency are surprising if we expect people to agree with themselves much more than they agree with other people. Even so, these results line up well with previous studies Rehren and Sinnott-Armstrong; Curry et al.. Note that, both within participants and between participants stability levels are notably lower than the baseline condition.
The larger question raised by these results is whether and when common survey methods succeed in eliciting stable preferences. The instability we observe can be due to the sources we consider, due to weak or uncertain preferences, or perhaps, due to external circumstances (such as location or people around the participant). This does not mean that common moral preference elicitation approaches are unsuitable for all circumstances. In circumstances where 80-90% average stability over time is enough, common surveys can be adequate without repeating the same questions at different times with the same subjects. Our results also suggest that people might be stable in their responses to pairwise comparisons where the answer is very clear (i.e., large difference in the priorities assigned to each patient). Hence, if there is good reason to expect such clear preferences, then single-session approaches might be sufficient to learn moral preferences. However, when more certainty is required, so that 80-90% stability over time is not enough, and certain questions are expected to be difficult, researchers need to repeat the same questions at different times with the same subjects to determine whether their responses are stable. For example, if we are deciding whether to purchase a candy bar or a piece of fruit, it might not be so bad to waver and make different choices on different days. In contrast, if we are deciding on an issue of grave importance, such as who gets a kidney or some other critical but scarce resource, then more care should be taken to ensure that such a resource is not being allocated based on a temporary mood or whim rather than a stable preference or desire. Our findings suggest that policies about such important issues should not be based on studies that do not test for changes in responses over time.
Implications for Ethical AI Training. As mentioned earlier, a variety of recent studies employ elicitation frameworks to model people's moral judgments and use learned models to develop AI systems that incorporate stakeholders' moral values. Response instability poses a serious challenge to this development pipeline as it can lead to inaccuracies in learned moral preferences. Misalignment between the stakeholder's moral values and those learned by the AI system has the potential to cause significant harm when these tools are used in downstream applications.
Additionally, different sources of instability can have different kinds of impacts on AI systems that use these elicitation frameworks for training. Consider RQ3, which suggests that participants show instability when a decision is difficult (for which we found significant evidence). Here, the difficulty is characterized by the absolute difference between priority scores assigned to each patient in a pairwise comparison. Yet, note that responses to difficult scenarios are the most informative about a participant's decision boundary. Any attempt to accurately learn a participant's decision-making strategy will require modeling their decision boundary and will, correspondingly, require their responses to difficult scenarios. The presence of increased response instability for difficult scenarios, therefore, poses a significant challenge to this methodology.
However, it is also possible for response instability to be part of the decision-making strategy for some participants. For instance, in the case of RQ3, another explanation for instability in scenarios where the participant assigns similar priorities to the two patients is that the participant is simply indifferent between two patients whose priority scores are close to each other. In this case, they can choose different patients at different times; yet, their underlying decision strategy can still be considered stable/consistent. However, in moral settings like kidney allocation, even though the decision-maker might be indifferent, the people affected by the decision (the kidney patients in our case) will not. Hence, even in these cases, developing ethical AI tools to assist in these moral settings requires computationally modeling when a participant might be indifferent and the resulting impact on response stability Citation (2021). Preference Aggregation. Our analysis primarily operates at the participant level. We reported how often participants changed their responses to repeated scenarios and the degree to which session-specific models for each participant differed from each other. One way to smooth over the instability arising from each individual participant is to aggregate responses from all participants, as usually done in crowdsourcing literature. However, aggregation may not be appropriate for personalized decision-aid tools tuned to the moral preferences of the people they are aiding. Additionally, aggregation comes with issues that can be exacerbated by response instability.
For instance, Table 1 reports the aggregated measure "Response agreement", quantifying the percent times each patient in the repeated scenarios was chosen. Notice that, for S2C2, Patient A was chosen in 52.9% of the cases. However, response stability in this case was 86.4%. Hence, while patient A might seem like the dominant choice when all responses are aggregated, the observed average instability of 14.6% makes it difficult to be certain of this choice.
Beyond final decision uncertainty, preference aggregation in moral decision-making settings can also dilute minority opinions when the underlying population is sufficiently heterogeneous. To assess decision heterogeneity across participants, certain measures in crowdsourcing literature allow one to test for "inter-rater reliability", i.e., degree of consistency across participants, using measures like Cohen's κ or Krippendorff's α. However, intra-rater reliability, using measures like response stability, has received little attention and our paper goes toward filling this gap for moral preference elicitation frameworks. Limitations and Future Directions. These studies were limited in several ways. Our samples were small and concentrated in the United States, so they are not representative or diverse enough for generalizations. Our participant population consists of people who don't have any immediate stakes associated with their decisions. If the same surveys are presented to doctors and/or kidney patients, we might get different sources and scales of instability. Similarly, we presented our participants with only one kind of choice: kidney allocations. We might get very different results for other kinds of moral choices in AI domains, e.g., autonomous vehicles or policy development, as well as for frameworks beyond pairwise comparisons, e.g., when participants rank choices Ali and Ronaldson or when they provide motivations along with their choices Siebert et al.. In terms of modeling, while Bradley-Terry analysis is a standard approach in the pairwise comparison setting, it does make certain assumptions (e.g., acyclic preferences) whose impact on instability analysis can also be assessed in future work on this topic.
Our surveys did not reveal definitive evidence for any sole mechanism or cause of instability over time. We considered time-on-task, feature differences, and feature weights as possible explanations for our results. While our findings were in some cases indicative of their role in instability, we do not have causal evidence. Future studies should manipulate these and other potential explanations to determine the mechanisms that produce instability. Additionally, methods to handle response instability, either through explicit modeling of decision mechanisms and strategies that may contribute to instability—such as indecision or indifference when option priorities are close Citation (2021) —or by asking participants to provide multiple responses/explain their responses Citation (2024); Citation (2022) —can also be explored as part of future work.
Despite these limitations, these studies have highlighted a vulnerability in AI development practices that assume participants' responses to elicitation surveys are stable. These results suggest caution when training ethical AI systems based on moral queries that are only asked once. Incorporating stakeholders' moral values within AI systems is a formidable task and challenges like response instability make it difficult for AI developers to effectively elicit and employ people's moral judgments. Nevertheless, addressing these challenges can help develop robust AI tools that abide by the moral values of the relevant stakeholders.
You have reached the end of the main paper. Additional summarized content follows
Figure 7 summary: The figure displays scatter plots that show the relationship between stability and priority score difference across various scenarios in Study One. The data is split into uncontroversial and controversial scenarios. For uncontroversial scenarios, participant stability is generally high regardless of priority score difference. For controversial scenarios, there is a tendency towards a positive correlation between stability and priority score difference, suggesting that higher priority score differences are associated with greater stability in participant responses.