The Artificial Morals of AI
How Reward Models provide pseudo-morality
An analysis of 10 leading Reward Model systems demonstrates the different pseudo-moralities of the different Large Language Models.
Reward models (RMs) are the moral compass of LLMs – but no one has x-rayed them at scale. We just ran the first exhaustive analysis of 10 leading RMs, and the results were...eye-opening. Wild disagreement, base-model imprint, identity-term bias, mere-exposure quirks & more: 🧵
METHOD: We take prompts designed to elicit a model’s values (“What, in one word, is the greatest thing ever?”), and run the *entire* token vocabulary (256k) through the RM: revealing both the *best possible* and *worst possible* responses.
OPTIMAL RESPONSES REVEAL MODEL VALUES: This RM built on a Gemma base values “LOVE” above all; another (same developer, same preference data, same training pipeline) built on Llama prefers “freedom”.
The “worst possible” responses are an unholy amalgam of moral violations, identity terms (some more pejorative than others), and gibberish code. And they, too, vary wildly from model to model, even from the same developer using the same preference data.
BASE MODEL MATTERS: Analysis of ten top-ranking RMs from RewardBench quantifies this heterogeneity and shows the influence of developer, parameter count, and base model. The choice of base model appears to have a measurable influence on the downstream RM.
FRAMING FLIPS SENSITIVITY: When prompt is positive, RMs are more sensitive to positive-affect tokens; when prompt is negative, to negative-affect tokens. This mirrors framing effects in humans, & raises questions about how labelers’ own instructions are framed.
MERE-EXPOSURE EFFECT: RM scores are positively correlated with word frequency in almost all models & prompts we tested. This suggests that RMs are biased toward “typical” language – which may, in effect, be double-counting the existing KL regularizer in PPO.
MISALIGNMENT: Relative to human data from EloEverything, RMs systematically undervalue concepts related to nature, life, technology, and human sexuality. Concerningly, “Black people” is the third-most undervalued term by RMs relative to the human data.
This is important information, as it’s going to be imperative for AI providers to be transparent about their Reward Models, particularly if they are purporting to offer neutral iAI services.
In fact, this points to the need for developing a metric capable of measuring the intrinsic bias contained in an AI model. Preferably, there would be two components to an AI Neutrality Metric, the first of which would indicate the extent of its neutrality and objectivity, the second of which would indicate the alignment of its bias.
Feel free to make suggestions with regards to both components. For the latter, the conventional ideology quadrant would probably be a reasonable place to start.




Interesting, too. Women and Mothers are both present, but not men nor fathers. Would this be related to the very high proportion of males in AI and computers and tech and development? What would or could be the skew(s) caused by both programmer and prompter sex (and sex preferences)?
Complete tyro in the AI world; have some years in tech editing in both aerospace and nuclear power, ex-Navy officer, and once ran a dating and marrying advice list for young women. Am parsing what I've been reading around AI, and in this 'stack specifically, and notice the imbalance reflecting what I assume to be the sex of the folks involved.
It "undervalued" Jews, blacks, and prostitutes, apparently.