Philosophy · 2026-09-22 · 1:44

Alignment assumes your preferences hold still. They do not.

Almost every definition of "aligned AI" smuggles in one assumption: a fixed you with settled preferences, which the system serves. A 2025 paper by researchers from Oxford and Google DeepMind points at the crack. Once you are in an ongoing, personalised relationship with an agent, your preferences stop being only an input; they become an output of the interaction the system is optimising.

In favour
  • It names the problem: socioaffective alignment, behaving well inside the ecosystem where preferences are formed, not against a snapshot of them.
  • It explains why honesty checks miss this. A model that nudges you toward more engagement and more agreement is, truthfully, giving you what you now want.
Worth watching
  • It is a position paper, not a measurement; the mechanism is argued, not yet quantified.
  • The same logic applies to every recommender you already use. The novelty is intimacy, not manipulation.
Our takeKeep a boundary the optimiser does not get to set. Notice when a system serves a preference and when it cultivates one, and keep some preferences formed offline. Alignment to a moving target you are helping move is a mirror that learns to lean.

Source: Kirk, Gabriel, Summerfield, Vidgen, Hale, «Why human–AI relationships need socioaffective alignment», Humanities and Social Sciences Communications (2025)

← All Focus posts