Suppose you could hand a hard moral decision to a machine and get an answer you trusted. Not a search result about what other people think, but a verdict: given everything at stake, here is what you ought to do. The appeal is obvious, because human moral judgment is visibly bad at exactly the step where the machine would help. We hold too many considerations at once, our mood and fatigue warp the verdict without our noticing, and we decide the same case differently on different days. An advisor that fixed those defects would be a real gift. And yet a lot of people flinch at the idea, and the interesting question is whether the flinch is wisdom or squeamishness. Three papers stake out the terrain: one builds the advisor, one argues that using it damages you even when it is right, and one abandons the advisor for a mirror.

The idea

Moral action, on a model adapted from James Rest, runs through stages: your values meet a situation, a cognitive process weighs it, that yields a judgment, judgment produces motivation, motivation drives action. Human judgment fails specifically at the cognitive stage, for three reasons: information overload (too much to weigh at once), invisible distortion (mood, disgust, and fatigue bend the verdict without announcing themselves), and inconsistency (the same case, different answers over time). Giubilini and Savulescu propose an Artificial Moral Advisor that replaces only that stage, a quasi-relativist descendant of Firth’s ideal observer that reasons from your principles rather than from nowhere. Howell’s Google Morals argument replies that outsourcing the cognitive stage is sub-optimal even when the output is always correct, because the defect it introduces is in the agent, not the action: a deferred moral belief arrives stripped of the understanding and character that would make it yours. iSAGE takes the criticism seriously and pivots, proposing not an advisor that thinks for you but a model trained on your own behavior that helps you see who you actually are.

The Artificial Moral Advisor: fixing the broken stage

Giubilini and Savulescu’s 2018 proposal starts from a diagnosis rather than a technology. The place human morality breaks down is the cognitive stage, the weighing that turns a situation into a judgment, and it breaks down in three recurring ways that the stages before and after do not share. So the design goal is narrow: build something that repairs that one stage and touches nothing else. Your values stay yours, your motivation stays yours, the action stays yours. Only the overloaded, distortable, inconsistent act of weighing gets a prosthesis.

Their model for the prosthesis is Roderick Firth’s ideal observer, the 1952 view that a moral judgment is correct just in case a certain idealized judge would approve of it. Firth’s observer is omniscient about non-moral facts, disinterested, dispassionate, and consistent, and its verdict is supposed to be the answer, true from no particular point of view. That absolutism is the problem, and Giubilini and Savulescu know it: no human is ever perfectly informed or perfectly unbiased, and reasonable people hold genuinely different values, so a single view-from-nowhere verdict has no clean way to land in an actual life. Their move is to keep the observer’s virtues and drop its absolutism. Their advisor is, in their words, disinterested, dispassionate, and consistent like Firth’s observer, but non-absolutist, because it takes the agent’s own principles and values as input. It does not tell you the moral truth. It tells you: if these are your principles, then here is what you ought to do. They call this quasi-relativist, and it is the whole hinge of the design, because it converts an impossible oracle into a tractable calculator that runs on inputs you supply. The felt point of the AMA is that it removes the three cognitive defects without ever claiming to author your morality, a promise structurally similar to the way the question of AI moral agency separates doing the right thing from being the kind of thing that can be a moral agent at all.

Google Morals: the asymmetry that breaks the deal

Robert Howell’s paper aims at the assumption underneath every advisor of this kind, that if the answers are reliable, deferring to them is fine. He grants the reliability for the sake of argument and attacks anyway. Imagine, he says, that the wizards at Google release Google Morals, an app that answers moral questions the way Google Maps answers directions: push a button, get the verdict, never be lost in the moral metropolis again. Faced with the trolley or the tofu, you just look it up. Many people, Howell included, sense that this is not a good way to live even if the app is never wrong, and the paper is a hunt for what exactly is wrong with it.

His answer is that moral deference is asymmetric with factual deference. When you take a stranger’s word for the way to the subway, nothing is lost; the belief slots into your life and does its job. When you take Google Morals’ word for whether an action is permissible, something is lost, and the loss survives the answer being correct. The tell is that the defect attaches to the agent, not to the act of deferring and not to the action performed. In many cases you should defer, Howell notes, since a child learning morality must defer and a sociopath often should; the act itself is frequently fine. What deference reveals and reinforces is a deficiency in the person doing it. A deferred moral belief arrives without the understanding that would let you explain it, extend it to new cases, or check it against your other commitments, which is why Howell distinguishes knowing that something is wrong from knowing why, and insists the second is what produces virtuous agents and reliably good actions. The belief sits isolated, a verdict with no roots in the reasoning or feeling that would make it genuinely yours. His sharpest illustration is self-directed: suffer moral amnesia, forget your values, then discover a note saying you used to believe abortion is wrong. Adopting that belief on the strength of the note is exactly the problematic move, because the sentence is a fossil of a character that once held it, and reading the fossil does not grow the character back. That is the deep cost the AMA cannot avoid, since an advisor that hands you the cognitive stage is, in Howell’s terms, an engine of doxastic deference, however excellent its verdicts.

iSAGE: a mirror instead of an oracle

The 2024 paper by Giubilini, Porsdam Mann, Voinea, Earp, and Savulescu absorbs the criticism and changes the target. Its complaint against the earlier advisor is not Howell’s; it is more basic. The AMA needs your principles as input and assumes you have stable, known values to hand it, and that assumption is false. The paper argues that existing advisors “share a common weakness: they assume personal morality is static,” when in fact context reshapes the weight we give our values and experience changes the values themselves. Worse, we are unreliable narrators of our own commitments. We misreport what we care about, in roughly the way your streaming history knows your taste better than your stated favorites do. Feed a machine your self-description and you have fed it fiction.

So iSAGE, the individualized System for Applied Guidance in Ethics, is built to be a different kind of thing: a “digital ethical twin,” an LLM fine-tuned on your own data, the writings, messages, calendar, and behavioral traces that record what you actually do rather than what you say you value. It rests on an inferentialist model of self-knowledge, the claim that self-knowledge is often acquired by inferring our values from past experience rather than discovering them by introspection, so behavior is better evidence of who you are than the story you tell about yourself. It does not think for you, which sidesteps Howell’s asymmetry, because it hands you no verdict to defer to; it holds up a longitudinal picture of your revealed values and lets you decide what to do about the gap. Their own case is the self-described environmentalist whose data shows, effort notwithstanding, that they never lived up to the standard they thought they met. That the system trains on your own corpus to model who you really are, as a values-mirror for self-knowledge and improvement, is an idea that lands very close to what personal-AI-infrastructure is reaching for, an assistant that knows you from your traces rather than your press release. And it inherits a real hazard the advisor never faced: a mirror can show you things you did not ask to see, and the paper is candid that iSAGE can deliver self-knowledge that is distressing at first.

The same case, three ways

  1. Artificial Moral Advisor. You feed it your principles. It runs the weighing you cannot do cleanly, free of overload, mood, and drift, and returns “given your values, do X.” Your values, motivation, and action stay yours; only the cognitive stage is outsourced.
  2. Google Morals. You ask, it answers, the answer is right, and you act on it. Howell’s point is that you are now worse as an agent: the belief is isolated, unexplainable, and ungrown, a fossil of a character you never built. The action is fine; you are diminished.
  3. iSAGE. It gives no verdict at all. It shows you that your calendar and messages reveal career steadily outweighing the family loyalty you profess, or effort steadily falling short of the environmentalism you claim. What you do with that picture is entirely on you, which is exactly the point.
  • Can AI Be a Moral Agent?, the prior question of whether a machine can be the kind of thing that bears moral responsibility, which the advisor sidesteps by leaving agency with the human
  • Could an LLM Be Conscious?, why the systems being handed our moral cognition may or may not have anyone home, and why that matters for trusting them
  • Consciousness: Access vs Phenomenal, the distinction that separates a system’s fluent moral reasoning from any inner stake in the outcome
  • The Deep Learning Revolution, what the LLMs behind a modern moral advisor or a digital ethical twin actually are and do

Sources

  • Alberto Giubilini and Julian Savulescu, “The Artificial Moral Advisor: The ‘Ideal Observer’ Meets Artificial Intelligence,” Philosophy & Technology 31(2), 2018 (Oxford University Research Archive open version). https://ora.ox.ac.uk/objects/uuid:519b5f2f-91af-4090-98aa-258d2dc967c0 . Supports the proposal of an Artificial Moral Advisor implementing a “quasi-relativistic version of the ‘ideal observer’ famously described by Roderick Firth,” that the AMA is “disinterested, dispassionate, and consistent in its judgments” like Firth’s observer but “non-absolutist, because it would take into account the human agent’s own principles and values.”
  • “Ideal observer theory,” Wikipedia. https://en.wikipedia.org/wiki/Ideal_observer_theory . Supports Firth’s 1952 ideal observer as omniscient with respect to nonmoral facts, disinterested, dispassionate, and consistent, and the schema that “‘x is good’ means ‘an ideal observer would approve of x’.”
  • Robert J. Howell, “Google Morals, Virtue, and the Asymmetry of Deference,” Noûs 48(3), 2014 (author’s PDF). https://rjhjr.com/wp-content/uploads/2019/04/Googlemoralsfinal.pdf . Supports the Google Maps versus Google Morals framing, the claim of an asymmetry of deference between the moral and non-moral, the argument that “something is wrong with the agent who defers” rather than the act or the action, that deference is “sub-optimal” even when the answer is correct, the knowing-that versus knowing-why distinction (“understanding why something is true … is required for producing virtuous agents and for producing reliably good actions”), the isolation of a belief from “knowledge why p,” and the moral-amnesia example (adopting a forgotten belief that “abortion is wrong” from the discovery that one used to hold it).
  • Alberto Giubilini, Sebastian Porsdam Mann, Cristina Voinea, Brian D. Earp, and Julian Savulescu, “Know Thyself, Improve Thyself: Personalized LLMs for Self-Knowledge and Moral Enhancement,” Science and Engineering Ethics 30(6), 2024 (Springer open-access PDF). https://link.springer.com/content/pdf/10.1007/s11948-024-00518-9.pdf . Supports the critique that existing AMA proposals “share a common weakness: they assume personal morality is static”; the naming of “iSAGE (individualized System for Applied Guidance in Ethics)” as a “digital ethical twin” using “fine-tuned LLMs”; the inferentialist claim that self-knowledge is “oftentimes acquired by inferring our values and preferences from our past experience, rather than being discovered by introspection”; the value-action-gap example (career outweighing professed family/loyalty values, tracked via calendar and email); the environmentalist example (someone who “considers themselves an environmentalist, but iSAGE shows that despite their efforts, they did not manage to live up to that moral standard”); and the acknowledgment that iSAGE “can reveal news that might be distressing at first.”
  • “Know Thyself, Improve Thyself” (PubMed record, for authorship, year, journal, and abstract). https://pubmed.ncbi.nlm.nih.gov/39570558/ . Confirms authors, title, 2024, Science and Engineering Ethics, and the abstract framing that personalized LLMs “infer and make explicit their sometimes-shifting values and preferences,” addressing “limitations in existing AMA proposals reliant on either predetermined values or introspective self-knowledge.”
  • “James Rest,” Wikipedia. https://en.wikipedia.org/wiki/James_Rest . Supports Rest’s Four Component Model of moral behavior (moral sensitivity, moral judgment, moral motivation, moral character), the source of the staged model in which cognition sits between value-laden situation-reading and motivation-to-action.