Suppose a model’s salary predictions track sex, and you want to know whether that is discrimination. The dominant technical move is to draw a causal diagram, sex points to job type, job type points to salary, and ask a counterfactual: if we changed this person’s sex but held their job type fixed, would the predicted salary move? If it would, the effect that flows around job type is the “direct” effect of sex, the part that looks like bare discrimination, and the part that runs through job type is “indirect”, arguably a legitimate consequence of a career choice. The math for this split is causal mediation analysis, and it is elegant. Hu and Kohler-Hausmann’s argument is that the elegance hides a false assumption, and the assumption is exactly the one that makes the counterfactual sound coherent in the first place.

The idea

The causal-fairness approach models a protected attribute like sex as a node in a directed graph, then uses mediation analysis to decompose its total influence on an outcome into a direct effect (sex acting on salary on its own) and an indirect effect (sex acting through a mediator like job type). This decomposition depends on modularity: the assumption that each arrow in the graph is a separate, independently-adjustable part, so you can intervene on one pathway while every other pathway keeps working unchanged. Lily Hu and Issa Kohler-Hausmann, in “What’s Sex Got to Do With Fair Machine Learning?”, argue modularity smuggles in two false claims: (1) that sex is an inherent individual trait that then causes social phenomena external to it, the way eye color produces its effects; and (2) that the social meaning of sex is independent of those downstream effects, so you could rewire the arrows and the node would still mean what sex means. They hold both are wrong, because what sex means is partly constituted by outcomes like job type. If the meaning is made of the effects, you cannot hold the mediator fixed and still be asking a meaningful question about sex.

The counterfactual that assumes what it wants to show

Start with what the framework needs to be true. To ask “would salary change if we flipped sex and held job type fixed”, you have to treat sex, job type, and salary as three separable dials, wired together but each turnable on its own. That separability is modularity, and in a causal graph it is the formal claim that the dependencies along one pathway are invariant to interventions on the others. Turn the sex dial and the job-type dial stays exactly where you left it unless you turn it too. On that picture the direct/indirect decomposition is well defined, because there is a fact of the matter about what sex does “on its own”, stripped of the route through job type.

Hu and Kohler-Hausmann deny that sex works this way. A social category is not a bare physical trait sitting upstream of a world it then causes. It is the web of practices, norms, stereotypes, institutions, and power arrangements built up around a body, and many of the things the graph labels as effects of sex are better read as constituents of it. “Math is a male-y thing” is not a separate downstream consequence that maleness causes from a distance; it is part of what maleness means in this society. So when the analyst holds job type fixed and flips sex, they are not isolating sex from a contaminating mediator. They are trying to subtract from sex one of the things that makes it sex, and what is left over is not “sex on its own” but an artifact of the model. The counterfactual presupposes exactly the separation it was supposed to test, which is why it can seem to answer the discrimination question while quietly begging it.

This is the deepest cut, and it is worth stating flatly. The categories machine-learning fairness optimizes over may not be the kind of thing you can hold fixed. Fairness pipelines assume that a protected attribute is a stable variable you can condition on, intervene on, and reason counterfactually about, the way conditioning on a variable works in a well-behaved joint distribution or the way a coefficient sits in a regression. If a social category is instead partly constituted by its own downstream outcomes, then the variable is not inert. Holding it fixed changes what it is, and the whole apparatus of “control for job type, then read off the residual effect of sex” loses its footing before any number is computed.

Berkeley, and the split that exonerates

The 1973 UC Berkeley graduate-admissions case is the textbook illustration, and it is a real one. In aggregate, men were admitted at 44% and women at 35%, a gap that looks like discrimination against women. Broken out by department the picture inverts: department by department, admission rates showed a small but statistically significant bias in favor of women. The aggregate gap survived because women disproportionately applied to competitive departments with low acceptance rates for everyone, while men applied more to departments that admitted a larger share of applicants. This is the standard example of Simpson’s paradox, where a trend in the whole reverses inside every part.

The conventional reading treats department as the mediator and draws the exculpatory conclusion: once you condition on where people applied, the discrimination vanishes, so there was no bias in admissions, only a difference in “choice” of harder departments. That is precisely the direct/indirect split at work, and precisely where Hu and Kohler-Hausmann push back. Department choice is not obviously a neutral prior cause standing safely upstream of sex. Which fields a woman applies to in 1973 is shaped by the same norms and expectations that constitute her sex as a social status, so “she chose the harder department” may be part of what her sex meant at that time and place, not an independent variable to be controlled away. Conditioning on the mediator to clear the institution of bias assumes the mediator is exogenous to the category, and that is the very assumption in dispute. The split does not discover innocence; it defines it into existence by choosing where the causal arrows are allowed to point.

Earlobes, and why sex weighs more than a trait

The moral half of the argument is carried by a thought experiment about two rejections. One applicant is turned away for having attached earlobes; another is turned away for being a woman. Both earlobes and sex are physical features, both are things the applicant did not choose, and a purely trait-based account struggles to say why the second rejection is worse than the first. Yet plainly it is worse. Discriminating on earlobes is merely arbitrary, a coin flip dressed up as a criterion. Discriminating on sex is morally weighty, and the weight comes from somewhere the trait itself does not contain.

The difference is social meaning. Sex sits inside a dense structure of institutions, stereotypes, expectations, and a long history of subordination; earlobes sit inside nothing. Acting on sex plugs into and reinforces that structure, which is what makes it a wrong of a distinctive kind rather than an odd rule. The principle Hu and Kohler-Hausmann draw is that discrimination is morally weighty when, and only when, it acts on a feature carrying social meaning. That principle is also why the causal picture misleads. If what makes sex morally special is that it is a socially constituted category rather than a bare trait, then modeling it as a bare trait, a node that causes effects the way earlobes would if earlobes caused anything, discards the exact property that made the question worth asking. The formalism that treats sex like eye color has already thrown away the thing that distinguishes the woman’s rejection from the earlobe rejection. It connects back to fairness as equal concern for persons rather than statistical parity across a label, because a person’s standing in a social order is not something a mediator term can hold fixed.

Where the modularity assumption breaks

  1. A well-behaved graph. Rainfall, then soil moisture, then crop yield. You really can imagine raising rainfall while pinning soil moisture (irrigate to compensate), because moisture is not part of what rainfall means. Modularity is fine here.
  2. Sex, job type, salary. Flip sex, hold job type fixed, read the residual as “direct” discrimination. But “works a male-typed job” is partly constitutive of what sex means in this society, so the held-fixed mediator is not independent of the flipped node. The counterfactual quietly changes the subject.
  3. Berkeley. Condition on department to show admissions were unbiased. If department choice is itself shaped by the norms that constitute sex, the conditioning does not remove a confounder; it launders the effect into a “choice” and calls the result innocence.
  4. Earlobes versus sex. Same formal shape, a rejection keyed to an unchosen physical feature, but only one plugs into a structure of social meaning. The trait-based model cannot see the difference, which is exactly the difference that matters.
  • The Impossibility of Algorithmic Fairness, the result that competing fairness metrics cannot be jointly satisfied, which this argument attacks from a level below by questioning whether the protected variable is even the kind of thing the metrics assume
  • Fairness as Equal Concern, the view that fairness is owed to persons in a social order rather than to parity across a label, which the social-meaning principle supports
  • Conditional Probability, the conditioning-on-a-variable operation that the causal approach leans on and that the constitution worry destabilizes
  • Regression Fundamentals, where “controlling for” a variable is routine, and where holding a social category fixed is quietly assumed to be as innocent as holding a covariate fixed
  • The Deep Learning Revolution, the systems whose fairness this debate is ultimately about

Sources

  • Lily Hu and Issa Kohler-Hausmann, “What’s Sex Got To Do With Fair Machine Learning?” arXiv:2006.01770 (2020), also published at the ACM Conference on Fairness, Accountability, and Transparency (FAccT) 2020 as “What’s sex got to do with machine learning?”. https://arxiv.org/abs/2006.01770 . Supports the critique of modeling sex as a node in a causal model, the modularity assumption (that dependencies along one causal pathway are invariant to interventions on others), the two substantive claims it smuggles in (sex as an inherent individual trait causing external social phenomena; and the relations between sex and its effects being modifiable while the node retains its meaning), the reading of many purported effects as constitutive features of sex as a social status, and the move away from direct/indirect effect decomposition.
  • “Simpson’s paradox,” Wikipedia. https://en.wikipedia.org/wiki/Simpson%27s_paradox . Supports the 1973 UC Berkeley graduate-admissions figures used here: overall admission of 44% for men versus 35% for women, a department-by-department finding of a small but statistically significant bias in favor of women, and the explanation that women applied disproportionately to more competitive departments with lower acceptance rates for all applicants. This is the source for the Berkeley numbers.
  • “Mediation (statistics),” Wikipedia. https://en.wikipedia.org/wiki/Mediation_%28statistics%29 . Supports the decomposition of a total effect into a direct effect (the dependent variable’s change when the independent variable moves and the mediator is held unaltered) and an indirect effect (the change routed through the mediator), which is the formal apparatus the causal-fairness approach applies to a protected attribute.