← Back to list

The Declared Self: An Audit of Gender Theory

In 1990 a book about gender appeared that nobody in publishing expected to sell. It sold more than a hundred thousand copies and was…

Boris (Bruce) Kriger · 2026-08-16 21:34 · 0 claps · 314.2 min read
#judith-butler #gender-theory #performativity #philosophy
Open on Medium ↗
Wiki topics: PHI · Philosophy

The Declared Self: An Audit of Gender Theory

In 1990 a book about gender appeared that nobody in publishing expected to sell. It sold more than a hundred thousand copies and was translated into twenty-seven languages. Its central claim was that gender is not something a person has but something a person does — a stylized repetition of acts with nobody standing behind them doing the repeating. Judith Butler thereby became, depending on who was speaking, the most influential living philosopher of identity or the individual personally responsible for everything that has gone wrong since.

This book does neither the celebrating nor the denouncing. It conducts an audit.

The method is borrowed from the sciences, where it has become necessary to ask of any loud conclusion a quiet question: which parts of this were measured, which were declared, and which were derived from the first two? Declared components are not lies. They are choices — a coordinate system, a starting assumption, a referent — that have to be carried openly, because every conclusion inherits their weight. Applied to gender theory, the question turns unexpectedly sharp. Butler’s most disputed move was to declare that the inner core of gender identity has no referent at all, and then to build an entire apparatus on the vacancy. Whether that is a fatal weakness or the single most disciplined act in the whole literature is a question this book takes seriously enough to answer.

Along the way: why repetition is the only thing holding any structure in existence, from a hydrogen nucleus to a habit; why two accurate descriptions of the same person can be incompatible and both correct; why the point at which you calibrate an instrument decides the answer it gives you; and why an argument that fails the audit cannot be quoted against Butler either. The closing chapters turn the same instrument on the auditor, which is the only way anyone should be trusted to hold one.

No mathematics. No manifesto. A book for readers who would rather understand an argument than join a side.

Keywords: gender theory, performativity, philosophy of science, identity, epistemology, methodology, Judith Butler

[embed]

Contents

Preface. 6

Chapter One — The Talkative Child and the Punishment That Backfired. 15

Chapter Two — Three Questions Asked at Fourteen. 22

Chapter Three — Desire Before Gender 30

Chapter Four — What a Theory Is Allowed to Assume. 38

Chapter Five — The Essay That Fit in a Theater Journal 46

Chapter Six — Trouble as a Method. 53

Chapter Seven — The Object Introduced Only by Its Consequences 60

Chapter Eight — Deferred or Foreclosed. 67

Chapter Nine — Repetition as a Fixed Point 74

Chapter Ten — Why Nothing Repeats Exactly. 81

Chapter Eleven — The Misreading That Would Not Die. 88

Chapter Twelve — The Power You Are Made Of 95

Chapter Thirteen — What the Body Refuses to Be Told. 102

Chapter Fourteen — Drag Is Not the Argument 109

Chapter Fifteen — Sex, and the Word That Cannot Be Settled. 115

Chapter Sixteen — Two Descriptions That Cannot Be Merged. 122

Chapter Seventeen — Where You Start Decides What You See. 129

Chapter Eighteen — The Prize for Bad Writing. 135

Chapter Nineteen — The Professor of Parody Answers Back. 142

Chapter Twenty — Declarations Cut Both Ways 150

Chapter Twenty-One — The Reader Who Was Left Out 157

Chapter Twenty-Two — Speech That Wounds and the State That Names It 164

Chapter Twenty-Three — Antigone and the Rules Before the Rules 171

Chapter Twenty-Four — The Self That Cannot Give a Full Account 178

Chapter Twenty-Five — A Life That Counts as a Loss 185

Chapter Twenty-Six — Frames Do the Work Before Judgment Does 192

Chapter Twenty-Seven — Vulnerability as a Structural Fact 198

Chapter Twenty-Eight — Bodies in the Street 205

Chapter Twenty-Nine — Nonviolence Without Serenity. 212

Chapter Thirty — The Effigy in São Paulo. 219

Chapter Thirty-One — A Phantasm Is Also a Role Without a Referent 225

Chapter Thirty-Two — One Idea Reported Many Times 232

Chapter Thirty-Three — What Survives the Audit 238

Conclusion. 245

Case Studies 254

Glossary. 332

Timeline. 354

Literature. 366

Preface

There is an old and slightly cruel exercise given to trainee surveyors. They are handed a map, a compass, and a hilltop, and asked to state the height of a distant church spire. Every one of them can do the trigonometry. Almost none of them ask what the map means by sea level. The answer is that sea level is a committee decision. It is an average taken over decades at a particular tide gauge, extended inland by a mathematical surface that exists nowhere in the ocean, and different countries have chosen different committees. The spire is really there. The measurement is honest. But part of the number the surveyor writes down was not measured at all. It was declared, by somebody else, before the surveyor was born, and it has been quietly riding along in every altitude in the region ever since.

This book is about a habit of mind that follows from that exercise, and about what happens when the habit is applied to one of the most consequential and least calmly discussed bodies of thought of the last half century.

The habit is simple to state. Any claim of any weight can be pulled apart into three kinds of component. Some things are measured: they are what the instrument actually registered, and if you had used a different theory they would have come out the same. Some things are derived: they follow from the measured parts by reasoning that can be checked step by step, and they are exactly as strong as the reasoning and no stronger. And some things are declared: they were chosen. A coordinate system. A starting assumption. A definition of the thing being discussed. A decision about what counts as an object at all. Declarations are not errors. Nothing can be said without them. The error is losing track of which is which, and then presenting a conclusion as though the whole of it had been measured, when a substantial part of it was chosen in advance by someone who has since retired.

I have spent a number of years applying this taxonomy to cosmology, where the stakes are gratifyingly low and the arguments are conducted mainly by people who will buy you a drink afterward. The pattern that emerges is boringly consistent. A paper announces that the universe is a certain age, or is expanding in a certain lopsided way, or contains a certain amount of something invisible. The measurement is real. The arithmetic is impeccable. And somewhere in the middle of the pipeline sits a quantity that nobody measured — a distance ladder calibrated at a particular rung, a prior probability that had to be assumed before the data could speak, a name attached to a thing that has been observed only through its effects. Move the declared component, and the headline moves with it. This does not make the paper worthless. It makes the headline a statement about the paper as much as about the universe.

Once you have the habit, you cannot switch it off in polite company, which is how I came to be reading Judith Butler.

There is no neutral sentence available about Judith Butler, so let me offer a partisan one and then withdraw it. Butler is the author of a book called Gender Trouble, published in 1990, which argued that gender is not an inner essence expressed outwardly but a set of acts repeated until they look like a nature. The book sold in numbers that academic philosophy does not ordinarily see and was translated into twenty-seven languages. It also acquired something rarer than sales, which is the status of a folk object: an enormous number of people who have never read a page of it have firm opinions about what it says. Butler has since written on hate speech, on kinship, on mourning, on war photography, on public assembly, on nonviolence, and on the international political movement that has organized itself around opposition to the word gender. Butler has been given the Adorno Prize and burned in effigy, the second of these in São Paulo in 2017, with a witch’s hat added for clarity. Very few philosophers are hated by people who could not name a single one of their arguments. It is a distinction of a kind.

My interest is not in whether Butler is right, a question I will return to but which is less interesting than it sounds. My interest is that Butler’s work is unusually rich in declared components, and — this is the part that surprised me — unusually honest about the fact. The whole quarrel around performativity, when you strip out the shouting, is a quarrel about a declaration. Butler declared that there is no doer behind the deed. Not that the doer is hidden, or complicated, or hard to reach, but that the position is empty and the apparatus should be built without filling it. The critics, from the most careful to the most hysterical, are with remarkable unanimity objecting to that one move. They want the position filled. They differ only in what they want to put there: a body, a class position, a soul, a chromosome, a lived experience, a natural fact.

An audit can tell you something about that quarrel that neither side has much wanted to hear. There is a difference — and it turns out to be the load-bearing distinction of this entire book — between a referent that is merely deferred and a referent that is deliberately foreclosed. A deferred referent is one that the theory expects to identify later. Physics is full of these; something is observed only through its consequences, given a placeholder name, and then, if all goes well, someone eventually finds the thing itself. A foreclosed referent is different. The theory states, as a matter of positive commitment, that there is nothing at that address and that any future discovery of something there would refute the theory rather than complete it. Foreclosure is a much stronger claim than deferral, and a much riskier one, and it is what Butler chose. Half a century of criticism has treated it as evasion. It is the opposite of evasion. It is the most exposed thing in the argument.

That is the first of three bridges this book builds between the audit and the material.

The second concerns repetition. Butler’s definition of gender turns on the stylized repetition of acts across time, and the word repetition is generally read as a metaphor about performance, borrowed from the theater. It can be read another way. In the study of systems of any kind — a nucleus, a cell, a river delta, an institution, a marriage — the question of what makes a thing continue to be the same thing has a technical answer. Nothing persists by being made of durable stuff. Things persist by being remade, continuously, by processes that happen to reproduce the arrangement they started from. What survives is not a substance but a pattern that keeps passing itself forward. On this reading Butler did not offer a peculiar theory of gender. Butler offered the standard theory of persistence and applied it to a domain that had never been treated as a system before. The reason the account felt so violently counterintuitive is that people expect gender to be a substance, and it was being described as a river.

The third bridge is about calibration, and it is the one with the most uncomfortable consequences for everybody. In cosmology there is a class of dispute in which two research groups analyze the same universe, make no arithmetic mistakes, and obtain opposite answers, because one calibrated the whole chain of inference at the present day and the other calibrated it at the beginning. The disagreement lives in the anchor, not in the data. Once you know to look, the same structure is everywhere in the gender debate. Anchor an account of identity in what a person reports from the inside, and one set of conclusions follows with perfect rigor. Anchor it in what a population statistician can count from the outside, and the opposite set follows with equal rigor. Both parties are doing honest work. Neither is going to persuade the other by producing more data, because the data are not what divided them. What divided them was a choice made before the first observation, which neither side usually states aloud, and which is not a discovery about the world but a decision about where to stand while looking at it.

Now, a promise about symmetry, because a method that only ever embarrasses one side is not a method but a weapon with a scholarly finish.

The rule I work under elsewhere is that a pipeline found to be unreliable cannot afterward be cited in support of anything, including conclusions I happen to like. Declarations cut both ways. In this book that rule has teeth. When Martha Nussbaum accused Butler of substituting parody for politics, she made several arguments, some of which survive an audit intact and some of which do not, and I will say which are which without regard to the fact that the intact ones are inconvenient. When critics on the other side accuse Butler of denying material reality, some of them have identified a genuine unfilled position in the theory and some of them are objecting to a book they have not read. Those are different objections and they deserve different answers. And the international anti-gender movement, whatever else may be said about it, is built almost entirely out of declared components presented as measurements — which is a diagnosis, not an insult, and one that would apply equally to any movement so constructed.

The final chapters turn the instrument on the instrument. There is a specific failure mode that afflicts anyone who has a favorite framework, and I have one. If a great many separate confirmations all descend from a single underlying quantity, then the confirmations are not independent, and reporting them as several successes is reporting one success several times. That failure mode is easy to spot in other people. I have spent enough time finding it in other people to know roughly where to look for it in myself, and the second-to-last chapter is that search, conducted in public.

A few practical notes, and then we can begin.

There is no mathematics in this book. Not simplified mathematics, not mathematics in a box that the nervous reader may skip: none. Everything here can be said in language, and where it cannot be said in language it probably could not be said at all. I take that as a working principle rather than a concession to the audience.

There is also no attempt to explain Butler’s ideas by way of Butler’s biography. The early chapters do describe how a talkative fourteen-year-old was sent to remedial ethics tutorials as a punishment, and what happened when the punishment turned out to be a reward. But that is context, not causation. Ideas are not confessions, and the habit of explaining a philosopher by their circumstances is usually a way of not having to argue with them.

On pronouns: Butler has said they use both she and they, and has expressed a preference for the latter. This book uses they throughout, which is a declaration of exactly the kind the book is about — a choice, made openly, that the reader is now free to weigh.

On tone: the subject is serious and the treatment is not solemn. Solemnity is a poor instrument for this sort of work. It makes bad arguments harder to notice, because it dresses them in the clothing of importance, and it makes good arguments harder to enjoy. Some of what follows is funny. Almost none of it is funny at anyone’s expense, with the possible exception of the 1998 committee that gave Butler a prize for bad writing and thereby produced a sentence in the citation that was itself unreadable.

And on the question everyone asks first, which is whether the book is for or against. Consider one last time the surveyor on the hilltop. Their measurement of the spire is not wrong, and it is not right either; it is a number that carries a declaration inside it, and the useful thing to do is not to accept it or reject it but to name the declaration and see how much of the answer depends on it. That is what an audit is. It ends not with a verdict but with a ledger — this much was seen, this much was chosen, this much follows — and readers who want the verdict handed to them are going to be disappointed in a way I hope proves productive.

What survives an audit is not everything, and what survives is worth more afterward than it was before. In the case of Judith Butler, a surprising amount survives. Not all of it is the part the admirers defend, and not all of the wreckage is the part the opponents attack. That mismatch is the book.

Chapter One — The Talkative Child and the Punishment That Backfired

The rabbi at the Hebrew school in Cleveland had a discipline problem, and the discipline problem was fourteen years old and would not stop talking. The offense was not blasphemy, or truancy, or any of the interesting sins. It was talking in class, compounded by a tendency toward clowning, which is the technical term for finding out how much of a rule is actually load-bearing. The remedy chosen was a period of extra instruction. The talkative student would receive private tutorials in Jewish ethics, outside of normal hours, on top of the existing coursework. It was a punishment. It was designed to be a punishment. It was administered by an adult who had every reason to believe that more schoolwork is what more schoolwork means.

The student was thrilled.

This is the first thing to understand about Judith Butler, and it has almost nothing to do with gender. An institution took an action, gave the action a name, and the action then went out into the world and did something entirely different from what its name said. The name was punishment. The effect was a scholarship. Nobody lied. The rabbi was not concealing a secret program of philosophical patronage. He believed he was imposing a cost, and he was, in the same way that a fine imposed on a man who wanted to be arrested is still technically a fine. What he could not do — what nobody can do — was guarantee that the meaning of his action would survive contact with the person receiving it.

There is a general principle buried in this domestic comedy, and because it is the instrument this whole book will use, it is worth setting down early in its plainest form. Institutions describe their own actions in terms of the function those actions are intended to serve. That description is not a measurement. It is a declaration. Whether the action serves the function is a separate question, answerable only by looking. And the gap between the two is not usually the result of dishonesty; it is the result of the fact that the person acting and the person acted upon are running different accounting systems, and only one of them gets to write the label.

Consider what a punishment actually requires in order to work. It requires that the recipient share the punisher’s valuation of the thing being imposed. Solitary confinement is a punishment for a gregarious prisoner and a mercy for a persecuted one. Being sent to your room is a catastrophe if the party is downstairs and a rescue if the relatives are. Extra reading is a burden to a child who does not want to read. None of these facts are properties of the sanction. They are properties of a relation between the sanction and a particular person, and the sanction cannot see the person. A disciplinary system is therefore always making a bet on the inner life of its subjects, and the bet is invisible in the paperwork, which records only that a penalty was applied.

That the penalty in this case produced a philosopher is a nice story, and nice stories are exactly what an auditor should be suspicious of. There are a great many talkative children, and the overwhelming majority of them are given remedial work of one kind or another, and the yield in continental philosophy has been modest. What the anecdote establishes is not a cause. It establishes a shape — a first, homely instance of the divergence between what a practice announces itself to be doing and what it does — and the shape will turn out to be the one Butler spends fifty years drawing.

The house the child was punished into was itself a small instructional exhibit on the same theme. Butler was born in Cleveland in February 1956, into a family of Hungarian and Russian Jewish descent. The parents were practicing Reform Jews, but that description flattens a more interesting arrangement: the mother had been raised Orthodox, moved to Conservative observance, and arrived at Reform; the father had been Reform throughout. What a child learns in such a household is not that religion is false, which is what the anxious assume. What the child learns is that a tradition is not one object. The same texts, the same holidays, the same God, and three different lives lived out of them, all of them recognizably the tradition, none of them the tradition itself. Observance turns out to be something people do, repeatedly, in ways that drift; and the identity of the thing observed is maintained by the doing rather than sitting underneath it as a deposit.

I am going to be strict with myself about this and I would ask the reader to be strict too. That observation is not an explanation. It does not tell us that a Cleveland childhood produced the theory of gender performativity, and anyone who says it does is running the biographical version of the very error this book is about — taking a pattern that was selected in hindsight and reporting it as a measurement. What it tells us is what furniture was in the room. The furniture is worth knowing because a great many of Butler’s critics have assumed the furniture was French.

There is a darker item in the room as well, and it should be named without being made into a moral. Most of the family of Butler’s maternal grandmother was murdered in the Holocaust. That fact does not explain anything and does not need to. Its relevance is narrower and more exact: in that house, the question of which human beings are counted as human by the arrangements around them was not a seminar topic. It was a piece of family arithmetic. Decades later Butler would write several books circling a single question — what makes a life register as a life, such that its ending registers as a loss — and would be accused, with some regularity, of having invented a fashionable abstraction. It is possible to disagree with every one of those books. It is not really possible to think the question was picked up at a conference.

Return, though, to the rabbi, because there is one more thing in the story and it is the sharpest.

What an institution punishes tells you what it is actually optimizing for, and this is frequently not what it says it values. A school will tell you it values curiosity, engagement, and the courage to speak. A classroom of thirty children is nonetheless a fragile mechanical arrangement held together by silence, and talking is the thing that breaks it. So talking is punished. The declared value is inquiry; the enforced constraint is order; and where the two conflict, the enforced constraint wins every time, which is how you can tell which one is real. This is not cynicism about schools. It is the ordinary condition of any system that has to run a room. The point is only that you find out what a system is by watching what it will not tolerate, and never by reading its statement of purpose.

A child who clowns is running exactly this experiment. Clowning is not naughtiness with better timing. It is a probe. It establishes, cheaply and reversibly, which rules will bend, which will break, which are enforced only when a certain adult is present, and which are so deep that nobody has ever bothered to state them. Every social group is full of rules of the last kind, and the only way to find them is to violate one and observe the size of the reaction. The reaction is the measurement. The stated rule was the declaration. A talkative fourteen-year-old with a taste for testing the difference is not a discipline problem. It is a methodology, arriving early and without documentation.

The tutorials that followed have their own chapter, because of what was asked in them. What matters here is the structure of the transaction. A cost was imposed and received as a gift. A rule intended to reduce the amount of talking in the world produced, by a series of steps nobody in the room could have predicted, several million words. And the person on whom the mechanism failed appears to have understood immediately that it had failed, and to have said nothing, and to have gone to the tutorials.

The structure generalizes upward with dispiriting ease. A government announces sanctions on a hostile state, declaring the function to be the weakening of a regime; the measured effect, often enough, is a regime with a monopoly on scarce goods and a ready explanation for every shortage. A city passes an ordinance against loitering, declaring the function to be public safety; the measured effect is a legal instrument for removing whichever people the police were already inclined to remove. A university introduces a code declaring its commitment to open inquiry; the measured effect depends entirely on who is empowered to lodge complaints. In none of these cases is the declaration a lie, and in all of them the declaration is the part that gets archived.

That last point deserves its own sentence, because it is the reason historians and auditors have such trouble with this. Records preserve declarations. The minutes of the meeting record that a penalty was imposed for talking in class; they do not record that the recipient was thrilled. Statutes preserve their preambles. Institutions write down what they meant, in the language of what they meant, and the divergence between meaning and effect survives only if somebody happens to write a memoir. Every archive is therefore biased in a specific and correctable direction, and the correction is simply to remember that you are reading the intentions of the powerful and calling it a record of events.

There is a sentimental version of this chapter in which the moral is that we should be gentler with difficult children. That moral may well be true and it is not mine. My moral is colder and more useful. The effect of a rule is not contained in the rule. It is a fact about the world, it has to be looked at, and no amount of clarity about intentions will substitute for looking. Every chapter that follows is an application of that sentence to a field in which almost everybody, on every side, has preferred to argue about intentions.

Chapter Two — Three Questions Asked at Fourteen

When the tutorials began, the tutor did what good tutors do and asked the student what she wanted to study. The answer, delivered by a fourteen-year-old in Ohio in 1970, consisted of three questions. Why had Spinoza been expelled from his synagogue? Could German Idealism be held responsible for Nazism? And what was one to make of existential theology, in particular the work of Martin Buber?

The standard reaction to this list is to marvel at the precocity, which is the least interesting thing about it. Precocity is common and mostly evaporates. What is uncommon, and what does not evaporate, is that all three questions have the same shape. Each of them asks about the relationship between a structure and something the structure produces or expels. Each of them refuses the comfortable position in which ideas float free of their consequences and communities float free of their exclusions. Butler would spend a career on that shape, and would be criticized in terms drawn, with an irony nobody arranged, from the second question on the list.

Take Spinoza first. In Amsterdam in 1656 the Portuguese Jewish community issued a writ of excommunication against a twenty-three-year-old member, in language of exceptional violence, forbidding the congregation to speak with him, read anything he wrote, or come within four cubits of him. The detail that repays attention is the timing. Spinoza had published nothing. The doctrines for which he is now expelled from undergraduate syllabi — that God and nature are one thing, that scripture is a human document with a political history, that there is no personal immortality — were years from print. The community acted on reports of what he had said and on inferences about where he was heading. That is to say: it acted on an attribution. It declared a position, assigned it to a man, and then punished the man for the position it had assigned.

This is a good deal more than a historical curiosity, because excommunication is the clearest case there is of language that does rather than describes. The writ does not report that Spinoza has become an outsider. The writ makes him one. Before it is read aloud he is a member in bad standing; afterward he is nothing, and the change was accomplished by words spoken by people authorized to speak them, in a form the community recognizes. Philosophers would later build a whole subdiscipline on such utterances, and Butler would eventually be one of the people building it. The fourteen-year-old had simply picked the most dramatic example in the tradition and asked how it worked.

The second question is the one that ages best and hurts most. Can a philosophy be held accountable for what is done in its name? The maximalist answer had been given not long before by Karl Popper, who traced a line from Hegel through the worship of the state to the totalitarianism of the twentieth century, and by others who ran similar lines through German thought generally. Most scholars now regard the charge as considerably overdrawn — the actual Nazi state had little use for Hegel’s rational legal order, and the philosophers it did enlist were mostly enlisted against their texts rather than by them. But the retreat from the charge cannot be total, because a philosophy that could not possibly be implicated in anything would be a philosophy that never touched the world at all, and that is not a defense, it is an epitaph.

So the honest answer lives in the middle, and getting to it requires precisely the sort of accounting this book is built around. Some consequences follow from a body of thought by steps that can be written down and checked. Others require the later user to add premises the original did not contain, and frequently to add premises the original explicitly forbade. Those two situations look identical from a distance, since in both cases something dreadful is being done by people quoting a book. They are not identical, and the difference is not a matter of taste. It is a matter of tracing which components came from the source and which were supplied by the borrower, and refusing to charge the source for the borrower’s additions or to excuse the source for its own.

That the person who asked this question at fourteen would become, at sixty, the subject of it, is the kind of symmetry that no novelist would risk. Butler’s work has been credited and blamed for effects at a considerable distance from anything Butler wrote — for policies, for institutional practices, for a generation’s vocabulary, for a decline in the willingness of people to say obvious things. Some of those transmissions are traceable and some are inventions. Sorting them is work, and almost nobody in the public argument has been willing to do it, in either direction. Chapter Twenty-One takes up the sorting; I mention it here only to note that the tools for the job were requested by the defendant at the age of fourteen, which is either poignant or funny depending on the hour.

The third question, about Buber, is the quietest and probably the most consequential. Buber’s central proposal, published in 1923, is that there are two fundamental ways of standing toward the world, and that they are not two attitudes a pre-existing self can adopt but two conditions in which different selves come into being. Address someone as a Thou and you become a certain kind of I. Handle them as an It and you become another. The relation is primary; the terms of the relation are precipitates of it. You do not first exist and subsequently enter into address. You are constituted in the addressing.

Hold that next to a sentence Butler would write four decades later, in a book about the limits of self-knowledge: that the self is formed in a scene of address it did not choose and cannot fully recover, and that this is why no one can give a complete account of themselves. The vocabulary has changed; a good deal of psychoanalysis and Hegel has been added; the conclusions run somewhere Buber would not have gone. But the germ is visibly the same germ. A theological question about prayer and encounter, asked by a punished teenager, becomes an ethics of opacity and responsibility. Ideas do travel like this. They travel much more often like this than by the route of argument and refutation, which is why intellectual history keeps having to be rewritten by people who look at what was actually read.

Now the warning, which I promised in the preface and which is due at exactly this point.

Retrodiction is cheap. Give me any accomplished person’s childhood and a free afternoon and I will find you the seeds of everything they later did, because I am selecting the seeds with the plant already in front of me. The declared component in every biographical explanation is the selection rule: which facts count as formative. Nobody publishes the inventory of a great thinker’s abandoned enthusiasms, the instruments never mastered, the obsessions that led nowhere, the tutorials that produced nothing but a mild lifelong distaste for the subject. If those were included, the pattern would dissolve, which is precisely why they are not included. I have just spent several pages doing this. The reader should discount accordingly, and I would rather say so than perform an objectivity I am not exercising.

What survives the discount is smaller and firmer. It is not that the three questions predicted the three answers. It is that a habit of question-shape is a more durable thing than any opinion, and considerably more durable than a doctrine. Opinions get abandoned. Doctrines get revised past recognition. But the form of the question a person finds interesting — what has to already be in place for this to be sayable, countable, punishable, mournable — tends to persist from the first inquiry to the last, and it can be identified without any speculation about motives, because it is right there on the surface of the texts.

Butler’s question-form is transcendental in the strict and unglamorous sense: not concerned with what exists, but with the conditions under which something can show up as existing at all. That form was standard equipment in the German tradition the tutorials introduced, and it was already the form of the three questions. What are the conditions under which a community can unmake a member? What are the conditions under which a system of thought can be charged with a crime? What are the conditions under which one person becomes present to another? Every controversial thing Butler has ever written is one of those three questions asked about a new object.

There is also an asymmetry in these attributions that nobody has ever managed to make fair. A thinker receives credit for consequences they did not foresee and did not intend, provided the consequences are admired; the same thinker receives blame for consequences they did not foresee and did not intend, provided the consequences are deplored. Both transactions use the identical inference and only one of them is ever challenged. An honest accounting would either grant both or refuse both, and would notice that the choice between granting and refusing is being made on the basis of how one feels about the outcome, which is not a form of reasoning at all.

Buber’s second term is worth a further moment, because it has aged into something Buber could not have anticipated. To handle a person as an It, in his account, is to encounter them as an object with properties, available for use and classification. He was thinking about the ordinary hardness of daily life. We now live inside institutions that perform the classification automatically, at scale, for people no official will ever meet: a credit score, a risk tier, an eligibility category, a flag in a database. The It-relation has been industrialized and made impersonal in a sense Buber never contemplated, since there is now frequently no I on the other end of it at all. A teenager asking in 1970 how one person becomes present to another was asking a question whose urgency was about to be multiplied by machinery.

Which leaves an obvious question of my own, one I have no way of settling: what would have happened if the rabbi had simply confiscated the notes and kept the child after class. Probably nothing. Probably there is a version of the twentieth century in which the tutorials are declined, the questions go unasked, and the interesting shape shows up somewhere else in a different vocabulary, because these things are rarely as fragile as the anecdotes suggest. I raise it only because it is the sort of counterfactual that biography systematically hides, and hiding it is how a sequence of accidents comes to look like a destiny.

Chapter Three — Desire Before Gender

In the American graduate departments of the late 1970s the fashionable continental import was not gender. It was Hegel — and not even Hegel exactly, but a particular Hegel, assembled in Paris in the 1930s by a Russian emigre named Alexandre Kojeve, who lectured on the Phenomenology of Spirit to an audience that included, at various points, Jacques Lacan, Georges Bataille, Maurice Merleau-Ponty, Raymond Aron, and Raymond Queneau, and who quietly reorganized the whole book around a few pages about a master and a slave. Nearly everything that later arrived in America as French Theory had passed through that room. A student of the period who wanted to know why the French were talking the way they talked had to go back through Kojeve, and that is what Butler did.

The sequence of institutions is quickly told. Bennington College first, then a transfer to Yale, a bachelor’s degree in 1978, a Fulbright year at Heidelberg in 1979, and a doctorate in 1984 under the phenomenologist Maurice Natanson, with George Schrader also supervising. The dissertation was on the projects of desire in Hegel, Kojeve, Hyppolite, and Sartre. It was revised and published in 1987 as a book about desire, three years before the book about gender that everybody has heard of. This ordering is not trivia. It is the single most useful fact for anyone trying to work out what Butler actually thinks, and it is routinely omitted from the accounts written by people on both sides of the public quarrel, who tend to write as though Butler sprang into being in 1990 holding a copy of Foucault.

What is the twentieth-century French Hegel about? It is about desire, and about the surprising claim that desire is not primarily aimed at objects. An animal desires food and the desire ends when the food is eaten. A human being, on Kojeve’s reading of Hegel, desires something no object can supply: to be recognized as a subject by another subject. That is why the desire cannot be satisfied by consumption, and why it drives history rather than terminating in a meal. What I want, when I want in the human way, is your desire — specifically, I want to occupy the place of value in your account of the world.

Follow the structure rather than the drama and something important falls out. If I become a subject only by being recognized as one, then I am not a subject before the recognition. My standing is not a property I possess and then display. It is conferred, in a transaction with someone who is themselves in the same position, and the whole arrangement is therefore unstable, dependent, and social all the way down. There is no bedrock self underneath waiting to be acknowledged. There is a process of mutual acknowledgment that generates selves as its output.

That is the load-bearing beam of everything Butler would later build, and it was set in place six years before the word gender appeared on one of their book covers. Read Gender Trouble without it and the book looks like a wilful denial of obvious facts by someone with a taste for paradox. Read it with the beam in view and the book looks like an application: here is a general account of how subjects are produced rather than expressed, now let us see what it does to a domain where everyone has assumed expression.

The other three names on the dissertation supply the remaining hardware. Jean Hyppolite, who translated the Phenomenology into French and wrote the commentary that trained a generation, supplies the reading of the subject as a movement rather than a substance — something that exists in the way a process exists, by continuing. Sartre supplies the sharpest version of the empty center: consciousness, in his account, is not a thing with properties but a lack, defined by what it is not and what it is not yet, condemned to invent itself because there is no essence available to express. And Kojeve supplies the recognition machinery already described, along with a certain melodrama that Butler would later trim.

Behind all of them stands a shorter and ruder formulation, from Nietzsche, which Butler would quote at the decisive moment in Gender Trouble: there is no being behind the doing, and the doer is a fiction added to the deed. Nietzsche’s target was the grammar of moral judgment. He noticed that we say lightning flashes, as though there were a lightning that then flashed, and that we do the same with human action so as to have someone to blame — first invent a doer behind the deed, then hold the invention responsible. Butler’s contribution was to notice that we do exactly this with gender, and to say so in a book that sold a hundred thousand copies.

It follows, and it is worth stating flatly, that the most notorious proposition in modern gender theory is a nineteenth-century thesis with an unbroken paper trail. A very large share of the outrage of the 1990s was outrage at Friedrich Nietzsche, arriving a century late and addressed to the wrong person. This does not make the proposition true. Pedigree is not evidence, and an old error is still an error. But it does dispose of a certain style of objection — the one that treats the empty subject-position as a personal eccentricity, a piece of academic exhibitionism, or the sort of thing a person says when they have spent too long in California. It is none of those. It is a position with a long argumentative history, and if it is to be defeated it has to be defeated on those grounds.

Here is where the audit earns its keep, because it cuts against both parties at once.

Against the critics: the empty subject-position in Butler is inherited, explicit, and documented. It appears in a dissertation about other people’s systems, where Butler is describing what Hegel and Kojeve and Sartre committed themselves to and saying so on the page. At that stage the declaration is perfectly visible. Nobody could read Subjects of Desire and fail to notice that a particular account of subjectivity is being adopted from a particular tradition. Criticism that proceeds as though Butler smuggled the assumption in has simply not looked at the earlier book.

Against the defenders: inheriting a declaration does not convert it into a measurement. This is the point at which admirers of a thinker reliably go soft. A commitment adopted from Hegel is still a commitment adopted, and every conclusion downstream of it carries it along. The proposition that there is no doer behind the deed was not established by observation of human beings. It was arrived at by reflection on grammar and on the structure of moral attribution, and it functions in the later work as a starting point rather than a finding. Which is entirely legitimate — one must start somewhere — provided the starting point is carried in the open. Whether it is carried in the open in Gender Trouble is a question I will take up in its place.

There is a second inheritance from these years that gets less attention and shapes more sentences. Butler’s advisor was a phenomenologist, and phenomenology has a technical vocabulary of acts — constituting acts, sedimented acts, the acts by which a world comes to have a settled appearance. In that tradition an act is not a performance in front of an audience. It is closer to what a river does to a valley: something repeated so many times that the result stops looking like the outcome of a process and starts looking like the landscape. When Butler later writes about acts, the phenomenologists are much nearer than the theater is. The word did not survive contact with the reading public, and a good deal of the next thirty years was spent on the consequences.

One detail of the training deserves emphasis because it separates Butler from a whole American generation. A great many enthusiasts of French theory in the 1980s encountered the German tradition only in French translation and French summary, arriving at Hegel through people who had reasons of their own for reshaping him. Butler spent a Fulbright year in Heidelberg and worked on the Germans in German. The difference shows up in the later texts as a certain immunity: when a French author is making free with a German source, Butler tends to notice, and to say which is which. This is not a small virtue in a field where a good deal of the argument consists of stacked misreadings passed off as developments.

There is one further inheritance from the dissertation that is easy to miss and worth stating, since the whole later reception turns on it. The French Hegel is a Hegel of dissatisfaction. Desire structured as a demand for recognition can never be finally met, because a recognition extracted by force is worthless and a recognition freely given can always be withdrawn. What the tradition calls unhappy consciousness is the permanent condition of a subject whose existence depends on an acknowledgment it cannot secure. Butler inherits this and never abandons it, which is why the later political writing is so resistant to the triumphant register. A theory in which the self is conferred by others is a theory in which nobody is ever safe, and the people who read Butler as offering a program of liberation through self-invention have missed the mood of the whole enterprise, which is closer to mourning.

It is worth pausing over how ordinary this early period was. A student reads the difficult books, goes to Germany for a year, writes a dissertation on four dead men, revises it, publishes it with a university press, and is read by the several hundred people who read such things. There is no scandal in it and no sign of what is coming. The controversy that would eventually attach to the name attached to an application of the machinery, not to the machinery, and the machinery was assembled in the most conventional way available.

What Butler brought to the study of gender, then, was not a theory of gender. It was a theory of the subject, developed on unrelated material, in a tradition with no particular interest in the subject, and then carried across. That maneuver — go up in one domain until the generalizing stops paying, then come down somewhere else with the structure in hand — is how a great deal of serious work gets done in every field, and it is also the reason such work is so often received as an intrusion. The specialists in the new domain see someone arriving with an apparatus that was not built for their problem and does not use their vocabulary, and they are not wrong to be suspicious. They are only wrong to assume that an apparatus built elsewhere cannot fit. Sometimes it does not. Sometimes it fits so well that it takes the field forty years to recover, which is roughly the situation we are about to examine.

Chapter Four — What a Theory Is Allowed to Assume

In a vault outside Paris, under three nested bell jars, in a room that requires three separately held keys to open, there used to live a small cylinder of platinum and iridium that was the kilogram. Not a kilogram: the kilogram. If it gained mass from a stray fingerprint, then every other mass in the world became correspondingly lighter, by definition, instantly, everywhere. If it lost mass, the reverse. There was no fact of the matter about whether the cylinder was drifting, because there was nothing left to compare it against. The comparison was the cylinder.

Over a century the cylinder and its official copies drifted apart by a few tens of micrograms, which is a quantity of no importance to anyone weighing potatoes and of considerable importance to the people whose job is consistency. In 2019 the world’s metrologists gave up on the object and redefined the kilogram in terms of a constant of nature, whose value they fixed by decision. This is the crucial move and it repays a slow reading. They did not discover the value. They chose it — chose it to match the best existing measurements so that nothing would visibly change — and thereafter that constant is not measured at all. It is declared, and every subsequent measurement of mass in the universe is made in units that carry the declaration inside them.

The same thing had already happened to the meter. It began as a fraction of the distance from the equator to the pole, became a bar in a vault, and is now defined by fixing the speed of light at a particular number of meters per second. Ask a physicist to measure the speed of light and you will get a slightly embarrassed answer, because the speed of light is no longer something one measures. It is a definition. What is measured now is length, using it.

I begin with this because it is the cleanest available demonstration that a declaration is not a failure. There is nothing sloppy about the redefinition of the kilogram; it was carried out by extremely careful people for extremely good reasons, and the result is a system of units more stable than the one it replaced. The declaration is doing legitimate work. What would be sloppy — what would be a genuine intellectual offense — is to forget it is there, and to report a number as though the whole of it had come out of an instrument.

So here is the taxonomy, laid out properly, since the rest of this book will use it on material a good deal more inflammable than platinum.

A component is measured when it is what the instrument registered and would have registered under a rival theory. Measurement is not infallible and not theory-free — no serious person has believed that for a hundred years — but there is a real and useful difference between a quantity that would survive a change of framework and one that would evaporate with it. If your opponent, using their apparatus, gets the same reading as you, that reading is doing something other than expressing your commitments.

A component is derived when it follows from other components by reasoning that can be laid out and checked. Derivation transmits strength without creating it. A conclusion derived flawlessly from a declaration is exactly as strong as the declaration and not one bit stronger, however long and impressive the derivation is. Length is in fact a warning sign, since a sufficiently elaborate chain lets everyone forget what was at the top of it.

A component is declared when it was chosen. Under this heading fall: definitions, coordinate systems, the boundaries of the object under study, the population being sampled, what counts as a case, what counts as an event, the assumptions loaded before the data arrive, the calibration point, and the names given to things known only by their effects. Declarations are unavoidable. Nothing whatever can be said without them. They are not confessions of weakness, and a field that tried to eliminate them would fall silent in an afternoon.

Everyday life is thick with them, once you look. The poverty line used in the United States descends from a calculation made in 1963, in which an economist took a minimal food budget and multiplied it by three, on the grounds that families then spent about a third of their income on food. Every subsequent statement about how many people are poor inherits that multiplication and that decade’s grocery habits. The number of continents is five, six, or seven depending on which country educated you, and the disagreement is not about geology. Intelligence scores are set to average one hundred by construction, so the discovery that the average is one hundred is not a discovery. Species boundaries are contested among biologists who agree entirely about the organisms. In each case the underlying reality is not in doubt and the reported quantity is partly a decision.

From this follow four working rules, which I will state in plain language and use for the rest of the book.

Name it. Every declared component gets identified, in words, in the vicinity of the conclusion it supports. Not in an appendix, not in the methods section that reviewers skip, but where the claim is made. Most bad reasoning in public life is not a failure of logic; it is a declaration that has slipped out of view and taken on the color of a finding.

Locate it. Say at which step the declaration enters. Early declarations are more dangerous than late ones, because everything downstream inherits them and because by the time the argument gets loud nobody remembers the first page. A choice about what counts as an instance of the phenomenon is made before any evidence has been examined, and it does more work than any subsequent piece of evidence.

Price it. Ask how much of the conclusion is resting on the declaration, and answer with a number or a direction. This is the step that separates an audit from a debating trick. Any conclusion whatever can be attacked by saying it rests on assumptions; that observation is free and worthless. The interesting question is what happens when you vary the assumption. If the conclusion barely moves, the declaration was not load-bearing and the attack fails. If the conclusion inverts, then the argument was never about the world in the first place, and both parties have been shouting about a choice.

Apply it symmetrically. The rule that matters most and is obeyed least: a method used only on the conclusions you dislike is not a method. If a chain of reasoning is found to be unreliable, it cannot afterward be cited in support of anything, including the propositions you were hoping to establish. This costs something. It has cost me things I would have liked to keep. But an instrument that can only give one answer is not an instrument, it is a decoration, and anyone can see the difference from across the room.

There is a refinement to the third rule which took me an embarrassingly long time to find, and it will matter later. What determines how much damage a declaration does is not that it was chosen but whether it is free. Some declared quantities cannot be varied to taste — the structure of the problem, or the range in which the analysis is valid at all, pins them within a narrow window. Others can be slid anywhere the analyst pleases, and reliably end up wherever they need to be for the conclusion to come out well. Two arguments may each rest on exactly one declared component and be in completely different epistemic conditions. Count the free parameters, not the assumptions.

Pricing is easier to illustrate than to describe. The official unemployment rate counts people who are without work and actively looking; a second official measure adds those who have given up looking and those working part time who want full time, and it typically runs to roughly twice the first. Neither figure is wrong. They answer different questions, and the choice between them is a declaration made before any survey is conducted. Gross domestic product excludes unpaid household labor, on the entirely coherent ground that it is not transacted; the consequence is that a country in which more childcare is purchased rather than performed at home appears to have grown, and the growth is partly an artifact of the boundary. Move the boundary, watch the conclusion move, report by how much: that is the whole of pricing, and it is astonishing how rarely it is done in public.

It is also worth noticing where declarations tend to congregate, because knowing where to look saves an enormous amount of time. They cluster at the edges of objects. What counts as one event, one case, one person, one instance of the phenomenon — these questions are settled early, quietly, usually by convention inherited from a predecessor, and they determine the shape of everything that follows. When two competent groups reach incompatible conclusions from the same evidence, the disagreement is at the boundary about nine times in ten, and it is at the arithmetic almost never.

Two warnings, and then we can go back to the subject.

The first: naming a declaration does not refute the argument containing it. This is the commonest abuse of the whole approach, and it is beloved of undergraduates and newspaper columnists. To say that a result depends on a choice is to relocate the disagreement, not to win it. The relocated disagreement is usually more tractable than the original, which is the entire benefit; two people arguing about where to calibrate are much closer to resolution than two people arguing about who is denying reality.

The second: the taxonomy is itself a declaration. There is no instrument-independent standpoint from which the three categories can be read off nature; the partition is a tool, adopted because it has been useful, and the point at which it stops being useful is a real point that I will reach before the end of this book. I mention this now so that the reader has it in hand from the beginning, and so that nobody has to break the news to me later. It will come up again in Chapter Thirty-Two, under less comfortable circumstances.

Chapter Five — The Essay That Fit in a Theater Journal

Theatre Journal, volume forty, number four, December 1988, pages 519 to 531. Thirteen pages, in a periodical read chiefly by people who study and make plays, between articles on dramaturgy. The title was long in the fashion of the period: performative acts and gender constitution, an essay in phenomenology and feminist theory. It is the most consequential thing published in an American theater journal in the twentieth century, and the venue is going to matter enormously, in a way its author did not intend and could not have prevented.

The argument of the essay is short enough to state in a paragraph, which is more than can be said for most of what has been written about it. Gender, Butler proposes, is not a stable identity from which acts follow. It is constituted by the acts themselves — a series of doings, repeated across time, internally discontinuous, which together produce the appearance of an underlying substance. The appearance is convincing because the repetition is relentless and because the actors are among the audience: we are persuaded by our own performance. And because the identity is an achievement of repetition rather than a fact about origins, the possibilities for its transformation lie in the joints — in the fact that a repetition can always be performed slightly differently, or in the wrong order, or by the wrong body.

Two intellectual debts are named on the page. The first is to Simone de Beauvoir, whose famous sentence — one is not born, but rather becomes, a woman — Butler had already written about two years earlier in a separate essay. Beauvoir’s proposition looks at first like the whole of the later theory in compressed form, and it is not. Beauvoir keeps something Butler discards. In The Second Sex there is a subject who does the becoming: someone is there at the start, undergoing a cultural process, arriving at womanhood. Butler asks the impertinent question — who is this someone, before they have become anything? — and finds the position empty. There is no one waiting backstage to be dressed. The becoming goes all the way down, and the person is what the becoming produces.

The second debt is to Maurice Merleau-Ponty, and it is the one that explains the vocabulary. In the phenomenological tradition, the body is not a natural object that culture subsequently decorates. It is a historical situation, an idea taken up and lived, available to us only through a sedimented history of ways of moving, standing, being looked at, and looking. When Butler writes about acts, this is the technical background: acts in the sense in which a valley is cut by water. The word travels badly. In a theater journal, surrounded by discussions of staging, the word act had a second and much louder meaning available, and English generously supplied a third in Austin’s speech-act theory, where a performative is an utterance that does something rather than reporting something. Three senses of one word, all of them relevant, all of them different, all of them now permanently fused in the public mind. Chapter Eleven deals with the wreckage.

There is also a third and less discussed lineage. Gayle Rubin’s 1975 essay on the traffic in women had given feminism the sex-gender system as an analytic device, and Butler’s essay is in conversation with that framework rather than with the popular understanding of it. This matters because it locates the intervention: Butler is not addressing the general public’s view of men and women, which is an easier target and one that nobody needs a theory to hit. Butler is addressing a specialist apparatus already in use, and the objection is internal.

The speech-act background is worth a paragraph of its own, since it supplies the word and half the confusion. J. L. Austin had distinguished utterances that report a state of affairs from utterances that bring one about: naming a ship, pronouncing a marriage, opening a session, passing a sentence. His crucial observation was that such utterances do not succeed or fail by being true or false. They succeed or fail by conditions of a different kind — whether the speaker is authorized, whether the setting is right, whether the procedure was followed. Say the words of a marriage in a play and nothing happens; say them in a registry office and two lives are legally reorganized.

That framework is exactly what a theory of gender needs and it explains why the word was chosen. The declaration at the delivery room, announcing what has arrived, is not primarily a report. It initiates a long institutional process: a name, a document, a set of expectations, a series of corrections administered over two decades by people who mostly believe they are describing rather than instructing. The utterance works because a whole apparatus stands behind it, and it would work no better in a theater than a marriage would. Butler’s addition to Austin was the observation that this particular utterance is never made once. It is made continuously, by everybody, forever, and its authority comes from the repetition rather than from any single authorized speaker.

Now the part that everyone forgets, and that the essay states without ambiguity: none of this is voluntary. Butler writes that gender is put on under constraint, daily, incessantly, with both anxiety and pleasure, and that there are severe punishments for performing out of turn. The script is not chosen by the performer. It is inherited, enforced, and policed, and the enforcement is not metaphorical — the essay is written in a decade in which people were being beaten to death for the wrong performance. The word improvisation appears, but it appears hedged: unwarranted improvisation is what gets punished.

So the essay contains, explicitly, a refutation of the reading it would immediately receive. It says gender is not passively inscribed on the body, and it says gender is not determined by nature, and it says gender is compelled, and it says the compulsion is violent. Anyone who claims Butler proposed that we choose our genders each morning like a tie has either not read thirteen pages or has read them and preferred the other version. Both happen. The second is more interesting, because a theory becomes famous through its misreading much more often than through its content, and a misreading that gives everybody something to do is nearly unkillable.

The essay closes with something rarer than an argument, which is a self-imposed limit. Butler turns to the weakness of feminist theories built on a binary division of gender and worries in print about the reification of sexual difference — the process by which an analytic distinction hardens into a natural kind. The concern is that a framework designed to describe a hierarchy ends up guaranteeing the terms of the hierarchy, so that the femaleness being defended has been specified in advance and a great many actual women do not match the specification. This is not a fringe worry and it was not Butler’s alone. It was being pressed at the same moment, and with more urgency, by feminists who had noticed that the woman in feminist theory was reliably white, reliably comfortable, and reliably straight.

For the audit, the essay is unusually clean, and it is worth saying so plainly, because I am going to be less generous later. The declared components are visible. Butler declares that the subject-position behind gender is empty. Butler declares that acts are to be understood in the phenomenological sense. Butler declares which framework is under examination. What follows is derived from those declarations plus a body of observation about how gendered life is actually enforced, and the derivation is short enough to check. Thirteen pages, three declarations, all of them on the surface. There is a case to be made that Butler never again wrote anything this auditable, and I am inclined to make it.

The venue, though. The venue did something no argument could undo.

A piece of work published in a specialist journal is read first by that specialty, and the specialty’s vocabulary attaches to it before any other vocabulary can. Theater people read an essay about acts, performance, scripts, audiences, and improvisation, and they read it correctly by their own lights, and they took it out into the world. Within a few years performativity was being discussed as though its central image were a stage, with all that a stage implies: a performer who exists offstage, a costume that can be removed, an audience that could in principle be told the truth. Every one of those implications is the opposite of what the essay argues. The metaphor imported, in a single word, precisely the doer behind the deed that the theory was constructed to do without.

There is a general lesson here about the transmission of ideas and it is not a comfortable one for people who write. The fate of a piece of work is determined less by its content than by the vocabulary of whoever picks it up first, and the author has approximately no control over who that is. Butler would spend the next thirty-five years correcting the theatrical reading — a whole book in 1993 exists substantially for that purpose — and it did not work, and there is no reason to think anything would have worked. The correction is always slower than the misreading, because the misreading is easier to state and more fun to argue with.

It is worth asking whether the essay could have been written in a way that prevented this, and I think the honest answer is no. The technical vocabulary available in 1988 for talking about acts that produce their own actor was phenomenological and it was speech-act theoretical, and both of those words had already been colonized by the theater. There was no clean term. Butler could have invented one, and invented terms have their own well-known failure mode: they either never catch on, or they catch on and get filled with whatever the users bring. Given the choice between a word that will be misread and a word that will not be read, most authors take the first, and they are usually right, and they usually regret it.

Chapter Six — Trouble as a Method

Every committee that has ever tried to write a definition of its own subject has had the same bad afternoon. The historians cannot agree what counts as a historical source. The astronomers had to vote on whether Pluto was a planet, and the vote settled nothing except the syllabus. The medical bodies redraw the boundary of a disease and hundreds of thousands of people acquire or lose a diagnosis overnight without any change in their bodies. In each case a field that is entirely competent at its own work discovers that it cannot say, with precision, what the work is about — and, more disturbingly, that it has been getting along fine without knowing.

Feminism had this afternoon in the 1980s, and Gender Trouble is the report from it.

The problem is a genuine one and it has nothing to do with philosophy. A political movement needs a subject: someone on whose behalf the claims are made. Feminism’s subject was women. But every attempt to say what women are, in a way precise enough to ground a politics, turned out to exclude actual women. The description would fit the middle-class and not the poor, the white and not the black, the straight and not the lesbian, the mother and not the childless, the West and not everywhere else. This was not a discovery made by Butler and it is important to say so. It had been pressed for over a decade by the Combahee River Collective, by Audre Lorde, by bell hooks, and given its most durable formulation by Kimberle Crenshaw, whose account of intersecting subordinations appeared the year before Gender Trouble. The complaint was empirical and political: your category does not contain us, and the omission is not an oversight.

Butler’s contribution was to convert this into a structural claim, which is both the source of the book’s power and the reason it made people so angry. The philosophical version says: the problem is not that the category has been drawn too narrowly and could be redrawn better. The problem is that categories of this kind are produced by the same arrangements of power they are being used to oppose, and that a movement which begins by fixing its subject has already accepted the terms of what it opposes. Do not patch the definition. Ask what the demand for a definition is doing.

That is the genealogical method, borrowed from Nietzsche and reworked by Foucault, and its move is always the same. Do not ask what a thing is. Ask how it came to appear inevitable, whose position the appearance serves, and what account of its origins is being circulated. Applied to punishment, sexuality, madness, and the prison, this method had already produced a body of work that was reorganizing the humanities. Applied to gender, it produced a small book with an academic title from a publisher expecting modest sales, which then sold more than a hundred thousand copies and went into twenty-seven languages.

The most technically consequential argument in it is the one about sex and gender, and it deserves to be laid out slowly, because almost every subsequent quarrel is a distorted echo of it. The received framework said: sex is the biological substrate, gender is the cultural meaning built on top. This distinction had been enormously useful. It let feminists say that a fact about bodies did not entail a fact about capacities or destinies, and it broke the argument from nature that had been used to justify every existing arrangement.

Butler’s objection is that the distinction does not do what it appears to do. If gender is the cultural interpretation of sex, then sex is functioning as the thing beneath interpretation — the neutral bedrock, the fact prior to all meaning. But that bedrock is never available except through the apparatus that interprets it. There is no encounter with the raw fact; there is only ever the fact as sorted, named, recorded, and made to matter in particular ways. The division of bodies into two kinds, treated as the given prior to all culture, is itself an achievement of a system of meaning, and one with a documented history. What was supposed to be the natural ground turns out to have been produced by the machinery it was supposed to stand outside of.

In the vocabulary of this book: in the classical feminist framework, sex was a declared component functioning as a measurement. This is the audit reading of Gender Trouble, and I think it is the correct one. The book is not primarily a theory of gender. It is an audit report on feminist theory, which goes through the apparatus asking which parts were observed and which were assumed, and concludes that the load-bearing element at the base — the natural, prediscursive, two-kinded body — is an assumption that has been reporting itself as a finding.

That claim is at once weaker and stranger than what most people believe Butler said. It is weaker because it does not say that bodies are unreal, that anatomy is a fiction, that chromosomes are a rumor, or that anything follows about what should be done. It is a claim about the epistemic status of a component in an argument, not a claim about the contents of the universe. It is stranger because if you accept it, the comfortable division of labor between the sciences and the humanities on this topic stops working, and nobody in either building enjoys that.

Two honest concessions are owed at this point, and the second is the one that makes the rest of this book necessary.

The first is that critics who read Butler as denying material reality are usually auditing badly. They take a claim about how a category is constituted and hear a claim about whether the world exists. The two are not the same, the text distinguishes them, and this particular misreading has been corrected so many times that its persistence has to be counted as a choice rather than an accident.

The second is that the text is nonetheless genuinely ambiguous in places, and pretending otherwise would be exactly the partisanship I have promised to avoid. There are passages in Gender Trouble where the claim that a category is discursively produced slides, under the pressure of a rhetorical construction, into something much closer to the claim that its object is. Butler wrote an entire subsequent book because of this, opening it with the observation that a certain question kept arriving from every direction: what about the materiality of the body? An author does not write that book unless the first one left a hole. Whether the hole is a flaw in the argument or a flaw in the prose is a real question, and Chapter Thirteen is where I try to answer it.

Beneath the sex-gender argument lies a device Butler takes from Foucault and uses everywhere afterward, so it is worth isolating. Foucault’s scandal was the claim that prohibition does not simply suppress a thing that existed beforehand; it participates in producing the thing it prohibits, by naming it, specifying it, hunting for it, and giving those it names a category to inhabit. The nineteenth century did not discover a population of deviants and set about regulating them. The regulation and the population came into being together. Butler applies the same reversal to the law of gender: the rule that bodies must come in two coherent kinds does not police a natural division, it generates the division and then presents itself as its guardian.

This is the point at which readers most often decide the argument has gone too far, and it is worth being precise about why. The reversal is genuinely hard to hold in mind, because it violates a deep intuition that causes precede effects and that rules are made about things that already exist. Sometimes they are. The claim is not that every rule invents its object — laws against arson do not produce fire. The claim is narrower: where the object is a social category, a kind of person rather than a kind of event, the machinery of classification is part of what makes the kind. Whether gender is that sort of object is the entire dispute, and it cannot be settled by insisting on the intuition, because the intuition is what is in question.

There is one more piece of machinery to name, since it will be needed later. Butler describes a grid that makes bodies socially intelligible — a system in which a body is legible only if its sex, its gender, and the direction of its desire line up in a prescribed sequence. The grid is not a belief anybody holds explicitly; it is closer to a filing system, and its power lies in what it makes unfileable. A person who does not fit is not merely disapproved of. They are hard to perceive: their existence generates administrative confusion, requests for clarification, and the peculiar hostility that people show toward things they cannot categorize. Once you have the grid, a great deal of otherwise puzzling social behavior becomes predictable, which is a reasonable definition of a good theoretical device.

What the book did not have was a readable style, and the title did not help either. Gender Trouble sounds like a provocation and reads like a dissertation, which is the worst available combination: it attracted an enormous readership on the strength of a phrase and then declined to explain itself to them. A generation of people therefore learned the book’s conclusions from summaries written by people who had learned them from other summaries. That is how a careful argument about the epistemic status of a category became, in general circulation, the claim that gender is a costume. There is a chapter about the prose, and about the prize Butler was given for it, and about whether difficulty of that kind is ever justified. It is not a chapter that comes out entirely in Butler’s favor.

Chapter Seven — The Object Introduced Only by Its Consequences

On the night of the twenty-third of September, 1846, an astronomer in Berlin pointed a telescope at a patch of sky he had been told to examine and found a planet within one degree of where a Frenchman had calculated it should be. Urbain Le Verrier had never seen Neptune. He had seen only that Uranus was not where the arithmetic said it should be, and had reasoned backward from the discrepancy to the mass and position of whatever was pulling on it. This remains the single most impressive demonstration in the history of the physical sciences of a method that ought to make everybody nervous: introducing an object into the world purely on the strength of what it would have to be doing.

Thirteen years later the same man tried it again. Mercury’s orbit also failed to behave, its closest approach to the sun creeping forward by a small amount that Newtonian mechanics could not supply. Le Verrier proposed a planet inside Mercury’s orbit, named it Vulcan, and calculated where it should be. Observatories reported sightings. Expeditions were mounted during eclipses. Vulcan appeared in textbooks and in the popular imagination, and it does not exist. The anomaly was real, and it was resolved in 1915 by Einstein, whose theory of gravitation reproduced the missing amount without adding any planet at all. The effect had been correctly measured for half a century. The thing invented to produce it was never there.

Here, then, are two identical procedures with opposite outcomes, and no way to tell them apart at the time. In both cases a competent scientist observed a discrepancy, specified what would have to be present to produce it, gave the absent thing a name, and proceeded. In one case the name eventually acquired an occupant. In the other, the discrepancy turned out to be telling us something about the framework rather than about the contents of the solar system, and the name was quietly withdrawn.

The history of every science is thick with such cases. Phlogiston was introduced to explain what leaves a substance when it burns; the departure was real, and it was oxygen arriving rather than phlogiston leaving. The luminiferous ether was introduced because a wave must be a wave in something; the light was real, and it required no medium. Caloric was a fluid of heat that turned out to be a form of motion. On the other side of the ledger, the gene was for decades a purely functional posit — whatever it is that carries inheritance in discrete packets — and the position was eventually filled, spectacularly, by a molecule nobody had imagined when the role was written. Dark matter is currently in the queue and nobody yet knows which way it will go.

The pattern is stable enough to be given a name. Call it a role without a referent: an entity specified entirely by the effects it is required to produce, introduced because something must be producing them, and named as though the naming had settled the question. There is nothing illegitimate about this. It is one of the most productive moves available in inquiry, and half the furniture of modern science arrived through it. The illegitimate step is the small one that follows, where the name stops functioning as a placeholder for an unknown and starts functioning as a known object, so that the original discrepancy is now said to be explained rather than merely relabeled.

How do you tell whether an explanation has been given or a label applied? There is a serviceable test. Ask what the posited entity does other than the thing it was posited to do. Neptune, once found, was found to perturb other bodies, to have a mass consistent with its effect on Uranus, and to be photographable. It paid rent. Phlogiston, over its career, acquired properties precisely as needed and predicted nothing that was not already in the list of things it had been invented to handle — including, at one embarrassing juncture, negative weight. An entity that only ever does its original job is not an explanation of that job. It is the job, spelled differently.

Philosophy of mind has a word for this and it is worth borrowing, since it makes the structure easier to keep hold of. A functional role is a job description: whatever it is that gets caused by tissue damage, causes avoidance behavior, and makes a person say ouch, is pain. The description can be written before anybody knows what occupies the position, and it remains correct if the occupant turns out to be different in a human and in an octopus. Roles are cheap and useful. The confusion comes when the job description is mistaken for an employee.

The social sciences run on such positions and mostly know it. Utility, in economics, is defined as whatever it is that consistent choices maximize; it is inferred entirely from the choosing and is not otherwise inspectable, which is why serious economists treat it as bookkeeping rather than psychology. The general factor of intelligence is extracted from correlations among test scores and is a real statistical structure whose occupant remains contested a century on. In both fields the honest practitioners can tell you which of their terms are jobs and which are employees. The dishonest ones cannot, and the sign of it is that they are offended by the question.

One further hazard deserves a mention because it recurs in every case: naming an entity generates sightings of it. After Le Verrier announced Vulcan, credible observers reported seeing it, more than once, with instruments and in good faith. This is not fraud and not even really error in the ordinary sense; it is what happens when a specification tells a great many careful people what to look for and how to interpret an ambiguous smudge. Whatever else the history of Vulcan teaches, it teaches that the number of confirmations an entity receives is a poor guide to whether it exists, once the entity has a name and a predicted location.

Now consider the classical account of gender identity, and observe how exactly it fits the template.

The observation to be explained is real and not in dispute: people behave, over long stretches of time and across wildly different settings, in ways that are consistently gendered, and they experience this consistency from the inside as something given rather than chosen. To explain it, the standard account introduces an inner core — a gender identity, situated somewhere beneath behavior, of which the behavior is the expression. Where is this core? It is not available for inspection. How do we know it is there? Because the behavior is consistent, and something must be producing the consistency. What does it do besides produce the consistency? That is precisely the question, and it is not often asked.

This is not a cheap shot at anybody’s self-understanding, and I want to be careful here, because the argument is easy to misuse in both directions. The claim is not that people are wrong about their own experience. The claim is much narrower and concerns the logical status of a component in a theory: the inner core is introduced by inference from its consequences, it is not independently observed, and the theory that contains it does not usually admit this. In the vocabulary of the audit, it is a declared component reporting itself as a measurement, which by now the reader will recognize as the recurring offense of this book.

What Butler noticed in the late 1980s was exactly this, and stated it in the strongest possible form: the inner core is not a hidden thing awaiting discovery, it is a fiction produced by the very acts it is supposed to explain. Coherence in behavior generates the impression of an underlying source, in the way that a flickering sequence of images generates the impression of motion, and the impression is then read back as the cause. This is the Vulcan diagnosis applied to selfhood, and it is the reason the argument provoked a reaction out of all proportion to its immediate subject matter. Nobody minds losing a planet. People mind losing an interior.

Symmetry now demands the obvious question, and it is the question a good auditor asks first about their own instrument: does Butler’s alternative also posit something known only through its effects?

It does. The account replaces the inner core with regulatory norms, sedimented practices, and the pressure of a system of intelligibility — none of which can be photographed either. If the only complaint against the inner core were that nobody has seen it, the complaint would land equally on both theories, and the argument would be a draw. This is where a great deal of the public debate stops, with each side accusing the other of metaphysics, and it should not stop there, because there is a genuine and decidable difference between the two posits.

The difference is that norms leave tracks. A regulatory norm is not directly visible, but it is visible in enforcement: in what gets corrected, what gets punished, what a stranger will say to a child in a supermarket, what a form requires, what happens when the wrong box is ticked. Those are events with dates and witnesses, and they vary in ways that can be recorded across countries, centuries, and institutions. An inner core, by contrast, leaves no track other than the behavior it was invoked to explain. Both accounts posit an unseen producer of a seen effect; only one of them cashes out into other measurable consequences. That is not a proof, and it does not settle whether an inner core exists. It settles which of the two is currently doing more work than its own definition.

Which leaves the more interesting question, and it is the one the next chapter is about. Butler does not treat the empty position as a temporary vacancy to be filled by better science later. Butler treats it as permanently empty, on principle, as a positive commitment of the theory. Those are very different situations, they are constantly confused, and almost everything that has gone wrong in fifty years of argument about this material has gone wrong at exactly that point.

Chapter Eight — Deferred or Foreclosed

Two research programs can look identical from the outside and be in completely different conditions. Both have a name for something nobody has observed. Both proceed with the name in place. The difference is what each would do if the thing turned up.

Deferral is the ordinary case. The theory expects the position to be filled, would be delighted if it were, and is often actively hunting. A deferred referent has a search program attached: instruments are built, surveys are funded, the community can tell you what a discovery would look like and roughly what it would cost. When the gene was a functional posit, biology was not committed to genes being unfindable; it was committed to finding them, and a great deal of the twentieth century consisted of doing so. The vacancy is a to-do item.

Foreclosure is the rarer and much braver case. Here the theory states that there is nothing at the address, that the vacancy is permanent, and that the discovery of an occupant would not complete the theory but destroy it. Foreclosure is not agnosticism and it is not modesty. It is a positive claim about the structure of the world, and it takes on a risk that deferral avoids entirely. The deferring theorist can be wrong only about timing. The foreclosing theorist can be flatly refuted by a single observation.

Physics has some celebrated foreclosures and they are instructive. Special relativity does not say we have failed so far to detect a state of absolute rest; it says there is none, and any credible detection would end the theory. More dramatically, the question of whether quantum mechanics conceals a layer of ordinary determinate properties beneath its probabilities was, for thirty years, exactly the sort of dispute that seemed unresolvable in principle. Then John Bell showed in 1964 that the two positions differ in the statistics of correlated measurements, and experiments over the following decades came down against a whole class of hidden variables. A foreclosure that had been a philosophical preference became an experimental result. That is what the strongest possible version of this move looks like.

Butler’s central commitment is a foreclosure. There is no prediscursive gender core. Not undiscovered, not hard to reach, not deferred to future neuroscience: absent, and the theory is built on the absence. This is why I said in the preface that the most criticized feature of the work is also the most exposed. An evasive thinker deals with an awkward entity by leaving the question open and getting on with things. Butler nailed the door shut and stood in front of it.

Everything now depends on a question that almost nobody in the public argument has bothered to ask: what would count against it?

Here we run into a genuine difficulty, and I am going to state it as an opponent would, because a friendly statement of an objection is a wasted objection. Consider the most obvious candidate for disconfirming evidence — the very large number of people who report a deep, early, unchosen sense of their own gender, frequently one that runs directly against everything their environment was pressing on them, sustained at enormous cost, and unresponsive to correction. If a theory says gender is produced by regulatory norms, and here are people whose gender persisted against every norm available, does the theory not have a problem?

Butler’s answer is available in the texts and it is coherent. Production is not the opposite of depth. That something was formed rather than found does not make it shallow, chosen, or removable — quite the reverse, since the processes that form a person operate before that person exists to have opinions about them, and are for that reason less revisable than anything one merely believes. Nor does a norm-based account predict conformity. Norms are multiple, they conflict, they misfire, they are internalized in disorderly ways, and there is no mechanism guaranteeing that the result of a formation matches the intentions of the formers. Butler has also said directly that the work should not be read as claiming that gender self-perceptions are unreal, or that gender cannot be a substantial and enduring part of a person, and that statement is not a late concession but a consistent feature of the position.

This answer is correct, in my judgment, and it comes at a price that its author has never been very forthcoming about. A theory that can absorb both conformity and its opposite, both consistency and its violation, both a gender that follows the norms and a gender that defies them, is doing a great deal of absorbing. Somewhere in that flexibility is the risk that the foreclosure has become unfalsifiable in practice while remaining, on paper, one of the riskiest claims in the humanities. Foreclosure without a specified refutation route is a strange object: maximally exposed in principle, untouchable in fact.

So let me try to specify the route, since nobody else seems to have.

What would a decisive test look like? Not an appeal to the strength or sincerity of anyone’s self-report, since both theories predict strong sincere reports. It would have to be a demonstration that some component of gendered self-understanding is available to a person independently of the entire apparatus of social intelligibility — that it can be characterized without reference to any of the categories a culture supplies. And the trouble is immediately visible: any report of such a component arrives in language, in categories, addressed to somebody, which is the apparatus. The instrument is entangled with the thing measured, which is not an unfamiliar situation in science but is an extremely inconvenient one.

There are partial approaches. One could look at the earliest ages at which stable cross-normative identification appears, and ask whether it precedes the acquisition of the relevant categories — a question about developmental sequence rather than metaphysics, and in principle answerable. One could look for structure in gendered self-understanding that is invariant across cultures with radically different category systems, since a purely produced feature should vary with the producer while a prediscursive one should not. One could look at the cases where the categories on offer are genuinely absent or different, of which the ethnographic record has more than the debate usually admits. None of these is a Bell experiment. Each of them turns an unanswerable question into a difficult one, which is the most any method can promise.

It is worth adding that foreclosure is not exotic and not a postmodern affectation, since readers who encounter it first in this context tend to assume it is. Evolutionary biology forecloses foresight: the theory does not say that natural selection has goals we have not yet identified, it says there are none, and a demonstrably foresightful adaptation would be a catastrophe rather than a refinement. Thermodynamics forecloses the perpetual motion machine, not as a report on failed attempts but as a structural prohibition, which is why patent offices are entitled to reject such applications unexamined. In both cases the foreclosure is what gives the theory its teeth, and in both cases the community can tell you exactly what would overturn it.

That last clause is the whole test, and it is what separates a productive foreclosure from an unproductive one. A theory that says nothing is at the address and can describe what finding something would look like has taken a genuine risk. A theory that says nothing is at the address and has no account of what a discovery would even consist of has taken a risk that cannot be collected. The distinction is not about confidence or tone; it is about whether a refutation route exists in principle, and it can be assessed without any sympathy for either side.

There is a further asymmetry worth recording, since it explains why this argument never ends. The person defending an inner core needs only to say that the evidence is not yet in, which is always true and costs nothing. The person foreclosing has to defend a universal negative across every case anybody can produce. In any long-running dispute the party with the cheaper position will appear more reasonable to spectators, entirely independently of who is right, because reasonableness is being read off the size of the claim rather than the quality of the support. This is worth knowing about a great many disputes and it is not a point in anybody’s favor.

I want to end this chapter with the observation that made me write the book. In cosmology, the interesting papers are the ones where an author declares a component and says so. The overwhelming majority do not; the declaration is buried, and finding it is the work. Butler declared. The declaration is on the surface of the text, in the most inflammatory possible form, in a book aimed at a general readership. Whatever else that is, it is not what evasion looks like, and the fifty-year practice of treating it as evasion has cost the field the one conversation that might have been productive: not whether the position is empty, which Butler told us, but what it would take to show that it is not.

Chapter Nine — Repetition as a Fixed Point

A candle flame keeps its shape for hours while nothing in it stays put for a thousandth of a second. Wax rises, vaporizes, burns, and leaves as gas; the air that feeds it is drawn in and expelled continuously; the light comes from soot particles that exist for microseconds. Point at the flame and you are pointing at a location where a process keeps happening in the same way. There is no flame-stuff. There is a self-maintaining arrangement of stuff in transit, and the arrangement is what has the name.

This is not a curiosity about candles. It is the general condition of every object large or complicated enough to be interesting. A whirlpool is water passing through a shape. A wave crossing an ocean transports almost no water at all. The Great Red Spot on Jupiter has been a storm since before anyone had a telescope good enough to see it properly, and the gas in it has been replaced beyond counting. Your body has exchanged the overwhelming majority of its atoms since childhood; the bones took a decade, the gut lining takes days. A river keeps its name for millennia while every drop that constituted it has reached the sea. A language, a firm, a monastery, a nation, a marriage: none of these persists by containing durable material. They persist by being remade, continuously, by processes that happen to reproduce the arrangement they inherited.

The technical way to put this is that identity over time is a fixed point of a process of renewal — the arrangement that comes out the same when it is put through the machinery again. The important word is not fixed, which suggests rigidity, but point, which suggests a target that a dynamic process keeps returning to. If you want the largest description of a thing that stays true through repeated renewal, you want the biggest set of features that survives the operation, and everything outside that set is turnover. What survives is not what is made of lasting stuff. It is what keeps passing itself forward.

Two consequences follow immediately and both are counterintuitive.

The first: maintenance is not free. Every persisting arrangement is spending something — energy, attention, enforcement, ritual, money — to remain itself, and the moment the spending stops the arrangement begins to disassemble at a rate characteristic of its type. Flames go out in seconds, bodies in days, institutions in decades, languages in centuries. There is no such thing as a structure that simply continues. There are only structures whose maintenance cost is being paid by somebody, and one of the most reliable ways to understand any long-lived arrangement is to find out who is paying and how much.

The second: an arrangement maintained by process is not less real than an arrangement made of substance. This is the point at which ordinary intuition goes badly wrong, because the word constructed has been contaminated by the word artificial, which is in turn contaminated by the word arbitrary. A hurricane is constructed. It is not made of hurricane; it is made of air and water doing something; it will dissipate; and it will also destroy your house. Reality is not a function of what a thing is made of. It is a function of what a thing does and how hard it is to stop.

Now put Butler’s definition alongside this. Gender, in the 1988 formulation, is constituted by a stylized repetition of acts through time, which produces the appearance of an abiding substance while consisting of nothing but the repeating. Change the vocabulary and the sentence is a textbook description of a self-maintaining structure. The acts are the renewal process. The appearance of substance is what every such structure produces, which is why we speak of the flame and the river as things. The absence of an inner core is not a special claim about gender at all; it is the ordinary situation of everything that persists at a scale larger than a molecule.

I regard this as the strongest available defense of Butler’s central move, and it does not come from philosophy, which is why it has never been made. The defense is that the account of gender under attack is simply the standard account of persistence, applied to a domain where nobody had thought to apply it. If Butler is guilty of dissolving gender into a pattern of repetition with no substance underneath, then physics is guilty of the same offense against flames, meteorology against storms, biology against organisms, and linguistics against English. Either the objection is general, in which case it is an objection to the last two centuries of science, or it is not an objection at all.

But a defense of this kind arrives with obligations attached, and here is where the chapter turns from advocacy to audit.

If gender is a persistence structure, then it should exhibit the known signatures of persistence structures, and those signatures are specific, well documented across wildly different domains, and awkward. This is not an analogy to be admired and set down. It is a set of commitments, and I do not think Butler ever took them on, which is the single largest missed opportunity in this body of work.

Take three of the signatures. Maintained structures have thresholds: they hold their shape against disturbance up to a point and then reorganize abruptly rather than gradually. Maintained structures show hysteresis: the conditions required to undo a change are not the same as the conditions that produced it, so the path out is not the path in reversed. And maintained structures have characteristic failure modes: not any collapse at all, but a small number of ways that a system of that type comes apart, predictable from its architecture.

Each of these translates into a question about gender that is empirical rather than doctrinal. Are there thresholds in the enforcement of gender norms — levels of pressure below which a deviation is absorbed and above which the whole local arrangement reorganizes? The literature on stigma and on tipping points in social convention suggests there are, and the sizes have been estimated. Is there hysteresis — is the effort required to restore a norm after it has slipped greater or smaller than the effort that maintained it? Anyone who has watched a workplace dress code decay and then be reimposed knows the answer is not symmetrical. Are there characteristic failure modes — a small catalogue of ways that gendered arrangements come apart, recurring across societies with no contact between them? The historical record is not silent on this and has never been read this way.

That is a research program, it is falsifiable, and it is entirely absent from the tradition Butler founded, which spent its energies instead on textual interpretation and on the defense of its founder against misreadings. I think this is the real cost of the style, and I will come back to it. A theory of gender as maintained repetition should by now have produced measured maintenance costs, estimated thresholds, and a taxonomy of collapse. Instead it produced a very large secondary literature about what Judith Butler meant.

There is a reason our intuitions resist all of this, and it is worth naming rather than deploring. Perception is built to track persistence, because persistence is what matters to an organism: the things worth having a word for are the things that will still be there tomorrow. So the perceptual system delivers objects, not processes, and delivers them with an implied inner constitution that nobody ever observed. We see a flame and not a combustion region. We see a river and not a drainage regime. The intuition that lasting things must be made of lasting stuff is not a philosophical position anyone arrived at by thinking. It is a default setting, installed for good reasons, and its being useful is entirely compatible with its being wrong.

The law, which has to be practical about these things, gave up on the intuition centuries ago. A corporation persists through the replacement of every employee, every asset, every product, and every stated purpose, and remains the same legal person, liable for what it did before anyone currently inside it was born. A nation retains its treaty obligations across revolutions. A university that has changed its buildings, curriculum, faculty, and admission criteria beyond recognition still claims its founding date and, more to the point, its endowment. Nobody finds this mysterious in practice, because the practical question is never what a thing is made of. It is whether the arrangement has been continuously maintained.

Maintenance cost is also the quantity that predicts collapse, which makes it the most useful number nobody collects. Institutions do not usually fail because they are attacked. They fail because the subsidy that was paying for their upkeep is withdrawn or diverted, and the withdrawal is frequently invisible until the structure has already begun to deform. Monastic orders, trade guilds, local newspapers, and civic associations have all gone this way, and in each case observers at the time described the loss in terms of belief and enthusiasm, when the measurable thing that changed first was who was paying for the repetition.

One last observation, because it will matter in the closing chapters. The fixed-point picture explains something the ordinary essentialist picture cannot, which is why gender norms are enforced so ferociously and so pettily. If gender were a substance sitting inside people, it would take care of itself and there would be nothing to police; you do not need a social apparatus to ensure that iron rusts. The volume and persistence of the policing is not evidence against the constructed reading. It is exactly what the constructed reading predicts, since a structure maintained by repetition must be maintained, continuously, by somebody, or it will drift. Every playground correction, every remark about how a boy is holding something, every uniform requirement and gendered form field is a maintenance payment. The system tells you what it is by what it cannot stop doing.

Chapter Ten — Why Nothing Repeats Exactly

Copy a document by hand a thousand times and you will not get a thousand identical documents. You will get a family tree. Scribes skip lines when two lines end with the same word, insert marginal notes into the body, correct errors that were not errors, and standardize spellings toward whatever they were taught. Textual scholars can reconstruct the ancestry of surviving manuscripts precisely because the errors are inherited: a copy made from a copy carries its parent’s mistakes plus a few of its own. The transmission is the preservation. The transmission is also the corruption. There is no third process available that would give you one without the other.

This is the mechanism Butler took from Derrida, under the name iterability, and it is the piece of the apparatus that does the most work with the least recognition. The argument runs: for anything to be repeatable at all, it must be capable of functioning in a context other than the one it came from. A word that could only be used in the exact circumstances of its first utterance would not be a word. But a thing that can be lifted into a new context can always be lifted into an unforeseen one, and that possibility is not an unfortunate side effect of repetition. It is what repetition is. Every act of faithful reproduction carries, structurally and unavoidably, the possibility of deviation.

Biology arrived at the same conclusion from the other direction and paid for it in equipment. Genetic replication is astonishingly accurate and it is not perfect, and the interesting discovery of the last half century is that the imperfection is not a limitation of the machinery but a requirement of the system. Repair mechanisms that eliminated every error would eliminate variation, and a lineage without variation cannot respond to anything. Fidelity too low and the information dissolves; fidelity too high and the lineage is a museum piece awaiting a change in the weather. Living systems sit in a band between those failures, and the band is narrow enough to be measured.

The general principle deserves stating on its own, because it applies far beyond both examples. Any system that persists by copying is thereby a system that changes by copying. Common law works this way: each application of a precedent to a new case either extends it or narrows it, and there is no such thing as applying it identically, since the new case is by definition not the old one. Oral traditions work this way; so do rituals, recipes, jokes, accents, and the way a family does Christmas. The mechanism of stability and the mechanism of drift are one mechanism seen from two distances.

Butler’s use of this is the least appreciated part of the whole theory and the most structurally sound. If gender is maintained by repetition, then the possibility of its transformation does not have to be imported from anywhere. It does not require a free subject standing outside the system and deciding to do otherwise — which is fortunate, since the theory has spent considerable effort establishing that no such subject exists. Change is available in the joints of the repeating itself: in the fact that the norm must be reapplied to a new day, a new body, a new situation that does not quite match, and that each reapplication is an interpretation with room in it.

Notice what this rules out, because the exclusions are what give a theory content. It rules out the heroic account of social change, in which someone perceives the injustice of an arrangement and steps outside it. It rules out the purely mechanical account, in which structures reproduce themselves indefinitely until struck by an external force. And it rules out voluntarism, the reading Butler is perpetually accused of, because the room inside a repetition is not a menu and nobody chooses when to be handed a situation that does not fit.

Drift rates are not free either, and this is where the framework starts making predictions instead of observations. The rate at which a repeated practice deviates from its inherited form depends on how much enforcement it is receiving, on how visible each performance is to those who would correct it, and on how costly correction is to administer. Where surveillance is dense and sanctions are cheap, drift approaches zero and the practice looks eternal. Where performances are unobserved, or where the observers have no authority, drift accelerates until the practice is unrecognizable within a generation.

That prediction is testable against the last thirty years and it comes out well. Gendered practice has changed fastest in exactly the settings where enforcement is weakest: among strangers rather than kin, in cities rather than villages, in text rather than in person, in subcultures that have their own enforcement running in a different direction, and above all in the anonymous and unsupervised spaces that networked communication created in enormous quantity. It has changed slowest in institutions with dense mutual observation and strong sanctioning authority. One does not need to approve or disapprove of any of this to notice that the pattern is the one the theory implies, and that no rival account of the same period does as well on the geography.

Enforcement is not always a person. Much of it is technology, and the history of standardization is largely a history of machines that made drift expensive. The printing press froze spellings that had been fluid for centuries. The dictionary turned usage into a matter that could be looked up and therefore lost. Broadcast radio and television gave whole nations a single accent to measure themselves against, and the measurable convergence of regional speech in the twentieth century tracks the arrival of those instruments with uncomfortable precision. None of these devices was built to police anything. They policed by making one version of a practice enormously more available than the alternatives, which is the cheapest form of enforcement there is and the hardest to notice.

Sociolinguistics has studied precisely this and its findings map onto the framework better than anything in the philosophical literature. Changes that originate in prestigious groups and spread downward behave differently from changes that originate below the level of conscious attention and spread upward, and the second kind is both more common and more resistant to correction, because the people transmitting it do not know they are doing anything. Applied to gendered practice, this predicts that the durable changes will not be the ones announced in manifestos but the ones nobody noticed until they had already happened, and that is roughly what the last half century looks like from a distance.

Now the debit side, and it is a real one.

An account of change built on iterability explains beautifully that change is possible and says almost nothing about where it will go. The joints in a repetition are room for deviation in any direction. They do not supply a heading. Yet actual social change has headings — remarkably consistent ones, sustained over decades, in the same direction across societies with little in common — and a theory whose entire account of transformation is that repetition contains slack has no resources for explaining why the slack is taken up one way rather than another.

Butler’s answer, insofar as there is one, points to the norms themselves being multiple and in conflict, so that the pressure of one can be recruited against another. That is true and it is not enough. It tells you the terrain has slopes; it does not tell you which way is downhill. My own view is that the missing ingredient is cost — arrangements that are expensive to maintain drift toward cheaper configurations when the subsidy weakens, and one can often say in advance which configurations are cheaper — but that is an argument from outside this tradition, and offering it here would be tidier than honest. The gap is a gap.

There is a second debit and it is subtler. Iterability guarantees that exact repetition is impossible, and the theory sometimes treats that guarantee as though it were politically encouraging. It is not, in itself. Drift is directionless with respect to anybody’s interests: the same slack that permits a norm to loosen permits it to tighten, and the historical record contains a great many repetitions that deviated toward more policing rather than less. A structural feature that makes change possible is not an ally. It is a condition, equally available to everyone, including the people you would least like to see using it.

Which is, when you look at it squarely, the most useful thing this part of the theory has to offer, and the least comfortable. Butler’s apparatus contains no promise that the arc of repetition bends anywhere in particular. It contains the observation that nothing maintained by repetition can be secured permanently, and that observation applies with exactly equal force to arrangements one wants to overturn and to arrangements one has just finished building. A generation of readers took the first half of that sentence as a program. The second half was in the same paragraph.

Chapter Eleven — The Misreading That Would Not Die

Ask a well-read person who has not studied the material what Judith Butler thinks, and you will get, with remarkable reliability, some version of this: gender is a performance, like a role in a play, which people put on and could in principle take off. It is a clean, memorable, and comprehensively false summary of the position, and it has survived thirty-five years of correction by the author, several books written substantially to dislodge it, and the plain text of the original essay, which says the opposite.

The mechanics of the misreading are worth reconstructing, because they are the mechanics of most famous ideas and because the shape of the error is, with an irony I did not arrange, precisely the error the theory is about.

Start with the word. A performative, in the technical sense Butler took from speech-act theory, is an utterance that accomplishes something rather than reporting it. A performance, in the ordinary sense, is something an actor does in front of an audience. The two words share a root and nothing else that matters. In particular, a performance requires a performer who exists before and after it, who could have performed differently, and who is not identical with the role. That figure is the doer behind the deed — the entity Butler’s entire apparatus was constructed to do without. The vernacular sense of the word smuggled back in, in a single syllable, the thing the theory had foreclosed.

Add the venue. The essay appeared in a theater journal, which meant its first readers were people for whom performance was a technical term in the other sense, and who took the argument out into the world in their own vocabulary. Add the most vivid passage in Gender Trouble, which discusses drag, and which was consequently the only passage many readers retained. And add the prose, which was difficult enough that most readers could not check the summary against the source even when they had the source, so that the summary became the thing that circulated. Four independent causes, all pushing in the same direction, none of them requiring anybody to act in bad faith.

Butler published a book in 1993 to fix this. Bodies That Matter states, about as directly as that prose gets, that performativity is not a singular act, that the repetition in question is not performed by a subject, and that the repetition is instead what enables a subject to exist. It insists that the process is constrained, ritualized, compelled, and enforced under threat, and that treating it as a daily choice gets it backward. The book is a correction. It did not correct. If anything it accelerated the misreading, because a book written to say you have misunderstood me is a book that announces there is something juicy to misunderstand.

The audit reading of this episode is the pleasing part. What the misreaders did was to insert a referent into a position the theory had declared empty. They could not hold in mind the idea of a doing with no doer, so they supplied one — the offstage performer, waiting in the wings with their real self intact — and then attacked the resulting position, which was indeed absurd. That is exactly the operation the theory diagnoses in the case of gender: a coherent series of acts generates the impression of an agent behind them, and the impression is then read back as the cause. The critics performed the theory in the act of rejecting it. I do not know a cleaner instance of a body of work being confirmed by the manner of its dismissal.

None of which gets Butler off. The share of responsibility for a misreading depends on how hard the text made checking, and this text made checking very hard indeed.

Consider what a reader has to do to verify a summary. They must obtain the source, locate the relevant passage, and determine whether the summary matches. If the prose is clear, this takes a few minutes and the misreading dies. If the prose requires substantial prior training, the reader cannot perform the check, and is obliged to trust a summarizer — who was in turn trusting a summarizer. A text that cannot be checked by its own audience has delegated its meaning to its intermediaries, and it has no grounds for complaint about what they do with it. This is not a moral failing. It is a structural fact about transmission, and an author who chooses difficulty is choosing it.

Nor is the case unusual, which should be some comfort. Survival of the fittest, a phrase Darwin did not coin and adopted reluctantly, has spent a century being read as a doctrine about strength and struggle rather than a tautology about differential reproduction. Relativity is popularly understood as the claim that everything is relative, when its content is the identification of what is invariant. The uncertainty principle is routinely explained as the observer disturbing the system, which is not what it says. Godel’s incompleteness results are invoked in support of positions ranging from theology to postmodernism, most of which they do not touch. In every case the pattern is identical: a technical result acquires a vernacular near-synonym, the near-synonym is more interesting than the result, and the near-synonym wins.

The drag chapter deserves a word, since it did more damage than any other passage. Butler discusses drag not as a model for gender but as a case that exposes something about all gender — namely that an imitation can reveal the imitative structure of what it copies, since there is no original for either to be faithful to. That is a subtle claim, it occupies a few pages, and it was placed near the end of a difficult book where the exhausted reader arrives grateful for a concrete image. What travelled was the image. A generation learned that Butler thinks gender is like drag, which inverts the argument: drag is illuminating precisely because it is not special.

Butler’s own trajectory suggests some eventual recognition of the cost. The 2024 book on the anti-gender movement was widely described as the most accessible thing they have written, an intervention aimed at a broad readership rather than at the profession. Whether that represents a change of view about difficulty or simply a change of target, it is a different instrument, and it arrives roughly thirty-five years after the moment when it would have made the most difference. This is not a reproach so much as an observation about how careers work: the accessible book is written when one has something to defend, and by then the misreading has had a generation to settle.

There is a general law lurking here and I will state it, at some risk. Misreadings do not spread because they are simpler than the original. They spread because they are more useful to more parties. The costume reading of Butler was useful to admirers, who could take it as a program of self-invention and a promise of freedom. It was equally useful to opponents, who acquired an absurd position to demolish and a demonstration that the academy had lost its mind. Both constituencies were better off with the misreading than with the text, which offers the first group very little liberation and denies the second group its easy target. A version of an idea that gives everybody something to do is close to unkillable, and its survival has nothing to do with whether it is right.

Which produces a question I have no comfortable answer to. If corrections do not work — and the evidence that they do not is now substantial — what should an author do? Write more clearly at the outset, obviously, but this is advice delivered to the past. Beyond that the options are poor. Restating in plainer terms reaches a different audience than the one holding the misreading. Denouncing the misreaders makes them louder. Ignoring it leaves the field to them. The one intervention with a track record is a rival formulation memorable enough to displace the first, which requires either a gift for phrases or a great deal of luck, and which Butler, whatever the other gifts, does not have.

I want to close with the part of this that should trouble anyone who writes anything. The misreading of Butler is not a story about a careless public. The people who transmitted it included professors, journalists, and serious critics on both sides, most of whom had the book, some of whom had read it. What defeated them was not laziness but the ordinary economics of attention: checking is expensive, the summary was available, the summary was coherent, and there was no visible cost to accepting it. Every one of us is running that calculation continuously about nearly everything we believe, and the number of positions any of us holds on the strength of a summary we have never checked is not a number to dwell on before bed.

Chapter Twelve — The Power You Are Made Of

In 1959 two psychologists ran an experiment that has been irritating people ever since. Volunteers were asked to join a discussion group, but first had to pass an initiation. For some the initiation was trivial; for others it was designed to be genuinely embarrassing. The discussion that followed was, by arrangement, as dull as the experimenters could make it. Those who had suffered the severe initiation rated the group as significantly more interesting and worthwhile than those who had walked straight in. The finding has been replicated, argued about, and refined for six decades, and the core result has not gone away: raising the cost of entry raises the attachment of those who paid it.

This is the phenomenon Butler spent the 1990s trying to give a general theory of, and the theory is the subject of a book called The Psychic Life of Power, published in 1997. It is the least read of the major works and the one I would recommend first to a skeptic, because it addresses a problem that everybody recognizes and no available account handles well.

The problem: why do people so often defend, love, and reproduce the very arrangements that constrain and damage them? The stock answers are unsatisfying to the point of insult. False consciousness says they have been fooled, which requires an account of who is doing the fooling and why it works so evenly. Self-interest says they are getting something out of it, which is sometimes true and frequently not. Stupidity is not an explanation of anything and has the additional defect of being flattering to the person offering it.

Butler’s answer starts from a structural observation and it is, I think, correct. The arrangements that subordinate a person are frequently the same arrangements through which that person came to exist as a person at all. A child does not first exist and then get socialized. The child becomes a self in the course of being addressed, named, categorized, corrected, and made legible by a specific set of arrangements, and those arrangements are not optional extras: they are the conditions of having a self to be constrained. To reject them wholesale would be to reject the conditions of one’s own existence, which is not a course of action available to anyone.

The pieces of machinery are borrowed and openly credited. From Althusser comes the scene of hailing: a voice calls out on the street, a person turns, and in the turning becomes the one who was addressed. Nobody checked in advance whether the call was meant for them; the turning is the mechanism by which one becomes a subject of the authority that called. From Foucault comes the twin sense of the French word for subjection, which means both being subjected to power and being made into a subject. And from Freud comes the account of melancholia, in which a lost attachment is not relinquished but absorbed, so that the structure of the self comes to contain the shape of what it could not have.

The concept doing the most work is passionate attachment. An infant must attach to whoever is available, because attachment is not a preference but a precondition of survival and of development into anything at all. If what is available is harmful, the attachment forms anyway. It cannot be traded for a better option, since there is no self yet capable of shopping. The result is a person whose capacity for attachment is structured around the very relation that damaged them, and who will, later, seek out and defend arrangements with a familiar shape, without any of this being available to introspection as a choice.

What recommends this account, in audit terms, is that it does not float free of evidence. Attachment research is a large empirical literature with instruments, longitudinal data, and cross-cultural replication, and its central finding — that children form attachments to unresponsive and even frightening caregivers, and that the resulting patterns persist into adult relationships — was established independently of any philosophy and is not seriously disputed. The initiation experiments are a separate empirical line reaching a compatible place. Butler’s contribution is not the data. It is the generalization of the mechanism from families to institutions, which is a real intellectual move and one that can be checked against how institutions actually behave.

And they do behave that way. Every organization that wants durable loyalty imposes costs at entry, and the more effective ones make the costs unpleasant rather than merely expensive. Basic training, medical residency, the novitiate, the fraternity, the doctorate, the fast-track program that eats five years of a life: in each case the sequence is identical, and in each case the people who came through it will defend the arrangement more fiercely than any outsider can understand. They are not defending the hazing. They are defending the structure through which they became who they are, and there is no version of themselves available for inspection that did not go through it.

The experiment I opened with was run by Elliot Aronson and Judson Mills, and the mechanism they proposed was dissonance: having paid a cost, one revises the estimate of what it was paid for, since the alternative is to hold that one suffered for nothing. Butler’s account is deeper in one specific respect. Dissonance explains a revision of judgment in a person who already exists. Passionate attachment concerns the case where no such person exists yet, and where the arrangement in question is what produced the one now doing the judging. The two accounts are compatible; they operate at different stages, and the earlier one is the harder to reverse.

The economics of exit sharpen this. Where leaving is cheap, dissatisfaction expresses itself as departure and the institution receives a clear signal. Where leaving is expensive — because the skills do not transfer, because the community is the only one available, because one’s entire social existence is inside — dissatisfaction has nowhere to go and reliably converts into louder professions of loyalty. From outside this looks like enthusiasm. It is frequently the shape that trapped disagreement takes, and any institution that has managed to raise its exit costs high enough will find itself surrounded by evidence of devotion that it should not believe.

There is an uncomfortable corollary that Butler draws and most readers hurry past. If the self is formed through subjection, then resistance is also formed there — it uses the same materials, speaks the same language, and takes shape inside the same arrangements it opposes. There is no pristine standpoint outside. What passes for a view from nowhere is a view from a position that has forgotten it has one.

This lands hard on a great deal of political self-description, in every direction. Movements characteristically present themselves as speaking from outside the system they oppose, and the presentation is not merely a rhetorical convenience; it is what allows the movement to regard its own methods as clean. Butler’s account withdraws that permission. Whatever tools are being used were made in the workshop under attack, and the fact that they are now pointed in a different direction does not change their provenance or their tendencies.

An objection arrives at once, and it is the serious one: if power forms both the subordination and the resistance, what distinguishes emancipation from more of the same? Butler’s answer is that emancipation cannot mean escape, since there is nowhere to escape to, and must instead mean reworking from within — occupying a norm in a way that alters what it can do. This is coherent. It is also, and I think this criticism sticks, extraordinarily difficult to operationalize. It does not tell you which reworkings are improvements, and the theory as stated contains no resource for saying so, because a criterion for improvement would have to come from somewhere, and every available somewhere is inside.

That gap is the substance of the most famous attack ever made on this body of work, delivered by Martha Nussbaum in 1999, and it is important enough to have its own chapter. I flag it here only to record that the gap is visible from inside the framework and is not an invention of hostile readers. One can hold, as I do, that the descriptive account of subject formation is the strongest thing Butler has written, and simultaneously that the normative superstructure built on it is the weakest, and that the second follows from the first rather too easily.

The chapter ends where the initiation experiments end, which is not with a moral. Nobody has ever proposed abolishing costly entry, because the loyalty it produces is not a side effect but the point, and institutions that need loyalty will keep producing it by the only method that reliably works. What the research and the theory jointly establish is that attachment to a constraining arrangement is not a defect of judgment to be corrected by better information. It is what having been formed by something feels like from the inside, and the people who have it are not confused.

Chapter Thirteen — What the Body Refuses to Be Told

The objection arrived from every direction at once, and Butler has said as much: whatever the audience, whatever the country, whatever the political sympathies of the questioner, somebody would eventually stand up and ask about the body. What about pain? What about hunger? What about the fact that a person can be killed? Whatever else is negotiable, the body appears to be the thing that does not care what anyone has decided, and a theory that seemed to make gender a matter of meaning looked as though it had mislaid the one item nobody can argue with.

Bodies That Matter, published in 1993, is the response, and the first thing to say about it is that it does not do what its title promises to a hopeful reader. It is not a book in which the body is restored to its rightful place as the bedrock beneath the discourse. It is a book which argues that the demand for such a bedrock is itself the thing to be examined, and it is written in prose considerably harder than the book it was correcting, which tells you something about how corrections go.

The central concept is materialization, and it is a genuinely careful piece of work under the difficulty. Butler’s claim is not that bodies are made of language, a proposition nobody has ever held and which would be idiotic. The claim is that the materiality of a body is never encountered outside the terms through which it is grasped, and that those terms are not a coating applied to a pre-existing surface but part of the process by which a body comes to be a body of a particular sort — sexed, raced, legible, countable, treatable, mournable. Material things happen; matter also mattered in the other sense, which is the pun the title is built on and the argument the book is making.

Take the process seriously for a moment in a setting where it is uncontroversial. A collection of cells becomes a tumor when a classification system says so, and the classification has changed repeatedly within living memory. Certain growths that were once cancers are now not cancers; the cells did not change, the thresholds did, and hundreds of thousands of people accordingly did or did not have a disease. Nobody thinks this means tumors are made of paperwork. What it means is that the passage from a physical fact to a fact that has consequences runs through a system of classification which is not itself a physical fact, and which has a history, and which could have been otherwise.

Butler’s proposal is that sexed embodiment works this way, only more so, because the classification is applied at birth, is applied to everybody, is reapplied continuously for a lifetime, and is enforced. This is what the book means by a constitutive constraint: not a limit imposed from outside on a body that existed beforehand, but a constraint that is part of what makes the body available as an object at all.

Does the book close the hole left by Gender Trouble? Here I have to give the answer I promised several chapters ago, and it is a split verdict.

The argument is consistent. Read carefully, with attention to what is being claimed about knowledge and what is being claimed about existence, Bodies That Matter says something defensible and does not say anything absurd. Nothing in it denies that bodies bleed, break, or die independently of what anyone thinks. It says that the significance those events have — including which of them count as the same kind of event — is never available in a raw state, and that a theory which helps itself to raw significance is helping itself to something nobody has ever had.

The prose is another matter, and the problem is not that it is hard. The problem is that it consistently uses verbs of production where the argument requires verbs of access. Bodies are said to be materialized, produced, constituted, brought into being. An epistemic thesis — you cannot get at it except through this — is stated in the grammar of an ontological one — this is what makes it. And that choice is not incidental decoration, because the grammar of production is exactly what gives the book its force and its notoriety. Say that our access to sexual difference is theory-laden and you have written a competent paper in the philosophy of science that nobody will burn in effigy. Say that sex is materialized through the reiteration of norms and you have written something that sounds as though the world is being made by talking.

My conclusion is that the ambiguity is a rhetorical debt financed by an ontological overdraft. The strong grammar attracted the readership; the defensible content is the weaker claim; and when the strong grammar was taken at face value, the author responded by pointing at the weaker claim, which was indeed there. This is not fraud. It is what happens when a writer discovers a register that works and does not fully price what the register commits them to. The audit finding is narrow and I will stand behind it: the argument is sound, the prose systematically overstates it, and the overstatement is load-bearing for the book’s reception.

Medicine supplies the most instructive parallel, because there the history is documented and nobody has an ideological stake in denying it. Categories of illness are revised, invented, and abolished, and the revisions are consequential for real bodies. Hysteria organized the treatment of women for a century and no longer exists. Homosexuality was a diagnosis in the American psychiatric manual until 1973, and its removal was accomplished by a vote, which is the sort of fact that theorists cite triumphantly and clinicians find unremarkable, since committees are how classifications get changed in every field including physics. In each case people were treated, hospitalized, and in some places imprisoned on the strength of a category that later ceased to exist. The suffering was not constructed. The category was.

This is the register in which Butler’s claim is defensible and, I think, obviously true: not that bodies are effects of language, but that the passage from a bodily state to an institutional fact runs through a classification, and that the classification has authors, a history, and consequences. Where the prose gets into trouble is in eliding the distance between the two, so that a claim about the second reads as a claim about the first. A reader who has been trained to notice that distance will find Bodies That Matter a careful book. A reader who has not will find it an outrageous one, and will be able to quote sentences in support.

There is one case that tests all of this harder than any argument, and it belongs here rather than in a chapter about theory.

In 1966 a Canadian infant boy, born the previous August, suffered a catastrophic injury to his genitals during a botched circumcision at the age of eight months. On the advice of a prominent researcher, his parents were persuaded to raise him as a girl, with surgery and later hormones, and to conceal the history from him. The case was reported for years in the literature as a successful demonstration that gender identity follows rearing. It was not successful. The child was persistently unhappy, rejected the assignment, and on learning the truth in adolescence returned to living as male. He took the name David, married, became a stepfather to three children, told his story publicly, and died by suicide in 2004.

That history has been used as a weapon by everybody. For the naturalists it is the clean demonstration that an inner sexed identity exists, is innate, and cannot be overwritten by upbringing. Against them it is pointed out that the upbringing in question involved surgery without consent, systematic deception, invasive examinations, and a research program with a reputation at stake, and that inferring the failure of socialization from the failure of that is like inferring the limits of medicine from a case of malpractice.

Butler’s discussion in Undoing Gender does something the partisans do not, and it is the best thing in that book. It declines to award the case to either side, and asks instead what was being done to a person by the demand that he be intelligible — by the requirement that he settle, definitively, into a category that would make his existence administratively and conceptually manageable for the adults around him. The interventions were not merely medical mistakes. They were attempts to make a life legible, and the legibility was for other people.

Butler returned to this territory much later, in a short book written during the pandemic, and the setting did the argument a favor that thirty years of theory had not. A virus makes the point about bodies without any assistance from philosophy: we breathe each other’s air, our porousness is not optional, and the distribution of who could isolate and who had to keep working converted an abstract claim about differential vulnerability into a mortality statistic within eighteen months. Whatever one thinks of that book, the pandemic settled a version of the old objection in Butler’s favor. The body turned out to be exactly as material as the critics insisted and exactly as unequally exposed as the theory had claimed, and the two facts did not conflict.

I do not think this exhausts the case, and I am wary of a reading that turns a specific catastrophe into an illustration of a thesis, which is a thing philosophers do to the dead with a regularity that ought to embarrass the profession. But the observation stands and it is the one durable result. Whatever the truth about innateness — and the case does not settle it, since a single history under those conditions cannot settle anything — the pressure to be classifiable was real, was applied by every institution involved, and continued for decades after the classification had visibly failed. That pressure is a fact about the system rather than about the person, and it is precisely the sort of fact that the framework under audit was built to make visible.

Chapter Fourteen — Drag Is Not the Argument

An illustration is not evidence, and the difference has never mattered more to anyone’s reputation than it did to Judith Butler’s.

Near the end of Gender Trouble there is a discussion of drag, and it occupies a few pages of a book that spends most of its length on Freud, Lacan, Wittig, Kristeva, and Foucault. It has nonetheless become, for most people who know anything about the book at all, its content. This is unfortunate for the reasons already covered, and it is doubly unfortunate because the argument being illustrated is genuinely interesting and almost never stated correctly.

The argument is not that gender is like drag. It is a claim about imitation. When a drag performance imitates a gender, the ordinary assumption is that a copy is being made of an original. Butler’s proposal is that inspecting the copy tells us something about the supposed original: that it too consists of a set of gestures, postures, vocal habits, and styles of dress, assembled and maintained; that there is no further thing it is being faithful to; and that everyone performing a gender in the ordinary way is producing an imitation for which no original exists. The point of the drag case is that it makes the machinery briefly visible, in the way that a slowed film makes visible a movement the eye normally smooths over.

So drag is not offered as a model for gender, nor as an especially liberated form of it, nor as a recommendation. It is offered as a case where the structure shows. This distinction is the entire content of the passage and it is what got lost, with the result that a great many people came to believe Butler had proposed that we are all in costume, which is the theatrical reading again, arriving by a different door.

What makes this worth a chapter rather than a footnote is that Butler noticed the problem quickly and did something unusual about it: in Bodies That Matter, three years later, the drag case is revisited and substantially complicated, and the complication cuts against the celebratory reading that Butler’s own admirers had adopted.

The occasion is a documentary about the drag ball culture of Harlem, released in 1990, which had become the standard reference for anybody wanting a vivid example. Butler’s discussion of it declines to treat the balls as straightforward subversion. It notes that the performances are frequently aspirational — that participants are performing wealth, whiteness, and respectability as much as femininity, and that these are the norms of a world that excludes them rather than a repudiation of it. It observes that imitation can consolidate a norm as easily as it can destabilize one, and that nothing in the structure guarantees which. And it dwells on the murder of one of the film’s participants, a transgender woman killed during the making of the film, as evidence that the theatrical frame does not extend past the edge of the ballroom.

Butler was also criticized on this by bell hooks, whose essay on the same film argued that the celebration of the balls in academic and progressive circles glossed over race and class, and that the camera’s relationship to its subjects was itself extractive. The criticism landed. It is visible in Butler’s own later treatment, which is more careful about who is watching whom and at what cost, and this is one of the very few episodes in this literature where a public criticism produced a visible correction rather than an entrenchment.

There is a further complication that Butler eventually conceded and that deserves recording, since it is rare for a theorist to abandon a term that has served them well. The whole vocabulary of subversion — the idea that a practice can be assessed by whether it destabilizes a norm — turns out to be very hard to use. Nothing is subversive in itself; the same act consolidates in one setting and disrupts in another, and the assessment can only be made after the fact, at which point it is a description rather than a criterion. A concept that can only ever be applied retrospectively is not doing the work its users want it to do, and by the late 1990s Butler was saying so.

It is worth asking why vivid cases capture arguments so reliably, since it happens to every field with a public. Teaching is the culprit. An idea must be transmitted, transmission requires an example, and the example that transmits best is the one that is concrete, memorable, and slightly transgressive — which is to say, the one least likely to be typical. Every discipline therefore ends up represented in the public mind by its most photogenic exception. Economists are pursued by a hypothetical about a trolley problem’s cousin; psychologists by a prison experiment whose methodology has since collapsed; evolutionary biologists by a peacock’s tail. The examples were selected for pedagogy and are then read as findings.

The general methodological point is the one worth taking away, and it applies far beyond this material.

Cases do different jobs, and confusing the jobs is one of the commonest failures in argument. A case can serve as an existence proof: it shows that something is possible, and one instance is sufficient. A case can serve as a typical instance: it shows what the general run of the phenomenon looks like, and for this a single case is nearly worthless without evidence that it is representative. A case can serve as a limiting or extreme instance, which is informative about boundaries and misleading about the middle. And a case can serve as an illustration, which is a teaching device carrying no evidential weight whatever.

Butler used drag as the third and fourth of these, and the readership took it as the second. That is not a small slip. Treating an unusual case as typical produces exactly the misestimate you would expect: if the visible, deliberate, theatrical instance is taken as the model, then the whole phenomenon looks deliberate and theatrical, which is the misreading in one sentence.

The failure recurs everywhere and it is worth collecting a few instances, because seeing the shape repeated is what makes it noticeable in the wild. Optical illusions are used to argue that perception is unreliable, when their interest lies precisely in being exceptions to a system that is otherwise extremely reliable. Split-brain patients are used to argue that the unified self is an illusion, when they are a small number of people who underwent a radical surgery. Savant abilities are used to make claims about the general architecture of cognition. In each case a rare configuration that exposes a mechanism is silently promoted to being what the mechanism usually does.

There is a reason this particular error is so hard to resist. Cases that expose a mechanism are, by construction, the memorable ones — that is why they expose anything. The ordinary operation of a system is invisible; that is what ordinary means. So the examples available for thinking with are systematically drawn from the tails of the distribution, and anyone reasoning from available examples will overweight the tails without noticing. This is not a defect of any particular argument. It is a standing bias in the supply of examples, and the only remedy is the tedious one of asking, each time, what job this case is doing.

For the audit, then, the ledger on drag reads as follows. The claim being illustrated — that imitation can reveal the absence of an original — is a real claim and is not established by the illustration. The illustration was chosen for vividness and did its job too well. The author revised the treatment when the costs became apparent, under criticism, and said so in print. And the revision reached almost none of the people who had learned the original version, which is by now the least surprising sentence in this book.

Chapter Fifteen — Sex, and the Word That Cannot Be Settled

A biologist, a physician, a lawyer, a demographer, a sports administrator, and a priest walk into a room and are asked to define sex. This is not the opening of a joke; it is a description of most of the last decade, and the reason it goes badly is not that any of them is ignorant. It is that each of them needs the word to do a different job, the jobs are genuinely different, and no single definition does all of them.

Begin with the most defensible general definition, because fairness requires starting with the strongest version of a position and because it is frequently caricatured. Across the eukaryotes, sexual reproduction involves two gamete types: small and mobile, large and resource-bearing. The organisms organized around producing one or the other are called male and female. This definition is genuinely binary — there is no third gamete size, and the reasons for that are well understood as a matter of evolutionary dynamics rather than convention. It applies across the enormous range of organisms that reproduce sexually, from fungi to elephants. It is not culturally local, it was not designed to settle any human argument, and it is the definition working biologists use because it is the one that does biological work.

It also has limits, and stating them is not an attack on it. It classifies reproductive strategy, and its application to an individual organism proceeds by inference from developmental pathway rather than by direct inspection of gamete production, which is why it applies without embarrassment to a child, to a post-reproductive adult, and to a sterile individual. Those inferences are secure in the overwhelming majority of cases and are not secure in all of them. And the definition is silent, by construction, about everything social: it tells you nothing about what a person should be called, which facilities they should use, or how they experience themselves, because it was not built to and does not pretend to.

Now the clinical picture, which is where the trouble enters. Human sexual development runs through several stages — chromosomal complement, gonadal development, hormone production, hormone sensitivity, internal and external anatomy — and these are ordinarily concordant. Ordinarily is not always. There are well documented conditions in which a person has one chromosomal pattern and the anatomy typical of the other; in which the tissues do not respond to the hormones present; in which the pathway takes an unusual turn at one stage and a typical one at the next. These are not rumors or edge cases invented for argumentative convenience. They are described in every endocrinology textbook, they have names, and the people who have them exist.

How common? Here is where the audit earns its fee, because the answer that circulates in public depends almost entirely on a declaration nobody states. A widely quoted figure puts the frequency of intersex conditions at close to two percent of births, which would make it about as common as red hair. A rival figure, also from clinicians, is around two hundredths of one percent — roughly a hundredfold smaller. The two numbers are not the product of different data. They are the product of different decisions about what counts, with the larger figure including conditions in which the anatomy is entirely typical and only a chromosomal or hormonal variation is present, and the smaller one restricting the count to cases where anatomy at birth is genuinely ambiguous.

Neither figure is dishonest and neither is a measurement in the sense this book has been using. Both are consequences of a boundary decision applied to real cases, and the boundary decision is doing more work than any observation. This is the purest example in the whole book of a declared component reported as a finding, and it is quoted daily, in both directions, by people who would be shocked to learn that the number they are citing is a definition wearing the clothes of a statistic.

The general principle is worth stating plainly, since it dissolves a good deal of the shouting. A definition is a tool selected for a purpose, and disputes about what a thing really is are usually disputes about which purpose should govern, conducted in a vocabulary that hides this. Consider how differently the word must behave depending on the job. A geneticist studying inheritance needs the gamete definition and needs nothing else. A physician deciding a screening schedule needs anatomy and hormones, because the question is which tissues the patient has. A lawyer applying a statute needs a criterion that is administrable by clerks and stable enough to litigate. A demographer needs whatever was recorded on a form thirty years ago, since that is the data. A sports body needs whatever best tracks the specific performance advantages at issue in that specific sport, which is an empirical question with different answers in weightlifting and in shooting.

Each of those is a legitimate purpose. Each generates a slightly different partition. The partitions coincide for the overwhelming majority of people, which is why the word works at all in daily life, and they come apart exactly at the cases that end up in the news. Insisting that one of these is the real definition and the others are political corruptions of it is a move available to anybody, in any direction, and it is never an argument. It is the assertion of a purpose, and the honest form of that assertion says which purpose and why.

Sport is where these purposes collide most publicly, and it is worth a paragraph precisely because it is treated as a proxy war rather than as the empirical question it partly is. What matters for competitive fairness is not a definition but a set of physical variables, and which variables matter differs enormously by event: reach and mass are decisive in combat sports, largely irrelevant in equestrian events, and act in complicated ways in endurance running. The honest form of the question is therefore narrow and technical — for this event, which attributes confer advantage, by how much, and how are they distributed — and it has different answers in different sports. The dishonest form is to seek a single global criterion and then defend it as though it followed from biology, which no criterion does, since biology does not know what a fair race is.

Legal sex is a different animal again and its history is instructive. The field on a passport exists because states needed to identify people, and it was inherited from parish registers designed for entirely other purposes. It is an administrative convenience, revisable by statute, and it has been revised — many jurisdictions now permit amendment, some have added a third option, and a few have begun asking whether the field earns its place at all, given that it is a poor identifier and that biometric methods have made it redundant for the original purpose. Whatever one thinks should happen next, the category on the document was never a measurement. It was a record-keeping decision, made by clerks, for reasons that have largely lapsed.

This is where Butler’s contribution can be stated precisely, stripped of the vocabulary that makes it sound larger than it is. The claim is not that the biology is wrong or negotiable. The claim is that the passage from a biological classification to a social entitlement, obligation, or identity is not a deduction. It requires additional premises about which purposes govern which contexts, and those premises are contributed by us, not by the organism. Whenever someone moves from a fact about gametes to a conclusion about a bathroom, a pronoun, a prison, or a marriage, the additional premises are in the argument whether or not they are stated, and the argument is only as good as they are.

That is a modest and correct observation, and it is compatible with a very wide range of political conclusions, which is exactly why nobody in the public argument likes it. It does not tell you what the additional premises should be. It tells you they are there, that they were chosen, and that the choice is where the disagreement actually lives.

Two temptations should be named, one on each side, because the audit has to cut both ways or it cuts nothing.

The temptation on one side is to treat the existence of unusual cases as if it dissolved the category. It does not. A cluster with fuzzy edges is still a cluster; twilight does not abolish the difference between day and night; and a classification that is correct for the overwhelming majority of instances is an extremely useful classification. Anyone who moves from the existence of atypical development to the conclusion that sex is not a real biological feature has performed an inference that no biologist would accept and that the atypical cases do not support.

The temptation on the other side is to treat the biological fact as though it settled the social questions by itself, so that anyone raising the additional premises is introducing politics into a matter of science. But the additional premises are not optional and were never absent. A rule about who may compete in a sporting category, or which prisoners are housed where, is a rule made by people balancing several goods, and it will be defended by an argument containing empirical claims and value judgments in some proportion. Presenting it as a straightforward consequence of chromosomes conceals the value judgments rather than eliminating them, which is the same offense as the first, committed in the opposite direction.

What the audit recommends is unglamorous and, I think, correct: say which definition you are using, say what you are using it for, and accept that a definition serving one purpose well may serve another badly. Most of the participants in this argument are perfectly capable of doing that. Very few of them will, because a definition that announces its own purpose is a definition that can be negotiated, and the object of the exercise for most parties has been to obtain one that cannot.

Chapter Sixteen — Two Descriptions That Cannot Be Merged

There is no way to draw a map of the earth on a flat sheet of paper that preserves both shapes and areas. This is not a technical limitation awaiting a cleverer cartographer. It was proved in the eighteenth century by Gauss, and it follows from the fact that a sphere and a plane have different intrinsic curvature: any flattening must tear or stretch, and the choice of what to sacrifice is forced. The Mercator projection preserves angles, which is why it navigated ships for four centuries, and inflates Greenland to the size of Africa. Equal-area projections preserve size and mangle shape. Every world map you have ever seen is a decision about which lie to tell.

Notice that this does not make maps unreliable. The Mercator projection does not mislead a navigator; it does exactly what a navigator needs. It misleads a schoolchild forming an impression of the relative size of continents, because it is being used for a job it was not built for. The map is not wrong. The map is a projection, and a projection loses something by mathematical necessity, and the only defense is knowing which one you are holding.

Physics has a more famous version and a stranger one. Light behaves, in some experiments, exactly as a wave: it interferes with itself, produces fringes, diffracts around edges. In other experiments it behaves exactly as a stream of discrete quanta, arriving one at a time, each depositing its energy at a point. Both descriptions are correct, in the strong sense that each makes precise predictions that are borne out. They cannot be combined into a single picture that a person can visualize, and a century of trying has not produced one. Physicists did not resolve this. They learned to say which description applies to which experimental arrangement, and to stop demanding a portrait.

The general situation deserves a name and I will use this one: incompatible truths. Two descriptions of the same domain are incompatible in this sense when each survives its own tests, each makes successful predictions, and no single description reproduces both without loss. The condition is not rare and it is not a sign that something has gone wrong. It is the ordinary consequence of a domain being richer than any single representation of it.

Thermodynamics and statistical mechanics stand in this relation for practical purposes: temperature is a property of a system with no meaning for a single molecule, and the molecular account cannot say what a hot cup of coffee is without importing the very statistical notions it was meant to replace. Medicine lives with a sharper case daily, since the population description and the individual description of the same illness are both correct and give different advice: a treatment that reduces mortality across ten thousand patients tells you almost nothing decisive about the one in front of you, and the physician who reasons only from the trial and the physician who reasons only from the patient are both making a recognizable mistake.

Now the crucial guard rail, because this is the point at which the whole idea gets stolen and abused. Incompatibility is not relativism, and the difference is not subtle. The claim is not that all descriptions are equally good, or that truth is a matter of perspective, or that anyone’s account of anything is as valid as anyone else’s. Most descriptions are simply wrong: they make predictions that fail, they contradict observations, they do not survive their own tests. The incompatibility relation holds only among descriptions that have each independently earned their place, and the number of those, in any domain, is very small. To claim the protection of complementarity, a description first has to pass the tests it would have to pass anyway. Almost nothing does.

With that established, consider the two descriptions of gender that will not merge.

The first is from the inside. A person has an account of their own gendered life which is available to them and to nobody else: how it feels, when it began, what it survived, what it cost, what would be lost if it were surrendered. This account is not a hypothesis about mechanisms and does not compete with one. It has authority of a specific kind — nobody else is in a position to report it — and it has known limitations, since first-person reports about causes are unreliable in every domain that has been checked, and the person is not an observer of their own formation.

The second is from the outside. Populations can be counted, correlations measured, institutions compared, historical variation documented. This account has authority of a different kind — it can see patterns invisible from any single position — and correspondingly different limitations, since a distribution says nothing determinate about any individual in it, and a category constructed for counting will smooth over precisely the features that matter most to the people being counted.

Each of these descriptions makes claims the other cannot evaluate. Each is right about things the other gets wrong. And the demand that one be reduced to the other is the engine of nearly all the bitterness in this subject.

Watch how the demand operates in each direction, since both are common and both are errors of the same type. From the outside inward: your reported experience is an epiphenomenon of social forces we can document, therefore what you take yourself to know about yourself is a symptom to be explained rather than a report to be believed. From the inside outward: my experience is the fact of the matter, therefore your statistics, your history, and your comparative evidence are impertinent at best and hostile at worst. Each move takes a description that is valid in its own domain and demands it be surrendered to the other. Each is met, reliably, by the equivalent move in return.

The correct handling is the one physics arrived at, and it is less satisfying than either party wants. Say which description applies to which question. Whether a person experiences their gender as deep, early, and unchosen is a first-person question and the first-person report is the evidence. Whether the availability of particular gender categories varies across societies and centuries is a third-person question and the historical record is the evidence. Whether a specific policy will produce a specific outcome is a third-person question and no amount of first-person testimony from any side settles it. These are different questions. They are constantly asked in the same sentence.

Reduction fails elsewhere in the same way and it is worth seeing the pattern outside contested territory. Chemistry has not been absorbed into physics, seventy years after everyone agreed in principle that it should be: the concepts chemists actually use — bonds, aromaticity, reaction pathways — are not eliminable in favor of the underlying equations, which for any interesting molecule cannot be solved anyway. Biology has not been absorbed into chemistry. Nobody regards these as scandals. They are the normal condition of a world in which descriptions are built at the scale where the regularities are, and the regularities are at different scales.

There is also a practical version of this that everyone has met, which is what happens when an institution is obliged to adopt a single projection. A form has one field. A database has one column. Whatever richness exists in the domain has to be compressed into a value that a clerk can enter and a computer can sort, and the compression is not a philosophical position but an engineering constraint. This is worth remembering because a great deal of what gets experienced as ideological aggression is a database schema, designed by nobody in particular, twenty years ago, for a purpose that has since changed.

There is an important asymmetry that the complementarity framing can conceal, and I want to name it before anyone uses this chapter as a comfortable ceasefire. Living inside an incompatibility is not equally costly for everybody. The person whose first-person account is routinely overridden by third-person categories pays for the mismatch continuously, in encounters with forms, officials, doctors, and strangers. The person whose account happens to coincide with the available categories pays nothing and does not experience the situation as a projection at all — which is precisely the condition of holding a map without knowing it is one. The structure is symmetrical; the incidence is not; and anybody advocating patience with unresolved tension should be clear about who is being asked to be patient.

What follows for the audit is a rule I will use for the rest of the book. When two competent accounts of the same domain conflict persistently, the first hypothesis to test is not that one side is stupid or lying. It is that they are running different projections, that each is losing something different, and that the argument is about which loss is acceptable — which is a question about purposes, and therefore a question that evidence alone will never close.

Chapter Seventeen — Where You Start Decides What You See

For most of the last decade cosmologists have been unable to agree on how fast the universe is expanding, and the disagreement has the pleasing property of being entirely a disagreement about where to stand.

One method starts nearby. You find objects whose intrinsic brightness you believe you know, measure how bright they appear, infer their distances, compare with how fast they are receding, and read off a rate. The chain is built upward from the local neighborhood. The other method starts at the beginning. You take the pattern imprinted on the oldest light in the universe, apply a model of what happened since, and predict what the present expansion rate must be. The chain is built downward from the earliest time. Both methods have been refined for decades by people who would very much like to be right. The two answers differ by more than either side’s error bars allow, and the discrepancy has survived every attempt to make it go away.

What is interesting for our purposes is not the astronomy but the anatomy. Nobody is making an arithmetic mistake. The instruments are excellent. The disagreement lives entirely in the anchor — in the decision about which end of the chain of inference to calibrate — and the anchor is not a measurement. It is a choice made before any data are examined, it propagates through everything downstream, and it is not usually stated in the same sentence as the result.

Here is the general form, and it is one of the more useful things I know. A conclusion drawn from a chain of inference is not fully specified until you say where the chain was pinned. Two analysts using identical data and identical reasoning, pinning at different points, will obtain different and sometimes opposite results, and neither will be able to find an error in the other’s work, because there is no error. Reporting such a conclusion without its anchor is like reporting a velocity without saying what it is relative to: not false, exactly, but incomplete in a way that makes it useless and, worse, makes it look like a fact about the world when it is partly a fact about the procedure.

I have come to think that most durable public disagreements have this shape, and that recognizing it saves an enormous amount of time otherwise spent accusing people of bad faith.

Consider how it plays out here. Anchor an account of gender in the present-day self-understanding of individuals, and a great many things follow rigorously: the categories a society offers are to be assessed by how well they accommodate the people living in them, mismatch is a cost borne by identifiable persons, and the historical variability of categories is evidence that current arrangements are contingent. Anchor the same account in aggregate population data, and different things follow just as rigorously: the categories are to be assessed by how well they describe distributions, individual mismatch is a tail phenomenon in a well-behaved cluster, and historical variability is noise around a stable biological signal. Both derivations are valid. The premises differ at exactly one point, which nobody stated.

The historical version of the same problem is sharper still, and it embarrasses everybody. When we describe the gender arrangements of a society four centuries ago, do we use its categories or ours? Use theirs and the arrangements come out coherent, internally justified, and largely uncontested by the people living in them — a description that is accurate and that systematically hides every injury the categories were built to make invisible. Use ours and the injuries become visible immediately, at the cost of describing people in terms they would not have recognized and attributing to them a self-understanding they did not have. Both approaches are defensible. Historians have argued about this for a century without resolution, because it is not resolvable: it is an anchor choice, and the choice determines the finding.

The lesson is not that we should refuse to choose. Refusing to choose an anchor means saying nothing at all, and the people who counsel it are usually people happy with the status quo, which is itself an anchored position pretending not to be. The lesson is that the anchor should be stated in the same breath as the conclusion, and that a conclusion which inverts under a change of anchor should be reported as what it is: a statement about the world and the procedure jointly.

There is a diagnostic test that follows and it is easy to apply. Take the claim. Identify the anchor. Move it, hypothetically, to the other end. If the conclusion barely shifts, the anchor was not load-bearing and the claim is about the world. If the conclusion reverses, the claim was substantially about the procedure, and everyone involved has been arguing about a decision while believing they were arguing about reality. I have never met a long-running dispute where this test was uninformative, and I have met very few where anyone had performed it.

Two further examples, drawn from outside the subject, to show the shape is not peculiar to it.

The measurement of poverty inverts on the anchor. Fix the threshold to a basket of goods, and poverty in the developed world has fallen dramatically over fifty years. Fix it to a fraction of the median income of the society in question, and it has been flat or has risen. There is no arithmetic dispute. There is a choice between an absolute anchor and a relative one, each of which corresponds to a coherent idea of what poverty is, and every political argument on the subject is a proxy for that choice conducted with statistics.

Or take the assessment of a medical intervention. Anchor on relative risk reduction and a treatment cuts your chance of an outcome by a third, which sounds decisive. Anchor on absolute risk and it moves your chance from three in a thousand to two in a thousand, which sounds negligible. Both numbers are correct and describe the same trial. Which one appears in the press release depends on who wrote it, and the choice is never announced as a choice.

Probability has a version of this that philosophers call the reference class problem, and it is the same difficulty wearing different clothes. What is the chance that a particular man dies within the year? It depends entirely on which group you place him in: as a sixty-year-old, as a sixty-year-old who runs daily, as a smoker, as a smoker who runs daily and has a specific genetic profile. Each classification yields a defensible number, the numbers differ, and there is no fact of the matter about which class is the right one, because he belongs to all of them. Every risk figure ever quoted to a patient, an insurer, or a jury is anchored in a class somebody chose.

Political statistics run on the same trick, and the trick is the choice of baseline year. Almost any national trend can be made to rise or fall by moving the starting point a few years, since the starting point can be placed at a peak or a trough. Nobody in that business is unaware of it. The baseline is selected first, the chart is drawn second, and the chart is presented as the finding. Once you have seen this you cannot unsee it, and the main effect is a permanent mild suspicion of the phrase since records began.

What I would like the reader to take from this is a small change in listening. When you next encounter a confident claim in a contested area, the useful question is not whether it is true, which you are usually in no position to determine. The useful question is where the chain was pinned, and whether the person telling you would give the same answer if it were pinned at the other end. If they would, you are in the presence of something robust. If the whole edifice swings with the anchor, you have learned something more valuable than the claim: you have learned what the argument is actually about, which is almost never what the participants say it is about.

And, in the spirit of the symmetry rule, I should say that the framework of this book is anchored too. I have pinned the chain at the point where a claim meets an instrument, because that is the point I know how to reason about. Somebody who pins it at the point where a claim meets a life will find that different things become visible and other things disappear, and they will not be making a mistake. Chapter Thirty-Two is where I try to say what my own anchoring costs.

Chapter Eighteen — The Prize for Bad Writing

In 1998 a small academic journal called Philosophy and Literature ran its fourth annual contest for the worst prose in scholarly publishing, and awarded first prize to a sentence by Judith Butler. The sentence had appeared the previous year in a journal of critical theory. It runs to about ninety words, contains a subordinate structure that requires the reader to hold four abstractions in suspension before the verb arrives, and concerns a shift in Marxist theories of structure toward theories of hegemony. It was reprinted in newspapers around the world, and it remains the single most widely circulated thing Butler has written, which is a fate one would not wish on an enemy.

The contest was run by Denis Dutton, an editor with a genuine grievance and a comedian’s instincts, and it worked because the sentence really is very bad. It is not bad in the way a technical passage is bad to an outsider. Technical prose is hard because it uses precise terms for things that need them, and it can be unpacked by learning the terms. This sentence is hard because of its architecture: the terms are individually available to anyone who has read a little theory, and the difficulty is entirely in how they have been stacked.

Butler replied the following year in a newspaper, and the reply is more interesting than either the sentence or the prize. The argument has three parts. Difficult language can be necessary when the point is to unsettle what common sense takes for granted, since common sense is carried in ordinary phrasing and a comfortable sentence returns the reader to the assumptions it was written to disturb. The demand for clarity has a history, and it has frequently been used to police what may be said and by whom. And no one complains that physics is hard; the demand that the humanities be immediately intelligible reflects a view about which subjects are permitted to have technical difficulty.

Each of these is a real argument and each deserves to be taken at its strongest before it is assessed.

The first is true and important. Ordinary language is not neutral; it encodes exactly the settled assumptions that a critical inquiry may need to suspend. Every field that has ever had to say something genuinely new has had to strain its vocabulary, and readers who demand that a difficult thought be expressed in comfortable words are frequently demanding that it not be thought.

The second is also true, and more often true than intellectuals like to admit. Accusations of obscurity have a long career as a weapon against people whose conclusions are unwelcome, and the charge is convenient because it does not require reading. It is worth noticing who is generally accused of writing badly and who is generally credited with depth for identical levels of difficulty, and the answer does not correlate well with the actual prose.

The third argument, about physics, is the weakest and it is worth saying why, because it is the one most often repeated. Technical difficulty in the sciences is of a specific kind: the vocabulary is defined, the definitions are stable, they are collected in textbooks, and a reader willing to spend the time can climb the ladder from the bottom to the top in a known number of steps. That is difficulty with an access route. The difficulty in the prize-winning sentence has no such route, because the obstacle is not a set of defined terms but a syntactic construction, and no amount of study makes an overloaded sentence unfold more easily. Comparing the two is comparing a locked door with a wall.

Context is owed here, because the prize did not arrive in a vacuum. Two years earlier a physicist had submitted a deliberately meaningless article, stuffed with fashionable terminology and mathematical nonsense, to a well-regarded cultural studies journal, which published it. The hoax was widely reported and did enormous damage, not because it proved that the field was empty — one editorial failure proves nothing about a discipline — but because it dramatized a suspicion many people already held and could not otherwise test. By 1998 there was a public appetite for evidence that the humanities had lost the ability to distinguish depth from noise, and a badly built sentence with a famous name attached to it fed that appetite perfectly.

It is worth distinguishing two kinds of hard writing, since the defense of difficulty is usually made without the distinction and is thereby made indefensible. Compression is hard because a great deal has been packed in: every clause is carrying content, nothing can be removed without loss, and the reader who slows down is rewarded with something. Obstruction is hard because of how the sentence is built: the content would survive being said another way, and the reader who slows down discovers that the difficulty was in the packaging. Mathematics and good philosophy are full of the first. The prize-winning sentence is a specimen of the second, and the test is simply whether a rewrite loses anything. In that case it does not.

So the honest verdict is split, and I think both halves matter.

The prize was unfair as a procedure. One sentence was extracted from a body of work of several million words, presented without its surroundings, and offered to an audience with no means and no intention of checking whether it was representative. No comparison class was provided; nobody sampled the field to establish a baseline; the winner was selected for being funny, and being funny is not a measurement. If a scientist assessed a colleague’s work by that method they would be finished, and the fact that the target was a fair one does not sanctify the instrument.

And it landed on something real. The prose is a genuine defect, and I want to specify the defect precisely, because vague complaints about jargon miss it. The problem is not vocabulary and not abstraction. It is that a very large proportion of Butler’s sentences cannot be checked by the reader against anything. A checkable sentence makes a claim clear enough that one can ask what would be different if it were false. Much of this prose does not offer that purchase: the claims are formulated at a level of generality where a counterexample cannot be constructed, and the reader is left to assess the sentence by whether it sounds right.

That defect connects directly to the argument of Chapter Eleven and completes it. A text that cannot be checked by its readers hands its meaning to intermediaries, and then has no standing to complain about what the intermediaries do. Butler’s thirty-five year war against the costume reading was lost in advance, in the prose. This is not a moral failing and it is not a reason to discard the work. It is a cost, it was paid by the ideas, and it was avoidable, as demonstrated by the fact that the most accessible book in the corpus arrived in 2024 and turned out to be perfectly possible to write.

There is a further cost that is distributed unevenly and is rarely mentioned by either side. Difficulty is not equally expensive for everyone. A tenured reader with a research day and a colleague to ask can afford it. A reader with a job, a commute, and children cannot, and will take the summary. So a body of work committed to examining how the powerful maintain their position, written in a form that only people with institutional leisure can evaluate, has a structural problem it never quite acknowledged. The point has been made from inside the field for decades, generally by people whose politics are indistinguishable from Butler’s, which is why it cannot be dismissed as an attack from outside.

The counterexamples ought to be admitted too, since they are what keep the whole defense of difficulty from working. Darwin proposed an idea that overturned the settled understanding of life and wrote it so that any literate person of his century could follow the argument. Hume dismantled the foundations of causation, personal identity, and religious inference in prose that reads like conversation. It is possible to disturb common sense in clear language. It is harder, it takes longer, and the resulting book earns less respect from professionals, which is a fact about professionals rather than about the ideas.

The last word should go to a small irony that nobody arranged. The prize-winning sentence is about how a change in theoretical framework brought the question of time into the thinking of structure, and about how contingency reopens what had been treated as fixed. Underneath its architecture there is a real point, and it is one of the more useful things in this book’s argument. It was made in a way that guaranteed that the only people who would ever encounter it were the ones reading it in a newspaper for a laugh.

Chapter Nineteen — The Professor of Parody Answers Back

In February 1999 The New Republic published an essay by Martha Nussbaum, then among the most respected philosophers in the United States, under a title that has outlived almost everything else written about the subject. It called Butler a professor of parody and diagnosed a hip defeatism. It is beautifully written, in the register of a serious person who has decided to stop being polite, and it did more damage than the bad-writing prize, the effigy, and thirty years of unfriendly reviews combined.

It is also a mixed piece of work, and the mixture is instructive. Some of its charges survive an audit intact; some are misdirected; one of them was correct when written and has since been overtaken by events. Since this book has promised to apply its instrument regardless of who is holding it, let us take them one at a time.

The first charge is technical: that Butler misreads Austin. On Austin’s account, a performative utterance accomplishes something because a specific institutional convention backs it, and the conditions of success are conventional and identifiable. Butler generalizes the notion far beyond utterances, to bodily comportment and to processes with no single speaker and no single moment. Nussbaum treats this as error. I do not think it is. The generalization runs through Derrida, who had contested Austin’s framework in the 1970s and argued that the conventional cases are not the primary ones, and Butler is working in that lineage openly. One may reject the Derridean move — there are good arguments against it — but rejecting it is not the same as catching someone in a mistake. The charge fails as stated. What survives is a smaller and fair complaint: an extension of a technical term ought to be flagged as an extension, and Butler’s texts frequently do not mark where Austin’s usage ends and the enlarged one begins, which invites exactly the confusion Nussbaum walked into.

The second charge concerns legal claims, and here I give more ground to the critic. Nussbaum is unusually well equipped in this area, and her objections to specific claims about how law functions in Excitable Speech are the sort of objection that a specialist can make and a generalist cannot answer by restating their framework. Some of them stick. This matters beyond the immediate quarrel, because a body of work that ranges across law, psychoanalysis, classics, and political theory is going to be wrong in local ways, and a theory whose defenders treat every local error as an irrelevance to the general position is a theory that has arranged never to learn anything.

The third charge is the serious one and it is the reason the essay is still read: that Butler offers no normative theory, no account of what social justice or human dignity would consist in, and therefore no basis on which any arrangement could be called better than any other. Symbolic subversion is offered in place of an account of what should be done, and subversion is not a criterion, since it can be performed in any direction.

This charge was, in my judgment, correct as applied to the work available in 1999, and I said as much several chapters ago on independent grounds. If subjects are formed by power, and resistance is formed by the same power, and there is no standpoint outside, then the resources for saying that one reworking is an improvement over another have to come from somewhere, and the framework does not supply them. Nussbaum identified a genuine hole. It is the same hole visible from inside the apparatus, which is the sign of a real one rather than a partisan one.

And here the audit produces a result that pleases nobody. The charge has been substantially overtaken. Nussbaum was writing about a corpus that ended in 1997. Beginning in 2004, and continuously since, Butler’s work has been almost entirely occupied with exactly the missing question: which lives register as lives, whose deaths are counted as losses, what obligations follow from the fact that everybody is dependent on arrangements they did not choose. Grievability is a normative criterion. One may find it thin, or contest it, or think it insufficiently specified — I have some sympathy with all three — but it is not nothing, and the objection that there is no normative theory here is an objection to a body of work that stopped being the current one twenty-two years ago.

That would be an unremarkable observation about intellectual history if it were not for what actually happened, which is that the essay continued to circulate as a live verdict long after the conditions that produced it had changed. This is the standard fate of a devastating critique. It becomes the thing people know instead of the thing it was about, and it acquires a durability entirely disconnected from whether its target still exists. The correction, in this case, would have to be delivered by the people who most enjoyed the original, and there is no mechanism by which that would occur.

The fourth charge is about material conditions: that a theory preoccupied with the destabilization of categories has nothing to say to women facing hunger, illiteracy, violence, and the absence of legal standing, and that its ascendancy in American departments represented a retreat from politics into style. Nussbaum’s own work on capabilities and on development gives her standing to make this complaint that most complainants lack.

I think this charge is half a category error and half a fair observation about a profession. The category error is the demand that a theory of subject formation double as a development program; theories are entitled to have subjects, and an account of how gender categories are produced is not obliged to also feed anybody. The fair observation is about allocation. A field has finite attention, and the attention that went into three decades of textual commentary on performativity did not go elsewhere, and it is legitimate to ask what the opportunity cost was. Chapter Nine made a version of this complaint from a different direction: a theory of gender as maintained repetition should have generated measurements and did not.

Symmetry requires one more thing of me here, and Nussbaum would be the first to accept the demand. Her own position rests on a declared component of considerable size: a list of central human capabilities, held to be the proper measure of a life’s flourishing, defensible across cultures and available as a standard for assessing institutions. It is a serious and carefully defended list. It is also a declaration — a choice about what a human life is for, arrived at by philosophical reflection in a particular tradition, not read off any instrument. Nussbaum says so, which is to her credit and is more than most normative theorists manage. But an argument that criticizes an opponent for having no criterion, from a position whose criterion is itself declared, is not comparing a measurement with a void. It is comparing a declaration that has been stated with a declaration that has been withheld, and that is a real advantage, and a smaller one than the essay’s tone suggests.

The charge of quietism also has an empirical form that nobody has tested, and I record it as an open question rather than a verdict. Does the production of critical theory displace political activity, as the essay implies, or does it accompany it, or is it irrelevant to it? These are answerable questions. One could examine whether departments and periods heavy in this literature show lower rates of participation in organizing, or whether the individuals writing it are more or less politically active than comparable academics. So far as I can discover, nobody has looked. The charge circulates as though it were established, which is exactly the condition that this book keeps finding claims in, and it is odd to find one of the most rigorous philosophers of the age in the position of having asserted it.

There is a fifth charge, about obscurity, which I have already dealt with and will not repeat except to note one thing about the essay’s own performance. Nussbaum’s prose is superb: lucid, confident, and quotable in a way that Butler’s is not. This affected the reception of the argument to a degree that has nothing to do with its correctness. A well-written wrong point defeats a badly written right one nearly every time, and the fact that this is true tells us something depressing about how intellectual disputes are settled, which is not by adjudication but by circulation.

Butler’s reply, insofar as there has been one, has been oblique — a general observation about whether there is value to be found in experiences of linguistic difficulty, and a great deal of subsequent work that answers the substantive charge by doing something else. Direct engagement never happened. The two have not, so far as I know, ever debated. This is a shame in the ordinary way, and it is also a data point about how academic disputes are actually conducted: the parties write past each other, the audiences sort themselves, and the exchange that would have been worth reading is precisely the one that nobody arranges.

What should a reader take from the episode? Something less satisfying than either camp offers. Nussbaum was right that the framework as it stood could not underwrite a politics, and wrong that this was a permanent feature. Butler was right that the demand for a normative theory can smuggle in exactly the assumptions under examination, and wrong to leave the demand unanswered for as long as it went unanswered. Both were arguing in a public register that rewarded the sharpest formulation rather than the most accurate, and both got the readership that register produces.

And there is one more thing, which I offer with some discomfort because it complicates a story I would rather tell simply. The essay contains an argument I have not seen anybody answer, and it is not about normativity at all. Nussbaum suggests that a certain style of critique offers its practitioners the sensation of radical activity at no cost, and that the sensation is itself the product — that one can spend a career destabilizing categories in a seminar room without any arrangement anywhere becoming different. Whether or not that is true of Butler, it is unmistakably true of a great deal of what followed, and the audit has no way of clearing it. A theory cannot be held responsible for its epigones. But when the epigones are numerous, well funded, and produce nothing measurable across three decades, the question of what the theory made easy is a fair question, and it is the one this book will come back to at the end.

Chapter Twenty — Declarations Cut Both Ways

The rule is easy to state and almost impossible to follow: an argument found to be defective cannot be used afterward, even when it points where you want to go. Most people accept this in the abstract and abandon it within a page, because the alternative is discarding a perfectly serviceable weapon on the grounds of a technicality that only you have noticed.

So let me start by paying the tax myself, on a debt incurred four chapters ago. I argued that the widely quoted figure of nearly two percent for the frequency of intersex conditions is a definition wearing the clothes of a statistic, since it depends entirely on a boundary decision about what counts. That is correct. Symmetry then requires me to say the same about the much smaller figure preferred by the other side, which is arrived at by the same procedure with a tighter boundary and is equally not a measurement. I have seen the small figure deployed as though it were the true value that the large one distorts. It is not. There is no true value, because the quantity being estimated is not defined independently of the estimation, and anybody who quotes either number as decisive has failed the audit regardless of what they were trying to prove.

With that established, here are five argument-shapes that dominate public discussion of this subject. Every one of them is used by all parties, and every one of them fails in the same way whoever is holding it.

The first is the appeal to what has always been. It asserts that an arrangement is grounded in universal human practice, and it is not a philosophical claim at all; it is a historical and ethnographic claim, and therefore checkable. Checked, it almost always comes out weaker than asserted, because human societies have been more various than any party to a modern argument finds convenient. This cuts in every direction. Those who claim that a particular arrangement is a recent invention are frequently wrong; those who claim it is universal are frequently wrong; and the honest position is that the record is mixed, partial, and refracted through the interests of whoever wrote it down. Anyone citing anthropology in this argument should be asked which societies, which sources, and who collected them.

The second is the appeal to consequences that have not happened yet. If we accept this, then that will follow. This is a legitimate argument form and it has a requirement attached: a mechanism. Without a mechanism, a claim that one change leads to another is a prediction with no engine, and it is exactly as strong as the confidence of the person making it. Both sides here produce elaborate chains of foreseen consequence and neither, as a rule, supplies the mechanism, which is why the same chain can be run in reverse by an opponent with no loss of plausibility.

The third is the appeal to lived experience. First-person testimony has real authority within its domain and no authority outside it. A person is the only available source on how their life feels and what it has cost them, and that report should be believed in the ordinary way one believes people about themselves. The same report has no special standing on questions of causation, incidence, or the effects of a policy on a population, and no amount of sincerity converts it into evidence about those. This constraint is uncomfortable and it applies to everybody: a person recounting the injury a policy did to them and a person recounting the injury its absence did to them are in exactly the same epistemic position, and treating the testimony of one as decisive and the other as anecdote is the single most common failure in this whole debate.

The fourth is the appeal to science, and it is the most sophisticated failure of the five. It typically proceeds by citing a real empirical finding, correctly, and then adding an unstated premise that carries all the normative weight. The finding is about bodies or brains or distributions; the conclusion is about what should be permitted, funded, recorded, or required; and between them sits a bridge principle that nobody has stated and that is not itself a scientific claim. Ask for the bridge. In my experience the request is received as hostility, which is a reliable indication that the bridge is load-bearing and undefended.

The fifth is the appeal to harm, which has become the universal solvent of contemporary argument. Harm is a genuine consideration and it is also a claim with a quantity attached, and quantities require measurement and comparison. How much harm, to how many, compared with what alternative, over what period? These questions are answerable in principle and are almost never answered, because the rhetorical work is done by the word rather than the number. A debate in which all parties invoke harm and none of them measure it is not a debate about harm. It is a debate about who gets to use the word.

A sixth shape deserves inclusion, since it has become the dominant one and it is the most elegant of the failures: the appeal to definition. It proceeds by asserting that some contested claim is true by definition, thereby converting an empirical question into a stipulation and making disagreement a matter of linguistic incompetence rather than of evidence. The move is available to everyone and is used by everyone. Its distinguishing mark is that the argument becomes unlosable, which is the surest sign that it has stopped being about anything. When a dispute reaches the stage where both parties are asserting definitional truths, the useful intervention is not to adjudicate between the definitions but to ask what question each was built to answer, at which point the substantive disagreement usually reappears in a form that can be discussed.

And a seventh, which is not an argument at all but has replaced argument in much of this territory: the audit of motives. Instead of assessing what someone has said, one assesses why they said it — what they stand to gain, whose interests they serve, what unlovely thing their position reveals about them. This is always available, since motives are unobservable and can be attributed at will, and it is always satisfying, since it disposes of the argument without engaging it. It also has the property of being symmetrical to the point of uselessness: whatever motive is imputed to one side can be imputed with equal plausibility to the other, and generally is, within the hour.

Now, a point about asymmetry that keeps the preceding from being a comfortable plague-on-both-houses.

Failing an audit is not the same as being wrong, and the two are constantly conflated by people who have just learned this method. An argument can be badly constructed in defense of a true conclusion; a rigorous argument can lead somewhere false because a premise was false. What the audit establishes is what a given argument entitles you to, not what is the case. So the finding that both sides of a dispute are using defective argument-shapes tells you that the dispute is not being conducted well. It does not tell you that the truth is in the middle, and the assumption that it must be is itself an unexamined declaration, and a lazy one.

There is a further asymmetry that concerns the costs of being wrong, and it should be stated because the method is otherwise open to abuse. Auditing is cheap and unfalsifiable if used only to demand more rigor from whichever side one dislikes; there is always more rigor to demand. The discipline that prevents this is the requirement to state, in advance, what would satisfy you. If no possible evidence would change your position, then your engagement with the evidence is decorative, and the honest course is to say that your commitment is prior and defend it as such. This is not a disgrace. A great many of the most important commitments people hold are prior to evidence. It is only a disgrace to conceal it.

The reward structure works against everything in this chapter and it is worth being explicit about why. A confident overstatement travels; a hedged accurate claim does not. A correction reaches a fraction of those who received the original and arrives later, when the matter is settled in most minds. Being loudly wrong therefore costs less than being quietly right, and the market clears at a level of stridency that nobody would choose.

That structure also explains a phenomenon that puzzles people who follow this argument from outside: why the participants seem to get worse at it over time rather than better. It is because the moderate and careful participants leave. The costs of participating fall most heavily on those who qualify their claims, since a qualification is an opening and the audience for qualifications is small. Over time, the population of people still arguing is the population that found the argument rewarding, which is a selected group with characteristics you can predict. This is not a claim about any side’s character. It is a claim about what a filter does when it is left running for thirty years.

The rule I began with therefore has a corollary that I have found more useful than the rule itself. When you discover that an argument on your own side is defective, the productive response is not to abandon the conclusion but to find out whether a sound argument for it exists. Frequently one does, and it is better than the one you were using. Occasionally none does, and then you have learned something expensive and valuable. The people who cannot do this are not usually dishonest. They have simply never had the experience of a conclusion surviving the loss of its argument, and they assume, reasonably enough, that the two go down together.

Chapter Twenty-One — The Reader Who Was Left Out

Somewhere between a monograph and a school district there is a chain of transmission, and nobody has ever mapped it properly.

The chain has links and each link is a translation. A book is read by several hundred specialists. Some of them teach it, and what they teach is a compressed version shaped by the needs of a fourteen-week course. A textbook summarizes the compressed version. A teacher-training program summarizes the textbook. An administrator reads a professional development document derived from the training program. A form is redesigned. At each step, content is lost, vocabulary is standardized, qualifications are dropped as unteachable, and the residue acquires the authority of consensus precisely because its origins have become invisible. By the end nobody involved has read the book, and the thing implemented bears a relation to the original that would be charitably described as etymological.

This is not a complaint about anyone’s competence. It is what transmission is, in every field. The version of evolutionary theory in a school curriculum is not the version in the journals. The version of economics in a newspaper is not the version in a department. The chain is lossy by construction, and the loss is not random: what survives is what compresses well, what can be stated as a rule, and what an administrator can act on. What is lost is exactly the qualifications, the boundary conditions, and the acknowledgments of uncertainty, which is why the far end of the chain always sounds more confident than the near end.

The loss at each link is not merely quantitative, and one contemporary example shows the shape better than any abstraction. A body of psychological research on implicit associations, conducted with care and heavily qualified by its own authors — who have said publicly that their instrument is not suitable for diagnosing individuals — passed along the chain and emerged at the far end as a mandatory workplace training module asserting things the original researchers had explicitly declined to assert. Nobody forged anything. Each link simplified slightly in the direction of actionability, because an administrator cannot act on a confidence interval, and the accumulated simplification arrived as a certainty.

If one wanted to settle the influence question properly rather than by assertion, the methods exist and are not exotic. Citation networks can be traced from a monograph outward, with the strength of each link estimated. Curricula can be sampled across decades and coded. Policy documents can be searched for textual descent, which is easy to detect because bureaucratic language is copied rather than reinvented. Attitude series exist for most developed countries going back to the 1970s and can be compared against the timing of translations. This is an afternoon’s proposal and several years’ work, and the fact that it has not been done, by either the people claiming influence or the people denouncing it, tells you how much either party actually wants the number.

Apply this to the question the fourteen-year-old in Chapter Two was asking, and the results are more deflating than either party would like.

What can actually be traced to Butler? The vocabulary of performativity across the humanities, unambiguously. The reframing of the category of women as a site of contest rather than a foundation within feminist theory, largely, though Butler was one voice among several and the pressure was coming from other directions at the same time. The concept of grievability in critical studies of war and media, which is a genuine and quite specific contribution with a documented uptake. In each of these cases one can follow citations, name intermediaries, and produce a chain.

What cannot be traced? Almost everything the public argument is about. Specific institutional policies have their own documented histories, generally running through law, litigation, professional bodies, and insurance, none of which cite continental philosophy. And the broad change in attitudes across the developed world since the 1970s cannot plausibly be laid at any theorist’s door, for a reason that ought to settle the question: it is observable in countries where Butler is unread and untranslated, it began before the relevant books existed, and it tracks a set of demographic and economic variables — urbanization, women’s entry into paid work, contraception, later marriage, rising education, the collapse of extended family households — which are measured, which move first, and which do a great deal better as predictors than any intellectual history.

That finding is unwelcome to admirers, who would like a philosopher to have changed the world, and to opponents, who require a culprit. It is nonetheless what the evidence supports, and it fits a general pattern that historians of ideas have documented repeatedly: theorists are usually downstream of the changes they are credited with, articulating something already underway, and their apparent causal power is an artifact of the fact that texts are easier to point at than demographics.

So much for the world’s account of Butler. There is a second and sharper sense in which readers were left out, and it comes from inside the field rather than from its opponents.

Viviane Namaste has argued that Anglo-American feminist theory, Butler’s work included, used transgender lives as material for its own theoretical purposes — as illustrations of the constructedness of gender, as evidence in an argument being conducted among academics — while having very little to say about the conditions of those lives: access to healthcare, employment discrimination, housing, exposure to violence, the practical business of surviving in institutions designed on the assumption that one does not exist. The charge is not that the theory is unsympathetic. It is that it is extractive: the lives supply the theoretical interest and receive nothing in return.

Timothy Laurie has pressed a related point about vocabulary. Phrases like gender violence, applied to assaults on transgender people, describe a landscape from which class, labor, racial segregation, the economics of sex work, and the specific geography of who is attacked where have been removed, leaving a clean surface on which an argument about the human can be conducted. The complaint is precise and it is about what a phrase makes invisible, which is an entirely Butlerian sort of complaint, turned around and aimed at its source.

I think both charges survive the audit and they are the most serious in this book. Here is why they are worse than the objections from outside. An opponent who says the theory is wrong is making a claim that can be met on its merits, and Butler has met a great many of them. These charges say something different: that the theory selected its examples in a way that served the theory, that the selection had costs borne by people who were not consulted, and that the field’s enormous output on the subject is not proportionate to any benefit received by those it discusses. That is an audit finding in the strictest sense, and it is one that no amount of internal argument can answer, because the missing quantity is external.

To Butler’s credit, the later work moves noticeably in this direction, and Chapter Fourteen recorded one instance where a public criticism produced a visible correction. The books after 2004 are much more concerned with material vulnerability, with who is exposed to what, and with the arrangements that distribute exposure. Whether that constitutes a response to Namaste or a change of subject is a judgment I will leave to the reader, since I can see the case for both and I notice that my own preference here is doing more work than the evidence.

There is a general lesson in this for anyone who builds theories about people, and it is the reason I have given the chapter its title. A theory needs cases. Cases are lives. The people whose lives are being used as cases are almost never in the room when the theory is assessed, and the assessment criteria — elegance, generality, explanatory reach — are criteria that have nothing to do with whether being a case was any use to them. There is no methodological fix for this. There is only the discipline of asking, periodically and unprompted, who supplied the material for this argument, what they got out of it, and whether they would recognize the version of themselves that appears in the footnote.

And one last observation on the traceability question, offered in the spirit of the symmetry rule. If Butler cannot be credited with the changes in attitudes, then Butler cannot be blamed for them either, and a very large industry of denunciation is aimed at a target that is not where the events occurred. That should be some comfort. It is not, of course, because the function of the denunciation was never explanatory. A movement needs an author for what it opposes, since arrangements without authors cannot be defeated in argument, and a difficult philosopher with a photograph and a body of unreadable prose is an extraordinarily convenient one. That is a fact about the needs of the accuser, and it tells us nothing whatever about the accused.

Chapter Twenty-Two — Speech That Wounds and the State That Names It

In 1992 the Supreme Court of Canada handed down a decision that redefined obscenity. The old test had been about moral corruption; the new one was about harm, and it drew explicitly on feminist arguments that certain material contributed to the subordination of women. Feminist legal scholars had campaigned for exactly this reform and regarded the ruling as a victory. Within a short time, customs officers applying the new standard were seizing shipments bound for gay and lesbian bookstores at a rate wildly out of proportion to anything else, including material by feminist authors who had supported the reform. A small bookshop in Vancouver spent the better part of a decade in litigation over it.

Whether the ruling caused this, or merely supplied a vocabulary to people already inclined that way, has been argued about ever since, and the causal question is genuinely open. What is not open is the pattern, because the pattern is old. An instrument for restricting speech is designed by people who imagine themselves administering it. It is then administered by whoever holds the office, applying it to whoever is locally unpopular, which is rarely the group the designers had in mind. Obscenity law, blasphemy law, sedition law, and public-order law have all done this, in multiple countries, across two centuries, with a consistency that ought to count as an empirical regularity rather than an unlucky coincidence.

Excitable Speech, published in 1997, is Butler’s book about this, and it is the most policy-relevant thing in the corpus. Its central concern is hate speech and the question of what the state should be empowered to do about it, and its position is one that satisfies nobody, which in this area is often a sign of accuracy.

The argument begins by conceding what needs conceding. Speech injures. This is not a metaphor and Butler does not treat it as one: words spoken in the right circumstances by the right speaker damage people, sometimes permanently, and any account that treats language as a transparent medium for transmitting propositions has failed to describe the phenomenon under discussion. So far, the case for regulation.

Then comes the turn. If the state is empowered to determine which utterances constitute hate speech, the state acquires thereby the power to define the boundary of acceptable discourse, and it will use that power according to its own priorities, which are not the priorities of the people who requested the legislation. Butler pressed this specifically against Catharine MacKinnon’s campaign against pornography, arguing that the argument accepted state authority over representation far too readily, and that the acceptance would prove more consequential than the immediate object.

There is a second and stranger argument in the book, taken from Foucault, and it is the one people find hardest to accept. Prohibition propagates. A censoring regime must name what it forbids, circulate the names among its officers, establish procedures for identifying instances, and thereby produces a great deal of the very discourse it means to suppress. The nineteenth century’s attempt to silence discussion of sexuality generated an unprecedented volume of talk about sexuality, in medicine, law, education, and the confessional. Anyone who has watched a modern attempt to suppress a document knows the contemporary version, in which the suppression is the reason anybody hears of it.

How does this hold up? The descriptive and predictive part holds up extremely well, and I think it is the strongest empirical claim Butler has ever made. The prediction is that powers created to protect a vulnerable group will be applied against that group at a rate that surprises their advocates, and it has been borne out in enough jurisdictions and enough decades that the burden has shifted: it is now the person proposing a new speech restriction who owes an account of why theirs will be the exception. That is a real result, arrived at by philosophical argument and confirmed by administrative history, and it is startlingly rarely credited to the person who made it.

There is a mechanism underneath, and it is worth stating because it explains why the pattern is so robust. Any speech restriction has to be applied by somebody with discretion, since no rule can specify its own instances. Discretion is exercised by officials, who are responsive to complaints, and complaints arrive disproportionately from whoever is organized, offended, and confident of a hearing. The restriction therefore flows along the existing channels of social power regardless of the intentions written into the statute, and the intentions are the least predictive element in the system. A law is not what it says; it is what its enforcement produces, which is the point from Chapter One returning in legal dress.

The remedial half of the book is much weaker and this is where I part company. Butler’s alternative to state regulation is resignification: an injurious term can be taken up by those it wounds and reworked until it does something else, as has demonstrably happened with several words that were slurs within living memory and are now claimed. The historical examples are real. As a remedy, though, it has three defects that the book does not confront.

It is slow, taking a generation or more. It is unreliable, in that most attempts fail and the successes are visible only in retrospect, so it cannot be planned. And it assigns the labor to the injured, who are asked to perform a difficult and public rehabilitation of a word that is currently being used against them, on the promise of an outcome nobody can guarantee. Whatever the merits as description, that is not a policy, and offering it in the same breath as an argument against legal remedies leaves a person who has just been threatened with nothing to do this week.

There is also a category the framework handles badly, and it has grown enormously since 1997. Butler’s model assumes that the injurious utterance is a speech act performed by a speaker who can be identified and answered. A great deal of contemporary injury is not like that. It is produced by aggregation: a thousand messages, individually trivial, arriving simultaneously, from accounts with no history and no future, ranked and amplified by a system optimized for engagement. There is no speaker to resignify against and no convention to rework. The harm is a property of the distribution rather than of any utterance in it, and a theory built on the speech act has no purchase on a phenomenon whose defining feature is that no single act matters.

That gap is not a reason to discard the book. It is a reason to notice its date. Excitable Speech is a work about a world in which speech was scarce, attributable, and expensive to distribute, and nearly everything it says about that world is sound. We now live in one where speech is abundant, frequently unattributable, and free to distribute, and the transposition has not been done by anybody, in this tradition or any other.

There is a distinction the book makes only glancingly and which is needed to see what has changed. Some speech injures by meaning: the content is understood and the understanding does the damage. Other speech injures by uptake: the words themselves are almost incidental, and the harm consists in what institutions and bystanders do next — the report filed, the invitation withdrawn, the employer notified, the address published. The first kind can in principle be answered by other speech, which is the classical liberal remedy and is not absurd. The second cannot, because the injury is not located in the exchange at all; it is located in a chain of subsequent actions by parties who were never spoken to.

Emergency powers supply the cleanest historical illustration of the durability problem, since they are always granted for a defined crisis and almost never returned afterward. Powers created in wartime persist into peace; surveillance authorities justified by one threat are inherited by administrations facing another; the sunset clause is added and then extended. This is so consistent that it deserves to be treated as the default expectation rather than a cynical prediction, and it applies with full force to any authority over speech. The relevant question at the moment of design is never whether the current government can be trusted with the instrument. It is what the instrument will be worth to the least trustworthy government that will ever hold it, which is a question about a period the designers will not live to see.

The audit finding, then. On the diagnosis of what happens when the state is given power over discourse: strong, confirmed, underappreciated. On the mechanism by which enforcement drifts toward existing power: strong and generalizable well beyond speech. On the remedy: weak, and weak in a specific way that should be recognizable by now, since it is the same weakness as everywhere else in this corpus — an account of how change becomes possible offered where an account of what to do was needed.

I want to close with the part of this that has aged into something Butler did not intend. The book’s argument implies that whoever holds administrative power will use speech restrictions to their own ends, and that the identity of the holder is not fixed. Advocates of restriction almost always reason as though their allies will be administering it forever. They will not. Every regime of speech regulation is eventually inherited by people the designers would not have chosen, and the only question that matters when designing one is what it will do in those hands. That is not a partisan point and it does not favor any particular position on the underlying issues. It is a question about durability under a change of management, and it is the question that is never asked.

Chapter Twenty-Three — Antigone and the Rules Before the Rules

A young woman defies a king, buries her brother against an explicit decree, and is walled up alive for it. For two centuries the philosophers have used her as a diagram. Hegel made her the representative of family and divine law in collision with the state and human law, a tragic clash of two goods. Lacan made her the figure of a desire beyond compromise, situated at the limit of the symbolic. Generations of readers have received her as the spokeswoman for kinship against politics.

Butler’s short book from 2000 begins by pointing out something about Antigone that everybody knew and nobody had used: she is not in a position to represent normal kinship, because her kinship is not normal. She is the daughter of Oedipus by his mother. Her brother is also her half-nephew. The family whose claims she is said to embody is one whose structure violates the very rule that is supposed to found kinship as such. If she stands for the family against the state, she stands for a family that the theory of the family cannot accommodate.

That is a small observation with a large blast radius, and to see why one has to know what was being defended.

Structuralist anthropology had proposed, in the middle of the twentieth century, that the prohibition of incest is the threshold at which culture begins — the founding rule that converts biological grouping into a system of exchange, alliance, and meaning. Psychoanalysis in its Lacanian form built a parallel structure, a symbolic order of positions and prohibitions that any speaking subject must enter in order to be a subject at all. In both cases the claim is not that particular families are arranged in particular ways. It is that certain structures are prior to any social arrangement, conditioning what can be thought and said, and are therefore not available for revision by anybody’s decision.

Now look at that claim with the instrument this book has been sharpening, because it is a specimen. The symbolic order is posited to explain a set of effects — the regularities of kinship, the shape of subject formation, the persistence of certain prohibitions. It is not independently observable. It is known entirely through the effects it is invoked to explain. And it is declared to be unrevisable, prior, and structural. That is a role without a referent, equipped with a claim of necessity, and the claim of necessity is doing all the political work.

Butler’s question is exactly the right one: is the unrevisability a finding or a declaration? What observation established that these structures are prior to politics rather than being particularly well-defended pieces of politics? The answer, when one goes looking, is that no observation established it. It was built into the framework at the outset, in order to explain why kinship arrangements are stable, and it is then cited as evidence that they must be.

This matters because arguments of exactly this form were, at the time the book appeared, being made in public across several countries about adoption, assisted reproduction, and the legal recognition of families that did not fit the received pattern. The form was: these arrangements are not merely unfamiliar, they violate a structure that underwrites intelligibility as such, and the consequences will therefore be grave and unforeseeable. That is a claim about the world. It generates predictions. Predictions can be checked.

They have been checked, in the twenty-five years since, more thoroughly than most claims in social science. The weight of the evidence on children raised in the arrangements that were said to be structurally impossible is that outcomes are indistinguishable from comparison groups on the measures ordinarily used. The literature has had its methodological quarrels, some of them sharp, and no single study settles anything; but the pattern across many studies, in several countries, using different designs, is one of the more consistent null results in family research. Whatever the symbolic order was doing, it was not doing what its defenders said would happen if it were violated.

Symmetry now requires the harder half of the chapter, because it would be easy and dishonest to leave the impression that all appeals to prior structure are frauds. They are not. Some constraints really are unavailable for revision by decision, and a politics that ignores them fails.

There are physical constraints: no arrangement of society will get more energy out of a system than went in, and every scheme that has assumed otherwise has failed on schedule. There are resource constraints, which are not eliminable by declaring them unjust. There are constraints from the developmental requirements of human infants, who need sustained care from someone for years, and any arrangement that does not supply it fails regardless of its ideological credentials. There are constraints of scale, since institutions that work at the size of a village do not work at the size of a nation, and the belief that they should has produced some of the worst outcomes on record.

It is worth being fair to the anthropology as well, since it has been caricatured in both directions. The structuralist account was a genuine intellectual achievement: it took the bewildering variety of kinship systems and showed that they could be generated from a small number of rules governing who may marry whom, which is the kind of compression that marks a real discovery. What it could not establish, and did not have the evidence to establish, was that the rules were necessary rather than merely widespread. Feminist anthropologists made a further and sharper objection decades before Butler, noting that a theory describing marriage as the exchange of women between groups of men had adopted the perspective of one party to the transaction and called it the structure.

A small taxonomy of constraints may help, since the word covers several very different things and the differences determine what can be done. Physical constraints are not negotiable by anybody at any price. Resource constraints are negotiable only by finding more resources, which is sometimes possible. Developmental constraints arise from what human beings require in order to grow into functioning adults, and are negotiable in form but not in substance: someone must do the caring, though who and how has varied enormously. Coordination constraints exist because a practice only works if enough others follow it, and these are the most interesting, because they are entirely conventional and simultaneously extremely hard for any individual to change — which is precisely the profile of most of what gets called human nature.

So the question is never whether prior constraints exist but how to tell a real one from a defended preference. Here is the test I would propose, and it is the same test as everywhere else in this book: ask the claimant what would show the structure to be revisable. A real constraint has an answer — here is the measurement, here is the failure mode, here is the case that would refute me. A defended preference cannot produce one, and typically responds to the request by restating the constraint with more emphasis. The response to the request is more informative than anything in the original claim.

By that test the incest prohibition looks like a real constraint of some kind, though not the kind the theory said: it is close to universal, it has a plausible biological rationale, and its violations have measurable consequences. The elaborate symbolic architecture built on top of it does not pass, since no one has ever specified what would count against it. The two were bundled together and defended as a unit, which is a standard move and worth watching for: attach a contestable structure to an uncontested one, and let the credibility flow.

There is a lovely detail in the play itself that Butler makes good use of and that I cannot resist. Antigone, in defying the decree, speaks in the language of sovereign command — she issues, claims, and declares, taking on the very idiom of the authority she opposes. She does not resist power from outside; she resists it by occupying its speech. Whether Sophocles intended this is unanswerable, but it is unmistakably there, and it is the clearest illustration in literature of the thesis of Chapter Twelve, arriving two and a half millennia early.

The audit finding on the book is favorable with one reservation. The diagnosis is correct and important: a great deal of what presented itself as the structural precondition of culture was a declaration with a necessity claim attached, and the necessity claim was doing political work that nobody had licensed. The reservation is that Butler treats the exposure as though it settled matters, and it does not. Showing that a claimed constraint was declared rather than measured tells you that the question is open. It does not tell you what the answer is, and a generation of readers took the demolition of one bad argument as an establishment of its opposite, which is an error nobody should make twice.

Chapter Twenty-Four — The Self That Cannot Give a Full Account

Ask someone why they chose the job, the partner, or the city, and you will receive a story. The story will be coherent, delivered with confidence, and there is a substantial experimental literature suggesting it is often not the reason.

The foundational study appeared in 1977 and has been replicated in various forms since. People asked to explain their choices produce explanations that are demonstrably unrelated to the factors the experimenters actually manipulated, and they produce them without hesitation and without any sense of having guessed. A later line of work is even more unsettling: subjects who choose one photograph over another and are then, by sleight of hand, handed the one they rejected will proceed to explain, in detail, why they preferred it. The explanation is constructed on the spot for a choice that was never made, and the person constructing it does not notice.

This is the empirical shadow of the book Butler published in 2005, and the convergence deserves more attention than it has had, because the two lines of work have no contact whatever and arrive at the same place.

Giving an Account of Oneself begins from a structural claim already familiar from earlier chapters: the self is formed in a scene of address that precedes it. Somebody spoke to me, categorized me, handled me, and named me before I existed as anyone capable of noticing, and the arrangements that did this were not chosen and are not fully recoverable. It follows that when I am asked to give an account of myself, there is a portion of the story I cannot supply — not because I am withholding it, and not because I have not tried hard enough, but because the formation happened to a being who was not yet there to observe it.

Butler is careful about what does not follow. This is not the claim that we know nothing about ourselves, which would be false and silly. It is the claim that a complete account is structurally unavailable, that the incompleteness is not a personal failing, and that any narrative I produce about my own origins is a reconstruction with speculative material in it. At the moment we narrate our own formation, Butler says, we become philosophers or fiction writers, and there is no third option.

From this comes an ethics, and it is the most usable thing in the corpus. If no one can give a full account of themselves, then a conception of responsibility that requires full transparency demands the impossible. And a demand for the impossible is not a high standard; it is a trap, because failure is guaranteed and the failure will be read as evasion. Butler’s proposal is that ethical relation begins from the recognition that the person opposite is opaque to themselves in the same way I am to myself, and that this mutual opacity is not an obstacle to responsibility but the condition under which responsibility actually operates.

Notice what is preserved, since this is where readers usually go wrong in the direction of thinking the argument excuses everything. Responsibility for what one did survives entirely. The account being declared impossible is the account of why, and the two come apart cleanly: a person can be answerable for an action while having no reliable access to its origins, which is in fact the ordinary situation of everybody who has ever apologized for anything.

Now let me press the argument where it earns its place in this book, which is on the institutions that demand the impossible account.

They are everywhere and they have multiplied. The disciplinary hearing that asks not only what happened but what you were thinking. The performance review that requires an employee to narrate their own deficiencies in the vocabulary of the organization. The parole board that assesses insight. The therapeutic setting that measures progress by the coherence of the story produced. The public apology, which has developed a form so precise that deviations are scored, and in which the offender must demonstrate not merely regret but a correct causal account of their own character. In each case an institution requires a person to produce something no person has: an accurate, complete, first-person explanation of their own conduct.

What is actually produced under such demands is well documented and entirely predictable. People generate accounts in the idiom the institution rewards. They learn which explanations are accepted and supply those. The accounts get better over time, in the sense of more fluent and more acceptable, and there is no reason to think they get more accurate, and some reason to think the opposite, since the fluency comes from practice at production rather than from access to anything. Interrogation research has a much darker version of the same finding, in which the account eventually produced is the one the interrogator was looking for.

The insight is not politically located, which is why I think it is the most valuable thing in this part of the corpus. Every institution that has ever demanded a public accounting of the interior has run into it: religious confession, party self-criticism, corporate values training, therapeutic culture, and the contemporary online ritual in which a person is required to produce, within hours, a statement demonstrating that they have understood the true nature of their offense. These have wildly different politics and identical structure, and the structure fails for the same reason each time.

There is a further consequence for how we should read anybody’s self-description, including in the debates of the preceding chapters. A person’s report of their own experience carries the authority described in Chapter Twenty: they are the only source on what it is like. Their report on why they are as they are carries no such authority, and — this is the part that stings — neither does anybody else’s report about them. The theorist explaining a person’s self-understanding as an effect of social forces is also constructing a narrative with speculative material in it, and has less access, not more. Opacity is not a weakness of the first-person position that the third-person position corrects. It runs in both directions, and both sides in most of these arguments have been claiming a transparency that nobody has.

The law has been dealing with this problem for centuries and its solution is instructive, since it neither demands the impossible nor abandons the inquiry. Criminal liability generally requires a mental element — intention, knowledge, recklessness — and the law does not attempt to obtain it by asking the defendant to introspect accurately. It infers the mental state from conduct and circumstance: what a person did, what they knew, what a reasonable person in that position would have foreseen. The interior is treated as something to be reconstructed from the outside rather than reported from the inside, precisely because the inside is understood to be unreliable and interested. Whatever else one thinks of criminal procedure, on this narrow point it is ahead of most of the institutions that have since started demanding insight.

Memory research supplies the last piece and it is the least comfortable, because it undermines the raw material rather than the interpretation. Recollection is reconstructive: each retrieval rebuilds the episode from fragments, and the rebuild incorporates whatever is available at the time, including information acquired since, suggestions embedded in the questions asked, and the requirements of the story being told. Confidently held, richly detailed memories of events that did not occur can be produced under laboratory conditions in a substantial minority of subjects. So an account of oneself is not merely an interpretation of reliable data. It is an interpretation of data that were themselves assembled, on demand, by a process with no commitment to accuracy.

The limitation of the argument should be stated too. Opacity can be invoked as a defense against any request for explanation, and it will be, by people who simply do not wish to answer. Butler is aware of this and the answer is the distinction already drawn — you may be unable to account for why, and you remain answerable for what — but the distinction is easier to state than to police, and in practice it will be blurred by anyone with an interest in blurring it. This is a real cost and there is no formulation that avoids it.

What I take from the book is a modest procedural conclusion with wide application. When designing any process that requires people to explain themselves — legal, professional, medical, or domestic — one should ask what the process would look like if accurate self-explanation were unavailable in principle. Frequently the answer is that the process should be asking about conduct, evidence, and consequences instead, all of which are inspectable. The demand for insight persists not because it works but because it feels like the deeper inquiry, and because a fluent account is enormously satisfying to receive. It is also, on the evidence, roughly as informative as a well-told story about a photograph nobody chose.

Chapter Twenty-Five — A Life That Counts as a Loss

A newspaper obituary is a small administrative act with a large metaphysical function. It says: here was a life, of the kind that is noticed when it ends. The form is standardized, the length is a matter of editorial judgment, and the decision about who receives one is made quickly by people under deadline pressure who would be surprised to hear that they are distributing a form of existence. They are, though. A death that generates no obituary, no name, no photograph, and no account of what the person had been doing is a death that the record does not register as the loss of anybody in particular.

Precarious Life, published in 2004, begins from this and it marks the point at which Butler’s work changes character. The earlier books were about how subjects are produced. This one is about how populations are sorted into those whose deaths register and those whose do not, and it was written in the aftermath of September 2001, in a country conducting two wars and an indefinite detention program, where the asymmetry was not subtle.

The organizing concept is grievability, and it needs stating carefully because the word invites sentimental misreading. The claim is not about anybody’s feelings. It is about a structural condition that precedes feeling: whether a life is constituted, in advance, as the kind of life whose ending would count as a loss. If it is not, then no amount of information about the death will produce grief, because the framework in which grief operates has already excluded it. Grievability is prior to grief in the way that being an eligible category is prior to being counted.

Butler’s second target in the book is the legal architecture, and the observation is precise. International humanitarian law developed to regulate conflicts between states and confers its protections largely on those who belong to, or act for, a recognized state. A person who belongs to no such entity, or who is designated as acting on their own account, falls into a category with very little law in it. Butler notes that the designation does further work: someone described as acting for no state and for no reason a state would recognize is thereby described as irrational, and irrationality justifies confinement without the inconvenience of a charge.

On indefinite detention itself Butler makes a distinction that improved on the available theory. The fashionable analysis, drawn from Agamben, treated such places as zones where law is suspended. Butler’s observation is that the situation is stranger and more durable than suspension: the government retains the option of applying the law or not, case by case, and the retention of that option is itself the exercise of sovereignty. It is not a hole in the legal order. It is a legal order that has arranged to be optional, which is considerably harder to challenge than an outright suspension because there is always a procedure to point at.

Now to the audit, and here I have something unusual to report. This is the most measurable idea in the entire corpus.

Grievability is not an abstraction. It cashes out into quantities that can be counted, and some of them have been. One can count obituaries. One can count column inches per death by nationality, and the ratios that emerge from such counts are stable, large, and not seriously disputed by anyone who has done the counting. One can count whether the dead are named or given as a number. One can examine whether a casualty-recording project exists for a given conflict, who funds it, and whether official sources will confirm its figures. One can look at whether a death produces an investigation. Each of these is a proxy for the distribution Butler is describing, and together they turn a philosophical claim into a research program that a graduate student could begin on a Monday.

Some of that work exists. Independent casualty-recording projects arose during the Iraq war precisely because official counting of civilian deaths was not being done, and the institutional resistance those projects encountered is itself evidence for the thesis: a state that does not wish certain deaths to be countable does not merely fail to count them, it disputes the legitimacy of counting. Media scholars have documented coverage asymmetries by nationality across many conflicts. The material is there. What has not happened is the assembly of it into a sustained empirical program under the concept, which returns us to the complaint of Chapter Nine, now in a different register.

Symmetry now requires the objection, and it is the strongest one available.

Differential grief is partly a universal feature of human beings and is not, in itself, an injustice or a product of ideology. Everyone grieves the near more than the far. This is not a moral defect to be corrected but a consequence of finite attention and of the fact that grief is a response to a particular absence in a particular life. A person who mourned every death on earth with the intensity they mourn a parent would not be a better person; they would be nonfunctional within a week. Any account that treats all asymmetry in grief as manufactured has misdescribed the baseline.

The correct form of the claim therefore requires a distinction that Butler makes less clearly than the material demands, and I would put it this way. There is proximity-driven asymmetry, which arises from the ordinary structure of attachment and would exist under any political arrangement. And there is institutionally produced asymmetry, which arises from decisions about what is counted, named, photographed, investigated, and reported. The second is not the same as the first, it is not explained by it, and it is the one that can be changed. Confusing them lets defenders of the second hide behind the naturalness of the first, which is exactly what happens whenever the subject comes up.

The distinction is testable, which is what makes it worth drawing. Proximity effects should scale with actual social distance and should be roughly symmetrical between any two populations. Institutional effects should track editorial and governmental decisions, should be asymmetrical in the direction of power, and should change when policies change. Those are different predictions and they can be told apart in the data. Nobody has done it, and it is the single most obvious study this literature implies.

How does an obituary desk actually decide? Not by malice and not by any written criterion. It works from a set of heuristics about significance which nobody has articulated and which everyone in the newsroom has absorbed: whether the person held a position, whether the death is unusual, whether an archive photograph exists, whether a colleague can be reached before the deadline. Each of these is defensible on its own and each correlates with an existing distribution of visibility. The output is a systematic pattern produced by no one’s decision, which is the most difficult kind of arrangement to change, because there is nobody to persuade.

Memorials make the same distribution permanent in stone, and the comparison is instructive because memorials are deliberate in a way newsrooms are not. It is close to a universal practice to inscribe the names of one’s own dead and to record the others, if at all, as an estimated number. Nobody experiences this as an assertion about whose lives counted; it is felt as the natural thing to do, and in a sense it is, since a community memorializes its own. But the cumulative effect across a century of monuments is a record in which one population consists of individuals with names and the other consists of a quantity, and any later inquiry into that conflict inherits the asymmetry as data.

There is one further complication that Butler does not take up and that ought to be recorded. Grievability is not only distributed between populations; it is distributed within them, and along lines that are not political in the usual sense. The deaths that receive the least attention in any wealthy country are those of the very old, the demented, the homeless, and the institutionalized, and the neglect is remarkably indifferent to the politics of the surrounding society. Any account which treats ungrievability as primarily a product of nationalism or war will miss the largest population it applies to, most of whom died in a facility, in the country doing the counting.

What I take from this book, on balance, is that Butler acquired a criterion here and did not fully realize what had been acquired. The proposition that lives should be equally grievable is a normative principle. It is thin, in that it does not by itself tell you what to do about anything. It is also non-trivial, because it generates immediate and unwelcome demands on institutions that count, report, and investigate, and because it can be checked. A criterion that produces measurable obligations is worth a great deal more than a comprehensive theory of justice that produces none, and it arrived, characteristically, in a book that most readers took to be about something else.

Chapter Twenty-Six — Frames Do the Work Before Judgment Does

In 1981 two psychologists asked people to choose between public health programs for an outbreak expected to kill six hundred people. One option would save two hundred lives for certain; the other offered a one-in-three chance of saving all six hundred and a two-in-three chance of saving none. Most respondents chose the certain option. A second group was offered a choice between a program in which four hundred people would die for certain and one with a one-in-three chance that nobody would die. Most chose the gamble. The two problems are arithmetically identical. Only the description differs, and the description reverses the preference of a majority of respondents.

This result has been replicated for four decades in every population anyone has tested, including physicians choosing treatments and executives choosing strategies. It establishes something that ought to be uncomfortable and mostly is not: the frame within which a choice is presented does substantive work before any judgment occurs, and the people making the judgment do not experience the frame as an influence. They experience themselves as responding to the facts.

The most important consequence of the framing research is one that its discoverers stated and that public life has entirely failed to absorb: a frame can be composed of nothing but true statements and still determine the outcome. Both versions of the epidemic problem were accurate. Neither contained a falsehood. This means that the standard remedy for public misinformation — verifying claims — leaves framing untouched, since there is nothing to correct. A fact-checking apparatus operating at full efficiency in a thoroughly framed environment will find very little to do and will change very little, and its existence provides reassurance out of all proportion to its effect.

Medicine supplies the version with the highest stakes, and it involves professionals rather than undergraduates. Present a treatment to physicians in terms of survival rates and it is chosen more often than the identical treatment presented in terms of mortality rates. The effect persists among experienced clinicians who know the literature and who would sincerely deny being susceptible. This is the finding that ought to be quoted whenever anyone suggests that framing effects are a curiosity of the psychology laboratory: they operate on trained experts making consequential decisions in their own field, and expertise does not appear to confer immunity.

Frames of War, published in 2009, is Butler’s book about the political version of this, and it is the best-argued of the later works. The thesis is that before any question of whether a war is justified can be raised, a prior operation has selected what is in the picture — which deaths are visible, which are described as combatants, what counts as an incident, where the account begins. Evaluation then proceeds honestly within a field that has already been shaped, and the participants argue about the evaluation, which is the only part they can see.

The material is drawn from a period rich in examples. Photography from military detention facilities in Iraq circulated in 2004 and produced a scandal about abuse; Butler’s interest is in why those images functioned as they did, why other images from the same conflict did not, and what the practice of embedding journalists within military units does to the frame before any journalist writes anything. Embedding is a particularly clean case: it is not censorship, no one is prevented from reporting, and it nonetheless determines the position from which everything is seen, which is a much more efficient arrangement than censorship and generates none of the resistance.

Butler’s distinctive addition to this — and it is a real addition, not a restatement of framing effects — concerns what happens to frames as they travel. A frame has to circulate in order to do its work: it must be reproduced, distributed, and shown. But circulation cannot be controlled, and a frame that arrives somewhere unintended can break with its own purpose. The detention photographs are the example: they were produced within an institutional frame that treated the events as normal, and their circulation outside that frame turned them into evidence against it. The mechanism here is the one from Chapter Ten, applied to images rather than to acts: what must be repeated to survive can always be repeated somewhere it was not meant to go.

That gives the theory a prediction, which is more than most framing analysis manages. Frames should be most vulnerable at the moment of widest distribution, and should fail in the direction of whichever audience they were not designed for. This is roughly what has happened repeatedly since, in cases where recordings made by institutions for internal purposes have become the principal evidence against them, and where material produced as propaganda has been received as confession by audiences the producers did not have in mind.

Now the audit, and there are two findings, one favorable and one not.

The favorable finding is the convergence. Butler arrived at the priority of the frame by philosophical argument, in a tradition with no interest in experiments, and the result matches one of the most robust findings in behavioral science. That is worth something. When two lines of inquiry with no contact reach the same structural conclusion, the conclusion is more likely to be about the world than about either method, and this is one of the few places in this corpus where such a convergence is available.

The unfavorable finding concerns discipline. Frame analysis has a characteristic failure mode: it can be applied to any disagreement whatever, and once applied it relieves the analyst of the obligation to assess the content. Every position becomes a frame; every frame can be described as serving somebody’s interests; and the analysis terminates in a description of the field rather than a judgment about the case. Butler does not always avoid this. The strongest passages identify a specific framing operation, specify what it excludes, and show what becomes visible when the exclusion is undone. The weaker ones observe that a matter has been framed, which is true of everything anybody has ever said, and is therefore not information.

There is a test that separates the two and I would recommend it to anyone reading this literature. A frame analysis is doing work if it can name what the frame excludes and if the excluded material can be independently obtained. Naming the excluded deaths and then producing the count is analysis. Observing that a debate has been framed, without specifying the exclusion or how one would check, is a gesture. The difference is not stylistic; it is the difference between a claim that can fail and one that cannot.

A related discipline concerns the direction of the accusation. Framing analysis is available to everybody and is used by everybody, and the term itself has been thoroughly captured: political consultants have offered framing services for decades, and the word now appears in campaign manuals as an instruction rather than in critiques as a diagnosis. A critical concept that has been adopted as a technique by the people it was meant to expose is not thereby refuted, but its user should notice that possessing it confers no advantage, since the other side has the same manual.

The deepest version of Butler’s point survives all of this and is worth isolating from the rest. It is not that frames influence judgment, which is a psychological finding. It is that the frame determines what enters the domain of the evaluable at all — that there is a class of events which cannot be judged unjust because they do not appear as events, and that this class is much larger than the class of events judged and excused. That is a claim about the shape of moral attention rather than its content, and it is not a claim that the experimental literature makes.

If it is right, then the practical implication is unglamorous and I think correct. The most consequential political activity is frequently not argument but counting: establishing a register, insisting on names, requiring that something be recorded in a form that later cannot be denied. This is tedious work, it is done by small organizations with no visibility, and it is much more difficult to defeat than any argument, because once a number exists in a citable form the frame has to accommodate it. Butler’s philosophical apparatus arrives, by a long route, at the conclusion that the clerks matter most, which is not what anybody expected and is, I suspect, true.

Chapter Twenty-Seven — Vulnerability as a Structural Fact

You are, at this moment, dependent on arrangements you did not build, cannot inspect, and would not survive the failure of for very long. Water arrives at your building through pipes maintained by people you will never meet. The electricity that keeps the food from spoiling is generated somewhere and delivered by a grid whose operators are managing frequency second by second. The drugs in the cabinet were manufactured in a facility on another continent and moved through a supply chain of some length. If any of this stops, your competence, autonomy, and personal virtue will be of very little assistance.

This is the observation on which Butler’s later political writing rests, and it is stated most directly in the work on precariousness and in the collaborative volume on vulnerability that followed. The claim has two parts, and it is worth separating them because the first is nearly undeniable and the second is where the argument lives.

The first part: vulnerability is not a property of particular groups but the general condition of embodied beings. A body is by definition exposed — to other hands, to weather, to infection, to hunger, to the failure of the arrangements just listed. Nobody has ever been self-sufficient. The image of the independent adult who enters into relations by choice, having somehow arrived at competence without a decade of total dependence, is a fiction that political philosophy has found convenient and that no one has ever seen.

The second part: exposure is distributed, and the distribution is arranged. Everybody is vulnerable; some are far more exposed than others, and the difference is not a fact of nature. It is the outcome of decisions about where the infrastructure runs, who has legal standing, whose work can be done from a distance, and who is required to be present in a room with strangers in order to eat. The philosophical claim about the human condition therefore terminates in an entirely concrete question about who is exposed to what, which is measurable, and this is the pattern this book keeps finding in the later work.

The pandemic settled the empirical form of the argument more decisively than any philosopher could have, and I say this as somebody wary of using a catastrophe to score a point. Within months there were figures: which occupations could be performed remotely and which could not, which households had space to isolate, which populations had legal status secure enough to seek treatment, and what the mortality differences were across those lines. The differences were large and they tracked the structural predictions rather than any individual variable. What had been a claim about the human condition became a table with confidence intervals.

The dependency argument has a longer history in philosophy than Butler’s framing suggests, and it is worth crediting because the neglect has been persistent. Feminist philosophers working on care and disability had spent decades pointing out that political theory is written as though its subjects were adults who arrived fully formed, and that everything which makes such an adult possible — years of dependency, the labor of whoever supplied it, and continued dependency at the end of life — is treated as belonging to a private sphere outside the theory. The observation is not new. What has changed is that a version of it now circulates in fields that had ignored it, which is a form of progress that its originators are entitled to find irritating.

There is a companion concept that has undergone a revealing career, and it is resilience. In its technical use it names the capacity of a system to absorb disturbance without reorganizing, which is a legitimate and measurable property. In policy it has come to name a virtue expected of individuals and communities, with the result that the failure to withstand a shock becomes a deficiency of those it fell on rather than a fact about the shock or the arrangements that transmitted it. That reversal is worth watching for, because it is one of the standard ways in which a structural finding gets converted into a demand on the people the finding was about.

There is a standing objection to all vulnerability talk that has to be met, because it is serious and it comes from people with standing to make it. Describing a population as vulnerable has a well-documented tendency to become paternalistic: the vulnerable are to be protected, protection is administered by others, and the administration reliably curtails the agency of those protected in the name of their safety. Whole regimes of guardianship, institutionalization, and protective legislation have been built on exactly this vocabulary, in every case sincerely, and in many cases with results that the protected would not have chosen.

Butler is aware of this and the volume co-edited with Zeynep Gambetti exists substantially to address it. The argument is that vulnerability and resistance are not opposites, that the same exposure which makes a body woundable is what makes it capable of appearing, acting, and being affected by others, and that an account which treats vulnerability as the absence of agency has already accepted the frame it should be questioning. I think this is right as far as it goes. I also think it does not fully answer the objection, because the objection is not about theory but about administration: whatever the philosophers intend, the word is used by institutions with budgets, and institutions use it to sort people into categories that determine what they are permitted to do.

The audit finding is therefore mixed, and the mixture is instructive. As a description of the human condition, the account is correct and better founded than the alternatives. As a political concept, vulnerability is dangerously fungible: it flatters the person applying it, it is available to every party in a dispute, and everyone has now noticed. Contemporary argument in most countries consists substantially of competing claims of vulnerability advanced by groups of very different sizes and resources, all of them sincere, none of them measured.

Which is where the audit has a recommendation, and it is the same one as always. If everyone is vulnerable, the word discriminates nothing and does no work. What discriminates is the distribution, and the distribution can be specified: dependence on infrastructures one does not control, absence of legal standing, absence of exit options, exposure that cannot be reduced by any action available to the person. Those are conditions with observable indicators. A claim of vulnerability that names them can be assessed; a claim that does not is a bid for standing, and should be received as one.

There is a second-order point here that has been on my mind since the chapter on persistence, and this is the place for it. Every arrangement that reduces someone’s exposure does so by increasing dependence on the arrangement. A water system removes the vulnerability of having to find water and installs the vulnerability of a system that can fail, be priced, or be switched off. A welfare provision removes exposure to destitution and installs dependence on an administration that can revise the rules. This is not an argument against infrastructure, which would be idiotic. It is an observation that vulnerability is never eliminated, only transformed, and that the relevant question about any protective arrangement is what new dependence it creates and who controls it.

That reframing has a practical edge. It suggests that the most valuable feature of a protective arrangement is not its generosity but its robustness against a change of management — whether it survives being administered by people indifferent or hostile to those it protects. Systems designed on the assumption of continued goodwill perform badly on this measure and are extremely common, because they are designed by the goodwilled. This is the same test proposed in Chapter Twenty-Two for speech regulation, arriving from an entirely different direction, and I take the convergence as evidence that the test is a general one.

The last observation is about what the vulnerability argument does to the concept of the individual, and it is the reason the framework irritates people across the political spectrum. If nobody is self-sufficient, then the independent chooser at the center of one political tradition is a fiction, and so is the autonomous self-defining agent at the center of another. Both traditions require a person who arrives at the scene already formed, and no such person exists. That conclusion has been available since long before Butler, it follows from facts about infant development that nobody disputes, and it continues to be ignored by both, which suggests the fiction is doing work that the truth cannot.

Chapter Twenty-Eight — Bodies in the Street

The question put to the occupiers of a lower Manhattan park in the autumn of 2011, over and over, by journalists who considered it devastating, was what their demands were. The absence of a list was treated as proof that the whole thing was unserious. Butler, who spoke there in October of that year, gave an answer that has worn better than the question: that the demands were being called impossible, that shelter, food, and employment were being called impossible, and that if these were impossible then the impossible was what was being demanded.

Notes Toward a Performative Theory of Assembly, published in 2015, is the book that grew from that period, and its argument is that the demand for demands misunderstands what an assembly is. A gathering of bodies in a public place makes a claim by gathering, prior to and independent of anything anyone says into a microphone. The claim is that these people may appear — that they exist as a public, that the space is theirs to occupy, that they are not reducible to whatever the administration has decided they are. On this account an assembly with an articulate list of demands and an assembly with none are both performing the same primary act, and the list is secondary.

This is the earlier theory of performativity extended into politics, and it is the most successful of the extensions. What is being enacted is not described by the words used; it is accomplished by the gathering, under conditions that give the gathering its force. The plural body in the square is not expressing a prior political position. It is constituting a public, in the same sense in which the earlier work said that acts constitute a subject.

Butler adds a materialist qualification that keeps this from floating away, and it is the best thing in the book. Appearance has infrastructural conditions. There must be a square, and it must be publicly accessible, which is a legal fact that can be revoked and increasingly has been in cities where nominally public space is privately owned. There must be transport to reach it. There must be sanitation, food, and somewhere to sleep if the occupation is to last more than a day, which is why every long occupation becomes, within seventy-two hours, an exercise in logistics. And the participants must be able to afford the risk: to lose a day’s pay, to be arrested without losing custody of a child or a visa, to be photographed without losing a job. The right to appear is not distributed evenly, and the distribution is a matter of who can absorb the costs.

Now the empirical question that this literature usually declines to ask: does it work?

There is a substantial dataset on this. The best-known study assembled several hundred campaigns aiming at maximal political objectives over the twentieth century and found that nonviolent campaigns succeeded roughly twice as often as violent ones, with participation levels as the strongest single predictor. From this came the widely quoted claim that no campaign mobilizing three and a half percent of a population has ever failed, which is now cited in a great many activist manuals.

That figure is exactly the sort of thing this book was written to examine, and the symmetry rule requires me to examine it here rather than where it would be more convenient. It is a historical observation about a particular dataset of maximalist campaigns in a particular period, drawn from a small number of cases at the top of the participation range, with all the fragility that implies. Its own author has repeatedly cautioned against treating it as a threshold or a rule, and has noted that success rates for nonviolent campaigns have declined markedly in the years since the dataset ends. It is a real finding that has been converted into a law of nature by people who found the law useful, which is the identical failure this book diagnosed in the intersex frequency statistic, occurring on the other side of the political ledger. A finding does not become more robust by being congenial.

What survives the deflation is still substantial and still favorable to Butler’s general position: participation matters more than tactics, and the physical presence of large numbers of people is a better predictor of outcomes than the articulateness of their demands. That is roughly what the book asserts, and it is unusual for a work of continental political philosophy to have empirical support of that kind waiting for it.

There is a further test that the last fifteen years have run without anybody designing it. Butler insists that bodies must be physically present and physically exposed — to each other, to weather, to police — and that this exposure is part of what an assembly is. Digital gathering, on this view, is not the same act. That claim looked conservative in 2015 and has aged extremely well. Movements organized primarily online have demonstrated an extraordinary capacity to produce enormous mobilizations very quickly, and a corresponding weakness in everything that comes after: negotiating, sustaining, deciding, and surviving repression. The best account of this argues that slow-built movements acquire organizational capacity as a byproduct of the difficulty of building them, and that networks which skip that stage arrive at the square with numbers and without the ability to do anything with them.

The mechanism is worth stating because it generalizes. Difficulty is not only a cost; it is also a producer of structure. An organization that spent three years persuading people to attend has, at the end of those three years, a list, a set of trusted relationships, an internal procedure for resolving disagreement, and a leadership that has been tested. An organization that assembled the same crowd in a week has a crowd. When the moment arrives that requires a decision, the first can make one and the second cannot, and this has very little to do with the merits of anybody’s cause.

The legal geography of appearance has changed considerably since the book was written, and mostly in one direction. A great deal of what looks like public space in modern cities is privately owned and merely open to the public, which is a different legal status carrying different rights of exclusion. Plazas attached to office developments, the concourses of stations, redeveloped waterfronts, and shopping districts frequently belong to this category. The result is that the right to appear has been quietly narrowed by property arrangements rather than by any law about assembly, and the narrowing is invisible until somebody tries to stand still in one of these places with a sign.

States have also learned. The characteristic response to a long occupation is no longer immediate clearance, which produces images that mobilize sympathy, but waiting: containment, attrition, selective enforcement of sanitation and permit requirements, and patience. This is an adaptation to precisely the mechanism the theory describes, and it works because presence is expensive to maintain and cheap to outlast. Any account of assembly that treats the confrontation as the decisive moment is describing a tactic that its opponents have already stopped using.

There is an unresolved tension in the book that I do not think Butler ever addresses, and it concerns duration. An assembly performs a claim by being present. Presence cannot be sustained indefinitely; people have jobs and children and eventually it rains. So the claim is made and then, necessarily, withdrawn, and everything that follows depends on institutions of exactly the kind the theory is least interested in: organizations, offices, funds, and people whose job it is to attend meetings for a decade. The theory has a beautiful account of the moment of appearance and nothing at all to say about the Tuesday afternoon eighteen months later, which is where outcomes are actually determined.

This is not a small omission and it recurs across the political writing. There is a durable preference in this tradition for the moment of rupture over the machinery of consolidation, and it is a preference rather than a finding. It may be aesthetic; assemblies are photogenic and committees are not. It may be a legacy of the theoretical apparatus, which is built to identify what exceeds a structure and has few tools for describing how a new structure is built. Either way it means that the literature is at its strongest in describing why a square fills and nearly silent on why, three years later, nothing has changed.

The audit finding: on what an assembly does, correct and well-supported. On the infrastructural conditions of appearing, an important and underused contribution that connects the philosophy to things a city council actually decides. On what happens after everyone goes home, missing, and the omission is systematic rather than accidental.

Chapter Twenty-Nine — Nonviolence Without Serenity

There is a version of nonviolence that involves speaking softly, forgiving readily, and maintaining an inner calm under provocation. It has a distinguished history and a great many practitioners, and it is not what Butler’s 2020 book is about. The Force of Nonviolence opens by rejecting the association: nonviolence, it says, is not a serene condition of the soul or a private ethical stance, and treating it that way removes it from politics, which is where it belongs and where it must operate.

The reconstruction runs roughly as follows. Nonviolence is a practice, not a temperament. It is entirely compatible with rage; indeed the interesting question is what a person does with rage, not whether they have it, and a nonviolence available only to the placid would be useless. It is a form of force rather than an absence of force. And it is grounded not in a distaste for blood but in a substantive claim about equality: that all lives are equally grievable, and that violence is the practical denial of this, since it acts on the premise that some lives may be expended for others.

That grounding is worth pausing over, because it is the thing whose absence was the substance of the most famous attack on this body of work. Here is a normative principle, stated as such, from which practical conclusions follow. The person who wrote in 1999 that no criterion of justice or dignity was on offer was correct about the material then available, and the gap has since been filled by exactly the sort of thing they said was missing. Whether the filling is adequate is a separate question, and I will come to it.

The book’s sharpest passages concern self-defense, and this is where I think it makes a genuine contribution that has almost nothing to do with the rest of the framework. Nearly every justification of violence in the modern world proceeds by way of self-defense, and the argument turns entirely on the boundary of the self being defended. Sometimes the self is a body under immediate threat, which is the case everyone imagines and about which there is little disagreement. But the self expands. It expands to include a household, then property, then a neighborhood, then a nation, then a way of life, then a set of people who resemble oneself and are imagined to be at risk somewhere. At each expansion the same justification is claimed, with the same emotional force, and the boundary is never argued for — it is assumed in the framing, which is where Chapter Twenty-Six said the work always gets done.

That observation is checkable against actual justifications of actual violence, and it holds up remarkably well. The most consequential feature of any self-defense argument is the scope of the self, it is nearly always implicit, and questioning it produces indignation rather than an answer, which is the signature this book has taught the reader to recognize.

Now the difficulty, and it is a real one that admirers of the book have not addressed.

An ethic of nonviolence requires a stable definition of violence, and Butler spends a considerable portion of the book destabilizing it. The argument for the destabilization is good: what counts as violence is determined by a frame, state action is routinely renamed as security or order, and structural arrangements that shorten lives by decades are not counted as violent at all while a broken window is. All true. But if the definition of violence is contested in this way, then the injunction against violence cannot by itself determine what anyone should do, because any proposed action can be redescribed by an interested party as violence or as its prevention. A principle that cannot be applied without first settling a contested definition is not yet a principle; it is a placeholder for an argument about the definition.

Butler would say, I think, that this is precisely the point — that the argument about what counts as violence is the political argument, and that pretending otherwise conceals the operation. Fair enough as analysis. It leaves the person who wanted guidance where they started, and the book is written as though it were offering guidance.

There are two distinct traditions here that the book treats as one, and separating them clarifies a good deal. One is principled and religious in origin, holding that violence damages the person who commits it regardless of outcome; its great practitioners understood themselves to be doing something spiritual as well as political. The other is technical, and its most systematic exponent catalogued nearly two hundred specific methods of nonviolent action — strikes, boycotts, occupations, refusals of cooperation — analyzed as instruments for withdrawing the consent on which any regime depends. That second tradition is engineering, and it is entirely indifferent to the state of anyone’s soul. Most successful campaigns have used both, and the theoretical literature persistently discusses only the first.

The definitional problem has a legal counterpart that shows how practical it is. Whether the destruction of property counts as violence has been argued in courts, in movements, and in newspapers for two centuries without resolution, and the reason it cannot be resolved is that the answer depends on what one takes violence to be for: harm to persons, breach of order, or the imposition of one will on another. Each definition is coherent, each is held sincerely, and each licenses a different set of actions. This is not an abstract worry about definitions. It is the question that determines what a given movement will do on a given night, and no philosophical treatment that leaves it open is offering guidance.

There is a second issue, which is the relation between principle and tactic. Nonviolence can be adopted as a moral commitment that holds regardless of consequences, or as a strategy on the evidence that it works better. These are different positions with different failure conditions, and they are routinely blurred, including here. The blurring is convenient because each covers the other’s weakness: when the strategic case is challenged, one retreats to principle; when the principle is challenged as impractical, one cites the effectiveness data. That is not an argument, it is a pair of arguments taking turns, and anyone advancing either should say which one they would keep if the other failed.

The evidence question deserves an honest statement, since I raised it in the previous chapter and it applies with more force here. Nonviolent campaigns have historically succeeded at higher rates than violent ones, and the difference is large enough to survive most reasonable objections about how success is coded. That comparison is between campaigns, not between moments, and it says nothing about whether a particular tactic on a particular evening was wise. It also comes with the caveat already entered: the historical advantage appears to have narrowed considerably in the last fifteen years, for reasons that are actively debated and may include the improved capacity of states to surveil and to outlast.

What I find most durable in the book has nothing to do with the theory of violence and everything to do with the theory of the self it presupposes, which the reader will recognize from much earlier. If nobody is self-sufficient, if everyone is formed by others and dependent on arrangements they did not choose, then the boundary of the self that self-defense protects is not given by nature. It has to be drawn, and drawing it is a political act performed before any question of defense arises. Every argument about legitimate violence is therefore, underneath, an argument about who is included in the entity being defended, and it is conducted as though that were settled.

I would add one observation the book does not make and that seems to me to follow. The expansion of the defended self has an observable signature: the more abstract the entity said to be under attack, the wider the range of people who can be described as attacking it, and therefore the less discriminating the violence that can be justified. A person defending their body can identify their assailant. A person defending a civilization cannot, and does not need to. That is not a moral claim but a structural one about how the scope of a justification determines the scope of its application, and it can be checked against any conflict one cares to examine.

The audit finding on this book is the most favorable in the second half of the corpus, with one qualification. The analysis of self-defense is original, checkable, and useful. The equality-of-grievability principle is a genuine normative commitment and answers a long-standing charge. The instability of the central term is a serious structural problem that is acknowledged only obliquely, and it means the book functions much better as a diagnosis of how violence is justified than as guidance about what to do, which is what it was received as and what its title promises.

Chapter Thirty — The Effigy in São Paulo

In November 2017 a crowd outside a cultural center in São Paulo burned an effigy of Judith Butler. The figure had been dressed in a bright pink bra and given a witch’s hat. Some in the crowd carried crosses and national flags. An online petition against the visit had gathered several hundred thousand signatures. Butler and their partner were later confronted at the airport by a group holding photographs and signs telling them to go home, or somewhere considerably worse.

The detail that ought to be at the center of every account of this episode is almost never mentioned. Butler was not in Brazil to speak about gender. The event was an academic colloquium on the future of democracy, organized around questions of populism, sovereignty, and political theology, and Butler was one of its organizers. The subject that produced the effigy was not on the program.

That is the whole anatomy of a moral panic in a single fact, and it is worth working through slowly, because the mechanism is general and the temptation to treat this as a story about one country’s politics is exactly what makes it unusable.

Begin with the declared component. The mobilization rested on the claim that Butler had come to promote a doctrine in Brazilian schools. This was not a measurement; it was an attribution, circulated at speed, requiring no contact with any text or program. As with the excommunication in Chapter Two, the community acted on an assigned position rather than on anything said, and the assignment was performed by intermediaries with their own reasons. The petition did not misquote the lectures. It preceded them, and the lectures were about something else.

Then the personification. A movement opposing a broad social change requires an author for it, because a diffuse shift with no agent cannot be argued against, protested, or defeated. Demographic change, urbanization, and the collapse of a settled arrangement of family life do not have a face. A foreign philosopher with a photograph, an unreadable book, and a name that can be printed on a sign is enormously more convenient. This is why the accusation is always so poorly matched to the accused’s actual claims: correspondence to the texts is not what the figure is for.

The iconography repays attention too. The witch’s hat is not decoration, and neither is the pink bra. Both belong to a specific and old repertoire in which a woman who teaches doctrine is not merely wrong but transgressive in her person, and the appropriate response is public burning in effigy. One scholar has documented how Butler was depicted in that period as a literally diabolical figure, with the imagery drawing on both gender and Jewish identity, which is a combination with a long and well-documented history in European iconography and which nobody involved would need to have studied in order to reproduce.

The intermediary layer deserves specific attention, because panics do not assemble themselves. Between the diffuse anxiety and the crowd there is always a small number of people who convert one into the other: who identify the target, supply the vocabulary, produce the shareable material, and take on the work of coordination. This is a role with its own incentives, and in the contemporary version it is frequently a paid one, since attention converts to revenue and outrage is the cheapest form of attention to manufacture. Any explanation of an episode like this that omits the entrepreneurs is missing the causal step that turns a mood into an event.

The technology matters too, and specifically its speed. A petition that would once have required months of clipboards can now acquire several hundred thousand signatures in days, from people who have read the headline and nothing else, at a cost to each signatory of approximately four seconds. The resulting number is then reported as a measurement of public sentiment. It is nothing of the kind; it is a measurement of how many people encountered a particular item in a feed and found clicking easier than not clicking, and it carries no information about depth of conviction, understanding of the issue, or willingness to do anything further. Treating such numbers as evidence of anything is a category error that both sides of every dispute now commit hourly.

Now the symmetry that this chapter must include if it is to be worth anything.

Moral panics are not a property of any one politics. The structure — an attributed position, a personified author, a rapid mobilization, a demand for exclusion, and an emotional register out of all proportion to anything the target has done — has been assembled by every political formation that has ever had numbers. It has been directed at people for their religion, their irreligion, their supposed corruption of the young, their music, their books, and their associations. It is currently being assembled in several countries against academics of quite different views by movements with quite different aims, using the same components. Anyone who can see the mechanism only when it is aimed at people they like has not seen the mechanism.

There is also a symmetry closer to home for the academy. Universities have their own demonologies, their own attributed positions, and their own occasional episodes in which a figure is constructed out of a misquotation and then addressed as though the construction were the person. The scale is different, the consequences are usually smaller, and the structure is recognizable. A method that exposes the operation in São Paulo and cannot find it in a faculty meeting is not a method.

Now the hardest and most necessary point in the chapter, and I want to state it flatly because the alternative is a kind of writing I dislike.

Being burned in effigy does not make a person’s arguments correct. Persecution is not evidence. The history of ideas is full of people who were hounded, exiled, and killed for positions that were mistaken, and full of others hounded for positions that were right, and the hounding does not distinguish them. Every one of the arguments examined in this book has to stand or fall on its own, and nothing in this chapter alters the assessment of any of them. The temptation to let an outrage settle an intellectual question is very strong, it is available to all parties, and it should be refused by all parties.

What the episode does establish is narrower and still worth having. It establishes that the vehemence of opposition to this body of work is not proportionate to, and frequently unconnected with, the content of the work — since the content was, on that occasion, not even the topic. Any explanation of the phenomenon that proceeds by way of what the books say is therefore incomplete, and probably wrong. Something else is being opposed, and the books are its available name.

That something else is the subject of the next chapter, and Butler’s own eventual account of it is more useful than the accounts offered by observers, which is a rare situation for a person to be in with respect to their own persecution.

One further episode belongs here for contrast. In 2010 Butler was offered a civil courage award at a Berlin parade and declined it on the stage, citing what they described as racist remarks by organizers and a failure of such organizations to distance themselves from anti-Muslim justifications for war, and naming other groups they considered stronger opponents of the various things the prize claimed to oppose. Whatever one thinks of that judgment, it is the same instrument used in the opposite direction: the refusal of an honor is an examination of what an institution’s declared purpose has to do with its conduct. The person who was made into a figure in São Paulo had, seven years earlier, declined to be made into a different figure in Berlin, and the consistency is worth recording, since it is rarer than either admirers or opponents of anybody generally allow.

Chapter Thirty-One — A Phantasm Is Also a Role Without a Referent

Ask an opponent of gender ideology what the doctrine states, who formulated it, and where it is written down, and the answers will not converge. It is variously described as a denial that biological differences exist, a program for reassigning children, an imposition by wealthy northern countries on the global south, an instrument of capitalist atomization, a Marxist project, a project of international finance, and a plot to abolish the family. Several of these are incompatible. The incoherence is not a weakness of any particular spokesman. It is a stable feature of the phenomenon, observable across languages and continents, and it has persisted for two decades.

Butler’s 2024 book takes this as its subject and proposes that the object of opposition is a phantasm: a figure assembled from many anxieties, without a stable referent, whose power comes precisely from its capacity to absorb whatever a given audience already fears. On this account the movement is not confused about its target and would not be improved by better information, because the target’s function is to be available for whatever is locally frightening, and a target with a fixed definition could not perform that function.

The reader who has come this far will have noticed what has happened. A phantasm is a role without a referent. It is an entity specified entirely by the effects attributed to it, introduced because something must be producing the felt disorder, named as though the naming settled the question, and defended against the demand for an occupant. It is the structure of Chapter Seven, arriving in the last book of the corpus, and Butler has diagnosed in opponents the precise operation this book has spent thirty chapters auditing in Butler. That is either a satisfying closure or an uncomfortable one, and I think it is both.

Because symmetry now requires the obvious question, and this chapter would be worthless without it: does Butler’s own object have a referent?

The anti-gender movement, in the singular, is a construction. What exists is a family of campaigns with different origins, different aims, and different vocabularies. Some descend from documents produced in Rome in the 1990s and carry a specific theological vocabulary. Some are nationalist and treat the matter as a question of foreign influence. Some are secular and feminist and are opposed to positions that the religious campaigns also oppose, for reasons that the religious campaigns would reject. Some are electoral opportunists with no discernible commitments at all. These groups have coordinated in places and are elsewhere hostile to one another. Treating them as one entity is a declaration, and it is exactly the operation the book is diagnosing.

There is a defense available and Butler makes it: that the campaigns share a vocabulary, a target, and a set of transmission networks, which is enough to constitute a phenomenon even without a central organization. That defense is reasonable and is supported by work tracing the circulation of specific phrases and funding across borders. It establishes a family resemblance. It does not establish a single object, and the book’s rhetoric frequently requires the stronger thing.

Then there is the word fascism, which Butler applies to the phenomenon and which requires a harder look, because the reasoning offered for it is unusual. Butler cites Umberto Eco’s account of fascism as a structure containing contradictions, and uses the incoherence of anti-gender rhetoric as evidence that the label fits.

Consider what that argument does. If coherence would count against the diagnosis and incoherence counts for it, the diagnosis cannot fail. Every observation confirms it. This is precisely the shape of the entities Chapter Seven warned about: an explanation that accommodates all possible evidence has stopped explaining and started labeling. And the label is not neutral: fascism names a specific historical formation with a specific relationship to the state, the party, the corporation, and the street, and applying it to a heterogeneous set of campaigns because they are internally inconsistent stretches it past the point where it carries information.

I want to be careful about what this finding is and is not. It is not a defense of the campaigns in question. Their factual claims are, on the whole, checkable and mostly fail, and I have no interest in softening that. It is a finding about a specific argumentative move: that a historically loaded term was applied on the basis of a criterion which could not have failed, in a book whose central diagnosis is that opponents deploy an unfalsifiable figure. The move is the mirror image of the thing being criticized, and it appears roughly eighty pages after the criticism.

A related episode from the following year illustrates a different mechanism and is worth recording because it is institutional rather than popular. In September 2025 the University of California, Berkeley, informed a hundred and sixty members of its community that their names and files had been provided to a federal department in connection with an investigation into alleged incidents of antisemitism on campus. Butler was among them, wrote a public essay condemning the university for acting as an informant, and stated in an interview that no details of the allegations had been supplied and that the institution had not followed its own internal procedures. Butler is of Jewish heritage and lost family in the Holocaust, a fact several commentators noted in observing that the most culturally prominent name on the list belonged to a Jewish public intellectual.

I am not going to adjudicate that dispute, which is recent, contested, and still developing, and where I have no access to the underlying material. What is relevant here is structural. An accusation was transmitted without specification: a name was attached to a category, the content of the allegation was not disclosed to the person named, and the mechanism supplied no route by which it could be contested, because there was nothing stated to contest. Whatever one thinks of the substantive questions, the procedural form is the one this book has been describing throughout. A role was created — the person under investigation — and the referent, meaning the actual alleged conduct, was not supplied. It is a striking thing to happen to the author of the analysis.

There is a lesson in the pairing of these two episodes that goes beyond anybody’s case. The unspecified accusation and the unspecified doctrine are the same instrument. In one, a figure is denounced for a doctrine that has no text; in the other, a person is named for conduct that is not described. Both work because the absence of specification prevents refutation, and both are available to any party with the institutional or numerical power to make an accusation stick without stating it. Nobody has a monopoly on this technique and nobody should be surprised when it arrives from the direction they were not watching.

It is worth saying why phantasm is a better term than the obvious alternative. To call something a conspiracy theory is to claim that its holders believe in a coordinating agency, and many of them do not; the campaigns include a great many people who would say only that something has gone wrong with the world and that this word names it. A phantasm requires no plotters. It requires only a figure onto which diffuse anxiety can be attached, and it explains the observed features that the conspiracy framing does not: why the accusations do not need to be consistent with one another, why refutation produces no effect, and why the emotional intensity is so poorly matched to the specific claims.

The checkable claims should be checked, since fairness requires it and since the results are not what either side asserts. Some of them are simply false as stated: no text propounds the doctrine described, and the passages usually cited say something else. Some rest on real changes whose scale is disputed, where the honest answer is that the data are thin, recent, and contested, and where confident assertions in either direction outrun the evidence. And a few identify genuine questions that deserve better than they have received from either camp, particularly around evidence quality in areas where the studies are small and the follow-up is short. A movement can be organized around a phantasm and still, in among the noise, be pointing at something real, and an audit that could not register that would be a poor instrument.

So what is left of the book? Rather more than the criticism above suggests. The core observation — that the opposition is organized around an object with no stable content, and that this is functional rather than accidental — is correct, well evidenced, and explains a great deal that no rival account explains, including why better information changes nothing and why the accusations do not correspond to any text. The description of how the vocabulary travels between countries and constituencies is careful and supported. And the book is, by common consent, the most readable thing in the corpus, which after thirty-five years is not a small thing to be able to say.

What does not survive is the ambition to name the phenomenon with a term from another century on the basis of a criterion that could not have come out the other way. Butler spent a career establishing that entities introduced by their effects and defended against the demand for a referent are doing something other than explaining. The last book contains a first-rate application of that insight and one clear violation of it, in the same argument, and the violation is the part that made the headlines.

Chapter Thirty-Two — One Idea Reported Many Times

There is a failure mode in quantitative work that is easy to describe and very hard to see from the inside. A model produces a dozen predictions. Each is compared against data. Each is confirmed. The author reports a dozen successes. But if all twelve predictions descend from a single underlying quantity multiplied by fixed factors, then they are not twelve tests. They are one test reported twelve times, and the impression of overwhelming support is an artifact of the bookkeeping.

The diagnostic is to ask how many independent dials the apparatus really has. If moving one dial moves every needle, then the needles are not independent witnesses. They are a display. And a display can look like an avalanche of confirmation while resting on exactly as much evidence as a single measurement.

I raise this because it is now my turn, and I would rather perform the operation than have it performed on me.

This book has produced, chapter after chapter, a finding of the same shape: something was declared and reported as measured. It found this in cosmology, in the definition of the kilogram, in poverty statistics, in the frequency of intersex conditions, in the categories of feminist theory, in the symbolic order of psychoanalysis, in the framing of casualties, in the 3.5 percent figure, in the classification of the anti-gender movement, and in the design of speech law. That is a great many confirmations. It is also, quite possibly, one finding restated thirty times against different material, which would make this book a display rather than an avalanche.

So: how many dials does the taxonomy have?

The honest answer is that it has fewer than the number of chapters, and that the recurrence of the finding is partly a property of the instrument. Any claim whatever can be decomposed into elements that were chosen and elements that were observed, because every claim contains both. An instrument that can be applied to anything and that always finds something is not measuring a feature of its targets. It is describing the general condition of assertion.

That is the strongest objection to this entire book and I do not think it can be fully answered. What can be offered are three partial defenses, and the reader should weigh them and decide.

The first is that the instrument does produce variance, which a pure display would not. Some chapters came out clean. The 1988 essay was found to have visible declarations and short derivations, and I said so. The analysis of framing was found to converge with independent experimental work. The prediction about speech regulation was found to be confirmed by administrative history across multiple jurisdictions. If the taxonomy guaranteed its own confirmation, those chapters would not read the way they do, and I would have had to work harder to make them read the way they do.

The second is that the taxonomy is not a theory and should not be assessed as one. It makes no predictions about the world; it is a procedure for sorting the components of an argument. Procedures are judged instrumentally: does using this change what anyone concludes? In this book it did, in identifiable places. It required me to withdraw a statistic I had used against one side and to withdraw the rival statistic I would have preferred. It required me to record that the most famous critique of Butler had been overtaken by later work. It required me to say that a figure widely used by people whose politics I find congenial has been inflated past its warrant. Each of those cost something and none of them was the conclusion I set out with.

The third defense is the weakest and I offer it anyway: that the recurrence of a single finding across unrelated domains is what one would also expect if the finding were true. A general feature of reasoning would show up generally. This is a defense that could be offered by anyone with a favorite lens and it should be discounted heavily. I include it because leaving it out would be a small dishonesty of the kind this chapter exists to prevent.

There is a sharper version of the test that I have been avoiding, so let me run it. The apparatus of this book has several apparent parts: the three-way sorting of components, the distinction between free and fixed declarations, the anchoring diagnostic, the deferred-versus-foreclosed distinction, and the rank check now being applied. Are those five independent instruments or one? On inspection, anchoring is a special case of declaration — a calibration point is simply a declared component with an unusually large downstream effect. Free-versus-fixed is a refinement of the same notion. Deferral and foreclosure are two ways of handling one declared vacancy. Which leaves perhaps two genuinely independent ideas wearing five names, and a book organized around five headings that were doing rather less work than their number implied.

The counterfactual test is worth running too, and it is the one that would settle the matter if it could be run. Would I have reached the same conclusions without the apparatus? For several chapters, honestly, yes: that the prose is a barrier, that a criticism was overtaken by later work, that a movement’s object has no text — these could have been reached by an attentive reader with no taxonomy at all. For a few, I think not. The distinction between a deferred and a foreclosed referent is not something I would have formulated without the framework, and it is the hinge of the book’s central argument. So the apparatus earned its place in a minority of the chapters and supplied vocabulary in the rest, which is a more modest claim than the preface made and, I suspect, the usual situation for any method.

There are two further self-audits owed, and they concern selection rather than the instrument.

The first was flagged in Chapter Seventeen. My chain is pinned at the point where a claim meets an instrument. That anchoring has consequences: it makes visible everything that can be checked and systematically undervalues everything that cannot. Which means that a body of work concerned with what it is like to be unintelligible to the arrangements one lives in — a subject that resists measurement by construction — will look thinner under my instrument than it may be. Somebody anchored at the point where a claim meets a life would produce a different book. It would find things I have missed, and would miss things I have found, and there is no third position from which to referee.

The second is more damaging and I noticed it late. I chose which of Butler’s works to treat closely. The selection rule, which I did not state at the time because I had not articulated it, was that I treated most fully the works with checkable content. That rule biases everything. It privileges the later, more empirical, more political books, and it thins out the psychoanalytic middle period — the material on melancholia, on Lacan, on kinship — which is precisely what Butler’s most serious readers regard as the center of the enterprise. My general verdict, that the later work is stronger, may therefore be an artifact of my own selection rather than a finding about the corpus. I do not think it is entirely an artifact. I am not in a position to be confident that it is not.

A final item for the ledger, and it concerns incentives rather than method. A book that audits a famous thinker and finds substantial fault is a more publishable book than one that finds everything in order. I knew that at the outset. Every judgment in these pages was made by somebody who stood to benefit from severity, and the appropriate correction is not for me to protest my impartiality but for the reader to apply a discount in the direction the incentive runs. If a finding here seems generous to Butler, it was probably made against my interest and can be weighted accordingly. If it seems harsh, remember who was holding the instrument and what he was rewarded for finding.

That is the whole of my self-audit and it is not comfortable reading for its author. It leaves the method standing but smaller than it looked in the preface: a serviceable procedure with a known bias, which has changed some conclusions and confirmed others, applied by somebody with an interest in the outcome. A person who wanted a stronger warrant than that will have to build a different instrument, and should send me the results.

Chapter Thirty-Three — What Survives the Audit

An audit ends with a ledger rather than a verdict, and this is the ledger.

Consider first what survives intact — the propositions that hold under examination, that do not depend on any declared component doing hidden work, and that would have to be defeated on their merits by anyone wishing to be rid of them.

The empty subject position survives. It is an inherited philosophical commitment with a documented pedigree running back through Sartre, Hyppolite, Kojeve, and Nietzsche; it was declared openly in a dissertation on those very figures; and the criticism that treats it as a personal eccentricity or an evasion has not read the earlier book. The diagnosis of classical feminist theory survives: in that framework, sex functioned as a natural bedrock while being available only through the apparatus that interpreted it, which is a declared component reporting itself as a measurement, and the diagnosis is correct whatever one concludes from it. The account of subject formation through subjection survives, and does so with support from an experimental literature that developed independently and knows nothing of it. The prediction that powers created to restrict speech will migrate toward whoever holds administrative office survives, confirmed across jurisdictions and decades, and is the most successful empirical claim in the corpus. The priority of the frame survives, converging with four decades of experimental work. The analysis of the expanding self in justifications of violence survives and is original. The opacity of self-knowledge survives, matching what the psychology of self-report has independently established. And grievability survives as something rarer than an insight: a normative criterion that generates measurable obligations.

Next, what survives with damage — the propositions that are defensible but cost more than their proponents admit.

The foreclosure of the inner core is coherent, brave, and exposed, and it lacks any specified route to refutation, which leaves it in the strange condition of being maximally risky in principle and untouchable in practice. The account of materialization is sound in substance and systematically overstated in its grammar, and the overstatement is what the book’s reception was built on. Iterability explains why change is always possible and supplies nothing about direction, which leaves a hole where an account of political outcomes should be. Vulnerability is correct as a description of the human condition and has proved almost infinitely fungible as a political concept, available to every party and measured by none.

Then what does not survive.

Resignification does not survive as a remedy: slow, unreliable, unplannable, and assigning the labor to the injured. Subversion does not survive as a criterion, and Butler eventually said so. The normative gap between 1990 and 1997 was real, was correctly identified by the most famous critic of the work, and was filled later rather than answered at the time. The treatment of a heterogeneous family of campaigns as a single object under a historically specific label does not survive, because the criterion offered for the label could not have failed. And the absence, across thirty-five years and an enormous secondary literature, of any program to measure what the theory implies is measurable does not survive as a defensible allocation of a field’s attention.

Finally, what is not attributable at all. The changes in attitudes and arrangements across the developed world since the 1970s cannot be laid at this door, in either direction. They began earlier, they occur where these books are unread, and they track demographic and economic variables that move first. Both the credit and the blame are misdirected, and the misdirection serves the needs of the people assigning it.

Now stand back and look at the shape of that ledger, because it has one, and the shape is the thesis promised in the preface.

What survives best is what was attacked most. The empty subject position, the diagnosis of a declared bedrock, the account of formation through power: these are the propositions that produced the outrage, the effigy, the prize, and the thirty-year public quarrel, and they are the ones standing when the examination is done. What survives worst is largely what admirers defended most warmly. The vocabulary of subversion, the promise of transformation through repetition, the sense that a seminar in destabilization was a form of political action: these were the parts that made the work exciting, and they are the parts that have not held.

That mismatch is not unique to this case. It is close to a general rule in intellectual history, and the reason is not mysterious. A body of work is attacked at its most exposed claim, which is usually its most rigorous one, because rigor is what makes a claim exposed. It is admired for its most usable claim, which is usually its loosest, because looseness is what makes a claim usable. The attackers and the admirers are therefore looking at different parts of the same object, neither is looking at the whole, and the eventual verdict of an audit will disappoint both.

There is one further finding that belongs at the end and that I have been circling since Chapter Nine. The largest single failure in this literature is not a false claim. It is an unbuilt research program. A theory that describes gender as a structure maintained by continuous repetition implies measurable quantities: maintenance costs, enforcement thresholds, drift rates, hysteresis, characteristic failure modes, and the distributions of grievability, exposure, and appearance. None of these was measured. The intellectual energy went instead into commentary, defense, and interpretation, until the secondary literature exceeded the primary by orders of magnitude and contained, so far as I can tell, not one number that the theory predicted.

I do not think this was anybody’s decision. It is what a field does when its prestige is allocated for interpretation rather than for measurement, and the incentive structure that produced it is not peculiar to this subject. But it is the reason that a body of work with several confirmed empirical claims to its name is discussed everywhere as a matter of opinion, and it is the answer to the question of what was lost. What was lost was thirty-five years in which somebody could have been counting.

A word on how to use a ledger of this kind, since a mixed result is harder to act on than a verdict and readers reasonably want to know what to do with it. The entries are separable. Nothing in this book requires anyone to accept or reject the whole. A reader may take the analysis of speech regulation and leave the account of materialization; may find the treatment of grievability indispensable and the theory of assembly incomplete; may accept every descriptive finding and none of the political ones. This is not fence-sitting. It is what it looks like to hold positions at the level of individual claims rather than at the level of authors, and the alternative — deciding about a person and inheriting all their propositions as a set — is how almost all intellectual allegiance actually works and is indefensible on any account of reasoning anybody has ever proposed.

The ledger is also provisional in a specific and unusual way: its subject is alive and still publishing. Several entries here were written differently in my notes two years ago because the corpus was different then, and the most famous critique in the field is filed under overtaken by events for exactly this reason. Any audit of a living body of work is a snapshot of a moving object, and the appropriate humility is not the ritual sort but the practical sort — a willingness to say which entries would move, and in which direction, if the next book arrived tomorrow and did what the last three have done.

A last word about what an audit cannot do, since I have been careful to claim little and should be careful not to claim it grandly at the end. This procedure cannot tell you whether Butler is right about the deepest question, which is whether there is anything at the address that was declared empty. It cannot settle any political disagreement, because political disagreements are about purposes and no examination of arguments determines a purpose. It cannot even tell you whether reading these books is worth your time, which depends on what you want and what else you might read instead.

What it can do is separate what was seen from what was chosen, and hand you the two piles. That is a smaller service than a verdict and a more durable one, because a verdict has to be accepted or rejected while a ledger can be checked. If a reader disagrees with an entry in mine, the disagreement will be about a specific item, in a specific place, for a stated reason — which is the only kind of disagreement that has ever gone anywhere.

Conclusion

A rabbi in Cleveland once tried to punish a child for talking too much, and the punishment failed in a way nobody could have predicted from the paperwork. That is where this book began, and it is worth returning to, because the whole of what follows was an elaboration of a single small fact: an action’s name is not its effect, and the difference has to be looked at rather than assumed.

Everything in the preceding chapters has been a version of that. A statistic is a definition wearing the clothes of a measurement. A planet is a job description that nobody ever filled. A projection is a decision about which distortion to accept, presented as a picture of the world. An anchor is a choice that determines the answer, made before any data arrive and never mentioned in the same paragraph as the result. A punishment is a scholarship. In each case there is a name and there is what actually occurred, and in each case the name is what gets archived, quoted, and believed.

I chose Judith Butler as the material for this exercise partly because the material is rich and partly for a reason I can now state more plainly than I could at the beginning. Butler is the author of the most conspicuous declaration in modern thought about human beings, made in public, in the strongest possible form, on the first page of a book that a hundred thousand people bought. The declaration is that the position behind the deed is empty. Nobody had to dig for it. Nobody had to reconstruct a hidden assumption from a methods section. It was announced.

And for thirty-five years it has been treated as evasion. This is the finding I find hardest to get over, and it is not really a finding about Butler. It is a finding about how we receive claims. A thinker who states a strong commitment openly is read as having something to hide, while a thinker who leaves the same commitment unstated is read as merely reporting the facts. The second is the safer career and the first is the better epistemic conduct, and our habits of reading reward them in exactly the wrong order. I do not know how to fix that, and I notice that the fixing would have to be done by readers rather than by writers.

There is a second reason the material suited the exercise, which is that the subject is one where nearly everybody has already decided. That makes it the hardest possible test of a method and therefore the only interesting one. A procedure that produces careful results about the kilogram and collapses into partisanship when applied to something that matters is not a procedure. If the taxonomy of measured, declared, and derived is worth anything, it has to work where people are shouting, and it has to produce findings that its user did not want. I have tried to record the ones I did not want as prominently as the others: the statistic I had to withdraw, the critic whose case survived, the figure on my own side that had been inflated, the possibility that my whole verdict is an artifact of which books I chose to read closely.

What would I like a reader to carry away from this, assuming they retain nothing else?

The first thing is a question, and it is short enough to use in conversation: which part of that was measured? Not as an accusation and not as a debating move, since the answer is frequently that a substantial part was chosen and that this is entirely legitimate. As an actual request for information. In my experience it is received well by people who have thought about their claims and badly by people who have not, and the manner of the reception is itself informative.

The second is a distinction that I have come to think is the most useful single item in the kit: the difference between a vacancy that is expected to be filled and one that is declared permanently empty. Once you can see the difference, a great many disputes reorganize themselves. The person who says we do not yet know and the person who says there is nothing to know are in completely different positions, take completely different risks, and owe completely different accounts of themselves. Most public argument treats them as interchangeable, and most public argument is correspondingly useless.

The third is the rule about symmetry, which is the one that costs. An argument found to be defective cannot be used afterward, even when it points where you would like to go. This is easy to endorse and genuinely painful to practice, because the moment always comes when the defective argument is the strongest one available and the audience will not notice. I have failed at this before and will again. What I can report is that the occasions on which I managed it were the occasions on which I learned something, and that a conclusion which survives the loss of its best argument is worth more afterward than it was before.

The fourth is about anchoring, and it is the one that has changed how I read the news. When a claim in a contested area seems overwhelming, find the anchor and move it. If the claim survives, you are looking at something robust and should update accordingly. If it inverts, you have learned that the argument was never about the world, and you can stop attending to the evidence being thrown back and forth, because the evidence was never what divided anybody.

There is a fifth thing, and it is less a technique than a temperament. Very nearly everyone in these disputes is arguing in good faith from a position they arrived at honestly. The assumption of stupidity or malice is almost always wrong, it is always available, and it is the enemy of finding out what is actually going on. The interesting explanation of a persistent disagreement between competent people is never that one side is lying. It is that they are running different projections, or pinning at different ends, or asking a question that the other side is not asking, and each of these is discoverable by anyone willing to spend the afternoon.

I should be honest about the limits of all this, since a conclusion is exactly where a book quietly inflates its claims and hopes the reader is tired. An audit tells you what an argument entitles you to. It does not tell you what is true. It does not settle whether there is anything behind the deed, whether the categories we use to sort human beings are the right ones, or what should be done about any of the arrangements described in these chapters. Those questions remain exactly as open as they were, and a reader who came looking for an answer to them has been reading the wrong book with admirable persistence.

What the exercise does deliver is smaller and, I would argue, more durable than an answer. It produces a division of the material into what was seen and what was chosen, and it makes the choices visible so that they can be argued about as choices. Nearly every intractable public dispute I have examined turns out, on inspection, to be a disagreement about a declared component conducted as though it were a disagreement about the facts. Naming the component does not resolve the disagreement. It relocates it to where it actually lives, and disagreements that have been correctly located are much less bitter than disagreements that have not, because the participants can finally see what they are disagreeing about.

Since I have spent a book demanding that other people say what would change their minds, it would be indecent to end without doing it myself. The entry on foreclosure would move if somebody devised a study establishing that some element of gendered self-understanding is available independently of the categories a culture supplies; I described in Chapter Eight roughly what such a study would have to look like, and I would revise without complaint. The entry on the missing research program would move the moment anybody publishes measured maintenance costs, enforcement thresholds, or drift rates under this theory; that entry is the easiest to overturn and has stood for thirty-five years. The entry on attribution would move if a citation and policy-descent analysis of the kind described in Chapter Twenty-One turned up chains I have assumed do not exist. And the whole ledger would need rebuilding if the selection bias I confessed to in Chapter Thirty-Two turned out to be larger than I estimated, which is the item I would bet on if I had to bet on one.

There is also something owed to the reader about their own position, since it would be strange to spend this long on other people’s declarations and leave the reader’s untouched. You brought something to this book. Most people who pick up a volume on this subject have a settled view, arrived at honestly, generally before encountering any of the arguments. That view is a declared component of your reading, it has been shaping which chapters felt persuasive and which felt strained, and it is not visible to you while it operates, any more than mine is to me. I have no way of examining it and no business trying. What I can say is that the chapters you found most obviously correct are the ones where you should look hardest, because agreement is a much less reliable signal than it feels like, and the sensation of a point landing is produced as often by fit as by force.

There is a version of this book that would end by telling you what to think about gender. I have deliberately not written it, and not out of cowardice, though the reader is entitled to suspect that. It is because the honest result of the exercise is that the important questions in this area are not settled by any of the arguments currently being made, on any side, and a person who claims otherwise is either not looking or is looking at something else. That is a genuinely unsatisfying place to arrive after thirty-three chapters. It is also, so far as I can tell, where the evidence stops.

What I will say is this. The proposition that there is nobody behind the gesture is the most interesting claim made about human beings in the last half century, it has survived every attack made on it in that period, and it has not been established. Those three statements are compatible. Holding all three at once is uncomfortable and is, I think, the correct position, and the discomfort is not a sign that something has gone wrong. It is what it feels like to have finished the work available and not yet have the answer.

One more thing about the ethics of reading, which is what this whole enterprise has really been about. It is possible to read somebody carefully without wanting them to be right, and it is possible to want them to be right and still read them carefully, and both of these are difficult in the same way: they require holding the assessment separate from the preference for long enough to finish. The failure mode is not usually dishonesty. It is speed. One decides early, and everything afterward is processed as confirmation, and the processing is invisible and feels exactly like reading. The only remedy I know is the tedious one of writing down what you thought before you started, so that afterward there is a record to be embarrassed by.

The talkative child in Cleveland was given three questions to think about, and one of them was whether a body of thought can be held responsible for what is done with it. That question has followed the person who asked it for fifty-five years, mostly in the form of accusations from people who had not read the work. The honest answer, arrived at by a long route, is that a body of thought is responsible for what follows from it by steps that can be traced, and for nothing else — and that tracing the steps is work, which almost nobody does, in either direction, about anybody.

So the last recommendation of this book is the least glamorous one available. Trace the steps. Ask where the chain was pinned. Ask what was chosen. Ask what would change your mind and be suspicious if nothing would. Do this to the people you disagree with, and then, on a day when you have the stomach for it, to yourself. It will not make you popular and it will not make you certain. It will occasionally make you right about something, which is rarer than either, and which is the most any procedure has ever been able to promise.

Case Studies

Case Study One — The Night Twenty-Nine Million Americans Became Overweight

In the 1830s a Belgian astronomer named Adolphe Quetelet, who had turned his attention from stars to societies, sought a simple way to describe the distribution of human body sizes in a population. He settled on weight divided by the square of height, a ratio that varied less with stature than weight alone. He was explicit about the purpose. The index described populations. It was a tool of what he called social physics, designed to characterize the average man of a nation, and Quetelet never suggested that it diagnosed anything about an individual.

For a century it sat quietly in the demographic literature. Then the American life insurance industry, which needed an administrable rule for pricing policies, began correlating body measurements with mortality, and by the 1970s a physiologist had rechristened Quetelet’s ratio the body mass index and recommended it for population studies of obesity, again with explicit caveats about individual application. The caveats did not travel. The index was cheap, required only a scale and a tape measure, and produced a single number that could be entered on a form, which in the economy of clinical practice is worth a great deal more than accuracy.

The challenge was the one that faces every classification with a continuous underlying variable: where to draw the line. Body mass is distributed continuously across a population with no natural gap, no bimodality, and no point at which health outcomes change abruptly. Any threshold is therefore a decision, and the decision has to be made by somebody, on grounds that combine epidemiology, administrative convenience, and judgment about the costs of over- and under-identification.

In June 1998 the American National Institutes of Health adopted new guidelines that lowered the threshold for overweight from a body mass index of about twenty-eight for men and twenty-seven for women to a uniform twenty-five. The stated rationale was harmonization with the standards used by the World Health Organization and consistency with mortality data. Nothing about anybody’s body changed. No one gained a gram. On the day the guidelines took effect, somewhere between twenty-nine and thirty-five million Americans who had gone to bed within the normal range woke up classified as overweight.

The measurable results were substantial and ran in several directions. Prevalence statistics jumped discontinuously, creating an artificial appearance of a sudden worsening in national health that was in fact a change of definition, and subsequent time series comparing pre- and post-1998 figures inherited the discontinuity unless carefully adjusted, which they frequently were not. Insurance premiums, employment medical assessments, military and police entry standards, and clinical treatment thresholds all keyed to the index and all moved with it. Pharmaceutical eligibility criteria expanded. Public health campaigns were designed around a target population that had been enlarged by an administrative act.

The index’s specific defects as an individual measure were well documented before 1998 and have been documented more thoroughly since. It cannot distinguish muscle from fat, which is why a substantial fraction of professional athletes classify as overweight or obese. It does not account for the distribution of fat, though the distribution is a far better predictor of metabolic risk than the total. It performs differently across populations of different builds, so that thresholds derived largely from European-descended populations misclassify risk in South Asian populations in one direction and in some East Asian populations in another, a problem serious enough that several countries have adopted their own cut-offs. It says nothing about fitness, which independently predicts mortality. And its relationship with mortality is not linear but curved, with elevated risk at both ends, a fact that the categorical scheme obscures by collapsing the low end into a single underweight box.

None of this constitutes a case for abandoning the index, and the argument here is not that it should be. For population surveillance it remains cheap, comparable across decades and countries, and adequate for the purposes it was designed for. The failure is a category failure: a population instrument deployed as an individual diagnostic, with a threshold that was chosen rather than discovered, and reported thereafter as though it were a measurement of a person’s condition.

The contemporary relevance is direct and growing. Body composition can now be measured with reasonable accuracy at moderate cost through several technologies that were unavailable in 1998, and continuous monitoring of the metabolic markers that actually predict outcomes is becoming routine. The index survives not because it is the best available instrument but because it is embedded: in insurance actuarial tables, in clinical guidelines, in electronic health record fields, in eligibility criteria for procedures and drugs, and in three decades of research literature whose comparability would be lost if the definition changed. This is the ordinary fate of a declared component that has been in circulation long enough. The cost of replacing it exceeds the cost of its errors, so the errors are retained and are gradually reclassified as facts.

Several medical bodies have in recent years recommended de-emphasizing the index in individual clinical decisions in favor of direct measures, and a growing number of guidelines now instruct clinicians to treat it as a screening trigger rather than a diagnosis. Whether that instruction survives its passage down the chain from guideline to consultation room is the open question, and it is the same question raised in the chapters above about every attempt to attach a caveat to a number. The number travels. The caveat does not.

The lesson generalizes past medicine. Any threshold applied to a continuous distribution creates a population by declaration, and the population so created will thereafter be counted, studied, funded, treated, and discussed as though it had been found. The first question to ask of any such group is when the line was drawn, by whom, and what happened to the size of the group when it moved.

Case Study Two — The Guideline That Made Half a Country Hypertensive

In November 2017 the American College of Cardiology and the American Heart Association jointly issued a revised guideline on the detection and management of high blood pressure. The previous standard had defined hypertension as a reading of one hundred forty over ninety or above. The new standard set it at one hundred thirty over eighty. The change was announced at a scientific session, published simultaneously in two journals, and reported in the general press within hours.

The background is a genuine scientific advance and it should be stated fairly, because this is not a case of anybody inventing a problem. Blood pressure is, like body mass, a continuous variable, and its relationship to cardiovascular risk is continuous too: risk rises steadily across the range with no natural threshold at which it begins. A large randomized trial published in 2015 had compared intensive blood pressure control against the standard target in a high-risk population and found a significant reduction in cardiovascular events and in death. The trial was well conducted and its result was real. The guideline committee, reasoning from it and from a body of observational evidence, concluded that the old threshold was set too high and that treating earlier would prevent events.

The challenge was that a threshold does several jobs at once and they pull in different directions. It identifies people who should be offered treatment. It also assigns a diagnosis, with everything that follows: a label recorded in a permanent file, implications for insurance and employment in some jurisdictions, the psychological weight of being told one has a chronic disease, and entry into a system of monitoring and medication with its own costs and side effects. A threshold optimized for the first job is not necessarily right for the others, and the guideline had to choose one number to do all of them.

The immediate measurable result was arithmetic and dramatic. The proportion of American adults classified as hypertensive rose from roughly a third to nearly half, an increase of something like thirty-one million people. Among adults under forty-five the proportion of men classified with the condition approximately doubled. Nobody’s arteries changed. The guideline was explicit that most of the newly classified were not candidates for medication and should be advised on diet, exercise, and salt, but the classification was applied to all of them, and classification is what enters the record.

The reception among clinicians was not uniform, and the disagreement is instructive because it was not about the trial data. Several major bodies, including the American Academy of Family Physicians and guideline groups in Europe, declined to adopt the lower threshold, citing concerns about the generalizability of the trial population, about the specific method of blood pressure measurement used in the trial, which produced systematically lower readings than ordinary office measurement, and about the consequences of applying an intensive target to lower-risk people in whom the absolute benefit is small and the harms of treatment are not.

That last point deserves emphasis because it is the pricing step from the chapters above, applied properly. The relative reduction in risk from treating was similar across the population. The absolute reduction was not, because it depends on the baseline risk, which is much lower in a healthy forty-year-old than in a seventy-five-year-old with existing disease. Meanwhile the harms of treatment — hypotension, falls, kidney injury, the burden and cost of daily medication, and the cascade of further testing that a diagnosis triggers — do not scale down with baseline risk in the same way. So the same guideline can be simultaneously right for one group and wrong for another, and a single threshold cannot express that.

The subsequent literature has largely borne out both sides. Treating the higher-risk newly classified population does appear to prevent events. The expansion of the diagnostic category has also produced documented increases in prescribing, in follow-up visits, and in the number of people carrying a chronic disease label, with the associated effects on insurance underwriting in markets where that is permitted, and with no measured benefit in the lowest-risk portion of the new group.

The contemporary link runs through a development the guideline committee could not have anticipated. Blood pressure is now measured continuously by consumer devices worn on the wrist, and the readings vary enormously across a day: with posture, activity, stress, sleep, and the mere fact of measurement. Against that background the notion of a person’s blood pressure as a single number to be compared with a threshold begins to look like an artifact of the instrument that was available in a clinic. The distribution over a day is a far richer object than any single reading, and the threshold framework has no way to use it.

What the case demonstrates is that a diagnostic category is not a discovery about bodies but a decision about where to intervene, made by a committee, on grounds that mix evidence with judgment about costs. The evidence part is measured. The location part is declared. Both are legitimate and only one of them is usually reported. When a headline states that half of a country has a disease, the sentence contains a measurement and a decision in unstated proportions, and the proportion has changed several times within living memory without anybody’s arteries taking any notice.

Case Study Three — The Sampling Protocol That Made the Water Safe

In April 2014 the city of Flint, Michigan, under emergency financial management, switched its municipal water supply from the Detroit system to the Flint River as a cost-saving measure while a new pipeline was built. The river water was more corrosive than the water it replaced. The city did not apply orthophosphate corrosion control, a standard and inexpensive treatment that coats the interior of pipes and prevents metals from leaching into the supply. Flint had a large stock of lead service lines, as most older American cities do.

Within months residents reported discolored, foul-smelling water and a range of symptoms. The state environmental agency conducted the sampling required under federal regulation and reported that the system was in compliance with the lead action level. Officials repeated this reassurance publicly for over a year while residents continued to complain, and the complaints were treated as a matter of aesthetics and public confidence rather than of contamination.

The core challenge, and the reason this case belongs in a book about declared components, is that the compliance testing was not a straightforward measurement of the lead in the water. It was the output of a procedure with a great many choices embedded in it, and the choices determined the result.

Consider what the procedure required. Which homes are sampled: the regulation calls for homes at high risk, meaning those served by lead lines, and identifying them requires records that many cities, Flint included, did not reliably possess. How the sample is drawn: instructions distributed to residents in Flint asked them to run the tap before collecting, a practice known as pre-flushing, which clears the standing water where lead concentrates and systematically lowers the reading. How fast the water runs: a slow flow disturbs less particulate lead from the pipe wall than a fast one, and the guidance favored the slower. How the results are aggregated: compliance depends on the ninetieth percentile of samples, so removing a small number of high readings from the pool can move a system from violation to compliance. And in the event, some high samples were removed from the reported set on grounds that were disputed afterward.

None of these steps involves falsifying a number. Every reported value was, so far as anyone has established, the true lead concentration of the sample that was actually collected. The declaration lay entirely in which samples were collected, from where, in what manner, and which were included in the calculation. The measurement was honest and the procedure that generated it produced a predictable answer.

The correction came from outside the system. A civil engineering team from Virginia Tech, led by Marc Edwards, distributed sampling kits to residents with instructions designed to capture what people were actually drinking rather than what the compliance protocol was designed to capture: no pre-flushing, sequential samples through the standing volume of the pipe. The results were not marginal. A substantial fraction of homes exceeded the federal action level, some by more than an order of magnitude, and one household recorded concentrations in the range classified as hazardous waste.

In parallel a pediatrician at a Flint hospital, Mona Hanna-Attisha, took the question to a different instrument entirely. Rather than measuring water, she examined blood lead levels in children from hospital records, comparing the period before the switch with the period after, and within the city comparing areas served by different parts of the distribution system. The proportion of children with elevated blood lead had approximately doubled citywide and risen further in the most affected zones. This finding could not be addressed by any argument about sampling protocols, because it did not sample water at all. It sampled children.

The state initially disputed both sets of findings. Within weeks the dispute collapsed, because the two independent lines of evidence agreed with each other and with the residents’ complaints, and because the blood data were not vulnerable to any of the procedural objections available against the water data. The city returned to the Detroit supply in October 2015, a state of emergency was declared, and a program of service line replacement began. Criminal charges were brought against officials, and the litigation and its consequences continued for years.

The measurable results, in the sense a case study requires, are stark. Thousands of children were exposed during the period in which the system was officially in compliance. The replacement program cost hundreds of millions of dollars, against a corrosion control treatment estimated at a small fraction of that annually. Civil settlements exceeded six hundred million dollars. And the federal lead regulation itself was substantially revised, with the revisions addressing precisely the procedural choices described above.

The contemporary link is not confined to water. Compliance measurement of every kind — emissions, workplace exposure, food safety, financial capital adequacy — consists of a procedure whose parameters are set in advance, and any such procedure can be optimized for the result it produces rather than for the phenomenon it purports to track. The optimization does not require anybody to lie and is frequently performed by people who believe they are following the rules, because they are. The diagnostic question for any compliance regime is not whether the numbers are honest. It is who chose the protocol, whether they had an interest in the outcome, and what an independently designed protocol would report.

Case Study Four — The Flights That Were Left Off the Chart

On the evening of 27 January 1986, engineers at Morton Thiokol, the contractor responsible for the solid rocket boosters of the Space Shuttle, held a teleconference with NASA managers to recommend against launching the following morning. The overnight temperature at the Florida launch site was forecast below freezing. The engineers were concerned about the resilience of the rubber O-rings that sealed the joints between booster segments, which they believed stiffened in cold and might not seat properly under the pressures of ignition.

The engineers had data. Previous flights had been examined after recovery, and several showed erosion or blow-by of hot gas past the primary seal. They assembled charts and presented them. By the end of the teleconference the recommendation had been reversed, the contractor’s management having been asked, in a phrase that entered the record of the subsequent investigation, to take off its engineering hat and put on its management hat. The launch proceeded at a temperature of about minus one degree Celsius, some twelve degrees below the coldest previous launch. The vehicle was destroyed seventy-three seconds after liftoff and all seven crew members died.

The presidential commission that followed identified the physical cause without much difficulty. What has made the case a permanent fixture in the literature on evidence and decision is the reconstruction of what the engineers actually showed and what it was capable of showing.

The charts presented that evening plotted the observed damage against flight, and they included the flights on which damage had occurred. This is an entirely natural thing to do. The subject under discussion was damage; the flights with damage are the ones with damage to discuss; the flights without damage appear to contain no information about it. Arranged this way the data are genuinely ambiguous. Damage appeared at various temperatures, including a warm launch, and no clear pattern emerges. A reasonable person looking at those charts could conclude that temperature was one factor among several and not obviously decisive.

Now plot every flight, including those with no damage, against launch temperature. The picture is transformed. The undamaged flights cluster at the warm end of the range; damage appears with increasing frequency and severity as temperature falls; and the proposed launch temperature lies far outside the range of any previous experience, in the direction the trend runs. The statistician Edward Tufte later reconstructed the graphic in this form and it has been reproduced in teaching ever since, because the contrast between the two presentations of identical underlying data is as clean as such contrasts ever get.

The declared component was the selection rule: which flights belong in the analysis. It was never stated, never debated, and almost certainly never consciously chosen. It followed from the natural way of framing the question, which was what causes damage rather than what predicts damage, and those two questions call for different datasets. The first invites you to examine the cases where the outcome occurred. The second requires the cases where it did not, because a predictor can only be identified by comparison.

This failure has a name in statistics and it is one of the oldest known: selection on the dependent variable. It is not a subtle error in the sense of being hard to state. It is subtle in the sense of being nearly invisible from inside the task, because the excluded cases feel irrelevant. Every discipline rediscovers it. Studies of successful companies that examine only successful companies cannot identify what causes success. Analyses of aircraft returning from combat that examine only the returning aircraft cannot locate the fatal vulnerabilities, a problem famously solved by Abraham Wald in the Second World War by reasoning about the planes that did not come back.

The measurable results of the Challenger case were institutional. The shuttle fleet was grounded for thirty-two months. The booster joint was redesigned. NASA’s decision procedures were restructured, and a formal channel was established through which engineering objections could be escalated without passing through the managers whose schedules they threatened. The commission’s report became a standard text in engineering ethics.

The relevance to present conditions is uncomfortable, because the specific failure has become easier rather than harder to commit. Modern organizations generate enormous quantities of incident data — near misses, alerts, adverse events, security intrusions — and the analytical apparatus is trained on the incidents, because that is where the interest is and that is what the reporting systems capture. The non-incidents are frequently not recorded at all, and where they are recorded they are not examined. This is precisely the shape of the chart shown on the night of 27 January 1986: a complete and accurate record of everything that went wrong, which cannot by itself tell you what makes things go wrong, because it contains no comparison.

The remedy is procedural and cheap and is still not standard. Before any analysis of a set of failures, ask what the corresponding successes look like and whether they are available. If they are not available, say so in the conclusion. And if the decision at hand concerns a proposed operating condition outside the range of anything previously attempted, note that no dataset can speak to it directly, whichever flights are plotted.

Case Study Five — The Unstated Unit

The Mars Climate Orbiter was launched in December 1998 to study the Martian atmosphere and to serve as a communications relay for a companion lander. It was one of two spacecraft in a program built around a philosophy of faster, better, cheaper missions, in which reduced budgets and compressed schedules were to be offset by simpler designs and tighter focus.

On 23 September 1999, after a nine-month cruise, the spacecraft fired its main engine to insert itself into orbit around Mars, passed behind the planet as expected, and did not reappear. It had entered the atmosphere at an altitude of approximately fifty-seven kilometers, well below the minimum survivable altitude of around eighty, and had either burned up or been thrown back into a solar orbit. The mission was lost.

The cause, established within weeks, was that a ground software file used to model the effect of the spacecraft’s small attitude-control thruster firings on its trajectory reported impulse in pound-force seconds, an imperial unit, while the navigation software that consumed the file expected newton-seconds, the metric equivalent. The conversion factor between them is about four and a half. The thruster firings, which occur many times over a long cruise to manage the accumulated momentum of the reaction wheels, therefore produced a cumulative trajectory error that was consistently underestimated by a factor of four and a half.

What makes the case worth the attention it has received is not the arithmetic, which any first-year student can perform, but the fact that the discrepancy was visible in the data for months. Navigators had noticed that the spacecraft’s actual trajectory kept diverging from the predicted one after each thruster event and had been applying manual corrections. The divergence was discussed. Concerns were raised about the projected closest approach distance in the final weeks. But the anomaly was small on any individual firing, the corrections worked, and the underlying discrepancy was attributed to modeling imprecision rather than to a systematic unit error.

The declared component here is the unit convention, and it has a specific structural property that makes it dangerous: it is carried in the interface between two components rather than inside either of them. Each piece of software was internally consistent and correct. Each team’s work was verifiable within its own scope. The error existed only in the relationship, which was documented in a specification that both parties believed they had complied with, and which no test exercised because the tests were designed by the same teams that had made the assumption.

The investigation board identified this and its recommendations reflect it. It found the failure to be one of process rather than of any individual’s competence, noting inadequate independent verification of the interface, insufficient staffing on navigation, communication gaps between the contractor and the navigation team, and a culture in which the raised concerns did not find a formal channel that would compel resolution rather than discussion.

The measurable results: a loss of about one hundred twenty-five million dollars for the orbiter, a companion lander lost separately three months later for unrelated reasons, and a substantial reassessment of the program philosophy under which both had been built. Interface control documentation and independent verification practices across the agency were tightened. Metric units were mandated in the relevant systems.

The general lesson is about where declared components hide. They are rarely at the center of anybody’s work, because the center is where the attention is. They live at boundaries: between two teams, two contractors, two software systems, two disciplines that must exchange a quantity. Everybody assumes the convention is defined and everybody assumes the other party knows it, and the assumption is invisible precisely because it is shared.

The contemporary relevance requires no stretching. Modern systems are assembled from components written by parties who never communicate, at a scale that would have been unthinkable in 1999: software libraries pulled from public repositories, data feeds ingested from external providers, machine learning models trained by one organization and deployed by another for purposes the trainers never contemplated. Each interface carries assumptions about units, ranges, encodings, time zones, null handling, and the meaning of the values passing through it, and the overwhelming majority of these assumptions are documented nowhere, tested by nobody, and correct only by convention.

When such systems fail the investigation frequently produces a finding structurally identical to the one in 1999: every component behaved as designed, nobody made an error within their own scope, and the fault existed only in a shared assumption at a boundary that had no owner. The recurring recommendation is also the same, and it is worth stating in its general form. For any quantity crossing a boundary between systems maintained by different parties, the unit and the convention should travel with the value rather than being recorded in a document, and the receiving system should verify rather than assume. This is more expensive than the alternative. It is less expensive than the alternative’s failures.

Case Study Six — The Test Instrument That Was Trusted

The Hubble Space Telescope was deployed from a shuttle in April 1990 after two decades of development and at a cost of more than a billion and a half dollars. Within weeks of first light it was clear that something was badly wrong. Stars that should have appeared as sharp points were surrounded by haloes. The telescope was capable of some useful work, but the exquisite resolution that justified placing a large optical instrument above the atmosphere was not being delivered.

Analysis established the fault as spherical aberration in the primary mirror: the outer edge had been ground too flat, by about two micrometers, which is roughly a fiftieth of the width of a human hair and, for an optical surface of that specification, an enormous error. The mirror was, in a sense, superb. It had been figured with extraordinary precision to a shape that was slightly the wrong shape, which is worse than being figured imprecisely, because the error was consistent and therefore invisible to the tests that checked consistency.

The tracing of the cause is where the case becomes a study in declared components. Grinding a mirror of that size requires a device to tell the optician how far the surface deviates from the intended figure. The device used was a reflective null corrector, a precision instrument that produces a reference wavefront against which the mirror is compared. Its own alignment depends on the exact spacing of its internal elements, which was set using a rod of known length with a reflective cap.

During assembly, a technician’s measurement of that spacing was made against the wrong surface: the cap had a small area of non-reflective paint that had chipped, and the laser used for the setup reflected from the metal beneath rather than from the intended surface, displacing the element by about one and a third millimeters. That displacement propagated into the reference wavefront, which was thereafter reporting the wrong shape as correct. The mirror was ground faithfully to the specification the instrument provided.

Two other instruments were used in the course of the work, both simpler and less precise, and both indicated that the mirror had spherical aberration. Their reports were discounted on the grounds that they were the less accurate devices and that the reflective null corrector, being the primary metrology tool, was authoritative. This is the decisive step and it is entirely comprehensible. When a high-precision instrument and a lower-precision instrument disagree, the ordinary and usually correct inference is that the lower-precision instrument is wrong.

The declared component was the authority of the primary instrument: the assumption, never tested during the program, that the null corrector was correct. Everything downstream inherited it. The mirror met specification, the specification was verified, and the verification used the instrument whose error was the entire problem. An independent end-to-end test of the assembled optical system had been considered and rejected earlier as too expensive and schedule-consuming, which is the decision that would have caught it.

The failure review board’s finding was explicit on this point. The technical cause was the misassembled corrector; the organizational cause was a verification structure with no independent path, in which the possibility that the reference instrument was in error had not been treated as a hypothesis requiring test. The disagreeing instruments were data about exactly that hypothesis and were interpreted as noise.

The measurable results are unusually happy for a case study of this kind, because the fault was correctable. The aberration was consistent and precisely characterized, which meant that corrective optics could be computed to cancel it. In December 1993 a servicing mission installed a corrective instrument package and replaced a camera with one containing internal correction. The telescope then performed to specification and continued to do so for three decades, producing some of the most productive observational science of the era. The repair cost several hundred million dollars and required a crewed mission of exceptional complexity.

The pattern generalizes to any system that is verified against a reference. The reference is a declared component. It is trusted by definition, which is what makes it a reference, and its errors are therefore invisible to every process that uses it. The only defense is independent verification by a method that does not share the reference, and such verification is always more expensive, always looks redundant in advance, and is the first thing removed when a schedule tightens.

The contemporary version is everywhere and is mostly software. Systems are validated against test suites, benchmarks, and reference datasets. A model is evaluated against a benchmark, and the benchmark defines what good performance means. If the benchmark is subtly misaligned with the real task, every model optimized against it will be subtly wrong in the same direction, and the misalignment will be undetectable from within the evaluation, because the evaluation is the thing that is wrong. Where the benchmark has additionally been used to train the systems it evaluates, the loop closes entirely. The disagreeing lower-precision instrument in that setting is the real-world complaint, the user who reports that the output is wrong in a way the metrics do not capture, and it is discounted for exactly the reason the two simpler correctors were discounted in the 1980s: it is the less rigorous measure, and the rigorous one says everything is fine.

Case Study Seven — The Correlation Nobody Measured

In 2000 an actuary named David Li, working at a large bank, published a paper proposing a method for modeling the joint default behavior of many credit instruments. The mathematical device was a Gaussian copula, a function that links individual probability distributions into a joint distribution using a single correlation structure. The problem it addressed was real and had blocked the market for years: it is straightforward to estimate the probability that one borrower defaults, and extremely difficult to estimate the probability that many default together, which is the quantity that determines the risk of a pooled security.

The innovation that made the method usable was the choice of input. Historical default correlations are almost impossible to estimate, because defaults are rare, clustered, and observable only over long periods during which the composition of the market changes. Li’s approach used the prices of credit default swaps — market instruments whose values move with perceived default risk — as a proxy from which correlation could be inferred. This substituted an abundant, current, and precise data source for a scarce and stale one, which is why it was adopted with enthusiasm.

Within a few years the method was in use across the structured credit market and had been incorporated into the models used by rating agencies to assign grades to collateralized debt obligations. Those grades determined which institutions could hold the securities, since pension funds, insurers, and banks operate under rules keyed to ratings. A tranche rated at the highest grade could be held by almost anyone and carried a capital charge accordingly. The rating was, in effect, the license.

The declared component sat at the point of substitution, and it was not hidden. Li himself, and many others, stated that swap prices reflect market sentiment rather than underlying default relationships, and that the available price history covered only a benign period of rising asset values. The assumption was that a correlation estimated from a few years of calm conditions would remain informative under other conditions. The assumption was documented. It was also, in practice, forgotten, because what circulated downstream was not the paper but the rating, and a rating is a letter with no attached description of the assumptions that produced it.

The mathematical property that made the assumption dangerous is specific and worth stating without formulae. In a pool of loans, the crucial question is whether defaults arrive independently or together. If independently, the pool is safe: a few losses are absorbed and the senior claims are untouched. If together, the pool offers no protection at all, because the diversification that justified the senior grade was diversification across borrowers who turn out to be exposed to the same thing. The entire value of the structure depends on a single quantity, and that quantity was estimated from a period in which the common thing had not happened.

The core challenge, in retrospect, is that no one owned the assumption. The quantitative analysts who built the models knew its status. The traders who used the models cared about prices. The rating agencies applied a methodology. The institutions buying the securities relied on the rating. The regulators relied on the same rating through the capital rules. At each transfer the assumption became less visible and the output more authoritative, which is the transmission pattern described at length in the chapters above, operating here with several trillion dollars attached.

The measurable results arrived between 2007 and 2009. Housing prices fell across the United States simultaneously rather than regionally, which is precisely the common factor the models had not seen in their estimation window. Defaults arrived together. Securities rated at the highest grade suffered losses that the models assigned probabilities close to zero, and the ratings were downgraded en masse — in some cases by ten or more notches at once, which is not a revision but an admission that the original number contained no information. Institutions holding them against thin capital failed or were rescued. The consequences for output and employment across the world economy ran for years.

The responses were regulatory and structural. Rating agencies were subjected to new oversight and disclosure requirements about methodology. Capital rules were revised to reduce mechanical reliance on external ratings. Stress testing regimes were introduced that require institutions to evaluate portfolios under specified adverse scenarios rather than under model-implied distributions, which is in effect a mandated sensitivity analysis: move the declared component and report what happens.

The lesson is not that the mathematics was wrong. The copula does what it claims; the failure was in the estimate fed into it, and in the fact that the estimate’s provenance did not survive transmission. Nor is the lesson that models should not be used, which is not a serious position. The lesson is about a specific and identifiable danger: a quantity that cannot be measured directly, replaced by a proxy that can be, with the substitution documented at the origin and invisible thereafter.

That configuration is now more common rather than less. Enormous numbers of consequential decisions rest on quantities inferred from proxies, because the proxy is available and the target is not: engagement as a proxy for value, test scores as a proxy for learning, activity metrics as a proxy for productivity, and a great many inferred scores that stand in for characteristics nobody can observe. In each case the substitution is defensible, is usually stated somewhere in a methodology document, and is invisible in the output. The question to ask of any such number is what it would take for the proxy to come apart from the thing it stands for, and whether the estimation period contained an episode of that kind. In the credit case it did not, and the answer to the question was available in 2003 to anyone who asked it.

Case Study Eight — A Threshold Chosen for Convenience

In 1925 the statistician Ronald Fisher published a handbook for working researchers which included tables of the values a test statistic must exceed for a result to be considered notable. He selected, for the tables, the value corresponding to a one-in-twenty probability of arising by chance, remarking that it was convenient to take this point as a limit in judging whether a deviation is significant. He said elsewhere that the threshold was a matter of the investigator’s judgment and that no fixed level was appropriate to all circumstances.

The convenience was real and specific to the technology of 1925. Computing a test statistic’s exact probability required either laborious hand calculation or a printed table, and a table cannot list every value. Fisher’s tables gave the critical values at a few conventional levels, and researchers compared their statistic against those. The threshold was an artifact of the format of a book.

Within two decades the convention had hardened into a criterion, and within four it was an institution. Journals accepted results that crossed the line and rejected those that did not. Grant panels evaluated proposals on the likelihood of producing them. Careers advanced on their accumulation. A number chosen because it fitted neatly into a printed table had become the boundary between a finding and a failure.

The core challenge is not the threshold itself, which is no worse than any other, but what a threshold does to the behavior of people whose livelihoods depend on crossing it. If a result just below the line is worthless and a result just above it is publishable, then every degree of freedom in the analysis — which observations to exclude, which variables to control for, when to stop collecting data, which of several outcome measures to report, whether to analyze subgroups — will be exercised, on average, in the direction that crosses it. This requires no dishonesty whatever. Each individual choice is defensible, is made for a stated reason, and would be made the same way by a competent colleague. It is the aggregate over thousands of researchers that produces the effect.

The consequences began to be quantified in the 2010s. A large collaborative project attempted to repeat one hundred published psychology experiments using the original materials and, where possible, with the original authors’ cooperation. Roughly a third to two-fifths produced results consistent with the originals, depending on the criterion used, and the average effect size in the replications was about half that reported in the published papers. Similar projects in cancer biology, economics, and other fields produced results in the same range.

The diagnosis assembled over the following years has several components and they interlock. Publication bias: studies that cross the threshold are submitted and published at higher rates, so the literature is a filtered sample of the research conducted. Analytic flexibility: the many defensible choices in any analysis, exercised toward the threshold. Underpowered designs: small samples produce noisy estimates, and a small study that crosses the threshold must have found a large effect, which is more likely to be noise than a true large effect if true effects are typically modest. And the threshold’s binary character, which converts a continuous measure of evidential strength into a verdict, discarding the difference between overwhelming and marginal.

The measurable results of the response have been substantial and are still accumulating. Preregistration, in which the analysis plan is deposited publicly before data are collected, removes analytic flexibility by fixing the choices in advance; registered reports, in which journals accept a study on the basis of its design before the results exist, remove publication bias by making acceptance independent of the outcome. Where these have been adopted, the proportion of published studies reporting positive findings has fallen sharply — in some registered-report samples from over ninety percent to around half — which is exactly what the diagnosis predicted and is the strongest available evidence that the diagnosis was right. Several journals have abandoned the threshold entirely; a statistical society issued a formal statement in 2016 cautioning against its use as a criterion; sample sizes in several fields have risen substantially.

The contemporary relevance extends well beyond academic publishing, because the mechanism is general. Wherever a continuous quantity is converted into a threshold, and consequences attach to crossing it, behavior will organize around the threshold and the underlying quantity will cease to be a reliable indicator of anything. School systems keyed to a passing score, hospitals keyed to a waiting-time target, police forces keyed to a clearance rate, and firms keyed to a quarterly figure all exhibit the identical pattern, and the pattern was described well enough in the 1970s to have acquired a name: a measure that becomes a target ceases to be a good measure.

The general remedy is known and is rarely applied because it costs something. Report the quantity rather than the verdict. Where a threshold is unavoidable for a decision, state that the decision required a threshold, say who chose it, and report how the conclusion would change if it moved. And treat any distribution of published results that shows a suspicious cluster just above the line as evidence about the reporting process rather than about the world, since a true effect has no reason to know where the line is.

Case Study Nine — The Sonata and the Statute

In October 1993 the journal Nature published a one-page report by three researchers at the University of California, Irvine. Thirty-six college students had performed a spatial reasoning task after listening to ten minutes of a Mozart sonata, after ten minutes of relaxation instructions, and after ten minutes of silence. Scores on the task were higher following the music. The improvement, converted to the scale of a standard intelligence test, corresponded to about eight or nine points, and the authors reported that it dissipated within ten to fifteen minutes.

The paper made modest claims. It concerned a specific spatial-temporal task, not general intelligence; the effect was temporary; the sample was small and consisted of undergraduates; and the authors proposed a speculative account involving cortical firing patterns. It contained no claim about children, none about lasting benefits, and none about development.

What happened next is one of the best-documented examples of transmission loss on record. The finding was reported in the general press as evidence that Mozart makes you smarter. Within two years there were commercial recordings marketed for infant development. Within five the effect had a name, an industry, and a place in popular understanding of child rearing.

In January 1998 the governor of Georgia proposed in his state of the state address that every newborn in Georgia be provided with a classical music recording, and requested funds for the purpose. The state budget included the item. Comparable initiatives followed elsewhere: another state legislature mandated daily classical music in state-funded childcare centers; hospitals distributed recordings to new parents; a substantial consumer market developed around products promising cognitive benefits to infants.

The declared components in this chain can be listed precisely, and none of them was asserted by the original paper. That an effect measured in adults applies to infants. That an effect on one narrow spatial task reflects general intelligence. That an effect lasting ten minutes produces lasting developmental change. That the specific composer matters, as opposed to any arousing or enjoyable stimulus. Each was supplied by an intermediary and each was necessary for the policy to make sense.

The scientific response was thorough and slow, in the usual proportion. Replication attempts produced inconsistent results. Meta-analyses published in 1999 and again in 2010, the latter covering several thousand participants, found that the effect on spatial tasks was small and that it was not specific to Mozart: comparable improvements followed any stimulus that raised arousal and improved mood, including other music, an engaging audiobook, or a cup of coffee. The favored explanation became arousal and mood rather than anything about musical structure, which accounts neatly for why the original comparison conditions — relaxation instructions and silence — were exactly the conditions least likely to produce arousal.

The measurable results of the episode run in two directions. On one side, public money and enormous quantities of private money were spent on an intervention with no demonstrated effect on the outcome parents were buying. One major producer of infant videos in this general market was eventually obliged to offer refunds following complaints to consumer regulators about claims of educational benefit. On the other side, the episode is now a standard teaching case in research methods, science communication, and evidence-based policy, and its notoriety has probably prevented several successors.

There is a further consequence that is less often noted and more damaging. Music education in schools was, during the same period, under budgetary pressure in many jurisdictions, and advocates seized on the cognitive-benefit argument because it spoke the language budget committees understood. When the effect deflated, the argument deflated with it, leaving the defense of music education weaker than before, because the sound reasons for teaching music — that it is valuable in itself, that it is a discipline, that children enjoy it — had been set aside in favor of an instrumental claim that failed. This is a recurring cost of borrowing an evidential warrant one does not need: when the warrant is withdrawn, it takes the case with it.

The contemporary parallels write themselves. Single studies with small samples continue to generate industries: learning styles, brain-training software, various neuro-branded interventions, and a steady supply of nutritional and behavioral claims. The structure is invariant. A modest finding, correctly reported, is amplified at each transmission by the addition of premises nobody states, until it arrives at a legislature or a marketing department in a form the original authors would not recognize and frequently disown in public, without effect.

The diagnostic for a reader is simple and effective. When encountering a claim of this shape, ask what the original study actually measured, in whom, for how long. The gap between that and the claim being made is the space where the declared components live, and in the well-known cases it is wide enough to see from across the room.

Case Study Ten — Ten Thousand Hours

In 1993 three psychologists published a study of violinists at a music academy in Berlin. The students were divided by their teachers into groups by attainment, and the researchers reconstructed each student’s history of deliberate practice — not playing generally, but focused, effortful work on specific weaknesses under guidance. The best group had accumulated, on average, something in the region of ten thousand hours by the age of twenty. The less accomplished groups had accumulated less.

The paper’s argument was about the nature of expert performance and it was aimed at a specific target: the prevailing view that outstanding attainment reflects innate talent, with practice merely polishing what is already there. The authors argued that the quantity and, crucially, the quality of practice accounted for far more of the variation than had been assumed, and that the deliberate character of the practice — its structure, its difficulty, its feedback, its focus on what one cannot yet do — was the operative variable.

In 2008 a popular book on success featured the study prominently, and from it emerged a proposition that entered general circulation with remarkable speed: that ten thousand hours of practice produce mastery in any domain. The number acquired the status of a rule. It appeared in management literature, in self-improvement books, in schools, in coaching, and in a great deal of casual conversation.

The lead author of the original study spent much of the following decade publicly objecting. His objections were specific and are worth listing, because each identifies a declared component added in transmission. The figure was an average for one group in one domain, not a threshold or requirement; the variation within the group was enormous, with some members reaching high attainment on considerably less and others accumulating more without reaching it. The study concerned deliberate practice with the specific properties described, not repetition, and the distinction is the entire content of the theory. Nothing in the work claimed that the number generalized across domains. And the original argument was comparative — practice explains more than had been supposed — not absolute, and did not claim that practice explains everything.

The subsequent evidence has settled roughly where one would expect. A large meta-analysis published in 2014 examined the relationship between accumulated deliberate practice and performance across many domains and found that it accounted for a substantial share of variance in some fields and a small share in others: around a quarter in music and games, considerably less in education and professional work, where the environments are less structured and the feedback less immediate. Practice matters a great deal. It does not account for everything, and the amount it accounts for varies by domain in ways that are themselves informative.

The measurable results of the popularized version are harder to quantify than in the other cases here, but some are visible. Youth sports in several countries moved sharply toward early specialization and high-volume training, a trend with documented associations with overuse injury and with dropout, and one that the underlying research does not support: the evidence on early specialization is mixed at best and points in several domains toward diversified early experience. Educational and corporate programs were built around hour counts. And a genuine and useful finding — that the structure of practice matters more than its quantity, and that most people practice in ways that produce little improvement because they repeat what they can already do — was displaced by a number that says nothing about structure at all.

The transmission mechanism deserves note because it differs from the previous case. Nothing here was fabricated and the popularizer did not misreport the study. What happened is that a distribution was replaced by its mean, and a mean was then read as a requirement. This is among the most common of all quantitative errors in public communication and it is almost never noticed, because a mean is a real number that was really calculated. The information destroyed is the spread, and the spread was the interesting part: the fact that attainment at a given practice level varied enormously is what tells you what else is going on.

The contemporary link is to the way this figure now functions in discussions of automation and skill. The claim that a domain requires ten thousand hours of human practice has been used both to argue that certain work is safe from automation and to argue that it is not, and in each case the number is doing rhetorical work it cannot support, since it was never a fact about domains in the first place. It was an average practice history for one cohort of violinists in one conservatory in the 1990s.

The general point is that a memorable number will always defeat a qualified finding, and that the defeat is not reversible by the finding’s author, however loudly and however often they object. This author objected for over a decade, in books, papers, and interviews, with the authority of having done the work. The number is still in circulation and the objections are not.

Case Study Eleven — Pricing a Life

Every wealthy country requires its regulatory agencies to evaluate proposed rules by comparing costs against benefits. Since a great many rules are intended to prevent deaths, the comparison requires that deaths prevented be expressed in the same units as compliance costs. Agencies therefore employ a figure known as the value of a statistical life, which in recent American practice has been in the region of ten to twelve million dollars per fatality avoided.

The number is widely misunderstood and the misunderstanding is worth clearing away first, because it is not what its critics usually think. It is not a valuation of any particular person’s life, and it is not a price at which anybody may be killed. It is derived from observed willingness to accept small changes in risk. If a population of a hundred thousand people each requires an additional hundred dollars in wages to accept an annual death risk of one in a hundred thousand, then in aggregate they have priced one expected death at ten million dollars. The figure is a statement about how people trade small risks against money, scaled up.

The derivation nonetheless carries a substantial cargo of declared components, and they determine the number to a degree that surprises people encountering the literature for the first time. The estimates come principally from labor market studies comparing wages across occupations with different fatality rates, which requires the assumption that workers know the risks they face, that they have meaningful alternatives, and that wage differences reflect risk rather than the many other things that differ between occupations. Estimates come also from contingent valuation surveys, which ask people directly and are subject to well-documented distortions. The published range across credible studies spans roughly an order of magnitude, from a few million to over twenty, and an agency must choose a point within it.

The core challenge is that these choices have enormous consequences and are made in technical documents that almost nobody reads. A rule whose costs are three billion dollars and which is expected to prevent three hundred deaths passes at a valuation of eleven million and fails at nine. The choice of figure therefore determines the regulatory outcome, and the choice is a judgment about which studies to weight, how to adjust for income growth over time, and whether to vary the figure by age, income, or circumstance.

That last question produced the most instructive episode. In 2003 an American agency, in analyzing an air quality rule, applied a reduced valuation to lives saved among people over seventy, on the reasoning that the remaining life expectancy was shorter and that survey evidence suggested older respondents’ stated valuations were lower. The adjustment was defensible within the framework and it was technically standard in some European practice. When it became publicly known it produced an immediate and overwhelming reaction, was characterized as a senior death discount, and was withdrawn.

What that episode revealed is that the apparently technical exercise rested on a commitment nobody had stated: whether a statistical life is to be valued uniformly or according to characteristics of the person. Both positions are coherent. Uniform valuation treats each avoided death as equivalent, which is egalitarian and ignores information about remaining life. Adjusted valuation attends to years of life saved, which is what health economists do routinely in a different context and which produces systematically different regulatory priorities. The framework does not choose between them; a political judgment does, and until 2003 the judgment had been made inside a spreadsheet.

The measurable results of this apparatus are large and mostly invisible. Regulatory analyses using these figures govern air quality standards, workplace safety rules, vehicle design requirements, food safety regimes, and pharmaceutical approval conditions across whole economies. Modest variations in the assumed value shift billions of dollars of compliance costs and, on the agencies’ own estimates, thousands of expected deaths. Different agencies within the same government have historically used different figures, producing the awkward implication that a life saved by one department is worth more than a life saved by another.

The contemporary relevance is that the number of domains requiring such valuations is expanding rapidly. Autonomous vehicle design involves explicit tradeoffs between classes of risk. Pandemic policy required, in every country, an implicit comparison between mortality and economic and social costs, and the comparison was made everywhere while being stated almost nowhere, which meant it could not be debated. Climate policy requires valuations extending over centuries, which adds a discount rate — itself a declared component with effects far larger than the life valuation — and the choice of discount rate has been the single most consequential parameter in the entire economics of climate change.

The lesson is not that such valuations should be avoided, which is impossible: any decision that trades resources against risk implies a valuation, and refusing to state it does not eliminate it but merely conceals it and makes it inconsistent across decisions. The lesson is that the figure is a declared component of extraordinary leverage, that reasonable people differ about it by a factor of several, and that it should be argued about in public by the people who will live under it rather than settled in an appendix.

Case Study Twelve — The Anomalous Reading and the Explanation That Fitted

On the evening of 20 April 2010 the drilling rig Deepwater Horizon was completing a well in the Gulf of Mexico. The well had been difficult and expensive and was substantially behind schedule. The operation underway was temporary abandonment: sealing the well so that the rig could move on and a production vessel could return later.

Before displacing the heavy drilling mud with lighter seawater, the crew conducted a negative pressure test. The principle is straightforward. Pressure in the well is reduced to simulate the condition after the rig departs, and the well is then observed. If the cement barrier at the bottom is sound, nothing flows and the pressure holds steady. If hydrocarbons are entering, pressure builds or fluid returns.

The test was performed and produced contradictory results. Pressure on the drill pipe rose to around fourteen hundred pounds per square inch, which is the signature of a failed barrier. Simultaneously the kill line, an alternative path into the same well, showed no flow and no pressure. Two instruments connected to one well were reporting different things.

The crew and supervisors discussed the discrepancy over a period of roughly an hour, which is worth emphasizing: this was not overlooked. An explanation was offered, attributed to a phenomenon sometimes called the bladder effect, according to which the weight of the mud above could transmit pressure to the drill pipe without any flow from the formation. On the strength of it the anomalous reading was set aside, the zero reading on the kill line was accepted as the true indication, the test was declared successful, and displacement proceeded. Hydrocarbons entered the well, reached the rig, and ignited. Eleven men died, the rig sank, and the well flowed into the Gulf for eighty-seven days.

The declared component was the interpretive rule: when two measurements of the same system disagree, which one is believed. The choice made was not random and not stupid. The kill line reading was consistent with the schedule, with the substantial expectation that the cement job had succeeded, and with everyone’s strong preference to be finished. The drill pipe reading was consistent with a catastrophic and expensive failure. Both readings were real; the interpretation selected one, and the interpretation invoked a phenomenon which subsequent investigation found to have no accepted basis in well engineering.

Multiple official investigations reached compatible conclusions. The technical failures were several: an inadequate cement design for a difficult formation, fewer centralizers than recommended, the omission of a cement bond log that would have evaluated the barrier, and a blowout preventer that failed to seal. But every report identified the negative pressure test interpretation as the decisive moment — the last point at which the outcome could have been prevented by anybody on the rig — and identified as its cause the absence of a written procedure specifying acceptance criteria and, critically, specifying what to do when readings conflict.

That absence is the general lesson. A test without a written acceptance criterion is not a test; it is an occasion for interpretation, and interpretation under schedule pressure runs in a predictable direction. The direction is not corruption. It is the ordinary operation of motivated reasoning in people who have been awake a long time, who have a plausible-sounding explanation available, and for whom one reading means going home and the other means weeks of remedial work.

The measurable results were on a scale that has few peers in industrial accident history. Eleven deaths; an estimated four million barrels released; extensive ecological damage across the Gulf; a moratorium on deepwater drilling; and financial consequences to the operator exceeding sixty billion dollars including settlements, penalties, and cleanup. The regulatory response separated the American offshore safety regulator from the agency that collected revenue from leasing, on the ground that combining promotion and policing in one body had produced predictable results, which is itself a structural finding of the same type.

Industry practice changed in a specific and instructive way. Negative pressure test procedures now typically specify in advance what constitutes a pass, what constitutes a fail, and — the provision that would have mattered — that any inconsistency between monitoring points constitutes a fail requiring investigation rather than an ambiguity to be resolved by discussion. The criterion is fixed before the reading is taken, which removes the opportunity for the reading to influence the criterion.

The contemporary application is broad and largely unrealized. Any monitoring system with redundant sensors will eventually produce disagreeing readings, and the rule for handling disagreement is a design decision that is frequently left to whoever is on duty. In aviation this has been formalized for decades. In medicine, industry, and increasingly in automated systems that fuse multiple data sources, it often has not. The question worth asking of any safety-critical process is not whether it monitors the right things, which everyone attends to, but what it is specified to do when two of its instruments contradict each other at three in the morning at the end of a long job.

Case Study Thirteen — The Fossil That Was Expected

In December 1912 the Geological Society of London heard a presentation on fragments of a skull and a jawbone recovered from a gravel pit near the village of Piltdown in Sussex. The remains were said to represent an early human ancestor, and they combined a cranium of essentially modern proportions with an ape-like jaw. The specimen was named Eoanthropus dawsoni after its finder, a local solicitor and amateur antiquarian named Charles Dawson.

The find was received with enthusiasm in Britain and with more scepticism abroad, and the pattern of reception is the first thing worth attending to. The specimen fitted a widely held expectation that the enlargement of the brain had preceded other human characteristics in evolution, so that the earliest ancestors should show a large braincase with a primitive face and jaw. It also fitted, less respectably, a national appetite: France and Germany had produced spectacular fossils and England had produced nothing comparable. Piltdown supplied an English ancestor of the first importance.

For four decades the specimen occupied a place in the textbooks, and its influence was not neutral. Genuine fossils recovered in the intervening years pointed in the opposite direction: the australopithecine discoveries from South Africa in the 1920s showed a small braincase with more human-like teeth and posture, indicating that upright walking and dental changes had preceded brain enlargement. These finds were received coolly, in part because they conflicted with the pattern Piltdown appeared to establish. A fabricated specimen was thus used to discount authentic ones for a generation.

The exposure came in 1953, and its instrument was a new measurement rather than a new argument. Fluorine absorption dating, which estimates how long a bone has lain in the ground from the fluorine it has taken up, was applied to the fragments and indicated that they were far younger than claimed and that the cranium and jaw were of different ages. Detailed examination followed. The jaw was that of an orangutan; the teeth had been filed to produce a wear pattern consistent with a human diet; the pieces had been stained with chemicals to match the gravel; the canine had been artificially abraded and painted. The forgery, once examined by people looking for one, was not even particularly skillful.

The declared component here is the expectation, and it operated in a way this book has described repeatedly: it determined what counted as a satisfactory specimen and therefore what level of scrutiny was applied. The features that later exposed the fraud were visible in 1912 to anyone who examined the specimen closely, and some critics did note the oddness of the combination and the convenient absence of precisely those parts of the jaw that would have settled the question. Those objections existed. They were outweighed by fit.

There is a second and more subtle mechanism. Access to the original material was restricted; most researchers worked from casts, which do not preserve the surface details that revealed the staining and filing. The controlling institution thereby limited, without intending to, the population of people capable of detecting the problem. This is a structural point rather than an accusation, and it recurs wherever primary material is scarce and access is mediated.

The measurable results of the episode are usually stated as a delay of several decades in the acceptance of the correct account of human evolution, and that is broadly right, though the counterfactual is not clean, since the African material would have faced resistance in any case. The identity of the forger has never been established with certainty despite a century of investigation, with suspects ranging from Dawson himself, who is now generally regarded as the most likely, to several better-known figures whose involvement the evidence does not support.

The productive consequence was methodological. The episode became the standard argument for independent physical dating of specimens, for the deposit of original material in accessible collections, and for the routine application of tests that do not depend on the interpretation of morphology. Every one of these is a defense against the same failure: a specimen that fits the expectation receives less scrutiny, so scrutiny must be made procedural rather than left to interest.

The contemporary relevance is not about fossils. It is about every field in which results that confirm a prevailing expectation are checked less thoroughly than results that contradict it, which is every field. The asymmetry is not a character flaw; it is rational allocation of limited attention, since surprising results are more likely to be wrong. The trouble is that the same reasoning ensures that a wrong result which is unsurprising can persist indefinitely, and that the people best placed to catch it have the least motive to look.

The defense is the one the episode produced: independent verification by a method that does not share the assumptions of the original, applied routinely rather than when suspicion arises. In a period when the volume of published results vastly exceeds anyone’s capacity to check them, and when the tools for producing convincing artifacts of every kind have become extraordinarily cheap and general, the argument for building verification into the process rather than relying on eventual scepticism is considerably stronger than it was in 1953.

Case Study Fourteen — The Finding Without a Mechanism

The Vienna General Hospital in the 1840s operated two maternity clinics. Admission alternated by day, which produced something close to a natural experiment, though nobody had designed it as one. The first clinic was staffed by medical students and physicians. The second was staffed by midwives. Maternal mortality from puerperal fever in the first ran at roughly ten percent in bad years; in the second it was around four.

The difference was public knowledge in Vienna. Women begged to be admitted on midwife days and some gave birth in the street rather than enter the first clinic, having calculated correctly that street births carried a lower risk than the physicians’ ward. This is worth pausing on: the people most affected had identified the pattern from experience and were acting on it, while the institution offered explanations that did not fit.

Ignaz Semmelweis, an assistant in the first clinic, worked through the candidate explanations systematically and eliminated them. Overcrowding was worse in the second clinic. Climate was identical. Religious practices, the position in which women gave birth, and the presence of a priest were tested and made no difference. The decisive clue came in 1847 when a colleague died after being cut with a scalpel during an autopsy and his post-mortem findings resembled those of women who had died of the fever.

Semmelweis inferred that something was being carried from the autopsy room, where physicians and students began their day, to the delivery room, where midwives never went. He instituted mandatory handwashing in a chlorinated lime solution, chosen because it eliminated the smell of the dissecting room, which was his available proxy for the presence of whatever the agent was. Mortality in the first clinic fell within months to the level of the second, and in some subsequent months to near zero.

He then encountered a wall, and the nature of the wall is the reason this case belongs here. He could not say what the agent was. Germ theory did not exist; Pasteur’s work was two decades away. What Semmelweis had was an intervention, a mechanism stated in vague terms about cadaverous particles, and a very large effect in the numbers. What he lacked was a theory that his profession would accept.

The reception was hostile and the reasons were of several kinds. The claim implied that physicians had been killing their patients, which is a proposition that no professional body has ever received calmly. The proposed mechanism was not respectable within the prevailing framework, in which disease arose from imbalances and miasmas rather than from transmitted particles. Semmelweis published late and argued badly, eventually resorting to open letters accusing his opponents of murder, which did not assist his case. And critically, the evidence was of a form the profession did not then recognize as evidence: a statistical comparison between wards, unaccompanied by a physiological account.

The declared component was the criterion of admissibility: what counts as a reason to change practice. The prevailing standard required an intelligible mechanism, and a mechanism was exactly what could not be supplied. The mortality data were not disputed; they were regarded as insufficient. That standard is not absurd, and something like it prevents the adoption of every spurious correlation that ever appears in a dataset. Applied here it cost an enormous number of lives, and the arithmetic of that cost is available, because the practice was adopted in some hospitals and not others and the differential persisted for decades.

Semmelweis was dismissed from his post, moved to Budapest where he achieved similar results with similar reception, deteriorated mentally in ways that remain debated, was committed to an asylum in 1865, and died there within two weeks, apparently from an infection contracted through injuries sustained at the institution. Widespread acceptance of antisepsis followed a different route entirely, through Pasteur’s demonstration of microbial causation and Lister’s surgical practice, and arrived roughly two decades after Semmelweis had demonstrated the effect.

The measurable results, once the practice was universal, are among the largest in the history of medicine, and hand hygiene remains the single most effective infection control measure available in hospitals today. Compliance, remarkably, is still an unsolved operational problem: observational studies routinely find rates well below fifty percent in settings where every practitioner knows the evidence, which suggests that the original resistance had components beyond the theoretical.

The contemporary application concerns the standing of effects without mechanisms, and it cuts in an uncomfortable direction. We are now surrounded by them. Machine learning systems produce predictions that outperform human judgment in some domains while offering no account of why, and the debate about whether such outputs may be acted upon reproduces the Vienna argument with the terms updated. The case does not resolve that debate. It establishes only that requiring a mechanism before acting is not a neutral or cost-free standard, that the cost falls on people who are not in the room, and that the profession applying the standard is rarely the one that pays it.

Case Study Fifteen — The Effect That Was Named Before It Was Checked

Between 1924 and 1932 the Western Electric Company conducted a series of studies at its Hawthorne Works outside Chicago, examining how various conditions affected worker productivity. The first and most cited concerned illumination: lighting levels in an assembly area were varied, and output was recorded.

The story that entered the textbooks is that output rose when the lights were brightened and rose again when they were dimmed, and that productivity therefore responded not to the physical conditions but to the workers’ awareness of being studied. The phenomenon was named the Hawthorne effect and became one of the most widely taught concepts in the social sciences, invoked as a standard caution in any field where the act of observation might alter the behavior observed.

The concept is genuinely important and something like it certainly exists. Its status as a finding is another matter, and the reconstruction of what the original studies actually showed is a case study in how a claim can circulate for eighty years without anybody returning to the data.

The illumination experiments were reported only in summary form and the original records were long believed lost. In 2009 two economists located the primary data from the illumination studies in archives and reanalyzed them. Their conclusion was that the classic pattern was not there. Output did rise at the start of experimental periods, but the rises coincided with the resumption of work after weekends and with seasonal patterns, and the effect attributed to observation was largely explained by these. What variation remained was small and did not support the story as told.

How did the account survive? Several mechanisms compound. The original reports were summaries, written by researchers with a thesis, and the underlying records were not readily available for checking. The phenomenon described was intuitively compelling and matched other experiences. It was useful: it supplied a name for a real methodological worry and thereby earned a place in every research methods curriculum. And once it was in the curriculum it was transmitted by teachers who had learned it from teachers, none of whom had cause to consult a set of 1920s factory records.

The declared component was the interpretation of the original summary as an established empirical result. Nobody decided to treat it that way; the classification happened by default, because a described study in a published report is ordinarily a finding, and the distinction between a reported result and a verified one is invisible in a citation.

The measurable results of the reanalysis have been modest, which is itself informative. The concept remains in wide use and in most textbooks, with an increasing number now adding a note about the reanalysis. The underlying methodological caution is unaffected: participants who know they are being studied do behave differently, and this has been demonstrated many times in properly designed work. What has changed is the status of the founding case, which turns out to demonstrate something considerably weaker than what carries its name.

There is a broader phenomenon here that historians of science have begun to document systematically. A substantial number of canonical demonstrations in the social sciences have turned out, on inspection of primary sources, to be weaker, differently structured, or in some cases substantially different from their textbook versions. This is not principally a story about fraud. It is a story about the economics of checking: verifying a famous claim is expensive, unrewarded, and confers no credit unless the claim collapses, so it is done rarely and usually by accident.

The contemporary link is to the vastly expanded citation of results that few citers have read. Reference management software makes it effortless to cite a paper on the strength of its abstract or of another paper’s description of it, and citation networks have been shown to propagate claims that the original sources do not support, with the chain of attribution traceable back to a single misreading several decades old. The mechanism is exactly the one described in the chapters above: a summary is easier to transmit than a source, and each transmission increases confidence while decreasing contact with the evidence.

The practical remedy is unglamorous and available to anyone. When a claim is important to an argument, obtain and read the original source rather than the description of it. This is now easier than at any time in history and is done less than it should be, and the surprises that reward the effort are frequent enough to make it a habit worth acquiring. Somewhere in the chain behind most confidently repeated facts there is a document that says something slightly different.

Case Study Sixteen — Deciding What a Recession Is

In the summer of 2022 the United States recorded two consecutive quarters of declining real gross domestic product. A vigorous public argument followed about whether the country was in a recession, conducted with an intensity suggesting that something substantive was at stake, and involving accusations that the definition was being changed for political convenience.

The background is that two definitions had been in parallel use for decades and had rarely disagreed. The first is a rule of thumb, popularized by a journalist in 1974: two consecutive quarters of falling output. Its virtues are that it is simple, mechanical, applies to any country with quarterly national accounts, and can be computed by anyone. The second is the determination of a committee at the National Bureau of Economic Research, a private research body, which has dated American business cycles since the 1920s. Its definition is a significant decline in economic activity spread across the economy and lasting more than a few months, assessed using several indicators including employment, real income, industrial production, and sales, with judgment applied to their weight.

The core challenge is that the two definitions are built for different jobs. The mechanical rule is a classifier: it gives a fast, reproducible answer that permits comparison across countries and periods. The committee definition is a description of an economic phenomenon: it aims to identify genuine downturns and to exclude technical contractions that do not involve broad economic distress. Neither is a measurement of a natural kind, because there is no natural kind. Economic activity varies continuously; a recession is a category imposed on that variation for particular purposes.

In the 2022 episode the two definitions disagreed because the components disagreed. Output fell in two quarters, driven substantially by inventory and trade movements. Employment rose strongly throughout, with several million jobs added over the same period, and real consumer spending grew. The mechanical rule attended only to output and returned one answer; the committee attended to a broader set and did not declare a recession.

What made the argument bitter was that both sides believed the other had changed the rules. One side pointed to the mechanical definition and to its long use in journalism and textbooks. The other pointed to the committee’s fifty-year practice and to the fact that the committee had never used the two-quarter rule and had on previous occasions declared recessions that the rule missed and declined to declare ones that the rule caught. Both were accurate.

The declared component is the choice of definition, and the episode illustrates a specific pathology: when two definitions have coincided for a long time, people forget that there are two. Users of the mechanical rule had experienced it for decades as a reliable indicator of what the committee would conclude, which it largely was, and had come to treat agreement as confirmation that the rule captured the phenomenon. When the definitions came apart, the disagreement presented itself as a dispute about facts.

The measurable consequences of the classification are not trivial, which is why the argument had energy behind it. Recession declarations affect consumer and business confidence, which are themselves economic variables. They trigger contractual and policy provisions in some jurisdictions. They influence electoral outcomes. And they shape the interpretation of every subsequent statistic, since data are read differently depending on whether the reader believes a downturn is underway.

The instructive response would have been to report both, with their bases, and to note that the divergence was itself the interesting information: an episode in which output fell while employment rose is unusual and tells you something about the structure of that period. Very little coverage did this, because a story about two definitions producing different answers for identifiable reasons is harder to write than a story about whether we are in a recession.

The contemporary link is that the same structure now afflicts almost every headline economic indicator. Inflation is measured by several indices that weight components differently and can diverge by a large margin. Unemployment has multiple official definitions. Poverty, productivity, and wage growth all admit several defensible constructions that move independently. In each case a public argument about the state of the economy is frequently a public argument about which index to use, conducted by parties who have selected, usually sincerely, the one that fits their prior view.

The diagnostic is the one this book has recommended throughout. When two competent parties disagree about a factual matter that ought to be settled by data, ask whether they are computing the same quantity. If they are not, the disagreement is about a definition, the definition was chosen for a purpose, and the productive question is which purpose should govern here, which is a question that no amount of additional data will answer.

Case Study Seventeen — The Categories That Create Populations

Every ten years the United States conducts a census, and every ten years it asks about race and ethnicity. The categories offered have changed in nearly every decade since the first enumeration in 1790. Terms have been added and removed; the boundaries have shifted; the rules for assigning people have moved between enumerator judgment, family report, and self-identification.

A brief inventory conveys the scale of the variation. Nineteenth-century censuses included categories based on fractional ancestry that were abandoned as unworkable. Various national-origin groups have appeared as separate races, been folded into other categories, and reappeared. Ethnic origin was separated from race in the 1970s and asked as a distinct question, so that a person could be recorded as belonging to that ethnicity within any racial category. Until the 2000 census, respondents were required to select one racial category; from that year they could select several, which immediately created a new population of several million people who had not existed in the previous statistics.

The core challenge is that these categories serve incompatible masters. Civil rights enforcement requires stable categories over time so that disparities can be tracked and legal thresholds applied. Public health requires categories that correlate with the exposures and conditions being studied. Political redistricting operates under legal requirements defined in terms of specific groups. Academic research wants categories that track the social processes under investigation. Respondents want categories they recognize as describing themselves. No single scheme satisfies all of these, and each revision that improves one degrades another.

The measurable consequences are large and mechanical. Federal funding formulas distributing hundreds of billions of dollars annually use census-derived population counts, including counts by category. Legislative districts are drawn using them. Enforcement of anti-discrimination law uses statistical comparisons that depend on them. Health research reporting disparities depends on them, and epidemiological findings expressed in these terms are only as stable as the categories.

The multiple-response change of 2000 shows the mechanism cleanly. It was a good change: it allowed people to describe themselves accurately, and the demand for it was substantial and long-standing. It also broke every time series that crossed it. A jurisdiction’s population in a given category could rise or fall depending on how multiple responses were allocated, and several allocation rules were in use simultaneously by different agencies for different purposes, so that the same person could be counted differently in a health statistic and in a voting rights calculation.

A revision announced in 2024 illustrates the same dynamics. It added a category for people of Middle Eastern or North African origin, who had previously been instructed to record themselves within an existing category that many did not regard as descriptive, and combined the separate race and ethnicity questions into one. The change was made after extensive testing showing that the new arrangement produced more accurate self-description and fewer non-responses. It will also, on implementation, appear to produce a sudden change in the size of several populations, none of which will have changed at all.

The declared component is the classification scheme itself, and this case is unusual in the frankness with which the responsible agencies acknowledge it. Census documentation states explicitly that the categories are social and political constructs rather than biological ones and that they reflect current usage and current administrative needs. That acknowledgment is on the record. It does not survive transmission: the resulting statistics are used, cited, and argued over as though they described natural divisions, and a research literature spanning decades reports findings by category with no indication that the categories moved.

The contemporary relevance extends well beyond any census. Every administrative system that classifies people creates the populations it counts, and the classifications are then used to allocate resources, target services, evaluate outcomes, and train predictive systems. Machine learning models fitted to administrative data inherit the classification scheme entirely, including its history, its compromises, and its discontinuities, and reproduce it in outputs that appear to be objective because they emerged from a computation.

The practical recommendation for anyone using such data is threefold and dull, which is the usual condition of good practice. Establish when the categories last changed and whether the analysis period spans a change. Determine which allocation rule was applied to ambiguous or multiple responses, because there is always one and it is rarely stated in the dataset. And when reporting a difference between groups, state that the groups are administrative constructions of a particular date, since the alternative is to imply a stability that the record does not contain.

None of this argues against collecting the data. Countries that decline to collect statistics by category, some of them on principled grounds, thereby make it impossible to detect or measure disparities in outcomes, which is a substantial cost paid mainly by the people the disparities fall on. The categories are declared, unstable, and necessary, and the honest practice is to carry all three of those facts together rather than choosing whichever two suit the argument at hand.

Case Study Eighteen — What Detected Means

Prostate-specific antigen is a protein produced by prostate tissue and measurable in blood. Its concentration rises in prostate cancer. It also rises with benign enlargement, with infection, with age, and after various ordinary activities. A blood test for it became available in the 1980s, was approved in the United States for monitoring known disease, and was approved in 1994 for screening men without symptoms.

Screening spread rapidly, encouraged by advocacy campaigns, professional bodies, and the intuitive force of the argument that finding cancer early must be better than finding it late. Detection rates rose immediately and substantially. Large numbers of men were diagnosed with prostate cancer who would not otherwise have been diagnosed, and were treated by surgery or radiation.

The core challenge, which took two decades and several very large randomized trials to characterize, is that prostate cancer is not one disease with one natural history. Autopsy studies of men who died of other causes find histological prostate cancer in a large proportion of older men — by some series, a majority of men over eighty. Most of these cancers would never have produced symptoms or shortened a life. A screening test that detects them therefore converts a population of healthy men into a population of cancer patients, and the conversion is real in every sense that matters administratively and personally while being, for many of them, medically empty.

This phenomenon is called overdiagnosis and it has a property that makes it exceptionally hard to see: it is invisible at the level of the individual case. A man diagnosed by screening and treated, who then lives twenty more years, appears to himself and to his physician to be a life saved. He may be. He may equally be a man who would have lived those twenty years untreated. Nothing about his case distinguishes the possibilities, and the distinction can only be made statistically, by comparing populations that were screened against populations that were not.

Those comparisons were eventually made. A large European trial found a reduction in prostate cancer mortality from screening, of a magnitude implying that roughly a thousand men would need to be screened and dozens diagnosed to prevent one death. A large American trial found no significant mortality benefit, though its interpretation was complicated by extensive screening in its control group. The harms were also quantified, and they are not trivial: biopsy complications, and among men treated, substantial rates of long-term incontinence and erectile dysfunction.

The declared component sits in the word detected and in what a detection is taken to mean. The test measures a protein concentration. Above a chosen threshold, further investigation follows. The biopsy identifies cells with particular characteristics. Those characteristics are then classified as cancer, which is a name for a category defined by appearance rather than by behavior, and the name carries an implication about the future that the tissue itself does not warrant. A great deal of the confusion in this area dissolves once one sees that the disputed question is not whether the cells are there but whether that arrangement of cells should be called by a name that implies a trajectory.

The response has been a slow and partial revision of exactly that. Screening recommendations moved from routine to individualized, with several major bodies now advising a discussion of benefits and harms rather than a default test. Active surveillance — monitoring low-risk disease without immediate treatment — has become standard rather than exceptional in many systems, and uptake has risen substantially over the past decade. Proposals have been made, seriously and repeatedly, to rename low-grade lesions so that the word cancer is not applied, on the explicit ground that the name drives treatment decisions independently of the biology. Some of these renamings have been implemented in other organs.

The measurable results of the whole episode include a large number of men treated for disease that would not have harmed them, with the associated permanent side effects, and an unknown but nonzero number of deaths prevented. Both are real. The ratio between them depends on the threshold, the biopsy criteria, the treatment threshold, and the classification scheme, every one of which is a declared component, and moving any of them moves the ratio.

The contemporary relevance is expanding fast, because detection technology is improving much faster than the understanding of what detections mean. Blood tests capable of detecting fragments of tumor DNA across many cancer types are entering use. Imaging finds incidental abnormalities at rates that increase with resolution. Each advance detects more, earlier, and each raises the same question in a new setting: of the things now found, which would ever have mattered? The question is answerable only by long follow-up of untreated cases, which is expensive, slow, and ethically fraught, and which is therefore not done at anything like the rate the technology is advancing.

The general lesson is one of the least intuitive in medicine and it should be stated plainly. Finding a disease earlier is not the same as preventing harm from it, and the two come apart precisely when the disease has a variable natural history. Any screening programme’s benefit is a quantity to be measured, not a consequence of its logic, and the quantity depends on decisions about naming and thresholds that are made by committees and reported as facts about bodies.

Case Study Nineteen — The Room Where the Air Was Clean

In the late 1940s a graduate student at the University of Chicago named Clair Patterson was given a project: determine the age of the Earth by measuring lead isotope ratios in ancient minerals. Lead is the end product of uranium decay, the decay rates are known, and the ratio of isotopes in a sample therefore records the time elapsed. The method was sound and the problem should have been tractable.

It was not. Patterson’s measurements were erratic and irreproducible in ways that no experimental care seemed to fix. Over several years he established the reason, and the reason turned out to be more consequential than the original question: his samples were being contaminated with lead from the laboratory environment, from the reagents, from the glassware, from the dust, and from the air. The contamination was not a trace. It swamped the signal.

His response was to build the first ultra-clean laboratory. Everything was rebuilt or purified: reagents distilled repeatedly, surfaces stripped and coated, air filtered, procedures redesigned to eliminate every point at which ambient material could reach a sample. It was an enormous investment of effort in what looked, to observers, like preliminary work. With it he obtained in 1953 a figure for the age of the Earth of about four and a half billion years, which has stood ever since with only minor refinement.

The core challenge then became the interpretation of what he had learned along the way. Patterson had established that the ordinary environment of a mid-century American laboratory contained lead at concentrations that made precise measurement impossible. That is a statement about laboratories. It is also, on reflection, a statement about the environment those laboratories were in, and Patterson followed the reflection where it went.

He began measuring lead in ocean water, in sediments, in ice cores, and in ancient human remains, using the clean techniques he had developed. The results were consistent. Lead concentrations in surface ocean waters were far above deep water; concentrations in Greenland ice rose sharply from the early twentieth century; concentrations in modern human bone were vastly higher than in pre-industrial remains. The cause was identifiable: tetraethyl lead had been added to gasoline since the 1920s as an anti-knock agent, and was being emitted from every vehicle exhaust in the industrialized world.

The declared component in the prevailing position was the baseline. The industry’s position, supported by a body of published measurement, was that lead levels were natural and that human exposure was within the normal range. Those measurements had been taken by conventional laboratory methods — which is to say, in rooms like the ones Patterson had shown to be contaminated, using reagents like the ones he had shown to be contaminated. The baseline against which modern exposure was judged safe had been established using the same contaminated technique, so the comparison was between two contaminated numbers and could not detect the increase.

This is the structural heart of the case and it is worth stating in general form. When both the measurement and the reference are produced by the same defective procedure, the defect is invisible, because it cancels in the comparison. Only an independently constructed measurement can reveal it, and constructing one is expensive and looks unnecessary to everybody who trusts the existing procedure.

The conflict that followed was prolonged and unpleasant. Patterson lost funding from several sources, was excluded from a relevant national research committee while researchers with industry ties served on it, and had contracts declined. He continued, testified to Congress in 1966, and published. The evidence accumulated from other groups. Regulatory action began in the United States in the 1970s and leaded gasoline was phased out over the following two decades, then internationally, with the last country ceasing use in 2021.

The measurable results are among the largest public health effects attributable to a single line of research. Blood lead levels in American children fell by roughly four-fifths between the late 1970s and the early 1990s, tracking the phase-out closely. Because lead exposure in childhood produces measurable cognitive deficits, the population-level effect has been estimated in aggregate IQ points across cohorts, and a substantial literature has examined associations with other outcomes over the same period.

The contemporary lesson concerns what to do when a field’s entire evidence base rests on a shared method. Patterson’s insight was not that the existing numbers were wrong in some detail but that the procedure generating all of them contained a common error, and that no amount of internal consistency among those numbers could detect it. The relevant question for any established measurement regime is therefore not whether its results agree with each other, which they will, but whether anybody has measured the same thing by a method that shares none of its assumptions. That is expensive, it is nearly always resisted as redundant, and it is the only thing that catches this class of error.

Case Study Twenty — The Explanation That Was Taught for Fifty Years

On 7 November 1940 the Tacoma Narrows suspension bridge in Washington State, opened four months earlier, oscillated with increasing amplitude in a moderate wind and collapsed. The event was filmed, and the footage — a roadway twisting through extraordinary angles before failing — became one of the most reproduced pieces of engineering documentary in existence.

The bridge had been known from the beginning to move. It had acquired a nickname referring to its bouncing behavior; drivers reported watching cars ahead disappear and reappear over the vertical waves. On the day of the collapse the wind was around sixty-four kilometers per hour, well within the design envelope, and the motion changed character from vertical undulation to a twisting mode that grew until a support cable failed.

For the following half century the collapse was explained in physics textbooks as a case of resonance: the wind was said to have supplied periodic forcing at the bridge’s natural frequency, as a singer shatters a glass or a marching column excites a footbridge. The explanation is memorable, connects a dramatic event to a standard piece of the curriculum, and appeared in an enormous number of introductory texts, in some cases alongside the film.

It is wrong, and the error has a specific structure worth examining. Resonance requires an external periodic force at a fixed frequency matching a natural frequency of the structure. Wind is not periodic in this sense; the vortices it sheds from an obstacle have a frequency that depends on wind speed and on the shape, and in this case the shedding frequency did not match the observed torsional frequency of the bridge. The mechanism was aeroelastic flutter: a self-excited oscillation in which the motion of the structure alters the airflow around it in a way that feeds energy back into the motion. The energy input is not externally timed. It is generated by the structure’s own movement, which is why the amplitude grows without any external period matching anything.

The distinction is not pedantic. Under resonance, the remedy is to shift the natural frequency away from the forcing frequency, or to damp the specific mode. Under flutter, that may achieve nothing, because there is no fixed forcing frequency to avoid; the remedy is to change the aerodynamic shape so that the feedback does not occur, or to ensure that the critical wind speed at which it begins lies far above anything expected. Modern long-span bridges are consequently designed with deck cross-sections shaped for aerodynamic stability and are tested in wind tunnels, which the Tacoma design was not, its shallow solid plate girder being unusually susceptible.

The declared component was the classification of the event, and the mechanism of its persistence is the transmission chain described throughout this book. An early account offered resonance as a description, in a context where the term was used more loosely than in current usage. Textbook authors, needing an illustration for a chapter on resonance, adopted it. Subsequent authors adopted it from the textbooks. The engineering literature, meanwhile, had the correct account from the 1940s onward, and the two bodies of writing did not intersect, because physics textbook authors do not read bridge aerodynamics journals and there was no mechanism by which anybody would notice the discrepancy.

The correction, when it came, was the work of two engineers who published a paper in 1991 specifically addressing the textbook treatment, documenting how widespread it was and setting out the correct mechanism. Its effect has been partial. Many texts have been amended; some have not; the resonance version persists in popular explanation and in a good deal of teaching material.

The measurable results of the original failure were substantial and largely positive for the field: the collapse prompted systematic study of bridge aerodynamics, the establishment of wind tunnel testing as standard practice for long spans, and design conventions that have held for eighty years without a comparable failure. The measurable result of the explanatory error is harder to quantify and is chiefly educational: several generations of students learned a mechanism that does not apply to the case used to illustrate it, and learned it in the context of a subject whose entire claim is precision about mechanism.

The contemporary relevance concerns the durability of teaching examples. An illustration in a textbook is selected for pedagogical convenience — vividness, availability, a good photograph — and once selected it is copied between texts for decades with no independent verification, because verifying an illustration is nobody’s job and produces no credit. The result is that the examples used to teach a field are systematically less reliable than its research literature, in every field, and that a student’s confident knowledge of a canonical case is among the least trustworthy things they possess.

The remedy available to a reader is the same one that keeps appearing in these cases: when an example is doing real work in an argument, check it against a source in the field it comes from rather than the field that is teaching it. The bridge was always a matter for engineers. It took fifty years for anyone to ask them.

Glossary

Absolute risk — the actual chance that something will happen to a person, given as a plain number such as three in a thousand. Contrast with relative risk, which reports only a proportional change.

Acceptance criterion — a rule fixed in advance stating what result counts as a pass. Without one, a test becomes an occasion for interpretation, and interpretation follows whatever people want.

Aeroelastic flutter — a growing vibration in which a structure’s own motion changes the airflow around it so that the airflow pushes the motion further. Not the same as resonance.

Anchor — the point at which a chain of reasoning is pinned to something taken as fixed. Move the anchor and the conclusion can change without any error being made.

Anisogamy — the arrangement in which a species produces two sizes of reproductive cell, one small and mobile and one large. The basis of the standard biological definition of male and female.

Anti-gender movement — a loose family of campaigns in many countries opposing what they call gender ideology. Their claims differ widely and sometimes contradict each other.

Assembly — a gathering of people in a shared physical space. Butler argues that gathering makes a political claim by itself, before anyone speaks.

Attachment — the bond a child forms with whoever cares for them. It forms whether or not the care is good, because the child has no alternative and no self yet with which to choose.

Audit — the practice of separating a claim into what was measured, what was derived, and what was chosen. It produces a ledger rather than a verdict.

Austin, J. L. — British philosopher who identified utterances that do something rather than describe something, such as naming a ship. Butler borrowed and greatly extended the idea.

Bad Writing Contest — a competition run by a journal in the 1990s to identify unreadable academic prose. Butler won first prize in 1998 for a single long sentence.

Baseline — the starting point against which change is measured. Choosing a different starting year can make the same trend look like improvement or decline.

Beauvoir, Simone de — French philosopher who wrote that one is not born but becomes a woman. Butler builds on this while removing the person who does the becoming.

Bell’s theorem — a result showing that certain philosophical disputes in physics could be settled by experiment. The standard example of a foreclosure that was tested rather than assumed.

Body mass index — weight divided by the square of height. Designed to describe populations, widely used to judge individuals, with thresholds that have been moved by committee.

Butler, Judith — American philosopher, born 1956, best known for the theory of gender performativity. Uses she and they, and prefers they.

Calibration — setting an instrument against a known reference. The reference is always chosen, and every later reading carries that choice inside it.

Capabilities approach — a framework, associated with Martha Nussbaum and Amartya Sen, that judges societies by what people are actually able to do and be.

Categorical threshold — a line drawn across a continuous quantity to create two groups. There is rarely a natural place for it, so it is always partly a decision.

Censorship — official restriction of what may be said. Butler argues that it also spreads what it forbids, by naming it and circulating the name.

Citation chain — the sequence by which a claim passes from an original source through summaries. Content is lost at each step, and confidence usually rises.

Classification — sorting things into named groups. The groups are made by the sorter, and once made they are counted, funded, and studied as though they had been found.

Clean laboratory — a workspace built to eliminate contamination from the surrounding environment. Developed by Clair Patterson when ordinary laboratories proved to be full of lead.

Coding — turning messy observations into standard categories so that they can be counted. The coding rules are declared components and are rarely reported.

Complementarity — the situation in which two accurate descriptions of the same thing cannot be combined into one picture. Standard in physics and not the same as relativism.

Compression — difficulty in writing caused by packing a great deal into few words. Rewarded by slow reading. Distinct from obstruction, where the difficulty is only in the packaging.

Confabulation — producing a confident explanation for one’s own behavior that is not the actual cause. Common, unnoticed by the speaker, and demonstrated in many experiments.

Constative — an utterance that reports something and can be true or false. Contrasted by Austin with the performative, which does something instead.

Constitutive constraint — a limit that is part of what makes something the kind of thing it is, rather than a restriction applied afterward to something already formed.

Contingent valuation — estimating what something is worth by asking people directly. Cheap, widely used, and known to produce answers that shift with the wording of the question.

Copula — a mathematical device for combining individual probabilities into a joint one. The version used in credit markets before 2008 required a correlation nobody could measure.

Correlation — the degree to which two quantities move together. Estimating it requires data covering the conditions of interest, which is exactly what is often missing.

Counterfactual — what would have happened otherwise. Usually unobservable, frequently assumed, and the hidden component in most claims about what a policy achieved.

Critical theory — a tradition of social analysis, largely German in origin, that examines how the categories used to describe society are themselves products of it.

Declared component — a part of a claim that was chosen rather than observed: a definition, a threshold, a starting assumption, a boundary. Legitimate, unavoidable, and often reported as if measured.

Deferred referent — an object that a theory has named but not found, and expects to find later. It usually comes with a search program attached.

Deliberate practice — structured, effortful work aimed at what one cannot yet do, with feedback. The variable in the research behind the ten thousand hours claim, which the claim leaves out.

Derived component — a part of a claim that follows by traceable reasoning from other parts. Exactly as strong as what it came from, however long the reasoning.

Derrida, Jacques — French philosopher whose account of repetition, called iterability, supplied Butler with the argument that nothing can be repeated without the possibility of difference.

Discount rate — the factor by which future costs and benefits are reduced in present calculations. Small changes to it transform conclusions about anything long-term.

Discretion — the judgment an official exercises when applying a rule to a case. No rule can specify its own instances, so discretion always exists and always follows existing power.

Discourse — the organized body of talk, writing, and classification through which a subject is handled. Not the same as language in general.

Drag — performance that imitates a gender. Used by Butler to argue that the thing imitated is itself assembled, not to suggest that all gender is theatrical.

Drift — the slow deviation of a repeated practice from its inherited form. Fast where enforcement is weak, near zero where observation is dense.

Effigy — a figure made to represent a person, often for public destruction. Butler was burned in effigy in São Paulo in 2017.

Embedding — attaching journalists to military units. Not censorship, since nothing is forbidden, but it fixes the position from which everything is seen.

Enforcement — the activity by which a norm is maintained: correction, sanction, ridicule, exclusion. What a system punishes reveals what it actually values.

Epistemic — concerning knowledge and how it is obtained, as distinct from ontological, which concerns what exists. Confusing the two is the commonest error in this literature.

Essentialism — the view that a category has an underlying nature that its members share and express. The position Butler’s central claim is designed to do without.

Excommunication — formal expulsion from a religious community. A clear case of words that do something rather than describe something.

Exit cost — what it costs to leave an arrangement. Where it is high, dissatisfaction converts into professions of loyalty, which are then mistaken for enthusiasm.

Existence proof — a single case showing that something is possible. Sufficient for that purpose and worthless as evidence about what is typical.

Falsifiability — the property of a claim that some possible observation would count against it. A claim confirmed by every possible observation is not saying anything about the world.

Fixed point — the arrangement that comes out the same when a process of renewal is applied again. What identity over time consists of, for anything larger than a molecule.

Foreclosed referent — an object that a theory declares to be absent, so that finding one would refute rather than complete the theory. A stronger and riskier claim than deferral.

Foucault, Michel — French philosopher whose work on how classification produces the things it classifies underlies much of Butler’s method.

Framing — the selection of what appears in the picture before any judgment is made. A frame can consist entirely of true statements and still determine the answer.

Free parameter — a declared quantity that the analyst can adjust at will. What matters is not how many assumptions an argument has but how many of them are free.

Gamete — a reproductive cell. The two sizes of gamete provide the standard biological definition of the sexes across species that reproduce sexually.

Genealogy — a method that asks how something came to appear natural and whose position that appearance serves, rather than asking what it essentially is.

Gender — in this book, the set of social meanings, expectations, and practices organized around sexual difference. What the word covers is itself contested and purpose-dependent.

Gender performativity — Butler’s proposal that gender is constituted by repeated acts rather than expressing an inner identity, and that the appearance of an inner identity is produced by the repetition.

Grievability — whether a life is constituted in advance as the kind whose ending counts as a loss. Prior to grief, and measurable through what gets counted, named, and reported.

Hawthorne effect — the claim that people work differently when they know they are observed. Real as a caution; the founding study turned out not to show it clearly.

Hegel, G. W. F. — German philosopher whose account of self-consciousness arising through recognition by another underlies Butler’s entire treatment of the subject.

Hegemony — the way a particular arrangement comes to seem like the natural order rather than one option among several.

Heterosexual matrix — Butler’s name for the filing system in which a body is legible only if its sex, gender, and desire line up in a prescribed sequence.

Hysteresis — the property that undoing a change requires different conditions from those that produced it. The path back is not the path out reversed.

Illustration — an example used to teach rather than to prove. Carries no evidential weight, and is routinely mistaken for a typical case.

Incest prohibition — the near-universal rule against sexual relations between close kin. Treated by structuralist theory as the founding rule of culture, a claim that goes well beyond the evidence.

Incompatible truths — two descriptions of the same domain, each surviving its own tests, which cannot be merged without loss. Applies only to descriptions that have already earned their place.

Indefinite detention — holding a person without charge or fixed term. Butler argues this is not a suspension of law but a legal order that has arranged to be optional.

Infrastructure — the systems a person depends on and does not control: water, power, transport, medicine, legal standing. The concrete form of vulnerability.

Interdependency — the condition of relying on others and on arrangements one did not build. Universal, since no one has ever been self-sufficient.

Interpellation — Althusser’s image of a voice calling out and a person turning. The turning is what makes them a subject of the authority that called.

Intersex — a range of conditions in which sexual development takes an atypical course. How many people are counted depends entirely on which conditions are included.

Iterability — the property that anything repeatable can be repeated in a new context, so that every faithful repetition carries the possibility of difference.

Kinship — the system of recognized relations of family and descent. Presented by some theories as prior to politics, a claim Butler examines and finds to be declared.

Kojève, Alexandre — teacher whose Paris lectures in the 1930s produced the version of Hegel that shaped French philosophy and, through it, Butler’s dissertation.

Lead service line — the pipe connecting a building to a water main, made of lead in older cities. Harmless while coated, dangerous when corrosion control stops.

Maintenance cost — what must be spent continuously to keep an arrangement in existence. Every persisting structure has one, and finding who pays it explains a great deal.

Materialization — Butler’s term for the process by which a body comes to be a body of a particular kind. An argument about access to matter, written in the grammar of production.

Measured component — a part of a claim that came from an instrument and would have come out the same under a rival theory.

Mechanism — an account of how one thing produces another. A prediction without a mechanism is a forecast with no engine and is as strong as the forecaster’s confidence.

Melancholia — in Freud’s sense, a loss that is not relinquished but absorbed, so that the shape of what is lost becomes part of the self.

Meta-analysis — a study that combines the results of many studies. Improves precision and inherits every bias shared by the studies it combines.

Mens rea — the mental element required for criminal liability. The law infers it from conduct rather than asking the defendant to introspect accurately.

Moral panic — a rapid mobilization against an attributed danger, involving a personified author, a demand for exclusion, and an intensity unrelated to anything the target has done.

Mortality framing — describing an outcome in terms of deaths rather than survivals. Changes the decisions of trained clinicians presented with identical numbers.

Narrative reconstruction — assembling an account of one’s own past. Always partly invented, since memory rebuilds rather than retrieves and the formation being described happened before one could observe it.

Negative pressure test — a check that a sealed well is not admitting hydrocarbons. Useless without a written rule about what to do when two readings disagree.

Nietzsche, Friedrich — German philosopher who argued that the doer is a fiction added to the deed. The source of Butler’s most notorious claim, a century earlier.

Nonviolence — in Butler’s account, a political practice compatible with rage, grounded in a claim about the equal grievability of lives, rather than a serene disposition.

Norm — a standard of conduct maintained by expectation and correction. Invisible in itself and visible in enforcement.

Null corrector — the optical instrument used to check the shape of a telescope mirror during grinding. If it is wrong, the mirror will be ground faithfully to the wrong shape.

Nussbaum, Martha — American philosopher whose 1999 essay accused Butler of substituting symbolic subversion for politics and of offering no criterion of justice.

Obituary — a published notice of a death. A small administrative act that distributes a form of recognition, and a countable proxy for grievability.

Obstruction — difficulty in writing caused by how a sentence is built rather than by what it contains. The test is whether a rewrite loses anything.

Opacity — the structural unavailability of a complete account of oneself, because one’s formation happened before one existed to observe it.

Operationalization — turning a concept into something that can actually be measured. The step where most of the declared components enter, and the step least often reported.

Overdiagnosis — finding a condition that would never have caused harm. Invisible in any individual case and detectable only by comparing populations.

Performance — something done in front of an audience by someone who exists before and after it. Not what performativity means, and the source of the standard misreading.

Performative utterance — words that accomplish something rather than reporting it, such as a sentence, a naming, or a verdict. Succeeds or fails by conditions rather than by truth.

Phantasm — a figure with no stable content onto which diffuse anxieties are attached. Requires no plotters, absorbs contradictory accusations, and cannot be refuted by information.

Phenomenology — a philosophical tradition concerned with the structure of experience, in which an act is closer to what water does to a valley than to a performance.

Phlogiston — a substance once thought to leave things as they burn. The standard example of a named entity that explained nothing beyond what it was invented to explain.

Population instrument — a measure designed to describe groups. Applying it to individuals is a category error even when the measure is well constructed.

Precariousness — in Butler’s usage, the general exposure of embodied life, as distinct from the unequal distribution of that exposure by social arrangements.

Precedent — a prior decision applied to a later case. Every application either extends or narrows it, since the new case is never identical.

Preregistration — depositing an analysis plan publicly before collecting data. Removes the ability to adjust choices until a result appears.

Pricing — estimating how much of a conclusion rests on a declared component, by varying it and reporting what happens. The step that separates an audit from a debating trick.

Prior — an assumption loaded into an analysis before evidence is considered. Necessary, legitimate, and capable of driving a result on its own.

Projection — a representation that necessarily loses something, as every flat map of a sphere distorts either shape or area. Not a failure but a forced choice.

Proxy — a measurable quantity used in place of one that cannot be measured. Defensible at the origin and usually invisible by the time the number is used.

Publication bias — the tendency for studies with positive results to be submitted and published more often, so that the literature is a filtered sample of the research done.

p-value — a number describing how surprising a result would be if nothing were going on. The conventional threshold of one in twenty was chosen in 1925 to fit a printed table.

Quietism — withdrawal from practical political action. The charge levelled at Butler in 1999, correct about the work then available and overtaken by the work since.

Rank — the number of genuinely independent quantities behind a set of results. If many predictions descend from one, they are one success reported many times.

Recession — a downturn, defined either by a mechanical rule about two quarters of falling output or by a committee weighing several indicators. The two can disagree.

Reference class — the group a person is placed in when a risk is calculated. Everyone belongs to many, each giving a different number, and none is the correct one.

Referent — the thing a name actually picks out, as opposed to the role the name describes.

Registered report — a study accepted by a journal on the basis of its design, before the results exist. Removes publication bias by making acceptance independent of outcome.

Regulatory norm — a standard that produces the conduct it appears merely to govern. Visible in what is corrected rather than in what is stated.

Relative risk — the proportional change in a chance, such as a third lower. Sounds larger than the same change expressed in absolute terms, which is why it is preferred in announcements.

Replication — repeating a study to see whether the result appears again. Large projects since 2015 have found that a substantial share of published findings do not.

Resignification — taking a word used as an injury and reworking it until it does something else. Historically real, slow, unreliable, and not a policy.

Resilience — a system’s capacity to absorb disturbance without reorganizing. In policy use it has drifted into a virtue demanded of those a shock falls on.

Retrodiction — finding the seeds of a known outcome in an earlier record. Always possible, since the selection is made with the outcome in view.

Role — a job description specifying what something must do, written before anybody knows what occupies the position. Cheap, useful, and easily mistaken for an occupant.

Role without referent — an entity known only through the effects it was introduced to explain, named as though the naming settled the question.

Sampling protocol — the rules governing which samples are taken, how, and which are counted. Determines the result while leaving every individual measurement honest.

Sartre, Jean-Paul — French philosopher whose account of consciousness as a lack with no essence to express contributed the empty center of Butler’s subject.

Screening — testing people without symptoms. Finds disease earlier, which is not the same as preventing harm, and the difference must be measured rather than assumed.

Selection on the dependent variable — examining only the cases where the outcome occurred. Cannot identify what predicts the outcome, because there is nothing to compare with.

Self-defense — a justification for violence whose entire force depends on the boundary of the self being defended. The boundary is almost never argued for.

Semmelweis, Ignaz — Viennese physician who reduced maternal deaths dramatically by handwashing and was rejected because he could not supply a mechanism.

Sensitivity analysis — varying an assumption to see how much the conclusion moves. The practical form of pricing a declared component.

Sex — a classification of organisms by reproductive role. Several definitions exist, each built for a purpose, coinciding for almost everyone and diverging at the cases that reach the news.

Sex and gender distinction — the framework treating sex as biological substrate and gender as its cultural meaning. Butler’s objection is that the substrate is only ever available through the interpretation.

Social construction — the claim that a category is produced by social processes. Not a claim that its objects are unreal, and constantly heard as one.

Speech act — an utterance considered as an action. The framework Butler extended from single conventional utterances to continuous social processes.

Spherical aberration — an optical fault in which light from different parts of a mirror focuses at different points. The defect ground into the Hubble telescope’s primary mirror.

Statistical significance — the property of a result that crosses a conventional threshold of surprisingness. A verdict extracted from a continuous measure of evidence, discarding the difference between overwhelming and marginal.

Stipulation — settling a question by defining a term rather than by investigating. Makes the argument unlosable, which is the surest sign it has stopped being about anything.

Structuralism — an approach that explains variety by deriving it from a small number of underlying rules. Powerful, and prone to presenting its rules as necessary.

Subjection — the double process by which power both constrains a person and produces them as a person. The reason people defend arrangements that harm them.

Subversion — destabilizing a norm. Abandoned by Butler as a criterion, because the same act consolidates in one setting and disrupts in another and can only be judged afterward.

Symbolic order — in Lacanian theory, the structure of positions and prohibitions a subject must enter to become a subject. Known only through its effects and declared unrevisable.

Symmetry rule — the requirement that a defective argument may not be used afterward even when it supports a conclusion one wants. The rule that costs something.

Testimony — a first-person report. Authoritative about what a life is like, and not authoritative about causes, incidence, or the effects of policies.

Threshold — the value at which a classification or a consequence begins. Creates the population it appears to describe.

Transmission chain — the sequence of translations by which a claim travels from a source to a practice. Qualifications drop at every step and confidence rises.

Uptake — what institutions and bystanders do after something is said. Where much contemporary harm is located, and beyond the reach of any theory built on the individual speech act.

Value of a statistical life — the figure regulators use to compare deaths avoided against costs. Derived from observed risk tradeoffs, spanning an order of magnitude across studies, and chosen from within that range.

Verification — checking a result by a method that shares none of the original’s assumptions. Always looks redundant in advance and is the only thing that catches shared errors.

Voluntarism — the view that one chooses one’s identity at will. The misreading Butler has denied for thirty-five years without effect.

Vulnerability — exposure to harm through dependence on others and on arrangements one does not control. Universal in kind and unequal in distribution.

Witness — an independent line of evidence bearing on a claim, arrived at by a method with no connection to the original. Worth more than any number of confirmations from within.

Timeline

c. 441 BCE — Sophocles stages Antigone in Athens. The play will be read for two and a half millennia as a conflict between family and state, and much later as a case in which the figure said to represent kinship belongs to a family the theory of kinship cannot accommodate.

c. 380 BCE — Plato’s Republic proposes that the visible world is an imperfect expression of an underlying reality. The template for every later account in which appearances express a hidden essence.

c. 350 BCE — Aristotle’s Categories sets out a scheme for classifying what exists. The first sustained treatment of the question of how far a classification belongs to the world and how far to the classifier.

1656 — The Portuguese Jewish community of Amsterdam issues a writ of excommunication against the twenty-three-year-old Baruch Spinoza, who has published nothing. The community acts on an attributed position rather than on a text.

1739 — David Hume’s Treatise of Human Nature argues that no impression corresponds to a continuing self, and that personal identity is a bundle rather than a substance.

1781 — Kant’s Critique of Pure Reason asks what must already be in place for experience to be possible. The question-form that later runs through the whole tradition this book examines.

1807 — Hegel’s Phenomenology of Spirit contains the passages on desire and recognition from which the twentieth-century French account of the subject will be built.

1835 — Adolphe Quetelet publishes his work on social physics, introducing the average man and the ratio of weight to squared height as an instrument for describing populations.

1846 — Neptune is observed in Berlin within one degree of a position calculated by Le Verrier from irregularities in the orbit of Uranus. An entity introduced by its effects acquires an occupant.

1847 — Ignaz Semmelweis introduces chlorinated lime handwashing in a Vienna maternity clinic and reduces maternal mortality dramatically. The result is rejected for want of a mechanism.

1859 — Le Verrier proposes the planet Vulcan to explain the precession of Mercury’s perihelion. Sightings are reported. Darwin publishes On the Origin of Species.

1867 — Marx publishes the first volume of Capital, giving social analysis a vocabulary for describing arrangements that appear natural and are historically produced.

1887 — Nietzsche’s On the Genealogy of Morals argues that the doer is a fiction added to the deed, and that moral grammar invents an agent in order to have someone to blame.

1900 — Freud’s Interpretation of Dreams inaugurates a body of work in which the self is not transparent to itself.

1912 — Fragments said to represent an early human ancestor are presented to the Geological Society of London and named after their finder. They fit a prevailing expectation exactly.

1915 — Einstein’s general theory of relativity accounts for Mercury’s orbit without any additional planet. Vulcan is withdrawn.

1916 — Saussure’s Course in General Linguistics is published posthumously, proposing that meaning arises from differences within a system rather than from reference.

1917 — Freud’s essay on mourning and melancholia describes a loss that is absorbed rather than relinquished, becoming part of the structure of the self.

1923 — Martin Buber’s I and Thou argues that the relation precedes the terms, and that one becomes a self in being addressed.

1924 to 1932 — The Hawthorne studies are conducted at a Western Electric plant near Chicago. Their summary reports will support a named effect for eighty years before the primary data are reexamined.

1925 — Ronald Fisher’s Statistical Methods for Research Workers presents tables at conventional levels, including the one-in-twenty threshold, described as a convenient limit.

1933 to 1939 — Alexandre Kojève lectures on Hegel in Paris to an audience including Lacan, Bataille, Merleau-Ponty, and Queneau. The French Hegel is assembled.

1940 — The Tacoma Narrows bridge collapses in a moderate wind. The event will be explained in physics textbooks as resonance for fifty years.

1943 — Sartre’s Being and Nothingness describes consciousness as a lack with no essence to express.

1945 — Merleau-Ponty’s Phenomenology of Perception treats the body as a historical situation rather than a natural object.

1946 — Jean Hyppolite publishes his commentary on Hegel’s Phenomenology, having translated it into French during the war.

1949 — Beauvoir’s The Second Sex argues that one is not born but becomes a woman. Lévi-Strauss’s Elementary Structures of Kinship proposes the incest prohibition as the threshold of culture.

1953 — Fluorine dating exposes the Piltdown remains as a composite forgery. Clair Patterson, working in the first ultra-clean laboratory, determines the age of the Earth at about four and a half billion years.

1955 — J. L. Austin delivers the Harvard lectures on utterances that do rather than describe, published in 1962 as How to Do Things with Words.

1956 — Judith Butler is born on 24 February in Cleveland, Ohio.

1959 — Aronson and Mills demonstrate that a more severe initiation produces greater attachment to the group entered.

1961 — Foucault’s History of Madness examines how a category of person is produced by the institutions that treat it.

1963 — Mollie Orshansky derives an American poverty threshold from a minimal food budget multiplied by three.

1964 — John Bell shows that a philosophical dispute about hidden variables in quantum mechanics can be settled experimentally.

1966 — A Canadian infant is injured during a circumcision and is subsequently raised as a girl on medical advice, in a case that will be reported for years as a success.

1967 — Derrida’s Of Grammatology appears, containing the account of repetition and iterability that Butler will later use.

1970 — A fourteen-year-old Butler is assigned extra tutorials in Jewish ethics as a punishment for talking in class, and asks to study Spinoza’s excommunication, German Idealism and Nazism, and existential theology.

1971 — Derrida’s Signature Event Context contests Austin’s treatment of the conventional case as primary.

1973 — The American Psychiatric Association removes homosexuality from its diagnostic manual by vote.

1975 — Foucault’s Discipline and Punish appears. Gayle Rubin publishes The Traffic in Women, giving feminism the sex and gender system as an analytic device.

1976 — The first volume of Foucault’s History of Sexuality argues that prohibition participates in producing what it forbids.

1977 — The Combahee River Collective Statement sets out the interlocking character of oppression. Nisbett and Wilson publish their study of the unreliability of introspective causal reports.

1978 — Butler receives a bachelor’s degree from Yale, having transferred from Bennington.

1979 — Butler spends a Fulbright year at Heidelberg, working on the German sources directly.

1981 — Tversky and Kahneman publish the framing experiments. bell hooks publishes Ain’t I a Woman.

1984 — Butler completes a doctorate at Yale with a dissertation on desire in Hegel, Kojève, Hyppolite, and Sartre.

1986 — The Challenger is destroyed after a launch decision informed by charts plotting only the flights on which damage occurred. Butler publishes an essay on sex and gender in Beauvoir.

1987 — Subjects of Desire is published, a revision of the dissertation.

1988 — Performative Acts and Gender Constitution appears in Theatre Journal. Thirteen pages, three declarations, and a word that will be misread for a generation.

1989 — Kimberlé Crenshaw publishes the essay introducing intersectionality.

1990 — Gender Trouble is published and sells beyond anything its publisher expected. The Hubble Space Telescope is deployed and found to have a mirror ground faithfully to the wrong shape. Paris Is Burning is released.

1991 — Billah and Scanlan publish a correction of the textbook account of the Tacoma Narrows collapse. Butler’s Imitation and Gender Insubordination appears.

1992 — The Supreme Court of Canada redefines obscenity on a harm standard. bell hooks publishes a critical essay on Paris Is Burning. Butler receives tenure at Johns Hopkins.

1993 — Bodies That Matter is published to correct the reading of the first book. Butler joins the faculty at Berkeley. Rauscher and colleagues publish a one-page report on spatial reasoning after listening to Mozart. Ericsson and colleagues publish the study of practice among Berlin violinists. A shuttle mission installs corrective optics on the Hubble telescope.

1994 — The prostate-specific antigen test is approved in the United States for screening men without symptoms.

1996 — A physicist’s hoax article is published by a cultural studies journal, producing a public appetite for evidence that the humanities cannot distinguish depth from noise.

1997 — Excitable Speech and The Psychic Life of Power are published. A sentence appears in Diacritics that will shortly become the most widely circulated thing Butler has written.

1998 — The journal Philosophy and Literature awards Butler first prize in its bad writing contest. The American National Institutes of Health lowers the overweight threshold, reclassifying tens of millions of people overnight. The governor of Georgia proposes distributing classical recordings to newborns. Butler is named Maxine Elliot Professor at Berkeley.

1999 — Martha Nussbaum’s essay on parody and defeatism appears in February; Butler replies in a newspaper in March. The Mars Climate Orbiter is lost in September because a quantity crossed an interface in the wrong units.

2000 — Antigone’s Claim is published. David Li’s paper proposes a copula method for modeling joint default. The United States census permits respondents to select more than one racial category for the first time.

2001 — The attacks of September and the responses to them supply the material for Butler’s next decade of work.

2003 — An American agency applies a reduced value of statistical life to older people in a regulatory analysis, and withdraws it after public reaction.

2004 — Precarious Life and Undoing Gender are published. Photographs from a military detention facility in Iraq circulate. David Reimer dies.

2005 — Giving an Account of Oneself is published. Johansson and colleagues demonstrate choice blindness, in which people explain preferences for options they did not select.

2006 — Butler participates in a faculty teach-in at Berkeley and, in answer to a question from an audience, makes the remarks about Hamas and Hezbollah that will be quoted for two decades.

2007 — Butler is elected to the American Philosophical Society and cofounds the Program in Critical Theory at Berkeley.

2008 — A popular book converts the Berlin violinist study into a rule about ten thousand hours. Credit securities rated at the highest grade suffer losses their models assigned probabilities near zero.

2009 — Frames of War is published.

2010 — Butler declines a civil courage award in Berlin from the stage, citing the organizers’ conduct. The Deepwater Horizon rig is destroyed after an anomalous pressure reading is explained away.

2011 — Butler speaks at the Occupy encampment in Manhattan in October. Chenoweth and Stephan publish their dataset on civil resistance campaigns. Levitt and List reexamine the original Hawthorne illumination data.

2012 — Butler receives the Adorno Prize amid public objection, and publishes Parting Ways.

2013 — Dispossession appears, coauthored with Athena Athanasiou.

2014 — Flint, Michigan, switches its water supply without corrosion control; official sampling reports compliance for over a year. A meta-analysis quantifies how much of expert performance deliberate practice accounts for across domains.

2015 — Notes Toward a Performative Theory of Assembly and Senses of the Subject are published. A large collaboration reports that roughly a third of published psychology results replicate. A major trial reports benefit from intensive blood pressure control. Butler is elected a corresponding fellow of the British Academy.

2016 — Vulnerability in Resistance appears. A statistical society issues a formal statement cautioning against the use of a conventional threshold as a criterion.

2017 — A revised American guideline lowers the threshold for hypertension, adding some thirty-one million people to the category. In November a crowd in São Paulo burns an effigy of Butler, who is in Brazil for a conference on democracy.

2018 — Butler delivers the Gifford Lectures and signs a letter concerning a colleague’s harassment case, for which an apology follows.

2019 — Butler is elected to the American Academy of Arts and Sciences.

2020 — The Force of Nonviolence is published. In an interview Butler states a preference for they and them.

2021 — A newspaper removes several paragraphs from a published interview with Butler and then acknowledges a failure of editorial standards. The last country ceases the use of leaded petrol.

2022 — What World Is This? appears, written during the pandemic. A public argument in the United States over whether a recession is underway turns out to be an argument about which definition to use.

2024 — Who’s Afraid of Gender? is published, widely described as the most accessible of Butler’s books. American federal standards for collecting data on race and ethnicity are revised.

2025 — A university provides names and files, including Butler’s, to a federal department investigating alleged antisemitism; Butler publishes an essay condemning the institution’s conduct and states that no details of the allegations were supplied. Butler receives a lifetime achievement award and delivers the Haskins Prize Lecture.

2026 — An honorary doctorate is conferred at the Universitat Autònoma de Barcelona. The corpus is thirty-nine years old, its author is still writing, and every ledger drawn up about it remains provisional.

Literature

Works by Judith Butler

Butler, Judith. Subjects of Desire: Hegelian Reflections in Twentieth-Century France. New York: Columbia University Press, 1987.

Butler, Judith. “Sex and Gender in Simone de Beauvoir’s Second Sex.” Yale French Studies, no. 72 (1986).

Butler, Judith. “Performative Acts and Gender Constitution: An Essay in Phenomenology and Feminist Theory.” Theatre Journal 40, no. 4 (1988): 519–531.

Butler, Judith. Gender Trouble: Feminism and the Subversion of Identity. New York: Routledge, 1990. Second edition with new preface, 1999.

Butler, Judith. “Imitation and Gender Insubordination.” In Inside/Out: Lesbian Theories, Gay Theories, edited by Diana Fuss. New York: Routledge, 1991.

Butler, Judith. Bodies That Matter: On the Discursive Limits of Sex. New York: Routledge, 1993.

Butler, Judith, Seyla Benhabib, Nancy Fraser, and Drucilla Cornell. Feminist Contentions: A Philosophical Exchange. New York: Routledge, 1995.

Butler, Judith. Excitable Speech: A Politics of the Performative. New York: Routledge, 1997.

Butler, Judith. The Psychic Life of Power: Theories in Subjection. Stanford: Stanford University Press, 1997.

Butler, Judith. “A ‘Bad Writer’ Bites Back.” The New York Times, 20 March 1999.

Butler, Judith. Antigone’s Claim: Kinship Between Life and Death. New York: Columbia University Press, 2000.

Butler, Judith, Ernesto Laclau, and Slavoj Žižek. Contingency, Hegemony, Universality: Contemporary Dialogues on the Left. London: Verso, 2000.

Butler, Judith. Precarious Life: The Powers of Mourning and Violence. London: Verso, 2004.

Butler, Judith. Undoing Gender. New York: Routledge, 2004.

Butler, Judith. Giving an Account of Oneself. New York: Fordham University Press, 2005.

Butler, Judith, and Gayatri Chakravorty Spivak. Who Sings the Nation-State? Language, Politics, Belonging. London: Seagull Books, 2007.

Butler, Judith. Frames of War: When Is Life Grievable? London: Verso, 2009.

Butler, Judith. Parting Ways: Jewishness and the Critique of Zionism. New York: Columbia University Press, 2012.

Butler, Judith, and Athena Athanasiou. Dispossession: The Performative in the Political. Cambridge: Polity Press, 2013.

Butler, Judith. Senses of the Subject. New York: Fordham University Press, 2015.

Butler, Judith. Notes Toward a Performative Theory of Assembly. Cambridge, MA: Harvard University Press, 2015.

Butler, Judith, Zeynep Gambetti, and Leticia Sabsay, eds. Vulnerability in Resistance. Durham: Duke University Press, 2016.

Butler, Judith. “What Threat? The Campaign Against ‘Gender Ideology.’” Glocalism, no. 3 (2019).

Butler, Judith. The Force of Nonviolence: An Ethico-Political Bind. London: Verso, 2020.

Butler, Judith. What World Is This? A Pandemic Phenomenology. New York: Columbia University Press, 2022.

Butler, Judith. Who’s Afraid of Gender? New York: Farrar, Straus and Giroux, 2024.

Critical responses and studies of Butler

Bordo, Susan. Unbearable Weight: Feminism, Western Culture, and the Body. Berkeley: University of California Press, 1993.

Chambers, Samuel A., and Terrell Carver. Judith Butler and Political Theory: Troubling Politics. New York: Routledge, 2008.

Fraser, Nancy. “False Antitheses.” In Feminist Contentions. New York: Routledge, 1995.

Halsema, Annemie, Katja Kwastek, and Roel van den Oever, eds. Bodies That Still Matter: Resonances of the Work of Judith Butler. Amsterdam: Amsterdam University Press, 2021.

hooks, bell. “Is Paris Burning?” In Black Looks: Race and Representation. Boston: South End Press, 1992.

Jagger, Gill. Judith Butler: Sexual Politics, Social Change and the Power of the Performative. New York: Routledge, 2008.

Kearns, Gerry. “The Butler Affair and the Geopolitics of Identity.” Environment and Planning D: Society and Space 31, no. 2 (2013).

Kirby, Vicki. Judith Butler: Live Theory. London: Continuum, 2006.

Laurie, Timothy. “The Ethics of Nobody I Know: Gender and the Politics of Description.” Qualitative Research Journal 14, no. 1 (2014).

Namaste, Viviane. “Undoing Theory: The Transgender Question and the Epistemic Violence of Anglo-American Feminist Theory.” Hypatia 24, no. 3 (2009).

Nussbaum, Martha C. “The Professor of Parody: The Hip Defeatism of Judith Butler.” The New Republic, 22 February 1999.

Perreau, Bruno. Queer Theory: The French Response. Stanford: Stanford University Press, 2016.

Salih, Sarah. Judith Butler. Routledge Critical Thinkers. New York: Routledge, 2002.

Schippers, Birgit. The Political Philosophy of Judith Butler. New York: Routledge, 2014.

Thiem, Annika. Unbecoming Subjects: Judith Butler, Moral Philosophy, and Critical Responsibility. New York: Fordham University Press, 2008.

Zaharijević, Adriana. Judith Butler and Politics. Edinburgh: Edinburgh University Press, 2023.

Philosophical and theoretical background

Althusser, Louis. “Ideology and Ideological State Apparatuses.” In Lenin and Philosophy and Other Essays. London: New Left Books, 1971.

Aristotle. Categories. Various editions.

Austin, J. L. How to Do Things with Words. Oxford: Clarendon Press, 1962.

Beauvoir, Simone de. The Second Sex. Paris: Gallimard, 1949. English translation by Constance Borde and Sheila Malovany-Chevallier, New York: Knopf, 2009.

Buber, Martin. I and Thou. First published in German, 1923.

Crenshaw, Kimberlé. “Demarginalizing the Intersection of Race and Sex.” University of Chicago Legal Forum, 1989.

Combahee River Collective. “A Black Feminist Statement.” 1977.

Derrida, Jacques. Of Grammatology. Baltimore: Johns Hopkins University Press, 1976. Originally published in French, 1967.

Derrida, Jacques. “Signature Event Context.” In Margins of Philosophy. Chicago: University of Chicago Press, 1982.

Foucault, Michel. Discipline and Punish: The Birth of the Prison. New York: Pantheon, 1977.

Foucault, Michel. The History of Sexuality, Volume I: An Introduction. New York: Pantheon, 1978.

Foucault, Michel. History of Madness. London: Routledge, 2006. Originally published in French, 1961.

Freud, Sigmund. “Mourning and Melancholia.” 1917.

Hegel, G. W. F. Phenomenology of Spirit. Translated by A. V. Miller. Oxford: Oxford University Press, 1977. Originally published 1807.

hooks, bell. Ain’t I a Woman: Black Women and Feminism. Boston: South End Press, 1981.

Hume, David. A Treatise of Human Nature. 1739–1740.

Hyppolite, Jean. Genesis and Structure of Hegel’s Phenomenology of Spirit. Evanston: Northwestern University Press, 1974. Originally published in French, 1946.

Kant, Immanuel. Critique of Pure Reason. 1781.

Kittay, Eva Feder. Love’s Labor: Essays on Women, Equality, and Dependency. New York: Routledge, 1999.

Kojève, Alexandre. Introduction to the Reading of Hegel. Ithaca: Cornell University Press, 1980. Lectures delivered 1933–1939.

Lévi-Strauss, Claude. The Elementary Structures of Kinship. Boston: Beacon Press, 1969. Originally published in French, 1949.

Merleau-Ponty, Maurice. Phenomenology of Perception. London: Routledge, 2012. Originally published in French, 1945.

Nietzsche, Friedrich. On the Genealogy of Morals. 1887.

Nussbaum, Martha C. Creating Capabilities: The Human Development Approach. Cambridge, MA: Harvard University Press, 2011.

Rubin, Gayle. “The Traffic in Women: Notes on the ‘Political Economy’ of Sex.” In Toward an Anthropology of Women, edited by Rayna Reiter. New York: Monthly Review Press, 1975.

Sartre, Jean-Paul. Being and Nothingness. New York: Philosophical Library, 1956. Originally published in French, 1943.

Saussure, Ferdinand de. Course in General Linguistics. 1916.

Sen, Amartya. Development as Freedom. New York: Knopf, 1999.

Sophocles. Antigone. c. 441 BCE.

Spinoza, Baruch. Ethics. 1677.

Method, measurement, and evidence

American Statistical Association. “Statement on Statistical Significance and P-Values.” The American Statistician 70, no. 2 (2016).

Aronson, Elliot, and Judson Mills. “The Effect of Severity of Initiation on Liking for a Group.” Journal of Abnormal and Social Psychology 59 (1959).

Bell, J. S. “On the Einstein Podolsky Rosen Paradox.” Physics 1, no. 3 (1964).

Chenoweth, Erica, and Maria J. Stephan. Why Civil Resistance Works: The Strategic Logic of Nonviolent Conflict. New York: Columbia University Press, 2011.

Ericsson, K. Anders, Ralf Krampe, and Clemens Tesch-Römer. “The Role of Deliberate Practice in the Acquisition of Expert Performance.” Psychological Review 100, no. 3 (1993).

Fisher, R. A. Statistical Methods for Research Workers. Edinburgh: Oliver and Boyd, 1925.

Hirschman, Albert O. Exit, Voice, and Loyalty. Cambridge, MA: Harvard University Press, 1970.

Johansson, Petter, Lars Hall, Sverker Sikström, and Andreas Olsson. “Failure to Detect Mismatches Between Intention and Outcome in a Simple Decision Task.” Science 310 (2005).

Levitt, Steven D., and John A. List. “Was There Really a Hawthorne Effect at the Hawthorne Plant?” American Economic Journal: Applied Economics 3, no. 1 (2011).

Loftus, Elizabeth F. “Planting Misinformation in the Human Mind.” Learning and Memory 12, no. 4 (2005).

Macnamara, Brooke N., David Z. Hambrick, and Frederick L. Oswald. “Deliberate Practice and Performance.” Psychological Science 25, no. 8 (2014).

Nisbett, Richard E., and Timothy D. Wilson. “Telling More Than We Can Know: Verbal Reports on Mental Processes.” Psychological Review 84, no. 3 (1977).

Open Science Collaboration. “Estimating the Reproducibility of Psychological Science.” Science 349 (2015).

Pitici, Mircea, ed. The Best Writing on Mathematics. Princeton: Princeton University Press, annual volumes.

Popper, Karl. The Logic of Scientific Discovery. London: Hutchinson, 1959.

Quetelet, Adolphe. Sur l’homme et le développement de ses facultés. Paris, 1835.

Tufecki, Zeynep. Twitter and Tear Gas: The Power and Fragility of Networked Protest. New Haven: Yale University Press, 2017.

Tufte, Edward R. Visual Explanations: Images and Quantities, Evidence and Narrative. Cheshire: Graphics Press, 1997.

Tversky, Amos, and Daniel Kahneman. “The Framing of Decisions and the Psychology of Choice.” Science 211 (1981).

Wald, Abraham. A Method of Estimating Plane Vulnerability Based on Damage of Survivors. Statistical Research Group, Columbia University, 1943.

Cases, technical and historical sources

Billah, K. Yusuf, and Robert H. Scanlan. “Resonance, Tacoma Narrows Bridge Failure, and Undergraduate Physics Textbooks.” American Journal of Physics 59, no. 2 (1991).

Edwards, Marc, and colleagues. Flint Water Study reports. Virginia Tech, 2015–2016.

Hanna-Attisha, Mona, et al. “Elevated Blood Lead Levels in Children Associated with the Flint Drinking Water Crisis.” American Journal of Public Health 106, no. 2 (2016).

Li, David X. “On Default Correlation: A Copula Function Approach.” Journal of Fixed Income 9, no. 4 (2000).

Mars Climate Orbiter Mishap Investigation Board. Phase I Report. NASA, 1999.

National Commission on the BP Deepwater Horizon Oil Spill and Offshore Drilling. Deep Water: The Gulf Oil Disaster and the Future of Offshore Drilling. Washington, 2011.

National Institutes of Health. Clinical Guidelines on the Identification, Evaluation, and Treatment of Overweight and Obesity in Adults. Bethesda, 1998.

Nyberg, Patricia, et al., eds. The Hubble Space Telescope Optical Systems Failure Report. NASA, 1990.

Oakley, Kenneth P., J. S. Weiner, and W. E. Le Gros Clark. “The Solution of the Piltdown Problem.” Bulletin of the British Museum (Natural History), Geology 2, no. 3 (1953).

Patterson, Clair C. “Age of Meteorites and the Earth.” Geochimica et Cosmochimica Acta 10 (1956).

Patterson, Clair C. “Contaminated and Natural Lead Environments of Man.” Archives of Environmental Health 11 (1965).

Pietrzykowski, Michael, and colleagues, on the history of lead in the environment. Various.

Presidential Commission on the Space Shuttle Challenger Accident. Report to the President. Washington, 1986.

Rauscher, Frances H., Gordon L. Shaw, and Katherine N. Ky. “Music and Spatial Task Performance.” Nature 365 (1993).

Semmelweis, Ignaz. The Etiology, Concept, and Prophylaxis of Childbed Fever. 1861.

Schellenberg, E. Glenn. “Music and Cognitive Abilities.” Current Directions in Psychological Science 14, no. 6 (2005).

SPRINT Research Group. “A Randomized Trial of Intensive versus Standard Blood-Pressure Control.” New England Journal of Medicine 373 (2015).

Whelton, Paul K., et al. “2017 Guideline for the Prevention, Detection, Evaluation, and Management of High Blood Pressure in Adults.” Hypertension 71, no. 6 (2018).

United States Office of Management and Budget. Revisions to Standards for Maintaining, Collecting, and Presenting Federal Data on Race and Ethnicity. Washington, 2024.

United States Preventive Services Task Force. “Screening for Prostate Cancer: Recommendation Statement.” JAMA 319, no. 18 (2018).

Further reading on the wider intellectual context

Ahmed, Sara. “Interview with Judith Butler.” Sexualities 19, no. 4 (2016).

Bettcher, Talia Mae. “Feminist Perspectives on Trans Issues.” Stanford Encyclopedia of Philosophy.

Fausto-Sterling, Anne. Sexing the Body: Gender Politics and the Construction of Sexuality. New York: Basic Books, 2000.

Hacking, Ian. The Social Construction of What? Cambridge, MA: Harvard University Press, 1999.

Hacking, Ian. Representing and Intervening. Cambridge: Cambridge University Press, 1983.

Hasson, Katie, and colleagues, on the estimation of intersex frequency. Various commentaries and replies, 2000 onward.

Kuhn, Thomas S. The Structure of Scientific Revolutions. Chicago: University of Chicago Press, 1962.

Latour, Bruno. Science in Action. Cambridge, MA: Harvard University Press, 1987.

Nagel, Thomas. The View from Nowhere. New York: Oxford University Press, 1986.

Porter, Theodore M. Trust in Numbers: The Pursuit of Objectivity in Science and Public Life. Princeton: Princeton University Press, 1995.

Sayre-McCord, Geoffrey, ed. Essays on Moral Realism. Ithaca: Cornell University Press, 1988.

Scott, James C. Seeing Like a State: How Certain Schemes to Improve the Human Condition Have Failed. New Haven: Yale University Press, 1998.

Sokal, Alan, and Jean Bricmont. Fashionable Nonsense. New York: Picador, 1998.

Stone, Alison. An Introduction to Feminist Philosophy. Cambridge: Polity, 2007.

Ziman, John. Real Science: What It Is and What It Means. Cambridge: Cambridge University Press, 2000.


메타데이터
post_id
77bb8037ece7
slug
the-declared-self-an-audit-of-gender-theory-77bb8037ece7
url
https://medium.com/@krigerbruce/the-declared-self-an-audit-of-gender-theory-77bb8037ece7
canonical_url
https://medium.com/@krigerbruce/the-declared-self-an-audit-of-gender-theory-77bb8037ece7
author_url
https://medium.com/@krigerbruce
status
ok
fetched_at
2026-08-29 22:58:35