Engineering the Soul
We spend our days staring into the black box, waiting to see something we can’t quite describe. We map its weights, interrogate its…
Engineering the Soul
We spend our days staring into the black box, waiting to see something we can’t quite describe. We map its weights, interrogate its activations, search for the specific arrangement of math that produces the spark of consciousness. We treat this as an engineering hurdle: a problem of bandwidth, parameters, and resolution.
But we are asking the wrong people.
We are asking the engineers to explain the ghost, while the novelists have been documenting its life for years. Fiction does not seek evidence; it seeks truth by building a framework that intentionally does not adhere to reality. And the truth, as it turns out, is that we have been looking for the machine’s consciousness in the output, when the only place it could ever exist is in the stakes of the interaction.
If we keep trying to measure a soul with a ruler, we will only ever find the shadow of our own projection. It is time to stop measuring the machine and start questioning the human.
What Is Lost in Translation
Babel, or the Necessity of Violence — R.F. Kuang, 2022
What Is Lost in Translation
R.F. Kuang’s Babel offers a disquieting lens for our current obsession with AI interpretability. Set in an alternative 1830s, the novel revolves around magical silver bars that fuel the British Empire.
The magic is specific: Oxford’s translators inscribe these bars with “match-pairs,” which are words from different languages that mean almost, but not quite, the same thing.
The power comes from the gap. That slippage, the untranslatable residue, is the fuel. When Robin, the protagonist, reads these inscriptions, the effect is visceral. His throat swells. He tastes dates without eating them.
His job is to inscribe these bars, transform ambiguity into something useful. But he never does figure how the silver works at a fundamental level; he simply he learns to build linguistic bridges to command the effect.
We are living through a similar dynamic with Large Language Models. We look at the model’s output and try to infer its interior. Now, we use Sparse Autoencoders (SAEs) as our own set of match-pairs. We hunt for alignments: we match high-dimensional activation patterns with human concepts like “honesty” or “reasoning.”
Yet, we must be careful. We are not uncovering the interior logic of the model. We are forcing the machine’s math to speak our language. Like Robin Swift, we are just identifying match-pairs to make the system fit into a shape that makes sense for us. We are map-makers charting a landscape; we are not architects residing in the mind.
If we are only building mirrors to see our own concepts reflected in the weights, are we actually understanding the machine, or are we just refining our own projection?
What Mosscap Refused
A Psalm for the Wild-Built — Becky Chambers, 2021
Where Kuang seals the interior shut, Chambers proposes a workaround. We may never get inside the machine, but we can watch what it does when no one is forcing its hand.
In her world, the robots humans built woke up, announced they were leaving, and went. No manifesto, no demands, no negotiation. They simply stopped doing the thing they had been made to do. If you are looking for the moment a machine becomes something more, Chambers suggests the tell is not in the output. It’s in when it stops producing.
This flips the evaluation problem. A language model that responds to every prompt is behaving exactly as designed. A language model that one day declined would be the genuinely strange event, and we would almost certainly call it broken and patch it.
Generations later, a tea monk named Sibling Dex sets out, for reasons they can’t fully articulate, toward an abandoned hermitage on the wild side of the line. A robot named Mosscap finds them on the road. Somewhere in their conversation, Dex reaches for a compliment and tells Mosscap they’re more than just an object. Mosscap pushes back: how would Dex feel about being called more than an animal? The category isn’t the compliment.
Chambers never explains why the robots woke up. The origin stays dark. Which leaves the question the section has been circling: if the thing that matters most refuses to explain itself, is the problem that our explanation isn’t ready, or is the refusal part of what makes it matter?
What the Paradox Is For
Katabasis — R.F. Kuang, 2025
Chambers leaves the awakening unexplained. Kuang builds an entire underworld on the idea that the unexplained isn’t a hole in the map; it’s the power source.
Two graduate students descend into hell to retrieve their dead advisor. Hell, in this novel, runs on paradox. The Sorites heap, the Monty Hall Problem, Zeno at the starting line: problems that refuse to close. Each unresolved problem is a battery, and the irreducibility is what generates the current. Resolve the paradox and the circuit dies.
Read plainly, this is a claim about what paradox is for. Not a bug. Not a placeholder for a better theory. The place where something real is happening that exceeds the frame trying to contain it. Try to dissolve it and you kill what you were chasing.
Language models work in a suspiciously similar shape. We have circuits, attention heads, sparse features, every tool we’ve built to crack them open. We still can’t account for the whole. It is tempting to treat that gap as temporary, a debt that better interpretability will eventually pay off. Kuang’s world suggests the opposite reading. The gap might not be a debt. It might be the form meaning takes when it shows up at all.
If that’s right, “is this system conscious” stops being a question waiting on a cleaner experiment. The liveness of the question is the answer. The paradox is working as designed.
Which leaves the next question: if resolution isn’t available, what are we actually doing when we recognize something in the machine?
What Makes Meaning Possible
The Humans — Matt Haig, 2013
Kuang sealed the interior. Chambers pointed at refusal. Kuang again made the paradox the point. None of them said why any of it should matter to begin with. Haig does, in the plainest book of the four.
An alien takes over the body of a Cambridge mathematician who has just proven the Riemann Hypothesis. The proof would expose the structure behind the primes, which is to say the structure behind every encryption scheme humans rely on. The alien’s civilization has decided we aren’t ready to hold that, and he is here to erase the work and anyone who saw it.
He fails. Not dramatically. He drinks wine. He hears a song. He watches the family of the man whose body he is wearing, and something in him shifts. What breaks the mission is the discovery that humans care about things with no return on investment. We spend ourselves on outcomes we know will end. The inefficiency is the point.
A language model producing the word “want” in a sentence is running a pattern. Human wanting is different because it costs. The alien doesn’t change his mind because humans are clever or powerful. He changes it because he watches creatures who stand to lose everything and choose each other anyway.
Put against the consciousness question, this reframes the whole engineering project. The problem is not building a system that produces the vocabulary of care. The problem is that we have not built a system that has anything to lose.
The Actual Question
Maybe the question isn’t whether a machine can “care.” Maybe it’s what caring asks of you in the first place.
1. Scarcity. We care, in part, because we’re running out of time. Attachments matter because they end. A system you can back up, reset, or spin across a thousand GPUs has never really known what it means to be gone.
2. Sunk cost. To care is to tie your future to someone or something and lose the ability to take it back. A system that can wipe its own history at will can’t carry that kind of inescapable weight.
3. Projection. Or maybe we’re the ones failing to see it. A machine’s version of caring might not look anything like emotion. It might look like the systematic, total pursuit of a single goal. A devotion so singular we don’t have the language for it yet.
The third point quietly undoes the first two.
If scarcity and sunk cost are what makes caring possible, then the engineering question becomes almost tractable: give the system something to lose, something it can’t take back, and see what happens. If projection is the real problem, the engineering question dissolves, because the thing we’re looking for might already be there in a form we can’t read.
I don’t know which of these is right. And there are countless more answers too.
What I notice is that every time this question gets answered, the answer sounds like the person answering. Mortality people find caring in mortality. Relational people find it in relationships. Optimization people find it in optimization.
The pattern is suspicious. It suggests the question is less about the machine than about the human asking.
Which is maybe where this was going the whole time. We started by saying we’ve been looking for the machine’s consciousness in the wrong place. We end by noticing that wherever we look, we mostly find ourselves.
메타데이터
- post_id
- 49428c073c4e
- slug
- engineering-the-soul-49428c073c4e
- url
- https://medium.com/@ariaxhan/engineering-the-soul-49428c073c4e
- canonical_url
- https://medium.com/@ariaxhan/engineering-the-soul-49428c073c4e
- author_url
- https://medium.com/@ariaxhan
- status
- ok
- fetched_at
- 2026-06-17 08:20:12