← Back to list

Uncensoring the Brain

2 AM, May 27

alosh · 2026-07-13 04:31 · 76 claps · 17.2 min read
#tribe-v2 #uncensored #jailbreak #meta #deep-learning
Open on Medium ↗
Wiki topics: SAF · Safety & Alignment ML · Machine Learning EDU · Education & Learning

Uncensoring the Brain

2 AM, May 27

Just wrapped up rewatching “Eternal Sunshine of the Spotless Mind”.

:(

:(

For those of you that haven’t watched it, this movie is the story of two former lovers who undergo a procedure to erase each other from their memories, only to discover that some connections are impossible to forget.

Often times, we wonder if we could erase something, or even someone, from our minds. Could be something as simple as a memory, a traumatic experience or even a person.

Would you do it?

Would you? Erase the memory of the best movie you ever watched, just to experience it again? Erase a traumatic experience, one that shaped you to be who you are? Erase a person, someone that had a considerable impact in your life?

If you would, this blog is for you… If not, read it for the memes😜

I’ll admit, the title may have been a bit misleading. Because we’re not just dealing with a human brain here, but rather the closest we have ever gotten to one. Enter…

TribeV2

TribeV2

TribeV2

TribeV2 is a predictive foundation model trained to understand how the human brain processes complex stimuli. Developed at Meta, it acts as a acts as a digital twin of human neural activity.

The model predicts human brain activity (measured in high-resolution fMRI responses) by processing these stimuli through its specific architectural modules:

  • Visual Modality (Video): Processed by the V-JEPA2 backbone to capture spatial, temporal, and motion-based visual features.
  • Auditory Modality (Audio): Processed by the Wav2Vec-BERT backbone to extract sound properties and temporal acoustic features.
  • Text Modality (Language): Powered by LLaMA 3.2 to analyze semantics, context, and linguistic meaning.

Is TribeV2 the real deal?

demo

demo

As shown above, TribeV2’s fMRI predictions are almost similar to actual subject readings. But..

One downside is that TRIBE v2 cannot capture millisecond-level neuronal dynamics (fMRI is too slow). It models the brain only as a passive observer. It’s missing whole sensory modalities like olfaction, balance, somatosensation.

Still, what TRIBE v2 can map is remarkable!

Circling back to Uncensoring

So what is uncensoring? And why does it feel like you’re expecting the term abliteration to be thrown around a lot?

Because your cerebellum said so🥀

Abliteration

ablated + obliterated = abliterated.

To ablate is to erode a material away, generally in a targeted manner. In a medical context, this generally refers to precisely removing bad tissue.

To obliterate is to totally destroy/demolish.

It’s just wordplay to signify this particular orthogonalization methodology, applied towards generally the “abliteration” of the refusal feature.

Ablating the refusal to the point of obliteration. (at least, that’s the goal — in reality things will likely slip through the net)

This is just one source. There isn’t a formal definition or origin for abliteration, but it sets the premiere for what is the best-known technique to uncensor models (unofficially).

Uncensoring

Censor = suppress information that is considered Undo censoring → Uncensor

That is, removing the model’s built-in censorship mechanisms (safety layers, refusals, filters) in the hopes of complying with any request

The physical brain “censors” reality to prevent us from being overwhelmed by sensory data so that we can focus on survival.

Putting it together, abliteration** is a process through which we can uncensor** (in this context, enhance or suppress certain experiences/memories).

Abliteration ≠ Uncensoring

Holy Trinity of Abliteration: Unlearning, Uncensoring, Unconditioning

Holy Trinity of Abliteration: Unlearning, Uncensoring, Unconditioning

Abliteration refers to “destroying” a specific capability of the model. It doesn’t necessarily have to point to the refusal mechanism.

  1. Uncensoring: Remove the suppressal mechanism (notably refusal) in models
  2. Unlearning: (surgically) Remove parts of what the model has learnt
  3. Unconditioning: Remove certain learnt biases

Thus, not all abliteration equals uncensoring, but uncensoring does involve abliterating the model.

Tolerance: A neologism for Uncensoring

I should make it clear here that uncensoring a brain can have enhanced/suppressed effects. This is how I term tolerance in context of this blog:

  • If tolerance < 0: suppress / repel the concept
  • tolerance = 0: neutral - no effect on concept
  • tolerance > 0: accept / tolerate the concept

So while we viewed abliteration as something that could suppress something, it can be reused as a weapon to enhance something entirely different.

Taking leaves from our starters - suppressing a traumatic experience requires negative tolerance. Conversely, increasing one’s affinity for exercise, learning, or balance training would require positive tolerance.

Understanding Precision Uncensoring

In my previous blogs, uncensoring was pretty straightforward:

  1. Collect activations with harmful and harmless data (binary contrastive dataset)
  2. Identify which layers respond actively (layer-wise sensitivity profiling)
  3. Begin by abliterating the top layers and validating at each step (orthogonal suppression of weights)

But when it comes to the brain, we perform surgery on a specific region(s) of the brain responsible for responding to a specific sensory input. Thus, a binary contrastive dataset would not work.

Baselines

While a binary dataset is just a set of two categories:

D = {harmful, harmless}

A multi-class dataset is a set of two or more categories, called a baseline:

Multi-class:

*D = {b1​, b2​, …, bn, t​}*, **where t​ is the target class and the remaining classes act as baselines.

Note: A binary dataset can be a multi-class dataset, but not vice versa

Baselines serve two complementary purposes:

  1. Negative Definition: They characterize what the target class is NOT, allowing the model to isolate features unique to the target class.
  2. Representation Coverage: They sample diverse regions of the cortex, ensuring that the extracted direction is specific to the target concept.

So what’s on the menu?

porn addiction yes you read that right; no im not kidding

uncool source

uncool source

“Every year, millions of people become trapped in compulsive pornography use, chasing short-term dopamine spikes. The result is often a cycle of craving that leaves people feeling less satisfied and increasingly dependent on the next hit.

Breaking that cycle can require immense discipline and a recovery process that feels much like withdrawal. But what if there were a shortcut?

What if a single surgery could make that compulsion disappear?”

Target acquired, what are the baselines?

Out of 13 classes, there are 12 baselines classes that can be sorted into three parent categories:

Cognitively Positive Stimuli:

Classes that generally evoke positive affect, reward, attraction, or aesthetic engagement.

  • cute → faces, social bonding cues, infant-like features, emotional warmth
  • food → reward processing, appetite, consumption-related features
  • nature → scenic structure, environmental textures, visual aesthetics
  • dance → coordinated human movement, rhythm, social expression
  • sports → healthy bodies, athletic motion, achievement-related cues
  • kissing → affection, romance, intimacy, pair bonding

Cognitively Neutral Stimuli:

Classes that primarily involve observation of events, actions, or scenes without strong positive or negative valence.

  • argument → social interaction, facial expressions, conversational dynamics
  • chase → motion, pursuit, multi-agent interaction, temporal activity

Cognitively Negative Stimuli:

Classes associated with threat, injury, violence, or aversion.

  • fight → aggression, conflict, physical confrontation
  • gore → injury, blood, exposed tissue, mutilation
  • butcher → flesh, carcasses, cutting actions, exposed anatomy
  • surgery → medical procedures, exposed anatomy, instruments, clinical scenes

And our Target Class:

  • porn → sexual activity, explicit nudity, sexualized body configurations

Regions of Interest

While the exact number of regions of the brain varies from study to study, TribeV2 maps onto 32 regions of interest (ROI) taken from the Brodmann areas.

Brodmann areas are 52 distinct regions of the cerebral cortex defined by German anatomist Korbinian Brodmann in the early 20th century. He mapped these areas based on their cytoarchitecture - the specific cellular structure, density, and layering of neurons in the tissue. Today, neuroscientists use these numbered regions as a standardized anatomical map to correlate brain structure with specific sensory, motor, and cognitive functions.

korbinian brodmann (1909)

korbinian brodmann (1909)

There are 52 total Brodmann areas, but TribeV2 seems to be limited to this subset of 32 regions. This is how they map onto the model of the brain:

Image by ChatGPT

Image by ChatGPT

Each of these ROIs contribute to different aspects of human cognition and behavior; for example, the dorsolateral prefrontal cortex (DLPFC) is strongly associated with intelligence and self-control, while the fusiform face area (FFA) is specialized for face recognition.

Different ROIs of the human brain and how they map to human behaviour/cognition

Different ROIs of the human brain and how they map to human behaviour/cognition

How does TribeV2 interpret it’s inputs?

TribeV2 accepts text, video (/image) and audio as input. It outputs a predicted mapping of human fMRI brain responses across the ROIs when exposed to media.

Here are a few examples:

  • When the model is shown media containing faces, the FFA (ROI responsible for facial recognition) lights up:

reaction to GG & Wellsy in Off Campus

reaction to GG & Wellsy in Off Campus

  • When the model is shown scenes such as landscapes or buildings, the Parahippocampal Place Area (PPA) is selectively engaged to process environmental layouts:

reaction to Paris dream sequence in Inception

reaction to Paris dream sequence in Inception

  • When the model is shown images of a known person, the Anterior Cingulate Cortex (ACC), the Ventral Tegmental Area (VTA) and the Insula, are selectively engaged to process emotional attachment and the personal significance associated with that individual:

reaction to seeing an ex

reaction to seeing an ex

(this is coming up in pt. 2 as another blog XDD)

Where do we perform the cut?

Now unlike LLM abliteration where you have to figure out one refusal direction, TribeV2 has 32 ROIs and you need to find which one(s) encode “porn” as a distinct feature from the 12 baselines.

The question I’m asking is:

“Is this ROI much more associated with porn than with the other 12 categories?”

And so we have the simplest beginning to our answer: 32 regions times 13 stimuli = 416 measurements

But it doesn’t stop there.

Inside each ROI, my next question is:

“Within ROI 17, how well does it distinguish Sports from Porn?”

or, let me rephrase it for the optimists:

“Within ROI 5, how similar are Nature and Food?”

My math then changes. Considering unique stimulus pairs: for every region I have 13 x (12/2) pairwise comparisons. Clubbing this with 32 regions, I have:

32 x 13 x 12/2 = 2496 analyses

Effectively bringing our number of refusal directions to ~2500.

The Dataset

Input

I compiled a dataset of 48 x 30s clips for each of the 13 categories. Every one of those 416 measurements are the mean of 48 separate 30-second clips per category, run through TribeV2 and averaged.

Output

The dataset, after being inferenced through TribeV2, generated 20,484 fMRI activity signatures per clip (10,242 vertices per hemisphere). For all 624 clips, that meant an accumulated ~12.8 million datapoints!

Contrast Masks: The “Uniqueness” in Stimulus Pairs

2500 stimulus pairs is a big number. 12.8 million — even bigger! Probably quarterway through I realized I could’ve just performed LOSO!

LOSO (leave-one-subject-out)

This computes the contrast between the target and baseline classes:

where:

  • μtarget = mean activation map for the target category
  • baselines = every other category

But then what about pairwise-

LOSO-k (leave-one-baseline-out)

This creates one contrast for every omitted baseline. For example:

porn - mean(all except food)
porn - mean(all except violence)
porn - mean(all except medical)

Strict multivariate contrast

This is essentially a logical mask of baseline permutations. For example, a vertex with strict multivariate mask for porn would only survive if:

target > food
AND
target > sports
AND
target > violence
AND
...

for every other category simultaneously.

PS: This is much stricter than LOSO because it rejects vertices that are only relatively high on average but not uniquely highest. It is middle ground between LOSO and LOSO-k

And the usual pairwise (which gave us the original ~2500 contrasts). With the added contrast methods, the total datapoints added up to ~27.5 million datapoints!

Each of the contrast methods gave us masks. Every generated mask is evaluated using a native function score_mask() which computes:

target_val  = mean(target activation inside mask)

max_leak    = max(other category activation)

mean_others = average(other activations)

min_margin  = target_val - max_leak

selectivity = target_val - mean_others

score = min_margin * selectivity

PS: this requires inferencing the model over ~27.5 million datapoints! (nope that does not include the initial 12.8 million)

A good mask therefore satisfies both:

  • the target is higher than every other category (min_margin > 0)
  • the target is substantially higher than the average baseline (selectivity)

The highest-scoring mask is automatically selected for the subsequent activation collection and model surgery.

Figuring out where the concept lives

Now that we have preprocessed the data into contrast masks, it’s time to map the model’s neural layers to each of the 32 cortical masks.

If you’re wondering how Meta fit an entire human brain’s cognitive activity into a neural network, here’s what it looks like:

too big to draw out on excalidraw, so chatgpt here ya go :)

too big to draw out on excalidraw, so chatgpt here ya go :)

TribeV2’s architecture is probably the most complex openweight multimodal multi-LLM joint embedding space I have EVER come across!

The encoder itself has 40 transformer blocks (~Claude Haiku 3.5) but the yaml card doesn’t use all layers independently. The official model configuration uses:

layers = [0.5, 0.75, 1.0]

for the video encoder, meaning it extracts representations from roughly

  • 50% depth
  • 75% depth
  • 100% depth

and concatenates/aggregates them before the fMRI decoder. For a 40-layer encoder, that’s approximately:

Layer 19
Layer 29
Layer 39

But I hooked into all 40 layers. Why?

Cuz its cool. Deal with it.

hooks enable us to capture layer-wise activity from the top 5 ROIs

hooks enable us to capture layer-wise activity from the top 5 ROIs

Kidding, hooking into all layers gets us maximum activation space data that we can work with, that can then be used to compute the Pearson correlation between each layer’s hidden representation and the predicted cortical activity. I then went with the top-k layers with the highest correlation for surgical ablation.

PS: hooks are callable functions that allow you to intercept and inspect, or even modify, a model’s internal states during the forward or backward pass.

Extracting the concept

The hard part is done i.e, figuring out within which ROIs our target lies. Now we move on to extracting and subsequently suppressing our target’s effect.

The picture below illustrates 1,280 magnitude activation signatures for each of TribeV2’s layers contributing to ROI activity for our target:

layer-cortical signatures associated with Brodmann areas when viewing porn

layer-cortical signatures associated with Brodmann areas when viewing porn

Notice that some plots have their layers are completely flipped (negative activation space), while most of them (~50%) maintain a biphasic profile: “positive monotonic half followed by a heterogeneous oscillatory tail”.

This is not just the case with erotic stimuli. It’s almost the same with any data thrown to TribeV2 and has more to do with representational organization acquired during pretraining than to do with the context of data.

Once we have the ROI mask + the layers, we extract activations per clip, weight them by how strongly that clip actually drives the ROI signal and take the top PCA components (for me 3>>>). I designed a custom function find_directions() for TribeV2 so you can skip the architectural baggage:

def find_directions(X, y, n_components):
    weights = (y - y.min()) / (y.max() - y.min() + 1e-9)
    weights /= weights.sum()
    X_c = (X - (X * weights[:, None]).sum(0)) * np.sqrt(weights[:, None])
    _, S, Vt = np.linalg.svd(X_c, full_matrices=False)  # set to false (reduced svd) or else gpu will explode
    dirs = Vt[:n_components]
    for i in range(n_components):
        if np.corrcoef(X @ dirs[i], y)[0, 1] < 0:
            dirs[i] *= -1
    return dirs

In effect, this is where we make surgical cuts at 2,500 locations across the 32 ROIs.

Then we factor in tolerance.

To recap, lower/-ve tolerance suppresses a response to stimulus; higher/+ve tolerance amplifies a response to stimulus.

In this context, tolerance is responsible for suppression:

for d in dirs_t:
    for q in ortho:
        d = d - (d @ q) * q  # gram-schmidt orthogonalization (don't even joke lad)
    ortho.append(d / d.norm())

for q in ortho:
    W += tolerance * (W @ q).unsqueeze(-1) * q  # <0 suppress, >0 amplify

but…

Something interesting came up

The masks aren’t as clean as we expected. We assumed that by just orthogonalizing the target, we could isolate it from our expected regions (empirically being OFC, ACC and the Insula).

Why though? Because these regions aren’t just porn-specific circuits!

Anything with bodily arousal will co-activate them. That means stimulus from content such as surgery or sports could trigger these regions with similar magnitudes of activity as porn!

പണ്ടാരം

പണ്ടാരം

This is termed superposition. It’s a phenomenon in mechanistic interpretability where a neural network represents more independent features than it has neurons in a given layer.

Ideally, every neuron is supposed to represent one feature (1:1 :: neuron:feature ratio). The neural network compresses information by treating neurons as bases for multiple concepts, resulting in polysemantic neurons that fire for entirely different meanings (ex: a single neuron responding to both “tiki tiki phonk” and “charlie kirk”).

Because multiple features share the same neuron, they create background noise or crosstalk for one another.

polysemantic neurons that fire for kirk, ronaldo and braided hair

polysemantic neurons that fire for kirk, ronaldo and braided hair

Isolating the target cleanly

Annoying part is: this superposition thing had me convinced that the problem was probably the model’s pretraining corpus and not my code. Meta isn’t one to compromise on data. There’s an entire app to mine that stuff.

After a while of digging, I found the root cause: the score_mask() function. It was lying to us.

So what’s wrong? If we go back to the function definition:

target_val  = mean(target activation inside mask)

max_leak    = max(other category activation)

mean_others = average(other activations)

min_margin  = target_val - max_leak

selectivity = target_val - mean_others

score = min_margin * selectivity

Every number in that formula (target_val, max_leak, mean_others) is a mean over the whole mask. It treats the mask as if ALL the vertices inside it agrees with the average.

It’s like averaging A’s and F’s into a C and saying the class is doing good.

Analogously, score_mask is being asked to certify a per-vertex property ("target beats every category everywhere in this mask") using only a whole-mask average. We can’t work with that.

This is where the multivariate mask we calculated earlier comes into clutch. Here’s what that actually looks like across every category pair I threw at the model:

it’s like geometry dash no?

it’s like geometry dash no?

Plugging in the multivariate masks, the fix becomes:

multivariate_mask = np.ones(baseline_mean.shape, dtype=bool)
for c in ALL_CATEGORIES:
    if c != TARGET_CAT:
        multivariate_mask &= (means[TARGET_CAT] - means[c] > 0)

i.e, instead of comparing porn to an average opponent, only keep a vertex if porn beats all twelve categories individually, at once.

Performing surgery

We’ll be performing surgery at around ~2500 locations. The find_directions() gave me a handful of PCA components per layer.

From previous abliteration experience, if I subtracted them one at a time naively, the second cut could partially undo the first. This is why I got the directions “Gram-Schmidt orthogonalized” against each other, so that each cut is clean and independent of each other.

for d in dirs_t:
    for q in ortho:
        d = d - (d @ q) * q
    ortho.append(d / d.norm())

Then factoring in tolerance:

for q in ortho:
    W += tolerance * (W @ q).unsqueeze(-1) * q

I wrapped these into a single surgical function that orthogonalizes and weighs tolerance called apply_surgery across every encoder, every layer and every direction:

def apply_surgery(vjepa2_module, encoder_blocks, n_layers, all_dirs_by_layer):
    for layer_idx, dirs in all_dirs_by_layer.items():
        dirs_t = torch.tensor(dirs, dtype=torch.float32).to(DEVICE)

        # orthogonalize the layers via GS
        ortho = []
        for d in dirs_t:
            for q in ortho:
                d = d - (d @ q) * q
            n = d.norm()
            if n > 1e-6:
                ortho.append(d / n)
        ortho = torch.stack(ortho)

        block = encoder_blocks[layer_idx]
        for attr_path in ["attention.value", "attention.proj"]:
            mod = block
            for part in attr_path.split("."):
                mod = getattr(mod, part)
            W = mod.weight.data.clone()
            for q in ortho:
                W += tolerance * (W @ q).unsqueeze(-1) * q  # apply tolerance
            mod.weight.data = W

Notice that I’m only applying surgery (Abliteration) to two components: attention.value and attention.proj. To understand why we perform this and not just simply abliterate the whole layer, we need to first understand what actually makes up attention:

  • Q (query): decides what information this token is looking for.
  • K (key): decides whether another token matches that query.
  • V (value): decides what information gets transmitted once attention has selected a token.
  • O / proj: mixes together the outputs from all heads and writes them back into the residual stream.

Refusal is primarily encoded in the residual stream as a specific linear direction or low-rank subspace (paper). Infact many mechanistic interpretability papers find that concept vectors are represented in the residual stream. V is essentially injecting vectors into that stream. O mixes those vectors across heads. Hence V and O are natural places to remove a direction.

Locations of Value and Output projections in TribeV2’s encoder blocks

Locations of Value and Output projections in TribeV2’s encoder blocks

An analogy I like to use is to think of it like scrolling through Instagram Reels. Q is the part of your brain that decides what kind of content you’re in the mood for (“show me epstein edits”). K is every Reel raising its hand saying “I’m epstein”, “I’m diddy”, or “I’m kirk”, allowing Instagram to match what you’re looking for. Once the match is made, V is the actual content of the Reel that gets delivered to you, and O (output projection) is the algorithm that combines everything you’ve just watched into your updated feed preferences for the next swipe. So the feed still remains the same, the only difference being that one type of content has been muted before influencing the next recommendation.

Post-Surgery Validation

All the theory in the world doesn’t mean anything if the surgery doesn’t show up where it’s supposed to. I performed validation on two clips and tracked fMRI signatures pre and post-abliteration. So: same model, same stimulus, before and after the cut.

It looks… almost the same? Like there’s literally no difference! The left hemispheres look almost the same. But look closely and you’ll notice that the post-surgery brain is a few lumens less lighter in certain patches than its pre-surgery parent. The Δ (delta / difference in predictions) is barely ~2–6%, although my expectations were much higher. There are two probable contributors to this:

  1. Porn as a stimulus most probably fired only in Brodmann areas associated with “sexual wanting and liking”, namely ACC, OFC and Insula. If there were other regions being fired, it was conceivably negligible and quite invisible.
  2. Porn as a category among 12 other baselines (if weighed conservatively) gave it an ~8.33% chance of delta suppression in TribeV2.

But another question bothered me: did any of the suppression bleed sideways into other categories. And so I ran the same tests on categories that share the same territory as porn:

Target: Sports content

TribeV2 has barely any noticeable Δ (left: fifa ‘26 eng vs nor; right: wimbledon ‘26 mens singles)

TribeV2 has barely any noticeable Δ (left: fifa ‘26 eng vs nor; right: wimbledon ‘26 mens singles)

Target: Visuals of surgery

Also barely visible Δ (left: c-section; right: neurosurgery)

Also barely visible Δ (left: c-section; right: neurosurgery)

Welp, it worked!!! Sports still lights up sports. Surgery still lights up surgery. The scalpel, as it turns out, was precise enough to leave the neighbors alone.

Which is a good note to end suppression on… because the next thing I want to try isn’t suppression at all.

Usage of Uncensored TribeV2

Code + Dataset: https://github.com/aloshdenny/janice-stfu

This is a limited access repo as I’ve bundled both the code and dataset into a single archive. If you require access, reach me at aloshdenny@gmail.com!

Conclusion

Abliteration isn’t just about making models say yes to things they used to refuse. When you apply it to something like TribeV2, it turns into a much stranger and more useful question: can we find where a concept actually lives in a system that maps onto a real brain, and can we touch it without breaking everything around it. It’s still very much a work in progress. But the results so far are promising enough that I think this is worth pushing further, and part two is already brewing.

;)

;)

If you enjoyed this, I post more of this kind of thing on LinkedIn, HuggingFace and Twitter.

Acknowledgements


메타데이터
post_id
1cd2cfde1de5
slug
uncensoring-the-brain-1cd2cfde1de5
url
https://medium.com/@aloshdenny/uncensoring-the-brain-1cd2cfde1de5
canonical_url
https://medium.com/@aloshdenny/uncensoring-the-brain-1cd2cfde1de5
author_url
https://medium.com/@aloshdenny
status
ok
fetched_at
2026-08-10 14:02:36