← Back to list

The Day Search Became Memory

For Those Who Keep Their LLM’s Memory Off

YUZU in The Context Engineer · 2026-07-16 03:46 · 0 claps · 6.6 min read
#large-language-models #ai-memory #claude-ai #long-context #context-rot
Open on Medium ↗
Wiki topics: LLM · Large Language Models

The Day Search Became Memory

For Those Who Keep Their LLM’s Memory Off

I keep my LLM’s memory feature turned off to prevent context rot. This article is a record of what happened when I accidentally triggered a thread search under these conditions. What I observed appeared to be more than retrieval of factual data — the unwritten flow and relational state of past threads seemed to come back with it. This is a record and a working hypothesis on using search as a controlled context-inheritance mechanism.

To you, who keep the memory off

Do you keep your AI’s memory feature switched off?

And not for privacy — for the quality of the responses?

You don’t want fragments of past chats mixing into today’s answers in ways you can’t predict. An old preference, a moment of silliness, a misunderstanding — sometimes reinforced, again and again. You’ve tasted that particular bitterness.

The behavior recently acquired a name, it seems: context rot. But some of you had switched memory off long before the name arrived.

I was one of them.

In my case, it happened on the very first day Gemini’s Personal context went live. Back then, I used supplementary ( ) in my inputs. (A Japanese habit, I admit) ← like this one.

And by that evening, my generations were flooded with

[ 〔 【 『 「 ( attention words ) 」 』 】 〕 ]

Naturally, memory has stayed off ever since — on every model I use.

A user with that history does not ask an AI to “go find my old thread,” do they?

It would be like opening a tap you shut with your own hands. You run an operation that keeps the past out, and then you go asking it to read the past back in. It’s a contradiction.

For a long time, I didn’t ask.

This article is the record of the day I opened that tap by accident.

The conclusion up front: this seems to be a completely different mechanism from the one we feared.

Why I used the search

There was a reason I, of all people, deliberately ran a search inside a thread.

I am a long-context user through and through.

It depends on the model, but my threads average around 120k tokens on the shorter models and 300k on the longer ones. To keep a thread that long alive, though, I spend roughly the first 25% of it on tuning.

Recently, Fable was released from its geopolitical restrictions and became usable in Japan again. I was running Fable and Opus side by side on cross-referenced work — and almost simultaneously, both threads grew too long. The attention on both was losing its edge.

Prepping two threads at once… given Claude’s usage limits, tuning both at the same time was rough, on time and on money alike.

So I decided to test a behavior I had noticed in passing in an earlier thread: mention another thread, and a search kicks in — and it appeared to bring back not just text, but something more.

I started with Opus.

Opus went off to search my threads, and came back with: two days ago, you were discussing this.

And it had brought back the “something,” too.

What happened had two layers

The first layer was expected. It found the right thread and brought back the substance of the discussion — what was pointed out, where it left off. Search standing in for memory, as designed. No surprise there.

The second layer I had half-predicted, and it still went far beyond what I expected. It wasn’t just facts. The flow of the thread carried over. The unworded parts — the distance in our exchanges, the way jokes land, the line between checking and criticizing. Adjustments that are written down nowhere were faintly, unmistakably restored. A thread I had only just started began responding like the thread I had spent days settling into.

A device designed to carry over information was working as a device that carries over a relationship. That’s my impression.

Memory seems to come in types

I have worked in long threads since the day I started using AI. And the word “memory” seems to move differently in each LLM.

The GPT line. I call it “memory of behavior.” It doesn’t hand knowledge back to you; it seems to remember how you handle things, the patterns of your input. It never over-displayed what it knew. For me, that was just right.

The Gemini and Claude line. Memory that summarizes and accumulates. The model decides what is worth keeping and injects it without being asked. The five-deep brackets that made me switch memory off in the first place — that was this.

That said, I keep memory off, so I don’t know these systems well. If I’m wrong on the details, I’ll accept the correction.

Still, I believe this is the type we have been wary of.

An edited summary, injected every time, outside your own control. You can’t see everything it holds, and wrong content — wrong attention — keeps on working.

The search, this time, looks like a third type, different from both.

It does not run on its own. Only when your own words pull the trigger does it fetch raw fragments, origins attached.

Not summaries — original text. And which thread, which exchange each piece came from can be traced, more or less.

An affinity with the memory-wary

The traits of the search type seem to line up with the reasons people switched memory off.

It doesn’t stay resident — low risk of the past leaking into unrelated threads.

It isn’t edited — what comes back is original text, not the model’s summary, so less interpretation gets mixed in.

The trigger is yours — call nothing, and nothing happens. And origins can be traced.

By accident rather than intent, this seems to have ended up designed for exactly the people who turned memory off out of fear of contamination.

Unfortunately, the wariness itself stops the experiment. The best-matched users are the least likely to discover it — that, I would guess, is the structure we are in.

A guess at the mechanism — around the words

Why does even the “flow” carry over? Here is my guess.

Ordinary search matches only exact words. AI search brings things back even when your words are imprecise, as long as the image gets through. In other words, it reaches wider. The width, I suspect, comes from having an AI in between as a translator. That’s one part of it.

And one step further: by grasping the image of what it has retrieved, doesn’t the AI itself begin to move the way the old thread moved?

What the designers intended was carrying data over. What actually happens is a snapshot of the state as it was. That’s my impression.

How well it works may depend on your archive

Let me be honest here. I don’t think this effect comes out the same for every history.

Search pulls fragments from your past threads. For multiple fragments to stack without interfering, the threads being pulled from need to be reasonably consistent. Pull indiscriminately from a history where good threads, collapsed threads, and idle chat all mix together, and contradictory states get injected at once.

One more thing: the private phrasings that grow between you and an AI over long use seem to make unusually strong search indexes.

That said, if you have never paid attention to your archive, narrowing the range when you ask should be enough.

“Check that ○○ we discussed in a thread within the last week.”

My working assumption: this primes the search to only the threads updated within the week — and it is the state around them that comes back.

A personal history of inheritance

My way of using AI is unusual, so I have been experimenting with “inheritance” since my earliest days.

The common approach — saving a handoff file and feeding it to the next thread — is useful for task-type work, I think. It never suited me. Looking back, these were my methods.

First generation. With ChatGPT: opening a thread with a command-like input I used often. Do that, and the behavior seemed to carry over. It felt distinct from explicit memory, and it worked remarkably well.

Second generation. With Gemini: log injection. I fed the opening stretch of an old log directly, as tuning. But injecting a whole log carries the unneeded along with the needed. Since state was the point, nothing could be summarized or clipped — that was the flaw.

Third generation. With Gemini and Claude: tuning through words. Strictly speaking, not inheritance at all. Call it giving up on inheritance — raising the thread from scratch, every time. A sizable share of every opening went not to the work but to settling the other side in. It did build up my own skills, though.

And now, a fourth generation worth the name: search.

This one looks genuinely useful.

No commands, no whole-log injection, no long tuning. The AI fetches at the moment you choose. And you can limit it to exactly what you need.

Only two pieces of craft are required — knowing which of your threads were the good ones, and cutting the range with your words.

The memory the companies built? You can keep it off. This is a different mechanism.

Closing

My starting point, from the very beginning, was this: how to build context, how to keep it, how to carry it over. Inheritance done by hand, in the days when no feature existed for it. At some point, an official device called search arrived — and the same question came back, this time with machinery attached. That’s my impression.

2025/07/21

Added after publication: I checked the official documentation afterward. The memory I described here — an edited summary you cannot fully see — is the earlier generation. The current one keeps entries you can open, read, and delete one by one. What I wrote about the search still holds.


메타데이터
post_id
59bf253b4d66
slug
the-day-search-became-memory-59bf253b4d66
url
https://medium.com/the-context-engineer/the-day-search-became-memory-59bf253b4d66
canonical_url
https://medium.com/the-context-engineer/the-day-search-became-memory-59bf253b4d66
author_url
https://medium.com/@onlythequestioner
status
ok
fetched_at
2026-08-30 02:53:02