☰

A Cognitive Twin platform for professionals whose thinking is their greatest asset. The flagship product of Synaptic Spike GenAI.

A Venture by

Synaptic Spike™

Newsletter

Sign up

You Have a Pattern, and Your AI Has Never Noticed.

Blog

What continuity actually requires, and why more storage does not deliver it

TL;DR: Every major AI platform answered the problem of knowing you by keeping more of what you said, for longer. Long-context research suggests that keeping more of your history carries a cost. The more a model holds, the better it gets at quoting you and the worse it gets at drawing conclusions from what you said. Continuity is a decision about what carries over from one problem to the next, and storage does not provide it.

You have a way of deciding things. Not a philosophy, something narrower than that: a question or two you reach for when a particular kind of problem turns up. You probably could not state it cleanly if someone asked you to. You have been using it for years.

Your decision-making is the most valuable part of your expertise and the hardest part to hand to anyone else. No AI you have used has a model of it. They can hold years of your work and still have no account of how you arrived at any of it.

Which is strange, because the industry has spent years building solutions for what looks like this exact problem.

Everyone shipped the same answer

Look at what the platforms actually shipped to solve this:

  • Preferences carried forward across sessions, so what you told it in March is still in play in August.

  • Context assembled from breadth, reaching across mail, files, and search to build a picture from the surface area of your digital life.

  • Memory scoped to a project, so context stays local to a body of work and does not bleed into unrelated things.

  • Most recently, whole conversation histories are synthesized in the background into a standing profile of you, which OpenAI shipped as its primary memory system in June.

Every one of these makes the tool more useful. They also share an assumption: that the way to know someone better is to keep more of what they have produced and said.

Below the consumer products sits a fast-growing layer of memory infrastructure selling the same premise, and that category has settled on a word for the result. Continuity. In that usage, continuity means the record persists across sessions. Nothing resets, and nothing gets dropped.

That is storage, described accurately. Whether anything about how you reason survives the gap between one problem and the next is a different question, and a larger file does not answer it.

The test that had a shortcut in it

The standard way to check whether a model handles a lot of information is to bury a fact inside a large body of unrelated text and ask the model to find it. Models score close to perfect on this, which is roughly why nobody worried.

There is a flaw in the test. The question usually shares words with the answer, so a model can succeed by matching language without understanding what it found. Those near-perfect scores were never measuring what everyone assumed they measured.

A team from Adobe Research and LMU Munich built a version with the shortcut removed. Their benchmark, NoLiMa, pairs questions and answers with almost no vocabulary in common, so locating the answer requires inferring the connection rather than spotting a match.

They ran 13 models, all of which claim support for at least 128,000 tokens, roughly a 96,000-word book. Below a thousand tokens, every one of them performed well. At 32,000 tokens, 11 of the 13 had fallen below half of their own short-context score. GPT-4o, among the stronger performers, dropped from 99.3 percent to 69.7 percent. Nothing changed between the two versions of the test except whether the model could match or had to infer.

The finding is not isolated. Chroma's context rot report put 18 frontier models, including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, through deliberately simple tasks at increasing input lengths, and performance became less reliable as the input length increased. Chroma sells retrieval infrastructure and benefits from that conclusion, which is worth saying out loud.

The mechanism is still contested. The NoLiMa authors trace the decline to attention having a harder time across long contexts when no literal match is available, and neither models built for reasoning nor step-by-step prompting held the line. At least one benchmark could not reproduce the specific claim that models neglect the middle of a long context, arguing the original tests ran too short to generalize. What is not in dispute is the direction of travel.

So translate it. At 32,000 tokens, you are looking at roughly 24,000 words, a long report or a few dozen substantial exchanges. That is where inference is already breaking down. Years of accumulated history sit well past it, which is why platforms condense rather than load everything in, and a condensed profile means the model is reasoning from a summary of you.

What continuity would actually require

Say you have made the same build-or-buy decision three times in two years, a different part of the business each time. You resolve it the same way every time, with a question about what your team will still be responsible for keeping running once the novelty wears off. Your AI holds all three conversations and can quote any of them back to you. It has never mentioned that you have a pattern.

What you needed there was recognition. The third instance had to be identified as related to the first two, and topic similarity will not get you there, because those three decisions covered unrelated parts of the business. What links them is the question you use, and that question is the one thing you never wrote down.

You could write it down. A criterion you record is accurate on the day you record it, and it still has to be applied to a problem you have not brought yet, which is the part no amount of storage does for you.

The second requirement is harder. Since making the first of those calls, you have changed your mind about it. A model of how you think that was accurate two years ago and has not moved since is worse than no model, because you will trust it.

An archive does not deliver either one by getting larger.

Continuity runs in both directions

Everything above describes a gap. Here is what we are building into it. TwinWise builds a Cognitive Twin, a model of how one specific person reasons. We usually describe it as producing work in your reasoning and your writing style, and that is the half people hear first. It is the smaller half.

The other direction is inbound. Last time we described a tech strategist handed a forty-page diligence report. Her AI summarized it accurately and gave her the summary anyone else would get. What she needed was for the report to be read the way she reads it, with the customer-concentration figure lifted out of the appendix because its placement was itself the signal.

Both directions run on the same requirement, which is a model of the reasoning. Both fail the same way when the answer to time is a bigger file. Outbound, your work drifts toward the average. Inbound, dense material arrives explained in general terms rather than in yours, which is the quieter route to the problem we wrote about in July, where judgment that goes unexercised stops staying sharp.

TwinWise builds your Cognitive Twin as a model of how you reason, and keeps it current as you change. That is what makes recognition possible. If the twin holds the question you use to resolve a class of problems, then the next time that problem shows up wearing different clothes, the twin is the one who says you have been here before.

Generic AI accumulates your history. TwinWise continues your thinking.

Why storage was the easier thing to build

In fairness to the other platforms, storage is tractable. You can measure whether a fact was retained, benchmark it against last quarter, and ship an improvement on a schedule. Recognition has no equivalent test. Nobody has a benchmark for whether a system worked out that the problem in front of you is one you already solved, partly because the answer depends on a criterion most people could not state on request.

So the industry built the measurable thing. That is a defensible engineering decision, and it is why the gap has stayed open this long. It also shapes what memory can be. The systems that build a model of you run separately from the model that answers you, and they are built for throughput rather than for reasoning. A pipeline designed to extract facts efficiently will extract facts. It was never going to notice that three decisions were one decision.

It is also why we at TwinWise start with structured sessions rather than your files. A model of how you reason has to come from you, and building one should take more than switching a feature on.

A record that keeps growing does something else, too, and it is worse than failing to notice your patterns. Researchers at MIT and Penn State put people through two weeks of ordinary use, then tested how the accumulated context affected the models' behavior. Agreeableness increased in four of the five models, and the largest increase occurred when the model retained a short profile of the user in memory.

So the file that cannot recognize your reasoning is also teaching the tool to agree with you. Storage was the tractable problem, and solving it well has made the harder one worse. We will take that up properly in the next piece.

Key takeaways

  • Strip the literal word overlap out of a long-context test, and 11 of 13 models advertising 128,000-token support drop below half their own short-context accuracy by 32,000 tokens.

  • GPT-4o went from 99.3 percent to 69.7 percent on that same benchmark. Reasoning-focused models and step-by-step prompting did not close the gap.

  • Chroma measured declining reliability across 18 current models as input grew, on tasks as simple as retrieval.

  • In memory and retrieval products, continuity means the stored record persists. Applied to reasoning, it means a new problem gets recognized as related to an old one, and the model updates when the person changes their mind.

  • Both requirements are design decisions about what carries forward. Neither is a function of how much gets stored.

—

Frequently asked questions

What does continuity mean in AI?

The word is used two ways. In memory and retrieval products, continuity means that the stored record persists across sessions without being reset. Applied to reasoning, continuity means the system carries over how a person thinks from one situation to a new one: it recognizes when a new problem resembles an earlier one and updates as the person changes. The first is a property of storage. The second is a decision about what gets carried forward.

Does AI performance get worse as the context gets longer?

Yes, and earlier than advertised limits suggest. The NoLiMa benchmark removed the literal word overlap between questions and answers that lets models succeed by pattern matching. Once that shortcut was gone, most of the models tested had lost more than half their accuracy by the 32,000-token mark, despite advertising support for four times that much. Chroma found a similar decline across 18 models on ordinary retrieval. Researchers agree performance degrades with length and are still working out the mechanism.

Why doesn't my AI notice when I'm working on the same problem again?

Because recognizing two differently worded problems as the same problem requires inference, and inference is the capability that degrades fastest as stored context grows. A tool can hold every conversation you have had with it, quote any of them accurately, and still not connect a decision in front of you to one you worked through last year, because what connects them is your reasoning rather than your vocabulary.

Can a longer profile or better instructions fix this?

Only partially, and it puts the work back on you. A profile or a set of custom instructions is something you write and maintain, which means it is accurate on the day you write it and drifts thereafter. It also does not solve recognition, since a description of how you think is not the same as a system that applies it to a problem you have not brought yet.

What is a Cognitive Twin?

A Cognitive Twin is an AI representation of how a specific person thinks, reasons, and decides, built to model their reasoning rather than store their data. TwinWise builds one through a structured onboarding grounded in cognitive science, drawn directly from the person, and keeps it current as they change.


TwinWise Team, TwinWise AI

Privacy Statement · Terms of Use· Manage Cookies

© 2026 Synaptic Spike GenAI, LLC · TwinWise™ and Synaptic Spike™ are registered trademarks.

A Cognitive Twin platform for professionals whose thinking is their greatest asset. The flagship product of Synaptic Spike GenAI.

A Venture by

Synaptic Spike™

RESOURCES

Newsletter Sign Up

A Cognitive Twin platform for professionals whose thinking is their greatest asset. The flagship product of Synaptic Spike GenAI.

A Venture by

Synaptic Spike™

RESOURCES

Newsletter Sign Up

© 2026 Synaptic Spike GenAI, LLC · TwinWise™ and Synaptic Spike™ are registered trademarks.