Thought Leaders

Your Company’s Memory Should Outlast Its AI Models

mm
Add Unite.AI to your preferred sources on Google

People often ask which model powers my company. The answer matters: it sets quality, cost and limits, and I am happy to spend an hour on it. I also want to know a second thing: what the organisation will still have when that model changes.

Put it as a date. If your provider doubled its price on a Tuesday, or retired the endpoint you built on, what would you still own on Wednesday morning?

For a lot of companies, the honest answer is: an API key, an invoice, and a very long conversation history sitting on someone else’s servers under someone else’s retention schedule.

My background is chemistry. I spent years building grey-box models for chemical plants, wired into SCADA, lab results, and the operator’s own notes. That work teaches one rule fast: a number without its conditions is not a result. If someone swapped a sensor last week and nobody wrote it down, the reading on the screen explains nothing.

That is a lab notebook problem. It arrives now as a procurement problem, and it is rarely on the sheet where the AI budget is decided.

Three Questions That Get Answered as One

A larger context window lets a model use more information inside a single request. Durable storage determines what survives between requests. Portability determines whether that information remains usable when the provider changes. These are separate architectural questions, and the million-token window encouraged everyone to treat them as one.

Google’s own documentation offers the analogy directly: “An analogy for the context window is short-term memory.” Take that literally. Short-term memory is the thing you lose.

So every vendor selling an enormous window also sells something separate to persist state. Read those layers narrowly, in their own words, and note what each one actually covers.

OpenAI stores Response objects for 30 days by default; Conversation objects and their items are exempt from that TTL. Amazon’s 30-day deletion applies to the Bedrock Session Management API, and covers that API alone. Anthropic’s client-side memory tool illustrates one useful boundary: the model requests memory operations, the application controls the storage, while the same company also ships managed memory stores for its hosted agents.

All of these are ordinary engineering decisions, published openly. They also mean the retention rules for your company’s working context live in someone else’s release notes, and you should be able to name which rule applies to which of your data.

Where the Switching Cost Actually Lives

Swapping one text-generation call for another is a weekend. That is why “we’re multi-model” is such a cheap thing to say. The expensive layers are the ones nobody demos.

Embeddings. Changing the embedding model typically requires re-embedding the corpus and migrating or rebuilding the index. A change of generation model carries no such requirement, and the two get confused constantly — usually by whoever is promising the migration will be quick. The 2025 literature describes the standard path as “re-encoding the entire corpus and rebuilding the Approximate Nearest Neighbor (ANN) index, leading to significant operational disruption and computational cost”; that paper’s own contribution, Drift-Adapter, is a method for deferring the rebuild by learning a transform between embedding spaces. Either way, the migration cost covers retrieval validation and operational work as well as the embedding calls themselves.

Fine-tuned behaviour. One sentence from Cohere’s September 2025 deprecation notice belongs above every procurement desk: “Previously fine-tuned models will no longer be accessible.” Behaviour you paid to create, retired alongside the endpoint that hosted it. Everybody followed the published process. Your model still went away.

Prompt and tool behaviour. The same instructions produce a different application on a different model. In one enterprise application, researchers reported the regression pass rate falling from 100% on the original model to 97.3% on a newer one with identical prompts, recovered only after deliberate prompt redesign. Test scenarios appear in production at their own frequencies, so that figure supports no estimate of how often live workflows fail. The operational rule it supports is simpler: validate every model upgrade at the application level, yourself, before it reaches a customer.

All of this reaches the invoice eventually, long after the pricing pages were compared.

Somebody Else Is Holding Your Calendar

Deprecation runs on a published schedule, and the schedules are not yours. OpenAI gave a year’s notice before shutting down the Assistants API in August 2026 – generous by the standards of the same page, where GPT-4.5 Preview got about three months. Mistral’s policy is six months for GA models and one month for preview and third-party ones, with a warning that a rolling alias “may expose you to silent updates in model behavior and pricing.”

AWS says the quiet part out loud in its lifecycle policy: “Migrate to an Active model before the EOL date; migration will not happen automatically.”

Your roadmap has a co-author. It never attends your planning meetings and does not care what quarter you are in.

Price Is a Moving Part Too

At Gemini 3.7 Flash’s listed Standard rates, five billion uncached input tokens and one billion billed output tokens cost $7,500 a month. The rates currently scheduled for 1 January 2027 would take that token bill to $15,000, before other charges or discounts, while your product does exactly what it did before.

A frozen price list can also produce a bigger bill. Anthropic’s pricing documentation notes that the tokenizer used by its newer model generations “produces approximately 30% more tokens for the same text” – content-dependent and specific to those generations- which makes the effect on any given invoice something you have to measure. So measure the tokens you are billed for. A rate card on its own will not tell you.

For a product that runs workflows, the useful economic measure is the cost of a successfully completed task, including retries and human review. That number moves when a model changes even if the rate card does not.

And none of it is negotiable from a position where leaving takes a year. Capgemini surveyed 1,300 executives at billion-dollar organisations in spring 2026 about critical technology suppliers generally: 36% said moving off one would take more than twelve months, and one in ten had no viable alternative at all. That is a description of a supplier relationship with no second source.

Take It to Procurement

I run a French company, registered in Marseille, hosted in Europe, and I will still skip the sovereignty speech in the abstract. Independence from every technology supplier is out of reach for a normal company. The same Capgemini study found 59% of executives consider full digital sovereignty unrealistic, and I agree with them.

Understanding and controlling your critical dependencies is a smaller job, and a possible one. Three questions, all answerable in a meeting: where is our data stored and processed, who can access it, and what would it take to keep operating with another provider?

Every company knows how to ask that about logistics, payments, and cloud storage. Asked about working context, European dependency arrives in the meeting as a second-source discussion with a number attached, which is a form procurement can act on.

One caution on the numbers people quote here. That an enterprise uses several models tells us little about whether it can move its working context between providers. That requires a separate test: can the records, the permissions and the unfinished work be carried across and still be used?

What Actually Makes a Model Swappable

A shared interface across models is useful, and I would build one. It should make model-specific capabilities and dependencies visible. Portability requires no pretence that every model behaves the same way, and an abstraction that flattens the differences quietly throws away the reason you paid for a good model.

The heavier work is owning the state no provider should be the sole custodian of. Four things, in the order they break.

An authoritative record under your control. Messages, retrieved sources, approvals, execution checkpoints. Record what was completed in external systems and what can safely be retried: the email may have gone out while the confirmation failed to save, and a recovery procedure that does not know this will send it twice. Any provider state you cannot reconstruct is a dependency. Write it down as one, explicitly, before an incident does the documenting for you.

The source alongside every vector. A vector is a compiled artefact; the chunk is the source code. Store the source document and its version, document ID, chunk boundaries, chunker version, embedding model with version and dimensions, index version, and what processing produced the text. A stored text fragment alone will not reconstruct a table, an image or an OCR result. Keeping all of it turns a re-index into a planned migration, which is as much comfort as I would promise.

Records that distinguish source, inference, and decision. Memory has to separate what a source said, what a model inferred, and what a person approved. A client promising to pay on Thursday, a model’s guess that the payment will slip, and an agreed extension to Friday are three different records with three different statuses, and the system must know which one it may act on. Each needs its origin, scope, access rules, and current status, including whether it has been corrected or superseded. A transcript may well contain the evidence; replaying it does not reliably identify which conclusions are still current and which decisions still hold.

Portability you have actually tested. Make it testable two ways: an export-and-restore exercise, and an evaluation suite defining what the application must accomplish. Then “we could switch” becomes a measurement somebody can run: can the replacement continue representative workflows inside your requirements for quality, permissions, latency and cost? And preserve the training data you are entitled to reuse, together with its versions, tuning configuration and acceptance tests. Those make retraining possible, with no guarantee of identical behaviour, and the fine-tuned model ID expires on a date set by someone you have never met.

Then use hosted sessions, prompt caches and provider retrieval wherever they help. They are often excellent, and refusing them on principle is its own kind of expensive. The requirement is that the records they hold can be recovered and used elsewhere with their access and retention rules intact. Rent the computation; keep control of the record.

I would not oversell the architecture either. Before product-market fit, shipping six weeks earlier often beats any abstraction layer, and a startup that builds three retrieval backends before it has customers has built a museum. Whether that trade is right depends on the product and the price of the eventual rebuild. Keep the raw material recoverable and a shortcut stays a shortcut.

The Part That Changed

While AI only answered questions, losing context was an inconvenience. You retyped the prompt and moved on. Now it books the meeting, files the ticket and touches the CRM, and missing context leads to duplicated or half-finished work in systems that belong to your customers.

The model vendors build extraordinary things and I use them daily. The responsibility I am describing sits with those of us building the platforms in between: making the dependencies visible, and preserving a reliable record of what was authorised and what was completed.

Every completed workflow should leave the organisation with something it can use again — a verified result, a decision with a traceable history, a clearer starting point for the next task. That accumulated value should outlast the model that helped create it.

Ask what you would still own on Wednesday morning.

Ilia Razvin is the founder and CEO of IOSYA, a business productivity and AI workflow automation platform. He previously built grey-box process models and decision-support systems for chemical manufacturing, and holds a Master 2 in microsensors and detection systems from Aix-Marseille Université.