Short answer: Short-term memory is live context that dies with the current call or sitting. Long-term memory is durable storage that survives across sessions.
This split cuts across episodic, semantic, and other content types. Short-term holds what is needed now. Long-term holds what should still be available next week. Agents need both, and they need clear rules for what moves from one to the other.
Every category covered so far, working memory, episodic memory, semantic memory, procedural memory, reflective memory, describes a different kind of content. Underneath all of them sits a simpler, more structural split that cuts across every one of those categories rather than adding another type alongside them: short-term memory, meaning whatever exists inside the model’s live context right now, and long-term memory, meaning anything stored outside the model entirely, persisting independently of any single call. Understanding this split clearly matters because it explains why these two halves of memory have to be architected in completely different ways, not just sized differently.
What Structurally Separates Short-Term From Long-Term Memory?
The dividing line isn’t about content type at all. A fact can exist as short-term memory, sitting in the current context, or as long-term memory, stored externally waiting to be retrieved. The same is true of a past event, a learned procedure, or a reflective judgment about confidence. What actually separates the two categories is location and lifespan: is this piece of information present inside the one call happening right now, or does it exist independently, outside that call, available to be pulled back in later.
This is why the two terms describe an architectural boundary rather than a content classification. Short-term memory is defined entirely by presence in the current call. Long-term memory is defined entirely by persistence outside it. Everything else, what kind of content it happens to be, is a separate question layered on top of this more basic distinction.
Why Is Short-Term Memory Inherently Self-Limiting in a Way Long-Term Memory Isn’t?
Short-term memory has a hard capacity ceiling built into it, and every piece of information it holds gets paid for again on every single call it’s part of, whether that payment is measured in cost, latency, or the model’s own limited attention across a long input. None of this changes no matter how carefully the content inside it is chosen. The ceiling and the recurring cost are structural properties of being inside a single call, not something that improves with better content selection.
Long-term memory doesn’t share either constraint in the same way. A store of long-term memory can grow indefinitely without making any individual future call more expensive, as long as only what’s actually relevant gets pulled back into short-term memory on demand rather than the entire store being resent every time. This asymmetry, one side structurally bounded and expensive per use, the other side structurally unbounded and cheap to hold, is the real reason these two halves of memory need genuinely different architecture rather than just different amounts of the same kind of storage.
What Has to Happen for Something to Cross From Long-Term Into Short-Term Memory?
Nothing in long-term storage automatically shapes a response just by existing there. A retrieval step has to run, decide what’s actually relevant to the current request, and insert the result into that call’s context, at which point it functions exactly like any other piece of short-term memory for the duration of that one call. This crossing is a deliberate, triggered event rather than something continuous or automatic. If the retrieval step doesn’t run, or doesn’t find anything relevant, whatever’s sitting in long-term storage contributes nothing to that particular response, no matter how directly relevant it might have been.
This is the bridge moving in one direction: from persistent, external storage into the temporary, bounded space where the model can actually see and use it.
Does Information Ever Move in the Other Direction, From Short-Term Into Long-Term?
It does, and this direction is just as deliberate as the first. Something said or reasoned through within one call’s short-term context has to be actively identified as worth keeping and extracted into long-term storage before it can survive past that call. Nothing in short-term memory automatically becomes long-term memory simply by having existed there for a while; someone, or something, has to make the decision that it’s worth the promotion, exactly the write-control judgment already covered when deciding what deserves to become a memory in the first place.
Together, these two directions form a complete bridge: retrieval carries relevant long-term memory into a specific short-term context when it’s needed, and extraction carries worthwhile short-term content back into long-term storage so it isn’t lost once that particular call ends. Neither direction happens by default. Both require a deliberate step designed specifically to make that crossing happen.
How Does Weaviate Engram Implement This Two-Way Bridge Between Short-Term and Long-Term Memory?
Weaviate Engram is built around exactly these two deliberate crossings rather than treating memory as one undifferentiated pool. Consider a customer-loyalty-program concierge handling point-redemption requests, where something learned mid-conversation is worth keeping well past that one exchange. The write path carries content from the current short-term exchange into long-term storage:
from engram import EngramClient
client = EngramClient(api_key=os.environ["ENGRAM_API_KEY"])
client.memories.add(
[
{"role": "user", "content": "I always want to redeem points for travel, never merchandise, even if merchandise has better value per point."},
],
user_id="member-31820",
)
Weeks later, in a completely new conversation with no memory of the earlier one still lingering in any context window, the read path carries that same fact back from long-term storage into this new call’s short-term context:
preferences = client.memories.search(
query="How does this member prefer to redeem their loyalty points?",
user_id="member-31820",
)
Once retrieved, that preference behaves exactly like any other piece of short-term memory for the length of this new call, even though moments earlier it was sitting untouched in long-term storage. Nothing about the underlying fact changed by crossing that bridge. What changed is which side of the boundary it’s currently on, and that’s the entire relationship between short-term and long-term memory in practice: two structurally different spaces, connected by exactly two deliberate operations, each one doing the work of moving information across a boundary that never crosses itself on its own.
The short-term and long-term split describes where memory lives and how long it lasts, but cognitive science offers another, older distinction worth borrowing: memory that can be stated outright versus memory that only shows up in behavior, never as an explicit fact anyone could simply say out loud. Our next chapter, What is the difference between declarative and non-declarative memory?, takes up exactly that distinction.