Most AI agents have the same problem: they forget everything the moment a session ends. You can spend ten minutes explaining your preferences, your project context, your constraints — and the next conversation starts from zero. That’s not a model capability gap. It’s a missing memory layer.
Cognee is one of the more interesting solutions to this problem. It processes text, documents, and conversation history through a pipeline that extracts entities and relationships, builds a knowledge graph in Neo4j or NetworkX, and adds vector embeddings for semantic retrieval. The whole stack is open source (GitHub: topoteretes/cognee), self-hostable, with a managed cloud option starting at $35/month for the Developer tier.
But Cognee isn’t the only answer, and depending on what you’re building, it might not even be the right one. The memory layer landscape has matured significantly in 2026. Here’s how the main alternatives actually compare.
Why the Memory Layer Choice Is an Architectural Decision
Before the comparison: picking a memory layer isn’t like picking a logging library. The data structure it imposes — flat facts vs. temporal graphs vs. knowledge graphs — determines how your agent reasons about history. Switching later means migrating stored memory, not just swapping an API. Get this decision right early.
Cognee’s specific constraint worth knowing upfront: even in fully self-hosted deployments, embedding generation requires an OpenAI API key. Your data goes through OpenAI’s servers for vectorization. For teams with strict data residency requirements, that’s a real consideration that often gets missed until late in evaluation.
Mem0 — Simplest Path to Per-User Personalization
Mem0 (GitHub: mem0ai/mem0, ~48k stars, Apache-2.0) takes the opposite architectural bet from Cognee. Rather than building a knowledge graph, it abstracts memory into a clean CRUD interface: add a memory, search memories, delete memories. The underlying storage — vector DB, graph DB, key-value — is implementation detail the developer doesn’t need to manage.
This design choice has a major practical upside: Mem0 integrates with any agent framework in a few lines of code, and swapping it out later doesn’t require touching your agent logic. The boundary is clean.
Pricing: self-hosted is fully free. Managed cloud runs from a free Hobby tier (10k memory writes/month, 1k retrievals) up to Pro at $249/month, which unlocks graph memory, advanced retrieval modes, and SOC 2/HIPAA compliance. The $24M Series A in early 2026 and the AWS SDK partnership have accelerated enterprise feature development significantly.
One caveat on benchmarks: third-party comparisons show wildly different LoCoMo scores for Mem0 (ranging from 68.5 to 92.5 depending on configuration and version). Treat published benchmark numbers as directional signals, not gospel — run your own evaluation on a representative sample of your actual data.
Best for: Teams that need fast, framework-agnostic memory integration for conversational personalization. Especially good when you want to stay flexible about your underlying stack.
Letta (formerly MemGPT) — Agents That Manage Their Own Memory
Letta’s conceptual model is the most interesting of the bunch. It draws from operating system design: an agent has a fixed “in-context” memory block (like CPU registers — always present, always fast) and an archived memory layer (like disk — retrieved on demand). The key difference from other tools: the agent itself can edit its own core memory blocks through tool calls, deciding what to keep in its working memory and what to offload.
This makes Letta less of a memory library and more of a stateful agent runtime. You’re not just adding memory to an existing agent — you’re building agents that intrinsically know how to manage their own state. The tradeoff is the tighter coupling: adopting Letta means building around Letta’s agent primitives, not just plugging in a memory layer.
Open source under Apache-2.0 (GitHub: letta-ai/letta), with a managed cloud platform for teams that don’t want to run the server themselves.
Best for: Long-running autonomous agents where memory management is a first-class design concern, not an afterthought. If you’re building agents that need to operate continuously over days or weeks and self-organize what they remember, Letta’s model fits well. If you just need a chatbot to remember user preferences, it’s overkill.
Zep / Graphiti — Temporal Knowledge Graphs
Zep’s approach to memory is built around a specific insight: facts change over time, and most memory systems ignore this. A user’s job title, a contract status, an order state — these aren’t permanent facts. They’re facts with a validity period. Zep’s Graphiti engine builds a bi-temporal knowledge graph that tags every fact with when it became true and when it stopped being true.
This makes Zep uniquely good at questions like “what did we know about this user as of last Tuesday?” or “when did this customer’s subscription status change?” No other tool in this list handles temporal fact tracking at this level of precision.
The tradeoff: Graphiti requires Neo4j, which adds infrastructure complexity. In independent benchmark evaluations, Zep ranks near the top for temporal reasoning tasks — which is exactly what it was built for. For purely conversational personalization without time-sensitivity, the added complexity isn’t worth it.
Best for: Applications where facts change over time and you need to track those changes accurately. CRM systems, contract management, customer support with long account histories.
LangMem — The LangGraph Native Option
LangMem is the memory layer built specifically for LangChain’s LangGraph framework. It supports semantic, episodic, and procedural memory types — meaning it can remember facts, past events, and even how an agent should behave (not just what it knows).
If you’re already using LangGraph for orchestration, LangMem is the path of least resistance. The integration is native, the API is familiar, and you don’t need to manage cross-framework compatibility.
The honest limitation: LangMem’s benchmark numbers are middling, and retrieval latency in some configurations is high enough to make it impractical for real-time conversational agents. It works well for batch-mode or offline agents where latency isn’t a constraint.
Best for: LangGraph-first teams that want memory without adding new infrastructure dependencies. Not the right choice for latency-sensitive real-time interactions.
Graphiti Standalone
Worth a separate mention: Graphiti is also available as a standalone library independent of Zep’s managed platform. It’s the same bi-temporal knowledge graph engine, but you manage the Neo4j backend yourself. Good option if you want Zep’s temporal reasoning capabilities without the managed service dependency.
Comparison Table
| Tool | Architecture | Best Use Case | Self-Hosted | Key Constraint |
|---|---|---|---|---|
| Cognee | Knowledge graph pipeline | Multi-source heterogeneous data | Yes (needs OpenAI for embeddings) | OpenAI API dependency for embeddings |
| Mem0 | Abstracted CRUD interface | Fast per-user personalization | Yes, fully free | Graph memory requires paid tier |
| Letta | Stateful agent runtime | Long-running autonomous agents | Yes (Apache-2.0) | Tight coupling to Letta’s agent model |
| Zep/Graphiti | Bi-temporal knowledge graph | Time-sensitive fact tracking | Yes (needs Neo4j) | Neo4j infrastructure overhead |
| LangMem | LangGraph-native memory | LangGraph stacks, batch agents | Yes | High retrieval latency (~60s p95) |
How to Choose
The decision mostly comes down to what problem you’re actually solving:
- Need memory fast with minimal lock-in? Mem0. Clean API, huge community, framework-agnostic.
- Building agents that self-manage state over long runs? Letta. Design your agent around it from the start.
- Facts in your domain change over time and you need to track that? Zep/Graphiti. Nothing else handles bi-temporal tracking this well.
- Already deep in LangGraph? LangMem, unless you have real-time latency requirements.
- Ingesting documents, PDFs, and conversations into a unified knowledge base? Cognee’s pipeline architecture handles heterogeneous sources well — just plan for the OpenAI dependency.
One Cognee 1.0 note: the September 2026 release added significant improvements to pipeline configurability and reduced the Neo4j requirement for lighter deployments. If you evaluated Cognee pre-1.0 and dismissed it on complexity grounds, it’s worth a second look.



