Never losecontext.Never again.
Cencori Memory gives any AI app persistent memory. Two calls — recall() and remember() — give any model context across new chats, new devices, new providers, with your inference running wherever it already runs. Already on Cencori? It collapses to one flag. No vector store to build. No context to re-paste.
Works with any model
Memory for any LLM.
Not just ours.
recall() pulls what you know about the user; remember() extracts and stores the new facts. Wrap them around your own OpenAI, Anthropic, or local model call — your inference runs wherever it already does. Moving off Mem0 or Zep takes an afternoon.
// Works with any model — memory around your own model call.
const context = await cencori.memory.recall(userId, message);
const reply = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [
{ role: 'system', content: context },
{ role: 'user', content: message },
],
});
// Extract the durable facts from the exchange and store them.
await cencori.memory.remember(userId, {
user: message,
assistant: reply.choices[0].message.content,
});Already on Cencori
Or skip the wiring.
One flag.
If your inference already routes through Cencori, recall and remember fuse into a single field on the request. The gateway handles retrieval, injection, PII redaction, extraction, embedding, storage, and forget in-line. Something no other memory API can do — because none of them are the gateway.
// Without memory: every new chat starts from zero.
await cencori.chat.completions.create({
model: 'gpt-4o',
messages,
});// On the gateway: recall + remember collapse into one flag.
await cencori.chat.completions.create({
model: 'gpt-4o',
messages,
memory: { userId: session.user.id },
});The pipeline
Four steps.
One engine, either path.
Whether you call recall/remember yourself or hand the gateway one flag, the same pipeline runs underneath. p95 retrieval overhead: under 150ms.
Retrieve
The latest user message is embedded. pgvector query, scoped to (org, userId), returns top-K memories above the relevance threshold. Runs alongside the input guards, so the chat rarely waits on it.
Inject
Retrieved memories are formatted as a system block ahead of the user turn. The model reads them as prior context, not as a separate lookup.
Complete
Provider dispatch — OpenAI, Anthropic, Google, whoever. BYOK if set. Response streams back; retrieval overlaps the input pipeline so memory adds no separate round trip.
Extract & write
Async fact extraction on the completed exchange. PII redaction before writeback. Embedded and stored region-pinned. Audit entry emitted.
The new-chat moment
Session ends.
Context doesn’t.
Every AI product ships context-loss as an accepted default. Hit the context limit, start a new chat, re-paste your project context like a caveman. Cencori Memory fixes it out of the box.
- New chat — session-scope resets, user-scope carries forward.
- New device — same userId, same memory.
- New provider — memory lives above the model layer.
- MCP bridge — Cursor and Claude Desktop join in.
session #47 — earlier this week
─────────────────────────────────────────────
user Help me redo the pricing tiers on my app.
model [200 messages of back-and-forth. You land
on tiered Stripe billing, 14-day trial,
metered overages at $0.01/unit.]
session #48 — new chat, moments ago
─────────────────────────────────────────────
user Let's keep going.
model Picking up from where we left off with the
tiered Stripe billing — the metered overage
rate you set was $0.01/unit. Want to add
annual discounting next?Same user. Two sessions. Zero re-paste.
What’s in the box
Everything you’d otherwise
build yourself. Handled.
Scoped memory
session · user · workspace · org. Each scope has its own storage and lifecycle.
Region-pinned
Memories written in EU stay in EU. Enforced at write, not query. No accidental leakage.
PII redaction pre-write
Regulated identifiers detected and redacted before anything hits storage.
Immutable audit log
Every read and write recorded. Cryptographic timestamps. Exportable to your SIEM.
Hard-forget
Real deletion, not a soft flag. Every forget is recorded in the audit log.
Multi-provider
Memory lives above the provider layer. Swap OpenAI for Claude for Gemini — memory unchanged.
MCP bridge
Cursor, Claude Desktop, and any MCP client can read from your memory. Distribution moment.
React components
<Chat memory> for stateful chat; the useMemory hook for search, forget, and GDPR-export panels.
We ship the UI too
A memory-aware
UI, out of the box.
<Chat memory /> turns any surface stateful. The useMemory hook — list, search, forget, export — is everything you need to build a right-to-be-forgotten panel, no backend.
import { Chat, useMemory } from 'cencori/react';
// One flag, memory-aware chat.
<Chat model="gpt-4o" memory={{ userId }} />
// Build a "what do you remember about me" panel from the hook:
// list · search · forget(id) · exportAll (GDPR export).
const { memories, forget, exportAll } = useMemory({ userId });// Hard-delete a memory by id. Real deletion, audit-logged.
await cencori.memory.forget(memoryId);
// Surface stale, low-value memories to prune (candidates only).
const { suggestions } = await cencori.memory.forgetSuggestions({ userId });Pricing — the fill gauge
One bar. Zero to 100%.
You know what to do.
Same shape as Vercel bandwidth or Supabase storage. Free tier fills fast — that’s the upgrade signal. Reads keep working at 100%; only new writes block, with a clean 429 memory_quota_exceeded and an upgrade URL in the error payload.
Start free
The starter tank. Enough to prototype a real product and hit the demo bar.
- 1,000 memories per project
- session + user scope
- 30 / 90-day retention
