The Dual-Engine Pattern: MongoDB + Redis as a Unified Runtime for AI Agents
Every AI agent tutorial shows you one database. The real world needs two. We learned this building StreakAI — Polystreak's production AI infrastructure agent that handles vector search, semantic caching, multi-layer rate limiting, session management, and LLM orchestration. It runs on MongoDB Atlas and Redis Cloud simultaneously, and neither one could do the other's job.
We call it the Dual-Engine Pattern: Redis is the speed engine — sub-millisecond reads, atomic counters, in-memory vector search, ephemeral session state. MongoDB is the intelligence engine — durable knowledge, session archives, metrics observability, long-term memory. Together they form a single AI agent runtime where every component is placed in the database optimized for that specific workload.
MongoDB stores what your AI agent has learned. Redis serves what your AI agent needs right now. Together, they're the complete runtime.
1. The Architecture — Who Does What
Here's the exact split we run in production with StreakAI. This isn't theoretical — these are live Redis keys and MongoDB collections running right now.
| Workload | Engine | Why | Data Structure |
|---|---|---|---|
| Live session state | Redis | Sub-ms reads, sliding TTL, atomic appends | RedisJSON — session:{id} |
| Message counter per session | Redis | Atomic INCR, no race conditions | JSON numIncrBy on $.message_count |
| Semantic cache | Redis | KNN vector similarity in <1ms | RedisJSON + HNSW index — cache:{hash} |
| Knowledge base vectors | Redis | HNSW vector search, top-K in 2-5ms | RedisJSON + FT index — kb:{slug}:{n} |
| Rate limiting (5 layers) | Redis | Atomic INCR + EXPIRE, zero-lock | String counters — ratelimit:* |
| Global token budget | Redis | Atomic INCRBY across all sessions | String counter — ratelimit:global_tokens:{day} |
| Token usage per session | Redis → MongoDB | Accumulate in Redis, persist on eviction | JSON $.token_usage → MongoDB field |
| Session archives | MongoDB | Durable, queryable, indexed by date | agent_sessions_archive collection |
| Per-request metrics | MongoDB | Time-series analytics, cost tracking | agenti_ai_metrics collection |
| Blog/knowledge content | MongoDB → Redis | Ingested, chunked, embedded, indexed in Redis | Source docs → HNSW vectors |
The rule is simple: if the data needs to be read in under 1ms during an inference call, it lives in Redis. If the data needs to survive a server restart and be queried for analytics, it lives in MongoDB. Some data — like session state — lives in Redis during its active life, then graduates to MongoDB when the session ends.
2. Session Lifecycle — Redis Hot, MongoDB Cold
This is the pattern that changed how we think about agent state. A session isn't stored in one place — it migrates between engines based on its lifecycle stage.
Phase 1: Birth (Redis)
When a user starts a conversation, we create a RedisJSON document with a 30-minute sliding TTL. Every interaction resets the TTL. The session stores messages, timestamps, message count, and cumulative token usage — all in a single JSON document that Redis can read in 0.1ms.
// Redis key: session:{uuid}
{
session_id: "a1b2c3d4-...",
created_at: 1712345678,
last_active: 1712345890,
message_count: 3,
messages: [
{ role: "user", content: "What is semantic caching?", timestamp: 1712345700 },
{ role: "assistant", content: "Semantic caching stores...", timestamp: 1712345702 },
// ...
],
token_usage: { prompt: 2840, completion: 312, total: 3152 }
}Why RedisJSON and not a plain hash? Because we need atomic array appends (arrAppend on $.messages), atomic counter increments (numIncrBy on $.message_count), and sub-path reads (get only $.messages[-8:] for the LLM context window). RedisJSON gives us document-database semantics at in-memory speed.
Phase 2: Active Life (Redis, with sliding TTL)
Each message triggers four atomic Redis operations in sequence: append the message to $.messages, update $.last_active, increment $.message_count, and reset the TTL. All four complete in under 2ms total. The session stays hot in Redis as long as the user is active, with the 30-minute TTL acting as an automatic garbage collector for abandoned sessions.
// Atomic session update — 4 ops, <2ms total
await redis.json.arrAppend(key, "$.messages", message);
await redis.json.set(key, "$.last_active", timestamp);
await redis.json.numIncrBy(key, "$.message_count", 1);
await redis.expire(key, 1800); // reset 30-min TTLPhase 3: Eviction (Redis → MongoDB)
When the session hits the message limit (5 messages in StreakAI), we execute the eviction pipeline: read the full session from Redis, persist it to MongoDB's agent_sessions_archive collection with proper date conversions and token usage data, then delete the Redis key. The user sees 'Session archived to MongoDB Atlas, evicted from Redis' — because that's literally what happens.
// Eviction pipeline — Redis → MongoDB → Redis DEL
const session = await redis.json.get(`session:${id}`);
await mongodb.collection("agent_sessions_archive").updateOne(
{ session_id: id },
{ $set: {
...session,
evicted_at: new Date(),
token_usage: {
prompt_tokens: session.token_usage.prompt,
completion_tokens: session.token_usage.completion,
total_tokens: session.token_usage.total,
}
}},
{ upsert: true }
);
await redis.del(`session:${id}`); // free memoryThis is the Dual-Engine lifecycle: Redis handles the real-time state (fast reads/writes during conversation), MongoDB handles the historical state (durable archives for analytics and observability). Neither database does both jobs well. Redis would waste memory on cold sessions. MongoDB would add latency to hot reads.
3. Five-Layer Rate Limiting — All Redis, All Atomic
Rate limiting for AI agents isn't one counter — it's a wall of counters at different granularities. We run five layers in StreakAI, all backed by Redis atomic operations. Every check completes before any expensive work (embedding, vector search, LLM call) happens.
| Layer | Redis Key Pattern | Limit | TTL | Purpose |
|---|---|---|---|---|
| Session messages | Checked against session JSON | 5 per session | Session lifetime | Prevent unbounded session growth |
| Hourly messages | ratelimit:msgs:{sessionId}:{hourBucket} | 30 per hour | 3600s | Throttle burst usage |
| Daily sessions per IP | ratelimit:sessions:{ip}:{dayBucket} | 10 per day | 86400s | Limit session creation |
| Daily messages per IP | ratelimit:ip_msgs:{ip}:{dayBucket} | 50 per day | 86400s | Prevent token waste from scrapers |
| Global daily token budget | ratelimit:global_tokens:{dayBucket} | 500,000 tokens | 86400s | Hard cost ceiling across all users |
The pattern for each counter is identical: Redis INCR + conditional EXPIRE. If the INCR returns 1 (first hit), we set the TTL. If the value exceeds the limit, we reject with a 429 before touching the LLM. The global token budget uses INCRBY to add actual token counts after each LLM response.
// Hourly rate limit — INCR + EXPIRE
const hourBucket = Math.floor(Date.now() / 3600000);
const key = `ratelimit:msgs:${sessionId}:${hourBucket}`;
const current = await redis.incr(key);
if (current === 1) await redis.expire(key, 3600);
if (current > 30) return { allowed: false, reason: "Rate limit exceeded" };
// Global token budget — INCRBY after LLM response
const dayBucket = Math.floor(Date.now() / 86400000);
const tokenKey = `ratelimit:global_tokens:${dayBucket}`;
await redis.incrBy(tokenKey, promptTokens + completionTokens);Why Redis and not MongoDB for rate limiting? Because INCR is O(1), atomic, lock-free, and returns in under 0.1ms. A MongoDB findOneAndUpdate with $inc on the same counter takes 2-5ms and requires a write concern round trip. At 5 checks per request × 1,000 requests per second, that's the difference between 0.5ms total overhead and 10-25ms. The rate limiter must be invisible to the user — and Redis makes it invisible.
4. Semantic Caching — The $0.00 LLM Call
Every LLM call costs money and adds latency. Semantic caching eliminates both for repeated or similar questions. StreakAI caches every LLM response in Redis with its query embedding, then checks for KNN similarity before calling the LLM on the next request.
The flow: embed the user's query → KNN search the cache index for the nearest cached query → if similarity ≥ 0.80, return the cached response (0ms LLM latency, $0.00 cost). If no hit, call the LLM, then store the response + embedding in Redis with a 24-hour TTL.
// Check semantic cache — KNN 1 on HNSW index
const result = await redis.ft.search(
"idx:semantic_cache",
"*=>[KNN 1 @query_embedding $query_vec AS similarity]",
{ PARAMS: { query_vec: queryBuffer }, DIALECT: 2 }
);
const similarity = 1 - parseFloat(result.documents[0].value.similarity);
if (similarity >= 0.80) {
// Cache hit — skip LLM entirely
return cached.response; // 0ms LLM latency, $0.00
}
// Cache miss — call LLM, then store
const llmResponse = await streamFromLLM(query);
await redis.json.set(`cache:${hash}`, "$", {
query, query_embedding: embedding, response: llmResponse,
model: "deepseek.v3.2", created_at: timestamp
});
await redis.expire(`cache:${hash}`, 86400); // 24h TTLThe cache index is a Redis HNSW vector index with 1024-dimensional Amazon Titan embeddings. A KNN 1 search completes in under 1ms. The cache stores both the embedding (for similarity matching) and the full response (as RedisJSON for efficient retrieval). At a cache hit rate of 20-30% on a typical agent, this saves hundreds of dollars per month in LLM API costs.
The fastest LLM call is the one you don't make. Semantic caching in Redis turns repeated questions into sub-millisecond lookups at zero cost.
5. Context Retrieval — Redis Vectors, MongoDB Knowledge
The RAG pipeline in StreakAI follows a clear dual-engine split. The source content (blogs, docs, case studies) is authored and stored as JSON files — the system of record. During ingestion, each document is chunked (word-window splitting), embedded (Amazon Titan, 1024 dimensions), and indexed into Redis as HNSW vectors. At query time, the agent never touches the source files — it hits Redis exclusively.
| Stage | Engine | Operation | Latency |
|---|---|---|---|
| Ingest: chunk + embed | AWS Bedrock → Redis | Store chunks as RedisJSON with HNSW-indexed embeddings | Batch, offline |
| Query: embed user question | AWS Bedrock | Embed query with Amazon Titan | 50-80ms |
| Query: vector search | Redis | FT.SEARCH with KNN top-5 on HNSW index | 2-5ms |
| Query: build context | Redis results → LLM | Format top-K chunks into <context> prompt | <1ms |
| Archive: persist knowledge | MongoDB | Store original docs, chunk metadata, embeddings for audit | Async |
# Redis HNSW vector index for knowledge base
FT.CREATE idx:knowledge_base ON JSON
PREFIX 1 "kb:"
SCHEMA
$.embedding AS embedding VECTOR HNSW 6
TYPE FLOAT32 DIM 1024 DISTANCE_METRIC COSINE
$.title AS title TEXT
$.source AS source TAG
$.url AS url TEXT
$.text AS text TEXT
# Query: find top-5 similar chunks
FT.SEARCH idx:knowledge_base
"*=>[KNN 5 @embedding $query_vec AS score]"
PARAMS 2 query_vec <binary_embedding>
RETURN 5 text title source url score
SORTBY score ASC
LIMIT 0 5
DIALECT 2The context window is capped at the last 8 messages from the Redis session. This bounds the token cost per LLM call while keeping the conversation coherent. The context prompt includes chunk titles, source types, and URLs — so the LLM can cite its sources with clickable links in the response.
6. Token Accounting — Track, Cap, Persist
Token usage flows through both engines. During a session, Redis accumulates token counts in real-time using JSON numIncrBy operations. The global daily budget uses a Redis atomic counter. When the session is evicted, the cumulative token usage is persisted to MongoDB alongside the full conversation archive.
- Per-request: After each LLM call, promptTokens and completionTokens are recorded in the request metrics (MongoDB) and added to the session's cumulative $.token_usage (Redis).
- Per-session: The Redis session JSON accumulates total prompt, completion, and combined token counts. On eviction, this is written to MongoDB as token_usage.prompt_tokens, token_usage.completion_tokens, token_usage.total_tokens.
- Global: The ratelimit:global_tokens:{dayBucket} Redis counter is incremented by the total tokens of each request. When it crosses 500,000, all users see 'daily limit reached' until the bucket expires at midnight.
- Observability: MongoDB's agenti_ai_metrics collection stores per-request metrics including estimated cost (DeepSeek pricing: $0.0014/1K input, $0.0028/1K output), letting us build dashboards on spend per session, per user, per day.
The key insight: use Redis for real-time token math (atomic counters are fast and correct under concurrency), and MongoDB for historical token analytics (aggregation pipelines, time-series queries, cost dashboards). Trying to do real-time accounting in MongoDB adds latency. Trying to build dashboards from Redis keys is painful.
7. The Redis Key Map — A Living Blueprint
Here's every Redis key pattern in StreakAI's production instance, with data types and TTL policies. This is the real-time engine's footprint.
| Key Pattern | Type | TTL | Purpose |
|---|---|---|---|
| session:{uuid} | RedisJSON | 1800s (sliding) | Active conversation state, messages, token_usage |
| cache:{md5hash} | RedisJSON | 86400s (fixed) | Semantic cache — query, embedding, LLM response |
| kb:{slug}:{chunk_n} | RedisJSON | None | Knowledge base chunks with HNSW-indexed embeddings |
| ratelimit:msgs:{sessionId}:{hourBucket} | String (counter) | 3600s | Hourly message rate limit per session |
| ratelimit:sessions:{ip}:{dayBucket} | String (counter) | 86400s | Daily session creation limit per IP |
| ratelimit:ip_msgs:{ip}:{dayBucket} | String (counter) | 86400s | Daily message limit per IP |
| ratelimit:global_tokens:{dayBucket} | String (counter) | 86400s | Global daily token budget across all users |
8. The MongoDB Collection Map — Durable Intelligence
And here's MongoDB's footprint — the durable store that survives restarts, enables analytics, and feeds observability.
| Collection | Purpose | Key Fields |
|---|---|---|
| agent_sessions_archive | Evicted sessions with full conversation history | session_id, messages[], token_usage, evicted_at, total_user_messages |
| agenti_ai_metrics | Per-request pipeline metrics and cost tracking | session_id, promptTokens, completionTokens, estimatedCost, cacheHit, vectorSearchMs |
| knowledge_base_meta | Source document metadata for ingestion auditing | slug, title, chunk_count, embedded_at, dimensions |
9. Why Not Just One Database?
We hear this question every week. 'Can't MongoDB do vector search now? Can't Redis persist data?' Yes to both — but being capable isn't the same as being optimal.
- MongoDB-only: Atlas Vector Search adds 5-15ms per similarity query (vs 1-3ms in Redis HNSW). Rate limiting with findOneAndUpdate adds 2-5ms per check × 5 checks = 10-25ms overhead. Session reads from disk-backed storage add 1-3ms. Total added latency per request: 20-50ms. Feels sluggish for a chat agent.
- Redis-only: No aggregation pipelines for analytics. No durable storage without AOF/RDB complexity. No rich query language for 'show me all sessions from last week with >1000 tokens'. Session archives would consume expensive memory indefinitely. Token analytics would require external tooling.
- Dual-Engine: Redis handles the 6-8 operations per request that must be fast (session load, cache check, vector search, rate limit checks, session update, cache store). MongoDB handles the 1-2 operations that must be durable (metrics insert, session archive on eviction). Total request overhead from the fast path: <10ms. Durability and analytics: complete.
A Dual-Engine architecture isn't complexity for complexity's sake. It's putting each operation in the database that can execute it 10x faster. The result is an AI agent that feels instant and an analytics layer that misses nothing.
10. The Request Pipeline — Both Engines, One Flow
Here's the exact execution order for a single user message in StreakAI. Every step is annotated with which engine it touches and why.
- 1. Load session from Redis (0.1ms) — RedisJSON GET. Need instant access to conversation history and message count.
- 2. Check 5 rate limits in Redis (0.5ms total) — Five INCR/GET operations. Must reject before any expensive work.
- 3. Embed the query via AWS Bedrock (50-80ms) — Amazon Titan, 1024 dimensions. This is the expensive step we protect with rate limiting.
- 4. Check semantic cache in Redis (0.5-1ms) — KNN 1 on HNSW index. If hit (similarity ≥ 0.80), skip steps 5-6 entirely.
- 5. Vector search knowledge base in Redis (2-5ms) — KNN 5 on HNSW index. Retrieve top-5 relevant chunks with titles and URLs.
- 6. Stream LLM response from DeepSeek via Bedrock (500-3000ms) — The only step that takes real time. Everything else is optimized around it.
- 7. Store in semantic cache in Redis (1-2ms) — RedisJSON SET with embedding + response for future cache hits.
- 8. Update session in Redis (1-2ms) — Append messages, increment counter, accumulate token usage, reset TTL.
- 9. If session limit reached: Evict to MongoDB (5-10ms) — Read full session, persist to agent_sessions_archive, DEL from Redis.
- 10. Log metrics to MongoDB (fire-and-forget) — Insert per-request metrics for observability dashboards.
Redis touches: steps 1, 2, 4, 5, 7, 8, 9 (7 of 10 steps). MongoDB touches: steps 9, 10 (2 of 10 steps, both non-blocking or end-of-session). This is the Dual-Engine Pattern in action — Redis dominates the hot path, MongoDB handles the cold path.
11. Getting Started — Build the Dual-Engine Stack
If you're building an AI agent today, here's the minimum viable Dual-Engine setup:
- Week 1 — Redis: Set up RedisJSON for session state with sliding TTL. Add a semantic cache with HNSW vector index. Implement INCR-based rate limiting. Index your knowledge base as HNSW vectors. This gives you a fast, protected agent.
- Week 2 — MongoDB: Add session archiving on eviction. Store per-request metrics with timestamps. Build a knowledge base metadata collection for ingestion auditing. This gives you observability and durability.
- Week 3 — Connect them: Implement the eviction pipeline (Redis → MongoDB on session end). Add token accumulation in Redis, persist on eviction. Set up global token budgets in Redis, report spending from MongoDB. This gives you the full Dual-Engine runtime.
- Week 4 — Optimize: Add source-based filtering to vector search (@source TAG filters). Tune cache similarity threshold (we use 0.80). Add context window limits (last 8 messages to LLM). Build cost dashboards from MongoDB metrics. This gives you a production-grade system.
The total Redis memory footprint for StreakAI with ~200 knowledge chunks, semantic cache, active sessions, and rate limit counters is under 50MB. The MongoDB storage for session archives and metrics grows at roughly 2KB per session and 500 bytes per request — trivial at any scale.
12. The Bottom Line
The Dual-Engine Pattern isn't a compromise — it's a deliberate architectural decision. MongoDB and Redis aren't competing for the same job. They're collaborating on different jobs within the same AI agent runtime.
- Redis is the speed engine: sessions, cache, vectors, counters, real-time state. Anything the agent reads during inference.
- MongoDB is the intelligence engine: archives, metrics, knowledge metadata, analytics. Anything the team reads after inference.
- The eviction bridge connects them: sessions graduate from Redis (hot) to MongoDB (cold). Token budgets accumulate in Redis (fast) and report to MongoDB (durable). Knowledge is ingested into MongoDB (source of truth) and indexed in Redis (search engine).
We run this exact architecture in StreakAI — and you can try it live right now on our agent page. Ask it about semantic caching, then ask again to see the Redis cache hit. Watch the X-Ray panel to see latency numbers from both engines in real-time. That's not a demo — that's the Dual-Engine Pattern in production.
Stop choosing between MongoDB and Redis. Use both. That's not complexity — it's engineering each layer for what it does best. MongoDB + Redis = the complete AI agent runtime.