All posts
Redis CloudMongoDB AtlasAI AgentsVector SearchSession ManagementRate LimitingSemantic CachingArchitecture

The Dual-Engine Pattern: MongoDB + Redis as a Unified Runtime for AI Agents

Polystreak Team2026-04-0518 min read

Every AI agent tutorial shows you one database. The real world needs two. We learned this building StreakAI — Polystreak's production AI infrastructure agent that handles vector search, semantic caching, multi-layer rate limiting, session management, and LLM orchestration. It runs on MongoDB Atlas and Redis Cloud simultaneously, and neither one could do the other's job.

We call it the Dual-Engine Pattern: Redis is the speed engine — sub-millisecond reads, atomic counters, in-memory vector search, ephemeral session state. MongoDB is the intelligence engine — durable knowledge, session archives, metrics observability, long-term memory. Together they form a single AI agent runtime where every component is placed in the database optimized for that specific workload.

MongoDB stores what your AI agent has learned. Redis serves what your AI agent needs right now. Together, they're the complete runtime.

1. The Architecture — Who Does What

Here's the exact split we run in production with StreakAI. This isn't theoretical — these are live Redis keys and MongoDB collections running right now.

WorkloadEngineWhyData Structure
Live session stateRedisSub-ms reads, sliding TTL, atomic appendsRedisJSON — session:{id}
Message counter per sessionRedisAtomic INCR, no race conditionsJSON numIncrBy on $.message_count
Semantic cacheRedisKNN vector similarity in <1msRedisJSON + HNSW index — cache:{hash}
Knowledge base vectorsRedisHNSW vector search, top-K in 2-5msRedisJSON + FT index — kb:{slug}:{n}
Rate limiting (5 layers)RedisAtomic INCR + EXPIRE, zero-lockString counters — ratelimit:*
Global token budgetRedisAtomic INCRBY across all sessionsString counter — ratelimit:global_tokens:{day}
Token usage per sessionRedis → MongoDBAccumulate in Redis, persist on evictionJSON $.token_usage → MongoDB field
Session archivesMongoDBDurable, queryable, indexed by dateagent_sessions_archive collection
Per-request metricsMongoDBTime-series analytics, cost trackingagenti_ai_metrics collection
Blog/knowledge contentMongoDB → RedisIngested, chunked, embedded, indexed in RedisSource docs → HNSW vectors

The rule is simple: if the data needs to be read in under 1ms during an inference call, it lives in Redis. If the data needs to survive a server restart and be queried for analytics, it lives in MongoDB. Some data — like session state — lives in Redis during its active life, then graduates to MongoDB when the session ends.

2. Session Lifecycle — Redis Hot, MongoDB Cold

This is the pattern that changed how we think about agent state. A session isn't stored in one place — it migrates between engines based on its lifecycle stage.

Phase 1: Birth (Redis)

When a user starts a conversation, we create a RedisJSON document with a 30-minute sliding TTL. Every interaction resets the TTL. The session stores messages, timestamps, message count, and cumulative token usage — all in a single JSON document that Redis can read in 0.1ms.

typescript
// Redis key: session:{uuid}
{
  session_id: "a1b2c3d4-...",
  created_at: 1712345678,
  last_active: 1712345890,
  message_count: 3,
  messages: [
    { role: "user", content: "What is semantic caching?", timestamp: 1712345700 },
    { role: "assistant", content: "Semantic caching stores...", timestamp: 1712345702 },
    // ...
  ],
  token_usage: { prompt: 2840, completion: 312, total: 3152 }
}

Why RedisJSON and not a plain hash? Because we need atomic array appends (arrAppend on $.messages), atomic counter increments (numIncrBy on $.message_count), and sub-path reads (get only $.messages[-8:] for the LLM context window). RedisJSON gives us document-database semantics at in-memory speed.

Phase 2: Active Life (Redis, with sliding TTL)

Each message triggers four atomic Redis operations in sequence: append the message to $.messages, update $.last_active, increment $.message_count, and reset the TTL. All four complete in under 2ms total. The session stays hot in Redis as long as the user is active, with the 30-minute TTL acting as an automatic garbage collector for abandoned sessions.

typescript
// Atomic session update — 4 ops, <2ms total
await redis.json.arrAppend(key, "$.messages", message);
await redis.json.set(key, "$.last_active", timestamp);
await redis.json.numIncrBy(key, "$.message_count", 1);
await redis.expire(key, 1800); // reset 30-min TTL

Phase 3: Eviction (Redis → MongoDB)

When the session hits the message limit (5 messages in StreakAI), we execute the eviction pipeline: read the full session from Redis, persist it to MongoDB's agent_sessions_archive collection with proper date conversions and token usage data, then delete the Redis key. The user sees 'Session archived to MongoDB Atlas, evicted from Redis' — because that's literally what happens.

typescript
// Eviction pipeline — Redis → MongoDB → Redis DEL
const session = await redis.json.get(`session:${id}`);

await mongodb.collection("agent_sessions_archive").updateOne(
  { session_id: id },
  { $set: {
    ...session,
    evicted_at: new Date(),
    token_usage: {
      prompt_tokens: session.token_usage.prompt,
      completion_tokens: session.token_usage.completion,
      total_tokens: session.token_usage.total,
    }
  }},
  { upsert: true }
);

await redis.del(`session:${id}`); // free memory

This is the Dual-Engine lifecycle: Redis handles the real-time state (fast reads/writes during conversation), MongoDB handles the historical state (durable archives for analytics and observability). Neither database does both jobs well. Redis would waste memory on cold sessions. MongoDB would add latency to hot reads.

3. Five-Layer Rate Limiting — All Redis, All Atomic

Rate limiting for AI agents isn't one counter — it's a wall of counters at different granularities. We run five layers in StreakAI, all backed by Redis atomic operations. Every check completes before any expensive work (embedding, vector search, LLM call) happens.

LayerRedis Key PatternLimitTTLPurpose
Session messagesChecked against session JSON5 per sessionSession lifetimePrevent unbounded session growth
Hourly messagesratelimit:msgs:{sessionId}:{hourBucket}30 per hour3600sThrottle burst usage
Daily sessions per IPratelimit:sessions:{ip}:{dayBucket}10 per day86400sLimit session creation
Daily messages per IPratelimit:ip_msgs:{ip}:{dayBucket}50 per day86400sPrevent token waste from scrapers
Global daily token budgetratelimit:global_tokens:{dayBucket}500,000 tokens86400sHard cost ceiling across all users

The pattern for each counter is identical: Redis INCR + conditional EXPIRE. If the INCR returns 1 (first hit), we set the TTL. If the value exceeds the limit, we reject with a 429 before touching the LLM. The global token budget uses INCRBY to add actual token counts after each LLM response.

typescript
// Hourly rate limit — INCR + EXPIRE
const hourBucket = Math.floor(Date.now() / 3600000);
const key = `ratelimit:msgs:${sessionId}:${hourBucket}`;
const current = await redis.incr(key);
if (current === 1) await redis.expire(key, 3600);
if (current > 30) return { allowed: false, reason: "Rate limit exceeded" };

// Global token budget — INCRBY after LLM response
const dayBucket = Math.floor(Date.now() / 86400000);
const tokenKey = `ratelimit:global_tokens:${dayBucket}`;
await redis.incrBy(tokenKey, promptTokens + completionTokens);

Why Redis and not MongoDB for rate limiting? Because INCR is O(1), atomic, lock-free, and returns in under 0.1ms. A MongoDB findOneAndUpdate with $inc on the same counter takes 2-5ms and requires a write concern round trip. At 5 checks per request × 1,000 requests per second, that's the difference between 0.5ms total overhead and 10-25ms. The rate limiter must be invisible to the user — and Redis makes it invisible.

4. Semantic Caching — The $0.00 LLM Call

Every LLM call costs money and adds latency. Semantic caching eliminates both for repeated or similar questions. StreakAI caches every LLM response in Redis with its query embedding, then checks for KNN similarity before calling the LLM on the next request.

The flow: embed the user's query → KNN search the cache index for the nearest cached query → if similarity ≥ 0.80, return the cached response (0ms LLM latency, $0.00 cost). If no hit, call the LLM, then store the response + embedding in Redis with a 24-hour TTL.

typescript
// Check semantic cache — KNN 1 on HNSW index
const result = await redis.ft.search(
  "idx:semantic_cache",
  "*=>[KNN 1 @query_embedding $query_vec AS similarity]",
  { PARAMS: { query_vec: queryBuffer }, DIALECT: 2 }
);

const similarity = 1 - parseFloat(result.documents[0].value.similarity);
if (similarity >= 0.80) {
  // Cache hit — skip LLM entirely
  return cached.response; // 0ms LLM latency, $0.00
}

// Cache miss — call LLM, then store
const llmResponse = await streamFromLLM(query);
await redis.json.set(`cache:${hash}`, "$", {
  query, query_embedding: embedding, response: llmResponse,
  model: "deepseek.v3.2", created_at: timestamp
});
await redis.expire(`cache:${hash}`, 86400); // 24h TTL

The cache index is a Redis HNSW vector index with 1024-dimensional Amazon Titan embeddings. A KNN 1 search completes in under 1ms. The cache stores both the embedding (for similarity matching) and the full response (as RedisJSON for efficient retrieval). At a cache hit rate of 20-30% on a typical agent, this saves hundreds of dollars per month in LLM API costs.

The fastest LLM call is the one you don't make. Semantic caching in Redis turns repeated questions into sub-millisecond lookups at zero cost.

5. Context Retrieval — Redis Vectors, MongoDB Knowledge

The RAG pipeline in StreakAI follows a clear dual-engine split. The source content (blogs, docs, case studies) is authored and stored as JSON files — the system of record. During ingestion, each document is chunked (word-window splitting), embedded (Amazon Titan, 1024 dimensions), and indexed into Redis as HNSW vectors. At query time, the agent never touches the source files — it hits Redis exclusively.

StageEngineOperationLatency
Ingest: chunk + embedAWS Bedrock → RedisStore chunks as RedisJSON with HNSW-indexed embeddingsBatch, offline
Query: embed user questionAWS BedrockEmbed query with Amazon Titan50-80ms
Query: vector searchRedisFT.SEARCH with KNN top-5 on HNSW index2-5ms
Query: build contextRedis results → LLMFormat top-K chunks into <context> prompt<1ms
Archive: persist knowledgeMongoDBStore original docs, chunk metadata, embeddings for auditAsync
redis
# Redis HNSW vector index for knowledge base
FT.CREATE idx:knowledge_base ON JSON
  PREFIX 1 "kb:"
  SCHEMA
    $.embedding AS embedding VECTOR HNSW 6
      TYPE FLOAT32 DIM 1024 DISTANCE_METRIC COSINE
    $.title AS title TEXT
    $.source AS source TAG
    $.url AS url TEXT
    $.text AS text TEXT

# Query: find top-5 similar chunks
FT.SEARCH idx:knowledge_base
  "*=>[KNN 5 @embedding $query_vec AS score]"
  PARAMS 2 query_vec <binary_embedding>
  RETURN 5 text title source url score
  SORTBY score ASC
  LIMIT 0 5
  DIALECT 2

The context window is capped at the last 8 messages from the Redis session. This bounds the token cost per LLM call while keeping the conversation coherent. The context prompt includes chunk titles, source types, and URLs — so the LLM can cite its sources with clickable links in the response.

6. Token Accounting — Track, Cap, Persist

Token usage flows through both engines. During a session, Redis accumulates token counts in real-time using JSON numIncrBy operations. The global daily budget uses a Redis atomic counter. When the session is evicted, the cumulative token usage is persisted to MongoDB alongside the full conversation archive.

  • Per-request: After each LLM call, promptTokens and completionTokens are recorded in the request metrics (MongoDB) and added to the session's cumulative $.token_usage (Redis).
  • Per-session: The Redis session JSON accumulates total prompt, completion, and combined token counts. On eviction, this is written to MongoDB as token_usage.prompt_tokens, token_usage.completion_tokens, token_usage.total_tokens.
  • Global: The ratelimit:global_tokens:{dayBucket} Redis counter is incremented by the total tokens of each request. When it crosses 500,000, all users see 'daily limit reached' until the bucket expires at midnight.
  • Observability: MongoDB's agenti_ai_metrics collection stores per-request metrics including estimated cost (DeepSeek pricing: $0.0014/1K input, $0.0028/1K output), letting us build dashboards on spend per session, per user, per day.

The key insight: use Redis for real-time token math (atomic counters are fast and correct under concurrency), and MongoDB for historical token analytics (aggregation pipelines, time-series queries, cost dashboards). Trying to do real-time accounting in MongoDB adds latency. Trying to build dashboards from Redis keys is painful.

7. The Redis Key Map — A Living Blueprint

Here's every Redis key pattern in StreakAI's production instance, with data types and TTL policies. This is the real-time engine's footprint.

Key PatternTypeTTLPurpose
session:{uuid}RedisJSON1800s (sliding)Active conversation state, messages, token_usage
cache:{md5hash}RedisJSON86400s (fixed)Semantic cache — query, embedding, LLM response
kb:{slug}:{chunk_n}RedisJSONNoneKnowledge base chunks with HNSW-indexed embeddings
ratelimit:msgs:{sessionId}:{hourBucket}String (counter)3600sHourly message rate limit per session
ratelimit:sessions:{ip}:{dayBucket}String (counter)86400sDaily session creation limit per IP
ratelimit:ip_msgs:{ip}:{dayBucket}String (counter)86400sDaily message limit per IP
ratelimit:global_tokens:{dayBucket}String (counter)86400sGlobal daily token budget across all users

8. The MongoDB Collection Map — Durable Intelligence

And here's MongoDB's footprint — the durable store that survives restarts, enables analytics, and feeds observability.

CollectionPurposeKey Fields
agent_sessions_archiveEvicted sessions with full conversation historysession_id, messages[], token_usage, evicted_at, total_user_messages
agenti_ai_metricsPer-request pipeline metrics and cost trackingsession_id, promptTokens, completionTokens, estimatedCost, cacheHit, vectorSearchMs
knowledge_base_metaSource document metadata for ingestion auditingslug, title, chunk_count, embedded_at, dimensions

9. Why Not Just One Database?

We hear this question every week. 'Can't MongoDB do vector search now? Can't Redis persist data?' Yes to both — but being capable isn't the same as being optimal.

  • MongoDB-only: Atlas Vector Search adds 5-15ms per similarity query (vs 1-3ms in Redis HNSW). Rate limiting with findOneAndUpdate adds 2-5ms per check × 5 checks = 10-25ms overhead. Session reads from disk-backed storage add 1-3ms. Total added latency per request: 20-50ms. Feels sluggish for a chat agent.
  • Redis-only: No aggregation pipelines for analytics. No durable storage without AOF/RDB complexity. No rich query language for 'show me all sessions from last week with >1000 tokens'. Session archives would consume expensive memory indefinitely. Token analytics would require external tooling.
  • Dual-Engine: Redis handles the 6-8 operations per request that must be fast (session load, cache check, vector search, rate limit checks, session update, cache store). MongoDB handles the 1-2 operations that must be durable (metrics insert, session archive on eviction). Total request overhead from the fast path: <10ms. Durability and analytics: complete.
A Dual-Engine architecture isn't complexity for complexity's sake. It's putting each operation in the database that can execute it 10x faster. The result is an AI agent that feels instant and an analytics layer that misses nothing.

10. The Request Pipeline — Both Engines, One Flow

Here's the exact execution order for a single user message in StreakAI. Every step is annotated with which engine it touches and why.

  • 1. Load session from Redis (0.1ms) — RedisJSON GET. Need instant access to conversation history and message count.
  • 2. Check 5 rate limits in Redis (0.5ms total) — Five INCR/GET operations. Must reject before any expensive work.
  • 3. Embed the query via AWS Bedrock (50-80ms) — Amazon Titan, 1024 dimensions. This is the expensive step we protect with rate limiting.
  • 4. Check semantic cache in Redis (0.5-1ms) — KNN 1 on HNSW index. If hit (similarity ≥ 0.80), skip steps 5-6 entirely.
  • 5. Vector search knowledge base in Redis (2-5ms) — KNN 5 on HNSW index. Retrieve top-5 relevant chunks with titles and URLs.
  • 6. Stream LLM response from DeepSeek via Bedrock (500-3000ms) — The only step that takes real time. Everything else is optimized around it.
  • 7. Store in semantic cache in Redis (1-2ms) — RedisJSON SET with embedding + response for future cache hits.
  • 8. Update session in Redis (1-2ms) — Append messages, increment counter, accumulate token usage, reset TTL.
  • 9. If session limit reached: Evict to MongoDB (5-10ms) — Read full session, persist to agent_sessions_archive, DEL from Redis.
  • 10. Log metrics to MongoDB (fire-and-forget) — Insert per-request metrics for observability dashboards.

Redis touches: steps 1, 2, 4, 5, 7, 8, 9 (7 of 10 steps). MongoDB touches: steps 9, 10 (2 of 10 steps, both non-blocking or end-of-session). This is the Dual-Engine Pattern in action — Redis dominates the hot path, MongoDB handles the cold path.

11. Getting Started — Build the Dual-Engine Stack

If you're building an AI agent today, here's the minimum viable Dual-Engine setup:

  • Week 1 — Redis: Set up RedisJSON for session state with sliding TTL. Add a semantic cache with HNSW vector index. Implement INCR-based rate limiting. Index your knowledge base as HNSW vectors. This gives you a fast, protected agent.
  • Week 2 — MongoDB: Add session archiving on eviction. Store per-request metrics with timestamps. Build a knowledge base metadata collection for ingestion auditing. This gives you observability and durability.
  • Week 3 — Connect them: Implement the eviction pipeline (Redis → MongoDB on session end). Add token accumulation in Redis, persist on eviction. Set up global token budgets in Redis, report spending from MongoDB. This gives you the full Dual-Engine runtime.
  • Week 4 — Optimize: Add source-based filtering to vector search (@source TAG filters). Tune cache similarity threshold (we use 0.80). Add context window limits (last 8 messages to LLM). Build cost dashboards from MongoDB metrics. This gives you a production-grade system.

The total Redis memory footprint for StreakAI with ~200 knowledge chunks, semantic cache, active sessions, and rate limit counters is under 50MB. The MongoDB storage for session archives and metrics grows at roughly 2KB per session and 500 bytes per request — trivial at any scale.

12. The Bottom Line

The Dual-Engine Pattern isn't a compromise — it's a deliberate architectural decision. MongoDB and Redis aren't competing for the same job. They're collaborating on different jobs within the same AI agent runtime.

  • Redis is the speed engine: sessions, cache, vectors, counters, real-time state. Anything the agent reads during inference.
  • MongoDB is the intelligence engine: archives, metrics, knowledge metadata, analytics. Anything the team reads after inference.
  • The eviction bridge connects them: sessions graduate from Redis (hot) to MongoDB (cold). Token budgets accumulate in Redis (fast) and report to MongoDB (durable). Knowledge is ingested into MongoDB (source of truth) and indexed in Redis (search engine).

We run this exact architecture in StreakAI — and you can try it live right now on our agent page. Ask it about semantic caching, then ask again to see the Redis cache hit. Watch the X-Ray panel to see latency numbers from both engines in real-time. That's not a demo — that's the Dual-Engine Pattern in production.

Stop choosing between MongoDB and Redis. Use both. That's not complexity — it's engineering each layer for what it does best. MongoDB + Redis = the complete AI agent runtime.