Databases 26 min read

Redis 8.0 Adds Native Vector Search, Agent Memory, and Semantic Caching for AI Apps

Redis 8.0 introduces native vector search with Vector Set data type, HNSW indexing, quantization, and the Iris platform for agent memory and context retrieval, enabling semantic caching that cuts LLM costs by 70% and integrates with LangChain4j and Spring AI.

java1234
java1234
java1234
Redis 8.0 Adds Native Vector Search, Agent Memory, and Semantic Caching for AI Apps

What Is Redis

Redis (REmote DIctionary Server), open-sourced in 2009 by Salvatore Sanfilippo, is an in-memory data-structure store. Beyond simple strings, it natively provides List, Hash, Set, Sorted Set, Stream, Bitmap, and HyperLogLog. For over a decade it has served as a cache in front of MySQL, a distributed lock ( SET key value NX PX 30000), a message queue (List then Stream), a leaderboard/counter (Sorted Set, INCR), and a session store — all roles that share two traits: microsecond latency and flexible data structures.

Why Redis for the AI Era

LLMs are stateless; every call forgets prior context unless the application re-injects it. This requires three steps within tens of milliseconds: (1) embed the user question, (2) find the semantically closest items from knowledge/history, (3) assemble them into the prompt. Low-latency, high-volume, structured retrieval is exactly what Redis has optimized for 15 years, so Redis moved vector capabilities into the kernel and layered agent memory on top.

Core Foundation: Vector Capabilities in the Kernel

3.1 Native Vector Set Data Type

Redis 8.0 adds Vector Set , a first-class data type analogous to Sorted Set. Where Sorted Set attaches a scalar score to each member, Vector Set attaches a vector; similarity replaces score ordering. The underlying algorithm is HNSW (Hierarchical Navigable Small World), delivering logarithmic-time approximate nearest-neighbor search.

Basic commands:

# VADD: add elements to vector set "movies"
# Syntax: VADD key VALUES <dim> <v1> <v2> ... <member>
VADD movies VALUES 4 0.12 0.88 0.35 0.91 "流浪地球"
VADD movies VALUES 4 0.15 0.85 0.31 0.88 "星际穿越"
VADD movies VALUES 4 0.91 0.12 0.77 0.05 "疯狂动物城"

# SETATTR: attach scalar attributes for later filtering
VSETATTR movies "流浪地球" '{"year": 2019, "type": "科幻"}'
VSETATTR movies "星际穿越" '{"year": 2014, "type": "科幻"}'
VSETATTR movies "疯狂动物城" '{"year": 2016, "type": "动画"}'

# VSIM: similarity search, return top 2 with scores
VSIM movies VALUES 4 0.13 0.86 0.33 0.90 COUNT 2 WITHSCORES
# Returns: 1) "流浪地球" 2) "0.998" 3) "星际穿越" 4) "0.995"

# Vector search + scalar filter in one call
VSIM movies VALUES 4 0.13 0.86 0.33 0.90 COUNT 2 \
  FILTER '.year > 2015 and .type == "科幻"'

The FILTER clause executes vector retrieval and scalar filtering in a single round-trip, eliminating application-layer post-filtering — critical for recommendation and permissioned semantic search.

3.2 Built-in Vector Search Engine

Previously, vector search required the separate RediSearch module; now it is merged into the core. It supports both HNSW and FLAT indexes and combines vector KNN with full-text, numeric range, and tag filters in a single hybrid query.

# Create a vector index on Hash data
FT.CREATE idx:docs ON HASH PREFIX 1 doc: SCHEMA
  title TEXT          # full-text search
  category TAG        # exact filter
  views NUMERIC SORTABLE  # range query + sort
  embedding VECTOR HNSW 6  # vector field, HNSW index
    TYPE FLOAT32
    DIM 1536          # matches embedding model output
    DISTANCE_METRIC COSINE

# Hybrid query: filter first, then KNN
FT.SEARCH idx:docs
  "(@category:{技术} @views:[1000 +inf])=>[KNN 5 @embedding $vec AS score]"
  PARAMS 2 vec "<binary vector>"
  SORTBY score ASC
  RETURN 2 title score
  DIALECT 2

The “filter-then-KNN” syntax lets business filters and semantic search run atomically.

3.3 Quantization: Trade a Little Precision for 32× Memory Savings

A 1536-dim FLOAT32 vector consumes 6 KB; 10 million vectors = 60 GB. Redis offers three quantization levels:

No quantization (FP32) – 1× storage, highest recall; for small data, extreme precision.

8-bit (Q8) – ~4× compression, minimal recall loss; default for most production workloads .

Binary (BIN) – up to 32× compression, noticeable recall drop; for massive scale where coarse recall is acceptable.

Common practice is a two-stage pipeline: coarse recall with BIN, then re-rank candidates with full-precision vectors.

# Enable 8-bit quantization at insert
VADD movies Q8 VALUES 4 0.12 0.88 0.35 0.91 "流浪地球"

# Binary quantization for maximum compression
VADD movies BIN VALUES 4 0.12 0.88 0.35 0.91 "流浪地球"

Agent Memory & Context Engine: Redis Iris

Iris is a purpose-built context and memory platform for AI agents. It addresses the production gap where demo agents work but production agents hallucinate due to fragmented, stale data.

① Redis Context Retriever

Developers semantically model business entities (Customer, Order, Product) and their relationships; the system auto-generates safe data-access tools (via MCP) for the agent. This prevents the agent from writing raw SQL — avoiding syntax errors, full-table scans, and privilege escalation.

② Redis Agent Memory

Dual-layer architecture mirroring human memory:

Working Memory (Session) – short-lived, stores current dialogue, expires via TTL.

Long-term Memory – persistent, stores user preferences, historical conclusions, key facts; reusable across sessions.

Example: a peanut allergy mentioned three months ago is still honored today.

import com.redis.vl.extensions.session.ChatMessage;
import com.redis.vl.extensions.session.SemanticSessionManager;
import redis.clients.jedis.JedisPooled;
import java.util.List;

public class AgentMemoryDemo {
  public static void main(String[] args) {
    JedisPooled client = new JedisPooled("localhost", 6379);
    SemanticSessionManager session = SemanticSessionManager.builder()
      .name("assistant_session")
      .sessionTag("user:1001")
      .redisClient(client)
      .build();

    // Write session memory (paired Q/A)
    session.addMessage(new ChatMessage("user", "Recommend a nearby Sichuan restaurant"));
    session.addMessage(new ChatMessage("llm", "Recommend \"Shuxiangju\", rating 4.8"));
    session.addMessage(new ChatMessage("user", "Reminder: I'm allergic to peanuts"));

    // Semantic retrieval: fetch only relevant memories
    List<ChatMessage> context = session.getRelevant(
      "What's good for dinner tonight",  // current question
      3,                                 // max 3 memories
      0.3                                // semantic distance threshold
    );
    // Result includes "peanut allergy", enabling safe recommendation
    for (ChatMessage item : context) {
      System.out.println(item.getRole() + " : " + item.getContent());
    }
  }
}

The key is getRelevant(): it selects only semantically pertinent memories instead of stuffing the entire history into the limited, expensive context window.

③ Redis Data Integration

Near-real-time sync from databases and warehouses into Redis. Solves the “stale data” problem — e.g., inventory already zero but agent still recommends the item.

Developer Tools & Ecosystem Integration

5.1 RedisVL: Client Library for AI Apps

RedisVL (Redis Vector Library) wraps index creation, vector writes, and similarity queries into ergonomic APIs. Java artifact: com.redis:redis-vl-java:0.4.0.

<dependency>
  <groupId>com.redis</groupId>
  <artifactId>redis-vl-java</artifactId>
  <version>0.4.0</version>
</dependency>

Minimal RAG retriever example:

import com.redis.vl.index.SearchIndex;
import com.redis.vl.query.VectorQuery;
import com.redis.vl.schema.IndexSchema;
import java.util.HashMap;
import java.util.List;
import java.util.Map;

public class RagRetrieverDemo {
  private static final String REDIS_URL = "redis://localhost:6379";

  public static void main(String[] args) {
    // 1. Define index schema (or load from YAML)
    IndexSchema schema = IndexSchema.builder()
      .name("kb_index")
      .prefix("kb")
      .storageType(IndexSchema.StorageType.HASH)
      .addTextField("content")
      .addTagField("source")
      .addVectorField("embedding", builder -> builder
        .dims(1536)
        .algorithm("hnsw")
        .distanceMetric("cosine")
        .dataType("float32"))
      .build();

    // 2. Create index (false = reuse if exists)
    SearchIndex index = new SearchIndex(schema, REDIS_URL);
    index.create(false);

    // 3. Bulk load documents (embeddings from your model)
    index.load(List.of(
      buildDoc("Redis 8.0 introduces native Vector Set data type", "release-note"),
      buildDoc("HNSW is a graph-based approximate nearest neighbor algorithm", "wiki")
    ));

    // 4. Semantic search with scalar filter
    VectorQuery query = VectorQuery.builder()
      .vector(embed("What is Redis's vector type called?"))
      .vectorFieldName("embedding")
      .returnFields("content", "source")
      .numResults(3)
      .filterExpression("@source:{release-note}")  // filter + vector search simultaneously
      .build();

    for (Map<String, Object> doc : index.query(query)) {
      System.out.println("[" + doc.get("vector_distance") + "] " + doc.get("content"));
    }
  }

  private static Map<String, Object> buildDoc(String content, String source) {
    Map<String, Object> doc = new HashMap<>();
    doc.put("content", content);
    doc.put("source", source);
    doc.put("embedding", embed(content));
    return doc;
  }

  private static float[] embed(String text) {
    // TODO: plug in OpenAI, Qwen, or local embedding model
    return new float[1536];
  }
}

5.2 LangCache: Semantic Caching — the ROI Winner

Traditional caches key-match exactly; “how to return”, “return process”, “I want a return” are three misses → three LLM bills. LangCache matches on semantic similarity: vectors close enough → cache hit.

Official numbers: up to 70% LLM cost reduction, ~15× faster response on cache hit .

import com.redis.vl.extensions.cache.CacheHit;
import com.redis.vl.extensions.cache.SemanticCache;
import java.util.List;

public class SemanticCacheDemo {
  private final SemanticCache cache;

  public SemanticCacheDemo() {
    this.cache = SemanticCache.builder()
      .name("llm_cache")
      .redisUrl("redis://localhost:6379")
      .distanceThreshold(0.15)  // typical range 0.1–0.2; larger = looser match
      .ttl(3600)                // 1 hour TTL
      .build();
  }

  public String ask(String question) {
    // 1. Check semantic cache
    List<CacheHit> hits = cache.check(question, 1);
    if (!hits.isEmpty()) {
      System.out.println("Cache hit — saved one LLM call");
      return hits.get(0).getResponse();
    }
    // 2. Miss → call LLM
    String answer = callLlm(question);
    // 3. Store for future semantic matches
    cache.store(question, answer);
    return answer;
  }

  private String callLlm(String question) {
    // TODO: integrate OpenAI, Qwen, etc.
    return "Answer from LLM";
  }

  public static void main(String[] args) {
    SemanticCacheDemo demo = new SemanticCacheDemo();
    demo.ask("How does Redis do vector search?");      // miss → LLM
    demo.ask("How does Redis implement vector search?"); // semantic match → cache hit
  }
}

5.3 LangChain4j / Spring AI Integration

Redis is an officially supported vector store and chat memory backend in both frameworks, serving three roles simultaneously: vector store, semantic cache, and conversation history persistence.

<dependency>
  <groupId>dev.langchain4j</groupId>
  <artifactId>langchain4j-redis</artifactId>
  <version>0.36.0</version>
</dependency>
import dev.langchain4j.data.embedding.Embedding;
import dev.langchain4j.data.message.AiMessage;
import dev.langchain4j.data.message.ChatMessage;
import dev.langchain4j.data.message.UserMessage;
import dev.langchain4j.data.segment.TextSegment;
import dev.langchain4j.memory.chat.MessageWindowChatMemory;
import dev.langchain4j.model.embedding.EmbeddingModel;
import dev.langchain4j.model.openai.OpenAiEmbeddingModel;
import dev.langchain4j.store.embedding.redis.RedisEmbeddingStore;
import dev.langchain4j.store.memory.chat.redis.RedisChatMemoryStore;
import java.util.List;

public class LangChain4jRedisDemo {
  private static final String REDIS_HOST = "localhost";
  private static final int REDIS_PORT = 6379;

  public static void main(String[] args) {
    // Role 1: Vector store (knowledge base)
    RedisEmbeddingStore embeddingStore = RedisEmbeddingStore.builder()
      .host(REDIS_HOST)
      .port(REDIS_PORT)
      .indexName("langchain_kb")
      .dimension(1536)
      .build();

    EmbeddingModel embeddingModel = OpenAiEmbeddingModel.withApiKey("YOUR-API-KEY");
    List<String> docs = List.of(
      "Redis 8.0 supports native vector type",
      "Iris is a context platform for agents"
    );
    for (String doc : docs) {
      TextSegment segment = TextSegment.from(doc);
      Embedding embedding = embeddingModel.embed(segment).content();
      embeddingStore.add(embedding, segment);
    }

    // Role 2: Chat history persistence (survives restarts)
    RedisChatMemoryStore memoryStore = RedisChatMemoryStore.builder()
      .host(REDIS_HOST)
      .port(REDIS_PORT)
      .ttl(86400L)  // 1 day retention
      .build();

    MessageWindowChatMemory memory = MessageWindowChatMemory.builder()
      .id("user:1001")
      .maxMessages(20)
      .chatMemoryStore(memoryStore)
      .build();

    memory.add(UserMessage.from("Introduce Vector Set"));
    memory.add(AiMessage.from("Vector Set is Redis 8.0's native vector data type..."));
    for (ChatMessage message : memory.messages()) {
      System.out.println(message);
    }
  }
}

5.4 MCP Server: Let AI Assistants Operate Redis Directly

Redis ships an official MCP (Model Context Protocol) server. Once configured, assistants like Claude can converse with your Redis data in natural language — “show keys about to expire”, “display this Hash structure” — without manual commands.

{
  "mcpServers": {
    "redis": {
      "command": "docker",
      "args": [
        "run", "--rm", "-i",
        "-e", "REDIS_HOST=host.docker.internal",
        "-e", "REDIS_PORT=6379",
        "-e", "REDIS_PWD=123456",
        "mcp/redis"
      ]
    }
  }
}

The Redis AI Incubator continues to release experimental projects (filesystem-in-Redis, Google ADK integration, Microsoft Agent Framework integration), signaling a clear strategic direction.

A Complete Production-Grade RAG Request Flow

Combining the pieces, a production AI request chain looks like this:

RAG request flow diagram
RAG request flow diagram

The dashed box highlights the convergence: semantic cache, agent memory, and vector search all land in a single Redis instance . Previously you might need Pinecone for vectors, PostgreSQL for sessions, and Redis for caching — three systems, three ops burdens, data shuffling. Now they collapse into one component with millisecond latency and a drastically simpler stack.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

QuantizationRAGRedisVector SearchSpring AIAgent MemoryLangChain4jSemantic Caching
java1234
Written by

java1234

Former senior programmer at a Fortune Global 500 company, dedicated to sharing Java expertise. Visit Feng's site: Java Knowledge Sharing, www.java1234.com

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.