Tagged articles

semantic cache

10 articles · Page 1 of 1
Su San Talks Tech
Su San Talks Tech
Aug 27, 2026 · Artificial Intelligence

How Redis Becomes the Real‑Time Data Backbone for AI

Redis has evolved from a pure cache middleware into a full‑stack AI data infrastructure, offering vector search, native vector sets, semantic caching, and an AI Agent context engine, with sub‑millisecond latency, high throughput, and detailed performance benchmarks that illustrate its strengths and trade‑offs.

AIAgent MemoryJava
0 likes · 15 min read
How Redis Becomes the Real‑Time Data Backbone for AI
AI Architecture Path
AI Architecture Path
Aug 27, 2026 · Artificial Intelligence

A Free, No‑API‑Key Agent Web Toolkit: Local Search, Deep Crawling, and Site Monitoring in One Package

Wigolo is a 4.7k‑star open‑source toolbox that brings internet search, page fetching, full‑site crawling, structured extraction, local caching, research reporting, and web‑monitoring to AI agents entirely on the developer’s machine, eliminating API keys, usage limits, and privacy leaks.

AI AgentMCP protocolWeb Crawling
0 likes · 14 min read
A Free, No‑API‑Key Agent Web Toolkit: Local Search, Deep Crawling, and Site Monitoring in One Package
ITPUB
ITPUB
Jul 24, 2026 · Databases

Interview with Yang Yu: Databases Face Their Most Dramatic Role Shift in 50 Years

The article examines how databases, after five decades of human‑centric design, are undergoing a fundamental transformation driven by Agentic AI, requiring new semantics, multimodal storage, memory capabilities, and integrated engines, illustrated through insights from Yang Yu of KuKe Data.

AIAgentMultimodal
0 likes · 13 min read
Interview with Yang Yu: Databases Face Their Most Dramatic Role Shift in 50 Years
ThinkingAgent
ThinkingAgent
Jul 20, 2026 · Artificial Intelligence

AI Infra in Practice Part 12: Cross‑Cutting Cost Governance with Token Economics

The article presents a comprehensive AI FinOps framework that attributes every AI expense to specific apps, users, and tasks, normalizes diverse cost units, and applies token economics, smart routing, semantic caching, and budget controls to ensure sustainable AI operations and measurable ROI.

AI FinOpsCost attributionGPU utilization
0 likes · 33 min read
AI Infra in Practice Part 12: Cross‑Cutting Cost Governance with Token Economics
Tencent Cloud Middleware
Tencent Cloud Middleware
Jul 15, 2026 · Cloud Computing

How AI Gateway Makes Token Costs Visible, Controllable, and Cost‑Effective

A CTO discovers a three‑fold surge in AI model token bills with no clear usage breakdown, prompting a deep dive into Tencent Cloud's AI Gateway, which offers multi‑dimensional quota governance, real‑time cost monitoring, semantic caching, and a unified cost‑management dashboard to make every token expense transparent and controllable.

AI GatewayCloud Computingquota management
0 likes · 11 min read
How AI Gateway Makes Token Costs Visible, Controllable, and Cost‑Effective
Architect Practice
Architect Practice
Jun 1, 2026 · Artificial Intelligence

When AI Rate Limiting Goes Wrong: A Four‑Dimension Framework and Three‑Layer Gateway in Practice

A midnight alarm at a fintech AI platform revealed that traditional QPS throttling missed a runaway Agent that consumed hundreds of times more tokens, prompting a detailed analysis of four token‑based limiting dimensions, three‑layer gateway design, agent‑specific controls, semantic caching, and tool selection to prevent similar “ghost avalanche” failures.

AI rate limitingAgent SafetyLLM Operations
0 likes · 20 min read
When AI Rate Limiting Goes Wrong: A Four‑Dimension Framework and Three‑Layer Gateway in Practice
Su San Talks Tech
Su San Talks Tech
May 11, 2026 · Artificial Intelligence

Designing a Production‑Ready LLM Gateway: Architecture, Routing, Fallback, and Observability

This article outlines a production‑grade LLM Gateway design, detailing a three‑layer architecture, capability‑, cost‑, latency‑ and semantic‑based routing strategies, multi‑level fallback mechanisms, specialized load balancing, unified API adaptation, semantic caching, observability, and compares popular open‑source implementations.

LLMfallbackgateway
0 likes · 17 min read
Designing a Production‑Ready LLM Gateway: Architecture, Routing, Fallback, and Observability
Linyb Geek Road
Linyb Geek Road
May 5, 2026 · Artificial Intelligence

Optimizing Retrieval and Generation Latency in High‑Concurrency RAG Agents

The article dissects latency in high‑concurrency RAG Agent pipelines, showing how retrieval, re‑ranking, and LLM generation each contribute milliseconds of delay, and presents system‑level tactics—from ANN index tuning and partitioned search to vLLM PagedAttention, continuous batching, speculative decoding, model quantization, routing, semantic caching, and pipeline parallelism—to dramatically cut end‑to‑end response time.

ANNLLMRAG
0 likes · 15 min read
Optimizing Retrieval and Generation Latency in High‑Concurrency RAG Agents
Linyb Geek Road
Linyb Geek Road
Apr 27, 2026 · Artificial Intelligence

Designing a Production LLM Gateway: Architecture, Routing, and Fallback

The article outlines a production‑grade LLM Gateway architecture divided into ingress, decision, and egress layers, detailing capability‑based, cost‑aware, latency‑aware, and semantic routing, multi‑stage fallback mechanisms, specialized load‑balancing, protocol unification, semantic caching, observability, and evaluates open‑source solutions such as LiteLLM, RouteLLM, and Portkey.

LLM gatewayfallbackload balancing
0 likes · 18 min read
Designing a Production LLM Gateway: Architecture, Routing, and Fallback