Xiaohongshu Tech REDtech
Author

Xiaohongshu Tech REDtech

Official account of the Xiaohongshu tech team, sharing tech innovations and problem insights, advancing together.

134
Articles
0
Likes
913
Views
0
Comments
Recent Articles

Latest from Xiaohongshu Tech REDtech

100 recent articles max
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Aug 14, 2026 · Artificial Intelligence

dots3-note Preview: A First Step Toward Long‑Term Agents for Real‑World Service

The open‑source dots3-note Preview model, a 280B‑parameter multimodal agent with 512K context, introduces the TEMPO training scheme to improve long‑term reinforcement learning, achieves benchmark gains of up to 31.5% over baselines, and is evaluated on new VibeSearchBench and VibeLifeBench suites while acknowledging current limitations.

Agentic AIbenchmarklong-term planning
0 likes · 27 min read
dots3-note Preview: A First Step Toward Long‑Term Agents for Real‑World Service
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Aug 13, 2026 · Artificial Intelligence

dots.tts: Open‑Source Continuous Autoregressive TTS Model for Sustainable, Scalable Voice Synthesis

dots.tts is a 2‑billion‑parameter, fully continuous, end‑to‑end autoregressive TTS model released by the Xiaohongshu dots team, offering state‑of‑the‑art zero‑shot voice cloning, low‑step inference, 1‑to‑1 audio‑text streaming, and an extensible training pipeline for research and deployment.

Speech SynthesisTTSaudio-text streaming
0 likes · 22 min read
dots.tts: Open‑Source Continuous Autoregressive TTS Model for Sustainable, Scalable Voice Synthesis
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Aug 7, 2026 · Artificial Intelligence

Can AI Translation Miss Memes? CULTURE‑MT Benchmark at ICML 2026

The authors introduce CULTURE‑MT, the first Chinese‑English social‑media translation benchmark that evaluates cultural effectiveness, define a new metric, release the JUDGER automatic evaluator (86 % accuracy, κ = 0.72), and show that even top models like Gemini 3 pro achieve only 38 % perfect cultural translations.

AI translationbenchmarkcultural evaluation
0 likes · 10 min read
Can AI Translation Miss Memes? CULTURE‑MT Benchmark at ICML 2026
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Aug 5, 2026 · Artificial Intelligence

Accelerating Xiaohongshu ‘Ask’ Inference: Slim Vision Tokens, Focused MoE

The article details how Xiaohongshu’s multimodal “Ask” service reduces visual token bloat and MoE expert compute by applying dynamic Vision Token compression (VisionZip) and similarity‑based expert re‑routing (SERE), achieving up to 13% lower first‑token latency and 16.5% faster end‑to‑end throughput while preserving answer quality.

MoE expert routingQwen modelsSERE
0 likes · 19 min read
Accelerating Xiaohongshu ‘Ask’ Inference: Slim Vision Tokens, Focused MoE
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jul 24, 2026 · Artificial Intelligence

UniNote: Unifying Multimodal Representation and Ranking in a Single Embedding Model

UniNote introduces a unified multimodal embedding model that combines representation learning and ranking optimization in a single forward pass, using a two‑stage SFT‑then‑RL training paradigm and Matryoshka Representation Learning to achieve competitive Item‑to‑Item retrieval performance while reducing latency.

Item2ItemMatryoshka Representation LearningSFT
0 likes · 12 min read
UniNote: Unifying Multimodal Representation and Ranking in a Single Embedding Model
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jul 23, 2026 · Artificial Intelligence

HELMSMAN: Redefining Large-Scale Vector Retrieval on All‑Flash Servers (OSDI 2026)

HELMSMAN replaces DRAM‑heavy graph indexes with a clustering‑based, SSD‑first ANN system that uses a custom SPDK storage stack, adaptive LLSP pruning, and a GPU‑CPU construction pipeline, achieving 2‑16× throughput, up to 85% of in‑memory performance, and over 90% hardware cost reduction for billion‑scale search workloads.

Approximate Nearest NeighborClusteringGPU
0 likes · 13 min read
HELMSMAN: Redefining Large-Scale Vector Retrieval on All‑Flash Servers (OSDI 2026)
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jul 16, 2026 · Artificial Intelligence

Cut First‑Token Latency by 3.25×: Introducing HYPIC for Position‑Independent Caching in Hybrid‑Attention LLMs

HYPIC combines position‑independent caching with hybrid‑attention LLMs, reducing first‑token latency by 3.25× and sustaining QPS by 1.66× while keeping quality loss under 2 points, through a segment‑cumulative transition operator, a seam‑window fix for full‑attention layers, and segment‑parallel execution.

HYPICKV CacheLLM inference
0 likes · 13 min read
Cut First‑Token Latency by 3.25×: Introducing HYPIC for Position‑Independent Caching in Hybrid‑Attention LLMs
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jul 14, 2026 · Artificial Intelligence

How Xiaohongshu Built an Enterprise AI Personal Assistant from Zero to Full‑Staff Coverage

Xiaohongshu’s AI team describes a month‑long, three‑person effort that leveraged an AI‑Native project model, isolated Kubernetes clusters, a custom sandbox (NEX), token‑saving Self‑GC, cost‑aware routing, and a three‑layer memory architecture to roll out a secure, low‑cost AI personal assistant used by every employee.

AI agentEnterprise AIMemory Architecture
0 likes · 14 min read
How Xiaohongshu Built an Enterprise AI Personal Assistant from Zero to Full‑Staff Coverage