grep Showdown: rg vs tgrep vs rawgrep vs zg
This article compares four code search tools—ripgrep, tgrep, rawgrep, and zg—analyzing their distinct technical approaches: full-scan SIMD regex, trigram indexing, raw disk access, and semantic vector search, with benchmarks and guidance for choosing based on repository size, frequency, and AI agent integration.
Introduction
When developers search codebases, the default answer is rg (ripgrep). Since its 2016 release, ripgrep has become the de facto standard with 68k stars, praised for speed and stability. However, as repositories grow to where rg takes several seconds per search, newer tools have emerged with fundamentally different architectures.
1. rg: Full-Scan Optimized to the Extreme
rg reads every file from start to finish, using a SIMD-accelerated regex engine. Its performance hinges on bytes scanned per unit time. Key features:
Default recursive search with automatic filtering: respects .gitignore, skips hidden and binary files; rg -uuu disables all filters.
Unicode enabled by default without speed penalty — GNU grep still cannot do this.
File-type filtering: rg -tpy foo searches only Python files, rg -Tjs foo excludes JavaScript; custom types supported.
Optional PCRE2 via -P for lookahead/backreferences; default engine stays fast.
Search compressed files directly with -z (gzip, xz, bzip2).
Preprocessor pipeline to extract text from PDFs before searching.
Encoding support for UTF-16, GBK, Shift_JIS, etc.
Benchmarks from rg's README: searching Linux kernel for [A-Z]+_SUSPEND takes 0.082s (i9-12900K), 5x faster than ag, 30x faster than ack or git grep. Search time scales linearly with total bytes of searched files (O(n)); no index or cache is built.
2. tgrep: Prebuilt Trigram Index, Touch Only Candidate Files
Microsoft's tgrep changes the paradigm: prebuild a trigram (three-character substring) index, then search only files that might match. Its motto: "Start a server once, search instantly forever."
Usage
tgrep index . # build trigram index
tgrep serve . # start background server, watches file changes
tgrep "fn main" . # instant results, connects to running serverArchitecture: server/client model.
Benchmark Results (from README)
gecko-dev (388K files): rg 33.4s vs tgrep 643ms — 51.9x faster.
chromium (504K files): 15.8x faster.
Linux kernel: 21x faster.
tgrep wins 17 of 18 test configurations.
Integrated into GitHub Copilot CLI for AI-assisted code search in huge repos, validating the approach at scale.
Architecture Highlights
Two-layer index: IndexReader (mmap'd on-disk index, zero-copy, binary search) + LiveIndex (in-memory overlay for files changed after server start); queries merge both, overlay wins.
Parallel background indexing via rayon (batches of 1024 files); queries fall back to filesystem scan until indexing finishes — no waiting.
Periodic flush: every 5 minutes or 50k files, memory index written to disk and read endpoint swapped; memory stays bounded.
Memory-bounded indexing: default external strategy uses external merge sort to cap peak memory at ~160MB vs 2–3.6GB for pure in-memory, a 17x reduction without speed loss.
Best Fit & Caveats
Ideal for huge monorepos searched repeatedly; index cost amortizes over many queries. For occasional searches on small projects, indexing overhead may exceed full scan — tgrep provides --no-index to fall back.
Two gotchas from README: tgrep index and tgrep serve must use identical parameters (e.g., --max-filesize, --exclude), otherwise server treats oversized files as deleted. .gitignore only works inside git repos (like rg); for non-git directories, add --no-require-git or index will be larger than expected.
3. rawgrep: Bypass Filesystem, Read Raw Disk
If tgrep changes the algorithm, rawgrep changes the medium: it reads raw block devices directly, skipping the filesystem layer. Tagline: "Grep at the speed of raw disk."
Setup
# One-time grant capability (bypass read permission check only, read-only, never writes)
sudo setcap cap_dac_read_search=eip ./target/release-fast/rawgrep
# Afterwards, no sudo needed
rawgrep "TODO" /var/logAccessing /dev/sda1 requires either sudo per run or granting the cap_dac_read_search capability to the binary.
Speed Keys
Bypass filesystem: stream directly from device, eliminating one layer of overhead.
Fragment cache (inspired by nowgrep): learns which file fragments can be skipped across repeated searches; cache value grows with search volume.
Benchmarks
Chromium (500K files) searching TODO: hot cache + fragment cache 131.8ms vs rg 363.5ms (2.76x); cold cache 2.72s vs 11.90s (4.38x).
Author reports ~60x speedup on personal 1.27M-file home directory.
Trade-offs (explicit in README)
Linux only; requires ext4/ntfs filesystems — macOS and Windows unsupported.
Requires root or setcap capability; many users cannot or should not grant this.
Memory usage not yet optimized; author acknowledges high RSS, currently focused solely on speed.
Verdict: a cool technical demo and a reference for grep's performance ceiling; practical for niche scenarios (read-only partitions, massive logs, repeatedly searched huge disks), but the "raw disk" barrier excludes most developers.
4. zg: Semantic Search Inside grep, Built for AI Agents
While the first three optimize keyword matching, zg (zvec-grep) steps out of the keyword box: unifies ripgrep + BM25 + vector retrieval in a single local-first entry point.
Usage
# Requires Node.js 22+, install globally
npm install -g @zvec/zvec-grep
# Build index (default uses local embedding model, data never leaves machine)
zg index --embedding local/potion-retrieval-32m
# Human query
zg query --human "An unseen creature left a few marks. What did the detective infer?" --limit 3
# Install into AI Agent
zg install --target opencode --yesCore Philosophy
"Find by meaning first, verify by keyword." Semantic recall surfaces relevant snippets ranked by relevance; precise confirmation falls back to text/regex. Supports source code, docs, structured data; results retain structure and source locations.
Two Key Differentiators
Local-first is not a slogan: files, indexes, local models all stay on-device; remote embeddings only with explicit user consent — critical for sensitive code.
Designed for Agents from day one: MCP support integrates with Codex, Claude Code, Qwen Code, Cursor, OpenCode. Benchmarks use Agent workflows: SWE-QA-Bench with Claude Code + Opus, BrowseComp-Plus with Codex — measuring answer quality, input tokens, tool calls, latency, not raw human speed. Conclusion: semantic recall significantly reduces Agent's futile searches and token consumption.
Authored by Alibaba's zvec team; bilingual README, active community via DingTalk, WeChat, Discord.
5. What Coding Agents Actually Use: All Choose ripgrep
Investigating three leading AI coding agents reveals unanimous convergence on ripgrep:
Claude Code: Built-in Grep tool is ripgrep, bundled in the distribution — no separate rg install needed. Cross-platform behavior consistent; default parallel search; regex semantics identical to rg.
Codex (OpenAI): Packaging script hard-requires "ripgrep is required for all package targets" — every platform release must bundle rg. Sandbox code search uses this bundled ripgrep.
Pi (earendil-works/pi): Built-in grep and find tools; grep spawns rg with --json --line-number --color=never --hidden and parses JSON output. Auto-downloads rg if missing.
Reason: Agents need search that never fails, never lags, respects .gitignore, skips binaries, and outputs parseable JSON ( --json). Ripgrep is the battle-tested default after ten years.
This explains the newer tools' purpose: rg is the Agent's default first search; when repos grow so large that rg takes seconds, tgrep replaces repeated full scans with index-based candidate filtering, and zg adds semantic recall to cut wasted searches and token burn.
6. Selection Guide: A Decision Matrix
The article summarizes the four tools in a comparison table (converted to list below):
rg — Full-scan, algorithm speed; Rust; all platforms; daily grep everything; zero setup.
tgrep — Prebuilt trigram index; Rust; all platforms; huge monorepos, high-frequency repeated search; medium (index first).
rawgrep — Raw disk bypass; Rust; Linux only (ext4/ntfs); read-only partitions, massive logs; high (root/capability required).
zg — Semantic (vector + BM25 + rg); TypeScript; all platforms; AI Agents, semantic retrieval; medium (Node 22+).
Practical Recommendations
Daily coding search: rg remains default — rock-solid, no fuss.
Repo so large rg takes seconds, and you (or your AI) search dozens of times a day: tgrep pays off — index cost amortizes quickly; Copilot CLI endorsement proves the path.
rawgrep as a performance ceiling reference; only adopt if you accept Linux + capability constraints.
Using AI coding agents and tired of them rummaging through irrelevant code, burning tokens: zg is the one to watch — semantic + keyword combo fills the gap when exact names are unknown.
Final thought: four blades, four distinct technical bets — algorithm, index, hardware, semantics. Pick the one that fits your scenario; no need to obsess over which is "strongest."
References
rg: https://github.com/BurntSushi/ripgrep tgrep: https://github.com/microsoft/tgrep rawgrep: https://github.com/rakivo/rawgrep zg: https://github.com/zvec-ai/zvec-grep
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
BirdNest Tech Talk
Author of the rpcx microservice framework, original book author, and chair of Baidu's Go CMC committee.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
