Tagged articles

URL deduplication

4 articles · Page 1 of 1
Subtle Storm
Subtle Storm
Jun 14, 2026 · Fundamentals

Master Bloom Filters in 10 Minutes: How They Prevent Cache Penetration

A Bloom filter is a probabilistic data structure that answers set‑membership queries with guaranteed no false negatives but possible false positives, using a bit array and multiple hash functions; the article demonstrates its mechanics with examples and shows its use in preventing cache penetration, URL de‑duplication, accelerating distributed databases, and handling blacklists, while noting drawbacks like no deletion and limited scalability.

Bloom filterURL deduplicationcache penetration
0 likes · 8 min read
Master Bloom Filters in 10 Minutes: How They Prevent Cache Penetration
Full-Stack Internet Architecture
Full-Stack Internet Architecture
Sep 13, 2020 · Backend Development

URL Deduplication Techniques in Java, Redis, and Databases

This article reviews six practical URL deduplication methods—including Java Set, Redis Set, database queries, unique indexes, Guava Bloom filter, and Redis Bloom filter—explaining their principles, providing complete implementation code, and recommending the most suitable approach for different system scales.

Bloom filterRedisSet
0 likes · 13 min read
URL Deduplication Techniques in Java, Redis, and Databases
Sohu Tech Products
Sohu Tech Products
Dec 5, 2018 · Backend Development

Overview of Web Crawler Types and the Architecture of the Mole Crawler System

This article explains the evolution and classification of web crawlers, describes the design and components of the Mole distributed crawler—including scheduler, fetcher, processor, rate‑limiting, URL deduplication, and Elasticsearch storage optimization—and outlines common anti‑anti‑crawling strategies.

ElasticsearchURL deduplicationWeb Crawler
0 likes · 12 min read
Overview of Web Crawler Types and the Architecture of the Mole Crawler System