Why Filename-Based Deduplication Fails in Knowledge‑Base Incremental Uploads
The article shows that relying on filenames for deduplication creates both duplicate entries and missed updates, explains why a two‑level identity (stable document ID and chunk ID) is required, and presents practical tests and design guidelines for reliable incremental knowledge‑base ingestion.
