Tagged articles

warmup

5 articles · Page 1 of 1
Java Architect Handbook
Java Architect Handbook
Aug 17, 2026 · Backend Development

Interview Question: What Is Flash‑Sale Warmup and Why Does It Matter?

The article explains warmup (prewarm) as the practice of moving a Java system from a cold to a hot state before traffic arrives, covering JVM JIT compilation, cache preloading, connection‑pool and thread‑pool initialization, service gray‑release, and how these steps prevent cold‑start failures in high‑concurrency scenarios.

High ConcurrencyJVMThread Pool
0 likes · 13 min read
Interview Question: What Is Flash‑Sale Warmup and Why Does It Matter?
ITPUB
ITPUB
Feb 10, 2024 · Backend Development

How to Warm Up Your Cache to Boost High‑Concurrency System Performance

Cache warming, a technique used in high‑concurrency systems, involves preloading frequently accessed data into memory before traffic spikes to improve hit rates, reduce cold‑start latency, prevent cache breakdowns, and lessen backend load, with various strategies such as startup loading, scheduled jobs, manual triggers, Redis tools, and Caffeine loaders demonstrated through Spring Boot code examples.

BackendCaffeineSpring Boot
0 likes · 10 min read
How to Warm Up Your Cache to Boost High‑Concurrency System Performance
Baobao Algorithm Notes
Baobao Algorithm Notes
Jan 14, 2022 · Artificial Intelligence

BERT Interview Q&A: Decoding CLS, Masks, Complexity, and More

An in‑depth Q&A breaks down core BERT concepts—from the purpose of the [CLS] token and masking strategies to self‑attention complexity, sparse attention tricks, subword handling of OOV words, warm‑up learning rates, GPT’s unidirectional nature, and ALBERT’s parameter sharing—providing concise explanations for each.

BERTSubword TokenizationTransformer
0 likes · 7 min read
BERT Interview Q&A: Decoding CLS, Masks, Complexity, and More
iQIYI Technical Product Team
iQIYI Technical Product Team
Nov 27, 2020 · Artificial Intelligence

Optimizing TensorFlow Serving Model Hot‑Update to Eliminate Latency Spikes in CTR Recommendation Systems

By adding model warm‑up files, separating load/unload threads, switching to the Jemalloc allocator, and isolating TensorFlow’s parameter memory from RPC request buffers, iQIYI’s engineers reduced TensorFlow Serving hot‑update latency spikes in high‑throughput CTR recommendation services from over 120 ms to about 2 ms, eliminating jitter.

Model Hot UpdateTensorFlow Servingjemalloc
0 likes · 11 min read
Optimizing TensorFlow Serving Model Hot‑Update to Eliminate Latency Spikes in CTR Recommendation Systems