18 Battle-Tested Java Coding Techniques for High-Performance, Low-CPU Systems (Part 1)

This article presents 18 practical coding techniques derived from million-QPS production systems to reduce CPU consumption and improve performance, covering type unification, string handling, loop optimization, caching strategies, and data structure selection with concrete before/after code examples.

JD Tech
JD Tech
JD Tech
18 Battle-Tested Java Coding Techniques for High-Performance, Low-CPU Systems (Part 1)

1. Reduce Data Type Conversions

Core principle: type conversions involve memory parsing, format recalculation, and byte reorganization, each consuming CPU. Reducing conversions lowers CPU instruction overhead, temporary object creation (GC pressure), and error rates (null pointers, format exceptions, precision loss).

Implementation

Unify data types across the full chain: interfaces, DB storage, memory calculation, cache storage. Example: use Long for IDs everywhere, not String + Long mix.

Ban meaningless formatting: no number-to-string just for logging/concatenation, no string-to-number just for comparison, no repeated byte-array-to-string conversions.

Use primitive types for computation: long, int, double instead of wrapper classes to avoid autoboxing/unboxing.

Cache/store raw types: Redis stores numbers, hashes, binary directly, not JSON strings requiring repeated parsing.

Serialize/deserialize only once at API entry; pass raw objects/types internally.

Use numeric types for enums/state machines; int comparison is 10-100x faster than string equals.

Key Cases

Case 1: ID Type Unification — Bad: String → Long → String round-trips. Good: keep as Long throughout, saving two conversions.

Case 2: Status/Enum — Bad: string equals checks. Good: int status = 1; if (status == 1) — numeric comparison 10-100x faster.

Case 3: Loop Conversion Ban — Bad: for(String idStr : list) Long id = Long.parseLong(idStr). Good: convert once to List<Long> before loop.

Case 4: Cache Serialization — Bad: object ↔ JSON string ↔ object. Good: serialize once or use Kryo/Protobuf (benchmark shown).

Case 5: Primitive Calculation — Bad: Long amount = 100L; if(amount > 0) triggers unboxing. Good: long amount = 100L; no object creation.

2. Simpler Types Are Better

Prefer primitive numbers, booleans, fixed-length types over complex objects, collections, long strings, nested structures. Benefits: minimal memory (fixed-size, contiguous), highest compute efficiency (CPU native ops), zero GC (stack allocation), lowest transport/storage cost (numbers vs JSON), aligns with CPU funnel model.

Implementation

Use numbers not strings for status, type, switch, ID, enum.

Use primitives not wrappers in computation, judgment, loops.

Use single variables not collections for single values.

Use boolean not 1/0, Y/N, on/off.

Use fixed types not dynamic/generic structures.

Use constants not variables/config for fixed values.

Key Cases

Case 1: Status Check — Bad: String status = "WAIT_PAY"; if("WAIT_PAY".equals(status)). Good: int status = 1; if(status == 1) or enum int.

Case 2: Boolean Flag — Bad: String flag = "Y"; if("Y".equals(flag)). Good: boolean flag = true; if(flag).

Case 3: User ID — Bad: String userId = "1001"; Long.parseLong(userId). Good: long userId = 1001.

Case 4: Cache Storage — Bad: redis.set("key", "{status:1,userId:1001}"). Good: redis.set("key", 1).

Case 5: Core Method Params — Bad: public Result calc(UserDTO dto) { if(dto.getStatus().equals("SUCCESS")) }. Good: public Result calc(long userId, int status) { if(status == 1) }.

3. Reduce String Concatenation

String immutability makes + concatenation create many temporary objects, triggering GC. Use StringBuilder, log placeholders, avoid loop concatenation.

Implementation

Ban + in loops — #1 CPU/GC killer.

Use StringBuilder for multi-line, multi-variable, loop, high-frequency concatenation; pre-size capacity.

Log with {} placeholders: log.debug("user:{} order:{}", userId, orderId) — no concatenation if log level disabled.

Return primitives not strings; avoid assembly.

Static text as final static constants, not runtime concatenation.

Large templates: use dedicated tools (JSON builders), not manual concatenation.

Key Cases

Case 1: Loop Concatenation — Bad: for loop with +. Good: StringBuilder with initial capacity.

Case 2: Logging — Bad: log.debug("user:" + userId). Good: log.debug("user:{}", userId).

Case 3: High-Frequency Key Building — Bad: return "user:info:" + userId. Good: StringBuilder with pre-sized capacity.

Case 4: Multi-Variable Concatenation — Bad: type + ":" + id + ":" + status. Good: new StringBuilder().append(type).append(":").append(id)...

Case 5: Static Constants — Bad: public static final String KEY = "user" + ":" + "info". Good: "user:info".

4. Reduce Repeated Data Calls

Same request/method/period: fetch identical data/query results once, cache and reuse. Eliminates duplicate DB/RPC calls, repeated calculations, parsing.

Implementation

Method-local reuse: store in local variable after first call.

Request-level reuse: ThreadLocal/RequestContext (author prefers global Context over ThreadLocal).

Local cache (Caffeine/Map) for semi-static data (config, dict, user basics) with TTL.

Distributed cache for cross-service shared data.

Batch query instead of loop single-query: one SELECT ... WHERE id IN (...) then in-memory map lookup.

Key Cases

Case 1: Method-Local Reuse — Bad: multiple userService.getUser(userId) calls. Good: User user = userService.getUser(userId); reuse variable.

Case 2: Loop Batch Query — Bad: for(orderId : list) orderMapper.selectById(orderId). Good: List<Order> orders = orderMapper.selectBatchIds(orderIdList); then map by ID.

5. Split Multiple Judgments, Fail Fast

Break complex validation into ordered, independent, front-loaded checks. Invalid requests terminate early, reducing wasted computation, resource usage, and improving readability/throughput.

Implementation

Order checks: light → heavy (null/format/range first, then business/DB/RPC).

Each check on its own line, no nesting; immediate return/throw on failure.

Happy path at top level, no indentation.

Ban "wrapping big if" that nests all logic.

Common validation utilities throw exceptions directly, don't return status codes.

Key Cases

Case 1: Deep Nesting → Flat Checks — Bad: nested if-else 4 levels. Good: sequential if checks with early returns.

Case 2: High-Concurrency RPC Order — Bad: RPC call first, then business validation. Good: validate params first, fail fast before RPC.

Case 3: Loop Fail-Fast — Check CollectionUtils.isEmpty(list) before loop; inside loop, null/id checks with immediate return false.

6. Array Traversal Over Collections

Arrays (and ArrayList) use contiguous memory, CPU cache-friendly, O(1) random access, zero iterator overhead, zero GC. Outperform LinkedList/iterators by 10-100x in high concurrency.

Implementation

Prefer arrays for fixed-length, primitive data.

Use ArrayList not LinkedList for frequent traversal.

Use classic for-loop (index) not iterator/enhanced-for; cache length: int size = list.size(); for(int i=0;i<size;i++).

Convert non-array collections to array before traversal.

Primitives: long[] > Long[] > List<Long>.

Core paths (seckill, order, payment) must use array traversal.

Key Cases

Example 1: int[] arr = new int[100]; for(int i=0;i<arr.length;i++) { ... } — fastest.

Example 2: ArrayList + classic for — good.

Example 3: LinkedList + iterator — banned in high concurrency.

Example 4: Seckill inventory traversal with array.

Example 5: Large batch matching: array traversal 50x faster than LinkedList.

ArrayList for vs Iterator: Classic for 10-30% faster; no iterator object, no hasNext/next calls, no boundary checks per iteration. JIT optimizes classic for to near-native instructions.

7. Ban forEach/Lambda/Stream in High-Concurrency Paths

Lambda creates functional interface instances, iterators, Consumer objects — heavy GC. Virtual calls prevent JIT optimization. Enhanced-for = iterator. Classic for on array/ArrayList is zero-allocation, direct indexing, JIT-friendly.

Implementation

Core paths (seckill, order, payment, inventory, hot queries): mandatory classic for.

Ban: list.forEach(...), list.stream().forEach(...), for(Item item : list).

Cache loop length outside condition.

Prefer primitive arrays: long[], int[] over List<Long> + Lambda.

Batch/large loops (10K-1M): classic for mandatory; Lambda overhead 300-1000%.

Key Cases

Example 1 (Banned): orderList.forEach(order -> doProcess(order)).

Example 2 (Banned): orderList.stream().forEach(order -> doProcess(order)) — slowest.

Example 3: High-concurrency inventory check with classic for on array.

8. Reduce String Splitting

split() involves char traversal, regex compilation, array allocation, substring creation — high CPU, many temp objects, GC pressure. Avoid in loops/hot paths.

Implementation

Fixed-format data: use substring(index) instead of split.

Pre-structure storage: separate fields/entities/arrays instead of concatenated strings.

Ban split in loops, batch jobs, hot interfaces.

If splitting needed: use lightweight char-matching (indexOf + substring), not regex split.

API/config: pass structured objects/JSON/multi-field, not concatenated strings.

Single split per request: parse once, store in local variable, reuse.

Key Cases

Example 1 (Banned): Frequent split in hot path.

Example 2: Fixed-format: substring by index instead of split.

Example 3: Loop split → pre-parse to structured objects before loop.

Example 4: Special char regex trap: split(".") compiles regex; use split("\\.") or indexOf+substring.

Example 5: Source-level fix: pass multiple parameters instead of concatenated string requiring split.

9. Optimize Date Parsing

Date parsing, current time, time arithmetic are high-frequency CPU ops. Reuse formatters, avoid heavy time objects, reduce syscalls, avoid timezone recalculation, reuse immutable instances.

Implementation

Parsing: Discard SimpleDateFormat (unsafe, slow). Use static final DateTimeFormatter with fixed zone (Asia/Shanghai). Parse in batch outside loops.

Current Time: Performance: System.currentTimeMillis() > Instant.now() > LocalDateTime.now() > new Date(). Use long timestamp in loops/high-concurrency; fetch once per request/batch.

Time Arithmetic: Discard Calendar. Best: timestamp long arithmetic (pure math). Second: LocalDateTime.plusHours/minusDays (immutable, lightweight).

Key Cases

Global static formatter example with ZoneId. Wrong: loop creates new formatter + parse each iteration. Right: static formatter, batch parse outside loop. Time arithmetic: now + 30*60*1000L vs LocalDateTime.plusHours(2).

10. Reduce Exception Handling Overhead

Exception creation captures full stack trace — heavy CPU. try-catch adds bytecode branches, hinders JIT. Exceptions for "unexpected", not business branching.

Implementation

Pre-validate: null, range, format, boundary checks upfront; return/throw early.

Ban try-catch in loops; single outer catch (except batch partial-failure cases).

Custom business exceptions: disable stack trace filling.

Happy path mainline, no try-catch; global handler at top.

Minimize active throws; return boolean/enum for validation/parsing.

Avoid nested try-catch.

Key Cases

Example 1: Bad: try { Integer.parseInt(str) } catch -> isNumber. Good: pre-check with regex or Apache Commons NumberUtils.isCreatable.

Example 2: Bad: loop with try-catch per iteration. Good: loop pure logic, outer try-catch.

Example 3: Bad: high-frequency throw new BizException. Good: return boolean false.

Example 4: Custom exception overriding fillInStackTrace() to return this (no stack trace).

RPC Template: Unified RPCUtils.invoke(() -> rpc.call()) handles null, code!=0, data null, exceptions once. Eliminates duplicated checks/try-catch per RPC call. Benefits: single JIT-optimized code, 30-50% faster due to branch prediction, lower GC, cleaner business code.

11. Avoid Duplicate Code

Duplicate logic = duplicate CPU instructions, JIT cannot optimize scattered copies, repeated object creation, branch prediction failures. Extract common checks, calculations, RPC validation, loop-invariant code, exception handling into static utilities/templates.

Implementation

Common checks → static utility methods.

Common calculations → static methods.

RPC calls → unified wrapper (as in #10).

Loop-invariant code → hoist outside loop.

Repeated object creation → singleton/static reuse.

Exception handling → single outer layer.

Key Cases

Example 1: Three methods each duplicate null/code/data checks → extract to RPCUtils.invoke.

Example 2: Loop creates SimpleDateFormat and calls list.size() each iteration → hoist both outside.

12. Use Set.contains Over Chained equals

HashSet.contains() is O(1) hash lookup; chained equals is O(n) sequential with multiple method calls and branch mispredictions.

Implementation

Fixed whitelists/status sets → static final HashSet initialized once.

Replace eq1 || eq2 || eq3 with set.contains(value).

Must use HashSet, not ArrayList (List.contains is O(n) traversal).

Elements must be immutable constants.

Null-check before contains.

Best for small fixed sets (5-20 items).

Key Cases

Example 1: Bad: status.equals("A") || status.equals("B") || ... Good: static final Set<String> VALID = Set.of("A","B","C"); VALID.contains(status).

Example 2: Numeric status codes with HashSet<Integer>.

Example 3 (Wrong): List.contains — no performance gain.

13. Convert Multi-Level if-else to O(1) Map Lookup

if-else chain is O(n) sequential with branch mispredictions; HashMap.get(key) is O(1) hash lookup, zero branches, JIT-friendly.

Implementation

Static final Map built at class load (read-only, no sync, HashMap faster than ConcurrentHashMap for read-only).

Key = condition (status, type, enum); Value = result or strategy (Lambda/Function).

Runtime: single map.get(key), zero if-else.

Only worthwhile for 3+ branches, high-frequency calls.

Handle null key/value upfront.

Key Cases

Example 1: Order status mapping: 5-level if-else → static Map<Status, String> STATUS_DESC; return map.get(status). Performance up 50-200%, stable under load.

14. Avoid Unnecessary Object Creation

Object creation: class check, memory alloc, header init, zero-fill, reference bind — many CPU instructions. Frequent short-lived objects → frequent Minor GC → STW pauses. Each object has header/alignment overhead. Reusable objects (formatters, collections, constants) should be created once. Primitives avoid wrapper overhead.

Implementation

High-frequency utilities → static final singletons (DateTimeFormatter, Pattern, fixed Sets/Maps).

Loop-invariant objects → hoist outside loop.

Prefer primitives: int/long/double over Integer/Long/Double.

String constants → static final, avoid runtime new String().

Collection reuse: clear() and reuse instead of new ArrayList().

Eliminate autoboxing in math/conditions.

Ban split/substring/format in loops; batch parse beforehand.

Key Cases

Example 1: Loop creates SimpleDateFormat + calls list.size() → hoist both.

Example 2: Loop: Long num = i (autobox) → use long num = i.

Example 3: Method creates new HashSet each call → static final Set.

Example 4: Loop: String msg = "ID:" + i → static final prefix + i (still creates String but avoids double concat).

15. Use More Efficient Data Structures (fastutil)

JDK containers (HashMap, ArrayList) store wrappers (Integer/Long) — 16B header each, hash collisions, linked lists/trees, load factor, boxing/unboxing, low memory density. fastutil provides primitive-specialized collections: direct primitive storage, pure arrays, no hashing overhead, 50-90% less memory, 3-15x speed, zero GC, 100% CPU cache hit.

Key Replacements

HashMap<Integer, T> → TIntObjectHashMap<T>

HashMap<Long, T> → TObjectLongHashMap<T> (or TLongObjectHashMap)

Integer/Long → int/long

Set<Integer> → TIntHashSet

List<Integer> → TIntArrayList

Enum-keyed Map → EnumMap (JDK, pure array, 10-20x HashMap)

Maven dependency: it.unimi.dsi:fastutil:8.5.12

16. Reduce Null Fields

Null fields still occupy reference slots (4B on 64-bit compressed oops). Object with all-null fields still consumes header + all field slots + padding. Only making the whole object null (no instance) saves heap.

Key Cases

Scenario 1: User object with String name, Integer age, List list — all null. Still 28B (header 12 + 3*4 refs + 4 padding).

Scenario 2: String s = null (4B ref) vs String s = "" (4B ref + shared empty string). int a = 0 (4B) vs Integer a = null (4B ref).

Scenario 3: User user = null — only 4B reference, no heap object.

Scenario 4: new User() with all null fields — full object allocated with all field slots.

Summary: field-level null saves nothing; object-level null saves everything. Benefits: fewer null checks, no NPE, better branch prediction, better JIT, direct collection/string usage.

17. Prefer keySet Over entrySet When Only Keys Needed

entrySet() creates Map.Entry object per iteration (key+value wrapper) → massive temp objects, GC pressure. keySet() returns keys directly, zero extra objects. Never do keySet + map.get(key) — double hash lookup = O(n²).

Implementation

Only need keys → keySet().

Need key+value → entrySet().

Never: for(key : map.keySet()) value = map.get(key).

Key Cases

Example 1 (Bad): Only need user IDs, but uses entrySet → creates Entry per loop.

Example 1 (Good): map.keySet() — zero objects, 30-100% faster.

Example 2: Need both → entrySet (Entry holds value reference, no extra lookup).

Example 3 (Disaster): keySet + map.get(key) — two full hash lookups per iteration.

18. Reduce JSON String Parameter Passing

JSON is text protocol: serialize (object→string) + transmit large string + deserialize (parse+build object). Binary protocol (protobuf, Hessian, JDK serialization) sends compact bytes, 50-100x lower overhead. Never use JSON for intra-service method calls.

Comparison

JSON String: Text protocol, Object→JSON serialization, verbose text payload, strong cross-language, average performance.

Object Transfer: Binary protocol, Object→bytes serialization, compact binary payload, weak cross-language, extreme performance.

JSON costs: 1 serialization + 1 deserialization + 1 large String object = 50-100x direct object passing. Ban JSON for internal A→B calls.

Code example

int
status =
1
;
if
(status ==
1
) { ... }
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

JavaPerformance OptimizationMemory Optimizationbackend developmentCPU OptimizationHigh ConcurrencyCoding Best PracticesGC Reduction
JD Tech
Written by

JD Tech

Official JD technology sharing platform. All the cutting‑edge JD tech, innovative insights, and open‑source solutions you’re looking for, all in one place.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.