Java Stream API in Practice: From 80 Lines to 30 — Deduplication, Grouping, Sorting & Statistics

A hands-on guide demonstrating how Java Stream API reduces a real-world reporting task from 78 lines to 17, covering deduplication strategies, multi-level sorting pitfalls, grouping with downstream collectors, and one-pass statistics using IntSummaryStatistics.

Java Tech Workshop
Java Tech Workshop
Java Tech Workshop
Java Stream API in Practice: From 80 Lines to 30 — Deduplication, Grouping, Sorting & Statistics

A Real Requirement That Shrinks 80 Lines to 30

The article opens with a concrete reporting scenario: fetch 2,000 employee records once, then in memory (1) deduplicate by employee number keeping the latest record, (2) filter out probationary employees, (3) group by department, (4) compute per-department headcount, average salary, max salary, and total salary, (5) sort members within each department by salary descending, and (6) output the top 3 departments by average salary descending, each showing only its top 3 earners.

Traditional vs. Stream Implementation

The traditional approach spans 78 lines with four intermediate containers, three nested loops, and numerous temporary variables. The Stream version accomplishes the same in 17 lines:

public List<DeptReport> buildReport(List<Employee> list) {
    return list.stream()
        .filter(e -> Boolean.FALSE.equals(e.getProbation()))
        .collect(Collectors.toMap(
            Employee::getEmpNo,
            Function.identity(),
            (a, b) -> a.getJoinDate().isAfter(b.getJoinDate()) ? a : b,
            LinkedHashMap::new))
        .values().stream()
        .collect(Collectors.groupingBy(Employee::getDept))
        .entrySet().stream()
        .map(e -> {
            IntSummaryStatistics st = e.getValue().stream()
                .collect(Collectors.summarizingInt(Employee::getSalary));
            List<String> top3 = e.getValue().stream()
                .sorted(Comparator.comparingInt(Employee::getSalary).reversed())
                .limit(3)
                .map(m -> m.getName() + "(" + m.getSalary() + ")")
                .collect(Collectors.toList());
            return new DeptReport(e.getKey(), st.getCount(), st.getSum(),
                st.getMax(), st.getAverage(), top3);
        })
        .sorted(Comparator.comparingDouble(DeptReport::getAvg).reversed())
        .limit(3)
        .collect(Collectors.toList());
}

Six Core Concepts Explained with Pitfalls

1. Stream Pipeline Anatomy

Every stream is a three-stage pipeline: Source → Intermediate Operations → Terminal Operation. The author emphasizes identifying stages by return type: operations returning Stream<T> are intermediate (lazy), others are terminal (eager). A table lists typical methods for each stage.

2. Lazy Evaluation

Without a terminal operation, no intermediate operation executes. Example:

Stream<String> stream = Stream.of("张", "李", "王")
    .filter(s -> { System.out.println("filter 执行了: " + s); return true; });
System.out.println("=== 到这里还没动 ==="); // only this prints
stream.count(); // now all three filter logs appear
peek()

is strictly for debugging; business logic inside it never runs without a terminal operation.

3. One-Time Consumption

A stream cannot be reused. Calling count() twice on the same stream throws IllegalStateException. Solution: invoke collection.stream() anew each time.

4. No Side Effects on Source

sorted()

returns a new stream; the original list remains unchanged. For in-place sorting use List.sort(). Warning: mutating element fields inside lambdas (e.g., e.setSalary(...)) does affect the original objects because streams hold references.

Deduplication: Beyond distinct()

Four approaches compared:

distinct() — uses equals()/hashCode(), keeps first occurrence. Pitfalls: fails silently if equals/hashCode not overridden; BigDecimal scale differences ( 100.0 vs 100.00) break equality.

toMap with merge function — most common in business. Four-parameter overload: key mapper, value mapper, merge function (required, else IllegalStateException on duplicate keys), map factory (use LinkedHashMap::new to preserve order). Value cannot be null.

filter + external Set — enables early limit and mid-pipeline deduplication. Stateful lambda; avoid with parallelStream() unless using concurrent set.

TreeSet collector — deduplicates by compare() == 0 (not equals) while sorting. Requires nullsLast for nullable fields.

A decision table summarizes when to use each.

Sorting: Comparator Chaining Traps

Prefer comparingInt/Long/Double over generic comparing to avoid boxing overhead.

Classic trap: reversed() at the end of a chain flips all preceding conditions. Correct: wrap only the specific condition:

thenComparing(Comparator.comparingInt(Employee::getSalary).reversed())

.

Null handling: Comparator.nullsLast(Comparator.reverseOrder()) or nullsLast(Comparator.comparing(...)).

Chinese sorting: use Collator.getInstance(Locale.CHINA) instead of default Unicode order. List.sort() mutates in place; stream().sorted() caches all elements in memory.

Grouping: groupingBy Deep Dive

Basic: groupingBy(Employee::getDept) returns HashMap (unordered). Add LinkedHashMap::new for insertion order.

Null keys throw NPE; pre-process with Optional.ofNullable(dept).orElse("未分配").

Downstream collectors (2nd arg) avoid retaining full lists: counting(), summingInt, averagingInt, maxBy/minBy, mapping, filtering (JDK 9+), summarizingInt (count/sum/avg/max/min in one pass).

Multi-level grouping: nest groupingBy calls. Beyond two levels, prefer a composite key record DeptLevel(String dept, String level). partitioningBy for boolean splits — guarantees both true and false keys present. collectingAndThen for post-processing (e.g., keep top 3 per group). teeing (JDK 12+) runs two collectors in one pass and merges results (e.g., max + min).

Statistics: One-Pass Metrics

Primitive streams ( mapToInt) provide sum(), average(), max(), min(), summaryStatistics() without boxing. Collectors.summarizingInt(Employee::getSalary) returns IntSummaryStatistics with getCount(), getSum(), getAverage(), getMax(), getMin().

Precision traps: summingInt returns Integer (overflow risk for large sums — use summingLong); averagingInt returns Double (compare with epsilon).

No BigDecimal support in summarizing collectors. Use reduce(BigDecimal.ZERO, BigDecimal::add) for sums; compute average manually with divide and RoundingMode. For grouped BigDecimal sums, use groupingBy with reducing.

Always unwrap Optional via orElse / orElseGet; never call get() directly.

Complete Runnable Demo

The article provides a full StreamDemo class with sample data (7 employees, including duplicate E001 with newer record and one probationary employee). Output confirms correctness: Zhang San replaced by 32000 record; Wang Wu (probation) excluded; departments ordered by average salary descending with top 3 members each.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

statisticsdeduplicationsortingComparatorCollectorsgroupingByIntSummaryStatisticsJava Stream API
Java Tech Workshop
Written by

Java Tech Workshop

Focused on Java backend technologies, sharing fundamentals, multithreading, JVM, the Spring ecosystem, microservices, distributed systems, high concurrency, source‑code analysis, and practical experience. Continuously delivers high‑quality original content, interview guides, and learning roadmaps to help Java developers progress from beginner to advanced, enhancing technical skills and core competitiveness.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.