Streams, Collectors & Optional

Streams, Collectors & Optional

Post 14 gave you lambdas — small functions you can pass around. This post shows where lambdas earn their keep: streams for bulk data processing, and Optional for saying "this value might be missing" without null landmines.

The stream pipeline mental model

Forget the individual method list for a moment. A stream pipeline always has exactly three parts:

Source Intermediate ops (lazy) Terminal (eager)

Hold onto one rule that explains almost everything in this post:

Intermediate operations are lazy — they describe the work. One terminal operation is eager — it does the work, in a single fused pass.

List<String> names = List.of("anita", "bob", "farah", "bo", "kunal");

long count = names.stream()              // source
    .filter(n -> n.length() >= 4)      // intermediate (lazy)
    .map(String::toUpperCase)            // intermediate (lazy)
    .sorted()                            // intermediate (lazy)
    .count();                            // terminal (eager) — NOW it runs

System.out.println(count);               // 3

Nothing happens until count() is called. The lazy ops just stack up a description of the pipeline. When the terminal runs, the stream fuses the whole pipeline into one pass over the data: element by element, filter → map, not "filter everything, then map everything". That's why the source collection can be huge or even infinite — the pipeline only pulls what it needs.

Streams are single-use

A stream can be consumed exactly once. This is not a style suggestion; it's a runtime rule.

Stream<String> s = names.stream();
s.count();
s.forEach(System.out::println);  // IllegalStateException: stream has already been operated upon or closed

Decision rule: never store a stream in a variable to use twice. Either build it fresh each time, or store the supplier — Supplier<Stream<String>> fresh = names::stream; — and call fresh.get() per use.

flatMap: the one-to-many flattener

map turns one element into one element. flatMap turns one element into a stream of elements, then flattens them all into one stream. The classic use: nested collections.

record LineItem(String sku, int qty) {}
record Order(String id, List<LineItem> items) {}

List<Order> orders = List.of(
    new Order("o1", List.of(new LineItem("A", 2), new LineItem("B", 1))),
    new Order("o2", List.of(new LineItem("A", 3)))
);

// All SKUs sold across all orders
List<String> skus = orders.stream()
    .flatMap(o -> o.items().stream())  // Order -> Stream<LineItem>, flattened
    .map(LineItem::sku)
    .distinct()
    .toList();

System.out.println(skus);   // [A, B]

When to reach for it: any time you have a list inside a list (or arrays, files-per-directory, words-per-line). map would give you Stream<List<LineItem>> — a stream of lists, which is almost never what you want.

Collectors: what you build with the results

The workhorse terminal is collect(), powered by the Collectors utility class. These six cover the vast majority of real code:

// Continuing with the orders list from the previous example:

// toList / toSet
List<String> ids = orders.stream().map(Order::id).toList();

// joining
String csv = orders.stream().map(Order::id).collect(Collectors.joining(", "));
// -> "o1, o2"

// groupingBy — the one that pays the rent
Map<String, List<LineItem>> byOrder = orders.stream()
    .flatMap(o -> o.items().stream())
    .collect(Collectors.groupingBy(LineItem::sku));

// partitioningBy — two buckets by a predicate
Map<Boolean, List<Order>> bigSmall = orders.stream()
    .collect(Collectors.partitioningBy(o -> o.items().size() > 2));

// summarizingInt — count, sum, min, max, average in one pass
IntSummaryStatistics stats = orders.stream()
    .flatMap(o -> o.items().stream())
    .collect(Collectors.summarizingInt(LineItem::qty));
System.out.println(stats);  // IntSummaryStatistics{count=3, sum=6, min=1, average=2.000000, max=3}

Decision rule: groupingBy whenever you need "bucket things by some key" — reports, dashboards, category rollups. partitioningBy when the bucket test is a yes/no. summarizingInt/Long/Double when you want all the statistics without writing five separate streams.

reduce, done carefully

reduce folds the stream into a single value: sum, product, max. The two-argument form takes an identity (the starting value) and a combiner:

int totalQty = orders.stream()
    .flatMap(o -> o.items().stream())
    .mapToInt(LineItem::qty)
    .reduce(0, Integer::sum);

System.out.println(totalQty);  // 6

Two cautions:

  • The identity must be a true identity. 0 for addition, 1 for multiplication. reduce(0, Integer::sum) on an empty stream returns 0 — so you lose the difference between "no data" and "sum is zero". If that distinction matters, use the one-argument reduce(BinaryOperator), which returns an Optional.
  • The combiner must be associative ((a + b) + c == a + (b + c)). Subtraction is not associative — reduce(0, (a, b) -> a - b) gives wrong answers, especially in parallel streams. If you need ordered, non-associative math, write a plain loop instead.

Short-circuiting: do less work, not more

Some terminals don't need the whole stream. findFirst and anyMatch stop as soon as they have an answer — and limit(n) caps how much data flows through the pipeline:

Optional<String> firstLong = names.stream()
    .filter(n -> n.length() >= 4)   // maybe expensive in real life
    .findFirst();                    // stops at "anita", skips the rest

boolean hasBo = names.stream()
    .anyMatch(n -> n.equals("bo"));  // stops at "bo"

This matters most with expensive sources: the pipeline pulls elements lazily, so with findFirst + limit + an infinite or costly source, work stops early instead of churning through everything. findFirst also returns an Optional — which leads us to the second half of this post.

Parallel streams: mostly don't

One honest paragraph. .parallelStream() (or .parallel()) splits the work across the common ForkJoinPool. It helps only when all of these hold: the dataset is large enough that splitting beats the overhead, the per-element work is CPU-bound (real computation, not waiting), the lambdas are stateless (no shared mutable variables), the source splits cheaply (ArrayList, arrays — not LinkedList), and you're not doing blocking I/O inside it (you would starve the shared pool that other code uses too). If any of those fails — small data, stateful lambdas, blocking calls — a parallel stream is slower, buggier, or both. Default to sequential; reach for parallel only with a measured performance problem, and benchmark before and after.

The two classic traps

Trap 1: a pipeline with no terminal operation does nothing. This compiles fine and silently no-ops:

names.stream()
    .filter(n -> n.length() >= 4)   // describes work...
    .map(String::toUpperCase);       // ...that never runs. No terminal!

No terminal, no execution. If your stream "isn't working", check whether a terminal is actually attached.

Trap 2: side effects in intermediate operations are unreliable. Lazy pipelines can invoke your lambda zero, one, or many times per element — fusion and short-circuiting make that unpredictable:

List<String> upper = names.stream()
    .map(n -> { System.out.println("mapping " + n); return n.toUpperCase(); })
    .limit(2)
    .toList();   // "mapping" may print once, twice, or more times — don't depend on it

Decision rule: keep lambdas in intermediate operations pure — compute and return, no printing, no mutating outside state. Side effects belong in terminal operations like forEach, where execution is explicit.

Optional: making absence explicit

Some questions have no answer: "find the first admin" in a list with none. Returning null forces every caller to remember a null-check; forgetting one is a NullPointerException at 3 a.m. Optional<T> makes "might be absent" part of the type signature, so the compiler — not your memory — enforces the check:

Optional<String> firstAdmin = names.stream()
    .filter(n -> n.startsWith("admin"))
    .findFirst();

firstAdmin.ifPresentOrElse(
    a -> System.out.println("Found: " + a),
    () -> System.out.println("No admin present"));  // runs this branch

String label = firstAdmin
    .map(String::toUpperCase)                 // Optional<String> — skips work if empty
    .orElse("DEFAULT");

String strict = firstAdmin
    .orElseThrow(() -> new IllegalStateException("admin required"));  // throws here

Note the vocabulary: orElseThrow instead of get() — never call get() without an isPresent() check first; on an empty Optional it throws NoSuchElementException, which is null-dereference bugs wearing a nicer suit. And flatMap on Optional matters when the mapping function itself returns an Optional — it avoids the awkward Optional<Optional<String>>.

Where NOT to use Optional

Optional is a return type for "value may be absent". Three places it doesn't belong:

  • Fields. Optional is not Serializable, and it clutters every access. Use a nullable field with a documented contract, or model absence with a proper type.
  • Method parameters. void register(Optional<String> email) forces callers to wrap and unwrap pointlessly. Prefer an overloaded method (register() / register(String email)) or a sentinel/nullability contract.
  • Collections. A method returning "zero or more results" should return an empty collection, never Optional<List<T>>. return List.of(); is the convention — callers get a uniform API with no unwrapping.

The one-line summary of the whole post: streams describe bulk work lazily and execute it eagerly in one pass; keep the lambdas pure, always attach a terminal, and use Optional to make "might be missing" visible in the type system — as a return type, not as a field.

Continue: Java Learning Roadmap 2026

Comments

Popular posts from this blog

JSP Servlet Interview Questions For Freshers Series 1

Java Banking Finance Services and Insurance (BFSI) domain interview questions

Java program to check even or odd number