Streams, Collectors & Optional
Streams, Collectors & Optional
Post 14 gave you lambdas — small functions you can pass around. This post shows where lambdas earn their keep: streams for bulk data processing, and Optional for saying "this value might be missing" without null landmines.
The stream pipeline mental model
Forget the individual method list for a moment. A stream pipeline always has exactly three parts:
Hold onto one rule that explains almost everything in this post:
Intermediate operations are lazy — they describe the work. One terminal operation is eager — it does the work, in a single fused pass.
List<String> names = List.of("anita", "bob", "farah", "bo", "kunal");
long count = names.stream() // source
.filter(n -> n.length() >= 4) // intermediate (lazy)
.map(String::toUpperCase) // intermediate (lazy)
.sorted() // intermediate (lazy)
.count(); // terminal (eager) — NOW it runs
System.out.println(count); // 3
Nothing happens until count() is called. The lazy ops just stack up a description of the pipeline. When the terminal runs, the stream fuses the whole pipeline into one pass over the data: element by element, filter → map, not "filter everything, then map everything". That's why the source collection can be huge or even infinite — the pipeline only pulls what it needs.
Streams are single-use
A stream can be consumed exactly once. This is not a style suggestion; it's a runtime rule.
Stream<String> s = names.stream();
s.count();
s.forEach(System.out::println); // IllegalStateException: stream has already been operated upon or closed
Decision rule: never store a stream in a variable to use twice. Either build it fresh each time, or store the supplier — Supplier<Stream<String>> fresh = names::stream; — and call fresh.get() per use.
flatMap: the one-to-many flattener
map turns one element into one element. flatMap turns one element into a stream of elements, then flattens them all into one stream. The classic use: nested collections.
record LineItem(String sku, int qty) {}
record Order(String id, List<LineItem> items) {}
List<Order> orders = List.of(
new Order("o1", List.of(new LineItem("A", 2), new LineItem("B", 1))),
new Order("o2", List.of(new LineItem("A", 3)))
);
// All SKUs sold across all orders
List<String> skus = orders.stream()
.flatMap(o -> o.items().stream()) // Order -> Stream<LineItem>, flattened
.map(LineItem::sku)
.distinct()
.toList();
System.out.println(skus); // [A, B]
When to reach for it: any time you have a list inside a list (or arrays, files-per-directory, words-per-line). map would give you Stream<List<LineItem>> — a stream of lists, which is almost never what you want.
Collectors: what you build with the results
The workhorse terminal is collect(), powered by the Collectors utility class. These six cover the vast majority of real code:
// Continuing with the orders list from the previous example:
// toList / toSet
List<String> ids = orders.stream().map(Order::id).toList();
// joining
String csv = orders.stream().map(Order::id).collect(Collectors.joining(", "));
// -> "o1, o2"
// groupingBy — the one that pays the rent
Map<String, List<LineItem>> byOrder = orders.stream()
.flatMap(o -> o.items().stream())
.collect(Collectors.groupingBy(LineItem::sku));
// partitioningBy — two buckets by a predicate
Map<Boolean, List<Order>> bigSmall = orders.stream()
.collect(Collectors.partitioningBy(o -> o.items().size() > 2));
// summarizingInt — count, sum, min, max, average in one pass
IntSummaryStatistics stats = orders.stream()
.flatMap(o -> o.items().stream())
.collect(Collectors.summarizingInt(LineItem::qty));
System.out.println(stats); // IntSummaryStatistics{count=3, sum=6, min=1, average=2.000000, max=3}
Decision rule: groupingBy whenever you need "bucket things by some key" — reports, dashboards, category rollups. partitioningBy when the bucket test is a yes/no. summarizingInt/Long/Double when you want all the statistics without writing five separate streams.
reduce, done carefully
reduce folds the stream into a single value: sum, product, max. The two-argument form takes an identity (the starting value) and a combiner:
int totalQty = orders.stream()
.flatMap(o -> o.items().stream())
.mapToInt(LineItem::qty)
.reduce(0, Integer::sum);
System.out.println(totalQty); // 6
Two cautions:
- The identity must be a true identity.
0for addition,1for multiplication.reduce(0, Integer::sum)on an empty stream returns0— so you lose the difference between "no data" and "sum is zero". If that distinction matters, use the one-argumentreduce(BinaryOperator), which returns anOptional. - The combiner must be associative (
(a + b) + c == a + (b + c)). Subtraction is not associative —reduce(0, (a, b) -> a - b)gives wrong answers, especially in parallel streams. If you need ordered, non-associative math, write a plain loop instead.
Short-circuiting: do less work, not more
Some terminals don't need the whole stream. findFirst and anyMatch stop as soon as they have an answer — and limit(n) caps how much data flows through the pipeline:
Optional<String> firstLong = names.stream()
.filter(n -> n.length() >= 4) // maybe expensive in real life
.findFirst(); // stops at "anita", skips the rest
boolean hasBo = names.stream()
.anyMatch(n -> n.equals("bo")); // stops at "bo"
This matters most with expensive sources: the pipeline pulls elements lazily, so with findFirst + limit + an infinite or costly source, work stops early instead of churning through everything. findFirst also returns an Optional — which leads us to the second half of this post.
Parallel streams: mostly don't
One honest paragraph. .parallelStream() (or .parallel()) splits the work across the common ForkJoinPool. It helps only when all of these hold: the dataset is large enough that splitting beats the overhead, the per-element work is CPU-bound (real computation, not waiting), the lambdas are stateless (no shared mutable variables), the source splits cheaply (ArrayList, arrays — not LinkedList), and you're not doing blocking I/O inside it (you would starve the shared pool that other code uses too). If any of those fails — small data, stateful lambdas, blocking calls — a parallel stream is slower, buggier, or both. Default to sequential; reach for parallel only with a measured performance problem, and benchmark before and after.
The two classic traps
Trap 1: a pipeline with no terminal operation does nothing. This compiles fine and silently no-ops:
names.stream()
.filter(n -> n.length() >= 4) // describes work...
.map(String::toUpperCase); // ...that never runs. No terminal!
No terminal, no execution. If your stream "isn't working", check whether a terminal is actually attached.
Trap 2: side effects in intermediate operations are unreliable. Lazy pipelines can invoke your lambda zero, one, or many times per element — fusion and short-circuiting make that unpredictable:
List<String> upper = names.stream()
.map(n -> { System.out.println("mapping " + n); return n.toUpperCase(); })
.limit(2)
.toList(); // "mapping" may print once, twice, or more times — don't depend on it
Decision rule: keep lambdas in intermediate operations pure — compute and return, no printing, no mutating outside state. Side effects belong in terminal operations like forEach, where execution is explicit.
Optional: making absence explicit
Some questions have no answer: "find the first admin" in a list with none. Returning null forces every caller to remember a null-check; forgetting one is a NullPointerException at 3 a.m. Optional<T> makes "might be absent" part of the type signature, so the compiler — not your memory — enforces the check:
Optional<String> firstAdmin = names.stream()
.filter(n -> n.startsWith("admin"))
.findFirst();
firstAdmin.ifPresentOrElse(
a -> System.out.println("Found: " + a),
() -> System.out.println("No admin present")); // runs this branch
String label = firstAdmin
.map(String::toUpperCase) // Optional<String> — skips work if empty
.orElse("DEFAULT");
String strict = firstAdmin
.orElseThrow(() -> new IllegalStateException("admin required")); // throws here
Note the vocabulary: orElseThrow instead of get() — never call get() without an isPresent() check first; on an empty Optional it throws NoSuchElementException, which is null-dereference bugs wearing a nicer suit. And flatMap on Optional matters when the mapping function itself returns an Optional — it avoids the awkward Optional<Optional<String>>.
Where NOT to use Optional
Optional is a return type for "value may be absent". Three places it doesn't belong:
- Fields.
Optionalis notSerializable, and it clutters every access. Use a nullable field with a documented contract, or model absence with a proper type. - Method parameters.
void register(Optional<String> email)forces callers to wrap and unwrap pointlessly. Prefer an overloaded method (register()/register(String email)) or a sentinel/nullability contract. - Collections. A method returning "zero or more results" should return an empty collection, never
Optional<List<T>>.return List.of();is the convention — callers get a uniform API with no unwrapping.
The one-line summary of the whole post: streams describe bulk work lazily and execute it eagerly in one pass; keep the lambdas pure, always attach a terminal, and use Optional to make "might be missing" visible in the type system — as a return type, not as a field.
Continue: Java Learning Roadmap 2026
Comments
Post a Comment