Files, I/O & NIO: Reading and Writing Data
Almost every program you've written so far has worked with data that vanishes when the program ends. Real programs keep data around: they read configuration files, load CSV exports, write logs, generate reports. Java has two file APIs — the 1996 original (java.io) and the modern java.nio.file package, introduced in Java 7 and known as NIO.2. In 2026, NIO.2 is the default: it is shorter, safer, and harder to misuse. You'll still meet java.io in legacy code and libraries, so this post teaches NIO.2 first and gives you just enough java.io to read old code without fear.
This is the final post of the Core Java track — it leans on exceptions and try-with-resources (post 10), collections (post 11), and streams (post 15). The next track, Tooling & Testing, assumes you can read and write files comfortably.
The decision rule: which API do I reach for?
Writing new code? Use
java.nio.file(Path+Files). Reading old code or a library's streams? Learn thejava.ioshapes below, then wrap or convert.
What breaks if you default to java.io? Mostly verbosity and traps: new FileReader(path) silently uses the platform's default charset (your file reads fine on your Mac and corrupts on a production Linux box), resource leaks unless you remember to close everything by hand, and directory traversal that requires recursive helper methods. NIO.2 fixes all three: explicit charsets, one-line whole-file reads, and Files.walk for trees. The only common reason to use java.io directly today is when an API hands you a stream — then you wrap it, which we'll cover.
Paths: how Java names a file
A Path is just a structured name for a file or directory — it does not have to exist yet. You build one with Path.of(...) (the older Paths.get(...) does the same thing):
import java.nio.file.Path;
Path config = Path.of("config", "app.properties"); // relative: config/app.properties
Path home = Path.of(System.getProperty("user.home"), "data", "sales.csv"); // absolute-ish
Path here = Path.of(".").toAbsolutePath().normalize(); // where am I?
System.out.println(config.toAbsolutePath());
System.out.println(config.getFileName()); // app.properties
System.out.println(config.getParent()); // config
System.out.println(home.resolve("archive")); // join: .../data/sales.csv/archive
Two habits that prevent real bugs:
- Resolve relative paths explicitly. A relative
Pathresolves against the JVM's working directory, which is wherever the program was launched — your IDE, a terminal, a Docker container, and a CI server can all disagree. For anything important, anchor on a known base (user.home, an env var, a config value) withresolve. normalize()away the noise.Path.of("data", "..", "data", "sales.csv").normalize()collapses todata/sales.csv. Useful when joining user input or config values into paths.
Checking what's there: Files
Path names things; Files does things. Before touching a file, inspect it:
import java.nio.file.*;
Path p = Path.of("data", "sales.csv");
if (Files.notExists(p)) {
System.out.println("Missing: " + p.toAbsolutePath());
} else if (Files.isDirectory(p)) {
System.out.println("It's a directory");
} else if (Files.isRegularFile(p)) {
System.out.println("File, " + Files.size(p) + " bytes");
}
Note Files.notExists rather than !Files.exists: if the JVM can't even check (permissions), both exists and notExists return false — neither claims knowledge it doesn't have. In that ambiguous case, notExists is the honest negative.
Reading a whole file in one line
For config files, templates, and small data files, this is the whole story:
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
String text = Files.readString(Path.of("config", "app.properties"), StandardCharsets.UTF_8);
System.out.println(text);
Output:
app.name=CityOps
app.region=ap-south-1
retry.max=3
The platform-default charset trap
Files are bytes; Strings are characters; something has to translate. That something is a charset, and the trap is letting Java pick it for you. The classic new FileReader(path) uses the platform default — UTF-8 on most Linux/macOS systems, but historically windows-1252 on Windows. A file written as UTF-8 and read as windows-1252 turns café into café, and it works perfectly on your machine while corrupting data in production.
The rule is simple and absolute: always pass StandardCharsets.UTF_8 explicitly when converting between bytes and text. (Files.readString without a charset argument happens to default to UTF-8, but sibling methods like Files.readAllLines default to the platform charset — so be explicit everywhere and never think about it again.)
Writing files: create, overwrite, append
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
import static java.nio.file.StandardOpenOption.*;
Path report = Path.of("out", "summary.txt");
Files.createDirectories(report.getParent()); // create out/ if needed — no error if it exists
// Default: CREATE the file if missing, TRUNCATE it if present, then WRITE.
Files.writeString(report, "CityOps daily summary\n", StandardCharsets.UTF_8);
// Append a line instead of wiping the file:
Files.writeString(report, "orders=42\n", StandardCharsets.UTF_8, CREATE, APPEND);
Common option combinations to know:
CREATE— create the file if it doesn't exist.TRUNCATE_EXISTING— wipe existing content (this is the default behavior ofwriteString, made explicit).APPEND— add to the end instead of wiping.CREATE_NEW— create, but fail if the file already exists (throwsFileAlreadyExistsException). Use this when overwriting would mean losing data — e.g. generating a report that must never silently replace yesterday's.
What breaks if you misuse these? Writing logs with the default truncate behavior deletes your log on every restart — that's the APPEND use case. Conversely, appending to a report file that should be regenerated fresh gives you doubled numbers — that's the truncate use case. Match the option to the intent.
Big files: stream the lines
readString loads everything into memory. For a multi-gigabyte log file, that's an OutOfMemoryError. Files.lines gives you a Stream<String> — one line at a time, lazily — which plugs straight into everything from post 15:
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
import java.util.stream.Stream;
Path log = Path.of("logs", "app.log");
try (Stream<String> lines = Files.lines(log, StandardCharsets.UTF_8)) {
long errors = lines
.filter(line -> line.contains("ERROR"))
.count();
System.out.println("error lines: " + errors);
} // the stream (and the file) closes here, even on exception
Output:
error lines: 17
The try-with-resources is not optional decoration: Files.lines opens a file handle, and streams don't close themselves. This is the same rule from post 10 applied to I/O — anything that opens a resource gets declared in the try header. Forget it on a long-running server and you leak file descriptors until the OS refuses to open anything.
Walking directory trees
Listing one directory is Files.list(dir) (also a stream, also needs try-with-resources). Recursing through a tree used to mean writing your own recursion; now it's one call:
import java.io.IOException;
import java.nio.file.*;
import java.util.stream.Stream;
try (Stream<Path> tree = Files.walk(Path.of("data"))) {
tree.filter(Files::isRegularFile)
.filter(p -> p.toString().endsWith(".csv"))
.forEach(System.out::println);
}
Output:
data/sales.csv
data/archive/sales-2026-09.csv
data/archive/sales-2026-08.csv
Two things to know: Files.walk(dir, maxDepth) caps how deep it goes (Files.walk(dir, 1) is just the top level), and Files.walk follows the stream lazily — the tree isn't fully loaded into memory first. If you only need one level and want file attributes cheaply, Files.list is the lighter call; walk is for real recursion.
Copying, moving, deleting
import java.nio.file.*;
import static java.nio.file.StandardCopyOption.*;
Path src = Path.of("data", "sales.csv");
// Copy: needs REPLACE_EXISTING or it throws if the target exists.
Files.copy(src, Path.of("backup", "sales.csv"), REPLACE_EXISTING);
// Move/rename: ATOMIC_MOVE guarantees all-or-nothing on filesystems that support it.
Files.move(src, Path.of("data", "archive", "sales.csv"), ATOMIC_MOVE);
// Delete one file; throws NoSuchFileException if it's already gone.
Files.deleteIfExists(Path.of("tmp", "scratch.txt"));
Decision rules: prefer deleteIfExists over delete unless a missing file is genuinely a bug you want to hear about. Prefer CREATE_NEW / no-REPLACE_EXISTING when overwriting would destroy data; use REPLACE_EXISTING deliberately when the target is disposable (a backup, a temp file). There is no recursive delete in Files — deleting a directory tree means walking it in reverse order (files before their parent directories) or using a small utility; the absence is deliberate, because recursive delete is the operation most likely to ruin someone's day.
Classic java.io: what you're looking at in old code
You don't need to write new code with java.io, but you will read plenty of it. The whole package is two ladders:
- Bytes ladder —
InputStream/OutputStream: raw bytes.FileInputStreamreads a file's bytes;BufferedInputStreambatches reads for speed. You'll still see this ladder for non-text data: images, zip files, and wrappers likeGZIPInputStream. - Chars ladder —
Reader/Writer: bytes decoded to characters.FileReaderis the charset trap — it uses the platform default.BufferedReader.readLine()is the classic line-by-line reader. - Bridges between ladders:
InputStreamReaderwraps bytes with an explicit charset;OutputStreamWriterdoes the reverse.
The translation rule: BufferedReader br = new BufferedReader(new FileReader(p)) becomes Files.newBufferedReader(p, StandardCharsets.UTF_8) — same shape, explicit charset, NIO path handling. When a library hands you an InputStream (a common case: HTTP responses, compressed files), wrap it rather than fighting it:
// A library gives you bytes (inputStream); you want lines of text.
InputStream inputStream = null; // supplied by the library (e.g. an HTTP response body)
try (var reader = new java.io.BufferedReader(
new java.io.InputStreamReader(inputStream, StandardCharsets.UTF_8))) {
String line;
while ((line = reader.readLine()) != null) {
System.out.println(line);
}
}
Serialization: the warning
Java has a built-in object serialization mechanism (Serializable, ObjectOutputStream) that converts objects to bytes and back. Do not use it for new code. Three reasons: the binary format is Java-only and version-fragile (rename a field and old data may stop deserializing); it's a notorious security hole (deserializing untrusted bytes can execute attacker code — this has produced real CVEs for years); and it bakes your class's private internals into the stored format. For persistence or wire formats, use JSON (or another text format) with an explicit schema — the libraries are named in the next section, and the Data tracks will go deep.
CSV and JSON: the teaser
Real data files are rarely plain lines — they're CSV, JSON, or similar. Java's standard library deliberately does not ship a CSV or JSON parser (CSV dialects are messier than they look; JSON handling belongs to libraries). The standard choices:
- JSON — Jackson (
jackson-databind) is the industry default; Gson is the lighter alternative. You'll meet Jackson properly in the Spring track, where it serializes every REST response. - CSV — Jackson has a CSV module, and Apache Commons CSV is a solid standalone. For quick scripts, splitting lines on commas works until a field contains a quoted comma — which is exactly when you switch to a library.
The worked example below parses a simple CSV by hand to keep the focus on I/O; treat the hand-rolled parsing as a teaching scaffold, not a production pattern.
Worked example: CSV in, summary report out
Putting it together end-to-end: read data/sales.csv, total the amounts per region, and write out/summary.txt. Every resource opened here is managed by try-with-resources; the charset is explicit at both ends.
Input file data/sales.csv:
order_id,region,amount
1001,west,250.00
1002,east,99.50
1003,west,75.25
1004,north,310.00
1005,east,40.00
import java.io.IOException;
import java.math.BigDecimal;
import java.math.RoundingMode;
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
import java.util.Map;
import java.util.TreeMap;
import java.util.stream.Stream;
import static java.nio.file.StandardOpenOption.*;
public class SalesSummary {
public static void main(String[] args) throws IOException {
Path in = Path.of("data", "sales.csv");
Path outDir = Path.of("out");
Files.createDirectories(outDir);
// region -> total, kept sorted for a stable report
Map<String, BigDecimal> totals = new TreeMap<>();
int orders = 0;
try (Stream<String> lines = Files.lines(in, StandardCharsets.UTF_8)) {
var dataLines = lines.skip(1) // drop the header
.map(String::trim)
.filter(l -> !l.isEmpty())
.toList();
for (String line : dataLines) {
String[] cols = line.split(",", -1); // -1 keeps trailing empties
if (cols.length != 3) {
System.err.println("skipping malformed line: " + line);
continue;
}
String region = cols[1].trim();
BigDecimal amount = new BigDecimal(cols[2].trim()); // never double for money (post 17)
totals.merge(region, amount, BigDecimal::add);
orders++;
}
}
StringBuilder report = new StringBuilder("Sales summary\n");
totals.forEach((region, total) ->
report.append("%-8s %10s%n".formatted(region,
total.setScale(2, RoundingMode.HALF_UP))));
report.append("orders: %d%n".formatted(orders));
Path out = outDir.resolve("summary.txt");
Files.writeString(out, report.toString(), StandardCharsets.UTF_8,
CREATE, TRUNCATE_EXISTING);
System.out.println("wrote " + out.toAbsolutePath());
}
}
Output — out/summary.txt:
Sales summary
east 139.50
north 310.00
west 325.25
orders: 5
Notice the deliberate choices: skip(1) for the header, malformed lines logged and skipped instead of crashing the whole run, BigDecimal for money (post 17's rule), and the write uses TRUNCATE_EXISTING because a summary report should be regenerated fresh, not appended to. In production you'd replace the hand-rolled split with a CSV library and wrap the parsing in tests — which is exactly what the Tooling & Testing track is for.
Where Core Java ends — and what's next
That closes the Core Java track: 18 posts from your first program to reading and writing real data. You now have the language fundamentals — types, objects, collections, generics, lambdas, streams, records, time, money, exceptions, and files — that everything else in the roadmap builds on.
Next is Tooling & Testing: Maven and Gradle (because real projects aren't single files), JUnit 5 and Mockito (because the sales-summary parsing above deserves a test), Testcontainers, Docker, and CI/CD. The code you've learned to write is about to become code you can build, verify, and ship.
Continue: Java Learning Roadmap 2026
Comments
Post a Comment