Spring Data JPA: Entities, Repositories & Relationships
The "customer order history" screen was the checkout service's most ordinary feature: given a customer, show their orders. The first version was a 40-line DAO with hand-built SQL — string concatenation, a hand-rolled row mapper, and a WHERE clause assembled from three optional filters. It worked for months. Then a refactor dropped one filter condition on a code path nobody tested, and for eleven minutes the screen showed every customer's orders to every customer. No data was modified, but the incident report's root cause was one sentence long: the relationship between customers and orders existed only inside a SQL string.
JPA — the Jakarta Persistence API — makes that relationship explicit, in Java, where the compiler and your tests can see it. You describe your domain as classes; Hibernate (the JPA provider under Spring Boot 4.1) translates between those classes and your tables. And Spring Data JPA goes one step further: for most queries, you don't even write the translation — you declare a repository interface, and Spring generates the implementation, including the SQL. This post maps the checkout domain (Customer, Order), derives queries from method names, and models the one-to-many relationship properly.
Your first entity: a class that is a table
An entity is a plain Java class that JPA knows how to persist — one instance per row. Here's the checkout service's customer:
@Entity
@Table(name = "customers")
public class Customer {
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(nullable = false)
private String name;
@Column(nullable = false, unique = true)
private String email;
@OneToMany(mappedBy = "customer", cascade = CascadeType.PERSIST)
private List<Order> orders = new ArrayList<>();
protected Customer() {
}
public Customer(String name, String email) {
this.name = name;
this.email = email;
}
// Keep both sides of the relationship in sync. JPA only looks at the
// owning side (Order.customer) when writing the foreign key — if you
// forget order.setCustomer(this), the order row gets a NULL customer_id.
public void addOrder(Order order) {
orders.add(order);
order.setCustomer(this);
}
public Long getId() { return id; }
public String getName() { return name; }
public String getEmail() { return email; }
public List<Order> getOrders() { return orders; }
}
Each annotation is a mapping decision, and there are no defaults you should leave unexamined:
@Entity— "this class maps to a table." The table name defaults to the class name;@Table(name = "customers")pins it explicitly so a future class rename doesn't silently remap your data.@Id+@GeneratedValue(IDENTITY)— the primary key, generated by the database's auto-increment.IDENTITYis the right choice for Postgres, MySQL, and H2; other strategies (sequences, UUIDs) exist for other databases and other trade-offs.@Column(nullable = false, unique = true)— column constraints declared next to the field they protect. Theuniqueon email is doing real work: it's the database-level backstop behind "one account per email."- The protected no-arg constructor — Hibernate instantiates entities via reflection and requires it. Make it
protected, notpublic, so application code can't accidentally create half-initialized customers. - Annotations on fields, not getters — JPA uses whichever placement you choose (field access here). Pick one per project and stay consistent; mixing them confuses both Hibernate and humans.
Note the imports: everything is jakarta.persistence.*. Spring Boot 4 sits on Jakarta EE 11 — if you see javax.persistence in a tutorial, that tutorial predates the namespace migration and its code will not compile against Boot 4.
The other side: who owns the foreign key
Every one-to-many relationship has an owning side — the side whose table holds the foreign key — and JPA only reads the owning side when writing that key. For customers and orders, the orders table holds customer_id, so Order is the owning side:
@Entity
@Table(name = "orders")
public class Order {
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(nullable = false, precision = 12, scale = 2)
private BigDecimal total;
@Column(nullable = false, length = 20)
private String status = "NEW";
// The owning side: this field owns the foreign key column customer_id in
// the orders table. Many-to-one is EAGER by default; we ask for LAZY
// because an order screen rarely needs the whole customer object.
@ManyToOne(fetch = FetchType.LAZY, optional = false)
@JoinColumn(name = "customer_id")
private Customer customer;
protected Order() {
}
public Order(BigDecimal total) {
this.total = total;
}
public Long getId() { return id; }
public BigDecimal getTotal() { return total; }
public String getStatus() { return status; }
public void setStatus(String status) { this.status = status; }
public Customer getCustomer() { return customer; }
public void setCustomer(Customer customer) { this.customer = customer; }
}
The relationship annotations form a pair: @ManyToOne + @JoinColumn(name = "customer_id") on the owning side says "my table has this FK column"; @OneToMany(mappedBy = "customer") on the other side says "I'm the mirror — look at Order.customer for the truth." The mappedBy value is a field name, not a column name — get it wrong and Hibernate tells you at startup, which is exactly the kind of loud failure you want.
Repositories: queries without writing SQL
A Spring Data repository is an interface. You declare method signatures; Spring generates the implementation — including the SQL — at startup:
public interface CustomerRepository extends JpaRepository<Customer, Long> {
// SELECT c FROM Customer c WHERE c.email = ?1
Optional<Customer> findByEmail(String email);
// SELECT c FROM Customer c WHERE c.name LIKE ?1
java.util.List<Customer> findByNameContaining(String fragment);
}
public interface OrderRepository extends JpaRepository<Order, Long> {
// SELECT o FROM Order o WHERE o.status = ?1 AND o.customer.id = ?2
List<Order> findByStatusAndCustomerId(String status, Long customerId);
// Property traversal: o.customer.email — joins customers behind the scenes.
List<Order> findByCustomerEmailAndStatus(String email, String status);
// OrderBy, top-N and range keywords compose the same way.
List<Order> findTop5ByTotalGreaterThanOrderByTotalDesc(BigDecimal minTotal);
// When the name gets silly, say it in JPQL instead.
@Query("SELECT o FROM Order o WHERE o.customer.id = :customerId AND o.status IN :statuses")
List<Order> findActiveForCustomer(Long customerId, List<String> statuses);
}
JpaRepository<Customer, Long> already gives you save, findById, findAll, deleteById, paging, and sorting for free — the derived methods above are the custom part. The naming rules are a small grammar worth learning exactly:
- Prefix +
By:find…By,read…By,get…By,query…By, pluscount…By,exists…By,delete…By. The verb chooses the return shape;Bystarts the conditions. - Property names in camelCase:
Statusmaps to thestatusfield.And/Orcombine conditions. - Nested traversal:
findByCustomerEmailwalkscustomer.email, generating the join for you. - Keywords:
Containing,StartingWith,LessThan/GreaterThan(and…Equal),Between,In,IsNull,OrderBy,Top/First+ count,Distinct. - The ambiguity gotcha: Spring Data resolves
CustomerIdby trying the literal propertycustomerIdfirst, then the nested pathcustomer.id. OurOrderhas nocustomerIdfield, sofindByStatusAndCustomerIdmeanscustomer.id— but if someone later adds acustomerIdfield, the same method silently changes meaning. When a name feels ambiguous, it is: use@Query.
Decision rule: derived queries for conditions you can read aloud; @Query with JPQL the moment the method name needs a second line. JPQL looks like SQL but names entities and fields (Order o, o.customer.id), never tables and columns — that indirection is what keeps queries working when the schema evolves.
Cascading: saving the whole graph at once
Registering a customer with their first order touches two tables. Without cascading, you'd save the customer, then save the order, keeping the order of operations straight yourself. With cascade = CascadeType.PERSIST on Customer.orders, one save() persists the whole graph inside one transaction:
@Service
@Transactional
public class CheckoutService {
private final CustomerRepository customers;
private final OrderRepository orders;
public CheckoutService(CustomerRepository customers, OrderRepository orders) {
this.customers = customers;
this.orders = orders;
}
// cascade = PERSIST on Customer.orders means saving the customer also
// inserts the order — one save() call, two rows, one transaction.
public Customer registerWithFirstOrder(String name, String email, BigDecimal total) {
Customer customer = new Customer(name, email);
customer.addOrder(new Order(total));
return customers.save(customer);
}
public Order placeOrder(Long customerId, BigDecimal total) {
Customer customer = customers.findById(customerId)
.orElseThrow(() -> new IllegalArgumentException("unknown customer " + customerId));
Order order = new Order(total);
order.setCustomer(customer);
return orders.save(order);
}
public List<Order> openOrdersFor(String email) {
return orders.findByCustomerEmailAndStatus(email, "NEW");
}
}
Cascading is powerful enough to deserve a strict rulebook, because the wrong cascade deletes data:
PERSISTfor creation flows — parent created together with its children (register customer + first order). Safe: it only ever inserts.- Never
REMOVEacross a shared reference.CascadeType.ALLincludesREMOVE: deleting one order would delete its customer, and every other order with them. ReserveREMOVE(withorphanRemoval = true) for true composition — an order and its line items, where the child cannot meaningfully exist without the parent. - Never cascade from the many side to the one side.
@ManyToOne(cascade = ALL)onOrder.customermeans deleting an order deletes the customer. The many side should almost never cascade at all.
Decision rule: cascade down the composition hierarchy (customer → orders on create), never up and never across shared references. If you're unsure whether a cascade is safe, leave it off and save explicitly — two save() calls in one @Transactional method are just as atomic.
Schema: who creates the tables
Hibernate can generate your schema from the @Entity classes (spring.jpa.hibernate.ddl-auto=create). That is a wonderful development convenience and a production landmine: create drops your tables on every restart, and even the gentler update makes schema changes that were never reviewed, never versioned, and occasionally lossy. Real schema changes travel as versioned migrations — SQL files with names like V1__create_customers_and_orders.sql, applied in order by Flyway (or Liquibase), reviewed like code:
-- V1__create_customers_and_orders.sql
CREATE TABLE customers (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
name VARCHAR(255) NOT NULL,
email VARCHAR(255) NOT NULL UNIQUE
);
CREATE TABLE orders (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
total DECIMAL(12, 2) NOT NULL,
status VARCHAR(20) NOT NULL,
customer_id BIGINT NOT NULL REFERENCES customers (id)
);
The application properties for a grown-up setup:
spring.datasource.url=jdbc:h2:mem:checkout;DB_CLOSE_DELAY=-1
spring.datasource.username=sa
spring.datasource.password=
spring.jpa.hibernate.ddl-auto=validate
spring.jpa.show-sql=true
spring.jpa.properties.hibernate.format_sql=true
Decision rule: ddl-auto: create (or create-drop) for scratch development only; validate or none in every shared environment — and none in production, where Flyway owns the schema. A migration that adds a column gets the same review as the code that reads it.
Wiring it up
The Maven setup — spring-boot-starter-data-jpa brings Hibernate 7, Spring Data JPA, and the transaction machinery; the H2 version below is the current release I verified on Maven Central:
<project>
<modelVersion>4.0.0</modelVersion>
<parent>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-parent</artifactId>
<version>4.1.1</version>
</parent>
<groupId>com.javamakeuse</groupId>
<artifactId>checkout-service</artifactId>
<version>0.0.1-SNAPSHOT</version>
<properties>
<java.version>25</java.version>
</properties>
<dependencies>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-data-jpa</artifactId>
</dependency>
<dependency>
<groupId>com.h2database</groupId>
<artifactId>h2</artifactId>
<version>2.5.252</version>
<scope>runtime</scope>
</dependency>
<dependency>
<groupId>org.flywaydb</groupId>
<artifactId>flyway-core</artifactId>
</dependency>
</dependencies>
</project>
(Check for newer versions than the one pinned above — the coordinates are the stable part. The Flyway version is managed by Spring Boot's dependency management, so no pin needed.) Under the hood, Boot's auto-configuration builds exactly what the previous post built by hand: a HikariCP DataSource from the spring.datasource.* properties, an EntityManagerFactory scanning your @Entity classes, and a JpaTransactionManager behind every @Transactional. The JDBC post's pool knobs and transaction rules all still apply — JPA didn't replace that layer, it just gave you a nicer API over it.
One testing note: H2 is perfect for learning, but your CI should eventually run these repositories against the database you actually ship. This track's Testcontainers post shows how to spin up real Postgres in tests — the repository interfaces don't change, only the JDBC URL does, which is exactly the portability JPA promises.
Field check before you move on: add a findByEmailAndName method to CustomerRepository, then deliberately misspell one property (findByEmial) and start the app. Watch it fail at startup with a clear "no property emial" error — that fail-fast parsing is the safety net hand-written SQL never gave you. Then write a Flyway V2 migration adding a phone column, add the field to the entity, and run with ddl-auto=validate to feel the migration and the mapping check each other.
Continue: Java Learning Roadmap 2026
Comments
Post a Comment