Spring Data JPA

Spring Data JPA is the most common way Spring Boot applications talk to relational databases. You define entities (Java classes mapped to tables with JPA annotations) and repository interfaces, and Spring Data generates the implementations: CRUD operations, paging, and queries derived from method names like findByEmailAndStatus. Underneath, Hibernate (the default JPA provider) handles SQL generation, change tracking, caching, and relationships.

That productivity comes with the classic ORM trade-offs. Lazy loading causes N+1 queries, transactions define when changes are flushed, and entity design choices ripple into performance. Knowing what SQL your repositories actually run is essential for production apps.

TL;DR

Quick Example

Core Concepts

Entities and Mapping

Relationships

cascade and orphanRemoval propagate persistence operations to children. Use them for true aggregates (order → items), not for shared references.

Repositories and Queries

Projections

Loading full entities (and their associations) for a list view wastes memory and queries. Projections fetch only what's needed:

Transactions and the Persistence Context

Within a @Transactional method, the persistence context tracks loaded entities. Modifying a managed entity schedules an UPDATE at flush or commit (dirty checking), so no explicit save() is needed. Repository methods are transactional by default (reads are readOnly). Service methods should define the business transaction boundary. See database transactions and the proxy caveats in dependency injection.

Set spring.jpa.open-in-view=false: the default open-session-in-view keeps a database connection through view rendering and hides lazy-loading problems until production load.

Fetching and the N+1 Problem

Loading 50 orders and then accessing order.getItems() in a loop triggers 1 + 50 queries. Fixes:

Log SQL during development (spring.jpa.show-sql or, better, a datasource proxy with query counts) and assert query counts in tests. See query optimization.

Schema Management

Use Flyway or Liquibase for versioned migrations (db/migration/V12__add_order_version.sql), run automatically at startup or as a separate deployment step, and set spring.jpa.hibernate.ddl-auto=validate so Hibernate checks mappings against the real schema. ddl-auto=update is convenient for prototypes, but it can't handle renames, data migrations, or safe production changes. See database migrations.

Best Practices

Make Associations Lazy by Default

Set @ManyToOne(fetch = LAZY) explicitly and load what each use case needs with fetch joins or entity graphs. Eager associations cascade into huge, unexpected queries.

Keep Transactions in the Service Layer

Define transaction boundaries around business operations in services. Avoid long transactions that call slow external APIs while holding database connections and locks.

Use Projections for Read Paths

APIs and list pages rarely need full entity graphs. DTO projections reduce query size, memory, and accidental lazy loading.

Tune the Connection Pool

HikariCP (the default) should be sized to database capacity, not thread count. Keep maximumPoolSize modest (for example 10–30 per instance), and monitor wait times. See PostgreSQL connection pooling.

Common Mistakes

LazyInitializationException

Accessing a lazy association after the transaction ended (for example in a controller or a JSON serializer) throws LazyInitializationException. Fetch the needed data inside the transaction via an entity graph, fetch join, or projection, rather than re-enabling open-in-view or switching to EAGER.

Returning Entities From REST Controllers

Serializing entities directly exposes internal fields, triggers lazy loads (N+1) during serialization, and risks infinite recursion on bidirectional relationships. Map to DTOs.

equals/hashCode on Generated IDs or All Fields

Using a generated ID that's null before persist, or including lazy collections, breaks Set membership and triggers loads. Base equality on a stable business key, or use ID-based equality that handles transient state carefully.

FAQ

What's the difference between JPA, Hibernate, and Spring Data JPA?

JPA (Jakarta Persistence) is the specification: annotations and the EntityManager API. Hibernate is the most popular implementation of it. Spring Data JPA sits on top, generating repository implementations and query methods so you write less boilerplate. Spring Boot auto-configures all three.

When should I use native queries?

When you need database features JPQL doesn't support (window functions, CTEs, full-text search, JSON operators, upserts), or hand-tuned SQL for performance. Map results to projections or DTOs, and keep them covered by integration tests against the real database.

Do I need to call save() after modifying an entity?

Not for entities loaded within the current transaction. Hibernate detects changes and flushes them on commit. You do need save() for new (transient) entities, or for detached entities from outside the transaction.

How do I avoid N+1 queries?

Keep associations lazy, and fetch what each use case needs using join fetch, @EntityGraph, batch fetching, or DTO projections. Monitor query counts in tests and logs to catch regressions.

Related Topics

References