JVM Performance Tuning
The Java Virtual Machine does a lot automatically. It compiles hot code to optimized machine code at runtime (JIT), manages memory with sophisticated garbage collectors, and adapts to the hardware it runs on. Modern JVMs (Java 17, 21, and later) have sensible defaults, so most Java applications need little tuning, and many "tuning" efforts copied from decade-old blog posts make things worse.
When tuning is needed (latency spikes from GC pauses, out-of-memory kills in containers, slow startup, high CPU), the approach is the same as any performance engineering: measure first with GC logs and profilers, change one thing at a time, and validate under realistic load.
TL;DR
- Use a current LTS JDK (21 or later); many performance improvements come free with upgrades.
- In containers, size the heap with
-XX:MaxRAMPercentage(for example 70–75%), leaving room for non-heap memory. - G1 is the default and a good balance. ZGC (generational) gives sub-millisecond pauses for latency-sensitive services; Parallel GC maximizes batch throughput.
- Enable GC logging (
-Xlog:gc*) and use Java Flight Recorder plus async-profiler to find real bottlenecks. - Diagnose leaks with heap dumps (Eclipse MAT) and native memory tracking.
- Improve startup with Class Data Sharing (AppCDS), CRaC checkpoints, or GraalVM native images.
Quick Example
A pragmatic container configuration for a latency-sensitive Spring Boot service:
(On JDK 23+, generational ZGC is the default mode, so -XX:+ZGenerational is unnecessary.)
Quick diagnostics on a running JVM:
Core Concepts
Memory Layout
- Heap: objects, managed by the garbage collector. Sized by
-Xms/-Xmx, or by percentages of available RAM. - Metaspace: class metadata.
- Thread stacks: per platform thread (virtual threads use heap). See virtual threads.
- Code cache: JIT-compiled code.
- Direct and native memory: NIO buffers (Netty), JNI, and GC internal structures.
Total process memory is the heap plus all of these. In containers, setting -Xmx to the full container limit gets the process OOM-killed. Leave 20–30% headroom.
Container Awareness
Modern JVMs read cgroup limits for CPU and memory. Use percentage-based sizing (-XX:MaxRAMPercentage=70–80, -XX:InitialRAMPercentage) rather than hard-coded -Xmx, so the same image works at different sizes. Very small CPU limits (under 1 CPU) change GC and JIT thread counts, and can hurt performance. Give JVM services at least 1–2 CPUs. See Kubernetes workloads.
Garbage Collectors
Choose based on the goal (throughput, tail latency, or footprint), then tune minimally: heap size first, then collector choice, then specific flags only if measurements justify them. See garbage collection.
GC Logs
Unified logging (-Xlog:gc*) records pause times, causes, heap occupancy before and after, and allocation rates. Analyze with GCeasy, GCViewer, or JDK Mission Control. Red flags: frequent full GCs, pauses exceeding SLOs, heap that never drops after GC (a leak), and humongous allocations in G1.
JIT Compilation and Warm-Up
The JVM starts by interpreting bytecode, then tiered compilation (C1 quick compilation, then C2 or Graal optimization) compiles hot methods using runtime profiles. Consequences:
- Warm-up: the first seconds or minutes of traffic are slower. Benchmark after warm-up, and consider warm-up traffic or readiness delays in deployments.
- Microbenchmarks must use JMH, since naive timing loops mislead because of JIT effects like dead-code elimination. See benchmarking.
Profiling
- Java Flight Recorder (JFR): built-in, low-overhead (around 1–2%) recording of CPU samples, allocations, locks, GC, I/O, and exceptions. Safe to run continuously in production, and analyzed with JDK Mission Control.
- async-profiler: accurate CPU, allocation, lock, and wall-clock flame graphs, without safepoint bias.
- APM agents (OpenTelemetry Java agent, Datadog, New Relic) connect JVM metrics to traces. See OpenTelemetry instrumentation.
See profiling.
Memory Leaks
Java leaks are objects unintentionally kept reachable: unbounded caches, static collections, listeners never removed, ThreadLocals in thread pools, or classloader leaks on redeploy. Diagnose with:
- GC logs showing post-GC heap growing over time.
- A heap dump (
-XX:+HeapDumpOnOutOfMemoryError, orjcmd GC.heap_dump) analyzed in Eclipse MAT for dominator trees and retained sizes. - Native Memory Tracking (
-XX:NativeMemoryTracking=summary) when the process grows but the heap doesn't.
Startup Optimization
These matter for serverless, CLIs, and scale-to-zero workloads.
Best Practices
Upgrade Before You Tune
Each JDK release improves GC, JIT, and libraries. Moving from Java 11 or 17 to 21+ often yields meaningful gains with zero code changes.
Tune for an Explicit Goal
Define the target (p99 latency under X ms, throughput Y requests/s, memory under Z), measure the baseline under realistic load, and change one parameter at a time.
Keep JFR and GC Logging On in Production
Continuous low-overhead recordings mean you have data when an incident happens, instead of trying to reproduce it later.
Reduce Allocation, Not Just GC Time
High allocation rates drive GC frequency. Profile allocations, reuse buffers where sensible, avoid needless boxing and intermediate collections in hot paths, and stream large data instead of loading it all.
Common Mistakes
Setting -Xmx Equal to the Container Limit
The JVM needs memory beyond the heap. A 2 GiB container with -Xmx2g gets OOM-killed by the kernel, often without a Java stack trace. Use MaxRAMPercentage with headroom.
Copying Old Tuning Flags
Flags like -XX:+UseConcMarkSweepGC (removed), large NewRatio tweaks, or -XX:+AggressiveOpts from old guides may be obsolete or harmful. Start from defaults, and add only what measurements support.
Benchmarking Without Warm-Up or JMH
Timing code in main with System.nanoTime measures interpreter and JIT noise. Use JMH for microbenchmarks, and load tests with warm-up for services.
FAQ
Which garbage collector should I use?
Start with the default, G1, which suits most applications. Choose generational ZGC if you need consistently low pause times (latency-sensitive APIs, large heaps), Parallel GC for maximum throughput in batch processing, and Serial for very small memory footprints. Validate with GC logs under realistic load.
How much heap should a Java service in a container have?
Commonly 60–80% of the container memory limit via -XX:MaxRAMPercentage, leaving the rest for metaspace, thread stacks, code cache, and direct buffers. Services with heavy off-heap usage (Netty, large direct buffers) need more headroom.
What is Java Flight Recorder?
A profiling and event recording framework built into the JVM, with very low overhead, suitable for continuous production use. It captures CPU samples, allocations, GC activity, locks, I/O, and exceptions. Recordings are analyzed in JDK Mission Control or converted for other tools.
Should I use GraalVM native image?
For workloads where startup time and memory footprint dominate (serverless functions, CLI tools, scale-to-zero services), native images can be transformative. For long-running, throughput-heavy services, the JIT-compiled JVM typically achieves higher peak performance, with fewer restrictions on reflection and dynamic class loading.
Related Topics
- Java — The language overview
- Garbage Collection — GC algorithms in depth
- Java Virtual Threads — Concurrency and memory implications
- Profiling — Finding hot spots
- Spring Boot Actuator — JVM metrics in production
- Performance Engineering — Methodical performance work