JVM Performance Tuning

The Java Virtual Machine does a lot automatically. It compiles hot code to optimized machine code at runtime (JIT), manages memory with sophisticated garbage collectors, and adapts to the hardware it runs on. Modern JVMs (Java 17, 21, and later) have sensible defaults, so most Java applications need little tuning, and many "tuning" efforts copied from decade-old blog posts make things worse.

When tuning is needed (latency spikes from GC pauses, out-of-memory kills in containers, slow startup, high CPU), the approach is the same as any performance engineering: measure first with GC logs and profilers, change one thing at a time, and validate under realistic load.

TL;DR

Quick Example

A pragmatic container configuration for a latency-sensitive Spring Boot service:

(On JDK 23+, generational ZGC is the default mode, so -XX:+ZGenerational is unnecessary.)

Quick diagnostics on a running JVM:

Core Concepts

Memory Layout

Total process memory is the heap plus all of these. In containers, setting -Xmx to the full container limit gets the process OOM-killed. Leave 20–30% headroom.

Container Awareness

Modern JVMs read cgroup limits for CPU and memory. Use percentage-based sizing (-XX:MaxRAMPercentage=70–80, -XX:InitialRAMPercentage) rather than hard-coded -Xmx, so the same image works at different sizes. Very small CPU limits (under 1 CPU) change GC and JIT thread counts, and can hurt performance. Give JVM services at least 1–2 CPUs. See Kubernetes workloads.

Garbage Collectors

Choose based on the goal (throughput, tail latency, or footprint), then tune minimally: heap size first, then collector choice, then specific flags only if measurements justify them. See garbage collection.

GC Logs

Unified logging (-Xlog:gc*) records pause times, causes, heap occupancy before and after, and allocation rates. Analyze with GCeasy, GCViewer, or JDK Mission Control. Red flags: frequent full GCs, pauses exceeding SLOs, heap that never drops after GC (a leak), and humongous allocations in G1.

JIT Compilation and Warm-Up

The JVM starts by interpreting bytecode, then tiered compilation (C1 quick compilation, then C2 or Graal optimization) compiles hot methods using runtime profiles. Consequences:

Profiling

See profiling.

Memory Leaks

Java leaks are objects unintentionally kept reachable: unbounded caches, static collections, listeners never removed, ThreadLocals in thread pools, or classloader leaks on redeploy. Diagnose with:

  1. GC logs showing post-GC heap growing over time.
  2. A heap dump (-XX:+HeapDumpOnOutOfMemoryError, or jcmd GC.heap_dump) analyzed in Eclipse MAT for dominator trees and retained sizes.
  3. Native Memory Tracking (-XX:NativeMemoryTracking=summary) when the process grows but the heap doesn't.

Startup Optimization

These matter for serverless, CLIs, and scale-to-zero workloads.

Best Practices

Upgrade Before You Tune

Each JDK release improves GC, JIT, and libraries. Moving from Java 11 or 17 to 21+ often yields meaningful gains with zero code changes.

Tune for an Explicit Goal

Define the target (p99 latency under X ms, throughput Y requests/s, memory under Z), measure the baseline under realistic load, and change one parameter at a time.

Keep JFR and GC Logging On in Production

Continuous low-overhead recordings mean you have data when an incident happens, instead of trying to reproduce it later.

Reduce Allocation, Not Just GC Time

High allocation rates drive GC frequency. Profile allocations, reuse buffers where sensible, avoid needless boxing and intermediate collections in hot paths, and stream large data instead of loading it all.

Common Mistakes

Setting -Xmx Equal to the Container Limit

The JVM needs memory beyond the heap. A 2 GiB container with -Xmx2g gets OOM-killed by the kernel, often without a Java stack trace. Use MaxRAMPercentage with headroom.

Copying Old Tuning Flags

Flags like -XX:+UseConcMarkSweepGC (removed), large NewRatio tweaks, or -XX:+AggressiveOpts from old guides may be obsolete or harmful. Start from defaults, and add only what measurements support.

Benchmarking Without Warm-Up or JMH

Timing code in main with System.nanoTime measures interpreter and JIT noise. Use JMH for microbenchmarks, and load tests with warm-up for services.

FAQ

Which garbage collector should I use?

Start with the default, G1, which suits most applications. Choose generational ZGC if you need consistently low pause times (latency-sensitive APIs, large heaps), Parallel GC for maximum throughput in batch processing, and Serial for very small memory footprints. Validate with GC logs under realistic load.

How much heap should a Java service in a container have?

Commonly 60–80% of the container memory limit via -XX:MaxRAMPercentage, leaving the rest for metaspace, thread stacks, code cache, and direct buffers. Services with heavy off-heap usage (Netty, large direct buffers) need more headroom.

What is Java Flight Recorder?

A profiling and event recording framework built into the JVM, with very low overhead, suitable for continuous production use. It captures CPU samples, allocations, GC activity, locks, I/O, and exceptions. Recordings are analyzed in JDK Mission Control or converted for other tools.

Should I use GraalVM native image?

For workloads where startup time and memory footprint dominate (serverless functions, CLI tools, scale-to-zero services), native images can be transformative. For long-running, throughput-heavy services, the JIT-compiled JVM typically achieves higher peak performance, with fewer restrictions on reflection and dynamic class loading.

Related Topics

References