Insecure Deserialization & Integrity Failures
Deserialization turns bytes back into objects. When an application deserializes untrusted data with a format that can instantiate arbitrary types and run their code (Java's native serialization, Python's pickle, PHP's unserialize, .NET's BinaryFormatter, unsafe YAML loaders), an attacker can craft a payload that executes commands on the server the moment it's parsed. Insecure deserialization has produced some of the most severe vulnerabilities in widely used software, often with full remote code execution (RCE).
The 2021 OWASP Top 10 folds it into A08: Software and Data Integrity Failures, a broader category about trusting data or code without verifying its integrity. That includes unsigned updates, compromised CI/CD pipelines, dependencies from untrusted sources, and tampered cookies or tokens. The unifying lesson: never let untrusted input decide what code runs, and verify integrity before trusting data.
TL;DR
- Native serialization formats can instantiate arbitrary classes; with gadget chains in the classpath or environment, that means RCE.
- Never deserialize untrusted data with Java
ObjectInputStream, Pythonpickle/marshal, PHPunserialize, .NETBinaryFormatter, or YAMLload(unsafe modes). - Prefer data-only formats: JSON, Protobuf, or MessagePack, mapped to explicit types with schema validation.
- If native deserialization is unavoidable: authenticate the data (HMAC or signatures), and use class allowlists (Java serialization filters).
- Integrity failures also include unsigned software updates, tampered CI/CD, and untrusted dependencies. Sign and verify artifacts.
- Watch for deserialization in cookies, session stores, caches, message queues, and ML model files.
Quick Example
Python pickle is code execution by design:
Vulnerable pattern: storing a pickled object in a cookie or accepting uploaded pickles:
YAML:
Java deserialization filter (JEP 290) as defense in depth:
Core Concepts
Why Deserialization Can Execute Code
Native serialization formats encode type information along with data. During deserialization, the runtime creates instances of the named classes and invokes special methods: Java's readObject/readResolve, Python's __reduce__/__setstate__, PHP's __wakeup/__destruct, .NET serialization callbacks. Attackers chain together existing classes in the application's libraries (gadgets) whose methods, invoked in sequence during deserialization, end up executing commands, writing files, or making requests. Tools like ysoserial automate gadget chains for popular Java libraries, which is why "we don't have dangerous classes" is rarely true.
Dangerous APIs by Ecosystem
Where Untrusted Serialized Data Appears
- Cookies, hidden form fields, and "view state".
- Session stores and caches (Redis or Memcached entries an attacker can influence).
- Message queues and event streams consumed from less-trusted producers.
- File uploads (documents, archives, machine learning model files, which are often pickles).
- RMI, JMX, and custom binary protocols; inter-service calls using native serialization.
Software and Data Integrity Failures
The broader A08 category also covers:
- Unsigned or unverified updates: auto-update mechanisms that don't check signatures.
- CI/CD compromise: pipelines pulling unpinned actions, scripts from the internet (
curl | sh), or untrusted dependencies. See GitHub Actions security. - Dependency integrity: typosquatting, compromised packages, and dependency confusion. Use lockfiles with hashes and trusted registries.
- Tampered client data: trusting unsigned cookies, JWTs without signature verification, or client-side prices.
Supply chain controls include signing artifacts (Sigstore/cosign), verifying provenance (SLSA attestations), SBOMs, and pinned, hash-verified dependencies. See container security.
Best Practices
Use Data-Only Formats Across Trust Boundaries
Anything crossing a trust boundary (client ↔ server, service ↔ service, uploads, queues) should use formats that can't encode behavior: JSON, Protobuf, Avro, CBOR, or MessagePack, deserialized into explicit, validated types.
Authenticate Data You Must Round-Trip
If you store state in clients or other less-trusted locations (cookies, tokens, URLs), sign it with HMAC or digital signatures, and verify the signature before parsing, or use framework-provided signed or encrypted sessions. Signing doesn't make unsafe formats safe if the key leaks, so still prefer data-only formats.
Restrict Polymorphism
Disable polymorphic type handling driven by input (Jackson default typing, Json.NET TypeNameHandling), or constrain it to an allowlist of known subtypes, such as discriminated unions with fixed type names.
Treat Model Files as Code
Loading a pickled ML model is running code. Use safetensors or ONNX for weights, load models only from trusted, verified sources, and scan or sandbox third-party models. See MLOps.
Common Mistakes
"It's Internal, So It's Trusted"
Internal queues and caches become attack paths once any producer, network segment, or cache entry is compromised, including via SSRF into an unauthenticated Redis. Apply the same rules internally.
Blocklisting Known Gadget Classes
Blocklists of known-dangerous classes are bypassed by new gadget chains in newly added libraries. Use allowlists, or better, eliminate native deserialization of untrusted data.
Encrypting Instead of Authenticating
Encrypting a serialized blob without integrity protection (for example AES-CBC without a MAC) can still allow tampering and padding oracle attacks. Use authenticated encryption or signatures.
FAQ
What is insecure deserialization?
It's deserializing attacker-controlled data with a mechanism that can instantiate arbitrary objects and trigger their methods. Attackers craft payloads that abuse existing classes (gadget chains) to execute code, tamper with application logic, or cause denial of service when the application parses the data.
Is JSON deserialization safe?
Plain JSON parsing into primitive types or explicitly declared classes is generally safe from code execution. Risk returns when libraries let JSON specify which class to instantiate (polymorphic type handling driven by input), so keep that disabled or strictly allowlisted, and validate the resulting data.
Why is Python pickle dangerous?
Pickle is a small program that the unpickler executes. It can call any importable function with any arguments via __reduce__, so loading a malicious pickle is equivalent to running attacker code. Python's documentation explicitly warns never to unpickle untrusted data.
How does this relate to supply chain security?
Both are integrity failures: trusting data or code without verifying where it came from and that it wasn't modified. The defenses are similar: verify signatures and hashes, restrict trusted sources, and avoid mechanisms that execute whatever they're given.
Related Topics
- OWASP Top 10 — The web application risk list
- API Security — Validating input across API boundaries
- GitHub Actions Security — CI/CD integrity
- Container Security — Signed images and provenance
- SSRF — Reaching internal stores that hold serialized data
- Cryptographic Failures — Signatures and authenticated encryption