Neo4j
Neo4j is the most widely used graph database. Instead of tables and foreign keys, it stores nodes (entities), relationships (typed, directed connections), and properties on both. Relationships are first-class data stored alongside nodes, so traversing connections — friends of friends, the supply chain behind a product, the accounts linked to a fraudulent device — is fast even many hops deep.
You query Neo4j with Cypher, a declarative language whose ASCII-art patterns ((a)-[:KNOWS]->(b)) look like the graph you're describing. Neo4j is used for fraud detection, recommendations, identity and access graphs, network and IT operations, knowledge graphs, and increasingly as the knowledge layer for GraphRAG with large language models.
TL;DR
- Data is a labeled property graph: nodes with labels, typed relationships, properties on both.
- Cypher matches patterns:
MATCH (u:User)-[:FOLLOWS]->(f) RETURN f. - Index-free adjacency makes traversals cost proportional to the neighborhood, not the table size.
- Model around the questions you'll ask; relationships are nouns worth naming.
- Use indexes and constraints for entry points; use
MERGEcarefully for idempotent writes. - Choose a graph database when relationships and multi-hop queries dominate; otherwise a relational database is usually simpler.
Quick Example
Model users, products, and purchases, then find recommendations ("people who bought what you bought also bought"):
Find the shortest connection between two accounts (common in fraud investigations):
Core Concepts
The Property Graph Model
Relationships are traversed in either direction at query time, so direction is about meaning, not performance.
Cypher Essentials
MATCH— find patterns;OPTIONAL MATCH— like a left join.WHERE— filter;RETURN/WITH— project and chain query parts.CREATE— always creates;MERGE— match or create the whole pattern.- Variable-length paths:
-[:KNOWS*1..3]->;shortestPathandallShortestPaths. - Aggregations (
count,collect),UNWINDfor lists, andCALL {}subqueries. - Always use parameters (
$userId) for performance and injection safety.
Why Traversals Are Fast
In a relational database, each hop is a join whose cost grows with table and index sizes. Neo4j stores direct pointers from each node to its relationships (index-free adjacency), so a traversal only touches the connected records. Five-hop queries that time out as SQL joins can return in milliseconds.
Data Modeling
- Start from the questions: "Which accounts share devices with a flagged account?"
- Turn important connections into relationships with specific types (
:PURCHASED, not:RELATED_TO). - Promote things you traverse through into nodes (an
Addressnode linking many customers). - Avoid supernodes (a node with millions of relationships, like a "Country" everyone links to) on hot paths, or add intermediate nodes.
Graph Data Science and GraphRAG
The Graph Data Science library runs algorithms — PageRank, community detection (Louvain, Leiden), similarity, shortest paths, node embeddings — directly on the graph for fraud rings, influencer detection, and recommendations. See Graph Algorithms. GraphRAG combines a knowledge graph with vector search so LLMs retrieve connected facts rather than isolated text chunks. See Retrieval-Augmented Generation.
Deployment Options
Neo4j Community Edition (open source, single instance), Enterprise Edition (clustering, role-based security, online backup), and AuraDB (fully managed cloud). Clustering uses Raft for primaries and read replicas for scale-out reads.
Best Practices
Create Constraints and Indexes First
Uniqueness constraints and indexes on lookup properties make MATCH entry points fast and prevent duplicate nodes from concurrent MERGEs.
Use MERGE on Keys, SET for Everything Else
MERGE (u:User {id: $id}) SET u.name = $name is idempotent; merging on every property creates duplicates when any property changes.
Profile Queries
EXPLAIN and PROFILE show whether queries use indexes and how many database hits each step makes.
Bound Variable-Length Paths
Always set an upper bound (*..6). Unbounded traversals can explore huge parts of the graph.
Batch Large Imports
Use neo4j-admin database import for initial loads and batched UNWIND with CALL {} IN TRANSACTIONS for incremental ones.
Keep the Source of Truth Clear
Many teams keep transactional data in a relational database and project relationships into Neo4j via change data capture for graph queries.
Common Mistakes
Modeling Tables as Nodes
Copying relational schemas (join tables as nodes, foreign keys as properties) loses the benefits of a graph.
Generic Relationship Types
:RELATED_TO for everything forces property filters on every traversal. Use specific types.
Cartesian Products
Two disconnected MATCH patterns multiply every row by every row. Connect patterns or use WITH to reduce first.
String-Concatenated Queries
Building Cypher by concatenating user input invites injection and prevents plan caching. Use parameters.
Choosing a Graph for Simple CRUD
If queries are mostly single-entity lookups and simple joins, a relational database is simpler to operate.
Comparison
FAQ
What is Neo4j used for?
Problems dominated by relationships: fraud detection, recommendation engines, identity and access graphs, supply chain and network analysis, master data, and knowledge graphs for AI.
What is Cypher?
Neo4j's declarative graph query language, now the basis of the ISO GQL standard. It describes graph patterns using ASCII-art syntax like (a)-[:KNOWS]->(b).
When should I use a graph database instead of SQL?
When your most important queries traverse many relationships — variable-depth paths, networks, or recommendations — and relational joins become slow or complex.
Is Neo4j free?
Community Edition is open source. Enterprise Edition and AuraDB (managed cloud, with a free tier) add clustering, security features, and support.
Can Neo4j do vector search?
Yes. Neo4j supports vector indexes on node properties, which enables combining semantic similarity with graph traversals for GraphRAG.
Related Topics
- SQL vs NoSQL — Choosing data models
- Graph Algorithms — The algorithms behind graph analytics
- Retrieval-Augmented Generation — GraphRAG and knowledge graphs for LLMs
- PostgreSQL — The relational alternative
- Databases — Choosing the right store