Skip to content
Tech Interview Prep home
Technical interview guide

NoSQL, Graph & Key-Value Data Stores

When a relational database isn't the right fit — document, key-value, graph, and vector stores, and how to choose between them.

Read
23 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: SQL principles with PostgreSQL 18 examples; vendor-specific behavior must be verified.

Overview

Curated: · Written: · Reviewed:

NoSQL is a family of trades, not a category of database

"NoSQL" groups together stores that abandoned the relational model for different reasons, so the term says almost nothing on its own. What matters is which trade a particular store made, and whether that trade fits your problem.

The common thread is that each gives up some generality — a flexible query language, joins, multi-row transactions, strong consistency — in exchange for something specific: horizontal scale, latency, flexible schema, or a data structure the relational model expresses badly.

The trade has also narrowed. Relational databases now have native JSON with indexing (PostgreSQL's jsonb with GIN indexes, MySQL's JSON type with generated columns), and distributed SQL systems designed for horizontal write scaling with transactions (Spanner, CockroachDB, YugabyteDB) exist as managed services. "Relational cannot do JSON" and "relational cannot scale" were once fair; today both need checking rather than assuming. The honest starting position in an interview is: relational is the default, and a NoSQL store needs a reason — and the first probe you will get is "why not just Postgres?"

The mental model: model around access patterns, not around data

Relational modelling normalises the data into whatever shape minimises redundancy, then trusts the optimiser to make queries fast. Most NoSQL stores have no such optimiser, so the design method inverts: start from the queries, then shape the data to serve them, accepting duplication as the price. This is the single idea that organises everything below — embedding versus referencing in documents, table-per-query in wide-column stores, key design in key-value stores. It is also the source of the main failure mode: when a new query arrives that the data was not shaped for, you are not adding an index, you are re-shaping data and backfilling.

The decision inputs are access patterns, consistency needs, and relationship depth — not data size, not "modern vs legacy". A 4 TB dataset that is always fetched by primary key is a key-value problem; a 40 GB dataset with ad-hoc reporting is a relational problem.

The four families, by the query shape each serves

Key-value stores

A key maps to a value, retrieval is by key and nothing else — no querying by value, no ranges over content, no joins. That constraint buys predictable low latency and easy horizontal partitioning, because any key routes to a node by hashing, independently of every other key.

Redis is the common example, and it is really a data-structure server: values can be strings, hashes, lists, sets, sorted sets or streams, each with native operations. A sorted set gives leaderboards and rate limiters in one operation (ZADD/ZREVRANGE); a list gives a queue; a hash gives partial updates without rewriting the whole value.

Serves: cache, session storage, rate limiting, leaderboards, ephemeral state — access by a known key, latency dominates, state is rebuildable or has an explicitly acceptable durability model. Breaks on: any query by content ("find all sessions for user X" needs a second index structure you build yourself) and any need for multi-key atomic updates (Redis's multi-key transactions require all keys on the same slot in clustered mode). Treating an in-memory store as a durable system of record is a decision needing explicit justification, not a default.

Document stores

Documents are self-describing structures — JSON or similar — stored under a key and queryable by their contents. MongoDB is the common example.

The central modelling decision is embed or reference. Embedding a child inside its parent means one read fetches everything and the update is atomic at the document level; it suits data that is always accessed together, bounded in size, and owned by the parent — order lines within an order. Referencing suits data accessed independently, shared between parents, unbounded in growth, or updated on a different cadence.

Getting this wrong is the classic document-store failure. Embedding an unbounded array — every comment on a popular post — hits the document size limit (MongoDB: 16 MB per document), and long before that, every read of the post transfers all of them. Over-referencing recreates relational joins in application code, without the database's ability to optimise them — which is worse than just using a relational database.

"Schemaless" deserves scepticism. There is always a schema; the only question is whether the database enforces it or the application assumes it. Without enforcement, documents written by different code versions accumulate, every reader must handle every historical shape, and drift is invisible until something breaks. MongoDB's schema validation ($jsonSchema in a collection validator) is worth using.

Serves: self-contained aggregates with genuinely variable shape per record, read-mostly workloads with known access paths. Breaks on: ad-hoc queries across many documents (aggregation pipelines exist but are clumsier and slower than SQL), and multi-document transactions (supported in MongoDB since 4.0, but at a cost and against the model's grain).

Wide-column stores

Cassandra, ScyllaDB, HBase, Bigtable. Data is partitioned by a partition key and sorted within the partition by clustering columns, and queries that do not specify the partition key are either rejected or catastrophically expensive — Cassandra requires the partition key for a SELECT unless you allow a full-cluster scan with ALLOW FILTERING.

This forces the inverted design method at its most extreme: one table per query, duplicating data across them. Denormalisation is not a compromise here; it is the design. Writes are cheap and expected to go to several tables, and the price of a wrong guess about access patterns is high, because a new query pattern means a new table and a backfill.

The partition key must also distribute evenly. A key with skew creates a hot partition that limits the whole cluster, and a partition that grows unboundedly — everything keyed by one tenant — becomes unmanageable.

Serves: write-heavy time-series and event data, high-ingestion feeds with fixed, known read paths. Breaks on: ad-hoc analytics (you export to a warehouse instead), evolving query patterns, and anything needing secondary indexes as a routine tool rather than a last resort.

Graph databases

Nodes, relationships, properties — and relationships are first-class objects rather than a join to resolve. The consequence is index-free adjacency: once a start node is found, reaching its neighbours follows stored relationship pointers rather than repeatedly performing global join lookups. Cost still grows with the nodes and edges visited, and is affected by cache and network placement, but it need not scan the whole graph at each hop.

That matters when relationships are the point: social networks, recommendations, fraud rings, dependency analysis, access-control hierarchies. In SQL, a variable-depth traversal is a recursive CTE; a graph store makes the same traversal a native operation and a more natural query (Neo4j's Cypher: MATCH (a)-[:FRIENDS*1..3]->(b)), though both still do more work as the explored subgraph grows.

The honest boundary: a graph database is justified when traversal is central and deep. A single join to fetch a user's friends is not a reason. Many teams adopt one for a problem two joins would solve, and pay in operational unfamiliarity and a smaller ecosystem.

Consistency: what you actually give up, in product terms

Distributed NoSQL stores replicate data across nodes, and a write acknowledged on one replica is not instantly visible on all of them. The spectrum, as product consequences:

  • Eventual consistency: replicas converge, but reads can be stale. Product consequence: a user un-likes a post, refreshes, and sees the like still there for a second. DynamoDB's default reads and Cassandra reads below quorum sit at this end.
  • Read-your-writes: the user who just wrote always sees their own write. Usually the minimum acceptable guarantee for user-facing state. The implementations are concrete: DynamoDB strongly consistent reads on the same item; sticky routing of a session's reads to the same node; or an application-level version token the client sends with subsequent reads.
  • Monotonic reads: a user never sees time go backwards — a like count that was 47 must not read 45 on the next refresh. Broken when reads round-robin across replicas with different lag; fixed by sticky replica routing per session.
  • Strong/linearizable consistency: every read sees the latest write, at the cost of cross-node coordination on every operation. DynamoDB's strongly consistent reads return the most recent data but, per AWS documentation, may have higher latency than eventually consistent reads and are not supported on global secondary indexes.

The PACELC framing is the interview-ready extension of CAP: even without a network partition you are choosing between latency and consistency on every write. A quorum write (W + R > N in Cassandra/Dynamo terms) means waiting for the slowest of several replicas; W = 1 means fast acks and stale reads. This is a dial, not a theorem — say which setting you'd pick for a like counter (eventual, W = 1) versus a payment (strong, or relational).

Partitioning, replication, and the viral key

Two distribution strategies, with opposite failure modes:

  • Hash partitioning spreads writes evenly by default, but loses range queries — you cannot scan "all of October" without probing every partition.
  • Range partitioning keeps ranges together (good for time-series: one partition per day), but is vulnerable to skew — all writes for the current minute land on one partition. Most systems combine both: hash the partition key, range the clustering/sort key within it.

When one key goes viral — a celebrity's profile, a flash-sale item — every read and write for it lands on one partition. Whether replicas help depends on the replication architecture:

  • Leader-based systems (HBase regions, MongoDB replica sets): reads can spread across replicas, but writes go through the single leader for that partition, so a write-hot key saturates the leader no matter how many replicas exist.
  • Leaderless systems (Cassandra, DynamoDB-style): any replica can accept a write, so write load spreads — but the consistency dial from the previous section is what makes that safe, and a single hot key still saturates the one partition's storage and network.

Mitigations, in order of preference:

  1. Spread the key: append a random suffix ("salting") across N sub-keys and read from a random one — trades read-your-writes and atomic updates for write throughput.
  2. Cache in front: the viral key is hot precisely because it is cacheable; a Redis layer with a short TTL absorbs the read storm.
  3. Split the entity: move the counter or comment stream out of the hot document into its own keyed items.

Failover follows the same split: a leader-based system fails over by electing a replica, typically seconds of unavailability; a leaderless system keeps serving through node loss at the cost of the consistency dial above. Know which kind you are defending.

Choosing, and the traps interviewers wait for

The question is never "SQL or NoSQL" but "what does this workload need that my default cannot provide". Concretely:

  • Do the access patterns fit a key? Known identifier, latency dominates → key-value.
  • Is the data genuinely variable per record? Real variability suits documents. Fixed-but-tedious-to-migrate does not.
  • Are relationships the primary object of study? Deep, variable-depth traversal justifies a graph store.
  • Is write volume beyond what one node can take, with access always by a known partition? That is the wide-column case.
  • Do you need multi-entity transactions, ad-hoc queries, or strong constraints? Those favour relational, and giving them up should be deliberate, not a side effect.

Two traps recur, and interviewers probe both. Choosing for scale you do not have — most systems never exceed what a well-indexed relational database on one machine handles; the flexibility given up is paid immediately while the scale benefit never arrives. Choosing for schema flexibility to avoid migrations — this trades a one-time cost for permanent ambiguity about what the data contains.

Polyglot persistence — several stores, each for what it suits — is legitimate but not free: each is another system to operate, monitor, back up and staff, the consistency problems between them become yours, and there is no join across them. A second store should earn its place against that cost.

A weak answer sounds like: "we used MongoDB because it's more scalable" with no access pattern named, no consistency requirement stated, and no answer to "why not Postgres?" A strong answer names the specific query shape, the specific relational pain (measured, ideally), and the specific trade accepted.

Worked example: one tenant as the partition key

Orders table, 8M items. Access is GetItem(tenant_id, order_id). Tenant Acme is 41% of traffic. (Figures from a load-test scenario, not a specific benchmark.)

partition keyp99 GetItemthrottled requests / minnotes
tenant_id only820 ms140Acme is one hot partition
tenant_id#yyyy-mm-dd + GSI on order_id18 ms2even shards; list-by-tenant needs the GSI
Postgres 18, 8M rows, btree on (tenant_id, order_id)9 ms0you did not have a scale problem

The interview point is whether you have an access pattern that needs this trade, or a tenant-shaped hot key dressed as "we needed NoSQL to scale."

Likely follow-ups

  • "Why not just Postgres?" — your first probe. Have a measured answer: the specific access pattern, the specific pain.
  • "What consistency does feature X need?" — be ready to classify like counts (eventual), shopping carts (read-your-writes), payments (strong) and say how you'd get each.
  • "A key goes viral — walk me through what happens and what you do." — hot partition mechanics plus the mitigation ladder above, and whether your system is leader-based or leaderless.
  • "Your read paths changed; now what?" — the re-shaping cost: new table or embedding change, dual-write backfill, cutover. This is where access-pattern modelling shows its price.
  • "Design the schema" given read paths first — demonstrate the inverted method: list the queries, one table/collection per query, denormalise deliberately, state the duplication you accepted.