Skip to content
Tech Interview Prep home
Technical interview guide

Blockchain & Distributed Ledger Fundamentals

What a blockchain actually is — a replicated, append-only ledger, and why that structure trades off against a conventional database.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: NISTIR 8202; Bitcoin white paper; Ethereum developer documentation and proof-of-stake behavior reviewed 2026-09-04; Hyperledger Fabric latest documentation; IPFS documentation reviewed 2026-09-04.

Overview

Curated: · Written: · Reviewed:

A blockchain is a replicated state machine with an adversarial history

Review status: rewritten after review; Merkle proof size and Raft's fault model corrected. Score: pending re-review.

The mental model: a database with no single writer of record

A blockchain is an append-only ledger replicated across mutually distrusting nodes, where the replicas agree on one ordered transaction history through a consensus protocol rather than a designated primary. Compare that with a conventional database: there, one accountable operator owns write authority, can update or delete any row, and is the trust anchor. If you dispute a row, your recourse is against that operator. On a blockchain, no single party can rewrite committed history without either out-spending the network's security budget or coordinating a social fork — and either event is observable.

The interview-grade one-liner: a blockchain is a replicated state machine. Transactions are deterministic state transitions; consensus picks the order; every node re-executes the same transitions and must land on the same state. Everything else — blocks, hashes, mining, staking — exists to make that agreement work among parties who don't trust each other.

"Immutable" means tamper-evident and tamper-expensive, not physically impossible. A majority coalition can rewrite a proof-of-work chain's recent history; a coordinated social fork (Ethereum's DAO response in 2016) can rewrite anything. What the design guarantees is that alteration is detectable and costly, not that it cannot happen.

Block anatomy: why editing one transaction breaks the chain

A block is a header plus an ordered list of transactions. The header carries the fields that make tampering cascade:

  • previous block hash — chains this block to all prior history
  • Merkle root — a single hash committing to every transaction in the block
  • timestamp and difficulty/nonce — consensus-relevant metadata

Hash pointers are the mechanism. Each header contains H(prev_header), so the headers form a chain where each one commits to everything before it. Change one transaction anywhere and its Merkle path changes, the Merkle root changes, that header changes, and every subsequent header — which embeds the previous header's hash — changes too. An attacker rewriting block N must produce a valid replacement for every block after N, and validity includes the consensus rule (proof-of-work difficulty or validator signatures), which is where the cost comes from.

A worked trace of the Merkle mechanism, with 4 transactions (labeled hypothetical values):

tx1..tx4
leaves:  h1=H(tx1)  h2=H(tx2)  h3=H(tx3)  h4=H(tx4)
level2:  h12=H(h1||h2)   h34=H(h3||h4)
root:    r = H(h12||h34)   → goes in the header

If tx3 is altered, h3 changes, so h34 changes, so r changes, so the header changes, so every later header is now invalid. A light client holding only the block header can verify tx3's inclusion from the two sibling hashes h4 and h12: recompute h3 = H(tx3), then h34 = H(h3||h4), then r = H(h12||h34), and check it equals the root in the header. A 4-leaf tree needs exactly two hashes per proof — log₂(4) — and that's why light clients don't need the full block.

What a hash does not do: it proves later bytes match committed bytes. It doesn't publish the document, preserve it, or prove it was true when committed. "We put the hash on-chain" is a commitment, not a record.

Consensus: what PoW, PoS and BFT actually guarantee

Replicas receive transactions at different times, in different orders, and some replicas are faulty or malicious. Consensus is the layer that makes them agree on one sequence despite that.

Proof of work (Bitcoin): block production requires finding a header hash below a difficulty target; the chain with the most accumulated work is canonical. Finality is probabilistic — an attacker controlling a majority of hash rate can build a competing chain, and their chance of overtaking the honest chain decays with each confirmation, but it never reaches exactly zero. The security budget is the cost of acquiring majority hash rate.

Proof of stake (Ethereum since the 2022 Merge): validators sign blocks after committing 32 ETH, and misbehavior — signing two conflicting blocks at the same height, surrounding votes — is slashable. Ethereum gives explicit finality: after two checkpoint epochs (~13 minutes), a block is finalized and reversing it requires at least one-third of staked ETH to be burned through provable protocol violations. That's economic finality: not physically impossible to revert, but reverting costs a quantifiable fortune and is attributable.

The nothing-at-stake problem is why PoS chains penalize equivocation: without slashing, a validator could sign every competing fork for free. The longest-chain rule alone doesn't fix it; the penalty does.

Byzantine-fault-tolerant protocols (Tendermint/Cosmos, PBFT descendants): a known validator set of n nodes tolerates f Byzantine faults, with n ≥ 3f+1. Finality is absolute: once a block commits, it does not revert absent a protocol bug. The trade is membership: you must know and govern the validator set, which is why these appear in permissioned and consortium settings.

Crash-fault-tolerant protocols (Raft, used as Hyperledger Fabric's default ordering service) tolerate node crashes and network partitions but make no claim about Byzantine behavior: a malicious ordering node can equivocate or drop transactions, and Raft will not detect it. Raft gives fast, simple ordering among nodes you already trust to be non-malicious — appropriate inside one organization, not for a consortium of rivals. Fabric also offers a BFT ordering option where the trust model demands it.

PoWPoS (slashing)BFT (permissioned)Raft (CFT ordering)
Finalityprobabilisticeconomic, explicit checkpointsabsolute once committedabsolute only if all nodes honest
Faults toleratedcrashes + minority hash equivocationcrashes + <⅓ Byzantine stakecrashes + <⅓ Byzantine validatorscrashes and partitions only
Open membershipyes, via hash-rate competitionyes, via stakeno, fixed validator setno, fixed cluster
Cost to rewriteacquire majority hash rateburn ≥⅓ of stakecompromise ≥f+1 identified orgscompromise any ordering node
Latency to finalityminutes-to-never-certain~13 min (2 epochs)one commit round, secondsone append, sub-second

The threat model: what the design prices

Every security property is a cost imposed on an attacker, and you should be able to name the price for each threat:

  • Double spend — the same funds spent twice. Answer: global ordering. Once one spend is in the canonical history, the other is invalid, no matter when it arrived. The receiver's exposure is the gap before finality, which is why you wait for confirmations proportional to value at risk.
  • 51% / majority attack — rewrite recent history by out-producing the chain. Answer: the chain's security budget. On Bitcoin this means majority hash rate; on Ethereum, a supermajority of stake that gets slashed for provable equivocation; on a BFT chain, corrupting a third of identified organizations. A 51% attack on a small PoW chain has actually happened (Ethereum Classic, repeatedly, e.g. 2019 and 2020 reorgs) — the model is real, not theoretical.
  • Sybil attack — one actor spawns thousands of identities to dominate a vote. Answer: make identities expensive. PoW prices them in energy and hardware; PoS in capital at risk; permissioned systems in vetted legal identity. Without Sybil resistance, any open consensus is trivially captured.

What none of these defend against: buggy validation rules, compromised keys, malicious smart contracts, or bad data fed in from outside. Consensus guarantees replication and ordering of whatever the rules accept — not correctness of the rules or the inputs.

Distribution, decentralization, and the permission spectrum

These get conflated constantly; interviewers probe the distinction.

  • Distribution is a topology fact: many nodes, replicated data. A single company running 50 nodes is distributed and fully centralized.
  • Decentralization is a control fact: no single party can unilaterally rewrite, censor, or upgrade. It's a spectrum measured by who controls consensus, client code, governance and key custody.
  • Permissionless (public Bitcoin/Ethereum): anyone can join, validate, and propose; Sybil resistance is economic. You buy censorship resistance and no-membership-governance; you pay with probabilistic or slow finality, throughput limits, and full data visibility.
  • Consortium (e.g. a trade-finance network of 8 named banks): membership is a governed set; BFT ordering gives sub-second absolute finality. You buy performance and known counterparties; you give up trustlessness — you now trust the governance of the member set.
  • Private/enterprise (Hyperledger Fabric inside one org or supply chain): Fabric separates the blockchain — an append-only transaction log — from the world state, a database holding current values for efficient queries. Channels partition confidentiality between subgroups; endorsement policies specify which organizations must execute and sign a transaction before ordering. You buy auditability and controlled access; if one org owns it all, ask honestly whether a database with an audit log would do the same job cheaper.

Permissioned does not mean centralized (a consortium genuinely distributes control across rivals), and permissionless does not mean trust-free: every chain trusts its software, cryptography, incentive design, governance process and key custody.

State models, nodes, and what applications still own

Two state models dominate. UTXO (Bitcoin): spendable records are discrete outputs; a transaction consumes existing outputs and creates new ones; parallel transaction verification is natural. Account/nonce (Ethereum): balances, nonces, code and storage per address; a transaction's nonce orders that account's transactions and prevents replay of the same signed transaction. Neither removes concurrency concerns — applications still handle pending transactions, ordering races, reorgs and idempotency.

Node roles: a full node independently validates every protocol rule and derives state from history — trust nothing, verify everything. An archive node additionally retains historical state. A light client verifies headers and Merkle proofs against a trusted root, trading weaker assumptions for less work. A wallet manages keys and signs transactions; it is not the ledger, and losing a non-custodial key is losing the account.

On-chain storage is expensive and replicated to every validator — store only what needs shared verification. Never put secrets or personal data on-chain because "we'll encrypt it": ciphertext and metadata persist forever, keys leak, and deletion rights conflict with replication. IPFS content identifiers address content but guarantee nothing about availability — someone must still pin and serve it.

Smart contracts are deterministic programs invoked by transactions: public, composable state, but fees on every call, ordering hazards, and a trust boundary at every oracle that feeds in external facts. Upgradeability adds admin keys and governance risk; immutability makes bugs permanent. Keep privileged roles minimal, document pause/upgrade powers, and treat contract events as signals, not guaranteed off-chain completion.

When to use it — and the answer interviewers are fishing for

Use a blockchain when: multiple parties need a shared ordered record, no single party can be the accountable writer of record, the parties can agree on deterministic validation rules and a governance process, and they accept replicated cost and reduced confidentiality. Use a conventional database when any one of those fails: one accountable operator exists, low latency and confidential queries dominate, or records must be corrected or deleted. "We need auditability" is not sufficient — an append-only table in Postgres gives auditability at a fraction of the cost.

At scale, the constraints shift: throughput is bounded by how much data every validator must process (block size × block rate), so public chains push execution off-chain or into rollups while anchoring commitments on-chain; fee markets price contention, so p99 costs spike under demand; and indexer/derived views lag, fork and need reorg-aware ingestion with idempotent writes and reconciliation against an authoritative node.

What interviewers probe, and what a weak answer sounds like:

  • "Explain immutability." Weak: "blocks can't be changed." Strong: names the assumptions — tamper-evident under hash chaining, expensive to rewrite under the security budget, reversible by majority collusion or social fork.
  • "What does a signature prove?" Weak: "it proves who sent it." Strong: the private-key holder authorized these bytes under this scheme — not civil identity, not authority over the off-chain asset, not truth of the payload.
  • "Why not just a database?" Weak: no answer. Strong: names the trust problem that requires removing the single writer of record, or concedes a database wins.
  • "What happens on a reorg?" Weak: blank. Strong: pending vs. included vs. finalized, application rollback, idempotent re-ingestion.
  • "Is Raft Byzantine-fault-tolerant?" Weak: "it's a consensus protocol, so yes." Strong: Raft is crash-fault-tolerant only; a malicious ordering node can equivocate undetected, so it suits trusted clusters, not adversarial consortia.

Likely follow-ups: how finality is defined on the specific chain you cite; what a 51% attack costs and what it can and cannot do (it cannot steal keys or mint outside the rules — it can double-spend and censor); how oracles break the trust model; and who governs the protocol upgrade path. The strong through-line: a blockchain is a bounded, testable trust model whose replication cost is justified by a specific coordination problem — not decentralization as a slogan.