Overview
Curated: · Written: · Reviewed:
Consensus chooses one ordered history under explicit failure assumptions
Consensus is the protocol by which distributed participants agree on a sequence of valid state transitions despite delay, crashes and, in some systems, adversarial behavior. It does not make invalid application rules correct, make submitted facts true, keep private keys safe or guarantee that every client can reach the network. A production design begins with safety—honest nodes do not finalize conflicting decisions—and liveness—the system eventually makes progress—then states precisely when each property is expected to hold.
Fault models and the thresholds that follow
Failure models are the first thing an interviewer will pin down, because every number in the design follows from them. Crash-fault-tolerant protocols such as Raft assume faulty replicas stop or lose messages but do not equivocate with valid credentials; a majority of survivors is enough, so n = 2f + 1 replicas tolerate f crashes. Byzantine-fault-tolerant protocols admit arbitrary behavior, including signing conflicting votes or sending different messages to different peers, and the threshold drops to f < n/3: with 3f + 1 replicas, quorums larger than 2f overlap in honest weight no matter how the f Byzantine nodes vote. A system cannot claim Byzantine tolerance merely because it has several replicas—quorum size, authenticated voting, independent operators, membership rules, synchrony assumptions and recovery determine the real fault threshold. A weak answer names a protocol and stops there; a strong one states the failure model, the replica count, and the quorum arithmetic.
Network timing: where progress guarantees end
In a fully asynchronous network, a deterministic protocol cannot guarantee termination with even one crash fault (FLP). Practical protocols use synchrony or partial-synchrony assumptions: safety survives arbitrary delay, while liveness resumes after message delays become bounded and an honest leader or proposer is selected. Timeouts are progress mechanisms, not evidence a peer is malicious—delay, loss, overload and partition look identical from a timeout. Aggressive timeouts cause needless view or round changes; slow timeouts extend outage recovery. Expect the follow-up: "what does your system guarantee during a partition?" The correct answer is that safety holds and liveness does not, and any system claiming both under partition is claiming something consensus theory rules out.
Quorum intersection and split brain
Quorum intersection is the core of voting protocols. With 3f + 1 equally weighted Byzantine replicas, two quorums of size greater than 2f must share at least 2f + 1 − f = f + 1 replicas, so at least one honest replica sits in both—and that replica's locking and voting rules prevent two conflicting certificates. The same majority reasoning is what prevents split brain in Raft: two majorities of n = 2f + 1 nodes intersect in at least one node, so two conflicting decisions cannot both be committed. Stake-weighted systems replace node count with voting power, which cuts both ways: running ten validators under one administrator, cloud account or signing service does not create ten independent fault domains.
Leader-based BFT: HotStuff and CometBFT
Leader-based BFT protocols rotate leaders across views or rounds. A leader proposes; validators validate and vote; a quorum certificate justifies progression or commitment. Locking rules stop honest validators from supporting conflicting histories, while view-change evidence lets a new leader preserve the safest known proposal. HotStuff restructures this family for linear communication in its steady path and responsiveness after synchrony, but deployments still depend on pacemaker, networking, storage and membership correctness.
CometBFT illustrates explicit rounds: proposal, prevote and precommit repeat until more than two-thirds voting power precommits a block. Validators may lock after a qualifying prevote set and only unlock under justified later evidence. Nil votes and timeouts permit progress without endorsing a missing or invalid proposal. Signed votes bind height and round; otherwise a valid vote could be replayed into a different decision.
Probabilistic finality: proof of work and fork choice
Proof of work makes proposing history costly through computation. Bitcoin nodes validate rules and follow the chain with the most accumulated proof of work—not whichever chain has the most block objects or arrived first. Confirmation depth increases confidence probabilistically because an alternative history must outpace the honest chain's work rate to catch up; it never reaches certainty. Security depends on honest hash rate, incentives, propagation, difficulty adjustment, implementation diversity and the economic value at risk. Hash rate is not finality, and a fixed confirmation count is not equally safe for every value and threat model—an exchange moving six figures needs deeper confirmation than a coffee purchase.
Fork choice answers "which head do I build on now?": longest-chain picks the most accumulated work, GHOST/heaviest-subtree follows the subtree with the most weight, which dampens the advantage a withholder gets from withholding blocks on a stale branch. Exchanges and bridges care because a reorg means credited funds can vanish: a bridge that releases on inclusion rather than finality is the classic exploit story. Expect to be asked how a protocol resolves two valid competing blocks at the same height and why that rule converges.
Deterministic finality: proof of stake and finality gadgets
Proof of stake replaces external computation with slashable or otherwise accountable stake. Validators propose and attest; fork choice selects a head while a finality gadget can justify and finalize checkpoints after supermajority votes. Slashing punishes provable equivocation or conflicting votes, while inactivity penalties help the chain recover when participants are offline. Stake concentration, custodians, correlated clients, censorship and governance remain first-class risks even if the nominal validator count is large. BFT-style voting gives deterministic finality in rounds rather than depth: once a certificate exists, reversion requires violating stated assumptions (a third of voting power equivocating) rather than luck. But deterministic finality is still conditional—compromised keys, incorrect membership, long-range history, client bugs or governance recovery can challenge its operational meaning. Light clients require a trusted checkpoint, authenticated validator-set transitions, a trust period and correct time assumptions; verifying signatures against an attacker-selected historical validator set is insufficient.
The two must be reported separately: fork choice selects the current head; finality marks a prefix whose reversion would violate protocol assumptions or incur defined economic consequences. A transaction can be broadcast, included, confirmed and finalized at different times, and applications must model these states separately, choose risk-based settlement thresholds, handle reorganizations, and delay or compensate irreversible off-chain work.
Consensus and execution
Consensus orders bytes; the replicated state machine must parse and execute those bytes deterministically. Nondeterministic clocks, random generators, floating-point differences, external HTTP calls or unordered iteration can make honest nodes derive different state from the same block. Applications must define canonical encoding, deterministic execution, state commitments and versioned upgrades. Interviewers probe this boundary: a common weak answer conflates ordering with execution validity and cannot say what happens when two nodes disagree about the result of a transaction everyone agreed to include.
Mempools, membership and economics
Mempools are not consensus. Nodes can see different pending transactions, apply different admission policy and reorder or drop candidates. A proposer may censor, extract value or choose a subset while still proposing a protocol-valid block. Products should never treat mempool observation as acceptance; they need durable transaction identity, replacement and nonce semantics, canonical receipt checks and finality-aware reconciliation.
Membership changes are consensus-critical. Adding, removing or reweighting validators must preserve quorum intersection across the transition. Joint-consensus, delayed activation or protocol-specific validator updates prevent two disjoint configurations from each believing they can decide. Key rotation, unbonding periods, weak-subjectivity windows and recovery procedures must align so old credentials cannot fabricate an alternative history acceptable to new clients.
Economic consensus adds incentives but not automatic alignment. Rewards, penalties, slashing, fee markets and delegator behavior can deter attacks, yet bugs, bribery, derivatives, external positions and correlated infrastructure change incentives. Analyze the cost and profit of censorship, reorganization, finality delay and safety violation; do not reduce the threat model to one percentage threshold.
Operating and choosing
Operate consensus as a distributed safety system. Persist signed-vote state before transmission to prevent double signing after crashes; isolate validator keys; monitor height, round, head and finalized lag, missed votes, equivocation evidence, peer diversity, quorum participation, reorg depth and client/version concentration. Test partitions, asymmetric delay, invalid proposals, offline and Byzantine leaders, disk rollback, clock skew, validator-set transitions, software upgrades and disaster recovery. Incident response must state whether safety is intact, which liveness assumption failed, and what authenticated evidence permits restart.
The right protocol follows requirements. A single accountable organization may need a replicated database with crash tolerance, not a public blockchain. A consortium may need identified BFT voting and deterministic finality. An open network may accept probabilistic or economic finality for permissionless participation. Compare latency, throughput, adversary threshold, membership, censorship, privacy, cost, governance, observability and recovery against credible simpler alternatives. Consensus is justified only when its trust distribution solves a real coordination problem. When an interviewer asks "why not Postgres with two replicas here?", that is the real question being graded.
