Overview
Curated: · Written: · Reviewed:
Web3 delivery is reproducible state-machine engineering
Web3 tooling connects source code, compiler configuration, deployment artifacts, remote nodes, wallets and immutable production state. A green unit test is only one link. A production-grade workflow makes the whole chain reproducible: reviewed source and locked dependencies produce known bytecode; deployment uses an authorized signer on the intended network; verification matches the deployed artifact; monitoring reads canonical chain state; and recovery is rehearsed, not improvised.
The interview angle: senior Web3 roles probe whether you treat a contract system as a state machine with financial consequences, not a library with tests. Weak answers sound like a feature checklist — "we use Hardhat and Slither" — with no decision criteria, no failure modes, and no idea what a passing test actually proves. Strong answers name the tool, the trade-off it makes, and the incident it prevents.
Tooling decision: Foundry vs Hardhat
On a new protocol repo the first decision is the test stack, and it is a real trade-off, not a preference.
Foundry is Solidity-native: tests are contracts, cheatcodes (vm.prank, vm.deal, vm.expectRevert, vm.roll, mockCall) replace most mocking boilerplate, and fuzz and invariant testing are built in. Fuzz tests are declared by giving a parameter a type; the runner generates inputs and shrinks failures. That tight loop — write a property, get a shrunk counterexample in seconds — is why protocol teams doing heavy invariant work reach for it first. It loses where you need a rich off-chain ecosystem: TypeScript integration tests against a full app stack, mature plugin variety, or a team whose strength is JS.
Hardhat (with its EDR-based node and the ethers/viem test stack) wins when the deliverable includes off-chain code: indexers, keepers, relayers, a frontend. You write tests in TypeScript against the same contracts, and integration with JS tooling is frictionless. Its weaknesses are the mirror image: fuzz/invariant testing comes via plugins or extra setup, tests are slower to iterate, and property-based testing is less first-class.
Many production repos run both: Foundry for unit, fuzz and invariant suites; a thin Hardhat or viem layer for end-to-end wiring. If you say that, say why — the boundary is usually "property tests live in Solidity, integration with off-chain actors lives in TypeScript."
What interviewers probe: "Why did you pick your stack?" A weak answer is "Foundry is faster." A strong one names what each is bad at and where the team's boundary sits.
What contract tests must actually cover
A test suite for a contract is a specification of its state machine. The checklist a reviewer expects:
- Access control and auth paths. Every privileged function tested with an authorized caller and at least one unauthorized one. A role system tested only through the admin is untested.
- Revert conditions and custom errors. Assert the exact custom error and its arguments, not just that the call reverted.
vm.expectRevert(TransferRejected.selector)catches a revert happening for the wrong reason; a bareexpectRevertdoes not. - State transitions and invariants. Before/after state for every transition, plus the properties that must hold across all of them (next section).
- Boundaries. Zero, type max (
type(uint256).max), empty sets, first and last elements, rounding at 1 wei. Most production bugs live at boundaries, not on the happy path. - Reentrancy and cross-contract behavior. Test the state-before-external-call ordering explicitly: a malicious receiver that re-enters, and the assertion that the guard or checks-effects-interactions ordering holds.
- Events and off-chain integrators. Events are the indexer's API. Assert emission and indexed-argument correctness, because a broken event silently corrupts every downstream consumer.
A worked example of the difference between a weak and a strong assertion, in a Foundry test (Solidity 0.8.x):
function test_Withdraw_RevertWhenNotOwner() public {
vm.prank(alice);
vault.deposit(1 ether);
vm.prank(bob);
vm.expectRevert(NotOwner.selector); // asserts the *reason*, not just failure
vault.withdraw(1 ether);
assertEq(vault.balances(alice), 1 ether); // state untouched by the failed call
}
If withdraw reverted because of an unrelated underflow, the bare-expectRevert version passes and this one fails. That is the level of specificity interviewers look for.
Property-based, fuzz and invariant testing
An invariant is a property that must hold after any sequence of valid calls. For a vault: sum(userBalances) == address(this).balance. For a DEX pool: constant-product conservation after any swap. For a lending market: solvency — no user's collateral value ever goes negative. State the invariant mathematically first; the test just checks it.
Example-based tests check the paths you thought of. Fuzz tests check paths you didn't: randomized inputs find rounding errors, overflow-adjacent values and assumption violations. Stateful (invariant) fuzzing goes further — it generates call sequences across actors through handler contracts and asserts the invariant after every step. This is what catches the "individually-safe functions compose into an unsafe sequence" class of bug that unit tests structurally cannot.
Practical discipline:
- Bound only genuinely invalid inputs with
vm.assume; overuse starves the fuzzer of real cases. - Use handler contracts with bounded, realistic actions and ghost variables for accounting the contract doesn't store.
- Persist the failing seed. Every regression should preserve the smallest failing seed, sequence and environment.
- CI budget: run a few hundred to a few thousand fuzz runs per invariant in CI (Foundry's default is 256 runs per test; teams raise it for release branches), and schedule deep overnight runs with millions of runs locally or on dedicated hardware. Say the numbers you use.
Weak answer: "we do fuzz testing." Strong answer: "our three invariants are conservation, solvency and authorization; handlers bound actions to realistic actors; CI runs 1,000 runs per invariant on every PR and 100k overnight, and failures shrink to a seed we commit as a regression test."
Forked-mainnet and local-node testing
Local nodes (Anvil, Hardhat node) give fast, deterministic tests of your own contracts. Fork tests give you the rest of the world: real deployed liquidity, real oracle state, real adversary contracts. A fork is a view of an existing chain served by an RPC — an archive node when you pin a historical block, since pre-pinned state needs historical queries.
Rules that make forks trustworthy:
- Pin the block and chain ID. An unpinned
latestfork is not reproducible; a suite that passed Monday can fail Tuesday because a pool rebalanced. Record the fork URL, chain ID and block number with the release. - Impersonate only documented actors (
vm.prank/ Anvil'sanvil_impersonateAccount) and reset state between tests. - Know what a fork does not reproduce: mempool competition, validator/mev behavior, oracle heartbeat timing, bridge delays, future upgrades. A passing fork suite is evidence about that block hash, not about next week's mainnet. Re-pin when an integration's bytecode or storage layout changes.
Interview probe: "Your fork test passed but mainnet failed — why?" Weak answer: "flaky RPC." Strong answer walks the gap list above and names which one bit, then how the drill changes.
Static analysis, linting and where humans still win
Slither (static analysis), Aderyn, and Solhint (linting) run in CI on every PR, pinned to a version with triaged baselines. What they give you: pattern detection — reentrancy shapes, uninitialized storage, uninitialized proxy variables, unprotected initializers, arbitrary-jump/call patterns — across the whole codebase in seconds, including paths no test executes.
What they don't give you: correctness. A tool flags suspicious shapes; it cannot know that a delegatecall is your intended proxy pattern or that a rounding direction is the specified one. So the workflow is triage, not trust:
- Pin tool and rules; fail CI on agreed severities, document the rationale for every suppressed finding.
- Treat coverage runs as review prompts, not scores — coverage measures which code executed, not whether the right properties were asserted. 100% branch coverage with weak assertions is a green suite that encodes the wrong expected behavior. Mutation testing (or deliberate fault injection) reveals assertions that never detect semantic changes.
- Tool findings feed human review; they never replace specification. Formal verification (e.g., Certora-style, or Foundry's symbolic execution via
vm.assume+ Halmos-style solvers) proves properties relative to a spec — and inherits every bug in the spec. The hard part is the specification, not the solver.
The audit relationship: tooling narrows what an auditor must look at; it does not outsource the judgment. Weak answer: "Slither said it was fine." Strong answer: "Slither's 14 findings triaged to 2 real issues, 12 documented suppressions; the invariant suite covers the conservation property Slither can't reason about."
RPC reads, simulation and provider lifecycle
RPC reads are views from one endpoint at one chain head. Use explicit block tags and chain IDs, validate response shape, and handle reorgs, pruning, rate limits and endpoint disagreement. An eth_call at latest is one node's current head: it acquires no nonce, wins no block auction, and does not survive a reorg. Simulate with an explicit block tag, submit with a durable operation ID, reconcile from the canonical receipt. Treating a simulation revert as proof mainnet will revert — or a simulation success as proof it will land — is how teams double-spend nonces and ship "it worked in the console" incidents.
Wallet providers are eventful: accounts and chains change mid-session, connections drop, and user rejection is not an outage. Re-read state after provider events and before signing; never silently switch networks or treat a UI-remembered address as current authorization.
Typed signing (EIP-712) should bind domain, chain, verifying contract, nonce, deadline and action-specific fields, and display decoded intent. A signature proves key authorization over bytes, not that the signer understood them. Tests must cover replay across contracts and chains, nonce reuse, expired authorization, malleability handling and contract-wallet validation.
Deployment, upgrades and secrets
Deployment is a state transition with financial consequences. Generate a plan from reviewed artifacts — expected chain, deployer, nonce, artifact hashes, addresses, parameters, starting state — simulate it against the intended state, require environment and chain assertions, use least-privilege signers, record every transaction, verify postconditions before transferring authority. Deterministic addresses (CREATE2) help only when factory, deployer, salt and init-code hash are verified and squatting risk is addressed.
Source verification: published compilation inputs must reproduce the deployed bytecode. An exact match is stronger than a partial match, and neither proves the program is safe. Verify proxy and implementation separately; compare the onchain runtime hash.
Upgradeable systems need storage-layout compatibility, initializer safety and authorization tests. Diff layouts and selectors in CI, rehearse upgrade and rollback on a pinned fork, monitor admin changes. A rollback may be impossible after irreversible state migration — recovery is designed, not assumed.
Secrets never belong in source, command history, logs or frontend bundles. Separate deployer, upgrader and operational roles; prefer hardware or threshold signing for high-impact actions; simulate the exact payload and verify destination and chain on the signer; CI uses short-lived credentials and protected environments.
Events, indexing and production diagnostics
Events are an indexing interface, not the source of consensus truth. Indexers deduplicate by block and log identity, retain block hashes, rewind on reorg, and reconcile materialized views against contract state. APIs expose lag and confirmation tier. Missing logs, overloaded RPCs and decoder-version changes must not silently corrupt balances.
When something breaks in production, the diagnostic bundle is: chain, block, tx hash and receipt, decoded calldata and revert reason, trace, artifact hash, and relevant state — not a screenshot without a transaction hash. Reproduce on a pinned fork of the incident block, and distinguish contract failure, provider failure, wallet failure and indexing failure before proposing a fix.
The production invariant
Operate with evidence: locked tools and dependencies, deterministic builds, property-oriented tests, pinned fork scenarios, gas and size budgets, deployment manifests, exact source verification, independent RPC checks, alerts for privileged or anomalous activity. The invariant to defend in an interview: another qualified engineer can rebuild, simulate, deploy, verify and diagnose the system from retained inputs — no single laptop, no single hosted service, no "it worked on mine."
Likely follow-ups to be ready for: "Walk me through your CI pipeline for a contract repo." "A fuzz test found a conservation violation — what do you do in the next hour?" "How do you test an upgrade against live state?" "What does your deployment runbook say when the third of five transactions fails?" Each of those is answered with the artifacts above, not with a tool name.
