Top 100 Quantum Software Engineer Interview Questions and Answers
The questions most likely to actually come up in your Quantum Software Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1Your three-qubit circuit excites qubit 0 and the histogram key reads 001. Is that right?(show answer)
The first thing I would pin down about reading a measured bit string in the framework's qubit order is which convention the numbers are stated in.
Whether it is right depends entirely on a convention the circuit diagram does not show. A framework that treats qubit 0 as the least significant bit prints it rightmost, so 001 is correct there and wrong on a stack that orders the key by wire index from the top.
Concretely, fix the convention in one place, assert it in a test with an asymmetric fixture rather than a symmetric one, and convert at the boundary where results leave the framework instead of at each call site.
The reason for that specificity is a failure I have seen: A team compared a device histogram against a reference distribution generated in another framework and reported a fidelity of 0.31 for a circuit that was working; the two stacks disagreed about bit order and nothing raised an error.
One excited qubit, two conventions.
| Excited qubit | Little-endian key | Wire-order key |
|---|---|---|
| 0 | 001 | 100 |
| 1 | 010 | 010 |
| 2 | 100 | 001 |
I would not consider it settled without evidence: Excite exactly one qubit of a three-qubit register and assert the printed key, the statevector index and the classical bit mapping against values worked out by hand.
A convention that is never asserted is a convention that will differ somewhere.
Curated: · Written: · Reviewed:
QA-2A colleague describes a qubit in superposition as secretly being 0 or 1 with equal odds. What is wrong with that?(show answer)
I would answer what a superposition is not by separating what the mathematics guarantees from what the device delivers.
A coherent superposition and a classical coin flip produce identical statistics in one basis and different statistics in another, so the classical story is not a simplification but a prediction that fails. Amplitudes carry relative phase; probabilities do not.
Concretely, show the difference experimentally rather than by argument: prepare the plus state and a classical mixture, measure both in the computational basis, then apply a Hadamard to each and measure again.
The reason for that specificity is a failure I have seen: A demonstration explained superposition as a hidden fifty-fifty choice, and the same team later could not explain why an interference-based algorithm returned a deterministic answer, so they spent a week looking for a bug in a correct circuit.
Same computational-basis counts, different physics.
| Preparation | P(0) in Z | P(0) in X |
|---|---|---|
| plus state | 0.50 | 1.00 |
| classical mixture | 0.50 | 0.50 |
I would not consider it settled without evidence: Measure both preparations after a basis change and require the coherent state to give a deterministic outcome where the mixture stays at fifty-fifty.
The basis change is the experiment that separates the two.
Curated: · Written: · Reviewed:
QA-3A compiler emits a block equal to minus your intended unitary. Does it matter?(show answer)
This is an area where the ideal circuit and the circuit that executes global phase versus relative phase are different programs.
Global phase is unobservable on its own, so the block is a legitimate replacement anywhere it is applied directly. It stops being legitimate the moment the block is placed under a control, because the phase then becomes relative between the two branches and changes the interference the enclosing algorithm relies on.
Concretely, record global phase explicitly on any circuit object that may later be controlled, and test synthesized blocks in the controlled form they will be used in rather than standalone.
The reason for that specificity is a failure I have seen: A synthesized block equal to minus the target passed every standalone equivalence check, then produced a phase-estimation result off by exactly one bin because the controlled version differed from the intended one by a Z on the control.
Where the harmless phase becomes visible.
| Use | U vs minus U | Observable difference |
|---|---|---|
| applied directly | equivalent | 0 |
| under a control | not equivalent | phase estimation off by 1 bin |
I would not consider it settled without evidence: Compare the full unitary including phase for any block destined for a controlled construction, not the measurement distribution.
Phase that is global today is relative tomorrow.
Curated: · Written: · Reviewed:
QA-4You need a probability estimated to one percent. How many shots do you budget?(show answer)
My answer to the Born rule and shot noise starts from the resource count rather than from the asymptotic claim.
A measured probability is a binomial estimate, so its standard error is the square root of p times one minus p over the shot count. One percent absolute accuracy near p equal to one half needs about 2,500 shots, and one tenth of a percent needs 250,000.
Concretely, derive the shot count from the accuracy the decision requires, then state the resulting confidence interval next to every reported probability instead of quoting the point estimate alone.
The reason for that specificity is a failure I have seen: A team declared a circuit fix successful because a probability moved from 0.512 to 0.523 over 1,024 shots, a difference well inside the 0.016 standard error, and the change was later shown to have no effect at all.
Shots against accuracy near p = 0.5.
| Shots | Standard error | Useful for |
|---|---|---|
| 1,024 | 0.016 | coarse checks |
| 4,096 | 0.008 | routine comparison |
| 250,000 | 0.001 | calibration-grade claims |
I would not consider it settled without evidence: Recompute the standard error from the shot count and require the observed change to exceed several times that interval before it is called a change.
Quote the interval or the number means nothing.
Curated: · Written: · Reviewed:
QA-5Does the no-cloning theorem stop you from preparing two qubits in the same state?(show answer)
I would treat no-cloning and what it does not forbid as a claim that has to survive being recomputed from the raw counts.
No-cloning forbids a single unitary that copies an arbitrary unknown state, because such a map cannot be linear. It says nothing about preparing two qubits from the same known classical description, which is what a state-preparation routine does every time it runs.
Concretely, distinguish the three cases explicitly in design discussions: re-preparing from a description is free, copying a basis state with a CNOT works for basis states only, and duplicating an unknown state is impossible.
The reason for that specificity is a failure I have seen: A protocol design assumed a received unknown state could be duplicated so one copy could be measured and the other kept; the implementation used a CNOT, which entangles rather than copies, and the retained qubit was maximally mixed with fidelity 0.5 to the original.
What a CNOT does to two inputs.
| Input | Output | Copy achieved |
|---|---|---|
| zero state | both zero | yes |
| plus state | Bell pair | no |
I would not consider it settled without evidence: Apply the proposed copy operation to a state in the X basis and check the reduced state of each output rather than testing on basis states only.
A CNOT copies labels, not states.
Curated: · Written: · Reviewed:
QA-6How many qubits can you simulate exactly on a 256 GB machine?(show answer)
The useful question for the memory cost of a statevector simulation is what the hardware does that the simulator never showed.
A dense statevector holds 2 to the n complex amplitudes at 16 bytes each, so 30 qubits is 17.2 GB and 34 qubits is 275 GB. The limit is a hard wall rather than a slowdown, and it halves in qubit count when the object becomes a density matrix with 4 to the n entries.
Concretely, choose the simulation method from the question being asked, and state the width limit and the method together: stabilizer for Clifford circuits, tensor networks with a declared bond dimension for low-entanglement structure, and trajectory sampling with a stated trajectory count for noise.
The reason for that specificity is a failure I have seen: A validation plan promised noisy simulation of a 20-qubit circuit and discovered at run time that the density matrix needed 17.6 TB; the team quietly dropped the noise model and reported an ideal result under the original heading.
Exact simulation at 16 bytes per amplitude.
| Qubits | Statevector | Density matrix |
|---|---|---|
| 17 | 2.1 MB | 275 GB |
| 30 | 17.2 GB | beyond any machine |
| 34 | 275 GB | beyond any machine |
I would not consider it settled without evidence: Compute the memory requirement for the chosen representation before scheduling the run, and record which method produced each published number.
The method is part of the result.
Curated: · Written: · Reviewed:
QA-7When is a statevector the wrong object to describe your qubits?(show answer)
I would settle when a density matrix is required instead of a statevector against a shot budget before arguing about the algorithm.
A statevector describes a pure state only. Any probabilistic ensemble, any subsystem of an entangled register, and any state that has interacted with an environment requires a density matrix, which is Hermitian, positive semidefinite and trace one.
Concretely, use purity, the trace of rho squared, as the discriminator: it is exactly one for a pure state and lower for a mixed one, and it is cheap enough to assert in a test.
The reason for that specificity is a failure I have seen: A noise study normalized its output back to a unit vector after each channel application, which silently converted every mixed state into a pure one and reported a decoherence curve that stayed flat at fidelity 0.99 for 40 gate layers.
Purity separates the two objects.
| State | Trace of rho squared | Representable as a ket |
|---|---|---|
| pure | 1.00 | yes |
| slightly mixed | 0.92 | no |
| maximally mixed qubit | 0.50 | no |
I would not consider it settled without evidence: Assert that the simulated object's purity falls as noise is applied rather than assuming the representation can absorb it.
Renormalizing a mixed state hides the physics you were measuring.
Curated: · Written: · Reviewed:
QA-8Each qubit of your Bell pair looks maximally mixed. Was the state prepared badly?(show answer)
The judgement in partial trace and locally mixed subsystems is which assumption is load-bearing and whether it was checked.
No: a maximally mixed reduced state is the signature of maximal entanglement, not of a failed preparation. The partial trace discards the correlations, so the local description of half of a Bell pair is exactly the same as the description of a fair classical coin.
Concretely, judge the preparation on the joint state instead, through joint parity measurements in two bases or a fidelity witness, because no single-qubit measurement can distinguish an entangled pair from a mixture.
The reason for that specificity is a failure I have seen: An acceptance test measured each qubit of a Bell preparation separately, found 50/50 outcomes on both, and rejected the device; the pair was in fact entangled with a joint fidelity of 0.94.
What each view of a Bell pair shows.
| Observable | Ideal value | Tells you |
|---|---|---|
| single-qubit Z | 0.00 | nothing about the pair |
| ZZ parity | 1.00 | correlation in Z |
| XX parity | 1.00 | coherence, not just correlation |
I would not consider it settled without evidence: Measure ZZ and XX parity on the pair and require both correlations rather than testing the qubits one at a time.
Entanglement lives in the joint state and nowhere else.
Curated: · Written: · Reviewed:
QA-9Can you draw your two entangled qubits on Bloch spheres?(show answer)
Where candidates lose the interview on the Bloch sphere and its limits is treating an exact-state result as a device result.
You can draw their reduced states, but those two pictures do not contain the joint state. The Bloch representation is complete for one qubit and loses exactly the correlations that make a multi-qubit state interesting.
Concretely, use Bloch pictures for single-qubit calibration and gate tuning, and switch to joint quantities such as parity expectations, a density matrix or a witness the moment more than one qubit is in scope.
The reason for that specificity is a failure I have seen: A debugging session drew both halves of a Bell pair at the origin of their Bloch spheres, concluded the qubits were dead, and replaced a working calibration with one that produced a genuinely mixed state at fidelity 0.61.
Same local pictures, different joint states.
| Joint state | Bloch vector A | Entangled |
|---|---|---|
| Bell pair | origin | yes |
| classical mixture of 00 and 11 | origin | no |
I would not consider it settled without evidence: Show that two very different joint states can give identical single-qubit Bloch vectors before relying on the picture for a multi-qubit diagnosis.
Two spheres are not a two-qubit state.
Curated: · Written: · Reviewed:
QA-10How do you decide whether a two-qubit pure state is entangled?(show answer)
I would answer recognising a product state by naming the measurement that would contradict it.
A pure bipartite state is entangled exactly when it does not factor across the stated partition, which for two qubits is the condition that the determinant-like combination of its amplitudes is non-zero. The partition is part of the question: a state can be entangled across one cut and separable across another.
Concretely, compute the Schmidt rank across the declared partition, or equivalently check whether the reduced density matrix is pure, and state the partition next to the answer.
The reason for that specificity is a failure I have seen: A four-qubit state was called entangled without naming a cut; it was a product of two Bell pairs, and a protocol that assumed genuine four-party entanglement failed with a success probability of 0.25 instead of the expected 1.0.
Same register, different cuts.
| Cut | Schmidt rank | Entangled |
|---|---|---|
| qubits 01 vs 23 | 1 | no |
| qubit 0 vs 123 | 2 | yes |
I would not consider it settled without evidence: Report the Schmidt rank or reduced-state purity for the specific partition the protocol depends on.
Entanglement is a statement about a cut.
Curated: · Written: · Reviewed:
QA-11After you measure one qubit of a register mid-circuit, what happens to the rest?(show answer)
The engineering content of the post-measurement state is the compilation and the error budget, not the notation.
A projective measurement updates the whole state, not just the measured qubit: the register is projected onto the subspace consistent with the outcome and renormalized. For an entangled register that changes the description of qubits that were never touched.
Concretely, model mid-circuit measurement as an instrument with both an outcome and a state update, and write the conditional branches down explicitly when subsequent gates depend on the result.
The reason for that specificity is a failure I have seen: A circuit measured an ancilla to check a parity and then continued as though the data register were unchanged; the ancilla had not been uncomputed, so the measurement collapsed the data superposition and the algorithm's success probability dropped from 0.98 to 0.51.
Measuring one half of a Bell pair.
| Outcome on A | Probability | State of B |
|---|---|---|
| 0 | 0.50 | zero state |
| 1 | 0.50 | one state |
I would not consider it settled without evidence: Simulate both measurement branches and compare the conditional states against the assumption the rest of the circuit makes.
A measurement updates everything it was correlated with.
Curated: · Written: · Reviewed:
QA-12Your hardware only measures in the computational basis. How do you measure an X observable?(show answer)
Before trusting anything about choosing the measurement basis I would write down what a wrong answer would look like.
Basis change is part of the observable, not a separate concern: measuring X means applying a Hadamard and then measuring Z. Every expectation value a device reports is the computational-basis measurement of some rotated circuit.
Concretely, attach the required rotations to the observable definition so that the same Hamiltonian term always compiles to the same basis change, and group commuting terms so one rotated circuit serves several terms.
The reason for that specificity is a failure I have seen: An energy estimate omitted the basis rotation for the X terms of a Hamiltonian, so those terms were evaluated in Z and contributed near zero; the reported ground-state energy was off by 0.4 hartree and looked plausible.
Rotations that make a Z measurement an X or Y one.
| Observable | Pre-measurement gate | Shots shared with |
|---|---|---|
| Z | none | other Z terms |
| X | H | other X terms |
| Y | S dagger then H | other Y terms |
I would not consider it settled without evidence: Estimate a term whose exact value is known on a small instance and require agreement within the shot-noise interval before trusting the full Hamiltonian.
An observable includes the rotation that measures it.
Curated: · Written: · Reviewed:
QA-13Why can a quantum circuit not contain an AND gate as written classically?(show answer)
The first thing I would pin down about unitarity and reversibility is which convention the numbers are stated in.
Closed-system quantum evolution is unitary and therefore reversible, and a two-input AND destroys information because three of its four inputs map to the same output. Irreversible classical logic has to be embedded in a reversible form before it can appear in a circuit.
Concretely, implement classical logic with Toffoli-style constructions that keep the inputs and write the result into an ancilla, then uncompute the scratch so the ancilla returns to a known state.
The reason for that specificity is a failure I have seen: A predicate written as a chain of irreversible operations left 6 ancillas entangled with the search register, which destroyed the interference the algorithm depended on and produced a flat output distribution that was blamed on hardware noise for two weeks.
Classical AND embedded reversibly.
| Construction | Qubits | Ancillas returned to zero |
|---|---|---|
| irreversible AND | not expressible | not applicable |
| Toffoli with ancilla | 3 | yes, after uncompute |
I would not consider it settled without evidence: Assert that every ancilla returns to the zero state with probability one when the oracle runs on a superposition input.
Reversible means the scratch has to be cleaned.
Curated: · Written: · Reviewed:
QA-14Your device offers a fixed basis set. What does that cost your arbitrary rotations?(show answer)
I would answer universal gate sets and synthesis cost by separating what the mathematics guarantees from what the device delivers.
A discrete universal set approximates an arbitrary single-qubit rotation to accuracy epsilon in a number of gates that grows logarithmically in one over epsilon, so higher precision is affordable but not free. On hardware with continuous rotations the cost appears instead as calibration and pulse-level error.
Concretely, set the synthesis tolerance from the circuit's total error budget rather than per gate, since a thousand rotations each synthesized to 10 to the minus 3 accumulate an error of the same order as one uncorrected gate fault.
The reason for that specificity is a failure I have seen: A circuit with 1,200 synthesized rotations used a default tolerance of 10 to the minus 3 and drifted from its intended unitary badly enough that the expectation value it produced was wrong by 8 percent, with no gate individually at fault.
Synthesis tolerance against accumulated error.
| Per-gate tolerance | Rotations | Worst-case total |
|---|---|---|
| 1e-3 | 1,200 | 1.2 |
| 1e-6 | 1,200 | 0.0012 |
I would not consider it settled without evidence: Sum the per-gate synthesis error across the circuit and compare it against the accuracy the result requires.
Approximation error accumulates the way gate error does.
Curated: · Written: · Reviewed:
QA-15Which number would you ask for first when comparing two implementations of the same algorithm?(show answer)
This is an area where the ideal circuit and the circuit that executes two-qubit gate count as the cost metric are different programs.
The transpiled two-qubit gate count, because two-qubit error typically exceeds single-qubit error by an order of magnitude and dominates the failure probability of any circuit long enough to be interesting. Single-qubit counts and abstract depth are weak predictors by comparison.
Concretely, report the count after transpilation to the specific target rather than before, and pair it with the device's current two-qubit error so the number converts into an expected success probability.
The reason for that specificity is a failure I have seen: Two designs were compared on abstract gate count, which favoured the one that used 12 long-range interactions; after routing on a line topology it needed 39 CNOTs against the other design's 18, and its success probability was 0.76 against 0.88.
Abstract cost against routed cost.
| Design | Abstract CNOTs | Routed CNOTs | Success at 0.7% error |
|---|---|---|---|
| A | 12 | 39 | 0.76 |
| B | 16 | 18 | 0.88 |
I would not consider it settled without evidence: Transpile both candidates to the current target and compare two-qubit counts and estimated success rather than source-level metrics.
Compare the circuits that would actually run.
Curated: · Written: · Reviewed:
QA-16Your algorithm assumes all-to-all connectivity and the device is a lattice. What does that cost?(show answer)
My answer to qubit routing and SWAP overhead starts from the resource count rather than from the asymptotic claim.
Every interaction between qubits that are not physically coupled has to be routed, and each SWAP inserted by the router costs three CNOTs. Routing overhead is a property of the mapping between the algorithm's interaction graph and the device's coupling graph, not of the algorithm alone.
Concretely, spend the effort on the initial layout before optimizing gates, because placing the heaviest interaction pairs on physically coupled edges usually removes more error than any peephole optimization applied afterwards.
The reason for that specificity is a failure I have seen: A five-qubit circuit with 12 CNOTs and depth 18 was mapped without layout tuning onto a line, picked up 9 SWAPs, and executed as 39 CNOTs at depth 46, losing about a third of its signal to routing.
Same circuit, two layouts.
| Layout | SWAPs | CNOTs | Depth |
|---|---|---|---|
| default | 9 | 39 | 46 |
| interaction-aware | 2 | 18 | 24 |
I would not consider it settled without evidence: Compare the two-qubit count under the default layout against a layout chosen from the algorithm's interaction graph, on the same target.
Layout is a result, not a compiler detail.
Curated: · Written: · Reviewed:
QA-17How do you tell whether a circuit is too deep for a device before you run it?(show answer)
I would treat circuit depth against coherence time as a claim that has to survive being recomputed from the raw counts.
Compare the circuit's scheduled duration against the qubits' coherence times, not its gate count against a rule of thumb. A circuit whose critical path runs 60 microseconds on qubits with a T2 near 100 microseconds has spent most of its coherence budget before the measurement.
Concretely, ask the transpiler for the scheduled duration on the actual target, since gate durations differ by qubit pair, and treat idle qubits as accumulating decoherence rather than as free.
The reason for that specificity is a failure I have seen: A circuit was accepted because it used only 400 gates, but routing put its critical path at 210 microseconds against a T2 of 90; the output distribution was indistinguishable from uniform and three days went into looking for a logic error.
Two circuits with similar gate counts.
| Circuit | Gates | Duration | T2 | Usable |
|---|---|---|---|---|
| A | 400 | 18 us | 90 us | yes |
| B | 420 | 210 us | 90 us | no |
I would not consider it settled without evidence: Read the scheduled duration from the transpiled circuit and compare it with the coherence times in the same calibration record.
Duration, not gate count, is what decoherence sees.
Curated: · Written: · Reviewed:
QA-18Why do Z rotations often cost nothing on superconducting hardware?(show answer)
The useful question for virtual Z rotations is what the hardware does that the simulator never showed.
A Z rotation can be implemented as a frame change on subsequent pulses rather than as a physical operation, so it takes zero time and adds no gate error. That is why an arbitrary single-qubit rotation typically compiles to two calibrated half-pulses and three virtual Z rotations.
Concretely, let the transpiler decompose into the device's native form instead of hand-writing rotations, and read cost from the scheduled circuit, where virtual gates correctly show zero duration.
The reason for that specificity is a failure I have seen: An optimization pass was written to minimize total gate count and removed Z rotations at the cost of adding two physical pulses per qubit, increasing circuit duration by 18 percent while the reported gate count fell.
Cost of one arbitrary single-qubit rotation.
| Component | Count | Duration each |
|---|---|---|
| virtual Z | 3 | 0 ns |
| calibrated half-pulse | 2 | 20 ns |
I would not consider it settled without evidence: Compare scheduled duration and two-qubit count before and after an optimization rather than the raw instruction count.
Count the gates that take time.
Curated: · Written: · Reviewed:
QA-19Explain how a controlled operation writes information into the control qubit's phase.(show answer)
I would settle phase kickback in controlled operations against a shot budget before arguing about the algorithm.
When the target is prepared in an eigenstate of the controlled operation, applying the control leaves the target unchanged and multiplies the control's excited branch by the eigenvalue. The information moves into the control's relative phase, where interference can read it.
Concretely, prepare the eigenstate deliberately — the minus state for a controlled X, for instance — and check that the target register is unchanged at the end, since a target that has moved means the kickback did not happen as intended.
The reason for that specificity is a failure I have seen: An implementation prepared the target in the zero state instead of the minus state, so the controlled operation entangled the registers rather than imparting a phase; the interference step then produced a uniform distribution and the algorithm returned random answers at a 3 percent success rate.
Target preparation decides the outcome.
| Target prepared as | Effect on control | Registers entangled |
|---|---|---|
| minus state | phase of minus one | no |
| zero state | none | yes |
I would not consider it settled without evidence: Assert that the target register returns to its prepared eigenstate and that the control's phase carries the expected value on a two-qubit instance.
Kickback is a statement about eigenstates.
Curated: · Written: · Reviewed:
QA-20A circuit measures a qubit and applies a gate conditioned on the result. Can you avoid the measurement?(show answer)
The judgement in the deferred measurement principle is which assumption is load-bearing and whether it was checked.
Yes: measurement followed by classically controlled operations is equivalent to a coherent controlled operation followed by a measurement at the end. The equivalence is what lets a simulator run a dynamic circuit without mid-circuit collapse.
Concretely, use the deferred form for reasoning and simulation, and the measured form on hardware where deferring would extend the circuit's duration or need qubits the device does not have.
The reason for that specificity is a failure I have seen: A team assumed the two forms were interchangeable on hardware and deferred all measurements in an error-detection routine, which extended the circuit's duration from 30 to 140 microseconds and pushed the qubits past their coherence budget.
Two encodings of the same logic.
| Form | Extra qubits | Duration |
|---|---|---|
| deferred | 4 | 140 us |
| mid-circuit measurement | 0 | 30 us |
I would not consider it settled without evidence: Simulate the deferred form to establish correctness and then measure the dynamic form's duration and success rate on the target before choosing one.
Equivalent on paper is not equivalent in duration.
Curated: · Written: · Reviewed:
QA-21What does a device need to support before you can write a dynamic circuit?(show answer)
Where candidates lose the interview on mid-circuit measurement and feed-forward is treating an exact-state result as a device result.
Mid-circuit measurement alone is not enough: the device must also expose classical registers, conditional gates and a feed-forward latency short enough that the qubits still hold coherence when the conditional operation runs. Feed-forward latency is a hardware specification, not a compiler setting.
Concretely, read the target's supported instructions and conditional latency at submission time and fall back to a deferred formulation when the capability is absent, rather than discovering the gap in a failed job.
The reason for that specificity is a failure I have seen: A dynamic circuit was submitted to a backend whose feed-forward latency was 4 microseconds against a T2 of 60; the conditional branch executed on qubits that had lost most of their coherence and the protocol's fidelity came out at 0.55 against 0.93 in simulation.
What a dynamic circuit needs.
| Capability | Present | Consequence if absent |
|---|---|---|
| mid-circuit measurement | yes | no dynamic circuit |
| conditional gates | yes | no feed-forward |
| latency under 1 us | no | fidelity 0.55 |
I would not consider it settled without evidence: Measure the protocol's fidelity on the device with the conditional path exercised, not only in a simulator that treats feed-forward as instantaneous.
Conditional logic costs coherence while it waits.
Curated: · Written: · Reviewed:
QA-22Your algorithm needs a clean ancilla halfway through. Reset or allocate?(show answer)
I would answer reset versus allocating a fresh qubit by naming the measurement that would contradict it.
Reset reuses a physical qubit at the cost of time and residual excitation, while allocating another qubit costs width and possibly worse connectivity. Reset is not perfect: a device with 1 percent residual excitation leaves the ancilla in the wrong state one run in a hundred.
Concretely, choose by measuring both on the target: compare the reset's duration and residual excitation against the routing overhead the extra qubit introduces, and record which was chosen and why.
The reason for that specificity is a failure I have seen: An error-detection routine reset an ancilla 12 times per shot on a device with 2 percent reset error, which put the probability of a clean run at 0.78 and made the detection statistics unusable without anyone suspecting the reset.
Twelve resets per shot.
| Reset error | Clean-run probability | Usable |
|---|---|---|
| 0.2% | 0.98 | yes |
| 2% | 0.78 | no |
I would not consider it settled without evidence: Measure the ancilla immediately after reset over many shots and report the residual excitation rate alongside the algorithm's result.
A reset is an operation with an error rate like any other.
Curated: · Written: · Reviewed:
QA-23Why does forgetting to uncompute scratch qubits break an algorithm rather than just waste them?(show answer)
The engineering content of uncomputing ancillas is the compilation and the error budget, not the notation.
Leftover scratch stays entangled with the computational register, which makes the branches of the superposition distinguishable and destroys the interference that amplitude amplification and phase estimation depend on. The cost is not wasted qubits but a wrong answer.
Concretely, structure every reversible subroutine as compute, use, uncompute, and assert in simulation that the scratch register returns to zero with probability one on a superposition input rather than on basis states.
The reason for that specificity is a failure I have seen: An oracle left 6 ancillas dirty; the Grover iteration stopped amplifying and returned a distribution within 1 percent of uniform over 8,192 shots, which the team attributed to device noise until a 3-qubit simulation showed the same flat result.
Grover on 3 qubits with and without uncompute.
| Oracle | Marked-state probability after 2 iterations |
|---|---|
| uncomputed | 0.95 |
| scratch left dirty | 0.13 |
I would not consider it settled without evidence: Run the oracle on an equal superposition in simulation and assert every ancilla measures zero with probability one.
Dirty scratch is a correctness bug, not an efficiency one.
Curated: · Written: · Reviewed:
QA-24What does teleportation actually move, and what does it require?(show answer)
Before trusting anything about quantum teleportation as an engineering protocol I would write down what a wrong answer would look like.
It moves an unknown quantum state using a pre-shared entangled pair and two classical bits, consuming the entanglement and destroying the original. It is not faster-than-light communication, because the receiver's state is useless until the two classical bits arrive.
Concretely, implement it as a resource protocol with an explicit budget: one Bell pair and two classical bits per qubit teleported, plus the correction gates the receiver applies based on those bits.
The reason for that specificity is a failure I have seen: A distributed design assumed teleportation removed the classical channel, and its latency model omitted the round trip; measured end to end, the protocol was bounded by a 40 millisecond classical hop that the design had priced at zero.
Cost of teleporting one qubit.
| Resource | Quantity | Consumed |
|---|---|---|
| Bell pair | 1 | yes |
| classical bits | 2 | yes |
| original qubit state | 1 | destroyed |
I would not consider it settled without evidence: Trace the protocol end to end including the classical transmission and the receiver's conditional corrections, and time it.
The classical bits are on the critical path.
Curated: · Written: · Reviewed:
QA-25Can you send two classical bits by transmitting one qubit?(show answer)
The first thing I would pin down about superdense coding and its prerequisites is which convention the numbers are stated in.
Only if the sender and receiver already share an entangled pair, which had to be distributed earlier. Counting the distribution, the protocol does not beat sending two bits; what it does is let the entanglement be established at a convenient time and spent later.
Concretely, account for entanglement as a stored resource with its own generation rate and decoherence lifetime, so the protocol's throughput is bounded by pair distribution rather than by the transmission itself.
The reason for that specificity is a failure I have seen: A capacity estimate doubled a link's throughput on paper by assuming superdense coding, while the entanglement source produced usable pairs at 1.2 kHz against a 10 MHz classical channel, making the scheme four orders of magnitude slower.
Where the bits are actually spent.
| Step | Qubits sent | Bits conveyed |
|---|---|---|
| pair distribution | 1 | 0 |
| encoded transmission | 1 | 2 |
I would not consider it settled without evidence: Measure the pair generation rate and the pair fidelity after storage, and bound the protocol by those rather than by the qubit transmission.
The entanglement was the message's first half.
Curated: · Written: · Reviewed:
QA-26Your two qubits agree on every measurement. Is that entanglement?(show answer)
I would answer Bell states and what their correlations prove by separating what the mathematics guarantees from what the device delivers.
Agreement in one basis is reproduced perfectly by a classical mixture of both qubits being zero and both being one, so it proves correlation and nothing more. Entanglement shows in the second basis: a Bell pair stays correlated after both qubits are rotated, and the classical mixture does not.
Concretely, measure parity in at least two mutually unbiased bases and require both correlations, since one basis can always be reproduced classically.
The reason for that specificity is a failure I have seen: A demonstration reported a Bell state on the strength of 0.98 agreement in Z alone; the same preparation gave XX parity of 0.04, which is a classical mixture and not an entangled pair.
Two preparations, one indistinguishable basis.
| Preparation | ZZ parity | XX parity |
|---|---|---|
| Bell pair | 0.98 | 0.95 |
| classical mixture | 0.98 | 0.04 |
I would not consider it settled without evidence: Report ZZ and XX parity together and require both to be high before calling a state entangled.
One basis is a correlation; two are a claim.
Curated: · Written: · Reviewed:
QA-27You measure a CHSH value of 2.05. Have you violated the local bound?(show answer)
This is an area where the ideal circuit and the circuit that executes the CHSH bound and the statistics behind it are different programs.
Not yet: the local bound is 2 and the quantum maximum is about 2.828, so the whole window is 0.83 wide and the error bar decides. With 8,192 trials per setting the standard error on S is about 0.022, which puts 2.05 a little over two standard errors above the bound.
Concretely, compute the standard error from the four correlators and the trial counts, fix the number of trials before the run, and report S with its interval rather than as a point value.
The reason for that specificity is a failure I have seen: A result reported a violation at S equal to 2.03 from 1,024 trials per setting, where the standard error was 0.063; the same apparatus reproduced 1.98 on the next run and the claim was withdrawn.
The same S at two trial counts.
| Trials per setting | Standard error on S | S = 2.05 is |
|---|---|---|
| 1,024 | 0.063 | noise |
| 8,192 | 0.022 | marginal |
| 65,536 | 0.008 | a violation |
I would not consider it settled without evidence: State S, the per-setting trial count and the standard error together, and require several standard errors of margin above 2.
A violation is a distance measured in standard errors.
Curated: · Written: · Reviewed:
QA-28Your Bell test discards trials where a detector did not fire. What does that cost the claim?(show answer)
My answer to detection efficiency and the fair-sampling assumption starts from the resource count rather than from the asymptotic claim.
Discarding trials makes the analysed subset conditional on an event that can depend on the measurement setting, so a local model with a suitable detection strategy can reproduce apparent violations. For a maximally entangled pair the efficiency has to exceed about 82.8 percent before fair sampling can be dropped.
Concretely, report the trials attempted, the trials discarded and the reason for each discard, and label a result as conditional on fair sampling whenever efficiency sits below the threshold.
The reason for that specificity is a failure I have seen: An experiment at 70 percent detection efficiency reported a loophole-free violation; recomputed on all attempted trials rather than the coincidences, the same data supported no violation at all.
Same data, two denominators.
| Analysis | Trials counted | S |
|---|---|---|
| coincidences only | 43,120 | 2.61 |
| all attempted | 61,600 | 1.94 |
I would not consider it settled without evidence: Recompute S over every attempted trial, counting non-detections as outcomes, and compare it with the post-selected value.
The discarded trials are part of the result.
Curated: · Written: · Reviewed:
QA-29You need to certify a GHZ state on five qubits. Do you run tomography?(show answer)
I would treat entanglement witnesses versus full tomography as a claim that has to survive being recomputed from the raw counts.
Tomography needs 3 to the n measurement settings — 243 at five qubits — and reconstructs far more than the claim requires. A fidelity witness certifies genuine multipartite entanglement from two settings whenever the estimated fidelity exceeds one half.
Concretely, fix the witness before looking at the data, since choosing the witness that happens to violate its bound on the observed counts is the same error as choosing a hypothesis after seeing the sample.
The reason for that specificity is a failure I have seen: A team ran 243 tomography settings at 4,096 shots each — about a million executions and nine hours of device time — to establish a claim that two settings and 8,192 shots would have supported.
Certifying a five-qubit GHZ state.
| Method | Settings | Shots | Device time |
|---|---|---|---|
| full tomography | 243 | 995,328 | ~9 h |
| fidelity witness | 2 | 8,192 | ~4 min |
I would not consider it settled without evidence: Estimate the witness from its preregistered settings and report the fidelity with its confidence interval against the one-half bound.
Measure the claim, not the whole state.
Curated: · Written: · Reviewed:
QA-30Why does tomography stop being an option so quickly?(show answer)
The useful question for the cost of state tomography is what the hardware does that the simulator never showed.
The number of parameters in a density matrix grows as 4 to the n and the number of measurement settings as 3 to the n, so the shot budget grows exponentially in qubits while the useful conclusion usually does not. At ten qubits it is 59,049 settings.
Concretely, reserve tomography for one and two qubit characterization, and use randomized measurement techniques, witnesses or direct fidelity estimation for anything wider, stating the estimator and its assumptions.
The reason for that specificity is a failure I have seen: A characterization plan budgeted tomography for an eight-qubit register, discovered it needed 6,561 settings at 2,048 shots — over 13 million executions — and abandoned characterization entirely rather than switching estimator.
Tomography settings by width.
| Qubits | Settings | Real parameters |
|---|---|---|
| 2 | 9 | 16 |
| 5 | 243 | 1,024 |
| 10 | 59,049 | 1,048,576 |
I would not consider it settled without evidence: Compute the settings and shots the chosen estimator requires at the target width before committing device time.
Pick the estimator that scales with the question.
Curated: · Written: · Reviewed:
QA-31A vendor quotes 99.9 percent fidelity. What did they measure?(show answer)
I would settle state fidelity versus process fidelity against a shot budget before arguing about the algorithm.
The number is meaningless without the object and the protocol: state fidelity compares one prepared state with a target, process fidelity characterizes a gate over all inputs, and average gate fidelity from randomized benchmarking reports a Clifford-averaged quantity. They are different measurements and do not convert into each other casually.
Concretely, ask which protocol produced the figure, over which qubits and in which calibration window, and re-measure the quantity your workload depends on rather than adopting the headline.
The reason for that specificity is a failure I have seen: A design assumed a quoted 99.9 percent single-qubit figure applied to two-qubit operations; the two-qubit error was 0.8 percent, and a circuit with 200 two-qubit gates succeeded 20 percent of the time against an expected 82 percent.
Three numbers that all read as fidelity.
| Quantity | Typical value | Applies to |
|---|---|---|
| single-qubit average gate | 99.9% | one qubit |
| two-qubit average gate | 99.2% | one coupled pair |
| state preparation and measurement | 98.5% | readout chain |
I would not consider it settled without evidence: Reproduce the vendor's protocol on the qubits your circuit will use, and record the calibration timestamp with the number.
A fidelity without its protocol is a rumour.
Curated: · Written: · Reviewed:
QA-32Randomized benchmarking gives an excellent number and your algorithm still fails. Why?(show answer)
The judgement in the blind spot in randomized benchmarking is which assumption is load-bearing and whether it was checked.
Randomized benchmarking averages over random Clifford sequences, which converts coherent errors into effective depolarizing noise. A structured circuit repeats the same gates in the same order, so a small coherent over-rotation accumulates linearly rather than averaging away.
Concretely, complement the benchmark with a workload-shaped test — the actual circuit at several depths with a known answer — and look for error growing faster than the benchmark predicts, which is the coherent-error signature.
The reason for that specificity is a failure I have seen: A device reported 99.8 percent average gate fidelity while a circuit repeating one rotation 50 times drifted by 0.35 radians, because a 0.007 radian over-rotation per gate accumulated coherently and randomized benchmarking never saw it.
Coherent over-rotation across repetitions.
| Repetitions | Predicted error | Observed error |
|---|---|---|
| 10 | 0.002 | 0.005 |
| 50 | 0.010 | 0.121 |
I would not consider it settled without evidence: Run the repeated-gate sequence at increasing depth and compare the observed error growth against the linear prediction from the benchmark.
Averaging over random sequences hides errors that structure amplifies.
Curated: · Written: · Reviewed:
QA-33When is it legitimate to apply a readout correction matrix to your counts?(show answer)
Where candidates lose the interview on readout error mitigation and its limits is treating an exact-state result as a device result.
It is legitimate for characterizing a prepared state, where the readout channel is a known nuisance to be inverted. It is not legitimate inside a claim that tests the measurement apparatus itself, such as a loophole-free Bell test, because the correction assumes the model under examination.
Concretely, report raw and corrected counts side by side with the calibration matrix and its age, so a reader can see how much of the result came from the correction.
The reason for that specificity is a failure I have seen: A reported state fidelity of 0.96 was 0.88 before readout correction, and the calibration matrix used was 14 hours old; on recalibration the corrected figure moved to 0.91.
How much of the result is correction.
| Quantity | Value |
|---|---|
| raw fidelity | 0.88 |
| corrected, fresh matrix | 0.91 |
| corrected, 14-hour-old matrix | 0.96 |
I would not consider it settled without evidence: Publish raw counts, the correction matrix, its timestamp, and the corrected result rather than the corrected result alone.
Show the number before the correction.
Curated: · Written: · Reviewed:
QA-34A stakeholder asks whether entanglement gives you an instantaneous channel. What do you say?(show answer)
I would answer why entanglement does not permit signalling by naming the measurement that would contradict it.
No: the local reduced state of one party is unchanged by anything the other party does, so no measurement statistics available locally depend on the remote choice. Correlations only appear when the two records are compared over a classical channel.
Concretely, demonstrate it rather than assert it — show that the local outcome distribution is 50/50 regardless of the remote basis, and that structure appears only after the records are joined.
The reason for that specificity is a failure I have seen: A proposal budgeted a zero-latency control channel between two sites on the strength of shared entanglement; the design collapsed when the classical comparison step was added back, and three weeks of architecture work was discarded.
Local statistics under two remote choices.
| Remote basis | Local P(0) | Local P(1) |
|---|---|---|
| Z | 0.501 | 0.499 |
| X | 0.498 | 0.502 |
I would not consider it settled without evidence: Compare the local outcome distribution across remote measurement settings and require it to be statistically identical.
Correlation appears only in the joined record.
Curated: · Written: · Reviewed:
QA-35Which multi-qubit entangled state would you choose if one qubit may be lost?(show answer)
The engineering content of GHZ and W states under qubit loss is the compilation and the error budget, not the notation.
A GHZ state is maximally fragile to loss: tracing out one qubit leaves the rest in a classical mixture with no entanglement. A W state is robust in that specific sense, retaining bipartite entanglement among the survivors, at the cost of weaker correlations to begin with.
Concretely, choose from the failure mode the application actually faces — parity-based metrology favours GHZ, loss-tolerant distribution favours W — and state the assumption about loss explicitly.
The reason for that specificity is a failure I have seen: A three-party protocol built on GHZ states lost one party's qubit in 12 percent of rounds; the remaining pair had no entanglement to work with, and the protocol's success rate fell to 0.88 of its design value with no graceful degradation.
After losing one of three qubits.
| State | Residual entanglement | Usable |
|---|---|---|
| GHZ | none | no |
| W | present | yes |
I would not consider it settled without evidence: Trace out one qubit in simulation and measure the residual entanglement of the survivors for each candidate state.
Choose the state for the loss you expect.
Curated: · Written: · Reviewed:
QA-36What happens if you run Deutsch-Jozsa on a function that is neither constant nor balanced?(show answer)
Before trusting anything about the promise in Deutsch-Jozsa I would write down what a wrong answer would look like.
The algorithm returns an answer anyway, and the answer is meaningless. Its all-zeros outcome probability is the squared magnitude of the average of minus one to the f of x, so a function that is zero on three quarters of its inputs returns all zeros about 25 percent of the time and is reported as constant.
Concretely, treat the promise as an input precondition the caller must establish, and either verify it separately or use an algorithm that does not require it.
The reason for that specificity is a failure I have seen: A demonstration wrapped Deutsch-Jozsa around an unverified predicate; on inputs outside the promise it reported constant 1 run in 4 and the results were used to justify a design decision that was wrong.
Outcome probability against the function.
| Function | P(all zeros) | Reported |
|---|---|---|
| constant | 1.00 | constant |
| balanced | 0.00 | balanced |
| 75% zeros | 0.25 | constant 1 run in 4 |
I would not consider it settled without evidence: Feed the implementation a function outside the promise and confirm the output is unusable rather than merely noisy.
A promise is an assumption, not a check.
Curated: · Written: · Reviewed:
QA-37How many Grover iterations do you run over a million candidates?(show answer)
The first thing I would pin down about Grover's optimal iteration count is which convention the numbers are stated in.
About pi over 4 times the square root of N over M, which for N equal to 2 to the 20 with one marked item is roughly 804 iterations. The amplitude rotates in a two-dimensional subspace, so the schedule has a maximum rather than converging.
Concretely, compute the schedule from N and M before the run and stop at the computed optimum, since success falls again past it rather than plateauing.
The reason for that specificity is a failure I have seen: An implementation ran 1,608 iterations on a 2 to the 20 space "to be safe" and never once measured the marked item over 8,192 shots, against near certainty at the 804-iteration optimum.
Success against iterations at N = 2^20, M = 1.
| Iterations | Success probability |
|---|---|
| 400 | 0.50 |
| 804 | 1.00 |
| 1,608 | under 0.0001 |
I would not consider it settled without evidence: Plot success probability against iteration count on a small instance and confirm the peak lands where the formula predicts.
More iterations is not more search.
Curated: · Written: · Reviewed:
QA-38You do not know how many items are marked. How do you run Grover?(show answer)
I would answer unknown solution counts in amplitude amplification by separating what the mathematics guarantees from what the device delivers.
The fixed schedule is unusable because it depends on M, and guessing wrong rotates past the optimum. The standard answers are quantum counting to estimate M first, or an exponentially increasing randomized schedule that finds a solution in expected order square root of N over M queries without knowing M.
Concretely, pick the randomized schedule when a single solution suffices and counting when the number itself is the answer, and specify the zero-solution behaviour separately.
The reason for that specificity is a failure I have seen: A search assumed one marked item on an instance that had four; 804 iterations is exactly twice the 402 that four solutions require, so the amplitude landed back at a trough and the marked items were returned about once in a million shots, which was read as hardware noise.
Fixed schedule for M = 1 on other instances.
| Actual M | Optimal iterations | Run at 804 |
|---|---|---|
| 1 | 804 | 1.00 |
| 4 | 402 | 0.000001 |
| 0 | none | undefined output |
I would not consider it settled without evidence: Run the chosen strategy against instances with 1, 4 and 0 solutions and confirm the reported success rate in each case.
Unknown M is a different algorithm, not a tuning problem.
Curated: · Written: · Reviewed:
QA-39Where does the cost of a Grover search actually go?(show answer)
This is an area where the ideal circuit and the circuit that executes the cost of constructing an oracle are different programs.
Into the oracle and the diffuser, which the query-complexity statement counts as one call each. A 20-qubit diffuser contains a multi-controlled phase that decomposes into roughly 36 Toffolis, about 216 CNOTs, and the oracle usually costs at least as much again once its predicate is computed and uncomputed.
Concretely, compile the oracle for the real predicate and report the total two-qubit gate count of the full schedule, not the iteration count.
The reason for that specificity is a failure I have seen: A proposal quoted 804 queries as the cost of a search; compiled, the same schedule was about 350,000 two-qubit gates, which at 0.5 percent error has an essentially zero chance of running without a fault.
From queries to gates at N = 2^20.
| Quantity | Value |
|---|---|
| iterations | 804 |
| two-qubit gates per iteration | ~432 |
| total two-qubit gates | ~350,000 |
I would not consider it settled without evidence: Transpile one full iteration and multiply, then convert the total into a success probability at the device's error rate.
A query is a circuit, and the circuit is the cost.
Curated: · Written: · Reviewed:
QA-40How would you describe Grover to someone who is not searching a list?(show answer)
My answer to amplitude amplification as the general form starts from the resource count rather than from the asymptotic claim.
As amplitude amplification: given any state preparation with success probability p and a way to recognise success, roughly one over the square root of p repetitions of the reflect-and-reprepare pair raise the success probability near one. Search is the special case where preparation is a uniform superposition.
Concretely, report the technique with its preparation cost included, since the amplification multiplies the cost of the preparation circuit and its inverse by the repetition count.
The reason for that specificity is a failure I have seen: A sampling routine with p equal to 0.01 was described as gaining a hundredfold speedup; including the preparation and its inverse, the ten repetitions each cost twice the original preparation, so the real gain was a factor of five.
Amplifying a p = 0.01 preparation.
| Approach | Preparations run | Success |
|---|---|---|
| classical repetition | 100 | 0.63 |
| amplitude amplification | ~20 | 0.99 |
I would not consider it settled without evidence: Cost one amplification round as preparation plus inverse plus reflections, and compare the total against repeating the classical preparation.
Amplification multiplies whatever preparation costs.
Curated: · Written: · Reviewed:
QA-41Grover returns a candidate. What do you do with it?(show answer)
I would treat verifying a quantum search result classically as a claim that has to survive being recomputed from the raw counts.
Check it against the predicate classically before using it. The verification costs one evaluation, and it removes the entire class of failures where the oracle marked nothing, marked the wrong set, or the schedule overshot.
Concretely, make verification part of the algorithm rather than a test, so an unverified candidate is never returned to a caller and a failed verification triggers a retry with a recorded reason.
The reason for that specificity is a failure I have seen: An oracle with an inverted comparison marked the complement of the intended set; every returned candidate was wrong, and because no verification step existed the error surfaced three weeks later in downstream results.
What verification catches.
| Failure | Cost to detect | Detected by verification |
|---|---|---|
| oracle marks wrong set | 1 evaluation | yes |
| schedule overshoots | 1 evaluation | yes |
| zero solutions exist | 1 evaluation | yes |
I would not consider it settled without evidence: Assert that every returned candidate satisfies the predicate, and count verification failures as a monitored rate.
One classical check removes a class of silent failures.
Curated: · Written: · Reviewed:
QA-42The QFT is often called exponentially faster than the FFT. Is it?(show answer)
The useful question for the quantum Fourier transform's real cost is what the hardware does that the simulator never showed.
It applies the transform to amplitudes in order n squared gates against the FFT's n times 2 to the n operations, but the output is a quantum state whose amplitudes cannot be read out. The advantage exists only when a subsequent step extracts one classical answer, as phase estimation does.
Concretely, present the QFT as a subroutine inside an algorithm that produces a classical result, and cost the rotations, which are commonly truncated below a threshold to keep the gate count manageable.
The reason for that specificity is a failure I have seen: A proposal replaced a signal-processing FFT with a QFT and only then discovered that reading the transformed vector required tomography, whose cost exceeded the classical computation by orders of magnitude.
What each transform gives you.
| Transform | Cost | Output available |
|---|---|---|
| FFT | n 2^n | all amplitudes |
| QFT | ~n^2 gates | one sampled outcome |
I would not consider it settled without evidence: Name the classical quantity the algorithm extracts and show it does not require reading the full amplitude vector.
An unreadable answer is not an answer.
Curated: · Written: · Reviewed:
QA-43How many ancilla qubits does phase estimation need?(show answer)
I would settle quantum phase estimation and its precision qubits against a shot budget before arguing about the algorithm.
Roughly the number of bits of precision required, plus a few more to bound the failure probability: t bits of phase to accuracy 2 to the minus n with success probability 1 minus epsilon needs about n plus the logarithm of one over epsilon ancillas. Each additional bit doubles the number of controlled applications of the unitary.
Concretely, derive the ancilla count from the precision the downstream decision needs, and note that the controlled-U applications, not the ancillas, dominate the circuit.
The reason for that specificity is a failure I have seen: A design specified 16 bits of precision "for margin" where 8 sufficed, which multiplied the controlled-unitary applications by 256 and turned a feasible circuit into one 60 times longer than the coherence budget allowed.
Precision against circuit cost.
| Bits | Ancillas | Controlled-U applications |
|---|---|---|
| 8 | ~11 | 255 |
| 16 | ~19 | 65,535 |
I would not consider it settled without evidence: Compute the controlled-unitary count implied by the requested precision and compare it against the device's duration budget.
Precision is bought in doublings.
Curated: · Written: · Reviewed:
QA-44Which part of Shor's algorithm actually runs on the quantum computer?(show answer)
The judgement in order finding as the quantum core of factoring is which assumption is load-bearing and whether it was checked.
Only order finding: estimating the smallest r with a to the r congruent to one modulo N, via phase estimation on reversible modular multiplication. Choosing the base, computing the greatest common divisor, recovering r from the measured phase and verifying the factors are all classical.
Concretely, structure the implementation as a classical driver calling a quantum subroutine, with the driver owning retries, since failed bases and unusable samples are expected rather than exceptional.
The reason for that specificity is a failure I have seen: An implementation treated a single quantum run as the algorithm and reported failure when the first base gave an odd order; the classical driver should have retried, and roughly two bases in three succeed.
Where the work sits.
| Step | Runs on | Bases succeeding |
|---|---|---|
| choose base a | classical | 2 in 3 |
| order finding | quantum | retried |
| recover factors | classical | 1 attempt |
I would not consider it settled without evidence: Run the full classical wrapper over many bases and report the number of quantum runs needed per successful factorization.
The quantum part is a subroutine inside a classical loop.
Curated: · Written: · Reviewed:
QA-45Phase estimation gives you a noisy phase. How do you get an integer order from it?(show answer)
Where candidates lose the interview on continued fractions and discarded samples is treating an exact-state result as a device result.
By continued-fraction expansion of the measured phase, which yields a candidate denominator. Some samples give denominators that are not the order, so candidates must be tested by checking a to the r modulo N, and failures are part of the expected cost rather than a bug.
Concretely, record the number of samples drawn, the number that produced valid orders and the number of bases tried, so the algorithm's real cost is measured rather than assumed.
The reason for that specificity is a failure I have seen: A demonstration reported factoring 15 in one shot; the same code needed 6 samples on average across bases, and the single-shot claim came from a run that had been repeated until it worked.
Samples to a valid order.
| Run | Samples drawn | Valid order found |
|---|---|---|
| 1 | 3 | yes |
| 2 | 9 | yes |
| 3 | 5 | yes |
I would not consider it settled without evidence: Report the distribution of samples required over many independent runs rather than the best run.
Report the average run, not the lucky one.
Curated: · Written: · Reviewed:
QA-46A paper factors 21 on a small device. What does that tell you about RSA?(show answer)
I would answer compiled demonstrations of factoring by naming the measurement that would contradict it.
Usually nothing, because most small demonstrations compile the modular exponentiation using knowledge of the answer, which removes exactly the arithmetic that dominates cost at 2048 bits. The demonstration shows control and readout, not scalable factoring.
Concretely, ask whether the modular arithmetic was implemented generically or specialized to the known factors, and whether the qubit count scales with the number's bit length.
The reason for that specificity is a failure I have seen: A briefing cited a 4-bit factoring demonstration as evidence that RSA-2048 was years from falling; the circuit in question used 5 qubits and precomputed the period, and it would not have factored a number the team had not already factored.
Two ways to factor 21.
| Implementation | Qubits | Works on unknown input |
|---|---|---|
| compiled with known period | 5 | no |
| generic modular arithmetic | thousands | yes |
I would not consider it settled without evidence: Check whether the circuit's size grows with the modulus and whether it would run on an input whose factors were unknown.
A demonstration that needs the answer has not computed it.
Curated: · Written: · Reviewed:
QA-47What would it actually take to factor RSA-2048?(show answer)
The engineering content of resource estimates for breaking RSA-2048 is the compilation and the error budget, not the notation.
Published surface-code estimates put it in the millions of physical qubits running for hours to days, and they have moved substantially as arithmetic and factory constructions improved: the widely cited 2019 analysis gave about 20 million physical qubits for roughly 8 hours at a 0.1 percent physical error rate. Every such figure is conditional on its assumptions.
Concretely, quote a resource estimate with its assumptions attached — error rate, cycle time, code distance, factory design, arithmetic circuit — and treat an estimate stripped of them as unusable.
The reason for that specificity is a failure I have seen: A risk assessment quoted "20 million qubits" as a fixed threshold; the underlying assumption of a 0.1 percent physical error rate was never stated, and a later estimate under different assumptions moved the requirement by an order of magnitude.
What an estimate has to pin down.
| Assumption | Example value | Sensitivity |
|---|---|---|
| physical error rate | 1e-3 | high |
| cycle time | 1 us | high |
| code distance | 27 | high |
I would not consider it settled without evidence: List which of the logical qubit count, non-Clifford count, code distance and success probability were computed and which were assumed.
An estimate without assumptions is a number without meaning.
Curated: · Written: · Reviewed:
QA-48Does a quantum computer halve every key length?(show answer)
Before trusting anything about Grover against symmetric keys I would write down what a wrong answer would look like.
Grover reduces the query exponent for a symmetric key search from n to n over 2, so AES-128 falls to about 2 to the 64 oracle queries. Those queries are inherently sequential, and parallelizing across k machines improves the count only by the square root of k, so the practical margin is larger than the exponent suggests.
Concretely, treat symmetric primitives as a parameter review — move to 256-bit keys for margin — rather than as a replacement programme, and keep them separate from the public-key work in the plan.
The reason for that specificity is a failure I have seen: A migration plan spent two quarters replacing symmetric primitives while the public-key key-establishment path, which is the actually urgent item, remained on classical elliptic-curve Diffie-Hellman.
Exposure by primitive.
| Primitive | Quantum effect | Action |
|---|---|---|
| RSA / ECDH | broken by Shor | replace |
| AES-128 | 2^64 sequential queries | move to 256-bit |
| SHA-256 | modest margin loss | review |
I would not consider it settled without evidence: Separate the plan into key establishment, long-lived signatures and symmetric parameters, and show the sequencing rationale for each.
Not every primitive is equally exposed.
Curated: · Written: · Reviewed:
QA-49Why would you migrate key establishment before any quantum computer exists?(show answer)
The first thing I would pin down about harvest-now-decrypt-later exposure is which convention the numbers are stated in.
Because traffic captured today can be decrypted later, so the exposure of long-lived confidential data begins at capture rather than at the arrival of the machine. Authentication that is verified once and never again does not carry the same risk.
Concretely, rank systems by the confidentiality lifetime of what they protect, and migrate key establishment for the long-lived material first regardless of the estimated date of a capable machine.
The reason for that specificity is a failure I have seen: An estate migrated its public-facing signature algorithms first, leaving a research VPN carrying data with a 25-year secrecy requirement on classical key exchange for another two years.
Urgency by data lifetime.
| Data | Secrecy lifetime | Priority |
|---|---|---|
| research records | 25 years | first |
| session authentication | minutes | last |
I would not consider it settled without evidence: Produce a cryptographic inventory listing where keys live and how long each protected item must stay secret, and sequence from that.
The capture is the event, not the decryption.
Curated: · Written: · Reviewed:
QA-50How do you decide whether your migration is late without predicting a date?(show answer)
I would answer timing a post-quantum migration by separating what the mathematics guarantees from what the device delivers.
Compare the sum of your data's confidentiality lifetime and your migration duration against the time until a capable machine exists. If that sum exceeds the horizon, exposure has already begun, and the conclusion holds across every horizon in the plausible range.
Concretely, estimate migration duration from the slowest component — embedded devices, certificate re-issuance, vendor dependencies — rather than from the protocol change, and re-run the comparison annually.
The reason for that specificity is a failure I have seen: A plan assumed an 18-month migration based on server-side changes; the firmware signing chain on deployed hardware took 5 years, and the real duration made the estate exposed under every horizon the risk team had considered acceptable.
The subtraction that decides urgency.
| Term | Years |
|---|---|
| data confidentiality lifetime | 25 |
| migration duration | 7 |
| horizon assumed | 15 |
I would not consider it settled without evidence: Show the three numbers — data lifetime, migration duration, assumed horizon — and the subtraction, rather than asserting readiness.
The arithmetic survives disagreement about the date.
Curated: · Written: · Reviewed:
QA-51Which post-quantum algorithms do you deploy, and how do you choose?(show answer)
This is an area where the ideal circuit and the circuit that executes selecting post-quantum algorithms are different programs.
You deploy the finalized standards rather than choosing on cryptographic taste: ML-KEM for key encapsulation and ML-DSA or SLH-DSA for signatures, with parameter sets chosen against the sector guidance that applies. Inventing or tuning a scheme is not an option available to an implementer.
Concretely, plan for the size changes rather than only the algorithm change, since post-quantum keys and signatures are substantially larger and break protocols, hardware buffers and certificate chains built around classical sizes.
The reason for that specificity is a failure I have seen: A rollout replaced the key exchange and found that an embedded TLS stack rejected the larger handshake at a fixed 4 kilobyte buffer; the device fleet could not be updated remotely and 40,000 units needed a field visit.
What changes besides the algorithm.
| Property | Classical ECDH | ML-KEM |
|---|---|---|
| public key size | 32 bytes | over 1 KB |
| handshake buffer impact | none | often fatal |
I would not consider it settled without evidence: Test the chosen parameter set end to end against the real protocol stack and the smallest device in the fleet before committing to a schedule.
Size is the migration's hidden constraint.
Curated: · Written: · Reviewed:
QA-52Why combine a classical and a post-quantum key exchange rather than switching outright?(show answer)
My answer to hybrid key establishment and rollback starts from the resource count rather than from the asymptotic claim.
A hybrid construction derives the session secret from both, so the session stays secure if either component holds. That covers implementation flaws in the newer scheme without giving up quantum resistance.
Concretely, keep the combination in the key derivation rather than the transport, negotiate the algorithm so a rollback is possible, and record which combination each session used so an incident can identify affected traffic.
The reason for that specificity is a failure I have seen: A deployment hard-coded a single post-quantum scheme with no negotiation; when an implementation flaw was announced, the estate had no path back and ran degraded for six weeks.
Failure tolerance of each option.
| Construction | Components that must hold | Classical broken | PQ scheme flawed |
|---|---|---|---|
| classical only | 1 | insecure | secure |
| PQ only | 1 | secure | insecure |
| hybrid | 1 of 2 | secure | secure |
I would not consider it settled without evidence: Exercise the negotiation and the rollback path in a test environment and confirm sessions record the algorithm combination they used.
Agility is the property that survives the next surprise.
Curated: · Written: · Reviewed:
QA-53A device advertises T1 of 120 microseconds. What does that let you build?(show answer)
I would treat reading T1 and T2 as an operations budget as a claim that has to survive being recomputed from the raw counts.
It bounds how long a circuit can run, and the useful figure is operations per coherence time rather than the time itself. With two-qubit gates around 300 nanoseconds, 120 microseconds allows a few hundred sequential two-qubit operations before relaxation dominates, and T2 is usually the tighter constraint.
Concretely, convert coherence into an operation budget for the specific qubits your circuit will use, since T1 and T2 vary several-fold across a chip and across calibration windows.
The reason for that specificity is a failure I have seen: A workload was designed against the chip's best-qubit T2 of 140 microseconds; the qubits the router actually selected had T2 near 45, and the observed success rate was a third of the projection.
Coherence across one device.
| Qubit | T1 | T2 | Two-qubit ops available |
|---|---|---|---|
| best | 140 us | 140 us | ~460 |
| median | 90 us | 70 us | ~230 |
| worst | 40 us | 25 us | ~80 |
I would not consider it settled without evidence: Read T1 and T2 for the physical qubits in the chosen layout from the same calibration record as the run.
Coherence is per qubit, not per chip.
Curated: · Written: · Reviewed:
QA-54Two devices report the same error rate and one runs your circuit far worse. What could differ?(show answer)
The useful question for coherent versus stochastic errors is what the hardware does that the simulator never showed.
The error's character. Stochastic errors accumulate as the sum of probabilities, while coherent errors such as a systematic over-rotation add in amplitude and can accumulate quadratically in the worst case, so the same average rate produces very different circuit-level behaviour.
Concretely, distinguish them experimentally by running the same gate repeatedly and watching how error grows with depth, then correct coherent components through calibration rather than through more shots.
The reason for that specificity is a failure I have seen: Two backends both reported 99.8 percent average gate fidelity; on a 50-layer repeated sequence one drifted 0.35 radians from a systematic over-rotation while the other stayed within noise, and the difference was invisible in the published metric.
Error growth by character.
| Depth | Stochastic | Coherent |
|---|---|---|
| 10 | 0.020 | 0.005 |
| 50 | 0.095 | 0.121 |
I would not consider it settled without evidence: Plot error against repetition count and compare the growth with the linear prediction from the average rate.
Two errors with one number behave differently.
Curated: · Written: · Reviewed:
QA-55Your gates benchmark well individually and the circuit still fails. What do you check next?(show answer)
I would settle crosstalk under simultaneous operations against a shot budget before arguing about the algorithm.
Simultaneous operation. Isolated benchmarks measure a gate with its neighbours idle, and running gates in parallel on coupled qubits introduces crosstalk that no isolated number predicts.
Concretely, benchmark the parallel pattern the circuit actually uses — simultaneous randomized benchmarking or a layer-fidelity measurement — and let the scheduler avoid the worst parallel pairs when the cost is acceptable.
The reason for that specificity is a failure I have seen: A layer running four two-qubit gates in parallel measured an effective error of 2.1 percent against 0.6 percent for the same gates run in isolation, and a circuit budgeted on the isolated figure failed three times more often than projected.
Isolated against simultaneous.
| Pattern | Two-qubit error |
|---|---|
| one pair at a time | 0.6% |
| four pairs in parallel | 2.1% |
I would not consider it settled without evidence: Measure the error of the exact parallel layer pattern rather than summing isolated gate errors.
Gates in parallel are a different experiment.
Curated: · Written: · Reviewed:
QA-56Why is leakage worse for error correction than an ordinary bit flip?(show answer)
The judgement in leakage out of the computational subspace is which assumption is load-bearing and whether it was checked.
A leaked qubit is outside the computational subspace, so it is not a Pauli error and a decoder built on a Pauli noise model mis-corrects it. Leakage also persists across rounds and spreads through two-qubit gates, producing exactly the time-correlated faults the threshold analysis assumes away.
Concretely, detect leakage explicitly and insert leakage-reduction operations that return population to the computational subspace, then feed the detection into the decoder rather than hiding it.
The reason for that specificity is a failure I have seen: A distance-5 experiment saw its logical error rate stall at 1.1 percent per cycle regardless of added rounds; leakage on two qubits was persisting for tens of rounds and the Pauli decoder was compounding it.
Logical error with and without leakage handling.
| Configuration | Logical error per cycle |
|---|---|
| no leakage reduction | 1.1% |
| leakage reduction units | 0.32% |
I would not consider it settled without evidence: Measure the leakage population per qubit per round and show it is reset rather than accumulating over the run.
A decoder cannot correct an error its model excludes.
Curated: · Written: · Reviewed:
QA-57You transpiled yesterday and submit today. What can go wrong?(show answer)
Where candidates lose the interview on calibration drift and stale transpilation is treating an exact-state result as a device result.
The error map the transpiler optimized against is a snapshot, so a circuit routed onto a pair whose error has since doubled executes a layout chosen for a machine that no longer exists. Two shots either side of a recalibration ran on materially different devices.
Concretely, re-transpile against the current target rather than reusing a cached physical circuit, and record the calibration timestamp and backend version with every result.
The reason for that specificity is a failure I have seen: A cached physical circuit was reused for three weeks; the coupling it relied on degraded from 0.6 to 1.9 percent error, and the reported success rate declined steadily while the algorithm was blamed.
One coupling over three weeks.
| Date | Pair error | Circuit success |
|---|---|---|
| 2026-08-18 | 0.6% | 0.81 |
| 2026-09-08 | 1.9% | 0.44 |
I would not consider it settled without evidence: Compare the calibration timestamp of the transpilation with that of the execution and re-transpile when they differ.
A layout is only valid for the calibration it was chosen under.
Curated: · Written: · Reviewed:
QA-58Your qubits idle for long stretches. What can you do without changing the algorithm?(show answer)
I would answer dynamical decoupling as error suppression by naming the measurement that would contradict it.
Insert dynamical decoupling sequences into the idle windows, which refocus low-frequency dephasing and reduce the error accumulated while waiting. It is suppression, not correction: it lowers the exposure without creating any protected logical state.
Concretely, apply it where the scheduler shows real idle time, and measure the result, since a poorly chosen sequence adds pulses whose own error exceeds the dephasing they suppress.
The reason for that specificity is a failure I have seen: Decoupling was applied uniformly to a circuit with little idle time; the extra pulses added 0.4 percent error per qubit and the observed fidelity fell from 0.79 to 0.71.
Where decoupling helps.
| Circuit | Idle fraction | Fidelity change |
|---|---|---|
| long idles | 0.62 | 0.71 to 0.84 |
| dense schedule | 0.05 | 0.79 to 0.71 |
I would not consider it settled without evidence: Measure the workload's fidelity with and without the sequence on the same qubits in the same calibration window.
Suppression is measured, not assumed.
Curated: · Written: · Reviewed:
QA-59Zero-noise extrapolation improves your expectation value. What did it cost?(show answer)
The engineering content of the variance cost of zero-noise extrapolation is the compilation and the error budget, not the notation.
Shots and bias risk. Running the circuit at several amplified noise levels and extrapolating reduces bias while inflating variance, often by one to two orders of magnitude in shot count, and the extrapolation model itself is an assumption.
Concretely, report the mitigated value with its confidence interval and the total shots consumed, and state the extrapolation model, since a linear and an exponential fit to the same three points can differ by more than the effect being measured.
The reason for that specificity is a failure I have seen: A mitigated energy beat the classical reference by 0.004 hartree with a confidence interval of 0.011 after consuming 40 times the shots of the unmitigated run, and was reported as an improvement.
What mitigation bought.
| Run | Shots | Estimate | Interval |
|---|---|---|---|
| unmitigated | 100,000 | -1.094 | 0.002 |
| extrapolated | 4,000,000 | -1.137 | 0.011 |
I would not consider it settled without evidence: Compare the mitigated estimate against its own error bar and against the unmitigated result at equal total shot cost.
A result inside its own error bar is not a result.
Curated: · Written: · Reviewed:
QA-60A stakeholder asks whether mitigation gets you to fault tolerance. What do you say?(show answer)
Before trusting anything about error mitigation versus error correction I would write down what a wrong answer would look like.
No. Mitigation estimates a less-biased observable using extra samples and model assumptions; it never produces a protected logical state and its sampling overhead grows exponentially with circuit volume. Correction encodes information redundantly and can carry a state through an arbitrarily long computation.
Concretely, keep suppression, mitigation and correction in separate columns of any report, with the overhead each consumed, so nobody reads a mitigated observable as a logical one.
The reason for that specificity is a failure I have seen: A roadmap presented mitigation as an incremental path to fault tolerance; at the circuit volumes the applications needed, the sampling overhead exceeded 10 to the 6 and the plan had no reachable milestone.
Three techniques, three claims.
| Technique | Produces a logical qubit | Overhead |
|---|---|---|
| suppression | no | small |
| mitigation | no | exponential in volume |
| correction | yes | large but bounded |
I would not consider it settled without evidence: Compute the mitigation overhead at the target circuit volume and show whether the shot budget remains finite.
Mitigation buys accuracy on small circuits, not scale.
Curated: · Written: · Reviewed:
QA-61How can a code detect an error without measuring the logical state?(show answer)
The first thing I would pin down about stabilizer codes and syndrome extraction is which convention the numbers are stated in.
Stabilizer measurements ask about parities of groups of qubits, which commute with the logical operators, so the outcome reveals an error syndrome while leaving the logical amplitudes untouched. The decoder infers a correction from the syndrome, often as a frame update rather than a physical gate.
Concretely, extract syndromes repeatedly, because the extraction circuit is itself noisy and a single round cannot distinguish a data error from a measurement error.
The reason for that specificity is a failure I have seen: An implementation measured syndromes once per logical operation; measurement errors were indistinguishable from data errors, and the decoder applied corrections that introduced logical faults at 3 percent per round.
Why one round is not enough.
| Rounds | Distinguishes measurement error | Logical fault rate |
|---|---|---|
| 1 | no | 3% |
| d rounds | yes | 0.3% |
I would not consider it settled without evidence: Show that the logical state survives repeated syndrome extraction with no measurement of the logical operators themselves.
The syndrome is a question the logical state does not answer.
Curated: · Written: · Reviewed:
QA-62How many physical qubits does one logical qubit cost?(show answer)
I would answer code distance and physical overhead by separating what the mathematics guarantees from what the device delivers.
For a rotated surface code of distance d it is 2d squared minus 1: 49 at distance 5, 241 at distance 11, and 1,249 at distance 25. Distance buys exponential suppression below threshold, so the target logical error rate sets the distance and the distance sets the bill.
Concretely, work backwards from the algorithm's tolerable logical error rate to a distance, then to physical qubits, and add the magic-state factories separately, since non-Clifford gates usually dominate the budget.
The reason for that specificity is a failure I have seen: A resource plan assumed 100 physical qubits per logical qubit; the target logical error rate of 10 to the minus 10 required distance 25 and 1,249 physical qubits each, a twelvefold error in the machine size.
Rotated surface code overhead.
| Distance | Physical qubits | Suppression |
|---|---|---|
| 5 | 49 | modest |
| 11 | 241 | good |
| 25 | 1,249 | algorithmic |
I would not consider it settled without evidence: State the target logical error rate, the assumed physical error rate, the resulting distance and the qubit count, so each step can be checked.
The distance is chosen by the error target, not by preference.
Curated: · Written: · Reviewed:
QA-63Your device's error rate is below a quoted threshold. Are you fault-tolerant?(show answer)
This is an area where the ideal circuit and the circuit that executes the assumptions behind the threshold theorem are different programs.
Not by that fact alone. The threshold theorem applies to a complete scheme under a stated noise model — including syndrome extraction, decoding, state preparation, measurement and factories — and assumes largely independent faults. A device number below a threshold quoted for a different scheme and model says nothing.
Concretely, match the quoted threshold to the code, decoder and noise model being used, and check whether the device's correlated and leakage errors violate the model's independence assumptions.
The reason for that specificity is a failure I have seen: A team declared themselves below threshold on a 0.1 percent two-qubit error figure; the quoted threshold assumed independent depolarizing noise, while the device had correlated errors across a shared readout line, and the measured logical error rose with distance.
What distance scaling reveals.
| Distance | Logical error, independent noise | With correlated noise |
|---|---|---|
| 3 | 0.9% | 0.9% |
| 5 | 0.35% | 1.2% |
I would not consider it settled without evidence: Increase code distance and require the measured logical error rate to fall, which is the only direct test of being below threshold.
Below threshold is measured by distance scaling.
Curated: · Written: · Reviewed:
QA-64Your decoder is accurate but slow. Why is that fatal rather than merely inconvenient?(show answer)
My answer to decoder latency and the syndrome backlog starts from the resource count rather than from the asymptotic claim.
Syndrome rounds are produced continuously, so a decoder slower than the round period falls further behind every round and the backlog grows without bound. The Pauli frame needed to interpret a logical measurement is then never ready.
Concretely, report decoder throughput alongside decoder accuracy, and budget real-time decoding as an engineering deliverable rather than as post-processing.
The reason for that specificity is a failure I have seen: A decoder averaging 2 microseconds per round on a device producing rounds every microsecond accumulated a 1 millisecond backlog within a thousand rounds, and the experiment could not perform a conditional logical operation at all.
Backlog growth over 1,000 rounds.
| Decoder time per round | Round period | Backlog |
|---|---|---|
| 0.6 us | 1.0 us | none |
| 2.0 us | 1.0 us | 1.0 ms |
I would not consider it settled without evidence: Measure decoder time per round against the device's round period and require sustained headroom, not average parity.
A decoder that cannot keep up produces no logical result.
Curated: · Written: · Reviewed:
QA-65A result claims a logical qubit beat a physical one. What do you ask?(show answer)
I would treat logical break-even claims as a claim that has to survive being recomputed from the raw counts.
Which physical qubit, at what workload and for how long. Beating the average physical qubit is weaker than beating the best one, which is weaker again than beating the best one at matched duration and matched task — and only the last supports the sentence teams want to write.
Concretely, require the baseline to be named explicitly, and check whether the claim covers idle memory only or also logical operations, which need lattice surgery or transversal gates plus distillation.
The reason for that specificity is a failure I have seen: A break-even claim compared a distance-3 logical memory against the chip's median physical qubit; against the best physical qubit over the same duration the logical qubit was 1.8 times worse.
Three baselines, one experiment.
| Baseline | Physical error | Logical error | Break-even |
|---|---|---|---|
| median qubit | 0.9% | 0.5% | yes |
| best qubit | 0.28% | 0.5% | no |
I would not consider it settled without evidence: Report the logical error rate beside the best physical qubit's error over the same duration and the same task.
Name the baseline or the claim is unfalsifiable.
Curated: · Written: · Reviewed:
QA-66Why do non-Clifford gates dominate a fault-tolerant resource estimate?(show answer)
The useful question for magic state distillation as the dominant cost is what the hardware does that the simulator never showed.
Clifford operations are cheap in a stabilizer code, while T gates require magic states produced by distillation factories that consume many noisy states to yield one good one. In most algorithm estimates the factories occupy the majority of the physical qubits and much of the runtime.
Concretely, count T gates or T depth as the primary algorithmic cost metric under fault tolerance, and design circuits to reduce them even at the expense of Clifford count.
The reason for that specificity is a failure I have seen: A resource estimate counted total gates and concluded a circuit was modest; recounted by T gates, it needed 2.4 billion of them, and the factory footprint exceeded the data qubits by a factor of five.
Where the qubits go.
| Component | Physical qubits | Share |
|---|---|---|
| logical data | 320,000 | 17% |
| distillation factories | 1,600,000 | 83% |
I would not consider it settled without evidence: Report T count and T depth for any fault-tolerant estimate, separately from the Clifford count.
Under error correction, T gates are the currency.
Curated: · Written: · Reviewed:
QA-67Walk through what actually happens in one iteration of VQE.(show answer)
I would settle the structure of a variational loop against a shot budget before arguing about the algorithm.
A parameterized circuit prepares a candidate state, the device estimates the cost by measuring grouped Pauli terms over many shots, and a classical optimizer proposes new parameters from that noisy estimate. The optimizer never sees the landscape, only a sampled estimate of it.
Concretely, instrument each stage separately — circuit duration, shots per evaluation, estimator variance, optimizer steps — so a slow or non-converging run can be attributed rather than guessed at.
The reason for that specificity is a failure I have seen: A run was reported as 300 optimizer iterations; instrumented, each iteration consumed 1.2 million shots and 40 minutes of queue time, so the reported iteration count concealed 8 days of wall clock.
One reported iteration.
| Stage | Cost |
|---|---|
| circuit executions | 1,200,000 shots |
| queue and execution | 40 min |
| classical update | 0.4 s |
I would not consider it settled without evidence: Report shots and wall time per evaluation alongside the iteration count for any variational result.
Iterations are not the unit anyone pays in.
Curated: · Written: · Reviewed:
QA-68Why not use the most expressive ansatz you can fit on the device?(show answer)
The judgement in ansatz expressibility against trainability is which assumption is load-bearing and whether it was checked.
Expressibility and trainability trade against each other: a circuit that can reach any state approximates a random unitary, and random circuits have exponentially vanishing gradients. A problem-informed ansatz reaches fewer states, and reaches them from a landscape an optimizer can descend.
Concretely, prefer symmetry-preserving, problem-derived ansaetze that restrict the explored subspace, and measure gradient variance as the ansatz grows rather than assuming it will train.
The reason for that specificity is a failure I have seen: A 20-qubit hardware-efficient ansatz with 12 layers reached gradient variance near 10 to the minus 6, and 400 optimizer steps moved the energy by less than the shot noise on a single evaluation.
Gradient variance by ansatz.
| Ansatz | Qubits | Gradient variance |
|---|---|---|
| problem-informed | 20 | 1e-2 |
| deep hardware-efficient | 20 | 1e-6 |
I would not consider it settled without evidence: Sample gradient variance at random parameters for increasing width and check whether it decays exponentially before committing device time.
A landscape you cannot descend is not a model you can train.
Curated: · Written: · Reviewed:
QA-69Your optimizer makes no progress on a 20-qubit circuit. Is it the optimizer?(show answer)
Where candidates lose the interview on barren plateaus is treating an exact-state result as a device result.
Usually not. For deep circuits approximating random unitaries the gradient variance falls exponentially in qubit count, so a typical gradient component near 10 to the minus 3 requires on the order of 10 to the 6 shots per component to distinguish from zero. The landscape is flat rather than badly navigated.
Concretely, diagnose before optimizing by sampling gradient variance at random parameters across widths, and respond with structure — shallower problem-informed circuits, local cost observables, layerwise initialization — rather than with another optimizer.
The reason for that specificity is a failure I have seen: A team cycled through four optimizers over three weeks on a plateaued 20-qubit ansatz; the gradient variance was 10 to the minus 6 throughout and no optimizer could have helped.
Shots needed per gradient component.
| Qubits | Gradient variance | Shots to resolve |
|---|---|---|
| 8 | 1e-2 | ~100 |
| 20 | 1e-6 | ~1,000,000 |
I would not consider it settled without evidence: Measure gradient variance against qubit count and show whether it decays exponentially before attributing failure to the optimizer.
Changing the optimizer does not change the landscape.
Curated: · Written: · Reviewed:
QA-70How many shots does one VQE energy evaluation take?(show answer)
I would answer the measurement cost of an energy estimate by naming the measurement that would contradict it.
On the order of the squared sum of the Hamiltonian's coefficient magnitudes divided by the squared accuracy. A 12-qubit molecular Hamiltonian with a coefficient sum near 20 hartree, evaluated to chemical accuracy of 1.6 millihartree, needs about 1.6 times 10 to the 8 shots before any grouping.
Concretely, group commuting terms into simultaneously measurable sets and allocate shots proportionally to coefficient magnitude, which typically buys one to two orders of magnitude and turns an impossible run into an expensive one.
The reason for that specificity is a failure I have seen: A project budgeted 8,192 shots per energy evaluation for a Hamiltonian needing 10 to the 7 after grouping; the reported energies had an uncertainty of 0.05 hartree, thirty times the accuracy the conclusion required.
Shots per evaluation at 1.6 mHa accuracy.
| Strategy | Shots |
|---|---|
| term by term | 1.6e8 |
| commuting groups | 1e7 |
| weighted allocation | 2e6 |
I would not consider it settled without evidence: Derive the shot requirement from the coefficient sum and the target accuracy, and compare it against the budget before the run.
Chemical accuracy is a shot count before it is a chemistry result.
Curated: · Written: · Reviewed:
QA-71You want exact gradients on hardware. What do they cost?(show answer)
The engineering content of parameter-shift gradients and their price is the compilation and the error budget, not the notation.
Two additional circuit evaluations per parameter, since the parameter-shift rule evaluates the same circuit at shifted parameter values. A 200-parameter ansatz therefore costs 400 energy estimates per gradient step, each of which is itself a full shot budget.
Concretely, choose deliberately between exact gradients and stochastic approximations such as SPSA, which costs two evaluations per step regardless of parameter count at the price of a noisier trajectory, and state which was used.
The reason for that specificity is a failure I have seen: A 200-parameter run used parameter-shift gradients at 10 to the 6 shots per estimate, which put one optimizer step at 4 times 10 to the 8 shots; the run was cancelled after two steps.
Cost of one optimizer step.
| Method | Evaluations | Shots at 1e6 each |
|---|---|---|
| parameter-shift, 200 params | 400 | 4e8 |
| SPSA | 2 | 2e6 |
I would not consider it settled without evidence: Multiply parameters by two by the per-estimate shot cost to price a gradient step before choosing the optimizer.
Gradient exactness is bought per parameter.
Curated: · Written: · Reviewed:
QA-72You have a fixed shot budget and forty measurement groups. How do you split it?(show answer)
Before trusting anything about shot allocation across measurement groups I would write down what a wrong answer would look like.
Not evenly. The variance of the total estimate is the sum of the group variances scaled by their coefficients, so shots allocated proportionally to each group's coefficient magnitude and standard deviation minimize the total error for a fixed budget.
Concretely, allocate from the coefficients before the run and refine from measured group variances during it, rather than dividing the budget uniformly out of convenience.
The reason for that specificity is a failure I have seen: A uniform split gave a total uncertainty of 4.1 millihartree; the same budget reallocated by coefficient weight gave 1.4 millihartree, and the difference decided whether the result reached chemical accuracy.
Same 10 million shots, two allocations.
| Allocation | Total uncertainty |
|---|---|
| uniform | 4.1 mHa |
| coefficient-weighted | 1.4 mHa |
I would not consider it settled without evidence: Compare total estimator uncertainty under uniform and weighted allocation at the same total shot count.
The budget is a variance problem, not a fairness one.
Curated: · Written: · Reviewed:
QA-73Your QAOA run returns bitstrings that violate the problem's constraints. What happened?(show answer)
The first thing I would pin down about QAOA encoding and solution feasibility is which convention the numbers are stated in.
Constraints mapped into the objective as penalty terms are suggestions rather than guarantees, so the sampled bitstrings include infeasible ones. The reported objective silently scores them unless feasibility is checked separately.
Concretely, report the feasible fraction alongside the approximation ratio, and consider constraint-preserving mixers that keep the evolution inside the feasible subspace instead of penalizing departures from it.
The reason for that specificity is a failure I have seen: A run reported an approximation ratio of 0.88; 41 percent of the sampled bitstrings violated the cardinality constraint, and among feasible samples alone the ratio was 0.62.
Same run, two readings.
| Sample set | Fraction | Approximation ratio |
|---|---|---|
| all samples | 1.00 | 0.88 |
| feasible only | 0.59 | 0.62 |
I would not consider it settled without evidence: Filter samples by the constraint and report the objective over feasible samples with the feasible fraction beside it.
An infeasible sample is not a solution.
Curated: · Written: · Reviewed:
QA-74What do you compare your QAOA result against?(show answer)
I would answer choosing a classical baseline for a variational claim by separating what the mathematics guarantees from what the device delivers.
Against a tuned classical heuristic given the same wall-clock budget, and against the trivial bound the problem already guarantees. Comparing to a random assignment, or to an untuned solver, measures effort rather than capability.
Concretely, fix the baseline, its tuning budget and the instance set before running, and give the classical side compute comparable to the quantum side's device and queue time.
The reason for that specificity is a failure I have seen: A QAOA result on max-cut reported a 0.71 approximation ratio as evidence of promise; a random assignment guarantees 0.5 and a standard classical heuristic reached 0.93 on the same instances in under a second.
Same instances, three methods.
| Method | Approximation ratio | Time |
|---|---|---|
| random assignment | 0.50 | 0 s |
| QAOA depth 3 | 0.71 | 6 h |
| tuned heuristic | 0.93 | 0.8 s |
I would not consider it settled without evidence: Report the quantum result, the trivial bound and a tuned classical solver on identical instances with their runtimes.
A baseline nobody tuned is not a baseline.
Curated: · Written: · Reviewed:
QA-75How do you keep a quantum benchmark honest?(show answer)
This is an area where the ideal circuit and the circuit that executes preregistering a quantum benchmark are different programs.
By fixing the instances, the accuracy target, the baseline, the stopping rule and the metric before any outcome is inspected. Otherwise the reported result is selected from a family of analyses, and the selection is invisible in the write-up.
Concretely, record the protocol in the repository with a timestamp, run it, and report every trial including the restarts that failed and the instances that were dropped, with the reason.
The reason for that specificity is a failure I have seen: A study reported a 0.94 success rate over 8 instances; the protocol had run 31 instances and kept those where the ansatz converged, and over all 31 the rate was 0.42.
What the full ledger showed.
| Population | Instances | Success rate |
|---|---|---|
| reported subset | 8 | 0.94 |
| all attempted | 31 | 0.42 |
I would not consider it settled without evidence: Publish the preregistered protocol and the full trial ledger alongside the headline number.
Selection after the fact is the easiest result to produce.
Curated: · Written: · Reviewed:
QA-76Two frameworks both offer an expectation-value primitive. Are they the same thing?(show answer)
My answer to framework primitives and their differing semantics starts from the resource count rather than from the asymptotic claim.
Not necessarily. Primitives differ in whether they apply readout mitigation by default, how they group commuting terms, what they do with the shot count, and whether the returned value carries an uncertainty. Similar names do not imply identical estimators.
Concretely, read what the primitive does before comparing numbers across stacks, and reproduce one known value in both before treating a difference as physics.
The reason for that specificity is a failure I have seen: An expectation value differed by 0.06 between two stacks and was investigated as a device problem for a week; one primitive applied readout mitigation by default and the other did not.
Same circuit, two primitives.
| Stack | Mitigation default | Reported value |
|---|---|---|
| A | on | -0.734 |
| B | off | -0.672 |
I would not consider it settled without evidence: Run the same circuit and observable through both stacks with mitigation explicitly disabled and require agreement within shot noise.
Compare estimators before comparing results.
Curated: · Written: · Reviewed:
QA-77You export a circuit to OpenQASM and import it elsewhere. What do you check?(show answer)
I would treat round-tripping a circuit through an interchange format as a claim that has to survive being recomputed from the raw counts.
That the transformation preserved semantics rather than merely producing valid output. Each producer and consumer supports a subset of the format, so unsupported constructs are dropped or reinterpreted, and the imported circuit runs without complaint.
Concretely, round-trip through the format and compare the resulting unitary or sampled distribution against an independent reference on an asymmetric fixture, and require conversion to fail loudly on constructs it cannot represent.
The reason for that specificity is a failure I have seen: A conversion silently dropped a custom gate definition and substituted an identity; the imported circuit ran, produced plausible counts, and the discrepancy surfaced only when a colleague recomputed the expected distribution by hand.
What a round trip can lose.
| Construct | Instances in circuit | Survived | Silent |
|---|---|---|---|
| standard gates | 214 | yes | n/a |
| custom definition | 6 | no | yes |
| global phase | 1 | no | yes |
I would not consider it settled without evidence: Assert semantic equivalence after the round trip on a fixture with distinguishable amplitudes rather than a uniform one.
Valid output is not preserved meaning.
Curated: · Written: · Reviewed:
QA-78Your dynamic circuit behaves differently after export. Where do you look?(show answer)
The useful question for classical register aliasing across formats is what the hardware does that the simulator never showed.
At the classical registers. A format whose consumer flattens several classical registers into one changes which bit a conditional reads, so a branch that should fire on an ancilla parity fires on an unrelated measurement instead.
Concretely, name and index classical bits explicitly rather than relying on declaration order, and test each conditional branch by forcing the condition true and false.
The reason for that specificity is a failure I have seen: Two classical registers were flattened on import so a conditional X read bit 0 of the wrong register; the protocol's success rate fell from 0.91 to 0.52 and the circuit diagram looked correct throughout.
Register flattening.
| Before export | After import | Conditional reads |
|---|---|---|
| c0[0], c1[0] | c[0], c[1] | c[0] |
I would not consider it settled without evidence: Force each conditional branch in simulation and assert the gate fires on the intended measurement outcome.
Conditionals read bits, not intentions.
Curated: · Written: · Reviewed:
QA-79Does it matter that your rotation angles are serialized to six decimal places?(show answer)
I would settle parameter precision in serialization against a shot budget before arguing about the algorithm.
It matters in proportion to how many rotations there are. A truncation of about 10 to the minus 6 radians per gate is irrelevant once and comparable to a real error source after ten thousand gates, and it is invisible because the circuit remains valid.
Concretely, serialize parameters at full double precision or keep them symbolic until binding, and assert the bound circuit's unitary against the symbolic one for a representative case.
The reason for that specificity is a failure I have seen: A circuit with 12,000 rotations serialized at six decimals accumulated enough deviation to shift an expectation value by 0.9 percent, which was inside the team's tolerance and outside the accuracy their conclusion needed.
Truncation across a circuit.
| Rotations | Per-gate truncation | Effect on estimate |
|---|---|---|
| 10 | 1e-6 | negligible |
| 12,000 | 1e-6 | 0.9% |
I would not consider it settled without evidence: Compare results from full-precision and truncated parameter binding on the largest circuit in the workload.
Precision loss scales with gate count.
Curated: · Written: · Reviewed:
QA-80Your submission times out but the provider may have accepted it. What now?(show answer)
The judgement in idempotent job submission is which assumption is load-bearing and whether it was checked.
Without a durable idempotency key, a retry runs and bills the task twice, and on a device charging per task plus per shot a three-attempt retry loop triples the cost of every transient network error. The provider's acceptance and your knowledge of it are different events.
Concretely, generate an idempotency key before the request, persist it with the intended task, and on any ambiguous failure query by that key before resubmitting.
The reason for that specificity is a failure I have seen: A retry loop resubmitted through three gateway timeouts over a weekend, producing 214 duplicate tasks and a bill four times the expected one, with the duplicate results later averaged together as though independent.
One timeout, two behaviours.
| Client | Tasks created | Cost |
|---|---|---|
| naive retry | 3 | 3x |
| idempotent | 1 | 1x |
I would not consider it settled without evidence: Inject a timeout after the provider accepts and confirm the client recovers the existing task instead of creating a second.
Retry against a key, not against a hope.
Curated: · Written: · Reviewed:
QA-81How do you avoid discovering a device restriction from a failed job?(show answer)
Where candidates lose the interview on checking target capabilities at submission is treating an exact-state result as a device result.
By validating against the live target at submission rather than trusting the code path that worked yesterday. Basis gates, coupling map, maximum shots, supported instructions and available qubits are properties of the current target and they change.
Concretely, query the target, check the transpiled circuit against it, and fail locally with a specific message rather than sending a job that the provider will reject after an hour in a queue.
The reason for that specificity is a failure I have seen: A batch of 60 jobs used mid-circuit measurement on a backend that had it disabled during maintenance; all 60 failed after queueing for 90 minutes and the shot budget for the day was gone.
Where the failure surfaces.
| Check | Time to failure | Shots lost |
|---|---|---|
| local capability check | 0.2 s | 0 |
| provider rejection | 90 min | full batch |
I would not consider it settled without evidence: Assert that a deliberately unsupported instruction is rejected locally before submission rather than by the provider.
Fail before the queue, not after it.
Curated: · Written: · Reviewed:
QA-82What runs in CI for a quantum codebase, and what does not?(show answer)
I would answer simulator tiers in continuous integration by naming the measurement that would contradict it.
Almost every assertion belongs where execution is free. Ideal statevector checks run in milliseconds, a noisy tier exercises mitigation and post-selection logic, a provider-simulator tier validates credentials and serialization, and hardware runs only a small canary of known-answer circuits.
Concretely, keep hardware out of the pull-request path entirely and run the canary on deployment, so a provider default change fails a scheduled job rather than a developer's branch.
The reason for that specificity is a failure I have seen: A CI pipeline submitted three hardware jobs per pull request; queue times of 40 minutes made the pipeline unusable and the team disabled quantum tests altogether for two months.
Where assertions live.
| Tier | Runtime | Catches |
|---|---|---|
| ideal simulator | ms | logic and ordering |
| noisy simulator | s | mitigation logic |
| provider simulator | s | serialization, auth |
| hardware canary | min | provider changes |
I would not consider it settled without evidence: Show that each tier catches a class of defect the tier below cannot, and that the pull-request path uses no device time.
Hardware is a canary, not a test runner.
Curated: · Written: · Reviewed:
QA-83How do you assert on a sampled distribution without a flaky test?(show answer)
The engineering content of statistical tolerances in quantum tests is the compilation and the error budget, not the notation.
Derive the tolerance from the shot count. With 4,096 shots the standard error on a probability is under one percent, so an assertion at five percent passes through real regressions and an assertion on exact counts fails at random.
Concretely, assert within a stated number of standard errors, seed the sampler, and record the seed with the result so a failure can be reproduced exactly.
The reason for that specificity is a failure I have seen: A suite asserted counts within plus or minus 20 and failed roughly one run in six; the team responded by widening the tolerance to plus or minus 200, at which point it stopped detecting a real 8 percent regression.
Tolerance at 4,096 shots.
| Assertion | Flaky | Detects 8% regression |
|---|---|---|
| exact counts | yes | yes |
| +/- 5 percentage points | no | no |
| 3 standard errors | no | yes |
I would not consider it settled without evidence: Compute the assertion threshold from the shot count and demonstrate the test fails on an injected regression of the size that matters.
A tolerance is a statistics decision.
Curated: · Written: · Reviewed:
QA-84Why is pinning the SDK not enough to reproduce a run?(show answer)
Before trusting anything about pinning framework and provider versions together I would write down what a wrong answer would look like.
Because the executed circuit depends on the transpiler version, the optimization level, the router seed, the target snapshot and the provider API schema as well. A minor release can produce a different physical circuit from identical source.
Concretely, record the transpiled circuit itself with each result, alongside the versions and seeds, so reproduction does not depend on recreating the compiler's behaviour.
The reason for that specificity is a failure I have seen: A result could not be reproduced six weeks later: the same source transpiled under a newer release to a different layout with 22 percent more CNOTs, and only the source had been kept.
Same source, two releases.
| Release | CNOTs after transpile | Success |
|---|---|---|
| pinned | 96 | 0.71 |
| newer | 117 | 0.58 |
I would not consider it settled without evidence: Re-run a stored transpiled circuit and confirm it reproduces the recorded result, rather than re-transpiling from source.
Keep the artifact that ran.
Curated: · Written: · Reviewed:
QA-85Which seeds does a quantum experiment need to record?(show answer)
The first thing I would pin down about seeding and reproducing a stochastic run is which convention the numbers are stated in.
More than one: the sampler's shot seed, the transpiler's routing seed, and any classical optimizer's initialization seed. Recording one of the three makes a run partially reproducible, which in practice means not reproducible.
Concretely, capture all seeds in the result record together with the versions, and treat a result whose seeds are missing as an observation rather than an experiment.
The reason for that specificity is a failure I have seen: A variational run that reached a promising energy could not be repeated: the optimizer's initial parameters were drawn from an unseeded generator, and 20 subsequent restarts averaged 0.08 hartree worse.
Seeds a run depends on.
| Seed | Recorded | Effect if missing |
|---|---|---|
| shot sampler | yes | counts differ |
| transpiler router | no | different circuit |
| optimizer init | no | different result |
I would not consider it settled without evidence: Re-run from the recorded seeds on the same versions and require the same trajectory.
One unrecorded seed makes the run a story.
Curated: · Written: · Reviewed:
QA-86Your variational loop needs 100,000 circuit executions. Which modality?(show answer)
I would answer superconducting against trapped-ion trade-offs by separating what the mathematics guarantees from what the device delivers.
Gate speed decides this one. A transmon two-qubit gate runs in hundreds of nanoseconds against a trapped-ion gate in hundreds of microseconds, so a circuit that takes a few milliseconds on one takes seconds on the other, and the difference compounds across a hundred thousand executions.
Concretely, convert the modality comparison into total wall clock for the actual workload rather than comparing gate fidelities, and only then weigh the ion trap's connectivity and coherence advantages.
The reason for that specificity is a failure I have seen: A loop budgeted for a trapped-ion device at 3 seconds per execution needed 83 hours of pure execution time before queueing, against 2 hours on a superconducting device with worse per-gate fidelity.
100,000 executions of one circuit.
| Modality | Per execution | Total |
|---|---|---|
| superconducting | 70 ms | ~2 h |
| trapped ion | 3 s | ~83 h |
I would not consider it settled without evidence: Multiply per-execution duration by the required executions on each candidate target and compare the totals.
For a loop, throughput beats fidelity.
Curated: · Written: · Reviewed:
QA-87How much does device topology matter to your algorithm choice?(show answer)
This is an area where the ideal circuit and the circuit that executes connectivity topology as an algorithmic constraint are different programs.
As much as gate error, for any algorithm dense in long-range interactions. A device with all-to-all coupling executes those directly, while a degree-three lattice routes them at three CNOTs per SWAP, so the topology can dominate the comparison.
Concretely, compare the algorithm's interaction graph against each candidate coupling map before comparing error rates, and consider reformulating the algorithm to a locality-friendly form.
The reason for that specificity is a failure I have seen: An algorithm with all-to-all interactions among 12 qubits needed 214 CNOTs after routing on a lattice against 78 on a fully connected trap, and the lattice device's superior gate fidelity did not recover the difference.
Twelve qubits, all-to-all interactions.
| Topology | Routed CNOTs | Success at stated error |
|---|---|---|
| fully connected | 78 | 0.66 |
| degree-three lattice | 214 | 0.29 |
I would not consider it settled without evidence: Transpile the workload to each candidate topology and compare routed two-qubit counts and estimated success.
Topology is part of the algorithm's cost.
Curated: · Written: · Reviewed:
QA-88A vendor announces twice as many qubits. What have they told you?(show answer)
My answer to qubit count as a capability claim starts from the resource count rather than from the asymptotic claim.
Very little on its own. Usable width depends on which qubits are calibrated and connected, on two-qubit and readout performance, on coherence relative to gate duration, and on whether the compiler can place your workload without heavy routing.
Concretely, ask for the workload-shaped result instead: compile a representative circuit to the new target and compare success probability, shots and wall time against the previous one.
The reason for that specificity is a failure I have seen: A 127-qubit device replaced a 65-qubit one in a plan; the workload used 20 qubits, and after routing on the larger chip's sparser region the success probability fell from 0.62 to 0.41.
More qubits, worse result.
| Device | Qubits | Routed CNOTs | Success |
|---|---|---|---|
| previous | 65 | 84 | 0.62 |
| announced | 127 | 131 | 0.41 |
I would not consider it settled without evidence: Run the same workload on both targets and compare the end-to-end result rather than the announced width.
Width is not capability.
Curated: · Written: · Reviewed:
QA-89One device leads on quantum volume and another on a speed metric. Which do you pick?(show answer)
I would treat reading vendor benchmark scores as a claim that has to survive being recomputed from the raw counts.
Neither, from those numbers alone. Each score is defined by a protocol with its own compilation allowances, so a device can lead on one and trail on another without inconsistency, and benchmarks that permit full-circuit optimization measure the compiler as much as the hardware.
Concretely, use published scores to build a shortlist, then decide from your own compiled workload on each target under the accuracy and confidence your result requires.
The reason for that specificity is a failure I have seen: A selection made on quantum volume chose a device whose speed metric was four times lower; the variational workload needed 200,000 executions and the choice added eleven days of wall clock.
Two devices, two leaders.
| Device | Quantum volume | Executions per hour |
|---|---|---|
| A | 128 | 900 |
| B | 32 | 3,600 |
I would not consider it settled without evidence: Reproduce the vendor protocol or replace it with your own workload measurement before ranking devices.
Benchmarks shortlist; workloads decide.
Curated: · Written: · Reviewed:
QA-90Can you run your gate-model circuit on a quantum annealer?(show answer)
The useful question for annealers as a different computational model is what the hardware does that the simulator never showed.
No. An annealer evolves toward the ground state of an Ising or QUBO Hamiltonian and has no gate set to compile to; it returns low-energy samples rather than executing a program. It is a different computational model, not a slower device.
Concretely, decide first whether the problem is expressible as a QUBO with the connectivity available, and if it is, compare against classical solvers on the same objective and the same total time.
The reason for that specificity is a failure I have seen: A team spent six weeks trying to map a phase-estimation subroutine onto an annealer before establishing that the model cannot express it, and the effort produced nothing reusable.
What each model accepts.
| Model | Input | Gate set size | Output |
|---|---|---|---|
| gate model | circuit | 5 native gates | sampled bitstrings |
| annealer | QUBO | 0 gates | low-energy samples |
I would not consider it settled without evidence: Write the objective as an explicit QUBO and confirm the model can represent it before any hardware work begins.
The model decides what a device can express.
Curated: · Written: · Reviewed:
QA-91Your QUBO has 100 densely connected variables. How many physical qubits?(show answer)
I would settle minor embedding overhead on an annealer against a shot budget before arguing about the algorithm.
Far more than 100. Variables whose connectivity exceeds the hardware graph must be represented as chains of physical qubits coupled to act as one, so a dense 100-variable problem can consume thousands of qubits and the embedding step becomes the dominant constraint.
Concretely, compute the embedding before promising a problem size, and report embedding overhead and chain length distribution with any annealing result.
The reason for that specificity is a failure I have seen: A 100-variable dense problem was promised on a device advertising 5,000 qubits; the embedding needed chains averaging 38 qubits and did not fit, and the problem had to be decomposed.
Embedding a dense 100-variable problem.
| Quantity | Value |
|---|---|
| logical variables | 100 |
| mean chain length | 38 |
| physical qubits needed | ~3,800 |
I would not consider it settled without evidence: Produce the actual embedding for the target graph and report physical qubits used and maximum chain length.
Advertised qubits are not available variables.
Curated: · Written: · Reviewed:
QA-92Your annealing results are noisy and inconsistent. What tuning parameter do you check?(show answer)
The judgement in chain strength and broken chains is which assumption is load-bearing and whether it was checked.
Chain strength. Too weak and the qubits in a chain disagree, so the sample does not correspond to any assignment of the logical variable; too strong and the chain couplings dominate the problem couplings, flattening the energy landscape you meant to explore.
Concretely, report the chain break fraction with every run and treat it as a validity gate, resolving broken chains by an explicit stated rule rather than silently by majority vote.
The reason for that specificity is a failure I have seen: A run reported an excellent objective value with a 34 percent chain break fraction; the majority-vote repair had manufactured assignments that the annealer never sampled.
Chain strength sweep.
| Chain strength | Break fraction | Best objective |
|---|---|---|
| 0.5 | 34% | unreliable |
| 2.0 | 3% | usable |
| 8.0 | 0% | landscape flattened |
I would not consider it settled without evidence: Report the chain break fraction and the repair rule alongside the objective, and re-run at several chain strengths.
A broken chain is not a sample.
Curated: · Written: · Reviewed:
QA-93A neutral-atom platform advertises hundreds of atoms. What can it run?(show answer)
Where candidates lose the interview on neutral-atom arrays and analog simulation is treating an exact-state result as a device result.
It depends which mode. These platforms support programmable analog Hamiltonian evolution over large arrays as well as an evolving digital gate model, and a claim about one says little about the other.
Concretely, state the mode with the claim: an analog quench on 256 atoms and a gate-model circuit on the same hardware have different expressiveness, different error characterization and different verification methods.
The reason for that specificity is a failure I have seen: A plan assumed gate-model circuits at the array size quoted for analog operation; the available digital mode covered a fraction of the atoms, and the workload did not fit.
One platform, two modes.
| Mode | Scale quoted | Programs expressible |
|---|---|---|
| analog evolution | hundreds of atoms | Hamiltonian families |
| digital gates | smaller subset | general circuits |
I would not consider it settled without evidence: Confirm which mode the quoted size refers to and compile a representative workload in that mode.
The mode is part of the specification.
Curated: · Written: · Reviewed:
QA-94Your circuits take 70 milliseconds each. How long is the experiment?(show answer)
I would answer queue time in the total wall clock by naming the measurement that would contradict it.
Usually dominated by queueing rather than execution. On shared devices a job can wait minutes to hours, so an experiment of a thousand small jobs is a scheduling problem before it is a physics one.
Concretely, batch circuits into as few jobs as the provider allows, use session or reservation modes where available for iterative workloads, and report queue time separately from execution time.
The reason for that specificity is a failure I have seen: A variational loop submitted one job per iteration; 300 iterations at 40 minutes of queue each turned 6 hours of execution into 8 days of wall clock.
Same 300 iterations.
| Submission pattern | Queue total | Wall clock |
|---|---|---|
| one job per iteration | 200 h | ~8 days |
| session mode | 2 h | ~8 h |
I would not consider it settled without evidence: Record queue time and execution time separately for every task and report both with the result.
The device is fast; the queue is not.
Curated: · Written: · Reviewed:
QA-95A QPU task returns after your classical optimizer has moved on. What have you got?(show answer)
The engineering content of orchestrating a hybrid workload's failures is the compilation and the error budget, not the notation.
A correctness problem, not a latency one. A late result applied to a superseded parameter set corrupts the optimizer's trajectory silently, so the orchestration has to tie every result to the iteration that requested it.
Concretely, carry a request identifier through the loop, reject results whose identifier does not match the current iteration, and decide explicitly whether a timeout retries, skips or aborts.
The reason for that specificity is a failure I have seen: A loop applied a delayed gradient to the wrong parameter vector; the optimizer diverged over 40 iterations and the run was blamed on barren plateaus until the correlation with retries was noticed.
Late result, two handlings.
| Policy | Optimizer trajectory |
|---|---|
| apply whatever arrives | diverges |
| match on request id | unaffected |
I would not consider it settled without evidence: Inject a delayed response in test and assert the loop discards it rather than applying it.
A stale result is worse than no result.
Curated: · Written: · Reviewed:
QA-96Which parts of a hybrid workload run near the QPU and which do not?(show answer)
Before trusting anything about placing the stages of a hybrid architecture I would write down what a wrong answer would look like.
Placement follows latency requirements. Real-time feedback within a circuit must run on control hardware, near-time mitigation and parameter updates run beside the device, and data preparation or heavy post-processing runs wherever throughput is cheapest.
Concretely, draw the data movement and the latency requirement for each stage explicitly rather than treating the QPU as an opaque function call, and size the transfer as well as the compute.
The reason for that specificity is a failure I have seen: A design ran parameter updates in a service 90 milliseconds away from the control system; per-iteration overhead of 180 milliseconds exceeded the circuit execution time by a factor of two and doubled the run.
Latency by stage.
| Stage | Requirement | Placement |
|---|---|---|
| in-circuit feedback | sub-microsecond | control hardware |
| parameter update | milliseconds | beside the device |
| data preparation | seconds | anywhere |
I would not consider it settled without evidence: Measure per-stage latency and data volume in the real deployment topology rather than in a local prototype.
Latency decides the architecture, not the diagram.
Curated: · Written: · Reviewed:
QA-97Your dataset has 500 features. How do you get it into a quantum state?(show answer)
The first thing I would pin down about the cost of encoding classical data is which convention the numbers are stated in.
The encoding choice decides both qubit count and depth, and it is where most exponential-speedup intuitions quietly fail. Amplitude encoding packs 2 to the n features into n qubits but needs a preparation circuit exponential in n without structure or a QRAM that does not exist; angle encoding costs one qubit per feature and is cheap to prepare.
Concretely, state the encoding, its qubit count, its preparation depth and whether the claimed advantage survives with preparation included, before any model is trained.
The reason for that specificity is a failure I have seen: A proposal claimed an exponential advantage from amplitude encoding of 500 features into 9 qubits; the preparation circuit needed depth on the order of 2 to the 9 and consumed the entire claimed speedup.
Encoding 500 features.
| Encoding | Qubits | Preparation depth |
|---|---|---|
| amplitude | 9 | exponential |
| angle | 500 | shallow |
I would not consider it settled without evidence: Compile the state preparation for the real data and report its depth alongside the model's circuit.
Loading is part of the algorithm.
Curated: · Written: · Reviewed:
QA-98You want a quantum kernel model on 1,000 training samples. What does that cost?(show answer)
I would answer the sample scaling of a quantum kernel by separating what the mathematics guarantees from what the device delivers.
A Gram matrix over 1,000 samples has 499,500 distinct pairs, and at 4,096 shots each that is about 2.05 billion circuit executions. At a sustained ten thousand shots per second, the data preparation alone is roughly 57 hours of device time, and it scales quadratically in samples.
Concretely, price the Gram matrix before proposing the method, and consider reduced shot counts with an explicit noise model for the kernel entries rather than assuming precision is free.
The reason for that specificity is a failure I have seen: A project proposed a quantum kernel on 10,000 samples; the estimate came to over 200 billion executions, and the work was abandoned after two months that a five-minute calculation would have saved.
Gram matrix cost by sample count.
| Samples | Pairs | Shots at 4,096 each |
|---|---|---|
| 1,000 | 499,500 | 2.05e9 |
| 10,000 | 49,995,000 | 2.05e11 |
I would not consider it settled without evidence: Compute pairs, shots and device hours from the sample count before committing to the approach.
Quadratic in samples is the whole story.
Curated: · Written: · Reviewed:
QA-99Your estimated kernel matrix has negative eigenvalues. What do you do?(show answer)
This is an area where the ideal circuit and the circuit that executes a non-positive-semidefinite empirical kernel are different programs.
Finite shots and hardware noise make the empirical matrix noisy and asymmetric, and it can lose positive semidefiniteness, which is the property the convex solver assumes. The solver then fails or silently solves a different problem.
Concretely, symmetrize, then repair by eigenvalue clipping or a ridge on the diagonal, and treat the repair as part of the estimator: select its strength inside the cross-validation loop rather than against the test set.
The reason for that specificity is a failure I have seen: A model was trained on a repaired kernel whose ridge parameter had been tuned on the test split; the reported accuracy of 0.87 fell to 0.71 on a genuinely held-out set.
Where the accuracy went.
| Protocol | Reported accuracy |
|---|---|
| ridge tuned on test split | 0.87 |
| nested cross-validation | 0.71 |
I would not consider it settled without evidence: Report the repair method and its parameter, and show the selection happened inside nested cross-validation.
The repair is part of the model.
Curated: · Written: · Reviewed:
QA-100Your hybrid model beats the classical baseline. How do you show the quantum part did it?(show answer)
My answer to ablating a quantum machine learning claim starts from the resource count rather than from the asymptotic claim.
With a preregistered ablation, because a hybrid pipeline has at least four candidate sources of accuracy: the preprocessing, the feature map, the quantum-estimated quantities, and the classical model on top. Only replacing the quantum stage while holding everything else fixed isolates its contribution.
Concretely, substitute a random feature map of the same dimension and a classical kernel on the same preprocessed inputs, run several seeds, and report uncertainty over data splits rather than over shots.
The reason for that specificity is a failure I have seen: A model reported a 6-point accuracy gain over a classical baseline; replacing the quantum kernel with a random feature map of equal dimension recovered 5 of the 6 points, and a tuned classical kernel recovered all of them.
Where the gain came from.
| Pipeline | Held-out accuracy |
|---|---|
| quantum kernel | 0.81 |
| random feature map | 0.80 |
| tuned classical kernel | 0.82 |
I would not consider it settled without evidence: Report the held-out result of the full pipeline, the pipeline with the quantum stage replaced, and a tuned classical baseline given comparable tuning effort.
An unablated gain has no owner.
Curated: · Written: · Reviewed:
