Build the classical baseline first, every time
Two thirds of the quantum advantage claims we are asked to review disappear when you tune the classical solver properly. Here is the discipline we use instead.
Every quantum engagement we run starts with a stage that has no qubits in it at all. Before any circuit is written, we build the best classical solution we reasonably can and tune it properly. Clients sometimes read this as billable throat-clearing. It is the opposite: it is the only thing that makes the rest of the programme interpretable.
The failure mode it prevents
A quantum result on its own is not a result. 0.87 AUC means nothing until you know what a gradient
boosted tree gets on the same split. When we are asked to review an advantage claim, the pattern is
almost always one of the following:
- The classical comparison used library defaults, never tuned.
- The classical comparison was given less wall-clock time than the quantum run.
- The problem instance was small enough that the comparison was never the point.
- The metric was chosen after the results were in.
None of these require bad faith. They are what happens when the baseline is an afterthought rather than a deliverable with its own sign-off.
What "properly tuned" means here
We hold the baseline to the standard we would hold a production model to:
- a real hyperparameter search, not a hand-picked configuration
- the same feature engineering budget the quantum model received
- the same cross-validation splits, fixed and committed before any quantum run
- wall-clock and cost recorded for both sides
from sklearn.model_selection import StratifiedKFold
import numpy as np
# Splits are generated once, hashed, and committed to the repo before any
# device time is bought. Both the classical and quantum models see the
# identical folds — no exceptions, no re-splitting after a bad result.
SEED = 20260918
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=SEED)
folds = [(tr.tolist(), te.tolist()) for tr, te in cv.split(X, y)]
np.save("folds.npy", np.array(folds, dtype=object), allow_pickle=True)
print(f"committed {len(folds)} folds @ seed {SEED}")
Freezing the splits before the first hardware run is the single highest-leverage habit in this whole process. It removes the degree of freedom that quietly manufactures most advantage claims.
It is also the cheapest useful outcome
Here is the part that is awkward to put in a proposal:
A properly tuned classical baseline is, often enough, the thing the client ends up shipping.
That is a good outcome. The programme produced a working system, a measurement harness, and a documented answer about where quantum methods currently sit for this problem. The alternative — two years of circuit work with nothing to compare it against — costs more and tells you less.
The shape of the deliverable
| Artefact | Produced in | Survives the engagement |
|---|---|---|
| Frozen CV splits | Stage 02 | Yes — reused for every later run |
| Tuned classical model | Stage 02 | Yes — often goes to production |
| Evaluation harness | Stage 02 | Yes — the thing that scores everything after |
| Quantum circuit | Stage 03 | Only if it clears the baseline |
If the circuit never clears the baseline, you still keep three of the four rows. That asymmetry is the whole argument for doing it in this order.