Four things happen to your rule
This is the order the machine does them in, and each one answers a question you'd ask anyway.
You describe the rule
Say it the way it was taught to you and the engine reads it back, asking only for what you left out; or fill in five short steps yourself. No coding. Which market, what triggers a trade, when you get out. At the end you see your rule written back in plain words and you lock it.
Because the most common way people fool themselves is by changing the rule a little, running it again, and keeping the best result. Locking means the rule you described is the rule that gets tested. If you change it later, that's a new test, and the report says so. Submission freeze. The locked configuration is the only thing simulated; any change is a new submission counted in your lineage, and the lineage count feeds a multiplicity adjustment on the reported p-value. Without freeze, the test is just an optimizer with extra steps.We replay it on real history
One-minute bars back to 2008 on most markets. Your rule trades every day exactly as written, with commission and slippage charged on every trade.
Because every trade costs money even when you're right: a fee to the broker, and the small amount you lose because you buy at the asking price and sell at the bid. Lots of rules make money on paper and lose money after costs. We charge them up front so the number you see is the number you'd keep. For scale, Topstep's published round-turn on NQ, ES and RTY is $3.78, and $1.22 on the micros (help.topstep.com, "TopstepX Commissions and Fees", read Sep 3, 2026); the engine charges $4.20 plus a tick each way, on purpose the harsher of the two. $4.20 round-turn commission plus one tick adverse slippage per side, market orders filled at next bar's open. Then re-simulated at two and three ticks, and searched to the breakeven slippage. For bracketed rules, slippage changes which exits fire, so each rung is a full re-simulation, never an arithmetic subtraction. Reference: Topstep publishes $3.78 RT on ES/NQ/RTY, $1.22 on MES/MNQ/M2K, $4.32 on GC, $4.02 on CL (help.topstep.com, "TopstepX Commissions and Fees", read Sep 3, 2026, includes exchange and NFA fees); a sim fills the target on touch with no slippage, which the engine does not.We compare it to coin flips
Five thousand twins trade the same days at the same times with the same costs, but pick long or short by coin flip. If your rule can't beat them, it doesn't know anything.
Because markets mostly went up for fifteen years, so almost anything that bought looks smart. The coin-flip twins get every advantage your rule had except the one decision it claims to be good at: which direction. If your rule beats nearly all of them, that decision is worth something. If it doesn't, the profit came from the era, not from you. A random-side null: identical eligibility, entry and exit timestamps, and cost model, with direction drawn per trade. It isolates directional information from timing, frequency and drift. The p-value is the empirical share of 5,000 seeds that match or beat the real net. It's computed on the full window and separately on the unseen era, and the unseen-era p governs. True, and the point. The test gives your rule perfect discipline, honest costs, and no hindsight, then asks whether its one decision — direction — beats a coin flip. If it can’t under those conditions, execution won’t save it. If it can, the Livability lines on the report say what holding it would take. A backtest is the floor, not the ceiling, and a rule that fails the floor is not a rule. The engine measures directional information under idealized execution with a conservative cost model. Live execution adds variance and, usually, cost; it does not add information. A rule that fails the random-side null under these conditions has no edge for execution to preserve. Livability metrics (max months to new high, worst 12 months, trailing-drawdown survival) are reported so the reader can judge whether a validated rule is holdable, which is a separate question from whether it is real.We score it and tell you plainly
A score from 0 to 100 and one of three words: Validated, Marginal, or Rejected. Then the lines that explain it, each one in language you chose at the top of this page.
Six checks, each worth a fixed number of points: is the data clean, does the rule work on years you never looked at, is it steady year to year, does it beat the coin flips, does it survive costs, and does it fall apart if you nudge a setting. A rule has to do reasonably well on all six. One great number can't rescue five bad ones. Protocol v1.1: six weighted examinations (data integrity 5, out-of-sample 25, era consistency 20, Monte Carlo 15, cost stress 20, parameter sensitivity 15). Validated needs 70+ with no pillar under half its points and no gate tripped; gates include insufficient sample and era-manufactured edge. The exact thresholds are ours; the structure and every input are printed on your report.- Fills at the next bar’s open, one tick against you per side, $4.20 a round turn. A target limit has to be traded through, not touched.
- All times are US Eastern. Flat at 15:30 is 2:30 Chicago, forty minutes inside a Topstep close-out.
- The unseen era governs. Scored on the years your development never touched; a result that only exists in the years you looked at trips a gate.
- Every look at the data is counted. Your declared lineage is printed on the report and priced into the probability. Nobody gets to show you the winner of thirty.
- The exit you describe is the exit that runs. “No stop” on a public teardown is the baseline; the bracketed version is measured next, as its own strategy.
- Reproducible from the rule text. Same data plus the printed rule lands within tolerance of the printed number, and every page says so.