Defaults are deliberately pessimistic: 20 bps slippage, 5 bps commission, next-bar fills, delisted symbols included.
Is this an edge, or did you search until something looked good?
A guest notebook on seeded sample markets. Defaults are deliberately pessimistic — 20 bps slippage, 5 bps commission, next-bar fills, delisted symbols kept in — so what you read here is the unflattering version. Cells 5 to 7 are the ones most backtesters skip: whether your setting sits on a plateau or a spike, whether it beats coin-flip entries with the same footprint, and how much of the result is explained by how many variants you tried.
A backtest is only as honest as its data. Every choice here has a price, and this cell puts a number on each one.
| Property | Value | What it costs you |
|---|
See what the strategy actually returned against simply holding the market — scroll to zoom, drag to pan, double-click to reset.
The halt is not a forecast. It is the level at which the system stops trading and asks for a human, which is why the curve is drawn against it.
The numbers a serious allocator asks for — including the deflated Sharpe that discounts how many variants you tried.
| Metric | Value | Reads as |
|---|
The same backtest re-run across the neighbourhood of your two shape parameters. A result that only works at exactly these numbers is not a result.
Beating buy-and-hold is table stakes. The real test is beating coin-flip entries that trade exactly as often, hold exactly as long and pay exactly the same costs.
| Comparison | Sharpe | CAGR | Max DD | Verdict |
|---|
Every configuration you run this session is counted. The more you try, the higher the Sharpe you would expect from noise alone — so the bar this result has to clear moves up as you search.
The noise line is the expected maximum of N independent Sharpe estimates drawn from a strategy with no edge at all, on a sample this long. It is not a p-value; it is the score to beat before the word “edge” means anything.
How bad a year can get — before you find out with real money.
Proof the parameters still work on data they were never fitted on.
| Fold | In-sample window | Out-of-sample window | IS Sharpe | OOS Sharpe | OOS return | Lookback chosen |
|---|
Where the returns actually came from, and how thinly the risk is spread across them.
Every round trip with its own fill cost, so nothing hides inside an average.
| # | Opened | Asset | Side | Entry | Exit | Hold | Slippage | Tick fill | P&L |
|---|
The record that keeps you honest with yourself. Every run this session, with its parameters, its result and the moment it happened — exportable, so the discarded attempts are on the page too.
| # | Time | Strategy | TF | Lookback | ATR | Costs | Sharpe | Max DD | Trades | Kept |
|---|
The same record, second register: every check the standing review made on its own, what it was watching, and what it found — including the checks where nothing crossed, which is most of them. Anything it priced is a row in the table above, tagged, and counted in cell [7] like a configuration you set yourself.
| # | Time | Mandate | What it checked | What it found | What it did |
|---|
The validated configuration as your execution platform reads it — parameters, costs and the risk guards travelling with them, so nothing is re-typed by hand at the other end.