Everything a pessimistic backtest needs
Eight capabilities, one local engine, no tiers — and a quant reviewer that will run all eight for you and rule on the result. Each of them runs in the free notebook demo on simulated data; open it in a tab and follow along.
A quant reviewer, not a chatbot
It does the job a second pair of hands would do on a research desk: read the result back in plain language, then go and find out whether the edge is real or an artefact of the search. It runs the cells itself, narrates each step with the number that step produced, and ends in a verdict it is willing to make negative. Everything quoted on this page is printed by the guest notebook on its default run.
The same ✦ key in every cell rail
One assistant, entered with context — never a second one. On cell [7], the search audit, it offers “How much have I searched?”, “What does that cost my Sharpe?”, “Show me what I discarded” and “Is this a real edge, or did I search?”. On cell [9] it offers “Is the optimiser fitting noise?” and “Which fold broke it?”. Three to five actions written for the surface you are standing on.
Six questions a system developer actually has
Is this a real edge, or did I search until it looked good? Is it better than luck? Does it survive its costs? Which parameter actually matters? Where does it break? Can I make it more robust? It runs the 49-square sweep, the 400 matched random entries, the six walk-forward folds and the cost re-price itself, then rules.
A standing instruction, then it works alone
Four mandates in your own words: tell me when something I am running stops being an edge · keep watching my search count · re-validate anything I adopt · only speak when something actually breaks. It re-checks on the clock and on every change, and reports that nothing crossed as readily as it reports that something did.
“Tell me when something I am running stops being an edge.”
Watching the noise bar, the deflated Sharpe, walk-forward efficiency and the drawdown halt — on the clock and on every change you make.
- 2 checks made
- 1 crossed something
- 1 nothing crossed
- 0 configurations priced
-
00:44:35 Crossed the drawdown against the −20 % halt, after you set the mandate
Worst drawdown is −23.3 %, through the halt. On this sample the system would have stopped trading and asked for a human.
What it did: Reported it and stopped. I do not move a parameter you set.
-
00:44:41 Clear the noise bar, the deflated Sharpe, walk-forward efficiency and the drawdown halt, after a scheduled check
Nothing crossed. Sharpe 1.09 against a noise bar of 0.83, deflated 76.7 %, drawdown −23.3 %, walk-forward efficiency 0.52.
What it did: Nothing. That is the answer most of the time, and I would rather say it than manufacture a finding.
Both entries land in the research log in cell [12] as a second register, exportable as CSV alongside the runs you made yourself. Pause, resume, undo the last thing it did or change the mandate — one click each.
“Is this a real edge, or did I search until it looked good?”
- Count every configuration this session has already evaluated
49 configurations so far — 1 by hand, 48 from sweeps. Best-of-49 noise on 7.5 years is 0.83; you are holding 1.09. Deflated Sharpe 76.9 %.
- Sweep the 49 settings around yours for a plateau or a spike
A plateau. 1.09 here against a neighbouring median of 1.07, spread ±0.05 across the grid. Most robust square: lookback 60 at ATR 2.5×, worst neighbour 1.03.
- Race it against 400 random entries with the same footprint
100.0th percentile against 400 coin flips that traded 190 times and held 6 days each. Median coin flip −0.00, luckiest 0.97, buy-and-hold 0.48.
- Cut six walk-forward folds and test on data the parameters never saw
η = 0.52 across six rolling folds, 6 / 6 of them profitable out-of-sample. Borderline.
- Re-price the identical parameters at 1 bp a side
Gross 1.41 against your net 1.09, so the cost model is taking 0.31 — 22 % of what it makes before it pays to trade.
Not yet. It clears the noise bar, but not every check agrees.
confidence: medium · 4 of 5 checks passIt is above what luck would hand you, and the deflated Sharpe puts the chance it survives 49 looks at 76.9 %. What is against it: the deflated Sharpe is 76.9 %. That is a lead worth validating rather than a result worth trading, and the honest next move is data this notebook has not touched.
Its own re-price at 1 bp is a configuration, so cell [7] counts it. The run that answered your question also moved the bar the answer is measured against, and the reviewer says so before it starts rather than after.
An automation researcher that bills itself for looking
The reviewer answers when you ask. The AI Automation Researcher does not wait: you compose an agent in a six-question wizard, save it, and it schedules itself on the notebook’s own clock — building candidates, running them, pricing their costs, sweeping their neighbourhoods, racing them against coin flips, cutting folds and resampling, and reporting each one. In a product whose whole claim is that looking harder makes a result less believable, an agent that searches continuously is the most dangerous thing that could be added. So it is built the only way it honestly can be: every configuration it evaluates goes through the same counter in cell [7] as a slider you moved yourself, and the bill is on the screen while it runs.
- Risk profile — the same six the notebook already uses. Costs: the profile prices its own configurations when applied; the researcher is billed separately on top.
- Direction — long only, or long and short. Costs: shorting opens the two books that are long and short by construction (6 families instead of 4), and adds a borrow fee and a recall risk this engine does not model, plus a loss with no natural bound. It does not change the sizing.
- Timeframe — 1h, 4h, 1D, or let it decide. Costs: deciding means pricing all three and reporting which it picked and why — 13 configurations a cycle instead of 11.
- Edge families — trend, breakout, mean reversion, cross-sectional relative value, macro regime, filter stack; each one a real notebook. Costs: every extra family multiplies the search space and therefore the bar.
- Budget — 40, 80, 160 or 320 configurations. Costs: exactly that, drawn down in front of you. There is no unlimited option, and there will not be one.
- What it may not do — shown, never offered: quarter-Kelly, the 7.5 % per-asset ceiling, the −20 % halt, the 20 bps slippage cap, no parameter you set is ever moved, and no configuration is evaluated without going on the counter.
- 1 · Build it — family, lookback, stop width, breadth-first so no family gets a second setting before the others have had a first.
- 2 · Run it — the same engine and cost model the notebook uses; adopting the candidate reproduces the figures exactly.
- 3 · Price the cost model — gross at 1 bp against net, and again at 45 bps a side. 2 configurations charged.
- 4 · Sweep the neighbourhood — the eight settings one step either side, the same arithmetic cell [5] draws: plateau, ridge or spike. 8 configurations charged.
- 5 · Race the coin flips — 400 books entering at random with the same trade count, hold and exposure. Charges nothing: they are draws from the benchmark, not configurations.
- 6 · Cut six folds — walk-forward efficiency and the mean Sharpe the out-of-sample windows printed. Charges no configuration and is not free: it spends the sample.
- 7 · Resample — 1,000 twelve-month paths and the share that touch the −20 % halt.
- 8 · Report — six checks, and the verdict, which is usually no.
72 of 80 configurations spent, and the counter moved by exactly 72
Trend and mean reversion, long and short, 4h, sceptic profile, 80-configuration ceiling. Ten candidates built. 110 configurations evaluated, 38 of them already tried and therefore free. It stopped when the next cycle needed more than the 8 it had left, and said so.
Four candidates cleared all six checks when they were found; three still clear them at 121 configurations. The fourth fell off without changing at all — the counter rose underneath it — and the shortlist says so in its place. The results list is ranked on the out-of-sample Sharpe the folds printed minus the noise bar in force when the candidate turned up, not on raw Sharpe: raw Sharpe would crown the candidate that printed 2.36 in sample and kept only 0.54 of it out of sample.
Event-driven, point-in-time, next-bar fills
Vectorized backtesters are fast and wrong: they let indicators see the whole array. Backcast Labs feeds your strategy one bar at a time. Signals computed at the close of bar t fill at the open of t+1, plus slippage. Indexing a future bar raises LookaheadError at run time.
- The same engine code path serves live execution — no “backtest vs live” divergence
- Intrabar fill models: next-open, VWAP-of-bar, worst-of-bar (stress)
- Deterministic: seed + config gives an identical result, hash printed in every report
1,000 resampled paths, drawdown-halt aware
Bootstrap the trade sequence — plain, block or stationary — and read the full distribution of outcomes: P5/P50/P95 terminal equity, the drawdown distribution, and one sentence you can act on: the probability of hitting the −20 % halt within the horizon.
- Trade-level, block and stationary bootstrap
- Paths respect the halt: when the guard would fire, that path stops trading
- Deflated Sharpe ratio reported alongside the raw Sharpe
Anchored or rolling, 3–12 folds, one stitched OOS curve
Backcast Labs optimizes on each in-sample window, freezes the best parameters, runs them on the following out-of-sample window and stitches every OOS segment into one equity curve — the only curve worth comparing with live performance.
- Per-fold table: IS Sharpe, OOS Sharpe, OOS return, chosen parameters
- Efficiency ratio η = mean OOS Sharpe ÷ mean IS Sharpe, with a 0.5 threshold
- Parameter-drift chart: see when the optimum wanders
Multi-asset books with shared capital and caps
Test 18 assets as one portfolio, not 18 isolated curves summed together. Capital is shared, sizing is Kelly 0.25 by default, no single asset exceeds 7.5 % of the book, and slippage is charged at the portfolio level when several legs rebalance on the same bar.
- Monthly-returns heatmap, allocation donut, per-asset contribution
- Cross-sectional ranking strategies (long top-k / short bottom-k) built in
- Correlation-aware sizing and a portfolio drawdown halt (−20 % default)
Slippage, commission, funding and borrow — on by default
The single biggest reason strategies break live is that the backtest paid nothing to trade. Backcast Labs defaults to 20 bps slippage and exchange-realistic commission, and lets you model perpetual funding, borrow costs for shorts and liquidity-dependent slippage curves.
- Slippage: fixed bps, spread-based, or volume-participation curve
- Perpetual funding and short borrow from your data pack or your own keys
- Sensitivity sweep: Sharpe against slippage 0–50 bps in one chart
Survivorship-bias-free, checksummed, tick-level where available
History that only contains today's winners flatters every long strategy. Data packs include delisted symbols, apply corporate actions and normalize timestamps to UTC. Or bring your own keys: the engine pulls candles and trades straight from your exchange account and caches them locally.
- Planned: packs from €49 one-time — no data subscription
- Your own exchange and broker keys, plus CSV/Parquet
- SHA-256 manifest per pack; the engine refuses tampered files
Run an imported strategy honestly
Import a Python strategy and it becomes a notebook: entries and exits mapped onto engine orders, tunable constants promoted to param() declarations, and every read of data that was not available at that bar flagged before the first run. Then it is run again with slippage and commission on, so the optimistic version and the honest one sit side by side.
- Repainting audit: every higher-timeframe read that would not have been available at that bar
- Side by side: the optimistic run against the point-in-time, costs-on run
- Every tunable constant kept as a
param()so walk-forward can sweep it
Export the validated configuration — no lock-in
When walk-forward and Monte Carlo pass your thresholds, export one signed config: frozen parameters, sizing, per-asset caps, drawdown halt and slippage cap, stamped with the engine build hash. Hand it to your own execution platform as JSON, CSV or a webhook payload.
- Signed JSON config, versioned with the engine build hash
- Kelly 0.25 sizing and the −20 % drawdown halt travel with the strategy
- Import realised fills back in and re-measure your true slippage
Everything in the planned €299 licence
Planned pricing, not on sale yet. No tiers, no bar caps, no “pro” unlock. Data packs and the update contract are the only add-ons.
| Capability | Backcast Labs €299 | Typical cloud backtester |
|---|---|---|
| Event-driven engine, next-bar fills | ✓ | Varies |
| Bars / symbols per test | Unlimited (your RAM) | Capped by plan |
| Monte Carlo (trade / block bootstrap) | ✓ 1,000+ paths | Add-on or manual |
| Walk-forward (anchored + rolling) | ✓ | ✗ |
| Deflated Sharpe ratio | ✓ | ✗ |
| An agent that runs the validation, not one that describes it | ✓ six plans, real cells | ✗ |
| Standing instruction — say it once, it keeps checking | ✓ four mandates | ✗ |
| The agent's own configurations counted against your search bar | ✓ logged and tagged | ✗ |
| A record of the checks that found nothing | ✓ exportable as CSV | ✗ |
| Slippage & commission on by default | ✓ 20 bps | Off by default |
| Portfolio-level slippage | ✓ | ✗ |
| Survivorship-bias-free data | ✓ data packs | Some markets |
| Strategy import as a notebook | ✓ with repainting audit | ✗ |
| Strategies stay on your disk | ✓ always | Uploaded to the cloud |
| Export to live execution | ✓ signed config, CSV, webhook | Platform lock-in |
| Price after 3 years | €299 planned (+ optional updates) | A recurring fee every year |