Early access Backcast Labs is not on sale yet. The research notebook runs in your browser on simulated data. Join early access
Examples

Everything a pessimistic backtest needs

Eight capabilities, one local engine, no tiers — and a quant reviewer that will run all eight for you and rule on the result. Each of them runs in the free notebook demo on simulated data; open it in a tab and follow along.

The agent

A quant reviewer, not a chatbot

It does the job a second pair of hands would do on a research desk: read the result back in plain language, then go and find out whether the edge is real or an artefact of the search. It runs the cells itself, narrates each step with the number that step produced, and ends in a verdict it is willing to make negative. Everything quoted on this page is printed by the guest notebook on its default run.

Every view

The same key in every cell rail

One assistant, entered with context — never a second one. On cell [7], the search audit, it offers “How much have I searched?”, “What does that cost my Sharpe?”, “Show me what I discarded” and “Is this a real edge, or did I search?”. On cell [9] it offers “Is the optimiser fitting noise?” and “Which fold broke it?”. Three to five actions written for the surface you are standing on.

Send it

Six questions a system developer actually has

Is this a real edge, or did I search until it looked good? Is it better than luck? Does it survive its costs? Which parameter actually matters? Where does it break? Can I make it more robust? It runs the 49-square sweep, the 400 matched random entries, the six walk-forward folds and the cost re-price itself, then rules.

Say it once

A standing instruction, then it works alone

Four mandates in your own words: tell me when something I am running stops being an edge · keep watching my search count · re-validate anything I adopt · only speak when something actually breaks. It re-checks on the clock and on every change, and reports that nothing crossed as readily as it reports that something did.

Standing instruction · said once

“Tell me when something I am running stops being an edge.”

Watching the noise bar, the deflated Sharpe, walk-forward efficiency and the drawdown halt — on the clock and on every change you make.

  • 2 checks made
  • 1 crossed something
  • 1 nothing crossed
  • 0 configurations priced
  • 00:44:35 Crossed the drawdown against the −20 % halt, after you set the mandate

    Worst drawdown is −23.3 %, through the halt. On this sample the system would have stopped trading and asked for a human.

    What it did: Reported it and stopped. I do not move a parameter you set.

  • 00:44:41 Clear the noise bar, the deflated Sharpe, walk-forward efficiency and the drawdown halt, after a scheduled check

    Nothing crossed. Sharpe 1.09 against a noise bar of 0.83, deflated 76.7 %, drawdown −23.3 %, walk-forward efficiency 0.52.

    What it did: Nothing. That is the answer most of the time, and I would rather say it than manufacture a finding.

Both entries land in the research log in cell [12] as a second register, exportable as CSV alongside the runs you made yourself. Pause, resume, undo the last thing it did or change the mandate — one click each.

You asked it · it went and did the work

“Is this a real edge, or did I search until it looked good?”

  • Count every configuration this session has already evaluated

    49 configurations so far — 1 by hand, 48 from sweeps. Best-of-49 noise on 7.5 years is 0.83; you are holding 1.09. Deflated Sharpe 76.9 %.

  • Sweep the 49 settings around yours for a plateau or a spike

    A plateau. 1.09 here against a neighbouring median of 1.07, spread ±0.05 across the grid. Most robust square: lookback 60 at ATR 2.5×, worst neighbour 1.03.

  • Race it against 400 random entries with the same footprint

    100.0th percentile against 400 coin flips that traded 190 times and held 6 days each. Median coin flip −0.00, luckiest 0.97, buy-and-hold 0.48.

  • Cut six walk-forward folds and test on data the parameters never saw

    η = 0.52 across six rolling folds, 6 / 6 of them profitable out-of-sample. Borderline.

  • Re-price the identical parameters at 1 bp a side

    Gross 1.41 against your net 1.09, so the cost model is taking 0.31 — 22 % of what it makes before it pays to trade.

Not yet. It clears the noise bar, but not every check agrees.

confidence: medium · 4 of 5 checks pass

It is above what luck would hand you, and the deflated Sharpe puts the chance it survives 49 looks at 76.9 %. What is against it: the deflated Sharpe is 76.9 %. That is a lead worth validating rather than a result worth trading, and the honest next move is data this notebook has not touched.

Configurations counted4950
Deflated Sharpe76.9 %76.7 %

Its own re-price at 1 bp is a configuration, so cell [7] counts it. The run that answered your question also moved the bar the answer is measured against, and the reviewer says so before it starts rather than after.

The agent, unattended

An automation researcher that bills itself for looking

The reviewer answers when you ask. The AI Automation Researcher does not wait: you compose an agent in a six-question wizard, save it, and it schedules itself on the notebook’s own clock — building candidates, running them, pricing their costs, sweeping their neighbourhoods, racing them against coin flips, cutting folds and resampling, and reporting each one. In a product whose whole claim is that looking harder makes a result less believable, an agent that searches continuously is the most dangerous thing that could be added. So it is built the only way it honestly can be: every configuration it evaluates goes through the same counter in cell [7] as a slider you moved yourself, and the bill is on the screen while it runs.

The six questions, and what each answer costs
  • Risk profile — the same six the notebook already uses. Costs: the profile prices its own configurations when applied; the researcher is billed separately on top.
  • Direction — long only, or long and short. Costs: shorting opens the two books that are long and short by construction (6 families instead of 4), and adds a borrow fee and a recall risk this engine does not model, plus a loss with no natural bound. It does not change the sizing.
  • Timeframe — 1h, 4h, 1D, or let it decide. Costs: deciding means pricing all three and reporting which it picked and why — 13 configurations a cycle instead of 11.
  • Edge families — trend, breakout, mean reversion, cross-sectional relative value, macro regime, filter stack; each one a real notebook. Costs: every extra family multiplies the search space and therefore the bar.
  • Budget — 40, 80, 160 or 320 configurations. Costs: exactly that, drawn down in front of you. There is no unlimited option, and there will not be one.
  • What it may not do — shown, never offered: quarter-Kelly, the 7.5 % per-asset ceiling, the −20 % halt, the 20 bps slippage cap, no parameter you set is ever moved, and no configuration is evaluated without going on the counter.
Eight steps a candidate, each logged with what, why and the number
  • 1 · Build it — family, lookback, stop width, breadth-first so no family gets a second setting before the others have had a first.
  • 2 · Run it — the same engine and cost model the notebook uses; adopting the candidate reproduces the figures exactly.
  • 3 · Price the cost model — gross at 1 bp against net, and again at 45 bps a side. 2 configurations charged.
  • 4 · Sweep the neighbourhood — the eight settings one step either side, the same arithmetic cell [5] draws: plateau, ridge or spike. 8 configurations charged.
  • 5 · Race the coin flips — 400 books entering at random with the same trade count, hold and exposure. Charges nothing: they are draws from the benchmark, not configurations.
  • 6 · Cut six folds — walk-forward efficiency and the mean Sharpe the out-of-sample windows printed. Charges no configuration and is not free: it spends the sample.
  • 7 · Resample — 1,000 twelve-month paths and the share that touch the −20 % halt.
  • 8 · Report — six checks, and the verdict, which is usually no.
Measured, on one saved assignment in the guest notebook

72 of 80 configurations spent, and the counter moved by exactly 72

Trend and mean reversion, long and short, 4h, sceptic profile, 80-configuration ceiling. Ten candidates built. 110 configurations evaluated, 38 of them already tried and therefore free. It stopped when the next cycle needed more than the 8 it had left, and said so.

Configurations counted49121
Deflated Sharpe of your own result76.9 %65.8 %
Noise bar0.830.95
Your Sharpe, unchanged1.09

Four candidates cleared all six checks when they were found; three still clear them at 121 configurations. The fourth fell off without changing at all — the counter rose underneath it — and the shortlist says so in its place. The results list is ranked on the out-of-sample Sharpe the folds printed minus the noise bar in force when the candidate turned up, not on raw Sharpe: raw Sharpe would crown the candidate that printed 2.36 in sample and kept only 0.54 of it out of sample.

1

Event-driven, point-in-time, next-bar fills

Vectorized backtesters are fast and wrong: they let indicators see the whole array. Backcast Labs feeds your strategy one bar at a time. Signals computed at the close of bar t fill at the open of t+1, plus slippage. Indexing a future bar raises LookaheadError at run time.

  • The same engine code path serves live execution — no “backtest vs live” divergence
  • Intrabar fill models: next-open, VWAP-of-bar, worst-of-bar (stress)
  • Deterministic: seed + config gives an identical result, hash printed in every report
Engine trace · bar 14,201
Signal at close · fill at next open+20 bps
Figure 1. Signal at the close of bar 41, fill at the open of bar 42 plus 20 bps. The gap between marker and fill is where optimistic backtests hide their edge.
2

1,000 resampled paths, drawdown-halt aware

Bootstrap the trade sequence — plain, block or stationary — and read the full distribution of outcomes: P5/P50/P95 terminal equity, the drawdown distribution, and one sentence you can act on: the probability of hitting the −20 % halt within the horizon.

  • Trade-level, block and stationary bootstrap
  • Paths respect the halt: when the guard would fire, that path stops trading
  • Deflated Sharpe ratio reported alongside the raw Sharpe
Monte Carlo · 600 paths · 252 days
Resampled trade sequencesblock size 20
Figure 2. Fan of resampled paths over one trading year. The median path is drawn heavy; P5 / P50 / P95 terminal equity is printed top-right.
3

Anchored or rolling, 3–12 folds, one stitched OOS curve

Backcast Labs optimizes on each in-sample window, freezes the best parameters, runs them on the following out-of-sample window and stitches every OOS segment into one equity curve — the only curve worth comparing with live performance.

  • Per-fold table: IS Sharpe, OOS Sharpe, OOS return, chosen parameters
  • Efficiency ratio η = mean OOS Sharpe ÷ mean IS Sharpe, with a 0.5 threshold
  • Parameter-drift chart: see when the optimum wanders
Walk-forward · rolling · 6 folds · 70/30
In-sample → out-of-sampleOOS Sharpe printed per fold
Figure 3. Six rolling folds. Blue: in-sample optimisation window; green: the out-of-sample window that follows, with its Sharpe printed.
4

Multi-asset books with shared capital and caps

Test 18 assets as one portfolio, not 18 isolated curves summed together. Capital is shared, sizing is Kelly 0.25 by default, no single asset exceeds 7.5 % of the book, and slippage is charged at the portfolio level when several legs rebalance on the same bar.

  • Monthly-returns heatmap, allocation donut, per-asset contribution
  • Cross-sectional ranking strategies (long top-k / short bottom-k) built in
  • Correlation-aware sizing and a portfolio drawdown halt (−20 % default)
Portfolio · 18 assets · monthly returns %
2019–2026green above zero, red below
Figure 4. Monthly returns of the 18-asset book, years as rows. Intensity scales with magnitude.
5

Slippage, commission, funding and borrow — on by default

The single biggest reason strategies break live is that the backtest paid nothing to trade. Backcast Labs defaults to 20 bps slippage and exchange-realistic commission, and lets you model perpetual funding, borrow costs for shorts and liquidity-dependent slippage curves.

  • Slippage: fixed bps, spread-based, or volume-participation curve
  • Perpetual funding and short borrow from your data pack or your own keys
  • Sensitivity sweep: Sharpe against slippage 0–50 bps in one chart
Sensitivity · Sharpe vs slippage per side
EMA trend · 4hbps per side
Figure 5. The same strategy re-run at 0, 5, 10, 15, 20, 30 and 50 bps per side. A strategy whose Sharpe collapses before 20 bps is trading too often for its edge.
6

Survivorship-bias-free, checksummed, tick-level where available

History that only contains today's winners flatters every long strategy. Data packs include delisted symbols, apply corporate actions and normalize timestamps to UTC. Or bring your own keys: the engine pulls candles and trades straight from your exchange account and caches them locally.

  • Planned: packs from €49 one-time — no data subscription
  • Your own exchange and broker keys, plus CSV/Parquet
  • SHA-256 manifest per pack; the engine refuses tampered files
Data integrity · delisted US equities kept
Symbols retained per yearcount
Figure 6. Symbols that left the US equity universe each year and are retained in the pack. A survivor-only dataset deletes every one of these from your backtest.
7

Run an imported strategy honestly

Import a Python strategy and it becomes a notebook: entries and exits mapped onto engine orders, tunable constants promoted to param() declarations, and every read of data that was not available at that bar flagged before the first run. Then it is run again with slippage and commission on, so the optimistic version and the honest one sit side by side.

  • Repainting audit: every higher-timeframe read that would not have been available at that bar
  • Side by side: the optimistic run against the point-in-time, costs-on run
  • Every tunable constant kept as a param() so walk-forward can sweep it
Imported signal stack · side by side
Optimistic (dashed) vs realisticfour stacked filters
Figure 7. The same strategy: dashed with lookahead and zero costs, solid with point-in-time data and 25 bps of round-trip costs.
8

Export the validated configuration — no lock-in

When walk-forward and Monte Carlo pass your thresholds, export one signed config: frozen parameters, sizing, per-asset caps, drawdown halt and slippage cap, stamped with the engine build hash. Hand it to your own execution platform as JSON, CSV or a webhook payload.

  • Signed JSON config, versioned with the engine build hash
  • Kelly 0.25 sizing and the −20 % drawdown halt travel with the strategy
  • Import realised fills back in and re-measure your true slippage
Report · export → live
Stitched OOS hands over to live fillsdrift monitor from here
Figure 8. Stitched out-of-sample curve hands over to live fills at the export marker. The drift monitor compares the two from that point on.
What you get

Everything in the planned €299 licence

Planned pricing, not on sale yet. No tiers, no bar caps, no “pro” unlock. Data packs and the update contract are the only add-ons.

Table 1. Capabilities of the Backcast Labs licence against a typical cloud backtester.
CapabilityBackcast Labs €299Typical cloud backtester
Event-driven engine, next-bar fillsVaries
Bars / symbols per testUnlimited (your RAM)Capped by plan
Monte Carlo (trade / block bootstrap)✓ 1,000+ pathsAdd-on or manual
Walk-forward (anchored + rolling)
Deflated Sharpe ratio
An agent that runs the validation, not one that describes it✓ six plans, real cells
Standing instruction — say it once, it keeps checking✓ four mandates
The agent's own configurations counted against your search bar✓ logged and tagged
A record of the checks that found nothing✓ exportable as CSV
Slippage & commission on by default✓ 20 bpsOff by default
Portfolio-level slippage
Survivorship-bias-free data✓ data packsSome markets
Strategy import as a notebook✓ with repainting audit
Strategies stay on your disk✓ alwaysUploaded to the cloud
Export to live execution✓ signed config, CSV, webhookPlatform lock-in
Price after 3 years€299 planned (+ optional updates)A recurring fee every year
See it move. The notebook demo runs every method on this page on simulated data — no account, nothing installed.