Early access Backcast Labs is not on sale yet. The research notebook runs in your browser on simulated data. Join early access

Hand the validation to it.

Define the rules. It goes and finds out whether they are real.

Backcast Labs is an event-driven backtest studio that runs on your own hardware — point-in-time bars, slippage and commission charged by default, every result resampled a thousand times before anything is called an edge. Sitting inside it is a quant reviewer: it reads your result back in plain language, then goes and tests whether the number is an edge or an artefact of how hard you searched. Its most common answer is Nothing crossed, and it would rather say that than manufacture a finding. Hand it an automation researcher assignment and it goes looking on its own — on a budget in configurations that you set, and that it bills itself for.

Free notebook demo on simulated data. No account, no card, nothing installed.

Backcast Labs · EMA trend · BTC·ETH·SOL · 4h · 2019–2026
CAGR 14.2 %Sharpe 1.09Max DD −23.3 %η 0.52Trades 190
Equity vs buy-and-holdslippage 20 bps · commission 5 bps
Drawdownhalt −20 %
Point-in-time engine · next-bar fillsSimulated data · illustration only

Designed for the data you already have — bring your own keys, or buy history once

Crypto exchange keysBroker keysCSV / ParquetPublic macro seriesOne-time data packs
The quant reviewer

Say what matters once. It does the validating.

Every cell in the notebook carries the same key, and behind it is one reviewer — it translates the result into plain language, then goes and finds out whether the edge is real or an artefact of how hard you searched. Give it a standing instruction in your own words and it keeps checking on the clock and on every change you make, writing what it found into the research log whether or not it found anything. Every line below is printed by the notebook demo on its default run — simulated data · illustration only.

Standing instruction · said once

“Tell me when something I am running stops being an edge.”

Watching the noise bar, the deflated Sharpe, walk-forward efficiency and the drawdown halt — on the clock and on every change you make.

  • 2 checks made
  • 1 crossed something
  • 1 nothing crossed
  • 0 configurations priced
  • 00:44:35 Crossed the drawdown against the −20 % halt, after you set the mandate

    Worst drawdown is −23.3 %, through the halt. On this sample the system would have stopped trading and asked for a human.

    What it did: Reported it and stopped. I do not move a parameter you set.

  • 00:44:41 Clear the noise bar, the deflated Sharpe, walk-forward efficiency and the drawdown halt, after a scheduled check

    Nothing crossed. Sharpe 1.09 against a noise bar of 0.83, deflated 76.7 %, drawdown −23.3 %, walk-forward efficiency 0.52.

    What it did: Nothing. That is the answer most of the time, and I would rather say it than manufacture a finding.

Two entries out of a session. Both are exportable as CSV alongside the runs you made yourself, and either can be undone in one click — the reviewer is pausable, resumable and its mandate is changeable, all from the card above the sheet.

You asked it · it went and did the work

“Is this a real edge, or did I search until it looked good?”

  • Count every configuration this session has already evaluated

    49 configurations so far — 1 by hand, 48 from sweeps. Best-of-49 noise on 7.5 years is 0.83; you are holding 1.09. Deflated Sharpe 76.9 %.

  • Sweep the 49 settings around yours for a plateau or a spike

    A plateau. 1.09 here against a neighbouring median of 1.07, spread ±0.05 across the grid. Most robust square: lookback 60 at ATR 2.5×, worst neighbour 1.03.

  • Race it against 400 random entries with the same footprint

    100.0th percentile against 400 coin flips that traded 190 times and held 6 days each. Median coin flip −0.00, luckiest 0.97, buy-and-hold 0.48.

  • Cut six walk-forward folds and test on data the parameters never saw

    η = 0.52 across six rolling folds, 6 / 6 of them profitable out-of-sample. Borderline.

  • Re-price the identical parameters at 1 bp a side

    Gross 1.41 against your net 1.09, so the cost model is taking 0.31 — 22 % of what it makes before it pays to trade.

Not yet. It clears the noise bar, but not every check agrees.

confidence: medium · 4 of 5 checks pass

It is above what luck would hand you, and the deflated Sharpe puts the chance it survives 49 looks at 76.9 %. What is against it: the deflated Sharpe is 76.9 %. That is a lead worth validating rather than a result worth trading, and the honest next move is data this notebook has not touched.

The part nobody else prints

Its own work raises your bar, and it says so

The reviewer pays for its own looking. That last step — re-pricing the strategy at 1 bp a side — is a configuration, so it goes into the research log tagged analyst and onto the same multiple-testing counter as a slider you moved yourself. The result: the notebook that answered your question is measuring against a slightly harder bar than the one you asked it about.

Configurations counted4950
Deflated Sharpe76.9 %76.7 %
Noise bar to clear0.83
Your Sharpe, unchanged1.09

Cell [7] before the run: 49 distinct configurations · deflated Sharpe 76.9 %. After it: 50 distinct configurations · deflated Sharpe 76.7 %. Autonomy is not free here, and the page that pretends otherwise is the one to distrust.

AI Automation Researcher

Agents that go looking on their own — and pay for every look

One control on the notebook’s standing bar opens a wizard. Six answers compose a research agent: your risk profile, whether it may go short, which timeframe, which edge families it may try, how many configurations you are willing to spend, and the house limits it cannot touch. Save it and it schedules itself, working candidate after candidate in eight visible steps. Every configuration it evaluates goes onto the same multiple-testing counter in cell [7] as a slider you moved yourself. There is no unlimited-budget option, and there will not be one. Every figure below was read off the free notebook demo on one saved assignment — simulated data · illustration only.

One saved assignment

“Trend and mean reversion, long and short, 4h bars, inside the sceptic profile — and stop at 80 configurations.”

It rotates the families breadth-first, so it cannot spend the whole budget polishing the first thing that looked promising.

  • 10 candidates built
  • 72 configurations spent
  • 4 cleared every check
  • 3 still clear now
  • step 4 / 8 Neighbourhood the eight settings one step either side of the candidate

    A plateau. 1.54 here against a neighbouring median of 1.51, worst neighbour 1.43, spread ±0.05.

    Why: the setting worth shipping is the one that survives being slightly wrong, not the one that scored highest. This is the same arithmetic cell [5] draws. 8 configurations charged.

  • step 6 / 8 Folds six walk-forward folds on data the parameters never saw

    η = 0.54, mean out-of-sample Sharpe 0.81 against 1.08 in sample, 6 of 6 folds profitable out-of-sample.

    Why: this charges no configuration and it is not free — a fold I have tested on is no longer untouched data for the next question either of us asks. The sample is what is being spent here.

  • step 8 / 8 Report including reporting that it found nothing

    “Candidate 9 is a no. It failed 3 of 6: deflated Sharpe at least 90 %; still clears the noise bar at 45 bps a side; walk-forward efficiency at least 0.60 over six folds.”

    Why: this is the answer most of the time, and publishing it is the only thing that makes the rare yes worth anything.

Eight steps a candidate: build it, run it, price the cost model at 1 bp and at 45 bps, sweep the neighbourhood, race 400 random entries with its own footprint, cut six folds, resample a thousand paths, report. Pause, resume, undo the last candidate, change the assignment and stop it are one click each — and the undo says out loud that the budget is not refunded, because the configurations were still evaluated.

The wizard · six questions, each one priced

“What may it search, and what will you pay for it?”

  • Risk profile

    The same six profiles the notebook already uses — there is no second risk scale, because two scales that disagree are worse than one. Costs: each profile also prices its own configurations when it is applied; the researcher is charged separately, on top of that.

  • Long only, or long and short

    Costs: shorting opens the two books that are long and short by construction, so 6 families instead of 4. It adds a borrow fee and a recall risk this engine does not model, and a loss with no natural bound. It does not change the sizing.

  • Timeframe, or let it decide

    Costs: if it decides, it prices the candidate at 1h, 4h and 1D and picks the coarsest that still leaves 100 round trips — 13 configurations a cycle instead of 11. It reports which it chose and why, every time.

  • Edge families

    Trend, breakout, mean reversion, cross-sectional relative value, macro regime, filter stack — each one a real notebook in the catalogue. Costs: every extra family multiplies the search space and therefore the bar. Two families is 40 distinct settings before it repeats.

  • Budget, in configurations

    40, 80, 160 or 320. Costs: the counter reads 49 on a fresh notebook; 80 would take it to about 129. An agent that can search without a ceiling is a strategy-mining machine.

  • What it may not do

    Shown, not offered. Quarter-Kelly, the 7.5 % per-asset ceiling, the −20 % halt, the 20 bps slippage cap — and the one that matters here: evaluate a configuration without putting it on the counter in cell [7].

Then: what I will do, why, what it will cost you, and what it will never do.

nothing is priced until you press save

“Your existing result does not get better while this runs. It gets harder to believe. That is not a side effect of letting an agent look — it is what letting an agent look means, and this is the only place you get to watch it happen.”

The bill, on that assignment

Eighty configurations of budget, seventy-two spent, and the counter moved by exactly seventy-two

Cell [7] read 49 distinct configurations when the agent started. It built ten candidates, evaluated 110 configurations — 38 of which had already been tried and therefore cost nothing new — and stopped when the next cycle needed more than the budget had left. It reports 72 spent. The counter moved from 49 to 121.

Configurations counted49121
Deflated Sharpe of your own result76.9 %65.8 %
Noise bar it must clear0.830.95
Your Sharpe, unchanged1.09

Cell [7] names the share: 1 by hand · 48 from the sweep · 72 from the researcher. Of the four candidates that cleared all six checks when they were found, three still clear them at 121 configurations. The fourth fell off the shortlist without changing at all — the counter rose underneath it — and the list says so where it used to sit rather than quietly dropping it.

How the results are ranked

Ranking a search by raw Sharpe is the mistake this product exists to prevent

Raw Sharpe is measured in sample, on the same data the setting was chosen on, and it rises with how hard you looked — a high best-of-N Sharpe is exactly what a lot of looking produces whether or not there is an edge underneath it. So the results list defaults to the mean Sharpe the six walk-forward folds printed out of sample, minus the Sharpe a no-edge strategy would have been expected to print as the best of the configurations already evaluated when that candidate turned up. A candidate found after 300 configurations carries a bigger haircut than one found after 60, and a candidate that scored well in sample and did not travel falls.

On the run above that changes the answer. By raw Sharpe the winner is the candidate that printed 2.36 — which keeps only 0.54 of its in-sample Sharpe out of sample and fails the folds. By the default measure the winner is the one that printed 1.68, found after only 5 configurations, keeping 0.84 out of sample: 1.88 − 0.84 = 1.04 against 1.81 − 0.89 = 0.91. Switch the sort to raw Sharpe and watch the order change; that change is the mistake, drawn.

Beside the list there is a shortlist of what still clears every check, and — one click away — what it got wrong: candidates that cleared the noise bar and beat 95 % of the coin flips, then broke on the folds or turned out to be a spike rather than a plateau. That 2.36 is the first entry on it. They are published, not dropped: a search that quietly discards its own near-misses is reporting a biased sample of itself.

What it draws on

Fourteen estimators and every dataset, each stating where it fails

Sharpe, Sortino, profit factor, Calmar, maximum drawdown and its path, skewness and kurtosis, the expected maximum Sharpe under the null, the deflated Sharpe, the Gaussian tail approximation both of those lean on, a 400-draw random-entry permutation test, walk-forward efficiency over six folds, a 1,000-path Monte Carlo resample, a maximin plateau search over the parameter neighbourhood, and the cost-model decomposition at 1 bp against 45 bps. Every one of them is computed in the notebook — nothing is listed that cannot be run and reproduced here. Each entry states what it assumes and when it fails: “Calmar’s denominator is a single observation. A longer sample almost always produces a worse one, so Calmar improves by shortening the window — the opposite of what it should reward.”

The datasets are the ones the notebook actually reads — the sample research pack and the packs in the data catalogue — each with its coverage and what it does not contain: the macro pack is “macro series, not prices — you cannot backtest an entry on it”. Every candidate in the results list names the dataset it was measured on.

Configure nothing

One question, and you can stop configuring

The notebook opens by asking “What are you here to do?” Pick the sentence that fits and it sets the strategy, the parameters, the cost model, the panels and how much of the reviewing it does without being asked — then it tells you what it chose, why, and what it costs you. Five answers in the notebook; three of them here, with the cost line each one carries. There is no free answer on the list.

1

“I have a strategy and I want to know if it is real”

What it chose
The EMA trend book at its published defaults — lookback 50, ATR 2.5×, 20 bps slippage and 5 bps commission on the full 2019–2026 sample, delisted names left in.
Why
It is the only configuration on this page I did not choose for you. Anything I tuned would already be a search result, and the whole question you just asked is whether a number survives being searched for.
What it costs you
Validation spends the one budget you cannot get back: the Analyst prices six configurations to answer this, and every one of them lifts the noise bar the result has to clear. By the time it answers you, the bar is higher than it was when you asked. That is the honest arithmetic — and it is why the answer here is often no.
3

“I am porting something and want it validated”

What it chose
The imported signal stack — the weakest book bundled — at 20 bps slippage and 5 bps commission, with the standing review set to re-validate anything you adopt.
Why
A notebook that came in from somewhere else was tuned on a bench this engine never saw, so the first honest question is not what it returns but how wide the ground under it is. That is the sensitivity sweep, and the Analyst is going after the robust square.
What it costs you
This path is built to find the failure, and on an imported notebook it usually does. It also spends the sample: the folds it cuts leave less untouched data for the next question you ask, and the configurations it prices raise your own noise bar. If you would rather not know, do not pick this one.
5

“I just want to look around”

What it chose
The defaults, untouched, with only the Explainer on.
Why
You did not ask for an answer, so I have not gone looking for one. The eight-step tour is the fastest way to see what this notebook can be asked.
What it costs you
Nothing here is validated, so nothing on this screen has earned the word edge yet. What you are about to read is one draw from a distribution, and it is the flattering half of the product until you run cells [5] to [7].

Educational model output on sample markets. Sizing stays quarter-Kelly with a 7.5 % ceiling per asset and a −20 % halt whichever answer you pick.

How it works

The blueprint for a strategy you can trust

Four steps, in this order. Every one of them can fail your idea — that is the point. What comes out the other end is a frozen configuration you can hand to your execution platform.

1

Define the strategy

Start from one of six templates — EMA trend, cross-sectional momentum, mean-reversion pairs, macro regime switch, breakout with an ATR stop, or the imported signal stack — or write your own. Every strategy here is a notebook: one .fnb file of plain Python against the engine API, its parameter declarations and one bar function.

  • Parameters are declared, so the optimizer knows what it may sweep
  • pandas and numpy available inside the bar loop for research
  • Seed plus config gives an identical result, every time
Strategies · EmaTrend
BTCUSD · 4hEMA 20 / 50
2

Set entry & exit rules

Signals computed at the close of bar t fill at the open of t+1, plus slippage — the way it happens in production. Choose the fill model, the stop, the sizing and the guards; every one of them travels with the strategy.

  • Fill models: next-open, VWAP-of-bar, worst-of-bar stress, tick-level
  • Kelly 0.25 sizing and a 7.5 % per-asset cap by default
  • Indexing a future bar raises LookaheadError instead of flattering the curve
Engine trace · bar 14,201
Signal at close, fill at next open+20 bps
3

Backtest & validate

One equity curve is an anecdote. Backcast Labs resamples your trade sequence into 1,000 paths, walks the parameters forward over six out-of-sample folds, and deflates the Sharpe by the number of combinations you actually tried.

  • Monte Carlo: P5 / P50 / P95 and “probability of a −20 % drawdown within 12 months: 2.2 %”
  • Walk-forward: anchored or rolling, with the efficiency ratio η
  • Deflated Sharpe counts every trial, including the ones the reviewer ran itself
Monte Carlo · 1,000 paths · block bootstrap
Resampled trade sequences · 252 daysblock size 20
4

Export to your execution platform

When the numbers hold up, export one signed JSON file: frozen parameters, sizing, per-asset caps, drawdown halt and slippage cap, stamped with the engine build hash. Feed it to your own runner over CSV, JSON or a webhook — nothing is locked in.

  • Signed .forge-live.json with the validation summary attached
  • Export stays disabled below η = 0.5 until you acknowledge it, in writing
  • Import realised fills back in to re-measure your true slippage
Report · export → live
Stitched out-of-sample curvefrozen at export
What you get

Five ways to break your own strategy first

Each of these is one panel in the studio, and each of them runs in the free notebook demo on simulated data.

Engine

Read the drawdown before the equity curve

Point-in-time bars, next-bar fills and an honest drawdown series with your halt drawn on it. If the deepest drawdown breaches the halt, the strategy would have been switched off live — exactly where the backtest shows it recovering.

Costs

Find out what trading actually costs you

Slippage and commission are on by default. Sweep 0 to 50 bps per side in one chart and see where your Sharpe crosses 1.0. A strategy that collapses before 20 bps is trading too often for its edge.

Monte Carlo

Resample your luck a thousand times

Trade-level, block and stationary bootstrap, all drawdown-halt aware: when a path would have tripped the guard, it stops trading, the way production would. Read the P5, not the P50.

Walk-forward

Prove the parameters on data they never saw

Anchored or rolling, three to twelve folds, with a per-fold parameter-drift table and one stitched out-of-sample curve — the only curve worth comparing with live performance.

Portfolio

Test the whole book, not one symbol

Eighteen assets with shared capital, correlation-aware sizing, a 7.5 % per-asset cap and slippage charged at the portfolio level when several legs rebalance on the same bar.

Data

Include the symbols that died

History that only holds today's winners flatters every long strategy. Data packs keep 12,800 delisted US equities and seven dead crypto perps, apply corporate actions and ship a signed SHA-256 manifest.

Education

Learn the method, not the marketing

Three tracks written by the people who build the engine. No signals, no “X % per month”, and every lesson ends in a setting you can change yourself.

Beginner

Validation 101

8 lessons2 h 10 m

What a backtest can and cannot tell you, why pessimistic defaults exist, and the order in which to read a report.

  1. Why the drawdown panel comes first
  2. Repainting and lookahead, explained
  3. Reading a Monte Carlo fan honestly
Start the track
Intermediate

Statistics for traders

6 lessons1 h 45 m

Sharpe, Sortino, deflated Sharpe and the efficiency ratio — what each one is actually measuring, and how each one lies.

  1. Selection bias and the number of trials
  2. Bootstraps: plain, block and stationary
  3. Deflated Sharpe from first principles
Start the track
Advanced

Building a research process

7 lessons2 h 30 m

From an idea to a frozen configuration: how to keep a research log, how many parameters are too many, and when to stop.

  1. Designing a walk-forward that is not theatre
  2. Portfolio-level costs and correlated legs
  3. The export checklist, line by line
Start the track
FAQ

Questions system developers ask first

Does Backcast Labs repaint or look ahead?
No. The engine is event-driven: on each bar it hands your strategy only the data available at that bar's close, and orders fill at the next bar's open (or at a configurable intrabar model with slippage). Indicators that reference future bars raise an error instead of silently repainting.
What data will I be able to use?
Three sources are planned: your own exchange or broker API keys (BYOK), CSV/Parquet import, or one-time data packs that are survivorship-bias-free and checksummed. Keys are encrypted on disk and never uploaded.
Is there a bar or symbol limit?
No. Limits are your RAM and CPU. A six-year, 18-asset, one-hour test runs in about four seconds on a laptop; tick-level crypto tests use the Parquet streaming path and stay under a few GB of memory.
Can I run the validated strategy live?
Backcast Labs validates; your own execution platform trades. “Export” hands over the frozen parameters, sizing (Kelly 0.25 by default), per-asset caps and the drawdown halt as one signed config file, and the same data goes out as CSV, JSON or a webhook payload if you prefer to wire it yourself.
What can I actually hand to the reviewer?
Validation. It takes a standing instruction in one sentence — “tell me when something I am running stops being an edge”, “keep watching my search count”, “re-validate anything I adopt”, “only speak when something actually breaks” — and after that it re-checks on the clock and on every change you make. Send it a question and it runs the cells itself: the 49-square sweep, 400 matched random entries, six walk-forward folds and a cost re-price, then rules yes, not yet or no. What it will not do is move a parameter you set: it reports and stops. Pause, resume, undo the last thing it did, or change the mandate — one click each.
Does the reviewer's own work count against my result?
Yes, and it says so. Every configuration it prices is written into the research log tagged with who ran it, and counted on the same multiple-testing bar in cell [7] as one you set by hand. On the default guest run, asking “is this a real edge?” takes the session from 49 configurations to 50 and the deflated Sharpe from 76.9 % to 76.7 %. That is not a bug to hide — a validator whose own searching is free is a validator that is lying to you about the cost of searching.
What happens if I stop paying for updates?
The plan: nothing breaks. You keep the software at the last version released while your update contract was active, forever, on up to three devices. Resume later and you jump to the current release.
How honest are the numbers on this page?
Every chart on this site is drawn in your browser from seeded sample series and illustrates a method, not a result. The engine draws the same figures from your own history once you install it. Backtested performance is hypothetical — see Legal & risk.