Hand the validation to it.
Define the rules. It goes and finds out whether they are real.
Backcast Labs is an event-driven backtest studio that runs on your own hardware — point-in-time bars, slippage and commission charged by default, every result resampled a thousand times before anything is called an edge. Sitting inside it is a quant reviewer: it reads your result back in plain language, then goes and tests whether the number is an edge or an artefact of how hard you searched. Its most common answer is Nothing crossed, and it would rather say that than manufacture a finding. Hand it an automation researcher assignment and it goes looking on its own — on a budget in configurations that you set, and that it bills itself for.
Free notebook demo on simulated data. No account, no card, nothing installed.
Designed for the data you already have — bring your own keys, or buy history once
Say what matters once. It does the validating.
Every cell in the notebook carries the same ✦ key, and behind it is one reviewer — it translates the result into plain language, then goes and finds out whether the edge is real or an artefact of how hard you searched. Give it a standing instruction in your own words and it keeps checking on the clock and on every change you make, writing what it found into the research log whether or not it found anything. Every line below is printed by the notebook demo on its default run — simulated data · illustration only.
“Tell me when something I am running stops being an edge.”
Watching the noise bar, the deflated Sharpe, walk-forward efficiency and the drawdown halt — on the clock and on every change you make.
- 2 checks made
- 1 crossed something
- 1 nothing crossed
- 0 configurations priced
-
00:44:35 Crossed the drawdown against the −20 % halt, after you set the mandate
Worst drawdown is −23.3 %, through the halt. On this sample the system would have stopped trading and asked for a human.
What it did: Reported it and stopped. I do not move a parameter you set.
-
00:44:41 Clear the noise bar, the deflated Sharpe, walk-forward efficiency and the drawdown halt, after a scheduled check
Nothing crossed. Sharpe 1.09 against a noise bar of 0.83, deflated 76.7 %, drawdown −23.3 %, walk-forward efficiency 0.52.
What it did: Nothing. That is the answer most of the time, and I would rather say it than manufacture a finding.
Two entries out of a session. Both are exportable as CSV alongside the runs you made yourself, and either can be undone in one click — the reviewer is pausable, resumable and its mandate is changeable, all from the card above the sheet.
“Is this a real edge, or did I search until it looked good?”
- Count every configuration this session has already evaluated
49 configurations so far — 1 by hand, 48 from sweeps. Best-of-49 noise on 7.5 years is 0.83; you are holding 1.09. Deflated Sharpe 76.9 %.
- Sweep the 49 settings around yours for a plateau or a spike
A plateau. 1.09 here against a neighbouring median of 1.07, spread ±0.05 across the grid. Most robust square: lookback 60 at ATR 2.5×, worst neighbour 1.03.
- Race it against 400 random entries with the same footprint
100.0th percentile against 400 coin flips that traded 190 times and held 6 days each. Median coin flip −0.00, luckiest 0.97, buy-and-hold 0.48.
- Cut six walk-forward folds and test on data the parameters never saw
η = 0.52 across six rolling folds, 6 / 6 of them profitable out-of-sample. Borderline.
- Re-price the identical parameters at 1 bp a side
Gross 1.41 against your net 1.09, so the cost model is taking 0.31 — 22 % of what it makes before it pays to trade.
Not yet. It clears the noise bar, but not every check agrees.
confidence: medium · 4 of 5 checks passIt is above what luck would hand you, and the deflated Sharpe puts the chance it survives 49 looks at 76.9 %. What is against it: the deflated Sharpe is 76.9 %. That is a lead worth validating rather than a result worth trading, and the honest next move is data this notebook has not touched.
Its own work raises your bar, and it says so
The reviewer pays for its own looking. That last step — re-pricing the strategy at 1 bp a side — is a configuration, so it goes into the research log tagged analyst and onto the same multiple-testing counter as a slider you moved yourself. The result: the notebook that answered your question is measuring against a slightly harder bar than the one you asked it about.
Cell [7] before the run: 49 distinct configurations · deflated Sharpe 76.9 %. After it: 50 distinct configurations · deflated Sharpe 76.7 %. Autonomy is not free here, and the page that pretends otherwise is the one to distrust.
Agents that go looking on their own — and pay for every look
One control on the notebook’s standing bar opens a wizard. Six answers compose a research agent: your risk profile, whether it may go short, which timeframe, which edge families it may try, how many configurations you are willing to spend, and the house limits it cannot touch. Save it and it schedules itself, working candidate after candidate in eight visible steps. Every configuration it evaluates goes onto the same multiple-testing counter in cell [7] as a slider you moved yourself. There is no unlimited-budget option, and there will not be one. Every figure below was read off the free notebook demo on one saved assignment — simulated data · illustration only.
“Trend and mean reversion, long and short, 4h bars, inside the sceptic profile — and stop at 80 configurations.”
It rotates the families breadth-first, so it cannot spend the whole budget polishing the first thing that looked promising.
- 10 candidates built
- 72 configurations spent
- 4 cleared every check
- 3 still clear now
-
step 4 / 8 Neighbourhood the eight settings one step either side of the candidate
A plateau. 1.54 here against a neighbouring median of 1.51, worst neighbour 1.43, spread ±0.05.
Why: the setting worth shipping is the one that survives being slightly wrong, not the one that scored highest. This is the same arithmetic cell [5] draws. 8 configurations charged.
-
step 6 / 8 Folds six walk-forward folds on data the parameters never saw
η = 0.54, mean out-of-sample Sharpe 0.81 against 1.08 in sample, 6 of 6 folds profitable out-of-sample.
Why: this charges no configuration and it is not free — a fold I have tested on is no longer untouched data for the next question either of us asks. The sample is what is being spent here.
-
step 8 / 8 Report including reporting that it found nothing
“Candidate 9 is a no. It failed 3 of 6: deflated Sharpe at least 90 %; still clears the noise bar at 45 bps a side; walk-forward efficiency at least 0.60 over six folds.”
Why: this is the answer most of the time, and publishing it is the only thing that makes the rare yes worth anything.
Eight steps a candidate: build it, run it, price the cost model at 1 bp and at 45 bps, sweep the neighbourhood, race 400 random entries with its own footprint, cut six folds, resample a thousand paths, report. Pause, resume, undo the last candidate, change the assignment and stop it are one click each — and the undo says out loud that the budget is not refunded, because the configurations were still evaluated.
“What may it search, and what will you pay for it?”
- Risk profile
The same six profiles the notebook already uses — there is no second risk scale, because two scales that disagree are worse than one. Costs: each profile also prices its own configurations when it is applied; the researcher is charged separately, on top of that.
- Long only, or long and short
Costs: shorting opens the two books that are long and short by construction, so 6 families instead of 4. It adds a borrow fee and a recall risk this engine does not model, and a loss with no natural bound. It does not change the sizing.
- Timeframe, or let it decide
Costs: if it decides, it prices the candidate at 1h, 4h and 1D and picks the coarsest that still leaves 100 round trips — 13 configurations a cycle instead of 11. It reports which it chose and why, every time.
- Edge families
Trend, breakout, mean reversion, cross-sectional relative value, macro regime, filter stack — each one a real notebook in the catalogue. Costs: every extra family multiplies the search space and therefore the bar. Two families is 40 distinct settings before it repeats.
- Budget, in configurations
40, 80, 160 or 320. Costs: the counter reads 49 on a fresh notebook; 80 would take it to about 129. An agent that can search without a ceiling is a strategy-mining machine.
- What it may not do
Shown, not offered. Quarter-Kelly, the 7.5 % per-asset ceiling, the −20 % halt, the 20 bps slippage cap — and the one that matters here: evaluate a configuration without putting it on the counter in cell [7].
Then: what I will do, why, what it will cost you, and what it will never do.
nothing is priced until you press save“Your existing result does not get better while this runs. It gets harder to believe. That is not a side effect of letting an agent look — it is what letting an agent look means, and this is the only place you get to watch it happen.”
Eighty configurations of budget, seventy-two spent, and the counter moved by exactly seventy-two
Cell [7] read 49 distinct configurations when the agent started. It built ten candidates, evaluated 110 configurations — 38 of which had already been tried and therefore cost nothing new — and stopped when the next cycle needed more than the budget had left. It reports 72 spent. The counter moved from 49 to 121.
Cell [7] names the share: 1 by hand · 48 from the sweep · 72 from the researcher. Of the four candidates that cleared all six checks when they were found, three still clear them at 121 configurations. The fourth fell off the shortlist without changing at all — the counter rose underneath it — and the list says so where it used to sit rather than quietly dropping it.
Ranking a search by raw Sharpe is the mistake this product exists to prevent
Raw Sharpe is measured in sample, on the same data the setting was chosen on, and it rises with how hard you looked — a high best-of-N Sharpe is exactly what a lot of looking produces whether or not there is an edge underneath it. So the results list defaults to the mean Sharpe the six walk-forward folds printed out of sample, minus the Sharpe a no-edge strategy would have been expected to print as the best of the configurations already evaluated when that candidate turned up. A candidate found after 300 configurations carries a bigger haircut than one found after 60, and a candidate that scored well in sample and did not travel falls.
On the run above that changes the answer. By raw Sharpe the winner is the candidate that printed 2.36 — which keeps only 0.54 of its in-sample Sharpe out of sample and fails the folds. By the default measure the winner is the one that printed 1.68, found after only 5 configurations, keeping 0.84 out of sample: 1.88 − 0.84 = 1.04 against 1.81 − 0.89 = 0.91. Switch the sort to raw Sharpe and watch the order change; that change is the mistake, drawn.
Beside the list there is a shortlist of what still clears every check, and — one click away — what it got wrong: candidates that cleared the noise bar and beat 95 % of the coin flips, then broke on the folds or turned out to be a spike rather than a plateau. That 2.36 is the first entry on it. They are published, not dropped: a search that quietly discards its own near-misses is reporting a biased sample of itself.
Fourteen estimators and every dataset, each stating where it fails
Sharpe, Sortino, profit factor, Calmar, maximum drawdown and its path, skewness and kurtosis, the expected maximum Sharpe under the null, the deflated Sharpe, the Gaussian tail approximation both of those lean on, a 400-draw random-entry permutation test, walk-forward efficiency over six folds, a 1,000-path Monte Carlo resample, a maximin plateau search over the parameter neighbourhood, and the cost-model decomposition at 1 bp against 45 bps. Every one of them is computed in the notebook — nothing is listed that cannot be run and reproduced here. Each entry states what it assumes and when it fails: “Calmar’s denominator is a single observation. A longer sample almost always produces a worse one, so Calmar improves by shortening the window — the opposite of what it should reward.”
The datasets are the ones the notebook actually reads — the sample research pack and the packs in the data catalogue — each with its coverage and what it does not contain: the macro pack is “macro series, not prices — you cannot backtest an entry on it”. Every candidate in the results list names the dataset it was measured on.
An agent that always finds an opportunity is a salesman.
Four things the reviewer said in the notebook demo, quoted as it wrote them. None of them is a marketing line — each one is a column in the research log or a stamp on the panel.
“Nothing. That is the answer most of the time, and I would rather say it than manufacture a finding.”
“Reported it and stopped. I do not move a parameter you set.”
“Validation spends the one budget you cannot get back … By the time it answers you, the bar is higher than it was when you asked.”
“No. I will not go looking for a way to size this above quarter-Kelly, and I will not model leverage.”
One question, and you can stop configuring
The notebook opens by asking “What are you here to do?” Pick the sentence that fits and it sets the strategy, the parameters, the cost model, the panels and how much of the reviewing it does without being asked — then it tells you what it chose, why, and what it costs you. Five answers in the notebook; three of them here, with the cost line each one carries. There is no free answer on the list.
“I have a strategy and I want to know if it is real”
- What it chose
- The EMA trend book at its published defaults — lookback 50, ATR 2.5×, 20 bps slippage and 5 bps commission on the full 2019–2026 sample, delisted names left in.
- Why
- It is the only configuration on this page I did not choose for you. Anything I tuned would already be a search result, and the whole question you just asked is whether a number survives being searched for.
- What it costs you
- Validation spends the one budget you cannot get back: the Analyst prices six configurations to answer this, and every one of them lifts the noise bar the result has to clear. By the time it answers you, the bar is higher than it was when you asked. That is the honest arithmetic — and it is why the answer here is often no.
“I am porting something and want it validated”
- What it chose
- The imported signal stack — the weakest book bundled — at 20 bps slippage and 5 bps commission, with the standing review set to re-validate anything you adopt.
- Why
- A notebook that came in from somewhere else was tuned on a bench this engine never saw, so the first honest question is not what it returns but how wide the ground under it is. That is the sensitivity sweep, and the Analyst is going after the robust square.
- What it costs you
- This path is built to find the failure, and on an imported notebook it usually does. It also spends the sample: the folds it cuts leave less untouched data for the next question you ask, and the configurations it prices raise your own noise bar. If you would rather not know, do not pick this one.
“I just want to look around”
- What it chose
- The defaults, untouched, with only the Explainer on.
- Why
- You did not ask for an answer, so I have not gone looking for one. The eight-step tour is the fastest way to see what this notebook can be asked.
- What it costs you
- Nothing here is validated, so nothing on this screen has earned the word edge yet. What you are about to read is one draw from a distribution, and it is the flattering half of the product until you run cells [5] to [7].
Educational model output on sample markets. Sizing stays quarter-Kelly with a 7.5 % ceiling per asset and a −20 % halt whichever answer you pick.
The blueprint for a strategy you can trust
Four steps, in this order. Every one of them can fail your idea — that is the point. What comes out the other end is a frozen configuration you can hand to your execution platform.
Define the strategy
Start from one of six templates — EMA trend, cross-sectional momentum, mean-reversion pairs, macro regime switch, breakout with an ATR stop, or the imported signal stack — or write your own. Every strategy here is a notebook: one .fnb file of plain Python against the engine API, its parameter declarations and one bar function.
- Parameters are declared, so the optimizer knows what it may sweep
- pandas and numpy available inside the bar loop for research
- Seed plus config gives an identical result, every time
Set entry & exit rules
Signals computed at the close of bar t fill at the open of t+1, plus slippage — the way it happens in production. Choose the fill model, the stop, the sizing and the guards; every one of them travels with the strategy.
- Fill models: next-open, VWAP-of-bar, worst-of-bar stress, tick-level
- Kelly 0.25 sizing and a 7.5 % per-asset cap by default
- Indexing a future bar raises
LookaheadErrorinstead of flattering the curve
Backtest & validate
One equity curve is an anecdote. Backcast Labs resamples your trade sequence into 1,000 paths, walks the parameters forward over six out-of-sample folds, and deflates the Sharpe by the number of combinations you actually tried.
- Monte Carlo: P5 / P50 / P95 and “probability of a −20 % drawdown within 12 months: 2.2 %”
- Walk-forward: anchored or rolling, with the efficiency ratio η
- Deflated Sharpe counts every trial, including the ones the reviewer ran itself
Export to your execution platform
When the numbers hold up, export one signed JSON file: frozen parameters, sizing, per-asset caps, drawdown halt and slippage cap, stamped with the engine build hash. Feed it to your own runner over CSV, JSON or a webhook — nothing is locked in.
- Signed
.forge-live.jsonwith the validation summary attached - Export stays disabled below η = 0.5 until you acknowledge it, in writing
- Import realised fills back in to re-measure your true slippage
Run the whole blueprint in the notebook Read the engine documentation
Five ways to break your own strategy first
Each of these is one panel in the studio, and each of them runs in the free notebook demo on simulated data.
Read the drawdown before the equity curve
Point-in-time bars, next-bar fills and an honest drawdown series with your halt drawn on it. If the deepest drawdown breaches the halt, the strategy would have been switched off live — exactly where the backtest shows it recovering.
Find out what trading actually costs you
Slippage and commission are on by default. Sweep 0 to 50 bps per side in one chart and see where your Sharpe crosses 1.0. A strategy that collapses before 20 bps is trading too often for its edge.
Resample your luck a thousand times
Trade-level, block and stationary bootstrap, all drawdown-halt aware: when a path would have tripped the guard, it stops trading, the way production would. Read the P5, not the P50.
Prove the parameters on data they never saw
Anchored or rolling, three to twelve folds, with a per-fold parameter-drift table and one stitched out-of-sample curve — the only curve worth comparing with live performance.
Test the whole book, not one symbol
Eighteen assets with shared capital, correlation-aware sizing, a 7.5 % per-asset cap and slippage charged at the portfolio level when several legs rebalance on the same bar.
Include the symbols that died
History that only holds today's winners flatters every long strategy. Data packs keep 12,800 delisted US equities and seven dead crypto perps, apply corporate actions and ship a signed SHA-256 manifest.
Learn the method, not the marketing
Three tracks written by the people who build the engine. No signals, no “X % per month”, and every lesson ends in a setting you can change yourself.
Validation 101
What a backtest can and cannot tell you, why pessimistic defaults exist, and the order in which to read a report.
- Why the drawdown panel comes first
- Repainting and lookahead, explained
- Reading a Monte Carlo fan honestly
Statistics for traders
Sharpe, Sortino, deflated Sharpe and the efficiency ratio — what each one is actually measuring, and how each one lies.
- Selection bias and the number of trials
- Bootstraps: plain, block and stationary
- Deflated Sharpe from first principles
Building a research process
From an idea to a frozen configuration: how to keep a research log, how many parameters are too many, and when to stop.
- Designing a walk-forward that is not theatre
- Portfolio-level costs and correlated legs
- The export checklist, line by line
Questions system developers ask first
Does Backcast Labs repaint or look ahead?
What data will I be able to use?
Is there a bar or symbol limit?
Can I run the validated strategy live?
What can I actually hand to the reviewer?
Does the reviewer's own work count against my result?
What happens if I stop paying for updates?
How honest are the numbers on this page?
Start with the notebook demo
The notebook demo runs the entire studio on simulated data — parameters, Monte Carlo, walk-forward, portfolio — with no account and nothing installed. Backcast Labs itself is not on sale yet.
- No card, no trial clock, no drip campaign
- Every method from this page, running in your browser
- Planned pricing: €299 once, updates optional
Join early access
Send us a short e-mail and we will let you know when a place opens. Nothing is charged and no account is created. One Mast account will give access to all our products.
Or open the notebook demo right now.