AlphaAssay $ test my signal

Failure codes

Failure codes are the machine-readable half of every verdict — the part your agent branches on. They follow one pattern: died_at names the stage, failure_codes name the causes.

What does died_at tell you?

died_at names the first of the eleven graded stages a signal failed — most deaths happen at net_edge (nothing survives realistic costs), family_deflation (the family tried too often), placebo (random twins did as well) or capacity (the edge does not survive its own trading size). The full mapping, in trial order, is below.

valuethe stage asks…failing there means, in plain terms
net_edgedoes anything remain after realistic costs, spread and delay?after real-world costs there is nothing left to trade
funding_edgedoes the edge survive the perpetual funding leg, charged against every holding period?the profit lived in a cost the exchange would actually have collected — or IS the funding income, a premium any holder collects, not timing
family_deflationhow often has this family tried before, and does the score survive the deflation?this idea has been tried so often that a result this good is expected by luck alone
power_honestycan a record this short even carry a claim this size, at this trial count?too little history for how big the claim is and how hard it was searched — more history, not more conviction, is the only cure
significancedo the bootstrap, bootstrap-t and Newey–West brackets around the mean exclude zero?on resampled versions of this very record, no-edge remains a perfectly plausible reading
cpcvdoes the edge hold across purged combinatorial time-partitions?the performance hangs on a few lucky segments of history, not on the whole record
walk_forwarddoes the in-sample fit carry into anchored out-of-sample folds?whatever the fit found does not carry forward even one step
concentrationdoes the book survive without its single best bars?the entire edge is a handful of jackpot bars — a lottery ticket, not a repeatable process
placebodoes the timing beat 500 random twins with the same trading profile?random look-alikes did just as well — the timing added nothing
capacitydoes the edge survive its own market impact at realistic size?the edge is only real as long as nobody trades it at size
graveyard_priorwhat does the anonymised graveyard already know about this family?informational — it attaches the family's published prior mortality to sharpen the story; it never kills on its own

And when the data cannot support a judgment either way, the verdict itself is insufficient_evidence — in plain terms: not enough history to tell skill from luck, an honest „cannot know" instead of a guess.

Which failure codes will I meet most often?

Two codes dominate real signals: no_net_edge means nothing is left once realistic costs, spread and delay are charged — the signal never had a net-of-cost edge — and deflated_out_at_n=N means the strategy family has used up its honest tries, deflated out after N effective trials. Both, with what they mean for you, are below.

codein plain terms (register wording)what it usually means for you
no_net_edgeafter realistic trading costs the strategy loses money — there is no edge left to test, so every deeper question is mootthe edge was never net-of-cost — a whipsaw or an in-sample artefact; costs killed it before any deeper test
deflated_out_at_n=Nthe Sharpe looks good only because many variants were tried: after deflating for the family's cumulative trial count, the result is indistinguishable from the best of random noisestop tweaking; a new variant cannot be distinguished from luck anymore

Which codes did the battery learn most recently?

The register grows as the battery does. The plain-terms sentences below are quoted verbatim from the machine-readable register — the same text your agent receives — so this page and the API can never tell two stories.

codein plain terms (register wording)
BACKTEST_TOO_SHORT_FOR_Nthe backtest is too short for how many variants were searched: at this trial count, a Sharpe this size is expectable from selection alone, so the sample cannot support the claim
TRACK_RECORD_TOO_SHORTthe live record is shorter than the minimum needed to support a Sharpe of this size at the stated confidence — more history, not more conviction, is the only cure
TRIALS_UNDISCLOSEDno trial count was declared, so the verdict had to assume a single try. An undeclared search history is the oldest way to smuggle overfitting past a test — declare how many variants were tried
ALPHA_INDISTINGUISHABLE_FROM_BETAthe return series is shaped like the benchmark and the residual alpha is statistically indistinguishable from zero — the claimed edge is consistent with plain market exposure: beta, not alpha
CI_CONTAINS_ZEROthe bootstrap confidence interval of the mean return includes zero: on resampled versions of this very track record, no-edge is a perfectly plausible reading
WINRATE_NOT_SIGNIFICANTthe win rate is statistically indistinguishable from a coin flip — the profit rests on a few large outcomes, not on repeatable hit-rate skill
CAPACITY_MODEL_RANGE_EXCEEDEDthe requested size exceeds the range where the impact model is empirically valid — we refuse to invent a cost number out there, so the level fails closed instead of passing on a guess
edge_concentration_extremewithout the single best 1% of bars the strategy loses money — the entire edge is a handful of jackpot bars, a lottery ticket luck hands out once, not a repeatable process
cpcv_unstabletoo many purged combinatorial time-partitions see no positive Sharpe at all — the performance hangs on a few lucky segments of history instead of being a property of the whole record
cpcv_partitionthe adversarial partition attack recombines history into purged combinatorial segments; an edge that disappears in too many of them lives in specific time slices, not in the strategy
wf_oos_negativeacross the standard walk-forward folds the out-of-sample halves lose money in aggregate — whatever the in-sample fit found does not carry forward even one step
synthetic_null_indistinguishablehalf or more of the synthetic no-edge worlds — stochastic volatility, jump diffusion, symmetric drift bursts, all built with zero exploitable signal — perform as well as the strategy
RESULT_NOT_REPRODUCEDthe claimed headline number cannot be recomputed from the submitted trades and candles within the disclosed tolerance — the result does not follow from the underlying trades
FILL_OUTSIDE_BAR_RANGEat least one fill price lies outside the low–high range of the bar it was supposedly executed in — that trade was never physically available on the submitted data
FILL_AMBIGUITY_MATERIALstop and limit both sit inside the same bar, so OHLC data cannot say which fired first; priced worst-case, the book flips from profit to loss — the claimed profit rests on an unknowable fill order
NO_SURVIVORS_AT_FWERnone of the submitted variants survives family-wise error control: under stepwise multiple testing with the family's own error budget, every variant is statistically indistinguishable from noise
PBO_HIGHthe configuration that wins in-sample ranks at or below the out-of-sample median in at least half of all combinatorial train/test splits — the selection process is overfit, so the winner's edge is a property of the search, not of the market
missing_source_rightscaller-supplied evidence attests the usage rights to the underlying data are missing or unclear — a verdict computed on data the submitter may not use is a liability, not evidence, so the outcome is withheld until the data owner decides
funding_erases_edgethe backtest ignored the funding leg of a perpetual position: charging the supplied funding rates against every settled holding bar erases the entire net edge — the profit lives only in a cost the exchange would actually have collected
edge_is_funding_carrywithout the funding income there is no edge: the price leg of this strategy loses money and the apparent profit IS the funding carry — a premium available to any holder of the position, crash-prone and no evidence of timing skill
VAR_BREACH_RATE_EXCESSreality breached the VaR forecast far more often than the claimed tail level permits — the breach count sits in the red zone of the exact binomial Basel traffic light, so the model understates risk rather than measuring it
ES_TAIL_UNDERSTATEDon breach days the realised losses run systematically deeper than the promised expected shortfall — the joint (VaR, ES) e-process crossed Ville's anytime-valid 1% line, evidence the tail size is understated, not merely unlucky
CONFORMAL_COVERAGE_SHORTFALLthe prediction intervals miss the realised outcomes far more often than the claimed confidence level permits — the miss count sits in the red zone of the exact distribution the claim implies, so the stated coverage is a label, not a property of the model
RISK_FORECAST_DOMINATED_BY_BENCHMARKthe named benchmark forecasts predict the tail demonstrably better than the submitted risk model — Diebold-Mariano on a strictly consistent loss puts the benchmark ahead past the one-sided 5% line. A risk model that loses to its own naive benchmark is sophistication theatre, not risk measurement

Not everything the battery learns becomes a kill. The sharper confidence brackets from the significance battery announce themselves as advisories first — studentized_ci_contains_zero and hac_ci_contains_zero flag that the claimed precision was borrowed from optimistic assumptions, without flipping the verdict — because a new cell earns the right to kill on calibration data, not on enthusiasm. The same discipline applies on the risk side (var_breach_rate_sparse: a model that overstates its own risk is conservative, never a kill; risk_forecast_lags_benchmark: the naive benchmark scores better but not yet decisively — the modelling edge is unproven, not dead).

And before any of this runs, assay_preflight (free) lints the payload's form in the same vocabulary — dsl_invalid, timestamps_not_monotonic, ohlcv_non_finite_or_non_positive, ohlcv_series_misaligned, symbol_data_missing, trades_rows_invalid, sample_below_engine_floor — so no paid check dies of a typo.

Two of these have their own explainers with calculators: Minimum Backtest Length (the arithmetic behind BACKTEST_TOO_SHORT_FOR_N) and break-even AUM (the model whose honesty limit is CAPACITY_MODEL_RANGE_EXCEEDED).

What is the survival map?

When a signal passes the gauntlet, the forensics report runs eight adversarial attacks — execution_lag, cost_stress, time_jackknife, regime_split, parameter_neighbourhood, cpcv_partition, synthetic_null, drift_burst — each of which the signal survives or is killed by. A signal that survives overall but is killed by regime_split is telling you where its risk hides. (The free demo returns the gauntlet gradient; the survival map is part of the forensics path.)

The response is self-describing: every code your agent encounters arrives together with its human-readable diagnosis in the same document, and every failed verdict additionally carries died_at_plain — the killer, named in one plain sentence, straight from the register quoted above. Each code also maps to a manual check in the backtest overfitting checklist, so you can run the same diagnosis by hand before you ever call the API.