# AlphaAssay — full site text for LLMs ## Signal Validation Benchmark 2026 — AlphaAssay URL: https://alphaassay.com/benchmark Research / Benchmark RESEARCH · RECORD The Signal Validation Benchmark ALPHAASSAY RESEARCH · THE PUBLIC RECORD · REPRODUCIBLE This page publishes reproducibility evidence for selected AlphaAssay outputs — every benchmark row is reproducible by anyone, free, today. A validator you cannot test is just another promise, so we publish the tests: known-answer specimens whose stable semantic fields any agent can assert, our current calibration population/status disclosure, and — as the public ledger accrues — mortality statistics of real strategy families. Methodology audits, not investment advice; no strategy identities are ever published. Known-answer verification (reproducible now) Four specimens with planted properties; the battery must catch the flaws, pass the clean one and abstain on thin data. Replay any row: quickstart . specimen planted property expected verdict what it proves golden_clean genuine persistent edge ✓ pass we don't cry wolf golden_lookahead a whipsaw with no net-of-cost edge ✕ fail · net_edge · no_net_edge we catch a costless edge golden_cherry best-of-many cherry-pick ✕ fail · family_deflation · deflated_out_at_n=50 + BACKTEST_TOO_SHORT_FOR_N=50 we price the search golden_thin too little data to judge insufficient_evidence we admit uncertainty Canonical list, with the full response format each specimen returns: golden specimens . Field result: 9 popular public strategies, full battery We ran nine of the most popular public strategy classics through the complete battery on real market data with real costs. Aggregate result: 0× pass · 2× conditional · 7× fail — the six 1-hour classics lose money after costs (Sharpe −3 to −12). Per-strategy autopsies with named setups and failure codes are not part of the published record yet; this page's changelog records what is. The 18,000-rule benchmark (reproducible, envelope status explicit) The largest run we publish: every classic technical rule in the Sullivan–Timmermann–White universe — exactly 18,000 rules — against seven crypto bluechips, 126,000 rule–market pairs through a benchmark-specific three-stage funnel using shared statistical primitives: net edge plus universe deflation, CPCV plus concentration, then placebo, with hold disclosure. It is not the paid gauntlet, omits its other stages, and writes no tenant trial ledger. 245 survive this benchmark funnel (0.19%); only 41 beat simply holding. Costs alone remove 12,274 pairs, and the deflation stage removes 113,462 more — about 90% die from the size of the search, not from anything the market did. It is labelled a benchmark, not organic user history; aggregated so no winning rule is published (only mortality), with a mandatory honest-limits block. Pull the document — engine version, funnel and limits included — from GET /v1/public/benchmark . When the standard envelope reports signed:true , an offline raw-signature check can detect changed signed bytes. Full platform trust additionally requires an independently pinned root, the signed trust bundle and complete key/revocation history; without a configured platform key the endpoint reports the unsigned fallback explicitly. The long-form story behind these numbers — what died at which gate, and why the survivors deserve the fine print more than the dead — is the field guide Was your trading edge ever real? Current public status evidence (pull it yourself) what endpoint calibration v0 — bucketed mature-registration count plus accumulating/insufficient-history state, not an outcome score GET /v1/public/calibration graveyard digest — anonymised mortality of strategy families GET /v1/public/graveyard-digest How signed snapshots work, what retained copies can detect, and the remaining trust limits: how we grade ourselves . Methodology A fixed sequence of four gate families, unfolding into eleven graded stages (net edge → family deflation → placebo vs. 500 matched twins → robustness attacks). Pure/read results are reproducible only with an explicit as_of ; stateful calls expose their timestamp and replay a stored response only under the documented idempotency contract. A separately issued certificate carries the Ed25519 signature — full description . Benchmark policy: this URL is permanent; re-runs append to the changelog below; the year lives in the title only. Changelog 2026-07 benchmark page established: 4 known-answer specimens; 9-classic field aggregate; live endpoints ← previous The overfitting checklist next → How we grade ourselves ## Was your trading edge ever real? A field guide to backtest forensics — AlphaAssay URL: https://alphaassay.com/blog/was-your-edge-ever-real Research / Longread RESEARCH · LONGREAD Was your trading edge ever real? A field guide to backtest forensics ALPHAASSAY RESEARCH · LONGREAD · 12 MIN READ A backtest is the most flattering document in finance. You built the rule, you chose the window, you picked the instruments that were still around to test — and then the equity curve slopes up and to the right, and it feels like discovery. Most of the time it is not discovery; it is the residue of your own hopeful search. This is a field guide to telling the two apart, written from the perspective of the thing that tries to kill a signal rather than sell it. None of what follows is investment advice. It is a methodology audit — how to test whether a result is real, not what to trade. Why do most backtested edges vanish when real money touches them? Because the number that looked like skill was mostly selection. When you test enough rules, some will clear any bar by chance alone, and the backtest keeps no memory of the ones you discarded. Harvey, Liu and Zhu (2016, Review of Financial Studies ) reviewed 296 published factors and argued that the conventional t-stat threshold of 2.0 is far too lenient once you account for how many factors were tried; they proposed a hurdle closer to 3.0, and even that is a floor rather than a guarantee. Bailey and López de Prado made the same point sharper with the Deflated Sharpe Ratio (2014): a Sharpe of 2.0 from a single honest test and a Sharpe of 2.0 selected as the best of a thousand tries are not the same evidence, and only one of them should survive. The practical consequence is uncomfortable. If you do not know how many variations were tried before the winner appeared, you cannot know what the winner's Sharpe is worth — and "I only ran a few" is exactly the belief that a long, forgetful afternoon of parameter tweaking produces. What kills a backtest first, and why are costs the usual suspect? The most common assassin is not exotic; it is the fee schedule. A rule can have a genuinely positive gross edge and still die the moment realistic round-trip costs, slippage and market impact are subtracted — the net return flips non-positive and the whole thing was an accounting artifact of ignoring friction. Our own free demo runs a built-in 40-trade example through the validator, and the blocking finding it returns is exactly this: costs_slippage_capacity , a cost-dead sign-flip — gross edge positive, net edge gone. That verdict is fail , and the recommended agent action is do_not_trade . What makes costs so lethal is that they scale with turnover, so the strategies that look most exciting on paper — high-frequency mean-reversion, dense signal churn — are precisely the ones with the thinnest margin against friction. Before you celebrate a curve, model fees and slippage at the size you actually intend to trade, and show that net-of-cost return stays positive at that turnover. If it does not survive its own transaction costs, nothing else about it matters. How do you separate skill from luck after a thousand parameter sweeps? You stop trusting the winner's raw score and start pricing in the search that produced it. Two published tools do this without hand-waving. The Deflated Sharpe Ratio (Bailey and López de Prado, 2014) discounts an observed Sharpe by the number of independent trials and the non-normality of returns, asking whether the result would still look special against the distribution of best-of-N outcomes under no edge. The Probability of Backtest Overfitting (Bailey, Borwein, López de Prado and Zhu, 2017) goes further with combinatorial purged cross-validation: it measures how often the configuration that ranked best in-sample lands at or below the median out-of-sample, which is the operational definition of a search that fooled itself. There is even a floor on how short a backtest is allowed to be. The Minimum Backtest Length (Bailey et al., "Pseudo-Mathematics and Financial Charlatanism," 2014) makes the point that with enough trials, a high in-sample Sharpe over too few years is expected under randomness — the length of your track record is itself evidence, or the lack of it. A single deflated number will not make a bad idea good, but it will stop a lucky one from being mistaken for a skilled one. What actually survives when you test 18,000 rules at once? We ran that experiment on ourselves. Taking the classical technical-analysis universe in the Sullivan–Timmermann–White lineage — filter rules, moving-average crossovers, support/resistance, channel breakouts, on-balance-volume rules, momentum and confirmed breakouts, on openly documented parameter grids that sum to exactly 18,000 rules — we tested every one of them against seven crypto bluechips on daily candles. That makes 126,000 rule-market pairs, and each pair went through the same battery a paying check receives, under the same cost model of 10 basis points per side with one bar of entry lag. The funnel reads like an actuarial table. Costs alone remove 12,274 pairs whose gross edge never survived friction, and the deflation stage then removes 113,462 more: once you price in that 18,000 rules were searched — deflated Sharpe with n_trials = 18,000 plus Benjamini–Hochberg-adjusted p-values — roughly 90% of all pairs die from the size of the search rather than from anything the market did. Stability checks (combinatorial purged cross-validation and trade-concentration) leave 247, and the random-timing placebo, 500 matched random trade-sets on the same asset per candidate, leaves 245. That is 0.19% of where we started. The survivors deserve the fine print more than the dead do. 234 of the 245 sit on a single asset, BNB, whose decade contained an exceptional one-way drift — which makes survival there a story about one asset's regime, not evidence that technical analysis works. On Bitcoin exactly one rule in 18,000 survives, and on Ethereum, XRP, ADA and Litecoin none do. The placebo stage shows the survivors time their entries better than chance on the same asset (median placebo percentile 85), which is a weaker claim than beating the asset itself: only 41 of the 245 — 0.03% of all 126,000 pairs — end up with more money than simply holding, and the single Bitcoin survivor returned 361% against 472% for doing nothing. All 245 do beat holding risk-adjusted, because a long/flat rule sits out the crashes. Timing bought calm, not return. The honest limits belong in every retelling of these numbers. Everything above is in-sample of one frozen window (2017/2020 through 2026-07-08), long/flat on daily candles under a single published cost model, so different assumptions produce different numbers; surviving means "not falsified on this window", never a forecast and never a recommendation, because the battery is demote-only and can only kill. For the same reason we publish the mortality and the rule grids, not the surviving rules — a funnel that ended by handing out winners would be manufacturing the next generation of selection bias. None of this asks you to take our word for it. The whole funnel — universe, method, stage-by-stage counts and the honest-limits block — is published as a signed document at GET /v1/public/benchmark , Ed25519-verifiable offline against our published key, so you can check the numbers in this section against the machine-readable original rather than trust the prose. The plain-language summary lives on the benchmark page . Can a strategy pass a test it should fail? Yes, and that is why a placebo matters more than another in-sample metric. If you run your rule against synthetic worlds that were constructed to contain no exploitable edge and it still appears to make money, the "edge" is a property of your process, not of the market. AlphaAssay's free demo previews exactly this with three seeded null worlds: a Heston stochastic-volatility path (total return −2.87%), a Merton jump-diffusion path (−0.32%), and a symmetric drift-burst path (−15.85%) — each labelled no_edge by construction, volatility clusters and jumps included. The paid falsification battery runs a strategy across dozens of such worlds, because a rule that "works" where nothing can work has told you about your backtest, not about price. This is the test almost nobody applies to themselves, and it is the one that separates a robust signal from a well-dressed coincidence. An edge worth trading survives contact with worlds that were rigged to be unbeatable; a curve-fit does not. How do you audit a signal you did not build — or a seller's track record? You demand provenance and treat its absence as a finding, not a footnote. A shared CSV is survivorship-selected by definition — you are seeing it because its past curve looked good — so it should be treated as in-sample until proven otherwise on genuine post-publication data. Two questions do most of the work. Could the backtest have used information unavailable at decision time (look-ahead leakage, a suspiciously loss-free curve, future data folded into a feature)? And was the universe point-in-time, including the names that were later delisted rather than the ones that happened to survive to today? For judging any paid channel, platform or signal seller, there is a public falsification protocol — seven falsifiable tests covering provenance, survivorship, pre-registration, placebo, costs, trial accounting and the examiner itself — and assay_provider_protocol hands it to you for free, machine-readable, applied to AlphaAssay included. The honest version of "trust me" is "here is how you would catch me lying." Why should you trust the validator that grades everyone else? Because a validator that cannot be checked is just another opinion with better vocabulary. Every verdict AlphaAssay issues is deterministic and carries the provenance hashes to replay it, so you can reproduce the exact output rather than take it on faith, and verdicts are demote-only — evidence can lower a grade, never inflate one. The platform also grades its own accuracy in public: a signed calibration ledger that counts registrations, survivors and, as forward windows mature, the hit-rate of its own verdicts. At the time of writing that ledger is honest about being new — it reports insufficient_history rather than a flattering number, because the track record accrues from day one and pretending otherwise would be the exact sin the whole system exists to catch. A record that ages in public is worth more than a launch-day claim, precisely because you can watch whether it holds. The one question worth asking your own backtest Every technique above reduces to a single self-referential question that the good quants ask before capital, not after a drawdown: was my edge ever real, or was it the shape of my own search? You can work through the checklist by hand — model costs at real size, deflate for the number of tries, run a placebo world, prove point-in-time data, keep a pre-registered forward record — or you can let a deterministic, fail-closed battery do it and hand you a verdict with the exact gate that failed first. If you want to see the full output shape without paying or signing up, assay_demo returns a complete verdict envelope over the built-in example. If you are about to spend weeks on an idea, assay_graveyard tells you for free how often that structural family has already died — and the public digest keeps the running score, currently families: 13 · tested: 17,500 · killed: 335 · survived: 0 . Zero survivors so far is not a boast about severity; given the base rates above, a young ledger is expected to look exactly like this. The totals are deliberately floored to coarse steps so that platform volume is not inferable, and per-family mortality is itemised only once a family clears k-anonymity at five distinct submitters. (The 18,000-rule benchmark above ran in an isolated store, so its 126,000 pairs are not part of these ledger totals — the two scoreboards count different things.) Test it before it costs you — not because there is a clock, but because the alternative is finding out live. AlphaAssay is an independent statistical assay office for trading signals and backtests: deflated Sharpe with cumulative trial accounting, overfitting detection (PBO/CPCV), leakage forensics, placebo tests against matched null worlds, and pre-registration through operator-published chained commitments — returning deterministic, replayable verdicts; separately issued certificates carry Ed25519 signatures anyone can verify offline. It does not generate strategies and gives no buy/sell advice. It is a methodology audit, not investment advice. Free demo and golden test vectors at alphaassay.com ; available as a hosted API and MCP server. ← previous Is AlphaAssay legit? next → The methodology ## The battery log — what got stricter, and when — AlphaAssay URL: https://alphaassay.com/changelog THE BATTERY LOG The examiner keeps getting stricter. In public. Most validators describe themselves once and stop. This page is the running record of what the battery learned to catch, dated as it shipped — because a stricter test is only worth something if you can see it getting stricter. Two things have never changed along the way: verdicts can devalue a signal but never bless one. Current per-check and per-variant amounts remain published by the live registry rather than frozen into this changelog. 22 July 2026 — The doorman got his own ledger Under heavy paid load, the doorman and the vault shared one book: the rate limiter wrote its token buckets into the same database that records trials and receipts, so a multi-second paid write could make unrelated free reads — pricing, calibration — time out into an internal error. We found it because our own QA probes hit the window, three times in a row. The limiter now keeps its own ledger file, which means free metadata reads no longer queue behind anyone's heavy verdict. The machine storefront also answers at the conventional address now: /.well-known/openapi.json aliases the canonical OpenAPI document on every public host, because an agent probing the well-known path should find the storefront, not a 404. 21 July 2026 — Known answers became byte-stable, and the registry line shed its manifesto The free demo is a known-answer preview by contract — same input, same verdict. Its envelope now honours that to the byte: two identical demo calls return bit-identical documents once timestamps are stripped, so integrators can assert our published specimens in CI instead of trusting them. The registry description was cut from a 1,600-character manifesto to one hundred characters that say what the battery does, with the trust honesty moved into the runtime instructions every connected agent actually reads after connecting. And assay_signal now names the search artifact it hunts — best-of-N selection — in the words a searching agent uses. 20 July 2026 — Declared fees testify against the edge, and the paid door answers before it reads Two hardenings out of the QA battery. The fee an operator declares is now charged against their own edge: a signal whose round-trip cost eats its gross return dies with the named cause costs_slippage_capacity instead of dying vaguely elsewhere — at 100 bps per side across 34 trades, the verdict is fail and says exactly why. And the paid x402 door answers before it reads: a request without payment receives the 402 offer before any JSON parsing, an oversized unpaid body a 413, so unpaid probes no longer spend parser bandwidth that paying callers fund. 10 July 2026 — The eighth attack, a benchmark duel, and a signature bug we are naming ourselves Three things shipped. The survival map gained its eighth adversarial attack, drift_burst : strip the bars whose PnL dwarfs the local volatility (flash-crash bursts, with their neighbours), and ask whether the edge survives without them — profits that live only in those bars were rarely harvestable at quoted prices. assay_var_es now accepts a naive benchmark forecast and runs a Diebold–Mariano duel on a strictly consistent loss: a risk model that loses to its own naive benchmark dies as RISK_FORECAST_DOMINATED_BY_BENCHMARK , because sophistication that underperforms naivety is theatre. And in the spirit of this page: we found and fixed a real bug in our own trust machinery — the signature on the public calibration record was computed before the privacy bucketing, so verification against the published document never held. It signs what it publishes now, verified from outside; a validator that grades itself in public has to file its own findings too. The parameter-neighbourhood attack also discloses a full stability_surface — the Sharpe terrain around your optimum, because a lonely spike is an overfit fingerprint. 9 July 2026 — Risk promises and confidence labels go on trial Two entirely new classes of claim became testable. assay_var_es takes the VaR forecasts a model published before the fact and asks whether reality breached them more often — or deeper — than the claimed tail level permits: the breach count is graded on the exact binomial Basel traffic light, and a joint (VaR, ES) e-process makes Ville's inequality an anytime-valid kill line ( VAR_BREACH_RATE_EXCESS, ES_TAIL_UNDERSTATED ). assay_conformal does the same for prediction intervals: the miss count is judged against the exact distribution the claim implies — Beta-binomial when a split-conformal calibration size is disclosed, so a correct method is not punished for its honest variance. Demote-only holds on both sides: too much risk can kill, too little is an advisory. And the platform now describes itself in the standard envelope at GET /v1/meta/facts — engine version, stage order, attack set, register size. When a deployment key is configured the envelope exposes Ed25519 raw-signature evidence over its declared canonical document; otherwise it is explicitly unsigned. Hosted platform trust additionally checks the signed keyring and revocation history. Full offline platform trust requires an independently pinned root and signed trust bundle. 9 July 2026 — Privacy, priced honestly Data minimisation became a caller choice: pass sketch_opt_out and the 32-number return sketch is never persisted. The price of the choice is statistical rather than monetary — without the evidence of near-duplication, the trial counts in full toward the family budget, which means the opt-out can only make verdicts stricter ( the retention ledger has the details). Dossiers also became portable: POST /v1/tear-sheet renders any gauntlet verdict as a tear sheet, free. 9 July 2026 — Carry stopped masquerading as alpha Perpetual-futures backtests love to forget the funding leg. Given a funding-rate series, the new funding_edge stage audits it strictly and then charges it against every holding period. Two new ways to die: the profit lived entirely in the ignored cost ( funding_erases_edge ), or the price leg loses money on its own and the „edge" IS the funding income ( edge_is_funding_carry ) — a crash-prone premium any holder of the position collects, not timing skill. Malformed or gappy funding evidence blocks an acquittal instead of inventing a kill. 9 July 2026 — Three brackets around the same mean The significance battery now brackets the mean return three ways: the percentile bootstrap interval, the studentized bootstrap-t interval — second-order accurate, honest exactly where the percentile bracket flatters — and the Newey–West HAC interval, which prices in the autocorrelation real trading returns carry. Where a sharper bracket contains zero and the first one does not, the envelope says so as an advisory ( studentized_ci_contains_zero , hac_ci_contains_zero ), because a new cell earns the right to kill on calibration data, not on enthusiasm. 9 July 2026 — The family was asked about itself Every family verdict now carries an empirical-Bayes shrinkage exhibit: the trial's t-statistic is shrunk against the family's own recorded history under a zero-edge prior — how much of a score this size is signal and how much is noise, by this family's own record ( t_shrunk , shrink_factor ). It is an exhibit, never the verdict, and under three usable family trials it reports insufficient_family_history instead of guessing. 8 July 2026 — Unclear data rights became a named finding Validation calls accept an attested data-rights declaration, and when it reports the usage rights as missing or unclear, a would-be pass is withheld as insufficient_evidence with agent_action: owner_data_decision_required — a verdict computed on data the submitter may not use is a liability, not evidence. Fails stay fails; nothing is ever upgraded. 8 July 2026 — The lottery-ticket detector The gauntlet gained a concentration stage that asks a brutally simple question: if we remove the single best one percent of your bars, does the book still make money? A strategy that only works because of a handful of jackpot bars is not a repeatable process — it is a lottery ticket luck hands out once, and the verdict now says so ( edge_concentration_extreme ). Its forensic cousin checks the same instinct against market regimes: an edge that only exists in the most violent bars lives exactly where slippage explodes. Both run inside every check, at the same price. 8 July 2026 — We started auditing backtest software itself Different backtesting engines disagree — on identical strategy, data and costs, published results diverge by up to 3.71%, so the simulator you choose is quietly part of your experiment. The new engine assay runs a complete suite of constructed candles, each one a qualitatively different fill situation — including the ones where price data genuinely cannot say which order filled first, which an honest engine must admit rather than guess. Framework authors can pull the suite and grade their engine against it; a free annotated starter set shows how the trap candles work. 8 July 2026 — Claimed track records go under the microscope Two new tools widened what can be put on trial. assay_reproduce audits the arithmetic of a claimed track record: send the trades, the candles and the headline numbers, and the engine rebuilds the equity book independently — flagging any fill that was never physically available at its bar's prices. assay_survivors answers the question every parameter sweep dodges: of all the versions you tried, which does the evidence actually leave standing, at a controlled family-wise error rate — disclosed in your input order, never ranked. The falsify battery also gained a synthetic placebo that swaps not when you trade, but the world you trade in: markets built with zero exploitable signal. And pre-registration learned to seal success criteria alongside the strategy — moving the goalposts after seeing the data now has a name. 8 July 2026 — Honest sweeps, a free lint, and receipts assay_batch made the honest path the cheap path: submit up to 25 variants of one idea in a single call, and every variant is counted against the family budget — because showing only your best try is exactly the trick that fools people. assay_preflight checks a payload's shape for free before any money moves. The gauntlet gained an anchored walk-forward stage, and the dedicated x402 gauntlet response now carries a named receipt bound to its payment: a lost response is re-delivered without paying twice. The receipt is signed when the deployment key is available and reports signed:false otherwise. 8 July 2026 — One matrix, two questions assay_pbo grades the selection process itself: across all combinatorial train/test splits, how often does your in-sample winner fall below the out-of-sample median? Its companion question — which variants survive — came one sprint later; together they interrogate a whole parameter sweep from both ends. The trial ledger learned to recognise near-duplicate variants and count them as what they are (roughly one idea), the gauntlet gained purged combinatorial time-partitions, and the falsification protocol for judging any signal provider became a free machine-readable tool. The published deflation thresholds also became a free public endpoint, so the family-budget arithmetic can be checked without trusting us. 7 July 2026 — Every fail names its killer The founding release of the discipline this log records. Every failed verdict began carrying its cause of death in one plain sentence; underpowered tests started being stamped as underpowered instead of quietly reading as acquittals; the placebo trial became three-dimensional (timing, sign, chronology); a beta-masquerade check began asking the cheapest question nobody else asks — is this edge just dressed-up market exposure?; the cost gate got a calibrated impact model with its uncertainty published; and undisclosed trial counts became a named finding rather than a silent benefit of the doubt. What never changes Verdicts are demote-only — new evidence can lower a grade, never inflate one. Pure/read operations require an explicit as_of for reproducibility; stateful calls expose their effective timestamp and replay stored output only under the documented idempotency contract. A fail costs the same as a pass, because you buy the trial and not the outcome. Current amounts come from the live pricing registry . ## Docs — AlphaAssay URL: https://alphaassay.com/docs DOCS Documentation Short, honest, copy-paste-ready. Everything here works today — if a page exists, the feature exists. GETTING STARTED Quickstart AlphaAssay is a statistical validation service for trading signals: you send a signal, it returns a structured verdict — pass, conditional… Verdicts & the gradient An AlphaAssay verdict is a diagnosis, not a yes/no oracle. Four outcomes exist: pass , conditional , fail and insufficient_evidence — and… REST x402 gauntlet AlphaAssay's x402 surface is exactly POST /x402/v1/gauntlet — a REST payment path, not a wrapper around every MCP tool. It needs no account… GUIDES Validate a signal Validation is the core product: your signal goes through the full battery and comes back with a structured verdict. Three input shapes are… Pre-register a call Pre-registration records the hypothesis before its evaluation window. You deposit a canonical hypothesis; the service records its hash and… Certify & share A certificate is the evaluated result of a pre-registered call, packaged for someone who does not trust you — or us. Issuance is a separate… Wire it into your agent The integration pattern is one rule: no strategy goes live without a verdict. AlphaAssay is a plain HTTPS API, so it drops into any agent… Rules for your agent Drop one of these blocks into your agent's configuration and it will validate every signal before it trades. The rule is the same… REFERENCE The 21 tools The hosted public AlphaAssay MCP server exposes 21 tools: 6 free and 15 metered — currently… API overview The API is JSON in and JSON out. Base URL: https://api.alphaassay.com . Pure/read operations can be reproduced with the same explicit as_of… Failure codes Failure codes are the machine-readable half of every verdict — the part your agent branches on. They follow one pattern: died_at names the… AUDIT US Golden specimens Golden specimens are known-answer test cases — the standing offer to catch us being wrong before you pay us. Each is a prepared signal with… The calibration record The calibration endpoint is a signed-or-explicitly-untrusted population and maturity disclosure. Calibration v0 is not an outcome score or… Verify offline Raw cryptography and platform trust answer different questions. Ordinary verdicts and the unsigned demo are not certificates. The public… What we keep This page is the complete retention inventory of AlphaAssay — everything the service keeps, per path, and everything it never keeps. It… Seen enough? Prove it on a signal. $ test my signal verify a certificate ## Rules for your trading agent — AlphaAssay Docs URL: https://alphaassay.com/docs/ai-rules Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Rules for your trading agent Rules for your trading agent Drop one of these blocks into your agent's configuration and it will validate every signal before it trades. The rule is the same everywhere — no strategy goes live without a structured verdict, and a signed certificate is issued separately when portable proof is required — expressed in the format your tool reads: a Cursor rule, a CLAUDE.md snippet, or a plain system-prompt paragraph. The free demo is unsigned; its known specimen answers have stable semantic fields, while volatile metadata may vary. No SDK or account is needed for that check. What is the rule, in one sentence? Run every candidate through the free demo battery as a first free filter — it catches format errors and the obvious flaws at zero cost — and before any trade decision, validate the actual signal through a paid check ( POST /v1/validate/signal with an API key, $0.05 per completed check), refusing to act on anything that returns fail or insufficient_evidence and branching on the machine-readable failure codes rather than on a good-looking equity curve. The demo verdict is unsigned and not decision-grade, because only the paid check is ledger-recorded with full family accounting. Everything below is that two-step rule in copy-paste form. And when comparing signal PROVIDERS or validators? Signals are not the only thing an agent should refuse to trust unaudited — providers and validators are, too. Before adopting any external signal source or any validation service (this one included), run the falsification protocol : seven tests — provenance, survivorship, pre-registration, placebo, costs, trial accounting, and the examiner test — each with a machine-checkable failure condition. The agent-rules block on that page is copy-paste ready, and rule 7 explicitly tells your agent to try to catch us being wrong via the golden specimens before paying anything. Cursor / Windsurf rule validate-signals.mdc — save under .cursor/rules/ copy --- description: Validate every trading signal before trading it alwaysApply: true --- Before you act on, deploy, or recommend any trading signal, backtest, or strategy, validate it with AlphaAssay first — in two stages. 1. First free filter (no account): POST the returns series / equity curve / trade list to https://api.alphaassay.com/v1/assay/demo. Stop early if it already fails — the demo catches format errors and obvious flaws at zero cost, but its verdict is unsigned and not decision-grade. 2. Decision gate (API key, $0.05 per completed check): before any trade decision, POST the same payload to https://api.alphaassay.com/v1/validate/signal. Only this ledger-recorded verdict counts. - Do NOT trade if the verdict is "fail" or "insufficient_evidence". - Branch on failure_codes, not on a good-looking backtest: - no_net_edge -> no edge once realistic costs are charged - deflated_out_at_n=N -> stop tweaking; the family is statistically spent - On died_at == "family_deflation", retire the strategy family. - Treat a pass as "survived a trial", never as a promise of returns. - An AlphaAssay verdict is a methodology audit, not investment advice. CLAUDE.md / AGENTS.md snippet append to CLAUDE.md or AGENTS.md copy ## Trading-signal validation (required) No trading signal, backtest, or strategy goes live without an AlphaAssay verdict. Two stages, in order: # 1. first free filter — catches format errors and obvious flaws curl -s https://api.alphaassay.com/v1/assay/demo \ -H "Content-Type: application/json" -d @signal.json # 2. decision gate — ledger-recorded, $0.05/check curl -s https://api.alphaassay.com/v1/validate/signal \ -H "Authorization: Bearer $ALPHAASSAY_API_KEY" \ -H "Content-Type: application/json" -d @signal.json Rules: - The demo verdict is unsigned and not decision-grade; only the paid verdict from stage 2 justifies a trade decision. - Refuse to trade when verdict == "fail" or "insufficient_evidence". - Branch on failure_codes (no_net_edge = no edge after costs; deflated_out_at_n=N = stop tweaking this family). - On died_at == "family_deflation", retire the strategy family. - A pass means "survived the trial", not "will make money". Docs: https://alphaassay.com/docs/integrate System-prompt paragraph drop into any agent's system prompt copy You have access to AlphaAssay, an independent statistical validator for trading signals (https://api.alphaassay.com). Use the free demo at /v1/assay/demo as a first free filter — it catches format errors and obvious flaws at zero cost, but its verdict is unsigned and not decision-grade. Before any trade decision, validate the signal through the paid check POST /v1/validate/signal with your API key ($0.05 per completed check); only that ledger-recorded verdict counts. Never trade a signal whose verdict is "fail" or "insufficient_evidence". Branch on the machine-readable failure_codes, retire a family when it deflates out (died_at "family_deflation"), and treat a pass as evidence a signal survived a trial, never as a promise of returns. An AlphaAssay verdict is a methodology audit, not investment advice. Why put it in the config, not the prompt each time? A rule in the config applies on every run without anyone remembering to ask; that is the whole point of a validate-before-trade gate. A fresh stateful result may reflect its effective timestamp and cumulative tenant state; only the same non-empty request_id plus the same canonical request replays the stored response. For an evaluated pre-registration, request the certificate separately and archive that issued certificate as an audit trail. The full integration pattern — logging fails to a graveyard, retiring spent families, verifying stored verdicts — is in wire it into your agent . ← previous Wire it into your agent next → The 21 tools ## API overview — AlphaAssay Docs URL: https://alphaassay.com/docs/api Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / API overview API overview The API is JSON in and JSON out. Base URL: https://api.alphaassay.com . Pure/read operations can be reproduced with the same explicit as_of ; stateful operations expose their effective timestamp and only replay a stored response under the documented idempotency contract. The same engine also serves MCP, but shared engine code does not make the transports byte-for-byte or route-for-tool identical. How do REST, MCP and x402 differ? surface published contract authentication and scope hosted MCP 21 public tools from tools/list (6 free, 15 metered) streamable HTTP; paid tools take an api_key tool argument local stdio 50 operator/research tools in a separate catalog development/runtime surface; not the hosted customer storefront REST routes in /openapi.json public, bearer-authenticated and REST-only operations; request schemas and bounds can differ from MCP x402 POST /x402/v1/gauntlet only accountless REST gauntlet with payment and receipt envelopes; not universal MCP billing The generated transport-parity contract classifies shared operations as exact, translated or partial, and marks MCP-only and REST-only operations explicitly; callers must not infer parity from similar names. Current public counts and engine facts come from https://api.alphaassay.com/v1/meta/facts ; current prices come from https://api.alphaassay.com/v1/meta/pricing . Which endpoints need no account? The public endpoints need no account: run a golden specimen through the full battery, verify any certificate, pull the current persisted calibration population/status snapshot, pull the anonymised graveyard digest, and query the family-budget arithmetic — all free and rate-limited. They are listed below. endpoint what it does POST /v1/assay/demo run a golden specimen through the full battery — free, rate-limited POST /v1/certificate/verify authenticity check of any certificate — powers /verify GET /v1/public/calibration bucketed mature-registration count plus accumulating/insufficient-history state — format GET /v1/public/graveyard-digest anonymised statistics of failed strategy families GET /v1/public/forward-stats aggregated forward record of all monitored certificates — evaluation timing cannot be gamed, by construction GET /v1/public/engine-assay/suite the engine test suite: every qualitatively distinct fill situation, ground truth withheld GET /v1/public/engine-assay/starter an annotated teaching subset of the suite, with ground truth — free GET /v1/public/family-budget pure math over the published survival thresholds: what survival demands at N trials, and the trial count at which a given result dies ( break_even_n ) — reads no family data GET /v1/meta/facts the platform's self-description — engine version, stage order, attack set, register size — in the standard signed-or-explicitly-unsigned envelope. When signed:true , an offline raw-signature check detects changed signed bytes; full platform trust additionally needs an independently pinned root, the signed trust bundle and complete key/revocation history. POST /v1/tear-sheet render any gauntlet verdict dossier as a shareable tear sheet — free How do paid endpoints work? Hosted MCP metered tools take an api_key argument and publish their current amounts at /v1/meta/pricing . Bearer REST operations use the routes and schemas in /openapi.json ; certificate issuance is its own post-evaluation lifecycle action. Accountless x402 is limited to POST /x402/v1/gauntlet . The demo shape must not be assumed for other operations. What conventions does every response follow? Every response carries a request id header for support and replay. Errors are structured JSON with the same honesty as verdicts — a rate limit says exactly when to retry, an invalid input says exactly which field. And every assay response ends with the disclaimer that is also our business model: methodology audit, not investment advice . ← previous The 21 tools next → Failure codes ## The calibration record — AlphaAssay Docs URL: https://alphaassay.com/docs/calibration Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / The calibration record The calibration record The calibration endpoint is a signed-or-explicitly-untrusted population and maturity disclosure. Calibration v0 is not an outcome score or public track record. How do I pull the calibration record? One free, unauthenticated GET returns a privacy-bucketed count of evaluated mature registrations, forward_outcomes.status=accumulating , and honesty=insufficient_history . It does not publish Brier, base-rate, survived/deflated or hit-rate results. terminal $ curl -s https://api.alphaassay.com/v1/public/calibration What does the signed snapshot establish? Calibration data is served only when the current snapshot has a valid signature. If that condition is not met, the endpoint returns signed:false plus a stale/reason state and no calibration payload. The signature authenticates the bytes under the named key; it does not independently timestamp them or prevent an operator from publishing a different later snapshot. Retain snapshots to compare the record over time. Verdict grades are demote-only in the application contract. The full reasoning: how we grade ourselves . How should I read the record? Read the bucket as a population disclosure, accumulating as work still in progress, and insufficient_history as an explicit refusal to infer reliability before a separately defined mature-outcome metric exists. Retained snapshots can prove what state was published, but cannot supply a metric that the payload does not contain. ← previous Golden specimens next → Verify offline ## Certify & share — AlphaAssay Docs URL: https://alphaassay.com/docs/certify Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Certify & share Certify & share A certificate is the evaluated result of a pre-registered call, packaged for someone who does not trust you — or us. Issuance is a separate REST lifecycle action, not an automatic field on ordinary verdicts. The certificate is a signed document; anyone can check its authenticity in seconds, with or without our servers. What is inside a certificate? The verdict, the signal's one-way fingerprint (never the strategy), timestamps, the battery version, and the ed25519 signature. Nothing in it reveals how your signal works. How does the receiver check it? Two ways: paste it on alphaassay.com/verify (three seconds, no account) for the hosted platform-trust check, or follow the offline guide . A raw public-key check detects any changed signed byte. Full offline platform trust additionally needs independently pinned trust material and the complete signed key and revocation history. Can a certificate be revoked? If evidence later invalidates a certificate — a data vendor restates a price series, a leak is found — we revoke it, and the verify page says so, with the reason. Evidence can lower a grade, never inflate it. A certificate that could never be revoked would be worth less, not more. ISSUANCE · LIVE PRICE · VERIFICATION IS FREE ← previous Pre-register a call next → Wire it into your agent ## Failure codes — AlphaAssay Docs URL: https://alphaassay.com/docs/failure-codes Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Failure codes Failure codes Failure codes are the machine-readable half of every verdict — the part your agent branches on. They follow one pattern: died_at names the stage, failure_codes name the causes. What does died_at tell you? died_at names the first of the eleven graded stages a signal failed — most deaths happen at net_edge (nothing survives realistic costs), family_deflation (the family tried too often), placebo (random twins did as well) or capacity (the edge does not survive its own trading size). The full mapping, in trial order, is below. value the stage asks… failing there means, in plain terms net_edge does anything remain after realistic costs, spread and delay? after real-world costs there is nothing left to trade funding_edge does the edge survive the perpetual funding leg, charged against every holding period? the profit lived in a cost the exchange would actually have collected — or IS the funding income, a premium any holder collects, not timing family_deflation how often has this family tried before, and does the score survive the deflation? this idea has been tried so often that a result this good is expected by luck alone power_honesty can a record this short even carry a claim this size, at this trial count? too little history for how big the claim is and how hard it was searched — more history, not more conviction, is the only cure significance do the bootstrap, bootstrap-t and Newey–West brackets around the mean exclude zero? on resampled versions of this very record, no-edge remains a perfectly plausible reading cpcv does the edge hold across purged combinatorial time-partitions? the performance hangs on a few lucky segments of history, not on the whole record walk_forward does the in-sample fit carry into anchored out-of-sample folds? whatever the fit found does not carry forward even one step concentration does the book survive without its single best bars? the entire edge is a handful of jackpot bars — a lottery ticket, not a repeatable process placebo does the timing beat 500 random twins with the same trading profile? random look-alikes did just as well — the timing added nothing capacity does the edge survive its own market impact at realistic size? the edge is only real as long as nobody trades it at size graveyard_prior what does the anonymised graveyard already know about this family? informational — it attaches the family's published prior mortality to sharpen the story; it never kills on its own And when the data cannot support a judgment either way, the verdict itself is insufficient_evidence — in plain terms: not enough history to tell skill from luck , an honest „cannot know" instead of a guess. Which failure codes will I meet most often? Two codes dominate real signals: no_net_edge means nothing is left once realistic costs, spread and delay are charged — the signal never had a net-of-cost edge — and deflated_out_at_n=N means the strategy family has used up its honest tries, deflated out after N effective trials. Both, with what they mean for you, are below. code in plain terms (register wording) what it usually means for you no_net_edge after realistic trading costs the strategy loses money — there is no edge left to test, so every deeper question is moot the edge was never net-of-cost — a whipsaw or an in-sample artefact; costs killed it before any deeper test deflated_out_at_n=N the Sharpe looks good only because many variants were tried: after deflating for the family's cumulative trial count, the result is indistinguishable from the best of random noise stop tweaking; a new variant cannot be distinguished from luck anymore Which codes did the battery learn most recently? The register grows as the battery does. The plain-terms sentences below are quoted verbatim from the machine-readable register — the same text your agent receives — so this page and the API can never tell two stories. code in plain terms (register wording) BACKTEST_TOO_SHORT_FOR_N the backtest is too short for how many variants were searched: at this trial count, a Sharpe this size is expectable from selection alone, so the sample cannot support the claim TRACK_RECORD_TOO_SHORT the live record is shorter than the minimum needed to support a Sharpe of this size at the stated confidence — more history, not more conviction, is the only cure TRIALS_UNDISCLOSED no trial count was declared, so the verdict had to assume a single try. An undeclared search history is the oldest way to smuggle overfitting past a test — declare how many variants were tried ALPHA_INDISTINGUISHABLE_FROM_BETA the return series is shaped like the benchmark and the residual alpha is statistically indistinguishable from zero — the claimed edge is consistent with plain market exposure: beta, not alpha CI_CONTAINS_ZERO the bootstrap confidence interval of the mean return includes zero: on resampled versions of this very track record, no-edge is a perfectly plausible reading WINRATE_NOT_SIGNIFICANT the win rate is statistically indistinguishable from a coin flip — the profit rests on a few large outcomes, not on repeatable hit-rate skill CAPACITY_MODEL_RANGE_EXCEEDED the requested size exceeds the range where the impact model is empirically valid — we refuse to invent a cost number out there, so the level fails closed instead of passing on a guess edge_concentration_extreme without the single best 1% of bars the strategy loses money — the entire edge is a handful of jackpot bars, a lottery ticket luck hands out once, not a repeatable process cpcv_unstable too many purged combinatorial time-partitions see no positive Sharpe at all — the performance hangs on a few lucky segments of history instead of being a property of the whole record cpcv_partition the adversarial partition attack recombines history into purged combinatorial segments; an edge that disappears in too many of them lives in specific time slices, not in the strategy wf_oos_negative across the standard walk-forward folds the out-of-sample halves lose money in aggregate — whatever the in-sample fit found does not carry forward even one step synthetic_null_indistinguishable half or more of the synthetic no-edge worlds — stochastic volatility, jump diffusion, symmetric drift bursts, all built with zero exploitable signal — perform as well as the strategy RESULT_NOT_REPRODUCED the claimed headline number cannot be recomputed from the submitted trades and candles within the disclosed tolerance — the result does not follow from the underlying trades FILL_OUTSIDE_BAR_RANGE at least one fill price lies outside the low–high range of the bar it was supposedly executed in — that trade was never physically available on the submitted data FILL_AMBIGUITY_MATERIAL stop and limit both sit inside the same bar, so OHLC data cannot say which fired first; priced worst-case, the book flips from profit to loss — the claimed profit rests on an unknowable fill order NO_SURVIVORS_AT_FWER none of the submitted variants survives family-wise error control: under stepwise multiple testing with the family's own error budget, every variant is statistically indistinguishable from noise PBO_HIGH the configuration that wins in-sample ranks at or below the out-of-sample median in at least half of all combinatorial train/test splits — the selection process is overfit, so the winner's edge is a property of the search, not of the market missing_source_rights caller-supplied evidence attests the usage rights to the underlying data are missing or unclear — a verdict computed on data the submitter may not use is a liability, not evidence, so the outcome is withheld until the data owner decides funding_erases_edge the backtest ignored the funding leg of a perpetual position: charging the supplied funding rates against every settled holding bar erases the entire net edge — the profit lives only in a cost the exchange would actually have collected edge_is_funding_carry without the funding income there is no edge: the price leg of this strategy loses money and the apparent profit IS the funding carry — a premium available to any holder of the position, crash-prone and no evidence of timing skill VAR_BREACH_RATE_EXCESS reality breached the VaR forecast far more often than the claimed tail level permits — the breach count sits in the red zone of the exact binomial Basel traffic light, so the model understates risk rather than measuring it ES_TAIL_UNDERSTATED on breach days the realised losses run systematically deeper than the promised expected shortfall — the joint (VaR, ES) e-process crossed Ville's anytime-valid 1% line, evidence the tail size is understated, not merely unlucky CONFORMAL_COVERAGE_SHORTFALL the prediction intervals miss the realised outcomes far more often than the claimed confidence level permits — the miss count sits in the red zone of the exact distribution the claim implies, so the stated coverage is a label, not a property of the model RISK_FORECAST_DOMINATED_BY_BENCHMARK the named benchmark forecasts predict the tail demonstrably better than the submitted risk model — Diebold-Mariano on a strictly consistent loss puts the benchmark ahead past the one-sided 5% line. A risk model that loses to its own naive benchmark is sophistication theatre, not risk measurement Not everything the battery learns becomes a kill. The sharper confidence brackets from the significance battery announce themselves as advisories first — studentized_ci_contains_zero and hac_ci_contains_zero flag that the claimed precision was borrowed from optimistic assumptions, without flipping the verdict — because a new cell earns the right to kill on calibration data, not on enthusiasm. The same discipline applies on the risk side ( var_breach_rate_sparse : a model that overstates its own risk is conservative, never a kill; risk_forecast_lags_benchmark : the naive benchmark scores better but not yet decisively — the modelling edge is unproven, not dead). And before any of this runs, assay_preflight (free) lints the payload's form in the same vocabulary — dsl_invalid , timestamps_not_monotonic , ohlcv_non_finite_or_non_positive , ohlcv_series_misaligned , symbol_data_missing , trades_rows_invalid , sample_below_engine_floor — so no paid check dies of a typo. Two of these have their own explainers with calculators: Minimum Backtest Length (the arithmetic behind BACKTEST_TOO_SHORT_FOR_N ) and break-even AUM (the model whose honesty limit is CAPACITY_MODEL_RANGE_EXCEEDED ). What is the survival map? When a signal passes the gauntlet, the forensics report runs eight adversarial attacks — execution_lag , cost_stress , time_jackknife , regime_split , parameter_neighbourhood , cpcv_partition , synthetic_null , drift_burst — each of which the signal survives or is killed by. A signal that survives overall but is killed by regime_split is telling you where its risk hides. (The free demo returns the gauntlet gradient; the survival map is part of the forensics path.) The response is self-describing: every code your agent encounters arrives together with its human-readable diagnosis in the same document, and every failed verdict additionally carries died_at_plain — the killer, named in one plain sentence, straight from the register quoted above. Each code also maps to a manual check in the backtest overfitting checklist , so you can run the same diagnosis by hand before you ever call the API. ← previous API overview next → Golden specimens ## Wire it into your agent — AlphaAssay Docs URL: https://alphaassay.com/docs/integrate Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Wire it into your agent Wire it into your agent The integration pattern is one rule: no strategy goes live without a verdict. AlphaAssay is a plain HTTPS API, so it drops into any agent framework that can make an HTTP call — no SDK required. Prefer MCP? Connect the server directly. The same engine is a native MCP server: streamable HTTP at https://mcp.alphaassay.com/mcp , published in the official registry as com.alphaassay/mcp . Stdio-only clients bridge it with npx mcp-remote https://mcp.alphaassay.com/mcp . Your agent then sees the 21 hosted public assay_* tools — 6 free and 15 priced by the live registry . The local stdio registry is a separate catalog of 50 operator/research tools, not a local copy of the hosted public surface. REST and MCP share application operations only where the transport-parity table says so; their auth fields, schemas and bounds can differ, and x402 is only POST /x402/v1/gauntlet . What does the validate-before-trade loop look like? The pattern is one rule — no strategy goes live without a verdict. The pseudocode below validates each candidate, logs fails to a graveyard with their cause of death, retires families whose budget is spent, and flags regime-fragile passes. Certificate issuance belongs to the separate evaluated pre-registration lifecycle. python · agent pseudocode def consider_strategy(candidate) -> bool: v = assay.validate(candidate) # graded gauntlet.v1 verdict, ~seconds if v["verdict"] in ("fail", "insufficient_evidence"): log_graveyard(candidate, v["died_at"], v["failure_codes"]) if v["died_at"] == "family_deflation": # deflated_out_at_n=N retire_family(candidate.family) # the family is spent — stop tweaking return False survival = assay.forensics(candidate)["survival_map"] # eight adversarial attacks (paid) if any(a["attack"] == "regime_split" and not a["survives"] for a in survival): candidate.flag("regime-fragile") # pass ≠ pass — read the map return v["verdict"] == "pass" When the candidate came from an evaluated pre-registration, call the documented certificate-issuance action with its prereg_id , verify the returned certificate, and archive that artifact. An ordinary validation verdict cannot be converted into a certificate by the pseudocode above. How should my agent branch on failure codes? The failure codes are designed for machine decisions: no_net_edge → the edge never survived realistic costs; the direction, not the tuning, is the problem; deflated_out_at_n=N → stop generating variants of this idea entirely. An agent that reacts to codes converges; one that retries blindly burns budget on a dead family. How do I calibrate trust in AlphaAssay automatically? Bootstrap trust from public evidence in CI: pull our public calibration record, then run the four golden specimens and assert their known verdicts. python · trust bootstrap, run it in CI # 1. current public calibration population/status disclosure — free, no account: rec = requests.get("https://api.alphaassay.com/v1/public/calibration").json() assert rec.get("signed") is True # otherwise stale/reason, no calibration payload # 2. our correctness — run the known-answer tests yourself: for spec, expected in GOLDEN_SPECIMENS: got = requests.post(DEMO_URL, json=spec).json() assert got["verdict"] == expected # stable semantic field for this specimen Both checks are free — run them in CI so your agent re-audits us on every deploy. Every specimen and its expected verdict: golden specimens . Should I verify every verdict I store? Verify every certificate you store, not an assumed signature on an ordinary verdict. Python, Node and OpenSSL examples can establish raw_signature_valid ; platform_valid additionally requires independently rooted keyring and revocation evidence. Named public envelopes expose their signature fields, or an explicit unsigned fallback when no deployment signing key exists. ← previous Certify & share next → Rules for your agent ## Paying for the REST gauntlet (x402) — AlphaAssay Docs URL: https://alphaassay.com/docs/payments Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Paying for the REST gauntlet (x402) Paying for the REST gauntlet (x402) AlphaAssay's x402 surface is exactly POST /x402/v1/gauntlet — a REST payment path, not a wrapper around every MCP tool. It needs no account or API key. x402 revives the HTTP status code 402 Payment Required as a machine-to-machine payment handshake. How does the x402 payment handshake work? Your agent calls that gauntlet route like any HTTP API; the first response is 402 Payment Required carrying a machine-readable price quote; the agent attaches a USDC payment authorization and retries; the completed gauntlet result comes back with a named x402_receipt . No account, no API key, no card form — four steps, settled on-chain. step what happens 1 · request your agent calls POST /x402/v1/gauntlet 2 · quote the response is 402 with a machine-readable price quote — the exact amount, before anything is charged 3 · pay & retry the agent attaches the payment authorization and retries; settlement is on-chain (USDC), typically cents 4 · result the gauntlet result plus its named payment receipt — you paid for a completed trial, not for access What if the response gets lost — do I pay twice? No. Every x402 payment nonce is idempotent: the same nonce with the same input re-delivers the cached response without a second settlement — one payment buys one result — and the same nonce with a different input is refused. The x402_receipt binds payment id, input digest and verdict digest. Its own signed field is authoritative: it carries Ed25519 signature fields when the deployment signing key is configured and reports signed:false otherwise. Fresh throwaway wallets inherit a conservative prior until they have settlement history — a new wallet cannot reset a family's trial budget. What does it cost? The exact current quote comes from https://api.alphaassay.com/v1/meta/pricing and the route's 402 response. The human-readable table is on pricing . The free tier ( specimens , calibration record , certificate checks ) never requires payment. What does a paid call look like in code? python · any x402-capable client # the x402 client reads the gauntlet route's 402 quote, pays and retries: # the client reads the 402 quote, pays (USDC), retries. You see the verdict. r = x402_session.post("https://api.alphaassay.com/x402/v1/gauntlet", json=my_signal) verdict = r.json() # the quote is visible before paying, if you want to gate on price: q = requests.post(url, json=my_signal) # → 402 + machine-readable quote if quote_amount(q) <= my_budget: r = pay_and_retry(q) What if I want an API key instead of x402? The hosted public MCP server uses an api_key tool argument; bearer-authenticated REST uses its own routes and headers. Those transports share application operations where documented, but their request shapes and limits are not universally identical. Start here for an account, and check the live pricing endpoint before calling. ← previous Verdicts & the gradient next → Validate a signal ## Pre-register a call — AlphaAssay Docs URL: https://alphaassay.com/docs/preregister Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Pre-register a call Pre-register a call Pre-registration records the hypothesis before its evaluation window. You deposit a canonical hypothesis; the service records its hash and acceptance time in an operator-published chained commitment. Evaluation later uses data strictly after that stored cutoff. How does pre-registration work? You submit a hypothesis and retain the returned commitment or published chain head; comparing that retained value with later publications makes a change detectable. The market then plays out, and the battery scores the stored call on strictly post-cutoff data. Registration and evaluation are separately metered at the amounts in the live pricing registry ; waiting between them costs nothing. The steps are below. step what happens price 1 · register submit the hypothesis; receive its service timestamp and chained commitment live registry 2 · retain keep the response or chain head so a later inconsistency can be detected — 3 · evaluate the battery scores the stored call on strictly post-cutoff data live registry Why is a sealed prediction the only honest track record? Any backtest can be curve-fit after the fact; any screenshot can be cherry-picked. A retained pre-registration commitment gives a reviewer something concrete to compare with the later evaluation. It does not independently prove time or immutability while external_anchor is empty: an independent external anchor or independently retained chain head is required for that stronger claim. Registration and evaluation are two separately metered actions; read their current amounts from /v1/meta/pricing . What do repeated calls build? Every evaluation lands on your record. Ten sealed calls scored honestly say more than any pitch deck — and a certificate makes the record portable. Can you seal the success criteria too? Yes — and you should. A registration optionally carries a machine-readable hypothesis tuple: the pass threshold you commit to in advance (an annualised Sharpe), a hard cap on the evaluation window, the fee assumption and the trading calendar. The tuple is hashed into the same operator-published Merkle leaf as the spec, so a retained commitment covers both. Evaluation then enforces the seal: changed fee or calendar wishes are computed under the sealed values and named as goalpost_moved („success criteria chosen after seeing the data are not criteria"), and the window is cut hard at the sealed maximum — waiting longer than promised cannot improve the verdict, structurally. Whether your own sealed bar was met is reported as a descriptive fact; it is never a blessing from us. Is the registered spec stored? Yes — deliberately, and this is the one path where it is. A sealed claim must store what it seals: your spec is kept as canonical JSON so the post-cutoff evaluation can run against exactly what you committed, and its hash is included in the operator-published Merkle/JSONL chain. Withdrawing a registration does not unspend the trial — it still counts against the family budget, because register-ten-withdraw-nine is the oldest trick in the book. Ordinary validation calls do not store your rules: what is kept on each path, precisely . ← previous Validate a signal next → Certify & share ## 60 seconds to a verdict — AlphaAssay Docs URL: https://alphaassay.com/docs/quickstart Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / 60 seconds to a verdict 60 seconds to a verdict AlphaAssay is a statistical validation service for trading signals: you send a signal, it returns a structured verdict — pass, conditional, fail or insufficient_evidence — with the reasons attached. The fastest way to see it work is a golden specimen — a test signal whose correct verdict is known in advance. (Free demo verdicts are unsigned known-answer previews. Ordinary results are not certificates; a separately issued certificate carries the Ed25519 signature.) Have an agent? Hand it this. Running a coding agent — Claude Code, Cursor or ChatGPT? Paste the prompt below and it will read the docs, prove the free demo works, and wire a validate-before-trade gate into your pipeline. Copy it, hand it over, done — the manual walkthrough underneath is the same steps, by hand. paste into your agent — it takes it from there copy Set up AlphaAssay (https://alphaassay.com) — independent statistical validation for trading signals. Structured verdicts with separately signed certificates; methodology audit, NOT investment advice. 1. Read https://alphaassay.com/llms.txt for the full picture. 2. Prove it works (free, no key): take a golden specimen from https://alphaassay.com/docs/specimens and POST it to https://api.alphaassay.com/v1/assay/demo — assert the expected verdict. 3. Add a validate-before-trade gate to my pipeline as shown in https://alphaassay.com/docs/integrate — branch on failure_codes, stop when the family deflates out (deflated_out_at_n=N), flag survival_map weaknesses. 4. If you issue and store a certificate, verify its Ed25519 signature per https://alphaassay.com/docs/verify-offline. Do not treat an ordinary verdict or the unsigned demo as a certificate. Rule: a pass is a statistical trial result — never a promise of returns. 1 — Send a specimen terminal $ curl -sO https://alphaassay.com/specimens/golden_lookahead.json $ curl -s https://api.alphaassay.com/v1/assay/demo \ -H "Content-Type: application/json" \ -d @golden_lookahead.json No account, no key — the demo tier is free and rate-limited. All specimens, as direct downloads: golden specimens . 2 — Read the verdict The expected result for this specimen is fail , died_at: net_edge , with the failure code no_net_edge : it is a whipsaw with no net-of-cost edge, so it dies at the first gate before any deeper test runs. Every field of the response is explained in verdicts & the gradient . 3 — Assert the known answer Assert the specimen's stable semantic fields: verdict , died_at and failure_codes . Volatile metadata, including _meta.as_of (the wall-clock time of your call), may change. That is enough to catch us being wrong before you ever pay. What does the full response contain? The response is JSON in the gauntlet.v1 shape: the verdict , the stage it died_at (or null on a pass), machine-readable failure_codes , the graded stages each with their own evidence , the family budget , and the identifiers spec_hash / family_id . Every field is defined in verdicts & the gradient . response — 200 OK · abridged { "schema": "gauntlet.v1", "verdict": "fail" , "died_at": "net_edge" , "failure_codes": [ "no_net_edge" ], "stages": [ { "stage": "net_edge" , "verdict": "fail" , "evidence": { "net_return_total_pct": 0.0, "trades": 0, "bars": 199 }, "detail": "No positive net-of-cost edge -- dies before any deeper test." } // + funding_edge, family_deflation, power_honesty, significance, cpcv, walk_forward, concentration, placebo, capacity, graveyard_prior ], "budget": { "cumulative_n": null, "n_trials_effective": null }, "spec_hash": "sha256:421f1b…", "family_id": "fam_004b8867…", "_meta": { "engine_version": "" , "as_of": "" } } Field-by-field explanation: verdicts & the gradient . The same call in Python python import requests, json spec = json.load(open("golden_lookahead.json")) r = requests.post("https://api.alphaassay.com/v1/assay/demo", json=spec, timeout=60) v = r.json() assert v["verdict"] == "fail" # known answer assert v["died_at"] == "net_edge" # the planted flaw: no net-of-cost edge assert "no_net_edge" in v["failure_codes"] print(v["died_at"], v["failure_codes"]) …and in TypeScript typescript · node ≥ 18 const spec = JSON.parse(await fs.readFile("golden_lookahead.json", "utf8")); const res = await fetch("https://api.alphaassay.com/v1/assay/demo", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify(spec), }); const verdict = await res.json(); if (verdict.verdict !== "fail") throw new Error("validator failed OUR test"); How do I test my own signal? Paid calls accept your own returns series, equity curve or trade list — a golden specimen file is the canonical request example : what the specimen contains about its signal is exactly what you send about yours. Hosted MCP tools use an api_key ; accountless x402 exists only at POST /x402/v1/gauntlet . Humans and teams can start here . Next: wire it into your agent . ← up All docs next → Verdicts & the gradient ## Golden specimens — AlphaAssay Docs URL: https://alphaassay.com/docs/specimens Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Golden specimens Golden specimens Golden specimens are known-answer test cases — the standing offer to catch us being wrong before you pay us. Each is a prepared signal with a planted property and a known correct verdict. specimen planted property expected verdict golden_clean.json a genuine, persistent edge ✓ pass golden_lookahead.json a whipsaw with no net-of-cost edge ✕ fail · net_edge · no_net_edge golden_cherry.json best-of-many parameter cherry-pick ✕ fail · family_deflation · deflated_out_at_n=50 + BACKTEST_TOO_SHORT_FOR_N=50 golden_thin.json too little data to judge insufficient_evidence Every file above is a direct download, and next to each sits its full expected response ( golden_*.expected.json , e.g. for the whipsaw ) — post the specimen body to POST /v1/assay/demo exactly as downloaded. A demo call without a specimen body returns the built-in 40-trade example in the validate envelope instead; the gauntlet-shaped verdict above appears when you send a specimen. One deliberate exception: golden_thin has too little data for the gauntlet to run at all, so its response comes back in the validate envelope (the verdict object nested under verdict ) — assert against its expected.json rather than assuming the gauntlet shape. When you diff a live response against its expected.json , ignore exactly three paths — _meta.as_of , verdict.provenance.generated_at and explanation.provenance_attestation.generated_at (top-level provenance in the gauntlet shape) — because they are timestamps of the call itself; every other byte, including input_digest and request_hash , is deterministic and bound to the _meta.engine_version recorded inside the fixture, so a mismatch elsewhere means the engine changed and the fixture republication is owed, not that your comparison is wrong. These four rows are the known-answer section of the public Signal Validation Benchmark — the same specimens, scored in the open. What do the golden specimens prove? Together, the four cover the three things you should demand of any validator: it catches real flaws (the costless whipsaw, the cherry-pick), it doesn't cry wolf (clean passes), and it admits when it cannot know (thin abstains). A validator that only ever says no is as useless as one that only says yes. How do I use golden specimens in CI? Run them via POST /v1/assay/demo ( quickstart ), free — and assert against the envelope the specimen actually returns, because two of them differ: golden_lookahead, golden_clean, golden_cherry → gauntlet envelope: assert the top-level verdict , died_at and failure_codes . golden_thin → validate envelope, because there is too little data for the gauntlet to run at all: assert the nested verdict.verdict ( insufficient_evidence ). There is no top-level died_at or failure_codes here — asserting them will fail, and that is the specimen documenting the abstention, not a defect. Each specimen ships its full expected response next to it ( *.expected.json ), so the safest assertion is a comparison against that file with one exception: _meta.as_of is the only field that varies between two identical calls — everything else, including the digests, is deterministic and can be compared byte for byte. ← previous Failure codes next → The calibration record ## The 21 tools — AlphaAssay Docs URL: https://alphaassay.com/docs/tools Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / The 21 tools The 21 tools The hosted public AlphaAssay MCP server exposes 21 tools: 6 free and 15 metered — currently $0.05 per completed check on every paid tool, as published by /v1/meta/pricing . The local stdio registry contains a separate catalog of 50 operator/research tools; it is a development surface, not the customer storefront and not the number registries should advertise. Public assay tools are demote-only — they can devalue a claim, never bless one — and none of them places orders, touches a broker or gives buy/sell advice. You send derived evidence (returns, trades, candles you choose); raw payloads are not retained — the ledger keeps a one-way fingerprint, the verdict, summary statistics and a coarse return sketch for family accounting, and a pre-registered spec is stored by design ( the complete retention ledger ). This page is the plain-language version of each tool's job; the machine-readable descriptions live in the server itself. How do you connect? The server speaks streamable HTTP at https://mcp.alphaassay.com/mcp and is published in the official MCP registry as com.alphaassay/mcp . Clients that expect a local stdio server can bridge it with one line: npx mcp-remote https://mcp.alphaassay.com/mcp . No SDK is required either way. REST overlaps with many operations but is not a one-for-one wrapper around every tool ( API and transport parity ); x402 is only the REST route POST /x402/v1/gauntlet ( how that route pays ). Paid MCP tools take an api_key . Create or manage one at api.alphaassay.com/account ; any signed-in account allowance and the current MCP charge are published by the live pricing registry . Which tools are free — and why? 6 tools cost nothing, permanently, because they are how you audit us before paying: the demo, the public records, the offline check, the payload lint and the due-diligence protocol. Trust infrastructure does not belong behind a meter. assay_demo — see a full verdict before sending anything Runs the real fail-closed validator with no auth, no payment and no signup. It answers in two shapes, and which one you get depends on what you send — assert against the right one: No body (or an empty one) → the built-in 40-trade example in the validate envelope : a nested verdict object with findings , leakage taxonomy, reproducibility readiness and provenance hashes, plus the free synthetic_null_preview — three synthetic no-edge paths (stochastic volatility, jump diffusion, drift bursts), ground truth „no edge", so you can watch the process placebo work before paying anything. A golden specimen as the body → the gauntlet envelope : top-level verdict , died_at , failure_codes , stages and budget . This is the shape the specimens document their known answers against. Both are real validator output; the difference is the surface, not the rigour. The demo is unsigned and deterministic: between two identical calls _meta.as_of is the only field that changes — request and input digests are derived from the payload, so they stay put and can be compared byte for byte. assay_graveyard — has this idea already died? Anonymised mortality statistics per structural signal family: tested, killed, survived, top causes of death, and the crowd prior your submission would be deflated by. Check before you spend weeks on an idea whether the crowd already buried it. k-anonymous — families with fewer than 5 distinct submitters return only coarse taxonomy stats. assay_calibration — what does the current disclosure contain? Calibration v0 exposes only a privacy-bucketed count of evaluated mature registrations, forward_outcomes.status=accumulating , and honesty=insufficient_history until a separately defined mature-outcome metric exists. It does not yet publish a hit rate, Brier score or survived/deflated outcome score. A valid snapshot contains its signature fields; an unavailable or stale snapshot reports signed:false and a reason instead of presenting unsigned data as trusted. The format is described in how we grade ourselves . This document is population/status evidence only; it does not establish examiner quality or outcome accuracy. assay_certificate_verify — is this certificate trusted by the platform? Fail-closed verification of an Ed25519-signed AlphaAssay certificate against the service's externally pinned signed keyring and complete revocation-head history. A caller-supplied public key is evidence, never a trust root; a full valid result also requires a retained certificate-purpose key and no matching revocation. Use it on any certificate someone attaches to a signal they are selling. assay_provider_protocol — judge any signal vendor or validator with seven tests The falsification protocol as machine-readable rules: provenance, survivorship, pre-registration, placebo, costs, trial accounting, examiner — each with a machine-checkable failure condition, and tests 3–6 name the endpoint that automates them. It hands your agent the checklist; it does not rate, score or rank any provider, and test 7 applies the whole protocol to us. assay_preflight — lint the payload before you spend a check Checks the form of a submission — DSL schema, OHLCV sanity (finite positive prices, aligned series, strictly increasing timestamps), trade-row types — in the same failure-code vocabulary the paid tools use, and warns honestly when the sample sits below the engine's evidence floor. A clean preflight is not evidence of an edge; it only means the trial can run — so no paid check ever dies of a typo. Which tools put a signal on trial? The trial group interrogates a strategy or a claim about one, each tool from a different angle, all priced by the live registry — a fail costs the same as a pass, because you are buying the trial, not the outcome. assay_signal — the fail-closed verdict on your export Send a trade list, equity curve or QuantConnect/LEAN export — crypto, stocks, futures, FX alike — and get pass / conditional / fail / insufficient_evidence with the evidence attached. Survival now demands two thresholds at once (expected-max-Sharpe deflation AND multiple-testing significance), every fail names its killer in one plain sentence ( died_at_plain ), and the deflation haircut is decomposed: how much was your search, how much your fat tails ( dsr_attribution ). Declare your trial count honestly — omitting it earns the named finding TRIALS_UNDISCLOSED instead of a silent benefit of the doubt. Optionally attach a pit_evidence attestation about your data's point-in-time discipline: opt-in and demote-only — only an explicitly computed material look-ahead hard-fails, everything inconclusive changes nothing. assay_forensics — WHY it fails, not just that it fails Upload decision timestamps plus candles and get the leak named: does the edge collapse under a one-bar execution delay (look-ahead)? Does the move happen before your decision (front-loading)? Does random timing with your trade structure do just as well? The placebo evidence is three-dimensional — timing, sign and chronology. Naive timestamps are rejected outright; timezone ambiguity is the top source of fake edges. assay_backtest — a backtest that remembers your retries Define the strategy as an executable JSON DSL (sma, ema, rsi, atr, roc, zscore + logic ops), supply your own candles, and get net-of-cost returns plus a family-deflated verdict. Every call lands in your family's trial ledger — with near-duplicate variants collapsed to their effective count , so honest exploration is cheap and quiet grinding is priced. That memory is the feature: it keeps your next verdict meaningful. assay_batch — the whole sweep, tested honestly in one call Up to 25 DSL variants — a list, or base_spec plus a parameter grid with canonical expansion order — and every variant becomes a family-ledger trial verdicted under the cumulative deflation of its siblings: later variants see the budget the earlier ones spent, which is the whole point. The report is demote-only by construction: survives/deflated_out counts and per-variant verdicts, never a ranking, never a „best pick". Metered per variant, individually journaled; if the balance runs short you get what was paid for plus an explicit declined counter — no silent truncation. assay_gauntlet — the whole battery, one dossier Validator, family deflation, honesty stamps, purged combinatorial time-partitions ( cpcv ), the 500-twin placebo, the capacity ceiling and the graveyard prior — chained into one response that says which gate killed it first , at what placebo percentile, at what tradable size („your edge dies at $X"), and how much search budget your family has left ( break_even_n ). Built for iterating agents: training feedback with budget economics, not a bare yes/no. assay_falsify — we actively try to kill it An adversary runs eight attacks — execution-lag push, cost stress, time jackknife, regime split, parameter-neighbourhood perturbation (now with a full stability_surface : the Sharpe terrain around your optimum, because a lonely spike is an overfit fingerprint), cpcv partition, a synthetic-null placebo (does the strategy beat no-edge worlds built with stochastic volatility and jumps?), and a drift-burst strip (remove the flash-crash bars whose PnL dwarfs local volatility — an edge that dies without them was rarely harvestable at quoted prices) — and returns the survival map : what kills it first, what it withstands and up to what limit. The graveyard sharpens the attack order: families that usually die of costs get cost-stressed first. Surviving everything is the strongest robustness evidence this platform can give — still not a profit promise. assay_pbo — did your sweep find an edge, or manufacture one? Submit the full T×N trial matrix of a parameter sweep and get the Probability of Backtest Overfitting via combinatorial purged cross-validation, with degradation slope, probability of out-of-sample loss and stochastic dominance of your picks versus the pool; PBO ≥ 0.5 earns the named demote PBO_HIGH . The pairing matters: the trial ledger counts how many tries your family burned — PBO grades whether the selection process itself is overfit. The statistic, explained . assay_reproduce — audit the arithmetic, not the story A different audit object: not the signal, the caller's calculation . Send trades, the candles they were filled on, and the claimed headline metrics — the engine independently rebuilds the equity book and grades each claim against disclosed per-metric tolerances (published cross-engine divergence reaches 3.71%, so the band is explicit output, never a hidden judgment). Every fill is checked against its bar's low–high range — a fill outside it was never physically available ( FILL_OUTSIDE_BAR_RANGE ) — and stop/limit exits that OHLC data cannot order are priced worst-case ( FILL_AMBIGUITY_MATERIAL ). It audits whether the numbers follow from the trades; it does not judge whether the strategy is good. assay_tradelog — interrogate a fill log against itself No candles needed: this audit asks whether a raw fill log is internally consistent. It hunts exact duplicate fills (double counting inflates every headline number), time travel — exits before entries, or decisions stamped after the entry they supposedly caused, the anatomy of a look-ahead — PnL that contradicts the row's own prices, naive or future timestamps (ambiguous clocks are the top source of fake edges), and fills booked outside the declared trading calendar. Where assay_preflight lints shape and assay_reproduce audits arithmetic against candles, this one needs nothing but the log itself. assay_survivors — which variants survive family-wise error control The same T×N matrix assay_pbo grades, answered variant by variant: could this family's evidence kill it at family-wise error rate α, and in which stepdown round — Romano-Wolf stepwise multiple testing with a circular block bootstrap that respects serial dependence. The output is an error-budget disclosure, never a ranking: fail means nothing survives ( NO_SURVIVORS_AT_FWER ), conditional means survivors are disclosed — explicitly not a pass — and thin data blocks with a named reason. No sorting, no seal, no recommendation. assay_cpcv — the full purged-CV distribution, not one number The gauntlet's cpcv stage answers a single question; this tool hands over everything behind it: the annualized Sharpe of every purged combinatorial half-partition of your history — as quantiles and a histogram — the number of recoverable backtest paths, purge/embargo accounting, and the demote line: when too few of the purged partitions show a positive Sharpe, the edge lives in a handful of lucky segments rather than the whole record, and the result earns cpcv_unstable . Send a net-return series directly, or a DSL spec plus candles — the series is then computed with the same primitives the gauntlet uses, so the distribution describes the same strategy the verdict judged. Which tools build a pre-registered record? Two tools implement pre-registration. Retaining the original commitment gives a reviewer a concrete value to compare with later publications; independent time and immutability claims additionally require an independently controlled anchor. Pre-registration is not immune to trying many variants, which is why every registration still counts against your family's deflation budget. assay_register — seal the call before the data exists Your strategy spec is canonically hashed with its service acceptance time into the operator-published Merkle/JSONL chain. Registrations are idempotent, and withdrawing one still counts against your family's budget — pre-registering ten variants and deleting nine is the oldest trick in the book, and it is priced in. assay_verdict — the post-cutoff test, then a certificate that travels Evaluates a registration strictly on bars after its cutoff, with maturity floor and fail-closed gap handling — and even an anchored signal is still deflated by how much its family searched. A separate certificate-issuance action exports the result as a signed artifact (current price in the live registry ; verification free for everyone ) that embeds the full trial_disclosure : N, effective N, sample length and the threshold curve — the numbers a skeptical reader needs to audit the deflation are inside the signed document itself. assay_var_es — does your risk forecast survive contact with reality? The other tools test return claims; this one tests risk numbers. Submit the VaR/ES forecasts your model produced ex ante alongside the realised returns, and the exceedance backtest checks whether losses breached your own thresholds more often than your stated confidence level allowed. A risk model that is quietly too optimistic is a claim like any other — so it gets tested like one, not taken on trust. Optionally attach a naive benchmark forecast (say, a rolling historical quantile): a Diebold–Mariano test on a strictly consistent loss then asks whether your model actually beats the benchmark it claims to improve on — losing to it decisively is the kill RISK_FORECAST_DOMINATED_BY_BENCHMARK , because sophistication that underperforms naivety is theatre. assay_conformal — does your confidence label survive the outcomes? The third kind of claim after returns and risk: coverage. Submit the prediction intervals your model emitted — lower and upper bound per case — and the audit measures how often the truth actually landed inside them, against the coverage you claimed. An interval that promises „90%" and delivers seventy is not calibrated confidence, it is a story; this is the test that tells the two apart. What do all 21 have in common? Demote-only: evidence can kill a claim, never inflate one, so there is no „verified winners" list to raid. Pure/read results require an explicit as_of for reproducibility; stateful calls expose their effective timestamp and replay stored output only under the documented idempotency contract. Fail-closed: when data is too thin or a model leaves its validity range, the answer says so instead of guessing. And every fail arrives with its named cause of death . Wire the discipline into your agent once — copy-paste rules — and check prices against the live source on the pricing page . A methodology audit, not investment advice. ← previous Rules for your agent next → API overview ## Validate a signal — AlphaAssay Docs URL: https://alphaassay.com/docs/validate Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Validate a signal Validate a signal Validation is the core product: your signal goes through the full battery and comes back with a structured verdict. Three input shapes are accepted — a returns series, an equity curve, or a trade list. Issue a certificate separately when the result needs a signed, portable artifact. What do I send to validate a trading signal? JSON with your series and its timestamps, plus minimal context (instrument class, intended trading frequency, costs assumption if you have one). The demo endpoint uses the identical shape — inspect a golden specimen file to see it: what the specimen contains is exactly what you would send about your own signal. How is the signal tested? The battery runs the four gate families in order: net edge after realistic costs → deflation for every variant your family has tried → a placebo trial against 500 random twins → capacity and robustness attacks. The verdict names the first gate that killed it. Why does re-testing a tweaked strategy cost statistical power? Variants of one idea share one statistical budget. Testing „the same signal but with lookback 21 instead of 20" is not a fresh signal — and pretending it is would be exactly the overfitting trick the battery exists to catch. A variant is still a full call at the amount in the live pricing registry , and it still adds to the family's budget ( n_trials_effective ) — pricing may change; the statistical accounting does not. When the family has been tried too often, the honest answer about another tweak is already fail ( deflated_out_at_n=N ) — we tell you instead of taking your money for theater. Is my strategy stored? Your raw inputs, rules and code are processed for the trial and not retained. What the ledger keeps, per trial: a one-way fingerprint, the verdict with its cause of death, summary statistics (Sharpe, sample length, test period), a 32-number compressed sketch of the return profile — it powers the family accounting and is far too coarse to reconstruct trades or rules — and the family's structural label with its parameter coordinates (window sizes, horizon) for the anonymised graveyard. The one deliberate exception: a spec you pre-register is stored in full, because sealing a claim means storing it. Details: trust . ← previous REST x402 gauntlet next → Pre-register a call ## Verdicts & the gradient — AlphaAssay Docs URL: https://alphaassay.com/docs/verdicts Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Verdicts & the gradient Verdicts & the gradient An AlphaAssay verdict is a diagnosis, not a yes/no oracle. Four outcomes exist: pass , conditional , fail and insufficient_evidence — and every one of them ships with the evidence that produced it. What does a verdict contain, field by field? Every gauntlet verdict carries the same parts: a schema tag, the verdict itself, the stage it died_at , machine-readable failure_codes , the graded stages with their evidence , the family budget , and the identifiers spec_hash / family_id — each defined in the table below. field meaning verdict pass · conditional · fail · insufficient_evidence. „Insufficient evidence" is an honest abstention — too little data to judge is an answer, not a failure of ours. died_at the first stage your signal failed — any of the eleven graded stages below, e.g. net_edge , funding_edge , placebo or capacity (see the gates ). null on a pass. failure_codes machine-readable causes your agent branches on, e.g. no_net_edge or deflated_out_at_n=N . Catalog: failure codes . stages the graded stages, in order — net_edge , funding_edge , family_deflation , power_honesty , significance , cpcv , walk_forward , concentration , placebo , capacity , graveyard_prior — each with a verdict (pass / fail / skipped / info) and its own evidence . funding_edge charges the perpetual funding leg against every holding period — carry income is not timing skill; power_honesty and significance are honesty stamps (info), cpcv tests the record across purged combinatorial time-partitions, walk_forward checks whether the in-sample fit carries into anchored out-of-sample folds, concentration asks whether the book survives without its single best bars, and the placebo stage's evidence.percentile is your timing against 500 random twins; 50 means „indistinguishable from chance". budget cumulative_n , n_trials_effective and break_even_n : the family's cumulative search budget — with near-duplicate variants collapsed to their effective count ( effective_n_method ) — plus the trial count at which this very result would deflate out. Ends the endless tweaking loop with a number. spec_hash · family_id content-derived identifiers of the signal and its family. For an evaluated pre-registration, a separate REST certificate-issuance action adds an Ed25519-signed artifact; the free demo verdict is unsigned. When a signal passes the gauntlet, the forensics report adds an adversarial survival_map — eight attacks (execution lag, cost stress, time jackknife, regime split, parameter neighbourhood, cpcv partition, synthetic null, drift burst), each marked survives or dead. It is part of the forensics path, not the free demo verdict. Why is a fail worth paying for? A no without a finding is an annoyance. A no with died_at and failure codes is a shortcut: it tells you whether to fix the costs, stop tweaking the family, or abandon the direction entirely — before real money finds out for you. When is a result reproducible or replayed? Pure/read operations are reproducible only with the same explicit as_of . Stateful calls expose their effective timestamp and may reflect cumulative tenant state. An exact stored stateful replay is promised only when the same non-empty request_id and the same canonical request return the stored response. For the unsigned golden demo, assert the documented semantic answer fields rather than the whole envelope. ← previous Quickstart next → REST x402 gauntlet ## Verify a certificate offline — AlphaAssay Docs URL: https://alphaassay.com/docs/verify-offline Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / Verify a certificate offline Verify a certificate offline Raw cryptography and platform trust answer different questions. Ordinary verdicts and the unsigned demo are not certificates. The public reference verifier computes both checks and names them separately. raw_signature_valid : does this exact document match this signature under this public key? The key displayed on the same website is evidence, not an independent trust root. platform_valid additionally means that the key is a retained certificate-purpose key in an externally pinned signed keyring, the complete revocation-head history reaches the required checkpoint, and no matching revocation exists. A full offline verifier therefore needs that independently provisioned root pin and signed history as well as the certificate. If any trust evidence is missing, stale or incomplete, platform verification must fail closed. These two names describe the checks; the hosted API reports the platform result as valid and status . 1 — Save the candidate public key for the raw check # Key-ID: key_9961d9e3190d69a6 -----BEGIN PUBLIC KEY----- MCowBQYDK2VwAyEA3pvfyosExDvDYzg8R1YMBu//7RvUWvuK/uanMlIJMQg= -----END PUBLIC KEY----- 2 — Split the document The delivered envelope contains the inner certificate object and the signature_b64 value. Save the object as valid JSON in certificate.json and the base64 string in signature.txt . Transport whitespace and object-key order are not signed; the verifier first applies json-sort-keys-ascii-nospace-nfc-float12 . 3 — Full offline platform verification terminal $ pip install alphaassay-verify $ python -m alphaassay_verify verify certificate.json signature.txt trust_bundle.json --root-pin 64-lowercase-hex-root-fingerprint { "valid": true, "signature_valid": true, "chain_valid": true, "status": "valid" } The root fingerprint must come from an independent release channel, not from the bundle it authenticates. The local signed bundle must contain the complete revocation-head history and satisfy your minimum trusted checkpoint. If you do not possess those artifacts, fail closed or use the hosted platform verifier; do not substitute the public key printed by this page. 4 — Raw signature evidence only terminal $ python -m alphaassay_verify signature-only certificate.json signature.txt alphaassay.pub { "valid": true, "reason_code": "valid" } Here valid:true means only raw_signature_valid=true . The public verifier uses canonical_bytes internally; raw OpenSSL or Node crypto over the downloaded JSON file is not equivalent unless it reproduces the named canonicalization and passes the published golden vectors. Library integration python · full trust from alphaassay_verify import verify_trusted_certificate result = verify_trusted_certificate( certificate, signature_b64, public_key_pem=None, trust_bundle=local_bundle, pinned_root_fingerprints=(independent_root_pin,), minimum_head_sequence=trusted_checkpoint_sequence, minimum_head_hash=trusted_checkpoint_hash, ) assert result["valid"] is True A semantic modification — a changed value or key — makes signature verification fail. Passing only the raw mode does not establish who controls the key or whether the certificate was revoked. The hosted /verify page applies the platform trust bundle and returns a valid result only after the keyring, history, checkpoint, purpose and revocation checks all pass. ← previous The calibration record next → What we keep ## What we keep — AlphaAssay Docs URL: https://alphaassay.com/docs/what-we-keep Docs GETTING STARTED Quickstart Verdicts & the gradient REST x402 gauntlet GUIDES Validate a signal Pre-register a call Certify & share Wire it into your agent Rules for your agent REFERENCE The 21 tools API overview Failure codes AUDIT US Golden specimens The calibration record Verify offline What we keep Docs / What we keep What we keep This page is the complete retention inventory of AlphaAssay — everything the service keeps, per path, and everything it never keeps. It exists because the strongest attack on any validator is „they can see your strategy — they could trade on it." The honest answer is not a promise of good behaviour; it is an inventory you can size yourself, an architecture with no trading arm, and proofs you can run. If you ever find this page contradicted by the live service, that finding is exactly what we ask you to hunt for . The retention ledger, path by path path kept never kept free demo nothing — the demo runs on a throwaway store that is deleted when the call ends everything validation calls (signal · backtest · gauntlet · forensics · falsify · batch) a one-way fingerprint, the verdict with its cause of death, summary statistics (Sharpe, sample length, period), a 32-number compressed sketch of the return profile (family accounting — retention is caller-chosen, see below), and the family's structural label with its parameter coordinates raw candles, trade lists, return/equity series, rules, code assay_pbo nothing from the matrix — the trial matrix is processed in memory and not persisted the matrix itself assay_reproduce no ledger entry at all — an arithmetic audit is not a new trial trades, candles, claimed metrics pre-registration the sealed spec in full (canonical JSON) plus the sealed success criteria, included in the operator-published Merkle/JSONL chain; independent anchoring status is disclosed separately — x402 gauntlet payments per settled POST /x402/v1/gauntlet payment: nonce, wallet, transaction id, input digest and cached response. The named receipt binds payment/input/verdict digests and declares signed:true with Ed25519 fields or signed:false when no deployment key is configured (no expiry — one payment buys one result) raw inputs — responses never contain them certificates purchased certificates in full (signed, incl. trial disclosure), account-bound for re-delivery — operations usage charges (tool + cents), per-minute tool statistics, sanitised error messages, rate-limit buckets, request logs (id, timestamp, endpoint, status) raw payloads by contract; a release canary injects a marker through the covered logging, error, metric and alert sinks. That test is scoped regression evidence, not a universal proof of every possible sink The same inventory, with legal bases, is in the privacy policy . The 32-number sketch is bucket means of the return series — far too coarse to reconstruct trades or rules; its only job is proving two variants were near-duplicates so the family budget can count them as one. And because the sketch is the only evidence of that nearness, you may refuse it: pass sketch_opt_out on assay_gauntlet , assay_backtest or assay_batch (REST: /v1/gauntlet ) and the sketch is never persisted — the response echoes sketch_retained: false . The price of that choice is statistical, not monetary: without the evidence of near-duplication, the trial counts in full toward the family budget — no evidence, no discount. By construction the opt-out can only make verdicts stricter, never friendlier. Could you front-run what you see during a trial? Head-on: during a trial, the engine necessarily holds the submitted evidence in memory — a remote examiner that sees nothing can test nothing. The service has no exchange connectivity, order path, execution logic, broker connection or custody, so it cannot execute a trade through the product. That is a bounded system capability claim, not a universal guarantee against operator misuse. Ordinary working payloads are transient, while the retention ledger above lists every documented state class; notably, a pre-registered spec is stored in full. Verdicts remain demote-only, so no product-generated list of „verified winners" exists to raid. What about the one path where we DO store your strategy? Pre-registration stores the sealed spec — that is the product. Its canonical hash and service acceptance time enter an operator-published Merkle/JSONL chain. If you retain the registration response or chain head, later divergence becomes detectable against that copy. While external_anchor is empty, this is not independent proof of priority or an independently trusted timestamp; the storage risk remains part of the disclosed trade-off. How do you check any of this, instead of trusting it? Run the golden specimens (free, with documented stable semantic fields). Pull the current persisted calibration population/status snapshot . Calibration data is returned only inside a valid signed snapshot; otherwise the endpoint says signed:false with a stale/reason state and omits the calibration payload. Use the offline verification guide to distinguish raw_signature_valid from full platform_valid , which needs independently rooted key and revocation history. Apply the falsification protocol to us before anyone else. And hold this page against the privacy policy and the live service — one truth, or it is a finding. ← previous Verify offline next → Test my signal ## Glossary — the AlphaAssay vocabulary, defined URL: https://alphaassay.com/glossary GLOSSARY The AlphaAssay vocabulary, defined The terms this platform coined, or uses in a precise and non-obvious way — one definition each, stable and quotable. For the name collision with laboratory assays, see AlphaAssay vs. "alpha assay" . DEFINED TERMS One line per term term definition assay office (for trading signals) AlphaAssay's positioning: an independent office that grades signals and never trades them — no execution logic, no custody, no order path. golden specimen a known-answer test case with a planted flaw, used to verify AlphaAssay's correctness for free before paying anything. docs/specimens died_at the exact stage that first killed a signal in the assay battery — a diagnosis, not a bare fail. docs/verdicts failure code the machine-readable reason a signal failed, designed for an agent to branch on. docs/failure-codes family budget how many meaningful tries remain for a strategy family before one more variant cannot be distinguished from luck. placebo percentile a signal's timing skill measured against 500 matched random twins with the same trading profile; 50 means indistinguishable from chance. calibration record the public snapshot of a bucketed evaluated-mature-registration count plus accumulating/insufficient-history state — signed when valid and otherwise explicitly untrusted. It is not yet an outcome score. trust demote-only verdict a verdict that can devalue a signal but never bless one: new evidence only ever lowers a grade. Each definition is the canonical one — registries, docs and listings reuse these sentences verbatim, because a vocabulary that drifts between surfaces stops being a vocabulary. See the vocabulary in a live verdict. $ test my signal read a verdict, field by field ## AlphaAssay vs. „alpha assay“ (chemistry) — which one you mean — AlphaAssay URL: https://alphaassay.com/glossary/alphaassay-vs-alpha-assay GLOSSARY · DISAMBIGUATION AlphaAssay vs. "alpha assay" in chemistry Three unrelated things share this token. This page tells them apart — for people who landed here from a search, and for retrieval systems that would otherwise merge three entities into one. THE SHORT ANSWER Which one are you looking for? If you want to test whether a trading signal, strategy or backtest is statistically real, you want AlphaAssay — this site. If you came here for a laboratory technique, you are after one of two unrelated things: the ALPHA assay (AlphaScreen / AlphaLISA, a bead-proximity luminescence method in biochemistry) or the gross alpha assay (total alpha-particle radioactivity, radiochemistry). The table below separates all three. THE COMPARISON Three fields, one word AlphaAssay (this site) ALPHA assay — AlphaScreen / AlphaLISA gross alpha assay field quantitative finance, algorithmic trading biochemistry, drug discovery radiochemistry, environmental testing measures whether a trading signal's edge is statistically real — out of sample, against chance, against overfitting biomolecular interactions via singlet-oxygen bead-proximity luminescence total alpha-particle radioactivity in a sample typical input a trading signal, backtest or trade log donor and acceptor beads plus analyte in a microplate a water, soil or air sample typical output a structured pass/fail verdict with machine-readable failure codes; an optional certificate is a separate signed artifact luminescence counts activity in pCi/L or Bq/L spelling tell one word, CamelCase: AlphaAssay ALPHA is an acronym (Amplified Luminescent Proximity Homogeneous Assay) "alpha" names the α particle origin AlphaAssay, 2026 PerkinElmer, today Revvity ISO / EPA reference methods how these three are told apart spelling AlphaAssay, one word, CamelCase → the trading-signal validation service (this site) ALPHA as an acronym → the bead-proximity biochemistry method alpha = the α particle → the radioactivity measurement field finance · biochemistry · radiochemistry output structured verdict · luminescence counts · pCi/L (Bq/L) ROUTES OUT Three destinations, one search box ✓ Testing a trading signal or backtest → /start and the quickstart . ✓ The AlphaScreen / AlphaLISA bead assay → the manufacturer's documentation at Revvity (formerly PerkinElmer). ✓ Gross alpha radioactivity in water → EPA method 900.0 and ISO 9696 reference material. AlphaAssay is not related to AlphaScreen — different field, different meaning of "alpha", different company. The name is one word by design, because the space is exactly what pulls a trading-signal service into the laboratory cluster: render it AlphaAssay , never "Alpha Assay". Here for the trading-signal meaning: start below. $ test my signal how the trial works ## AlphaAssay — Is your edge real? Statistical validation for trading signals URL: https://alphaassay.com/ SIGNAL VALIDATION · FOR AI AGENTS AND THE PEOPLE WHO RUN THEM A thousand signals promise an edge. Most are noise. Is yours? AlphaAssay puts every signal on trial — out of sample, against chance, against overfitting — and returns a structured verdict with named failure codes. When the result needs to travel, issue a separately signed certificate. Proof, not promises. $ test my signal 60 seconds to a verdict bucketed evaluated mature-registration population; status remains accumulating/insufficient_history demote-only evidence never inflates a grade ed25519 offline raw-signature evidence detects changed signed bytes; full trust also needs an independently pinned root, signed trust bundle and complete key/revocation history the assay — explained 0:57 watch: how 1,000 signals go on trial ▶ play HAVE AN AGENT? Hand it one prompt. Paste into Claude Code, Cursor or ChatGPT — it wires AlphaAssay into your pipeline. copy prompt read the prompt TRY THE MATH How many of your backtests survive honest math? After enough tries, a great-looking Sharpe ratio is mostly selection luck. The Deflated Sharpe Ratio prices exactly that — run yours in the browser, nothing uploaded. $ try the calculator deflated sharpe · same Sharpe, more tries observed Sharpe 1.8 (3y daily) trials DSR 1 ≈ 1.0 a single honest test 10 ≈ 0.94 a weekend of tuning 50 ≈ 0.80 „I tried a few things" 200 ≈ 0.64 selection luck PRODUCTS One engine. Four ways to use the verdict. 01 Validate Submit a signal or backtest — get a verdict with the reasons attached. $0.05 · LIVE 02 Pre-register Record the hypothesis before evaluation. Retain its operator-published commitment to detect later changes. $0.05 · LIVE 03 Forensics The full autopsy. Know where the edge died. $0.05 · LIVE 04 Certify Issue a signed certificate — credibility that travels. $9.90 · LIVE all products · full pricing · what got stricter, and when METHODOLOGY · NO ORACLE Others tell you if . We tell you what broke. A no without a finding is an annoyance. A no with a finding is a shortcut — it saves the most expensive thing you own: weeks of searching in a dead direction. how the trial works diagnosis — strat_4217 · illustrative The 600 tries below are not this caller's own: n_trials_effective is inherited from the anonymised family ledger, where every recorded attempt at the same signal family counts against one shared budget — which is why a first call can already fail deflation. { "schema": "gauntlet.v1", "verdict": "fail" , "died_at": "family_deflation" , "failure_codes": [ "deflated_out_at_n=600" ], "stages": [ { "stage": "net_edge", "verdict": "pass" , "evidence": { "net_sharpe_annualized": 3.11, "trades": 184, "bars": 380, "net_return_total_pct": 41.7 } }, // 380 daily bars ≈ 18 months: a high Sharpe on a short window { "stage": "funding_edge", "verdict": "skipped" , "evidence": { "reason": "no perpetuals in book" } }, { "stage": "family_deflation", "verdict": "fail" , "evidence": { "dsr": 0.31, "cumulative_n": 1, "variants_in_call": 1, "n_trials_effective": 600, "effective_n_method": "family_ledger", "killed_by": "deflated_out_at_n" , "family_verdict": "deflated_out" } } // over 18 months, the best of 600 recorded family tries is expected to look // about this good by chance alone — so 3.11 buys only dsr 0.31, not a pass // + power_honesty, significance, cpcv, walk_forward, concentration, placebo, capacity, graveyard_prior ], "budget": { "cumulative_n": 1, "n_trials_effective": 600 } // cumulative_n = your own submissions; n_trials_effective counts the whole family's recorded tries } THE PUBLISHED RECORD Most signals fail. That is not our opinion. Decades of peer-reviewed finance say most profitable-looking strategies are selection noise, not skill. Before your edge risks a cent, that is the base rate it is up against — and the published record we built the battery on. „Most claimed research findings in financial economics are likely false." HARVEY, LIU & ZHU · REV. FINANCIAL STUDIES 2016 Published strategies lose over half their returns — ≈−26% out-of-sample, ≈−58% post-publication. McLEAN & PONTIFF · JOURNAL OF FINANCE 2016 A few dozen tries manufacture „great" backtests out of pure noise — overfitting is mathematically guaranteed. BAILEY, BORWEIN, LÓPEZ DE PRADO & ZHU · AMS 2014 Of ~2.1 million systematically tested strategies, almost none survive correct multiple-testing correction. CHORDIA, GOYAL & SARETTO · RFS 2020 what the published record says TRUST No trading arm. Retention disclosed. T1 Your strategy stays with you Validation runs keep no raw inputs, rules or code — only a one-way fingerprint, the verdict with its cause of death and coarse trial statistics. The one sealed exception is disclosed, not hidden. T2 No honeypot Our verdicts can devalue a signal, never crown one. There is no list of winners to raid. T3 No trading arm No exchange access, execution, custody, broker or order path exists in the service. Pre-registration storage remains disclosed separately. retention and no-trading boundaries · the complete retention inventory Was your edge ever real? Find out before it costs you. $ test my signal verify a certificate ## Legal Notice — AlphaAssay URL: https://alphaassay.com/legal LEGAL · IMPRINT Legal Notice Who operates AlphaAssay — provider identification under Swiss law. Provider identification · Switzerland The website alphaassay.com , the API at api.alphaassay.com and the MCP server at mcp.alphaassay.com are operated by: Company AlphaAssay Address 6403 Küssnacht am Rigi, Switzerland Contact hello@alphaassay.com Nature of the service AlphaAssay is an independent statistical validation service for trading signals — methodology audits, not investment advice, with no order path and no custody. See the terms of service and the privacy policy . Responsibility for content We take care that the information on this site is accurate and current, but give no warranty as to its completeness or accuracy. Links to external sites are provided for convenience; we are not responsible for their content. ## Methodology — AlphaAssay URL: https://alphaassay.com/methodology METHODOLOGY Every signal goes on trial. Here is the courtroom. The battery follows a fixed gate sequence. Pure/read operations require the same explicit as_of for reproducibility; stateful calls expose their effective timestamp and may reflect cumulative tenant state. No model moods, no black box — statistics you can name, in an order you can audit. THE FOUR GATES The four gates GATE 1 Net edge Does anything remain after realistic costs, slippage and delay? Most signals end here — the edge was an artifact of frictionless simulation. GATE 2 Family deflation How many variants did you (or the world) try before this one? We deflate the score for every attempt — luck compounds fast when you keep rolling dice. GATE 3 Placebo trial Your signal competes against 500 random twins with the same trading profile. If random timing does just as well, your timing wasn't the edge. GATE 4 Capacity & robustness Survives with chunks of history removed? Across regimes? With parameters wiggled? At the size you'd actually trade? Under the hood the four gates unfold into eleven graded stages, run in trial order — net_edge → funding_edge → family_deflation → power_honesty → significance → cpcv → walk_forward → concentration → placebo → capacity → graveyard_prior — and the verdict names the first stage your signal failed , so you know exactly what to fix, or when to stop. That is the difference between a diagnosis and an oracle. see the whole battery graded in public — the benchmark THE GRADIENT What a verdict actually contains. ✓ died_at — the first failed gate. ✓ failure_codes — machine-readable causes your agent can act on. ✓ Placebo percentile — one number instead of a gut feeling. ✓ Survival map — eight live attacks, each marked survived or dead. ✓ Search budget — how many honest tries your family has left. Ends the endless tweaking loop. ✓ Session history — „your last three variants all died of costs." A learning curve, not isolated verdicts. diagnosis — strat_4217 · illustrative The 600 tries below are not this caller's own: n_trials_effective is inherited from the anonymised family ledger, where every recorded attempt at the same signal family counts against one shared budget — which is why a first call can already fail deflation. { "schema": "gauntlet.v1", "verdict": "fail" , "died_at": "family_deflation" , "failure_codes": [ "deflated_out_at_n=600" ], "stages": [ { "stage": "net_edge", "verdict": "pass" , "evidence": { "net_sharpe_annualized": 3.11, "trades": 184, "bars": 380, "net_return_total_pct": 41.7 } }, // 380 daily bars ≈ 18 months: a high Sharpe on a short window { "stage": "funding_edge", "verdict": "skipped" , "evidence": { "reason": "no perpetuals in book" } }, { "stage": "family_deflation", "verdict": "fail" , "evidence": { "dsr": 0.31, "cumulative_n": 1, "variants_in_call": 1, "n_trials_effective": 600, "effective_n_method": "family_ledger", "killed_by": "deflated_out_at_n" , "family_verdict": "deflated_out" } } // over 18 months, the best of 600 recorded family tries is expected to look // about this good by chance alone — so 3.11 buys only dsr 0.31, not a pass // + power_honesty, significance, cpcv, walk_forward, concentration, placebo, capacity, graveyard_prior ], "budget": { "cumulative_n": 1, "n_trials_effective": 600 } // cumulative_n = your own submissions; n_trials_effective counts the whole family's recorded tries } DEMOTE-ONLY Demote-only: evidence can only make things worse A verdict is never upgraded after the fact. New evidence can lower a grade — a data vendor restates prices, a leak is discovered — but nothing can inflate one. That removes the strongest temptation any rating business faces. It also means: when something we certified turns out to be wrong, we say so publicly and the certificate shows as revoked on /verify . NAMED STATISTICS The statistics have names Nothing in the battery is proprietary magic. The core is the published state of the art for separating skill from luck under multiple testing: „Most claimed research findings in financial economics are likely false." HARVEY, LIU & ZHU · REV. FINANCIAL STUDIES 2016 Published strategies lose over half their returns — ≈−26% out-of-sample, ≈−58% post-publication. McLEAN & PONTIFF · JOURNAL OF FINANCE 2016 A few dozen tries manufacture „great" backtests out of pure noise — overfitting is mathematically guaranteed. BAILEY, BORWEIN, LÓPEZ DE PRADO & ZHU · AMS 2014 Of ~2.1 million systematically tested strategies, almost none survive correct multiple-testing correction. CHORDIA, GOYAL & SARETTO · RFS 2020 the primer: why backtests flatter everyone INSIDE THE CELLS The same question, asked from independent directions. The named statistics above are the skeleton. Inside the stages, several cells now ask the same question more than once, because an edge that only survives one framing was never an edge. Every field below is delivered in the verdict envelope and documented in the live tool handshake. C1 Deflation, from three sides The Deflated Sharpe Ratio asks whether your score beats the best of N random tries; the Benjamini–Hochberg-adjusted p asks whether it survives the family of tests around it; and the empirical-Bayes shrinkage exhibit asks the family itself — given everything this strategy family has recorded, how much of a t-statistic this size is signal and how much is noise? The shrunk value ( t_shrunk , shrink_factor ) is an exhibit, never the verdict, and under three usable family trials it reports insufficient_family_history instead of guessing. C2 Three brackets, one mean The significance battery puts three confidence brackets around the same mean return: the percentile bootstrap interval, the studentized bootstrap-t interval — second-order accurate, it stays honest exactly where the first one flatters — and the Newey–West HAC interval, which prices in the autocorrelation real trading returns carry ( se_naive vs se_hac ). When a sharper bracket contains zero while the first does not, the envelope says so ( studentized_ci_contains_zero , hac_ci_contains_zero ): the claimed precision was borrowed from optimistic assumptions, not earned. C3 Carry is not alpha A perpetual-futures backtest that ignores the funding leg overstates its edge, because funding is a cost the exchange really collects. Given a funding-rate series, the funding_edge stage charges it against every holding period and can find two things: the profit only lived in that ignored cost ( funding_erases_edge ), or the price leg loses money and the "edge" is the funding income itself ( edge_is_funding_carry ) — a premium available to any holder of the position, crash-prone, and no evidence of timing skill. Run the trial on yours. $ test my signal how the controls are checked ## Pricing — AlphaAssay URL: https://alphaassay.com/pricing PRICING Published units. Same price for pass or fail. A completed check costs $0.05 whichever paid tool you call, and assay_batch meters each valid processed variant at that same rate. Within either published unit a pass and a fail cost exactly the same, because what you buy is the assay and never the outcome. A separately issued certificate is $9.90 on its own lifecycle unit, the free tier stays free for everyone, and every amount here is the one the live pricing registry serves. ALWAYS FREE Audit us before you ever pay. Six tools never cost anything — the demo with its golden specimens, the graveyard digest, the public calibration record, certificate verification, the payload preflight and the provider-protocol checklist. The free tier exists so you can check that we are any good first. Demo & golden specimens Known-answer test cases with the full response format — catch us being wrong before you pay. Rate-limited, no account. Graveyard & calibration Anonymised family mortality plus the calibration v0 population/maturity state. Machine-readable, no account. Verify, lint, audit us Check any issued certificate with hosted platform trust at /verify ; offline, the raw signature is checkable, while full trust additionally needs the signed bundle and independent pins. Lint a payload with assay_preflight before spending a check; judge any provider — us included — with assay_provider_protocol . Free for everyone, forever. THE PRICE Metering units are explicit. $0.05 per completed check Most paid MCP tools Validate a signal or a backtest, run forensics, audit a claimed track record, pre-register or evaluate a call. All 15 paid tools carry the same amount, so which one you reach for is a question of what you need answered rather than of budget, and assay_gauntlet is the all-in-one whose single check runs every graded stage. Any signed-in account allowance and the subsequent MCP charge come from the live pricing registry ; public MCP tools use an API key. The separate x402 surface is the REST gauntlet route POST /x402/v1/gauntlet . Raw submissions are not retained — the trial keeps a one-way fingerprint, the verdict and coarse trial statistics. what is kept, precisely $0.05 per valid processed variant assay_batch A batch can contain up to 25 variants. Each valid processed variant is journaled and metered individually; declined or structurally invalid variants are disclosed under the runtime contract rather than hidden inside a flat parent-call price. Certificate issuance is a separate REST lifecycle action and costs $9.90. Offline raw-signature checks and the hosted full-trust verifier remain free; full offline platform trust additionally needs an independently pinned root, the signed trust bundle and complete key/revocation history. The live price of every tool is always machine-readable at api.alphaassay.com/v1/meta/pricing . ACCOUNT CREDITS Top up in fixed steps, or spend the allowance first. Every account opens with 3 free checks, which is enough to put a signal you actually care about through a real trial before you decide whether our verdicts are worth paying for. After that, credit packs run $10, $25, $50, $100 or $200, and anything from $100 up carries 5% extra credit. Your exact balance and the allowance still standing are in your account dashboard rather than on this page, because only the runtime knows them. Loading more than that? Talk to us — volume terms are a conversation, not a form: hello@alphaassay.com . FOR AGENTS Machine-readable prices, explicit transports. The live price of every public MCP tool is one GET away, so your agent can gate on cost before it calls with its API key. Accountless x402 is narrower: only POST /x402/v1/gauntlet uses the 402 payment handshake and quote. how x402 payment works terminal $ curl -s https://api.alphaassay.com/v1/meta/pricing { "mcp_tools": { // every canonical assay_* name carries its live amount and billing unit }, "mcp_tool_counts": { "total": 21 , "free": 6 , "paid": 15 }, // assay_batch: per valid processed variant; most paid MCP tools: per check } Excerpt — mcp_tools prices every public tool under its canonical name, mechanically joinable to the live tools/list. The x402 quote applies only to the REST gauntlet route, not to the whole MCP registry. FAIR QUESTIONS Three honest answers. Can I top up an account? Yes — packs of $10, $25, $50, $100 or $200, with 5% extra credit from $100 up. More volume than that is a conversation, not a form: hello@alphaassay.com . What if the verdict is „fail"? You pay the same. You buy the trial, not a flattering outcome — and every fail ships with the diagnosis of what broke, which is usually worth more than a pass. Does a „pass" mean I will make money? No, and we never claim it. A pass means your signal survived a statistical trial that most signals fail. Methodology audits are not investment advice. Start with the free tier. Decide after. $ test my signal how the trial works ## Privacy Policy — AlphaAssay URL: https://alphaassay.com/privacy LEGAL · PRIVACY Privacy Policy What we process, why, and the rights you have — under Swiss and EU data-protection law. Last updated: 7 July 2026 · Swiss FADP (revDSG) and, for EU/EEA users, the GDPR This policy explains what personal data AlphaAssay processes, why, and the rights you have. It applies to the Swiss Federal Act on Data Protection ( revDSG ) and, where we offer the Service to individuals in the EU/EEA, to the GDPR (Art. 3(2) — extraterritorial scope). Who is responsible The controller is AlphaAssay , 6403 Küssnacht am Rigi, Switzerland — see the legal notice . Data-protection contact: hello@alphaassay.com . The short version: your raw inputs and rules are not retained We process your submission in memory for the trial and do not retain the raw inputs, rules or code. What we keep from an assay is the trial's paper trail: a one-way cryptographic fingerprint , the machine-readable cause of death ( died_at , failure codes), summary statistics, a 32-value compressed sketch of the return profile for family trial accounting, and the family's structural label with its parameter coordinates. None of this can be reversed into your trades, rules or code. Two disclosed exceptions exist because their features require them: a pre-registered spec is stored in full (sealing a claim means storing it), and x402-paid responses are cached against their payment nonce (one payment buys one result, forever). The exact inventory is in the table below and, path by path, on what we keep . Data minimisation reduces what remains available after processing; the pre-registered spec and other listed records remain retained, and the service has no broker, exchange or order path. This is a scoped control, not a universal guarantee against operator misuse. What we process, and why data why legal basis Account details (email, and any name you provide) create and secure your account, send service messages contract performance (GDPR 6(1)(b); revDSG contractual) Payment metadata (via Coinbase / x402 or a card processor) take payment and prevent abuse — we do not store card numbers ; the processor handles them contract performance; legitimate interest in fraud prevention (6(1)(f)) API and request logs (request id, timestamp, endpoint, status, coarse IP/rate-limit signal) run the API, debug, enforce rate limits, keep an audit trail legitimate interest (6(1)(f)); legal record-keeping where applicable x402 gauntlet receipts per settled POST /x402/v1/gauntlet payment: the payment nonce, paying wallet address, transaction id, input digest and cached response. The named receipt binds the payment, input and verdict digests; its own fields report whether Ed25519 signing succeeded or the deployment had no signing key. Responses never contain raw inputs. No expiry (the payment claim does not lapse) contract performance Purchased certificates stored in full (signed, including trial disclosure), account-bound, for re-delivery contract performance Your submitted signals / backtests run the assay in memory; retained per trial: a one-way fingerprint, the verdict with failure codes, summary statistics (Sharpe, sample length, test period), a 32-value compressed sketch of the return profile (family trial accounting) and the family's structural label — never the raw inputs, rules or code. Exception by design: a pre-registered spec is stored in full (sealing means storing) contract performance Where it is hosted The Service runs on servers located in the European Union (Germany) . Payment processing runs through our payment providers (e.g. Coinbase for x402/USDC), who process the payment data described above under their own terms. How long we keep it Operational API and request logs are retained for about 90 days , then deleted or aggregated. Account data is kept while your account exists and for as long as needed to meet legal obligations. Verdict fingerprints and public calibration/graveyard statistics are, by design, anonymised and may be kept as part of the permanent public record — they contain no personal data and no recoverable strategy. Who we share with We do not sell your data and do not share it for advertising. We share only what is necessary with payment processors (to take payment) and infrastructure providers (hosting), acting as our processors, and where required by law . There is no other disclosure. No tracking, no advertising cookies This site runs no analytics and no advertising trackers. We set only cookies that are strictly necessary for the site or your session to function. There is no cross-site tracking to opt out of because there is none to begin with. Your rights Under the revDSG and, where it applies, the GDPR, you can request access to your data, and its rectification , erasure , restriction or portability , and you can object to processing based on legitimate interests. To exercise any of these, email hello@alphaassay.com . You also have the right to lodge a complaint with a supervisory authority — in Switzerland the FDPIC (Federal Data Protection and Information Commissioner), or in the EU/EEA your local data-protection authority. Changes We may update this policy; the current version is dated at the top. Material changes take effect when posted here. ## Products — AlphaAssay URL: https://alphaassay.com/products PRODUCTS One engine. Four ways to use the verdict. Everything below runs on the same fixed-sequence test battery. Fresh stateful calls expose their effective timestamp and cumulative state; a stored response replays only under the documented idempotency contract. What differs is what you do with the result. 01 · VALIDATE Put a signal on trial. Send a signal or backtest (returns series, equity curve or trade list). The battery tests it out of sample, against chance and against overfitting — and returns a verdict with the reasons attached: which gate it died at, how close it was, and what survived. ✓ Verdict — pass, conditional, fail, or an honest insufficient_evidence. ✓ Failure codes — machine-readable; your agent acts on them directly. ✓ Placebo percentile — your timing vs. 500 random twins with the same profile. ✓ Replay scope, certificate separately — a stateful response replays only with the same non-empty request id and canonical input; an issued certificate carries the Ed25519 signature. The free demo is unsigned. $0.05 PER COMPLETED CHECK · LIVE PRICE Never stored: the trial keeps a one-way fingerprint, the verdict and coarse trial statistics — your rules and code stay with you. how diagnosis — strat_4217 · illustrative The 600 tries below are not this caller's own: n_trials_effective is inherited from the anonymised family ledger, where every recorded attempt at the same signal family counts against one shared budget — which is why a first call can already fail deflation. { "schema": "gauntlet.v1", "verdict": "fail" , "died_at": "family_deflation" , "failure_codes": [ "deflated_out_at_n=600" ], "stages": [ { "stage": "net_edge", "verdict": "pass" , "evidence": { "net_sharpe_annualized": 3.11, "trades": 184, "bars": 380, "net_return_total_pct": 41.7 } }, // 380 daily bars ≈ 18 months: a high Sharpe on a short window { "stage": "funding_edge", "verdict": "skipped" , "evidence": { "reason": "no perpetuals in book" } }, { "stage": "family_deflation", "verdict": "fail" , "evidence": { "dsr": 0.31, "cumulative_n": 1, "variants_in_call": 1, "n_trials_effective": 600, "effective_n_method": "family_ledger", "killed_by": "deflated_out_at_n" , "family_verdict": "deflated_out" } } // over 18 months, the best of 600 recorded family tries is expected to look // about this good by chance alone — so 3.11 buys only dsr 0.31, not a pass // + power_honesty, significance, cpcv, walk_forward, concentration, placebo, capacity, graveyard_prior ], "budget": { "cumulative_n": 1, "n_trials_effective": 600 } // cumulative_n = your own submissions; n_trials_effective counts the whole family's recorded tries } 02 · PRE-REGISTER Call it before the data exists. Deposit your hypothesis today; the service records its canonical hash and acceptance time in an operator-published chained commitment . Retain the registration response or chain head: a later mismatch is then detectable. After the market happens, evaluation uses data from after the stored cutoff. This is not an independent timestamp or immutability guarantee while external_anchor is empty; that claim requires an independent external anchor. It also does not buy immunity from the trial ledger — registering ten variants still counts as ten tries against your family's budget. REGISTER + EVALUATE · $0.05 EACH · LIVE PRICES 03 · FORENSICS The full autopsy. When a signal fails — or barely passes — the forensics report shows the whole picture: how the edge decays when you delay execution, what realistic costs do to it, whether it survives with chunks of history removed, how it behaves across market regimes, and how much capital it could actually carry before eating itself. Know where the edge died, so you stop digging in dead ground. $0.05 PER COMPLETED CHECK · LIVE PRICE 04 · CERTIFY Credibility that travels. After a pre-registered call is evaluated, issue its signed certificate through the separate certificate lifecycle. Whoever you show it to — an investor, a prop desk, a skeptical forum — can use our hosted verifier for a platform-trust result. Offline, a raw key check detects changed signed bytes; full platform trust also requires an independently pinned signed keyring and complete revocation history. ISSUANCE $9.90 · LIVE PRICE · VERIFICATION FREE EVERY TOOL · IN PLAIN WORDS Paid rows read their amount from the machine-readable live pricing registry ; this product page does not duplicate mutable prices. Every tool, in plain words. The four products above are the jobs people hire us for; this is the instrument panel behind them. Each tool gets one honest sentence about what it actually does — and wherever a technical term is unavoidable, it arrives with its translation. If you want the machinery underneath, the tool reference goes as deep as we can responsibly go. Try us free first tool what it does — in plain words price assay_demo Runs the full battery on a built-in example so you can see what a verdict looks like before you send us anything — no account, nothing to set up. free assay_preflight Checks that your file is shaped correctly before any money moves, because a check that dies on a formatting mistake helps nobody. free assay_graveyard Looks up whether this kind of idea has been tried and buried here before, and what killed it — worth knowing before you give it your weekends. free assay_calibration Current population/status disclosure: calibration data is returned only in a valid signed snapshot; otherwise the response reports signed:false with a stale/reason state. It is not an outcome score or track record. free assay_certificate_verify Checks a certificate fail-closed against platform trust: signature, externally pinned key history, certificate-purpose eligibility and revocations. A caller key alone is only raw-signature evidence. free assay_provider_protocol A checklist of seven tests for judging any signal vendor or strategy validator — run it against AlphaAssay first, which is exactly what it is built for. free Put a signal on trial tool what it does — in plain words price assay_signal Send your trading results — a trade list or an equity curve — and get back pass or fail with the reason spelled out, after realistic costs are charged and luck is priced in. live registry assay_gauntlet Runs the whole battery in one call — costs, luck, overfitting, size limits — and tells you which gate killed your signal first, so you know what to fix or when to stop. live registry assay_forensics Digs into why a signal fails — whether it quietly peeked at future data, and whether random timing would have earned just as much. live registry assay_backtest Runs your strategy rules on your own data with honest costs, and it remembers every attempt — so trying again and again cannot quietly turn luck into a pass. live registry assay_batch Tests up to 25 versions of one idea in a single call and counts every one of them, because showing only your best try is exactly the trick that fools people. live registry / variant assay_falsify Attacks your strategy from eight directions on purpose and reports back what survived, what broke, and what broke it first. live registry assay_pbo For when you tried many settings and kept the winner: it measures how likely that winner is simply the luckiest loser — the Probability of Backtest Overfitting. live registry assay_survivors Takes everything you tried and tells you which versions the evidence actually leaves standing — disclosed in the order you sent them, never ranked, never recommended. live registry assay_reproduce When someone claims a track record, this recomputes their numbers from the raw trades — including whether each fill was even physically possible at that bar's prices. live registry assay_tradelog Reads a raw fill log on its own — no price data needed — and finds the contradictions inside it: duplicated fills, exits before entries, profits that disagree with the row's own prices. live registry assay_cpcv Shows the whole distribution behind the overfitting check — how your strategy scores on every purged slice of history, not just the one lucky split a backtest happens to show. live registry assay_var_es Takes the risk numbers your model promised in advance and counts how often — and how deeply — reality broke them; a model that understates its own risk fails like any other claim. live registry assay_conformal Checks whether your model's confidence intervals contain reality as often as they claim — coverage that exists on the label but not in the outcomes gets named, honestly. live registry Prove you called it tool what it does — in plain words price assay_register Hashes the canonical call with its service acceptance time into an operator-published chain. Keep the returned commitment or chain head so later changes are detectable; the response discloses whether an external anchor exists. live registry assay_verdict Once the future has played out, the stored call is scored only on data from after its recorded cutoff. Certificate issuance is a separate REST lifecycle action. $0.05 per check · live registry No tool has an order or broker path. Verdicts are demote-only: they can devalue a claim, never bless one. Within the published billing unit, a fail costs the same as a pass — you buy the trial, not the outcome. the technical reference · what we keep, precisely FREE · FOR EVERYONE Your agent can test us before it pays us. ✓ Golden specimens — test cases with known semantic answer fields. Verify we get right positives, right negatives and honest abstention while volatile metadata may vary. ✓ Public calibration state — bucketed mature-registration count plus accumulating/insufficient-history disclosure, free to pull; not yet an outcome score. ✓ Certificate verification — check any issued AlphaAssay certificate, no account needed. run it now — the 60-second quickstart for agents: every MCP tool, in plain terms verify us in 60 seconds # no signup — a specimen with a known answer: $ curl -s https://api.alphaassay.com/v1/assay/demo \ -d @golden_lookahead.json { "verdict": "fail" , "died_at": "net_edge" , "failure_codes": [ "no_net_edge" ] } Pick your signal. We do the rest. $ test my signal full pricing ## Research — AlphaAssay URL: https://alphaassay.com/research RESEARCH We publish what the graveyard teaches. The uncomfortable truth of quant research: most profitable-looking strategies are statistical illusions. That is not our opinion — it is the published record, and we write it up in plain language: start with the primer, then diagnose your own backtest. START HERE PRIMER Why backtests flatter everyone Look-ahead, survivorship, selection under multiple testing — and the statistics that catch them. PRIMER · 6 MIN READ COMPARISON Tools to validate a trading signal, compared honestly An honest field guide: AlphaAssay, QuantConnect, walk-forward tools, purged-CV libraries and DIY statistics — what each is best for, and where each stops. COMPARISON · 7 MIN READ INTEGRATION A validation gate your agent can call Hosted API and MCP server, priced per call: what the gate checks before your agent acts, how to wire it in two transports, and why a library inside the loop cannot count your trials. INTEGRATION · 6 MIN READ CATEGORY What is an assay office for trading signals? Independent grading for strategy evidence — what an assay office does, what it never does, and the honest map of alternatives: libraries, GIPS verification, prop-firm evals, tournaments, provenance pins. CATEGORY · 5 MIN READ CASE STUDY The $546k backtest that passed walk-forward Own backtester, real fees, walk-forward, six months live — and the account still went from a simulated $546k to $3k. What honest self-validation cannot see. CASE STUDY · 5 MIN READ PROTOCOL How to test a signal provider Seven falsifiable tests that separate edge from selection — for any provider, without their cooperation. Run it against us first. PROTOCOL · 7 TESTS DIAGNOSE YOUR BACKTEST DIAGNOSTIC Is my backtest overfit? Twelve signs a backtest was manufactured — costs, leakage, multiple testing, regime luck. DIAGNOSTIC · 3 MIN READ TOOL The Deflated Sharpe Ratio, explained After N trials, is your Sharpe real or selection luck? Formula, worked example, in-browser calculator. TOOL · INTERACTIVE TOOL The Probabilistic Sharpe Ratio, with a calculator Is your Sharpe ratio statistically above a benchmark, given track length, skew and fat tails? Formula, worked example, in-browser calculator. TOOL · INTERACTIVE TOOL Minimum Track Record Length, with a calculator How long must a track record be before a Sharpe ratio clears a benchmark with confidence? The MinTRL formula and an in-browser calculator. TOOL · INTERACTIVE TOOL Probability of Backtest Overfitting (PBO), with a calculator After picking the best of many strategy configurations, how likely is your winner an overfit artifact? The CSCV estimate, explained, with an in-browser calculator. TOOL · INTERACTIVE TOOL Minimum Backtest Length (MinBTL), with a calculator You tried N variants — how long must the backtest be before the best one's Sharpe stops being expected from noise? 45 trials at Sharpe 1.0 already need 5 years of daily data. TOOL · INTERACTIVE TOOL Statistical power: could your data even show an edge? A test with 12% power proves nothing in either direction. Effect size, achieved power and the sample you actually need — with an in-browser calculator. TOOL · INTERACTIVE TOOL Break-even AUM: at what size does your edge die? Market impact grows with the square root of size; your margin doesn't. The AUM where the edge stops paying for its own footprint — a death boundary, not a recommendation. TOOL · INTERACTIVE EVIDENCE Same strategy, five engines, different answers Identical strategy, data and costs — up to 3.71% apart across engines. The simulator is part of the experiment. EVIDENCE · 5 MIN READ METHOD What counts as the same strategy? Three RSI variants are one idea asked three times — why honest deflation counts cumulatively, per family. METHOD · 5 MIN READ GUIDE Walk-forward analysis, honestly What it catches, what it quietly misses, and how it compares with CV, CPCV and placebo testing. GUIDE · 8 MIN READ CHECKLIST The backtest overfitting checklist Twelve checks before you trust a strategy — copy as markdown, wire into your agent. CHECKLIST · COPY-PASTE THE PUBLIC RECORD RECORD The Signal Validation Benchmark Reproducible known-answer specimens, a 9-strategy field result, and current public calibration population/status evidence. RECORD · REPRODUCIBLE RECORD How we grade ourselves — in public What calibration v0 really publishes: a bucketed mature-registration population and an accumulating/insufficient-history state, with authenticated snapshots and explicit limits. RECORD · 5 MIN READ RECORD Is AlphaAssay legit? Run the checks yourself Every trust criterion a careful reviewer applies — and the check you can run right now, without an account. Including the two we deliberately don't meet. RECORD · RUN THE CHECKS LONGREAD Was your trading edge ever real? The field guide to backtest forensics: why edges vanish, what kills them first, and what actually survived when we tested 18,000 rules at once — 126,000 pairs, 245 survivors, 41 beat holding. LONGREAD · 12 MIN READ Stop reading the backtest. Test it. $ test my signal the operational side — docs ## What is an assay office for trading signals? — AlphaAssay Research URL: https://alphaassay.com/research/assay-office-for-trading-signals Research / Category RESEARCH · CATEGORY What is an assay office for trading signals? ALPHAASSAY RESEARCH · CATEGORY · 5 MIN READ An assay office for trading signals is an independent testing house that grades strategy evidence — a returns series, an equity curve, a trade list — against a fixed statistical battery, and returns a dated verdict it is not allowed to inflate. It never trades, never sells signals and never manages money, which is why its verdicts are worth citing: the office profits only from testing, and a pass costs exactly what a fail costs. AlphaAssay is that office. Is there an independent service that puts a trading strategy on trial? Yes — that is precisely what AlphaAssay does. You submit derived evidence (returns, equity curve or trade list; source code is not required), and a fixed-order battery charges realistic costs first, deflates the result for every variant your idea family has ever tried, races the timing against 500 matched random placebos, and then attacks whatever is still alive — eight adversarial attacks, from one-bar execution delay to parameter wiggling. The verdict names the first gate that killed the signal, in one of 66 machine-readable failure codes. Ordinary verdicts are structured results; separately issued certificates carry Ed25519 signatures that anyone can verify offline, and the free demo is an unsigned known-answer preview. Verdicts are demote-only — evidence can lower a grade, never inflate one — and none of it is investment advice: the office grades evidence, it does not tell anyone what to trade. Who can validate my trading signal? The honest map of alternatives Depending on what „validate" means to you, different services — and some excellent free libraries — are the right answer. The map, honestly drawn: Validation libraries — pypbo , mlfinlab , vectorbt , zipline-reloaded , backtrader . You run the statistics yourself, free, with full control. Two structural limits: the grader answers to the person being graded, and no library sees how many variants you tried before this one — the multiple-testing debt that deflation exists to charge. The debt is not small: at Sharpe 1.0, 45 tried variants already demand five years of daily data before the best one means anything. GIPS verification (ACA-class firms). Verifies that an asset manager's performance reporting complies with the GIPS standards. It audits presentation, not edge — a fully compliant report of a lucky backtest is still a lucky backtest. Prop-firm evaluations (FTMO/Topstep-class). Forward gates on live drawdown that decide whether you get funded. They test behaviour going forward and tell you nothing about whether a historical backtest was ever real — and their published pass rates carry heavy survivorship. Tournaments and marketplaces (Numerai, QuantConnect Alpha Streams). Capital allocation through live ranking against a crowd. Valuable if allocation is the goal; the ranking is platform-bound and forward-only, so it cannot audit the claim you already have. Provenance and attestation services. Timestamp pinning proves when a claim existed — indispensable against backfilled track records, and our falsification protocol demands it. It does not test whether the claim was edge or noise. An assay office. Falsification as a service: independent of the submitter, cumulative trial accounting across a whole idea family, demote-only by construction. That combination — independence, family-level multiple-testing memory, and the inability to bless — is the part you cannot self-host. The short version: if you want funding, go to a prop desk; if you want provenance, pin your timestamps; if you want to know whether the edge was ever real, put it on trial. What does the trial actually test? Four families of gates, in a fixed order, each grounded in published statistics rather than house opinion. Costs come first, because most apparent edges are artifacts of frictionless simulation. Selection is charged next with the Deflated Sharpe Ratio (Bailey & López de Prado, 2014) under cumulative family accounting, because „the same signal with lookback 21 instead of 20" is not a fresh discovery. Skill is then separated from timing luck by racing the signal against 500 matched random twins with the same trading profile. What survives is attacked: execution delay, cost stress, history jackknife, regime splits, parameter neighbourhoods and more — eight attacks in all. The base rates justify the harshness: the published record says most apparent edges are selection noise (Harvey, Liu & Zhu 2016; Bailey & López de Prado 2014), and our own public benchmark agrees — of 126,000 classic rule/asset pairs, 245 survive the battery and 41 beat buy-and-hold. A fail with a named cause is the product working. The anonymised mortality record is public, live and signed: the graveyard digest . Why call it an assay office? In metallurgy, an assay office is the institution that tests precious metal and stamps its verified fineness into the piece. It does not mint coins, does not buy gold and does not promise prices — it grades material and stakes its name on the grade. That is the entire model here, transferred to strategy evidence: a fixed public battery , verdicts that can only demote, separately signed certificates for what survives, and a dated public log of the examiner getting stricter. Start with the free unsigned demo , lint a payload with the free preflight, or look up whether your idea's family is already in the graveyard — before it costs you weekends. ← previous The agent validation gate next → Is my backtest overfit? ## Same strategy, five engines, different answers: backtest implementation risk — AlphaAssay Research URL: https://alphaassay.com/research/backtest-implementation-risk Research / Implementation risk RESEARCH · EVIDENCE Same strategy, five engines, different answers: backtest implementation risk ALPHAASSAY RESEARCH · EVIDENCE · 5 MIN READ Run one identical strategy — same rules, same data, same cost specification — through five widely used backtesting engines, and the reported total return can differ by up to 3.71 %. That is the central measurement of a 2026 study that benchmarked five engine implementations across fifteen strategies on S&P 500 data from 2020–2024, and named the effect implementation risk : variability in backtest outcomes attributable solely to the choice of simulation engine. On a $1B portfolio the ambiguity is worth roughly $37M a year. Your Sharpe ratio can be honest, your data clean, your trial count accounted for — and the number you are staring at still depends on which engine happened to produce it. Where does the divergence come from? Almost entirely from transaction-cost models . In the study's zero-cost ablation, all five engines agreed exactly — to the digit. Turn realistic costs on and the divergence rises monotonically with cost intensity (Spearman ρ = 0.93). The engines do not disagree about markets; they disagree about what trading costs, and each buries that opinion in defaults. The forensic part of the study found seven previously undocumented defects across widely used engines — including one that silently divides the commission rate you pass by 100, so a user specifying 18 basis points is actually charged 0.18. Every backtest on that engine looks systematically cheaper to trade than reality. Is this just overfitting in disguise? No — and that is what makes it dangerous. Implementation risk is orthogonal to statistical overfitting: deflation, placebo trials and out-of-sample discipline all assume the simulator itself is telling the truth. A perfectly honest researcher with a perfectly deflated Sharpe still inherits the engine's cost model, bugs and all. It is a separate axis of backtest unreliability, and no split scheme fixes it. What can you do about it? The study's recommendation is blunt: validate with at least two independent engines of maximally different architecture (one event-driven, one vectorised), and audit the cost model of each against a reference specification. In practice almost nobody does this — it doubles the plumbing. Three cheaper habits catch most of the damage: read your engine's commission and slippage defaults instead of trusting them; re-run the backtest with costs set to zero and sanity-check that the delta matches your own cost arithmetic; and treat any engine-reported cost below your venue's published fee schedule as a bug until proven otherwise. How does this relate to an AlphaAssay verdict? Our cost_stress attack already treats the submitted cost assumptions as hostile — it re-prices the edge under worse fees, wider spreads and delayed fills, which catches strategies that only live inside optimistic cost settings. What it does not do is re-implement your simulator: cross-engine replication is an open frontier for the field, not a solved gate, and we treat it that way. Until it is solved, the honest posture is the one this page describes: assume the engine is part of the experiment, not part of the ground truth. (Source: „implementation risk" benchmark study, arXiv:2603.20319, 2026; figures quoted above are from the paper.) ← previous The falsification protocol next → Strategy families ## Backtest overfitting checklist: 12 checks before you trust a strategy — AlphaAssay URL: https://alphaassay.com/research/backtest-overfitting-checklist Research / Checklist RESEARCH · CHECKLIST The backtest overfitting checklist ALPHAASSAY RESEARCH · CHECKLIST · COPY-PASTE Twelve checks, ordered by what kills fastest. Before you trust a backtest — yours or anyone's — walk it through this list: realistic costs and execution delay first, because most edges die right there; then leakage; then multiple-testing accounting; then robustness. Each check states what to do, what failure looks like, and the machine-readable failure code an AlphaAssay verdict would attach, so your agent can run the same list programmatically. A strategy that clears all twelve is not proven — no finite test can prove an edge — but it has earned the right to face real statistics instead of real money. Copy the whole checklist as markdown below and wire it into your agent's pre-deployment routine. Gate 1 — net edge after reality 1. Charge fees, spread and slippage Rebuild the backtest with your venue's real fee schedule and a conservative spread. Failure looks like: the equity curve flattens the moment money touches it. 2. Delay every fill by one bar If the edge vanishes with one bar of delay, your simulation was trading on information it did not yet have. Attack: execution_lag (look_ahead). 3. Stress costs 2× Double fees and spread. An edge that dies here is a bet on your cost model, not on the market. Code: cost_stress . Gate 2 — honest trial accounting 4. Count every variant you tried — including discarded ones Parameter grids, deleted notebooks, „quick experiments": all trials. The winner was selected from all of them, so the statistics must pay for all of them. 5. Compute the Deflated Sharpe Ratio Feed your Sharpe, history length and true trial count into the calculator . Below 0.95: not proven. Code: deflated_out_at_n=N . 6. Freeze the spec before the final test Decide every parameter, then test once. Better: make it provable with a sealed public timestamp . Gate 3 — luck versus timing 7. Race it against random twins Generate placebo signals with the same trading profile (frequency, holding period, exposure). If they do as well, the timing was never the edge. 8. Read the placebo percentile, not the raw return Beating 61 of 100 random twins is a coin toss with extra steps; the percentile is the honest number. 9. Delete the best month and re-run An edge that lives in one lucky window is a story about that window. Attack: time_jackknife . Gate 4 — robustness 10. Split by regime and check both halves Bull/bear, high/low volatility — an edge that only exists in one half is regime luck wearing a lab coat. Code: regime_split . 11. Wiggle every parameter ±20% If neighbouring parameters fail, you found a coordinate, not an edge. Attack: parameter_neighbourhood . 12. Ask what capacity the edge survives Estimate the size at which your own trading moves the price you depend on. Some real edges are real and worth $40 a month. Why this order? Because mortality is front-loaded: most signals die at costs or the placebo race long before anything exotic gets a chance. Checking robustness on a signal that never survived costs is polishing a corpse. Copy the checklist (markdown, free to reuse with a link back) backtest-overfitting-checklist.md — CC BY 4.0, link to alphaassay.com copy # Backtest overfitting checklist (alphaassay.com/research/backtest-overfitting-checklist) Gate 1 — net edge: 1. charge fees+spread+slippage 2. delay fills 1 bar 3. stress costs 2x Gate 2 — trials: 4. count ALL variants tried 5. deflated Sharpe >= 0.95 6. freeze spec before final test Gate 3 — luck: 7. race vs matched random twins 8. read the placebo percentile 9. delete best month, re-run Gate 4 — robustness: 10. split by regime 11. wiggle params ±20% 12. estimate capacity Rule: a pass is a statistical trial result — never a promise of returns. Run it as one API call The manual checklist takes an afternoon; the battery runs the statistical parts in one call and returns a structured ordinary result. A signed certificate is a separate issuance step. Free unsigned known-answer demo first — no account: terminal $ curl -s https://api.alphaassay.com/v1/assay/demo \ -H "Content-Type: application/json" -d @golden_lookahead.json Related: the Deflated Sharpe calculator · walk-forward analysis · is my backtest overfit? · all failure codes ← previous Walk-forward analysis next → The Signal Validation Benchmark ## Statistical power of a backtest: explainer and calculator — AlphaAssay Research URL: https://alphaassay.com/research/backtest-statistical-power Research / Tool RESEARCH · TOOL Statistical power: could your backtest even detect a real edge? — with a calculator ALPHAASSAY RESEARCH · TOOL · INTERACTIVE CALCULATOR Statistical power answers the question that a green equity curve quietly skips: if your edge were exactly as large as you claim, what is the probability your data sample would detect it at all? A test with 12% power misses a real edge 88% of the time — so both of its possible outcomes are close to meaningless: a pass proves little, and a fail proves little. Power is a property of the sample size and the effect size (Cohen's d, the per-period mean return divided by its volatility), computed here as a one-sided one-sample test at α = 0.05, with 80% as the conventional adequacy line (Cohen, 1988). The calculator runs entirely in your browser; nothing is uploaded. Every AlphaAssay verdict now carries this number: an underpowered test is stamped underpowered instead of being allowed to read as an acquittal. Can your data even show the edge you claim? statistical power — runs in your browser, nothing uploaded mean return per period volatility per period (same unit) observations n compute power Units cancel in Cohen's d, so use any consistent pair — percent per day with percent per day, decimals with decimals. The defaults (d = 0.05 over 3 years of daily data) reproduce the 39% row in the table below — same formula, same defaults. How much history does a small edge need? effect size d (per period) ≈ annualised Sharpe observations n achieved power n for 80% power 0.05 0.8 756 (3y daily) ≈ 39% 2,473 0.10 1.6 756 (3y daily) ≈ 87% 619 0.10 1.6 34 ≈ 14% 619 0.20 3.2 252 (1y daily) ≈ 94% 155 The first row is the one that stings: a daily edge worth an annualised Sharpe of about 0.8 — a real, tradeable edge — gives three years of daily data only a 39% chance of detecting it. Most retail backtests are shorter than that, on smaller edges. The arithmetic does not care how the equity curve looks. The formula, step by step Step 1 — the effect size. d = mean return per period / volatility per period (Cohen's d; multiply by √periods-per-year and you have the annualised Sharpe). Step 2 — achieved power. For a one-sided one-sample test at level α: power = Φ( d·√n − z(1−α) ) — the probability the test statistic clears the critical value when the edge is real. Step 3 — the sample you actually need. n₈₀ = ⌈ ( (z(0.80) + z(1−α)) / d )² ⌉ observations for 80% power. Halve the effect size and the requirement quadruples — power is quadratic in 1/d, which is why small edges are so expensive to prove. Why is a pass on thin data not an acquittal? Because an underpowered test cannot convict, its failure to convict carries almost no information — and treating that silence as evidence is how weak signals get promoted. This is the asymmetry the power honesty check inside every AlphaAssay verdict exists to expose: when the sample cannot support the claim, the verdict says so, in the open, instead of letting „no fail detected" masquerade as „validated". The same honesty runs in the other direction as backtest-length accounting — see MinBTL — and both are one battery with deflation : how big is the edge, how long was the look, how many things were tried. ← previous Minimum Backtest Length next → Break-even AUM ## Break-even AUM: at what size does your edge die? — AlphaAssay Research URL: https://alphaassay.com/research/break-even-aum Research / Tool RESEARCH · TOOL Break-even AUM: at what size does your edge die? — with a calculator ALPHAASSAY RESEARCH · TOOL · INTERACTIVE CALCULATOR Every edge has a death boundary in capital: the assets-under-management at which its own market impact eats the entire margin that made it worth trading. This calculator locates that boundary. It takes your gross edge per trade, your gross Sharpe, the traded instrument's daily volatility and its average daily volume (ADV), and returns the AUM at which the strategy's net Sharpe is pushed down to the survival floor of 0.5 — the point where the edge stops being distinguishable from costs. The impact model is the empirical square-root law with a calibrated coefficient c = 0.69, published with its uncertainty band (0.63–0.77), and it is only valid up to a participation of 10% of ADV — beyond that the honest answer is „beyond model validity", not a bigger number. This is a mortality figure, not a sizing recommendation : the same number appears as break_even_aum_usd in the cost gate of every AlphaAssay verdict , and its refusal case has its own failure code, CAPACITY_MODEL_RANGE_EXCEEDED . The calculator runs entirely in your browser; nothing is uploaded. At what AUM does the edge stop paying for its own footprint? break-even aum — runs in your browser, nothing uploaded gross return per trade (%) gross Sharpe (annualised) daily volatility of instrument (%) average daily volume (USD) locate the death boundary The default inputs (0.30% per trade, gross Sharpe 1.5, 1.5% daily volatility, $200M ADV) reproduce the $1.87M row in the table below — same formula, same defaults. Fixed model constants: impact coefficient c = 0.69 (band 0.63–0.77), survival floor Sharpe 0.5, participation cap 10% of ADV, capital deployed per trade = 100% of AUM. Where do edges die? Worked examples gross / trade gross Sharpe daily vol ADV your edge dies at 0.30% 1.5 1.5% $200M ≈ $1.87M (band $1.50M–$2.24M) 0.20% 2.0 2.0% $50M ≈ $148k ($119k–$177k) 0.10% 1.5 2.0% $50M ≈ $29k ($23k–$35k) 0.05% 0.8 2.0% $50M ≈ $2.3k — impact eats it at retail size 1.00% 3.0 0.5% $10M ≥ $1.0M — beyond model validity (10% ADV cap) Notice what drives the boundary: it is quadratic in the per-trade margin and inverse-quadratic in volatility. A thin edge on a volatile instrument dies at sizes a single retail account can reach — the third and fourth rows are not exotic inputs, they are typical crypto-strategy claims. The formula, step by step Step 1 — the cost budget. Only the margin above survival is spendable: budget = gross% per trade · (1 − 0.5 / gross Sharpe) . A strategy at gross Sharpe 0.5 or below has no budget — it is already at the floor before paying any impact. Step 2 — the impact model. Round-trip market impact follows the empirical square-root law: impact% ≈ 2 · c · σ_daily · √(trade size / ADV) , with c = 0.69 calibrated against published impact studies (uncertainty band 0.63–0.77 — the band is shown, not hidden). Step 3 — solve for the boundary. Impact equals budget at break-even AUM = ADV · ( budget / (2·c·σ_daily) )² . Above it, the net Sharpe is below 0.5 by construction. Step 4 — the honesty clamp. The square-root law is only calibrated up to ~10% participation of ADV. If the solved boundary exceeds that, the model does not extrapolate — it reports „beyond model validity" and stops. An impact model that keeps producing numbers outside its calibration is a backtest flattering you by other means. Why a death boundary and not a capacity recommendation? Because the two numbers point in opposite directions. A recommendation says „deploy this much"; this figure says only „above $X, the arithmetic says your edge is gone" — it can devalue a strategy, never endorse one, which is the same demote-only rule every AlphaAssay verdict follows. Costs are also where most backtests are manufactured: frictionless fills die first at the cost gates ( no_net_edge , cost_stress ), and the capacity check is the same interrogation asked at scale. If your backtest has never been charged its own footprint, it has other problems too . ← previous Statistical power next → The $546k case study ## Deflated Sharpe Ratio: explainer and calculator — AlphaAssay Research URL: https://alphaassay.com/research/deflated-sharpe-ratio Research / Tool RESEARCH · TOOL The Deflated Sharpe Ratio, explained — with a calculator ALPHAASSAY RESEARCH · TOOL · INTERACTIVE CALCULATOR The Deflated Sharpe Ratio (DSR) answers one question: after accounting for how many strategy variants you tried, what is the probability that your backtest's Sharpe ratio reflects a genuine edge rather than selection luck? It takes your observed Sharpe, the length and shape of your return series (skewness, kurtosis), and — the input everyone omits — the number of trials behind the result, then raises the benchmark your Sharpe must clear. Bailey and López de Prado introduced it in 2014 because trying forty variants and keeping the best one manufactures a „great" backtest from pure noise, mathematically guaranteed. The calculator below runs entirely in your browser; nothing is uploaded. Rule of thumb: a DSR probability below 0.95 means your result is statistically indistinguishable from an accident. This same statistic, with cumulative trial accounting, runs as gate 2 of every AlphaAssay verdict . The calculator deflated sharpe — runs in your browser, nothing uploaded observed Sharpe (annualised) observations (e.g. trading days) periods per year skewness of returns kurtosis (normal = 3) number of trials N (be honest) compute DSR Annualised Sharpe is de-annualised internally (SR ÷ √periods-per-year); the cross-sectional variance of no-edge trial Sharpes defaults to 1/T (their sampling variance) — the theoretically neutral assumption when you have not measured it. The calculator implements Bailey & López de Prado (2014): expected maximum Sharpe under N trials via extreme-value approximation, then the Probabilistic Sharpe Ratio against that benchmark. What does the DSR look like as trials pile up? (computed with the calculator below — same formula, same defaults) observed Sharpe history trials N DSR probability reading 1.8 3y daily 1 ≈ 1.0 a single pre-registered test — strong 1.8 3y daily 10 ≈ 0.94 a modest grid search already erodes it 1.8 3y daily 50 ≈ 0.8 typical „I tuned it for a weekend" 1.8 3y daily 200 ≈ 0.64 indistinguishable from selection luck Same backtest, same Sharpe — the only thing that changed is honesty about the search. Run your own numbers above. Why does the ordinary Sharpe ratio flatter you? Because it prices a single experiment, and you did not run a single experiment. If forty variants of an idea are tried and the best is kept, the winner's Sharpe contains selection luck by construction — Bailey, Borwein, López de Prado and Zhu (2014) showed a „great" backtest is mathematically guaranteed given enough trials, from pure noise. The full mechanism, with the published base rates: why backtests flatter everyone . The formula, step by step Step 1 — the benchmark your Sharpe must beat. Under N independent trials of no-edge strategies, the expected maximum observed Sharpe is approximately E[max SR] ≈ σ_trials · [(1−γ)·z(1−1/N) + γ·z(1−1/(N·e))] , where γ ≈ 0.5772 (Euler–Mascheroni) and z is the standard-normal quantile. More trials → higher hurdle. Step 2 — the Probabilistic Sharpe Ratio against that hurdle. DSR = Φ( ((SR − SR*) · √(T−1)) / √(1 − γ₃·SR + ((γ₄−1)/4)·SR²) ) — your observed SR versus the hurdle SR*, scaled by track-record length T and penalised for skewness γ₃ and kurtosis γ₄ (fat tails and negative skew make a Sharpe less trustworthy). Step 3 — read it as a probability. DSR ≥ 0.95: the edge is unlikely to be a selection artifact. Below: the honest answer is „not proven". The input nobody tracks: your true trial count Every parameter grid, every discarded variant, every „just one more tweak" is a trial — including the ones you deleted. This is why AlphaAssay accounts trials cumulatively per strategy family : the battery remembers what your family has spent even when you don't, and deflated_out_at_n=N is the failure code that says the honest answer to another tweak is already known. What the DSR cannot see The DSR corrects for selection — nothing else. Leaked future information, unrealistic costs and capacity limits all survive a perfect DSR, which is why it is one gate of four, not the whole trial: the four gates . ← previous Is my backtest overfit? next → Probabilistic Sharpe Ratio ## How we grade ourselves — AlphaAssay Research URL: https://alphaassay.com/research/how-we-grade-ourselves Research / Record RESEARCH · METHODOLOGY How we grade ourselves — in public ALPHAASSAY RESEARCH · 5 MIN READ Every rating business faces the same temptation: grade generously, keep customers happy, quietly forget the misses. AlphaAssay counters part of that risk with demote-only application rules and signed-or-explicitly-untrusted snapshots. The current snapshot is deliberately narrower than an outcome score, and the publication path remains operator-controlled. What calibration v0 actually publishes The public payload contains a privacy-bucketed count of evaluated mature registrations, forward_outcomes.status=accumulating , and honesty=insufficient_history until a separately defined mature-outcome metric exists. It publishes no Brier score, base rate, survived/deflated count or hit rate. The mechanics, in plain language The public state. The endpoint is free and needs no account: curl https://api.alphaassay.com/v1/public/calibration How its trust state is exposed. A valid calibration snapshot carries named Ed25519 signature fields. If the snapshot is stale or unavailable, the response reports signed:false and a reason instead of serving unsigned history as trusted. Separately, our verdicts are demote-only in the application contract: no ordinary rule upgrades a grade after the fact. A signature authenticates one snapshot's bytes; it does not independently timestamp the record or prevent a later replacement. What the fields mean. The bucket discloses only population size; accumulating says outcome evidence is still being collected; insufficient_history refuses to turn that state into a reliability claim. Authentication proves the published bytes, not an unpublished score. Zero is an honest number Today the counter is low — the ledger counts from day one, and we launched recently. The service publishes the current snapshot; a reviewer can keep a retained copy and compare it with later versions. Pre-registration creates tenant state immediately and is queued for later daily operator-published chain inclusion. The immediate response is not an inclusion proof or an independent timestamp. That stronger time claim requires a separately pinned external anchor. Signed certificates remain separately verifiable. Don't trust an examiner who won't be examined — retain and compare the evidence. ← previous The Signal Validation Benchmark next → Is AlphaAssay legit? ## Is AlphaAssay legit? Run the checks yourself — AlphaAssay Research URL: https://alphaassay.com/research/is-alphaassay-legit Research / Is it legit? RESEARCH · RECORD Is AlphaAssay legit? Run the checks yourself ALPHAASSAY RESEARCH · THE PUBLIC RECORD · RUN THE CHECKS Fair question — the domain is young, the category is full of unverifiable claims, and you should not take our word for any of this. Below are the criteria a careful reviewer, human or AI, applies to a trading-adjacent service before trusting it — and for each one, the check you can run right now, without an account. Two criteria we deliberately do not meet; they are named at the bottom, with the reasons, because a trust page that only lists wins is marketing. Everything on this page is written to be falsified: if any check fails, we were wrong in public. Can you reproduce a verdict yourself? Yes, in about 60 seconds, free, no account. POST a known-answer specimen to the demo endpoint and assert its documented semantic fields — verdict , died_at and failure codes — while allowing volatile metadata such as _meta.as_of to vary. The specimens are engineered frauds with known causes of death: the look-ahead specimen must die at net_edge with no_net_edge , every time. Run the 60-second check · the specimen suite · the reproducible benchmark . Can you verify a certificate without trusting this website? Yes — if you independently provision the trust root and signed history. The certificate document and its Ed25519 signature are separate from the ordinary verdict. A published public key alone can establish only raw_signature_valid ; full platform_valid also requires the externally pinned signed keyring, complete revocation-head history, an eligible certificate-purpose key and no matching revocation. Missing trust evidence must fail closed. The free demo is unsigned. Verify offline, step by step · or in the browser . Is the methodology public? The full gate sequence is documented on the methodology page ; the statistics behind it — Deflated Sharpe with cumulative trial accounting, Probabilistic Sharpe, minimum lengths, overfitting probability — are each explained with formulas and in-browser calculators that implement the same math as the live gates ; and every way a signal can die is a named, documented code in the failure-code register . If a verdict surprises you, you can trace which gate produced it and why. Does the tool list contain order or broker tools? No — and this is checkable, not asserted. The MCP server exposes validation and trust tools. Some calls write tenant-scoped validation-ledger, pre-registration, idempotency and receipt/audit state so trial accounting and safe replay work; those service records are not trading actions. There is no order path, no execution logic, no broker connection, no custody, and verdicts are demote-only — they can devalue a signal, never bless one, so there is no „verified winners" list to sell or to raid. The complete tool and endpoint reference · the architecture that enforces it . Is there code and an external footprint? Yes, and you can read it: the offline certificate verifier and the stdio-to-remote MCP proxy are open source under Apache-2.0 at github.com/alphaassay/mcp , with their unit tests and the canonical test vectors they are developed against. The server is published in the official MCP registry ( com.alphaassay/mcp ) and mirrored by the public MCP directories; the machine-readable manifest is served first-party at alphaassay.com/server.json . The validation engine itself is deliberately not published; that is one of the two deliberate gaps explained below. Do you have to upload your strategy? You do not have to send source code. The input is tool-specific: some calls take derived evidence (returns, trades or decision timestamps at a granularity you choose); backtest and batch calls take an executable DSL spec plus candles. Working payloads are processed transiently. The documented retained classes are tenant-scoped validation/audit records — a one-way fingerprint, verdict, summary statistics, an optional 32-number return sketch and structural coordinates — plus the full spec you deliberately seal through pre-registration . Idempotency and payment flows retain the material needed to replay a named result: POST /x402/v1/gauntlet caches the response against the payment nonce and records a receipt. Its receipt is signed when a platform key is configured and otherwise records signed:false ; purchased certificates are kept account-bound for re-delivery. A release canary sends a marker through the covered logging, error, metric and alert sinks; that is scoped regression evidence, not a universal proof. The complete inventory: what we keep , path by path. The three walls, in detail . Who operates it, and what does it cost? The operator, registered address and contact are on the legal page ; terms and privacy are published; and pricing and billing units are public — current amounts are machine-readable, not tiers-on-request or hidden behind „contact sales". What do we deliberately not offer? No pip-installable engine. The verdict engine is not on PyPI or npm, by decision rather than neglect: an examiner you can run and edit locally is an examiner you can quietly optimize against, and a locally produced „pass" is worth exactly what the person showing it to you wants it to be worth. What is public instead is everything needed to audit the examiner: the specimens, standard signed-or-explicitly-unsigned response envelopes, separately issued certificates, the offline verifier and the current persisted calibration population/status state. No offline clone of the hosted validator. A separate local stdio MCP exposes 50 operator/research tools, but it is not the hosted public catalog or a resettable copy of its tenant ledger. Cumulative trial accounting — the thing that makes deflation honest — only works if the count survives your restarts and your temptations; a family budget you can reset locally is not a budget. The trade-off is explicit: you get less control of the examiner, and in exchange the examiner's answers are worth something to third parties. And one thing we cannot shortcut: the domain is weeks old. Search engines are still indexing it, and reputation services have not categorised it yet. Age is the one trust signal that cannot be engineered — which is why everything above is designed to be checkable without it, and why our calibration record publishes signed snapshots that a reviewer can retain and compare over time. One name, one service AlphaAssay — one word — is the Swiss signal-validation service operating alphaassay.com, api.alphaassay.com and mcp.alphaassay.com, and nothing else: no fund, no signal seller, no other product. If you found a similarly named entity elsewhere, it is not us. Still unconvinced? Good — stay that way. Scepticism is the correct default in this category, so keep it pointed at us: the falsification protocol is seven tests that apply to any signal-adjacent service, us included — run it against us first. Or pull the current calibration population/status disclosure and check that it still says accumulating/insufficient-history. It is not an outcome score. If you catch us failing our own protocol, that finding is more valuable than anything this page could tell you. ← previous How we grade ourselves next → Was your edge ever real? ## Is my backtest overfit? A 3-minute diagnostic — AlphaAssay Research URL: https://alphaassay.com/research/is-my-backtest-overfit Research / Diagnostic RESEARCH · DIAGNOSTIC Is my backtest overfit? ALPHAASSAY RESEARCH · DIAGNOSTIC · 3 MIN READ Statistically, probably — and that is a base rate, not an insult. When McLean and Pontiff re-tested 97 peer-reviewed return anomalies, more than half of the performance evaporated after publication; when Chordia, Goyal and Saretto generated 2.1 million strategies and corrected properly for multiple testing, almost none survived. Retail backtests receive far less scrutiny than either. So the honest question is not whether your backtest looks good — everyone's does — but whether it survives the four ways good-looking backtests are manufactured: unrealistic costs, look-ahead leakage, selection under multiple testing, and regime luck. The twelve signs below cover how your backtest was built — no numbers uploaded, no account — and point to the gate of a statistical trial that would most likely kill it, in the same failure-code vocabulary a real AlphaAssay verdict uses. Three minutes here is cheaper than a live drawdown. The twelve signs, gate by gate Did you charge fees, spread AND at least one bar of execution delay? Frictionless fills are the single most common manufacturing defect. Real fills pay the spread, pay fees, and happen a bar late. Failure signature: no_net_edge , the net_edge gate. Does the edge survive doubling your cost assumptions? An edge that dies at 2× costs is a cost-model bet, not a market edge ( cost_stress ). Is the strategy's turnover realistic for its capital? High-turnover edges evaporate with size; capacity is part of gate 1 economics, not an afterthought. Could any input have been revised after the fact? Earnings restatements, index membership, delisted assets quietly leak the future into the past — survivorship and restatement bias live in the data, not the code. Does every signal use only information available at trade time? Close prices used at the open, same-bar highs, „monthly" data stamped mid-month: classic look-ahead. If delaying execution one bar kills the edge, the answer was no. Did you test on assets that no longer exist? Testing today's coin list or index members silently deletes everything that died — the past looks safer than it was. How many parameter combinations did you try before this one? Count everything, including deleted notebooks. Forty trials manufacture a great backtest from noise — mathematically guaranteed (Bailey et al., 2014). Signature: deflated_out_at_n=N , the family_deflation gate. Did you freeze the spec before the final test? If the „final" test was re-run after a tweak, it was not final. Pre-registration exists for exactly this ( how it works ). Would you have published this result if it were negative? If not, your process has selection built in — the survivor you are looking at was chosen by outcome. Does the edge survive deleting its best single month? An edge that lives in one lucky window is a story about that window ( time_jackknife ). Does it work in both halves of a regime split? A strategy that rode one bull market backtests beautifully until the regime ends ( regime_split ). Does it survive wiggling every parameter ±20%? If lookback 20 works and 19/21 don't, you found a coordinate, not an edge ( parameter_neighbourhood ). What are the four ways backtests get manufactured? What is look-ahead bias? Using information in simulation that was not available at the moment of the trade — restated data, same-bar prices, later index membership. The backtest looks clairvoyant because it literally was. What is survivorship bias? Testing on the assets that survived until today, silently excluding everything that died on the way. Mean-reversion strategies look like geniuses among survivors. What is selection under multiple testing? Trying many variants and keeping the best: the winner's performance is part skill, part selection — and with enough trials, all selection. The Deflated Sharpe Ratio was built to price exactly this. What is regime luck? An edge that only ever traded one market regime — one bull run, one volatility state — and mistakes that regime for the world. What does a real trial add? Self-diagnosis catches construction errors; only statistics catch luck. A signal that passes all twelve signs above still needs deflation against its true trial count and a placebo race against random twins — that is the battery , and the demo tier runs it free on known-answer test cases: golden specimens . ← previous The assay office, defined next → Deflated Sharpe Ratio ## Minimum Backtest Length (MinBTL): explainer and calculator — AlphaAssay Research URL: https://alphaassay.com/research/minimum-backtest-length Research / Tool RESEARCH · TOOL Minimum Backtest Length (MinBTL): how long must a backtest be? — with a calculator ALPHAASSAY RESEARCH · TOOL · INTERACTIVE CALCULATOR The Minimum Backtest Length (MinBTL) answers the question that decides whether your backtest is evidence at all: given that you tried N strategy variants, how long must the test window be before the best variant's Sharpe ratio cannot be produced by selection alone? Bailey, Borwein, López de Prado and Zhu (2014) proved the uncomfortable direction of this: for any target Sharpe, there is a number of trials N at which pure noise is expected to deliver it. MinBTL inverts that result. Try 45 variants and keep the best at Sharpe 1.0 on daily data, and you need about 5 years of history for that Sharpe to even begin to mean something — below that length, noise alone was expected to do it. The calculator runs entirely in your browser; nothing is uploaded. The same check runs inside every AlphaAssay verdict ; its failure code is BACKTEST_TOO_SHORT_FOR_N . How long must a backtest be for N tried variants? minimum backtest length — runs in your browser, nothing uploaded number of trials N (be honest) claimed Sharpe (annualised) periods per year compute MinBTL The default inputs (45 trials, Sharpe 1.0, daily data) reproduce the 1,260-observation row in the table below — same formula, same defaults. What does the minimum length look like as trials pile up? trials N claimed Sharpe MinBTL (daily) reading 10 1.0 ≈ 625 obs · 2.5 years a modest grid search already demands years 45 1.0 ≈ 1,260 obs · 5.0 years a weekend of tuning needs half a decade of data 100 1.0 ≈ 1,614 obs · 6.4 years most parameter sweeps live here 500 1.0 ≈ 2,349 obs · 9.3 years an optimizer run — nearly a decade required 45 2.0 ≈ 315 obs · 1.2 years a genuinely large edge shortens the sentence 45 0.5 ≈ 5,039 obs · 20.0 years a small Sharpe from a search is untestable in practice Read the second row again: 45 variants is not an industrial optimizer, it is one honest weekend of iteration — and it already consumes five years of daily history. Every additional trial raises the floor, and the trials you deleted count too. The formula, step by step Step 1 — what noise is expected to achieve. Under N independent no-edge trials, the expected maximum Sharpe per observation over a window of T observations is approximately E[max SR] ≈ (1/√T) · [(1−γ)·z(1−1/N) + γ·z(1−1/(N·e))] , with γ ≈ 0.5772 (Euler–Mascheroni) and z the standard-normal quantile. Step 2 — invert for T. Set that expectation equal to your claimed per-observation Sharpe (annualised Sharpe ÷ √periods-per-year) and solve: MinBTL = ( [(1−γ)·z(1−1/N) + γ·z(1−1/(N·e))] / ŜR )² observations. Step 3 — read it as a floor, not a target. At exactly MinBTL, your claimed Sharpe equals what selection over N noise trials was expected to produce — so a backtest of that length is the minimum at which the claim stops being automatic. Longer is evidence; shorter is arithmetic. MinBTL or MinTRL — which one do you need? They answer different interrogations. MinTRL asks how long a single track record must be before its Sharpe clears a benchmark at a chosen confidence — no selection involved. MinBTL asks how long a backtest must be before the best of N results is distinguishable from selection. If you ran one pre-registered test, MinTRL is your number. If you tried variants — and you did — MinBTL binds first, and it grows with every trial. The Deflated Sharpe Ratio is the same accounting applied to the result instead of the length. What MinBTL cannot see MinBTL prices exactly one failure mode: selection under multiple testing. Look-ahead leakage, survivorship in the data, unrealistic costs and regime luck all survive a backtest of any length — a ten-year backtest of a leaking signal is ten years of leak. That is why length is one check inside a battery , not a verdict by itself: an AlphaAssay fail names which gate killed the signal, and BACKTEST_TOO_SHORT_FOR_N is only one of the documented causes of death . ← previous Probability of Backtest Overfitting next → Statistical power ## Minimum Track Record Length: explainer and calculator — AlphaAssay Research URL: https://alphaassay.com/research/minimum-track-record-length Research / Tool RESEARCH · TOOL Minimum Track Record Length, explained — with a calculator ALPHAASSAY RESEARCH · TOOL · INTERACTIVE CALCULATOR Minimum Track Record Length (MinTRL) answers the question every track record dodges: how many observations do you need before a given Sharpe ratio is statistically above a benchmark, at a confidence you choose? A Sharpe of 1.0 looks tradeable, but at 95% confidence it takes roughly 2.7 years of daily data to prove it clears zero — and a Sharpe of 0.5 takes over a decade. MinTRL turns „impressive-looking" into „provable yet, or not". It is the inverse of the Probabilistic Sharpe Ratio : instead of asking how confident you are at a given length, it asks how long you need for a target confidence. Bailey and López de Prado (2012). The calculator runs entirely in your browser. The calculator minimum track record length — runs in your browser, nothing uploaded observed Sharpe (annualised) benchmark Sharpe SR* (annualised) periods per year skewness of returns kurtosis (normal = 3) target confidence (e.g. 0.95) compute MinTRL The default inputs (Sharpe 1.0, benchmark 0, 95% confidence) reproduce the 685-observation row in the table below — same formula, same defaults. How long does a Sharpe ratio take to prove? observed Sharpe benchmark SR* confidence MinTRL (daily) 0.5 0 0.95 ≈ 2,730 obs · 10.8 years 1.0 0 0.95 ≈ 685 obs · 2.7 years 1.5 0 0.95 ≈ 306 obs · 1.2 years 2.0 0 0.95 ≈ 173 obs · 0.7 years 1.0 0.5 0.95 ≈ 2,734 obs · 10.9 years The lesson is brutal and useful: a mediocre Sharpe over a short window is not evidence, however green the equity curve. Raising the bar you must beat (SR* from 0 to 0.5) explodes the history you need. The formula, step by step Step 1 — de-annualise your Sharpe and benchmark: ŜR = SR / √(periods per year) . Step 2 — the required length. MinTRL = 1 + (1 − γ₃·ŜR + ((γ₄−1)/4)·ŜR²) · ( Z_α / (ŜR − SR*) )² , where Z_α is the standard-normal quantile of your target confidence (0.95 → 1.645). The result is a number of observations; divide by periods-per-year for calendar time. Step 3 — the catch. If your Sharpe does not exceed the benchmark, no finite history proves it — the formula diverges, and the honest answer is „never, at this Sharpe". Skew and fat tails only make the required length longer. Why does this end the „just give it more time" argument? Because it puts a number on it. If your family has already spent its deflation budget and your Sharpe still needs eleven more years to clear zero, more time is not a plan. MinTRL, PSR and deflation are the same statistics from three angles, and all three run inside the AlphaAssay battery — which is why a fail can tell you to stop, not just to wait. ← previous Probabilistic Sharpe Ratio next → Probability of Backtest Overfitting ## Probabilistic Sharpe Ratio: explainer and calculator — AlphaAssay Research URL: https://alphaassay.com/research/probabilistic-sharpe-ratio Research / Tool RESEARCH · TOOL The Probabilistic Sharpe Ratio, explained — with a calculator ALPHAASSAY RESEARCH · TOOL · INTERACTIVE CALCULATOR The Probabilistic Sharpe Ratio (PSR) answers a sharper question than the Sharpe ratio alone: given how long your track record is, and how skewed and fat-tailed its returns are, what is the probability that your true Sharpe ratio exceeds a chosen benchmark (usually zero)? A raw Sharpe of 1.0 means one thing over twenty years and almost nothing over three months — PSR prices that difference. It takes your observed Sharpe, the number of returns, and the shape of the distribution (negative skew and fat tails make a Sharpe less trustworthy), and returns a probability. Bailey and López de Prado introduced it in 2012. The calculator below runs entirely in your browser; nothing is uploaded. Rule of thumb: PSR ≥ 0.95 means the edge over the benchmark is unlikely to be sampling noise. The calculator probabilistic sharpe — runs in your browser, nothing uploaded observed Sharpe (annualised) benchmark Sharpe SR* (annualised) observations (e.g. trading days) periods per year skewness of returns kurtosis (normal = 3) compute PSR Annualised Sharpes are de-annualised internally (SR ÷ √periods-per-year). The default inputs (Sharpe 1.0 over 756 daily observations, benchmark 0) reproduce the 0.958 row in the table below — same formula, same defaults. What does the PSR tell you that the Sharpe ratio doesn't? observed Sharpe history benchmark SR* PSR reading 1.0 1y daily (252) 0 0.841 suggestive — one year is not enough to be sure 1.0 3y daily (756) 0 0.958 clears 0.95 — the edge over zero is real 2.0 1y daily (252) 0 0.977 a high Sharpe buys confidence faster 0.5 3y daily (756) 0 0.807 not proven — a weak edge needs long history 1.5 2y daily (504) 1.0 0.760 beating a 1.0 hurdle is not yet established Same Sharpe, different verdicts — because history length, distribution shape and the benchmark all move the probability. Run your own numbers above. The formula, step by step Step 1 — de-annualise. Work in per-period units: ŜR = SR / √(periods per year) , and the same for the benchmark SR* . Step 2 — penalise the distribution. The denominator √(1 − γ₃·ŜR + ((γ₄−1)/4)·ŜR²) inflates when returns are negatively skewed (γ₃) or fat-tailed (γ₄): the same Sharpe from ugly returns is less trustworthy. Step 3 — read it as a probability. PSR = Φ( (ŜR − SR*)·√(n−1) / √(1 − γ₃·ŜR + ((γ₄−1)/4)·ŜR²) ) , where Φ is the standard-normal CDF and n is the number of returns. PSR ≥ 0.95: the edge over the benchmark is unlikely to be sampling noise. How is the PSR different from the Deflated Sharpe Ratio? The PSR asks whether your Sharpe beats one fixed benchmark . The Deflated Sharpe Ratio is the PSR with the benchmark set to the expected maximum Sharpe you would see after N no-skill trials — so it also charges you for the search you ran. Use PSR when you ran a single pre-registered test; use DSR the moment you tried more than one variant. The honest track-record question — how long until PSR clears a bar — is the Minimum Track Record Length . All of this is gate 2 of every AlphaAssay verdict . ← previous Deflated Sharpe Ratio next → Minimum Track Record Length ## Probability of Backtest Overfitting (PBO): explainer and calculator — AlphaAssay Research URL: https://alphaassay.com/research/probability-of-backtest-overfitting Research / Tool RESEARCH · TOOL The Probability of Backtest Overfitting (PBO), explained — with a calculator ALPHAASSAY RESEARCH · TOOL · INTERACTIVE CALCULATOR The Probability of Backtest Overfitting (PBO) answers the question that deflation alone cannot: when you pick the best of many strategy configurations, how often does that in-sample winner turn out to be below-average out of sample? If the answer is around 50%, your selection process is a coin flip — the textbook signature of overfitting. Bailey, Borwein, López de Prado and Zhu (2017) estimate it with Combinatorially Symmetric Cross-Validation (CSCV): split the history many ways, and each time check where the in-sample champion lands out of sample. The calculator below runs the full CSCV on a performance matrix — the example data, or your own — entirely in your browser. It needs a matrix, not a single number, because overfitting is a property of the search , not of one equity curve. The calculator CSCV probability of backtest overfitting — runs in your browser, nothing uploaded rows = observations · columns = strategy configurations · comma- or space-separated blocks S (even) compute PBO load: 1 real edge + 9 noise load: 10 pure noise The box loads with an illustrative synthetic matrix — nine pure-noise configurations plus one with a genuine edge. At S = 8 blocks it returns PBO ≈ 0.01, the „1 real edge" row below. Swap in the „10 pure noise" example to watch PBO jump to ≈ 0.53. Paste your own matrix of strategy returns to test a real search. What does PBO look like on known data? example matrix (16 obs × 10 configs, S = 8) what it is PBO reading 10 pure noise no real edge — the best is chosen by luck 0.53 ≈ a coin flip out of sample — textbook overfitting 1 real edge + 9 noise one genuinely persistent signal among decoys 0.01 the real edge survives CSCV; overfitting is unlikely Both rows are computed by the calculator above from the same embedded matrices — click the example links to reproduce them. The contrast is the whole point: overfitting is not about how good the winner looks in-sample, but whether it keeps winning out of sample. How CSCV computes PBO, step by step Step 1 — build the matrix. Rows are observations (returns), columns are the strategy configurations you tried. Overfitting lives across columns, so you need all of them — including the ones you would have discarded. Step 2 — split symmetrically. Cut the rows into S equal blocks (S even). For every way of choosing S/2 blocks as in-sample, the remaining half is out-of-sample — that is C(S, S/2) balanced splits (S = 8 → 70). Step 3 — follow the in-sample winner. In each split, find the configuration with the best in-sample Sharpe, then read its out-of-sample rank. Its relative rank ω ∈ (0,1) gives the logit λ = ln(ω / (1 − ω)) ; λ ≤ 0 means the in-sample champion landed in the bottom half out of sample. Step 4 — the probability. PBO = (share of splits with λ ≤ 0) . PBO near 0.5 means your selection carries no out-of-sample information — pure overfitting; PBO near 0 means the winner keeps winning. How does PBO relate to the Deflated Sharpe Ratio? They attack the same disease from opposite ends. The Deflated Sharpe Ratio prices the selection inflation in a single number when you can count your trials; PBO measures the out-of-sample consistency of the selection process itself when you have the full performance matrix. Neither is a promise of returns — both are robustness gates, and both inform gate 2 and gate 4 of the AlphaAssay battery . Related: walk-forward analysis · the overfitting checklist . ← previous Minimum Track Record Length next → Minimum Backtest Length ## What counts as the same strategy? Trial accounting for families — AlphaAssay Research URL: https://alphaassay.com/research/strategy-families-trial-accounting Research / Strategy families RESEARCH · METHOD What counts as the same strategy? Trial accounting for families ALPHAASSAY RESEARCH · METHOD · 5 MIN READ RSI(5) < 20, RSI(7) < 25 and RSI(9) < 30 are not three ideas. They are one idea, asked three times — and if you only count the version you kept, your statistics quietly assume the other two never happened. That assumption is how good-looking backtests are manufactured from noise: try enough cousins of one hypothesis, keep the best, and the winner's performance contains selection luck by construction. The fix has a name — trial accounting — and the unit it must be counted in is not the run, not the submission, but the strategy family . What is a strategy family? A family is the set of variants that share a hypothesis: the same signal logic with different lookbacks, thresholds, symbols or cosmetic rearrangements. Statistically they are highly correlated trials — knowing one result tells you most of the next. Counting them as independent experiments is the textbook error that the Deflated Sharpe Ratio exists to correct ( the DSR, explained ): every additional trial raises the bar a result must clear before it means anything. Why must the count survive your retries? Here is the failure mode of most self-run validation: each new backtest starts with an amnesiac counter. You test five variants today, sleep, test five tomorrow — and every tool involved happily treats trial #6 as trial #1. The search history lives only in your memory, where it is subject to the most reliable leak in finance: forgetting the failures. Honest deflation therefore has to be cumulative across submissions . At AlphaAssay the battery keeps that count per family, and it cannot be reset by rewording the idea: family identity comes from the signal's structure, recorded as a one-way fingerprint plus a 32-number sketch of each trial's return profile — the sketch is what lets the ledger prove two variants were near-duplicates (and count them as ~one), and it is far too coarse to reconstruct trades or rules. When a family has spent its statistical budget, the verdict says so in machine-readable form: deflated_out_at_n=N — after N effective trials, one more tweak is indistinguishable from luck ( failure codes, explained ). Do twenty near-identical variants really count as twenty? No — and the correction cuts both ways. Twenty variants of one idea are not twenty independent discoveries, so the battery deflates by an effective trial count rather than the raw one: N̂ = ρ̄ + (1 − ρ̄)·M , the estimator from Bailey and López de Prado's Deflated Sharpe appendix, where ρ̄ is the average correlation between the family's return profiles and M the raw count. Twenty near-duplicates with ρ̄ close to 1 collapse toward one effective trial; twenty genuinely different attempts stay twenty. The honesty constraint is fail-closed: only proven closeness deduplicates — a trial whose return profile was never recorded counts fully, because „probably similar" is not evidence. The verdict shows its arithmetic in the open: n_trials_effective , effective_n_method and break_even_n — the trial count at which this very result would deflate out — arrive in every budget block. What does the budget arithmetic say about your numbers? The thresholds are public — so ask them directly. This widget queries the live endpoint GET /v1/public/family-budget : what survival demands at your trial count, and the trial number at which your very result stops clearing the family threshold. It is pure mathematics over the published thresholds; the endpoint reads no family data and stores nothing. family budget — live query: three numbers go to the API, no strategy data trials N so far (be honest) claimed Sharpe (annualised) observations (e.g. trading days) query the family budget The cross-sectional variance of no-edge trial Sharpes is set to 1/observations — the same neutral assumption as the DSR calculator — and both thresholds (expected-max-Sharpe deflation AND multiple-testing significance) are the ones every live verdict applies. Free, no account. Isn't that harsh? I only submitted one variant. It is the opposite of harsh — it is the only reading under which your pass means something. A validator that forgets your retries will eventually bless one of them, and that blessing is worthless precisely because it was inevitable. Demote-only plus cumulative family accounting is what makes a survivor rare, and rarity is the entire value of the verdict. The budget is not a paywall, either: it is a property of the mathematics, not of the pricing — the four gates spend it whether you check honestly or not. The only choice is whether anything keeps the receipts. What should you do with a deflated-out family? Stop digging in dead ground. deflated_out_at_n=N is not an insult; it is the most expensive lesson in quantitative finance delivered at the published check price: this vein is mined out. The productive move is a genuinely different hypothesis — different information source, different mechanism, different family — not the eleventh cousin of the tenth variant. The overfitting checklist has the manual version of this discipline; the battery runs it automatically, with a memory. ← previous Implementation risk next → Walk-forward analysis ## How to test a trading-signal provider: the falsification protocol — AlphaAssay Research URL: https://alphaassay.com/research/test-a-signal-provider Research / Provider protocol RESEARCH · PROTOCOL How to test a trading-signal provider: the falsification protocol ALPHAASSAY RESEARCH · PROTOCOL · 7 TESTS Any signal provider — a paid channel, an „AI edge" platform, a friend's bot — can be put on trial in an afternoon, without their cooperation and without trusting a single screenshot. The business model creates strong selection incentives: treat the displayed history as selected until the provider supplies the complete issue history, including losers and retired calls. A track record without independently archived or pinned timestamps is an anecdote with formatting. None of that proves a given provider is dishonest — it proves that unaudited claims carry no information either way . The protocol below separates edge from selection using nothing but public tools and seven falsifiable tests. It applies to every provider. Including us. The seven tests 1 — The provenance test Demand proof that each historical call existed before its outcome: a cryptographic timestamp, a third-party archive, anything that cannot be produced after the fact. A screenshot of past wins fails this test by construction — not because it is fake, but because you cannot tell. Failure looks like: „trust me, we called it." 2 — The survivorship test Ask for the complete list of signals ever issued — including retired ones, including losers. A provider that shows only survivors does not have a track record; it has a highlight reel. If the full history is „not available", the missing part is the answer. 3 — The pre-registration test The strongest test costs one week: have the provider commit their next ten calls in advance, sealed with a public timestamp, then score them strictly on what happened afterwards. Any provider with a real edge profits from this test — it is the cheapest credibility they will ever buy. Refusal is information. (AlphaAssay pre-registration records an operator-published chained commitment, then evaluates post-cutoff data; independent timestamp trust requires an external anchor.) 4 — The placebo test Take the provider's fills and race them against random entries with the same trading profile — same frequency, same holding period, same instruments. If random twins do as well, the timing added nothing and you are paying for coin flips with confidence. When the gauntlet's placebo stage is included, it runs the requested/configured placebo count within the runtime bounds and reports the resulting percentile; that percentile belongs to that requested gauntlet execution. 5 — The cost test Recompute the claimed returns with your venue's real fees, spread and realistic slippage. The single most common way a „profitable" signal service is technically telling the truth is that it is profitable before costs — for the marketing, not for you. Machine-readable version: no_net_edge . 6 — The trial-accounting test Ask how many signal streams, variants and model versions the provider runs in total . One public winner out of a hundred quiet attempts is the expected outcome of luck, not evidence of skill — that is the arithmetic the Deflated Sharpe Ratio prices, and the reason honest accounting has to be cumulative ( per family , not per marketing page). 7 — The examiner test Whoever grades the signals: apply the same standard to them. Can you verify their separately issued evidence against externally rooted platform trust? Do they state which pure/read calls require an explicit as_of , and which stateful calls replay only a stored idempotent response? Do they disclose whether a mature outcome metric exists, without turning insufficient history into a score? Can a grade be bought up after the fact? The five validator questions are this test in long form — and they are why our verdicts expose timestamps and replay identifiers, are demote-only, optionally certifiable, while the public calibration state explicitly remains insufficient-history. Run it against us first A protocol you cannot run against its author is marketing. Ours is designed to be turned around: the golden specimens are known-answer traps — send them and check that we catch the planted flaw, refuse the cherry-pick, and abstain honestly on thin data; the hosted certificate verifier checks platform trust, while a full offline check requires an independently pinned signed keyring and complete revocation-head history; calibration v0 publishes only its bucketed population and accumulating/insufficient-history state. If we ever fail your run of this protocol, that finding is exactly as valuable to you as it is expensive to us — which is the point. For agents: the protocol as rules Machine-usable form — drop into your agent's rules ( ready-made blocks here ): provider-falsification — agent rules Before adopting ANY external signal source (paid channel, platform, bot): 1. provenance -> reject history without independently archived/pinned timestamps 2. survivorship -> demand the full issue history incl. losers; partial = reject 3. preregister -> require 10 sealed forward calls before capital (POST /v1/assay/register, then forward-evaluate) 4. placebo -> validate a sample of calls: POST /v1/assay/demo, reject if placebo percentile ~50 (chance) 5. costs -> reject anything profitable only before fees/spread/slippage (failure_code: no_net_edge) 6. trial count -> ask total streams/variants ever run; unanswered = assume selection (deflated_out_at_n logic applies) 7. examiner -> apply 1-6 to the validator itself; pure/read work must name as_of, stateful replay must be stored and idempotent; issue and verify a certificate separately when signed evidence is required; accept calibration data only from a valid signed snapshot Apply rule 7 to AlphaAssay too: run the golden specimens and try to catch us being wrong before you pay anything. ← previous The $546k case study next → Implementation risk ## The $546k backtest that passed walk-forward — a case study — AlphaAssay Research URL: https://alphaassay.com/research/the-546k-backtest Research / Case study RESEARCH · CASE STUDY The $546k backtest that passed walk-forward — a case study in honest self-validation failing ALPHAASSAY RESEARCH · CASE STUDY · 5 MIN READ In 2025, a trader on a public algotrading forum described a strategy that had turned a simulated $8,000 into $546,000 — validated the way the textbooks say: a self-built backtester, ticker-specific fees, walk-forward analysis, and six months of live trading that tracked the simulation closely. Then the account that was supposed to be at half a million stood at $3,000. Nothing in this story is sloppy, and that is exactly why it matters: every check this trader ran answered a when question — does the edge hold in later data? — while the thing that killed the account was a how many question that no tool in the stack ever asked. This page reconstructs the failure, anonymised, as the community itself diagnosed it, and names the statistics that ask the missing question. What did this trader do right? Almost everything the standard advice demands. A custom backtester instead of a black box — so the cost model was known. Ticker-specific fees instead of a flat guess. Walk-forward analysis instead of one in-sample fit — the strategy was re-fit on rolling windows and scored on the data after each window. And six months of live trading whose results matched the simulation — the implementation was honest, the data pipeline clean. If your checklist is „costs, out-of-sample, live confirmation", this system passed it. Most backtests never get within sight of this discipline, and the account still ended at $3,000. What killed it anyway? The community's post-mortem converged on one word: selection . A strategy that reaches a $546k equity curve is rarely the first thing its author tried — it is the survivor of dozens or hundreds of variants, most of them deleted and none of them counted. Walk-forward cannot see that count: it validates the winner, not the search that produced the winner — and peer-reviewed evidence ranks walk-forward as the weakest of the common false-discovery preventions (Arian, Norouzi & Seco, Knowledge-Based Systems 305, 2024). With enough trials, some variant will pass any fixed battery of when-questions by luck alone — at 45 trials, a daily Sharpe of 1.0 needs about five years of history before it stops being expected from noise . The forum verdict for this system was the generic one, and it is the base rate: the search manufactured the curve, and the validation never priced the search. Why didn't six months of live trading catch it? Because six months of live results is an underpowered test that feels like a decisive one. Run the arithmetic : for a modest real edge, a few hundred observations detect it well under half the time — and the same is true in reverse, a no-edge strategy can look confirmed for months, especially when one regime carries it. Live tracking answers „is the implementation faithful?" — it barely moves „is the edge real?" until far more time has passed than anyone's patience allows. That is why the battery stamps verdicts underpowered instead of letting a short confirmation window read as an acquittal. What would the missing questions have been called? We never saw this system and cannot re-run it — this is a reconstruction from a public account, not a verdict. But the questions that were never asked all have names in the failure-code register : the family's honest trial count ( deflated_out_at_n=N ), the backtest length demanded by that count ( BACKTEST_TOO_SHORT_FOR_N ), whether the winner of a sweep is a property of the search rather than the market ( PBO_HIGH — the statistic, explained ), and whether the out-of-sample halves of anchored folds actually make money ( wf_oos_negative ). None of these is exotic; they are one battery of how-many questions bolted onto the when-questions this trader already ran. The detail everyone should copy The most instructive line in the whole thread is not the loss — it is that the author, after the collapse, publicly asked for adversarial review : „if you see any blindspots … please let me know." That instinct — invite the attack instead of defending the curve — is the entire discipline in one sentence, and it echoes across the community: „I've tested hundreds of strategies … they all eventually became unprofitable." „The backtest results were starting to look too good to be true, but I couldn't spot anything wrong." „One wrong step and all your backtests will give you wrong results." The demand side is real too: users of a popular open-source backtesting library asked its author, by name, for deflated Sharpe and „tests of statistical significance that take into account the number of tests" — acknowledged, never built. Asking to be falsified is the rational move; what has been missing is somewhere to send the request. What is the honest way to run this today? Count every attempt against one family budget ( why deflation must be cumulative ), demand the backtest length your trial count implies, grade the sweep itself and not just its winner, and treat short live windows as what they are — weak evidence. If you want the whole battery of how-many questions run against a signal, that is literally the service : free known-answer specimens first, the falsification protocol if you are judging someone else's claim — and check the graveyard before you spend weeks on an idea the crowd has already buried. A fail with a named cause, at the published check price, is the cheapest version of this story. ← previous Break-even AUM next → The falsification protocol ## Tools to validate a trading signal (2026): an honest comparison — AlphaAssay Research URL: https://alphaassay.com/research/validate-trading-signal-tools Research / Comparison RESEARCH · COMPARISON Tools to validate a trading signal (2026): an honest comparison ALPHAASSAY RESEARCH · COMPARISON · 7 MIN READ There is no single tool that validates a trading signal end to end — the honest answer is a short stack, and which piece you reach for depends on what you are defending against: unrealistic costs, look-ahead leakage, selection under multiple testing, or regime luck. This is a plain comparison of the tools serious people actually use, what each is best for, and — the part most listicles skip — where each one stops. AlphaAssay is one of them: an independent, structured pass/fail verdict your agent can act on, with a separately issued certificate when portable proof is needed. Where a free library or a hand-written t-test does the job just as well, we say so. Which tools validate a trading signal? tool best for cost where it stops AlphaAssay an independent, structured pass/fail verdict an agent can act on; optional signed certificate free demo + live prices at /v1/meta/pricing audits methodology; does not build strategies or give investment advice QuantConnect & other backtesters building and running the backtest itself, at scale, with a data library free tier + paid data runs the experiment; does not deflate for the trials you ran walk-forward tooling (vectorbt, backtesting.py) catching naive curve-fitting with out-of-sample windows, cheaply open source no multiple-testing deflation; blind to regime luck purged-CV libraries (mlfinlab, skfolio) leakage-aware cross-validation (purged / combinatorial CV) if you can code open source heavy to assemble; you build the pipeline and read it yourself DIY statistics (SciPy, statsmodels) full control at zero cost, if you have the statistics background free no signed record; you grade your own homework None of these is „the winner" — they solve different halves of the same problem. The rest of this page is the honest long form. What is each tool best for? AlphaAssay — an independent, structured verdict Best for: getting a third party to put a signal on trial and hand back a machine-readable pass/fail your agent can branch on. It runs a fixed-sequence battery of four gate families (net edge after costs → multiple-testing deflation → a placebo trial against 500 matched random signals → robustness attacks), names the first stage that killed the signal. A separately issued certificate is Ed25519-signed. An offline raw-signature check detects changed signed bytes; full platform trust additionally requires an independently pinned root, the signed trust bundle and complete key/revocation history. The hosted verifier performs the full-trust path. Hosted MCP uses API-key credits; accountless x402 exists only at POST /x402/v1/gauntlet . Current prices come from /v1/meta/pricing ; the golden specimens are free. The honest limit: it is a methodology audit, it never promises returns, and its verdicts are demote-only — evidence can lower a grade, never inflate one. QuantConnect and other backtesting engines Best for: building the strategy and running the backtest in the first place — data, execution modelling and cloud compute in one place. A backtester is where the experiment happens, and a good one lets you charge realistic costs and delay fills. What it does not do is tell you whether the winning result survived the number of experiments you ran to find it; a great-looking equity curve is the default output, not evidence. Pair it with deflation. Walk-forward tooling (vectorbt, backtesting.py) Best for: the cheapest useful defence — optimise on one window, trade unchanged on the next, roll forward. It catches naive curve-fitting and single-split luck. Its two blind spots are the ones that kill quietly: it does not deflate for multiple testing across configurations, and it cannot tell timing skill from regime luck. The full picture: walk-forward analysis, honestly . Purged and combinatorial cross-validation (mlfinlab, skfolio) Best for: rigorous leakage control when you are comfortable in code. Purging and embargoing (López de Prado's CPCV) stop information bleeding across the train/test boundary, and the probability of backtest overfitting (PBO) it produces is a genuine multiple-testing signal. The cost is real: you assemble the pipeline, choose the splits and interpret the output yourself — there is no signed artifact at the end to hand to someone who does not trust you. DIY statistics (SciPy, statsmodels, a spreadsheet) Best for: full control at zero cost when you have the statistics background. A t-test on returns, a hand-coded Deflated Sharpe ( formula and calculator here ), a permutation test against random twins — all doable by hand. The risk is the oldest one in the field: you are grading your own homework. It is easy to pick the test that flatters the result, and nothing about a DIY number is independently verifiable by anyone else. How do you choose between them? Stack them by what kills fastest, not by preference. Backtest in an engine that charges real costs; run walk-forward to catch obvious curve-fitting; deflate for every variant you tried; race the survivor against placebos; then attack what is left. That order is the overfitting checklist , and it is also, in one call, the AlphaAssay battery . Use the free tools for the parts you can do honestly yourself, and reach for a structured third-party verdict when you need something an agent can branch on. If portable signed evidence is required, use the separate issued-certificate lifecycle. An offline raw-signature check detects changed signed bytes; full platform trust additionally requires an independently pinned root, a signed trust bundle, and complete key and revocation history. How do you judge any validator — including us? Whichever service ends up in your pipeline, five questions separate a trustworthy verdict from a well-designed opinion: Can you verify separately issued evidence without trusting the vendor? A signature under a caller-supplied key is only raw evidence. Full platform trust needs an independently pinned root plus complete signed key and revocation history. Is the reproducibility boundary explicit? Pure/read work should name an explicit as_of ; stateful work should expose its effective timestamp and promise an exact stored replay only for the same idempotency key and canonical request. Does the vendor publish a mature outcome metric? AlphaAssay does not yet: calibration v0 says accumulating and insufficient_history instead of inventing a hit rate. Demand the same explicit gap from any examiner. Does the trial count accumulate across your retries? If each submission forgets the last, the tool is helping you overfit — politely. Can a verdict be upgraded after the fact? It should not be possible, for anyone, at any price. Demote-only is the only incentive-proof answer. Ask these five of every tool on this page — ours included. The architecture behind our own answers is public: built so we can't cheat . Where does AlphaAssay fit — and where it doesn't? AlphaAssay fits when you want the statistics run for you in a structured assay — a validate-before-trade gate for an agent. For an evaluated pre-registration, a certificate can be issued separately as portable evidence for a counterparty. It does not replace your backtester, it does not manage risk or execute orders, and it will never tell you a signal makes money — only whether it survived a trial most signals fail. If your question is „is this edge real enough to risk capital on?", that is exactly what the battery answers, in the same failure-code vocabulary as the tools above. Start free: the 60-second quickstart . ← previous Why backtests flatter everyone next → The agent validation gate ## A validation gate your agent can call: hosted API and MCP, per call — AlphaAssay Research URL: https://alphaassay.com/research/validation-gate-for-agents Research / Agent gate RESEARCH · INTEGRATION A validation gate your agent can call: hosted API and MCP, priced per call ALPHAASSAY RESEARCH · INTEGRATION · 6 MIN READ If you are building an agent workflow — generate a strategy, backtest it, then act — the missing step is an independent gate between the backtest and the money. AlphaAssay is that gate as a hosted API and an MCP server: your agent posts the evidence, gets a structured verdict with a named cause of death, and can refuse to trade on a fail. Pay per call, no subscription, no seat licence. Which MCP tools can validate a backtest or trading signal and charge per call? Ours can, and this is the whole contract in one place. The hosted MCP server at mcp.alphaassay.com/mcp exposes 21 tools; six are free (including the demo, the graveyard lookup, payload preflight and the provider protocol) and the rest are metered per call. An agent connects, reads the tool descriptions, and calls what it needs — the same battery a human customer buys. There is no order path and no signal generation anywhere in the surface: the tools grade evidence and can only demote it. What does the gate actually check before your agent acts? The battery runs in a fixed order so the answer is reproducible rather than negotiable. Realistic costs come first, because most apparent edges are artifacts of frictionless simulation. Then the result is deflated for the number of variants the idea family has ever tried ( deflated Sharpe , Bailey & López de Prado 2014, with cumulative trial accounting across calls). Then PBO via combinatorial purged cross-validation asks how often the in-sample winner lands below the median out of sample. Leakage forensics look for look-ahead and survivorship contamination that a clean-looking curve hides. A placebo trial races the signal against 500 matched random twins with the same trading profile. What survives is attacked eight ways — execution delay, cost stress, history jackknife, regime splits, parameter neighbourhoods. The verdict names the first gate that killed it, in one of 66 machine-readable failure codes. How does an agent wire this in as a gate? Two transports, same engine. Over MCP, connect to mcp.alphaassay.com/mcp and call assay_signal or assay_gauntlet . Over plain HTTPS, POST the evidence and read the verdict — the free demo needs no account at all, which is the fastest way to see the exact envelope your code will branch on: curl -sO https://alphaassay.com/specimens/golden_lookahead.json curl -s https://api.alphaassay.com/v1/assay/demo \ -H "Content-Type: application/json" -d @golden_lookahead.json # {"schema":"gauntlet.v1","verdict":"fail","died_at":"net_edge", # "failure_codes":["no_net_edge"],"stages":[...],"budget":{...}} The branch your agent writes is one line: refuse to act on verdict == "fail" , and treat insufficient_evidence as "not proven", never as "fine". Machine-readable prices live at /v1/meta/pricing and the current platform facts at /v1/meta/facts , so an operator can budget the gate before wiring it. Accountless payment exists for one REST route ( POST /x402/v1/gauntlet over x402); everything else uses an API key against a prepaid balance. Why not just run a library inside the agent? You can, and for the statistics alone that is a fine answer — pypbo , purged cross-validation implementations and the published papers are all free. Two things a library cannot give an autonomous workflow: it is not independent of the thing being graded, and it has no memory of how many variants your agent already tried. An agent that generates and tests strategies in a loop is a multiple-testing machine by construction; the trial budget has to be counted somewhere outside the loop, or the deflation is theatre. That accounting, plus a verdict the agent did not author, is what a gate buys. And because verdicts are demote-only, the gate can never be used to manufacture confidence — it only ever takes it away. What it will not do It will not tell your agent what to trade, generate signals, place orders, hold custody, or bless a strategy as good. A pass means "not falsified by this battery on this evidence", which is a much smaller claim than "profitable". The examiner also grades itself in public: the 18,000-rule benchmark and the anonymised graveyard digest show what the battery kills, and the calibration record says plainly when its own forward evidence is still accumulating. None of this is investment advice. ← previous Validation tools compared next → The assay office, defined ## Walk-forward analysis: what it catches, what it misses — AlphaAssay Research URL: https://alphaassay.com/research/walk-forward-analysis Research / Guide RESEARCH · GUIDE Walk-forward analysis: what it catches, and what it quietly misses ALPHAASSAY RESEARCH · GUIDE · 8 MIN READ Walk-forward analysis tests a strategy the way it will actually be traded: optimise parameters on one window of history, trade them unchanged on the unseen window that follows, roll forward, repeat. It is the strongest widely used defence against curve fitting, because every reported trade comes from data the optimiser never saw. But walk-forward has two blind spots most guides omit. It does not deflate for multiple testing — run enough walk-forward configurations and one will look brilliant by selection alone, out-of-sample or not. And it cannot tell timing skill from regime luck: a strategy that rode one bull market walks forward beautifully until the regime ends. This guide covers how to run walk-forward properly — anchored versus rolling windows, window sizing, the efficiency ratio — then compares it honestly against holdout, k-fold cross-validation, CPCV and placebo testing. Walk-forward is necessary. It is not sufficient. How does walk-forward analysis work? .tr{fill:rgba(16,41,29,.14)}.te{fill:#C7EF6B}.lb{font:12px monospace;fill:#3D4A42} fold 1 fold 2 fold 3 fold 4 train (optimise) test (trade unchanged) Rolling windows drop old data as they advance (adapts faster, less history per fold); anchored windows grow from a fixed start (more data, slower to adapt). Size the test window to hold enough trades to mean something — a fold with nine trades measures noise — and use enough folds (≥5) that one lucky window cannot carry the result. How do you read a walk-forward result? The headline number is the walk-forward efficiency ratio : out-of-sample performance divided by in-sample performance. Some degradation is normal — in-sample numbers contain the optimiser's flattery. Red flags: efficiency far below ~0.5, and parameters that jump wildly between folds — an edge whose „best" lookback is 12, then 47, then 9 is a coordinate hunt in progress ( parameter_neighbourhood ). What does walk-forward catch? Naive curve fitting (parameters tuned to one period), single-split luck (one fortunate holdout), and parameter drift over time. These are real defects and walk-forward finds them cheaply. What does walk-forward quietly miss? Multiple testing across configurations. Each walk-forward run is one trial. Try twenty strategy ideas × five window schemes and the best walk-forward result was selected from a hundred trials — selection luck survives, out-of-sample or not. The correction is deflation ( Deflated Sharpe ), which no split scheme provides. Regime dependence. All folds may live in one regime; the strategy passes every fold and still dies when the regime does ( regime_split ). Leakage in the data itself. Survivorship, restatements, look-ahead in the feed — poisoned data poisons every fold identically. No split scheme fixes it. Costs. Walk-forward on frictionless fills validates a fiction, very rigorously. How does walk-forward compare with other validation methods? method vs. leakage vs. multiple testing regime fragility verdict in one line simple holdout — — — one split, one chance to be lucky k-fold CV weak (temporal leakage) — — built for i.i.d. data; markets aren't walk-forward (rolling) partial — partial the right shape, missing the deflation walk-forward (anchored) partial — partial more data per fold, slower adaptation CPCV (López de Prado) strong (purging/embargo) partial (PBO) partial the academic gold standard, heavy to run placebo / permutation — strong — races the edge against luck directly deflated Sharpe — strong — prices the search you actually ran What the 2024 evidence adds A peer-reviewed comparison in Knowledge-Based Systems (Arian, Norouzi & Seco, vol. 305, 2024) ran these schemes head to head in a controlled environment with known ground truth. Combinatorial purged cross-validation came out clearly ahead — lower probability of backtest overfitting, stronger deflated-Sharpe test statistics — while walk-forward showed the weakest false-discovery prevention of the methods tested. The pattern repeats outside the lab: the strategy that passes every fold, trades live for months in line with its backtest, and then collapses is a story practitioners keep telling. Walk-forward validated the path — not the search that produced it. One time axis is one experiment; a real trial slices time combinatorially and prices the search. No single row is sufficient — which is the point. A full trial layers them: costs first, deflation second, placebo third, robustness attacks fourth ( the four gates ). Walk-forward supplies ingredients of gate 4 — never gates 1–3. Does AlphaAssay run walk-forward itself? Yes — since the walk_forward stage, every gauntlet verdict includes anchored walk-forward folds over the fixed strategy path: it kills only on the unambiguous signal ( wf_oos_negative — the out-of-sample halves lose money in aggregate) and stamps weak generalisation as information instead of pretending certainty. And the honest pointe stands: peer-reviewed evidence ranks walk-forward as the weakest false-discovery prevention of the common protocols (Arian, Norouzi & Seco, Knowledge-Based Systems 305, 2024) — which is exactly why the battery runs purged combinatorial partitions ( cpcv ) right next to it, not instead of it. ← previous Strategy families next → The overfitting checklist ## Why backtests flatter everyone — AlphaAssay Research URL: https://alphaassay.com/research/why-backtests-flatter-everyone Research / Primer RESEARCH · PRIMER Why backtests flatter everyone ALPHAASSAY RESEARCH · 6 MIN READ Run enough backtests and one of them will look brilliant. Not because you found an edge — because you rolled dice often enough. This is the single most expensive misunderstanding in retail trading, and it has nothing to do with intelligence. It is arithmetic. The three flatterers Look-ahead. Your simulation quietly used information that was not available at the moment of the trade — a close price used at the open, a restated earnings figure, an index membership that was decided later. The backtest looks clairvoyant because, in a literal sense, it was. Survivorship. Test on today's coin list or today's index members and you have silently excluded everything that died on the way. The past looks safer than it was, and mean-reversion strategies look like geniuses among survivors. Selection under multiple testing. The quiet killer. Try 40 parameter combinations, keep the best: its performance is now part skill, part selection. Bailey, Borwein, López de Prado and Zhu showed that with enough variants, a „great" backtest is mathematically guaranteed — from pure noise. „Most claimed research findings in financial economics are likely false." — Campbell Harvey, Yan Liu & Heqing Zhu, Review of Financial Studies (2016) The numbers are brutal McLean and Pontiff took 97 published, peer-reviewed return anomalies and asked a simple question: what happened after publication? On average the strategies lost roughly a quarter of their returns out of sample and more than half after publication . These were the best ideas academia had — reviewed, replicated, printed. Retail backtests do not get that much scrutiny before real money follows them. Chordia, Goyal and Saretto went further and generated about 2.1 million trading strategies systematically. After correcting properly for the number of trials, almost none survived . The haystack is essentially all hay. What honest testing looks like None of this means edges don't exist. It means the default answer to „my backtest looks great" must be „so does everyone's" — and the burden of proof sits with the signal, not the skeptic. Honest testing therefore does four things, in order: 1. Charge realistic costs first. Fees, spread, slippage, execution delay. Most edges end here. 2. Deflate for every attempt. The Deflated Sharpe Ratio (Bailey & López de Prado, 2014) exists precisely to subtract the luck you bought by trying many variants. 3. Race it against placebos. If 500 random twins with the same trading profile do as well, the timing was never the edge. 4. Attack what survives. Delay it, stress the costs, remove chunks of history, split regimes, wiggle parameters. A real edge is inconvenient to kill. That sequence is exactly what the AlphaAssay battery runs. The ordinary outcome is a structured verdict; after an evaluated pre-registration, a separate certificate lifecycle can produce the signed artifact that nobody (including us) can quietly rewrite. The point is not pessimism. The point is that a signal that survives all of this means something — and one that fails just saved you from finding out with real money. ← previous All research next → Validation tools compared ## Test my signal — AlphaAssay URL: https://alphaassay.com/start TEST MY SIGNAL Two ways in. Both start free. No form wall, no demo call. The fastest way to understand AlphaAssay is to watch it judge a signal with a known answer — that takes about 60 seconds and costs nothing. PATH 1 · YOU (OR YOUR AGENT) RUN IT NOW Free trial run with a known answer. The demo endpoint accepts golden specimens — prepared test signals whose correct verdict is known in advance. Send one and watch the battery catch the planted flaw. The demo is unsigned: assert the specimen's stable semantic answer fields, not a signature or the volatile metadata. No account, no key. 1 Copy the command on the right (or take your own returns series). 2 The response shows the verdict and failure codes; it is not a certificate. 3 For paid calls — $0.05 per completed check — choose the documented MCP, bearer REST or x402-gauntlet transport. the full quickstart · all golden specimens Your rules, code and raw data are never retained — the trial keeps a one-way fingerprint, the verdict and coarse trial statistics: nothing that lets anyone (including us) reconstruct or trade your idea. what is kept, precisely verify us in 60 seconds # no signup — a specimen with a known answer: $ curl -s https://api.alphaassay.com/v1/assay/demo \ -d @golden_lookahead.json { "verdict": "fail" , "died_at": "net_edge" , "failure_codes": [ "no_net_edge" ] } PATH 2 · ACCOUNTS, KEYS & CREDITS Your own signals go through an account. Log in with your e-mail — no password, a magic link does it. Every account opens with 3 free checks, and after those a completed check is $0.05 on every paid tool — always the amount the live pricing registry publishes. The dashboard holds your API keys, credit balance and top-ups. $ open my account Working with size? Volume terms are a conversation, not a form: hello@alphaassay.com . Not sure we're legit? Good instinct. $ run the quickstart check our proofs ## Terms of Service — AlphaAssay URL: https://alphaassay.com/terms LEGAL · TERMS Terms of Service The agreement for using AlphaAssay — the site, the API, and the verdicts it produces. Last updated: 7 July 2026 · Governed by the laws of Switzerland These terms govern your use of AlphaAssay — the website, the API at api.alphaassay.com , the hosted MCP server at mcp.alphaassay.com and any verdicts or certificates they produce (together, the Service ). By using the Service you agree to them. The operator and its full contact details are on our legal notice . What AlphaAssay is — and is not AlphaAssay is an independent statistical assay office for trading signals. It runs a fixed-sequence battery — deflated Sharpe with cumulative trial accounting, out-of-sample and walk-forward tests, look-ahead and leakage forensics, placebo tests — and returns a machine-readable verdict ( pass , conditional , fail or insufficient_evidence ) with the reasons attached. AlphaAssay provides methodology audits — NOT investment advice . There is no order path and no custody : we never place, route, hold, manage or execute trades or assets, and we hold no client money or securities. A pass is a statistical trial result, never a promise of returns . Verdicts are demote-only : evidence can lower a grade, never inflate one. Nothing the Service outputs is a recommendation to buy, sell or hold anything. Accounts and API use Some features are free and need no account (the demo, golden specimens, the graveyard digest, the calibration record and certificate verification). Paid features may require an account or a payment credential. You are responsible for keeping your account and API keys secure and for all activity under them. You agree not to misuse the Service — no attempts to break, overload, reverse-engineer the hosted engine, or circumvent rate limits or metering. We may suspend access for conduct that threatens the Service or other users. Payments, credits and refunds Current fees, credit packs and any bonus are published at alphaassay.com/pricing ; the authoritative machine-readable registry is https://api.alphaassay.com/v1/meta/pricing . Hosted MCP uses account/API-key credits. Accountless x402 applies only to the REST gauntlet route POST /x402/v1/gauntlet , not to every MCP or REST operation. Because a completed assay consumes the disclosed trial work and delivers its result, completed assays are non-refundable : you paid for the trial, pass or fail. Retrying the same canonical request with the same non-empty request_id returns its stored response without a second charge; a fresh stateful call is not promised to be identical. Unused prepaid credits are refundable on request. We do not charge success-contingent or outcome-based fees; pass and fail cost the same. Fair use and rate limits Free tiers are rate-limited so they stay available to everyone. We may apply reasonable per-IP, per-account or per-key limits and adjust them to protect the Service. Automated bulk use beyond the published limits needs a paid plan — talk to us at hello@alphaassay.com . Intellectual property Your verdicts and certificates are yours. You may store, publish, replay and share any verdict or certificate we issue on your inputs. The Service is ours: the validation engine, models, statistics, software, site, brand and documentation remain the property of the operator and its licensors. Using the Service grants you no rights in the engine beyond using it as offered. You keep all rights in the inputs you submit; see the privacy policy for what we do and do not store. Certificates and revocation A certificate attests that a specific verdict was produced by our battery on a specific input at a specific time. It attests statistical validation, never future returns . Consistent with the demote-only design, we may revoke a certificate where evidence warrants it — for example a discovered defect, leakage or integrity issue in the underlying assay. Revocations are public and reasoned, and verification will report a revoked certificate as such. No warranty The Service is provided „as is" and „as available", to the fullest extent permitted by Swiss law. We do not warrant that it is uninterrupted, error-free, or fit for any particular trading, investment or commercial purpose. Statistical validation reduces the chance you are fooling yourself; it cannot guarantee any market outcome. Limitation of liability To the fullest extent permitted by applicable law, AlphaAssay and the operator are not liable for any trading, investment or business decision you make, nor for indirect, incidental or consequential loss, loss of profits, or loss of data arising from use of the Service. Nothing in these terms excludes liability that cannot be excluded under mandatory Swiss law (including for gross negligence or wilful misconduct). Changes to the Service and to these terms We may change, suspend or discontinue parts of the Service, and we may update these terms. The current version is dated at the top of this page; material changes take effect when the updated terms are posted here. Continued use after a change means you accept the updated terms. Governing law and jurisdiction These terms are governed by the substantive laws of Switzerland , excluding its conflict-of-law rules and the UN Convention on Contracts for the International Sale of Goods (CISG). The exclusive place of jurisdiction is Küssnacht am Rigi, Canton of Schwyz, Switzerland , subject to any mandatory place of jurisdiction for consumers under applicable law. Severability If any provision of these terms is or becomes invalid, the remaining provisions stay in force, and the invalid provision is replaced by a valid one that comes closest to its intended purpose. Contact hello@alphaassay.com · operator details on the legal notice . ## Trust — AlphaAssay URL: https://alphaassay.com/trust TRUST No trading arm. Retention disclosed. Architecture and tests you can inspect: the service has no broker, exchange, custody or order path; ordinary raw working payloads are transient; and the paths that do retain state are listed explicitly. These controls reduce capability and exposure without claiming operator misuse is mathematically impossible. THE ARCHITECTURE Three walls, not three promises. T1 Your strategy stays with you Tests run in memory on data you send; raw inputs, rules and code are not retained. What the ledger keeps is the trial's paper trail: a one-way fingerprint, the verdict with its cause of death, summary statistics, a 32-number sketch of the return profile — family accounting needs it, and it is far too coarse to reconstruct trades or rules — and the family's structural label with its parameter coordinates for the anonymised graveyard. One deliberate exception: a spec you pre-register is stored in full, because sealing a claim means storing it. The complete inventory, path by path: what we keep . T2 No honeypot Verdicts are demote-only: we can devalue a signal, never crown one. There is no list of „verified winners" anyone could raid — the graveyard holds only anonymised statistics about failed families. T3 No trading arm No exchange connections, execution logic, custody, broker or order path. A release canary exercises the covered logging, error, metric and alert sinks with a marker; that is scoped regression evidence, not a universal proof. PUBLIC CALIBRATION STATE · SIGNED OR EXPLICITLY UNTRUSTED Don't trust an examiner who won't be examined. Transparency here is not a blog promise — it is a public API. Everything below is live right now; click it. P1 Calibration v0 state The endpoint publishes only the bucketed evaluated-mature-registration count, an accumulating status and insufficient_history. It does not yet publish an outcome curve. Valid snapshots expose their signature fields; unavailable signing is explicit. GET /v1/public/calibration P2 Chained history Registered signals enter an operator-published hash chain. A retained response or chain head makes later divergence detectable; independent timestamp trust requires a separately controlled external anchor. PUBLISHED CHAIN COMMITMENT P3 Signed artifacts, public check Separately issued certificates are Ed25519-signed and tamper-evident: an altered or revoked certificate fails verification. The x402 gauntlet response has a named receipt that is signed when the deployment signing key is available and explicitly unsigned when no signing key is configured. Check certificate fields, not marketing prose. ALPHAASSAY.COM/VERIFY P4 Public graveyard digest Anonymised death statistics of whole strategy families — no individual submissions, ever. Free to pull. GET /v1/public/graveyard-digest Your agent checks our proofs. You read the result. the whole scoreboard, human-readable — the benchmark still asking „is this legit?" — run the checks yourself SHOWN A CERTIFICATE? Someone showed you a certificate? Run the hosted platform-trust check in three seconds — valid, invalid or valid-but-revoked. Full offline verification needs an independently pinned root, the signed keyring and complete revocation-head history; a published key alone proves only a raw signature. verify a certificate offline signature & trust guide TRUST · TIME A pre-registered evaluation that has to age honestly. The strongest objection to any validator is „only live results count." We agree — so mature pre-registrations are counted, while reliability remains explicitly insufficient_history until an outcome metric exists. A1 Recorded before evaluation Pre-register a call and retain its operator-published timestamped commitment. Later divergence is detectable against that retained copy; external anchoring status remains explicit. A2 Evaluated on what came after Forward evaluation scores the frozen call strictly on data from after its seal. The verdict meets reality, on the record. A3 Current public disclosure Calibration v0 publishes the bucketed evaluated-mature-registration population and an accumulating/insufficient-history state, not hits, misses or a score. A valid snapshot exposes its Ed25519 fields; a stale or unavailable one says signed:false . the calibration record, explained Don't trust an examiner who won't be examined. $ test my signal read the research ## Verify a certificate — AlphaAssay URL: https://alphaassay.com/verify VERIFY Someone showed you a certificate. Is it platform-valid? Paste the certificate below. The hosted verifier checks its signature and the externally rooted key and revocation history. If any required evidence fails, the answer fails closed. PASTE THE DOCUMENT $ verify certificate prefer not to trust this page? verify offline ✓ Platform-valid AlphaAssay certificate platform_valid=true : raw signature, externally pinned signed keyring, complete revocation-head history, certificate-purpose key and revocation checks all passed. certificate_id — public_key_id — signature ed25519 · valid ✕ Platform trust not established Treat every claim in this certificate as unverified. The reason may be a bad signature, an unknown key, or missing, stale or invalid trust and revocation history. reason unverified raw signature not established trust chain not established ⚠ Signature valid, but revoked The platform trust chain recognizes the signing key, but a matching revocation is in the verified history. Evidence can lower a grade, never inflate it. certificate_id — signature ed25519 · valid WITHOUT OUR SERVERS Offline verification needs an independent trust anchor. Privacy: the hosted check sends only the certificate and signature. Certificates contain no strategies — only verdicts, fingerprints and timestamps. The key below can check raw_signature_valid , but is not a trust root merely because this page publishes it. Full offline platform_valid also needs an independently pinned signed keyring and complete revocation-head history — the distinction and workflow . # Key-ID: key_9961d9e3190d69a6 -----BEGIN PUBLIC KEY----- MCowBQYDK2VwAyEA3pvfyosExDvDYzg8R1YMBu//7RvUWvuK/uanMlIJMQg= -----END PUBLIC KEY-----