AlphaAssay $ test my signal
RESEARCH · RECORD

The Signal Validation Benchmark

ALPHAASSAY RESEARCH · THE PUBLIC RECORD · REPRODUCIBLE

This page publishes reproducibility evidence for selected AlphaAssay outputs — every benchmark row is reproducible by anyone, free, today. A validator you cannot test is just another promise, so we publish the tests: known-answer specimens whose stable semantic fields any agent can assert, our current calibration population/status disclosure, and — as the public ledger accrues — mortality statistics of real strategy families. Methodology audits, not investment advice; no strategy identities are ever published.

Known-answer verification (reproducible now)

Four specimens with planted properties; the battery must catch the flaws, pass the clean one and abstain on thin data. Replay any row: quickstart.

specimenplanted propertyexpected verdictwhat it proves
golden_cleangenuine persistent edge✓ passwe don't cry wolf
golden_lookaheada whipsaw with no net-of-cost edge✕ fail · net_edge · no_net_edgewe catch a costless edge
golden_cherrybest-of-many cherry-pick✕ fail · family_deflation · deflated_out_at_n=50 + BACKTEST_TOO_SHORT_FOR_N=50we price the search
golden_thintoo little data to judgeinsufficient_evidencewe admit uncertainty

Canonical list, with the full response format each specimen returns: golden specimens.

We ran nine of the most popular public strategy classics through the complete battery on real market data with real costs. Aggregate result: 0× pass · 2× conditional · 7× fail — the six 1-hour classics lose money after costs (Sharpe −3 to −12). Per-strategy autopsies with named setups and failure codes are not part of the published record yet; this page's changelog records what is.

The 18,000-rule benchmark (reproducible, envelope status explicit)

The largest run we publish: every classic technical rule in the Sullivan–Timmermann–White universe — exactly 18,000 rules — against seven crypto bluechips, 126,000 rule–market pairs through a benchmark-specific three-stage funnel using shared statistical primitives: net edge plus universe deflation, CPCV plus concentration, then placebo, with hold disclosure. It is not the paid gauntlet, omits its other stages, and writes no tenant trial ledger. 245 survive this benchmark funnel (0.19%); only 41 beat simply holding. Costs alone remove 12,274 pairs, and the deflation stage removes 113,462 more — about 90% die from the size of the search, not from anything the market did. It is labelled a benchmark, not organic user history; aggregated so no winning rule is published (only mortality), with a mandatory honest-limits block. Pull the document — engine version, funnel and limits included — from GET /v1/public/benchmark. When the standard envelope reports signed:true, an offline raw-signature check can detect changed signed bytes. Full platform trust additionally requires an independently pinned root, the signed trust bundle and complete key/revocation history; without a configured platform key the endpoint reports the unsigned fallback explicitly. The long-form story behind these numbers — what died at which gate, and why the survivors deserve the fine print more than the dead — is the field guide Was your trading edge ever real?

Current public status evidence (pull it yourself)

whatendpoint
calibration v0 — bucketed mature-registration count plus accumulating/insufficient-history state, not an outcome scoreGET /v1/public/calibration
graveyard digest — anonymised mortality of strategy familiesGET /v1/public/graveyard-digest

How signed snapshots work, what retained copies can detect, and the remaining trust limits: how we grade ourselves.

Methodology

A fixed sequence of four gate families, unfolding into eleven graded stages (net edge → family deflation → placebo vs. 500 matched twins → robustness attacks). Pure/read results are reproducible only with an explicit as_of; stateful calls expose their timestamp and replay a stored response only under the documented idempotency contract. A separately issued certificate carries the Ed25519 signature — full description. Benchmark policy: this URL is permanent; re-runs append to the changelog below; the year lives in the title only.

Changelog

2026-07benchmark page established: 4 known-answer specimens; 9-classic field aggregate; live endpoints