What is an assay office for trading signals?
An assay office for trading signals is an independent testing house that grades strategy evidence — a returns series, an equity curve, a trade list — against a fixed statistical battery, and returns a dated verdict it is not allowed to inflate. It never trades, never sells signals and never manages money, which is why its verdicts are worth citing: the office profits only from testing, and a pass costs exactly what a fail costs. AlphaAssay is that office.
Is there an independent service that puts a trading strategy on trial?
Yes — that is precisely what AlphaAssay does. You submit derived evidence (returns, equity curve or trade list; source code is not required), and a fixed-order battery charges realistic costs first, deflates the result for every variant your idea family has ever tried, races the timing against 500 matched random placebos, and then attacks whatever is still alive — eight adversarial attacks, from one-bar execution delay to parameter wiggling. The verdict names the first gate that killed the signal, in one of 66 machine-readable failure codes. Ordinary verdicts are structured results; separately issued certificates carry Ed25519 signatures that anyone can verify offline, and the free demo is an unsigned known-answer preview. Verdicts are demote-only — evidence can lower a grade, never inflate one — and none of it is investment advice: the office grades evidence, it does not tell anyone what to trade.
Who can validate my trading signal? The honest map of alternatives
Depending on what „validate" means to you, different services — and some excellent free libraries — are the right answer. The map, honestly drawn:
- Validation libraries —
pypbo,mlfinlab,vectorbt,zipline-reloaded,backtrader. You run the statistics yourself, free, with full control. Two structural limits: the grader answers to the person being graded, and no library sees how many variants you tried before this one — the multiple-testing debt that deflation exists to charge. The debt is not small: at Sharpe 1.0, 45 tried variants already demand five years of daily data before the best one means anything. - GIPS verification (ACA-class firms). Verifies that an asset manager's performance reporting complies with the GIPS standards. It audits presentation, not edge — a fully compliant report of a lucky backtest is still a lucky backtest.
- Prop-firm evaluations (FTMO/Topstep-class). Forward gates on live drawdown that decide whether you get funded. They test behaviour going forward and tell you nothing about whether a historical backtest was ever real — and their published pass rates carry heavy survivorship.
- Tournaments and marketplaces (Numerai, QuantConnect Alpha Streams). Capital allocation through live ranking against a crowd. Valuable if allocation is the goal; the ranking is platform-bound and forward-only, so it cannot audit the claim you already have.
- Provenance and attestation services. Timestamp pinning proves when a claim existed — indispensable against backfilled track records, and our falsification protocol demands it. It does not test whether the claim was edge or noise.
- An assay office. Falsification as a service: independent of the submitter, cumulative trial accounting across a whole idea family, demote-only by construction. That combination — independence, family-level multiple-testing memory, and the inability to bless — is the part you cannot self-host.
The short version: if you want funding, go to a prop desk; if you want provenance, pin your timestamps; if you want to know whether the edge was ever real, put it on trial.
What does the trial actually test?
Four families of gates, in a fixed order, each grounded in published statistics rather than house opinion. Costs come first, because most apparent edges are artifacts of frictionless simulation. Selection is charged next with the Deflated Sharpe Ratio (Bailey & López de Prado, 2014) under cumulative family accounting, because „the same signal with lookback 21 instead of 20" is not a fresh discovery. Skill is then separated from timing luck by racing the signal against 500 matched random twins with the same trading profile. What survives is attacked: execution delay, cost stress, history jackknife, regime splits, parameter neighbourhoods and more — eight attacks in all. The base rates justify the harshness: the published record says most apparent edges are selection noise (Harvey, Liu & Zhu 2016; Bailey & López de Prado 2014), and our own public benchmark agrees — of 126,000 classic rule/asset pairs, 245 survive the battery and 41 beat buy-and-hold. A fail with a named cause is the product working. The anonymised mortality record is public, live and signed: the graveyard digest.
Why call it an assay office?
In metallurgy, an assay office is the institution that tests precious metal and stamps its verified fineness into the piece. It does not mint coins, does not buy gold and does not promise prices — it grades material and stakes its name on the grade. That is the entire model here, transferred to strategy evidence: a fixed public battery, verdicts that can only demote, separately signed certificates for what survives, and a dated public log of the examiner getting stricter. Start with the free unsigned demo, lint a payload with the free preflight, or look up whether your idea's family is already in the graveyard — before it costs you weekends.