← Papers

Catch Alpha:

Testing the Evaluator Before Testing the Trading Strategy

Joris Viaud

Independent Researcher, France

ICAIF '26 · Poster · Milan, 16–17 November 2026

Finite calibration on known truth. Left: none of 200 Gaussian zero-edge strategies passes (95% upper bound 1.9%). Right: detection rises from 0% at planted Sharpe 0.5–1.0 to 100% at 3.0, with 40 draws per planted effect; error bars are 95% Wilson intervals. Additional tested null generators also produce zero acceptances. All rates and bounds are conditional on the tested generators and sample sizes; they do not imply a universal false-positive rate.
Finite calibration on known truth. Left: none of 200 Gaussian zero-edge strategies passes (95% upper bound 1.9%). Right: detection rises from 0% at planted Sharpe 0.5–1.0 to 100% at 3.0, with 40 draws per planted effect; error bars are 95% Wilson intervals. Additional tested null generators also produce zero acceptances. All rates and bounds are conditional on the tested generators and sample sizes; they do not imply a universal false-positive rate.

Abstract

A trading result is only as trustworthy as its evaluation. We first test our procedure on synthetic data with known truth. In this design, 0/200 zero-edge strategies pass, while detection rises with the planted edge; these finite tests do not cover every market. We then audit learning in a controlled market. Starting every policy flat, behavior cloning and implicit Q-learning recover 90% and 73% of a risk-adjusted analytic reference without costs. With quadratic costs, all ten fits for both methods lose money; median turnover is 6.03× for cloning and 11.97× for implicit Q-learning relative to the reference. Conservative Q-learning fails the no-cost prerequisite, so its costly result is not adjudicated. On twelve liquid cryptocurrency futures, neither the tested long-loser/short-winner reversal nor funding-carry implementations produce reliable net profit. The contribution is a reproducible workflow that measures power, counts attempts, and limits final-period access.

TL;DR: Anyone who tries enough trading ideas will find one that looks great on past data, even if none of them works. We built a judge that counts every attempt and refuses to rule when the data are too short, and we tested the judge first: on simulated markets with no real edge it almost never accepted one, and it detected planted edges once they were large enough. It then judged four real strategies on twelve cryptocurrencies: the data show predictive patterns, but no tested portfolio established a reliable net profit after fees under the protocol.

Can the learner use a known signal?

Corrected learning test
Corrected learning test at IC = 0.05; every policy starts flat. Left: median score relative to GP without trading costs; the dotted line is the pre-specified 0.50 minimum. BC and IQL pass, while CQL does not. Right: median mean profit after quadratic costs. BC and IQL are negative for all ten fitted runs; their median turnovers are 6.03× and 11.97× that of GP, respectively. CQL is shown in gray because its failed no-cost check prevents a conclusion under costs. Adding the previous position to BC does not rescue it.

Real-data study

Cross-sectional signal screen
Cross-sectional signal screen on training and validation data. Points show average rank correlation with future returns. Bars are approximate intervals: for the price panel, mean fold correlation ±1.96·|mean|/|median fold t|, a heuristic rather than a 95% confidence interval; for the external panel, normal intervals that combine fold-level Newey–West standard errors treating folds as independent. Labels give the number of folds whose Newey–West interval excludes zero, out of six. The on-chain panel precedes market-cap residualization. Left: short-horizon momentum is negative (reversal). Right: the largest stable non-price relations are slow on-chain size measures. Open points are unstable in more than two folds.
Cash-and-carry fee sensitivity
Cash-and-carry net return over the training-and-validation window versus fee per leg (dotted: funding received; band: the frozen five-basis-point assumption). Labels: folds of six with a positive net-return lower bound.

Conclusion

We test the evaluator before trusting the backtest. In finite synthetic tests, none of 200 Gaussian zero-edge strategies passes, and detection rises with a planted edge. Other tested null generators also give zero acceptances; untested market processes remain outside this calibration.

In the controlled market, BC and IQL learn the frictionless signal, but all 10/10 fits lose under quadratic costs; their median turnovers are 6.03× and 11.97× the GP reference. CQL fails the no-cost prerequisite, so its costly case is not adjudicated. These failures are configuration-specific, not a general statement about RL.

On twelve liquid crypto perpetuals, no tested reversal or carry implementation produces a validated net-of-cost trade. Funding is positive in every fold, but the portfolios do not retain it after price risk and retail taker fees. Two on-chain relations remain provisional. The single planned final-period evaluation does not contradict the result, but is too short for stronger claims about small effects.

The reusable contribution is the workflow: minimum sample size, a count of every attempted variant, declared baselines, acceptance checks calibrated on known nulls and planted edges, and rationed final-period access.

Poster

ICAIF '26 poster
The ICAIF '26 poster (A0 portrait). The PDF is linked from the Poster button above.

Citation (BibTeX)

@inproceedings{viaud2026catchalpha,
  title     = {Catch Alpha: Testing the Evaluator Before Testing the Trading Strategy},
  author    = {Viaud, Joris},
  booktitle = {7th ACM International Conference on AI in Finance (ICAIF '26)},
  address   = {Milan, Italy},
  year      = {2026}
}