Catch Alpha:
Testing the Evaluator Before Testing the Trading Strategy
Independent Researcher, France
ICAIF '26 · Poster · Milan, 16–17 November 2026

Abstract
A trading result is only as trustworthy as its evaluation. We first test our procedure on synthetic data with known truth. In this design, 0/200 zero-edge strategies pass, while detection rises with the planted edge; these finite tests do not cover every market. We then audit learning in a controlled market. Starting every policy flat, behavior cloning and implicit Q-learning recover 90% and 73% of a risk-adjusted analytic reference without costs. With quadratic costs, all ten fits for both methods lose money; median turnover is 6.03× for cloning and 11.97× for implicit Q-learning relative to the reference. Conservative Q-learning fails the no-cost prerequisite, so its costly result is not adjudicated. On twelve liquid cryptocurrency futures, neither the tested long-loser/short-winner reversal nor funding-carry implementations produce reliable net profit. The contribution is a reproducible workflow that measures power, counts attempts, and limits final-period access.
TL;DR: Anyone who tries enough trading ideas will find one that looks great on past data, even if none of them works. We built a judge that counts every attempt and refuses to rule when the data are too short, and we tested the judge first: on simulated markets with no real edge it almost never accepted one, and it detected planted edges once they were large enough. It then judged four real strategies on twelve cryptocurrencies: the data show predictive patterns, but no tested portfolio established a reliable net profit after fees under the protocol.
Can the learner use a known signal?

Real-data study


Conclusion
We test the evaluator before trusting the backtest. In finite synthetic tests, none of 200 Gaussian zero-edge strategies passes, and detection rises with a planted edge. Other tested null generators also give zero acceptances; untested market processes remain outside this calibration.
In the controlled market, BC and IQL learn the frictionless signal, but all 10/10 fits lose under quadratic costs; their median turnovers are 6.03× and 11.97× the GP reference. CQL fails the no-cost prerequisite, so its costly case is not adjudicated. These failures are configuration-specific, not a general statement about RL.
On twelve liquid crypto perpetuals, no tested reversal or carry implementation produces a validated net-of-cost trade. Funding is positive in every fold, but the portfolios do not retain it after price risk and retail taker fees. Two on-chain relations remain provisional. The single planned final-period evaluation does not contradict the result, but is too short for stronger claims about small effects.
The reusable contribution is the workflow: minimum sample size, a count of every attempted variant, declared baselines, acceptance checks calibrated on known nulls and planted edges, and rationed final-period access.
Poster

Citation (BibTeX)
@inproceedings{viaud2026catchalpha,
title = {Catch Alpha: Testing the Evaluator Before Testing the Trading Strategy},
author = {Viaud, Joris},
booktitle = {7th ACM International Conference on AI in Finance (ICAIF '26)},
address = {Milan, Italy},
year = {2026}
}