← Alpha
Alpha Research · · 13 min read

Do trading strategies need an economic rationale?

Sixty Binance perpetual contracts, two arms of strategies matched on how good they looked in sample, and one question: does writing down the reason first change what survives?

Practitioners repeat a claim often enough that it has become folklore: a strategy you can explain will hold up out of sample, and one you found by searching will not. The claim is testable, and as far as we can tell it has never been tested on a matched corpus.

It is worth testing because it is load-bearing, deciding whether a researcher spends the next six months thinking about who is on the other side of a trade or building search infrastructure. Most desks have picked a side already, and usually without evidence.

One obstacle explains why the test is rare. The comparison is almost always made unfairly, because a strategy with a written mechanism and one found by enumeration rarely get the same search budget, the same costs, the same instruments or the same period. Whichever side wins, the reader cannot tell whether the reason mattered or the setup did.

This study holds all of that fixed and varies one thing: whether an economic mechanism was written down before the parameter search began. Sixty Binance perpetual contracts, seventeen strategy families split across the two arms that matter, an identical search budget on both sides and eighteen quarters of out-of-sample history, all of it on measured spreads, real exchange fees and observed funding.

Reproducibility

The universe rule, the family definitions, the measured-spread cost model, the matching procedure and the inference live on their own page, so they don't crowd the argument here.

Read the reproducibility →

What we tested

Two arms of strategies run through one engine, on one panel of Binance USDT-margined perpetual futures at 30-minute bars.

Catalog, arm C. Eight named technical indicators: ATR, EMA, SMA, MACD, PPO, RSI, a static RSI level and stochastic %K, each compared against its own reference. No claim is made about why any of them should work, because enumerating the catalog is itself the method.

Mechanism-first, arm M. Nine families whose entry rule was specified to express a written economic mechanism, one that names who is on the other side of the trade and why that counterparty transacts for reasons other than expected profit. Volatility-compression breakout, time-series momentum, session momentum, weekend reversal, volatility climax, relative strength against BTC, gap fade, range reversal and Keltner trend.

Both arms are price-only, so they see the same information. Both draw their variants from one shared library of 35 modifiers, 11 transforms and 24 confluence gates, copied verbatim from the generator that produced the two source corpora and hashed, so the claim that the arms share a vocabulary is checkable rather than asserted. Each family receives an identical budget of 945 attempted configurations per contract and per window, which leaves the base entry rule as the only thing separating the arms.

A third arm of eleven mechanism families that also read funding, open interest, basis, order flow or positioning is carried as a registered secondary. It exists so that "has a reason" is not silently confounded with "has more data".

Why the estimand is a difference

Both arms were expected to lose money net of costs, and both did. That is not a problem, because the question is not which arm is profitable.

Strategies are matched inside each contract and window on in-sample Sharpe, within a caliper of 0.05 annualised Sharpe units, greedily and without replacement. Two strategies that looked equally good in sample are paired against each other, and the estimand is the difference in what each kept out of sample, which makes the result a measurement of differential decay from a common starting point rather than a comparison of performance.

Matching only works where the arms overlap, so the overlap is a precondition rather than a result. It is also the reason no widening of the caliper was needed: 6,812,171 pairs matched at 0.05.

Fig. 1: In-sample Sharpe of every eligible strategy-window, by arm. Descriptive: this is the full score table the matcher reads from, not the matched set. The arms overlap across the whole range the caliper has to work in.

Universe and costs

Eligible contracts are USDT-quoted, traded on every day of the October-to-December ranking quarter and neither index products nor redenomination duplicates. 141 contracts qualify, and the universe is the top 30 and the bottom 30 by median daily notional, which separate by a ratio of 20.0x against a registered floor of 10.

Contracts are drawn from the bulk archive listing rather than the exchange's current contract list. exchangeInfo returns only currently-listed contracts, so a universe built from it silently deletes everything that died. Eighteen non-overlapping quarterly windows run from 1 January 2022 to 30 June 2026, each preceded by a 12-month in-sample block.

Costs are five basis points of taker fee per fill, spread measured per contract-month from aggTrades as the median inter-trade bounce, with each fill paying half the measured spread plus half again as a slippage allowance. Funding is the observed signed event at its true settlement bar, valued at the published mark price, so longs are debited when funding is positive.

Measuring the spread rather than assuming it is the single thing that keeps the liquidity result honest. A flat 3 bp slippage constant would have overcharged BTCUSDT by more than a hundredfold while roughly matching the thinnest contracts, which biases in exactly the direction of the liquidity result and would have manufactured it.

Fig. 2: Measured median spread per contract, log scale, coloured by the stratum each contract was assigned to at the selection date. Hover any bar for the contract. The two dashed lines are the stratum medians: 1.597 bp against 3.738 bp, a ratio of 2.3x.

Results

The registered primary contrast is the matched difference in out-of-sample net return, mechanism-first minus catalog, pooled across both liquidity strata.

D = 223.4 bp, with a paired week-clustered 95% confidence interval of [98.5, 344.8] basis points of fixed capital, p = 0.0018. The interval lies entirely above zero.

That estimate rests on 6,812,171 matched pairs across 250 week-fragment clusters and 1,642 trading days, matched at a caliper of 0.05 annualised Sharpe units with no widening required.

As a consistency check, the same quantity computed from the score table's per-window totals is 213.1 bp against 223.4 bp from the per-trade ledger, a relative gap of 4.6%. The two are independent routes to the same number and are not identical by construction, because the ledger excludes trades straddling a window boundary.

ContrastEstimate95% intervalpBH-adjusted p
Liquidity interaction, D_H minus D_L-114.9 bp[-197.0, -16.4]0.01380.0184
High-liquidity stratum, D_H171.2 bp[32.8, 295.6]0.0140.028
Low-liquidity stratum, D_L278.4 bp[131.4, 410.5]0.00040.0004
Information set, exogenous arm minus catalog139.2 bp[-82.1, 361.2]0.21080.8431

The four registered secondary contrasts share one Benjamini-Hochberg family at 5%. Three of the four reject after correction.

Fig. 3: The registered primary contrast and the four secondaries, with their paired week-clustered 95% intervals. The information-set contrast is the one that spans zero.

Survival rate, the share of matched strategies with positive out-of-sample net return, runs 29.74% in the catalog arm against 30.93% in the mechanism-first arm, a difference of 1.19 percentage points.

The registered ruling is that a written economic mechanism predicts greater out-of-sample retention. That is what the primary contrast says. Four qualifications bound how far it carries, and the first one is large.

+223.4
Mechanism minus catalog, bp of fixed capital
[98.5, 344.8]
Paired 95% interval, bp
6.81M
Matched pairs at a 0.05 caliper
33%
Of the effect reproduced on a driftless corpus
7.5e-05
Difference in decay slope
2.3x
Measured spread, thin stratum against liquid

A third of the effect belongs to the rule vocabulary

The same pipeline was run over a driftless corpus, where by construction no strategy can have an edge and any difference between the arms is the self-reference artifact described in the limitations below. That corpus produced a matched difference of 72.92 bp with an interval of [-15.11, 166.60].

Set against the observed 223.39 bp [98.52, 344.82], the artifact's point estimate accounts for 33% of the observed effect. Its upper bound of 166.60 bp sits above the observed effect's lower bound of 98.52 bp, so the two intervals overlap across the range 98.52 to 166.60.

Fig. 4: The observed effect against the same estimand measured on a driftless corpus, where nothing can have an edge. Drawn on one axis because the overlap between the two intervals is the point.

Stated honestly, the claim is narrower than the ruling. A matched difference of roughly 220 basis points was observed, of which something like a third is attributable to a property of the shared rule vocabulary rather than to the presence of an economic mechanism, and the separation between the two is not clean. The artifact-adjusted effect is around 150 bp, but that subtraction has no joint interval behind it and should be read as an order of magnitude rather than an estimate.

It is a level effect, not a slower decay

Regressing out-of-sample outcome on in-sample Sharpe over the matched set gives a slope of 0.0549 [0.0465, 0.0634] for the catalog arm and 0.0550 [0.0481, 0.0621] for the mechanism-first arm. The difference is 0.000075 against individual intervals roughly 0.015 wide, so the rate at which out-of-sample performance tracks in-sample performance is indistinguishable between the arms.

Fig. 5: How steeply out-of-sample outcome tracks in-sample Sharpe, per arm, with window-clustered intervals. The two sit on top of each other.

This matters for how the folklore should be restated. Mechanism-first strategies did not decay more slowly, and matched at equal in-sample Sharpe they simply delivered more out of sample at the same slope, so whatever the mechanism buys, it is not resistance to decay.

The effect is larger where liquidity is thinner

Interaction between the strata runs -114.94 bp [-197.04, -16.42], surviving Benjamini-Hochberg at an adjusted p of 0.0184. The low-liquidity stratum shows 278.45 bp against 171.24 bp in the high-liquidity stratum, so the advantage is roughly 60% larger among thinner contracts.

Read that with care. Stratum L here is low-mid-cap rather than the capacity-constrained tail, spreads there are 2.3 times wider and measured rather than assumed, and the artifact caveat above applies to both strata.

Exogenous data adds nothing detectable

The families that read funding, open interest, basis, order flow and positioning on top of a written mechanism return 139.15 bp [-82.10, 361.24], an interval spanning zero with a Benjamini-Hochberg adjusted p of 0.843. On this corpus, adding non-price information to a mechanism-first rule produced no measurable improvement over the catalog baseline.

Defects found before publication

Six defects were found and fixed during construction. Each is recorded because each would have produced a plausible number, and one of them did: the first completed analysis reported a positive finding in the direction this study hypothesises, and it was an artifact of corrupted inputs.

Funding was being charged at settlement bars that were themselves missing from the panel, where both the published mark price and the bar open are absent, and the charge became non-finite. On CVXUSDT that condition occurs on exactly one bar out of 67,608, and it produced 1,092 corrupted rows. Across the corpus, 1,238 out-of-sample returns of 25,251,675 were non-finite, which is enough to make every point estimate return nothing.

The more serious half is the guard. The ledger writer refuses to persist a ledger whose costs do not reconcile, by testing whether the residual exceeds a tolerance. A non-finite residual fails that comparison, so every poisoned row passed the check built specifically to catch bad costs, and 1.68 billion trades were certified as reconciling. Tested afterwards on a planted non-finite row, the old check reports a residual of 1.7e-18 and concludes the ledger reconciles cleanly, which is worse than missing the corruption: it produced positive evidence of correctness out of it.

What caught it was a consistency cross-check added earlier for unrelated reasons, comparing the matched difference computed from the score table against the same quantity computed from the per-trade ledger, which printed a divergence and made the corruption visible. That cross-check is the same one reported above as 213.1 against 223.4 bp.

Three other defects are worth naming because they are easy to reproduce elsewhere. Target exits were filling at the exit bar's open rather than at the target price, worth up to 12.5 basis points per trade on a driftless path and larger than the entire cost model, and it acted only on winning trades. Strategy ids restarted at zero for every contract-family, so joining matched pairs back to their returns matched ambiguously and turned 119,178 pairs into 8,580,816 rows. The bootstrap was clustering on window start dates, which collapsed the registered Monday-to-Sunday week fragments into one cluster per window, 18 instead of the 250 the realised data carries.

Matched at equal in-sample Sharpe, a written mechanism was worth roughly 220 basis points out of sample, and about a third of that was worth nothing at all: the same vocabulary produces it on data with no market in it.

Limitations

  • The mechanism statements are reconstructions. The mechanism families were authored by the same researcher after the catalog corpus was already known to fail, and their written mechanisms were reconstructed from the original design record rather than sealed contemporaneously. Arm M is mechanism-first with respect to its own parameter search, not blind with respect to the catalog result. Blind grading of the statements limits this without eliminating it.
  • A self-reference artifact is present in the rule vocabulary. A signal derived from the traded series, combined with barriers scaled by that same series' lagged volatility, does not earn zero on a driftless corpus, where both arms lose roughly 550 basis points from the artifact alone. It is a property of the vocabulary rather than a defect in the engine, so it does not vanish by fixing code, and it is the reason a third of the headline is not attributable to the mechanism.
  • The first four windows precede the selection date. The universe was ranked on the last quarter of 2022 while the claim period opens on 1 January 2022, so four of eighteen windows trade a panel chosen with knowledge of that year. Both arms trade the same contracts in the same windows, so the matched difference is not obviously biased by it, but the panel for those windows is not point-in-time.
  • One contract's out-of-sample medians were seen before the freeze. During a smoke test of the runner, a summary printed median out-of-sample net return by arm for BTCUSDT alone. Those numbers were produced before the target-fill defect was found and are therefore artifacts, but the record is that they were seen. The analysis now refuses to run outside a guarded mode.
  • The low-liquidity stratum is not the capacity-constrained tail. The thinnest contract in the frozen universe still traded $5.0M a day over the ranking quarter. Everything genuinely thinner listed after the selection date, so "low liquidity" here means low-mid-cap and no claim about structurally excluded institutional capital rests on it.
  • Resting depth would be a better liquidity instrument. daily/bookDepth gives resting notional at fixed bands from mid across full history, which is more direct than a trade-derived spread for a study stratified on liquidity. It was not adopted, because the cost source was already registered and validated and an out-of-sample value had been exposed by then.

Every chart on this page is drawn in your browser from the study's published result tables, the same numbers reported in the text and on the reproducibility page. There are no image files behind them.

Full methodology

Data sources, the universe rule, the shared variant library, execution and the measured-spread cost model, the walk-forward loop, matching, inference and the twenty blocking gates, with code snippets.

Reproducibility →