What matters more in crypto: the market or the strategy?
Six markets, six strategies, fourteen out-of-sample quarters, and one pre-registered contest between the two ways researchers usually start.
Most crypto research starts from one of two moves. You find a market that looks like it has something in it, then you throw several rules at it. Or you find a rule you believe in, then you spread it across whatever markets will take it. Both moves feel reasonable, and they make very different bets about where returns come from: one says the venue carries the edge, the other says the logic does.
Almost nobody tests which move pays better, because the two are rarely put on the same footing. Market selection usually gets tested with a favourite strategy attached, and strategy selection usually gets tested on a favourite market, so whichever axis the researcher already believes in ends up with the better hand.
We ran the contest with both arms built the same way. Six structural strategies and six liquid Binance USDT perpetual markets form a balanced grid; one picker gets to choose the market and must then hold every strategy on it, the other gets to choose the strategy and must then hold it on every market. Same capital, same held-forward in-sample evidence, same net P&L objective, same fourteen out-of-sample quarters.
The universe rules, the strategy definitions, the walk-forward boundaries, the cost model and the frozen hashes live on their own page, so they don't crowd the argument here.
What we tested
Six markets, six strategies, one grid. The markets are BTC, ETH, DOGE, XRP, BNB and SOL perpetual futures on Binance, ranked by traded notional and fixed before the test period opened. The strategies are the ordinary furniture of systematic crypto: EMA trend, Donchian breakout, time-series momentum, RSI reversal, Bollinger fade and a volatility-compression breakout. All 36 strategy-market cells trade hourly bars and every cell runs in all three portfolios, so no arm gets access to an instrument or a rule the other one lacks.
Numeric parameters are not strategies. Lookbacks, thresholds, stops, targets and maximum holding times are knobs refitted inside each walk-forward window, exactly as a desk would refit them, and each family gets the same fixed search budget in every market and every window. A family that happens to have more tunable surface than another cannot win by searching harder.
Each quarter of out-of-sample history is preceded by a year of in-sample history, split in two, with the first nine months fitting each cell’s numeric parameters. The last three months then replay the already-selected configuration untouched, and those held-forward scores, not the fitting scores, are what the pickers see when they choose.
That split is the reason the contest is fair, because a picker choosing on the same P&L that selected the parameters would be reading its own optimizer's output back to itself.
Three portfolios come out of every window and all three deploy the same fixed capital. The no-selection portfolio holds all 36 cells at equal weight and makes no choice at all. The market picker takes the market with the best held-forward mean across its six strategies and holds all six strategies on it. The strategy picker mirrors that, taking the best strategy across the six markets and holding it everywhere.
Capital does not compound and flat sleeves stay in cash, so an arm never quietly concentrates into whatever is left running. Costs are charged the same way everywhere: every fill is a taker fill paying 5 basis points of fee and 2 basis points of slippage, and funding is applied at its real Binance settlement events with the real sign. Stops and targets inside an hourly candle are resolved from one-minute data rather than assumed, since the order in which a high and a low arrived is the difference between a trade that stopped out and one that reached its target.
The primary comparison was registered before any out-of-sample number existed: cumulative market-picker net P&L minus cumulative strategy-picker net P&L, with a paired confidence interval built by resampling calendar weeks inside each fixed quarter. The ruling was fixed in advance too, where an interval entirely above zero means market selection won, entirely below means strategy selection won and an interval containing zero means the experiment produced no reliable winner.
Results
Neither axis won. Market selection finished 193.9 basis points of fixed capital ahead of strategy selection over the fourteen quarters, with a paired 95% confidence interval running from -7,868.3 to 8,194.8 basis points. The point estimate is about 1% of the width of its own interval, which is another way of saying the experiment cannot tell the two apart.
| Arm | Cumulative OOS net return | Versus no selection | BH-corrected p |
|---|---|---|---|
| Market picker | -45.22% | +13.20 pp | 0.7513 |
| Strategy picker | -47.16% | +11.26 pp | 0.7513 |
| No axis selection | -58.42% |
Cumulative out-of-sample net return on fixed capital, fourteen quarters. Capital does not compound, so these are arithmetic sums of sleeve returns rather than account drawdowns.
Every arm lost money. Both pickers lost less than the portfolio that made no choice, by 13.20 and 11.26 percentage points, and neither improvement survived the registered two-test Benjamini-Hochberg correction.
The quarterly record explains where the width of that interval comes from. Market selection was ahead in six windows, strategy selection in eight and the lead changed hands four times across the fourteen. The largest single-quarter gap in favour of market selection is 1,852.7 basis points; the largest in favour of strategy selection is 1,386.9.
Two pre-registered sensitivities moved the point estimate across zero without narrowing anything. Fitting parameters from each contract's launch instead of a rolling nine months flips the sign to strategy selection at -1,596.7 basis points, while replacing real funding with a constant adverse 1 basis point per event pushes it further toward market selection at 1,359.1. All three intervals span roughly 14,000 basis points of ground on either side of zero.
The grid underneath
Seven of the 36 cells made money over the full period. The rest lost, and the losses are large enough that every market mean and every strategy mean is negative.
Read the margins and the picker result stops being surprising. The spread across market means is 5,151 basis points, from BNB at -3,771 to SOL at -8,922, while the spread across strategy means is wider at 8,209 basis points, from time-series momentum at -3,123 to RSI reversal at -11,332. Both axes separate, yet neither separates the same way twice in a row, which is the one thing a picker needs.
Where the dispersion lived
A fixed-effects decomposition of the 36-cell grid across all fourteen windows splits the out-of-sample sums of squares into a persistent market effect, a persistent strategy effect, their interaction, a common time effect and everything left over.
Interaction is the largest of the three cross-sectional components, and the simultaneous bootstrap intervals put it above both main effects: market minus interaction runs from -11.45 to -1.61 percentage points and strategy minus interaction from -9.96 to -0.12, both entirely below zero, while market minus strategy runs from -6.40 to +3.44 and settles nothing.
In plain terms, the specific pairing of a rule with a market mattered more than either the rule or the market on its own. RSI reversal on DOGE lost 25,027 basis points while time-series momentum on the same market made 5,533, and neither the market nor the strategy decided that.
That finding does not rescue the pickers, because the pairing is not what they were allowed to choose, and because the residual is more than twelve times the interaction share. With one observation per cell-window, that residual holds time-varying market effects, time-varying strategy effects and everything idiosyncratic, none of it stable enough to select on from three months of held-forward evidence.
What the costs did
The grid was profitable before costs. Summed over all 36 cells and fourteen quarters, gross P&L came to 27,454.6 basis points of sleeve capital, then fees took 168,288.4, slippage took 67,315.4 and funding took 2,147.0, which turns +27,454.6 into -210,296.2.
Per trade, that is 1.63 basis points of gross against 14.13 basis points of cost across 16,826 out-of-sample trades: seven basis points a fill, two fills a trade and a median holding time of 11 hours. No selection rule fixes a ratio like that, which is the honest reason all three arms sit deep in the red. The axis contest was run on a population of hourly rules that trade far too often for their edge.
Where it looks fragile
Leaving one market out at a time is where the result strains hardest. Five of the six omissions leave market selection ahead, and dropping BNB reverses the difference to -4,494.8 basis points. A six-market comparison that changes sign when one market leaves is a comparison with one market inside it, not a general property of the venue.
Clone handling changed nothing at all, since signed daily position correlations formed six singleton groups in every window, so no two families were close enough to collapse into one. Transfer told the same story from another direction: replaying every window's frozen parameters across all six markets, twelve of fourteen windows show no positive aggregate cross-market edge under either total P&L or expectancy, one window passes the market-dependence check and one is held up by a single market.
The checks
The final corpus holds 31,770,479 candidate trade rows and 16,826 selected out-of-sample trades. Ledger costs reconcile to within 5.55e-17 return units, the Python reference and the Rust production engine agreed on 25,259 trades across 72 real-data parity cases with a maximum absolute P&L disagreement of 1.94e-16, and a 30-seed zero-drift random-walk corpus produced no family whose selected out-of-sample lower confidence bound cleared zero.
Two automated warnings were resolved by checking the data directly rather than by argument. The generic check read the shared in-sample and out-of-sample timestamp as a zero-day purge, when this engine purges at trade level, and a direct boundary count found zero trades crossing or falling outside their registered window. The Sharpe-oriented stage of the same check reported low power and a noise-scale spread across families, which neither selects nor decides anything here but is the reason no individual family in the grid is presented as an edge.
Deciding whether to search markets or search strategies allocates effort to an axis that carried under 2.5% of the dispersion between them, while the pairing carried 6.98% and time carried the rest.
For these six markets, these six strategies and these fourteen quarters, picking the market first did not reliably beat picking the strategy first, and both beat picking nothing by an amount too small to distinguish from noise. The persistent main effects on both axes were tiny, the pairing effect was several times larger than either, and everything cross-sectional together accounted for barely a tenth of the variation.
For this class of hourly perpetual system the argument was settled before either picker ran, by a cost line 8.7 times the size of the gross edge.
Limitations
- The result covers six Binance perpetual markets, six structural families and fourteen quarters of hourly data. It is not a statement about crypto in general, other venues or other holding periods.
- There is no untouched external market panel, so transfer beyond these six markets is unproven.
- The comparison changes sign when BNB is removed, which makes the six-market estimate dependent on one member of its own universe.
- Leave-one-market-out arms hold different sleeve counts after a market is dropped, so those estimates are descriptive dependence checks rather than replacement headline numbers.
- Every arm lost money, so the contest is between two ways of losing less. A profitable population of strategies might rank the two axes differently.
- The decomposition reports descriptive sums-of-squares shares on a deliberately selected grid, not population variance components.
Every chart on this page is drawn in your browser from the study's published result tables, the same numbers reported in the text and on the reproducibility page. There are no image files behind them.
Universe rules, the six strategy definitions, execution and cost model, the walk-forward loop, inference, the decomposition and all nineteen pre-registration gates, with hashes and code snippets.

