← Alpha
Alpha Research · · 12 min read

Does a wider, less liquid universe improve a systematic book?

Thirty Binance perpetuals instead of six, reaching down to contracts trading a two-hundredth of Bitcoin's daily notional. Everything else held fixed.

Most systematic books trade a handful of instruments. The reason is usually practical rather than considered: the data was easy to get, the contracts are liquid, and nobody ever went back to ask what the universe was costing. Two arguments push the other way. The first is diversification, that a book of thirty weakly correlated sleeves should carry a better Sharpe than a book of six. The second is crowding, that the majors are picked over by everyone with a terminal while the thin end of the exchange is not, so whatever edge remains should be sitting down there.

Both are testable, and they are different claims. One is about the arithmetic of averaging. The other is about where signal lives. A study that widens a universe and reports only the portfolio Sharpe cannot tell you which of the two moved.

So we took a set of classical trading rules that already loses money on the six most liquid Binance perpetual futures, gave it twenty-four more contracts, and changed nothing else. Same six strategy families, same frozen parameter grids, same walk-forward schedule, same costs, same fourteen out-of-sample quarters.

Reproducibility

The exact recipe behind every number here, including the death rules, the cost ladder and every leak check, lives on its own page.

Read the reproducibility →

What we tested

The universe is ranks 1 to 30 by median daily notional traded during October to December 2022, drawn from the Binance USDT-margined perpetuals with at least two years of genuine traded-volume history by the end of 2022. Ranks 1 to 6 are exactly the six markets of the earlier study this one extends, which is the first evidence that the ranking rule was reimplemented faithfully. The rank axis spans a factor of 205 in traded notional, from Bitcoin at the top to contracts turning over about $34 million a day at the bottom.

Six structural strategy families run on every market: EMA trend, Donchian breakout, time-series momentum, RSI reversal, Bollinger-band fade and volatility-compression breakout. Each family carries a frozen list of sixty-four numeric configurations, generated once and reused in every market and every window, so no market ever gets a wider search than another. Thirty markets by six families gives 180 sleeves, each with its own walk-forward loop, and every sleeve carries the same fixed capital whether it is trading or flat.

The schedule is the inherited one: twelve months in sample, three months out, advancing a quarter at a time, fourteen complete out-of-sample quarters covering January 2023 through June 2026. Selection inside each window reads the in-sample block only and never touches a Sharpe. A sleeve that cannot field an eligible configuration in a window holds cash for that window rather than being dropped, which keeps thin contracts from quietly filtering themselves out of the sample.

Two partitions of the same frozen ranking do the work. Four tranches, ranks 1 to 6, 7 to 14, 15 to 22 and 23 to 30, ask whether per-sleeve economics vary with liquidity. Six nested rungs, the top 6, 10, 15, 20, 25 and 30 markets, ask what breadth alone does to the book. Neither partition was chosen after seeing anything.

The primary measurement is the per-sleeve out-of-sample net expectancy of the twenty-four added markets minus the six incumbents, in basis points per live sleeve-day, with intervals from a paired week-clustered bootstrap. A live sleeve-day is a day on which the contract was actually trading, which matters because four of the added contracts were force-settled by Binance inside the out-of-sample period. Those sleeves hold cash from settlement onward, capital is never reallocated onto the survivors, and dead days never enter the denominator, so a contract death cannot flatter the added group by diluting it toward zero.

Costs are charged on every fill: five basis points of taker fee and two of slippage, on every market from rank 1 to rank 30. That flat treatment is deliberately optimistic at the thin end, and the specification says so, which is why it also registers a mandatory adverse model in advance that charges 2, 4, 6 and 8 basis points of slippage by tranche and re-selects every configuration under those costs.

The breadth question was registered as a model rather than as a comparison. Portfolio Sharpe for a book of equally weighted sleeves should be the mean per-sleeve Sharpe times the square root of N / (1 + (N-1) rho), where rho is the realized average pairwise correlation. Both the prediction and the realized value are computed inside the same bootstrap replicates, so the difference between them carries an interval and the ladder can be judged against its own theory rather than merely described.

Results

Under the primary cost model there is no measurable liquidity-rank effect. Per-sleeve net expectancy of the added twenty-four minus the incumbent six is -1.046 basis points per live sleeve-day, with a 95% interval from -2.996 to +0.934, and both groups lose, the incumbents at -4.574 and the added markets at -5.620.

The registered adverse cost model puts the interval entirely below zero, at -3.471 with an interval of -5.444 to -1.515, and the added markets fall to -8.046. That is the second of the three registered verdicts: widening the universe downward in liquidity makes the book worse. Which reading you believe depends entirely on what you think it costs to trade a contract turning over $34 million a day.

Fig. 1: Added twenty-four minus incumbent six, per-sleeve net expectancy, under each registered cost model. Two intervals span zero and the mandatory rank-scaled model's does not. Hover a row for its interval.

The decomposition, not the headline

Gross profit and loss, fee, slippage and funding are separate columns in the trade ledger and they reconcile to net exactly, so the difference between the two groups splits without residual.

Fig. 2: Where the difference sits. Under the primary model the gross term is essentially zero and the shortfall is turnover; under the rank-scaled model the fee difference collapses and almost all of it moves into slippage.

At -0.009, the gross term under the primary model says something specific. On a per-sleeve basis these rules extract essentially the same raw edge from a contract ranked thirtieth by liquidity as from Bitcoin. Every basis point of the shortfall is the cost of harvesting it. Slippage is flat at two basis points per fill on every market in that model, so a bigger slippage bill cannot mean worse fills, only more of them.

Under the adverse model the fee difference collapses to -0.006, because re-selection cuts turnover in the thin markets almost exactly enough to equalise fees, and the difference moves into slippage at -2.949. Gross falls to -0.501 because the lower-turnover configurations the optimiser now prefers capture less raw edge.

Sorted by liquidity, gross rises and cost rises faster

Fig. 3: Per-sleeve economics by liquidity tranche, primary cost model. The thinnest eight contracts carry the highest gross expectancy and the highest cost, and every tranche finishes net negative.

The thinnest eight contracts show four times the gross per-sleeve expectancy of the six most liquid ones, +2.422 against +0.597. It is tempting to read that as the crowding argument showing up as a measured quantity, and it does not support the weight. Run through the study's own paired week-clustered bootstrap, the T4 minus T1 gross difference is +1.825 with a 95% interval of -0.826 to +4.493 and p = 0.184, which contains zero. The specification states in advance that this design cannot resolve differences below roughly 2 basis points per day, which puts 1.825 under the line, and the comparison is not a member of the registered secondary family, which is defined on net rather than gross.

It fails a harder test as well. Under the adverse cost model, where configurations are re-selected because higher per-fill costs change which parameters win, T4's gross expectancy falls from +2.422 to +0.997 while T1 stays at +0.597. A pattern that halves when you re-price the fills was never a property of the instruments.

Cost is the part of the picture that is not fragile. It rises monotonically across all four tranches, from 5.17 basis points per sleeve-day in T1 to 6.74 in T4, and it does so under either cost model.

TestPrimary cost modelAdverse cost model
Sharpe(30) minus Sharpe(6)-1.066, p 0.0278, not rejected-2.430, p 0.0002, rejected
T2 minus T1-2.438, p 0.0362, not rejected-3.381, p 0.0034, rejected
T3 minus T1-1.006, p 0.4364, not rejected-2.962, p 0.0122, rejected
T4 minus T1+0.258, p 0.8409, not rejected-4.080, p 0.0018, rejected

The registered secondary family, corrected together with Benjamini-Hochberg at 5%. Under the primary cost model nothing survives correction. Under the adverse model all four reject, and every one points the same way.

Breadth did exactly what the correlation structure said it would

Fig. 4: The breadth ladder and its registered model. Realized Sharpe falls from -2.384 at six markets to -3.450 at thirty, and the multiplier model predicts the same fall from the measured correlation structure alone.
Fig. 5: The model gap at each rung, with its paired 95% interval. Every interval contains zero, so breadth delivered the multiplier its correlation structure implies and not one unit more.

This is why the ladder falls. The sleeves were already close to orthogonal at six markets, with an average pairwise correlation of 0.011, so the diversification multiplier was already 5.08 and the concave square root had flattened, and adding twenty-four markets moved it only to 7.60. A multiplier scales Sharpe without being able to change its sign, so applied to a per-sleeve expectancy that stays negative at every rung, a larger multiplier makes the book's Sharpe more negative, which is precisely what the ladder shows. Every rung has a negative cumulative net return, so none of them improved on any reading. They lost at different rates.

The specification registered a counterfactual that separates the two possible causes. Hold every added sleeve at the incumbent per-sleeve mean and volatility while keeping the realized 180-sleeve correlation structure, and the thirty-market book would have scored -3.572 against an actual -3.450. Splitting the -1.066 Sharpe deterioration on that basis gives -1.188 for the breadth multiplier acting on a negative per-sleeve mean against +0.122 for the added instruments' own expectancy, which pulls very slightly the other way. The residual is exactly zero, and the breadth component's interval of -1.778 to -0.649 excludes zero.

The ladder falls because averaging more losing bets sharpens the loss, not because the new instruments were worse.

The turnover story in two numbers

Fig. 6: Fills rise more than fivefold across the ladder while mean gross exposure stays near a quarter of capital. Hover a rung for both figures.

The binding constraint was already visible in the six-market book, where gross came to 7.63% of sleeve capital against 66.05% of fees, slippage and funding over the out-of-sample period. Widening the universe did not change that arithmetic, it bought more instruments to pay it on.

Per fill the two groups are almost indistinguishable, the six incumbents losing 6.248 basis points against the added markets' 6.408, so what differs is the number of fills, 33,657 against 148,881.

180
Sleeves walked forward, 30 markets by 6 families
-1.046
Added minus incumbent, bp per live sleeve-day
0.011
Average pairwise sleeve correlation at six markets
0 / 14
Profitable quarters at twenty markets or more
66.05%
Cost paid against 7.63% gross, six-market book
205x
Notional span of the rank axis

The window record is blunter. At twenty markets or more, not one of the fourteen out-of-sample quarters was profitable, while at six only one was, and removing the best window makes the cumulative result worse rather than better, because no good window is carrying anything. Anyone reading this as a route to a live book should read the levels rather than the differences, because every arm loses. The week-clustered bootstrap over 184 weeks puts mean weekly profit and loss at -0.633 with an interval of -0.824 to -0.434, which excludes zero.

What this rules out and what it does not

It rules out the claim the study was built to test, that instrument breadth was the binding constraint. Going from six markets to thirty, on a fixed capital base with everything else held constant, moved per-sleeve net expectancy by an amount whose interval spans zero under the primary cost model and sits entirely below zero under the adverse one, while moving portfolio Sharpe in the wrong direction for a reason a one-line formula fully explains.

Whether less-covered instruments hold more edge is a separate question this study does not settle, because the gross difference it measures is not distinguishable from zero at the resolution the design was pre-registered to have. What it does establish is the cost side, and that side is unambiguous.

Addendum: was it the cost model?

Everything above rests on charging every fill as a taker, five basis points of fee plus two of adverse slippage, on entries, stops, take-profits, reversals and time exits alike. It was a deliberate conservatism and it was stated as one, and it is the largest single assumption in the work. If it is wrong, the null above is an artifact of it.

It is wrong in one specific place. A take-profit is not a market order. It is a resting limit order, and on Binance USD-M perpetuals a maker fill pays two basis points rather than five and executes at its own limit price rather than through it, so it carries no adverse slippage. Of 91,277 out-of-sample trades, 17,970 exit at a take-profit, which is 19.7% and the only fill in this architecture that rests in the book.

Re-pricing those exits at the maker rate, holding every configuration exactly as selected, moves the thirty-market book from -64.68% to -59.69%. Five percentage points, recovered by changing nothing but what the exits are charged. That is not what you get. Net return is what the walk-forward optimiser scores, so making one exit type cheaper changes which configuration wins the nine-month parameter block, and re-running the registered selection under the corrected cost model gives -63.12%.

Fig. 7: What the correction promises against what survives. The mechanical saving is +5.00 percentage points; the realised saving is +1.57.

The optimiser gives back 69% of the correction as soon as it is allowed to respond to it, and very little has to change to produce that gap. Across 2,520 selection decisions, exactly 105 change, which is 4.2%, and that 4.2% destroys 3.43 of the 5.00 percentage points.

The optimiser is not behaving perversely. It does what a cheaper take-profit should make it do, with the realized share of exits taken at a target rising from 19.687% to 20.778%, so it leans into the newly cheap fill. It is simply not paid for leaning, because the configurations that look better in sample under the corrected objective transfer worse out of sample than the ones that looked better under the wrong one.

A cost-model correction cannot be evaluated by re-pricing a fixed book, because a fixed book is not what the strategy would have done under the corrected costs.

That also means the usual instinct is backwards. Discovering your cost model is too harsh feels like good news waiting to be collected, and most of it was not collectable.

Under order-type pricing the primary estimand moves from -1.046 to -0.993 basis points per live sleeve-day, an interval still spanning zero, so no ruling changes. Every rung of the breadth ladder stays negative under both cost models. Correcting the single largest cost assumption in the work, in the one place it is genuinely wrong, moves a sixty-five point hole by one and a half points.

Limitations

  • One venue, one contract type, one strategy set. Six families of classical trend, breakout, momentum and mean-reversion rules on hourly bars, so a rank effect measured under that set is a statement about that set. Thin contracts paying no better under these rules does not establish that nothing pays better there, and the gross result points the other way.
  • The out-of-sample period is shared with the six-market study that preceded this one. The incumbent contribution is a known quantity rather than a discovery. The genuinely unseen information is the twenty-four added markets, which is why the primary estimand is a difference dominated by them.
  • Rank-scaled costs are declared rather than calibrated. The intended estimator was degenerate on hourly perpetual data, and its only validation source begins inside the out-of-sample period, so routing on it would have leaked. No tick-level quote source available here covers the in-sample window for these contracts, so a properly calibrated cost ladder was not available at all.
  • The model check carries a declared bias. Correlation is computed on daily sleeve series that are exactly zero on roughly three-quarters of sleeve-days, because a sleeve is flat most of the time. Those zeros mechanically depress correlation, which raises the multiplier and makes the predicted Sharpe more negative. The bias runs in the same direction as the conclusion that realized tracks predicted, so the point estimates should be read with that in mind.
  • An earlier draft led with the tranche gross pattern as a positive finding. Rechecking the implementation from source, without reference to the write-up, established that the contrast contains zero, is not part of the registered estimand set, then showed that it does not survive its own registered sensitivity. That draft also described a mandatory sensitivity as though it had been run when it had not, and running it changed the ruling. Both are recorded rather than quietly fixed, because the failure happened at the writing stage after correct analysis, which pre-registration does not catch on its own.

Every chart on this page is drawn in your browser from the study's stored result tables, the same numbers reported in the text and on the reproducibility page. There are no image files behind them.

Full methodology

Universe construction, contract death, the walk-forward loop, the cost ladder, inference, the breadth model check, every leak check and the addendum's method.

Reproducibility →