← Alpha
Alpha Research · · 11 min read

Should a strategy's recent results change how it manages the next trade?

Size down after three losses, more room after a good run. The belief is old enough to have its own vocabulary. We held the entry signal fixed and tested what the management decision is worth.

Almost every trader believes some version of this. Size comes down after three losses. The next trade gets more room after a good run. The belief is old enough to have its own vocabulary: pressing a hot hand, respecting a drawdown, trading smaller until form returns. Systematic desks encode the same instinct in equity-curve filters, dynamic sizing and rules that widen a stop when the last few trades went well.

The claim is testable, and it is narrower than it looks. It is not a claim that a strategy works. It says that once the decision to enter has already been made, the strategy's own completed-trade history carries information about how that entry should be handled: where the stop goes, where the target goes, how big the position is or whether to take it at all.

We tested it on a fixed library of mechanical strategies across two markets and two speeds, holding the entry signal fixed so that any difference has to come from the management decision. One statistic the study is required to report then turned out to weaken its own evidence for three of the five management families, so a second protocol was registered as an addendum, frozen under its own configuration hash, to remove that weakness and re-test those three. Both are reported below.

Reproducibility

The exact recipe behind every number here, including what the addendum changes and the frozen hash each one carries, lives on its own page.

Read the reproducibility →

What we tested

Every comparison in this study holds the entry fixed. Two policies see the same opportunity at the same bar and differ only in what they do with it. That is what makes the test clean, and it is why the result is about management rather than about the strategies themselves.

Five families of management rule were tested, each reading a summary of the strategy's own recent performance and moving one lever: the profit target, the stop, both together, the position size or whether to take the trade at all. Eight different summaries of recent performance drove each family, among them cumulative result over the last ten and fifty trades, win rate over several lookbacks, the current winning or losing streak and the drawdown of the last fifty trades. Every summary was run in both directions, covering the hot-hand reading, where good results justify doing more, and the mean-reversion reading, where bad results do. That is 80 distinct management policies, and neither direction was chosen by us after the fact. Both were registered in advance and both were searched.

Underneath those policies sits a base library of 80 mechanical strategies drawn from eight standard families: moving-average crossover, momentum, mean reversion, breakout, volatility breakout, trend following, Bollinger reversal and RSI reversal, every one of them signalling in both directions. The universe is 10 crypto perpetual futures and 10 large-cap US stocks, each at 15-minute and hourly resolution, which gives four cells that are reported separately and never pooled, because a finding in crypto is not evidence about equities. Out-of-sample scoring runs from January 2023 to July 2026 in quarterly steps, each quarter scored by a policy selected on the preceding 24 months alone.

Strategies were selected before any out-of-sample number existed, on trade frequency and holding period only, never on profitability. Barrier width and reward-to-risk are tuned in sample per window, so the baseline the treatments have to beat is tuned static management rather than an arbitrary two-to-one strawman.

Nothing here is a costless backtest. Entry crosses the spread and pays the taker fee plus slippage. A profit target rests as a limit order, so it pays the maker fee with neither spread nor slippage, while a stop crosses and pays taker plus slippage. Crypto positions pay or receive funding at every settlement they cross, at the rate actually observed and with its real sign, and funding is reported on its own rather than folded into the net figure, because over this period it ran as a persistent cost against a long book and would otherwise be easy to mistake for a strategy effect.

A trade ends at its stop, at its target or when the strategy changes its mind and the position flips. There is no time exit and no maximum holding period, because a forced exit is an artifact of the harness rather than a decision a trader makes.

What counts as a finding

Crediting a management family in a cell requires clearing eight conditions. Seven are descriptive. They ask whether the improvement exists at all, survives on executed trades, leaves the risk ratio intact, wins more months than it loses, survives deletion of its single best month, beats a matched placebo and holds up when the trade sequence is perturbed.

The eighth is statistical: a paired bootstrap that resamples whole calendar weeks, so trades sharing a week move together and overlapping positions cannot fool the test, after which every comparison in the declared family is corrected together for multiple testing.

Among the seven, the placebo deserves description because it is the sharpest. It takes the same policy and rotates its actions across the same trades, preserving the exact count of large positions, small positions and skips while destroying only the alignment between the equity state and the trade it was reacting to. A policy that cannot beat its own scrambled actions is responding to how often it acts, not to when.

Results

Across 20 comparisons, being five management families in each of four market-and-timeframe cells, no family was credited anywhere. The statistical condition failed in all 20, with adjusted significance values starting at 0.700 against a threshold of 0.05.

Fig. 1: Every comparison in the study, plotted as its measured improvement against its adjusted significance. The dashed line is the bar a result has to clear to count, and nothing comes within an order of magnitude of it. Colour is the market and timeframe; hover a point to name the management family.

Measured differences are mostly positive and uniformly tiny, running from -0.0014 to +0.0094 units of risk per signal. A trade in this design risks a full unit, so the largest effect found anywhere in the main run is worth under 1% of one trade's risk.

20
Comparisons in the declared family
0
Management families credited
0.700
Best adjusted significance, against a 0.05 bar
+0.0094
Largest improvement, risk units per signal
19 / 20
Comparisons that beat their own scrambled actions
70.3%
Trades that ended at a reversal, not a barrier

A row that clears six or seven of the eight conditions is not nearly credited, and that is worth saying plainly, because the table invites the opposite reading. The other conditions are descriptive, saying only that the overlay moved the right way and held its shape. None of them tests whether the difference can be told apart from noise, and the one that does failed everywhere.

The more interesting half

The management policies did beat their scrambled placebos. In 19 of the 20 comparisons the real policy landed above the 95th percentile of its own rotated actions, with percentiles running from 0.807 to 1.00.

Fig. 2: Where each policy ranks against 1,000 rotations of its own actions. The rotation preserves how often the policy acts and destroys only when it acts, so clearing the line means the timing carries information.

Scrambling the alignment between recent results and the next trade reliably degrades performance. The alignment is therefore carrying something measurable and repeatable. What it carries is simply not worth acting on: the effect is real in sign, consistent in direction and far too small to separate from noise once you account for how many ways it was searched for. That is a more precise answer than a flat verdict that the idea fails.

The statistic that forced an addendum

One of the required reported statistics complicates the null above, and it is reported here rather than buried, because it changes how much of the result you should believe.

70.3% of trades ended at a reversal rather than at a barrier. The cause is mechanical. Barrier distances are tuned in sample, the search prefers the wider settings and a wide barrier is rarely reached before the next signal arrives.

For the two families that act on size or on whether to trade at all, this changes nothing, because they act on every candidate regardless of how it ends. For the three families that move a stop or a target it changes everything. Barriers that seldom execute give those families almost nothing to act on, so a null result for them is weak evidence: it says the treatment had little opportunity, not that recent results carry no information about barrier placement.

Addendum: the barrier-only protocol

The addendum is an addition to the study above, not a replacement for it. Nothing in the main run is re-scored, the two are never pooled and the addendum was specified in writing and frozen under its own configuration and implementation hashes before any of its out-of-sample numbers were produced or opened. It re-tests only the three families whose test was compromised, since the other two were already tested fairly.

Its single change: a trade ends at its take-profit or at its stop-loss and in no other way. No time exit, no maximum hold, no reversal.

Fig. 3: How trades ended. Under the study's exit rule the barrier families rarely got to decide anything; under the addendum's, a barrier decides essentially every trade.

That change raises a question the main study never had to answer, namely what happens when a signal arrives while a position is already open. Both answers were run, as separate branches, and both are reported.

BranchOn a new signal while already holding
Single positionthe signal is declined and the book keeps what it holds
Stackinga further position opens, so trades overlap

Both branches are offered an identical stream of signals, verified cell by cell, and both are scored on net result per signal offered. That shared denominator is what makes them comparable: the branch that declines signals is not flattered by being measured against a smaller count. Because the barriers now carry the whole burden, barrier distances were recalibrated per cell before the run, reading holding period and resolution only and never profit.

Across its 24 comparisons, being three families in four cells in two branches, no family was credited in either branch. The statistical condition failed in all 24, with adjusted significance values starting at 0.819.

Fig. 4: The addendum's comparisons, one panel per occupancy branch, coloured by market and timeframe on the same key as the first figure. Giving the barrier families a clean opportunity moved the effects around and left the significance untouched. Hover a point to name the management family.

Effects run from -0.0011 to +0.0144 units of risk per signal, so the largest anywhere is worth about 1.4% of one trade's risk. The placebo behaves as it did in the main run: 17 of the 24 comparisons beat the 95th percentile of their own scrambled actions.

With reversals removed, the only way a trade can fail to resolve is to still be open when the instrument's history runs out, and such a trade's result is set by where the data ends rather than by the strategy, which is the failure mode that could have sunk the design. Measured across every cell and both branches, it happens to between 0.08% and 0.46% of trades.

So the three families that move a stop or a target got a clean shot on essentially every trade, under two different occupancy rules, in two markets and at two speeds, yet they still produced nothing that survives correction. The main study's null for those three families was weak evidence, and it is not weak any more.

What each branch costs

Both branches are runnable books, and each pays for its rule in a different currency.

StatisticSingle positionStacking
Trades never reaching a barrier0.08% to 0.44%0.31% to 0.46%
Median holding period7.5h to 26h16h to 49h
90th percentile hold50.5h to 192h171h to 480.5h
Share of time holding a position23.7% to 44.1%31.0% to 74.6%
Signals declined because already occupied31.2% to 53.5%none
Long share49.6% to 50.0%50.0%

Ranges span the four market-and-timeframe cells. Every figure is measured out of sample except the exposure statistics below, which are measured in sample where the barrier settings are chosen.

The single-position branch turns down between 31% and 53% of the signals it is offered. That is the entire difference between the branches, and it is the price of never holding more than one position at a time.

The stacking branch places no cap on concurrent positions, which sounds unbounded, so signed exposure was measured directly. Counting open tickets overstates the risk, because a long and a short in the same instrument cancel; the honest measure is the signed sum, time weighted, over bars that actually hold something.

Fig. 5: Signed net exposure of the stacking branch, measured in sample and time weighted over the bars that hold something. A rule with no cap on concurrent positions turns out to hold one unit of direction most of the time.

On that basis the median is one unit of direction in every cell, the 90th percentile is one or two, and the maximum anywhere is five, so the extra tickets are transient. What the stacking rule actually spends is cost rather than risk capacity: both sides of an overlap are opened and closed, and its funding bill in 15-minute crypto came to roughly 18 times the single-position branch's.

The takeaway

A strategy's recent results do carry a faint signal about the next trade. Scrambling the alignment makes results measurably worse, which means the relationship is real rather than imagined.

Acting on it never produced an improvement large enough to distinguish from chance, across 44 tests spanning two exit rules, two occupancy rules, two markets and two speeds.

For a systematic book the implication is to fix the management rules, size them for the risk you intend to take and spend the research budget on the entry instead.

Limitations

  • The stacking branch barely stacks in half the cells. Measured in sample it holds a median of 1.05 and 1.07 positions in the two hourly cells, which makes it very nearly the same experiment as the single-position branch there. The comparison between branches is genuine only in the 15-minute cells, where the median reaches 1.33 and 1.61. This should have been measured before the out-of-sample was spent, while the design could still have been widened.
  • Coverage in the single-position branch is thinner. Declining roughly half the signals leaves fewer trades for the in-sample search, so in the 15-minute cells a management family is absent from up to 17% of windows because no policy was worth selecting. The stacking branch fields all three families in every window of every cell, so the single-position results rest on fewer windows.
  • One barrier setting exceeds its holding target. At the widest setting in 15-minute equities, 24.3% of trades run two weeks or longer against a 20% allowance. It is reported rather than removed, because deleting a setting to make a target look met is choosing the answer.
  • Equity borrow fees are not modelled. They are absent from the upstream data. At multi-day holds this understates the cost of short equity positions.
  • Six of ten gate suites were not re-run against the main study's engine. Four were, including the execution parity test, which is where that design actually changed. The addendum re-ran all four of its required suites against its own engine.
  • Scope is limited. Twenty instruments, two markets, two speeds and one definition of recent performance per family leave room for a different equity-state summary, a slower library or a longer horizon to behave differently. This is a null on this design, not a proof about the idea in general.

Every chart on this page is drawn in your browser from the study's stored result artifacts, the same numbers reported in the text and on the reproducibility page. There are no image files behind them.

Full methodology

Data, library calibration, the candidate grid, equity-state variables, the walk-forward loop, the cost model, inference, what the addendum changes and every frozen hash.

Reproducibility →