← Alpha
Alpha Research · · 11 min read

Does tick data beat OHLC for return forecasting?

900 symbol-days of nanosecond US equity trades and quotes, against the same bars reduced to open, high, low and close. Same folds, same models, same seeds.

Nanosecond trade and quote data is expensive, heavy and slow to work with. A single US equity symbol-day arrives as a gzipped CSV of tens of millions of rows, a modest five-name sample runs to 86.6 GB compressed, and everything downstream of it needs a different engineering budget than a folder of minute bars does.

Anyone paying that price is buying one belief: that the microstructure inside the tape carries information which open, high, low and close throw away. That belief is rarely tested as a controlled comparison.

Microstructure obviously describes the state of the market, since spread, depth, aggressor mix and order-flow imbalance are literally measurements of it, yet describing the present is a different claim from forecasting the future. The narrower question, and the one a desk has to answer before signing a data contract, is whether the extra columns survive into an out-of-sample return forecast at a horizon somebody would trade.

We built the comparison so the answer could not hinge on an arbitrary choice. Two datasets are constructed from the same filtered tape and evaluated on the same bars, the same folds, the same models and the same seeds, differing by exactly one thing: whether the tape-derived columns are present.

Reproducibility

The vendor schema, the two traps inside this feed, every formula and every leak check live on their own page, so they don't crowd the argument here.

Read the reproducibility →

What we tested

Three arms were fitted in a matched-pair design. Arm B is the OHLC baseline, 53 features that are all functions of open, high, low and close. Arm A is B plus everything derivable from the nanosecond tape, 204 features in total, which makes it a strict superset so any gap between the two is attributable to information content and not to a different modelling choice.

Arm T, 156 features, carries the tape plus raw bar geometry with none of B's technical indicators and answers whether the classical indicator set still earns its place once microstructure is available.

The universe is five tickers picked to span price level, spread regime and message rate rather than size: AAPL, SPY, JPM, XOM and F. That last name matters more than its market cap suggests, because at roughly $12 a share one cent of tick is 8.3 basis points where the same cent on SPY is 0.22, and spanning that range is what lets the cross-section test a mechanical explanation later on.

Three blocks of 60 consecutive trading days were taken from different years and different volatility regimes: February to April 2019, February to April 2022 and February to April 2024. That comes to 900 symbol-days of Algoseek nanosecond trades and quotes, 86.6 GB compressed, rebuilt into a 1-minute regular-trading-hours grid.

Five forecast horizons were evaluated on that grid so no conclusion depends on one lead time: 1 bar, 60 bars, 390 bars (a full session), 1,950 bars (a week) and 8,190 bars (a month). The target is the forward log return of the bar close throughout, chosen because both arms can observe it, whereas an NBBO midpoint target would hand arm A a quantity arm B cannot construct.

Three gradient-boosting implementations were trained side by side, LightGBM, XGBoost and CatBoost, all under one deliberately heavy regularisation setting fixed before the run from the target's signal-to-noise. Validation is a purged expanding walk-forward inside each symbol-block, discarding the final h bars of every training range so an overlapping target cannot reach into the test window.

Two things about the evidence base needed correcting before any comparison could be read. Overlapping h-bar targets do not supply n independent observations, and five tickers whose mean pairwise one-hour return correlation is 0.32 do not supply five independent series, which by the standard variance argument leaves 2.19.

Every result therefore carries a corrected effective sample size, with significance judged against an empirical null built by circularly rotating the target inside each symbol-block. That rotation destroys the feature-target correspondence while preserving every marginal distribution and the whole correlation structure.

Results

At the 1-minute horizon the tape wins decisively. Pooled out-of-sample rank IC roughly doubles, from about 0.045 on OHLC alone to about 0.103 with the tape, and arm A beats arm B in all 36 matched cells across three blocks, four folds and three model families at a Wilcoxon p of 2.9e-11.

HorizonIC, arm AIC, arm BNull IC 95th pctA minus BA winsWilcoxon p
1 minute0.100 to 0.1070.043 to 0.0480.014+0.05736 of 362.9e-11
1 hour0.012 to 0.0250.016 to 0.0360.096 to 0.108-0.00113 of 360.52
1 day0.029 to 0.0360.015 to 0.0300.15 to 0.20+0.04424 of 360.052
1 week0.152 to 0.1590.092 to 0.0960.23 to 0.30+0.06121 of 360.23
1 monthnot testablenot testablen/an/an/an/a

Pooled rank IC by horizon. Ranges span the three model families, so nothing here depends on which library is used. One month has no row: its test window is shorter than its own forecast horizon.

Fig. 1: Pooled out-of-sample rank IC by horizon. The third bar is the 95th percentile of the rotation null, taken as the higher of the two arms' nulls at that horizon. Only at one minute does either arm clear it.

The 95% bootstrap intervals at one minute do not come close to overlapping, with A at [0.095, 0.104] against B at [0.039, 0.048] for LightGBM, and against a rotation null whose mean IC is 0.002 with a standard deviation of 0.008 the observed 0.100 sits roughly 12 standard deviations out.

Past one minute nothing survives. At one hour the observed IC of 0.022 lands exactly on the null mean of 0.022, and neither arm clears its own null at any horizon beyond a minute, while the one-day paired difference of +0.044 sits at p = 0.052 on 24 effective observations, which is suggestive and no more than that.

Fig. 2: The headline paired difference, one bar per horizon. The muted bars are horizons whose effective sample size cannot support a test: the large one-week and one-month bars are not findings.

Given the tape, conventional OHLC engineering adds nothing measurable. Arm T runs 0.101 to 0.107 at one minute against arm A's 0.100 to 0.107, so the entire RSI, MACD, Bollinger, Stochastic and Donchian apparatus contributes no incremental out-of-sample skill once microstructure is present.

Only two of the five horizons can be tested at all

Correcting for overlapping targets and for the 2.19 effective series gives the honest sample size per fold, and it collapses fast.

HorizonEffective observations per foldVerdict
1 minute9,408testable
1 hour156testable
1 day24underpowered
1 week4.3not testable
1 month0.1not testable

Effective observations per fold after correcting for target overlap and for cross-ticker correlation.

The one-month test window is shorter than its own forecast horizon, 390 bars against an 8,190-bar target, so every target inside it points at essentially the same future instant. Its apparent 70% directional accuracy and its large positive paired delta are artifacts of a single overlapping observation, so week and month are reported as untestable instead of as weak evidence.

The one-minute advantage tracks tick size

At one minute the advantage holds on all five names, winning 94% to 100% of cells each, but the ordering carries the finding.

TickerPriceOne tick in bpsIC, arm AIC, arm BDifferenceCells A wins
F~$128.30.2110.137+0.074100%
XOM~$1001.00.0840.007+0.076100%
JPM~$1500.670.0730.015+0.05897%
AAPL~$1700.590.0460.010+0.03694%
SPY~$4500.220.0440.007+0.03797%

Rank IC at one minute by ticker, ordered from the largest relative tick to the smallest.

Fig. 3: Predictability tracks the tick as a fraction of price almost monotonically, from F at 8.3 bps per tick down to SPY at 0.22.

That is not the signature of a forecasting edge but of mechanical price discreteness, and at one hour the per-ticker differences change sign across names, which is what noise does.

What the models are actually using

SHAP attribution inside arm A at one minute puts 79.7% of the weight on tape columns, and the top of the list is unambiguous: a_micro_minus_close_bps alone takes 12.3%, followed by a_q_micro_close_bps at 6.5%, a_micro_minus_mid_bps at 2.2%, a_t_vwap_bps at 2.1%, a_q_mid_tw_bps at 1.9% and a_q_mid_close_bps at 1.8%.

All six measure one quantity, namely how far the last printed trade sits from a fair-value estimate. What the model has learned is that a print below the microprice tends to be followed by a higher close and the reverse, and that is the bid-ask bounce observed directly, not predicted.

Past one minute the attribution shifts to slow quantities: momentum over 1,950 bars, realised volatility, return skewness and trailing averages of spread and odd-lot share. None of those horizons beat their null, so this describes what the trees latched onto in sample and nothing about predictive content.

Fig. 4: Where the fitted trees put their weight inside arm A. At one minute the tape takes 79.7% of the attribution; past that the split reverts towards even, on horizons that never beat their null.

None of it is worth money

Whatever the one-minute signal is worth statistically, it is smaller than the cost of acting on it.

ArmGross bps per barCost bps per barBreakeven costGross SharpeNet Sharpe
A, OHLC + TAQ0.43 to 0.470.931.00 bps+16 to +18-17 to -19
B, OHLC only0.17 to 0.190.800.51 bps+6.5 to +7.1-23
T, TAQ only0.44 to 0.460.940.98 bps+16 to +17-18

The economic check at one minute, trading the sign of the forecast. Ranges span the three model families.

Mean effective spread over the sample, measured from the tape and not assumed, is 1.60 basis points. Read gross, arm A produces an annualised Sharpe near 17, a number that should provoke suspicion before it provokes a term sheet, since the signal flips side on roughly 45% of bars and every flip pays that spread.

Fig. 5: What each arm earns per bar, what it could afford to pay, and what the tape says the market actually charges. The tape roughly doubles the affordable cost and still lands under the spread.
900
Symbol-days of nanosecond TAQ
36 / 36
Matched cells the tape wins at one minute
2.9e-11
Wilcoxon p on the one-minute paired test
1.60
Measured effective spread, bps
1.00
Bps the tape arm can afford to pay
-17
Net Sharpe of the tape arm

The breakeven column carries the honest summary. The tape roughly doubles what the strategy can afford to pay, from 0.51 to 1.00 basis points, against a market charging 1.60, so both representations lose money net while the tape loses less.

Nanosecond trade and quote data locates the current transaction price relative to fair value far more precisely than OHLC does, and that precision decays within about a minute. For execution, market making and transaction cost analysis that minute is the whole game. For directional forecasting at any horizon beyond it, on this evidence, OHLC gives up nothing.

The negative result past one minute deserves as much weight as the positive one at a minute, and it is the more surprising of the two. Roughly 150 tape-derived features covering order-flow imbalance, aggressor mix, quoted and effective spreads, depth, microprice deviation, quote intensity, Kyle's lambda, Roll's spread, markouts, odd-lot share and venue concentration produced an out-of-sample IC that landed on the null mean at one hour.

Whatever the nanosecond feed knows about the following hour, three well-regularised gradient-boosting implementations could not extract it from these features on this sample.

Every leak, causality and null gate we ran against this result passed, including two gate bugs we found and fixed before the reported run; the exact checks and their numbers are on the reproducibility page.

Limitations

  • Only two of the five horizons carry enough independent evidence to test. Week and month are reported as untestable, and extending to them properly needs years of continuous history instead of 60-day blocks.
  • Five tickers is a small cross-section and they are correlated at a mean pairwise one-hour return correlation of 0.32. The design spans price level and tick regime deliberately, which is what made the tick-size ordering visible, though a wider universe would test that ordering rather than suggest it.
  • The target uses the trade close and not the midpoint, chosen so both arms could observe it, yet the close is precisely the quantity that carries the bounce, and a midpoint target would likely shrink the one-minute advantage substantially.
  • All off-exchange prints were removed through the single FINRA venue label this feed carries, which discards roughly half of consolidated volume for some names, so a design that kept them could reach different conclusions about liquidity-related features.
  • Hyper-parameters were fixed a priori and never tuned, which protects against selection bias while leaving neither arm at its best achievable configuration.
  • The cost model charges the measured effective spread plus half a basis point of fees without modelling market impact, queue position, partial fills or latency, all of which make the picture worse rather than better.
  • Every number rests on one vendor's consolidated SIP reconstruction, and two vendor-specific properties of that feed had to be established empirically before anything could be computed.

Every chart on this page is drawn in your browser from the study's real result tables, the same numbers reported in the text and the reproducibility page. There are no image files behind them.

Full methodology

Data provenance, the two traps in this feed, feature formulas, the walk-forward loop, the rotation null, the cost model and every leak check, with code snippets.

Reproducibility →