Research · market microstructure
Published · SSRNQuoting the Touch Does Not Pay Its Adverse Selection
A true-aggressor-signed entry-markout decomposition for a last-in-queue quoter across Bybit, Binance USD-M and Hyperliquid
When a market maker rests an order at the best price, it earns a slice of the spread as the order fills and then watches the price move against it, because whoever took the order often knew something. This paper simulates the simplest version of that maker on three crypto perpetual venues, across 9,983 coin-days and 125.8 million simulated fills, and finds that the move takes back more than the spread pays on all three. The sections below rebuild the argument in plain terms, with an order-book simulator, diagrams and the paper's figures redrawn from its shipped results.
View on SSRN
ssrn.com/abstract=7461901
Replication package
github.com/DaruFinance/crypto-adverse-selection
The result in three lines
9,983 coin-days · 125.8M fills
The spread does not cover the drift
At 10 seconds and before fees, the net entry markout is −0.84 bp on Bybit, −0.39 bp on Binance USD-M and −0.46 bp on Hyperliquid, with intervals clustered on both month and coin excluding zero on all three venues.
trade-sign counterfactual
Guessing the aggressor keeps the sign
Lee–Ready and the tick rule agree with the exchange flag on 89% to 98% of fills, never flip the result and move its size by less than a quarter. The centralized venues err in one direction and Hyperliquid in the other, so no constant correction exists.
scope
One leg of a round trip
The measurement stops at a markout horizon of 10 or 60 seconds after the fill, with no exit, inventory, fee or rebate, so it cannot say whether market makers on these venues make or lose money. The simulated quoter also queues behind all visible size at its price, the least favourable assumption short of never filling.
Overview
A companion paper on CME futures found that the two halves of a maker's fill almost cancel, with +0.654 bp of captured half-spread against −0.661 bp of drift and a net indistinguishable from zero on all 13 contracts. Running the same decomposition on crypto perpetuals turns that balance negative, since the drift after the fill is larger than the half-spread on every venue before any fee is charged. The two studies place the order at opposite ends of the queue (the front on CME, the back here), so the contrast cannot be pinned on the asset class alone.
What makes the measurement possible is that Bybit, Binance and Hyperliquid stamp every trade with the side that crossed the spread, a field most equity feeds leave out. With the true side known, the study reruns everything under the classifiers other researchers have to rely on and prices the error they introduce. Much of the remaining work goes into inference, because fill weight piles into a handful of coins and calendar months are few, and the paper reports which statements survive that instead of assuming they all do.
The measurement
A quoter at the back of the queue
The simulated maker posts one unit at the best bid and one at the best ask on every book snapshot, at most once every 100 ms, joining the back of the queue behind all visible size at that price. Its order fills only when an aggressor trades at that price with more size than the queue still ahead of it, in which case the fill is the overflow, capped by the amount still resting. Each order lives for one snapshot before being posted again at the back, so the quoter never carries priority forward.
Interactive · order-book simulator
Step through one fill on an illustrative book with a 1 bp tick. Each step is one book snapshot, and the blue block is the simulated order, posted again at the back of the queue every time, so only a trade larger than everything displayed ahead of it can fill it; after the fill the chart follows the mid for 10 seconds.
Bids
A book with five price levels per side. Press Step to post the quote.
Mid price over time
queue ahead
·
coins
filled
0.0
coins
capture
·
bp
adverse 10s
·
bp
net markout
·
bp
tick
1.0
bp
Illustrative scenario, not paper data. Each step is one snapshot; sizes in coins, prices in USDT.
s is the maker's side (+1 when the aggressor sold), f the fill price, m₀ the mid read strictly before the fill and m_τ the mid τ = 10 s later. Capture is fixed at the fill, and only the adverse leg depends on the horizon.
The reference mid cancels in the sum, so the net depends only on the fill price and the later mid. A staler m₀ shifts value between the two legs without changing the sign of the net.
Two choices in that rule decide how the numbers should be read. Queueing behind all displayed size is pessimistic, since a real maker holding time priority would capture more and be picked off less, and the paper makes no claim about such a maker. The mid is read strictly before the fill because a trade and the book update it causes often share a timestamp, which happens on 52.7% of Binance USD-M fills, and an at-or-before read would price the fill against the post-trade book.
How much that pessimism costs is visible on Hyperliquid, where Albers, Cucuringu, Howison and Shestopaloff measure observed maker fills on the BTC perpetual at a 5-second horizon. The average maker fill marks out at about −0.48 bp before rebate, while a group of traders who win first place in the queue earns +0.32 bp on the same basis, a gap of roughly 0.8 bp that is larger than two of the three nets reported here. For that reason the paper treats a front-of-queue arm on the same panels as the experiment that decides whether the negative sign belongs to quoting at the touch or to quoting at the back of it.
One fill, three prices
Capture is the gap between the mid just before the fill and the fill price. Adverse selection is how far the mid moves over the next 10 seconds, and nothing after that point enters the measurement, including the exit.
The central result
Negative on all three venues
Pooled over every fill, the captured half-spread is +0.21 bp on Bybit, +0.46 bp on Binance USD-M and +0.56 bp on Hyperliquid, while the drift over the next 10 seconds takes back 1.04, 0.85 and 1.02 bp. The net entry markout is therefore negative on every venue before fees, and the 95% interval clustered on calendar month and coin excludes zero on each of them. The three panels cover different dates and coin sets, so the bars are three separate measurements rather than a ranking of exchanges.
Thin capture on the biggest coins
Bybit's largest coins trade in books one tick wide for most of the window (no shipped panel records the quoted spread to confirm it), so a touch quoter earns about half a tick, which on BTC comes to 0.0124 bp per fill across 34.5% of the venue's fills. Dividing by a capture that small makes the adverse-to-capture ratio unstable, at 5.1 pooled on Bybit and 50.2 on Bybit BTC against 1.8 on Binance and Hyperliquid, and for that reason the paper calls the ratio its least defensible number. The net, built from the same fills, keeps its sign when any single coin is dropped on any venue.
Coin by coin
All 13 Bybit coins are negative with intervals that exclude zero, and all 15 Binance coins are negative, with 12 clearing zero, XRP falling short and two verdicts withheld. Hyperliquid is where the pattern frays: 13 of 20 coins are negative at 10 seconds, only BTC and SOL clear zero, 11 verdicts are withheld for having fewer than five effective months and seven coins carry positive point estimates, led by NEAR at +0.40 bp.
The pooled Hyperliquid figure also leans on four coins that carry 80% of its fills, and weighting every coin equally moves it from −0.46 to −0.15 bp. Together with its withheld verdicts, a month distribution dominated by one dense block of dates and the extraction gap described further down, that makes Hyperliquid the weakest of the three results even though its pooled interval excludes zero.
The trade-sign counterfactual
What guessing the aggressor costs
Most markout studies never see which side crossed the spread and have to infer it. Lee–Ready compares the trade price with the prevailing mid and falls back to the tick rule at the mid, whereas the tick rule compares each price with the last different one and carries the previous call through a run of equal prices. Because these three venues publish the true flag, the whole simulation can be rerun under each classifier so that the gap between runs measures what the missing flag costs.
Interactive · classifying a tape
Ten trades from an illustrative tape, each with its real aggressor side. Pick a rule to see which trades it mislabels; two of the trades share a timestamp with the book update they caused.
Up from the last different price is a buy, down is a sell and an unchanged price repeats the previous call, which fails when a sell lands on a price that a buy just set.
| time | price | size | mid used | rule says | exchange flag | |
|---|---|---|---|---|---|---|
| 1000 ms | 100.01 | 0.5 | · | buy | buy | ✓ |
| 1350 ms | 100.00 | 0.4 | · | sell | sell | ✓ |
| 1720 ms | 100.00 | 0.3 | · | sell | sell | ✓ |
| 2100 mssame ts as book update | 100.01 | 2.8 | · | buy | buy | ✓ |
| 2500 ms | 100.01 | 0.6 | · | buy | sell | ✗ |
| 2900 ms | 100.02 | 0.4 | · | buy | buy | ✓ |
| 3300 mssame ts as book update | 100.01 | 0.9 | · | sell | sell | ✓ |
| 3650 ms | 100.01 | 0.2 | · | sell | buy | ✗ |
| 4000 ms | 100.00 | 0.5 | · | sell | sell | ✓ |
| 4300 ms | 100.01 | 1.0 | · | buy | buy | ✓ |
agreement
80%
mislabelled
2
Illustrative tape, not paper data.
On the real panels the classifiers agree with the exchange on 89% to 98% of fills, and under every rule on every venue the drift still exceeds the half-spread, so the sign of the result never changes. A mislabelled trade can also create or remove a fill, which puts fill counts under a classifier between 0.930 and 1.082 times the truth, while only three of the six paired errors have intervals that exclude zero.
Accuracy does not rank the error
Hyperliquid's Lee–Ready agrees with the exchange slightly more often than Bybit's (93.6% against 93.2%), yet its error is 20.4% of the effect against 3.8%. Binance's tick rule also beats Bybit's on agreement and still carries 13.4% against 3.8%, meaning an agreement rate says little about how far a classifier moves a markout. What does travel is a bound, since on these panels no classifier moved the pooled markout by a quarter of its size or changed its sign.
The clock matters as much as the rule
A fill shares its timestamp with a book update on 0.3% of Bybit fills, 8% to 10% of Hyperliquid fills and 52.7% of fills on millisecond-stamped Binance USD-M. On one heavy Binance day the quote rule scores 0.996 reading strictly before the trade and 0.620 reading at or before it, a difference that belongs to the feed's clock and would otherwise be blamed on the classifier. Unlike the other figures on this page, the collision shares and the two quote-rule scores are read directly from the raw tape, which sits outside the reproducible package.
Inference
How many independent observations are there?
Every row of the panel is a coin-day, and coin-days are correlated along two directions that do not nest: days in the same month share market-wide shocks, while days of the same coin resemble each other. Clustering on month alone ignores the second direction, so the headline intervals use the two-way estimator of Cameron, Gelbach and Miller and report the month-only interval beside them. On Bybit the two-way standard error is 6.0 times the month-only one (0.0882 against 0.0147 bp), on Binance 2.0 times and on Hyperliquid 1.2 times, with no verdict changing.
Two ways coin-days are correlated
Many clusters on paper can still be few in effect. When one coin carries a third of the fills its days dominate the average, and Kish's effective count (the squared total weight over the sum of squared weights) shrinks toward the number of coins that actually matter. Bybit's 13 coins count as 5.96, Hyperliquid's 20 as 5.43 and Binance's 15 as 10.73, and wherever an effective count falls below five the shipped code withholds the verdict instead of reporting a weak one.
Interactive · effective number of coins
The bars are each venue's real per-coin fill shares. Slide from equal weights through the measured weights to a more concentrated panel and watch the effective count move, while the t multiplier follows whichever of effective coins and effective months is smaller and the verdict is withheld below five.
effective coins
0.00 of 0
effective months
·
t multiplier
·
verdict
withheld
share of the venue's fills per coin
Beyond the choice of estimator, the paper tests its interval methods on simulated panels shaped like these. A coin effect that persists across months makes month-only clustering reject a true null 86.7% of the time against a nominal 5%, while the two-way estimator holds at 5.7%. When the heaviest clusters are also the unusual ones, as Bybit BTC is, the t and two-way estimators run at 18.3% and the guard returns a verdict on only 14% of runs, the condition the paper ranks as its first threat.
Quoting deeper
The rebate decides where to quote
On a separate sample of 300 Bybit coin-days, buy side only, the quote was placed at five price levels from the mid. Each deeper level loses less per fill, from −0.72 bp at the touch to −0.33 bp at the fifth level, but fills far less often, at 12% of the touch's rate. Because a rebate is paid per fill, the level that loses least per quoting opportunity at zero rebate is not the one that wins once the rebate grows.
Interactive · depth and rebate frontier
Value per quoting opportunity is the fill rate times the net per fill plus the rebate minus any residual fee. Move the rebate to see which level wins, then add a residual fee and watch every threshold shift by the same amount.
best level
none
net per fill
·
deepest breaks even
0.000 bp
touch overtakes
0.000 bp
x: rebate, bp per fill · y: bp per quoting opportunity
Below a 0.328 bp rebate no level is profitable per opportunity, between 0.328 and 0.771 bp the deepest level is best and above 0.771 bp the touch overtakes it. The crossing has a coin-clustered interval of [0.716, 0.834], yet a residual maker fee moves it one for one, so which side a real maker lands on depends on its own fee schedule rather than on this panel. None of these thresholds is a profit claim, since the simulated quoter is not shown to satisfy the uptime, coverage or size obligations attached to real rebate tiers.
An inventory-aware quote
On 91 Bybit coin-days an Avellaneda–Stoikov style arm lost less per fill than the touch, −0.54 against −0.91 bp, by taking only 9.2% as many fills, and it was still negative on 66 of those days. Clustered by coin the advantage is not significant, and per quoting opportunity the ordering reverses above a rebate of 0.908 bp. The arm's own diagnostics show it behaving as a depth rule, so the result is silent on inventory-aware market making in general.
What the panels cannot settle
- Venue comparison. On 325 matched coin-days Bybit's net is 0.355 bp below Hyperliquid's and lower on 13 of 13 coins, but the 2,000 ms re-quote rung behind that figure was fixed by hand and the gap reads −0.350 or −0.223 bp depending on which extraction is used, so only the direction is claimed. The two venues also differ in tick grid, fees, matching engine and publication speed at once, which leaves no architectural cause to claim.
- Lead-lag. Bybit's book leads Hyperliquid's by a median 500 ms, close to the 441 ms gap in how often the two venues publish, which leaves information leadership and publication rhythm inseparable.
- Volatility and session. On a separate Bybit subsample, the calm-to-stress contrast in the Asia, Europe and US blocks gives p = 0.79, 0.77 and 0.29 on 12 coins, a low-power failure to reject rather than evidence that the result is stable across regimes.
- Delisted coins. Every panel is a survivor set, and ten Bybit perpetuals delisted inside the window were 0.848 bp worse than survivors on matched dates, bounding the omission at 0.072 bp by fill weight and 0.126 bp by coin count, in the direction that flatters the reported net.
- Fill selection. Because the quote fills only on trades larger than the queue ahead, it is filled by the book-clearing trades that tend to carry the most drift afterwards. Sorting each coin's days into terciles of mean fill size leaves Bybit flat and runs Binance the opposite way, while Hyperliquid is ordered in the worrying direction at −0.42, −0.45 and −0.52 bp; realised fill size is capped at one coin, so this test bounds only the part of the channel visible below the cap.
- Hyperliquid extractions. Two shipped extractions of the same 500 Hyperliquid coin-days, run at the same 100 ms re-quote ceiling, differ twofold in fill count and give nets of −0.58 and −0.34 bp, whereas Bybit's two extractions agree fill for fill. One of the two must therefore be wrong by an amount the size of the estimate itself, which the paper names as the item that would most change its conclusions if it resolved against the panel; settling it needs the venue-specific programs that sit outside the package.
- The simulator itself. No simulated fill is checked against a real one. Albers, Cucuringu, Howison and Shestopaloff measure the same quantity on observed maker fills in Hyperliquid's BTC perpetual over July 2025, at roughly −0.48 bp before rebate on a 5-second horizon weighted by volume, near this paper's −0.58 bp for Hyperliquid BTC at 10 seconds weighted by fill, which the paper reads as directional agreement rather than a replication.
What would overturn it
Overturning the claim would take a venue, coin or state whose net entry markout reliably clears zero on the positive side under the same interval rules. No such cell exists in these panels, because of the seven Hyperliquid coins with positive point estimates five return no verdict and the two that were tested do not exclude zero. The per-coin result files are laid out so that a replication would spot such a cell immediately.
Method
- One queue simulator runs on all three venues, posting one unit on each side of the touch at every snapshot (at most every 100 ms), queued behind all visible size and filled only on overflow.
- Each fill is split into captured half-spread and 10- and 60-second adverse drift, signed by the exchange's aggressor flag, with the reference mid read strictly before the fill.
- Pooled intervals cluster on calendar month and coin with the t multiplier taken at one degree of freedom below the smaller Kish effective count, while verdicts are withheld below five effective clusters.
- The classifier counterfactual reruns the whole simulation under Lee–Ready and the tick rule, so each rule produces its own fill set, and errors are paired within coin-day.
- Per-coin results are corrected for multiplicity, and all 27 cells that clear zero at 10 seconds survive Benjamini–Hochberg.
- The survivorship bound follows a plan that fixed the prediction, coin set, date matching and weightings in advance and was hashed; because the repository is author-controlled, that ordering is recorded rather than independently verifiable.
Reproducibility
The replication package, crypto-adverse-selection, ships the ten aggregate coin-day panels, the analysis programs that produce every result file from them, the figure scripts and a machine-readable contract listing the claims the evidence permits. The panels rebuild byte for byte from hash-verified frozen artifacts that are not themselves shipped, so a reader can check the hashes but not rerun that step, while the raw exchange archives and the venue-specific extraction programs also sit outside the package.
Comparing shipped results with rebuilt ones reaches different depths per file, covering every value in three files but only 68 of 6,575 in the lead-lag file, and a fresh parse of Bybit data after January 2024 no longer reproduces the shipped panel because the order-book parser now prefers a different timestamp. The package also installs a small command-line tool that applies the entry-markout decomposition and abstention rule to any book and tape carrying a true aggressor flag, although it documents and tests the fill rule and is not the program that produced the paper's fills. The interactive figures on this page are drawn from the shipped result files and panels.
Cite
See also
The CME companion, Adverse Selection Consumes the Touch, runs the same decomposition on futures with the order at the front of the queue, while No Edge Without Information studies liquidity in decentralized-exchange pools, where there is no order book or queue at all.

