How to measure price impact from free daily data
The estimator, why the impact exponent is fitted rather than assumed, how the null is calibrated against synthetic data, and the point-in-time discipline that keeps a backtest honest.
The estimator
In a Kyle (1985) framework r = λ · flow, so λ — price
impact per unit of trade — is the conductance. With daily free data:
log G_t = log |r_t| − α · log v_t
v_t = V_t / median(V, trailing 252 days)
Normalising volume by a trailing median removes secular volume growth and contract-roll drift without touching sub-annual variation. The window is trailing only, so it introduces no lookahead.
Fitting the impact exponent
The square-root law of market impact assumes α = 0.5. Amihud
illiquidity assumes α = 1.0. Rather than pick, regress log|r| on
log v and read the exponent off the data: α̂ = 1.032 for gold, and a range
from 0.41 (euro) to 1.36 (S&P 500) across the universe. No single assumed exponent would have
been right.
Fitting α has a second benefit that matters more than accuracy. At the fitted value,
log G is orthogonal to log v by construction, so
persistence in conductance cannot be an artefact of volume persistence. An automated test verifies
the residual correlation stays below 0.02 on synthetic data. Results are reported at
α̂, 0.5 and 1.0; all three agree.
The null, and why it is sharp
Under the joint hypothesis that
Gis a constant, and- volume and volatility co-move only through information arrival (the mixture-of-distributions hypothesis),
log G is iid. That is the reason for this particular estimator: it
turns a vague hypothesis into a sharp, testable one.
The null distribution comes from 2,000 iid bootstrap resamples of the empirical series, preserving its exact fat-tailed marginal while destroying time structure. The identical statistic pipeline runs on real and simulated series, so anything the pipeline itself induces appears in the null too. Statistics: AR(1), AR(5), AR(22), Ljung-Box(22), variance ratios at 5, 22 and 66 days, and out-of-sample R² of a HAR forecast benchmarked against the expanding mean.
Calibrating size and power
The headline result rests on the test rejecting a constant-conductance world at roughly its nominal rate and having power against a varying one. That is checked, not asserted. Synthetic markets are built with known ground truth — persistent volume in both arms, conductance constant in one and an AR(1) state variable in the other — and the suite verifies the test does not over-reject under the null and does detect the alternative. It also verifies exponent recovery and the orthogonality property above.
Constructing signed Λ
Λ_down = MM_long × w(underwater_long) / absorption
Λ_up = MM_short × w(underwater_short) / absorption
absorption = OI × breadth × (1 − top-4 concentration)
Trigger distance uses a cost-basis tracker: a weighted-average entry price that moves toward the current price only when a position is added to, and stays put when trimmed. Distance from basis is scaled to a monthly volatility horizon and passed through a logistic.
Correction applied mid-run. The first version scaled by daily volatility (about 1%) while positions sit about 10% from basis. That saturated the logistic to exactly 0 or 1, turning Λ into a binary switch. It was fixed before the reported numbers were produced. Both versions fail their kill conditions; the fixed version fails less badly.
Point-in-time discipline
The CFTC Commitments of Traders report is a Tuesday snapshot published the following Friday at 15:30 ET — a three-day information lag. Every series stores both a reference date and a release date, and every join uses release date only. Timestamping by reference date instead produces a backtest that looks excellent and means nothing. An automated test asserts no row is ever consumed before the day it was published.
The scheduled-event design
Persistence alone cannot say why conductance persists. Two stories fit: a transmission operator varying with market structure, or autocorrelated shock magnitudes — clustered news. Scheduled FOMC decisions separate them, because the date is fixed years ahead and the surprise size is close to unpredictable by construction of an event study.
log|r_event| ~ β · log G_pre + γ₁ · log RV_pre + γ₂ · log IV_pre
G_pre and RV_pre are means over a strictly prior window; IV_pre
is the last GVZ close before the event. The benchmark is deliberately hostile: GVZ is forward-looking
and already knows the meeting is coming. With about 148 events an expanding window is too thin, so
out-of-sample uses leave-one-year-out cross-validation. A placebo runs the identical
specification on all non-FOMC days.
Only regularly scheduled meetings are used. Unscheduled and emergency actions happen because markets are stressed; including them would manufacture the exact correlation under test. Exclusion follows the Federal Reserve's own labels.
Structural breaks handled explicitly
| Date | Break | Consequence |
|---|---|---|
| 30 Jan 2015 | GOFO discontinued | Lease-rate series ends; anything post-2015 must be synthetic |
| 1 Jan 2022 | SA-CCR reclassifies gold derivatives | Inflates reported precious-metals notional in OCC data |
| 13 Jan 2026 | CME moves precious metals to percentage margin (gold 5%) | The compulsion mechanism itself changed; pre- and post-2026 Λ are not the same variable |
References
- Kyle, A. (1985). Continuous auctions and insider trading. Econometrica 53(6), 1315–1335.
- Amihud, Y. (2002). Illiquidity and stock returns. Journal of Financial Markets 5(1), 31–56.
- Gabaix, X. & Koijen, R. (2021). In search of the origins of financial fluctuations: the inelastic markets hypothesis. NBER 28967.
- Brunnermeier, M. & Pedersen, L. (2009). Market liquidity and funding liquidity. Review of Financial Studies 22(6).
- Corsi, F. (2009). A simple approximate long-memory model of realized volatility. Journal of Financial Econometrics 7(2).