2026 · Python · public repository

market-gnn

Does a stock-relationship graph add out-of-sample signal? Twelve leak-controlled experiments found little evidence of return-predictive value.

Finding
The tested ownership graph did not add a statistically significant weekly lead-lag return signal.
Scope
Twelve experiments on 90 large-cap US stocks, 2014 to 2024, using point-in-time graph construction and out-of-sample evaluation.
Boundary
The lead-lag estimate was +0.012 rank-IC with HAC t 1.4. The study MDE was 0.025; that detection limit does not prove the true effect is zero.

The question

Stocks co-move in groups through sectors, supply chains, and shared factors. Model that structure as a graph, run a GNN over it, and ask whether the graph improves how well the model ranks future returns out of sample. The reported statistic is rank information coefficient (rank-IC), compared with an identical feature set fed to a matched non-graph model, after point-in-time construction and purged walk-forward evaluation.

The graph is tested both as a correlation proxy and as an observed ownership-network relationship: institutional co-holding (Anton-Polk “Connected Stocks”) built from the SEC’s raw 13F filings with the 45-day filing lag included. The study combines direct signal reads, a GNN-vs-matched-MLP comparison, and a learned temporal model (GConvGRU).

The headline figure

Out-of-sample rank-IC vs minimum detectable effect

Solid bar: observed IC. Dashed box: 80%-power minimum detectable effect. An estimate below that threshold is not proof that the true effect is zero.

The controls show that the pipeline recovers the effects they were designed to contain: planted lead-lag reaches IC +0.089 (HAC t 9.6), and a future-injection canary approaches IC 1.0. For 13F lead-lag, the observed estimate was +0.012 (HAC t 1.4). The design had 80% power for an effect of 0.025 under its stated assumptions, and the estimate was not significant.

Controls and detection limits

Planted recovery and the canary test whether known injected effects survive the pipeline. Label shuffling and degree-preserving rewiring test whether the same machinery manufactures signal under controls. Purged walk-forward splits and point-in-time graph construction limit two common leakage paths.

Reported diagnostics include Newey-West/HAC t-statistics, a block bootstrap for autocorrelated labels, BH-FDR within endpoint families, and the minimum detectable effect next to each non-significant result.

Strategy check for the reversal control

The short-term-reversal control produced a significant cross-sectional association (IC +0.015, HAC t 3.4 after the stated FDR procedure). A separate strategy-level check then tested whether selecting among its configurations produced a reliable gross portfolio result:

PBO 0.437
near the repo's 4/9 noise reference; the in-sample-best configuration did not select reliably out of sample
DSR 0.74
reported deflated-Sharpe statistic at N=9 trials for a +0.37 gross annualized Sharpe
~0.9 bp
modeled one-way cost at which the estimated net Sharpe crosses zero, before a survivorship correction

Every configuration in the grid had a positive gross Sharpe (+0.16 to +0.52). These results do not support reliable configuration selection or a net trading claim.

The paper trail

The repository includes an internal review log. REVIEW.md records prior bugs, including a CUSIP map that silently isolated five names, an FDR that pooled null controls, a primary endpoint that was never actually computed, and a covariance result caused by an indefinite-matrix leverage artifact. Links on this page are pinned to the corrected revision. The repository does not include a byte-for-byte regeneration record for every real-data number; later market-data downloads can be readjusted by the provider. Full reported numbers are in RESULTS.md. Clone the repository and run pytest -q to inspect its automated checks.