London Paris Spurs Edge Lab

Spurs Edge Lab · quantitative systems · 2026

Live win probabilities that carry their own uncertainty.

Spurs Edge Lab is a research and trading system I built for NBA win-probability markets. It aligns play-by-play, live scoreboard feeds and Kalshi prices in time, computes strictly-prior features, runs a promoted model bundle with calibrated bounds, and keeps its order state correct when the exchange's answer never arrives.

Built solo, February to June 2026 · Python, LightGBM, pandas, pyarrow · 1,496 commits

Detroit at Oklahoma City Historical replay · 2026-03-30 · held-out game · overtime
Model decision probability, OKC winUncertainty boundsKalshi contract price

Recomputed from the promoted model over the 383 stored snapshots of a game it never trained on. Not the forecasts issued live that night.
6,565
completed NBA games with full play-by-play, 2021-22 through 2025-26
definition

One row per completed game in the promoted game universe; every game has both a final-score record and play-by-play that reconciles to it.

1.30M
offensive possessions, reconstructed and audited
definition

1,303,127 rows, one per offensive possession, with clock bounds, scoring outcome and correction flags. About 198 per game.

8.24M
time-state rows for training, two team perspectives each
definition

8,237,580 rows keyed by game, snapshot time and team perspective. Repeated observations within games, not independent games. 859,884 were captured live; the rest are reconstructed from play-by-play.

4,116
collected tests, including 18 calculator parity suites
definition

pytest collection count across 370 test files in the source repository on 2026-09-24. The parity suites check each live feature calculator against the frozen training-side values on sampled inputs.

01 / Explore one game

Held-out replay DET at OKC · 2026-03-30

A game the model never saw,
played back one snapshot at a time.

Detroit at Oklahoma City on 30 March 2026 was captured live at the time and held out of every training and calibration stage. Scrub the timeline, pick a moment, and read what the scoreboard, the market and the model said at that instant.

DET0
Q1
12:00
OKC0
Model, OKC win
–
Bounds
–
Kalshi price
–
Bid / ask · regime
–
383 snapshots
Move along the timeline to see the play recorded on each snapshot.
Diagnostic surfaces:
Decision probability (state-space point)Conformal boundsKalshi priceCalibrated model surface / raw composite (dotted)
Table view: every snapshot
Capture (UTC)ClockOKC–DETKalshiModelLowerUpperCalibratedRawRegime

How to read it. Time on the horizontal axis is game time; each snapshot also carries the capture clock stamped by the monitor. The market line is the home-win contract price as captured, in probability units. The solid model line is the surface the live runtime uses to decide, with its bounds; its point is a closed-form score-time estimate, and the evaluated, calibrated model surface is the dotted line. Nothing on the timeline uses information from later in the game; the final score appears only at the end.

02 / From raw observations to a decision record

Every record below is from this same game

Seven stages, one record each,
and the rule that holds them together.

The rule is strict priority in time. Pregame features are computed from games strictly before tip-off with a one-game shift on every rolling window; live features use only events before the snapshot; the model reads the same columns in training and in the live loop, and 18 calculator parity suites keep it that way.


      

03 / How the model is judged

Held-out period 2025-04-30 to 2026-04-25

Calibration first,
then the comparison that is fair.

The held-out set is the latest 1,310 games by date. The composite evaluation scores 1,304,150 team-side snapshot rows from 1,158 of them (games without a clean pregame anchor are excluded). Rows are repeated observations within games, not independent trials, so the exhibits below report calibration and a same-row comparison against the pre-tipoff market price rather than a single headline accuracy.

Reliability of the calibrated surface

Expected calibration error, raw → calibrated
–→–
Brier · log loss
––
AUC by game phase, calibrated surface
–early–final 5 min

What the bounds are. Per regime and probability bucket, a Wilson interval on the per-game hit rate of the calibrated probability, fitted on held-back training games at a 92% target and regenerated on 3 June 2026. No promoted coverage audit exists for these shipped bands, so no coverage figure is claimed here. Calibration itself is regime-gated isotonic, applied only where it improved on the raw surface (the midgame regime). The intervals are wide on purpose: the live runtime sizes positions on the lower bound, not the point.

Against the pre-tipoff market price, on the same rows

Baseline. The pre-tipoff Kalshi price, de-vigged so home and away sum to one, carried across every snapshot of its game. It is a static reference, so the gap measures what live information adds over the closing price; it is not a comparison against the in-game market, and not a claim about executable edge, fills or profit, none of which are established in the promoted evidence.

By regime, calibrated surface

RegimeRowsBrierLog lossECEAUC
Brier, pregame prior only
–
Brier, pregame + live state
–
Brier, full raw composite
–
Distributional heads and what is deliberately not claimed
TargetRowsMAEq05–q95 coverage (nominal 90%)Widthq25–q75 coverage (nominal 50%)

Not claimed. The promotion receipt blocks these until reconciled execution evidence exists:

    04 / Engineering under failure

    Traced from the current source and its tests

    Two places the system had to stay correct
    when the inputs were not.

    An order whose answer never came

    Symptom

    During a June 2026 session the exchange's gateway returned errors after orders had been sent. The client library had also, on other nights, crashed while parsing a successful reply. In both cases the order might exist with fills the process knew nothing about.

    Cause

    An exception after the request leaves the process is not evidence that nothing happened. Treating it that way would either retry into a duplicate or abandon a live position.

    Design

    Every attempt carries its own client order id. On any error after send, the thread queries the exchange for that id before believing the error. When the order is found, or the reply was unparseable, fills are cross-checked against the signed change in the account position; otherwise the error stands and any retry is bounded by a hard cap on contracts submitted per launch.

    Result

    Unknown outcomes are resolved from the exchange when it is reachable and stay explicitly unknown when it is not. Tests pin that the retry loop never submits more than requested when a fill is unknown.

    Money in flight, one writer of state

    Symptom

    Exposure accounting drifted: a fill could be visible in the exchange's positions before the process had recorded it, a zero-fill could mute a lane for a whole game, and callbacks on the order thread mutated dictionaries the decision loop was iterating.

    Cause

    Two sources of truth (local ledger and exchange) that needed reconciling, plus two threads writing the same risk state.

    Design

    Confirmed exposure comes from the exchange only; local memory tracks unresolved orders as in-flight, registered before the request leaves. Fill callbacks enqueue frozen events; the loop drains them as the single writer, releases in-flight on any outcome, refunds cap slots on terminal zero-fills, and holds new orders until the exchange confirms the fill.

    Result

    The in-flight ledger, the confirmation gate and the session caps have one writer. A restart rebuilds exposure from the exchange. Concurrency tests drain every event exactly once under contention, and per-event isolation means one failed application cannot drop the others.

    On the record. These sequences are reconstructed from the current source, its comments and its tests, not from exported incident logs. What the code guarantees and what it leaves unresolved are stated as such.

    05 / The film

    80 seconds · captioned · silent

    The whole idea,
    in a minute and twenty seconds.

    The held-out replay, one decision chain through the model, and one debugging insight from the order path. Rendered from the explorer above with the same data.

    07 / What I bring

    Correct data,
    stateful systems,
    honest evaluation.

    I built Spurs Edge Lab end to end: the data foundation and its audits, the feature calculators and their parity tests, the model bundle and its calibration, and the live loop that has to stay correct when a feed lags or an exchange fails mid-order.

    The habits that carried it: freeze what is proven and pin it by hash; compute every feature strictly before the moment it describes; evaluate on the same rows as the baseline and say what the sample unit is; treat an unanswered request as a question for the source of truth, not as a fact; and keep one writer for state that money depends on.

    That is the work I want to do on a team: data engineering where correctness is the product, systems that hold state under failure, and model evaluation that a skeptic can audit.

    London Paris · linkedin.com/in/londonparis