Moscow Transport Hackathon · 2026

See the delay
15 minutes
before it happens

The system listens to bus telematics over NDTP, matches every ping against the timetable and warns the dispatcher about a delay 10–15 minutes before the stop — with the cause and a recommendation for what to do.

on time at risk will be late 13 real routes from the dataset
© OpenStreetMap contributors
—
buses sending NDTP right now
—
telemetry packets received
—
active alerts
—
forecasts inside the 10–15 min window
—
ms per forecast (p50)

Numbers come straight from api.mowtransit.ru/health

01 · Horizon

Every forecast looks at the (T+10, T+15] min window

Every 5 seconds, for every bus, we take the first scheduled stop that falls inside the window — and ML forecasts the delay at that stop. As time moves, the target rolls forward and the forecast updates. No after-the-fact forecasts: an alert is only raised while the stop is still 10–15 minutes away.

T = 08:00:00
forecast window
T
T+5+10+15+20 min
Forecast target—
Scheduled—
Lead time—
Predicted delay—
100%of forecasts on the live stream land in the window — the service checks this itself and reports horizon_ok_share in /health
lead_minin every ML response: how many minutes before the scheduled stop the forecast was made
3 of 3verified alerts in a run on the real day turned out to be real delays (253–278 s)

02 · Incident card

Not just “it’ll be late”, but why and what to do

An alert is raised when the model puts the probability of being more than 2 minutes late at ≥ 0.7. The card shows the forecast with a range, the route segment, each cause’s contribution in seconds and a recommendation for the dispatcher. When the bus arrives, the alert is checked against the actual time.

  • The forecast range is calibrated: the actual delay falls inside it 80% of the time
  • Causes come from approximate SHAP, turned into plain language
  • The recommendation works out the speed needed to get back on schedule

The card is rendered straight from the live system, which talks to Moscow dispatchers — so its text is in Russian.

sample

03 · How it works

Stream → forecast → dashboard, end to end

Three independent services in Docker Compose. The backend decodes binary NDTP itself (no off-the-shelf parser); ML is a separate service that can be retrained and reloaded without stopping the backend.

📡

NDTP terminals

The organisers’ emulator replaying real tracks. Binary TCP to ndtp.mowtransit.ru:9201

13 buses live
⚙️

Backend · ingest

Our own NPL/NPH codec, CRC-16, handshake, G6CellNav00 cells. asyncio + uvloop

0.3 ms p50
🧭

Backend · state

Timetable matching, GPS-detected arrivals, current deviation, segment speed, dwell time

±3 s median
🧠

ML · CatBoost ×5

An ensemble on 56 features + range and probability models, causes, what-if

44 ms /predict
🗺️

Dashboard

MapLibre, risk traffic light, alert feed, what-if. WebSocket, updates every second

All history — telemetry, arrivals, every forecast and alert — is written to Postgres in the background (write-behind, COPY), so the database can never slow down the live path.

Python 3.12CatBoostFastAPIasyncio + uvloopONNX PostgresDocker ComposenginxMapLibreWebSocketSphinxSwagger

04 · Accuracy

Off by 40 seconds where the baseline is off by 93

MAE of the delay forecast on a held-out set. The hackathon metric hits its ceiling: score = 1.0, both locally and on the platform.

Predict “no delay” (0 s)
103 s
Organisers’ baseline (cur_dev_s)
93 s
Our CatBoost ensemble
40 s
1.0
hackathon score
6 of 6 points for accuracy
38.8 s

Shaped like validate

5 folds in 30-minute blocks: nearby moments of the same buses. Baseline: 88.4 s.

78.9 s

Honestly: an unseen bus

The whole bus is held out — the model has never seen it. Still beats the baseline (88.4 s).

0.97

AUC for “will it be late”

The late probability is calibrated: predicted ~0.89 → actually 0.95. It drives the risk traffic light.

Online = batch: the live service matches submission.csv on 151 of 151 points, maximum difference 0.00 s. The same function computes the features in both.

05 · Reliability

Break something. The service won’t go down

Pick a failure and see how the system degrades and recovers. All of it was tested on live containers.

— —

—


      
16,000terminals at once, not a single packet lost
25–30kNDTP packets per second on one core
7.6 µsfor the codec to decode one packet
2.2 sML cold start to a ready /health
20 msper point in batched /predict/batch
9.5 KBof memory per TCP connection

06 · Beyond the brief

Extra features

⇄

What-if

Release a reserve bus or a detour — forecasts for the route before and after, applied to the simulation.

⚡

ONNX

Model export: fp32 runs 6× faster than native CatBoost.

∑

Ensemble

5 models on different seeds + separate quantile and early / on-time / late class models.

±

Conformal ranges

Not one number but a range with a guaranteed 80% coverage.

⌖

Route matching

GPS-detected arrivals: 23,664 arrivals match the training pipeline exactly.

⟲

Hot reload

POST /reload picks up a retrained model without restarting the container.

07 · Team

Who built it