Abstract
Large R&E networks are full of partial signals: traceroutes show path changes, OWAMP tests show delay and loss, iperf3 shows throughput - each running at its own cadence, inspected with its own tool. When a data transfer slows down, the real question is simple but hard to answer fast: did the path change, and does that explain it?
We built a routing-aware incident triage tool on top of perfSONAR data from the WLCG/OSG mesh (170+ sites, ~4M tests/day). For each monitored corridor (source-destination pair), it detects routing shifts using graph-based path embeddings scored against that corridor's own recent baseline, then aligns them with delay, loss, and throughput evidence. Finally, each pair is assigned to an operator-facing class - confirmed impact, routing change with no measurable effect, tests interrupted, or needs review - so operators get a prioritized list.
We will walk through how the routing representation and correlation stage work, then through real incidents: a router migration, a hardware failure that rerouted traffic through commercial transit, and a submarine cable fault the tool caught live, hours before the NOC ticket was filed. The method is perfSONAR-native: no BGP feeds, no additional probing, no labelled training data, and directly reusable on any comparable measurement mesh.
Recording
Video will be added soon.
Speaker
Petya Vasileva
Rate this talk
Rating will open: Monday, 26 October 2026 09:00 (+0200).