Est.

Anycast Network Latency Benchmarking Methodology

Anycast latency requires measuring which server answered each probe.

Senior Writer · · 11 min read
Cover illustration for “Anycast Network Latency Benchmarking Methodology”
Network Performance · September 30, 2026 · 11 min read · 2,383 words

Anycast latency benchmarking gets treated as a variant of unicast benchmarking, run the same tests, just against a different kind of IP. That assumption is wrong, and it's wrong in a way that quietly wrecks the resulting numbers. BGP routing, catchment behavior, and path inflation all sit between the probe and the number it reports, and a simple ping result does not reveal any of them. Getting this right means treating probe placement, vantage point spread, metric choice, and timing as part of the measurement itself. This work presents an Anycast Network Latency Benchmarking Methodology.

Anycast latency benchmarking as distinct from unicast benchmarking

A unicast benchmark is about as close to a controlled experiment as networking gets: one IP, one destination, and the round-trip time reflects the path to that single host.

Anycast throws that out. That means the latency number reported isn't just a measure of infrastructure quality. It's a measure of routing policy wearing infrastructure's clothes. Papers on anycast performance make the point bluntly: BGP is a policy-routing system. It has no concept of round-trip time. The site it picks can be geographically distant, reached through an indirect path, poorly peered, or unstable when load balancing kicks in, and BGP will never flag any of that as a problem, because from BGP's point of view, nothing is wrong.

This matters more once you notice how many things run on anycast today: root DNS servers, public DNS resolvers, CDN edge ingress points, and general application delivery all lean on it, but each has a different relationship to latency. A methodology built for measuring root DNS performance will not transfer cleanly to measuring CDN edge performance, because the economics of a slow path are completely different in each case, and the next section gets into how BGP does the damage.

Silent performance corruption in anycast measurements caused by BGP routing decisions

Anycast works by advertising the same prefix from multiple physical locations at once, and every intermediate autonomous system along the way runs its own BGP best-path selection, typically ranking by AS-path length first, then IGP metric and MED, then whatever tie-breaking logic that router's implementation happens to use. None of those criteria have anything to do with round-trip time. A shorter AS-path can just as easily route a user's traffic to a site on the wrong continent as to the nearest one, and BGP will treat that as a perfectly good routing decision.

And the failure mode is silent by design. The service stays reachable, uptime monitors stay green, no alert fires, because nothing at the application layer knows or cares that a routing change just happened. It appears only in latency percentiles, and only for someone actually watching the right region at the right time.

The triggers for this kind of drift are mundane, which is part of what makes them dangerous. A remote peering arrangement changes. A transit provider adjusts local preference on a handful of routes. A new participant joins an internet exchange and starts announcing a shorter path. Traffic that used to land in one city now lands on a different continent, and there's no application-layer signal anywhere that says so. Once that's understood as a routing-layer problem, the fix has to start at the measurement layer: knowing, for every single probe, which physical site actually answered it.

Catchment mapping as the foundation of any valid anycast measurement

An anycast catchment is the set of clients, resolvers, or address prefixes that end up routed to one particular site, and that set is shaped by local preference, AS-path length, peering arrangements, export filters, and traffic engineering decisions the operator usually doesn't control. The instinct to picture catchments as tidy geographic regions, something like a Voronoi diagram drawn over a map, is exactly wrong. A client physically close to one anycast instance can still get routed to another, because its upstream ISP or transit provider prefers a different path for reasons that have nothing to do with distance.

Measuring this requires more than a stopwatch. One established method uses RIPE Atlas: each probe runs a traceroute to the anycast address, the penultimate hop gets correlated against internal BGP session data to work out which physical point of presence actually received the packet, and the round-trip time from that final hop gives the real latency between the probe and that specific site. This exercise carries real, practical stakes. That's the kind of scale it takes to say anything credible about global catchment behavior.

Two broad approaches to this kind of monitoring are worth telling apart. Client-side vantage points, RIPE Atlas being the standard example, identify which server each individual client reaches and come with rich per-probe metadata attached. Other approaches trade some of that per-probe detail for broader client coverage, a different set of tradeoffs depending on what the study needs to answer. Either way, without catchment attribution a latency number is fundamentally ambiguous: it might reflect a clean, well-peered path to the nearby site, or it might reflect a bad BGP decision routing the same probe to a distant one. The two scenarios produce numbers that look identical on paper. Knowing which site answered is what turns a raw round-trip time into something actually interpretable, which is the whole reason the next layer of methodology, controlling where and when probes fire, matters at all. A landmark anycast study used more than 7,000 vantage points worldwide from RIPE Atlas to study the relationship between latency and anycast deployment across four Root DNS servers (C, F, K, and L).

Vantage point placement and the geographic distribution requirement

Latency performance shifts substantially by region, and a benchmark built from probes clustered in a handful of well-connected cities isn't measuring the network, it's measuring those cities' routing topology and calling it a global result. Regional variation is the signal. It's the signal.

Good probe placement covers every major region where the actual user base sits, not just the regions closest to wherever the benchmark author happens to be located. It puts probes at or near major internet exchange points, since that's where peering decisions actually get made. And critically, it puts probes in regions with thin IXP density, because BGP path inflation tends to occur there first. Skipping those regions makes the benchmark look clean purely because it never went looking in the places most likely to show a problem.

CDN anycast raises the stakes further. Keeping path inflation small requires active, ongoing engineering: peering relationships, route policy, catchment scope, all tuned deliberately, with measurement feedback closing the loop. Vantage points thin enough to miss catchment boundaries will miss the artifacts that matter most.

Root DNS and CDN anycast don't carry the same stakes here, either. Root DNS can tolerate a fair amount of path inflation without users ever noticing, because recursive resolver caches absorb most of that delay before it reaches anyone. CDN anycast has no equivalent cushion. A 50-millisecond routing mistake repeats across every object on every page for every user caught in that catchment, and CDN probe density needs to be denser than root DNS probe density. Evaluate CDN options against where the actual user base sits, because a general-purpose benchmark built on someone else's preferred probe set may say nothing about the latency a specific organization's own traffic will actually see. Solving where to measure still leaves open what kind of data to collect there; synthetic and real-user monitoring address that.

Synthetic monitoring and real-user monitoring as complementary inputs, not alternatives

Synthetic monitoring runs automated test requests from fixed global locations on a regular schedule, measuring time to first byte, DNS resolution time, TLS handshake time, and availability, all under conditions the tester controls. The appeal is consistency: numbers that repeat, that can be compared across weeks, that make regression detection possible. The catch is that the probe set is a known, fixed list of autonomous systems, and any routing decision, adversarial or accidental, that treats known probe networks better than it treats everyone else will make the synthetic numbers look better than the real experience.

Real user monitoring fixes that blind spot by collecting performance data from actual traffic across whatever devices, ISPs, and locations real users happen to bring. It captures the long tail that controlled testing never will: the mobile carrier with unusual routing, the ISP with a weird peering arrangement, conditions no test harness would think to reproduce on purpose. Its own limitation is volume. RUM needs meaningful traffic to produce a stable regional distribution, and a service that hasn't launched yet has no RUM data at all, by definition.

Neither one substitutes for the other. RUM validates what synthetic testing predicts, while synthetic testing catches regressions early, well before RUM has accumulated enough volume to say anything statistically firm, and a rigorous anycast benchmark needs both running at once. Anycast adds one more wrinkle on top: RUM sessions have to be tagged with which point of presence actually served each one, since catchment attribution resurfaces in a new form, or the resulting latency distribution is just several different sites' numbers mashed into a single histogram that describes none of them accurately. Getting the data collection right is only half the job. What that data actually proves depends entirely on which metric gets pulled out of it.

Choosing the right latency metrics

Time to first byte is the workhorse CDN latency metric, and for good reason. It captures the full chain, DNS resolution, TLS handshake, routing time, and origin processing, in one number that runs from request to the first byte of the response. A sudden jump in TTFB is often the earliest visible symptom of catchment drift, appearing before anyone notices a routing table has quietly changed underneath the service. But TTFB numbers from different sources aren't interchangeable. Region, test methodology, and whether the response came from cache all move the number independently, and any benchmark that doesn't state all three alongside the figure isn't really reporting a number, it's reporting a mystery.

Percentiles matter as much as the headline figure. The median, p50, describes what a typical user experiences, while p95 and p99 describe the tail: users stuck on bad paths, on mobile networks, or unlucky enough to be caught in a BGP-inflated route. Systems like AnyOpt, built specifically around anycast SLO design, make the point explicitly: the goal isn't just pushing p90 down, it's confirming that improvement doesn't come at the cost of p99 or by quietly overloading one site to relieve another. A policy change that improves median TTFB while making p99 worse can end up hurting a meaningful chunk of users even as the topline metric looks like a win. Anyone optimizing purely for the median is optimizing for the users who were never going to complain anyway.

Cache hit ratio deserves its own line item, since it's an efficiency metric, not a latency metric proper. Every cache miss forces an origin fetch, and origin fetches are slower by construction, so hit ratio needs tracking alongside TTFB specifically to separate edge-layer latency from routing-layer latency. Error rates by status code round out the picture: a spike in 4xx codes usually points to something client-side or misconfigured, while a spike in 5xx codes points to server-side or CDN failure, and either kind of spike helps isolate misconfiguration, origin failure, or regional disruption in a way latency numbers alone can't. None of these metrics substitutes for the others. They triangulate.

Temporal sampling strategy and the threat of time-varying routing

None of the above holds up if the measurement window is too short. BGP changes, peering outages, attacks, congestion, route leaks, and maintenance windows can all reshape catchments between one measurement session and the next, and a point-in-time benchmark might be capturing a best-case or worst-case routing state, with no way for the operator to know which.

Before rolling out any BGP policy change, the traffic shift it's expected to cause should be estimated in advance, with rollback criteria defined before the change goes live, not improvised afterward. After the change ships, measurement needs to run long enough for BGP to actually converge and for the probe network to settle into steady-state routing, not just the first few minutes when everything is still in flux. Anything shorter is measuring the transition.

Known disruptions, maintenance windows, DDoS events, route leak incidents, need to be identified and either excluded from the dataset or clearly flagged. Folding them in without any annotation just makes real signal look like random noise. And single-day or single-session benchmarks shouldn't be presented as if they represent typical performance, because a benchmark that only ever measures one moment in time cannot tell a fluke from a pattern.

The reproducibility gap in public anycast benchmarks and the valid methodology required to close it

The most detailed CDN anycast research tends to rely on proprietary operator traces, the kind of internal data that includes real catchment ground truth, the exact history of route policy changes, and actual point-of-presence identifiers. Public studies almost never get that level of access, which creates a real epistemological gap: published TTFB figures, whether they come from CDN operators directly or from benchmarks built on operator-supplied data, can't be independently checked, because the measurement design, the probe set, and the time window behind them usually aren't disclosed.

Current research on the topic names this outright as an open problem: public, reproducible CDN anycast evaluation. A trustworthy benchmark would combine RIPE Atlas measurements, browser-based tests, openly published edge identifiers, and voluntary disclosure from operators about regional scoping and peering changes. That's a specific, achievable list, not a vague call for more transparency.

Putting everything from the preceding sections together brings the requirements for a genuinely reproducible methodology into focus. A published probe list, tagged with AS number and geographic location, so vantage point diversity can actually be checked rather than taken on faith. Catchment attribution recorded for every single measurement, so it's clear which physical site produced which number. A measurement window with explicit timestamps and an event log documenting anything unusual that happened during it. Separate reporting broken out by cache status and by region, instead of a single global average that flattens routing heterogeneity into a meaningless mean. Confidence intervals and sensitivity analysis attached to reported percentiles, not bare point estimates presented as if they were the whole truth.

Sources

  1. Anycast Performance in Context A Comparative Study of Root DNS and CDN Latency
  2. Anycast Latency: How Many Sites Are Enough? - The ANT Lab
  3. Locating and Enumerating Anycast: a Comparison of Two Approaches
  4. BGP Anycast Best Practices & Configurations | Noction
  5. Internet Anycast: Performance, Problems and Potential Paper #24. 13 pages
  6. AnyOpt: Predicting and optimizing IP anycast performance | APNIC Blog
  7. An Empirical Evaluation of Longitudinal Anycast Catchment Stability | Springer Nature Link
  8. Broad and Load-Aware Anycast Mapping with Verfploeter Wouter B. de Vries

More in Network Performance