Historical Traffic Volumes

AADT Accuracy and Validation 2025

1. Executive summary

This document reports the accuracy of the TomTom Historical Traffic Volumes annual averages, AADT and AAWHT, for 2025. Sections 2 to 6 describe the product and how we validate it and are the same in every year’s document; Section 7 and the key findings below are specific to 2025.

We measured accuracy against independent ground-truth traffic counts from permanent loop detectors and similar counting infrastructure in seven countries (Belgium, Netherlands, New Zealand, Norway, Sweden, United Kingdom, and United States) for 2025. The per-country metrics in this document come from k-fold cross-validation, which tests the model on data it was not trained on (Section 5). For markets beyond these seven countries, Section 6 describes the complementary leave-one-country-out (LOCO) validation, whose pooled results for 2025 are reported in Section 7.8. A market is a country where the product is available.

Key findings:

  • The United Kingdom is the strongest performer in this assessment, with an all-roads AADT MAPE of 4.0% and a Very good SQV (0.90 or better) sustained in every volume band; its high-volume band reaches a MAPE of 2.3%.
  • Across the seven countries, all-roads AADT MAPE ranges from 4.0% (United Kingdom) to 10.0% (New Zealand); Norway pairs a high MAPE of 9.3% with a Very good SQV of 0.90, showing that errors remain well-controlled in absolute terms.
  • For markets without local counter data, the pooled leave-one-country-out results (Section 7.8) give an all-roads AADT MAPE of 13.1%, above every k-fold country in this assessment; the SQV of 0.77 is Acceptable overall while low-volume roads stay Good (0.87).
  • Percentage errors are consistently higher on low-volume roads in every country, but on such roads they correspond to small absolute vehicle counts; Section 8 provides guidance on interpreting these values.

2. Introduction

How many vehicles pass a road segment on a typical day, and at a typical hour of the week? TomTom Historical Traffic Volumes estimates exactly that, for each road segment. We apply machine learning to probe data: the position and speed reports that millions of connected vehicles send (Sekuła et al., 2018; Zhan et al., 2017). The annual averages cover most roads of the network.

Historical Traffic Volumes delivers three quantities: AADT, AAWHT and hourly volumes. The product introduction defines them. This document covers AADT and AAWHT; the hourly volumes are documented in the Hourly volumes section of this documentation.

We improve the model continuously, and an improved model can recompute past periods, so a published figure for a past period can improve after publication (Section 9).

Traffic volume data informs decisions with real consequences:

  • Where to open a new store
  • How to allocate infrastructure budgets
  • How to assess road safety risk
  • Whether the surrounding road network can support a proposed development
  • Whether a new lane, a closure or another change to a road altered traffic, comparing the period before with the period after

Every one of those decisions rests on a volume number, so the useful question is not whether the data is reliable but how far off it can be on the roads you care about. This document answers that in vehicles, in percentages and as quality scores.

Traditional counting methods give precise counts at a small number of fixed points (Federal Highway Administration, 2022). TomTom Historical Traffic Volumes estimates the volume across the road network of a covered market. An estimate carries uncertainty, so we measure that uncertainty with the metrics defined in Section 4. With those numbers, customers can decide where and how to use the data.

We validate against counters: permanent loop detectors and similar counting stations whose independent ground-truth counts played no part in model training. The accuracy metrics in this document measure the gap between our estimates and those reference counts; a smaller gap means a more accurate estimate. Where results fall short of the quality thresholds in Section 4.4, we say so and give the available context.

3. How the model works

Turning raw reports from connected vehicles into reliable, network-wide volume estimates takes a sequence of steps, and each step solves a distinct problem. Unlike traditional counting methods, the product needs no physical equipment at each measurement point. We collect probe observations (Section 3.1), estimate the penetration rate (Section 3.2) and convert probe observations into volume estimates (Section 3.3). Section 3.4 describes the two annual averages the model produces.

3.1 From connected vehicles to traffic observations

Our primary data source is floating-car data: GPS and telematics signals from connected vehicles (Herrera et al., 2010). Each signal is a probe observation; together, the signals form the probe data. Connected vehicles include:

  • In-vehicle navigation systems
  • Smartphones running navigation apps
  • Connected commercial vehicles

The signals arrive passively and continuously from millions of devices worldwide. For each road segment, they form a continuous stream of speed and passage observations.

Probe observations are not traffic volumes, though. Only a fraction of the vehicles on a given road are connected and contributing data. We call this fraction the penetration rate. It varies by road type, geography and time of day, so converting probe observations into total volume estimates means accounting for that variation.

3.2 Estimating the penetration rate without counters

A probe count becomes a traffic volume only once we know the penetration rate: the share of vehicles on that road that report probe data. Permanent counters measure it directly, but they exist on a small fraction of roads, and in many markets on none. A method that needs counters everywhere cannot scale. Ours needs them only once, to learn.

The key is congestion. When a road operates at or near capacity (Transportation Research Board, 2022), physics constrains it: the relationship between the speed vehicles drive and the number of vehicles the road carries becomes tight and predictable, as the fundamental diagram of traffic flow describes (Greenshields, 1935; Treiber & Kesting, 2013). From probe speeds, read together with the road’s attributes, we can then estimate how many vehicles the road carries. We train a capacity model on roads with counters to learn this speed-to-flow relationship, together with map attributes such as road class, urban or rural context and lane configuration. Because the same congestion physics applies wherever a road runs near capacity (Kerner, 2004), the model transfers to roads that have never had a counter.

On any congested road we then hold two independent numbers: the total flow the capacity model estimates from speeds, and the probe flow we count directly. Their ratio is the penetration rate. Wherever probes meet congestion, we obtain a penetration-rate estimate, far beyond counter locations (Eisinga & Lorkowski, 2025). Individual estimates are noisy, so we aggregate them by region and road type into summaries that resist noise and fill the remaining gaps from similar surroundings. The result is a consistent picture of probe representativeness across countries and road classes. Roads that never congest inherit the estimate of their region and road type.

3.3 From penetration rate to traffic volume

The regional penetration picture tells us roughly what share of traffic the probes capture around a road. The volume model turns that into an estimate for the specific road. We train it on real ground truth: permanent counters, where they exist. From those counts it learns how the regional penetration rate, the observed probe data and the road’s own attributes combine into the volume of one road, in effect refining the regional penetration rate down to each segment. Once trained, it needs no counters. It runs wherever probe data and a map exist, which is what makes the product scalable to new regions, and why Section 6 tests it on countries it has never seen.

A motorway and a residential street sit at opposite ends of a wide range: among the counters the model learns from, the quietest carry fewer than 500 vehicles a day and the busiest more than 125,000. The relationship between a road’s attributes, its probe activity and its traffic load is not the same at the two ends. The estimate has to hold across that whole range, not only for the average road. So we train the model on counters across the full range of volumes, and we judge it on the same range. Section 7 reports its accuracy separately for high-, medium- and low-volume roads in every validated country, in the volume categories Section 5 defines. A reader can see where the estimates are strong and where they are weaker rather than take our word for it.

The model is tested only against counters it did not learn from. In k-fold cross-validation (Section 5), each counter is scored by a training round that did not include it. In leave-one-country-out validation (Section 6), a whole country is removed from training and scored as if it were a new market. Expect lower accuracy in a country the model has never seen than in one whose counters it trained on. Section 7.8 shows the difference for 2025; for markets without counters, it is the accuracy basis.

The two ends of the range err in different ways, and the metrics in Section 4 are chosen to show both. On a quiet road a handful of vehicles is a large share of the count, so percentage errors tend to be largest on low-volume roads even where the error in vehicles is small (Das & Tsapakis, 2020; Section 4.1). On a busy road the same percentage means many more vehicles. The Scalable Quality Value (Section 4.4) scales the error to the size of the count, so the two ends can be compared on one score. The median percentage error (Section 4.2) shows whether the estimates in a category lean high or low.

The volume model produces the two annual averages, AADT and AAWHT (Section 3.4).

3.4 What we publish: AADT and AAWHT

AADT, annual average daily traffic, is one value per road segment and year: the number of vehicles that pass the segment on an average day of that year. It is the standard yardstick of transport planning and the basis of network-wide totals such as vehicle kilometers traveled.

AAWHT, annual average week-hour traffic, is 168 values per segment and year in vehicles per hour, one for each hour of each day of the week: the typical Tuesday between 8am and 9am, for example. It describes the weekly rhythm of a road: the morning and evening peaks, the quiet of the night, the difference between a weekday and a Sunday.

The two are one profile at two resolutions. Summing the 24 AAWHT values of a day gives the typical volume of that day of the week, and the mean over the seven days gives the AADT. The hourly volumes documented in the Hourly volumes section are produced by starting from the typical value for that road, that day of the week and that hour, and adjusting it by how busy the road was in the period being estimated, as observed in probe data.

4. How we measure accuracy

One number cannot show both the typical error and its spread, so we report a set of complementary metrics. Each one highlights a different aspect of performance, and every one is measured against independent ground-truth counts. For AADT, each counter contributes one value per year; for AAWHT, one value for each of its 168 week-hours (Section 5 explains both).

MetricWhat it measures
MAPE (Mean Absolute Percentage Error)The average percentage by which estimates differ from actual counts, regardless of direction. The primary summary measure. Lower is better.
Median Percentage ErrorThe middle value of all signed errors. Indicates systematic bias: positive = tendency to over-predict; negative = tendency to under-predict. Values close to zero are ideal.
68th and 95th Percentile APEThe spread of errors across road segments. The 68th percentile covers roughly one standard deviation; the 95th captures the tail of the distribution where the model is most challenged.
SQV (Scalable Quality Value, 15th Pct)A bounded quality score (0 to 1) designed to be consistent across roads of all volumes. Reported at the 15th percentile: at least 85% of segments perform better than this value.

4.1 Mean absolute percentage error (MAPE)

MAPE measures the average size of the prediction error relative to the observed count, as a percentage. A MAPE of 10% means that estimates differ from actual counts by 10% on average. MAPE treats errors of all sizes equally, and it is a commonly reported accuracy measure in traffic estimation.

On very low-volume roads, a high percentage error can mean a small difference in vehicles (Hyndman & Koehler, 2006). A road with 300 vehicles per day and a MAPE of 15% has a typical absolute error of roughly 45 vehicles. For most planning and analytical purposes, a difference of that size is negligible. So read MAPE values for low-volume roads alongside the absolute volumes that matter for your use case.

4.2 Median percentage error

The median percentage error is the middle value of all signed errors: positive when our estimate exceeds the actual count, negative when it falls short. A value close to zero shows that the model has no strong tendency to over- or under-count. Low bias matters for applications such as aggregated network analysis or vehicle kilometers traveled (VKT) calculations, because systematic errors add up across many road segments.

4.3 68th and 95th percentile absolute percentage error

These percentiles describe how the errors spread across road segments. The 68th percentile corresponds roughly to one standard deviation in a normal distribution: about 68% of segments have an error at or below it. The 95th percentile captures the upper range, the error level of the most challenging segments. Together with MAPE, the percentiles show how the error is distributed, not only its average.

4.4 Scalable Quality Value (SQV)

A percentage error looks large on a quiet road and an absolute error looks large on a busy one. The Scalable Quality Value (Friedrich et al., 2019) handles both cases. It generalizes the GEH statistic used in transport model validation (Department for Transport, 2026). It is a bounded, scale-independent quality metric with a score between 0 and 1, where 1 is a perfect match. It measures the error against a yardstick that grows with the square root of the count, so it tolerates a larger percentage error on quiet roads, where a few vehicles make a large percentage, and a larger absolute error on busy roads. It then maps the result to a bounded score. SQV is therefore consistent across the full range of traffic volumes. The formula is

SQV = 1 / (1 + sqrt((M − C)² / (f × C)))

where M is the modeled value, C is the observed count and f is a scaling factor set by the order of magnitude of the quantity: 10,000 for daily volumes such as AADT and 1,000 for hourly volumes such as AAWHT. An error of zero gives a score of 1; the larger the error relative to the count, the lower the score. For example, an AADT estimate of 11,000 vehicles against a count of 10,000 scores about 0.91, and so does an AAWHT estimate of 1,100 vehicles in an hour against a count of 1,000. We report SQV as its 15th percentile across all segments in each group: at least 85% of the road segments in the group perform better than the stated value. The quality thresholds below come from Friedrich et al. (2019):

SQVAssessmentGuidance for use
≥ 0.90Very goodHigh confidence in segment-level comparisons. Suitable for precision analytics and granular planning.
≥ 0.85GoodSuitable for cross-segment analysis and most planning applications.
≥ 0.80FairSuitable for network-level planning and transport modeling. Validate individual segments where precision matters.
≥ 0.75AcceptableSuitable for aggregate and indicative use. Validate against local count data before relying on individual segments.
Below 0.75InsufficientUse with caution. Treat results as indicative and validate against local count data where possible.

5. Validation approach: K-fold cross-validation

The per-country accuracy metrics in this document come from k-fold cross-validation. The method tests the model on data it was not trained on, which gives a realistic measure of real-world performance.

We divide the counter-equipped road segments used in training into five groups, called folds (Roberts et al., 2017). In each round, we hold one fold back entirely, train the model on the remaining folds and evaluate it against the held-back segments. This repeats until every fold has served as the test set. We then aggregate the accuracy figures across all rounds, so every validation counter contributes to the reported results without ever training the model it is evaluated against.

The same rounds validate both annual averages. For AADT, each held-back counter contributes one value per year: its AADT estimate against the AADT derived from its counts. For AAWHT, it contributes one value per week-hour: the estimate for that hour of the week against the counter’s average for it over the year, so the AAWHT metrics are computed across segment-hours. Counters without usable hourly data are left out of the AAWHT figures, so the counter numbers in the AADT and AAWHT tables of Section 7 can differ slightly.

K-fold cross-validation is standard practice in machine learning (Hastie et al., 2009): it prevents overfitting to one test set, uses all available ground-truth data and gives a reliable estimate of how well the model generalizes.

We report results for three volume categories. The volume category is a different grouping from road class, the functional class of a road in the map, such as motorway, major road or local street:

  • High-volume roads: 55,000 or more vehicles per day (AADT)
  • Medium-volume roads: 5,000–54,999 vehicles per day
  • Low-volume roads: fewer than 5,000 vehicles per day

An All roads row gives the aggregated metrics across all counters, regardless of volume category. Section 7 reports the AADT and AAWHT results for each country and, in Section 7.8, pooled for countries held out from training.

K-fold cross-validation needs ground-truth counting data for every country in the validation set. For markets without counter data, Section 6 describes the complementary leave-one-country-out (LOCO) approach we use to estimate accuracy.

6. Leave-one-country-out (LOCO) validation

TomTom Historical Traffic Volumes is available in many markets with no permanent counting infrastructure. K-fold cross-validation (Section 5) needs ground-truth data, so it cannot be applied directly there. Leave-one-country-out (LOCO) validation fills that gap: in countries that do have counters, it measures how the model performs in a country it has never seen, and uses that result as the accuracy estimate for a market without counters (Roberts et al., 2017). This transfer assumes that the held-out countries resemble the unseen market in road network structure and probe coverage.

6.1 How LOCO validation works

In LOCO validation, we exclude one country entirely from model training. We train the model on all remaining countries and then evaluate it against the held-out country with the counter data available there. This repeats for each country in turn, so every country serves once as an unseen test market. We report the results jointly for the held-out countries in a year rather than per country, so they describe the accuracy to expect in a market the model has never seen, not the accuracy of any one country.

LOCO simulates deployment in a market where the model has never seen local data. Its results therefore estimate performance in new markets, where the model relies entirely on what it has learned from other countries.

6.2 Why LOCO matters for customers

If you use TomTom Historical Traffic Volumes in a market beyond the seven countries in Section 7, the LOCO results in Section 7.8 are your primary quality basis. The model learns traffic patterns that carry across countries: road network structure, speed-flow relationships, penetration-rate dynamics. Those learned patterns are what the model carries into a market it has never seen. The LOCO results show how much of its accuracy travels with them.

6.3 Where the LOCO results are reported

The LOCO results for 2025 are reported in Section 7.8, next to the per-country k-fold results and in the same tables and metrics, pooled across the countries held out from training. Because the held-out countries are evaluated jointly, the figures describe the accuracy to expect in a market the model has never seen rather than any specific market. They extend the k-fold results in Section 7 to every market where the product is available. The pooled AAWHT figures are computed in the hourly evaluation and cover a larger set of held-out counters than the pooled AADT figures, which is why their counter numbers differ.

7. Results (2025)

The tables below give the AADT and AAWHT accuracy results for each of the seven countries in the 2025 k-fold validation, followed by the pooled leave-one-country-out (LOCO) results for countries held out from training (Section 6). Section 4.4 defines the quality tiers for the SQV (15th Pct) values.

7.1 Belgium

AADT

Road CategoryCountersMAPE (%)68th Pct APE (%)95th Pct APE (%)Median PE (%)SQV (15th Pct)
All roads1,7307.18.120.1-0.40.88
High (55,000+)1284.14.514.2-0.90.85
Medium (5,000–54,999)1,1006.57.517.5-0.70.87
Low (under 5,000)5029.211.124.40.80.92
AADT k-fold validation plot for Belgium (2025)
Figure 7.1: AADT k-fold validation results for Belgium (2025).

AAWHT

Road CategoryCountersMAPE (%)SQV (15th Pct)
All roads1,6228.50.90
High (55,000+)1235.20.86
Medium (5,000–54,999)1,0257.90.89
Low (under 5,000)47411.40.93

Belgium reaches the Good tier nationally (SQV 0.88) and holds it in every volume band, with a median percentage error below 0.5% and a tight, unbiased scatter in the AADT plot.

7.2 Netherlands

AADT

Road CategoryCountersMAPE (%)68th Pct APE (%)95th Pct APE (%)Median PE (%)SQV (15th Pct)
All roads10,4076.26.619.6-0.70.89
High (55,000+)1,0283.53.611.5-0.60.86
Medium (5,000–54,999)7,2805.76.017.9-0.80.88
Low (under 5,000)2,0999.511.125.8-0.80.92
AADT k-fold validation plot for Netherlands (2025)
Figure 7.2: AADT k-fold validation results for Netherlands (2025).

AAWHT

Road CategoryCountersMAPE (%)SQV (15th Pct)
All roads9,6617.20.90
High (55,000+)1,0184.30.88
Medium (5,000–54,999)6,6896.60.90
Low (under 5,000)1,95411.30.93

The Netherlands combines the largest counter sample in the document (10,407 loops) with consistently strong results, holding an SQV of at least 0.86 in every band and 0.89 nationally. A slight negative median percentage error recurs across all bands and in the hourly breakdown, indicating a mild systematic underestimation, though it remains below 1% at the AADT level.

7.3 New Zealand

AADT

Road CategoryCountersMAPE (%)68th Pct APE (%)95th Pct APE (%)Median PE (%)SQV (15th Pct)
All roads52710.011.625.7-2.00.87
High (55,000+)215.86.915.6-2.60.82
Medium (5,000–54,999)3089.210.923.9-2.10.86
Low (under 5,000)19811.614.532.3-1.20.90
AADT k-fold validation plot for New Zealand (2025)
Figure 7.3: AADT k-fold validation results for New Zealand (2025).

AAWHT

Road CategoryCountersMAPE (%)SQV (15th Pct)
All roads51612.70.88
High (55,000+)207.30.81
Medium (5,000–54,999)30011.80.87
Low (under 5,000)19614.60.91

New Zealand carries the highest all-roads AADT MAPE in the document (10.0%) yet still reaches the Good tier (SQV 0.87), with its low band at Very good (0.90). A mild negative median percentage error (-2.0% at the all-roads level) recurs in every volume band, the most pronounced under-prediction tendency of all countries in this assessment, although the AADT scatter remains tightly grouped around the diagonal apart from a single strong mid-volume outlier. The high band’s Fair SQV (0.82) rests on only 21 counters, too few for a stable estimate.

7.4 Norway

AADT

Road CategoryCountersMAPE (%)68th Pct APE (%)95th Pct APE (%)Median PE (%)SQV (15th Pct)
All roads2,5939.311.024.8-1.00.90
High (55,000+)33.76.36.3-2.00.87
Medium (5,000–54,999)1,1618.39.722.4-1.40.87
Low (under 5,000)1,42910.212.126.1-0.30.92
AADT k-fold validation plot for Norway (2025)
Figure 7.4: AADT k-fold validation results for Norway (2025).

AAWHT

Road CategoryCountersMAPE (%)SQV (15th Pct)
All roads2,41411.70.91
High (55,000+)35.50.86
Medium (5,000–54,999)1,12410.30.89
Low (under 5,000)1,28713.00.92

Norway posts one of the highest national AADT MAPE values in this assessment (9.3%) yet still achieves a Very good SQV of 0.90, a pairing that indicates the percentage errors are inflated by the low absolute volumes that dominate the Norwegian network rather than by uncontrolled absolute errors. The high band contains only 3 counters, so its rows rest on too small a sample to be conclusive.

7.5 Sweden

AADT

Road CategoryCountersMAPE (%)68th Pct APE (%)95th Pct APE (%)Median PE (%)SQV (15th Pct)
All roads6187.88.824.3-0.40.85
High (55,000+)244.85.515.3-1.80.80
Medium (5,000–54,999)5317.38.221.5-0.20.85
Low (under 5,000)6313.014.833.8-0.70.90
AADT k-fold validation plot for Sweden (2025)
Figure 7.5: AADT k-fold validation results for Sweden (2025).

AAWHT

Road CategoryCountersMAPE (%)SQV (15th Pct)
All roads5448.60.88
High (55,000+)246.00.84
Medium (5,000–54,999)4838.40.87
Low (under 5,000)3714.40.91

Sweden’s national SQV of 0.85 sits exactly on the Good threshold, and its high band lands on the Fair boundary at 0.80. With 618 counters, among the smallest samples in the document, and only 24 high-band and 63 low-band loops among them, the band-level results should be read as noisy rather than as stable estimates. The AADT scatter shows no material bias, with a median percentage error below 0.5% in magnitude.

7.6 United Kingdom

AADT

Road CategoryCountersMAPE (%)68th Pct APE (%)95th Pct APE (%)Median PE (%)SQV (15th Pct)
All roads6,4654.04.312.40.00.90
High (55,000+)1,9612.32.66.40.10.90
Medium (5,000–54,999)4,1564.44.913.2-0.00.90
Low (under 5,000)3488.09.121.70.70.92
AADT k-fold validation plot for United Kingdom (2025)
Figure 7.6: AADT k-fold validation results for United Kingdom (2025).

AAWHT

Road CategoryCountersMAPE (%)SQV (15th Pct)
All roads6,3264.60.91
High (55,000+)1,9352.90.91
Medium (5,000–54,999)4,0495.20.91
Low (under 5,000)3429.90.93

The United Kingdom delivers the strongest results in this assessment: the lowest national AADT MAPE (4.0%), an essentially unbiased median error, and a Very good SQV of 0.90 or better sustained across all three volume bands. The high band is particularly accurate, with a MAPE of 2.3% and a 95th percentile APE of 6.4%.

7.7 United States

AADT

Road CategoryCountersMAPE (%)68th Pct APE (%)95th Pct APE (%)Median PE (%)SQV (15th Pct)
All roads7,8047.68.622.60.20.83
High (55,000+)2,2866.17.217.31.10.75
Medium (5,000–54,999)4,1057.78.623.3-0.20.85
Low (under 5,000)1,4139.811.227.3-0.40.92
AADT k-fold validation plot for United States (2025)
Figure 7.7: AADT k-fold validation results for United States (2025).

AAWHT

Road CategoryCountersMAPE (%)SQV (15th Pct)
All roads7,1218.70.86
High (55,000+)2,1336.70.80
Medium (5,000–54,999)3,6928.90.88
Low (under 5,000)1,29612.10.93

The United States records the lowest national SQV in the document (0.83, Fair), with the high band at 0.75 on the Acceptable boundary while the low and medium bands stay at or above 0.85. Median percentage errors are small throughout, and the AADT scatter remains well centred on the diagonal.

7.8 Unseen countries (LOCO)

AADT

Road CategoryCountersMAPE (%)68th Pct APE (%)95th Pct APE (%)Median PE (%)SQV (15th Pct)
All roads30,23513.116.332.11.20.77
High (55,000+)4,14111.314.526.35.80.66
Medium (5,000–54,999)19,23512.114.829.92.90.77
Low (under 5,000)6,85917.121.637.4-10.80.87

AAWHT

Road CategoryCountersMAPE (%)SQV (15th Pct)
All roads59,04217.10.79
High (55,000+)5,36012.90.70
Medium (5,000–54,999)41,29816.80.79
Low (under 5,000)12,38420.90.88

Pooled across the countries withheld from training, the all-roads AADT MAPE of 13.1% lies above that of every country in this assessment (4.0% to 10.0%), as expected without local training data, and the median percentage error of 1.2% indicates a slight tendency to over-predict that grows to 5.8% in the high band. The SQV is Acceptable on all roads (0.77) and in the medium band (0.77) and Insufficient in the high band (0.66); the low band remains Good (0.87) despite the highest AADT band MAPE (17.1%) and a median percentage error of -10.8%, a pattern confined to low-volume roads where absolute errors stay small.

8. What the results mean for your use case

8.1 For business decision-makers and analysts

Traffic volume data informs strategic decisions (Section 2). For most analytical applications, the question comes down to one thing: is the error range acceptable for the decision at hand?

A practical guide: a MAPE of 10% on a road carrying 20,000 vehicles per day means that estimates differ from the true figure by about 2,000 vehicles on average, and the 68th and 95th percentile columns in Section 7 show how wide the error gets on a single road. For retail site selection, insurance risk modeling or transport infrastructure planning, an error of this size is usually acceptable. For applications that need precise capacity calculations, such as junction design or traffic signal optimization, we recommend adding local count data to the volume estimates where it is available and practicable. For comparisons across time, such as a before-and-after study, Section 9 explains how to read a difference between two periods against the stated accuracy, and how model improvements reach past periods.

Low-volume roads (under 5,000 vehicles per day) tend to show higher MAPE values (Das & Tsapakis, 2020). On roads with very low daily volumes, a higher percentage error still means a small number of vehicles — the worked example in Section 4.1 (a MAPE of 15% on a road carrying 300 vehicles per day, about 45 vehicles) shows the scale. For use cases that depend on individual low-volume rural roads, treat the estimates as indicative and validate them against available count data where precision matters.

8.2 For data scientists and transport modelers

The SQV 15th percentile is a conservative quality indicator (Section 4.4 explains how to read it). When you integrate TomTom Historical Traffic Volumes into a transport model, the SQV shows which road categories you can use with confidence and where extra validation against local counts is advisable.

For road categories with SQV values at or above 0.80 (Fair or better), the data is suitable for transport models and analytical workflows that need segment-level accuracy. For categories between 0.75 and 0.79 (Acceptable), use the data for aggregate and indicative purposes and validate individual segments against local counts. For categories below 0.75 (Insufficient), treat the data as indicative and apply extra quality filters or local calibration where precision is required.

The AAWHT results carry the same guidance for the weekly profile. Where a road category reaches Fair or better on AAWHT, the profile is reliable enough for time-of-day analysis, such as peak-hour shares or the split between weekdays and weekends; where it does not, aggregate the profile to the day before using it. AAWHT describes the typical hour of the week; the accuracy of an estimate for a specific date and hour is documented in the Hourly volumes section.

Where the median percentage error (bias) of a category is close to zero, aggregate measures such as total vehicle kilometers traveled across a network, or the average AADT for a road class, are reliable even where individual segments carry errors. Where a category shows a negative median percentage error (a tendency to under-predict), account for it in applications where absolute volume totals matter.

9. Updates, versions and comparability

9.1 Estimates are updated

TomTom Historical Traffic Volumes is a modeled product. We improve the model continuously, and an improved model can recompute the periods we have already published. A figure for a past period can therefore improve after publication: the recomputed estimate reflects a better model, while the traffic that occurred is unchanged. Regenerating the history in this way keeps a series internally consistent, because every period in it comes from the same model.

9.2 Comparing periods

Every estimate in this product is a measurement with a stated accuracy: the quality figures in Section 7 give it for each country, and pooled for unseen markets, by road category. A comparison between two periods, in a before-and-after study, a year-over-year trend or network monitoring, is a comparison between two such measurements.

AADT and AAWHT aggregate a year of observations, so short-term variation in the probe data largely averages out, and they are the natural basis for year-over-year comparison.

Our guidance: read a difference between two periods against the accuracy figures published for both periods. When we regenerate the history, refresh both periods, so that they come from the same model.

We will extend this guidance as the product evolves.

10. Summary

This document reports the accuracy of the TomTom Historical Traffic Volumes annual averages, AADT and AAWHT, for 2025, measured against independent ground-truth counts in seven countries by k-fold cross-validation, and for markets beyond them by leave-one-country-out (LOCO) validation.

In the 2025 assessment the United Kingdom leads with an all-roads AADT MAPE of 4.0% and a Very good SQV in every volume band, and six of the seven countries keep their all-roads MAPE below 10%. Median percentage errors remain small throughout, so aggregate network-level measures can be used with confidence. For countries held out from training, the pooled LOCO results give an all-roads AADT MAPE of 13.1% (SQV 0.77), the accuracy to expect where the model has no local training data.

New Zealand pairs the highest national MAPE (10.0%) with a Good SQV of 0.87, indicating errors that remain well-controlled in absolute terms; readers using segment-level data should consult the band-level tables and the guidance in Section 8.2.

On low-volume roads, MAPE values tend to be higher in percentage terms; as Section 4.1 explains, the difference in vehicles on these roads is typically small.

For markets beyond the seven countries in this k-fold validation, the pooled leave-one-country-out (LOCO) results in Section 7.8 provide the accuracy basis (Section 6).

We publish the results for every validated country, including where they fall short of the quality thresholds in Section 4.4. With this document, customers have what they need to use TomTom Historical Traffic Volumes with a clear view of both its accuracy and its limits.

11. References

Das, S., & Tsapakis, I. (2020). Interpretable machine learning approach in estimating traffic volume on low-volume roadways. International Journal of Transportation Science and Technology, 9(1), 76–88. https://doi.org/10.1016/j.ijtst.2019.09.004

Department for Transport. (2026). TAG Unit M3.1: Highway Assignment Modelling (May 2026). Transport Analysis Guidance. https://www.gov.uk/government/publications/webtag-tag-unit-m3-1-highway-assignment-modelling

Eisinga, K., & Lorkowski, S. (2025). Network-Wide Traffic Volume Estimation Based on Probe Vehicle Data. Transportation Research Record, 2679(4), 264–277. https://doi.org/10.1177/03611981241289408

Federal Highway Administration. (2022). Traffic Monitoring Guide (Version 1.0, December 2022). U.S. Department of Transportation. https://www.fhwa.dot.gov/policyinformation/tmguide/tmg_2022/

Friedrich, M., Pestel, E., Schiller, C., & Simon, R. (2019). Scalable GEH: A Quality Measure for Comparing Observed and Modeled Single Values in a Travel Demand Model Validation. Transportation Research Record, 2673(4), 722–732. https://doi.org/10.1177/0361198119838849

Greenshields, B. D. (1935). A study of traffic capacity. Highway Research Board Proceedings, 14, 448–477. https://onlinepubs.trb.org/Onlinepubs/hrbproceedings/14/14P1-023.pdf

Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction (2nd ed.). Springer. https://doi.org/10.1007/978-0-387-84858-7

Herrera, J. C., Work, D. B., Herring, R., Ban, X., Jacobson, Q., & Bayen, A. M. (2010). Evaluation of traffic data obtained via GPS-enabled mobile phones: The Mobile Century field experiment. Transportation Research Part C: Emerging Technologies, 18(4), 568–583. https://doi.org/10.1016/j.trc.2009.10.006

Hyndman, R. J., & Koehler, A. B. (2006). Another look at measures of forecast accuracy. International Journal of Forecasting, 22(4), 679–688. https://doi.org/10.1016/j.ijforecast.2006.03.001

Kerner, B. S. (2004). The Physics of Traffic. Springer. https://doi.org/10.1007/978-3-540-40986-1

Roberts, D. R., Bahn, V., Ciuti, S., Boyce, M. S., Elith, J., Guillera-Arroita, G., Hauenstein, S., Lahoz-Monfort, J. J., Schröder, B., Thuiller, W., Warton, D. I., Wintle, B. A., Hartig, F., & Dormann, C. F. (2017). Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8), 913–929. https://doi.org/10.1111/ecog.02881

Sekuła, P., Marković, N., Vander Laan, Z., & Farokhi Sadabadi, K. (2018). Estimating historical hourly traffic volumes via machine learning and vehicle probe data: A Maryland case study. Transportation Research Part C: Emerging Technologies, 97, 147–158. https://doi.org/10.1016/j.trc.2018.10.012

Transportation Research Board. (2022). Highway Capacity Manual 7th Edition: A Guide for Multimodal Mobility Analysis. National Academies Press. https://doi.org/10.17226/26432

Treiber, M., & Kesting, A. (2013). Traffic Flow Dynamics: Data, Models and Simulation. Springer. https://doi.org/10.1007/978-3-642-32460-4

Zhan, X., Zheng, Y., Yi, X., & Ukkusuri, S. V. (2017). Citywide Traffic Volume Estimation Using Trajectory Data. IEEE Transactions on Knowledge and Data Engineering, 29(2), 272–285. https://doi.org/10.1109/TKDE.2016.2621104