Date of Graduation

7-2026

Document Type

Thesis

Degree Name

Master of Science in Civil Engineering (MSCE)

Degree Level

Graduate

Department

Civil Engineering

Advisor/Mentor

Hernandez, Sarah

Committee Member

Sasidharan, Lekshmi

Second Committee Member

Mitra, Suman

Keywords

similarity metrics, Arkansas, data vendors, origin-destination matrix

Abstract

The proliferation of commercial probe data vendors has posed both opportunities and challenges for transportation agencies regarding the choice of reliable vehicle probe data for decision making purposes. Probe data is a form of passive mobility data representing a small fraction of the traveling population. Vendors aggregate mobility data from different sources including GPS enabled smartphones, telematics systems, and connected vehicles. They possess unique data collection methodologies, sampling mechanisms, and expansion algorithms which introduce variability in derived traffic performance measures like speed and volume. The overall objective of this work is to assess similarities and differences among probe data vendors in terms of their ability to evaluate trip origin-destination, using a quantitative approach. This study presents a comprehensive normalization based comparative framework to evaluate three commercial probe data vendors for origin-destination data products. Two spatial scales are evaluated with case studies in Arkansas: a corridor level analysis of the Interstate I-30 corridor with seven exit locations (14 roadside sites) and an urban scale analysis of Harrison, Arkansas with 14 roadside sites. Passive mobility datasets are partial samples of the underlying travel patterns, and no universal ground truth exists to compute errors in different probe data vendors. This study evaluates consistence among vendors using individual matrix similarity metrics, and a composite similarity score. Total sum proportional normalization is applied to each vendor’s common-site OD matrix. For cross-vendor comparisons, the composite is an equal-weight average of agreement components on a 0–1 scale: inverted Frobenius distance, cosine similarity, CPC, and rescaled Spearman correlation. Three vendors are compared, referred to as Vendor A, B, and C. On the I-30 corridor (6 common exits, 15 directional OD pair), the all-day composite scores rank Vendor A–Vendor B highest (0.720), followed by Vendor A–Vendor C (0.608) and Vendor B–Vendor C (0.397). At the Harrison regional scale (8 common sites), the all-day scores are 0.480 for Vendor A–Vendor C, 0.458 for Vendor A–Vendor B, and 0.433 for Vendor B–Vendor C, while all period-specific cross-vendor composite scores range from 0.373 to 0.481. Vendor A–Vendor C has the highest Harrison composite in four of five periods, Vendor A–Vendor B is highest in the AM peak, and Vendor B–Vendor C is lowest in every period. Within-vendor temporal composite scores range from 0.811 to 0.976 on I-30 and from 0.872 to 0.991 in Harrison, confirming that each vendor’s internal temporal pattern is substantially more stable than its agreement with other vendors. The composite results show that the cross-scale change is vendor-pair specific: agreement decreases substantially for the pairs involving Vendor A, whereas Vendor B–Vendor C remains low and similar at both scales. VMT analysis further indicates that the vendors capture different travel patterns, with some datasets reflecting a greater share of longer through trips and others emphasizing shorter local movements. Individual metrics are retained alongside the composite because the same summary score can reflect different combinations of cell-level, directional, overlap, and rank agreement.

Share

COinS