Automated construction of living mobility datasets from public video: A YouTube case study

Alam, M. S., Hoggenmueller, M., Bazilinskyy, P.

Submitted for publication.
ABSTRACT Public video contains naturalistic mobility footage, but transforming it into reusable research data requires more than a search and download. We present an automated and resumable framework for constructing living mobility datasets from YouTube, instantiated through two domain specific pipelines: pedestrian walking environments and road traffic crashes and near collisions. The pipelines combine discovery, Qwen metadata screening, GPU temporal segmentation, Cosmos3 Nano visual review, geographic grounding, provenance tracking, and versioned output. The resulting corpora contain 3,868 walking segments and 34,295 crash related segments. The walking corpus includes 89 resolved locality records in 28 countries, while the crash corpus contains 17,588 geographically resolved segments that represent 5,270 resolved locality records. Because the framework is resumable and continuously executable, the collection can run to expand datasets over time while preserving traceable processing decisions and reproducible snapshots. Its modular design can be adapted to other large scale video collections beyond mobility and YouTube. artificial-intelligencepreprintdashcam