kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Population Synthesis Using Incomplete Microsample
KTH, School of Architecture and the Built Environment (ABE), Urban Planning and Environment, Transport and Systems Analysis.ORCID iD: 0009-0004-7594-1451
KTH, School of Architecture and the Built Environment (ABE), Urban Planning and Environment, Transport and Systems Analysis.ORCID iD: 0000-0001-8901-5978
KTH, School of Architecture and the Built Environment (ABE), Urban Planning and Environment, Transport and Systems Analysis.ORCID iD: 0000-0001-5290-6101
2025 (English)In: Proceedings 26th EURO Working Group on Transportation, EWGT 2024, Elsevier BV , 2025, p. 80-87Conference paper, Published paper (Refereed)
Abstract [en]

This paper presents a population synthesis model based on the Wasserstein Generative-Adversarial Network with Gradient Penalty (WGAN-GP) for training on incomplete microsamples. The proposed method aims to address the challenge of missing information in microsamples on one or more attributes due to privacy concerns or data collection constraints. By using a mask matrix to represent missing values, the study proposes a WGAN-GP training algorithm that lets the model learn from a training dataset that has some missing information. The paper compares the ability of WGAN-GP models trained on incomplete microsamples to those trained on complete microsamples to create a synthetic population. We conducted a series of evaluations of the proposed method using a Swedish national travel survey. We validate the efficacy of the proposed method by generating synthetic populations from all the models and comparing them to the actual population dataset. The results from the experiments showed that the proposed methodology successfully generates synthetic data that closely resembles a model trained with complete data as well as the actual population. The paper makes a contribution to the field by giving a strong solution for population synthesis using incomplete microsamples. It also opens up new research areas and shows how deep generative models can be used to improve population synthesis.

Place, publisher, year, edition, pages
Elsevier BV , 2025. p. 80-87
Keywords [en]
microsample, population synthesis, WGAN-GP
National Category
Computer Sciences Transport Systems and Logistics
Identifiers
URN: urn:nbn:se:kth:diva-364403DOI: 10.1016/j.trpro.2025.04.011Scopus ID: 2-s2.0-105007068225OAI: oai:DiVA.org:kth-364403DiVA, id: diva2:1968217
Conference
26th EURO Working Group on Transportation, EWGT 2024, Lund, Sweden, Sep 4 2024 - Sep 6 2024
Note

QC 20250613

Available from: 2025-06-12 Created: 2025-06-12 Last updated: 2026-05-08Bibliographically approved
In thesis
1. Deep Learning Methods in Transportation and Urban Planning: Advancing data collection and inference methods
Open this publication in new window or tab >>Deep Learning Methods in Transportation and Urban Planning: Advancing data collection and inference methods
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Rapid urban development and the overwhelming amount of data provided by connected sensors, mobile devices, and open-source datasets challenge traditional transport-analysis methodologies. These methodologies largely depend on rigid mathematical models, static assumptions, and sparse measurements. This thesis demonstrates how deep learning (DL) has the potential to build a systematically better data collection and inference capability across three critical domains of the LUTI cycle: traffic management, population synthesis, and workplace location choice. Through five research papers, this thesis demonstrates that DL methods can complement or outperform traditional approaches by extracting more comprehensive data, building better predictive models, and providing actionable implications for planners. The analysis is separated into two main themes: data acquisition and analytical inference.

The data acquisition theme proposes methods for transforming overlooked or missing data into valuable input for transport models. Paper 1 uses street-view video from vehicle-oriented cameras to integrate YOLOv5 object detection, StrongSORT tracking, photogrammetry, and geodesy, automating the creation of link-wide time-space diagrams. The analysis of the method indicates that ordinary vehicles can be used to construct a low-cost virtual sensor network with far more spatial coverage than fixed detectors or GPS floating-car data. Paper 3 deals with missing-attribute surveys by including a binary mask in a Wasserstein GAN (WGAN) training setup. The masked WGAN learns directly from incomplete microsamples and creates synthetic populations that closely match marginal and joint distributions similarly to models trained on complete data, thus recovering survey entries that would otherwise be wasted.

The analytical inference theme associates the novel data sources with DL-enabled and hybrid models. Paper 2 adapts the Cell Transmission Model (CTM) and a two-stage Genetic Algorithm to perform fundamental-diagram parameter calibration and boundary condition assignment from partial camera trajectories. It recovers free-flow speed and critical density across three simulated arterial flows, and estimates unobserved densities, suggesting that sparse vision data can promote robust, link-level traffic state estimation. Paper 4 uses a Conditional Tabular GAN (CT-GAN) to construct target-year synthetic populations out of marginals directly, alongside a hybrid CT-GAN and Fitness-Based Combinatorial Optimisation pipeline that fine-tunes the marginal fit. During experimentation, CT-GAN demonstrates better performance for conditioned attributes over a baseline IPF-style model, achieving higher convergence and fidelity across the whole sequence. Paper 5 presents a custom Deep Neural Network (DNN) with zone blocks that ingest occupation type, socio-economic variables, and multimodal accessibility to predict workplace choice for over 1,000 job options. The DNN achieves superior performance metrics and reproduces observed distance distributions more accurately than a two-level nested logit model, without needing to manually construct a utility specification.

Overall, this thesis makes four main contributions to the literature: (1) Identifying vehicle-mounted cameras and incomplete surveys as valid, high-resolution sources of data when combined with DL pipelines; (2) Extending traffic state estimation to use partial trajectories extracted from video sequences to extract useful traffic states by means of GA-calibrated CTMs; (3) Improving synthetic population generation, providing evidence that GANs can satisfy future marginal constraints and learn from sparse, masked data; and (4) Generalising deep choice modelling to thousand-alternative examples, showing that DNNs can equal or exceed traditional discrete-choice models in accuracy and behavioural realism.

This work also examines drawbacks such as sensor calibration mistakes, modelling assumptions, and the generalisation and scalability issues of DL models, presenting solutions to these problems. The thesis also discusses policy implications of the research and provides recommendations for the use of DL, with lessons that are transferable to sectors such as energy, public health, and spatial systems.

In summary, this thesis advances the state of the art by embedding modern deep learning approaches into the core of urban transport modelling, paving the way for more data-rich, flexible, and socially-driven transport-planning tools in an increasingly complex society. Finally, it demonstrates that by combining domain knowledge with scalable DL architectures, the potential of current urban data can be unlocked to support more efficient, sustainable, and equitable mobility systems.

Abstract [sv]

Snabb stadsutveckling och den överväldigande mängden data som tillhandahålls av uppkopplade sensorer, mobila enheter och dataset med öppen källkod utmanar traditionella metoder för transportanalys. Dessa metoder är till stor del beroende av stela matematiska modeller, statiska antaganden och glesa mätningar. Denna avhandling visar hur djupinlärning (DL) har potential att bygga en systematiskt bättre datainsamling och inferenskapacitet inom tre kritiska områden i LUTI-cykeln: trafikhantering, befolkningssyntes och val av arbetsplatsplats. Genom fem forskningsartiklar visar denna avhandling att DL-metoder kan komplettera eller överträffa traditionella metoder genom att extrahera mer omfattande data, bygga bättre prediktiva modeller och ge handlingsbara implikationer för planerare. Analysen är uppdelad i två huvudteman: datainsamling och analytisk inferens.

Temat datainsamling föreslår metoder för att omvandla förbisedd eller saknad data till värdefull input för transportmodeller. Artikel 1 använder gatuvideo från fordonsorienterade kameror för att integrera YOLOv5-objektdetektering, StrongSORT-spårning, fotogrammetri och geodesi, vilket automatiserar skapandet av länkomfattande tidsrumsdiagram. Analysen av metoden indikerar att vanliga fordon kan användas för att konstruera ett billigt virtuellt sensornätverk med betydligt större spatial täckning än fasta detektorer eller flytande GPS-bildata. Artikel 3 behandlar undersökningar av saknade attribut genom att inkludera en binär mask i en Wasserstein GAN (WGAN) träningsuppsättning. Den maskerade WGAN lär sig direkt från ofullständiga mikroprover och skapar syntetiska populationer som nära matchar marginal- och gemensamma fördelningar på liknande sätt som modeller tränade på fullständig data, och återställer därmed undersökningsposter som annars skulle vara bortkastade.

Det analytiska inferenstemat associerar de nya datakällorna med DL-aktiverade och hybridmodeller. Artikel 2 anpassar Cell Transmission Model (CTM) och en tvåstegs genetisk algoritm för att utföra fundamentaldiagramparameterkalibrering och randvillkorstilldelning från partiella kamerabanor. Den återställer fritt flödeshastighet och kritisk densitet över tre simulerade arteriella flöden och uppskattar oobserverade densiteter, vilket tyder på att gles syndata kan främja robust trafiktillståndsuppskattning på länknivå. Artikel 4 använder ett villkorligt tabellformat GAN (CT-GAN) för att konstruera syntetiska målårspopulationer direkt från marginaler, tillsammans med en hybrid CT-GAN och Fitness-Based Combinatorical Optimisation pipeline som finjusterar marginalanpassningen. Under experiment visar CT-GAN bättre prestanda för villkorade attribut jämfört med en baslinjemodell i IPF-stil, vilket uppnår högre konvergens och trohet över hela sekvensen. Artikel 5 presenterar ett anpassat djupt neuralt nätverk (DNN) med "zonblock" som tar in yrkestyp, socioekonomiska variabler och multimodal tillgänglighet för att förutsäga arbetsplatsval för över 1 000 jobbalternativ. DNN uppnår överlägsna prestandamått och reproducerar observerade avståndsfördelningar mer exakt än en tvånivås kapslad logitmodell, utan att behöva manuellt konstruera en verktygsspecifikation.

Sammantaget ger denna avhandling fyra huvudsakliga bidrag till litteraturen: (1) Identifiera fordonsmonterade kameror och ofullständiga undersökningar som giltiga, högupplösta datakällor i kombination med DL-pipelines; (2) Utvidga trafiktillståndsuppskattningen till att använda partiella banor extraherade från videosekvenser för att extrahera användbara trafiktillstånd med hjälp av GA-kalibrerade CTM:er; (3) Förbättra syntetisk populationsgenerering, ge bevis för att GAN:er kan uppfylla framtida marginalbegränsningar och lära av glesa, maskerade data; och (4) Generalisera djupinlärningsmodellering till tusentals alternativa exempel, vilket visar att DNN:er kan motsvara eller överträffa traditionella diskreta valmodeller i noggrannhet och beteendemässig realism.

Detta arbete undersöker också nackdelar såsom sensorkalibreringsfel, modelleringsantaganden och generaliserings- och skalbarhetsproblem med DL-modeller, och presenterar lösningar på dessa problem. Avhandlingen diskuterar också policyimplikationer av forskningen och ger rekommendationer för användningen av DL, med lärdomar som är överförbara till sektorer som energi, folkhälsa och rumsliga system.

Sammanfattningsvis för denna avhandling framåt genom att bädda in moderna djupinlärningsmetoder i kärnan av urban transportmodellering, vilket banar väg för mer datarika, flexibla och socialt drivna transportplaneringsverktyg i ett alltmer komplext samhälle. Slutligen visar det att genom att kombinera domänkunskap med skalbara DL-arkitekturer kan potentialen hos befintliga urbana data frigöras för att stödja mer effektiva, hållbara och rättvisa mobilitetssystem.

Place, publisher, year, edition, pages
KTH Royal Institute of Technology, 2026. p. 64
Series
TRITA-ABE-DLT ; 269
National Category
Transport Systems and Logistics
Research subject
Transport Science, Transport Systems
Identifiers
urn:nbn:se:kth:diva-381022 (URN)978-91-8106-555-8 (ISBN)
Public defence
2026-05-26, F3, Lindstedtsvägen 26. KTH Campus, public video conference link https://kth-se.zoom.us/j/63381308713, Stockholm, 08:00 (English)
Opponent
Supervisors
Note

QC 20260511

Available from: 2026-05-11 Created: 2026-05-08 Last updated: 2026-05-25Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Rastogi, TanayJonsson, R. DanielKarlström, Anders

Search in DiVA

By author/editor
Rastogi, TanayJonsson, R. DanielKarlström, Anders
By organisation
Transport and Systems Analysis
Computer SciencesTransport Systems and Logistics

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 275 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf