kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Fine-grained urban land use simulation: Integrating spatial dynamic modeling with a pre-trained vision-language model
KTH, School of Architecture and the Built Environment (ABE), Sustainable development, Environmental science and Engineering, Resources, Energy and Infrastructure. Department of Urban Planning, School of Architecture, Southeast University, Nanjing, China.ORCID iD: 0000-0001-7219-2222
Department of Architecture and the Built Environment, Lund University, Lund, Sweden.
Department of Urban Studies and Planning, Massachusetts Institute of Technology, Cambridge, MA, USA.
School of Geography, University of Leeds, Leeds, United Kingdom.
2026 (English)In: Computers, Environment and Urban Systems, ISSN 0198-9715, E-ISSN 1873-7587, Vol. 126, article id 102416Article in journal (Refereed) Published
Abstract [en]

Accurate prediction of urban land use changes at fine spatial scales is essential for developing healthy and sustainable cities, yet traditional simulation models struggle to capture local dynamics due to limited availability of fine-grained data and insufficient complexity in modeling urban systems. To address these limitations, we propose a novel approach that leverages advances in pre-trained vision-language foundation models combined with spatial dynamic modeling to forecast detailed urban land use patterns. Specifically, we collected a spatially dense collection of street view images (SVIs) throughout Shenzhen, China, and applied UrbanCLIP, a specialized vision-language prompting framework, to perform zero-shot inference of urban land use directly from images without labeled datasets and model retraining. The resulting fine-grained classifications delineate eight distinct urban land use types, producing a detailed urban functional map. These high-resolution patterns were then integrated into a spatial dynamic model enhanced by polynomial regression to simulate urban evolution toward 2035. This approach effectively captures neighborhood influences, socioeconomic drivers, and urban planning policies. Our simulation provides actionable insights for sustainable development in Shenzhen by identifying areas for balanced growth, targeted infrastructure investments, and ecological preservation. Compared to conventional methods, our methodology significantly improves predictive accuracy and spatial granularity. By incorporating foundation models, our approach addresses traditional data constraints, offering scalable and robust tools for informed urban governance and decision-making.

Place, publisher, year, edition, pages
Elsevier BV , 2026. Vol. 126, article id 102416
Keywords [en]
Foundation models, Land use change, Spatial dynamic modeling, Street view images, Vision-language models
National Category
Civil Engineering
Identifiers
URN: urn:nbn:se:kth:diva-377999DOI: 10.1016/j.compenvurbsys.2026.102416ISI: 001706512200001Scopus ID: 2-s2.0-105030933534OAI: oai:DiVA.org:kth-377999DiVA, id: diva2:2046127
Note

QC 20260316

Available from: 2026-03-16 Created: 2026-03-16 Last updated: 2026-03-16Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Cai, Zipan

Search in DiVA

By author/editor
Cai, Zipan
By organisation
Resources, Energy and Infrastructure
In the same journal
Computers, Environment and Urban Systems
Civil Engineering

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 28 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf