Traditional urban fire-risk assessments rely heavily on static proxies, such as point-of-interest (POI) density, creating a semantic gap between macro-level urban indicators and micro-scale physical hazards. To bridge this gap, we develop a visual-semantic risk indicator system that uses multimodal large language models (MLLMs) to extract fine-grained fire-hazard features from street-view and remote-sensing imagery and integrates these features into a geographically weighted XGBoost (GW-XGBoost) model. Applied to Wuhan, China, the framework identifies localized physical hazards within broader patterns of urban intensity and distinguishes urban areas that appear similar in terms of static indicators but differ substantially in their micro-scale hazard conditions. Spatial interpretation reveals a clear outward shift in dominant risk predictors: fire risk in the urban core is associated primarily with functional overload, particularly the mixing of dining and residential uses, whereas risk in transitional areas is more strongly associated with infrastructure deficits and informal practices such as illicit ebike charging. Scenario simulations indicate that interventions targeting such practices in transitional areas are projected to produce the greatest reductions in fire risk. By coupling direct visual-semantic sensing with spatially explicit modeling, this study provides an evidence-based framework for precision urban fire-risk governance.

Figure 1. Research framework.