Close×
    • Select all
      |
    • LI Yifan, LIU Changjun, ZHANG Chendi, YAO Shunyu, LI Qing
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Flash flood inducing backgrounds in complex mountainous regions are characterized by pronounced spatial heterogeneity and multi-factor coupling. Traditional clustering methods have difficulty simultaneously representing the spatial topological relationships among geographic units, while graph neural network (GNN)-based regionalization results often lack sufficient physical interpretability. This study aims to develop an integrated “regionalization-interpretation” framework for identifying flash flood inducing backgrounds, in order to quantitatively reveal the spatial heterogeneity and dominant driving mechanisms of flash flood inducing conditions in complex mountainous regions. [Methods] Taking the Hengduan Mountains Region as the study area, a comprehensive indicator system consisting of 36 environmental factors was constructed. The unsupervised GNN model Dink-Net was employed to capture spatial neighborhood characteristics and identify regionalization patterns. Subsequently, the DeepLIFT interpretation model was introduced to quantify both the magnitude and direction of single factor contributions. Compound factor analysis was further conducted by aggregating indicators to reveal differences in the dominant driving mechanisms among regionalization types with different runoff-generation patterns. [Results] ① Dink-Net effectively identified 12 secondary regionalization types with strong spatial continuity and geographic consistency. Based on the dominant inducing factors, typical environmental characteristics, and similarity in runoff-generation mechanisms, these secondary types were further generalized into 4 primary types: infiltration excess runoff type, saturation excess runoff type, compound runoff generation type, and retention and slow-release runoff type. ② The DeepLIFT interpretation results showed that extreme rainfall, elevation, temperature, and soil moisture were the key factors controlling regionalization boundaries at the regional scale. Moreover, the contribution directions of the same factor varied significantly among different types, suggesting that flash flood inducing processes in complex mountainous regions are controlled by different combinations of environmental conditions. ③ Compound factor analysis demonstrated that flash flood inducing backgrounds in the Hengduan Mountains Region are jointly constrained by extreme rainfall, hydrothermal background conditions, and large-gradient topography. Specifically, the infiltration excess runoff type is dominated by meteorological factors, the saturation excess runoff type is primarily controlled by antecedent hydrological conditions, the compound runoff generation type is characterized by pronounced terrain-driven subsurface stormflow processes and strong water-sediment coupling, while the retention and slow-release runoff type is jointly constrained by alpine climatic conditions and slow hydrological responses. [Conclusions] The proposed Dink-Net-DeepLIFT “regionalization-interpretation” framework realizes the integration of data-driven regionalization and physical mechanism interpretation, enabling complex black-box model outputs to be transformed into an intuitive understanding of flash flood-inducing mechanisms. This framework provides a scientific basis and methodological reference for differentiated flash flood prevention and risk management in complex mountainous regions.

    • WANG Ruifeng, WU Yunxia, ZHANG Hailan, XU Qian, LIU Jixin, ZHANG Han, WANG Zhihan, ZENG Jiachen
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Precipitation prediction, particularly for extreme precipitation events, remains a critical challenge in meteorological forecasting. Existing deep learning models predominantly optimize overall error metrics such as RMSE and MAE, which leads to insufficient sensitivity to low-frequency, high-intensity extreme precipitation events. To address this limitation, this study proposes a composite model named CWBiLSTMNets, which integrates a one-dimensional convolutional neural network (1D-CNN) and a bidirectional long short-term memory network (BiLSTM) to improve the accuracy of daily precipitation forecasting. In this model, CWB denotes the three core modules: CNN for local feature extraction, Weight Allocation Module for dynamic feature weighting, and BiLSTM for bidirectional temporal modeling. [Methods] To mitigate the insufficient model training caused by sparse extreme precipitation samples, we have designed a composite loss function that dynamically combines mean squared error with quantile loss. This function assigns higher weights to extreme precipitation samples while maintaining appropriate weights for conventional precipitation samples, thereby achieving implicit sample balancing and enhancing the model's sensitivity to extreme precipitation events. Additionally, an adaptive weight allocation module is introduced to dynamically adjust the weights of input meteorological features based on their relative importance across different precipitation intensities and climate periods. This mechanism enables the model to adapt to diverse precipitation patterns across different climate zones. Experiments were conducted using meteorological data from four cities in China (Shanghai, Shenzhen, Taipei and Tianjin), covering subtropical and temperate monsoon climate regions. Evaluation metrics include Root Mean Square Error (RMSE), Mean Absolute Error (MAE) for overall forecasting accuracy, as well as extreme precipitation forecasting metrics: Critical Success Index (CSI) and Probability of Detection (POD). [Results] In comparative experiments with CNN-BiLSTM, CNN-LSTM, LSTM, and Transformer models, CWBiLSTMNets demonstrates comprehensive superiority in extreme precipitation forecasting metrics. The average CSI reaches 0.139 6, representing a 138% improvement over CNN-BiLSTM (0.058 7); the average POD reaches 0.409 2, representing a 424% improvement over LSTM (0.078 1). Particularly in Shenzhen, the model achieves exceptional performance with CSI=0.296 3 and POD=0.571 4. For overall error metrics, the average RMSE is 13.97 mm and MAE is 7.25 mm, which are comparable to other baseline models, indicating that the proposed model maintains stable performance in conventional precipitation forecasting. [Conclusions] The proposed model effectively balances the prediction accuracy between conventional and extreme precipitation events. While maintaining the accuracy of conventional precipitation forecasting, the model significantly improves the capability to capture extreme precipitation events through its innovative composite loss function and adaptive weight allocation mechanism. This study provides a new approach for extreme precipitation forecasting and better meets the practical needs of disaster prevention and mitigation in meteorological applications.

    • WANG Jiawei, WANG Ke, CHEN Yuehong, LI Yunqiang, ZHANG Xiaoxiang, MA Qiang
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Current data-driven flash flood risk assessment methods heavily rely on large volumes of high-quality historical labeled samples, which limits their predictive performance in remote data-sparse or ungauged mountainous regions where historical disaster records are scarce. Although adversarial transfer learning provides a feasible and promising solution to alleviate such data-sparse problem by transferring knowledge from data-rich regions, existing methods often overlook the intrinsic geographical heterogeneity of flash flood disasters during the feature alignment process. This not only easily increases the risk of negative transfer by forcing the alignment of heterogeneous geographical backgrounds, but also fails to effectively characterize complex geospatial structures, such as river network topology. As a result, these issues limit the accuracy and reliability of flash flood risk assessment methods in complex geographical environments. [Methods] To address these limitations, this study proposes a Geography-constrained Domain Adversarial Neural Network (GCDANN) for flash flood risk assessment, which incorporates geographical constraints to enhance cross-domain adaptation and predictive accuracy. First, the source-domain is partitioned into multiple sub-environments based on the hierarchical clustering of flash flood-breeding environmental factors. Multi-domain discriminators tailored to specific geographical contexts are constructed accordingly. Second, the geographical similarity between target-domain and source-domain samples is calculated using a cosine similarity metric and utilized as dynamic attention weights for the corresponding domain discriminators. This weighting mechanism strategically strengthens deep feature alignment within homogeneous geographical environments while effectively suppressing adversarial interference and negative transfer from heterogeneous environments. Last, a graph neural network (GNN) based on spatial topology is employed as the core feature extractor to explicitly capture the upstream-downstream spatial relationships among sub-catchments. This is combined with a gradient reversal layer (GRL) to facilitate the robust end-to-end adversarial training between the feature extractor and the domain discriminators. [Results] The proposed GCDANN was applied to six major basins in the Hengduan Mountains (the Nu, Jinsha, Lancang, Yalong, Dadu, and Min Rivers), and evaluated under six distinct cross-domain transfer scenarios. Its performance was compared with a direct transfer model, two existing representative adversarial transfer models (DANN and MADA), and a fully supervised learning model. The results demonstrate that GCDANN consistently achieves superior performance in all evaluated scenarios, with accuracy and generalization levels closely approaching to the fully supervised model. Specifically, in the highly challenging transfer scenario from the Dadu River to the Min River basin, GCDANN improved the overall accuracy by 8.5%, 5.4%, and 4.7% over the direct transfer models, DANN and MADA, respectively, while remaining only 0.8% lower than the fully supervised model. Furthermore, the spatial distribution of high-risk zones identified by the GCDANN show significantly higher consistency with historical flash flood events compared to baseline models. [Conclusions] By integrating geographical constraints into the domain adversarial transfer learning, GCDANN effectively promotes feature alignment among homogeneous regions and explicitly characterizes the geospatial topological structures of catchments. Consequently, GCDANN significantly improves cross-domain transferability in complex geographically heterogeneous conditions. The proposed framework provides a reliable technical approach and decision-making support for flash flood risk assessment in data-sparse mountainous regions.

    • SHI Duoyuan, LIU Ying, ZHENG Hongwei, HUANG Yue, XIN Chenglong, LI Junli, JIAPAER Guli, BAO Anming
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] The Kashgar-Yarkand River Basin, located in the hinterland of the West Kunlun Mountains in China, is a typical alpine cold-region watershed. Floods in this area are characterized by sudden onset and high destructiveness. However, due to the scarcity of local hydrological data and flood susceptibility data that reflect the regional environmental characteristics and their response to flood-generating processes, flood risk assessment and prediction remain particularly challenging. Existing models struggle to achieve an effective balance between prediction accuracy and the physical interpretability of model results. This study develops a spatiotemporal learning model that enhances predictive performance while maintaining physical interpretability, thereby effectively improving the accuracy and timeliness of flood susceptibility assessment and spatiotemporal prediction in this region. [Methods] We propose a SHAP-guided extended long short-term memory framework (SHAP-xLSTM). The framework integrates a Contextual Feature Construction (CFC) module to capture local spatial information, employs an xLSTM architecture to strengthen the handling of time series data, and introduces a SHAP-guided feature weighting strategy to dynamically optimize and refine input features. The Kashgar-Yarkand River Basin was selected as the study area. The proposed model was systematically compared with five mainstream deep learning models, including CFC-LSTM, CFC-BiLSTM, CNN, CNN-LSTM, and CFC-xLSTM. Typical flood-prone areas along the Yarkand River were further used as validation cases to comprehensively evaluate model performance. [Results] The results indicate that: (1) areas with high and very high flood susceptibility are mainly concentrated in the central and northern parts of the study area, particularly within mountain-plain transition zones and along major river corridors, whereas areas with very low and low susceptibility are predominantly distributed in regions farther from river channels and at higher elevations; (2) compared with the second-best CFC-xLSTM baseline model, SHAP-xLSTM achieved systematic performance improvements, with AUC increasing by 0.71%, F1-score by 6.33%, and Kappa coefficient by 15.08%, while MSE and RMSE decreased by 21.74% and 11.68%, respectively; (3) SHAP-based interpretability analysis identified maximum air temperature, distance to rivers, and precipitation as the dominant driving factors of flood susceptibility, quantitatively supporting the characteristic “ice and snow accumulation-snowmelt-flood” disaster formation mechanism in high-altitude cold regions. Seasonal-scale spatiotemporal validation for 2024 further revealed an intensifying trend of flood susceptibility under climate warming, consistent with regional observations. [Conclusions] The proposed SHAP-xLSTM framework effectively reconciles predictive accuracy and physical interpretability in flood susceptibility modeling under complex cryospheric environments. It provides a robust methodological reference for flood risk early warning and adaptive management in high mountain cold regions.

    • MA Yufei, RUAN Yuli, LIU Cuishan, JIN Junliang, WANG Guoqing, HE Ruimin, TAN Shuchan
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] In recent years, extreme low-flow events (the hydrological manifestation of prolonged meteorological drought) have frequently occurred in the middle and lower reaches of the Yangtze River, posing a continual threat to regional water security and sustainable development. To analyze their causes and estimate future low-flow regimes, this paper conducts a systematic attribution analysis and future low-flow volume estimation for extreme low-flow events in this region. [Methods] The research clarifies the complex response mechanisms of extreme low-flow events to multiple driving factors under changing environmental conditions, reveals the evolution trends of low-flow regimes under different future climate scenarios, and provides a new perspective and scientific basis for adaptive management of regional water resources.The study is based on monthly runoff depth gridded data in China from 1961 to 2018, climate model data, and driving factor data. First, relying on runoff depth gridded data, the non-parametric kernel density estimation method is used to identify typical extreme low-flow years in the middle and lower reaches of the Yangtze River under different assurance conditions. Secondly, leveraging Pearson linear correlation analysis, grey relational nonlinear correlation analysis, and random forest models, an optimized deep learning method was introduced to improve the nonlinear regression model. This approach effectively eliminates interference from irrelevant factors while accurately identifying spatially differentiated driving factors of extreme low-flow conditions. Finally, the nonlinear regression model driven by coupled optimization deep learning was used to estimate the water resource availability for future extreme low-flow years. [Results] The research results show that 1966 (75% guarantee rate), 1978 (90% guarantee rate), and 2011 (95% guarantee rate) were typical extreme low-flow years in the region. Taking the extreme low-flow year under a 95% guarantee rate as an example, the representative extreme low-flow years in various decades were selected as 1968, 1972, 1984, 1994, 2006, and 2011, and the average of these six typical years was defined as the 'extreme low-flow' scenario. Under this scenario, Pearson analysis indicated that rainfall was the strongest positive correlation factor affecting extreme low-flow runoff depth, with a correlation coefficient of 0.712, while LAI was the strongest negative correlation factor, with a correlation coefficient of -0.119; grey relational analysis showed that rainfall had the highest grey relational degree; random forest analysis indicated that rainfall was the most critical factor affecting extreme low-flow runoff depth, with an importance of 42.45%, significantly higher than other factors. On this basis, based on a coupled optimized deep learning nonlinear regression model, the core driving factor combination affecting the spatial differentiation of extreme low-flow was obtained, consisting of precipitation, temperature, potential evapotranspiration, Leaf Area Index (LAI), and the Standardized Precipitation Evapotranspiration Index (SPEI). The optimized nonlinear regression model improved the NSE by 0.2, with a Nash-Sutcliffe Efficiency (NSE) of 0.841. Under the SSP1-2.6, SSP2-4.5, and SSP5-8.5 scenarios, the average water resources during future extreme drought years are 753.23, 741.29, and 766.51 mm, respectively. [Conclusions] The study revealed how extreme low-flow conditions respond to influencing factors in changing environments and the future amount of low-flow resources, which can provide scientific support for regional water resource adaptive management.

    • LIU Jun, CHEN Jijun, XIONG Junnan, YE Chongchong, CHENG Weiming
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Runoff simulation is fundamental for watershed water resource management and flood forecasting. However, traditional conceptual hydrological models typically rely on fixed parameters, limiting their ability to represent the temporal variability of watershed hydrological responses. [Methods] In this study, the SIMHYD (Simplified Hydrology Model) conceptual rainfall-runoff model was used as the physical framework and reconstructed in PyTorch to enable differentiable computation, thereby establishing a Differentiable dynamic Parameter Learning framework (DPL). A Long Short-Term Memory network (LSTM) was used to learn the dynamic mapping between hydro-meteorological forcings and model parameters, allowing the parameters to vary continuously over time, while runoff simulation errors were backpropagated to optimize the parameter-learning network. The simulation performance of different parameter learning methods was further evaluated at two scales, namely the overall runoff process and representative flood events. [Results] Daily runoff simulations were conducted in the Changjiang watershed from 2007 to 2022, and the performance of DPL was compared with the covariance matrix adaptation evolution strategy (CMA-ES) automatic calibration method and a differentiable Static Parameter Learning approach (SPL). Results showed that DPL achieved the best simulation performance, with validation-period Nash-Sutcliffe Efficiency (NSE) and Kling-Gupta Efficiency (KGE) values of 0.82 and 0.91, respectively, and full-period values of 0.88 and 0.94. Across four representative flood events, DPL achieved a mean absolute peak-flow error of 5.24%, lower than those of CMA-ES and SPL, and showed good performance in reproducing flood hydrographs and peak magnitudes, with more stable representation of peak flows and hydrograph variations under high-flow conditions. Analysis of dynamic parameters revealed pronounced seasonal variations, indicating that the framework can capture the adaptive adjustment of rainfall-runoff generation and routing processes under changing watershed hydrological states. Attribution analysis further showed that precipitation and soil moisture were the dominant drivers of dynamic parameter variations. The contribution of precipitation decayed relatively rapidly with increasing antecedent lag, whereas the influence of soil moisture persisted for a longer period, reflecting the memory effect of watershed moisture conditions and indicating that the dynamic parameters can respond to recent rainfall events while retaining the influence of antecedent soil moisture conditions on current runoff generation processes. [Conclusions] The proposed differentiable dynamic parameter learning framework based on the SIMHYD model can enhance the model’s adaptability to changing hydrological conditions and clarify the link between dynamic parameter variations and runoff generation and routing responses, providing a new perspective for interpreting the internal hydrological processes of conceptual models.

    • XUE Fengchang, ZHONG Guoyang, CHENG Yannian, TANG Jianzhong
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Urban waterlogging images usually contain different types of wading targets, such as pedestrians, motor vehicles and non-motor vehicles. A single visual reference is often insufficient for visual inundation-state discrimination in complex scenes, because the target may be absent, too small, partially occluded or affected by low illumination, water reflection and perspective changes. To address this problem, this study proposes a visual inundation-state recognition model for urban waterlogging wading targets by integrating multiple visual references. [Methods] A multi-visual-reference urban waterlogging image dataset was constructed from publicly available social media images and traffic surveillance video frames. The dataset contains 3 600 images and covers three common types of wading targets, namely pedestrians, motor vehicles and non-motor vehicles. According to the relative positions between the waterline and the type-specific key parts of different targets, four visual inundation states, S1-S4, were defined. For pedestrians, the state criteria were determined mainly by the relationship between the waterline and the foot, ankle, lower leg, knee, thigh and waist-hip regions. For motor vehicles, the key parts included the wheel, axle, hub, chassis and door region. For non-motor vehicles, the key parts included the wheel, axle, pedal, saddle and handlebar region. Based on these criteria, 12 joint labels in the form of “reference-object category-visual inundation state” were generated. On the basis of YOLO11n, dynamic snake convolution (DySnakeConv), the Attention-based Intra-scale Feature Interaction (AIFI) module and the DetectAux auxiliary detection branch were introduced to construct the DAD-YOLO model, which outputs the bounding box, confidence score and joint label of each wading target. [Results] Experimental results show that DAD-YOLO achieved Precision, Recall, mAP@50 and mAP@50:95 values of 80.1%, 56.5%, 64.1% and 50.8%, respectively. Compared with YOLO11n, the mAP@50 increased by 3.1%, while the computational cost increased from 6.3 GFLOPs to 7.2 GFLOPs. The recall rate of S3/S4 high-inundation-state targets increased from 61.2% to 68.4%, indicating improved recognition of targets with relatively severe visual inundation. [Conclusions] The proposed method can integrate visual evidence from pedestrians, motor vehicles and non-motor vehicles within a unified detection framework, thereby improving the recognition capability of visual inundation states for multiple types of wading targets in urban waterlogging images. The method can provide auxiliary visual information for rapid urban waterlogging perception and manual interpretation, but it should not replace field water-level measurements or traffic-control decisions.

    • SHI Hui, YOU Zhen, WANG Huihui, XIE Jun, LIANG Kai
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Against the backdrop of global warming and rapid urbanization, the frequency and intensity of extreme precipitation events have increased markedly, and urban flooding has evolved into a structural risk that constrains the safe functioning and sustainable development of cities. Existing urban flood resilience assessment frameworks predominantly rely on static structural indicators and fail to adequately account for the temporal heterogeneity of urban operations across different periods of the day, limiting their ability to capture the dynamic nature of urban resilience. [Methods] Using Shenzhen as the study area, this study collected real-time traffic speed data from the Amap Open Platform with a scheduled automated program and constructed a high-frequency speed dataset with city-wide road network coverage, a temporal span of approximately two months, and a 15 minute temporal resolution. The dataset covers daytime and evening activity periods from 07:00 to 23:00 and serves as the core data foundation for the behavioral dimension. Building on this, a three-dimensional indicator system integrating natural exposure (NEI), structural support (SSI), and behavioral speed (BSI) was established, capturing flood-related hazard exposure, infrastructure support capacity, and real-time traffic operational efficiency, respectively. Using a 30 m regular grid as the unified spatial analysis unit and a combined weighting method that integrates equal-weight priors with CRITIC-based objective correction, an Urban Flood Resilience Index (UFRI) was constructed to enable fine-grained resilience assessment across four representative time periods: morning peak, off-peak, evening peak, and weekends. [Results] Results indicate that all dimensional indices exhibit significant spatial clustering. Both the NEI and the SSI display an areal differentiation pattern characterized by higher values in the west and lower values in the southeast, while the BSI shows a contrasting spatial pattern in which peripheral areas outperform the urban core. The combined overlay of these three dimensions produces a composite spatial structure in which multi-dimensional vulnerability concentrates in the core urban area while resilience is more evenly distributed in peripheral zones. In terms of temporal dynamics, overall flood resilience is lowest during the evening peak period; the proportion of areas with negative UFRI differences relative to the off-peak period accounts for 10.87%, with highly vulnerable subdistricts concentrated in industrial-residential mixed-use zones. [Conclusions] The proposed framework is applicable to high-density urban contexts characterized by pronounced temporal variation in travel demand and the availability of high-frequency traffic speed data. The traffic speed data from commercial map APIs employed in this study are generated through real-time fusion of crowdsourced trajectories, offering high temporal resolution and broad road network coverage. Their programmatic accessibility across cities covered by major commercial mapping services provides a solid data foundation for the transferability and broader application of this method to comparable high-density cities, while also offering scientific support for the refined management of urban flood resilience.

    • ZHANG Yifei, YANG Ji, YIN Xiaoling, BAI Jingwen, LI Yong, HONG Jiayi, DENG Liming, XIAO Yibo
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Natural successional mangrove communities are highly dynamic ecosystems characterized by complex, multi-layered canopy structures and mixed-species compositions, possessing immense biodiversity and ecological value. However, due to the severe spectral convergence among co-occurring species and fuzzy canopy boundaries, fine-scale remote sensing classification remains highly challenging. Furthermore, the underlying classification mechanisms of conventional models are often unclear, limiting their ecological applicability. To address these issues, this study proposes a fine-scale classification framework for natural mangroves using low-altitude Unmanned Aerial Vehicle (UAV) multispectral imagery integrated with explainable machine learning techniques. [Methods] Several machine learning models, including XGBoost (eXtreme Gradient Boosting) and LCE (Local Cascade Ensemble), were introduced to systematically evaluate their comprehensive performance under conditions of high-dimensional complex features and severe sample imbalance. To capture both spatial and spectral heterogeneity, a comprehensive set of spectral, textural, and vegetation index features was extracted utilizing Object-Based Image Analysis (OBIA). Subsequently, a rigorous feature selection strategy combining decorrelation analysis and permutation importance was employed to eliminate redundant variables and optimize the feature space. Finally, the SHAP (SHapley Additive exPlanations) framework was utilized to quantify feature contributions and interpret the biologically sensitive features associated with different mangrove species. [Results] Under the respective optimal feature configurations of the four classifiers, LCE achieved an Overall Accuracy (OA) of 0.917, which was comparable to those of XGBoost (0.918) and RF (0.914) and higher than that of SVM (0.900). LCE obtained the highest Macro F1-score (0.918), exceeding those of RF (0.915), XGBoost (0.913), and SVM (0.901), indicating more balanced class-wise performance under class-imbalanced conditions. The feature selection results showed that LCE reached its peak performance using a refined subset of 10 key features, suggesting that removing redundant features contributed to improved model performance. The subsequent SHAP analysis revealed a strong response of Rhizophora stylosa to the Normalized Difference Water Index (NDWI), which was presumably associated with its physiological structure and residual water signals beneath the canopy. In addition, the high red-band reflectance of juvenile Aegiceras corniculatum may be related to its developmental stage and potential salt stress. [Conclusions] The combination of LCE and feature selection maintained high overall accuracy while improving the balance of class-wise recognition, demonstrating its applicability to the fine-scale classification of complex natural successional mangrove communities under class-imbalanced conditions. This framework can provide methodological support for the precise monitoring and conservation of mangrove ecosystems.

    • SU Shiliang, ZHOU Yujia, SHAO Aitan, XIANG Juan, KANG Mengjun, WENG Min
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Background] Narrative map design is a multidisciplinary field where, regardless of the breadth of theoretical perspectives, maps must ultimately be designed. Design thus serves as the nexus between theory and practice. Yet it is precisely at this nexus that a significant disconnect persists: academia has long focused on cartographic theories and methods while neglecting the product attributes and design logic of maps; industry, despite producing some outstanding works, remains largely confined to data presentation and surface-level visualization, with practical experience seldom systematically consolidated to inform theory. This mutual disconnection has prevented the full unfolding of interaction between theory and practice. [Methods] This paper adopts theoretical deduction and empirical induction as its primary methodologies. Proceeding from the ontological characteristics and internal logic of narrative map design, it proposes a concrete engineering pathway. On this basis, taking the narrative map of The Grand Canal of China as an empirical object, the study examines the key decisions and iterative processes at each stage within a real-world project setting, so as to test the operational feasibility of the proposed pathway. [Results] The study has achieved the following outcomes: ① It constructs an engineering pathway for narrative map design comprising five stages: inscription and delimitation, data engineering, prototype architecture, product design, and social intervention. ② It verifies the effectiveness and adaptability of this pathway in the Grand Canal of China case, demonstrating the successful transformation of cultural heritage spatial information into a perceptible narrative space. ③ It proves that this pathway can achieve the integration of theoretical awareness and product realization, providing empirical support for moving narrative map design from individual creation toward systematic practice. [Conclusions] The engineering pathway established in this paper effectively bridges the gap between theory and practice, offering a methodological tool that is both operational and flexible for narrative map design. Without compromising design creativity, this pathway enhances the controllability of the design process and the replicability of its outcomes. Future research may further extend its application to different types of narrative maps and explore its deeper integration with digital interactive technologies and user participation mechanisms.

    • LI Sijin, DING Hu, CHEN Qingsong, TANG Guoan
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Accurate extraction of gullies is crucial for understanding soil erosion and land degradation. High Spatial Resolution (HSR) imagery provides detailed surface feature information and is an important data source for geomorphological object extraction and mapping. However, the complexity of surface features makes HSR imagery prone to strong shadow effects and spectral confusion, posing significant challenges to existing methods. [Methods] To address these issues, this study introduces an adaptive attention mechanism and a shadow-balancing module designed to enhance the model's ability to capture spatial information and mitigate performance disparities between shadowed and non-shadowed areas. Building on these components, we developed an end-to-end gully mapping network, GMNet (Gully Mapping Net), for single-scene HSR optical imagery. [Results] Experiments show that GMNet outperforms baseline models such as UNet and Attention UNet (AUNet). Specifically, it improved the F1-score and Intersection over Union (IoU) by approximately 2.3% and 22.7%, respectively, while also achieving higher accuracy and recall. Validation across vegetated and bare surfaces of the Loess Plateau, under varying shadow conditions, demonstrated an IoU of 0.92 against reference data, indicating strong adaptability. Some misclassification persists in areas with intensive human activity, suggesting that additional diverse training data are needed to further enhance generalization. Further experiments also confirmed the contribution of the shadow-balancing module, showing that its adjustable weighting mechanism effectively optimizes mapping results. [Conclusions] The proposed GMNet shows strong potential for large-scale gully mapping and offers valuable applications in soil erosion monitoring and land management.

    • LIAO Zhuosen, XIE Ke, WANG Tao, ZHONG Shuxian, LI Xiaoyu
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Three-dimensional Geographic Information Systems (3D GIS) support more realistic, information-rich, and fine-grained modeling of geographic scenes. Nevertheless, user interfaces for and interactions with 3D GIS are more complex than those in 2D GIS. Interactions with 3D GIS involve multiple layers of information, including scene states, spatial objects, attributes, tools and related parameter constraints. Natural User Interfaces (NUIs) can improve both efficiency and user experience of 3D GIS interactions. Natural language serves as a core medium in NUIs and is characterized by flexible expression and implicit semantics. Accurate mapping from natural language instructions to 3D GIS interactions is highly challenging. The advances in Large Language Models (LLMs) offer new opportunities for understanding and transforming complex natural language instructions to 3D GIS interactions. To address these challenges, this study proposes a natural language instruction mapping framework for 3D GIS scene interactions. [Methods] To facilitate generic LLMs, a knowledge base was first built based on the available 3D spatial database and 3D GIS functionality for Retrieval Augmented Generation (RAG). To support the following natural language based interaction, a triplet model was then designed for 3D GIS interaction. The three components comprise spatial objects and the knowledge base, parameters of the in-situ 3D scene context, and task related tools. Subsequently, to map natural language instructions to 3D GIS operations, a natural language parsing and scene-interaction workflow was implemented based on an LLM and an AI agent. The workflow first translates user natural language instructions to 3D GIS tasks, then routes the parsed results to parameter extraction and standardized GIS operation execution. Four types of typical 3D GIS interactions were implemented for testing, including symbolic styling, view control, spatial analysis, and spatial simulation based on procedural modeling. To evaluate the effectiveness of the proposed approach, 220 natural language instructions were constructed and tested for the four types of interactions. The performance of the proposed framework was evaluated using three indicators: instruction recognition accuracy, parameter mapping conformity, and final task-completion rate. [Results] The proposed framework achieved satisfactory operation-mapping performance across the four typical 3D GIS interaction tasks. Overall, the instruction recognition accuracy reached 98.64%, the parameter mapping conformity was 86.59%, and the final task-completion rate was 85.45%. These results verified the feasibility of the triplet-based model for representing 3D GIS interaction with workflow-based processing of natural language instructions. [Conclusions] The proposed framework provides a generic solution for the semantic organization of 3D GIS interactions, the design of natural language interpretation workflows, and the performance evaluation in 3D GIS natural language interaction based on LLMs. The approach can also be extended to adapt to the design of human-computer interaction in digital twins, virtual geographic environments, and 3D spatial decision-support systems.

    • JIANG Qi, HAN Yong, YANG Jingyuan, WANG Jinhua, YAO Zhixin, LIU Jin
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Large Language Models (LLMs) have demonstrated great potential in intelligent tourism recommendation due to their strong semantic understanding and generation capabilities. However, their application in tourism scenarios remains challenging due to factual hallucinations, limited domain knowledge, and insufficient handling of complex multi-constraint requirements. Existing Retrieval-Augmented Generation (RAG) methods also suffer from fragmented evidence and semantic mismatches, as tourism information is distributed across heterogeneous textual descriptions, visual contents, and structured attributes, limiting the generation of reliable and constraint-consistent recommendations. To address these challenges, this paper proposes TMKG-LLM, a collaborative framework integrating a Tourism Multimodal Knowledge Graph (TMKG) with LLMs to enhance knowledge representation, evidence organization, and recommendation generation. [Methods] First, structured knowledge is extracted from heterogeneous tourism data, including textual descriptions and attraction images, through prompt engineering to construct an attraction-centric Tourism Multimodal Knowledge Graph. The constructed TMKG integrates textual semantics, visual characteristics, and contextual attributes into a unified knowledge representation, providing comprehensive domain knowledge support for LLM-based recommendation. Second, a hypergraph-based higher-order evidence organization mechanism is introduced to address the fragmentation problem of conventional retrieval methods. Specifically, attractions and their associated multidimensional attributes are aggregated into higher-order semantic evidence units, enabling the joint representation of complex relationships among attractions, themes, services, spatial constraints, and user preferences. Furthermore, a candidate evidence selection strategy combining topological priors and multidimensional scoring is designed to improve evidence retrieval accuracy under complex user intents. Finally, the retrieved structured evidence is injected into LLMs as factual constraints, allowing the model to perform knowledge-guided reasoning and generate tourism recommendations with improved factual consistency and interpretability. [Results] Experiments conducted on a Beijing tourism recommendation benchmark demonstrate the effectiveness of TMKG-LLM. Compared with the best baseline model, TMKG-LLM achieves relative improvements of 10.1%, 8.1%, 9.5%, and 7.4% in Context Precision, Context Recall, Faithfulness, and Answer Relevance, respectively. These results indicate that multimodal knowledge graphs and LLMs achieve effective collaboration, where the former provides structured and traceable knowledge support while the latter performs semantic reasoning and recommendation generation under knowledge constraints. Further itinerary constraint evaluation shows that TMKG-LLM achieves 81.2% and 62.0% micro-level and macro-level hard constraint satisfaction rates, respectively, demonstrating its effectiveness in multi-condition tourism planning scenarios. Ablation studies further confirm that hypergraph-based evidence organization, topological prior selection, and multidimensional scoring mechanisms jointly contribute to improving retrieval quality and generation performance. [Conclusions] By integrating multimodal knowledge representation, higher-order evidence modeling, and LLM reasoning capabilities, TMKG-LLM promotes the transition of tourism recommendation from fragmented information retrieval toward structured knowledge-driven scenario-aware generation. The proposed framework provides a reliable, traceable, and interpretable technical pathway for intelligent tourism recommendation under complex user requirements.

    • XU Can, QIAN Haizhong, GONG Xianyong, LIU Chengyi, XIONG Shun
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Significance] Identical entities refer to spatial geographic objects that represent the same real-world geographic object or set of objects across multiple geospatial datasets. Such entities are typically represented in an object-based geographic form, among which vector geospatial data are the most representative format. Due to differences in data acquisition purposes, measurement errors, map scales, and coordinate reference systems, identical entities in different vector datasets often exhibit substantial discrepancies in the geometric shapes and attribute descriptions. These inconsistencies introduce uncertainties such as positional displacement, geometric deformation, and object splitting or merging, posing substantial technical challenges to the integration, fusion, and updating of multi-source and multi-scale geospatial data. Identical entity matching for vector geospatial data aims to identify corresponding entities across heterogeneous datasets by analyzing multi-dimensional features of spatial entities and measuring their similarities, thereby enabling precise association and alignment among entities. As a fundamental technique, identical entity matching in vector geospatial data provides critical support for multi-source data fusion and updating, as well as change detection and dynamic monitoring. [Analysis] This paper systematically reviews the domestic and international research progress on identical entity matching in vector geospatial data. First, the study categorizes and discusses common approaches for spatial and semantic feature representation from three perspectives: geometric features representation, topological relational structures representation, and semantic representation enhancement. This paper further analyzes the roles of different features in candidate matching generation, similarity measurement, and structural consistency constraints. Second, it outlines the modeling mechanisms of multi-feature fusion, global optimization and consistency matching methods, with emphasis on their roles and applicability in identifying complex correspondence relationships among heterogeneous spatial entities. Finally, this paper reviews the research progress of learning-driven matching methods, including traditional supervised learning, deep representation learning, graph neural networks, and pre-trained language models. It compares different models in terms of their architectures, input features, and applicable objects, and evaluates their advantages and limitations across different matching scenarios. [Objectives] This study aims to analyze the technical characteristics, application scenarios, and limitations of existing identical entity matching methods for vector geospatial data, particularly under multi-source and multi-scale conditions. Furthermore, it discusses future research directions from four perspectives: unified geospatial entity representation and matching modeling, extension to homogeneous group and cross-category entity matching, large-model-driven spatial reasoning enhancement for geospatial entity matching, and unified representation and dynamic matching of heterogeneous spatiotemporal data. The review is expected to provide theoretical and methodological references for future studies on vector geospatial data integration, updating, and intelligent matching.

    • LI Yali, HUANG Guie, ZHANG Caili, YU Yingji, XIANG Longgang
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Interchanges serve as critical hubs for grade-separated traffic flow in urban road networks, and the accuracy of their digital representation is vital for reliable vehicle navigation. Although OpenStreetMap (OSM) is widely used as a crowdsourced geospatial data source offering global coverage and low costs, it frequently suffers from topological issues—such as missing roads, geometric inaccuracies, and directional errors—within complex segments like interchanges. These defects severely limit its applicability in high-precision navigation. [Methods] To address this, this paper proposes a method that integrates crowdsourced trajectory data clustering and filtering with a gated bidirectional recurrent neural network (BiRNN) for OSM interchange identification and automatic road error detection. First, interchange regions are identified by combining OSM and trajectory data using DBSCAN clustering and a sliding window technique with multifeature fusion (trajectory direction entropy and intersection trajectory density). Second, an enhanced Hidden Markov Model (HMM) that allows path interruption is applied for partial map matching, effectively mitigating the impact of inherent OSM errors (e.g., missing or disconnected roads) during the matching process. Third, a collaborative feature construction scheme is designed from both local and global perspectives, yielding 11 basic features (6 local and 5 global) that are further encoded into a 21 dimensional standardized vector; dimensionality reduction is then performed via roadrepresentationbased principal component analysis retaining the first 12 principal components that explain 92.3% of the cumulative variance, to optimize training efficiency and suppress noise. Finally, a BiRNN embedded with a bidirectional independent gating mechanism is designed to dynamically focus on trajectory points associated with road anomalies, using prediction entropy as an interestpoint signal; together with a classsensitive weighted loss and a feature difference regularization term, the model achieves precise identification and classification of missing roads, missing corner ramps, directional errors, and geometric inaccuracies through a multistage collaborative pipeline. [Results] Experimental results on realworld vehicle trajectory data from Beijing’s main urban area demonstrate that the proposed method achieves an accuracy of 97.7%, a recall of 96.3%, and an F1score of 97.4% in detecting interchange road errors, significantly outperforming comparative baseline methods such as Logistic Regression, CRF, RNN, GNN, and Transformer. Ablation studies confirm that the gating mechanism improves the F1score by 0.7% over standard BiRNN, and the improved HMM matching contributes a 2.7% gain over the standard HMM. [Conclusions] This approach provides an effective technical solution for enhancing the navigability of OSM data in complex interchange areas, demonstrating that a modular framework deeply integrating physical priors (map matching and road semantic weighting) with datadriven models is particularly advantageous under imbalanced error categories and strong trajectory noise.

    • WANG Yifan, XIONG Liyang, TANG Guoan
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] The cross-sectional morphology of loess gullies constitutes a core external expression for understanding the processes and mechanisms of loess landform development. Morphological attributes such as gully-slope gradient, width, and cross-sectional area directly reflect erosion intensity, developmental stage, and regional variability. However, research on the spatial heterogeneity of gully cross-sectional morphology across the Loess Plateau remains limited in both spatial extent and data volume, and a systematic understanding of its variation with gully order is still lacking. [Methods] Based on DEM data of multiple spatial resolutions, this study addressed the limitations of previous research, which has relied largely on manual extraction of sampled cross sections, by proposing an automated, order-adaptive method for gully cross-section extraction. This method integrates gully network hierarchy and a normal-vector-based sampling strategy to enable dynamic, batch extraction of cross-sectional morphological features of gullies with widths adapted to their Strahler orders across representative regions of the Loess Plateau. Using this approach, more than 17 000 cross sections were obtained. On this basis, a series of indices were calculated, including gully-wall slope gradient, total gully width, cross-sectional area, total eroded area, and mean depth of the eroded portion. Their interrelationships, regional differences, variations with gully order, and sensitivity to DEM resolution were then systematically analyzed. [Results] The results show that: (1) Gully morphology on the Loess Plateau exhibits pronounced north-south spatial differentiation. Gully-slope gradients become progressively gentler from south to north. The total eroded area and total cross-sectional area of gullies are relatively large in parts of the central and northern regions, generally increasing with gully order but showing fluctuations at higher orders. Total gully width is greatest in the central region and expands with increasing order. The mean depth of the eroded portion of the gully cross section is highest in the northern region and tends to decrease with increasing order, reflecting intensified downstream vertical incision. (2) Gully-slope gradient is significantly negatively correlated with total gully width and positively correlated with depth-related indices, indicating that it effectively captures the degree of gully development. Width- and area-related metrics jointly reflect stages of gully evolution. In contrast, the gully-floor width index shows weak correlations with most other metrics and appears to have limited association with overall erosional dynamics, suggesting that it may serve as an independent indicator for evaluating gully maturity. (3) Coarse-resolution DEM significantly underestimates gully wall slope steepness by 30%~50% and markedly reduces its variability. It not only increases the difficulty of extracting low-order gullies, but also reduces the overall variability of the extracted gullies. [Conclusions] Gully development in the northern Loess Plateau is generally more advanced than that in the southern region. This study reveals the spatial differentiation pattern and scale sensitivity of gully morphology on the Loess Plateau, providing a scientific basis for soil erosion monitoring, geomorphic evolution research, and regional soil and water conservation.

    • LI Jintao, ZHAO Mingwei
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Open-source DEM data provide an important basis for surface flow analysis, but their relatively low spatial resolution makes it difficult to represent intra-cell terrain variations. Most existing flow-direction algorithms determine flow directions or flow-allocation proportions mainly based on elevation relationships between a central cell and its neighboring cells, with insufficient consideration of sub-grid-scale terrain control. Therefore, incorporating intra-cell terrain information to improve flow-allocation accuracy is a key issue in DEM-based hydrological analysis using open-source DEM data. [Methods] To address this issue, this study proposes a multiple-flow-direction algorithm based on DEM downscaling and triangular-facet continuous tracking. The method first constructs a new grid using the center points of the original DEM cells as elevation control points, and then obtains a downscaled terrain representation within the original cells through triangular subdivision and local plane equations. Subsequently, the downscaled cells are used as starting points for flow-path tracking, and flow-allocation proportions are calculated statistically. This approach overcomes the limitation of traditional algorithms that rely only on elevation relationships between grid cells, and effectively incorporates sub-grid-scale terrain information into the flow-allocation process. [Results] The proposed method was evaluated using theoretical mathematical surfaces and a real-world study area in Fengyang County, Anhui Province, China. The artificial DEMs were constructed from convex-centered, concave-centered, and saddle-shaped mathematical surfaces, while the real-world experiment used 30 m resolution ASTER GDEM V3 data. The proposed method was compared with the D8, MFD-md, and D-inf methods. The artificial DEM experiments show that the proposed method performs stably under different terrain conditions, especially for saddle-shaped slopes involving convergence, divergence, and flow-direction transitions. Under low-relief conditions, the RMSE values of the proposed method for saddle-shaped slopes at resolutions of 5 m, 10 m, and 20 m are 59.2 m, 32.2 m, and 20.7 m, respectively, which are approximately 66.4%, 78.4%, and 83.8% lower than those of the best-performing comparison method, MFD-md. The real-world experiment shows that the channels extracted by the proposed method correspond well to valleys, footslopes, and low-lying flow-convergence areas. Its low-lying-area coverage rate, channel-low-lying-area consistency, and comprehensive evaluation index are 0.186, 0.725, and 0.296, respectively. The comprehensive evaluation index is approximately 7.6% and 1.7% higher than those of D8 and D-inf, respectively, while the extracted channel area is approximately 15.2% smaller than that of MFD-md. These results indicate that the proposed method can maintain good low-lying-area coverage while controlling channel expansion, making it suitable for complex natural terrain scenarios that require a balance between channel spatial concentration and low-lying flow-convergence area coverage. [Conclusions] The results show that the triangular-facet continuous tracking method achieves better overall performance than the existing methods, especially for open-source low-resolution DEM data. This indicates that the proposed method can incorporate the influence of terrain differences within original cells on water-flow routing into the flow-allocation process, thereby partially compensating for the inability of open-source low-resolution DEMs to adequately represent sub-grid-scale terrain variations within grid cells. The proposed method provides a new methodological approach for hydrological analyses such as flow accumulation, SCA calculation, and channel extraction under such DEM conditions.

    • GE Mingming, SUN Lijian, LI Panpan, JI Ling, LIU Jiping, GUO Qingsheng, JIANG Hao, LI Bao, KONG Haozhu, LIAO Xin
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Significance] Cultivated land serves as a foundational resource for ensuring global food security and sustainable agricultural development. Consequently, its dynamic monitoring and comprehensive assessment have emerged as a critical research frontier within natural resource management and agricultural remote sensing. Driven by the rapid evolution of Artificial Intelligence (AI) technologies, cultivated land monitoring is transitioning from traditional field surveys, statistical sampling, and manual interpretation toward a new paradigm characterized by multi-source data fusion, end-to-end deep modeling, and intelligent decision support. Fueled by advancing computational power and expanding geospatial data, the deep integration of AI and remote sensing has reshaped this field. In this context, this study presents a holistic review centered on the “perception-modeling-decision” technical continuum, providing a comprehensive forward-looking perspective and a strategic roadmap for multi-scale precision identification and dynamic management of cultivated land resources. [Progress] Utilizing a systematic bibliometric analysis combined with a literature synthesis of research from the past four decades, this paper traces the historical evolution of AI-driven remote sensing for cultivated land. It delivers an in-depth analysis of methodological advancements across four core tasks: cropland area extraction, boundary delineation, agricultural parcel segmentation, and comprehensive assessment. The findings reveal that across these core tasks, research paradigms and hotspots have progressively evolved from conventional machine learning to deep learning, and are currently converging toward foundational spatial-temporal models and multi-agent synergy. This transition underlines the transformative potential of AI in empowering cultivated land protection, fine-grained management, and spatial governance toward higher efficiency, intelligence, and sustainability. [Prospect] Despite rapid progress, the operational, large-scale deployment of AI in cultivated land management remains constrained by critical bottlenecks, including the scarcity of high-quality annotated samples, limited physical interpretability of black-box models, weak cross-regional generalization under complex landscapes, and challenges in integrating emerging AI paradigms into legacy operational workflows. To address these limitations, future research urgently needs to: (1) construct sample augmentation frameworks driven by generative AI and multimodal foundation models to overcome small-sample generalization barriers; (2) couple data-driven algorithms with physical/agronomic domain knowledge to enhance both estimation accuracy and physical interpretability; (3) deeply integrate autonomous agent strategies with spatiotemporal digital twin technologies to establish comprehensive evaluation systems encompassing cropland quantity, quality, and ecological resilience; and (4) leverage Human-Machine-Physical (HMP) collaborative governance platforms to ultimately construct a smart supervision ecosystem integrating real-time sensing, scientific assessment, and automated decision-making. Collectively, these advancements will establish a vital technological bedrock for fostering more efficient, resilient, and sustainable agricultural ecosystems.

    • WANG Jiedong, ZHOU Qianwen, ZHANG Zishi, REN Na, ZHU Changqing
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Significance] Textures of 3D real-scene models may contain sensitive targets, such as signs, logos, or other identifiable objects, which may hinder the secure sharing and compliant application of such data in smart-city construction. Existing privacy-preserving or desensitization methods mainly focus on the detection and removal of sensitive information in two-dimensional texture images. However, texture information and geometric structures are highly coupled in 3D real-scene models. Even after sensitive textures are removed or regenerated, the corresponding spatial geometric contours may remain in the mesh, resulting in a potential risk of structural information leakage. Therefore, it is necessary to jointly process texture and geometry information to improve the completeness and security of sensitive-target desensitization. [Methods] This study proposes a joint geometry-texture desensitization method for 3D real-scene models. First, YOLOv11 is employed to detect sensitive targets in texture atlases, and LaMa is used to regenerate the corresponding sensitive texture regions. Second, a barycentric-coordinate-based mapping relationship is established between two-dimensional texture regions and three-dimensional mesh triangles, thereby accurately locating the sensitive geometric regions associated with the detected texture targets. Finally, non-sensitive boundary vertices surrounding each sensitive region are extracted to fit a local reference plane. Internal sensitive vertices are then updated by projecting them along the normal direction of the fitted plane. This operation weakens or eliminates the local spatial geometric contours of sensitive targets while maintaining the continuity of mesh boundaries and preserving the original mesh topology. [Results] Experiments were conducted on three 3D real-scene models collected from different areas, containing a total of 687 sensitive targets. The recall values of sensitive-target detection achieved by the proposed method were 0.95, 0.94, and 0.96 for the three datasets, respectively, representing improvements of 0.09, 0.12, and 0.12 over the reference method. After desensitization, the Hausdorff distances of sensitive geometric regions decreased from 0.35, 0.38, and 0.42 m to 0.012, 0.018, and 0.015 m, respectively. In addition, the LPIPS values decreased from 0.28, 0.31, and 0.34 to 0.12, 0.11, and 0.12, respectively, indicating improved visual consistency between regenerated texture regions and their surrounding contexts. Compared with the method without the mesh-flattening module, the proposed method caused a boundary geometric smoothness variation of no more than 0.15°. The cumulative time consumed by three-dimensional localization and local mesh flattening accounted for only 0.82% of the total processing time. [Conclusions] The proposed method jointly reduces the texture-level and spatial-geometric recognizability of sensitive targets while preserving texture naturalness and mesh-boundary continuity. It provides technical support for the secure sharing, compliant release, and practical utilization of 3D real-scene models.

    • PENG Chengli, LI Jiarui, LIU Yang
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Remote Sensing Image (RSI) object detection is a fundamental task in the intelligent interpretation of remote sensing imagery. Due to complex backgrounds, dense object distributions, and significant variations in object scales, capturing multi-scale features is crucial for improving detection performance, particularly for small and densely distributed objects. For instance, in some RSIs, small objects such as vehicles or ships may occupy only a few pixels yet carry critical semantic information, making their reliable detection an urgent challenge. However, most existing methods rely on sophisticated architectures to encode and fuse multi-scale information, which inevitably introduces substantial computational overhead and limits their practical deployment in resource-constrained scenarios. [Methods] To address these challenges, this paper proposes an efficient and accurate RSI object detection framework, termed F-DETR, from a frequency-domain perspective. Unlike conventional spatial-domain approaches, frequency-domain representations can explicitly separate high-frequency and low-frequency components, providing a more effective and interpretable way to model multi-scale information and compensating for the inherent limitations of spatial-domain feature extraction. The proposed framework consists of two lightweight yet effective components for multi-scale feature encoding and fusion. Specifically, for multi-scale information encoding, a Frequency-domain Visual Mamba (FVM) backbone is developed based on the VMamba architecture. By enhancing high-frequency components in feature maps, FVM effectively captures fine-grained details and improves the representation of small-scale objects, without requiring additional computational overhead that conventional methods incur for small object perception. At the same time, for multi-scale information fusion, a Frequency-domain Fusion Block (FFBlock) is designed to perform frequency band decomposition, enabling efficient modeling of multi-level correlations across multi-scale feature maps. Furthermore, a multi-band parameter sharing strategy combined with a channel attention mechanism is introduced to adaptively emphasize informative features, ensuring sufficient feature interactions while maintaining low computational complexity, thereby addressing the inherent trade-off between fusion completeness and efficiency in existing methods. [Results] Extensive experiments are conducted on two widely used remote sensing object detection benchmarks, namely DIOR and USOD. The results demonstrate that all detection metrics of F-DETR are optimal, and the key metric mAP can be improved by 0.7%~7.1% on DIOR and 1.2%~8.3% on USOD compared to the average level of other methods such as YOLOv5-L and RT-DETR. In addition, it consistently outperforms several representative object detection models, including YOLOv5, YOLOv8, and DINO. From an efficiency perspective, F-DETR achieves lower GFLOPs and faster inference speed, demonstrating a superior trade-off between detection accuracy and computational cost. [Conclusions] The proposed F-DETR framework verifies the effectiveness of leveraging frequency-domain representations to jointly improve detection accuracy and efficiency. It provides a practical and scalable solution for real-time remote sensing applications regarding remote sensing object detection, particularly in compute-constrained scenarios including lightweight edge detection terminals and resource-limited emergency response environments.

    • XIA Changshuo, ZHAO Wei, RAN Guangtai, TAN Jianbo, WU Tianjun, YANG Bin
      Download PDF ( ) HTML ( )   Knowledge map   Save

      [Objectives] Artificial forests are vital for ecological restoration, soil erosion control, and desertification mitigation, particularly in the middle reaches of the Yarlung Zangbo River Valley on the Qinghai-Tibet Plateau, where they are extensively distributed. However, the region features fragmented plantation parcels and heterogeneous planting patterns, which collectively hinder conventional pixel-based remote sensing approaches from achieving both accurate boundary delineation and reliable temporal characterization simultaneously. [Methods] To address these challenges, this study proposes an integrated technical framework for planted forest monitoring that combines semantic segmentation, object-based boundary optimization, and time-series analysis. In the spatial domain, high-resolution remote sensing imagery is initially segmented using the DeepLabV3 model, followed by watershed transformation and level-set methods to refine parcel boundaries. In the temporal domain, a long-term Normalized Difference Vegetation Index (NDVI) time series spanning 1988-2024 is constructed from Landsat and Sentinel-2 imagery. An annual-difference-based change detection approach is then applied to identify abrupt vegetation transitions, enabling the inversion of planting years at the parcel scale and establishing a consistent linkage between spatial units and temporal attributes. [Results] The proposed method achieves a mean Intersection over Union (mIoU) of 88.72%, with a maximum of 91.33% in specific regions, an Overall Accuracy (OA) of 90.12%, and a Recall of 90.05%. For planting year inversion, comparison with field survey data yields a Root Mean Square Error (RMSE) of 2.78 years, with approximately 80% of parcels exhibiting errors within ±2 years. Compared with conventional pixel-based methods like LandTrendr, our approach produces more complete parcel boundaries and spatially coherent temporal information under complex valley conditions. [Conclusions] In general, by integrating semantic segmentation and time-series analysis, this study realizes the synergistic expression of spatial distribution and planting‑year attributes of artificial forests in the plateau river valley region. The proposed approach offers robust data support and a practical technical reference for regional plantation monitoring, ecological engineering evaluation, and fine‑scale management.