TerraMosaic Daily Digest: September 8, 2026
Daily Summary
An exact analytical treatment of erosion-induced landslide motion revises a basic assumption about entrainment: erosion rate alone does not determine whether mobility increases or decreases. The new mobility controller shows that volume bulking, erosion drag or thrust, and flow depth jointly regulate runout, providing a mechanical test that any physically valid erosion model must satisfy. Field evidence from the Baihetan Reservoir reaches a complementary conclusion at slope scale. After impoundment, the Renhe landslide accelerated from less than 10 mm yr⁻¹ to 35-70 mm yr⁻¹; reservoir-level cycling destabilized the lower mass, while a longer infiltration-softening-deformation-cracking feedback propagated failure retrogressively toward the crest. Together, the studies place material exchange and internally differentiated hydrology at the centre of landslide mobility and progressive failure.
The 2025 Hualien disaster demonstrates how rapidly a compound mountain cascade can exhaust engineering response time. A roughly 290 × 10⁶ m³ landslide formed a 200 m-high dam and a lake holding about 91 × 10⁶ m³; two months later, typhoon rainfall drove overtopping and a two-peaked breach hydrograph reaching 9309 m³ s⁻¹. Scenario analysis suggests evacuation reduced confirmed deaths from a potential 136 to 19, underscoring contingency planning when stabilization is infeasible. At national scale, a physics-embedded learning model simulates more than 800,000 US river reaches hourly, raising median Nash-Sutcliffe efficiency from 0.461 to 0.683 relative to National Water Model v3.0 and reducing extreme-flood timing error. A proposed five-level FLOSEV scale addresses the parallel need to compare floods when discharge measurements are absent by using documented human and material impacts.
Observation and learning methods increasingly target the transition from signal to actionable warning. S-SAR monitoring of an open-pit landslide combines spatial clustering, critical-slowing indicators and inverse-variance extrapolation to identify warning structure and estimate instability time. In curved-terrain granular avalanches, staged PINN training lowers held-out trajectory error by roughly two orders of magnitude, while four observations bracketing acceleration-to-deceleration perform nearly as well as eight broadly sampled points. Nainital PSI-InSAR links millimetre-scale uplift and subsidence to active structures and unstable hillslopes; physical debris-flow experiments resolve how bends, density and particle sorting generate asymmetric depth, pore-pressure and impact-stress responses. Synthetic-data studies add a second route around observation scarcity: diffusion models generate landslide imagery for augmentation, and OPMS-Seg releases 2,814 expert-validated UAV images with 15,334 slope instances for open-pit monitoring.
Key Trends
The dominant advance is a tighter connection between process continuity, observation design and decisions made under short warning times.
- Erosion and hydrology are becoming state variables rather than correction factors: The analytical mobility controller separates erosion rate from bulking, drag or thrust and flow depth, while the Renhe case resolves different reservoir and infiltration controls in lower and upper slide masses. Both replace single-factor explanations with evolving internal state.
- Cascade models are being evaluated against response time and consequences: Hualien links earthquake legacy, successive typhoons, landslide dam formation, overtopping, breach hydraulics and evacuation within one event chain. Flood-severity and building-level multi-hazard frameworks similarly translate physical hazards into comparable losses and interventions.
- Physics-guided learning is moving to operational spatial scales: The US flood model combines long hydrologic memory, short shocks and infiltration excess across a continental river network; CyberShake NZ organizes thousands of finite-fault simulations into national PSHA; flash-flood prediction decouples physics-based routing from transferable rainfall-loss learning.
- Observation placement matters as much as observation volume: Granular-avalanche PINNs gain most from measurements spanning the acceleration transition, S-SAR warning analysis reduces hundreds of points to representative deformation zones, and PSI-InSAR uses spatial coherence with active structures to distinguish unstable Himalayan slopes.
- Foundation models are being adapted through domain structure and verified data: Landslide diffusion models use semantic conditioning and LoRA for scarce regional imagery, OPMS-Seg supplies expert-validated mine-slope geometry, and remote-sensing vision-language models introduce hierarchical resolution alignment. The emphasis is shifting from generic scale to domain-specific grounding.
Selected Papers
The September 8 literature is led by a new analytical law for erosion-controlled landslide mobility, a national-scale physics-embedded flood simulator, the Hualien landslide-dam breach reconstruction, Baihetan reservoir-landslide mechanics and S-SAR critical-slowing warning analysis. Direct geohazard studies extend across granular avalanches, debris-flow bends, Himalayan InSAR, flood severity, CyberShake NZ, wildfire sensing and building-level multi-hazard appraisal. A broad methodological layer develops domain-grounded foundation models, remote-sensing reconstruction, uncertainty-aware weather prediction, point-cloud analysis and coupled geotechnical simulation, linking process understanding with monitoring and intervention across scales.
1. Kinematics of erosion-induced landslide mobility
Core Problem: Erosion is widely associated with enhanced mobility, but erosion rate alone cannot explain why entrainment sometimes accelerates and sometimes retards a moving mass.
Key Innovation: Derives an exact mobility controller coupling erosion rate with bulking, erosion drag or thrust and flow depth, establishing a mechanical criterion that distinguishes enhanced from reduced mobility.
2. Hourly U.S.-wide flood simulation beyond the limits of traditional and data-driven models
Core Problem: Operational flood models lose skill at rare peaks and in ungauged reaches where hourly prediction matters most.
Key Innovation: Combines multiple hydrologic timescales and infiltration-excess physics across more than 800,000 reaches, improving median gauge efficiency and capturing substantially more 50- and 100-year floods than operational baselines.
3. CyberShake NZ: Probabilistic Seismic Hazard Analysis in New Zealand Using Physics-Based Ground-Motion Simulations
Core Problem: National PSHA requires computationally tractable sampling of rupture variability and wave propagation at consistent sites.
Key Innovation: CyberShake NZ automates rupture, velocity-model, HPC and post-processing stages, aggregating 16,043 finite-fault simulations at 25,948 receivers into hazard maps, curves and disaggregation products.
4. Failure mechanism of a landslide triggered by combined reservoir water level fluctuations and rainfall in the Baihetan Reservoir, China
Core Problem: Reservoir-level fluctuations and rainfall act differently across a compound slide mass, obscuring the sequence of progressive failure.
Key Innovation: Separates lower-slope groundwater forcing from upper-slope infiltration feedback and reconstructs a retrogressive toe-to-crest failure process after deformation accelerated to 35-70 mm yr⁻¹.
5. The landslide dam breach disaster in Hualien, Taiwan on 23 September 2025
Core Problem: A giant landslide dam formed only two months before breach, leaving insufficient time for conventional engineering mitigation.
Key Innovation: Integrates earthquake and typhoon precursors, a roughly 290 million m³ landslide, overtopping and two breach peaks with evacuation scenarios that quantify the life-safety value of contingency action.
6. Measuring floods by their scars: toward a data-driven flood severity scale
Core Problem: Flood magnitude is difficult to compare when peak discharge is missing or hydrologically incomparable across settings.
Key Innovation: Derives the five-level FLOSEV scale from the long-running Italian AVI inventory, classifying events through affected municipalities, fatalities and property damage rather than discharge alone.
7. Research on the extraction of landslide features, identification of early warning signals, and prediction of instability time based on synthetic aperture radar monitoring: A critical slowing down theory perspective
Core Problem: Dense deformation monitoring needs a defensible route from spatially heterogeneous motion to warning onset and failure-time estimation.
Key Innovation: Combines clustering of more than one hundred S-SAR points with variance-autocorrelation critical-slowing signals and inverse-variance prediction to isolate the active zone and estimate instability timing.
8. Physics-Informed Neural Networks for Depth-Averaged Granular Avalanche Dynamics on Curved Topography
Core Problem: PINNs for depth-averaged mass flow struggle with curved geometry, strain-rate-dependent closure and sparse observations.
Key Innovation: Embeds Savage-Hutter dynamics and Mohr-Coulomb closure in a staged temporal curriculum, reducing held-out trajectory error by about two orders of magnitude and showing that transition-spanning observations dominate sample count.
9. The Capacity of Generative Models to Synthesize Regional Landslide and Non-Landslide Remote Sensing Imagery Under Data-Scarce Scenarios: Insights from Multimodal Foundation Models
Core Problem: Regional landslide classifiers are constrained by limited, imbalanced high-quality imagery.
Key Innovation: Fine-tunes multiple Stable Diffusion backbones with LoRA on Bijie landslide and non-landslide imagery, lowering minimum FID to 54.70 and motivating heterogeneous backbones for balanced augmentation.
10. OPMS-Seg: A UAV-Based High-Resolution Image Dataset for Semantic Segmentation of Open-Pit Coal Mine Slopes
Core Problem: Vision-based mine-slope warning lacks a public dataset that distinguishes operationally relevant slope geometries.
Key Innovation: Releases 2,814 UAV images with 15,334 expert-validated polygons across four slope classes in COCO, YOLO and mask formats, with reproducible segmentation baselines.
11. Landslide hazard assessment in Nainital hills of Indian Himalayas using Interferometric Synthetic Aperture Radar (InSAR)
Core Problem: Tectonic uplift, subsidence and mass movement overlap in the Nainital hills, complicating regional hazard interpretation.
Key Innovation: Maps 2017-2024 deformation with PSI and cross-checks GPS strain, linking localized uplift and subsidence patterns to faults and identifying Balia Nala as a landslide-prone sector.
12. Physical modeling of nonlinear dynamic responses of debris flows through topographically variable paths
Core Problem: Straight flumes cannot represent particle sorting, vortex formation and asymmetric loading through successive channel bends.
Key Innovation: Introduces a configurable 2D/3D laboratory system and resolves nonlinear velocity, depth, pore-pressure and bank-stress evolution along an S-shaped path across multiple debris densities.
13. Source-Unbiased Subsurface Monitoring With Seismic Noise
Core Problem: This study addresses a key limitation of seismic noise interferometry for subsurface monitoring: the impact of non-stationary sources on seismic velocity-change estimates.
Key Innovation: This study addresses a key limitation of seismic noise interferometry for subsurface monitoring: the impact of non-stationary sources on seismic velocity-change estimates. Analytical and numerical tests show that source-frequency variations can produce apparent velocity changes of equivalent magnitude.
14. Fault-Controlled Plumbing System Beneath Lipari Volcanic Island Revealed by Seismic Matrix Imaging
Core Problem: Lipari Island hosts an active magmatic-hydrothermal system, but its deep plumbing architecture remains poorly resolved due to strong seismic scattering.
Key Innovation: Together, these features define a structural framework in which tectonically guided vertical transfer may connect with laterally partitioned shallow storage and possible hydrothermal sealing.
15. Sparse Incident-Cluster Learning for 12-hour Port Flood Pre-warning in Digital-Twin Analytics
Core Problem: We formulate 12-hour port flood pre-warning as an incident-cluster learning problem and evaluate a digital-twin analytics module using eight-point water-level histories, prediction-time contextual covariates, and interpretable short-window dynamics.
Key Innovation: Row-level classification can therefore overstate performance by placing windows from the same event in both model-development and evaluation data. Across the Liverpool folds, the top-10 ElasticNet model achieves mean F2 = 0.696, compared with 0.633 without top-k truncation and 0.681 for full-feature weighted XGBoost.
16. Diffuse2Seg: Diffusion Models Can Segment Anything Without Supervision
Core Problem: Collecting such labels is expensive and time-consuming, however, which limits how far this recipe can scale.
Key Innovation: Building on this observation, we introduce Diffuse2Seg, which repurposes generative diffusion models for automatic mask generation by propagating a grid of point prompts through their self-attention representations in an edge-preserving manner. In this setting, SAM sets a strong standard: trained on SA-1B, comprising 11M images and over 1B carefully annotated masks, it achieves remarkable zero-shot performance.
17. The 2016 Mw 7.0 Kumamoto Earthquake Sequence, Japan revisited: Insights from Spatio-Temporal Analysis of Seismicity Parameters
Core Problem: Understanding how crustal faults accumulate strain, nucleate ruptures, and redistribute post-seismic stress is fundamental to seismic hazard assessment.
Key Innovation: This framework resolves asperity locking, release, and healing better than any single metric; however, since anomalies were identified retrospectively, they should be read as evidence of coherent behavior rather than a validated forecast tool alone. Depth-sliced volumes show the low-b locked core (b ≤ 0.65) was stratified at 10--12.5~km depth and sharpened within the final four months before failure.
18. ANADEF: A Nested-Permutation Alarm for Dual-Parameter Earthquake Forecasting
Core Problem: Spatially resolved stress proxies and rate-based seismicity models are increasingly combined for regional earthquake forecasting, yet formally testing their non-redundancy remains largely unaddressed.
Key Innovation: We present the Nested-Permutation Alarm for Dual-Parameter Earthquake Forecasting (ANADEF) pipeline, integrating a stress-sensitive Gutenberg--Richter b-value field, estimated via a penalized 2D B-spline inversion, with a stationary background rate (μ) from space--time ETAS stochastic declustering.
19. SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs
Core Problem: Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored.
Key Innovation: We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA) generated from a 9.7K-image subset, spanning 10 evaluation dimensions from basic perception to higher-order reasoning.
20. Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems
Core Problem: Machine learning emulators have become essential for accelerating expensive Earth-system simulations, but most existing approaches remain passive forecasters: they reproduce simulator trajectories under prescribed forcings without an explicit interaction mechanism for user-specified interventions.
Key Innovation: We propose an action-conditioned world-modeling framework for Earth-system emulation that reformulates simulator trajectories as supervision for controllable state-transition learning. Experiments show that the model preserves competitive long-horizon emulation accuracy while enabling controllable structural interventions and coherent responses in coupled ecosystem-cycle variables.
21. MCANet: A Multi-Scale Class-Specific Attention Network for Multi-Label Post-Hurricane Damage Assessment Using UAV Imagery
Core Problem: Unmanned aerial vehicle (UAV) imagery provides high-resolution observations of affected areas, but post-hurricane scenes are difficult to classify because multiple damage categories often co-occur within the same image, appear at different spatial scales, and include visually similar severity levels as well as rare but operationally important classes.
Key Innovation: To address these challenges, this study presents MCANet, a multi-label classification framework for post-hurricane UAV damage assessment. Evaluation on the RescueNet dataset, which includes 4,494 UAV images collected after Hurricane Michael and annotated with 10 damage categories, shows that MCANet achieves the highest mean average precision (mAP) among the evaluated models, with an mAP of 91.37%.
22. Knowledge-Guided Vision-Language Inference for Image-Based Urban Flood Depth Estimation
Core Problem: Timely floodwater depth estimates support road accessibility assessment and emergency response during urban flooding.
Key Innovation: This paper proposes FloodVision, a knowledge-guided framework for estimating flood depth from a single RGB image.
23. Resolving sources of uncertainty in AI weather forecasting
Core Problem: Weather forecast uncertainty arises from imperfect analyses and forecast models, but ensemble spread alone does not reveal how distinct sources relate to downstream targets.
Key Innovation: We introduce Pangu-Bayes, a probabilistic forecasting hierarchy that treats atmospheric-state and learned-model uncertainty as distinct stochastic variables, crossing flow-dependent perturbations of the evolving state with Bayesian parameter samples.
24. RAH-VLA: Resolution-Adaptive Hierarchical Vision-Language Alignment for Multimodal Remote Sensing Understanding
Core Problem: However, existing methods are limited by fixed-resolution visual processing and single-scale vision-language alignment, making it difficult to simultaneously preserve fine-grained details and maintain semantic consistency across different spatial granularities.
Key Innovation: To address these challenges, we propose RAH-VLA, a Resolution-Adaptive Hierarchical Vision-Language Alignment framework for multimodal remote sensing understanding. Extensive experiments on multiple remote sensing benchmarks demonstrate that RAH-VLA consistently improves image captioning, visual grounding, and cross-modal reasoning performance while reducing computational redundancy.
25. Global satellite gravity data products for prompt detection of short-term Mass Change (MC)
Core Problem: Abstract.
Key Innovation: We present the globally available dataset of Line-of-sight Gravity Differences (LGD) as a new data product to fill the long-standing gap of investigating sub-monthly surface mass change from the satellite gravimetry along-track perspective. We demonstrate its potential through case studies, including along-track diagnosis of flash drought evolution in the southeastern United States and the characterization of sub-monthly.
26. Extraordinary Tibetan Plateau early-winter heating drove record-breaking winter precipitation in California and adjacent regions
Core Problem: However, subseasonal to seasonal predictive skill for Californian winter precipitation has remained persistently low.
Key Innovation: This study shows that anomalous early-winter heating over the Tibetan Plateau (TP) played a key role in driving the extreme precipitation based on observational analyses and Earth system model experiments.
27. City-scale dynamic assessment of pedestrian travel risk in urban subway systems during flood events
Core Problem: Urban flooding is intensifying due to climate change and rapid urbanization, yet the flood risk faced by subway travellers remains inadequately understood.
Key Innovation: Taking the Beijing subway system as a representative megacity case, this study develops a city-scale framework for dynamically assessing travel-related flood risk, integrating flood hazard, pedestrian travel exposure, and pedestrians’ physical vulnerability to floodwater environments. The results identify four spatially clustered high-risk regions across the Beijing subway network, involving multiple critical lines and.
28. From soil loss to sediment delivery: GeoAI-enhanced RUSLE modeling of erosion and sediment dynamics in a hyper-arid watershed
Core Problem: In the delineated watershed containing Wadi Samnan near Az Zulfi, Saudi Arabia, soil erosion is a local management concern because sparse vegetation, erodible sandy surface materials, escarpment-influenced terrain, and episodic rainfall events can concentrate runoff and sediment movement along wadi channels and drainage corridors.
Key Innovation: This study presents a GeoAI-enhanced RUSLE-based framework to model water-induced soil loss and sediment delivery dynamics in the study watershed. The results show that RUSLE-derived potential soil-loss estimates ranged from 0 to 319.5 t ha -1 yr -1, with the slight erosion class covering 43.2% of the watershed and modeled severe-erosion hotspots covering 10.9%.
29. Linking Riverbank Erosion Dynamics and Livelihood Vulnerability in a Rapidly Urbanising Mekong Delta River Corridor
Core Problem: Accretion was more spatially extensive than erosion along both banks; however, the left bank experienced a greater total extent and magnitude of erosion.
Key Innovation: This study examines long-term bankline change from 2001 to 2025 and develops expert-informed priorities for assessing livelihood vulnerability within the Can Tho reach. The exploratory hydraulic observations varied among the five selected locations but showed no consistent correspondence with historical erosion magnitude and are therefore interpreted only as a snapshot of conditions on the survey date.
30. Geo-XAI Reveals Wildfire Risk Driver Differences and Spatial Heterogeneity Between Drought and Non-Drought Periods in Southwest China Mountains
Core Problem: However, the factors associated with wildfire occurrence and their nonlinear spatial responses under contrasting drought conditions remain poorly understood.
Key Innovation: Using historical wildfire records from 2006 to 2020 and 16 wildfire drivers, we developed three machine-learning models and applied GeoShapley to quantify the contributions of key predictors and characterize their spatial dynamics under different drought conditions. The results showed that the Extreme Gradient Boosting (XGB) model achieved the best predictive performance (AUC = 0.85-0.91) and effectively captured the spatial.
31. A microscale framework for harmonized flood and earthquake risk assessment and multi-hazard intervention appraisal
Core Problem: Multi-hazard risk reduction is increasingly recognized as essential for effective disaster risk management, yet its operationalization remains constrained by decision-making frameworks not designed for cross-hazard comparison.
Key Innovation: Multi-hazard risk reduction is increasingly recognized as essential for effective disaster risk management, yet its operationalization remains constrained by decision-making frameworks not designed for cross-hazard comparison. The framework is demonstrated through an application to a portfolio of residential buildings in Pesaro (Italy) exposed to riverine flooding and seismic hazard under multiple hazard and vulnerability.
32. Gender differences in disaster mortality and risk perception: evidence from floods and landslides in Italy
Core Problem: Studying human consequences caused by landslides and floods is important to understand their impact and to investigate public risk perception We analysed a catalogue of 1,188 landslide and 752 flood fatalities that occurred in Italy during the 60-year period 1965-2024, for which information on sex, age, and circumstances of death was available.
Key Innovation: Observed fatalities were compared with expected distributions modelled from census data by sex and age for two non-overlapping 30-year periods. Overall, the results highlight a persistent mismatch between gender patterns in mortality and perceived personal threat and support targeted, gender-sensitive risk communication and disaster risk reduction strategies.
33. DNFNet: a brightness-adaptive lightweight RGB-thermal network for day-night UAV wildfire segmentation
Core Problem: Wildfires significantly threaten human safety and property, and unmanned aerial vehicles (UAVs) have become a crucial tool for regional fire monitoring.
Key Innovation: To reduce false positives and negatives caused by diurnal lighting changes and modality differences in day-night RGB-Thermal (RGB-T) UAV fire segmentation, we propose DNFNet, a brightness-adaptive, lightweight dual-branch, multi-level complementary fusion network. Results show that DNFNet achieves an IoU of 0.854 and an F1 score of 0.921 on day-night data, while remaining computationally efficient (12.06 million parameters.
34. Toward transferable flash flood oriented high flow prediction in ungauged basins using a decoupled physics-DL framework
Core Problem: Accurate flash flood prediction in ungauged basins remains a fundamental challenge.
Key Innovation: To address these limitations, we propose a decoupled physics-DL framework that separates the rainfall-runoff process into two components: physics-based surface runoff routing and a DL-based estimation of event-scale effective rainfall. In this evaluation, the framework achieved a median critical success index (CSI) of 0.50 under the strictest warning criteria, with the CSI increasing under more relaxed tolerance settings, and.
35. From precipitation forecasts to optimal reservoir operation: an integrated downscaling-forecasting-operation framework for reservoir floodwater utilization
Core Problem: Floodwater utilization (FU) has emerged as a critical strategy to balance flood control and water conservation, yet existing operations struggle to systematically integrate high-resolution meteorological forecasts and their inherent uncertainties into real-time decisions.
Key Innovation: To address this gap, this study develops an integrated FU framework that establishes forecasting-and-control system coupling a deep learning precipitation downscaling model (DeepSD), multi-model hydrological forecasting, and a rolling horizon control (RHC) optimization strategy. For inflow forecasting, the Graph Neural Network (GNN) achieved the best overall performance among the three hydrological models, with NSE values of.
36. Detecting Irrigation From Spectral Differences Between Satellite and Modeled Soil Moisture Across the Contiguous United States
Core Problem: Irrigation alters the terrestrial water cycle, yet its spatial distribution and temporal variability remain poorly constrained because existing data sets often rely on indirect proxies or inventories rather than observations tied to land-surface water-balance dynamics.
Key Innovation: Here, we introduce a wavelet-based method to detect irrigation from spectral differences between modeled and satellite-observed soil-moisture time series, implemented using Noah-MP simulations and Soil Moisture and Ocean Salinity (SMOS) observations.
37. Resolving the Physical Ambiguity of Passive Infrared Cloud Optical Thickness Retrievals With Microwave Observations
Core Problem: Satellite thermal infrared (TIR) observations, combined with machine-learning methods, enable nighttime retrievals of cloud optical thickness (COT), but their cloud-top-dominated radiances provide ambiguous information for optically thick clouds.
Key Innovation: Here we quantify this benefit by combining TIR and MW brightness temperatures within a U-Net COT retrieval framework. Results demonstrate that TIR-MW synergy provides physically complementary, column-integrated constraints that improve passive COT retrievals, strengthen the physical basis of machine-learning cloud property retrievals, and reduce their tendency to underestimate optically thick clouds.
38. SWOT Discharge Accuracy Benchmarked in South America
Core Problem: The Surface Water and Ocean Topography (SWOT) satellite mission offers unprecedented estimates of global river discharge.
Key Innovation: This study provides a first continental-scale assessment of SWOT Consensus discharge performance across South America and compares it with a large-scale hydrodynamic model and in situ observations.
39. Estimating subsurface water mass changes with ground-based gravimetry
Core Problem: Gravitational force is proportional to the mass of an attracting body; therefore, changes in subsurface mass can be detected using gravimetry.
Key Innovation: In the near subsurface, mass variations are primarily driven by changes in water storage, meaning that gravity measurements provide direct information on variations in water mass.
40. Gaussian Linear Functional Manifold Method for Massive Point Cloud Data
Core Problem: Reconstructing continuous terrain manifolds from massive, unstructured airborne LiDAR point clouds remains challenging in complex Wildland-Urban Interface (WUI) environments, where deep neural networks require costly point-wise annotations and nonparametric surface reconstruction methods often lack structural interpretability.
Key Innovation: This paper introduces the Gaussian Linear Functional Manifold (GLFM), a physics-informed statistical framework that represents continuous surface topography using deterministic linear functional bases while modeling microscale diffuse laser backscatter as an isotropic Gaussian process.
41. GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims
Core Problem: GeoContext constructs a context ladder by stratifying nearby reference points according to distance and referenceability, allowing the image to remain fixed while the supplied context varies.
Key Innovation: We introduce GeoContext, a resource supporting two complementary tasks: GeoHint, open-ended localization given a true but coarse location hint, and GeoVerify, binary verification of whether an image was taken within 150 m of a claimed place. Our evaluation reveals three main patterns.
42. CoRe-SAM3: Conditional Semantic--Visual Reconciliation for SAM3 Crack Segmentation
Core Problem: Crack segmentation requires a model to recognize target semantics while accurately recovering thin, low-contrast, and topologically continuous local structures.
Key Innovation: Based on this finding, we propose Conditional Semantic--Visual Reconciliation, termed CoRe. The results show that the semantic representation already carries most task information for crack prediction, whereas the utility of the visual representation depends on the current semantic state.
43. Map the Possibilities: Spatial Belief Fields for Language-Goal Aerial Navigation
Core Problem: Existing methods often reduce language grounding to one single waypoint or action, prematurely collapsing the spatial uncertainty inherent in incomplete evidence and ambiguous relations.
Key Innovation: To address this limitation, we introduce SBFNav, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF). Experiments on both the original and revised CityNav benchmarks achieve the best reported overall performance.
44. ProtoRAG: Prototype-Based Retrieval Augmentation for Few-Shot Fine-Grained Remote Sensing Object Detection
Core Problem: Few-shot fine-grained object detection (FGOD) in remote sensing imagery is challenging because limited annotations must support both object localization and discrimination among visually similar subcategories.
Key Innovation: To address this limitation, we propose ProtoRAG, a prototype-based retrieval-augmented framework that decouples coarse localization from fine-grained recognition by equipping MLLMs with an external object-level visual memory. Extensive experiments show that ProtoRAG consistently surpasses representative baselines in nine few-shot settings, outperforming the strongest baselines by 14.80, 2.27, and 4.04 mAP50 on MAR20.
45. EgoNeMo: Transferable Map of Pedestrian Dynamics via Egocentric LiDAR Scan
Core Problem: This paper proposes a transferable Map of Dynamics (MoD) framework that generalizes to unknown environments using only egocentric 3D LiDAR point clouds to overcome the long-standing limitation of traditional MoD methods.
Key Innovation: This paper proposes a transferable Map of Dynamics (MoD) framework that generalizes to unknown environments using only egocentric 3D LiDAR point clouds to overcome the long-standing limitation of traditional MoD methods. Comprehensive experiments demonstrate that our method effectively reconstructs underlying motion maps even in unknown locations from a single instantaneous LiDAR scan, despite highly sparse training data.
46. Adapting Vision Foundation Models to Acoustics for Pose-Free 3D Sonar Reconstruction
Core Problem: Unfortunately, a lack of freely available large-scale sonar datasets makes training such a model from scratch impractical.
Key Innovation: An acoustic foundation model trained on large-scale sonar datasets could enable similar capabilities in the underwater domain, where turbidity and low-visibility conditions make conventional RGB foundation models inapplicable. In this work, we demonstrate that vision foundation models can be efficiently adapted to the sonar setting by (1) exploiting the geometric relationship between the two sensing modalities and (2).
47. Radiation, Rotation and Scale Invariant Feature Descriptor for Multimodal Image Matching
Core Problem: However, geometric distortions and nonlinear radiometric differences (NRD) severely limit performance, especially under radiometric, rotation, and scale variations.
Key Innovation: To address this issue, we propose a radiation, rotation, and scale invariant (RRSI) feature descriptor. Experiments on optical-infrared and optical-SAR datasets demonstrate highly competitive matching performance and strong robustness to rotation and scale variations.
48. AGSA-Net: Abundance-Guided Self-Attention Network for Spectral Unmixing-Aware Hyperspectral Remote Sensing Image Classification
Core Problem: However, its performance remains challenged by high spectral redundancy, noise sensitivity, and the difficulty of jointly modeling local material composition and long-range spectral dependencies.
Key Innovation: To address this, we propose AGSA-Net, an abundance-guided self-attention network that explicitly integrates spectral unmixing priors into the classification process. Experiments on Indian Pines, Augsburg, and Berlin demonstrate the benefit of incorporating abundance- guided contextual modeling, particularly in heterogeneous urban scenes.
49. Physico-Geospatial Grounded Scene Interpretation for Mobile Robotics
Core Problem: Recent advancements in deep learning allow robotic agents to interact with dynamic and unstructured environments.
Key Innovation: In the present work, we introduce an approach to augment the output of pre-trained, unmodified VLMs used for scene interpretation by integrating semantic descriptions, OpenStreetMap building data and street information with positional, temporal and metric information obtained from our sensory systems, fusing this information using LLMs.
50. Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach
Core Problem: However, existing methods exhibit two fundamental limitations. (1) Cross-Scope Context Entanglement.
Key Innovation: To address these challenges, we propose BRAIN, a unified model that focuses on graph context that combines neighborhood scope with modality composition. Experiments across nine datasets and four task families demonstrate its broad effectiveness, improving node-classification and link-prediction performance by up to 4.73% relative to the strongest baseline, while achieving an average relative improvement of 14.72% across four.
51. Drones as Annotators: Amodal 3D Auto-Labeling for Ground LiDAR with Aerial Priors
Core Problem: Existing auto-labeling methods reduce this burden, but most of them rely on onboard sensors, where a single ground-level viewpoint yields occluded and sparse observations and inaccurate object geometry.
Key Innovation: We introduce DAA (Drones as Annotators), a drone-assisted training-free framework for amodal 3D auto-labeling. We evaluate DAA on an air-ground cooperative perception dataset, where it consistently outperforms existing auto-labeling baselines.
52. PCSDiff: Diffusion-Based Bias Correction and Super Resolution Toward Practical Operational Medium-Term Precipitation Forecast
Core Problem: Existing AI correction techniques lack dedicated modeling for multi-day dynamic bias evolution and proper meteorological constraints, often generating over-smoothed rainfall structures, and cannot meet operational deployment demands.
Key Innovation: This work introduces PCSDiff, a cascaded task-decoupled diffusion framework targeting 10-day precipitation bias correction and downscaling. Evaluated against CMA-CRA observations over China after global-data training, PCSDiff cuts RMSE by 16.1% and lifts ACC by 13.9% relative to raw ECMWF forecasts at 3-10-day lead times, and consistently outperforms mainstream deep-learning baselines on both general and extreme-precipitation.
53. MSSP: Multi-Scale Spatially-Constrained Partition for Unsupervised Semantic Segmentation of 3D Point Clouds
Core Problem: Existing superpoint-based methods typically rely on spectral analysis at a fixed granularity, failing to capture the hierarchical semantic structures inherent in complex indoor scenes.
Key Innovation: To bridge this gap, we present a Multi-Scale Spatially-Constrained Partition (MSSP) framework that combines multi-scale spectral analysis with spatially-constrained clustering. Extensive experiments on S3DIS and ScanNet show that MSSP achieves the best mIoU among unsupervised methods on the main benchmarks, with particularly significant gains on S3DIS.
54. DPSF-Net: A Dual-Prior Spatial-Frequency Network for Real-World Remote Sensing Image Dehazing
Core Problem: Real-world remote sensing image dehazing (RSID) remains challenging because atmospheric scattering, spatially non-uniform haze and colour distortion jointly degrade structural and spectral information.
Key Innovation: Here, we propose DPSF-Net, a dual-prior spatial-frequency network built on MCAF-Net for real-world RSID. Extensive experiments demonstrate that DPSF-Net achieves state-of-the-art performance on the real-world RRSHID remote sensing image dehazing benchmark and remains competitive across multiple synthetic datasets.
55. AstraMoE-SR: Trajectory-Guided Diffusion for Blind Satellite Jitter Deblurring and Super-Resolution
Core Problem: Pushbroom satellite imaging couples limited spatial resolution with platform attitude instability.
Key Innovation: We present AstraMoE-SR, a single-image framework that jointly restores motion blur and spatial resolution without auxiliary measurements. We further show that the remaining point-wise trajectory error is consistent with intrinsic jitter-phase ambiguity that is not resolved by increasing estimator capacity.
56. PCFlow: Physics-Conditioned Flow Matching for GPR B-Scan Image Synthesis
Core Problem: Ground-penetrating radar (GPR) B-scan image synthesis is important for data augmentation, algorithm validation, and simulation acceleration, yet generating radargrams with both visual realism and physical consistency remains challenging.
Key Innovation: In this paper, we propose PCFlow, a physics-conditioned flow matching framework for fast GPR B-scan image synthesis. Experimental results show that PCFlow generates images with more accurate response geometry and high visual fidelity, demonstrating its effectiveness for controllable and physically faithful radar image synthesis.
57. Statistical versus machine learning-based spatial interpolation of post-processed ensemble weather forecasts
Core Problem: Statistical post-processing improves ensemble weather forecasts, but generating calibrated predictions at locations without observations remains challenging.
Key Innovation: This study compares statistical and machine-learning-based methods for post-processing ECMWF 2-m temperature and 10-m wind speed forecasts at observed and unobserved stations in Germany. The results show that post-processing improves upon the raw ensemble in most settings, but no single method performs best across all variables, station groups, and evaluation metrics.
58. Topologically Consistent Agricultural Parcel Vectorization with Semantic-Guided Diffusion and Topology-Aware Polygonization
Core Problem: Yet this requirement remains largely unresolved: segmentation-based methods mainly produce parcel masks or raster boundary cues and rely on heuristic raster-to-vector conversion, instance- and contour-based methods reconstruct parcels independently, and recent vector-oriented methods improve polygon regularity but do not explicitly recover adjacent parcels from a shared topological structure.
Key Innovation: To address this gap, we propose a semantic-guided diffusion framework for topologically consistent agricultural parcel vectorization. The results show strong and competitive performance, with zero measured intrusion ratio and the highest shared-edge recall, demonstrating the potential of the proposed framework for accurate, regular, and topologically consistent agricultural parcel vectorization.
59. Solution for UCF UrbanTwin LUMPI Track: Sim-to-Real Urban LiDAR 3D Object Detection
Core Problem: We present our solution to the LUMPI track of the UCF UrbanTwin Sim2Real LiDAR Challenge at the 6th DriveX Workshop, ECCV 2026.
Key Innovation: We present our solution to the LUMPI track of the UCF UrbanTwin Sim2Real LiDAR Challenge at the 6th DriveX Workshop, ECCV 2026.
60. Solution for UCF UrbanTwin V2X-Real Track: Sim-to-Real Urban LiDAR 3D Object Detection
Core Problem: Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, including scene geometry, sampling density, return patterns, and pedestrian scale.
Key Innovation: This report presents a multi-source collaborative training and class-aware fusion framework for Sim2Real 3D detection. On the UrbanTwin V2X-Real hidden test set, the unified system achieves a combined score of 0.7421, with 3D mAP@0.5 of 0.4518 and a realism score of 0.8871.
61. TFTrack: A Template-Free Framework for Efficient 3D Point Cloud Tracking
Core Problem: LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to localize dynamic objects across frames in sparse point clouds.
Key Innovation: Based on the above findings, we propose the first Template-Free Tracking framework (TFTrack). Our in-depth analysis reveals: (i) the template paradigm is redundant, as the previous bounding box center encodes sufficient historical context; (ii) complex motion modeling is unnecessary, as geometric alignment provides adequate motion priors.
62. Cross-modal learning for SAR target recognition using optical vision foundation models
Core Problem: However, Automatic Target Recognition (ATR) remains a challenging problem due to limited labelled data, the strong speckle in SAR images and the significant domain gap between SAR and more abundant optical imagery.
Key Innovation: We propose a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs.
63. JEDI: JEPA-to-Edge Distillation for Efficient Cropland Segmentation from Satellite Imagery
Core Problem: Large vision models provide useful representations for remote-sensing segmentation but are often too expensive for deployment at the satellite or field edge.
Key Innovation: We introduce JEDI (JEPA-to-Edge Distillation), a two-stage framework that transfers representations from a large I-JEPA Vision Transformer teacher to a compact SegFormer student. On CalCROP21, JEDI-B0 achieves 68.0 mean Intersection-over-Union (mIoU) with 4.04M parameters, improving over the standalone student by 16.0 points and coming within 2.0 points of the 70.0 mIoU achieved by the 639M-parameter teacher.
64. Solving the Elastic Wave Equation with Physics-Informed Neural Networks: A Robust and Critical Assessment
Core Problem: While promising, PINNs are not a panacea; they inherit challenges such as spectral bias and unstable convergence.
Key Innovation: This presents a new paradigm compared to traditional discretization methods and purely data-driven machine learning techniques. We find that integrating an understanding of wave physics into the network design significantly improves accuracy.
65. Hyperspectral Anomaly Detection via Group Sparse Low-Rank Tensor Factorization With Automatic Anomaly Grouping
Core Problem: However, existing methods still suffer from high computational cost and limited flexibility in characterizing spatially structured anomalies.
Key Innovation: For anomaly modeling, a latent grouping map is introduced to build an automatic anomaly grouping penalty, allowing anomaly groups to be adaptively inferred from the data rather than predefined at the pixel level. Experimental results on five real hyperspectral datasets demonstrate that the proposed method achieves superior detection performance and competitive computational efficiency compared with several state-of-the-art.
66. Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking
Core Problem: However, existing benchmarks do not jointly provide radar measurements, dense moving-instance masks, and temporally consistent identities for surveillance.
Key Innovation: We therefore introduce RGBTR-Motion, a synchronized and calibrated fixed-camera benchmark that pairs RGB, thermal, and radar streams with dense instance masks and temporally consistent identities across diverse surveillance scenes.
67. From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMs
Core Problem: Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual question answering over satellite imagery, yet localizing specific objects or changed regions remains challenging.
Key Innovation: In this work, we present an RS-specific formulation of the region selection paradigm, previously explored in natural-image MLLMs, and extend it to temporal change localization over multi-image sequences. Experiments show that our approach substantially outperforms coordinate-generation baselines on temporal change localization, while improving single-image visual grounding and maintaining competitive understanding performance.
68. Stochastically Perturbed Weights: Ensembles from Deterministic Machine-Learning Weather Models
Core Problem: Many deployed MLWMs are deterministic, producing a single forecast with no estimate of its own uncertainty, whereas a growing family of trained-probabilistic models generate calibrated ensembles directly, at the price of a dedicated training run.
Key Innovation: Many deployed MLWMs are deterministic, producing a single forecast with no estimate of its own uncertainty, whereas a growing family of trained-probabilistic models generate calibrated ensembles directly, at the price of a dedicated training run.
69. AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation
Core Problem: However, existing zero-shot methods typically operate at a single spatial scale, relying either on local representations constructed online from current observations or on global memories built offline from historical experience.
Key Innovation: To address this limitation, we propose AirAnchor, a new paradigm that bridges local and global spatial information through spatial anchors and integrates both into a shared navigation framework, enabling comprehensive spatial grounding for decision-making. Extensive experiments on AerialVLN demonstrate that AirAnchor substantially outperforms existing zero-shot baselines, validating the effectiveness and efficiency of the.
70. AXS-Net: Interpretable Deep Unfolding for Hyperspectral Image Denoising via Spectral Basis Unmixing and Structured Noise Refinement
Core Problem: We instead model HSI denoising as \Y=\A\X+\Snoise+\Nnoise, where \A\X is a low-rank spectral-subspace (unmixing) reconstruction, \Snoise is structured sparse noise and \Nnoise is residual Gaussian noise.
Key Innovation: The resulting regularized optimization problem is unrolled into AXS-Net, a K-stage alternating proximal-point framework. Across ICVL, CAVE, and Harvard datasets and five noise configurations, the proposed AXS-Net achieves strong in-domain accuracy and competitive zero-shot transfer, with consistent gains across all five noise regimes on ICVL and Harvard.
71. Interpretable Hyperspectral Unmixing Framework with Fixed Endmember Prior and Structured Residual Refinement
Core Problem: Hyperspectral unmixing decomposes mixed pixels into material endmembers and their abundances from contiguous spectral observations.
Key Innovation: This study presents an interpretable stage-wise hyperspectral unmixing framework (I-HyperSU) under fixed endmember priors, which is explicitly decomposed into a fixed endmember matrix \mathbf{A}, an abundance block \mathbf{X}, and a structural residual refinement block \mathbf{S}. Experiments on Samson, Urban, and Jasper Ridge datasets demonstrate that, under fixed and imperfect endmember priors, soft abundance relaxation.
72. DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models
Core Problem: We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders.
Key Innovation: We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. Extensive experiments on KITTI odometry and Boreas demonstrate strong performance and robustness across seasons, weather, and day/night.
73. Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild
Core Problem: However, visual information is encoded in an angular domain, whereas the physical world is represented in Cartesian coordinates.
Key Innovation: To delve into this challenge, we introduce Spheriverse, comprising 64,400 temporally aligned spherical image-LiDAR pairs organized into 644 sequences. SphereOcc achieves 13.91% mIoU and 24.65% GeoIoU, outperforming the respective best-performing methods, TPVFormer and SurroundOcc, by 1.70 and 2.10 percentage points.
74. GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting
Core Problem: Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D.
Key Innovation: This embeds points in a joint vision-language space known to behave like a bag-of-words on compositional tasks.
75. Data-driven rational function neural networks: a new method for generating analytical models of rock physics
Core Problem: However, construction of a theoretical model requires careful physical considerations and mathematical derivations, which means a long research process.
Key Innovation: Rock physics models have long been the focus of predicting wave velocity.
76. Reliable iToF Depth Sensing via Sensor-Intrinsic Uncertainty Modeling and State-Space Restoration
Core Problem: Indirect time-of-flight (iToF) cameras provide compact and cost-effective dense depth measurements, but their ranging accuracy is often degraded by sensor-intrinsic uncertainty under practical imaging conditions.
Key Innovation: To address this problem, we propose a joint depth-uncertainty modeling and restoration framework for reliable iToF sensing. Controlled comparisons with fixed and range-aware Gaussian noise, together with evaluations on U-Net, Restormer, and DVSS, further demonstrate that the proposed synthesis consistently benefits different restoration backbones.
77. Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts
Core Problem: However, the effects of data volume, feature selection, and data preprocessing on the performance of such power prediction models have not been thoroughly studied.
Key Innovation: Therefore, this study developed a baseline Linear Regression for performance comparison with a more complex Artificial Neural Network model to predict the power output of a wind turbine, using weather conditions only to enhance applicability. The best results from the different models showed that the Artificial Neural Network models provided the highest accuracy, with an R² score of 0.98 and a low Mean Absolute Error of 194.
78. CAVEAT: Recurrent Multimodal Diffusion Planning for Mapless Aerial Exploration
Core Problem: Can exploratory UAV waypoint sequences be generated from multimodal onboard observations and a fixed-dimensional recurrent internal state without maintaining a persistent global map in the deployed policy?
Key Innovation: We investigate this question through CAVEAT, a diffusion policy conditioned on a recurrent internal state updated from fused LiDAR, visual, and pose features and trained from trajectories generated by the map-based FUELv2 expert. Simulation results evaluate both inference mechanisms and compare CAVEAT with its demonstration-generating expert.
79. Diagnosing and Dynamically Filtering Occupancy World Models for Active Mapping
Core Problem: Active mapping requires a robot to select camera viewpoints that efficiently reconstruct an unknown 3D scene.
Key Innovation: To reason about unobserved regions, recent systems use pretrained occupancy networks as world models that complete missing geometry. Our experiments show that correcting false positives or false negatives alone does not consistently improve final coverage.
80. D3ARC: Time-Critical Distributed Disaster Detection for Asynchronous Cooperative Multi-Robot Systems
Core Problem: In time-critical crises such as wildfires, traditional monitoring practices remain limited by coverage, cost, and personnel risk, paving the way for autonomous and adaptive monitoring solutions.
Key Innovation: Within this context, this paper introduces D3ARC, an asynchronous distributed hierarchical framework for time-aware and reliable wildfire detection. The framework is evaluated through realistic robotics simulations, ablation studies, and baseline comparisons, achieving an overall mission success up to 94% with 89.4% detection confidence.
81. A Systematic Analysis of Automatic Differentiation versus Discretization-based Constraints for Physics-Informed PDE Solvers
Core Problem: However, the inherent limitations of AD, particularly in handling higher-order derivatives and discontinuous solutions, pose significant challenges for complex problems.
Key Innovation: Automatic differentiation (AD) plays a central role in this paradigm, which is mesh-free and replaces traditional iterative solvers with gradient-based optimization in continuous space. Our results reveal a consistent trend: as nonlinearity strengthens, the accuracy advantage of discretization-based constraints becomes increasingly pronounced, with smaller optimization errors compensating for the truncation errors.
82. Topographic Disorder, Wind Coupling, and Directional Fire Spread: Critical Behavior in a Terrain-Weighted Forest Fire Model
Core Problem: For rough terrain and low tree density, the fire fails to percolate even at zero suppression.
Key Innovation: We introduce the Terrain-Weighted Forest Fire Model (TFFM), a lattice model in which fire spreads on a spatially correlated Gaussian height field with the asymmetric bond probability p_i\to j}=clip[e^{-β+\gamma(h_j-h_i)},0,1], plus an additive wind bias.
83. Neural Posterior Estimation for Tomographic Weak Lensing Mass Mapping
Core Problem: Inferring shear and convergence from images is a challenging inverse problem.
Key Innovation: As an alternative, we propose a probabilistic approach to field-level weak lensing inference in which we train a deep neural network to directly map a multiband image to a variational distribution over the underlying tomographic shear and convergence fields.
84. Rethinking Learned Occupancy in Autonomous Active Mapping with Observation-Gated Filtering
Core Problem: Autonomous 3D active mapping requires a space robot to choose where to sense while building the geometry needed for navigation.
Key Innovation: Guided by this diagnosis, we introduce an observation-gated filter that retains completion in insufficiently observed regions and suppresses predictions only after repeated frustum exposure without nearby RGB-D support. These results motivate online revision of planner-facing geometry during autonomous intervals between communication windows.
85. Towards Vision-Language Geo-Foundation Model: A Survey
Core Problem: However, most methods rely on training with general image datasets, and the lack of geospatial data leads to poor performance on earth observation.
Key Innovation: In particular, we introduce the background and motivation behind the rise of VLGFMs, highlighting their unique research significance.
86. L2G-Map: Local-to-Global Mapping via Hierarchical Diffusion Refinement and Elliptical Bayesian Fusion
Core Problem: However, local-to-global mapping under visual conditions confronts two fundamental challenges: single-shot local observations are susceptible to viewpoint variation and environmental interference, leading to geometric deviations, while multi-source local information exhibits heterogeneous confidence, rendering globally consistent aggregation difficult.
Key Innovation: To address these, this paper proposes L2G-Map, a framework comprising hierarchical prior diffusion refinement and elliptical space Bayesian fusion. Under sensor-degraded conditions, a 3.27% mIoU gain is achieved.
87. Flow-Matching Guided Deep Unfolding for Hyperspectral Image Reconstruction
Core Problem: Reconstructing a three-dimensional hyperspectral cube from a two-dimensional compressed measurement is a severely ill-posed inverse problem.
Key Innovation: This paper proposes FMU, a deep unfolding framework that couples a measurement-conditioned flow-matching prior with a sensing-model-guided measurement update. On the KAIST 10-scene benchmark under the optical-filter setting, FMU obtains 42.13dB PSNR and 0.9900 SSIM, outperforming LADE-DUN by 1.16dB in PSNR under the same training data, sensing mask, and evaluation protocol.
88. DAGLFNet: Deep Feature Attention Guided Global and Local Feature Fusion for Pseudo-Image Point Cloud Segmentation
Core Problem: Efficiently processing unstructured point clouds while extracting structured semantic information remains a significant challenge.
Key Innovation: This work proposes DAGLFNet, a pseudo-image-based semantic segmentation framework designed to extract discriminative features. Experimental evaluations demonstrate that DAGLFNet achieves mean Intersection-over-Union (mIoU) scores of 69.9% and 78.7% on the validation sets of SemanticKITTI and nuScenes, respectively.
89. Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark
Core Problem: However, spatial misalignment commonly exists between RGBT image pairs.
Key Innovation: To address this, we propose a Dual-Correlation Hypergraph Network (DCHNet) that captures high-dimensional complementary information by explicitly modeling two types of correlations: temporal correlation across consecutive frames and spatial correlation from cross-modal features. Comprehensive experiments on VT-VOD50 and our DVT-VOD1000 demonstrate that DCHNet achieves state-of-the-art detection accuracy.
90. Nonlinear Response History Soil-Structure Interaction Studies on UCSF Health Helen Diller Hospital Building
Core Problem: However, SSI is rarely incorporated into nonlinear response history analyses, particularly in performance-based seismic design.
Key Innovation: However, SSI is rarely incorporated into nonlinear response history analyses, particularly in performance-based seismic design. Additionally, results demonstrate the important dynamic higher mode effects of a massive and large footprint pile foundation.
91. A global hourly ISIMIP3 climate forcing dataset for impact modeling
Core Problem: Abstract.
Key Innovation: Sub-daily climate data are increasingly important for climate-impact assessments because many processes, such as heat stress, hydrological extremes, land-surface energy balance, and renewable-energy production, respond non-linearly to intra-day variability.
92. 3D crustal model of Southern Apennines (Italy) by integrating geological, geophysical and petrophysical constraints
Core Problem: Abstract.
Key Innovation: In this study, we present a high-resolution 3D crustal model of the Southern Apennines, parameterized on a 5 × 5 × 1 km³ grid and constructed through the systematic integration of geological, geophysical, and petrophysical data. All datasets are openly accessible under FAIR principles, promoting transparency, reproducibility, and interdisciplinary reuse of the results.
93. Spatiotemporal Dynamics and Nonlinear Associations of Eco-Environmental Quality in Mining Areas Using MRSEI and an Object-Based XGBoost-SHAP Framework
Core Problem: Long-term mining activities can cause complex ecological and environmental problems, including mineral surface exposure, dust-related disturbance, and vegetation degradation, while conventional ecological indices may not fully characterize the distinctive ecological conditions of mining areas.
Key Innovation: This study developed a mining-area Modified Remote Sensing Ecological Index (MRSEI) by incorporating the Lithological Mineral Index (LMI), Normalized Dust Difference Index (NDDI), and Vegetation Health Index (VHI) into the Remote Sensing Ecological Index (RSEI) framework.
94. An Intelligent Method for Ice Thickness Identification Using Drone-Borne Ground-Penetrating Radar
Core Problem: However, they fail when the radar signal lacks a clear bottom reflection-a common condition in ice layers containing unfrozen water-and manual interpretation remains time-consuming.
Key Innovation: Existing algorithms can extract ice layer boundaries by tracking continuous bottom reflections in GPR images. Field validation against drilling measurements demonstrates that the model achieves Intersection over Union (IoU) of 97.12% and an F1-score of 98.54% for ice layer identification, with a relative error in ice thickness measurement below 3% based on five borehole measurements.
95. Stratified Spatiotemporal Residual Detection of Weak Active-Fire Anomalies from VIIRS 375 m Time Series
Core Problem: Satellite active-fire products may miss fires that occupy only a small fraction of a pixel and produce limited absolute thermal responses.
Key Innovation: In this study, weak thermal anomalies are operationally defined as confirmed fire-affected VIIRS pixels with relatively low absolute BT4 and ΔBT responses but positive deviations from their recent temporal and local spatial backgrounds. These results show that STAR-FD provides complementary detection of confirmed fire-related thermal anomalies with weaker absolute thermal signals.
96. Weak-Observation-Aware Multi-Object Tracking in Satellite Video with Temporal Evidence and Trajectory Reliability
Core Problem: Multi-object tracking (MOT) in satellite video is fundamentally limited by low target observability: genuine targets often produce weak and unstable responses, while structured backgrounds can generate persistent target-like interference, leading to missed detections, fragmented trajectories, and identity switches.
Key Innovation: To address this ambiguity, we propose a weak-observation-aware and reliability-guided framework with a layered two-stage design: the front end enhances weak observations over short temporal windows, whereas the back end controls their use for long-term trajectory association and state updating according to their reliability. On VISO, the proposed method achieves a multiple object tracking accuracy (MOTA) of 72.5% and an.
97. Station-Based Evaluation of AI Weather Models for Near-Surface Temperature, Pressure, and Wind Forecasts over Eastern Coastal China
Core Problem: However, the station-level performance of global artificial intelligence (AI) weather models remains insufficiently characterized in complex coastal environments.
Key Innovation: This study evaluated Pangu-Weather, FengWu, FuXi, and the Global Forecast System (GFS) against observations from 210 stations in eastern coastal China from July to December 2022. The models showed distinct spatial error patterns, and wind-speed errors were concentrated at several northern coastal and transition-zone stations.
98. Analysis and selection of seismic intensity measures for railway simply-supported-bridge-vehicle coupled systems subjected to crossing-strike-slip faulting
Core Problem: However, for the simply-supported-bridge-vehicle coupled system (SSBVCS) subjected to cross-strike-slipfaulting (CSSF), the optimal selection of IMs remains challenging due to the complex multi-directional effects of faulting, the influence of fault-bridge spatial relationships, and the scarcity of recorded ground motions.
Key Innovation: To address this, this study proposes a modified IM selection framework for cloud analysis (CA), which enhances the accuracy of efficiency evaluation through normalization of IMs. Application of this framework demonstrates that under the coupled fling-step and forward directivity effects of CSSF, peak spectral displacement (SDmax) and peak spectral velocity (SVmax) exhibit superior performance for probabilistic seismic demand.
99. Assessment of water seepage in the Orzepowice embankment, southern Poland, using repeated near-surface multi-method geophysical surveys
Core Problem: Water seepage through undetected subsurface heterogeneities is a leading cause of internal erosion and failure in ageing embankments.
Key Innovation: This study assesses potential seepage and under seepage, their changes after conservation works, and evaluates the structural integrity of the 680-m-long Orzepowice embankment in southern Poland using repeated near-surface geophysical imaging.
100. Physics-guided feature-enhanced clustering for rigorous stratification of ERT subsurface profiles
Core Problem: Electrical Resistivity Tomography (ERT) is vital for engineering geological investigations; however, transforming inverted resistivity profiles into reliable geological ground models remains challenging.
Key Innovation: To address these limitations and achieve rigorous geoelectrical stratification, we propose a computationally efficient, physics-guided feature-enhanced clustering framework.
101. Delineation of inaccessible coal barrier pillar and its stability assessment through hybrid ERT and numerical modeling approach
Core Problem: Demarcation of unapproachable coal barrier between two adjacent mines intersected by old mine workings is a key issue for mine safety and flood-risk management.
Key Innovation: Electrical Resistivity Tomography (ERT) integrated with numerical modeling was applied to evaluate the status, thickness, and stability of the coal barrier pillar between Moonidih and Bhagabandh collieries of Jharia Coalfield, India. ERT results reveal two prominent low-resistivity zones at surface distances of approximately 140-300 m and 470 m-670 m, extending to depths of 40 m − 157 m, which were interpreted as water-logged.
102. Multiscale deterioration of fault-zone tectonic soft rock under wetting-drying cycles: A coupled mineral-structure-seepage-strength perspective from the Qinghai-Tibet Plateau
Core Problem: However, its multiscale deterioration mechanism remains insufficiently understood.
Key Innovation: In this study, TSR from the Pingding-Huama fault zone was investigated using X-ray diffraction, scanning electron microscopy, X-ray computed tomography, nuclear magnetic resonance seepage tests, and triaxial compression experiments. The results show that WDC causes preferential depletion of calcite and clay minerals, whereas quartz remains relatively stable, leading to weakened cementation and a transition from.
103. Hybrid Mechanism- and Data-Driven Model Updating Method for Structural Health Monitoring
Core Problem: In structural health monitoring (SHM), unavoidable discrepancies between numerical simulations and real structural responses stem from multi-source uncertainties, which severely reduce model reliability and drive the demand for model updating.
Key Innovation: To address these issues, this paper proposes a hybrid mechanism- and data-driven model updating framework. Numerical and experimental validations demonstrate that the updated hybrid model achieves high prediction accuracy.
104. Impacts of Climate Change on Pavement Performance: A Review of Stressors, Modeling, and Adaptation Strategies
Core Problem: To address these challenges, this study presents a systematic review of 99 peer-reviewed publications from 2000 to 2025, synthesizing an integrated analytical framework that traces the causal chain from climatic stressors through modeling approaches to pavement-level performance impacts and adaptation strategies.
Key Innovation: To address these challenges, this study presents a systematic review of 99 peer-reviewed publications from 2000 to 2025, synthesizing an integrated analytical framework that traces the causal chain from climatic stressors through modeling approaches to pavement-level performance impacts and adaptation strategies.
105. How sensitivity can support infrastructure systems under deep uncertainty: From hazard agnostic vulnerability to flexibility
Core Problem: Critical infrastructure systems, from transport to energy systems, are becoming increasingly exposed to operational risks driven by evolving uncertain operational conditions.
Key Innovation: Many of these systems are still designed using fixed assumptions that may not hold in the future. Three illustrative examples, a parallel system, a wind turbine and a traffic network, are used to showcase the importance of understanding sensitivities in safety-driven systems.
106. Resilience enhancement of multi-carrier microgrids under extreme weather: A distributionally robust operational strategy with actuarial risk
Core Problem: Extreme weather can simultaneously perturb energy supply, demand, and infrastructure availability in multi-carrier microgrids (MCMGs), while physical service loss and its monetary consequence may differ across load classes.
Key Innovation: This paper develops a two-stage data-driven Wasserstein distributionally robust optimization (WDRO) framework that separates physical service adequacy, actuarial valuation, tail-risk preference, and meteorological distributional ambiguity. Sensitivity analyses show stable held-out service outcomes across most tested perturbations, whereas stronger external distribution shifts cause marked deterioration.
107. Out-of-plane seismic performance of precast composite sidewalls with grouted lap-splice connections in utility tunnels
Core Problem: It was found that the failure modes of the precast composite sidewall and the cast-in-place sidewall were different.
Key Innovation: This study investigates the out-of-plane seismic performance of precast composite sidewalls with spiral stirrup sleeve grouted lap-splice connections at the bottom joint. Full-scale cyclic tests were conducted on three precast composite sidewalls with axial compression ratios of 0.05, 0.10, and 0.15, and the results were compared with those of a monolithic cast-in-place sidewall.
108. Analytical and DEM-FDM numerical analysis of lateral deformation of an adjacent two-pile group induced by shield tunnel excavation
Core Problem: Shield tunnel excavation can induce ground loss and lateral soil movement, which may cause lateral deformation and additional internal forces in adjacent pile groups.
Key Innovation: This study proposes an analytical solution for the lateral response of an adjacent two-pile group induced by shield tunnel excavation while considering the shielding effect between piles. The proposed solution is validated against centrifuge test results and compared with a Winkler-based solution.
109. Energy-based prediction of blast-induced rock fragmentation using interpretable machine learning
Core Problem: Accurate prediction of blast-induced rock fragmentation is critical for optimizing downstream mining operations; however, conventional empirical models often fail to capture the nonlinear and coupled effects of blast design, rock properties, and energy distribution.
Key Innovation: This study presents an energy-based, interpretable machine learning framework for predicting mean fragment size (X50) in open-pit blasting. GLMNET demonstrated the most balanced performance, achieving a test R² of 0.79 with low prediction error.
110. Effects of water immersion and surface moistening on the shear behavior of rough sandstone fractures: Insights from true triaxial shear tests
Core Problem: Water-rock interactions alter the shear behavior and frictional stability of fractured rock masses, potentially inducing catastrophic failure such as landslide, tunnel collapse and water inrush.
Key Innovation: This study provides new insights into the micro-to-macroscale mechanisms governing the failure of fractured sandstone under different water immersion and surface moistening conditions. Results show that water-induced strength degradation is more sensitive to immersion extent at 24 h, while such sensitivity diminishes with prolonged immersion.
111. Mechanism of tightness enhancement and risk assessment for gas storage salt cavern based on mechanical-seepage coupling effects
Core Problem: The long-term tightness of gas storage salt caverns is governed by the creep and permeability evolution of rock salt, the dynamic evolution of which directly affects operational safety.
Key Innovation: This study investigates the associated mechanisms of tightness enhancement. Using the self-developed SalLeak-RGV simulation platform, the study reveals that gas migration is controlled by geological structures, with micro-permeable interlayers acting as the dominant preferential pathways.
112. Unraveling meteorological contributions and lagged effects on multi-depth soil moisture prediction using Transformer-SHAP framework
Core Problem: Here, meteorological contributions refer to the SHAP quantified contributions of individual predictors to SM predictions, whereas lagged effects refer to the historical period over which antecedent meteorological information remains useful for predicting current SM.
Key Innovation: We developed separate Transformer models for the 5, 20, 45, and 80 cm soil layers and identified the optimal historical input window for each depth. The models achieved NSE values of 0.82, 0.78, 0.70, and 0.90, respectively.
113. Depth-dependent freeze-thaw propagation and antecedent controls on post-thaw soil water storage in a seasonally frozen basin
Core Problem: Seasonally frozen soils undergo vertically asynchronous freeze-thaw transitions, but the relative roles of antecedent wetness and winter thermal conditions in shaping post-thaw profile storage remain unclear.
Key Innovation: We analyzed observations from 34 sites in the Shandian River Basin over three cold seasons (2019-2022), with soil temperature and water content monitored at 3, 5, 10, 20, and 50 cm. Between-site persistence dominated the 0-20 cm relationship (β = 0.857 versus 0.137 within sites), whereas the 0-50 cm profile showed a weaker between-site coefficient and a larger, window-sensitive within-site coefficient (β = 0.600 and 0.329).
114. Impact of high-speed train vibrations on subsoil and built environment: Engineering-geological case studies from Poland
Core Problem: This engineering-geological study evaluates the vibrational impact of high-speed rail traffic on subsoil and its implications for future buildings, in the context of environmental protection and sustainable infrastructure development.
Key Innovation: The research concerned a planned high-speed railway line from Warsaw to Vienna, on the section between Zawiercie and Grodzisk Mazowiecki in Poland. Results reveal a clear relationship between increasing train speed and elevated vibrational acceleration in the subsoil, exceeding 5 cm/s2, which suggests that current safety limits and mandatory measurement distances of 25 m may not adequately reflect real conditions.
115. Building the Andean Crust at ∼35°S: Arc Geochemistry Reveals Cretaceous Thinning and Rapid Neogene Thickening
Core Problem: Quantification of the timing and mechanisms of crustal thickening in Cordilleran orogenic hinterlands remains challenging.
Key Innovation: We address this in the Andes using whole-rock geochemistry and zircon petrochronology from a more than 70 Myr volcanic record in the Tinguiririca valley of Chile (∼35S), which captures a pivotal transition from extensional to contractional deformation. Zircon (Sm, Gd, Dy)/Yb ratios generally track trends in whole-rock geochemical crustal thickness (paleomohometry) models across this record, with Dy/Yb showing the most.
116. Freshwater-Induced Surface Cooling of the Subpolar North Atlantic in a Large-Ensemble, High-Resolution, Coupled Model
Core Problem: While much attention has been paid to the expected long-term weakening of the Atlantic Meridional Overturning Circulation, the impact of freshwater on surface fluxes and seasonal to decadal air-sea coupling remains poorly understood.
Key Innovation: In recent decades, the region has experienced pronounced freshening.
117. The Impact of ENSO on Euro-Atlantic Circulation a Year Later via Long-Lived Tropical SST Anomalies
Core Problem: The second winter response to ENSO in the Euro-Atlantic sector is a surprising recent discovery, however, the underlying mechanisms remain uncertain and are challenging to determine from observations.
Key Innovation: Here we revisit the second winter response to ENSO using multiple large ensemble climate model simulations to examine its robustness and possible causes. Almost all the models demonstrate a significant second winter response in the Euro-Atlantic sector, typically resembling positive phases of the North Atlantic Oscillation and/or East Atlantic pattern, but with a large variation across models.
118. How can the Lateral Profile of Velocity Be Predicted on a Sloping Riverbank With Vegetation?
Core Problem: However, the impacts of flow, vegetation and side sloping banks on the velocity profile are unclear, and predicting the velocity profile on a vegetated sloping bank is challenging.
Key Innovation: In this study, rigid cylinders were used to mimic bank vegetation, and laboratory experiments were performed to clarify the mechanisms through which the lateral profile of depth-averaged velocity changes in a trapezoidal channel with a vegetated sloping bank. The velocity prediction results were consistent with the measurements.
119. Bayesian Joint Velocity and Impedance Inversion via Diffusion Models Conditioned on Common Image Gathers
Core Problem: We present a multi-parameter simulation-based inference framework for joint Bayesian recovery of subsurface velocity and acoustic impedance from seismic data.
Key Innovation: We present a multi-parameter simulation-based inference framework for joint Bayesian recovery of subsurface velocity and acoustic impedance from seismic data. On the Compass benchmark the model achieves velocity SSIM of 0.967 (RMSE 0.050 km/s) and impedance SSIM 0.867 (RMSE 0.279 km/s g/cm³), with velocity quality confirmed by CIG focusing.
120. An overview of 3D Vision-Language Models
Core Problem: Traditional 3D deep-learning models, however, are typically trained for specific tasks, such as classification, segmentation, or detection, and do not naturally support cross-modal retrieval from their embedding spaces using text or images as queries.
Key Innovation: This tutorial provides an overview of 3D VLMs, ranging from basic definitions of 3D representations and their encoding into embeddings to cross-modal contrastive alignment, modern multimodal frameworks, and 3D Vision-Large Language Models (3D VLLMs).
121. AAMBERS-UAV: Acquisition-Aware Multimodal Backbone Evaluation and Ranking for UAV Weedy Rice Segmentation
Core Problem: Using the 734-sample WeedyRice-RGBMS-DB, we fix a 124-image target-acquisition test set and compare two protocols with identical train, validation, and test counts: target-held-out, which excludes the target acquisition from development, and target-exposed, which admits its remaining images.
Key Innovation: Such splitting can place samples from one acquisition in both model development and testing, obscuring transfer to a genuinely new survey. A supplied-split audit reveals strong near-sequential dependence, and corruption tests show that early fusion is substantially more sensitive to RGB--MS displacement than to moderate radiometric scaling.
122. Seismic P-wave attenuation estimation based on frequency-dependent AVO using Kramers-Kronig relations for gas reservoir prediction
Core Problem: Estimation of seismic attenuation (inverse quality factor) is important for gas reservoir prediction.
Key Innovation: To address these issues, this study, within the framework of isotropic linear viscoelastic media, starts from the Kramers-Kronig relations and expresses the viscoelastic stiffness matrix as a function of seismic attenuation. Synthetic seismic data tests show that the estimated P-wave attenuation attribute is sensitive to variations in reservoir gas saturation, with reservoirs of higher gas saturation exhibiting stronger.
123. Efficient and Robust Camera-independent Multiview 3D Geometric Reconstruction from Noisy Monocular Depth Estimation and Multiple Point Matching
Core Problem: We present an efficient and robust method for 3D geometric reconstruction that is based solely on the camera-independent linear relationships among a given set of points, which are stable over time and robustly estimated using multiple point matches.
Key Innovation: We present an efficient and robust method for 3D geometric reconstruction that is based solely on the camera-independent linear relationships among a given set of points, which are stable over time and robustly estimated using multiple point matches. We also show that the principal eigenvectors of \mathbf{W}, which all have eigenvalue 1, provide a homogeneous representation of the 3D point configuration.
124. Depth-to-Image Synthesis-Driven Generative Unguided Depth Completion
Core Problem: Guided depth completion methods heavily depend on RGB quality and alignment, while unguided ones often suffer from limited precision due to the absence of explicit visual cues.
Key Innovation: In this paper, we present Depth-to-Image Synthesis-Driven Generative Unguided Depth Completion (GUDC), a new completion paradigm that innovatively bridges advanced 2D generative models with unguided depth completion, enabling semantics-aware depth inference without real RGB inputs. Extensive experiments on KITTI and NYUv2 validate that our GUDC achieves superior accuracy and robustness over existing methods.
125. Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation
Core Problem: This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026.
Key Innovation: The challenge evaluates video segmentation in three complementary settings: complex semi-supervised video object segmentation on MOSEv2, text-guided referring video object segmentation on MeViSv2-Text, and audio-guided referring video object segmentation on MeViSv2-Audio.
126. PLSR: Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion
Core Problem: Existing state-of-the-art diffusion-based models remain constrained by fixed resolutions, limiting their ability to produce details.
Key Innovation: In this paper, we tackle the challenge of generating more detailed, higher-resolution 3D objects by introducing a 3D super-resolution (SR) framework built on existing 3D generative foundation models. To this end, we design PLSR, a progressive and localized super-resolution solution to achieve this goal effectively and memory efficiently.
127. One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware Waypoints
Core Problem: Our approach introduces a training-free structured waypoint generator and a novel abstract representation that projects sparse, history-aware candidate waypoints directly onto RGB images as visual markers.
Key Innovation: To address prohibitive inference latency and computational overhead, we propose O2C-Nav, an efficient zero-shot navigation framework that calls only a single large model once per decision step. Extensive evaluations on the R2R-CE and RxR-CE benchmarks demonstrate that O2C-Nav outperforms current state-of-the-art zero-shot methods, highlighting its great potential for real-time robotic deployment.
128. Role-Specific Predictive Geometries for Nonstationary Multivariate Graph-Signal Forecasting
Core Problem: In an error-correction representation, long-run equilibrium restoration and short-run transient propagation represent different predictive roles and need not share a common cross-feature geometry.
Key Innovation: We introduce role-specific predictive geometries in which directed Long relations act on estimated equilibrium coordinates, whereas directed Short relations act on lagged differences.
129. Reading Decoder Trajectories: Training-Free Counterfactual Query-Trajectory Reliability for Small-Object Detection
Core Problem: Small-object detection remains challenging because limited pixels cause information loss and suppress the scale knowledge encoded in pretrained detectors.
Key Innovation: To test this hypothesis, we propose Counterfactual Query-Trajectory Reliability (CQTR), a training-free framework that elicits latent responses through counterfactual scale interventions and interprets candidate reliability from decoder-internal spatial convergence, semantic persistence, and cross-scale conflicts. Closed-loop analyses further show that scale intervention activates latent responses, trajectory evidence.
130. Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models
Core Problem: Recent vision-language models (VLMs) appear capable of spatial reasoning, but their ability to infer a speaker's viewpoint from contextual cues and interpret situated spatial relations from that viewpoint remains unclear.
Key Innovation: Humans infer such perspectives from shared environmental knowledge, activity context, and commonsense. We find that explicit breakdowns of observer-relative spatial reasoning improve target localization.
131. NeuCME: Toward Dynamic Multimodal Continual Learning via Neural Combinatorics of Multiple Experts
Core Problem: However, existing methods typically assume that the set of modalities per task is predefined and fixed.
Key Innovation: To address these challenges, we propose NeuCME (as shorthand for Neural Combinatorics of Multiple Experts), a novel framework designed to effectively learn and integrate knowledge across tasks with varying modalities. Extensive experiments using four real-world datasets demonstrate that the proposed NeuCME outperforms state-of-the-art methods markedly.
132. Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer
Core Problem: Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fidelity meshes, yet their post-training strategy remains largely unexplored.
Key Innovation: Inspired by the practicability of on-policy distillation (OPD) in large language models and image generation, we propose Flow3D-OPD, a two-stage post-training framework that introduces multi-teacher distillation into 3D geometry generation. Without relying on elaborate modifications, our straightforward yet effective design achieves consistent improvements across all geometric quality dimensions and surpasses all teacher.
133. KODAMA: Multimodal Digital Twin Reconstruction for Urban RF Propagation Modelling
Core Problem: 3D reconstruction typically strives for geometric fidelity or visual plausibility.
Key Innovation: We present KODAMA, an automated pipeline that reconstructs ray tracing-ready RFDTs at city scale from off-the-shelf geospatial data alone: aerial imagery, LiDAR, and photogrammetry yield terrain and watertight building meshes, while exposure-weighted multi-view fusion of street-level imagery recovers fa\c{c}ade relief, electromagnetic materials, and clutter-all without site visits or calibration.
134. Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection
Core Problem: How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, weakly labeled, and resource-constrained?
Key Innovation: We introduce a lightweight federated MIL-VLM cascade in which only a compact MIL scorer is trained across clients, while a frozen VLM verifies high-scoring suspect segments post hoc. Experiments on UCF-Crime with InternVL3.5-2B and Qwen3-VL-2B-Instruct show that text-generation verification can improve frame-level AUC after diagnostic temporal post-processing, but remains sensitive to prompts, parsers, model choice, and.
135. PICANet: Physics-Informed Cascaded Asymmetric Network for Infrared Small Target Detection
Core Problem: Infrared small target detection (ISTD) is an important research direction in image processing.
Key Innovation: Unlike previous work, a multi-level cross-feature attention module with the cascaded asymmetric mechanism is introduced to achieve precise alignment between high-level semantics and low-level spatial details.
136. InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling
Core Problem: Multimodal large language models often capture visual-linguistic correlations but struggle to predict how local visual interventions propagate and affect downstream answers.
Key Innovation: We introduce InfluenceField, an intervention-aware latent field inserted between the visual encoder and language decoder. For a nonlinear finite-basis population model, we show that target-aligned interventional supervision, together with a one-step separation condition on the transition, restricts admissible representations to within-location reparameterizations, so that the directed dependency graph of the full transition.
137. Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation
Core Problem: We observe that pretrained image generators show measurable zero-shot perceptual competence, but with a clear trade-off: specialist models remain stronger for in-distribution accuracy and efficiency, while generative models are often more robust under distribution shift and better at compositional semantic reasoning.
Key Innovation: We introduce ProbeGen, a benchmark for zero-shot generative perception that casts monocular depth estimation, referring/reasoning segmentation, and object counting as conditional generation tasks specified through text prompts, and compares 20 models in total-including proprietary and open-weight image generators, specialist perception models, and MLLMs-across 11 published benchmarks.
138. TaskGuard: Task-Conditioned Restoration Utility for Risk-Aware Object Detection
Core Problem: Image restoration is commonly applied before object detection under adverse conditions, yet a visually improved image need not improve the downstream task.
Key Innovation: We introduce TaskGuard, a post-hoc controller for frozen restoration and detection pipelines. Exact regional counterfactuals reveal substantial within-image utility heterogeneity, while a deployable pseudo-gradient preserves statistically reliable directional information.
139. Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Core Problem: Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run.
Key Innovation: Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D.
140. Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models
Core Problem: Dream replay grounding replays this dream through captioning and re-imagination, training claim-level evidence to remain consistent across the replay while separating unrelated visual experiences.
Key Innovation: We introduce Generative Grounding Feedback(GGF), a self-evolving post-training framework that uses only text prompts and the model's own visual experience. Experiments across unified models with different understanding--generation integration designs show consistent improvements in text-to-image generation together with modest gains in visual understanding.
141. MARS-CLIP: Multi-Resolution and Attention Refined Zero-Shot Image Segmentation
Core Problem: Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in zero-shot transfer but often struggles with dense prediction tasks due to low spatial resolution and the loss of structural information.
Key Innovation: To address these limitations, we propose MARS-CLIP (Multi-resolution and Attention Refined Segmentation for CLIP), a novel framework for zero-shot semantic segmentation. A set of experiments on six public datasets demonstrates that MARS-CLIP significantly outperforms state-of-the-art methods.
142. HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting
Core Problem: Multi-scale modeling has become an effective approach for long-term time series forecasting, capturing temporal patterns that range from fine-grained local dynamics to coarse global trends.
Key Innovation: In this paper, we introduce HypLTSF, a framework that endows the multi-scale hierarchy with a concrete geometric form by embedding scale-wise representations into the Poincaré ball, whose exponentially expanding volume naturally accommodates hierarchical structures. Extensive experiments on long-term time series forecasting benchmarks show that HypLTSF achieves state-of-the-art performance, suggesting that explicitly modeling.
143. Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting
Core Problem: However, conflict and harm are not the same thing.
Key Innovation: We propose Per-Variable Surgery (PV-Surgery), an optimizer-side training strategy for backbones with cache-compatible layers. This averaging is convenient, but the optimizer sees only the aggregated gradient, which does not reveal whether the variable-wise contributions align or oppose one another.
144. MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models
Core Problem: However, embedding models are hard to scale up: increasing parameters directly tradeoffs for the large training batch size that contrastive learning needs, and retrieval has to be served under tight latency.
Key Innovation: In this work, we propose MOEMB, which instead scales UME along the expert axis through mixture-of-experts (MoE), growing encoder capacity while preserving single-vector, non-autoregressive encoding. Together, these results support expert scaling as an effective and efficient direction for UME, with adaptive computation further improving efficiency for MLLM-based embedding models towards large-scale retrieval and.
145. SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation
Core Problem: Existing MLLM-segmenter interfaces either use a special trigger or compress both signals into one context, although they receive different supervision and fail differently.
Key Innovation: We present SeGDeP, an explicit what-where interface. Controlled stage-wise ablations, gradient diagnostics, and prompt interventions further show that the two paths develop complementary semantic and geometric specialization rather than duplicating the same evidence.
146. SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
Core Problem: World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions.
Key Innovation: We introduce SyncWorld, an action-conditioned world model that serves as a zero-shot simulator across unseen environments without any additional training. SyncWorld leverages a visual calibration episode-paired frames and actions that showcase all the controllable degrees of freedom-to specify the setup-specific Action--Visual Mapping in context.
147. Rollcast: Proper-Score Gated Rolling Anchors for Adaptive Probabilistic Time-Series Forecasting
Core Problem: A state-dependent softmax gate learns anchor probabilities by minimizing negative log predictive density, while residual distributions retrieved from similar historical states provide local uncertainty.
Key Innovation: Rolling means, medians, extrema, regression endpoints, and quantiles define candidate forecast locations and a representation of the current state. The results indicate that Rollcast can construct competitive probabilistic forecasts from simple, interpretable local summaries, while also identifying limitations in calibration and recursive uncertainty propagation.
148. CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning
Core Problem: We present CUSP (Collective Uncertainty through Semantic Opinion Pooling), a training-free uncertainty quantification framework that maps multiple VLM responses to a shared semantic response space, pools them into a pooled semantic opinion, and reports two complementary system-level signals: collective uncertainty, the dispersion of the pooled opinion, and Jensen-Shannon divergence (JSD), the conflict among the model-level.
Key Innovation: We present CUSP (Collective Uncertainty through Semantic Opinion Pooling), a training-free uncertainty quantification framework that maps multiple VLM responses to a shared semantic response space, pools them into a pooled semantic opinion, and reports two complementary system-level signals: collective uncertainty, the dispersion of the pooled opinion, and Jensen-Shannon divergence (JSD), the conflict among the model-level.
149. Closed-Loop Evaluation of Bird's-Eye-View Maps from Cross-View Transformers as Inputs to Behavior-Cloning Policies
Core Problem: In autonomous driving, Bird's-Eye View (BEV) representations provide a structured, top-down abstraction of the vehicle's surroundings and have become a key input modality for Behavioral Cloning (BC) policies.
Key Innovation: While ground-truth BEV maps are readily available in simulation, real-world deployment requires replacing them with camera-predicted counterparts - a substitution that introduces perceptual errors whose downstream impact on closed-loop driving performance is not well understood. Closed-loop evaluation across two CARLA towns shows that the KDE-weighted model is the only predicted-BEV agent to complete a full episode without.
150. From Narrative to Auditable Forecasts: A Structured Scaffold for Agentic Forecasting
Core Problem: However, current agentic forecasting often relies on implicit narrative aggregation: agents collect evidence, discuss it in prose, and often assign a probability without an explicit update path from evidence to forecast.
Key Innovation: We propose AuditForecast, an agentic scaffold for structured probabilistic forecasting. Across multiple live forecasting benchmarks, AuditForecast improves forecasting accuracy and calibration relative to strong agentic baselines, surpasses market-implied references in several settings, and outperforms substantially more expensive deep-research agents while remaining Pareto-dominant in the cost--accuracy tradeoff.
151. Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
Core Problem: Reliable tool use in multimodal agents, however, remains challenging because models must interpret text and images while integrating noisy retrieved evidence, often under sparse outcome-level supervision without explicit verification signals.
Key Innovation: We present Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that trains multimodal agents to verify and filter retrieved evidence within their own reasoning traces, reducing reliance on external verifiers at inference time.
152. AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction
Core Problem: At its core, Ray-GPIS estimates direction-wise reconstruction uncertainty along candidate viewing rays and selects next-best-view targets using an uncertainty--novelty objective, which are realized through an axis-conditioned in-hand rotation policy.
Key Innovation: We propose AURORA, an active 3D reconstruction framework that closes the loop between online object-centric reconstruction and in-hand reorientation. Experiments demonstrate that AURORA improves reconstruction quality and information-acquisition efficiency over non-active rotation strategies, while Ray-GPIS also outperforms active view-planning baselines in reconstruction performance, action-ranking quality, and planning.
153. Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language Representation
Core Problem: The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of reasoning across heterogeneous sources such as text, tables, and images.
Key Innovation: In this paper, we present a comprehensive methodological comparison of three influential frameworks, namely Multimodal Adaptive Extraction (MAE), Solar, and UniMMQA, tracing the evolution of multimodal question answering from modality-adaptive pipelines to fully unified architectures. Our analysis demonstrates a clear shift from explicit modality-specific processing toward unified text-centric formulations enabled by.
154. ChatBEV: Empowering Traffic Scene Understanding and Simulation via Vision-Language Model
Core Problem: While VisionLanguage Models (VLMs) have demonstrated strong reasoning potential, their application to Bird's-Eye View (BEV) maps in traffic contexts remains limited by narrow task definitions and scarce annotated data.
Key Innovation: We introduce ChatBEV-QA, a large-scale BEV VQA benchmark of 137K+ QA pairs, designed to evaluate global scene understanding, vehicle-lane interactions, and vehiclevehicle interactions within complex traffic environments. While VisionLanguage Models (VLMs) have demonstrated strong reasoning potential, their application to Bird's-Eye View (BEV) maps in traffic contexts remains limited by narrow task definitions and scarce.
155. From Proxies to Fields: Spatiotemporal Reconstruction of Global Radiation from Sparse Sensor Sequences
Core Problem: Accurate reconstruction of latent environmental fields from sparse, indirect observations is a fundamental challenge across scientific domains, from atmospheric science and geophysics to public health and aerospace safety.
Key Innovation: Here we introduce the Temporal Radiation Operator Network (TRON), a spatiotemporal neural operator architecture that infers continuous global scalar fields solely from sequences of sparse, non-uniform proxy measurements. We demonstrate this approach on global cosmic radiation dose mapping: TRON, trained on daily reference fields spanning 2001 to 2023, generalizes across 65,341 spatial locations with input sequences ranging.
156. A Review of the Long Horizon Forecasting Problem in Time Series Analysis
Core Problem: The long horizon forecasting (LHF) problem has come up in the time series literature for over the last 35 years or so.
Key Innovation: This review covers aspects of LHF in this period and how deep learning has incorporated variants of trend, seasonality, fourier and wavelet transforms, misspecification bias reduction and bandpass filters while contributing using convolutions, residual connections, sparsity reduction, strided convolutions, attention masks, SSMs, normalization methods, low-rank approximations and gating mechanisms.
157. AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization
Core Problem: While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities across diverse domains, their application to specialized anomaly detection (AD) remains constrained by domain adaptation challenges.
Key Innovation: We propose a comprehensive framework addressing these limitations through two synergistic innovations. Comprehensive evaluation across multiple industrial datasets demonstrates substantial performance improvements in adapting general vision-language models to specialized anomaly detection.
158. Multi-Fidelity Physics-Informed Neural Networks with Bayesian Uncertainty Quantification and Adaptive Residual Learning for Efficient Solution of Parametric Partial Differential Equations
Core Problem: However, solving high-fidelity PDEs remains computationally prohibitive, particularly for parametric systems requiring multiple evaluations across varying parameter configurations.
Key Innovation: This paper presents MF-BPINN, a novel multi-fidelity framework that synergistically combines physics-informed neural networks with Bayesian uncertainty quantification and adaptive residual learning.
159. Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution
Core Problem: Real-world image restoration (IR) remains challenging due to complex and coupled degradations.
Key Innovation: While recent agentic IR frameworks leverage Large Language Models for flexible tool planning, they face two critical limitations. Extensive experiments on synthetic and real-world benchmarks demonstrate its strong perceptual and quantitative performance.
160. Adaptive Densification for High-Fidelity and Efficient Sparse Gaussian Splatting in Arbitrary-Scale Super-Resolution
Core Problem: To bridge this gap, we observe that a core capability of GS remains largely underexplored in ASR: the potential for dynamic densification, i.e., the spatially adaptive allocation of Gaussians based on image content.
Key Innovation: To address this challenge, we propose QuADA-GS, an approach that retains a powerful representational backbone but autonomously predicts where to allocate Gaussians relying strictly on the low-resolution input. While 2D Gaussian Splatting (GS) has recently shown great promise for ASR, current methods struggle to balance visual quality and computational cost.
161. Installation-method compatibility for floating offshore wind: A foundation-depth-metocean matrix applied to the Korean East Sea
Core Problem: Floating wind projects often commit to a foundation type without a transferable basis for matching installation methods to site conditions.
Key Innovation: This paper proposes a foundation-depth-metocean compatibility matrix that organises installation-method selection along four foundation types, three water depth zones, and three metocean severity classes. Applied to the Korean East Sea with GEBCO 2026 bathymetry and ERA5 2000-2024, the matrix reveals a practical installation space far narrower than the conceptual design space.
162. A viscoelastic-plastic model for ice over a wide range of strain rates based on Peridynamics
Core Problem: The mechanical behavior of ice varies with strain rates, exhibiting elastic, viscous, ductile and brittle responses.
Key Innovation: To describe the complex mechanical behavior of ice with Peridynamics, a viscoelastic-plastic model was developed in this study. The numerical results capture ice deformation, crack propagation and ice load (i.e. the force exerted by ice during ice-structure interaction) with good agreement with available experimental results.
163. Mechanical performance of corroded offshore wind turbine monopiles under extreme aerodynamic-hydrodynamic loads
Core Problem: The long-term structural integrity of monopile-supported offshore wind turbines (MOWTs) is continuously challenged by the combined action of progressive marine corrosion and extreme aero-hydrodynamic loads.
Key Innovation: To address this, we adopt a fully coupled aero-hydro-geotechnical framework to investigate the time-dependent mechanical performance of a 1.5 MW MOWT over a 50-year lifespan.
164. Investigation of the corrosion fatigue behavior of marine EH690 ultra-high-strength steel load-carrying cruciform welded joints
Core Problem: A dedicated corrosion-fatigue testing system was developed, and corrosion-fatigue tests were conducted on EH690 ultra-high-strength steel (UHSS) load-carrying cruciform welded joints (LCWJs) to investigate their fatigue failure behavior under free-corrosion conditions in seawater.
Key Innovation: A dedicated corrosion-fatigue testing system was developed, and corrosion-fatigue tests were conducted on EH690 ultra-high-strength steel (UHSS) load-carrying cruciform welded joints (LCWJs) to investigate their fatigue failure behavior under free-corrosion conditions in seawater. The results show that seawater corrosion reduces the fatigue life of EH690 UHSS LCWJs by approximately 50%.
165. EGO: a global 0.05° hourly GPP dataset for monitoring diurnal photosynthesis dynamics
Core Problem: At sub-daily scales, diurnal GPP dynamics reveal rapid adjustments to changing light, temperature and water conditions that are largely obscured in daily-to-annual aggregates, underscoring the need for developing global hourly GPP products.
Key Innovation: Here, we develop a causal knowledge-driven upscaling framework that couples the Peter and Clark Momentary Conditional Independence guided causal weights with ensemble learning strategies. At sub-daily scales, diurnal GPP dynamics reveal rapid adjustments to changing light, temperature and water conditions that are largely obscured in daily-to-annual aggregates, underscoring the need for developing global hourly GPP products.
166. Satellite-based inversion of global methane fluxes: capabilities and implications of GOSAT-2 measurements
Core Problem: Abstract.
Key Innovation: This study presents an evaluation of the GOSAT-2 Level 4 (G2L4) CH4 flux product, supported by analysis of the underlying Level 2 (L2) XCH4 retrievals, and summarizes key findings on global and regional CH4 budgets.
167. The ABRACOS dataset: a multidisciplinary marine ecosystem baseline for the western tropical Atlantic (2015-2017)
Core Problem: Preprint under review for ESSD (discussion: open, 0 comments) Understanding marine ecosystems requires observing environment, organisms, and interactions.
Key Innovation: Off Northeast Brazil, we conducted expeditions in 2015 and 2017 to measure ocean conditions, map distributions with acoustics, and sample from plankton to deep-sea fish.
168. High-Resolution Wide-Coverage Urban Canopy Parameters for Urban Simulations in Weather Research and Forecasting Models
Core Problem: Abstract.
Key Innovation: Cross-disciplinary researchers focusing on connected urban processes, especially those running numerical weather simulations at microscales, require reliable data on building footprints, heights and locations.
169. A novel Gauss-Hermite High-Order Sampling Hybrid ensemble filter for computationally efficient data assimilation in geosciences - Part 1: Application to Lorenz-96 in PythonDA v1.2.2
Core Problem: Providing an estimation of both state and uncertainty, ensemble algorithms are among the most successful data assimilation approaches.
Key Innovation: This work introduces a sampling method featuring a higher polynomial order of approximation, and an ensemble filter, the Gauss-Hermite High-Order Sampling Hybrid filter (GHOSH), which exploits the higher order of the novel sampling method.
170. Leveraging JEDI for atmospheric composition: a unified framework for evaluating observations and model forecasts
Core Problem: Modern data assimilation systems provide precise observation operators for mapping model variables into observation space, yet these capabilities remain underutilized outside assimilation.
Key Innovation: Here, we demonstrate how the Joint Effort for Data assimilation Integration (JEDI) framework addresses this gap by offering a unified, modular, and model-agnostic system that integrates data assimilation with systematic evaluation.
171. Wind and turbulence evaluation of the ICON model (icon-2026.04) using Doppler lidar observations
Core Problem: Turbulence parameterization in numerical weather prediction (NWP) models remains a challenge, particularly as resolution continues to increase.
Key Innovation: Existing turbulence evaluation methods often rely on high-resolution benchmark simulations which commonly depend on idealized assumptions and boundary conditions. The evaluation method is demonstrated here by applying it to the TKE scheme “Turbdiff” used in the ICOsahedral Nonhydrostatic (ICON) model, run with 2.1 km horizontal mesh size in the regional NWP configuration at the German Weather Service.
172. Effects of spatial soil moisture variability in forest plots on model parametrization and simulated groundwater recharge estimates
Core Problem: Despite their usefulness, the reliability of SVAT models is frequently compromised by uncertainties arising from incomplete or imprecise input data.
Key Innovation: These models are widely employed to predict hydrological responses to environmental change, including the impacts of shifting meteorological conditions on forested landscapes. The findings reveal that soil moisture variability at the plot characterized by a heterogeneous soil was greater, both horizontally and in depth, throughout the study period.
173. Advancing conflict research and response through satellite-derived data
Core Problem: Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11004-6 Integrating satellite-derived war-damage data with text-based fatality records through improvement, enrichment and fusion mitigates limitations inherent in each source, revealing complex violence dynamics beyond fatality-centric paradigms, as case studies from Ukraine and Myanmar illustrate.
Key Innovation: Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11004-6 Integrating satellite-derived war-damage data with text-based fatality records through improvement, enrichment and fusion mitigates limitations inherent in each source, revealing complex violence dynamics beyond fatality-centric paradigms, as case studies from Ukraine and Myanmar illustrate.
174. Environmental filtering shapes patch dynamics across isolated mesophotic reefs
Core Problem: Mesophotic coral ecosystems (MCEs; ∼30 to 150 meters) are major but poorly understood benthic habitats.
Key Innovation: We used Autonomous Reef Monitoring Structures (ARMS) and integrated metabarcoding (mtCOI and 18 S), image analysis, and hydrodynamic modeling across six mesophotic banks in the Texas-Louisiana Shelf to test whether community assembly is governed by environmental filtering or dispersal limitation. Hydrodynamic simulations revealed dispersal is variable but not limiting.
175. Buried fragments reveal a “Greater Gondwana” supercontinent
Core Problem: However, Gondwana is frequently denied supercontinent status due to its perceived spatial deficit (∼64% of continental landmass).
Key Innovation: Our results expand Gondwana’s reconstructed area to ∼80% of the global landmass, firmly establishing its identity as a true supercontinent-termed “Greater Gondwana.” We propose that the assembly and peripheral subduction of this “Greater Gondwana” played a pivotal role in driving the geodynamic and environmental perturbations that culminated in the Cambrian metazoan radiation.
176. Seasonal predictability of winter PM2.5 pollution severity in India through a strong pollution-extratropical storm connection
Core Problem: India suffers from severe fine particulate matter (PM 2.5) air pollution in winter that harms public health and disrupts economic activities.
Key Innovation: Seasonal prediction of pollution severity could facilitate proactive mitigation planning but is hindered by incomplete understanding of the drivers of pollution variations.
177. Increasing nitrogen loading in the Gulf of Mexico undermines coral reef resilience
Core Problem: The ongoing global decline of coral reefs underscores the need to identify key mechanisms driving reef deterioration.
Key Innovation: The Flower Garden Banks (FGB) reefs in the northern Gulf are widely regarded as some of the healthiest in the continental United States, but recent disease outbreaks in the FGB have raised concerns about their long-term resilience. Our results show that anthropogenic nitrogen began to rise around 1850, coinciding with the early use of fertilizers, and increased sharply after the 1960s with the onset of the Green Revolution.
178. Winter Wheat Yield Estimations Based on Multisource Remote Sensing Parameters and the BiLSTM-CNN Model
Core Problem: Winter wheat is a cornerstone of China’s grain production, contributing substantially to national food security and overall cereal output.
Key Innovation: This study modeled the nonlinear associations between multitemporal remote sensing variables and winter wheat yield. The BiLSTM-CNN model showed higher estimation accuracy than individual BiLSTM and 1-D CNN models, with an R² of 0.69 and root mean square error (RMSE) of 478.68 kg/hm2.
179. Inshore Analysis of the Morphostructural Evolution of the Coastal Cliffs of Bessin, Basse Normandie, France
Core Problem: The limestone-marl cliffs along France’s northern coast, particularly from Arromanches-les-Bains to Longues-sur-Mer, are experiencing significant erosion.
Key Innovation: This study investigates cliff retreat mechanisms and classifies instability types through image analysis, LiDAR measurements, and field observations. A review of images from 1947 to 2012 showed that cliff recession occurs in discrete events, averaging 0.06-0.30 m per year.
180. Ground-Based GNSS Atmospheric Remote Sensing for Ultra-Short-Term Wind-Power Forecasting: A Direction-Proxy-Guided Graph-Residual Approach
Core Problem: Ground-based Global Navigation Satellite System (GNSS) stations provide continuous atmospheric remote sensing through electromagnetic propagation delays.
Key Innovation: We propose a model combining a long short-term memory (LSTM) backbone, GNSS conditioning, and a graph neural network (GNN), denoted LSTM+GNN+GNSS, for 4 h wind-power forecasting. The complete system achieves a normalized mean absolute error (nMAE) of 4.53% (4.527 ± 0.132% across three power-model seeds), reducing nMAE by 7.57% relative to LSTM+GNN and 8.23% relative to LSTM.
181. Frequency-Guided Cross-Scale Refinement Network for UAV Detection
Core Problem: In recent years, the use of UAVs has become increasingly widespread, and the public safety risks posed by unauthorized UAV flights have become increasingly prominent, creating an urgent need for effective detection and identification of UAV targets.
Key Innovation: However, such targets are small in size, have low contrast, and exhibit an extremely low signal-to-noise ratio; conventional detection methods generally suffer from insufficient feature discrimination, missed detections, and false alarms in complex backgrounds. Based on an encoder-decoder architecture, this network achieves end-to-end collaborative optimization through cross-layer feature fusion, side-channel prediction.
182. Retrieval of Vegetation Nitrogen from Hyperspectral Remote Sensing: A Critical Review of Recent Methodological Advances
Core Problem: Particular attention is given to protein-sensitive RTMs, advanced ML approaches, and uncertainty-aware retrieval.
Key Innovation: We examine the evolution from classical parametric regression and nonlinear ML approaches towards physically based RTM inversion and hybrid RTM-ML frameworks that integrate the complementary strengths of physical modeling and statistical learning.
183. A Hybrid Super-Resolution and Object Detection Framework for Small Ship Recognition in Optical Remote Sensing Imagery
Core Problem: Continuous maritime surveillance increasingly relies on optical remote sensing, yet small ships are difficult to detect in low-resolution imagery because limited spatial resolution, complex sea backgrounds, and sensor noise suppress the fine-grained cues on which detectors depend, while routine acquisition of high-resolution imagery is prohibitively costly.
Key Innovation: To address this problem, a hybrid framework (RGT-YOLOv5Det) is proposed that couples transformer-based super-resolution (SR) with a lightweight object detector so that fine spatial detail is restored before detection.
184. Enhancing Our View from Above: A Downscaling Technique to Reveal Finer-Scale Urban Heat Patterns
Core Problem: Downscaling approaches have been developed to generate higher-resolution LST products; however, in the absence of independent, high-resolution surface temperature observations, it remains difficult to determine whether the additional spatial variation introduced by downscaling reflects meaningful landscape-related thermal structure or model-generated variability.
Key Innovation: Downscaling approaches have been developed to generate higher-resolution LST products; however, in the absence of independent, high-resolution surface temperature observations, it remains difficult to determine whether the additional spatial variation introduced by downscaling reflects meaningful landscape-related thermal structure or model-generated variability.
185. An Integrated Approach for Mapping Multi-Category Abandoned Cropland Based on Swin-Unet, Land-Use Trajectories and Phenological Features in a Double-Cropping Rice Region of Southern China
Core Problem: However, multi-category abandoned-cropland mapping remains challenging in southern China’s double-cropping rice regions because fragmented fields, frequent cloud/rain conditions, and complex phenology can obscure abandonment signals.
Key Innovation: Focusing on a typical double-cropping rice region in southern China, this study first used Swin-UNet to generate annual land-use maps from 2018 to 2022 and derive interannual trajectories. Results showed that annual classification accuracies from 2018 to 2022 ranged from 94.7 to 95.5%, supporting reliable cropland change trajectory extraction.
186. High-Spectral-Resolution Retrieval of Ultraviolet and Visible Aerosol Optical Depth Using the Pandora Spectrometer
Core Problem: Atmospheric aerosols from natural and anthropogenic sources are a major component of air pollution and substantially affect air quality, human health, and atmospheric chemistry.
Key Innovation: This study presents a framework for retrieving τaer in the UV (315-375 nm) and visible (400-530 nm) spectral ranges using the Pandora spectrometer. The retrieved UV and visible τaer values show good agreement with collocated AErosol RObotic NETwork (AERONET) data, with R² = 0.88-0.99, RMSE = 0.016-0.039, |MBE| = 0.002-0.023, slopes of 0.93-1.01, and |intercepts| less than 0.01.
187. Multi-Sensor Mapping of Aquatic Vegetation Using Sentinel-2, SAR-Optical Fusion and Derived Spectral and SAR Indices in South Florida
Core Problem: Accurate mapping of submerged aquatic vegetation (SAV) and emergent aquatic vegetation (EAV) is essential for hydrologic assessment and wetland management.
Key Innovation: This study developed and evaluated a multi-sensor aquatic vegetation mapping framework using Sentinel-1 SAR, Sentinel-2 optical imagery, and derived spectral and SAR indices in the Everglades stormwater treatment areas (STAs) of South Florida. On the internal test set, both classifiers achieved similar overall accuracies (OA = 0.83-0.85), with no clear performance advantage between RF and TabNet.
188. Hyperspectral Mineral Mapping of Drill-Core Samples at Sadiola Hill Gold Deposit, Mali, West Africa: The Paragenesis of the Polyphase Alteration Mineralogy
Core Problem: The methodologies used included high-resolution short-wave infrared (SWIR) hyperspectral imaging, electron microprobe analysis, and conventional petrography.
Key Innovation: We present the results of spectral and petrographic studies of 36 drill core samples from six diamond boreholes from the Sadiola Hill gold mine to (1) refine paragenetic studies and (2) define alteration assemblages accompanying the main gold event.
189. Mechanism of seawater intrusion in underground water-sealed oil storage caverns on islands: a DFN approach
Core Problem: However, compared with inland freshwater environments, the groundwater seepage characteristics in island environments are more complex and involve the risk of seawater intrusion.
Key Innovation: Taking an island-based UWSOS project as the research subject, this study adopts numerical modeling to investigate the mechanisms of seawater intrusion in such systems. Simulation results indicate that the groundwater pressure distribution in the DFN model closely resembles that in the EPM model, but the seepage velocity in the DFN model is approximately four orders of magnitude higher than that in the EPM model.
190. Wellbore Stability Analysis in Fractured Anisotropic Shale: A Chemo-Poro-Elastic Dual Medium Model
Core Problem: However, existing models typically consider these factors in isolation, and a fully coupled framework that simultaneously accounts for anisotropy, dual-porosity systems, and chemical effects remains lacking.
Key Innovation: However, existing models typically consider these factors in isolation, and a fully coupled framework that simultaneously accounts for anisotropy, dual-porosity systems, and chemical effects remains lacking. The model is implemented using the finite element method under a generalized plane strain framework and validated against available analytical solutions in limiting cases, demonstrating good agreement.
191. Foredune notch longevity and spatial evolution across northwest Europe
Core Problem: Yet understanding of their long-term morphological evolution, persistence, and the extent to which design and wind regime influence outcomes remains limited.
Key Innovation: Yet understanding of their long-term morphological evolution, persistence, and the extent to which design and wind regime influence outcomes remains limited. Results show that notch evolution is predominantly characterised by contraction.
192. Estimating sand wave migration from local hydrodynamics and sediment mobility
Core Problem: Their migration causes seabed level changes that can affect offshore infrastructure, motivating the need for decadal-scale predictions of migration rates.
Key Innovation: Their migration causes seabed level changes that can affect offshore infrastructure, motivating the need for decadal-scale predictions of migration rates. Results show that an estimator based on tidal current asymmetry provides a reasonable first-order estimate of site-specific migration rates with strong linear correlation.
193. Consistent enough: When does perceived conflict between weather agency and emergency services warnings impact protective action?
Core Problem: A challenge during hydrometeorological hazard emergencies is encouraging the community to take protective action in a noisy, information-rich environment.
Key Innovation: In this study, we examine community response to severe weather warnings released by a weather agency and an emergency services agency.
194. Accessibility and Resource Use in Crisis and Emergency Risk Communication and Risk Reduction: The Self-Reinforcing Responsive Model (SRRM)
Core Problem: Background Increasingly frequent and complex emergencies heighten the need for accessible risk communication that improves public safety and upholds the rights and dignity of all community members.
Key Innovation: This study examined how accessible risk communication is conceptualized and implemented across emergency management organizations, and developed tools to support practice. Results Using inductive qualitative analysis, we identified existing accessibility initiatives alongside persistent challenges related to resource allocation, time and staffing pressures, siloed workflows, inconsistent standards, and fragmented knowledge.
195. The role of alternative information sources and channels on wildfire and disaster perceptions and preparedness in a wildfire context: A comparison of the Great Plains and Mountainous Western states in the U.S.
Core Problem: Greater wildfire frequency and severity have increased the need for disaster preparedness and related education, yet household disaster preparedness and response, which remain low, rely on accessible information.
Key Innovation: Such information is limited in rural communities of the Great Plains (GP) of the U.S that lack resources.
196. Towards a better understanding of evidence pathways in multi-level disaster governance arrangements: A case of Nepal
Core Problem: In the context of growing risk and uncertainty, disaster risk reduction (DRR) policies are increasingly expected to be evidence-informed.
Key Innovation: This imperative has contributed to the expansion of research and knowledge infrastructures and to the global circulation of disaster discourse.
197. Analysis of causal factors of chemical accidents in China using an Association rule mining - Bayesian network based AcciMap approach
Core Problem: Association rule mining (ARM) and Bayesian network (BN) modeling are combined with AcciMap to address the limitations of AcciMap, such as subjective dependency and lack of quantitative analysis ability.
Key Innovation: In this study, a hybrid AcciMap-based approach is proposed to analyze causative factors of severe chemical industry accidents.
198. A data-driven thermal runaway prediction method for lithium-ion batteries based on a diffusion model
Core Problem: However, the inherent complexity and uncertainty of the TR process complicate efforts to predict it accurately.
Key Innovation: Therefore, developing an effective prediction method for this phenomenon is of paramount importance.
199. Arc-Level Reliability Integration in Connected Autonomous Vehicle Emergency Routing: A Multi-Objective Framework with Critical Service Quality Threshold Analysis
Core Problem: This paper formulates the Reliability-Aware Connected Autonomous Vehicle Routing Problem with Fuzzy Time Windows (CA-VRPFTW-NR) as a three-objective mixed-integer program that simultaneously minimizes priority-weighted patient dissatisfaction and cumulative route unreliability and maximizes priority-weighted service coverage.
Key Innovation: The framework is applied to the 2025 Iberian Peninsula blackout, where the four blackout scenario points map onto the predicted service-quality phase diagram without recalibration, indicating that the findings extend to a real infrastructure failure event. Multi-objective analysis shows that Pareto decision flexibility contracts at the same critical threshold, making robust optimization most valuable when networks are most.
200. Dynamic Inspection and Maintenance Planning for Cracked Pipes using Reinforcement Learning with Risk-informed Rewards
Core Problem: Pipeline integrity management involves crack inspections and subsequent repair or other operation actions to prevent leakage or burst failures.
Key Innovation: To address this challenge, this study develops a dynamic I&M planning framework based on reinforcement learning (RL). For the purpose of demonstration, I&M planning for a single cracked pipe joint is considered, and the dynamic I&M planning based on RL is compared with conventional static I&M planning.
201. A new adaptive approach for detecting crop sowing and harvesting dates from Sentinel-1 time series
Core Problem: However, their performance varies significantly depending on crop types and growing environments, making the SAR-based methods unreliable for large-scale applications.
Key Innovation: Hence, we developed an Adaptive Sentinel-1 Indicator Selection (ASIS) framework for CAD detection by comprehensively investigating the mechanisms influencing the capabilities of Sentinel-1 indicators. Our findings suggested that CR is more effective for CADs detection in rain-fed crops with rough soil, while VH performs better in irrigated or aqua fields.
202. GBOV-trained Gaussian process models for Sentinel-3 SYN vegetation traits with decomposed uncertainty quantification
Core Problem: Accurate and uncertainty-aware retrieval of vegetation biophysical variables is essential for reliable Earth observation-based monitoring.
Key Innovation: We developed Gaussian Process Regression (GPR) models trained on globally distributed Copernicus Ground-Based Observations for Validation (GBOV) reference data to retrieve leaf area index (LAI), fraction of absorbed photosynthetically active radiation (FAPAR), and fractional vegetation cover (FVC) from Sentinel-3 SYNERGY (SY2_SYN) surface reflectance.
203. Tree-neighborhood scale structure-radiation coupling analysis and modeling using multi-source remote sensing and ensemble learning
Core Problem: However, how neighborhood spatial configuration and competition dynamics control APAR distribution remains poorly understood, particularly in structurally heterogeneous and topographically complex natural forests.
Key Innovation: To address this, we proposed FASAR (Forest Architecture-Solar Absorption Response), a structure-radiation coupling framework for complex-terrain natural forests. Applied to 144 plots (5.76 ha, 3,417 trees) in a natural Qinghai spruce forest in western China, FASAR revealed that large trees absorb more light with vertically stable patterns, while canopy stratification, competition, and size disparity intensify vertical.
204. Conductivity-based tracking of thin-film freezing in microencapsulated phase change asphalt mixtures for black-ice mitigation
Core Problem: Black ice remains a difficult pavement hazard because it is thin, transient and visually elusive, yet it can rapidly compromise surface friction.
Key Innovation: This study develops a microencapsulated phase change asphalt mixture (MPCAM) for delaying black ice formation and, more importantly, establishes electrical conductivity monitoring as a quantitative means of tracking the freezing of thin surface water films. These findings show that MPCAM is best understood not as a material for preventing sustained icing, but as a passive pavement technology for delaying short-duration black.
205. Water film transport modeling and bottom icing susceptibility assessment for high-speed railway contact wires
Core Problem: Contact wire icing threatens current collection and operational safety on high-speed railways, and liquid water transported to the lower surface is a direct precursor to bottom icing.
Key Innovation: This study develops a water film transport model for a grooved contact wire profile and systematically examines the effects of meteorological parameters on film flow and distribution. A neural-network surrogate is further developed to predict M b, achieving an R² of 0.9465 with substantially improved inference speed.
206. Damage mechanism and fiber reinforcement of cold recycled asphalt mixture under salt-freeze coupling in seasonal frozen regions
Core Problem: However, its durability is severely compromised by the coupled action of salt erosion and freeze-thaw cycles.
Key Innovation: This study investigates the deterioration mechanism of CRAM under salt-freeze coupling and evaluates the reinforcing effects of three fibers with distinct functional characteristics: cotton straw fiber (SF), basalt fiber (BF), and polypropylene fiber (PF).
207. Tracing fluoride fate and mobilization in lake-groundwater systems using multi-isotopic (87Sr/86Sr, δ 11B and 222Rn) and microbial evidence
Core Problem: Despite its environmental and health relevance, the origin and mobilization of fluoride (F−) in lake-groundwater systems have been scarcely quantified.
Key Innovation: This study integrated hydrochemistry, Bayesian isotope mixing modeling (MixSIAR) based on 87Sr/86Sr and δ 11B, a 222Rn mass balance model (RMBM), and microbial ecological analysis to investigate F− dynamics in the Bahannao Lake Group (BLG), a representative lake-groundwater system in Inner Mongolia, China. Analysis of 291 water samples showed that high-F− waters were mainly associated with HCO − 3-Na+ hydrochemical types.
208. Contrasting hydrologic pathways controlling nitrate and phosphorus transport in an agricultural critical zone
Core Problem: Understanding the influence of hydrologic processes on nutrient transport remains a critical challenge in agricultural watersheds, where nitrate (N) and phosphorus (P) often exhibit contrasting export patterns.
Key Innovation: In this study, we developed a coupled surface-subsurface modeling framework using SWAT + integrated with the gwflow module to simulate hydrologic fluxes and nutrient transport in the Choptank River watershed, USA. These findings demonstrate the capability of the coupled surface-subsurface framework to diagnose hydrologic pathway partitioning governing nutrient transport within the critical zone.
209. A three-spring model for pile-soil interaction of offshore monopile in sand: numerical and analytical investigation
Core Problem: Existing methods, primarily p-y curve approaches, demonstrate significant limitations when applied to large-diameter monopiles, failing to adequately capture evolving soil failure mechanisms.
Key Innovation: This study establishes a unified three-spring analytical framework capable of characterizing the lateral response of flexible, semi-rigid, and rigid monopiles. Existing methods, primarily p-y curve approaches, demonstrate significant limitations when applied to large-diameter monopiles, failing to adequately capture evolving soil failure mechanisms.
210. The pore radius propagation approach: computational characterization of pore structures in granular media
Core Problem: Accurate quantification of pore structures and characterization of pore-scale fluid behavior are fundamental to the modeling of unsaturated flow and multiphase interactions in geotechnical systems.
Key Innovation: This study presents a computationally efficient Pore Radius Propagation (PRP) approach for estimating pore-size distributions from computed tomography (CT) datasets. The proposed framework also accurately reproduced the microscale spatial distribution of capillary water under specified suction conditions, demonstrating its capability to represent pore-scale hydraulic responses.
211. From compaction to separation: A mechanistic framework for vibration-driven segregation in unsaturated soils
Core Problem: These mechanisms are important in a wide range of geotechnical problems that include embankments, mine tailings, mineral ore handling and transport, and pavement subgrades.
Key Innovation: These mechanisms are important in a wide range of geotechnical problems that include embankments, mine tailings, mineral ore handling and transport, and pavement subgrades. The model is able to reproduce the pore pressure responses and the calibrated parameters show reasonable transferability across related tests, conditional upon the measured specimen-state histories supplied to the model.
212. Selection of modified mining waste materials for sustainable open-pit haul roads: a fuzzy AHP-TOPSIS framework incorporating extreme climate and life-cycle perspectives
Core Problem: Open-pit haul roads are increasingly exposed to extreme climate conditions, while life-cycle carbon emissions are rarely incorporated into material selection in engineering practice.
Key Innovation: This study presents a fuzzy AHP-TOPSIS-based evaluation approach for mine-derived materials, in which extreme climate factors are integrated into the hierarchical weighting process and life-cycle carbon emissions are incorporated into the ranking procedure. Results under the SSP5-8.5 scenario indicated that mechanical strength and heavy metal control dominated decision priorities, while climate-related performance criteria.
213. A field-parameterized 3-DOF coupled dynamic model for vibratory compaction of soil-rock mixture subgrade: characterizing compaction quality via the drum-soil separation ratio
Core Problem: This study develops a field-parameterized framework for evaluating the vibratory compaction quality of soil-rock fill subgrades considering discontinuous drum-soil contact.
Key Innovation: This study develops a field-parameterized framework for evaluating the vibratory compaction quality of soil-rock fill subgrades considering discontinuous drum-soil contact. The calculated drum accelerations showed reasonable consistency with the field measurements.
214. Fluid Inertia Rewires Flow Graphs and Selects Late Solute Paths in a Rough Fracture
Core Problem: At the high flow rate, a uniformly scaled low-inertia graph underpredicted the 99th percentile by 16.1%; matching the global Peclet number also failed to collapse transport.
Key Innovation: We tested this assumption in a measured three-dimensional rough fracture using steady Navier-Stokes flow, flux-closure recirculation zones, directed voxel graphs, and conservative graph transport. The results show that edge directions vary with Reynolds number and carry information about late retention associated with inertial recirculation.
215. The tetragonal-cubic transition of davemaoite: Implications for lower mantle seismic anomalies
Core Problem: Davemaoite (CaSiO3 perovskite) is the third most abundant mineral in Earth's lower mantle and a dominant phase in subducted oceanic crust.
Key Innovation: Its crystal structure is distorted (tetragonal or orthorhombic) at ambient conditions but is considered to transform to cubic at high temperatures. Previous experiments reported low, nearly pressure-independent transition temperatures (approximately 600 K), demonstrating a long-standing discrepancy with theoretical predictions that mostly exceed 1000 K.
216. Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance
Core Problem: Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications.
Key Innovation: Scaling laws have guided resource allocation for large-scale training, yet prior work treats total CLIP model size as a single variable, without exploring how the capacity split between encoders impacts downstream performance. Here, we train multiple CLIP models with different vision and text encoder sizes, revealing that for most vision encoders, there is an optimal text encoder size beyond which zero-shot performance.
217. Hierarchical Prompt Injector for Domain Generalization Segmentation
Core Problem: Domain Generalized Semantic Segmentation (DGSS) is a challenging task, as vision models often rely on low-level appearance cues that change across domains.
Key Innovation: Additionally, we propose the Hierarchical Prompt Injector (HPI), which enables spatially adaptive prompt injection in foundation models. We achieve 70.62% and 72.74% mIoU on synthetic-to-real and real-to-real benchmarks, respectively.
218. FreeTransformSR: Efficient Lightweight Image Super-Resolution via Free Low-Rank Learnable Transform
Core Problem: Single image super-resolution aims to reconstruct high-resolution images from low-resolution inputs.
Key Innovation: To further enhance high-frequency detail recovery, we introduce a local feature modulation branch that complements transform-domain processing with depthwise convolution. Extensive experiments on five benchmark datasets demonstrate that FreeTransformSR achieves competitive PSNR/SSIM performance with significantly fewer parameters and FLOPs.
219. FineHOI: Part-Aware Dense Representations for Zero-Shot Human-Object Interaction Detection
Core Problem: However, they often rely on global or detector-centric features that compress interaction cues and hinder fine-grained spatial reasoning.
Key Innovation: To overcome this limitation, we propose FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features. Extensive experiments demonstrate that FineHOI consistently outperforms existing zero-shot HOI methods, achieving particularly strong gains on unseen interactions.
220. IXPLORE: Bounded Ideal Point Estimation with Grid-Based Uncertainty Quantification
Core Problem: However, selecting the corresponding spatial model involves various trade-offs: while model-based approaches such as Item Response Theory (IRT) are based on utility functions rather than optimized for predictive accuracy, most Machine Learning (ML) alternatives struggle to generalize beyond training data when embedding sparse test responses.
Key Innovation: We introduce IXPLORE, a bounded ideal point estimation algorithm that combines a predictive fit objective with a sparsity-aware likelihood function. Furthermore, we show that non-linear feature transforms can further reduce the reconstruction error while remaining visually interpretable.
221. Spatial Attention Supervision for Defect Localization: Exploiting Ground-Truth Masks as Training Signal in Diffusion-Augmented Defect Detection
Core Problem: Ground-truth defect masks in industrial inspection datasets are typically reserved for evaluation.
Key Innovation: This paper repurposes them as spatial supervision signals during training of classification networks, teaching a model not just what to predict but where to look.
222. MSCA-UNet: Multi-Scale Context and Attention U-Net for Image Segmentation
Core Problem: U-Net remains a practical baseline for image segmentation because of its simple encoder-decoder structure and skip connections.
Key Innovation: The results support the view that multi-scale context enrichment and attention-based feature refinement provide complementary benefits within a U-Net framework. Under identical experimental settings, the baseline U-Net achieves 96.9% mIoU on a held-out test set.
223. NOVA: Normal-Side Modeling for Training-Free Zero-Shot Video Anomaly Detection
Core Problem: Existing CLIP-based methods often emphasize anomaly-side semantics, while the competing normality side remains less carefully formulated.
Key Innovation: We propose NOVA, a training-free ZS-VAD framework that strengthens the normal side at both linguistic and visual levels. NOVA achieves 89.86 percent AUC on UCF-Crime and 95.07 percent AUC and 84.82 percent AP on XD-Violence, reaching state-of-the-art performance among comparable training-free zero-shot methods.
224. Assessing Covariate-Informed Grid Load Forecasting with a Time-Series Foundation Model
Core Problem: However, real-world load forecasting often involves multiple target variables and requires the integration of exogenous variables, raising important questions about the utility of TSFMs in realistic settings.
Key Innovation: In this study, we position Chronos-2, a recently developed model by Amazon, as a representative multi-channel TSFM that supports univariate, multivariate, and covariate-informed forecasting, and conduct a systematic investigation of how such models can be used for real-world load forecasting.
225. Back to the Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features
Core Problem: We present B2TFPose, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images.
Key Innovation: We present B2TFPose, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images. On the seven core datasets of the BOP Benchmark, B2TFPose achieves 40.7 mean AR without refinement and 56.4 with refinement, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.
226. From Synthetic Priors to Model Behavior: Structural Coverage in Tabular Foundation Models
Core Problem: Tabular foundation models (TFMs) are commonly pretrained on large collections of procedurally generated synthetic tasks, yet it remains unclear how well these synthetic pretraining priors support the downstream tasks on which the models are evaluated.
Key Innovation: We study this question from a distribution-level attribution perspective. We find substantial differences across synthetic pretraining priors: some generators provide consistently broader and denser support for benchmark tasks than others.
227. Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models
Core Problem: However, many existing methods overlook data characteristics and simply reuse the training strategies adopted during pre-training.
Key Innovation: However, many existing methods overlook data characteristics and simply reuse the training strategies adopted during pre-training. Experiments across distribution shift, transfer learning, and few-shot settings demonstrate consistent improvements over existing approaches.
228. One for All: Generalist Foundation Model for Cross-Sensor Skeleton Representation Learning
Core Problem: However, existing approaches are primarily limited by the inherent heterogeneity of skeleton data-characterized by varying joint counts, indexing protocols, and topological structures across different sensors-which typically necessitates training separate, sensor-specific, or even entirely dataset-specific models.
Key Innovation: To overcome this, we introduce SOfA (Skeleton One for All), the first generalist foundation model designed to achieve sensor-unified skeleton representation learning across diverse sensors.
229. Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification
Core Problem: While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse.
Key Innovation: In this work, we propose a robust, data-centric framework to stabilize RL training. Our method outperforms strong baselines on complex reasoning tasks, offering a principled solution for stable and efficient RL fine-tuning.
230. Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection
Core Problem: However, due to object-centric bias, normal and anomalous text prototypes exhibit a high semantic overlap.
Key Innovation: While enforcing strict orthogonality between them improves discriminability, mapping highly contiguous visual inputs onto drastically orthogonal prototypes introduces a geometric dilemma, disrupting the pre-trained structural continuity. Extensive experiments demonstrate that Proximity-CLIP outperforms current state-of-the-art methods across multiple ZSAD benchmarks with minimal architectural modifications.
231. Federated Binary Gating with Server-Side Vision-Language Inference for Surveillance Anomaly Classification
Core Problem: In federated learning settings, this challenge is amplified by non-independent and identically distributed (non-IID) client data, which can make direct multiclass anomaly classification unstable, especially for rare categories.
Key Innovation: We propose a hybrid two-stage architecture that combines a federated binary convolutional neural network (CNN) gate with server-side zero-shot VLM inference. These results suggest that federation is better suited to coarse local screening, while routing rules can be adjusted to trade server-side VLM usage for higher anomaly sensitivity.
232. RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
Core Problem: Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions.
Key Innovation: To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.
233. Latent-to-Latent Flow for Volumetric Stochastic Segmentation
Core Problem: Uncertainty arising from inter-observer variability in medical image segmentation plays an important role in developing treatment plans.
Key Innovation: Flow matching has emerged as a powerful framework for generative modelling and has also been demonstrated to maintain strong performance when working with latent representations of images.
234. Improving Multivariate Time Series Classification with Class-Wise Training and Model Aggregation
Core Problem: In this paper, we propose a class-wise dimension (channel) selection framework for Multivariate Time Series Classification (MTSC).
Key Innovation: In this paper, we propose a class-wise dimension (channel) selection framework for Multivariate Time Series Classification (MTSC). Experimental results indicate that class-wise dimension selection improves the quality of extracted representations and can enhance classification performance, particularly in high-dimensional settings.
235. CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models
Core Problem: We challenge this assumption.
Key Innovation: We propose CrACK (Cross-model Adversarial Consistency attack), an inference-time attack that exploits this interface without modifying any input pixel, model weight, or training data. Experiments on four collaborative pipelines across eight benchmarks show that CrACK causes catastrophic degradation while every individual model continues to produce its unchanged standalone output, rendering per-model defenses structurally.
236. TeMo: Temperature Modulation for Multimodal Contrastive Learning
Core Problem: However, most existing methods either fix this hyperparameter or learn a global value during training.
Key Innovation: In this paper, we introduce TeMo, Temperature Modulation framework, a similarity-based modulation approach that adaptively adjusts the temperature for each positive-negative pair according to their similarity, enabling more fine-grained multimodal contrastive learning. Extensive experiments demonstrate that each component of TeMo consistently enhances performance across diverse zero-shot retrieval and classification tasks.
237. Re-engineering SORT-based algorithms for low-cost small object tracking from omnidirectional footage
Core Problem: Multi-object tracking (MOT) has advanced rapidly in urban surveillance and autonomous driving, yet many trackers rely on ReID- and transformer-based appearance encoders and are designed for standard FoV cameras.
Key Innovation: These assumptions break down for low-cost omnidirectional deployments, where equirectangular projection introduces seam discontinuities and targets appear to be small and fast-moving.
238. SphereSOD: Geometry-Structure Coupled Learning for 360 Salient Object Detection
Core Problem: However, equirectangular projection (ERP) introduces severe spatial distortion when mapping the spherical domain onto a planar representation.
Key Innovation: However, equirectangular projection (ERP) introduces severe spatial distortion when mapping the spherical domain onto a planar representation. Extensive experiments on three public 360° SOD benchmarks demonstrate state-of-the-art performance and a favorable accuracy-efficiency trade-off, supporting structurepreserving inference directly in ERP space as a promising alternative to projection-heavy panoramic pipelines.
239. Zero-Shot 3D Plant Organ Segmentation with SAM3 and Semantic NeRFs
Core Problem: Accurate 3D plant organ segmentation is fundamental to automated phenotyping.
Key Innovation: We present an annotation-free pipeline for 3D plant organ segmentation, combining text-prompted SAM3 segmentation with semantic neural radiance fields (NeRFs). On a controlled Begonia maculata testbed the SAM3 pipeline achieves 92.6% mIoU, reaching 95.9% of the oracle upper bound established with perfect ground-truth masks.
240. DroneGround: Open-Vocabulary Drone Payload Characterization Using Synthetic Data and Grounded Vision-Language Models
Core Problem: While existing vision-based systems achieve strong performance for drone detection and tracking, reliable payload characterization remains highly challenging under long-range imaging conditions due to limited availability of annotated real-world datasets, and substantial distribution shifts encountered during deployment.
Key Innovation: To address these challenges, we generate a photorealistic synthetic drone-payload dataset using Unreal Engine 5 and Cosys-AirSim and propose DroneGround: Grounded Vision-Language Payload Characterization, a two-stage framework for robust open-vocabulary payload analysis.
241. VoT: Vision-of-Thought for Unified Multimodal Representation Alignment
Core Problem: Despite their success, these methods lack an explicit, interpretable intermediate representation that effectively bridges high-level linguistic semantics and low-level visual signals.
Key Innovation: In this paper, we propose Vision-of-Thought (VoT), a framework that introduces a discrete visual-thinking layer between vision-language models (VLMs) and diffusion transformers (DiTs). Experimental results demonstrate that VoT improves semantic alignment and provides a structured interface for interpretable and controllable generation.
242. Kalman Delta Networks: Uncertainty-aware Associative Memory
Core Problem: Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require.
Key Innovation: To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs).
243. Streaming Hierarchical Inference with Tabular Foundation Models
Core Problem: Tabular Foundation Models (TFMs) have recently demonstrated strong predictive performance through in-context learning, but their deployment in high-throughput data streams remains challenging due to communication overhead and latency.
Key Innovation: We propose \textit{HINT}, a hierarchical inference framework that combines edge-based retrieval with cloud-based TFM inference. Experiments show \textit{HINT} consistently identifies favorable trade-offs.
244. Heat Field Signatures: From Point Clouds to Smooth Geometry
Core Problem: Bringing multiscale geometric analysis directly to irregular point clouds remains difficult: quantities such as local dimension, anisotropy, density variation, and geometric transitions are typically estimated through explicit neighborhood, manifold, or graph constructions, or left for neural networks to infer from coordinates.
Key Innovation: We introduce Heat Field Signatures (HFS), which lift a point cloud to a multiscale family of smooth ambient heat fields, providing a direct interface from discrete samples to geometric analysis. Across synthetic and real-world benchmarks spanning subcellular, neuronal, tree, and protein data, HFS outperforms strong point-cloud and multiparameter-persistence baselines while substantially reducing end-to-end cost.
245. Geodesic-informed Generative Diffusion Model For Topology-preserved Image Video Generation
Core Problem: Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, including but not limited to synthesis, reconstruction, and segmentation.
Key Innovation: To address these challenges, we introduce IGG (Image Generation informed by Geodesic dynamics), a novel framework that integrates topology-preserving geodesic principles into the diffusion-based generative process.
246. Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action Segmentation
Core Problem: However, prevailing methods struggle with intrinsic uncertainty in boundary localization due to less discriminative features in action-transiting regions.
Key Innovation: In our paper, we introduce the boundary voting network that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions. Extensive experiments demonstrate the effectiveness of our method on GTEA, 50Salads, and Breakfast datasets.
247. Revisiting Spectral Representations in Generative Diffusion Models
Core Problem: Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains insufficiently understood.
Key Innovation: Motivated by this, we propose a self-supervised spectral representation alignment method to facilitate diffusion model training. Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains insufficiently understood.
248. Tracking-by-detection in Multi-object Tracking: Survey and Experiments
Core Problem: Despite recent progress, fair evaluation of TBD-based methods remains a challenge.
Key Innovation: Many studies introduce modules such as similarity metrics, data association strategies, or motion models, but they are often evaluated under inconsistent protocols, with different baseline trackers, hyperparameters, and datasets. Our findings establish a strong baseline tracker and provide a foundation for the principled design of robust and versatile MOT systems suitable for real-world deployment.
249. SAM3-O2D2: Zero-Shot Object Out-of-Distribution Detection by Object Class Prompting of the SAM3-Image Model
Core Problem: However, they are prone to overconfidence when encountering unseen objects in real-world deployments, causing potential safety issues.
Key Innovation: However, they are prone to overconfidence when encountering unseen objects in real-world deployments, causing potential safety issues. Experimental results show that our method significantly surpasses the so-far zero-shot SOTA method.
250. To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models
Core Problem: We conduct a per-sample analysis of model predictions before and after adaptation, and observe two failure modes in existing TTA methods that echo previous work.
Key Innovation: In this work, we introduce a new problem of selective adaptation, which aims to determine whether a given test sample should undergo adaptation or be skipped.
251. IPM-FM: A Foundation Model with Consensus Feature Selection for Industrial Process Monitoring
Core Problem: Current practice remains a one-task-one-model paradigm that is label-inefficient and prone to degradation under operating drift.
Key Innovation: We propose the industrial process monitoring foundation model (IPM-FM). On a seven-year hydrotreater dataset for diesel flash-point soft sensing, IPM-FM attains an RMSE of 2.99, R² of 0.50, and 97% coverage of its 95% predictive interval, outperforming the strongest classical and from-scratch sequence baselines by 8.3% and 14.6% in RMSE respectively, supporting the viability of a unified pretraining--adaptation framework for.
252. GSComplete: Gaussian Splat Completion with 2D Diffusion Priors
Core Problem: Existing completion methods either do not preserve the original splats or require scarcely available 3D training data.
Key Innovation: We propose GSComplete, which combines 3D generation based on Score Distillation Sampling with a novel preservation loss that encourages the original splats to be preserved where they should be visible. To evaluate our approach, we introduce a new dataset of partial Gaussian splat objects and show that GSComplete achieves significantly more accurate preservation of the input than existing methods with comparable plausibility.
253. Layer Selection in VLMs for Zero-Shot OOD Detection via Multi-Resolution Entropy Estimation
Core Problem: Out-of-distribution (OOD) detection is crucial for safe deployment of medical AI systems, where domain shifts arise across institutions, acquisition protocols, and patient populations.
Key Innovation: To address this instability, we propose a multi-resolution entropy estimation strategy that aggregates histogram statistics across multiple discretization scales, enabling robust and stable intermediate-layer selection. We first show that this assumption does not hold in medical imaging: intermediate layers provide complementary OOD signals, and the optimal representational depth depends on the respective image modality.
254. From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video
Core Problem: To address these challenges, we introduce Coherent4D, a large-scale egocentric dataset for continuous 4D interaction forecasting, comprising approximately 233K samples across three domains.
Key Innovation: To address these challenges, we introduce Coherent4D, a large-scale egocentric dataset for continuous 4D interaction forecasting, comprising approximately 233K samples across three domains. Extensive experiments across all three domains demonstrate consistent improvements over representative baselines on both location and pose forecasting, while ablations validate the contributions of the proposed components.
255. TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection
Core Problem: Onboard object detection in Earth observation is constrained by limited computational resources and the absence of fully corrected imagery.
Key Innovation: We introduce TriCCOT, a tri-part architecture for robust and deployable onboard object detection. Experiments on the DIOR and VDVRaw datasets demonstrate competitive detection performance and improved robustness to spatial blur and signal-dependent noise when compared to FPGA-compatible architectures.
256. Beyond Gait: Person Identification from Millimeter-Wave Point Clouds Across Activities of Daily Living
Core Problem: Indoor walking, however, is often brief and interrupted, while other activities of daily living (ADLs) may provide complementary identity information.
Key Innovation: This extension introduces heterogeneous states and transitions whose spatial and temporal characteristics vary with activity. Under a matched gallery partition, activity-specific experts also outperform a shared embedding, showing that the gain extends beyond restricting the gallery.
257. Length Generalization for Transformers via Compression
Core Problem: While this hypothesis has strong empirical validation, theoretical problems arise from the fact that no computable length generalization bounds exist for C-RASP, alongside the discovery of seemingly contradictory experiments.
Key Innovation: In particular, the C-RASP hypothesis (a formalized version of the so-called RASP-l conjecture) posits that transformers length-generalize on a task if and only if a solution is expressible in the C-RASP language. In doing so, we show a polynomial length generalization bound for transformers if we adopt compressed strings, via a novel connection to power words.
258. Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling
Core Problem: A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates.
Key Innovation: Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression.
259. Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks
Core Problem: In such settings, early discrete errors can be difficult to undo.
Key Innovation: To reduce this train-test mismatch, we further introduce self-correction training, which exposes the model to its own predictions, improving robustness to errors that arise during inference.
260. "World Knowledge" in the Weights: Reading Concept Circuits of Vision Transformers
Core Problem: To address this gap, we use Cross-Layer Transcoders (CLTs) to read concept circuits from ViTs: directed graphs whose nodes correspond to sparse, interpretable concepts and edges capture concept interactions across layers.
Key Innovation: To address this gap, we use Cross-Layer Transcoders (CLTs) to read concept circuits from ViTs: directed graphs whose nodes correspond to sparse, interpretable concepts and edges capture concept interactions across layers.
261. Multi-Task Learning for Sparsely-Labeled Time Series: A Case Study on Cold-Hardiness Modeling
Core Problem: We present a real-world case study of multi-task learning (MTL) for temporal process modeling from limited data with temporally sparse labels.
Key Innovation: We present a real-world case study of multi-task learning (MTL) for temporal process modeling from limited data with temporally sparse labels. Our results show significant differences between architectures and that certain architectures are able to consistently outperform single-task learning and state-of-the-art scientific models.
262. Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Core Problem: Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism.
Key Innovation: To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.
263. Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Core Problem: Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations.
Key Innovation: These evaluations do not fully capture how visual tokens behave when modeled jointly with text.
264. Damage-Aware Bandit Pruning for Vision and Language Transformers
Core Problem: Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation.
Key Innovation: We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fixed candidate-evaluation budget. Matched-evaluation results for ViT-B/16 and Swin-Tiny indicate that their gains are not explained solely by a larger candidate-evaluation budget.
265. ModularPhaseNet: Finite-Cyclic Phase Geometry for Computable Semantic Hierarchy, Direction, and Context Consistency in Standard Transformers
Core Problem: We propose ModularPhaseNet, a classical and integer-computable discretization of the continuous complex phase geometry introduced in QuantumPhaseNet.
Key Innovation: We propose ModularPhaseNet, a classical and integer-computable discretization of the continuous complex phase geometry introduced in QuantumPhaseNet.
266. Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model
Core Problem: We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3) with an adaptive arithmetic coder, guaranteeing |\hat{x}_t-x_t|≤\tau on every sample.
Key Innovation: We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3) with an adaptive arithmetic coder, guaranteeing |\hat{x}_t-x_t|≤\tau on every sample.
267. LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies
Core Problem: However, existing VLA interfaces offer limited flexibility in representation access: VLM information is exposed through fixed layer assignments for each action layer, while intermediate action states are only propagated implicitly through residual streams without explicit reuse.
Key Innovation: We introduce LayerRoute, an action-conditioned representation routing interface that enables adaptive access to VLM layers and action representations. Across diverse simulation and real-world benchmarks, LayerRoute consistently improves StarVLA-\pi and \pi0.5}, achieving up to 7.2 gains on LIBERO Long with only 0.31% / 3.87% additional parameters.
268. Decomposing LLM-Judge Uncertainty to Target Expert Labels
Core Problem: Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which labels do reduce.
Key Innovation: Experts should label only where it is least sure. We demonstrate we can estimate where a judge is ignorant rather than where experts genuinely disagree, and propose using this to direct expert labelling.
269. TD-STGT: A Spatio-Temporal Graph Transformer for Mobile Traffic Demand Forecasting
Core Problem: Fine-grained mobile traffic demand forecasting is essential for long-term planning of 5G and future 6G networks, including radio upgrades, site densification, backhaul expansion, and spectrum activation.
Key Innovation: This paper proposes the Traffic Demand Spatio-Temporal Graph Transformer (TD-STGT), a graph neural forecasting framework for predicting changes in wireless mobile traffic demand across fine geographic grids. Experiments across five Canadian metropolitan regions show that TD-STGT achieves the best performance in forecasting grid-level demand changes, reaching a \Delta R² of 0.462 and reducing \DeltaRMSE by 5.7% relative to the.
270. Distributed Dexterous Manipulation with Spatially Conditioned Multi-Agent Transformers
Core Problem: Distributed Dexterous Manipulation (DDM) is a novel paradigm that presents significant control challenges due to high action-space redundancy, inter-robot cooperation, and dynamic object-robot interactions.
Key Innovation: This paper introduces a framework based on spatially conditioned Multi-Agent Transformers (MATs) to efficiently learn robust control policies for a DDM system grounded in an array of 64 soft delta robots arranged in an 8x8 grid. Our experiments show that MATs iteratively refine their actions through the stacked attention blocks.
271. Weakly supervised neural network: segmentation of complex structures in X-ray microCT
Core Problem: Segmentation of complex structures in X-ray tomographic data is a fundamental task in biomedical research, but it often requires large amounts of precisely annotated data, making fully supervised approaches costly and difficult to scale.
Key Innovation: In this study, weakly supervised deep learning is investigated as a strategy to reduce annotation effort while maintaining accurate segmentation. Results indicate that weak supervision provides a meaningful learning signal, enabling reliable localization of glomeruli even in the absence of dense labels.
272. Graph neural networks and the energetic cavity method for combinatorial optimization
Core Problem: Efficiently finding these ground states is of broad significance because many combinatorial optimization problems can be formulated as an Ising model with the appropriate choice of couplings and fields.
Key Innovation: Efficiently finding these ground states is of broad significance because many combinatorial optimization problems can be formulated as an Ising model with the appropriate choice of couplings and fields.
273. The Role of Uncertainty in Assessing the Fairness of Machine Learning Models
Core Problem: A rigorous risk assessment of possible fairness violations requires quantifying the uncertainty associated with selecting and estimating such models.
Key Innovation: Verifying whether their outputs are biased against disadvantaged groups or individuals is crucial to ensuring they are fair and allowing their use in such settings.
274. Bayesian Matrix-Valued Graphs for Context-Dependent Multivariate Relationships
Core Problem: Many scientific graphs attach several variables to each node, so a single scalar edge weight cannot describe direction-dependent interactions.
Key Innovation: These results establish posterior matrix-valued edge geometry as a unified framework for quantifying and interpreting context-dependent multivariate reconfiguration.
275. Inclusive electron-nucleus cross section models from domain adaptation
Core Problem: We apply transfer learning (TL) to construct data-driven models of inclusive electron-nucleus cross sections.
Key Innovation: Starting from an ensemble of deep neural networks pretrained on ¹²C data, we fine-tune the models separately for ³He, ⁶Li, ¹⁶O, ²⁷Al, ⁴⁰Ca, and ⁵⁶Fe. The layer-wise analysis shows that oxygen requires only shallow adaptation, whereas helium, calcium, and iron require substantially deeper fine-tuning.
276. Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
Core Problem: However, preference labels provided in existing datasets are blended with layout and aesthetic opinions, which would disagree with aesthetic preference.
Key Innovation: To improve aesthetics economically, this paper uses existing generic preference data and introduces step-by-step preference optimization (SPO) that discards the propagation strategy and allows fine-grained image details to be assessed.
277. DUA-D2C: Dynamic Uncertainty Aware Method for Overfitting Remediation in Deep Learning
Core Problem: Overfitting remains a significant challenge in deep learning, often arising from data outliers, noise, and limited training data.
Key Innovation: Building upon this foundation, we introduce Dynamic Uncertainty-Aware Divide2Conquer (DUA-D2C), an advanced technique that refines the aggregation process. In this work, we provide a rigorous theoretical justification for this approach, analytically demonstrating how dynamic parameter fusion reduces model variance.
278. ControlTac: Scaling Tactile Data with Physically Controlled Tactile Image Generation
Core Problem: Vision-based tactile sensing is widely used in perception, reconstruction, and robotic manipulation, yet collecting large-scale tactile data remains costly due to diverse sensor-object interactions and inconsistencies across sensor instances.
Key Innovation: We propose \name, a two-stage controllable tactile image generation framework that generates realistic tactile images conditioned on a single reference tactile image, contact force, and contact pose. Across a series of downstream tasks and real-world experiments, such as object insertion, imitation learning, and object weighting, the augmented datasets using our approach consistently improve performance and demonstrate.
279. ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
Core Problem: To address these issues, we propose ID-Align, which alleviates these problems by reordering position IDs.
Key Innovation: To address these issues, we propose ID-Align, which alleviates these problems by reordering position IDs. Our experiments conducted within the LLaVA-Next framework demonstrate that ID-Align achieves significant improvements, including a 6.09% enhancement on MMBench's relation reasoning tasks and notable gains across multiple benchmarks.
280. Adaptive Nonlinear Vector Autoregression: Robust Forecasting for Noisy Chaotic Time Series
Core Problem: However, their reliance on fixed nonlinear transformations - polynomial expansions in NVAR or random feature maps in RC - limits their adaptability to high noise or complex real-world data.
Key Innovation: We propose a data-adaptive NVAR model that combines delay-embedded linear inputs with features generated by a shallow, trainable multilayer perceptron (MLP). Initial experiments across multiple chaotic systems, tested under noise-free and synthetically noisy conditions, showed that the adaptive model outperformed in predictive accuracy the standard NVAR, a leaky echo state network (ESN) - the most common RC model - and a.
281. DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
Core Problem: Existing text-to-image diffusion models excel at generating high-quality images, but face significant efficiency challenges when scaled to high resolutions, like 4K image generation.
Key Innovation: To bridge this gap, this paper introduces DC-Gen, a general framework that accelerates text-to-image diffusion models by leveraging a deeply compressed latent space. The resulting DC-Gen-SANA and DC-Gen-FLUX models achieve quality comparable to their base models but with a significant speedup.
282. Cascaded Diffusion Framework for Probabilistic Coarse-to-Fine Hand Pose Estimation
Core Problem: Existing cascaded approaches progressively refine pose predictions in a coarse-to-fine manner, but their deterministic nature prevents them from modeling pose uncertainty.
Key Innovation: To address these limitations, we propose a coarse-to-fine cascaded diffusion framework that combines probabilistic modeling with cascaded refinement. Experiments on FreiHAND, HO3Dv2, and DexYCB show that our method achieves state-of-the-art performance and remains stable under occlusion and pose ambiguity.
283. CylindTrack: Depth-Aware Cylindrical Motion Modeling for Panoramic Multi-Object Tracking
Core Problem: To address these challenges, we propose CylindTrack, a depth-aware cylindrical tracking-by-detection framework for panoramic MOT.
Key Innovation: Panoramic cameras offer wide surrounding coverage, but equirectangular projection introduces a periodic horizontal domain in which conventional planar motion models and IoU-based association become unreliable near the 0°/360° seam. Experiments on QuadTrack and JRDB achieve 33.67/31.12 HOTA and 40.45/34.33 IDF1 at 28.56/21.34 FPS, demonstrating the effectiveness and practical online efficiency of CylindTrack as a persistent.
284. Learning Subgroup Relations Using Siamese Graph Neural Networks
Core Problem: Determining whether one finite group is isomorphic to a subgroup of another is a fundamental problem in computational group theory.
Key Innovation: In this work, we propose a Siamese Graph Neural Network (Siamese GNN) for subgroup prediction using Cayley graph representations of finite groups. Experimental results on an expanded and more diverse dataset of 308 finite-group pairs drawn from 11 group families demonstrate the effectiveness of the proposed architecture, achieving a test BA of 91.67% on an independent test set.
285. On Generalisation Error Bounds for Transformers
Core Problem: In this paper, we establish a collection of covering number bounds for linear function classes under various norm constraints on the inputs and matrices.
Key Innovation: We then combine these results with existing covering number bounds to derive improved estimates and, based on these estimates, develop generalization error bounds for single-layer Transformers.
286. RoMu4o: A Robotic Manipulation Unit For Orchard Operations Automating Proximal Hyperspectral Leaf Sensing
Core Problem: Driven by the need to address labor shortages and meet the demands of a rapidly growing population, robotic automation has become a critical component in precision agriculture.
Key Innovation: This work introduces RoMu4o, a robotic manipulation unit for orchard operations offering an automated solution for proximal hyperspectral leaf sensing. Leaf-level hyperspectral spectroscopy is shown to be a powerful tool for phenotyping, monitoring crop health, identifying essential nutrients within plants as well as detecting diseases and water stress.
287. Fewer yet critical: Reducing Redundant Token Dependencies for Transformer-based Time Series Forecasting
Core Problem: However, existing methods are usually trained mainly with prediction error losses, which may cause models to exploit both critical and redundant token dependencies.
Key Innovation: Such redundant dependencies can introduce irrelevant information and weaken generalization. Recently, Transformer-based methods have achieved strong performance by modeling token dependencies through attention mechanisms.
288. Synthesizability Prediction of Crystalline Structures with Structure-Aware Feature Learning and Uncertainty Quantification
Core Problem: Predicting which hypothetical inorganic crystals can be experimentally realized remains a central challenge in accelerating materials discovery.
Key Innovation: SyntheFormer is a positive-unlabeled framework that learns synthesizability directly from crystal structure, combining Fourier-transformed crystal properties (FTCP) representation with structure-aware feature extraction, Random-Forest feature selection, and a compact deep MLP classifier. Under this temporally separated evaluation, SyntheFormer achieves a test AUC of 0.735, AUPRC of 0.099 and 97.6 percent recall at 94.2.
289. Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
Core Problem: Autonomous robotic navigation in real-world environments requires exploration to acquire environmental information as well as goal-directed navigation in order to reach specified targets.
Key Innovation: Active inference (AIF) based on the free-energy principle provides a unified framework for these behaviors by minimizing the expected free energy (EFE), thereby combining epistemic and extrinsic values. Real-world navigation experiments, including baseline comparisons, component ablations, and robustness evaluations, demonstrated that our framework achieved higher success rates and fewer collisions, particularly in.
290. Sustained Performance and Energy Accounting for Nonlinear Forecasting Across Classical and Simulated Quantum Models
Core Problem: Energy-efficient AI should be evaluated across the full application pipeline, not only by lowest error or shortest training time.
Key Innovation: We study this through nonlinear time-series forecasting using simulated quantum reservoir computing (QRC) as an emerging-computing case study. The best mean NRMSE is achieved by continuous-variable QRC with a Transformer readout (0.332), followed closely by classical TCN (0.341) and QRC+TCN (0.340).
291. A novel measurement approach for harbor ship monitoring based on time-domain electromagnetics
Core Problem: With increasing density of vessel traffic in port environments, reliable and high-resolution monitoring of ship movements has become a critical measurement challenge for maritime management and safety applications.
Key Innovation: Transient electromagnetic (TEM) techniques provide a non-contact sensing tool by measuring secondary electromagnetic fields induced by conductive targets to offer high temporal resolution and sensitivity to metallic structures.
292. Hello Equal Earth: what the UN’s new world map will change
Core Problem: Nature, Published online: 09 September 2026; doi:10.1038/d41586-026-02820-x What the UN vote to ditch the Mercator projection means for education, navigation and public perception.
Key Innovation: Nature, Published online: 09 September 2026; doi:10.1038/d41586-026-02820-x What the UN vote to ditch the Mercator projection means for education, navigation and public perception.
293. Bridging Canopy Light Interception and Absorption: Toward a Multi-Scale Radiometric Framework for Precision Irrigation in Woody Crops
Core Problem: Optimizing irrigation in orchards is increasingly challenged by climatic variability, water scarcity, and structural heterogeneity, requiring indicators that robustly link canopy architecture with transpiration and energy exchange processes.
Key Innovation: We suggest a multi-scale framework in which fIPAR is a structural-radiative constraint on potential canopy water demand. The conditions under which fIPAR approximates fAPAR are examined, revealing that their divergence is more relevant in discontinuous orchard systems than in homogeneous canopies.