社区所有版块导航
Python
python开源   Django   Python   DjangoApp   pycharm  
DATA
docker   Elasticsearch  
aigc
aigc   chatgpt  
WEB开发
linux   MongoDB   Redis   DATABASE   NGINX   其他Web框架   web工具   zookeeper   tornado   NoSql   Bootstrap   js   peewee   Git   bottle   IE   MQ   Jquery  
机器学习
机器学习算法  
Python88.com
反馈   公告   社区推广  
产品
短视频  
印度
印度  
Py学习  »  机器学习算法

机器学习学术速递[9.14]

arXiv每日学术速递 • 1 周前 • 100 次点击  

2026-09-14 | CS.LG机器学习 | 共 99 篇

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 深度学习架构与训练方法 8 篇

2. 表示学习、自监督与对比学习 2 篇

3. 强化学习与序列决策 17 篇

4. 生成模型与概率建模 3 篇

5. 优化、泛化与理论分析 8 篇

6. 联邦学习、隐私与安全 5 篇

7. 鲁棒性、不确定性与可信学习 3 篇

8. 图学习与结构化数据 2 篇

9. 迁移、元学习与持续学习 5 篇

10. 数据集、基准与评测 5 篇

11. 机器学习应用 5 篇

12. 其他/综合机器学习 36 篇

1. 深度学习架构与训练方法 | 8 篇

1. DCRA: Diffusion-Conditioned Representation Alignment for Robust Time-Series Learning

DCRA:扩散条件表示对齐用于鲁棒时间序列学习

AI 总结:DCRA通过将扩散前向过程作为结构化损坏调度器并引入特征级一致性目标,在噪声下对齐时间序列表示,提升EEG癫痫检测的鲁棒性和灵敏度。

链接:https://arxiv.org/abs/2609.11997

机构:University of Minnesota, Twin Cities(明尼苏达大学双城分校)

作者:Wenrui Xu, Anas Enanaa, Keshab K. Parhi

英文摘要:Learning robust representations for time-series signals under noise and distribution shifts remains challenging, especially in clinical applications such as electroencephalogram (EEG) and electrocardiogram (ECG) analysis. We propose Diffusion-Conditioned Representation Alignment (DCRA), a training framework that repurposes the forward diffusion process as a structured corruption scheduler for representation learning. Different from conventional augmentation and consistency-based methods that rely on independently sampled perturbations, DCRA introduces a structured corruption trajectory via the diffusion forward process, which enables continuous and controlled representation evolution across noise levels. We introduce a feature-level consistency objective that aligns representations across noise levels while preserving class-discriminative structure. This mechanism promotes structure-preserving consistency, which enables smooth and semantically coherent feature trajectories in latent space. The proposed framework is encoder-agnostic and can be integrated with state space models and Transformer architectures. The seizure detection experiments on the CHB-MIT EEG dataset show that DCRA consistently improves performance under multiple noise conditions and achieves higher sensitivity at low false-positive rates. Analysis reveals that DCRA produces more balanced and structured representations compared to baseline and diffusion-only models. These findings highlight the benefit of combining structured corruption with representation alignment for robust time-series learning.

2. Predicting Collision Cross Sections with GRACE: Geometric Residual Adduct Conditioning via Early-fusion

使用GRACE预测碰撞截面:基于早期融合的几何残差加合物条件化

AI 总结:GRACE通过早期融合的几何残差加合物条件化,利用预训练几何编码器预测碰撞截面,在多个数据集上取得最优精度,证明残差学习与编码器级条件化的有效性。

链接:https://arxiv.org/abs/2609.12223

机构:IBM Research(IBM研究院); Center for Computational Life Sciences, Cleveland Clinic(克利夫兰诊所计算生命科学中心); Department of Chemistry, Michigan State University(密歇根州立大学化学系)

作者:Parthasarathy Suryanarayanan, Susanta Das, Shreyans Sethi, Kenneth M. Merz Jr., Joseph A. Morrone

英文摘要:Collision cross section (CCS), derived from ion mobility mass spectrometry, is a common descriptor for molecular annotation. Prediction is challenging for machine learning models because it reflects the size, shape, and ionization state of a gas-phase molecular ion. Most predictors either ignore explicit 3D structure or treat adduct identity as a late categorical feature, which limits their ability to capture adduct-dependent geometric effects. We present GRACE (Geometric Residual Adduct Conditioning via Early-fusion), a 3D CCS predictor that adapts a pretrained molecular geometry encoder using geometric residual adduct conditioning via early fusion. GRACE combines two inductive biases: a residual objective relative to an adduct-aware physical descriptor baseline and adduct conditioning within the encoder via a learned adduct token and low-rank attention adapters. We evaluate the model on a curated set of over 9,000 experimental molecule-adduct CCS records with random, scaffold, and adduct-sensitive splits designed to separate interpolation, scaffold generalization, and adduct-driven generalization. GRACE achieves the best mean percentage difference among the evaluated learned models on all three splits: 1.67% on the random split, 2.11% on the scaffold split, and 2.36% on the adduct-sensitive split. Diagnostic analyses suggest that residual learning stabilizes training by removing the dominant mass-CCS trend, while early fusion improves adduct-sensitive prediction relative to late fusion. Across four independent external test sets, GRACE shows consistently lower error than the other evaluated models. On a held-out set, GRACE also attains the lowest mean percent difference when compared with four previously reported physics-based workflows. These results support residual learning and encoder-level adduct conditioning as practical inductive biases for fast, accurate CCS prediction.

3. Patient-Reported Survey Data Improve Prediction of Opioid Use Disorder

患者报告调查数据改善阿片类药物使用障碍的预测

AI 总结: 本研究利用All of Us数据,证明患者报告调查数据可显著提升阿片类药物使用障碍(OUD)的预测性能,其中24个月LightGBM模型PR-AUC提升至0.6603,并强调调查可用性的重要性。

链接:https://arxiv.org/abs/2609.12224

机构:Stony Brook University(石溪大学); Stony Brook Medicine(石溪医学中心)

作者:Xiyue Jiang, Zihan Ding, Grace Han, Yinan Liu, Richard N. Rosenthal, Fusheng Wang

英文摘要:Electronic health records (EHRs) may incompletely capture patient-reported factors associated with opioid use disorder (OUD). We evaluated whether survey data improve prediction of a first recorded OUD diagnosis among 267,747 All of Us participants with documented opioid exposure, including 15,287 OUD cases. We compared EHR-only and EHR+survey models across 6-, 12-, and 24-month look-back windows using logistic regression, random forest, XGBoost, LightGBM, multilayer perceptron, LSTM, GRU, and Transformer. Survey augmentation improved PR-AUC across all 24 model-window combinations by 0.0087-0.0505; the best 24-month LightGBM model improved from 0.6219 to 0.6603. Survey coverage increased with longer windows and differed by OUD status (24 months: 21.7% OUD-positive vs. 60.7% OUD-negative). Permutation analysis ranked survey features as the second most important information domain at 24 months in both evaluated models. Patient-reported data provide complementary predictive signals beyond structured EHRs while highlighting the importance of survey availability.

4. RiPPLE: Cross-Space Performance Prediction from Early Training for Neural Architecture Search

RiPPLE:基于早期训练的跨空间性能预测用于神经架构搜索

AI 总结:针对NAS评估昂贵问题,提出RiPPLE方法,利用早期训练锚点外推标签并传播,实现低成本跨空间性能排序。

链接:https://arxiv.org/abs/2609.12418

机构:University of New South Wales(新南威尔士大学); Korea Advanced Institute of Science and Technology(韩国科学技术院)

作者:Yifan Yang, Zhaoyan Wang, Zheng Gao, Xiaoyu Li, Jiaojiao Jiang

英文摘要:Neural architecture search (NAS) evaluates candidate networks, but fully training enough architectures to rank an entire space is expensive. Zero-cost proxies score architectures at initialization, yet their ranking quality varies across search spaces. Learned predictors reduce evaluation cost but typically require fully trained labels or partial-training features for individual candidates. We introduce $\textbf{RiPPLE}$, $\underline{\textbf{R}}$anking v$\underline{\textbf{i}}$a $\underline{\textbf{P}}$refix-$\underline{\textbf{P}}$ropagated $\underline{\textbf{L}}$abel $\underline{\textbf{E}}$xtrapolation, which treats partial training as a source of labels for a small coverage set of anchors. RiPPLE trains these anchors to an early prefix, extrapolates their learning curves to surrogate labels, and propagates the labels over label-free architecture features. The early-training signal remains a label on the anchors rather than a per-candidate feature. Feature, readout, and encoding rules are selected without held-out accuracy and reused across search spaces. We evaluate the method on twelve benchmark cells from four search-space families and on the larger DARTS space. The results examine ranking quality, label efficiency, architecture selection, and the roles of readout, coverage, and propagation. RiPPLE provides a whole-space ranking from a fractional anchor-training budget, with comparisons interpreted under their respective evaluation and cost protocols.

5. Quality-Constrained Routing over a Fixed Pool of Quantized Mixture-of-Experts Instances

固定量化混合专家实例池上的质量约束路由

AI 总结:针对固定量化MoE实例池,提出FWP脆弱性加权困惑度预测请求级质量风险,并通过窗口级线性规划路由,在质量预算下最大化吞吐,较静态W4提升28.4%吞吐。

链接:https://arxiv.org/abs/2609.12550

机构:The Hong Kong University of Science and Technology(香港科技大学)

作者:Zhenghong Huang, Hongfan Wu, Jiheng Zhang

英文摘要:Quantized Mixture-of-Experts (MoE) services can hold several pre-materialized instances of one base model, but quantization damage varies sharply across requests and bitwidths. Because instance materialization and replica counts consume memory and require slow reconfiguration, we treat them as upstream provisioning decisions and study routing within a fixed resident pool. Within this fixed-pool boundary, we route each request to maximize modeled throughput under a class-level expected quality-degradation budget and measured instance capacities. To predict this request-specific risk, we introduce FWP (Fragility-Weighted Perplexity), computed from prompt tokens on a reference-instance prefill and calibrated to candidate-instance degradation. Underlying FWP is an exact two-expert affinity--fragility decomposition and a conditional multi-layer top-$k$ expansion whose bias, interaction, route-change, separability, and higher-order terms remain explicit. Using these calibrated risks, a window-level linear program yields a signed reduced-reward score that is KKT-consistent with the LP optimum under optimal prices and primal-feasible tie allocation. On 88 extended Qwen prompts, complete W2, W3, and W4 instances quantizing all 6,144 expert blocks incur mean $\Delta$NLL of $0.9437$, $0.1832$, and $0.0513$. Under the same population and $\tau=0.1513$, FWP allocation reaches a $1.284\times$ offline model-based multiplier versus $1.253\times$ for request-agnostic mixing and $1.000\times$ for static W4, an incremental $2.5\%$ relative FWP gain.

6. Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery

解码器余弦相似度在SAE特征流发现中的失效场景

AI 总结:本文发现SAE特征流发现中解码器余弦相似度在MLP更新场景下失效,通过构建转移图谱和消融验证,在Pythia和Gemma模型中揭示了大量低余弦相似度但强效应的转移,为表示-更新机制诊断提供新工具。

链接:https://arxiv.org/abs/2609.12591

机构:Hasso Plattner Institute(哈索·普拉特纳研究所); IT University of Copenhagen(哥本哈根信息技术大学)

作者:Hendrik Droste, Christian Medeiros Adriano, Kathrin Korte, Holger Giese

英文摘要: Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities. Understanding the internal computations that support these adaptations is therefore becoming increasingly important for continual model evolution. Sparse autoencoders (SAEs) provide interpretable feature dictionaries for residual-stream activations and sublayer outputs, but it remains unclear how state features and update features interact to produce downstream residual features. In this work, we focus on MLP updates as a first test case. We construct a transition atlas of triples $s_k + u_j \rightarrow t_\ell$, where a residual-state feature and an MLP-update feature jointly predict a target residual feature, and validate candidate triples by ablating the decoded update feature. In a 20M-token Pythia-160M $L_7 \rightarrow L_8$ run, we find 38,125 strong ablation-effect transitions, but 88.0% have both state-target and update-target decoder cosine similarity below 0.7. As a preliminary cross-model check, a run of 20M-token Gemma-3-4B $L_{21} \rightarrow L_{22}$ causally validates only the top 30,000 ranked candidate triples by ablating the decoded update feature, and 53.6% of strong-effect triples have both state-target and update-target decoder cosine similarity below 0.7. The Gemma result is directionally consistent with Pythia, but weaker, since update-target cosine recovers many of the strongest Gemma effects and the run is not a full-atlas causal validation. Ultimately, our results suggest that feature flow atlases can serve as diagnostics of representation-update mechanisms and thereby inform tools for steering model updates. Future work will validate more complex patterns across layers, models, and SAE families.

7. SIFPBPNet: A Dual-Path Network for Wearable and Cuffless Blood Pressure Estimation via Individualized Steady-state Representation

SIFPBPNet:一种通过个体化稳态表示进行可穿戴无袖带血压估计的双路径网络

AI 总结:针对PPG血压估计中人群异质性和一对多映射问题,提出双路径网络SIFPBPNet,通过稳态与瞬时特征分离及交叉注意力融合,在可穿戴数据集上MAE达8.57/5.97 mmHg,且SFP模块可即插即用提升多种模型性能。

链接:https://arxiv.org/abs/2609.12690

机构:Shenzhen Technology University(深圳技术大学); OPPO Health Lab(OPPO健康实验室); Peking University(北京大学); Southeast University(东南大学)

作者:Shuailong Tang, Xiaoyu Li, Donglin Xie, Wei Chen, Guangpu Zhu, Yelei Li, Yali Zheng

英文摘要:Continuous and cuffless blood pressure (BP) monitoring using photoplethysmography (PPG) is of great interest for low-cost and personalized cardiovascular health management. However, significant population heterogeneity and the "one-to-many mapping" problem, where similar waveforms across individuals correspond to different BP levels, limit the accuracy of conventional population-based models. To address this challenge, we propose a dual-path architecture termed SIFPBPNet, which separately represents steady-state and instantaneous features, through a Steady-state Feature Path (SFP) and an Instantaneous Feature Path (IFP). The SFP employs a Graph Attention Network (GAT) to extract individual-specific and long-term characteristics from multi-day historical PPG trajectories. In parallel, the IFP captures short-term dynamics from current PPG segments and incorporates the steady-state prior via a cross-attention mechanism. Experiments on a large-scale wearable dataset demonstrate that SIFPBPNet achieves a Mean Absolute Error (MAE) of 8.57 and 5.97 mmHg for systolic and diastolic BP, respectively, outperforming state-of-the-art models. Furthermore, the SFP module consistently improves performance when integrated into various backbone architectures, yielding 2.8-13.1% relative MAE reductions for systolic BP. These results highlight the strong generalizability and plug-and-play transferability of the SFP module, underscoring its great potential for accurate cuffless BP monitoring.

8. RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States

RunningTensor:将线性注意力推广到高阶循环状态

AI 总结:RunningTensor将线性注意力的矩阵记忆推广到高阶张量,通过秩1外积更新和向量查询收缩,在保持线性时间的同时提升记忆容量,并在多查询关联回忆及语言理解任务上优于基线。

链接:https://arxiv.org/abs/2609.12814

机构:ISIR Sorbonne(索邦大学ISIR); Berkeley Lab(伯克利实验室)

作者:Luca Herranz-Celotti, Vincent Guigue

英文摘要:Linear attention and state-space models provide linear-time sequence modeling, but their recurrent memory remains a second-order tensor (a matrix), limiting the order of interactions that can be represented in the state. We introduce the RunningTensor, which generalizes this memory to an order-$o$ tensor, updated by a rank-1 outer product and read by contracting against $o-1$ vector queries. Order $2$ recovers linear attention; we study order $3$ as a proof of concept, retaining both recurrent and parallel forms while remaining linear in sequence length $T$ and improving working memory capacity from $\mathcal{O}(W^2)$ to $\mathcal{O}(W^o)$. On synthetic multi-query associative recall, RunningTensor outperforms linear-attention and SSM baselines. After pretraining, it also improves performance on language-understanding and non-synthetic retrieval tasks, suggesting that higher-order recurrent state can provide useful additional memory capacity beyond matrix-valued state.

2. 表示学习、自监督与对比学习 | 2 篇

9. GUIDE: Generative Utility Inference and Decision Engine

GUIDE:生成式效用推断与决策引擎

AI 总结:GUIDE是一种LLM驱动的偏好引出架构,通过贝叶斯自适应采样和符号表示学习,在对话中高效推断多维用户偏好,并在投资组合优化实验中改善冷启动、降低推荐遗憾。

链接:https://arxiv.org/abs/2609.12137

机构:University of Chicago(芝加哥大学); Centaur AI Institute(Centaur AI研究所); Carnegie Mellon University(卡内基梅隆大学); Booth School of Business(布斯商学院)

作者:Anagha Tiwari, Alexander G. Gray, Nick Feamster, Brian Jabarian, Alex Imas, Alex Kale

英文摘要: Measuring the preferences of human users remains a fundamental challenge of AI alignment. Existing elicitation approaches struggle to efficiently discover multidimensional preferences or accurately ground these inferences in domain knowledge. To address this, we introduce GUIDE, an LLM-driven elicitation architecture that infers user preferences through conversations by combining Bayesian adaptive sampling for question selection and symbolic representation learning to initialize domain-specific preference models. GUIDE generalizes adaptive sampling to diverse elicitation questions through an extensible type system of transforms on a parameterized preference state. GUIDE produces domain-specific preference representations through an initialization process using symbolic rule-based learning to capture world knowledge and set priors over preference dimensions grounded in data about decision alternatives. The architecture provides observability and steerability to facilitate deployment and analyze elicitation processes. In silico experiments on investment portfolio optimization demonstrate that GUIDE improves cold-start and minimizes recommendation regret consistently within early elicitation interactions across user personas compared to prior work, LLM-only baselines, and ablated GUIDE versions.

10. FRIST: FMRI Representation Informed Shared-space Training Improves EEG-only Individual-Finger BCI Decoding

FRIST:基于fMRI表征信息的共享空间训练提升仅用EEG的单指脑机接口解码

AI 总结:FRIST利用fMRI空间信息指导EEG表征学习,通过两阶段框架提升仅用EEG的单指BCI解码准确率,在多项实验中显著优于基线。

链接:https://arxiv.org/abs/2609.12298

作者:Jintao Zhang, Yidan Ding, Joshua Kosnoff, Maxim Karrenbach, Hanwen Wang, Bin He

英文摘要:Finger-level motor decoding is important for naturalistic brain-computer interface (BCI) control, yet individual-finger decoding from scalp electroencephalography (EEG) remains challenging because finger representations are spatially close in the sensorimotor cortex and blurred by volume conduction. Leveraging the high spatial resolution of functional MRI (fMRI), we introduce fMRI Representation-Informed Shared-Space Training (FRIST), a two-stage EEG decoding framework that first learns fMRI-informed spectral projections from simultaneous EEG-fMRI recordings and then uses fMRI-derived class geometry to guide residual refinement of EEG predictions. FRIST transfers information across recordings through shared finger labels without requiring paired trials and uses only EEG at inference. We evaluated 12 able-bodied participants during movement execution (ME) and motor imagery (MI) under two-class and three-class chronological session-held-out decoding simulating the online scenario. Using EEGNet as the EEG feature extractor, FRIST increased group average accuracy from 66.93% to 74.53% for two-class ME, from 44.83% to 56.58% for three-class ME, from 80.78% to 85.63% for two-class MI, and from 60.93% to 69.90% for three-class MI compared with the EEG-only EEGNet baseline. FRIST is also shown to improve EEG-only decoding when the target participant's own fMRI data were unavailable. FRIST also generalized across multiple EEG decoding backbones, reaching 87.40% in two-class MI and 72.54% in three-class MI with EEG Conformer as the EEG feature extractor. These findings indicate that fMRI provide useful spatial constraints for EEG representation learning. FRIST improves noninvasive EEG-based finger-level BCI decoding, offering a multimodal strategy for integrating the spatial specificity of fMRI with real-time applicability of EEG.

3. 强化学习与序列决策 | 17 篇

11. Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs

性能、效率与崩溃——代码大语言模型离线后训练的优势与挑战

AI 总结:本研究探讨代码LLM的离线RL后训练,利用现有数据集避免在线采样,数小时内显著提升零样本代码生成性能,并在0.5B至7B参数模型上普遍有效。

链接:https://arxiv.org/abs/2609.11956

机构:TU Darmstadt(达姆施塔特工业大学); Hessian Center for Artificial Intelligence(黑森人工智能中心); National Research Center for Applied Cybersecurity ATHENE(国家应用网络安全研究中心 ATHENE)

作者:Abhinav Anand, Sanjana Reddy Pachika, Shweta Verma, Mira Mezini

英文摘要:Post-training with reinforcement learning (RL) is a critical phase in the development of code-generating large language models (LLMs), as it ensures adherence to instructions and the production of functionally correct code. This process typically requires computationally intensive code sample generation from Transformer-based LLMs and substantial GPU-CPU communication for sequence verification. To address these computational challenges, this work examines whether RL-based post-training can be performed entirely offline by leveraging existing datasets rather than generating new samples. The findings indicate that, with only a few hours of training, zero-shot code generation performance of LLMs can be substantially improved without online sampling. Additionally, offline RL produces performance gains across models ranging from 0.5B to 7B parameters, although the extent of improvement varies among model families.

12. Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning

认证安全策展:安全离线强化学习的无分布保证

AI 总结:针对仅能通过片段比较和偶发预算询问判断安全性的离线强化学习,提出认证安全策展流程,以无分布保证选择安全轨迹并克隆策略,在DSRL任务上多数满足预算且可预测拒绝。

链接:https://arxiv.org/abs/2609.12014

机构:Iowa State University(爱荷华州立大学)

作者:Adam Haroon, Cody Fleming

英文摘要: Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free $(\alpha, \delta)$ bound on the unsafe fraction of the selection, and behavior cloning follows. We are not aware of prior work certifying the composition of a training set for offline RL or imitation. Oracle controls justify the design: reweighting individual transitions fails even with an exact value, so the value selects whole trajectories. The policies satisfy the cost budget on eleven of fifteen DSRL tasks, one short of cloning the ground-truth safe subset, which needs a label on every trajectory; the uncertified variant reaches twelve. Retrained on the certified selection, the strongest full-label method becomes safe where no setting of its own cost target rescues it. Refusal is predictable: the certificate's probability has a closed form in the purity the pool attains, which the calibration sample estimates and the scorer enters only through.

13. Inverting Self-Triggered Control: Adversarial Reinforcement Learning for Sparse Denial-of-Service Attacks

反转自触发控制:用于稀疏拒绝服务攻击的对抗强化学习

AI 总结:本研究将自触发控制反转,提出对抗强化学习生成最稀疏的拒绝服务攻击调度,以李雅普诺夫递增条件破坏闭环稳定性,实验证明其在多种环境下以100%成功率优于基线。

链接:https://arxiv.org/abs/2609.12016

作者:Adam Haroon, Erick J. Rodríguez-Seda, Tristan Schuler, Cody Fleming

英文摘要:Self-triggered reinforcement learning control (RL-STC) learns the sparsest control schedule that preserves Lyapunov-decreasing stability under a Run-Time Assurance (RTA) override. We invert this: an adversarial RL agent learns the sparsest jamming or Denial-of-Service (DoS) schedule that destabilizes the closed loop, with a Lyapunov-increase admissibility predicate mirroring the defender's safety certificate. We prove a plant-property lower bound on the minimum jam count required for an immediate hold-last medium-access-control adversary to force a crash against a self-triggered controller (STC) satisfying a Lyapunov contract, and recover a certificate-level analog of the consecutive-grouping optimality of prior count-budget DoS scheduling as a corollary. This extends the DoS-scheduling count-budget analysis from periodic and linear-time-invariant to STC controllers. Empirically, we train against four fixed defenders per plant (one Linear Quadratic Regulator (LQR) and three RL-STC) on Pendulum, CartPole, and Quadrotor2D. The learned adversary is the only adversary that crashes every defender on every plant at $100\%$: greedy misses Quadrotor2D LQR on $42\%$ of episodes and periodic misses Pendulum LQR on $97\%$. On jam-time-per-failure it beats baselines by up to $2.8\times$, and shows its widest absolute margin on Quadrotor2D LQR. Robustness ablations show that Gaussian observation noise exceeding the initial-state magnitude and position-only observation both preserve $100\%$ failure rate and keep the learned adversary strictly ahead of both baselines on jam-time-per-failure.

14. Reinforcement Learning for Syndrome Extraction

用于综合征提取的强化学习

AI 总结:本文提出基于强化学习和重要性采样的方法,用于高效搜索低逻辑错误率的量子错误综合征提取方案,在多个规模上显著优于现有工具。

链接:https://arxiv.org/abs/2609.12020

机构:University of California, Los Angeles(加州大学洛杉矶分校)

作者:John Zhuoyang Ye, Aarav Pabla, Jens Palsberg

英文摘要:A key subtask of quantum error correction is to extract a syndrome that, if nontrivial, signals an error. The number of possible ways to extract a syndrome grows exponentially with the syndrome size, and these implementations vary greatly in fault tolerance, as measured by their logical error rates. This creates a natural search problem: find an implementation with a low logical error rate. Previous work solves this problem but sacrifices either solution quality or scalability. In this paper, we use reinforcement learning and importance sampling to outperform previous work at all scales. Compared with the state of the art automatic scheduling tools AlphaSyndrome and PropHunt, our tool reduces the logical error rate by 25.9\% and 71.7\% on average, respectively, culminating with a reduction of 97.8\% for a surface code with distance 15.

15. Adaptive Chemotherapy Control under Tumor Heterogeneity via Reinforcement Learning

肿瘤异质性下的自适应化疗控制:基于强化学习的方法

AI 总结:针对肿瘤异质性和耐药性,本研究比较了基于强化学习(TD3与DQN)的闭环化疗给药策略与开环最优控制基准,在100名虚拟患者队列中验证了其有效性,并揭示了疗效与给药一致性之间的权衡。

链接:https://arxiv.org/abs/2609.12264

机构:University of Texas at Arlington(德克萨斯大学阿灵顿分校)

作者:Bereket Sitotaw Kidane, Md Samiul Haque Motayed, Shuo Wang

英文摘要:Designing effective chemotherapy regimens is hindered by tumor heterogeneity and drug resistance, which complicate the deployment of patient-specific model-based optimal control across diverse populations. We develop and compare closed-loop deep reinforcement learning (DRL) dosing policies with continuous (TD3) and discrete (DQN) action spaces trained on a high-dimensional heterogeneous tumor model. The DRL policies are benchmarked against a Pontryagin's Maximum Principle (PMP)-derived open-loop benchmark. We assess generalization under parametric heterogeneity using a 100-patient virtual cohort with plus or minus 10 percent uniform perturbations in growth and drug-sensitivity parameters. Across this cohort, TD3 achieves higher average tumor reduction, while DQN yields tighter inter-patient dosing consistency, revealing a clear efficacy-consistency trade-off in this study. Our simulations assume full observation of all tumor subpopulations; translation to sparse and noisy clinical measurements will require partial-observability formulations and/or state estimation. Overall, the results show that simulation-trained DRL can learn state-dependent feedback dosing policies that complement open-loop optimal control benchmarks.

16. Reinforcement Learning over Patient Trajectories for Clinical Reasoning in EHR Foundation Models

基于患者轨迹的强化学习用于EHR基础模型中的临床推理

AI 总结:针对EHR基础模型临床推理受限的问题,提出强化学习微调框架,通过时间感知奖励优化轨迹生成,使小模型在数据有限时超越大模型并实现跨任务正向迁移。

链接:https://arxiv.org/abs/2609.12277

机构:MIT(麻省理工学院); Microsoft Research(微软研究院)

作者:Yuxin Xiao, Sheng Zhang, Chandan Singh, Tristan Naumann, Hoifung Poon, Jianfeng Gao, Xiaodong Liu

英文摘要:Electronic health record (EHR) foundation models trained on longitudinal patient trajectories have demonstrated strong performance across diverse clinical prediction tasks. However, their clinical reasoning capabilities remain constrained by next-token prediction on limited and incomplete EHR data. To address this, we propose a reinforcement learning (RL) fine-tuning framework that treats EHR foundation models as generative policies over patient trajectories. We formulate common clinical prediction problems (e.g., hospital readmission) as event-conditioned, time-windowed reasoning tasks. We then design time-aware, rollout-sensitive rewards to account for finite rollout lengths and temporally inconclusive outcomes. We find that RL fine-tuning consistently improves over pre-trained backbones and strong baselines. Notably, it enables smaller models to surpass larger pre-trained models in data-limited regimes and induces positive transfer across tasks. Further analysis shows that RL fine-tuned models generate trajectories with stronger structural and semantic alignment to ground truth and greater downstream utility.

17. Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

通过决策流采样:大语言模型中无需训练的潜在推理路径提取

AI 总结:提出无需训练的决策流采样(DF-Sample)框架,通过构建层次化推理树并执行全局轨迹评估,解锁基座模型中潜在的高质量推理路径,在GPQA上以45.6%准确率超越GRPO等基线。

链接:https://arxiv.org/abs/2609.12317

机构:Stevens Institute of Technology(史蒂文斯理工学院)

作者:Zhendong Mi, Shaoyi Huang

英文摘要:A central question in LLM reasoning is whether reinforcement learning (RL) instills genuinely new capabilities or merely reshapes how existing knowledge is expressed during inference. Building on the distribution-sharpening hypothesis, which holds that RL reallocates probability mass toward high-reward trajectories already latent in base models, we ask: can we unlock those latent paths without costly RL fine-tuning? We present Decision-Flow Sampling (DF-Sample), a training-free, data-free inference-time framework that constructs a hierarchical reasoning tree, scores terminal nodes for quality, and back-propagates utilities to inform each intermediate branching decision. Unlike conventional sampling strategies that make purely local step-wise choices, DF-Sample performs explicit global trajectory evaluation before committing to a path, recovering high-quality but low-probability reasoning chains that standard decoding overlooks. On GPQA, DF-Sample achieves 45.6% accuracy, surpassing power sampling (38.9%) and GRPO (39.9%), showing that a training-free method can outperform a trained one. Across three models and four benchmarks, DF-Sample consistently outperforms baselines, indicating substantial latent reasoning potential in pretrained base models.

18. MInTRL: Off-policy Intervention can boost On-policy RL

MInTRL:离策略干预可以提升在策略强化学习

AI 总结:本文提出最小干预强化学习(MInTRL),通过在在策略轨迹生成中引入稀疏局部干预,以序列级优势回归目标训练,在不牺牲可学习性下扩展探索,在数学和代码基准上优于在策略和离策略基线。

链接:https://arxiv.org/abs/2609.12419

机构:Amazon Web Services(亚马逊云服务); Boston University(波士顿大学)

作者:Mingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong

英文摘要:Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the base model's capabilities, but may suffer from large distribution shift. The key challenge is thus to expand exploration without sacrificing learnability. In this work, we introduce Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts. During generation, a judge-intervention policy periodically reviews the current policy's output, replaces erroneous suffixes with short corrections, and immediately returns control to the policy. During training, MInTRL adopts a sequence-level advantage-regression objective that eliminates the need for importance sampling. We show that sparse, local interventions can substantially improve coverage beyond finite-budget on-policy sampling while preserving the overall on-policy nature of the resulting trajectories. Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines. Ablations show that MInTRL remains effective with self-intervention and across different judge policies, while performance peaks at moderate intervention intensity, highlighting the importance of intervening minimally. These results establish minimal intervention as an effective paradigm for enhancing on-policy RL.

19. Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning

长程LLM智能体强化学习中的粒度自适应信用分配

AI 总结:提出GACA,一种基于不确定性关键性代理的自适应粒度信用分配方法,在长程LLM智能体强化学习中优于GRPO和GiGPO。

链接:https://arxiv.org/abs/2609.12424

机构:Nankai University(南开大学); Peking University(北京大学); Shanghai Waybot Technology Co., Ltd.(上海微步科技有限公司); Wuhan University(武汉大学)

作者:Taoran Liang, Yang Liu, Shang Luo, Yingguang Yang, Rongrong Zhang, Yingzong Min, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong

英文摘要: Reinforcement learning is now the standard way to train large language model agents on long-horizon tasks, where dozens of interdependent actions precede a single sparse reward. Critic-free, group-relative methods such as GRPO suit this regime, but they broadcast one trajectory-level scalar to every step and cannot say which decision drove the outcome. GiGPO recovers a step-level signal by grouping time steps that share an anchor state, yet it merges the step- and episode-level estimates under one fixed weight, spending the same resolution on a pivotal branching decision as on a routine, near-deterministic transition. We argue that the right resolution is state-dependent, and propose GACA, a critic-free estimator whose granularity follows an uncertainty-based criticality proxy. GACA scores every step by the negative log-likelihood its own rollout already records, then blends the two advantages with a per-step weight that grows with that score, so the gradient places more weight on the fine-grained signal at above-average NLL and on the episode-level signal below it. We derive an exact risk decomposition for the implemented mixture and show that sufficiently small modulation improves on fixed mixing under positive directional alignment. A separate conditional result bounds local action-value variation using expected NLL, while an error-projection analysis characterizes when mixing adds value beyond scalar uncertainty reweighting. On ALFWorld and WebShop, GACA improves task success over GRPO and GiGPO at both 1.5B and 7B scales.

20. Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

重新审视AI对齐的失真:RLHF是一个体面的功利主义对齐器

AI 总结:本文重新分析RLHF的失真问题,证明其指数退化源于分布不匹配而非算法缺陷,并给出紧界,表明无失配时RLHF可达最优失真。

链接:https://arxiv.org/abs/2609.12651

机构:University of California, Berkeley(加州大学伯克利分校); Center for Advanced Intelligence Project, RIKEN(RIKEN先进智能项目中心); The Institute of Statistical Mathematics and the Graduate University for Advanced Studies(统计数理研究所与综合研究大学院大学); Tohoku University(东北大学)

作者:Kazusato Oko, Annie Ulichney, Nika Haghtalab, Han Bao

英文摘要:While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by Gölz et al. (2025) demonstrated that the \textit{distortion} -- defined as the multiplicative gap between the average user utility of the RLHF policy and the optimal average utility -- can scale exponentially with the Bradley-Terry temperature parameter $\beta$ when users have heterogeneous preferences. In this work, we present a fine-grained analysis of the distortion of RLHF with reward clipping and demonstrate that such exponential degradation is not a fundamental property of the algorithm but rather a consequence of distribution mismatch between the distribution generating preference data ($\mu$) and the KL reference policy ($\pi_{\mathrm{ref}}$). To this end, we establish tight upper and lower bounds on the distortion of RLHF across multiple regimes of the KL regularization strength. We show that in a representative regime, under the Bradley-Terry model, the distortion is $\tilde{\Theta}(\beta B + \beta)$, where $B$ is an upper bound on the log density ratio between $\mu$ and $\pi_{\mathrm{ref}}$. In particular, when there is no distribution mismatch (i.e., $\mu = \pi_{\mathrm{ref}}$), RLHF achieves the optimal distortion of $O(\beta)$ up to a constant. Our results suggest that, to reasonably maximize average utility with RLHF, it is preferable to use on-policy sampled preference data or to fine-tune before RLHF on data from a source close to $\mu$.

21. Curriculum-Based Adversarial Heterogeneous Agent Reinforcement Learning for Autonomous Quad-Copter Landing in Maritime Settings

基于课程学习的对抗异构智能体强化学习用于海上环境自主四旋翼无人机降落

AI 总结:针对海上四旋翼无人机回收难题,提出基于课程学习和对抗风智能体的HAPPO强化学习方法,在分布外海况下显著提升成功率并降低坠毁率。

链接:https://arxiv.org/abs/2609.12758

机构:Maastricht University(马斯特里赫特大学)

作者:Allan Minh-Tam Nguyen, Sree Showrya Kotala, Stefan Banioi-Crijman, Kurt Driessens, Rico Möckel

英文摘要:Recovering unmanned aerial vehicles (UAVs) in maritime environments is challenging due to wind turbulence and ship-deck motion, making it a valuable test case for alternative control and learning approaches as conventional landing approaches often become unreliable. We study simulated mid-air capture of quadrotor UAVs by a ship-mounted robotic arm, learning robust cooperative control policies with Heterogeneous-Agent Proximal Policy Optimization (HAPPO) Reinforcement Learning. We train with HAPPO using a curriculum and an adversarial wind agent (HARL-AC) in NVIDIA Isaac Lab, and compare the obtained control policies against those generated through curriculum-based domain randomization and a benchmark trained on a single sea state. In-distribution evaluation on sea states $0/4/5$ shows comparable success for HARL-AC and domain randomization of up to $97.5\%$. On out-of-distribution sea states $7/8/10$, HARL-AC generalizes better, achieving up to $16\%$ higher median success rate at sea state 10, and substantially lower crash rates of up to $14\%$ compared to the domain randomization policy. Furthermore, we show that the adversarially trained policy shows more cautious behavior, slightly increasing timeouts by $<3\%$, but yields safer recovery behavior in severe, unseen conditions.

22. Offline Reinforcement Learning for Wind Farm Control: A Wind Tunnel Study under Dynamic Wind Directions

风电场控制的离线强化学习:动态风向下的风洞实验研究

AI 总结:针对动态风向下的风电场功率最大化问题,提出离线强化学习算法MTD3-BC,通过偏航控制减轻尾流效应,风洞实验表明其比贪婪策略提升约10%功率,且无需尾流模型、训练成本低,为首次实验验证。

链接:https://arxiv.org/abs/2609.12905

机构:University of Warwick(华威大学); Technical University of Munich(慕尼黑工业大学)

作者:Yuhan Su, Hongyang Dong, Simone Tamaro, Filippo Campagnolo, Carlo L. Bottasso, Xiaowei Zhao

英文摘要: This paper addresses the wind farm power maximization problem in the presence of wind direction changes. Specifically, a model-free Modified Twin Delayed Deep Deterministic Policy Gradient with Behavior Cloning (MTD3-BC) algorithm is proposed to tackle this task through yaw control under varying wind direction conditions. MTD3-BC is an offline reinforcement learning (RL) algorithm that aims to infer good behavior from only a precollected offline dataset. Additionally, to ensure smooth and moderate yaw adjustments, a new action consistency term is introduced into the policy optimization objective. Unlike online RL methods, MTD3-BC does not require extensive interactions with a wind farm simulator during training, significantly reducing computational costs and training time. A wind tunnel experiment is conducted to validate the effectiveness of the algorithm under varying wind directions. The results demonstrate that MTD3-BC successfully mitigates wake effects, delivering farm-level power gains of approximately 10\% over the baseline greedy strategy and performance on par with a data-calibrated model-based wake-steering benchmark, while requiring no wake model and only a small fraction of the training cost of online RL. To our knowledge, this is the first time an offline RL wind farm control policy has been validated and demonstrated experimentally.

23. Groupoid-Based Internal State Representations for Reinforcement Learning with Local Symmetries

基于群胚的内部状态表示用于具有局部对称性的强化学习

AI 总结:针对现有强化学习难以利用局部对称性的问题,提出基于群胚的框架,动态发现等价结构,在缩减空间学习,提升样本效率和收敛性,优于标准Q学习。

链接:https://arxiv.org/abs/2609.13035

机构:City St George’s, University of London(伦敦城市圣乔治大学)

作者:Ben Opperman, Eduardo Alonso, Esther Mondragón

英文摘要:Symmetries play a central role in reducing the complexity of reinforcement learning problems, yet most existing approaches rely on fixed group actions or predefined state abstractions. Classical reinforcement learning algorithms typically assume a globally structured Markov decision process with uniformly applicable actions and transitions, an assumption that limits their ability to exploit modularity and local, context-dependent regularities present in many realistic environments. We propose a reinforcement learning framework using groupoids to capture local, state-dependent symmetries and support the dy- namic discovery of equivalence structures during interaction. The agent maintains orbit representatives together with transporters that map raw states to canonical forms, enabling learning and decision-making to be performed in a symmetry-reduced space while preserving local distinctions. Empirical results demonstrate that the proposed groupoid-based approach improves sample efficiency and convergence in dense and large-scale environments exhibiting strong partial symmetries, yielding substantial performance gains over standard Q-learning. These findings show that dynamically exploiting local symmetry provides a practical and mathematically principled route to scalable and generalisable reinforcement learning.

24. Robust Policy Optimization via Adversarial Importance Sampling

通过对抗重要性采样的鲁棒策略优化

AI 总结:本文提出Advis方法,通过对抗重要性采样优化鲁棒策略,并引入advrl库和更全面的评估,在连续控制任务中验证了有效性。

链接:https://arxiv.org/abs/2609.13044

机构:Mohammed VI Polytechnic University(穆罕默德六世理工大学); Khalifa University(哈利法大学); Concordia University(康考迪亚大学)

作者:Amine Andam, Jamal Bentahar, Mustapha Hedabou

英文摘要:Significant progress has been made in safeguarding deep reinforcement learning (DRL) policies against input perturbations. Developing robust DRL involves three main stages: algorithm design, implementation, and evaluation. In this work, we identify and address a key limitation at each stage. First, we introduce Adversarial Importance Sampling (Advis), a method that uses importance sampling over trajectories from standard training to estimate and optimize verifiable worst-case returns. Advis satisfies three desirable criteria not jointly achieved by prior work: it requires no additional environment interactions, no auxiliary networks, and captures long-term robustness. Second, we introduce advrl, a modular PyTorch library that provides clean, single-file implementations of existing robustness methods and adversarial attacks, facilitating rapid prototyping and enabling reproducible and traceable evaluations. Third, we revisit evaluation under learned adversaries and show that optimal adversarial hyperparameters do not transfer across agents, which can lead to an overestimation of robustness when using a limited set of attacker configurations. Accordingly, we evaluate policies against a large and diverse set of attackers, using 6-14x more configurations than prior work. Finally, we evaluate our approach on continuous control environments, demonstrating its effectiveness relative to existing baselines. The code is available at: this https URL

25. MCRL2: Multi-resource Cross-attention-based Representation Learning-augmented Reinforcement Learning for Cloud Microservice Scheduling

MCRL2:基于多资源交叉注意力的表示学习增强强化学习用于云微服务调度

AI 总结:针对云微服务调度中资源动态失衡与异构耦合问题,提出MCRL2,利用多资源交叉注意力表示学习增强强化学习,在真实集群轨迹上显著提升负载均衡、调度成功率并降低平均完成时间。

链接:https://arxiv.org/abs/2609.13048

机构:Wuhan University(武汉大学); Beijing University of Posts and Telecommunications(北京邮电大学)

作者:Tiangang Li, Shi Ying, Xiangbo Tian, Chuan Shi, Ding Xiao

英文摘要:Efficient microservice scheduling is crucial for maintaining load balance across nodes in data centers and ensuring high quality of service. However, achieving this in practice remains challenging due to dynamic resource imbalance under fluctuating workloads, nonlinear coupling across multiple resource dimensions, and the heterogeneity of microservice resource demands. While reinforcement learning-based approaches have shown promise, they struggle to capture the complex interdependencies among heterogeneous resources and neglect the importance of learning informative system representations. To address these limitations, we propose MCRL2, a novel reinforcement learning approach augmented with multi-resource cross-attention-based representation learning for microservice scheduling. Specifically, we first propose MCRL, a novel representation learning approach that captures structured and informative interactions among nodes, resources, and microservices via a multi-resource cross-attention mechanism. Then, MCRL2 augments reinforcement learning through MCRL-enhanced actor-critic architecture combined with a maximum entropy objective, improving system state expressiveness and leading to more stable and effective scheduling decisions. Extensive experiments on real production cluster traces demonstrate that MCRL2 significantly outperforms existing baselines in load balancing, scheduling success rate and average completion time across diverse workload patterns.

26. A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning

基于正则化的鲁棒强化学习:一个统一且受约束的视角

AI 总结:本文统一了基于正则化的鲁棒强化学习方法,推导出性能差距上界,并将鲁棒训练转化为约束优化问题,通过联合更新拉格朗日乘子自动调整正则化权重,在连续控制任务上验证了理论分析。

链接:https://arxiv.org/abs/2609.13050

机构:Mohammed VI Polytechnic University(穆罕默德六世理工大学); Khalifa University(哈利法大学); Concordia University(康考迪亚大学)

作者:Amine Andam, Jamal Bentahar, Mustapha Hedabou

英文摘要:Regularization-based methods have become a standard approach for training Deep Reinforcement Learning policies against adversarial input perturbations. In this paper, we unify these methods by deriving new upper bounds on the performance gap between the nominal and worst-case policies. Each upper bound is expressed as an existing regularization objective plus a KL-divergence penalty between the nominal and worst-case policies, which further explains why adding a KL penalty improves robustness in practice. Building on these bounds, we formulate robust training as a constrained optimization problem, showing that existing methods correspond to the special case of a fixed Lagrange multiplier. We instead update the multiplier jointly with the policy to automatically tune the regularization weight. Finally, we conduct extensive adversarial evaluations across several continuous control tasks to validate our theoretical analysis.

27. CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models

CanvasAnneal:用于扩散语言模型的课程强化学习

AI 总结:CanvasAnneal通过课程式教师引导注入与逐步退火,缓解扩散语言模型强化学习的探索瓶颈,在数学推理和工具使用基准上优于标准方法并加速收敛。

链接:https://arxiv.org/abs/2609.13060

机构:Google DeepMind(谷歌DeepMind)

作者:Blake Olson, Yuhang Song, Emmett McQuinn, Yuan Shangguan

英文摘要:Diffusion Language Models (DLMs) offer promising parallel generation capabilities but lag behind autoregressive models in complex reasoning and tool-use tasks. While Reinforcement Learning (RL) has recently been applied to enhance DLMs, standard RL approaches suffer from an exploration bottleneck. To address this, we inject reasoning priors from a stronger teacher model to guide RL exploration. In this paper, we introduce CanvasAnneal, a curriculum-guided diffusion RL framework. During the initial RL phase, we warm-start exploration by injecting teacher-generated reasoning traces into the initial diffusion canvas. As training progresses, we gradually remove this guidance and require the model to generate more of the reasoning trajectory independently. Across mathematical reasoning and tool-use benchmarks, CanvasAnneal improves over standard diffu-GRPO on MATH500, Countdown, and Tau2 and substantially accelerates reward improvement on several tasks, while gains are task-dependent. Our results suggest that structured training-time guidance can alleviate exploration bottlenecks in diffusion RL and speed up convergence on harder tasks.

4. 生成模型与概率建模 | 3 篇

28. Score-based Outlier Generation via Controlling the Radon-Nikodym Derivative

基于控制Radon-Nikodym导数的分数式异常值生成

AI 总结:本文提出一种基于控制Radon-Nikodym导数的分数式异常值生成方法,通过似然重加权修改扩散分数,无需重新训练即可生成低似然样本,并保持数据几何一致性。

链接:https://arxiv.org/abs/2609.12113

机构:University of Waterloo(滑铁卢大学); Royal Bank of Canada(加拿大皇家银行)

作者:Amartya Mukherjee, Tristan Milne, Kry Yik-Chau Lui, Stephanie Hazlewood, Jun Liu

英文摘要:Outliers are important for stress-testing algorithms and understanding system behaviour under rare conditions. Despite being commonly described as low-likelihood events, existing generative approaches rarely control likelihood explicitly. In this work, we introduce a measure-theoretic notion of outliers based on the distribution of log-likelihood values, which is guaranteed to assign higher probability mass to low-likelihood events with a specifiable magnitude. Building on this formulation, we derive how likelihood reweighting modifies the diffusion score and use this relation to motivate a controlled modification of the reverse-time dynamics. In particular, likelihood reweighting implies a scaling of the score function with a control term derived from the Radon-Nikodym derivative of the likelihood distributions. Correspondingly, the updated score function can be obtained with no retraining of the diffusion model. We exploit the Ornstein-Uhlenbeck semigroup underlying diffusion models to motivate an exponentially interpolated controller which approximates the true control. Experiments demonstrate controlled generation of low-likelihood samples while remaining consistent with the data geometry.

29. Certifying Concept Unlearning in Text-to-Image Diffusion Models

认证文本到图像扩散模型中的概念遗忘

AI 总结:针对T2I扩散模型概念遗忘评估依赖攻击成功率而忽视残余泄漏的问题,提出结合统计认证与最坏情况分析的认证框架,在三大概念类别和六种方法上验证,泄漏界超过攻击成功率16.2%,证明认证是可靠审计的必要补充。

链接:https://arxiv.org/abs/2609.12163

机构:Imperial College London(伦敦帝国理工学院); TU Wien(维也纳工业大学)

作者:Mansi, Luca Marzari, Francesco Leofante

英文摘要: Existing evaluations of concept unlearning in text-to-image (T2I) diffusion models primarily rely on attack success rates obtained through automated adversarial prompt search. However, these metrics provide only empirical evidence over a finite set of queries and leave residual leakage over the broader prompt space largely unquantified. This limitation can lead to overestimating unlearning effectiveness and underestimating safety risks. To address this gap, we introduce a novel certification framework for T2I concept unlearning that provides high-confidence guarantees with bounded error on residual concept leakage. Our approach combines statistical certification with worst-case analysis along concept-relevant embedding directions to derive explicit upper bounds on leakage probability under user-specified confidence levels. We evaluate our framework across three major concept categories namely NSFW content, artistic styles, and celebrity identities, and six state-of-the-art unlearning methods. Certified leakage bounds consistently exceed standard attack success rates by 16.2%, uncovering substantial residual risks missed by existing evaluation protocols. Crucially, our results demonstrate that empirical attack-based evaluations can significantly underestimate residual leakage and establish certification as a necessary complement for reliable auditing of concept unlearning in T2I diffusion models.

30. Theoretical Guarantees for One-Shot Magnitude Pruning and Compute-Adaptive Early Exit

一次性幅度剪枝与计算自适应早退的理论保证

AI 总结:本文从统一视角研究神经网络计算缩减,证明一次性幅度剪枝的集中定理,并引入条件感知器早退,其泛化误差随计算差距幂次衰减,扩展至深度网络并验证标度律。

链接:https://arxiv.org/abs/2609.12337

机构:University of Illinois Chicago(伊利诺伊大学芝加哥分校)

作者:Erdem Koyuncu

英文摘要:We study compute reduction in neural networks through a unified partial versus full computation view, captured by one-shot magnitude pruning in the static regime and early exit in the adaptive regime. In an asymptotic single-neuron model, we prove a concentration theorem for one-shot magnitude pruning with explicit rates. We also introduce the conditional perceptron for early exit and show that its excess generalization error decays as a power of the compute gap, with an exponent that grows to infinity as the alignment between partial and full computations tends to one. We then extend the analysis to deep networks, characterizing how pruning-induced distortions accumulate with depth and deriving a corresponding compute-accuracy tradeoff for frozen-backbone early exit under a neural network Gaussian process model. Numerical simulations corroborate the predicted scaling laws.

5. 优化、泛化与理论分析 | 8 篇

31. Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise

带裁剪和加性噪声的随机梯度方法的几乎必然收敛分析

AI 总结:本文证明带裁剪和加性高斯噪声的随机梯度下降在光滑性和噪声有界假设下几乎必然收敛,并推广至动量变体,为裁剪方法的路径行为提供理论基础。

链接:https://arxiv.org/abs/2609.12119

机构:University of Waterloo(滑铁卢大学)

作者:Amartya Mukherjee, Jun Liu

英文摘要:Stochastic gradient descent (SGD) with gradient clipping and additive noise has become a standard technique for training machine learning models, particularly in applications requiring robustness or privacy guarantees. However, clipping introduces a bias in stochastic gradients, while additive noise introduces additional variance, making the long-run behaviour of individual optimization trajectories difficult to characterize. In this work, we prove that SGD with clipping and additive Gaussian noise (SGD-CN) converges almost surely (a.s.) under smoothness and uniformly bounded stochastic-gradient noise assumptions, provided the step sizes satisfy some standard decaying conditions. Our analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, where we show that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods and suggest that, despite the bias and noise introduced by clipping and perturbation, the algorithm remains stable in both convex and nonconvex regimes.

32. Estimating Pedestrian Volumes from GIS-Derived Built-Environment Features: A Machine Learning Framework

从GIS衍生的建成环境特征估计行人流量:一种机器学习框架

AI 总结:本研究提出一种机器学习框架,利用GIS衍生的建成环境特征预测交叉口行人流量,通过特征选择和梯度提升显著降低预测误差,优于传统GLM基线。

链接:https://arxiv.org/abs/2609.12173

机构:Portland State University(波特兰州立大学)

作者:Bahareh Golchin, Banafsheh Rekabdar, Sirisha Kothuri, Joseph Broach

英文摘要:Transportation agencies need pedestrian volume estimates across entire road networks to prioritize safety investments, yet manual counts are expensive and cover only a small share of intersections. We present a machine learning pipeline that predicts 2-hour PM peak pedestrian volume at 101 urban intersections in Portland, Oregon, from built-environment, land-use, and street-network features drawn from open GIS data. Starting from the Negative Binomial GLM used in practice, we add feature selection, count-aware gradient boosting, and repeated cross-validation, selecting one configuration by a combined rank over RMSE, MAPE, and SMAPE across four cross-validation strategies. The winner, a histogram-based gradient boosting model with Poisson loss and L1 Lasso feature selection, reduces cross-validated RMSE by 12% over the GLM baseline (89.8 to 78.7) and holdout RMSE by 19% (108.0 to 87.9). Code is released on GitHub.

33. The Rank the Task Demands: A Causal Rank Law for Matrix Memories Trained on Group Composition

任务所需的秩:群组合训练矩阵记忆的因果秩定律

AI 总结:本文通过群组合测试平台,证明梯度下降在矩阵记忆中招募的秩等于任务所需的最小忠实表示维数,并验证了该秩定律的因果性及可解/不可解群的等价性。

链接:https://arxiv.org/abs/2609.12259

机构:Pebble ML

作者:Samuel Larson

英文摘要:Matrix-valued memories make rank the natural budget of a learned representation: the number of independent directions a state spans bounds what it can bind, compose, and track. We report causal evidence, on a group-composition testbed trained under a hard single-state bottleneck with a fixed decoder that cannot launder rank, that gradient descent recruits precisely the rank the task's algebra demands. A companion paper [Larson, 2026a] establishes the analogous recruitment and causal necessity pattern on a $K$-pair associative-binding testbed, where exact recovery provably requires state rank at least $K$; this paper inherits that instrument and extends the rank law from a scalar capacity bound to a representation-theoretic one. We train toward chosen minimal faithful reference representations embedded in larger matrices. On group-composition state tracking over five finite groups spanning the solvable/non-solvable divide, the recruited rank equals the group's minimal faithful real representation dimension $d_{\min}$ (Spearman $\rho = 0.9747$, the design's tie-capped maximum), the dimension-matched solvable/non-solvable pair $S_4$/$A_5$ is statistically equivalent under a pre-registered test, and a pre-registered force-rank test separates a guaranteed similarity ceiling from empirical recovery at the target dimension: one rank below $d_{\min}$, cosine similarity is capped by the target's tied unit spectrum at $\sqrt{(d_{\min}{-}1)/d_{\min}} \le 0.894$, below the $0.9$ threshold in every group by construction, with observed cells at 86-95% (mean 91%) of that ceiling; at $d_{\min}$, not guaranteed a priori, recovery clears the pre-registered anchor-relative bar at four seeds per group in all five groups. Within this testbed, measured effective rank tracks representation dimension; the matched-dimension $S_4$/$A_5$ comparison establishes equivalence within the pre-registered tolerance.

34. Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and Hölder Smoothness

重尾噪声与Hölder光滑条件下随机梯度方法的收敛性

AI 总结:本文在Hölder光滑和重尾噪声下,证明标准SGD、δ-GClip和G-Clip的收敛速率,其中G-Clip在极重尾区域首次获得收敛保证。

链接:https://arxiv.org/abs/2609.12785

机构:Indian Institute of Science Education and Research Kolkata(印度科学教育与研究学院加尔各答分校); The University of Manchester(曼彻斯特大学)

作者:Misbah Uz Zaman, Anirbit Mukherjee

英文摘要:Classical convergence guarantees for stochastic gradient methods typically assume Lipschitz-smooth objectives and finite-variance gradient noise, both frequently violated in practice. In contrast, we study nonconvex stochastic optimization under the joint relaxation of these assumptions: objectives with $(L,s)$-Hölder continuous gradients, $s\in(0,1]$, and gradient noise satisfying only a bounded $\alpha$-th moment condition for $\alpha\in(1,2]$. We establish three convergence results. Firstly, that standard SGD converges at rate $O(T^{-s/(1+s)})$ whenever $\alpha\ge1+s$, extending the classical nonconvex SGD rate to heavy-tailed noise and Hölder smoothness simultaneously. Secondly, we analyze $\delta$-regularized gradient clipping ($\delta$-GClip), a provable trainer of wide and deep nets, and establish a stationarity rate of $O(T^{-2s(\alpha-1)/[(1+s)(2\alpha-1)]})$ under the same condition. Thirdly, we analyze standard gradient clipping (G-Clip) and show that it recovers the above rate for $\alpha\ge1+s$ while in the very heavy-tailed regime $\alpha<1+s$, it has a convergence rate $O(T^{-2s(\alpha-1)/[(\alpha-1)+s(2\alpha-1)]})$ --- the first convergence guarantee in this regime for any stochastic gradient based method.

35. Quantifying the Value of Privileged Information Using a PAC-Bayesian Approach

量化特权信息价值:一种PAC-Bayesian方法

AI 总结:本文提出一种基于PAC-Bayes的算法无关信息论方法,通过比较有无特权信息时的最紧风险界来量化其潜在价值,并引入训练时指标,无需测试数据即可预测性能提升。

链接:https://arxiv.org/abs/2609.12891

机构:Leiden University(莱顿大学); Honda Research Institute Europe GmbH(本田欧洲研究院有限公司)

作者:Vasily Bokov (1 and 2 and 3), Sebastian Schmitt (3), Vedran Dunjko (1 and 2), Hao Wang (1 and 2) ((1) aQa, Leiden University, The Netherlands (2) LIACS, Leiden University, Leiden, The Netherlands (3) Honda Research Institute Europe GmbH, Offenbach, Germany)

英文摘要:In practice, various learning scenarios provide access to auxiliary features exclusively during training. Incorporating such data to enhance model performance gave rise to a paradigm known as Learning Using Privileged Information (LUPI). While this extra information is intended to improve the resulting model, establishing a generalized, cohesive understanding of how privileged information (PI) transfers useful knowledge remains a challenge. Vapnik's original theory and subsequent works offer performance guarantees in certain cases, but these results are inherently per-algorithm and rely on setting-specific proof approaches. Consequently, a more general framework explaining how and when PI transfers useful knowledge is still missing. To bridge this gap, we introduce an algorithm-agnostic, information-theoretic approach based on the PAC-Bayes framework. Rather than asking whether a particular algorithm exploits PI, we ask how much value it could offer: comparing the tightest achievable risk bound with and without PI yields its potential - an upper limit on the extractable gain. We introduce a metric that quantifies this potential directly from empirical training risk, bypassing the need for test-time data access, and validate our findings in both supervised and unsupervised settings. The results demonstrate a robust correspondence between our training-time metric and true test-time performance gains. Ultimately, this work takes a necessary step toward an information-theoretic understanding of LUPI, and quantifying the potential of privileged features before committing to a model.

36. Dual-guided Hierarchical Edge Localization for Large-scale Optimal Transport Across Dimensions

跨维度大规模最优传输的双引导分层边缘定位

AI 总结:提出HELLO分层求解器,利用对偶势引导边缘定位,高效求解大规模离散最优传输,在百万点规模上显著提升速度与精度,并支持多种OT变体。

链接:https://arxiv.org/abs/2609.13010

机构:Shanghai Jiao Tong University(上海交通大学); Institute of Natural Sciences, Shanghai Jiao Tong University(上海交通大学自然科学研究院)

作者:Wenzhou Xia, Qiaoqiao Ding, Jingwei Liang, Xiaoqun Zhang

英文摘要:Optimal transport (OT) compares distributions and aligns datasets in machine learning, yet unregularized discrete OT requires a linear program with quadratically many transport variables. We propose HELLO, a hierarchical solver that casts large-scale discrete OT as edge localization and uses dual potentials to guide both coarse-to-fine initialization and within-level refinement. Initialization propagates coarse dual potentials across a recursive subsampling hierarchy to assign candidate edges. Refinement then iteratively inserts the largest dual violators in each row and column until the relative KKT residual meets a prescribed tolerance, while budgeted pruning ensures linear memory complexity. For exact-arithmetic refinement, we prove finite termination at a global optimum under a symbolic lexicographic rule. At the million-point scale, HELLO attains lower transport objectives with order-of-magnitude runtime improvements over strong baselines across feature dimensions from single digits to thousands. It further scales to 1.28 million samples per marginal in 8192 dimensions on a single H100, using 41.6 GiB peak GPU memory while satisfying a full relative KKT residual below $10^{-6}$. Beyond standard discrete OT, the framework supports general pairwise costs and serves as a scalable balanced-OT oracle for semi-discrete OT, Gromov--Wasserstein, unbalanced OT, and OT-based Flow Matching.

37. Quantile-based Loss Filtering for Outlier-Robust Stochastic Gradient Descent

基于分位数的损失过滤用于抗离群值随机梯度下降

AI 总结:本研究提出分位数-k损失SGD(QkL-SGD)框架,通过采样损失并选择较低分位数索引更新,在凸假设下证明线性收敛,实验显示中间分位数优于标准SGD和min-k-loss,兼具鲁棒性与信息量。

链接:https://arxiv.org/abs/2609.13040

机构:Harvey Mudd College(哈维穆德学院); University of California, Irvine(加州大学尔湾分校); University of Oxford(牛津大学)

作者:Jamie Haddock, Anna Ma, Elizaveta Rebrova

英文摘要:We study loss-based filtering for finite-sum optimization with a subset of corrupted component functions whose gradients may be highly unreliable. Motivated by minimum-loss-based SGD (min-$k$-loss) and quantile-based methods for corrupted linear systems, we propose and analyze a general loss-filtering framework -- Quantile-\(k\)-Loss SGD (Q\(k\)L-SGD) -- that samples \(k\) component losses at each iteration and updates using an index chosen uniformly from the lower empirical \(q\)-quantile. We prove linear convergence of this family of methods under standard convexity assumptions, requiring the sample size to scale with the number of corruptions and a subset strong-convexity threshold. For the cases when large enough sampling is impossible or undesirable, we give a complementary small-sample probabilistic analysis that covers any sample size $k$ and the convergence behavior depends on the probability of selecting an outlier and on the curvature of the selected good step. Experiments on polynomial regression, regularized logistic regression, and regularized hinge loss show that intermediate quantiles often outperform both standard SGD and min-\(k\)-loss SGD. In particular, min-\(k\) often stalls by repeatedly selecting nearly solved components, while intermediate quantiles retain robustness and produce more informative updates.

38. Benign Loss Landscapes Can Coexist with Worst-Case Hardness

良性损失景观可以与最坏情况下的困难共存

AI 总结:该研究证明树张量网络虽包含难以学习的目标,但其损失景观对可实现目标是良性的,困难源于高阶退化鞍点,为理解深度学习的泛化提供了新视角。

链接:https://arxiv.org/abs/2609.13057

机构:University of Melbourne(墨尔本大学); Iliad

作者:Zach Furman, Stephan Wäldchen, Yangda Bei, Liam Hodgkinson

英文摘要:Deep neural networks are expressive enough to contain worst-case targets that can be evaluated in polynomial time but cannot be learned in polynomial time by gradient descent. For practical tasks they nonetheless learn well, raising the question of what non-generic structure of real-world targets enables this. Existing surrogate models cannot pose this question because they either lack hard-to-learn targets entirely (deep linear networks) or cannot evaluate such targets efficiently (kernel methods, infinite-width limits). We study tree tensor networks (TTNs), a model class that generalizes deep linear networks and Tucker decompositions. We show they embed arbitrary read-once Boolean formulas, and thus contain polynomial-size targets that cannot be learned by gradient descent in polynomial time under the same mechanism as neural networks. Despite this, we prove that their loss landscapes are conditionally benign for every realizable target: every local minimum that is minimum-norm is global. Thus, surprisingly, bad local minima are not what distinguishes between typical and worst-case problems in TTNs. Instead, learning difficulty in TTNs can arise from high-order degenerate saddle points, which we show are caused by rank-deficiency. This is explored through a case study of the parity function, illustrating the potential for TTNs to relate landscape geometry to computational hardness.

6. 联邦学习、隐私与安全 | 5 篇

39. Fed-Equilibrium Framework for Topological Pareto Control in Robust and Fair Clinical Federated Learning

Fed-Equilibrium框架:面向鲁棒且公平的临床联邦学习中的拓扑帕累托控制

AI 总结:提出Fed-Equilibrium框架,通过两阶段梯度控制实现拓扑均衡,解决临床联邦学习中的知识主导问题,在跨国模拟中兼顾鲁棒性与公平性,使少数节点达到深度收敛。

链接:https://arxiv.org/abs/2609.11937

作者:Ting Xu, Henry Leung

英文摘要: The deployment of Federated Learning (FL) in multi-center clinical networks faces the challenge of "knowledge dominance," where high-volume hubs naturally overwhelm minority community nodes, implicitly treating the distinct clinical patterns of smaller cohorts as outliers. Existing geometric defenses provide a security baseline but leave this efficiency-fairness dilemma unresolved. To bridge this gap, we propose Fed-Equilibrium, a framework that advances the paradigm from simple defense to topological equilibrium. Unlike traditional aggregators, Fed-Equilibrium implements a sequential architectural synergy. It utilizes a two-stage gradient control cascade: Stage I (geometric quality assurance) enforces directional consistency via a cosine similarity funnel to filter malicious noise, creating a stabilized manifold; Stage II (topological Pareto control) then actively modulates verified contributions by identifying the optimal Pareto knee point. We validated this framework on a bi-national simulation integrating Canadian (CNODES) and U.S. (SyntheticMass) registries. Experimental results demonstrate that the system simultaneously secures the network against adversarial divergence while accommodating underrepresented signals. Notably, the minority U.S. spoke (representing less than 3% of data volume) achieved deep convergence comparable to the data-rich Canadian hub. This confirms that Fed-Equilibrium effectively counters "knowledge dominance," establishing a true "knowledge commons" where global generalizability does not come at the cost of local clinical representation.

40. Scalable Discrete-to-Continuous Channel Simulation for Compression and Privacy

面向压缩与隐私的可扩展离散到连续信道模拟

AI 总结:提出一种固定随机样本数、运行时间与信道无关的离散到连续信道模拟方案,通过潜在置换和指数竞赛实现压缩与隐私应用,可扩展到长块长度。

链接:https://arxiv.org/abs/2609.12067

机构:University of Toronto(多伦多大学)

作者:Joseph Rowan, Buu Phan, Ashish J. Khisti

英文摘要:Channel simulation has recently emerged as a useful component in machine learning systems where samples from a prescribed probability distribution are to be compressed. Yet, general channel simulation algorithms often suffer from high computational costs, random stopping times or, in the worst case, can require generating an infinite number of shared random samples. We introduce a scheme for both exact and approximate simulation of discrete-to-continuous channels which conversely uses a fixed number of random samples, and therefore has a runtime independent of the channel and the input. Unlike existing channel simulation schemes which generate a sequence of independent samples from a proposal distribution, our approach generates one sample, or alternatively a fixed number of samples, from each potential target distribution. We then apply a latent permutation to the samples before performing sample selection using an exponential race. Our scheme provides a flexible tradeoff between the number of generated samples and the compression rate. Using polar and multilevel coding, we scale our approach to handle long blocklengths in $O(n \log n)$ time in order to benefit from reduced per-symbol overhead. We conclude by demonstrating applications to variable-rate compression with stochastic VQ-VAEs and communication-efficient differentially private distributed mean estimation via exact simulation of the Gaussian mechanism.

41. A Differentially Private Federated Proximal Optimization Framework for Customer Churn Prediction in Heterogeneous Federated Telecom Networks

一种用于异构联邦电信网络客户流失预测的差分隐私联邦近端优化框架

AI 总结:针对电信异构数据下客户流失预测的隐私问题,提出DP-FedProx差分隐私联邦近端优化框架,在公开数据集上以略降精度实现与集中式相当的预测性能,并兼顾隐私保护。

链接:https://arxiv.org/abs/2609.12470

机构:Bangladesh University of Engineering and Technology(孟加拉国工程技术大学); University of New England(新英格兰大学); Ulster University(阿尔斯特大学)

作者:Joydeb Kumar Sana, Subrata Chakraborty, M M Manjurul Islam

英文摘要:Customer churn is one of the major issues in the telecommunication industry. To predict customer churn, conventional centralized machine learning approaches have been widely used. This centralized approach requires customer data to be stored in a central repository, which raises privacy concerns and may violate data protection regulations. Federated learning addresses this problem by allowing multiple telecom operators to collaboratively train a global model without transferring their raw customer data. However, real-world customer data are often heterogeneous (non-IID), which may negatively affect the performance of standard federated learning. Trained models can also suffer from privacy attacks. To address those issues, we propose a Differentially Private (DP) based Federated Proximal optimization (FedProx) framework. All experiments were performed on two publicly available telecom churn datasets. We trained Federated Averaging (FedAvg), DP-FedAvg, FedProx, and the proposed DP-FedProx framework. For baseline comparison, we also used several centralized and local models. To evaluate the models, we employed seven widely used evaluation metrics. The experimental results show that the FedProx based models consistently outperform the FedAvg based models. Compared with the best centralized model, the proposed DP-FedProx framework achieves competitive prediction performance with only a small reduction in accuracy while providing privacy guarantees. To explain our model, we conducted SHAP analysis which shows that DP-FedProx method priorities revenue group features. These results indicate that the proposed DP-FedProx framework provides a practical balance between prediction performance and data privacy protection.

42. Correlation-Guided Fast Machine Unlearning via Hessian Analysis

基于Hessian分析的相关性引导快速机器遗忘

AI 总结:针对安全系统中机器遗忘计算开销大的问题,提出基于Hessian分析的相关性引导遗忘框架,利用闭式更新规则实现82倍加速,并在多个数据集上保持模型效用与遗忘效果。

链接:https://arxiv.org/abs/2609.12620

机构:Indian Institute of Technology (Banaras Hindu University) Varanasi(印度理工学院(巴纳拉斯印度教大学)瓦拉纳西校区); Halmstad University(哈尔姆斯塔德大学)

作者:Ayushi Thakur, Ruchir Gupta, Amit Kumar Jaiswal, Prayag Tiwari

英文摘要: The increasing adoption of machine learning in network and distributed security systems has created an urgent need for mechanisms that can selectively and efficiently remove the influence of specific training data to eliminate compromised or adversarial data points from production models. Privacy regulations such as GDPR's \emph{right to be forgotten} also pose similar requirements. However, existing approximate unlearning techniques remain computationally prohibitive for deployment in real-world security systems, as they require repeated expensive Hessian-inverse-vector computations for each data point removal, creating a bottleneck when processing multiple related requests in scenarios such as intrusion detection systems, spam filters, and threat intelligence platforms. Thus, we introduce a computationally efficient unlearning framework that identifies correlated data points in the training set and applies a theoretically derived closed-form parameter update rule, achieving an $82\times$ wall-clock speedup over standard influence function unlearning while preserving model utility with a $10^{-2}$ improvement in accuracy over state-of-the-art baselines. Our method establishes theoretical guarantees and ensures numerical stability through Hessian damping. Our evaluation across seven diverse dataset architecture combinations, including large-scale CIFAR-100 with ResNet-50, demonstrates superior forgetting effectiveness, with membership inference attack success rates of 0.660 and tug-of-war scores of 0.950.

43. Hidden in Rounds: Predicting the Time Cost of 802.11 Contention in Federated Learning

隐藏在轮次中:预测联邦学习中802.11竞争的时间成本

AI 总结:本研究通过ns-3模拟和Bianchi模型估计,预测联邦学习在802.11无线网络中的通信时间成本,发现通信时间随客户端密度增加约两个数量级,而轮次数变化不大。

链接:https://arxiv.org/abs/2609.12903

作者:Satwat Bashir, Tasos Dagiuklas

英文摘要:Federated learning over IEEE~802.11 shares the wireless channel among clients that send model updates. We use ns-3 to measure the frame-delivery ratio and saturation throughput for different client densities and offered loads. A separate FedAvg trainer uses the frame-delivery ratio as a first-order proxy for the update-admission probability and uses an equation to estimate communication time. The method does not simulate the delivery of a complete model update or measure end-to-end training time. Across 720 evaluated runs with two datasets, two data partitions, six client densities, six offered loads, and five seeds, all runs reached their predefined target accuracy within the round budget. Rounds-to-target changed little with offered load, while communication time-to-target increased by about two orders of magnitude across the client-density range. A Bianchi-anchored estimator produced a mean absolute percentage error from $2.3\%$ to $10.2\%$ on held-out configurations. This error is measured against communication time constructed from the same round-duration equation, not against independently measured completion time. We also compare uniform participation with persistent heterogeneous participation. The study does not detect a statistically distinguishable excluded-class accuracy gap over five seeds, but the confidence intervals are wide. The results apply only to the evaluated configurations and do not provide a general convergence or fairness guarantee.

7. 鲁棒性、不确定性与可信学习 | 3 篇

44. Physics-Informed Conformal Prediction: Embedding PDE Consistency into Distribution-Free Uncertainty Quantification for Neural Operators

物理信息保形预测:将PDE一致性嵌入神经算子的无分布不确定性量化

AI 总结:针对神经算子求解PDE缺乏严格不确定性估计的问题,提出物理信息保形预测框架,将PDE残差嵌入非一致性得分,生成具有无分布覆盖保证且空间自适应的预测区间,并揭示FNO平移等变性的近似障碍,在六个物理场景中验证了89-91%的稳定覆盖率。

链接:https://arxiv.org/abs/2609.11935

作者:Michael Chin

英文摘要:Neural operators such as the Fourier Neural Operator (FNO) achieve remarkable accuracy in approximating solutions to partial differential equations (PDEs). However, providing rigorous uncertainty estimates remains an open challenge. We propose Physics-Informed Conformal Prediction (PI-CP), a framework that embeds PDE residuals into the nonconformity score of split conformal prediction, producing prediction intervals that are (i) distribution-free with provable coverage guarantees, and (ii) spatially adaptive when the PDE residual correlates with prediction error -- tighter where physics is well-satisfied, wider where it is violated. Additionally, we prove that FNO's translation equivariance creates a fundamental approximation barrier for PDEs with Dirichlet boundary conditions, and show that coordinate channels resolve this with up to 63x error reduction. We validate PI-CP across six physics scenarios -- heat conduction (2D/3D), structural mechanics (2D/3D), Darcy flow, and Navier-Stokes -- demonstrating consistent 89-91% coverage for all four Conformal methods, while MC Dropout and Deep Ensembles are unstable (82-100%). FNO outperforms CNN and DeepONet by 10-12x.

45. PLSP (Pre-hoc Liminal Space Profiling): OOD Prediction over Detection -- An Anticipatory Approach for Machine Learning Model Reliability

PLSP(事前极限空间剖面分析):超越检测的OOD预测——一种面向机器学习模型可靠性的预期性方法

AI 总结:本文提出PLSP框架,将OOD检测转向事前预测,引入数据集无关的CREDS指标、可信度曲线和热图,以提升模型在分布偏移下的鲁棒性。

链接:https://arxiv.org/abs/2609.12225

机构:University of Wisconsin Madison(威斯康星大学麦迪逊分校); IMC University of Applied Sciences(IMC应用科学大学); Indian Institute of Technology Roorkee(印度理工学院鲁尔基分校)

作者:Vipul Bansal, Himanshu Buckchash, Balasubramanian Raman, Deepak Dhungana

英文摘要:Out-of-Distribution (OOD) data poses a significant threat to machine learning models, often leading to model failure during deployment. All existing OOD detection methods are post-hoc, relying on evaluation metrics such as accuracy and AUC-ROC during inference to indirectly assess the model's response to OOD data by measuring deviations. In contrast to existing approaches, the proposed work shifts the paradigm from OOD detection to OOD prediction by proposing a pre-hoc anticipatory framework called PLSP for OOD prediction. We make several key contributions: (a) a dataset-independent metric called the CREDibility Score (CREDS) is proposed for OOD prediction; (b) credibility curves are introduced to study the maximum credibility a model can attain; and (c) credibility heat maps (and volume under surface) are introduced to characterize pre-hoc model behavior across different datasets. This work provides a novel perspective on signal processing under distributional shifts. Experiments across multiple datasets demonstrate that the proposed metric serves as a valuable measure for improving the robustness of machine learning models toward OOD prediction.

46. Split Conformal Prediction with Label-Shift-Adjusted Bayesian Scores

分割共形预测与标签偏移调整的贝叶斯分数

AI 总结:针对标签偏移下共形预测失效问题,提出标签偏移调整的贝叶斯分数(LSA分数),通过后验预测倾斜恒等式修正贝叶斯分数,在分子性质预测中生成更短区间且覆盖相当。

链接:https://arxiv.org/abs/2609.12386

机构:MOGAM Institute for Biomedical Research(MOGAM生物医学研究所); CROID Research(CROID研究); aSSIST University(aSSIST大学)

作者:Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi, Hyunjin Shin

英文摘要:Conformal prediction provides distribution-free uncertainty quantification under exchangeability. However, this assumption is violated by label shift, where the marginal distribution of labels changes while the conditional distribution of inputs given labels remains stable. Under such shifts, standard conformal procedures no longer maintain their intended coverage behavior. Existing approaches address this via importance weighting. They pair the reweighting with residual-based nonconformity scores that ignore predictive uncertainty. The resulting intervals have uniform width. Bayesian conformal methods produce adaptive intervals by leveraging predictive distributions. They evaluate conformity under the source predictive, which is misaligned with the target domain under label shift. We propose the \emph{Label-Shift-Adjusted Bayesian Score} (LSA score), a nonconformity score derived from a posterior predictive tilting identity. This identity shows that the target predictive is an importance-weighted transformation of the source predictive. We use it to derive a direct correction to the Bayesian score. We evaluate the method on molecular property prediction under controlled label shift. The LSA score consistently yields shorter intervals than residual-based and source-based Bayesian scores. Coverage in the target domain remains comparable. Under stronger shift, all methods incur some coverage loss due to pseudo-label-based density-ratio estimation. The LSA score is defined for any source predictive with a tractable log-density. We instantiate it with Bayesian Ridge Regression, where the correction admits a closed form.

8. 图学习与结构化数据 | 2 篇

47. $\text{GSF-}χ$: Global Stereochemical Fields for Chiral Graph Transformers

$\text{GSF-}χ$:手性图变换器的全局立体化学场

AI 总结:提出GSF-$\chi$图变换器,利用立体生成单元调制所有原子对相互作用并引入手性RoPE,在中心ECD上全面领先,轴向任务提升12.6%和7.9%。

链接:https://arxiv.org/abs/2609.12532

机构:Shanghai Innovation Institute(上海创新研究院); Fudan University(复旦大学)

作者:Jiaqing Xie, Yuxin Wang, Xipeng Qiu

英文摘要:Enantiomers share atoms, bonds, and pairwise distances yet can behave differently in chiral environments, so molecular encoders must respect atom relabelings and proper rotations without becoming blind to reflection. We introduce GSF-$\chi$, a graph transformer in which stereogenic units modulate all pairwise interactions rather than single out one atom as special. Each central or axial stereogenic unit creates a reflection-even phase field over all atoms, a handedness pseudoscalar $\chi$ sets the direction of a relative rotation on latent query--key blocks, giving a \textbf{Chiral-RoPE} that reflection inverts rather than leaves fixed. A $C_2$ projection separates mirror-even ECD peak counts and positions from mirror-odd peak signs. We prove the operator's even--odd decomposition and its annotation-inversion, permutation, and unit-order identities under explicit canonical-role conditions; property tests and a coordinate-reflection audit verify the laws end to end. GSF-$\chi$ leads every central-ECD output and improves axial Rotation and Symbol by $12.6\%$ and $7.9\%$ over the strongest baseline. Equal-budget controls attribute the Rotation advantage to global signed support rather than parameter or edge count; the $C_2$ projection yields exact enantiomer-pair consistency at a small raw-accuracy cost under complete supervision and becomes predictive when mirror supervision is scarce.

48. InRTL: Effective Intra-Inter Interaction Learning for Relational Tables

InRTL:关系型表格的有效表内-表间交互学习

AI 总结:本文提出InRTL,一个统一框架,通过列感知编码器、Transformer自注意力与交叉注意力,以及线性化注意力和异构图神经网络,显式建模关系型表格的表内与表间交互,在24个真实任务上验证了有效性。

链接:https://arxiv.org/abs/2609.12712

机构:Shanghai Jiao Tong University(上海交通大学)

作者:Weichen Li, Ken Zhong, Zheng Wang, Li Pan, Jianhua Li

英文摘要:Relational table learning has recently emerged as an important research direction for modeling multiple tables connected through primary key-foreign key (PK-FK) relationships. Despite recent advances, a principled modeling framework tailored to this task remains underexplored. In this paper, we propose Intra-Inter Relational Table Learning (InRTL), a unified framework that explicitly models dependencies both within and across relational tables. Specifically, InRTL formalizes two complementary interaction patterns: intra-table interactions, describing associations among rows within the same table, and inter-table interactions, describing dependencies between rows across PK-FK-linked tables. To model these dependencies, we develop a column-aware table encoder to generate initial row representations, followed by Transformer-based self-attention and cross-attention modules for intra-table and inter-table learning, respectively. To further improve scalability, InRTL incorporates linearized attention and heterogeneous graph neural networks to simplify the self-attention and cross-attention operations. Extensive experiments on ten datasets covering 24 real-world tasks demonstrate the effectiveness of our approach. Code is available at this https URL.

9. 迁移、元学习与持续学习 | 5 篇

49. Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature

各向同性曲率下通过联合切空间优化的秩高效LoRA

AI 总结:针对LoRA秩利用率不足问题,提出ISO-LoRA优化器,通过谱下降耦合因子更新,提升有效秩与下游性能,在0.1B-7B模型上验证有效性。

链接:https://arxiv.org/abs/2609.12123

机构:University of Pennsylvania(宾夕法尼亚大学); DRW Associates LLC(DRW联合有限责任公司); University of Chicago(芝加哥大学)

作者:Zihan Zhu, Zhehang Du, Xuyang Chen, Tim Tsz-Kit Lau, Jiayuan Wu, X. Y. Han, Qi Long, Weijie Su

英文摘要:Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapter's parameter budget and representational capacity. We show that this view is incomplete: while the nominal rank determines the representational capacity, the optimizer shapes how much of that capacity is used in the induced weight-space updates. In a case study of GPT-2 adaptation with LoRA, we observe a strong rank-dependent optimizer effect. Despite using the same nominal rank, AdamW often produces per-step updates with concentrated singular spectra and low effective rank, whereas Muon uses a richer set of directions and benefits more consistently from increasing LoRA rank. These observations motivate ISO-LoRA, an optimizer that couples the LoRA factor updates through spectral descent on the induced tangent perturbation in weight space. ISO-LoRA promotes updates that distribute energy more evenly across singular directions, improving rank utilization while preserving compatibility with the LoRA parameterization. We complement this design with theoretical guarantees showing that ISO-LoRA can achieve higher effective rank than standard factor-wise optimizers through a one-step analysis under a stylized spiked-gradient model. We validate this design on language-model adaptation across 0.1B-7B-parameter models, where ISO-LoRA improves effective rank and downstream performance, with the strongest gains at moderate-to-large LoRA ranks. Our results highlight rank utilization as a key factor in LoRA optimization and suggest that optimizer design offers an important path toward stronger parameter-efficient adaptation.

50. Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

基于摊销低秩适应的模型强化学习

AI 总结:针对世界模型在测试环境中的适应问题,提出CLAW方法,利用超网络生成低秩适配器,在少量数据下实现高效适应,优于梯度适应和上下文学习。

链接:https://arxiv.org/abs/2609.12278

机构:University of Texas at Austin(德克萨斯大学奥斯汀分校)

作者:Fernando Palafox, David Fridovich-Keil

英文摘要:World models let agents plan by predicting the consequences of their actions, but changes in the environment can make them inaccurate. We study the problem of adapting a world model to an unknown test-time environment, drawn from a known environment family, using only a few episodes of interaction. Existing approaches trade off computational cost against expressivity, i.e., the range of models a method can produce. For example, in-context learning is computationally cheap but limited in expressivity, and gradient-based adaptation is expressive but computationally expensive. We present CLAW (Context-conditioned Low-rank Adaptation of World models), which addresses this tradeoff by using a hypernetwork to generate low-rank (LoRA) adapters at test time. During pretraining, we simulate adaptation to a variety of environments and jointly train the hypernetwork and base world model. At test time, we freeze the base model and use a forward pass of the hypernetwork to generate adapters from a small batch of test-time transitions. We evaluate CLAW in locomotion and manipulation environment families that vary in dynamics, embodiment, and reward. We show that, using only seconds of test-time data, CLAW outperforms gradient-based adaptation and in-context learning during online adaptation. We also show that CLAW avoids overfitting in data-scarce regimes, that its advantage comes from the expressive adapters rather than context conditioning, and that pretraining the hypernetwork jointly with the base model outperforms training it post hoc.

51. Geometric-to-Semantic Spherical Transfer Learning for Cortical Sulci Labeling

皮层沟标注的几何到语义球面迁移学习

AI 总结:针对皮层沟标注数据稀缺导致过拟合的问题,提出几何到语义球面迁移学习框架,利用大规模无标注数据预训练并注入拓扑先验,平均Dice达0.77,在罕见沟上提升达14.8%。

链接:https://arxiv.org/abs/2609.12627

机构:Paris-Saclay University(巴黎-萨克雷大学); CEA(法国原子能委员会); NeuroSpin(NeuroSpin神经影像研究中心); Baobab(Baobab实验室); LTCI, Télécom Paris, Institut Polytechnique de Paris(巴黎综合理工学院电信学院LTCI实验室)

作者:Saeb Tounsi, Joël Chavas, Pietro Gori, Vincent Frouin, Denis Rivière, Jean-François Mangin

英文摘要:Deep learning on cortical surfaces faces a dilemma: capturing the complex topology of over 60 nomenclature-dependent sulci per hemisphere requires high-capacity models, yet the extreme scarcity of expert annotations ($N=62$ subjects) inevitably causes overfitting. Standard supervised approaches fail to generalize in this data-scarce regime, particularly for variable and small sulci where topological ambiguity is high. To overcome this limitation, we introduce a Geometric-to-Semantic Spherical Transfer Learning framework. First, we leverage massive unlabeled data (UK Biobank, $\approx$30,000 subjects) to pre-train a spherical encoder using a locally-optimized strategy. By relying solely on continuous surface features (curvature and depth), the relevance of this pre-training is confirmed by the model's ability to detect localized and rare topological traits, such as sulcal interruptions. The downstream labeling task, however, introduces extracted sulcal fundi (lines) as an explicit semantic input. To bridge this dimensional domain gap (from purely geometric to semantic) without causing catastrophic forgetting, these anatomical lines are integrated into the pre-trained backbone via a soft-initialized Topological Prior Injector. Our experiments demonstrate that this approach outperforms fully supervised baselines trained from scratch, achieving a mean Dice of 0.77. Crucially, a local analysis reveals that the self-supervised geometric priors yield the largest performance gains on variable and tertiary sulci (up to 14.8%), confirming that learning the cortex shape is highly beneficial for identifying its rarest parts.

52. Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents

行为商学习用于LLM智能体的低秩适配

AI 总结:针对LLM智能体多LoRA适配器存储与路由开销问题,提出BQ-LoRA框架,通过行为商平衡与决策保持压缩在固定秩下组织轨迹更新,实验验证其优于标准LoRA及近期方法。

链接:https://arxiv.org/abs/2609.12896

机构:Alibaba Cloud(阿里云)

作者:Pengyang Zhou, Xiaobin Tu, Zhengxi Liu, Rongkun Xue, Haochen Li, Miancan Liu, Ziyuan Chen, Yinggui Wang, Jinkui Ren, Xiantao Zhang

英文摘要:LLM-based agents rely on heterogeneous interaction capabilities to accomplish complex tasks. Existing approaches often distribute these capabilities across multiple LoRA adapters, which increases adapter storage requirements and introduces routing overhead during inference. A single LoRA avoids this overhead, but learning from diverse agent trajectories under a fixed rank budget presents two challenges. First, trajectories with different interaction traces and parameter gradients can induce equivalent changes in decision distributions, causing repeated updates to overemphasize redundant behavioral changes. Second, an aggregated update may exceed the rank budget of the adapter, and approximating it in weight space can distort the decision changes that it is intended to produce. We propose BQ-LoRA, a low-rank adaptation framework that organizes trajectory updates through a local behavior quotient manifold. It contains two modules, i.e., behavior quotient balancing (BQB) and decision preserving compression (DPC). BQB constructs the quotient manifold from decision distributions and reweights trajectory update directions according to their local density in the quotient tangent space. DPC projects the balanced gradient onto the intrinsic fixed rank tangent space and refactorizes the resulting target by jointly controlling effective weight error and distortion of decision distributions. Experiments on AppWorld and BrowseComp-Plus compare BQ-LoRA with standard LoRA and recent low-rank adaptation methods, while separate ablations evaluate the complementary contributions of both components.

53. Transfer Learning for Evolving Domains

演化领域的迁移学习

AI 总结:本文提出演化领域迁移学习(TrED)问题,将经典迁移学习设置统一为动态数据可用性轨迹,并论证其为一个适定且未解决的重要研究方向。

链接:https://arxiv.org/abs/2609.13039

机构:Feedzai; University of Porto(波尔图大学)

作者:Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Pedro Ribeiro, Pedro Saleiro, Pedro Bizarro, Carlos Soares

英文摘要:Transfer learning explores how to leverage knowledge from various tasks or domains (sources) to enhance predictive performance in related tasks or domains (targets). Typically, transfer learning research is segmented into several isolated sub-areas (such as domain generalisation, domain adaptation, or multi-domain learning), each making distinct assumptions about target data availability, namely how much data and how many labels are available at training time. However, in many real-world applications, data availability is not fixed but evolves over time, as instances and labels are progressively collected from a new domain. Each of the classical settings then describes only a snapshot of a trajectory that a deployed system must traverse in full. We formalise this trajectory as a transfer learning problem in its own right, Transfer Learning for Evolving Domains (TrED), specified by a data availability process fixed by the environment, a learning protocol that the method is free to choose, and an evaluation criterion that scores the whole trajectory of models rather than a single one. Within this formalism, the classical settings are recovered as regimes that a learner may pass through, rather than as separate problems that TrED concatenates. We then examine the transfer learning literature to identify mechanisms that are promising building blocks for a solution, and find that most methods are tailored to a single regime and that even the strongest existing candidates do not yet optimise the whole trajectory. We argue that TrED is a well-posed and unsolved problem, and an important direction for future research.

10. 数据集、基准与评测 | 5 篇

54. Look Before You Leap: Pre-Action Verification for LLM Agents

三思而后行:面向LLM智能体的行动前验证

AI 总结:本文提出行动前验证框架,通过确定性检查在动作生效前捕获LLM智能体的静默失败,在shell命令和代码编辑上分别达到95.8%和99.9%的准确率,并发布基准与验证器。

链接:https://arxiv.org/abs/2609.11957

作者:Asaad Althoubi

英文摘要:An LLM agent acts on the world by emitting actions: shell commands to run, edits to apply. A wrong action does not always fail loudly; it can fail silently, producing a plausible but incorrect effect that raises no error. We argue that a cheap deterministic check, run before an action takes effect, is an effective and underused form of agent oversight, and we study it across two action modalities in one framework. The idea is to fix an action's correct effect by construction, before any executor runs, so that silent failure is measured directly and the verifier may abstain rather than guess. For shell commands, a static verifier over 9930 commands and 482 tools catches 95.8% of invalid commands at a 10.0% false-positive rate. Its syntax and binary checks are oracle-exact, giving zero false positives while catching half of all errors; the flag check is bounded only by help-text coverage and accounts for every false positive. For code edits, a benchmark of 640 edits over 224 files isolating the apply step exposes a sharp split. Content-anchored formats such as search/replace and diff fail cleanly, whereas location-anchored formats fail silently: line numbers corrupt 99.1% of files under a one-line shift, and function-name edits hit the wrong function 12.7% of the time. In both settings a refuse-when-unsure policy turns silent failures into recoverable ones at a tunable cost in applicability: selective grounding reaches 0.958 recall at 7.0% false positives, and an anchor-and-verify applier records one silent misapplication in 8320 trials (0.01%). We release both benchmarks, the verifiers, and the guards.

55. ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents

ParaRecover:并行工具使用智能体中错误定位与恢复的过程级基准

AI 总结:ParaRecover是一个过程级基准,通过14种错误类型和10,626个实例评估多轮并行工具使用智能体的错误定位与恢复能力,并提出SDE评分标准以提升其反思性恢复能力。

链接:https://arxiv.org/abs/2609.12345

机构:Dalian University of Technology(大连理工大学)

作者:Bowen Guan, Zhentao Yin, Yanming Shen

英文摘要:Existing agent benchmarks mainly evaluate final task success or tool-call correctness, providing limited insight into whether agents can reliably diagnose and recover from intermediate execution failures. This limitation becomes particularly critical in multi-turn parallel tool-use scenarios, where errors may propagate across dependent branches and trigger cascading failures. We introduce ParaRecover, a process-level benchmark for evaluating error localization and recovery in multi-turn parallel tool-use agents. Built upon a fine-grained taxonomy of 14 error types covering planning dependencies, tool selection, and argument matching, the benchmark comprises 10,626 instances spanning two difficulty levels. To enable finegrained, process-oriented evaluation, we further propose the SDE rubric, which measures structural integrity, diagnostic reasoning, and evolutionary strategy during agent this http URL across more than ten mainstream LLMs reveal that even state-of-the-art models still struggle with multi-turn error propagation,implicit tool-use failures, and precise replanning. Moreover, we demonstrate that the SDE rubric provides effective supervision signals for improving agents' reflective recovery capabilities. Our data and code are available at this https URL.

56. Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning

基于聚类的平衡采样与数据并行分配用于高性能微调

AI 总结:针对指令微调数据冗余与不平衡问题,提出聚类感知的平衡采样框架CluSTER,通过梯度空间聚类和DP感知分配实现高效数据缩减,最多减少69.6%训练时间且几乎无精度损失。

链接:https://arxiv.org/abs/2609.12584

机构:KAIST(韩国科学技术院)

作者:Hyunjin Kim, Youngeun Nam, Jaemin Han, Wonhyeok Choi, Jae-Gil Lee

英文摘要:Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation. Naive large-batch fine-tuning repeatedly includes overrepresented sample groups while weakly covering underrepresented but informative ones, especially under data parallelism (DP) across multiple GPUs. We propose CluSTER, a Cluster-aware balanced Sampling framework for Training Efficient data Reduction in DP instruction tuning. CluSTER curates a representative reduced dataset through gradient-space clustering and DP-aware balanced allocation, ensuring dual-level coverage across clusters and workers, while preserving the original data distribution by weighted update. As a result, CluSTER reduces redundant computation and improves training stability without compromising model quality. Across multiple instruction-tuning datasets, CluSTER reduces training time by up to 69.6% with almost no accuracy loss compared to prior sampling and data reduction methods. Code is available at this https URL.

57. A Large-Scale AIS Dataset from Finnish Water

来自芬兰水域的大规模AIS数据集

AI 总结:本文发布了一个包含2.29亿数据点的波罗的海及芬兰湖泊AIS数据集,并整理了现有数据集,提供船舶类型分析与交通可视化,以促进海事研究。

链接:https://arxiv.org/abs/2609.12938

机构:Åbo Akademi University(奥博学术大学)

作者:Debayan Bhattacharya, Ikram Ul Haq, Carlos Pichardo Vicencio, Sebastien Lafond

英文摘要:This research paper contributes to the maritime research community by introducing a comprehensive AIS dataset from Finnish waters, specifically the Baltic Sea region. AIS data, initially designed for collision prevention, have evolved into a versatile tool with applications across diverse maritime domains. Our paper not only curates and categorises existing AIS datasets but also introduces a collected AIS dataset from the Baltic Sea area, renowned for its intercontinental cargo routes, military activities, and frozen water expanses. This dataset includes 229 millions data points and provides researchers with a resource for studying maritime activities and vessel behaviour in this dynamic region, notably distinguished by the inclusion of data from Finnish lakes. Our analysis includes a detailed listing of ship types and relevant features, empowering researchers to explore various maritime domains. To enhance comprehension and analysis, we provide visualisations of maritime traffic patterns. By publishing this AIS dataset, we aim to catalyse innovation and collaboration in maritime research, offering a gateway to deeper insights into maritime activities in the Baltic Sea and Finnish lakes.

58. MAxBench: A Multinomial Concept Recovery Benchmark

MAxBench:一种多项式概念恢复基准

AI 总结:MAxBench是一个与几何无关的多项式概念恢复基准,通过比较10种定位方法,发现仿射子空间引导更可靠且召回率更高,且优势源于非零偏移,流形引导有竞争力,但无方法优于提示。

链接:https://arxiv.org/abs/2609.13072

机构:Boston University(波士顿大学); Technion – Israel Institute of Technology(以色列理工学院); Kempner Institute, Harvard University(哈佛大学肯普纳研究所)

作者:Divya Appapogu, Freya Behrens, Yonatan Belinkov, Aaron Mueller

英文摘要: Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research. For binary concepts such as refusal, a single direction in activation space often suffices for steering. However, many concepts are not binary: Animals and Countries contain many subcategories, each with multiple instances. For these concepts, the search space over possible representation geometries is far larger than for binary concepts; it is thus not clear what geometries are most appropriate, nor what methods are most effective at recovering them. In this work, we introduce MAxBench, a geometry-agnostic evaluation framework for multinomial concept representations based on sampling from the recovered concept representation. We use MAxBench to compare 10 localization methods (covering 5 geometry types) across 6 concepts and 4 models. Using this framework, we find that (i) affine subspaces steer more reliably and have greater recall than rank-one or linear subspaces; (ii) much of this advantage is due to better non-zero offsets rather than the choice of bases; (iii) manifold steering is competitive with the best methods when applicable; and (iv) no method consistently outperforms prompting, in alignment with prior findings on binary concepts. These findings underscore the importance of expanding the scope of interpretability research and meta-evaluation to concepts with more varied structure.

11. 机器学习应用 | 5 篇

59. FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences

FINESSE:面向多模态金融事件序列的基于智能体的模拟器与基准数据集

AI 总结:针对金融服务领域开源数据稀缺问题,提出基于智能体的模拟框架FINESSE,生成多模态事件流数据集,并配套基准FINESSE-Bench,支持四项预测任务,加速结构化事件序列建模研究。

链接:https://arxiv.org/abs/2609.11993

机构:Capital One(第一资本)

作者:Tyler Farnan, Benjamin Eng, Adam Abate, Xirui Hou, Rizal Fathony, Nam H. Nguyen, Senthil Kumar

英文摘要:Machine learning research in financial services is limited by the scarcity of representative open-source datasets. Existing resources are often narrowly focused on a single modality or task and fail to reflect the structured, multimodal, and dynamic nature inherent to many problems in financial services. In this paper, we introduce FINESSE, a Financial Event Sequence Simulation Environment, an agent-based simulation framework for generating synthetic, structured datasets composed of multiple interdependent event streams. Each stream corresponds to a distinct financial behavior such as transactions, payments, account status changes, and policy interventions, each with unique action spaces, schemas and variable types. These streams are coupled through agents' latent evolving states, enabling the simulation of temporally rich interactions. We also introduce FINESSE-Bench, a benchmark dataset generated by the simulator, supporting four representative tasks: balance forecasting, transaction fraud detection, missed payment prediction, and next event prediction. We report baseline results using methods from time series forecasting, event sequence modeling, temporal graphs, and temporal point processes. We release the FINESSE framework, including the simulator and dataset to accelerate research on structured, multimodal event sequence modeling challenges in financial services.

60. Observation-Anchored Selective Assimilation for Longitudinal Tumor-State Proxy Forecasting in Post-Treatment Glioma

基于观测锚定的选择性同化用于胶质瘤治疗后纵向肿瘤状态代理预测

AI 总结:本研究提出观测锚定选择性同化(OASA)方法,利用中间观测锚定患者状态并选择性更新,在胶质瘤术后MRI预测中实现与持久化相当的Dice性能,并改善校准。

链接:https://arxiv.org/abs/2609.12435

机构:Yonsei University(延世大学)

作者:Yeonjae Jung, Minwoo Shin

英文摘要:Post-treatment MRI in patients with glioma provides serial observations for updating patient-specific tumor-state proxy estimates, but variable appearances and trajectories complicate forecasting. We formulate forecasting as an observation-aware digital-twin update in which an intermediate observation anchors the patient-specific state. Among 203 patients and 594 follow-up time points, a predefined no-new-treatment criterion retained 120 of 236 candidate triplets, split into 81/24/15 training/validation/test triplets at the patient level. Each time point was represented by a continuous voxel-wise tumor-state proxy map in [0,1] derived from MRI lesion labels. A SegMamba-based single-step forecaster predicted update proposals from multimodal source-state tensors. Observation-Anchored Selective Assimilation (OASA) retained the observed intermediate proxy as the state anchor and selectively applied updates through a validation-selected tiered case-level rule and voxel-wise soft gate. We compared initial-scan forecasting, rollout without assimilation, latest-observation persistence, direct prediction, OASA, OASA + calibration, and morphological dilation. Checkpoints, OASA rules, and calibration thresholds were selected using validation data only. Across three seeds on 15 held-out test triplets, OASA maintained Dice at $\tau$ = 0.2 comparable to persistence (0.6071 $\pm$ 0.0025 vs. 0.6070) while yielding numerically higher Dice at $\tau$ = 0.5 (0.4269 $\pm$ 0.0079 vs. 0.3981), with a small RMSE increase. Calibration increased Dice at $\tau$ = 0.2 to 0.6178 $\pm$ 0.0025, increased false-positive (FP) support (11,836$\rightarrow$18,663), and reduced false-negative (FN) support (22,107$\rightarrow$17,536). This reflects near-threshold support calibration rather than improved biological predictive capability. Code is publicly available at this https URL.

61. Explaining Time Series Forecasting with Horizon-Resolved Attribution

用水平分辨率归因解释时间序列预测

AI 总结:针对时间序列预测中不同预测步骤依赖不同过去值的问题,提出水平分辨率解释(HRX)框架,为每个预测步骤生成独立重要性图,并通过实验验证其有效性。

链接:https://arxiv.org/abs/2609.12639

机构:LG AI Research(LG AI研究院)

作者:Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn

英文摘要: Recent advances in explaining time series (TS) models have produced methods that identify which past values a prediction depends on. However, most existing methods return a single importance vector, assuming that every predicted step depends on the same past values. In this paper, we show that this assumption does not hold, as different forecast steps depend on different past values. Motivated by this observation, we propose Horizon-Resolved eXplanation (HRX), which adds a horizon axis to the explanation, so that every forecast step receives its own importance map. HRX is a simple yet effective plug-in framework with three components: 1) an estimator that reads these maps out of any differentiable forecaster without modifying the TS backbone, 2) an evaluation protocol that validates the horizon axis by measuring how much a single forecast step changes when the inputs an importance map ranks highest are removed, and 3) a rank criterion that predicts in advance whether the axis is worth resolving on a given TS. We further show that this step-wise dependence is low-dimensional, as the explanations of all steps are built from a few shared maps whose number does not grow with the forecast length. Extensive experiments across various backbones and datasets show that the improvement comes from the horizon axis and holds for estimators of previous explanation methods. Code is available at this https URL.

62. VertiFuseX: Generalizable Financial Forecasting via Multi-Stream Temporal Fusion

VertiFuseX:通过多流时间融合实现可泛化的金融预测

AI 总结:VertiFuseX提出一种混合LSTM架构,通过倒数第二层垂直融合多尺度时间表示,在固定超参数下提升金融预测的准确性和泛化能力,显著降低预测误差并支持轻量部署。

链接:https://arxiv.org/abs/2609.12793

机构:Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校)

作者:Aashish Bohra, Vivek Vijay

英文摘要:Stock price prediction remains challenging due to the non-stationary and noisy nature of financial time series. Existing deep learning models often rely on rigid decision-level fusion, ad hoc hyperparameter tuning, and compressed final-layer outputs, causing information loss, overfitting, and limited cross-market generalization. We propose VertiFuseX, a hybrid LSTM architecture using penultimate-layer vertical fusion of multi-scale temporal representations. VertiFuseX stacks and reweights penultimate features from LSTM, Bi-LSTM, and St-LSTM branches, integrates a parallel DNN stream, and jointly optimizes all components via backpropagation under a fixed hyperparameter configuration. This preserves richer intermediate temporal information across scales. Evaluated on 15 years (2010-2024) of closing prices from 10 global equity indices using strict chronological out-of-sample testing with the final 365 trading days held out, VertiFuseX achieves 30-54% MAPE reductions and over 40% improvements in MAE and RMSE versus LSTM-based baselines, and outperforms seven state-of-the-art models across 33 metric-dataset comparisons. Ablation studies confirm penultimate-layer fusion drives these gains over final-layer fusion and decision-level ensembling. Gradient-based saliency analysis shows consistent emphasis on mid-range dependencies at lags 9-15 days. Economic validation via algorithmic trading simulation under extreme market regimes shows reduced maximum drawdowns and superior risk-adjusted returns. With 675k parameters, a 2.6 MB memory footprint, and 1.5 ms/sample inference latency, VertiFuseX offers a lightweight, interpretable, deployment-ready framework for robust financial forecasting.

63. Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting

大距离梯度未必可靠:面向长时程自回归预测的可靠性加权信用分配

AI 总结:针对长时程自回归预测中远距离梯度可能不可靠的问题,提出Internal-DW方法,通过可靠性加权内部梯度路径,在多个测试平台上降低预测误差5.2%-13.8%,优于现有方法。

链接:https://arxiv.org/abs/2609.12890

机构:University of Maryland, College Park(马里兰大学学院公园分校)

作者:Junhao Zhao, David Michael Simberg, Jacob Kang, Colin Connor Kurniawan, Nan Xu

英文摘要:In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make distant gradients dominate the update while amplifying predictable signal and unpredictable noise together; a large distant gradient therefore need not carry reliable learning signal. Motivated by this observation, we introduce Internal Dual-Wiener routing (Internal-DW), a principled backward-only intervention that preserves the full forward rollout and all horizon losses while reliability-weighting internal gradient routes. At each residual block, we derive bounded Wiener gains for the identity and nonlinear routes that balance preserving predictable learning signal against suppressing unpredictable variation, and estimate them from route-level gradient statistics and an explicit noise model. In a controlled system with known gradient signal-to-noise ratio (SNR), we show that distant gradients can grow even as their SNR falls, and that Internal-DW reduces held-out error in recovering predictable gradient signals and improves forecasting. On four history-dominated, weak-drive testbeds, Internal-DW reduces forecast error by 5.2%-13.8% relative to full BPTT, outperforms gradient clipping and Jacobian regularization on all four, and outperforms validation-selected truncated BPTT (TBPTT) on three. It also extends or preserves the fitted optimal training-horizon range across these four testbeds. Across the full benchmark suite, the current Internal-DW estimator has a clear applicability boundary: its benefit diminishes or reverses when usable history is limited or when the selected sampler fails to represent dominant drive-dependent variation. The results show that retaining long-horizon supervision does not require trusting every backward contribution equally.

12. 其他/综合机器学习 | 36 篇

64. Fundamental Dynamical Units for Physics-Informed Structural Inference from Perturbation Time-Series in Networked Systems

网络系统中基于扰动时间序列的物理信息结构推断的基本动力学单元

AI 总结:针对网络动力系统从扰动时间序列推断带符号交互结构的问题,提出以基本动力学单元为可组合基元,结合物理信息神经ODE,实现结构恢复与干预设计的统一框架。

链接:https://arxiv.org/abs/2609.11934

机构:Artificial Intelligence for Science Innovation, AstraZeneca(阿斯利康科学创新人工智能部门)

作者:Nima Nouri

英文摘要: In networked dynamical systems, the parameter of primary mechanistic interest is signed interaction structure. Recovering this structure from perturbation time-series data is a fundamental identification problem, compounded by three coupled obstacles: the combinatorial complexity of interaction architectures, ambiguity of causal attribution under limited interventions, and state-dependent dynamics that confound structural inference. Each obstacle is structural in origin and calls for a structural solution. We address these challenges by adopting a reductionist approach, introducing Fundamental Dynamical Units (FDUs): signed three-node interaction patterns as composable primitives that convert the interaction hypothesis space into a finite, constructive, and tractable representation. We show that local interaction structure determines the perturbation conditions required to disentangle direct from relayed influence, making intervention design a structural consequence of the FDU representation. We embed FDU-regularized structural inference within a physics-informed neural ordinary differential equation (ODE) whose governing-equation constraint transforms structural hypotheses into verifiable dynamical predictions, enabling joint recovery of interaction structure and perturbation-resolved trajectories. Validated on synthetic benchmarks with known ground truth, the framework supports structural commitment, expressed through FDU primitives, motif-prescribed intervention design, and physics-informed learning, as a principled basis for mechanistically interpretable inference in networked dynamical systems.

65. Efficient AI Model Deployment Using Quantization Analysis Tool

高效AI模型部署:基于量化分析工具

AI 总结:本文介绍一种基于ONNX的量化分析工具,通过逐层敏感性分析与分布可视化,指导精度选择,在保持准确率的同时优化模型大小和延迟,实现高效AI模型部署。

链接:https://arxiv.org/abs/2609.11954

作者:Dwith Chenna, Kanishka Macherla

英文摘要:As deep learning models are increasingly deployed on resource constrained devices, the demand for efficient model optimization techniques continues to grow. Effective deployment of AI models on edge and low power platforms requires optimization methods that reduce model size and computational cost while maintaining high accuracy. This paper presents Quantization Analysis Tool, a practical system designed to streamline quantization workflows and support performance efficient model deployment. Built on the ONNX framework for broad interoperability, the tool provides detailed layer-wise sensitivity analysis, visualization of weight and activation distributions, and insights to guide precision selection. By identifying layers that are resilient or sensitive to reduced precision, the tool enables developers to make informed trade-offs between model size, latency, and accuracy. Experimental evaluations across multiple neural network architectures demonstrate that the tool effectively improves the quantized accuracy, leading to improved efficiency in real-world deployment scenarios. The tool also provides developers valuable insights into the effects on quantization on the model and its accuracy. This work highlights the tools capabilities, practical applications, and its role in enabling efficient AI model deployment through robust quantization analysis

66. Decoding Mixture Perception through Computational Modeling of Component Interactions

解码混合物感知:基于组分相互作用的计算建模

AI 总结:本研究提出仿生深度学习框架,通过融合注意力加权多受体与浓度依赖多分子响应曲线,建模组分相互作用,实现混合物气味感知识别,准确率达92.2%,可集成于具身系统。

链接:https://arxiv.org/abs/2609.11958

作者:Fei Wang, Xiaoya Xie, Junfei Liu, Huihao Wang, Yixiao Wang, Yintao Wang, Yi Li, Hao Dong, Xing Chen

英文摘要:Olfaction played an indispensable role throughout human evolution and civilization. Even in the contemporary era of advanced technology, olfaction remains a critical channel for person to conduct danger discrimination, emotional experience, and memory formation. However, most substances in nature exist as multi-molecule mixtures. The complexity of mixture compositions, as well as concentration dependent saturation effects and receptor specific activation thresholds, pose substantial challenges in identifying olfactory characteristics. In this study, we proposed a novel bio inspired deep learning framework for accurate odor perception recognition of mixtures. We robustly constructed neural response curves for molecule-receptor interactions, and developed a fusion strategy that integrates attention-weighted multi-receptor curves with concentration-dependent multi-molecule curves, replicating the competitive activation and synergistic integration of mixture components. Furthermore, by comparing the consistency of response curve patterns, the model can transfer knowledge from the semantically rich space of molecular associations to guide recognition of mixture perception characteristics. Therefore, we established a complete computational pathway from chemical blending, neural encoding, to perceptual formation. Finally, we conducted comprehensive evaluation, and results demonstrated exceptional superiority, achieving an accuracy of 92.2%. Consequently, our work provides a generalizable solution to the long standing mixture perception challenge. More importantly, it can be integrated into embodied cognitive systems to enhance the agents perceptual and interactive capabilities in complex scenarios.

67. Space as an Interventional Invariant: Cross-Modal Predictive Geometry for Stratified Cities and Em-Spaced Intelligence

空间作为干预不变量:分层城市与仿真空间智能的跨模态预测几何

AI 总结:本文提出将空间定义为干预不变量,构建跨模态预测几何框架,在理论上证明潜在空间可识别性,并扩展到分层城市系统,为空间认知与具身智能提供统一基础。

链接:https://arxiv.org/abs/2609.11959

机构:Tsinghua University(清华大学); University College London(伦敦大学学院); Cardiff University(卡迪夫大学)

作者:Tao Yang, Xuhui Lin, Kunyao Li, Haijiang Li

英文摘要: Space is a foundational concept across mathematics, physics, spatial cognition, urban science, and embodied intelligence, yet these fields often treat spatial structure either as a shared geometric container or as a collection of disconnected representations. Such approaches struggle to explain how heterogeneous sensory and urban processes can jointly reveal a common spatial structure, particularly when different modalities do not share the same metric or representation. This paper addresses this gap by defining space as an interventional invariant: the minimal relational structure that preserves local compatibility and the conditional laws of future observations under admissible actions. We develop a cross-modal predictive geometry that integrates local state spaces, modality-specific observation maps, an action groupoid, and a canonical predictive-state quotient, with explicit causal conditions for identifying interventional rather than merely observational structure. The key theoretical result shows that, under joint point separation, equivariance, and interventional faithfulness, the latent space is identifiable up to the centraliser of the intervention group, thereby reducing representational ambiguity to residual coordinate freedom. The framework is further extended to stratified urban systems using sheaf-valued representations, allowing geometric, physical, mobility, social, and economic layers to coexist without being reduced to a single metric. Synthetic experiments under noise evaluate equivariance, predictive sufficiency, holonomy, restriction-map recovery, cross-scale consistency, and context saturation. The resulting framework provides a unified and falsifiable foundation for spatial cognition, urban science, embodied AI, and em-spaced intelligence.

68. On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health

面向隐私保护压力预测的设备端语言模型:移动健康的多模态评估

AI 总结:本研究评估设备端语言模型在移动健康中利用零样本提示进行多模态压力预测的可行性,发现客观传感器特征略优且轻量级模型低延迟,凸显其潜力与约束。

链接:https://arxiv.org/abs/2609.11961

作者:Ibukunoluwa Soyebo, Alyssa Donawa, Rodrigo Aguilar Barrios, Brice Patchou, Corey E. Baker

英文摘要:Stress is a pervasive determinant of mental health and a key target for mobile health interventions. On-device language models (ODLMs) offer privacy-preserving inference without cloud dependency, yet their feasibility for health prediction under mobile resource constraints remains underexplored. We evaluate ODLMs for multi-modal stress prediction using zero-shot prompting, measuring predictive accuracy alongside latency and throughput. Our results show that objective sensor features marginally outperform subjective self-reports on average, and that lightweight sub-2B models achieve low latency with predictable resource usage. Our findings highlight both the promise and the practical constraints of ODLMs for mobile mental health.

69. Explainable Prediction from Mobile Sensing Data through LLM-guided Concept Integration

基于LLM引导概念整合的移动感知数据可解释预测

AI 总结:针对小规模移动感知研究中预测难和可解释性差的问题,提出概念整合Transformer(CIT),利用大语言模型生成概念监督,在两个数据集上取得最优F1分数,并揭示可解释的行为生理模式。

链接:https://arxiv.org/abs/2609.11995

机构:University of Turku(图尔库大学); University of California, Irvine(加利福尼亚大学尔湾分校)

作者:Yuning Wang, Iman Azimi, Amir M. Rahmani, Pasi Liljeberg

英文摘要:Mobile sensing enables longitudinal monitoring of behavioral and physiological patterns in everyday settings. However, accurate prediction remains challenging in small-cohort health-sensing studies, where task-specific outcome supervision is limited relative to heterogeneous sensing data. Interpretability is also important, as model outputs should reflect meaningful behavioral and physiological patterns rather than predictive scores alone. We develop a Concept-Integrated Transformer (CIT) with LLM-guided concept supervision for explainable prediction from mobile sensing data. CIT uses a pretrained large language model to generate baseline-aware concept abnormality targets with confidence weights without manual concept annotation. Across two longitudinal datasets, CIT achieves the highest F1 score on AFFECT (0.756) and ties for the highest on a PHQ-9 dataset (0.765). The learned concept scores also reveal interpretable behavioral and physiological patterns; in AFFECT, sleep quantity and quality show the clearest difference between high and low negative affect groups. These findings support LLM-guided concept integration for accurate and interpretable prediction in small-cohort mobile sensing studies.

70. Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

固定状态,长程可达:常量大小缓存为大规模块扩散带来的收益

AI 总结:本文提出利用状态空间缓存实现O(1)内存的块扩散解码,在256k上下文下相比注意力缓存降低11倍内存、提升14倍聚合吞吐,并支持8-16倍训练长度的检索。

链接:https://arxiv.org/abs/2609.11998

机构:Mila(米拉研究所); Concordia University(康考迪亚大学); ServiceNow Research(ServiceNow研究院)

作者:Vaibhav Singh, Pierre-André Noël, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko

英文摘要:Diffusion language models decode tokens in parallel, but their bidirectional denoiser rules out the naive key--value (KV) cache behind fast autoregressive inference. Block diffusion restores caching by decoding block-by-block, and the block caches deployed on it so far are tied to attention: O(L)in memory and, if used as training-free retrofits, only an approximation of the model's computation. Both constraints can be overcome: sequence mixers that summarize finalized blocks into a reusable state support block caching, and the corresponding block-causal training objective makes the cache exact. We study this recipe at scale, pretraining three 3B block-diffusion denoisers (attention, mamba, and hybrid) on 300B tokens under one single-frontier objective and decoding all three through a single cached interface. Only the state-space cache is O(1) in sequence length: its memory and per-step latency stay constant at any context length, while an attention cache remains O(L). At 256k tokens (where attention has grown to 82GB and 29 ms/step), the Mamba cache delivers 4.3x lower latency, 11x less memory, and 2.6x higher single-stream throughput; and because that footprint is constant it scales with batch as well, reaching 14x the aggregate throughput, where attention cannot run beyond a single stream. The same linear-state bias lets the Mamba and hybrid backbones keep retrieving out to 8-16x their training length, whereas attention's retrieval collapses at 2x, at no measured quality cost.

71. Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

我们能信任LLM裁判吗:能力相关偏见与多裁判集成用于偏见校准的研究

AI 总结:本研究揭示LLM裁判存在能力相关偏见,提出基于分歧估计的加权多数投票集成方法,无需标签即可校准偏见,显著提升评估准确性与公平性。

链接:https://arxiv.org/abs/2609.12002

机构:Microsoft(微软); Massachusetts Institute of Technology(麻省理工学院)

作者:Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal

英文摘要:LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use cases. Across four benchmarks and six models (36 judge-examinee pairs), we show that a model's task accuracy strongly predicts its judging accuracy (Pearson $r \geq 0.90$ on most models) and inversely predicts its directional bias ($r \leq -0.83$), but that accuracy alone does not ensure fair evaluation: more capable examinee models consistently receive more lenient judgments from all judges ($r \geq 0.83$). To address this, we propose calibrated weighted majority voting (WMV), an ensemble evaluation method that aggregates multiple LLM judges weighted by online estimates of their false-positive and false-negative rates. We introduce a disagreement-based estimator that derives these error rates purely from inter-judge agreement patterns, requiring no ground-truth labels or task metadata. In a simulated experiment with shifting task distributions, our label-free WMV tracks an oracle with perfect error-rate knowledge to within 0.5 percentage points on average, outperforming both individual judges and unweighted majority voting. These results demonstrate that principled multi-judge calibration can simultaneously improve accuracy and correct for systematic leniency without requiring labeled data, offering a scalable path to reliable automated evaluation as model capabilities increase.

72. QTrans: A Quantum Transformer for Sentiment Classification

QTrans:用于情感分类的量子Transformer

AI 总结:针对轻量级模型难以捕捉情感线索非线性耦合的问题,提出量子Transformer模型QTrans,利用参数化量子电路构建注意力机制,在三个数据集上显著优于经典基线。

链接:https://arxiv.org/abs/2609.12011

机构:Xiangtan University(湘潭大学); Central South University(中南大学); Hunan University(湖南大学)

作者:Ren-Xin Zhao, Xinjie Huang, Yahong Liu, Maoyu Ye, Jinjing Shi, Shi Wang, Yaonan Wang

英文摘要:In small-scale binary sentiment classification scenarios, factors such as negation, contrastive shifts, and cross-word dependencies lead to the non-linear coupling of sentiment cues, making it difficult for conventional lightweight models to fully capture the contextual relationships between tokens. To address this issue, we propose a model named QTrans, which uses parameterized quantum circuits to construct query, key, and value features and derives attention coefficients from Gaussian distances between quantum measurements. By further integrating a quantum feed-forward neural network, residual connections, and layer normalization, the model establishes an end-to-end trainable quantum-classical hybrid framework for sentiment classification. Experimental results on the MR, CR, and MPQA datasets show that QTrans achieves test accuracies of 72.13\%, 69.51\%, and 63.45\%, respectively, representing improvements of 2.88, 3.17, and 3.79 percentage points over the best-performing classical baselines for each dataset. Overall, QTrans expands the application of parameterized quantum circuits in lightweight sentiment analysis and lays an experimental foundation for further research into quantum multi-head self-attention for modeling textual relationships.

73. Toward Reliable Railway-Bogie Response Prediction Using Multifidelity TDNN and Physics-Informed Residual Learning

基于多保真度TDNN与物理信息残差学习的铁路转向架响应预测可靠性研究

AI 总结:针对铁路转向架响应预测,提出结合多保真度TDNN与物理信息残差学习的修正方法,利用仿真与试验数据,在385 km/h工况下实现高精度预测。

链接:https://arxiv.org/abs/2609.12018

机构:Hanyang University(汉阳大学); Korea Railroad Research Institute(韩国铁道技术研究院)

作者:Gyeolhee Lee, Moosun Kim, Taewook Kwon, Jaehun Kim, Dongjin Lee

英文摘要:Railway engineers need simulation models that predict vehicle responses across operating scenarios that cannot be tested exhaustively. Agreement with representative measurements provides essential evidence, but calibration at a limited set of conditions does not guarantee accuracy elsewhere. We present a multifidelity railway-bogie response-correction method that treats multibody simulation histories as low-fidelity information and roller-rig measurements as high-fidelity evidence. This method combines an experiment-anchored fidelity assignment with physics-informed discrepancy learning for multichannel bogie-response histories. A time-delay neural network (TDNN) represents the condition-dependent simulation trend, and development-fitted amplitude alignment defines the low-fidelity baseline. A residual-correction network then models the reproducible response component not explained by this baseline and adds it to the baseline. An effective dynamic-balance equation constrains the learned discrepancy by representing differences in inertia, damping, stiffness, and external forcing between the simulated and physical systems. The training objective combines this constraint with residual matching, temporal smoothness, and selectively applied displacement-acceleration consistency terms. For the evaluated reconstruction case, the corrected response gives a mean coefficient of determination of 0.8197, a mean normalized root-mean-square error (NRMSE) of 4.6055 %, and a mean normalized mean absolute error (NMAE) of 1.9297 %. These results provide initial evidence of accurate response prediction at the held-out 385 km/h condition.

74. Explanations-Driven Active Feature Acquisition for Algorithmic Recourse

解释驱动的主动特征获取用于算法补救

AI 总结:本研究提出解释驱动的特征获取(EDFA)方法,联合优化算法补救与特征获取,利用马尔可夫毯统一多种解释,按解释价值选择特征,在7个数据集上以更少特征实现同等准确率并保证补救有效性。

链接:https://arxiv.org/abs/2609.12179

机构:Western University(西安大略大学)

作者:Vinura Galwaduge, Jagath Samarabandu

英文摘要: Algorithmic recourse methods typically assume that a predictive model has access to all features of an individual. In practice, decisions are often made with partial information, because features are costly to acquire. Active feature acquisition addresses cost-constrained prediction, but existing methods are explanation-agnostic: prior work provides explanations only after acquiring additional features, rather than using explanations to drive acquisition. This work flips that and treats algorithmic recourse and feature acquisition jointly. We use Markov Blanket theory to unify counterfactual, semifactual, and alterfactual explanations and to characterize how available recourse grows as features are acquired. Building on this framework, we propose an Explanation-Driven Feature Acquisition (EDFA) method that selects features by explanatory value per unit cost. The framework is further extended with distribution-free validity guarantees for recourse issued from partial information, which signal trustworthy, lower-cost recourse, along with a lower bound on the calibration data required to certify them. Experiments on 7 publicly available datasets with neural network-based predictive models show that EDFA acquires substantially fewer features than state-of-the-art AFA baselines while maintaining comparable accuracy and yielding more decision-relevant, actionable recourse. The implementation is available on GitHub.

75. CRFCAN: A Complex-Valued Cross-Domain Residual Network for Joint Channel and Phase Noise Estimation in Sub-THz OFDM Systems

CRFCAN:用于亚太赫兹OFDM系统中联合信道与相位噪声估计的复数域跨域残差网络

AI 总结:针对亚太赫兹OFDM系统中联合信道与相位噪声估计的高复杂度问题,提出复数域残差FFT卷积注意力网络CRFCAN,通过物理启发的跨域结构实现端到端联合恢复,显著优于现有算法并具备良好泛化性。

链接:https://arxiv.org/abs/2609.12244

机构:University of Victoria(维多利亚大学)

作者:Ruilin Wang, Xiaodai Dong

英文摘要:In sub-terahertz (sub-THz) communications, the coupling of ultra-wide bandwidth and severe phase noise (PN) impairments renders conventional joint channel and PN estimation highly complex and computationally prohibitive. To address this, we propose CRFCAN, a complex-valued residual FFT convolutional attention network designed for joint channel and PN estimation. Unlike existing deep learning schemes that rely on cascaded networks or hybrid frameworks combining neural networks with conventional iterative estimators, CRFCAN performs joint recovery in a truly end-to-end fashion through a physics-inspired cross-domain structure. Specifically, Fast Fourier Transform (FFT) and inverse FFT modules are embedded within residual groups to enable iterative feature interaction across the time and frequency domains, thereby capturing both frequency-selective fading and time-varying phase distortions. In addition, two dedicated residual blocks are introduced for complex feature extraction and multiplicative phase-distortion modeling, respectively. A physics-aware PN output tail with soft normalization is further employed to improve estimation stability while preserving the physical characteristics of the effective PN process. Simulation results demonstrate that CRFCAN significantly outperforms conventional algorithms and state-of-the-art deep learning models in terms of normalized mean square error (NMSE) and bit error rate (BER). Notably, CRFCAN achieves superior performance with single-shot, fixed-complexity inference and generalizes well to unseen PN models without fine-tuning, highlighting its robustness and practicality for sub-THz receivers.

76. Simulating Disengaged Students to Evaluate LLM-based Tutors

模拟不参与学习的学生以评估基于LLM的导师

AI 总结:提出DAS2协议,模拟五种学习者参与状态以评估AI导师,验证了模拟有效性并揭示导师表现随状态变化。

链接:https://arxiv.org/abs/2609.12331

作者:Xianghui Meng, Jionghao Lin

英文摘要:Simulated students generated by computational models provide a practical way to evaluate tutoring strategies and pedagogical approaches used by human and AI tutors. However, such simulations should account for disengaged behaviors, including gaming the system, wheel-spinning, and off-task behavior, because tutors may need different responses for different learner states. We present Disengagement-Aware Student Simulators (DAS2), a reproducible pre-deployment protocol that models five learner-engagement states: engaged, gaming, wheel-spinning, off-task, and mixed, and evaluates AI tutor performance across these states. Using ASSISTments09, two coders independently labeled 100 sampled tutoring sessions based on anonymized interaction-log summaries. They achieved 84% agreement (Cohen's kappa = 0.78), and among agreed cases, human consensus labels matched DAS2 rule-based labels in 81% of cases (kappa = 0.75). Conditioning simulations on intended learner states reduced the correctness-rate gap between simulated and authentic sessions from 0.54 to 0.20 for gaming and from 0.51 to 0.18 for wheel-spinning. Fine-tuned Qwen2.5-7B better matched authentic response-time distributions, while prompt-only GPT-4o generated more distinguishable learner states. Evaluation of five AI tutors from the Claude, Llama, Gemini, Qwen, and GPT families showed that relative rankings remained stable across learner states and interaction lengths, while absolute performance varied, revealing state-specific differences in tutor support. Human validation further showed that automated tutor evaluation does not fully align with human judgment. DAS2 provides a pre-deployment framework for evaluating how AI tutors respond to diverse learner-engagement states before deployment.

77. When Connected Does Not Mean Similar: Charting the Homophily Boundary of SNAP-KG for Streaming Entity Integration

当连接并不意味着相似:绘制SNAP-KG在流式实体集成中的同质性边界

AI 总结:本文探讨SNAP-KG在无同质性关系时聚类性能骤降,指出同质性假设是方法家族共性,并提出未来通过蒸馏异质性感知教师模型以扩展其适用性。

链接:https://arxiv.org/abs/2609.12356

机构:Rensselaer Polytechnic Institute(伦斯勒理工学院)

作者:Jui-Chien Lin, Oshani Seneviratne

英文摘要: SNAP-KG is a framework for assigning newly arriving entities to semantic communities in a growing knowledge graph (KG) using only their raw features, with no graph access and no retraining at inference time. It was evaluated on five multi-view benchmarks and a 2.4M-node OGB-WikiKG2 KG. In each of these datasets, at least one graph view is homophilous, meaning that connected nodes usually belong to the same class, and SNAP-KG performs well on all of them. This paper asks what happens outside that setting. We extend the evaluation to three heterophilous graphs (Texas, Wisconsin, Chameleon) and measure the edge homophily of every view. When no homophilous view is available, clustering quality drops sharply for both SNAP-KG and the transductive baselines used in its original evaluation. What decides this is the homophily of the relation, not the number of relations. Multi-view fusion still helps, but only when at least one homophilous relation provides a reliable foundation. The homophily assumption is therefore shared by the whole method family, not specific to SNAP-KG. We argue that heterophilous multi-view clustering is a separate research problem, outside the scope of this work. As future work, we outline how a heterophily-aware teacher could be distilled into SNAP-KG's projector to serve both homophilous and heterophilous KGs.

78. LatentVerse: A Framework for Understanding Shared and Modality-Specific Information in Multimodal Latent Representations

LatentVerse:理解多模态潜在表示中共享与模态特定信息的框架

AI 总结:LatentVerse是一个结合可视化平台与命令行界面的表示分析框架,通过分解共享与模态特定组件,实现多模态潜在表示的质量评估,提升可理解性与可复现性。

链接:https://arxiv.org/abs/2609.12364

机构:Broad Institute of MIT and Harvard(麻省理工学院和哈佛大学布罗德研究所); Massachusetts Institute of Technology(麻省理工学院); The Schmidt Center, Broad Institute of MIT and Harvard(麻省理工学院和哈佛大学布罗德研究所施密特中心)

作者:Majd Alafrange, Samuel Friedman, John Kitonyo, Sana Tonekaboni, Mahnaz Maddah

英文摘要:Latent embeddings have become a central data abstraction in modern machine learning, especially in biomedicine, where foundation models are increasingly used to encode multimodal data like clinical text, medical images, omics, and physiological signals. However, the utility and value of these representations depends on understanding their quality, structure, and the information they encode. Existing analysis workflows for evaluating representations remain fragmented across custom scripts, isolated metrics, and most importantly lack multimodal analysis, limiting accessibility and reproducibility. We present LatentVerse, a representation analysis resource that combines a web-based visual analytics platform for accessible, report-driven exploration with a command-line interface for scalable technical workflows. LatentVerse unifies diagnostics for various representation quality metrics and extends to multimodal settings by decomposing embeddings into shared and modality-specific components. We evaluate LatentVerse through controlled unimodal and multimodal simulations, discovery-oriented analyses on real biomedical embeddings, and a user study across diverse use cases. By supporting thorough and interpretable evaluation of latent spaces, LatentVerse makes foundation model representations more understandable in biomedical and data science applications.

79. Certified AI Triage of ICU Alarms

ICU警报的认证式AI分诊

AI 总结:本研究将ICU警报削减重构为三向分诊,提出认证式方法在95%置信度下控制真实警报静音比例,在VTaC基准上抑制74.8%虚假警报且仅静音1.5%真实警报,并揭示多重性校正对认证结果的代价。

链接:https://arxiv.org/abs/2609.12365

机构:Roshan AI; University of Arizona(亚利桑那大学)

作者:Mohammed Sameer Syed, Rozhin Yasaei

英文摘要:In the VTaC benchmark 71% of ventricular-tachycardia alarms are false, but silencing a real one can delay recognition of a dangerous arrhythmia. We reframe alarm reduction as three-way triage (retain, suppress, or defer) and bound the decision this analysis treats as harmful: among suppressed alarms, the fraction that were genuine stays below a user-set budget with 95% confidence, under i.i.d. event sampling. Alarms sharing a waveform record are dependent, so the clustered analysis is a sensitivity check. On the official split a 5% budget certifies in all three seeds, suppressing 74.8% of false alarms while silencing 1.5% of genuine ones, at AUROC 0.953 and Challenge Score 83.33, numerically comparable to the strongest of the eleven published systems. Our central finding measures what multiplicity costs: the correction charges for every candidate, so a finer grid can certify strictly less. Under held-out calibration the 885-cell grid we declared certifies 1 of 15 fold-runs, while choosing the grid on a separate selection partition certifies 8. We project the calibration volume each budget needs, making an uncertifiable budget a design parameter. Finally, adding a learned reliability dimension to the policy grid did not sharpen the certified frontier.

80. Beyond the Query: Do Retrieval Signals Improve Adaptive Multimodal RAG Routing?

超越查询本身:检索信号能否改进自适应多模态RAG路由?

AI 总结:本研究在文档、音频和视频RAG中对比仅查询与查询+检索路由器,发现检索信号未带来可靠的路由改进,强调其增量价值需通过匹配对照实验验证。

链接:https://arxiv.org/abs/2609.12437

机构:Kennesaw State University(肯尼索州立大学)

作者:Qiaomu Li, Qiuyuan Zhang, Nong Ming

英文摘要:Adaptive RAG often uses retrieval-time signals to decide whether another retrieval, reranking, or multimodal step should run. We ask whether these signals add routing value once the query itself is already known. Across document, audio, and video RAG, we compare matched query-only and query+retrieval routers while holding the optional actions, router family, training procedure, and evaluation fixed. On the held-out final evaluation, adding the tested retrieval signals does not produce a reliable routing improvement over the query-only baseline. Some retrieval signals are associated with whether a later step will help, but that predictability does not consistently lead to bet- ter RUN/SKIP decisions. The main lesson is therefore methodological: retrieval-state features should not be credited with routing value unless they improve over a matched query-only control. Our results do not show that routing or retrieval state is generally useless; they show that the incremental value of retrieval signals must be demonstrated rather than assumed.

81. 3D Digital Twin Visualization of Multiclass GRF-Based Gait Disorder Classification

基于多类GRF的步态障碍分类的3D数字孪生可视化

AI 总结:提出一个集成框架,利用双侧GRF和COP信号分类健康及多种肌肉骨骼损伤步态,结合ε-LRP可解释性和Blender 3D可视化,实现高准确率与透明性。

链接:https://arxiv.org/abs/2609.12442

机构:Yonsei University(延世大学)

作者:Nayoung Son, Minwoo Shin

英文摘要:Automated gait analysis requires accurate classification and interpretable outputs. We propose an integrated framework for classifying healthy gait and multiple musculoskeletal impairment groups using bilateral ground reaction force (GRF) and center-of-pressure (COP) signals. The signals were normalized over the stance phase and standardized using training-set statistics. The model achieved a validation accuracy of 99.00\% and a test accuracy of 90.07\% under a session-level split. Class-specific $\epsilon$-LRP identified positive and negative contributions across both sides, multiple signal components, and different stance phases. Separately, the processed GRF signals and model predictions were synchronized within a Blender-based 3D visualization, enabling sample-level inspection of gait trials and classification results. The proposed framework integrates classification, explainability, and 3D visualization to improve model transparency. The source code is available in the following repository: this https URL

82. SAGE-Loop: Reliable Closed-Loop LLM-Driven AutoML with Trial-and-Correction and Adaptive Ensembling

SAGE-Loop:具有试错修正与自适应集成的可靠闭环LLM驱动AutoML

AI 总结:针对LLM驱动AutoML缺乏闭环修正与静态集成的问题,提出SAGE-Loop框架,通过多轮试错修复与自适应集成,在20个数据集上提升分类、回归和聚类任务的性能与稳定性。

链接:https://arxiv.org/abs/2609.12455

作者:Junquan Gu, Shibo Cui, Xiangfeng Luo, Hang Yu

英文摘要:Automated machine learning (AutoML) is reshaping data-driven science and industrial practice, and as large language models are introduced into AutoML, pipeline reliability becomes as important as automation efficiency. However, existing AutoML still struggles to realize instant feedback and adaptive optimization during execution, so once a run drifts into a suboptimal or failed state, it lacks a process-level correction mechanism. The fundamental pathology lies in its one-way pipeline: intermediate failures are typically terminated or bypassed, while fixed paradigms often strengthen model generation but leave ensemble decisions static, weakening both execution reliability and the controlled use of structural diversity. This indicates that LLM-driven AutoML needs a closed-loop ability for trial-correction-improvement together with evidence-based use of model diversity. To this end, we propose SAGE-Loop, a reliable closed-loop, self-adaptive, LLM-driven AutoML framework that performs multi-round generation and validation for trial-and-repair, and adaptively selects ensemble strategies in both supervised and unsupervised tasks, thereby unifying how to generate with how to use models. Across 20 public datasets, SAGE-Loop consistently improves performance and stability on classification, regression, and clustering tasks. Additional results further show its ability to recover from execution failures and maintain robust pipeline behavior.

83. Temporal Recurrence Favors Fewer Layers

时间递归有利于更少的层数

AI 总结:该研究将递归模型中的深度需求视为计算分配问题,发现在流式任务中,时间递归可将最优计算分配转向更少的层数,且性能相当或更好。

链接:https://arxiv.org/abs/2609.12531

机构:Mila – Québec AI Institute(米拉-魁北克人工智能研究所); Université de Montréal(蒙特利尔大学); Sakana AI

作者:Ivan Anokhin, Johan Obando-Ceron, Irina Rish, Sebastian Risi

英文摘要:In streaming tasks, recurrent models can carry latent computation across time, allowing each update to build on representations produced earlier. This raises a basic question: once temporal recurrence provides sequential computation across steps, how much depth is still needed within each step? Prior work has shown that recurrence can make shallow models competitive. We instead study this question as a compute-allocation problem, varying within-step depth, expert width, and the number of parallel experts per layer across several compute budgets. For each budget, we compare the best observed recurrent and non-recurrent allocations and the performance they achieve under approximately matched per-step computation. Across Sokoban and autoregressive FineWeb language modeling, we find that temporal recurrence shifts the best observed compute allocation toward substantially fewer layers, with comparable or better performance.

84. TokenMapper: A Step Toward Interoperable Speech Token Translation

TokenMapper:迈向可互操作语音标记翻译的一步

AI 总结:TokenMapper提出方向感知的离散域标记翻译框架,实现异构语音分词器间直接映射,降低延迟并保持性能,迈向跨模型语音标记互操作。

链接:https://arxiv.org/abs/2609.12563

机构:Ben-Gurion University of the Negev(内盖夫本-古里安大学)

作者:Tal Kozakov, Tal Rosenwein, Eliya Nachmani

英文摘要: Neural audio codecs discretize speech into token sequences, but the resulting token spaces differ in vocabulary and codebook structure, preventing direct communication across models. This limitation affects applications such as conversational voice agents and speech to speech translation systems where multiple speech models must interact. As a result, transferring information between speech systems typically requires decoding to waveform audio and re-encoding with a second tokenizer, increasing latency and introducing potential information loss. To address these limitations, we present TokenMapper, a direction aware framework for direct token to token translation between heterogeneous speech tokenizers in the discrete domain. TokenMapper supports structurally mismatched token spaces, including mappings between single codebook and multi codebook representations, under a shared effective token rate. Experiments on GLM-4-Voice, MiMi and DualCodec show consistent cross model performance. Specifically, translation WER approaches native reconstructions within 2.5-6.8% absolute WER, human MOS for TokenMapper outputs ranges from 2.29 to 4.39, following the same direction level trends as UTMOS and end to end latency is reduced by 4.8-94.5% relative to waveform bridging, reaching up to 972 ms per utterance. These results provide a practical step toward cross model speech token interoperability without intermediate waveform reconstruction.

85. SCOPE-OPSD: Fisher-Conditioned Privileged Subspaces for On-Policy Self-Distillation

SCOPE-OPSD:用于在线策略自蒸馏的Fisher条件特权子空间

AI 总结:SCOPE-OPSD通过将特权教师-学生残差投影到Fisher条件秩64子空间,为在线策略自蒸馏增加第二监督通道,在Qwen3模型上多数检查点超越纯OPSD和随机对照,验证了数据相关方向的有效性。

链接:https://arxiv.org/abs/2609.12579

机构:Chongqing Ant Consumer Finance Co., Ltd.(重庆蚂蚁消费金融有限公司); Alibaba Cloud Computing Co., Ltd.(阿里云计算有限公司); Chongqing University of Posts and Telecommunications(重庆邮电大学)

作者:Yunmeng Chen (1), Kunyu Wang (2), Peihan Li (1), Yi Wang (1), Shuyin Xia (3), Yi Liu (1), Xinyong Cheng (2), Dehui Wang (2), Xiangyong Zhai (2), Yanxing Liu (1), Song Liu (1) ((1) Chongqing Ant Consumer Finance Co., Ltd., (2) Alibaba Cloud Computing Co., Ltd., (3) Chongqing University of Posts and Telecommunications)

英文摘要:On-policy self-distillation (OPSD) scores student-generated prefixes with a solution-conditioned self-teacher, yet transfers supervision only through next-token probabilities. We ask whether the aligned final-layer discrepancy offers a useful second channel, and how to test that channel without confusing its geometry with auxiliary strength. SCOPE-OPSD projects the privileged teacher-student residual onto a frozen rank-64 factor estimated from residual covariance and language-model-head Fisher sensitivity. It reuses the forwards already required by OPSD and adds neither rollouts nor inference-time modules. A matched Random control preserves the structured factor's rank and nonzero spectrum and uses per-arm gradient-RMS calibration, isolating the effect of the data-dependent orientation. Across the complete 25/50/75/100-step trajectories for Qwen3-1.7B, 4B, and 8B, Structured is never below Pure OPSD, with strict gains in 11 of the 12 model-checkpoint combinations and an exact tie at 4B step 25. Structured also exceeds matched Random in 10 of the 12 combinations. At step 75 on Qwen3-1.7B, Structured exceeds matched Random by 1.39 Macro Avg@12 points in each of two independent training reruns. A cross-fitted diagnostic also shows 4.40 times greater held-out privileged-gap capture than the matched random orientation. The results support a compact, Fisher-conditioned privileged subspace for short-budget OPSD.

86. Poisson-Corrector Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling

Poisson-Corrector 复杂度界用于 Moreau--Yosida 非调整 Langevin 采样

AI 总结:研究MYULA采样算法,证明其与Moreau平滑目标间的Wasserstein距离误差界,结合离散Poisson校正子等方法,实现O(ε^{-4/3})次迭代达到精度ε。

链接:https://arxiv.org/abs/2609.12594

机构:School of Mathematical Sciences, Peking University(北京大学数学科学学院)

作者:Yuchen Xin, Zhihua Zhang

英文摘要:We study the classical Moreau--Yosida unadjusted Langevin algorithm (MYULA) for $\pi(\,\mathrm{d} x)\propto e^{-f(x)-g(x)}\,\mathrm{d} x$, where $f\in C^2(\mathbb{R}^d)$ is $m$-strongly convex with $L_f$-Lipschitz gradient and $g:\mathbb{R}^d\to\mathbb{R}$ is convex and globally $G$-Lipschitz. For the Moreau-smoothed target $\pi_\lambda$ and the MYULA invariant law $\widehat\pi_{\lambda,h}$, we prove \[ \sqrt m\,W_2(\pi_\lambda,\widehat\pi_{\lambda,h}) =O(h)+\widetilde O(h^{3/4}) \] under $0

87. SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning

SIMS:基于尺度不变的 Merit 函数标量化方法用于多任务学习

AI 总结:针对多任务学习中目标尺度差异导致优化失衡的问题,提出尺度不变的 merit 函数标量化方法(SIMS),通过对数变换实现尺度不变性,保持弱 Pareto 最优性,并在多任务基准上取得最先进性能。

链接:https://arxiv.org/abs/2609.12599

机构:Southern University of Science and Technology(南方科技大学); City University of Hong Kong(香港城市大学); Shanxi University(山西大学)

作者:Zebin Chen, Fei Xing, Yang Chen, Hua Liu, Andy HF Chow, Yuhua Qian, Yu Zhang

英文摘要:Multi-task learning (MTL) requires navigating unavoidable trade-offs among competing objectives. This paradigm is frequently formulated as multi-objective optimization (MOO), where the scalarization is favored to reduce an MOO problem to a single objective. We empirically find that existing merit-function-based scalarization approaches are sensitive to the relative scales of different objectives in practical MTL, where task losses commonly differ by orders of magnitude. The optimization process often favors objectives with larger scales even though the underlying Pareto optimal solutions remains invariant to rescaling (i.e., multiplying an objective by a positive constant). To address this issue, we propose Scale-Invariant Merit-function-based Scalarization (SIMS) for MTL. Specifically, SIMS adopts a transformation-induced merit function to convert the MOO problem of MTL to a single objective that renders optimization invariant to the magnitudes of losses. Theoretically, we prove that the requirement for scale invariance uniquely determines this transformation to be logarithmic. We further show that this general transformation-induced merit function preserves weak Pareto optimality and admits a smooth surrogate with controllable approximation error. Extensive experiments on representative multi-task benchmarks demonstrate that SIMS consistently outperforms existing scalarization methods and achieves state-of-the-art performance.

88. ProactiveBench: Can Streaming Video Models Really Interact Like Humans?

ProactiveBench:流式视频模型真的能像人类一样交互吗?

AI 总结:针对流式视频模型缺乏主动交互评估的问题,提出ProactiveBench基准,以一秒间隔无提示评估模型,发现多数系统过早响应多于遗漏,揭示时间决策差距。

链接:https://arxiv.org/abs/2609.12658

机构:Beihang University(北京航空航天大学); Alibaba(阿里巴巴)

作者:Kaixuan Du, Xin Wan, YuKun Wang, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, Ni Li

英文摘要:Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context. Existing evaluations are predominantly reactive: they query a model at a selected timestamp and therefore do not assess when it should respond. Proactive interaction instead requires monitoring a standing request, responding within an appropriate interval after the target event, and otherwise remaining silent. We introduce ProactiveBench, which evaluates models at one-second stream intervals without an explicit response cue. Its six subtasks vary trigger ambiguity and timing tolerance. Event Sensitivity geometrically combines response and silence rates on the same recording; four window-based subtasks distinguish early, in-window, and missed responses; and Duplicate Counting penalizes omissions and repetitions. Premature responses outnumber missed responses for four of the six evaluated systems, revealing a substantial gap in the temporal decision-making required for human-like interaction.

89. Write on Paper and Get the Online Digital Trace:\newline A New Era for Handwriting

在纸上书写,获取在线数字轨迹:手写的新时代

AI 总结:本研究提出结合数字笔、AI算法与自适应技术,实现普通纸上手写轨迹的实时数字化重建,无需外部参考系统,推动手写数字化的新纪元。

链接:https://arxiv.org/abs/2609.12702

机构:IRISA, Université de Rennes, INSA Rennes(IRISA,雷恩大学,INSA雷恩); IRISA, Université Rennes 2(IRISA,雷恩第二大学); Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院); STABILO International GmbH(STABILO国际有限公司); LTU, Machine Learning Group, Luleå University of Technology(吕勒奥理工大学机器学习组)

作者:Florent Imbert, Yann Soullard, Eric Anquetil, Tanja Harbaum, Alexey Serdyuk, Fabian Kress, Tim Hamann, Peter Kampf

英文摘要:Capturing the digital trace of handwriting usually requires a specific stylus and a compatible substrate, be it a capacitive touchscreen, an ElectroMagnetic Resonance (EMR) tablet as used in Wacom systems or special paper. While writing on regular paper offers rich haptics, no latency and is well known for improving information retention, no low-cost and widely accepted, effective solution exists to digitize such a pen trace. The challenge is to accurately track the pen's trajectory without an external reference system while allowing unrestricted freedom of pen movement across a surface. We propose an innovative solution that combines a digital pen, advanced artificial intelligence algorithms, and adaptive AI techniques to reconstruct the digital trace of handwriting. Our approach integrates hardware development, focusing on a sensor-equipped pen, with software innovations to optimize trajectory reconstruction and processing in real time using an embedded AI. This work aims to advance the state-of-the-art in automated trace reconstruction of handwriting, enabling a seamless connection between traditional handwriting on paper and capturing the trace digitally.

90. Physics-Guided Synthetic High-Frequency Ultrasound Generation for Skin Layer Segmentation

物理引导的高频超声合成生成用于皮肤层分割

AI 总结:针对高频超声皮肤层分割标注稀缺问题,提出物理引导合成HFUS生成框架,通过k-Wave模拟生成配对数据用于预训练,在真实数据上微调后性能与真实训练相当,多数架构Dice/IoU提升。

链接:https://arxiv.org/abs/2609.12735

机构:Yonsei University(延世大学)

作者:Junkyung ju, Kyungho Yoon, Minwoo Shin

英文摘要:High-frequency ultrasound (HFUS) enables noninvasive visualization of superficial skin structures, but automated skin-layer analysis is limited by the scarcity of densely annotated data. Existing real HFUS datasets commonly provide annotations for superficial targets such as the epidermis and subepidermal low-echogenic band (SLEB), while dense labels for deeper structures such as dermis, subcutaneous tissue, fascia, and muscle are rarely available. We propose a physics-guided synthetic HFUS generation framework for skin layer segmentation. The framework constructs multilayer acoustic skin phantoms, assigns layer dependent acoustic properties, and uses k-Wave simulation to generate paired synthetic HFUS images, dense layer masks, and simulation metadata. To evaluate whether the generated data provide transferable supervision, we use it for downstream segmentation pretraining and fine-tune the models on real Mendeley HFUS data. Synthetic pretraining followed by real fine-tuning achieved real-domain performance comparable to real-only training and improved mean Dice/IoU in three of four evaluated trainable architectures. These results suggest that physics-guided synthetic HFUS images contain transferable anatomical and textural cues for real-domain skin layer segmentation, although further reduction of the synthetic-real appearance gap is needed to enable greater gains. The code and data are available at: this https URL.

91. Optimizing for the decision not the prediction: an exploration of Smooth Net Benefit as a training objective

优化决策而非预测:平滑净收益作为训练目标的探索

AI 总结:本研究探索平滑净收益作为训练目标,发现其在逻辑回归中略有收益,但对灵活模型无益,不能替代负对数似然训练。

链接:https://arxiv.org/abs/2609.12752

作者:Koen M.F. Gorgels, Lasai Barreñada, Maarten van Smeden, Ben Van Calster, Ewout W. Steyerberg, Wouter A.C. van Amsterdam

英文摘要: Objective Prediction models are commonly trained using objectives such as Bernoulli negative log-likelihood (NLL), although downstream clinical decisions may depend on specific risk thresholds. We introduce Smooth Net Benefit ($\sigma$NB), a differentiable approximation of Net Benefit designed to align model training with threshold-specific clinical utility. Materials and Methods We evaluated $\sigma$NB as a training objective for logistic regression, generalized additive models (GAMs), and XGBoost with three Hessian implementations. Experiments used the Framingham cardiovascular risk dataset and 44 TabZilla datasets comprising 72 dataset-threshold combinations. Results $\sigma$NB training did not consistently improve Net Benefit in Framingham. Across the TabZilla benchmark, mean standardized Net Benefit for logistic regression increased from 0.5669 with NLL to 0.5765 with $\sigma$NB (mean difference 0.0096, 95% CI -0.0001 to 0.0193). For GAMs, mean standardized Net Benefit decreased from 0.5921 to 0.5625 (mean difference -0.0296, 95% CI -0.0721 to 0.0129). For XGBoost, NLL achieved 0.6745 compared with 0.6723--0.6735 across $\sigma$NB implementations. In logistic regression, $\sigma$NB gains were positively associated with the performance advantage of XGBoost over NLL-trained logistic regression. Discussion The effect of $\sigma$NB was context dependent, with modest gains concentrated in logistic regression and little benefit for more flexible model classes. This suggests that decision-focused optimization may be most useful when limited model flexibility leaves greater scope for improvement. Conclusion Our results do not support $\sigma$NB as a general replacement for NLL training, but support further investigation of decision-focused objectives in settings where conventional likelihood-based training may not adequately capture decision-relevant structure.

92. GenOR-Twin: A Semantic Middleware for Integrating Operational Discourse with Mathematical Optimization

GenOR-Twin:一种将运营话语与数学优化集成的语义中间件

AI 总结:GenOR-Twin是一个神经符号语义中间件,利用大语言模型作为语义翻译器,通过动态约束注入和自适应决策策略,实现运营日志与数学优化的双向耦合,从而构建适应现实不确定性的弹性数字孪生系统。

链接:https://arxiv.org/abs/2609.12863

作者:Rahimeh Neamatian Monemi, Shahin Gelareh, Lubin Cui, Nelson Maculan

英文摘要:We introduce GenOR-Twin, a neuro-symbolic framework that bridges the translation gap between unstructured operational logs and rigorous mathematical optimization. Our architecture uniquely positions Large Language Models as semantic translators rather than direct solvers, ensuring that the system retains the feasibility guarantees of exact combinatorial methods. { \color{red}We design a dynamic constraint injection mechanism (the runtime translation of qualitative disruption events into formal mathematical constraints) that allows the system to structurally modify the optimization problem's feasibility region in real-time based on qualitative human inputs. The resulting bidirectional coupling---where operational observations update the virtual model state and optimized decisions are reflected back into the Knowledge Graph---satisfies the synchronization requirement of a proper Digital Twin. The framework features an adaptive decision policy} that automatically selects between low-complexity schedule repair and full re-optimization by analyzing the available system slack. Finally, we demonstrate the generalization of this approach across six distinct optimization domains, {\color{red}turning static models into resilient systems that adapt to the operational uncertainty and variability of real-world environments.}

93. What an odour descriptor corpus can and cannot measure: valence, attenuation, and the ceiling of the public record

气味描述符语料库能测量与不能测量的内容:效价、衰减与公共记录的天花板

AI 总结:本研究审计四个气味描述符语料库,发现其不一致性无法通过汇总修复,且模型性能存在天花板;缺失的效价需直接测量,单次愉悦度评分显著提升预测,并发布交叉映射与定理。

链接:https://arxiv.org/abs/2609.12875

机构:Tesseract Academy(泰瑟拉克学院)

作者:Stylianos Kampakis, Fabio Rovai

英文摘要:Machine olfaction trains on pooled public descriptor corpora, but whether a shared descriptor word measures the same thing across corpora has not been tested, nor has the ceiling of what any of them can measure. We audit four corpora from Pyrfume. Conditioning on the molecule makes McNemar's test the exact conditional test of the corpus effect. Corpora disagree heterogeneously across descriptors ($I^2 = 80\%$) and non-uniformly with labelling breadth ($z = 17.2$), so no single offset repairs pooling. Median tetrachoric agreement is 0.795 against median $\kappa$ of 0.413: sources largely concur on which molecules deserve a word and differ on how readily they apply it. Of 109 descriptors with an estimable effect, 36 show large differential functioning on the ETS scale. Against a human panel's reliability, Morgan fingerprints with the full RDKit descriptor block reach 32.9\% of achievable; adding every label from two merged corpora reaches 33.9\%. The gap does not close with model capacity, encoding choice, more molecules, or more words. The missing variance is valence. One pleasantness rating per molecule reaches 54.6\% of achievable (57.1\% on an independent older instrument). Valence recovered from descriptors ($\rho = 0.457$) yields only 16.6\%, so it must be measured. Five raters exceed structure plus the full descriptor record; fifteen to twenty saturate. We release a descriptor crosswalk and twenty machine-checked theorems.

94. Physical-State-Guided Diffusion Sampling for Full-Waveform Inversion

物理状态引导的扩散采样用于全波形反演

AI 总结:针对全波形反演中初始化依赖和物理引导不可靠的问题,提出物理状态引导的扩散采样(PSG),通过高斯桥耦合物理速度与扩散先验,在多个数据集上优于基线并支持大模型反演。

链接:https://arxiv.org/abs/2609.12899

机构:School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院); CMA-Shanghai, Shanghai Jiao Tong University(上海交通大学CMA-上海); School of Mathematics and Statistics, Lanzhou University(兰州大学数学与统计学院)

作者:Chen Min, Haowen Jiang, Zheng Ma, Xiongbin Yan

英文摘要: Full waveform inversion (FWI) estimates subsurface velocity from seismic recordings, but its ill-posedness and nonlinearity make accurate reconstruction strongly dependent on initialization and prior information. Diffusion posterior sampling provides a learned geological prior, yet directly coupling its denoiser to the nonlinear wave solver can yield unreliable physical guidance. We propose Physical-State-Guided Diffusion Sampling (PSG), which couples a persistent physical velocity to the diffusion prior through a Gaussian bridge. The physical state is refined by waveform fitting regularized by the denoised velocity, and in turn guides the reverse diffusion process. This formulation separates the wave-equation and denoiser gradients while preserving conventional FWI initialization and accumulated optimization history. On four OpenFWI families, PSG's terminal denoised estimates outperform classical and diffusion-based baselines under clean and missing-trace acquisitions and maintain strong structural recovery under measurement noise. Repeated stochastic runs preserve the dominant geological structures, with ensemble variability concentrated near geological interfaces and positively associated with local inversion error. A frozen OpenFWI-trained prior further supports inversion of the larger Marmousi, Overthrust, and BP2004 Salt models, recovering complex geological structures without retraining.

95. Information-Induced Training Geometry: Exact Reduction, Canonical Completion, and Structured Expressivity

信息诱导的训练几何:精确约化、典范完备化与结构化表达性

AI 总结:本研究在仿射不变黎曼几何下,证明信息诱导的压缩映射存在唯一完备化,实现精确变分约化,并刻画了正定锥上的分层结构与结构化表达性。

链接:https://arxiv.org/abs/2609.12991

机构:Xidian University(西安电子科技大学)

作者:Zavier Li

英文摘要:Training data constrains optimizer geometry through the covectors visible to a declared information channel. We study how such partial information determines a full positive cometric relative to a reference and which degrees of freedom remain unidentified. Our central result resolves full-column-rank positive-definite compression under affine-invariant Riemannian geometry. The compression map is a split-Hadamard metric submetry and admits an explicit unique completion that is the affine-invariant nearest full geometry realizing a visible target and yields exact full-to-visible variational reduction. When the channel moves, the completions form a gauge-invariant rank stratification of the positive-definite cone. Its closed-form pullback pair metric separates visible-metric motion from subspace rotation through a reference-mismatch weight, yields an explicit positive-semidefinite multi-direction Gram matrix, and exposes the precise singularity of reference-valued modes. The mechanism is explained by a metric theorem equating ball submetry, attained fiber distance, and lossless reduction of every monotone radial visible decision problem. A smooth split-Hadamard theorem supplies coherent information sheets, proximal commutation, and solution-wise gradient-flow lifting. The positive-definite realization also gives closed-form prior-data shrinkage. Diagonal and block optimizer families reduce to relative-interior conic image tests with valid facial certificates, while deterministic and finite-sample bounds quantify recovery of the visible geometry and its subspace. Together these results characterize exact reduction, reference-dependent completion, and structured expressivity for the stated finite-dimensional affine-invariant model.

96. Dimension-Corrected Hitting Times for Heavy-Tailed Spectral Emergence in Neural Optimizer Dynamics

神经优化器动力学中重尾谱涌现的维度校正击中时间

AI 总结:本研究将神经网络权重谱重尾涌现建模为右删失击中时间问题,提出维度校正谱隙定律,并通过理论与实验验证,为隐式自正则化提供可复现的定量描述。

链接:https://arxiv.org/abs/2609.12994

机构:Stanford University(斯坦福大学)

作者:Zongmin Liu

英文摘要:Heavy-tailed empirical spectral densities of neural-network weight matrices are widely used as diagnostics of implicit self-regularization, but the step complexity of heavy-tail emergence remains poorly understood. We formulate spectral heavy-tail formation as a right-censored hitting-time problem: a run that does not reach a heavy-tail diagnostic within the observation horizon is treated as censored rather than discarded. In controlled full-batch teacher--student dynamics, we find that the first-step spike--bulk gap alone does not explain onset time. Instead, finite-onset regression supports a dimension-corrected spectral-gap law, (\tau_{\mathrm{HT}}\approx C\Delta_1^{-\gamma}d^\rho), with (R^2=0.683), (\gamma=0.626), and (\rho=0.772) across 330 completed runs. Right-censored lognormal accelerated-failure-time models further favor the dimension-corrected model over a gap-only model, improving AIC from 706.62 to 628.70. Theoretically, we prove that exact early loss dynamics in linear networks do not determine factor spectral tails, that Adam recurrences alone do not imply spectral redistribution, and that projected singular-basis spreading implies contraction of a spectral-tail potential and hence a dimension-corrected hitting-time bound. Empirically, projected-kernel profiles support the sufficient spreading mechanism, Adam and AdamW agree under tested grids, GD and signGD do not reach onset in the same regimes, and real pretrained Qwen2.5-0.5B and Pythia-70M transformer weights show non-Gaussian spectral-tail structure relative to matched Gaussian nulls. The result is a reproducible spectral hitting-time law with rigorous conditional theory, not a claim that Adam necessarily generates heavy tails from first principles.

97. A Full Adam Theorem for Spectral Heavy-Tail Onset

谱重尾起始的完整Adam定理

AI 总结:该论文在封闭高斯Stein-Hermite教师-学生模型中证明了谱重尾起始的完整Adam定理,给出了精确的命中时间定律,并证明了更强定理的不可能性及线性网络动力学的局限性。

链接:https://arxiv.org/abs/2609.12996

机构:Stanford University(斯坦福大学)

作者:Zongmin Liu

英文摘要:We prove a full Adam theorem for spectral heavy-tail onset in a closed Gaussian Stein-Hermite teacher-student state-evolution model. The theorem begins with the actual full-batch Adam recurrences, derives the population gradient by Stein-Hermite calculus, proves finite-width covariance concentration, converts multi-step Adam momentum into an exact non-centered Gaussian sign kernel, controls the diagonal Adam denominator by a basis-homogenization theorem, derives a regularly varying projected update response from a Hermite edge-transfer theorem, pushes the response through the exact Gram update, and proves approximate-target KL contraction with matching upper and lower hitting bounds. The final law is (\tau_\varepsilon=\Theta(\Delta_1^{-\gamma}d^\rho\log(\Psi_0/\varepsilon))), where (\Delta_1) is the first spike-bulk spectral gap. The result is full in the following precise sense: every step from Adam's momentum and denominator to the spectral hitting law is formalized inside the closed state-evolution model. We also prove that a stronger arbitrary-gradient Adam theorem is impossible, and that exact two-step linear-network loss dynamics do not identify factor spectra or heavy-tail hitting times.

98. Attention Quantization for Tabular Foundation Models

表格基础模型的注意力量化

AI 总结:针对表格基础模型,提出FP8注意力量化策略,强调训练与测试行量化误差对齐,Triton内核实现1.7倍加速且精度无显著损失。

链接:https://arxiv.org/abs/2609.13031

机构:Prior Labs

作者:Jonas M. Kübler, Benjamin Jäger, Klemens Flöge, Noah Hollmann, Frank Hutter

英文摘要:With the recent rise and adoption of tabular foundation models, optimizing their inference performance becomes an emerging field for efficiency research. While the models are architecturally similar to transformer-based large language models (LLMs), the size and serving patterns differ significantly. We show that the focus should be on the attention calculation and less on weight or KV cache quantization, which are more popular in LLMs. We develop a quantization strategy for queries, keys, and values to FP8 and use explicit FP8 matrix multiplication instructions to speed up the attention calculation. We find that it is crucial to align the quantization error in the test rows with the quantization error in the training rows, as otherwise the accuracy drops drastically. Our Triton kernel achieves a speedup up to 1.7x over regular 16-bit kernels, and we show that on TabPFN-v3 and TabICLv2 there is no relevant accuracy loss across TabArena and BeyondArena.

99. DynSHAP: Towards Explainable Dynamic Survival Analysis

DynSHAP:面向可解释动态生存分析

AI 总结:DynSHAP提出面向动态生存分析的SHAP解释框架,通过时间-特征对博弈和条件采样处理纵向不规则数据,在合成及真实临床数据上实现忠实归因。

链接:https://arxiv.org/abs/2609.13042

机构:University of Cambridge(剑桥大学)

作者:Nastasya Anokhina, Jonas Jürß, Pietro Liò

英文摘要:Deep learning models for dynamic survival analysis (DSA) achieve strong predictive performance by incorporating longitudinal patient data, but their black box nature limits clinical trust and adoption. Existing explainability methods cannot handle longitudinal, irregular inputs and functional survival outputs simultaneously, which limits their usability in DSA. We propose DynSHAP, a SHAP framework suited specifically for dynamic survival analysis. It extends common marginal SHAP estimators to this setting by treating time--feature pairs as players in the Shapley game. We further introduce Temporal DynSHAP, which learns linear dependencies in features over time and uses conditional sampling to address them in explanations. When applied to synthetic data with known ground-truth attributions, Temporal DynSHAP recovers temporally dependent features more accurately than marginal estimators for a given state-of-the-art model. Applied to two real-world clinical datasets and two DSA architectures, DynSHAP produces attributions faithful to model learning, allowing medical experts to see which patient information drove the prediction and when.

Python社区是高质量的Python/Django开发社区
本文地址:http://www.python88.com/topic/201045