Py学习  »  机器学习算法

机器学习学术速递[9.21]

arXiv每日学术速递 • 1 周前 • 64 次点击  

2026-09-21 | CS.LG机器学习 | 共 83 篇

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 深度学习架构与训练方法 5 篇

2. 表示学习、自监督与对比学习 1 篇

3. 强化学习与序列决策 10 篇

4. 生成模型与概率建模 6 篇

5. 优化、泛化与理论分析 4 篇

6. 高效学习、压缩与部署 2 篇

7. 联邦学习、隐私与安全 2 篇

8. 鲁棒性、不确定性与可信学习 2 篇

9. 图学习与结构化数据 1 篇

10. 迁移、元学习与持续学习 3 篇

11. 数据集、基准与评测 4 篇

12. 机器学习应用 6 篇

13. 其他/综合机器学习 37 篇

1. 深度学习架构与训练方法 | 5 篇

1. FairLMs: A Turnkey Library for Fairness in Language Models

FairLMs:一个用于语言模型公平性的即用型库

AI 总结:FairLMs是一个Python库,通过声明模型能力与输入要求,整合偏见衡量、缓解与证据检查,提供33个指标、14个缓解组件及诊断工具,支持多种架构与API,便于方法比较与工作流扩展。

链接:https://arxiv.org/abs/2609.21296

机构:Indiana University(印第安纳大学); Carnegie Mellon University(卡内基梅隆大学); Florida International University(佛罗里达国际大学)

作者:Jiale Zhang, Michael Larionov, Zichong Wang, Zhipeng Yin, Wenbin Zhang

英文摘要:Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing tools offer complementary functionality through different interfaces, so combining them requires reconciling model interfaces, evidence formats, access constraints, and result types before applicability can be checked or methods compared. We introduce \textbf{FairLMs}, a Python library that connects these activities through explicit declarations of model capabilities and input requirements. It provides 33 intrinsic and extrinsic metrics, 14 mitigation components spanning four intervention categories, 14 dataset and scoring-instrument diagnostics, adapters for the three Transformer architectures and supported hosted completion APIs, and benchmark loaders. Declarations are checked before execution and results carry the configuration under which they were obtained, so that compatible components can be combined, methods compared under a common protocol, and workflows extended to new models and datasets. The source code is available at: this https URL.

2. IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

IntBMoE:将块级条件引入专家组合以实现全参与的混合专家模型

AI 总结:IntBMoE通过块级条件与稀疏执行解耦参与度、执行度和物化度,实现全参与MoE,在图像分类、语言建模和推荐任务上超越基线,并已部署于高德地图推荐系统。

链接:https://arxiv.org/abs/2609.21346

机构:DreamX, Alibaba Group(阿里巴巴集团 DreamX)

作者:Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu

英文摘要:Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many experts contribute knowledge to its output, execution is how many are actually computed (compute cost), and materialization is how many expert-sized parameter sets must be built and stored (memory cost). Sparse routing keeps execution and materialization low, but shrinks participation: for each token, only a few experts contribute. Dense output-mixing restores full participation, but its execution grows with the number of experts. Parameter-merging keeps execution at one expert, but its materialization grows with the number of routing decisions. We propose IntBMoE, a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution. Its blocks come from a small learned codebook, one per entry. At each internal layer, a lightweight hypernetwork merges all expert bases in that layer's pool into one composed expert. Participation is full, because every composed expert draws on the entire pool. Execution stays sparse, because a router sends each token to only a few blocks. Materialization is bounded, because the codebook, not the input, fixes how many blocks exist. Dual-Path Residual Gating (DPRG) further couples two independently composed paths through multiplicative gating. Experiments on image classification show consistent gains over representative sparse and dense MoE baselines. Additional experiments on language modeling and sequential recommendation validate its generalization beyond vision. IntBMoE is fully deployed in AMap's generative recommendation system, serving hundreds of millions of users under a 60ms latency budget, with a 2.4% relative UVCTR gain in online A/B testing. Our code is available at this https URL.

3. Efficient Architecture Search under Leave-One-Subject-Out Evaluation

留一受试者评估下的高效架构搜索

AI 总结:本文提出PainNAS,一种基于块且防泄漏的神经架构搜索方法,在留一受试者评估中通过共享搜索将计算复杂度从O(N^2)降至O(B),在BioVid数据集上以更少参数和FLOPs保持相当准确率。

链接:https://arxiv.org/abs/2609.21457

机构:IU International University of Applied Sciences(IU国际应用科学大学); Ulm University(乌尔姆大学)

作者:Heinke Hihn, Friedhelm Schwenker

英文摘要: Deep neural architectures are widely used for signal processing in automated pain assessment systems. However, architecture design has remained largely a manual task despite the potential efficiency benefits of Neural Architecture Search (NAS). Embedding NAS in a Leave-One-Subject-Out (LOSO) evaluation is computationally demanding because a fully nested implementation requires $N$ independent architecture searches and, assuming approximately linear training cost, scales as $\mathcal{O}(N^2)$. We propose a block-based, leakage-controlled approach that shares NAS runs between subjects, reducing the number of searches from $N$ to $B$, where $B \ll N$, dubbed PainNAS. On the BioVid Heat Pain dataset, PainNAS yields comparable subject-level accuracy with substantially fewer parameters and FLOPs.

4. OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

OneBid:面向多样化oCPX广告场景的统一自动出价基础模型

AI 总结:OneBid提出统一自动出价基础模型,通过双信号条件化、序列级MoE和CROP离线优化,在快手oCPX广告中实现整体+2.2%、ROAS场景+13.1%的ADVV提升。

链接:https://arxiv.org/abs/2609.21550

机构:Kuaishou Technology(快手科技)

作者:Yewen Li, Peng Jiang, Yitian Li, Pengfei Lv, Xialong Liu, Peng Jiang, Qingpeng Cai

英文摘要:Auto-bidding is central to computational advertising, where strategies must maximize advertisers' conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevailing optimized cost-per-X (oCPX) paradigm, which spans heterogeneous scenarios (e.g., registration, purchase), each served by a separate model, leading to fragmented pipelines and underexploring cross-scenario modeling. Inspired by foundation models like LLMs, unifying these oCPX scenarios into one model raises three challenges: multi-objective control, scalable capacity under strict latency, and safe offline policy improvement. We present OneBid, a unified auto-bidding foundation model that learns a reusable backbone from heterogeneous oCPX logs and adapts it to scenario-specific deployments via offline post-training. Building on DT, OneBid extends single Return-to-Go conditioning to two atomic signals, Return-to-Go for conversion value and Cost-to-Go for cost ratio, plus value-aware regularization on next-action prediction. To absorb distributional heterogeneity, we design a sequence-level Mixture-of-Experts architecture, where shared experts encode cross-scenario knowledge and sparsely-routed experts capture scenario-specific patterns at low latency, yielding consistent scaling with model size and data. During post-training, we align the backbone with scenario preferences via Critic-guided Relative Offline Policy optimization (CROP): a learned critic scores candidate actions group-relatively, avoiding the unsafe online exploration of GRPO-style fine-tuning while constraining policy shift to reduce OOD risk. Validated via online A/B tests and fully deployed at Kuaishou, OneBid delivers an overall +2.2% ADVV gain on oCPX Ads, peaking at +13.1% in the ROAS scenario.

5. Trading Depth for Time in Recurrent Transformers

在循环Transformer中用时间换取深度

AI 总结:本研究通过潜在循环Transformer比较时间递归与物理深度增加,发现插入思维词元可在减少约48%参数的情况下恢复双深度模型67%-81%的性能改进,表明时间思考是参数高效的深度替代方案。

链接:https://arxiv.org/abs/2609.21605

机构:Microsoft(微软); University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

作者:Zeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen

英文摘要:Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Latent Recurrent Transformers (LRTs), which retain one backbone forward pass per vocabulary token during decoding and provide a controlled setting for comparing these two ways of adding computation. Specifically, we insert a latent thought token between consecutive vocabulary tokens. Each thought token passes through the same $L$ layers as a vocabulary token, sharing the backbone parameters and providing an additional stage of hidden-state refinement before predicting the next token. We compare this $L$-layer LRT against a $2L$-layer LRT without thought tokens. Both execute $2L$ Transformer blocks per vocabulary token during decoding, but the thought-token model uses fewer parameters. On 16- and 20-layer mixture-of-experts NanoChat backbones, one thought token brings the shallower model within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with approximately 48% fewer total parameters. These results suggest that temporal thinking offers a parameter-efficient alternative to increasing physical depth in recurrent Transformers.

2. 表示学习、自监督与对比学习 | 1 篇

6. Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

Matrix AdaGrad:行式与列式自适应次梯度方法

AI 总结:本文提出行式与列式矩阵AdaGrad优化方法,通过在线镜像下降框架利用矩阵结构,实现更紧的遗憾界,并在矩阵分解和深度网络训练中提升稳定性与可训练性。

链接:https://arxiv.org/abs/2609.21815

机构:Shanghai Jiao Tong University(上海交通大学)

作者:Wenpeng Zhang, Runsheng Yu, Peilin Zhao

英文摘要:Adaptive optimization methods such as AdaGrad and Adam are widely used in modern neural-network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop a general Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled approach to deriving matrix-aware adaptive optimization through online regret minimization. By introducing row-wise and column-wise matrix proximal functions and analyzing the resulting regret trade-off, we derive Row-wise Matrix AdaGrad (Row-AdaGrad) and Column-wise Matrix AdaGrad (Column-AdaGrad), with adaptive scaling determined by the accumulated row-wise or column-wise gradient norms. We establish regret guarantees and show that these matrix-aware bounds can be strictly tighter than those of entry-wise AdaGrad under structured gradients. Experiments on matrix factorization and deep neural-network training further demonstrate the benefits of aligning adaptive scaling with matrix structure, including improved optimization stability and trainability at larger learning rates and greater network depths.

3. 强化学习与序列决策 | 10 篇

7. BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

BI-Agent 与 BI-Bench:迈向端到端商业智能的自动化

AI 总结:本文提出 BI-Agent 与 BI-Bench,首个系统评估 LLM 端到端 BI 能力的基准,通过工具增强与后训练显著提升准确率。

链接:https://arxiv.org/abs/2609.20886

机构:UIUC(伊利诺伊大学厄巴纳-香槟分校); Microsoft Research(微软研究院); Microsoft(微软)

作者:Chuxuan Hu, Yeye He, Penny Zhou, Wee Hyong Tok, Daniel Kang, Surajit Chaudhuri

英文摘要:Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions. These steps can be complex and time-consuming, making BI challenging. Given the strong capabilities of large language models (LLMs) in working with data, we study their ability to answer BI questions end-to-end, without requiring users to manually perform the tedious preparation steps. To do this, we harvest a large collection of real-world BI projects from public sources, and manually extract pairs of (questions, ground-truth answers) from real user dashboards. The resulting benchmark, BI-Bench, is the first benchmark to systematically study LLMs' ability on end-to-end BI. We find that even frontier LLMs perform poorly on BI-Bench, with less than 50% accuracy. To address their limitations, we design a tool-augmented BI-Agent that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages. Furthermore, we develop a post-training framework that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL). BI-Agent achieves substantial accuracy gains of up to 40 percentage points with vanilla LLMs, and post-trained BI-Agent yields gains of up to 30 points. Our results highlight the importance of combining tool-augmented reasoning with domain-specific post-training in complex BI workflows, and point to promising directions for future research.

8. Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies

连续延迟记忆随机梯度下降与基于天体物理时间序列历史的连续时间强化学习

AI 总结:本文提出连续延迟记忆随机梯度下降,在二维景观上比普通SGD探索更广、收敛更精确,并设计无需求解HJB方程的连续时间强化学习结构,其最优性条件恢复吉布斯策略。

链接:https://arxiv.org/abs/2609.20906

作者:Debartha Paul, Juncheng Yi

英文摘要:Quasars are luminous objects in the universe that exhibit stochastic brightness variations encoding information about the supermassive black holes powering them, and modeling these variations from ground-based survey data time series, known as light curves, is a statistical challenge. This paper reviews how stochastic differential equations (SDEs) have been adapted with neural network parameterizations to overcome this challenge in history. We create the Continuous-Delayed-Memory Stochastic Gradient Descent which depend on the past state of the discrete iteration process. We performed the simulation on some 2-dimensional landscape and observed some wider-exploration and more precise convergent behavior compared to Vanilla SGD by adjusting hyperparameters. Besides, we proposed a reinforcement learning structure with continuous time policy gradients for exploratory policies without solving HJB PDE, and we show that its optimality conditions recover the Gibbs policy of previous works.

9. Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

具有时序逻辑规范的高效贝叶斯自适应强化学习

AI 总结:本文提出一种基于模型的贝叶斯自适应强化学习算法,通过同步LDBA与BAMDP并采用新颖的BAMCP规划,在未知环境中高效满足LTL规范,提升样本效率并减少训练违规。

链接:https://arxiv.org/abs/2609.20954

机构:University of Oxford(牛津大学)

作者:Jonathan Hau, Alessandro Abate

英文摘要:We present a novel end-to-end model-based Reinforcement Learning (RL) algorithm for efficient policy synthesis under given Linear Temporal Logic (LTL) specifications (e.g., safety or reachability) in unknown environments. To do so, a Limit-Deterministic B{ü}chi Automaton (LDBA) representation of the LTL task is synchronised with a Bayes-Adaptive Markov Decision Process (BAMDP) representation of the environment, which allows us to leverage an enhanced exploration-exploitation trade-off that is achieved via Bayesian RL, as opposed to traditional non-Bayesian approaches. We further propose a novel Bayes-Adaptive Monte-Carlo Planning (BAMCP) algorithm to allow for approximate Bayes-optimal strategy synthesis in the synchronised BAMDP construct. A range of finite- and infinite-horizon task experiments demonstrate the effectiveness of our approach in terms of both property satisfaction and sample efficiency, when compared to traditional model-free approaches. Additional ablation studies also successfully highlight the value of the novel BAMCP algorithm in comparison to classical BAMCP for LTL task satisfaction. Finally, we also showcase a successful application of our approach for \textit{cautious} RL, namely to reduce the number of task violations incurred during policy training.

10. ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning

ASGARD:基于强化学习的无人机韧性动作空间防护

AI 总结:针对无人机强化学习控制器易受动作空间攻击的问题,提出ASGARD两阶段师生管道,通过编码器融合状态与攻击特权信息训练监控器,在运行时修正动作命令,实现对多种攻击的韧性与泛化。

链接:https://arxiv.org/abs/2609.20982

机构:The University of British Columbia(不列颠哥伦比亚大学)

作者:Mohsen Salehi, Karthik Pattabiraman

英文摘要: Reinforcement learning (RL) controllers have been recently adopted for Unmanned Aerial Vehicles (UAV) navigation and control. However, they are susceptible to action-space attacks that overwrite the action commands after the policy generates them and before the actuators execute them. While most existing defenses target attacks on the policy's inputs, those addressing action-space attacks retrain the policy at training time and are not resilient to corrupted actions at runtime. We propose ASGARD, a two-phase teacher-student pipeline for making RL-based UAV control resilient to action-space attacks. In the teacher phase, an encoder combines the UAV's physical state with action-attack-related privileged information to produce an action-attack-aware latent that trains the RL control policy and a monitor that outputs corrected action commands to the actuators. In the student phase, both the encoder and the monitor are trained via supervised learning from their teacher counterparts to run on-board using only the UAV's physical state history. We evaluate ASGARD across attack scenarios targeting different action commands on UAV. We find that ASGARD is resilient to action-space attacks and completes the missions despite the attack. We further find that ASGARD generalizes to unseen attacks and remains resilient against stealthy attacks.

11. REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement

REFINEPPO:通过迭代动作细化学习连续控制策略

AI 总结:本文提出REFINEPPO,通过迭代动作细化机制让策略逐步改进动作,结合PPO算法,在14个基准任务上达到或超越标准PPO,并加快收敛。

链接:https://arxiv.org/abs/2609.21108

机构:Northeastern University(东北大学)

作者:Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs

英文摘要:Deep reinforcement learning (DRL) has achieved strong performance across a wide range of continuous-control problems. These continuous-control policies, however, are often defined as direct mappings from an observed state to an action or action distribution, requiring a single feed-forward network to construct an optimal control decision in one pass. While effective, this formulation leaves little opportunity for the policy to reconsider or progressively improve an action once an initial prediction has been formed. In this work, we explore an alternative approach: rather than learning only to directly predict an action, can a policy learn to iteratively improve one, and can this iterative process provide advantages during policy learning? We introduce Iterative Action Refinement (IAR), an iterative action-construction method that constructs control actions through a sequence of learned residual corrections. Starting from an initial proposal, a shared refinement network repeatedly conditions on the observed state and the current action proposal, allowing each refinement step to revise the action constructed by preceding steps. The final refined proposal is then used to determine the action executed by the agent. We integrate this iterative action-construction mechanism with Proximal Policy Optimization (PPO), yielding REFINEPPO. We evaluate REFINEPPO across 14 benchmark control tasks, complemented by controlled ablations of refinement depth and update schedules and analyses aimed at understanding why iterative refinement is effective. Across these environments, REFINEPPO matches or exceeds the performance of standard PPO while demonstrating faster convergence on several tasks.

12. Deep Reinforcement Learning with Buffered Quantile Objectives

带缓冲分位数目标的深度强化学习

AI 总结:提出无模型深度强化学习框架Deep-BQRL,通过缓冲分位数目标和集成探索实现风险敏感决策,在资产出售等任务中优于PPO和TRPO。

链接:https://arxiv.org/abs/2609.21327

机构:Virginia Tech(弗吉尼亚理工大学)

作者:Mohammad Alipour-vaezi, Sajad Khodadadian

英文摘要:Quantile-based reinforcement learning provides an interpretable approach to risk-sensitive decision-making by optimizing a prescribed quantile of the cumulative-return distribution. Despite this appeal, learning under a point quantile objective is challenging: quantiles can change abruptly under small perturbations of the return distribution, and exact quantile-sensitive planning requires computationally demanding distributional optimization. Lower-buffered quantiles alleviate the former difficulty by averaging neighboring quantiles immediately below the target level, providing a smoother surrogate while preserving the underlying point-quantile objective. Existing methods based on this principle, however, remain model-based and rely on explicit return-law planning, limiting their applicability beyond small tabular problems. We develop Deep-BQRL, a model-free distributional reinforcement-learning framework that extends buffered-quantile learning to neural function approximation. The method learns conditional return quantiles directly from sampled transitions, constructs buffered action scores from the relevant region of the learned quantile function, and uses ensemble disagreement to guide exploration. An augmented input representation allows the learned policy to respond to trajectory information without explicitly reproducing the quantile-state recursion required by exact planning. Experiments on an asset-selling optimal-stopping problem and slippery FrozenLake compare Deep-BQRL with model-based UCB-BQRL and tabular PPO and TRPO implementations. In asset selling, Deep-BQRL attains smaller mean cumulative point-quantile policy gaps than PPO and TRPO at the reported target levels, while UCB-BQRL retains the smallest gaps. The learned stopping decisions also vary with the target quantile, providing an interpretable illustration of the method's risk-sensitive behavior.

13. IncentRL: The Trade-Off Between Preference Guidance and Task Performance

IncentRL:偏好引导与任务性能之间的权衡

AI 总结:IncentRL通过KL惩罚引入偏好引导并量化其对任务性能的影响,在MiniGrid上以0.01系数将成功率从90.5%提升至98%,平衡了偏好引导与原始任务目标。

链接:https://arxiv.org/abs/2609.21525

机构:Fudan University(复旦大学)

作者:Xuening Wu, Yanlan Kang, Shenqin Yin

英文摘要:Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performance. IncentRL adds a Kullback--Leibler (KL) penalty between a specified outcome distribution and a preferred distribution. For finite discounted Markov decision processes with bounded shaping costs, we derive an external-value perturbation bound, establish a sufficient strict-action-gap condition for preserving the original optimal policy, and characterize the large-weight regime through discounted cumulative preference cost. Exact examples clarify the limits of these guarantees, including tied optima and support mismatch. We study a practical implementation using a hand-designed, distance-based outcome proxy, a fixed preference distribution, and score-weighted coefficient search. On MiniGrid DoorKey-8x8, the reported three-seed mean success rate after two million training steps reaches 98\% with coefficient 0.01, compared with 90.5\% for the reported zero-coefficient baseline, while the search progressively shifts toward smaller coefficients. Together, these results provide a principled view of the central trade-off in preference-based RL: using additional guidance to improve learning without excessively distorting the original task objective. The current experiments remain descriptive and do not yet isolate KL shaping from simpler alternatives.

14. GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

GEM-MPC:通过专家引导规划平衡探索与利用

AI 总结: GEM-MPC提出一种基于MPPI的强化学习方法,通过结合克隆规划器的策略与KL正则化探索策略,并引入门控先验蒸馏,在较低计算预算下平衡探索与利用,显著提升连续控制基准性能。

链接:https://arxiv.org/abs/2609.21735

机构:Leiden University(莱顿大学)

作者:Alvaro Serra-Gomez, Thomas Moerland

英文摘要:Effective exploration in high-dimensional continuous control remains a central challenge in reinforcement learning. Planning-based methods address this by combining online planning with learned policies and value functions, but their components can become misaligned during training: learned sampling policies may diverge from planner behavior, while planning distributions stored in replay become stale as the model and value function evolve. Reanalysis can refresh these targets, but at substantial computational cost. We propose GEM-MPC, an MPPI-based reinforcement learning method that improves the interaction between planning and learning. GEM-MPC uses MPPI to combine a policy trained to clone the planner with a KL-regularized policy that explores around it, providing complementary exploitation and guided exploration within planning. We further introduce Gated Prior Distillation, which selectively learns from stored planning distributions only when they provide a better target than the current prior, reducing the impact of stale planning data without requiring full reanalysis. Across continuous-control benchmarks, GEM-MPC consistently outperforms existing planning-based baselines under lower computational budgets.

15. Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning

超越运动学:肌肉驱动模仿学习的仿真保真度基准测试

AI 总结:本研究系统比较了基于HyFyDy和MuJoCo的两种肌肉驱动模仿学习流程,发现HyFyDy的肌肉激活更接近实验EMG数据,更适于肌肉骨骼建模,但两者均需提升生理真实性。

链接:https://arxiv.org/abs/2609.21909

机构:Georgia Institute of Technology(佐治亚理工学院)

作者:Ayah G. Ahmad, Claire E. Borden, Maegan Tucker

英文摘要:In this work, we conduct a systematic comparison of two state-of-the-art motion-imitation reinforcement learning (MIRL) pipelines, one built on SCONE/HyFyDy and one built on MuJoCo/MyoSim. HyFyDy emphasizes physiological realism through detailed musculotendon modeling, while MuJoCo prioritizes computational efficiency and scalable policy learning. While recent work has demonstrated that both pipelines reproduce human kinematics with high fidelity, it remains unclear if they accurately capture the underlying neuromuscular behavior that produced the movement. This limitation is particularly important for robotic assistive-device design and control, where outcome measures such as muscle activation patterns and metabolic cost are often used as optimization targets. To conduct a systematic comparison, our work compares both pipelines using a common set of human motion-capture and electromyography (EMG) measurements. The results find that while both pipelines produce similar kinematics with relative accuracy, the muscle activations from HyFyDy are more aligned with the experimental EMG, as supported by the average pooled (RMSE, r) values for muscle activations from HyFyDy and MuJoCo: (0.164, 0.4) and (0.344, 0.11), respectively. While we conclude that the more advanced physiological realism of HyFyDy currently makes it more suitable for musculoskeletal modeling, both require further development to bring physiological realism to GPU-parallelizable simulation environments and advance robotic assistive device design.

16. Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks

学习移动城市:用于城市网络校准与控制的深度元模型和强化策略

AI 总结:本文提出共享潜在空间框架,结合MLP-自编码器与深度Q学习,实现城市交通模拟器校准和动态控制,在基准网络上将系统出行时间降低最多51%。

链接:https://arxiv.org/abs/2609.21945

机构:George Mason University(乔治梅森大学); University of South Dakota(南达科他大学); Syracuse University(雪城大学); Federal University of Technology Akure(联邦理工大学阿库雷分校)

作者:Adewumi Augustine Adepitan, Christopher J. Haruna, Oluwasegun Adegoke, Ayooluwatomiwa Ajiboye, Oluwatobi Oluwasakin

英文摘要:Urban transportation networks present complex optimization challenges spanning calibration of high-fidelity simulators and real-time operational control. This paper presents a shared latent-space framework that connects simulator calibration and reinforcement learning control through a common learned representation of urban traffic dynamics. First, we develop a combinatorial MLP-autoencoder architecture that learns low-dimensional manifolds linking simulator inputs (origin-destination demand, network parameters) to outputs (travel times, congestion patterns), enabling efficient Bayesian optimization for calibration. This approach demonstrates superior sample efficiency compared to traditional dimension reduction methods, achieving better fit to observational data within fixed computational budgets. Second, we implement a deep Q-learning agent with experience replay and target networks to optimize dynamic traffic assignment through scheduling and routing adjustments. In empirical evaluations on benchmark networks, our approach reduces system-wide travel times by up to 51% compared to baseline operations. The learned latent representation is not only used to reduce the dimensionality of Bayesian calibration, but is also incorporated into the reinforcement learning state representation, allowing the control policy to operate on compressed and calibrated traffic dynamics. This shared latent-space formulation provides a unified pathway from simulator calibration to adaptive operational control within intelligent transportation systems. Our results highlight the transformative potential of deep learning methods in urban mobility planning and management, particularly for large-scale networks where traditional optimization approaches face computational bottlenecks.

4. 生成模型与概率建模 | 6 篇

17. Generative inversion for early ranking of competing geologic interpretations

竞争性地理解释的早期排序的生成性反演

AI 总结:提出一种生成性反演工作流,通过将竞争性地质解释转化为空间先验并依据水头观测进行排序,利用文本到图像模型和变分自编码器生成图像,经反演网络和流动模拟评估兼容性,在合成基准和实际案例中验证了有效性。

链接:https://arxiv.org/abs/2609.20978

机构:Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)

作者:Harun Ur Rashid, Daniel O'Malley

英文摘要: High-consequence subsurface decisions are often made under severe data scarcity. Experts may arrive at competing interpretations of the same subsurface system, yet early in a project there is rarely a practical way to determine which one is most realistic. This uncertainty can persist until several wells are drilled, often costing millions of dollars. Existing approaches for evaluating geologic interpretations rely either on subjective judgment or on dense data that are rarely available in early-stage investigations. We present a workflow that addresses this challenge by translating competing geologic interpretations into alternative spatial priors and ranking them according to their consistency with hydraulic-head observations. For each interpretation, a text-to-image foundation model generates an ensemble of 1600 geologic images, and a separately trained variational autoencoder provides an interpretation-specific latent representation. A supervised inverse network maps the head observations into this latent space, and the frozen decoder produces an image that is mapped to a log-conductivity field. Steady-state flow simulation then provides predicted heads, and the resulting mismatch is converted into a Gaussian-form compatibility score. We evaluate the framework using a synthetic benchmark based on the Johansen Formation and three interpretations of decreasing consistency with the reference representation. Across 925 test cases, the mean head RMSE increases from 0.197 for the Precise \& Accurate interpretation to 0.227 for the Accurate interpretation and 0.280 for the Mismatched interpretation. We subsequently apply the workflow to two published conceptual models of the Culebra Dolomite Member at the Waste Isolation Pilot Plant. The revised model receives a compatibility weight of 0.991, compared with 0.009 for the original model, consistent with the independent evidence.

18. HMB-GAN: Hybrid Multi-Bézier GAN for Vector Shape Synthesis

HMB-GAN:用于矢量形状合成的混合多贝塞尔生成对抗网络

AI 总结:本文提出HMB-GAN,一种混合量子-经典生成对抗网络,用于合成CAD矢量几何形状,通过多段贝塞尔表示和构造性几何连续性,实验表明量子生成器虽收敛快、参数少,但受模拟器开销限制,验证了混合量子架构的可行性。

链接:https://arxiv.org/abs/2609.21158

作者:Elian Hugh Thiele-Evans, Binh Duong Pham, Hani Omar M Alharbi, Liibaan Aaden, Syed Umer Hasnain Zaidi, Prem Prakash Jayaraman, Muhammad Saeed, Boris Eisenbart

英文摘要:We explore the use of hybrid quantum-classical generative adversarial networks for synthesising CAD-ready vector geometries. Unlike prior work that operates in rasterised or single-Bézier domains, we introduce HMB-GAN (Hybrid Multi-Bézier GAN), an end-to-end differentiable generative framework that constructs closed shapes through stitched multi-segment Bézier representations with geometric continuity enforced by construction. We compare a quantum-enhanced generator with a classical generator within this architecture and evaluate them across point cloud distribution metrics and geometric shape statistics. Results show that despite faster convergence, a reduction in model parameter count, and slightly improved performance on point cloud metrics, the quantum generator suffers from excessive simulator overhead and thus classically-simulated evaluation suffers from hardware constraints. These results demonstrate the feasibility of modelling structured geometries through hybrid quantum architectures whilst highlighting contemporary hardware limitations.

19. Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability

黎曼神经哈密顿流:测地辛传输与可解释性

AI 总结:本文提出黎曼神经哈密顿流,结合流形固定动能、学习势与测地积分器,实现可解释的生成建模,并在多种空间验证其性能与可解释性。

链接:https://arxiv.org/abs/2609.21647

机构:CEA, DAM, DIF(法国原子能委员会)

作者:Vincent Souveton

英文摘要:Hamiltonian normalizing flows are attractive generative models because their phase-space maps are invertible and volume preserving, but most neural constructions are formulated in Euclidean space. We introduce Riemannian Neural Hamiltonian Flows, which combine the fixed kinetic energy of a Riemannian manifold, a learned scalar potential, and an explicit geodesic leapfrog integrator. Our analysis explains how the learned Hamiltonian can be made interpretable. Every normalizable potential defines an implicit profile, and the position marginal initially accelerates along the relative score between that profile and the base. The matched potential is the interpretable specialization for which the implicit profile is the target. In the isotropic Gaussian case, the mechanism corresponds to a phase-space rotation. A local harmonic analysis extends this result around each mode of a general target on a manifold. The gap between the learned and the matched potential is the sum of a residual memory of the base and a bias of the model, and the two potentials agree when the position base has been transferred to the momentum. This can be achieved when the former is broader than the target. Numerical experiments on Euclidean, hyperbolic, and spherical spaces show competitive sample quality and numerical cost against a Riemannian continuous normalizing flow, and confirm the interpretability of the learned potential.

20. Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources

扭结与光滑性:拉普拉斯类源分布下实解析nICA的可辨识性

AI 总结:本研究证明实解析生成函数在源分布具有有限一阶导数不连续点(如拉普拉斯分布)时可辨识,利用扭结与光滑性对比,适用于归一化流和变分自编码器,并在CelebA数据上恢复可解释潜在因子。

链接:https://arxiv.org/abs/2609.21926

机构:University of Florida(佛罗里达大学)

作者:Isaac Manring, Kejun Huang

英文摘要:Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central challenge of nonlinear Independent Component Analysis (nICA). We prove identifiability (exact recovery) up to trivial ambiguities for real analytic generating functions when source probability density functions have a finite number of discontinuities in the first derivative. The Laplace distribution is the most prominent example satisfying this assumption. Our proof relies on the contrast between kinks in the source distribution and the smoothness of real analytic functions. Real analytic functions comprise a broad class of generating mechanisms, and can be approximated with Normalizing Flows or Variational Autoencoders with standard activation functions (e.g., tanh, softplus, GELU), so our result applies with minimal changes to existing training pipelines. We perform experiments on real and synthetic data with both Normalizing Flows and Variational Auto-Encoders demonstrating their identifiability properties. In experiments on CelebA data we recover several interpretable latent factors controlling unique attributes across the dataset.

21. Time series generation with spectrally aligned latent flow matching

时间序列生成与谱对齐潜在流匹配

AI 总结:本文提出谱对齐潜在流时间序列生成器,通过傅里叶、小波及签名变换的微调损失改善谱失配,在真实基准上验证了其生成真实性与效率优势。

链接:https://arxiv.org/abs/2609.21989

机构:Imperial College London(伦敦帝国理工学院)

作者:Camilo Carvajal Reyes, Felipe Tobar

英文摘要:Latent flow models have proven to be a reliable and cost-effective method for time series generation. However, the latent compression induces unwanted artefacts, such as a spectral mismatch with respect to the underlying dataset, thus hindering their use as training surrogates. In this article, we propose a spectrally-aligned latent-flow time series generator, where the latent space for flow matching is trained to preserve dynamical properties that are relevant for the suitability of synthetic samples. We find that incorporating fine-tuning losses based on canonical signal representations such as the Fourier, wavelet and signature transforms helps overcome these issues. The interpretability of these transformations allows us to ensure that the synthetic signals are aligned with the true ones in terms of relevant features, such as smoothness or targeted spectral content, as opposed to relying on pointwise reconstruction losses only. We compare the proposed aligned models against a base latent-flow model and the state of the art over real-world long-range univariate and multivariate benchmark datasets. Our quantitative results validate the superiority of the proposed method in terms of its performance on metrics reflecting signal realness and computational efficiency, while being aligned to the training set with respect to its local structure.

22. $λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

$λ$-Controlled GRPO:将流匹配比率不稳定性转化为预算化资源

AI 总结:本文提出$λ$-Controlled GRPO,将流匹配强化学习中的路径方差作为可预算资源,按预测规律校准重要性比率并分配梯度,提升文本准确性和偏好奖励。

链接:https://arxiv.org/abs/2609.22041

机构:Stony Brook University(石溪大学); Rivian and Volkswagen Group Technologies(Rivian和大众集团科技公司)

作者:Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa

英文摘要:Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback. Training in this setting is unstable in a way specific to multi-step denoising: the policy update changes systematically across denoising steps, with importance ratios drifting below one, becoming increasingly dispersed, clipping at different rates, and leaving fewer usable samples late in training. Prior work treats these effects as separate failure modes and addresses each with a hand-tuned stabilizer. We show instead that they arise from a single per-step quantity, which we call path variance. This quantity is determined exactly by the sampler's Gaussian transition kernel and can be estimated cheaply during training. This reframes instability as a resource that can be measured and budgeted rather than a collection of symptoms to repair. Our method, $\lambda$-Controlled GRPO, calibrates importance-ratio behavior from this predicted law rather than from noisy empirical statistics, and allocates gradient effort across denoising steps according to their predicted cost. The two scales governing the update are fixed by standard policy choices rather than introduced as free tuning parameters. On a text-to-image model under two reward settings, rendering difficult target text scored by optical character recognition and matching human preferences scored by a preference model, $\lambda$-Controlled GRPO improves both text accuracy and preference reward over the strongest empirical stabilizer. It also keeps late-step path variance within its intended budget, precisely where the baseline systematically overshoots. The result is a Flow-GRPO update calibrated by its own transition law rather than stabilized after instability appears.

5. 优化、泛化与理论分析 | 4 篇

23. From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences

从切换遗憾到动态遗憾:一种通过无偏随机序列的简单归约

AI 总结:本文提出一种简单归约框架,将动态遗憾最小化转化为切换遗憾最小化,通过构造无偏随机序列,为强凸、指数凹和一般凸损失分别建立了匹配极小极大最优的动态遗憾界。

链接:https://arxiv.org/abs/2609.20968

机构:State Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术全国重点实验室); School of Artificial Intelligence, Nanjing University(南京大学人工智能学院); School of Software Technology, Zhejiang University(浙江大学软件学院)

作者:Yibo Wang, Wenhao Yang, Sifan Yang, Yuanyu Wan, Lijun Zhang

英文摘要:In non-stationary online learning, dynamic regret has attracted increasing attention as a measure of how well an online learner performs against a time-varying comparator sequence. Despite considerable advances, attaining optimal bounds for strongly convex and exp-concave losses often involves intricate analysis. In this paper, we present a \textit{simple} framework that reduces dynamic regret minimization to switching regret minimization. As a result, we can derive dynamic regret bounds by using off-the-shelf algorithms with switching regret guarantees. The key idea of our reduction is to construct, for \textit{any} comparator sequence, an auxiliary random sequence that is unbiased at each round, with the controlled variance and a manageable number of switches. Combining this construction with suitable surrogate losses, we can decompose dynamic regret into the expected switching regret against the random sequence and its controlled variance. Theoretically, for strongly convex and exp-concave losses, we establish the $\widetilde{O}(T^{1/3}P_T^{2/3})$ dynamic regret bounds, where $T$ denotes the time horizon and $P_T$ denotes the path-length of the comparator sequence. Moreover, for general convex losses, the same reduction also recovers the $O(\sqrt{T(1+P_T)})$ dynamic regret bound. Notably, all our findings match the minimax optimal results for these three types of losses, highlighting the versatility of our proposed framework.

24. Multi-Domain Clustering via Measure Quantization

基于测度量化的多域聚类

AI 总结:提出基于测度量化的多域聚类框架,通过最小化概率度量学习共享原型,结合最优传输分配,在小批量优化下高效且优于基线。

链接:https://arxiv.org/abs/2609.21664

机构:Instituto Federal de Educação, Ciência e Tecnologia do Ceará(塞阿拉联邦教育与科学技术学院); Federal University of Ceara(塞阿拉联邦大学); Sigma Nova Science(Sigma Nova Science公司)

作者:Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante

英文摘要:Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples from multiple domains, we learn a shared set of cluster prototypes by minimizing a probability metric, such as the Sinkhorn divergence or the Maximum Mean Discrepancy, between each domain's probability measure and the measure of prototypes. Data points are then assigned to clusters either via nearest centroid, or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini-batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance. Experimental results on 5 multi-domain benchmarks spanning image, audio and sensor data show that our Sinkhorn-based method consistently outperforms classical and multi-domain clustering baselines, and that this advantage persists when scaling to hundreds of thousands of samples.

25. Optimization Geometry of Equivalent Brownian RKHS Representations

等价布朗再生核希尔伯特空间表示的优化几何

AI 总结:该研究通过有限布朗再生核希尔伯特空间中的节点、增量和谱坐标,揭示了等价参数化下优化轨迹与条件数的坐标依赖效应,并验证了相关理论预测。

链接:https://arxiv.org/abs/2609.21693

机构:Free University of Bozen-Bolzano(博尔扎诺自由大学); Universitat de València(瓦伦西亚大学)

作者:Mahdi Mohammadigohari, Gustau Camps-Valls

英文摘要:Equivalent finite parameterizations can represent the same functions and intrinsic norm yet induce different optimization algorithms. We study this effect in a controlled finite Brownian RKHS with nodal, increment, and spectral coordinates. Classical finite-element, RKHS-interpolation, Brownian-covariance, and mixed-boundary DCT identities make the shared hypothesis class, Brownian energy, approximation operator, and coordinate maps explicit. Our main results concern the optimization geometry of this fixed model. With mapped initialization, identical scalar steps, and identical minibatches, nodal and spectral GD/SGD have exactly the same mapped trajectories. Increment GD is an explicit Euler step for the constant Brownian/Sobolev metric, with factor $1/h$. For Brownian-regularized least squares, $\kappa_2(\mathbf H_{\mathrm{inc}})\le1+A/\rho$, independently of grid resolution $G$ for fixed $A$, $\rho>0$, and the stated normalization. Under the stated standard-Adam convention, the universal orthogonal equivariance group is exactly the signed permutations; the block DCT-VIII transform is not one. Float64 tests over five grids numerically verify the finite identities, mapped one-layer and recursive trajectories, conditioning predictions, and theorem-matched Adam separation. Thus coordinate effects are isolated without changing the represented functions, intrinsic regularizer, or approximation space.

26. RACER: Role-Aligned Competence Estimation for Human-AI Routing

RACER:面向人机路由的角色对齐能力估计

AI 总结:RACER提出角色对齐的能力估计框架,利用上下文估计未见专家能力,结合模型后验实现贝叶斯最优的人机路由,在合成和医学影像基准上表现优异。

链接:https://arxiv.org/abs/2609.21953

机构:University of Oxford(牛津大学)

作者:Joshua Strong, Emma Sun, Alexander Capstick, Pramit Saha, Cheng Ouyang, J. Alison Noble

英文摘要:Learning to defer asks a predictive system when to act autonomously and when to defer to a human expert. Population-adaptive deferral extends this problem to unseen experts using a small context set of expert behavior. Neural context encoders such as L2D-Pop can be query-dependent, but may learn routing shortcuts tied to absolute class coordinates. Identity-Free Deferral (IFD) removes such shortcuts through role-indexed classwise competence profiles, but its estimates are constant within each class and cannot capture instance-level expert specialization. We propose RACER---Role-Aligned Competence Estimation for Routing---a role-relative framework for estimating an unseen expert's competence from context. RACER estimates the posterior-predictive probability that the expert is correct on a query under each candidate class role, then combines these estimates with the model posterior to obtain the Bayes-relevant expert-correctness probability. Nonparametric and neural kernel-pooling estimators use candidate-role relations, shared aggregation, and symmetric summaries, excluding absolute class-identity channels. We prove coherent class-relabelling invariance, derive a Bayes-aligned deferral surrogate, and give a plug-in regret bound relating routing regret to classifier and competence-estimation error. On controlled synthetic benchmarks, including a PathMNIST histopathology context-scaling study with simulated experts, RACER benefits from additional context under hidden subtype dependence and gives the strongest aggregate performance on a separately sampled unseen-expert split in the CIFAR-100 synthetic experiments. On the radiologist and human--AI chest-radiography benchmarks (VinDr-CXR and CheXpert), the RACER family is competitive or best in budget-swept deferral, with calibration results varying across metrics and datasets.

6. 高效学习、压缩与部署 | 2 篇

27. MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery

MOSAIC-SR:基于Transformer引导的符号回归用于科学方程恢复

AI 总结:MOSAIC-SR利用预训练Transformer生成初始草图,引导符号回归搜索,通过尺度感知常数优化和局部符号修复,在多个基准上实现最高符号解率与高预测精度。

链接:https://arxiv.org/abs/2609.20997

机构:University of Waterloo(滑铁卢大学)

作者:Peiyi Zheng, Yanming Kang, Hans De Sterck, Giang Tran

英文摘要: Symbolic regression aims to recover closed-form equations from observations, providing interpretable models for scientific discovery. Existing approaches struggle to combine flexible structural search with efficient inference. Search-based methods can refine expression structure but often rely on costly combinatorial optimization with random initialization. Pretrained neural models generate formulas almost instantly, but their predictions often contain symbolic errors. We introduce MOSAIC-SR, which uses a pretrained Transformer to propose multiple initial sketches. These sketches initialize searches in several promising regions, avoiding random starts in the vast expression space. Each search jointly recovers structure and constants through scale-aware constant optimization and local symbolic repair. We evaluate MOSAIC-SR on the SRSD-Feynman dataset with and without dummy variables and on six additional benchmarks. MOSAIC-SR obtains the highest symbolic solution rate on every dataset while ranking among the top two methods in predictive accuracy. This advantage persists in the presence of irrelevant dummy inputs. The results show that learned priors can focus search on promising equation structures, and that numerical optimization and symbolic repair are important for recovery.

28. SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

SpecQuant:基于多父量化的推测解码实现自适应LLM推理

AI 总结:SpecQuant是一种无需训练的框架,结合推测解码与多父量化,根据查询复杂度动态选择量化变体,在Qwen2.5模型上实现35-43%加速且精度损失不超过2%,支持设备端高效部署。

链接:https://arxiv.org/abs/2609.21704

机构:Vellore Institute of Technology(韦洛尔理工学院)

作者:Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T

英文摘要:Running large language models (LLMs) locally continues to be limited by restrictions of compute and memory on consumer hardware. The popular acceleration technologies, such as quantization, speculative decoding, and adaptive inferencing, offer substantial speed boosts but usually necessitate retraining, per architecture tuning, or draft models. SpecQuant is a trainingfree framework, that combines speculative decoding with multiparent quantization to perform adaptive, efficient inference of LLMs. SpecQuant derives multiple quantized variants (INT4, FP8, FP16) from a shared base model, and dynamically routes queries based on predicted complexity; lightweight variants are used for simple or factual tasks, and full-precision models are used for complex reasoning tasks or long-context inputs. The shared-weight design of SpecQuant ensures sufficient token acceptance for speculative decoding without compatibility issues using separate draft parent models. We evaluate SpecQuant on Qwen2.5 based models on the MMLU, AlpacaEval, and GSM8K datasets, or benchmarks, demonstrating 35-43% speedups without degrading accuracy greater than 2%, substantial within the LLM community. SpecQuant enables practical on-device LLM deployment across diverse hardware without special infrastructure or expertise.

7. 联邦学习、隐私与安全 | 2 篇

29. FedeRage: Provably Convergent Agnostic Federated Learning under General Client Drift

FedeRage:一般客户端漂移下可证明收敛的不可知联邦学习

AI 总结:针对客户端参与概率未知且数据异构的联邦学习,提出风险规避平均算法FedeRage,嵌入CVaR提升高损失和低频客户端权重,实现可证明收敛并在准确率、公平性和速度上超越现有方法。

链接:https://arxiv.org/abs/2609.21057

作者:Herlock Rahimi, Dionysis Kalogerias

英文摘要:Federated learning (FL) enables collaborative model training without sharing raw data, but its performance degrades under non-IID data and stochastic client participation. Remedies built on classical Federated Averaging (FedAvg) typically presuppose that client participation probabilities are known to the server, which is rarely the case in deployed systems. We first discuss and then characterize the optimization problem that \emph{distributionally agnostic} FedAvg actually solves when participation is entirely unknown, possibly highly skewed, and of variable size across rounds: uniform aggregation is shown to minimize a well-defined stochastic objective, weighted by the participation-induced marginal, at a standard $\mathcal{O}(1/\sqrt{T})$ rate for convex and possibly nonsmooth losses. Building on this characterization, we propose \emph{Federated Risk-Averse Averaging} (\textsc{FedeRage}), a risk-averse extension of FedAvg that embeds the \emph{Conditional Value-at-Risk} (CVaR) into the local objective within a natural distributionally robust optimization (DRO) framework. \textsc{FedeRage} implicitly upweights high-loss and infrequently participating clients while adding only a \emph{single scalar per-client}, and admits an $\mathcal{O}(\kappa/\sqrt{T})$ rate in which the factor $\kappa$ is the upper bound on the ``price" of risk aversion. In contrast with aggregation-alignment schemes based on optimal transport, which require the availability distribution as an input, \textsc{FedeRage} remains agnostic to it. Several experiments on three heterogeneous benchmarks indicate consistent improvements over state-of-the-art methods in accuracy, fairness, and convergence speed.

30. Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data

面向高维异构数据的联邦深度聚类网络

AI 总结:针对联邦学习中高维异构数据聚类问题,提出FedDCN方法,通过联合优化重建与聚类损失并引入几何正则化,在非独立同分布场景下实现鲁棒聚类。

链接:https://arxiv.org/abs/2609.21829

机构:Maastricht University(马斯特里赫特大学)

作者:Morris Stallmann, Charalampos S. Kouzinopoulos, Marcin Pietrasik, Anna Wilbik

英文摘要:Clustering high-dimensional data is a fundamental task in unsupervised machine learning with applications to a variety of domains. In the centralized data scenario, this task is commonly solved using deep clustering methods that utilize deep neural network architectures to learn clustering-friendly latent space representations. In Federated Learning, where data is distributed between clients and is private, deep clustering methods are less explored. In particular, recently introduced federated deep clustering methods, despite showing very promising performance, still fall short in reliably providing good performance if data across clients are non-identically-independently distributed. In this work, we introduce a generalization of Deep Clustering Networks to the federated scenario, named FedDCN, that simultaneously optimizes a reconstruction loss and a clustering loss. To ensure robustness and latent space alignment in non-identically-independently distributed data scenarios, FedDCN generates synthetic data augmentations, and its learning objective includes a geometric regularization for latent space alignment. Through experimental evaluation, the effectiveness of the approach under IID and non-IID assumptions is demonstrated, and future research directions are identified.

8. 鲁棒性、不确定性与可信学习 | 2 篇

31. On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation

最大编码率降低用于分布外泛化的局限性

AI 总结:本文揭示最大编码率降低(MCR²)在分布外泛化中的两大局限:其目标可导致完全预测失败,且结合不变性原理仍无法消除,需新假设或原则确保跨环境稳定预测。

链接:https://arxiv.org/abs/2609.21001

机构:University of Sheffield(谢菲尔德大学)

作者:Menghui Zhou, Gaoshan Bi, Vitaveska Lanfranchi, Po Yang

英文摘要:Substantial efforts have been devoted to making deep learning objectives, representations, and architectures interpretable, with the goal of improving the safety, robustness, and generalisation of learning systems in diverse real-world applications. The recently proposed maximal coding rate reduction ($\mathrm{MCR}^{2}$) offers a promising information-theoretic framework for learning structured, discriminative representations of class-wise submanifolds and has inspired interpretable white-box architectures. However, we observe that $\mathrm{MCR}^{2}$ can completely fail under distribution shift, motivating our study of its out-of-distribution (OOD) generalisation limits. We establish two limitations of $\mathrm{MCR}^{2}$ for OOD generalisation. First, the $\mathrm{MCR}^{2}$ objective alone can admit complete prediction failure: a representation based entirely on unstable environmental features can achieve the global coding optimum yet fail completely after correlation reversal, despite an available perfectly stable feature. This exact-optimum example includes test inputs that cannot occur during training. Even when every possible test input can also occur during training, coding quality can be arbitrarily close to optimal while prediction error is arbitrarily close to 100%. Second, directly incorporating the invariance principle underlying widely successful invariant risk minimisation (IRM) and risk extrapolation (REx) does not eliminate this failure. The failing representation admits the same optimal coding operator across training environments, showing that shared coding optimality does not ensure stable prediction. Reliable OOD guarantees for $\mathrm{MCR}^{2}$ therefore require additional new assumptions or learning principles that establish stable predictive relationships across environments.

32. ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

ExpBoN:用于高效测试时大语言模型对齐的指数噪声最佳n采样

AI 总结:本文提出ExpBoN,一种基于指数噪声机制的软BoN采样方法,实现指数级快速收敛,并集成到GSI框架形成ExpGSI,在保持精度的同时大幅降低测试时对齐的计算成本。

链接:https://arxiv.org/abs/2609.21899

机构:Imperial College London(伦敦帝国理工学院); University of Washington(华盛顿大学)

作者:Yanxiao Liu, Sicheng Wan, Deniz Gündüz

英文摘要:Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of-$n$ (Verdun et al. 2025) provides smoother control and converges to the optimal distribution associated with KL-regularized reward maximization. In this paper, we introduce ExpBoN, an alternative soft BoN method based on the exponential-noise report-noisy-max mechanism. It admits an exact finite-$n$ decomposition, which yields exponentially fast convergence in total variation, expected reward, and both directions of KL divergence. We provide comprehensive theoretical analyses of its convergence and regret behavior. We further integrate ExpBoN into the guided speculative inference (GSI) framework (Geuter, Mroueh, and AlvarezMelis 2025), resulting in ExpGSI, for efficient reward-guided LLM alignment. ExpGSI yields substantial reductions in computational cost while maintaining comparable accuracy. Experiments on MATH500, MMLU-STEM, and Minerva Math with the Qwen2.5-Math and Qwen3 model families show that ExpGSI reduces estimated computation by $14\%$-$39\%$ across candidate budgets for Qwen2.5-Math and by up to $45\%$ at $n=16$ for Qwen3. Overall, our results provide a theoretical and algorithmic foundation for exponential-noise BoN and efficient test-time LLM alignment.

9. 图学习与结构化数据 | 1 篇

33. EnSol: an environment-aware graph neural network for molecular solubility prediction

EnSol:一种用于分子溶解度预测的环境感知图神经网络

AI 总结:EnSol是一种环境感知概率图神经网络,通过交叉注意力捕捉溶质-溶剂相互作用并纳入温度调制,在多个基准上实现领先的溶解度预测,支持可靠溶剂选择。

链接:https://arxiv.org/abs/2609.21151

机构:Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校西贝尔计算与数据科学学院); Carl R. Woese Institute for Genomic Biology, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校卡尔·R·沃斯基因组生物学研究所); NSF Molecule Maker Lab Institute, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校NSF分子制造实验室研究所); Department of Chemical and Biomolecular Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校化学与生物分子工程系); DOE Center for Advanced Bioenergy and Bioproducts Innovation, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校DOE先进生物能源与生物制品创新中心)

作者:Thao Nguyen, Saman Shafaei, Zhengyi Zhang, Huimin Zhao, Heng Ji

英文摘要:Molecular solubility directly affects key aspects of molecular development such as reaction feasibility, formulation performance, separation efficiency, and solvent selection. However, experimental measurement across solutes, solvents, and temperatures remains costly and sparsely sampled. Existing computational models often rely on fixed-solvent assumptions, deterministic formulations, or simplified representations of solute-solvent interactions, limiting their ability to capture complex molecular interactions, continuous temperature effects, and experimental uncertainty. Here, we introduce EnSol, an environment-aware probabilistic framework for molecular solubility prediction. EnSol represents the solute and solvent as molecular graphs and learns separate representations for each before bringing them together through cross-attention to capture solute-solvent interactions. Temperature is incorporated directly into the solvent environment through feature-wise modulation, and a mixture density network predicts full solubility distributions to capture both temperature-dependent behavior and experimental uncertainty. On the independent SolProp and Leeds benchmark datasets, EnSol achieved Spearman correlations of 0.876 and 0.601, respectively, outperforming state-of-the-art solubility prediction models across both benchmarks. Beyond computational benchmarking, experimental validation across chemically diverse solute-solvent pairs showed that EnSol maintained strong predictive performance and supported reliable solvent ranking, achieving a Spearman correlation of 0.715. These results show that EnSol can support reliable solubility prediction and solvent selection across diverse chemical systems while accounting for predictive uncertainty.

10. 迁移、元学习与持续学习 | 3 篇

34. Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks

Stiefel-AdamW:用于线性分解块的几何感知AdamW

AI 总结: 针对线性分解块中分解不唯一导致的训练不稳定问题,提出Stiefel-AdamW优化器,通过将一因子约束于Stiefel流形以消除规范对称性,在保持AdamW效率的同时提升稳定性,并在多种模型上验证了其有效性。

链接:https://arxiv.org/abs/2609.21039

机构:Gran Sasso Science Institute(格兰萨索科学研究所); University of Edinburgh(爱丁堡大学)

作者:Emanuele Zangrando, Marco Sutti, Francesco Tudisco

英文摘要:A pervasive structural pattern in modern deep learning is the linear factorization block: a submodule of the form $W = BA$ in which two parameter matrices are multiplied directly, with no intervening nonlinearity. Such blocks appear in LoRA adapters, low-rank compressed layers, query-key products of self-attention, and share a common pathology: the factorization is non-unique, which can destabilize training and limit usable learning rates. Despite this, factorization blocks are typically optimized with standard Euclidean methods that ignore the underlying geometry. We introduce Stiefel-AdamW, a near drop-in replacement for AdamW for use wherever such blocks appear. By constraining one factor on the Stiefel manifold while leaving the other Euclidean, Stiefel-AdamW relaxes the full $\mathrm{GL}(\mathbb{R}^r)$ gauge symmetry to a compact orthogonal symmetry, ruling out factor blow-up while retaining the coordinate-wise diagonal preconditioning that gives AdamW its practical strength. Moment estimation is performed in the ambient Euclidean space, with geometry entering only through a tangent-space projection and a manifold retraction. The implementation overhead over AdamW is minimal, and we show that the resulting optimizer inherits both the stability benefits of Riemannian methods and standard convergence guarantees. We validate Stiefel-AdamW on LoRA-style fine-tuning of GPT2, ViT, and Mistral 7B and on full pretraining of GPT2 on OpenWebText, showing consistent improvements over strong baselines at essentially no additional cost over AdamW.

35. Neural Cellular Automata Learn General Features in their Hidden Channels

神经细胞自动机在其隐藏通道中学习通用特征

AI 总结:本文研究神经细胞自动机隐藏通道的内部动态,提出将预训练教师隐藏状态注入学生模型的迁移学习机制,在少样本MNIST上以约9800参数超越循环和前馈架构,证明隐藏通道捕获通用尺度不变拓扑特征,实现参数高效迁移学习。

链接:https://arxiv.org/abs/2609.21870

机构:Østfold University of Applied Sciences(东福尔郡应用科学大学)

作者:Etienne Guichard, Stefano Nichele

英文摘要:Modern deep learning models achieve impressive generalization through over-parameterization, but this paradigm often struggles with overfitting and memorization in few-shot regimes. Neural Cellular Automata (NCAs) offer a highly parameter-efficient alternative, yet research has focused primarily on their output, leaving the role of their internal hidden channels largely unexplored. In this paper, we investigate the internal dynamics of NCA hidden channels and introduce a novel transfer-learning mechanism that injects a pretrained teacher's hidden states into a student model to guide early optimization. Evaluated on few-shot and scale-variant MNIST benchmarks, NCAs outperform comparable recurrent and feed-forward architectures, demonstrating superior generalization with a minimal parameter budget (~9,800 parameters). Mechanistic analysis reveals that the hidden channels decouple feature extraction from uniform classification consensus by absorbing morphological complexity and converging to mutually orthogonal states. Furthermore, we demonstrate that these hidden channels capture general, scale-invariant topological primitives rather than class-specific templates. This allows a student model to achieve strong few-shot performance on unseen classes using features transferred from a teacher trained only on a subset of digits (0-5). Our results highlight the potential of utilizing hidden-state dynamics as a robust, decentralized computational substrate for parameter-efficient transfer learning

36. Benchmarking World Models for Continual Learning on Compositional Tasks

组合任务持续学习的世界模型基准测试

AI 总结:针对机器人操作中的世界模型,提出组合持续学习基准,以分离知识重用与学习新任务,评估显示模块化模型优于传统方法但仍有改进空间。

链接:https://arxiv.org/abs/2609.22055

机构:University of Oxford(牛津大学)

作者:Haoyu Zhou, Joe Watson, Anson Lei, Ingmar Posner

英文摘要:A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what the agent has already learnt. In particular, the ability to retain and reuse knowledge obtained from prior experiences underpins an agent's ability to efficiently adapt to novel environments, as the dynamics of the physical world can often be described in recurring mechanisms. However, the world model's measure of adaptation entangles two abilities: the speed and capacity to learn unseen tasks, and the reuse of knowledge already acquired, since incoming tasks carry novel content alongside what recurs. In order to isolate knowledge reuse from prior experiences, we propose a compositional continual learning benchmark for world models in robot manipulation. Specifically, we design each task curriculum with compositional tasks that combine aspects of the tasks seen in the sequence. We further factorise this composition along the axes of action and perception to better understand how different input modalities bottleneck knowledge reuse. We evaluate state-of-the-art world models under canonical continual learning methods, alongside a modular world model whose dynamics backbone contains explicitly reusable components. Results show that modularity balances reuse against forgetting better than conventional methods, but none solve the problem fully, leaving clear room for continual world models built to reuse without forgetting. More details are available on our project website: this https URL.

11. 数据集、基准与评测 | 4 篇

37. Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization

基于可靠性中心的稀疏纵向CT病灶大小预测评估:结合共形区间校准与Gompertz启发正则化

AI 总结:本研究基于稀疏纵向CT数据,构建基准并评估多种预测方法,发现额外历史观测收益有限,且准确性、可靠性与轨迹一致性需联合评估。

链接:https://arxiv.org/abs/2609.21197

作者:Lingfei Kong

英文摘要: Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker (DLT), yielding 205 trajectories from 129 patients. We compared an exploratory conventional sparse-to-final analysis with a primary fixed visit-index horizon design predicting the common log change from T3 to T4 while progressively adding earlier observations, evaluating predictive accuracy, uncertainty reliability, post-hoc conformal interval calibration, subgroup performance, and Gompertz-inspired trajectory regularization. The evaluated methods showed partially overlapping point-prediction accuracy but distinct uncertainty behavior. Mean held-out RMSE across ten training seeds was 0.4726, 0.4305, 0.4499, and 0.4513 for m = 1, 2, 3, 4, indicating the lowest mean RMSE at m = 2; additional history did not improve RMSE. At m = 4, raw Cohort-Level Feature GP coverage was near the 95% nominal level, whereas MC Dropout, Deep Ensemble, and residual-scale intervals were conservative. Patient-level conformal calibration generally produced near-nominal or conservative coverage at the cost of wider intervals. Patient-grouped development cross-validation selected lambda* = 0 for the Gompertz-inspired term. A global population reference frequently opposed lesion-level change directions, and prediction difficulty varied across anatomical subgroups. Overall, additional historical observations provided limited predictive benefit once the prediction horizon was controlled, while predictive accuracy, uncertainty reliability, and trajectory consistency did not necessarily improve together, and should be evaluated jointly in sparse longitudinal imaging.

38. OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

OpenMAS-GCom:面向图增强多智能体系统的诊断基准

AI 总结:针对图增强多智能体系统性能归因困难的问题,提出OpenMAS-GCom诊断基准,通过受控干预评估组件影响,在29个数据集上测试17种配置,发现不同组件干预导致不同性能变化。

链接:https://arxiv.org/abs/2609.21527

作者:Kairui Yang, Xunkai Li, Kaixiang Zhang, Minghao An, Zekai Chen, Yuxuan Ba, Rong-Hua Li

英文摘要:Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score comparisons across systems combine differences in models, communication patterns, roles, and computation costs, making performance differences difficult to attribute to specific communication structures, role assignments, and information flows. To address this evaluation attribution problem, we introduce OpenMAS-GCom, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions. We represent systems through collaboration units, communication links, shared intermediate information, and execution rules. OpenMAS-GCom compares original systems with versions modified by changing one component while keeping tasks, models, prompts, and budget limits fixed. We rewire communication edges, remove specialist or critic agents, replace intermediate messages with incorrect content, and disable workers during execution. The benchmark evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains. We add 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers. Experiments show larger mean losses after specialist removal than after critic removal, different performance degradation under incorrect messages and worker failures despite similar original scores, and different configurations achieving the highest accuracy and accuracy per token on G-MAS-Complex.

39. GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

GraphSkillEvo:图结构智能体技能的进化优化

AI 总结:GraphSkillEvo提出将智能体技能表示为图结构,并通过进化优化框架(含变异和交叉算子)在结构化空间中高效搜索,在五个基准上超越SkillOpt,平均准确率提升最高达4.01%。

链接:https://arxiv.org/abs/2609.21749

机构:City University of Hong Kong(香港城市大学); National University of Singapore(新加坡国立大学); Southern University of Science and Technology(南方科技大学)

作者:Rui Sun, Zhi Zheng, Zhenkun Wang, Zhichao Lu

英文摘要:Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. However, existing skill optimization methods typically represent skills as unstructured natural-language instructions, creating two key challenges: 1) Unstructured skills often lack explicit workflow-level guidance and contain substantial redundancy, making them difficult for LLMs to execute; 2) the vast search space of unconstrained natural-language skills makes skill optimization ineffective. To address these challenges, we propose representing skills as graph-structured natural-language artifacts. In graph-structured skills, each node represents an execution step together with its operational guidance, while directed edges encode context-dependent transitions between steps. Compared to unstructured skills, graph-structured skills can provide clear workflow-level guidance. Moreover, the proposed graph-structured skill can also facilitate skill optimization. Building on this structured representation, we introduce GraphSkillEvo, a population-based evolutionary optimization framework with mutation and crossover operators for graph-structured skills. By maintaining multiple candidate skills and combining effective components, GraphSkillEvo enables broader and more comprehensive exploration of the structured skill space than purely LLM-based iterative self-refinement. Extensive experiments across five agent benchmarks demonstrate that GraphSkillEvo consistently outperforms the strong skill optimization baseline SkillOpt, improving average accuracy by 4.01% on GPT-5.4-nano and 1.76% on GPT-5.4. Our code is available at this https URL.

40. BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings

BrainWideBench:多区域神经记录中的大规模预训练与跨动物迁移基准

AI 总结:BrainWideBench基于跨276个脑区、139只小鼠的数据,提出三套任务基准,系统评估预训练方法在跨动物迁移中的表现,发现当前方法在行为、动态和解剖结构上的泛化仍面临挑战。

链接:https://arxiv.org/abs/2609.22064

机构:University of Pennsylvania(宾夕法尼亚大学); Mila(米拉研究所); Université de Montréal(蒙特利尔大学); Stanford University(斯坦福大学); Columbia University(哥伦比亚大学); Allen Institute(艾伦研究所); William James Center for Research(威廉·詹姆斯研究中心); ISPA - Instituto Universitário(ISPA 大学研究所); University of Geneva(日内瓦大学); Karolinska Institutet(卡罗林斯卡学院); UCLA(加州大学洛杉矶分校); University College London(伦敦大学学院); Champalimaud Foundation(尚帕利莫基金会); Lingang Laboratory(临港实验室); The Chinese University of Hong Kong(香港中文大学); Donders Institute(唐德斯研究所); University of Minnesota(明尼苏达大学); Princeton University(普林斯顿大学); Leiden University(莱顿大学); McGill University(麦吉尔大学); IBM

作者: Alexandre Andre, Shivashriganesh P. Mahato, Vinam Arora, Keshav Balaji, Divyansha Lachi, Nanda H. Krishna, Jingyun Xiao, Yizi Zhang, Ximeng Mao, Wenrui Ma, Han Yu, International Brain Laboratory, Daniel Birman, Niccolò Bonacchi, Gaelle A. Chapuis, Joana A. Catarino, Felicia Davatolhagh, Mayo Faulkner, Laura Freitas-Silva, Fei Hu, Julia M. Huntenburg, Anup Khanal, Inês Laranjeira, Petrina Lau, Guido T. Meijer, Nathaniel J. Miska, Jean-Paul Noel, Alejandro Pan-Vazquez, Georg Raiser, Cyrille Rossant, Karolina Z. Socha, Anne E. Urai, Miles J. Wells, Steven J. West, Olivier Winter, Blake Richards, Guillaume Lajoie, Cole Hurwitz, Mehdi Azabou, Matthew R. Whiteway, Liam Paninski, Eva L. Dyer

英文摘要:Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural representations transferable across diverse downstream tasks. Yet, progress toward this goal has been limited by fragmented evaluation protocols and a narrow focus on individual task domains. Here, we present BrainWideBench, a benchmark for evaluating across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset of neural and behavioral recordings spanning 276 brain regions from 139 mice performing a sensory-guided decision-making task. The benchmark is organized around three complementary task suites that evaluate whether learned representations support downstream decoding of behavior, can predict masked or future neural activity, and can recover biologically meaningful anatomical organization. With this benchmark, we systematically evaluate pretraining methods across transfer settings, including finetuning on downstream objectives and zero-shot generalization to unseen animals. Our results confirm pretraining improves performance over matched single-session baselines, but we show current methods exhibit heterogeneity in transfer capabilities: gains depend strongly on the alignment between pretraining objectives and downstream tasks. No single approach performs uniformly well across all three suites, and most methods are designed to only address a subset of them. Together, these findings suggest that learning representations that jointly generalize across behavior, dynamics, and anatomy remains an open challenge. By providing a unified and reproducible evaluation suite, BrainWideBench establishes a framework for measuring progress toward general-purpose models of the mouse brain.

12. 机器学习应用 | 6 篇

41. A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters

一种用于基于Transformer的时间序列预测器的轻量级即插即用门控模块

AI 总结:本文提出一种轻量级编码器前门控模块,作为即插即用接口调节协变量表示,在零额外调参下提升TimeXer、iTransformer和PatchTST的预测性能,并验证了使用正则化的有效性。

链接:https://arxiv.org/abs/2609.21044

机构:Minjiang University(闽江学院)

作者:Hongkai Zhuang, Tao Huang, Chen Hou

英文摘要:Covariate-rich time-series forecasting requires deciding how external variables enter the target forecasting path. Existing Transformer-based forecasters usually build a covariate representation and pass it to the encoder without an explicit admission stage. This paper studies pre-encoder covariate admission as an input-side interface that regulates that representation immediately before encoder processing. We implement the interface with a lightweight representation-level pre-encoder gate that assigns sigmoid scores to representation units, and we also study a usage-regularized variant that penalizes average admission. The interface is evaluated as a plug-in module for TimeXer, Inverted Transformer (iTransformer), and Patch Time Series Transformer (PatchTST) under a zero-extra-tuning protocol, where each gated model inherits the corresponding baseline configuration. Experiments on the Electricity Transformer Temperature minute-level (ETTm1 and ETTm2) datasets, Traffic, Energy, and influenza-like illness (ILI) include paired forecasting comparisons, gate-placement ablation, initialization ablation, controlled covariate-admission analysis, and a variance inflation factor (VIF)-informed permutation feature importance (PFI) diagnostic case study. In the tested settings, the gate is competitive with the corresponding baselines, and the usage penalty reduces average admission scores while keeping forecasting errors close to the unpenalized TimeXer setting.

42. Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers

全连接层的分层解耦稳定结构化稀疏化方法

AI 总结:提出一种逐层解耦的结构化稀疏化方法,通过提取两层子网络并施加组惩罚,实现更稳健的剪枝,具有更宽的正则化范围和更低的过度剪枝率。

链接:https://arxiv.org/abs/2609.21126

机构:The University of Scranton(斯克兰顿大学); University of California, Santa Barbara(加州大学圣塔芭芭拉分校)

作者:Charles Kulick, Armenak Petrosyan, Sui Tang

英文摘要:We propose a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks. Rather than penalizing all layers jointly, our approach extracts shallow two-layer subnetworks, normalizes the inner weights, and applies a structured group penalty to the outer weight matrix of each block, processing layers sequentially to prune neurons and reduce the width of each layer. We prove that the constrained decoupled objective is equivalent at optimality to a specific joint penalty on the inner and outer weights, for any positively homogeneous activation, and thus admits a clean projected and proximal formulation. Our central finding is that this decoupled reformulation is more robust than coupled methods. In numerical experiments it provides a wider usable range of the regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. We establish these properties in controlled classification and sparse-recovery studies, and examine their scope in a high-dimensional PINN stress test and in the feed-forward layers of OPT-1.3B.

43. M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection

M2G-LLM:通过多模态图推理和LLM上下文注入增强临床预测

AI 总结:M2G-LLM通过图神经网络整合多模态数据并注入LLM中间层,在MIMIC-IV和MIMIC-CXR上提升临床预测性能,结合语言理解与关系推理。

链接:https://arxiv.org/abs/2609.21164

机构:University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校); University of Pennsylvania(宾夕法尼亚大学)

作者:Inyoung Choi, Sukwon Yun, Jiayi Xin, Jie Peng, Tianlong Chen, Qi Long

英文摘要: Integrating diverse data modalities --- such as clinical notes, laboratory results, and medical imaging --- is essential for advancing clinical decision-making. While Large Language Models (LLMs) have shown remarkable performance in processing unstructured clinical text, their limited capacity to incorporate non-text modalities hinders their broader utility in healthcare applications. Here, we introduce M2G-LLM (Multimodal MedGraph-LLM), a novel framework that enhances LLMs with multimodal integration and alignment via Graph Neural Networks (GNNs). Our approach models temporal relationships between patient visits, propagates information across clinically similar patients, and aligns heterogeneous data sources to construct enriched multimodal context vectors. These vectors are injected into the intermediate layers of the LLM, enabling joint reasoning over textual and non-textual modalities. We evaluate M2G-LLM on the MIMIC-IV and MIMIC-CXR datasets, demonstrating improvements in clinical prediction tasks over strong baseline models. Our results highlight the promise of combining the language understanding of LLMs with the relational reasoning capabilities of GNNs for comprehensive, multimodal healthcare analysis.

44. Knowledge-Graph-Augmented Chronos-2 for HEC-RAS Surrogate Forecasting

知识图谱增强的Chronos-2用于HEC-RAS代理预测

AI 总结:本研究提出KG-Chronos-2,将冻结的时间序列基础模型与水利工程知识图谱结合,通过残差解码和图检索,在HEC-RAS水面高程预测中显著降低RMSE,验证了知识增强的有效性。

链接:https://arxiv.org/abs/2609.21381

机构:Canizaro-Livingston Gulf States Center for Environmental Informatics(卡尼扎罗-利文斯顿墨西哥湾沿岸各州环境信息学中心); Naval Research Laboratory(海军研究实验室)

作者:Edward Holmberg, Elias Ioup, Mahdi Abdelguerfi

英文摘要:We investigate whether coupling a time-series foundation model to hydraulic project knowledge improves surrogate forecasting of HEC-RAS water-surface elevation (WSE). We present KG-Chronos-2, which combines a frozen Chronos-2 predictor with exact-state residual decoding, graph-conditioned historical retrieval, and input-aligned correction. We compare the method with persistence, a residual LSTM, project-conditioned recurrent GeoFNO, a hydraulic DCRNN-style model, and frozen Chronos-2. Task-specific fitting uses the 2008 simulation. Evaluation covers 64 fixed 24-hour windows from the 2011 and 2002 simulations at 4,675 cross sections in 71 reaches on a shared geometry. KG-Chronos-2 achieves event-balanced root-mean-square error 0.246970 in native WSE units. It reduces RMSE by 14.13% relative to frozen Chronos-2, 29.38% relative to the hydraulic DCRNN-style model, and 39.54% relative to recurrent GeoFNO. The 95% hierarchical-bootstrap interval for its event-balanced RMSE difference from frozen Chronos-2 is [-0.075177, -0.016317]. KG-Chronos-2 also achieves the lowest active-window and final-lead RMSE among the six completed systems. These results support coupling a frozen temporal predictor to project knowledge for warm-start HEC-RAS forecasting on the fixed benchmark.

45. Probabilistic Forecasting of Business Process Executions with Neural Temporal Point Processes

基于神经时间点过程的业务流程执行概率预测

AI 总结:针对业务流程执行预测,提出基于标记时间点过程的生成式模型,结合变换器与混合解码器处理并列时间戳,在保持点准确性的同时提升剩余时间分布校准性并降低推理成本。

链接:https://arxiv.org/abs/2609.21382

机构:Paris Dauphine University-PSL(巴黎多菲纳大学-PSL); University of Vienna(维也纳大学)

作者:Jiaxin Yuan, Daniela Grigori, Han van der Aa

英文摘要:Operators of service-based systems act on forecasts of how a running execution will continue, and such a forecast is actionable only if its reliability is known. Mainstream deep-learning models for this task are discriminative and deterministic: they emit a single next activity and a single remaining-time estimate, without a distribution to reason over. We instead cast the problem as generative sequence modelling with marked temporal point processes, which define a joint density over the next mark and its inter-event time and therefore deliver predictive distributions by construction. Real event logs violate the simple-point-process assumption these models rest on, since consecutive events frequently carry identical timestamps; we handle such ties explicitly and combine a transformer encoder with a mixture decoder over inter-event times, trained by exact log-likelihood. On ten public logs, the resulting model matches discriminative baselines on point accuracy, dominates them on the calibration and sharpness of remaining-time distributions, and is the cheapest at inference, since a full predictive distribution is obtained in a single forward pass without sampling.

46. Tracing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework

零样本时间序列预测的证据溯源:一种以来源为先的分类法与审计框架

AI 总结:本文提出以来源为先的分类法与审计框架,将零样本时间序列预测视为证据访问声明,区分三种证据来源并明确审计问题,以提升基准可审计性。

链接:https://arxiv.org/abs/2609.21425

机构:Technical University of Munich(慕尼黑工业大学)

作者:Delun Kong, Wanyun Ling, Chenxi Liu, Ziyue Li

英文摘要:Zero-shot time-series forecasting (TSF) is often described as forecasting without target-specific parameter updates, but that training-status condition does not specify what evidence the system may use. A frozen language model prompted with serialized values, a time-series model pretrained on broad forecasting corpora, and a retrieval-augmented forecaster may all satisfy the no-update condition while drawing on different transferable evidence. This paper argues that zero-shot TSF should therefore be governed as an evidence-access claim. We propose a source-first taxonomy that separates three primary evidence sources---frozen LLM prior reuse, parametric time-series pretraining, and retrieval-augmented external memory---from the architectures that implement them. After the source is identified, four additional audit questions remain: task interface, forecast object and scoring, prediction-time context, and resource budget. The resulting agenda is to make zero-shot leaderboards auditable by reporting evidence boundaries and interface assumptions alongside scores, so that benchmark progress reflects transferable forecasting capability rather than undisclosed changes in context, memory, or budget.

13. 其他/综合机器学习 | 37 篇

47. Sparse Priors for Efficient Distribution Learning

稀疏先验用于高效分布学习

AI 总结:针对高维分布学习样本复杂度随维度恶化的问题,提出稀疏先验类及稀疏维度度量,证明在k-稀疏先验下可实现O(√(k/n))的贝叶斯风险界,并扩展到学习采样,从而克服维数灾难。

链接:https://arxiv.org/abs/2609.20883

机构:Carnegie Mellon University(卡内基梅隆大学)

作者:Saumya Goyal, Barnabás Póczos

英文摘要: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal. We hypothesize that present bounds are too pessimistic because smoothness assumptions are not enough to capture the structure of distributions that often appear in real applications. Consequently, we introduce the class of sparse priors and define the "Sparse Dimension" as a measure of sparsity of a prior over the space of all distributions. We show that distribution learning under a $k$-sparse prior achieves a Bayesian risk lower bound of $\Omega(\sqrt{k/n})$ under common distance metrics, and show a matching (up to logarithmic terms asymptotically in $n,k$) upper bound for the TV distance under mild additional assumptions. We show the statistical equivalence of distribution learning and learning to sample in the Bayesian setting so that our results apply to learning to sample as well. While $k$ can still depend on the dimension $d$, or a notion of intrinsic dimension, our results show that learning under an appropriate prior overcomes the curse of dimensionality with respect to the dependence on $n$.

48. Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

弹性阈值注意力:面向长上下文解码的学习型上下文稀疏性

AI 总结:针对长上下文解码中KV缓存导致的内存带宽瓶颈,提出弹性阈值注意力(ETA),通过端到端学习动态上下文阈值实现硬件加速解码,在不损失稠密模型质量下,以约38%活跃密度媲美稠密注意力,并带来高达2.5倍解码加速。

链接:https://arxiv.org/abs/2609.20888

机构:Boston University(波士顿大学)

作者:Themistoklis Haris, Henry Li, Maryam Karimzadehgan

英文摘要:Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this via selective loading, but that comes at a cost: rigid heuristics drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ETA)}, an end-to-end trainable architecture that achieves hardware-accelerated decoding speed without sacrificing dense model quality. ETA predicts dynamic, contextual thresholds directly from query representations, allowing the model to allocate dense-like context to difficult retrieval or reasoning steps while pruning routine tokens. To learn this policy from scratch without representation collapse, ETA \emph{multiplicatively suppresses} sub-threshold logits toward zero during training rather than deleting them. Training against this smooth uniform attention floor provides a distributed probability reservoir that \textbf{causes localized attention sinks on initial tokens to disappear}. It also enables the model to hard-prune uninformative KV blocks at inference time and absorb incidental tokens co-admitted by coarse GPU block selection. As a result, a 1.45B pretrained ETA model rivals dense attention across language modeling, commonsense reasoning, and long-context needle retrieval at $\approx 85\%$ training sparsity and $\approx 38\%$ active decode density. At inference time, we implement a custom decode kernel in Triton that screens KV blocks in $O(1)$ time using cached geometric-probabilistic bounds, delivering up to $2.5\times$ wall-clock decode speedups over FlashAttention-2 on sequences up to 512K tokens. Finally, we introduce an offline calibration algorithm for domain-specific deployments that freezes per-head constant thresholds to eliminate predictor overhead, cutting attention compute by an additional $27\%$.

49. Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain--Computer Interfaces

Bio-MF:用于混合运动想象脑机接口的低延迟高保真EEG到fNIRS跨模态生成

AI 总结:本文提出Bio-MF,一种无潜变量的一步MeanFlow框架,实现低延迟高保真EEG到fNIRS跨模态生成,通过直接信号空间预测和多种正则化技术,在保持生成质量的同时实现857倍加速,提升混合MI-BCI解码性能。

链接:https://arxiv.org/abs/2609.20904

机构:Shaanxi Normal University(陕西师范大学); Microsoft(微软)

作者:Boyuan Zhao, Sifan Zhang, Luping Chen

英文摘要:Hybrid motor-imagery brain-computer interfaces (MI-BCIs) combining EEG and fNIRS can outperform EEG-only systems by exploiting complementary electrophysiological and hemodynamic information. To obtain such hybrid information when paired EEG-fNIRS acquisition is unavailable or inconvenient, recent studies have focused on EEG-to-fNIRS cross-modal generation. However, existing methods still suffer from slow generation and often require pretraining, limiting their use in real-time MI-BCI scenarios. Although one-step generative models offer an attractive route to low-latency synthesis, removing the iterative refinement process can reduce generation fidelity and introduce non-physiological artifacts. To address these problems, this paper proposes Bio-MF, a latent-free one-step MeanFlow framework for EEG-conditioned fNIRS generation. Bio-MF performs direct signal-space x-prediction, converts this signal-space output into MeanFlow velocity supervision, and completes inference with one network evaluation. To preserve task-relevant hemodynamic structure under heterogeneous sensor layouts, Bio-MF integrates Spatial-Temporal Interactive 4D Encoding, cross-modal classifier-free guidance, and noise-level-gated FFT regularization. On Dataset 1, EEG + synthetic fNIRS improves ACC over EEG-only by 3.37 and 4.15 percentage points for HbR and HbO, respectively. On Dataset 2, the corresponding gains remain 2.98 and 2.50 percentage points under the unseen 64-channel EEG montage. On an RTX PRO 6000 GPU, Bio-MF generates one fNIRS trial in 7.0 ms, corresponding to an 857x speedup over the 1000-step SCDM latency. These results show that Bio-MF enables fast EEG-to-fNIRS synthesis while preserving task-relevant generation quality for downstream hybrid MI decoding. Our code is available at this https URL.

50. Do Quantum Models Scale Like LLMs?

量子模型能否像LLM一样扩展?

AI 总结:本研究通过RydbergGPT模型研究量子数据上的神经标度律,发现近临界统计类似自然语言,支持多尺度依赖性对稳定神经标度的作用。

链接:https://arxiv.org/abs/2609.20912

机构:Queen Mary University of London(伦敦玛丽女王大学); National Taiwan University(国立台湾大学); University of Waterloo(滑铁卢大学); Perimeter Institute for Theoretical Physics(圆周理论物理研究所)

作者:David S. Berman, Ying-Jer Kao, Roger G. Melko, Alexander G. Stapleton

英文摘要: In this work, we study the neural scaling laws of RydbergGPT, an autoregressive transformer model trained on qubit projective measurement data gathered from interacting Rydberg atom arrays. The quantum system is known to exhibit a finite-size remnant of a critical point as the laser detuning parameter is varied. We find that near the critical point the transformer loss as a function of training dataset size is well described by a power-law with a loss floor correction. However, away from criticality the quality of the power-law description is substantially reduced. We then compare the statistical structure of both Rydberg measurements and natural-language corpora using an entropy-normalised, finite sample corrected mutual information "two-point" function. We find that near-critical statistics of the two point functions are closest to those observed in natural-language, whilst other qubit configurations far from the critical point have two-point functions that decay more rapidly. This supports the hypothesis that multi-scale dependence contributes to stable neural scaling, and that scaling behaviour should be viewed as a property of the model-data pair.

51. When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

当AI评审训练AI评审者:科学判断崩溃及其缓解

AI 总结:本文研究AI评审递归训练导致科学判断崩溃的问题,提出TrustReviewer系统,通过训练时精选语料和测试时激活引导缓解该问题,保持判断多样性。

链接:https://arxiv.org/abs/2609.20942

机构:University of Maryland, College Park(马里兰大学学院公园分校)

作者:Sy-Tuyen Ho, Minghui Liu, Furong Huang

英文摘要:Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer review can become recursive: later reviewers learn from judgments produced by earlier models. We study one step of this feedback loop in a controlled setting. Starting from Llama 3.1 8B, we first fine-tune a reviewer on official ICLR reviews from 2018--2023 and then train four successor models on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews. Our study shows that introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity. We call this pattern $\textbf{scientific-judgment collapse}$. To mitigate this failure mode, we introduce $\textbf{TrustReviewer}$, an open-source LLM-based system for generating peer reviews of AI and machine learning papers. TrustReviewer intervenes at two complementary stages. For training-time prevention, we train the core reviewer in a single stage on a curated corpus designed to reduce low-quality and semantically degenerate supervision. For test-time correction, paired activation steering aims to further mitigate residual tendencies toward collapsed judgments without further training or additional expert annotation. Together, these results characterize a concrete risk of recursive reviewer training and provide practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted scientific evaluation.

52. From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

从压力到情感:跨可穿戴传感器模态的生理情感识别多模态深度学习

AI 总结:本研究比较了LSTM、TCN和Transformer在WESAD和EmoWear数据集上的生理情感识别性能,发现架构优劣依赖数据集,多模态传感优于单模态,4 Hz采样频率为实用选择。

链接:https://arxiv.org/abs/2609.20991

机构:Howard University(霍华德大学)

作者:Desta Haileselassie Hagos, Saurav Keshari Aryal, Legand L. Burge

英文摘要:Physiological emotion recognition using wearable sensors has important applications in mental health monitoring, affective computing, and human-computer interaction. However, existing studies typically evaluate a single model, sensing configuration, or dataset, limiting our understanding of how these factors influence recognition performance. We present a comparative study of temporal deep learning architectures for physiological emotion recognition using two multimodal wearable datasets: WESAD and EmoWear. Bidirectional long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer models are evaluated under wrist-only, chest-only, and multimodal sensing configurations using participant-independent leave-one-subject-out cross-validation (LOSO-CV). We also investigate soft-voting ensembles, sensor ablation, sampling frequency, and gradient-based saliency. The Transformer achieved the highest multimodal accuracy on WESAD (99.02% +/- 0.51%), whereas the LSTM achieved the best multimodal accuracy on EmoWear for both arousal (91.80% +/- 1.06%) and valence (89.96% +/- 0.36%). These results show that relative architecture performance depends on dataset characteristics rather than one architecture being uniformly superior. Multimodal sensing consistently outperformed wrist-only and chest-only configurations across both datasets. Sampling-frequency analysis showed that 4 Hz provides a practical operating point, with performance comparable to higher frequencies at substantially lower training cost. These findings provide guidance for selecting architectures, sensing modalities, and sampling frequencies for wearable physiological emotion recognition.

53. Scaling Discovery through Test-Time Communication

通过测试时通信扩展科学发现

AI 总结:本文证明测试时通信能显著提升多智能体在难题上的表现,团队成功率相当于四倍独立智能体,且优势随规模增长,并迁移至研究任务,超越已知最佳人类方案。

链接:https://arxiv.org/abs/2609.21032

机构:UC Berkeley(加州大学伯克利分校); Microsoft Research(微软研究院)

作者:Jongho Park, Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy, Dimitris Papailiopoulos

英文摘要:Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.

54. Toward individual-level calibration in affect recognition with perceptual adjustment queries

面向个体水平的情感识别校准:基于知觉调整查询的方法

AI 总结:针对面部情感识别中个体知觉差异导致任务难度不均的问题,提出基于知觉调整查询(PAQ)估计个体恰可察觉差(JND)以校准刺激距离,实验证明其优于未校准和群体校准,能显著均衡个体难度并降低反应时间。

链接:https://arxiv.org/abs/2609.21073

机构:Georgia Institute of Technology(佐治亚理工学院)

作者:Xuanzhou Chen, Sankaraleengam Alagapan, Ashwin Pananjady

英文摘要:Behavioral tasks measuring facial affect perception assume that identical stimuli impose equivalent perceptual difficulty across participants. However, this assumption is systematically violated by individual differences in perceptual sensitivity. Using an affective perception task as our testbed, we propose a framework to normalize for perceptual difficulty that directly estimates each participant's Just Noticeable Difference (JND) along the facial affect spectrum via cognitively lightweight perceptual adjustment queries (PAQs). We use these PAQ-inferred JNDs to re-express stimulus distances, constructing difficulty-equated tasks in perceptual space. We validate the framework in a Two-Alternative Forced-Choice (2AFC) task using two complementary behavioral measures: binary metacognitive difficulty judgments and response time variance decomposition. We find that PAQ calibration significantly equalizes perceived task difficulty at an individual level when compared to both the non-calibrated baseline and population-level Weibull calibration, while also reducing mean response time and between-subject variance in response time. These results establish PAQ as a principled and practical instrument for individualized perceptual calibration in facial affect recognition.

55. Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars

Talk to Me, Jarvis:一个面向自动驾驶赛车的开源可边缘部署语音助手框架

AI 总结:本文提出Jarvis,一个离线边缘部署的语音助手框架,通过微调Mistral 7B实现低延迟命令分类,在自动驾驶赛车中达到97.63%意图识别准确率和1.39秒平均延迟,并开源支持进一步研究。

链接:https://arxiv.org/abs/2609.21109

机构:Technical University of Munich(慕尼黑工业大学); Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所)

作者:Daniel Henel, Frederik Werner, Alexander Langmann, Johannes Betz

英文摘要:Recent advances in large language models have improved their effectiveness as back-end components for voice assistants, particularly in intent understanding and context-aware input classification. However, online-hosted models introduce network dependency and variable inference latency, limiting their suitability for time-critical autonomous driving applications. In this work, we address these issues by developing Jarvis, an offline voice assistant for high-level behavioral commands of autonomous vehicles. Its architecture integrates speech recognition and synthesis with natural language command classification into a lightweight, local framework. Jarvis core component is a text-to-command classifier, built using a domain-specific fine-tuning of the Mistral 7B model, demonstrating low-latency inference. Our experimental evaluation demonstrates that our solution outperforms larger online-hosted models, achieving 97.63 % intent recognition accuracy with an average processing latency of 1.39 s, making it well-suited for operations requiring quick response times. To support further research and fine-tuning, we provide an open-source implementation.

56. Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications

面向机器学习驱动应用的信号中心遥感:替代预处理与声学处理方法

AI 总结:针对声纳数据处理的图像方法耗时问题,提出基于CSV格式的替代预处理与声学处理方法,实验显示处理时间减少91.18%,并提升目标检测精度及SNR、PSNR等指标。

链接:https://arxiv.org/abs/2609.21123

机构:Georgia Institute of Technology(佐治亚理工学院); Embry-Riddle Aeronautical University(安柏瑞德航空大学)

作者:Logan Luna, Sirio Jansen-Sánchez, Ilteris Demirkiran, Leo Ghelarducci

英文摘要:The dominant method of processing sonar data is using image-based representations, requiring the preprocessing of image data on autonomous systems. We propose an alternative data processing method for remote sensing applications via the use of data in Comma-Seperated Value format. Experimentation on our alternative approach shows a reduction of processing time by 91.18%, an improvement in accurate object detection by Machine Learning, and an increase in SNR (Signal-to-noise ratio), PSNR (Peak signal-to-noise ratio), and other evaluation metrics.

57. TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

TierKV:通过预测式多层KV缓存实现长上下文端侧大语言模型

AI 总结:TierKV提出预测式多层KV缓存优化框架,通过预填充状态预测缓存需求并分层分配,在移动设备上实现高达17.6倍预填充加速和12.5-34%内存节省,支持更长上下文。

链接:https://arxiv.org/abs/2609.21172

机构:University of Georgia(佐治亚大学); Peking University(北京大学); Western Digital Research(西部数据研究院)

作者:Zhihao Shu, Md Musfiqur Rahman Sanim, Jie Hu, Kun Yuan, Minghai Qin, Gagan Agrawal, Wei Niu

英文摘要:Large language models (LLMs) are moving onto mobile devices for increasingly diverse workloads over text, images, video, and audio. These applications often require long contexts, making the Key-Value (KV) cache a dominant memory bottleneck because it grows linearly with sequence length and is accessed at every decoding step. Prior work reduces KV-cache footprint through low-rank compression, token eviction, or flash offloading, but the resulting reconstruction overhead, irreversible token loss, or I/O stalls can offset the benefit of saving memory. We present TierKV, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO). Before decoding starts, PMCO predicts future cache demand from prefill hidden states and jointly assigns tokens to exact, low-rank, and flash-offloaded tiers under the device memory and accuracy budgets. This formulation retains access to the full context, removes the circular dependency of reactive eviction, and admits a closed-form solver that selects tier boundaries and per-layer ranks at runtime. Across eight text, vision, and audio models on three mobile SoCs, TierKV improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, while incurring only minor accuracy degradation.

58. SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

SWE-Proof:语言模型能否用机器检查的证明解决现实世界问题?

AI 总结: 本文提出Benchproofer流水线,将真实代码任务转化为形式化验证问题,构建SWE-Proof基准,证明形式化规范能捕获测试遗漏的错误并提升模型解决率,同时揭示规范忠实性合成的开放挑战。

链接:https://arxiv.org/abs/2609.21190

机构:UC Berkeley(加州大学伯克利分校); Georgia Tech(佐治亚理工学院); UIUC(伊利诺伊大学厄巴纳-香槟分校); AWS AI Labs(亚马逊云科技人工智能实验室)

作者:George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras

英文摘要:Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and increasingly susceptible to memorization. Formal verification avoids both problems, but existing work covers only standalone tasks whose specifications are given as input, not real issues, which touch large repositories and state intent in vague natural language. We present Benchproofer, a pipeline that turns a coding task with a known correct patch into a formally verified one: it writes a specification for the new code, summarizes the existing functions that code calls with axioms, and admits an instance only after mechanical and adversarial gates agree. Applying it to SWE-bench Verified yields SWE-Proof, 500 real issues whose correctness is formally verified rather than tested, and it extends to SWE-bench Pro. Across two frontier models, verification catches what tests miss: a quarter to a half of test-passing patches admit counterexamples, which a structured natural-language specification does not fix, while a correct formal one lifts resolution from 85% to 95% for Opus 4.8. Writing that specification is the hard part: models that must write their own gain nothing over an unaided baseline, and only 62% of their specifications pass our audit. The usual failure is faithfulness, a specification that constrains part of the required behavior and leaves the rest free. Specification quality still tracks the outcome, failing on 89% of unresolved instances against 47% of resolved ones, making faithful specification synthesis a concrete open problem.

59. MIRCID: Inferred Hub-miRNAs Drive Cross-Task Improvements in Drug Mechanistic Modeling

MIRCID:推断的Hub-miRNAs驱动药物机制建模中的跨任务改进

AI 总结:MIRCID框架利用推断的HubmiRs增强药物机制建模,在通路分类和MoA检索中优于TF活性,实现跨任务改进。

链接:https://arxiv.org/abs/2609.21280

机构:The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)); Warshel Institute for Computational Biology(瓦谢尔计算生物研究院); Guangdong Provincial Key Laboratory of Digital Biology and Drug Development(广东省数字生物与药物开发重点实验室); Peking Union Medical College Hospital(北京协和医院); Chinese Academy of Medical Sciences & Peking Union Medical College(中国医学科学院 北京协和医学院)

作者:Xin Cao, Yigang Chen, Jiatong Xu, Ziyue Zhang, Xiang Cheng, Shenyu Wang, Yangyi Zhang, Xiaoxuan Cai, Shidong Cui, Zihao Zhu, Xiang Ji, Hsi-Yuan Huang, Yang-Chi-Dung Lin, Hsien-Da Huang

英文摘要:Drug mechanism-of-action (MoA) modeling commonly relies on perturbational transcriptomes, but matched microRNA (miRNA) measurements are often unavailable. Inferred regulatory features offer a scalable way to reuse these data. Here, we present MIRCID, a framework comparing gene expression with inferred transcription factor (TF) activity and miRNA expression across pathway classification and similarity-based MoA retrieval. HubmiRNet infers 414 pan-cancer hub miRNAs (HubmiRs) from 977 L1000 landmark genes, achieving a Pearson correlation coefficient of 87.72\%; its 1,298-output variant also outperformed SiCmiR on the full-miRNA task (71.21\% versus 67.30\%). In the evaluated comparisons, miRNA augmentation provided more consistent gains than TF activity. Generic embedding controls showed model-dependent utility, while complementarity analyses identified a distinct, partially linearly recoverable representation that retained gene-derived structure. Illustrative rescue cases linked improved classification to biologically plausible miRNA patterns in samples with weak transcriptional signatures. These findings support inferred HubmiRs as a biologically informed recoding of transcriptomic data for perturbational drug modeling, while leaving recovery of measured perturbational miRNA responses to further validation.

60. Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding

多受试者预训练实现封闭语料库表面肌电语音解码的短校准个性化

AI 总结:本研究通过多受试者预训练和检查点初始化,在封闭语料库中实现表面肌电语音解码的短校准个性化,显著降低字符和词错误率。

链接:https://arxiv.org/abs/2609.21288

机构:New York University Tandon School of Engineering(纽约大学坦登工程学院); New York University Grossman School of Medicine(纽约大学格罗斯曼医学院)

作者:Chenqian Le, Beatrice Fumagalli, Yasamin Esmaeili, Xupeng Chen, Tianyu He, Nikasadat Emami, Adeen Flinker, Yao Wang

英文摘要:Surface electromyography (sEMG)-based silent speech interfaces are limited by cross-user variability and calibration burden. We study a limited-data setting in which each of 27 speech-typical participants contributed less than 0.5 h of data (21.3 min on average) across Aloud and Mimed speech. Within a closed 50-sentence corpus, we used leave-one-subject-out evaluation, initializing from a released single-subject checkpoint, pretraining on non-held-out participants, and fine-tuning on the target participant. This pipeline achieved 21.7% character error rate (CER) and 31.9% word error rate (WER), compared with 49.3% CER without target-subject calibration and 68.0% CER for direct checkpoint fine-tuning. Multi-subject pretraining from random initialization followed by fine-tuning reached 44.9% CER and did not converge under the fixed schedule in 5 of 27 folds, indicating substantial optimization and accuracy benefits from checkpoint initialization. Macro-averaged CER declined from 74.4% with one pretraining participant to 21.7% with 26. Three minutes of target-subject calibration achieved 20.5% CER and 31.7% WER, with no statistically significant difference from the full approximately 13-min pool (21.7% CER and 31.9% WER). A subject-specific adapter provided no detectable benefit. Excluding the five evaluation sentences from all sEMG model-training data increased CER and WER to 78.6% and 99.9%. These results support short-calibration personalization in a standardized-montage, closed-corpus setting.

61. Fast And Accurate Text Content File Type Identification

快速且准确的文本内容文件类型识别

AI 总结:本研究提出一种神经网络模型,用于快速准确地识别文本内容文件类型,在开源文件上比 Magika 准确率更高、速度快约四倍且体积小 28%。

链接:https://arxiv.org/abs/2609.21306

机构:CrowdStrike, Inc.(CrowdStrike公司); Univ. of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

作者:Manu Nandan, Michael Brautbar, Edward Raff

英文摘要:A common requirement across organizations is to have a tool that can identify file types based on their contents, particularly in the cybersecurity domain where magic numbers and file extensions can not be trusted. While existing tools work well in practice, there is plenty of room for improvement either in terms of computational load and time for detection in the case of model based tools like Magika or in terms of accuracy of detection in the case of file parsing tools that use programming language constructs. In this study, we propose a neural network model for identification of types of text content files, especially source code, that is more accurate and faster than other available tools. Our experiments on open-source files indicate that it is not only more accurate on average for text-content file-type identification, but also approximately four times faster than Magika, while being 28% smaller in size.

62. An Introduction to Compression-Based Machine Learning

基于压缩的机器学习导论

AI 总结:本文系统梳理压缩与机器学习的双向转换关系,提出并验证基于压缩的机器学习设计框架,在恶意软件任务上表现优于传统基线,准确率提升可达0.62。

链接:https://arxiv.org/abs/2609.21309

机构:University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校); CrowdStrike(CrowdStrike公司); Syracuse University(雪城大学)

作者:John Hurwitz, Edward Raff, Charles K. Nicholas

英文摘要:Any lossless compression algorithm (like gzip) may be converted into a machine learning method, via either Normalized Compression Distance or the Minimum Description Length principle. Any auto-regressive model may be converted into a lossless compression method via entropy coding. This seemingly circular dependence has unrealized potential in modern artificial intelligence and machine learning, and we survey and formalize the various strategies that have been used to leverage compression for machine learning. We introduce and empirically validate a design framework for compression-based ML, finding compression-based methods competitive with conventional baselines and decisively stronger on malware. We find that varying these design choices yields accuracy gains of up to 0.62.

63. Routine Blood Tests Outperform CRP for Distinguishing Bacterial From Viral Infection in Children

常规血液检测在区分儿童细菌性与病毒性感染方面优于CRP

AI 总结:本研究利用906名儿童患者数据,评估CBC对区分病毒与细菌感染的预测价值,发现包含全特征的XGBoost模型(AUC 81.7%)优于仅基于CRP的决策规则,建议抗生素处方决策应综合考虑多种因素。

链接:https://arxiv.org/abs/2609.21332

机构:Vector Labs(Vector 实验室); Zdraveto Hospital(Zdraveto 医院)

作者:Mihaela Demireva, Zhecho Mitev, Djuna Chinareva-Klimentova, Svetoslav Ivanov, Georgi Nalbantov, Dimitar Mitev

英文摘要:Acute infectious diseases are among the leading causes of medical consultations and hospitalizations in children worldwide. These infections are predominantly caused by viruses or bacteria, yet differentiating between the two remains a common clinical challenge. As a result, pediatricians often default to the safer option of prescribing antibiotics contributing to the growing problem of antimicrobial resistance. The objective is to assess the additional predictive value of CBC towards determining the current infection. This retrospective study used data from 906 pediatric patients aged between 2 and 14 years who were tested positive either for viral or bacterial infection between 2022 and 2026. Inclusion criteria further required availability of CBC results and CRP level measurements. These laboratory parameters as well as age were used as input features for several supervised classification models. Model performance was evaluated using AUC, sensitivity and specificity. The best performing model is XGBoost, which included all features, achieving out of-sample performance of AUC of 81.7% and sensitivity of 70.8%, specificity of 79.2%. All trained models outperform a CRP-based only decision-rule model in terms of AUC. We suggest that the decision to prescribe antibiotics should be based on a number of factors, including but not limited to CBC, some of which are not currently incorporated into routine practice.

64. Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation

面向均值-方差投资组合优化的基于KKT重构的决策聚焦学习

AI 总结:针对均值-方差投资组合优化中预测与决策目标不一致的问题,提出基于KKT条件的单层决策聚焦学习公式,显式保留约束,在真实ETF数据上取得最优投资表现。

链接:https://arxiv.org/abs/2609.21427

机构:University of Tsukuba(筑波大学)

作者:Kensei Nosaka, Shunnosuke Ikeda, Yuichi Takano

英文摘要:Mean-variance portfolio optimization (MVO) is a central framework in data-driven asset management. A widely adopted approach is a two-stage framework that first predicts expected returns and then solves the optimization problem based on these predictions, with the predictive models trained by minimizing prediction errors. However, this objective of prediction is not aligned with the quality of the downstream portfolio decision. Decision-focused learning (DFL), which directly minimizes the downstream decision loss within the learning process, has thus emerged as a promising direction. However, existing DFL approaches to MVO rely on surrogate losses or constraint relaxations for tractability, creating a structural mismatch between predictive model training and the constrained MVO solved at evaluation. We propose a single-level optimization formulation that incorporates the Karush-Kuhn-Tucker (KKT) optimality conditions of the lower-level MVO into the upper-level learning problem. This formulation explicitly preserves the budget and short-sale constraints while remaining tractable for standard nonlinear optimization solvers. Rolling-window experiments on real-world ETF (Exchange Traded Funds) data across two asset universes with different correlation structures show that our method achieved the best performance on multiple investment metrics and also demonstrated performance improvement due to the proposed regularization.

65. Optimal Randomized Proper Online Learning

最优随机适定在线学习

AI 总结:本文证明随机适定在线学习的最优期望错误界为 O(L(H) log T),改进了先前 O(L(H) log^6 T) 的结果,并达到常数因子内的最优性。

链接:https://arxiv.org/abs/2609.21445

机构:Rutgers University(罗格斯大学); The Hebrew University(希伯来大学)

作者:Zachary Chase, Idan Mehalel

英文摘要:We prove that the optimal expected mistake bound of online learning a function class $\mathcal{H}$ by a randomized proper learning algorithm is $O(\mathtt{L}(\mathcal{H}) \log T)$, where $\mathtt{L}(\mathcal{H})$ is the Littlestone dimension of $\mathcal{H}$ and $T$ is the time horizon. Our result improves upon the previously best known bound of $O(\mathtt{L}(\mathcal{H}) \log^6 T)$ given by Daskalakis and Golowich (STOC 2022), and is optimal up to a universal constant for worst-case classes.

66. Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals

通过激活引导补偿和正交残差理解LLM量化

AI 总结:本文通过将量化误差分解为激活引导补偿和正交残差,推导出旋转、符号选择和缩放的实际指南,在八个Llama和Mistral模型上实现与SpinQuant竞争的性能。

链接:https://arxiv.org/abs/2609.21450

机构:The University of Tokyo(东京大学)

作者:Yamato Narita, Issei Sato

英文摘要:Post-training weight-activation quantization reduces the memory and inference costs of large language models, but aggressive W4A4 quantization remains difficult because activation outliers degrade effective quantization resolution. Although weight optimization, channel-wise scaling, and orthogonal rotation mitigate this problem, the error components they address and their relationship remain unclear. Using an exact decomposition of local weight-activation quantization error into an activation-guided weight compensation term and an orthogonal residual, we bound the residual using persistent channel-wise outlier and regular activation quantities. This decomposition clarifies which error components can be addressed by weight compensation and which require transformation design. We then use the residual bounds to derive practical guidelines for applying randomized Hadamard rotation, sign selection, and channel scaling. In particular, the analysis explains how random signs suppress constructive interference among persistent outlier channels, how sampling multiple sign patterns can improve transformation selection, and how second-moment balancing leads to an $L_2$ scaling rule while a further relaxation recovers SmoothQuant-style $L_\infty$ scaling. We evaluate these guidelines through backpropagation-free configurations across eight Llama and Mistral models, obtaining performance competitive with gradient-trained SpinQuant.

67. What Must Survive? Exact Task-Information--State Frontiers for Resource-Sufficient Learning

什么必须留存?资源充足学习中的精确任务信息-状态前沿

AI 总结:本文研究下游任务未知时系统压缩所需保留的状态量,提出精确前沿公式,证明最优建议分区为强NP困难,并通过三个实例展示任务信息可大幅减少所需状态。

链接:https://arxiv.org/abs/2609.21523

机构:Kabale University(卡巴莱大学)

作者:Ronald Katende

英文摘要:A system may be compressed before its downstream task is fully known. We ask how much retained state is then necessary and how much can be saved by limited advance task information. For a finite family of linear tasks, a task message is revealed before state formation and the exact task only afterwards. For an advice alphabet of size $K$, the exact frontier is \[ p^*(K)= \min_{\substack{\Pcal\text{ partition of }\U\\|\Pcal|\le K}} \max_{C\in\Pcal}\rank(T_C), \] with the $b$-bit frontier obtained by setting $K=\min(2^b,|\U|)$. Thus advance task information reduces state through partitions whose joint task operators have low rank. We also give an approximate singular-value frontier, a common-core lower bound and exact direct-sum law, and strong NP-hardness of finding an optimal advice partition. The hardness persists at every fixed positive approximation tolerance. Three examples illustrate the result. A well-conditioned softmax attention construction gives an exact $524{,}288\to1{,}024$ coordinate frontier when nine bits resolve one of $512$ continuations. A domain-decomposed digital twin yields an interface-plus-local-state law and a weighted partition problem for heterogeneous regions. A hierarchical multi-task model gives a two-stage frontier in which three bits reduce the required state from $3136$ to $448$ coordinates, with further task information approaching the irreducible $328$-coordinate single-task floor.

68. MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems

MACE:面向多智能体系统的记忆-智能体协同进化与自适应记忆图

AI 总结:MACE提出记忆-智能体协同进化框架,通过执行反馈自适应组织记忆并优化智能体使用,在八个基准上以81.11%平均分超越最强基线SAGE。

链接:https://arxiv.org/abs/2609.21533

作者:Kairui Yang, Minghao An, Xunkai Li, Ziheng Yi, Zekai Chen, Guangyuan He, Rong-Hua Li

英文摘要:LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action's prerequisites and the outputs needed by subsequent agents. Our empirical studies show that grouping these dependencies into functional memory units improves their retention, while connecting units increases retrieval of the units and links jointly required by a task. The preferred combination of units also changes between instructions and checklists, even when each combination's content is fixed across formats. Updating choices from the outcomes of each combination and format pairing outperforms scoring combinations and formats separately. These findings motivate MACE, a memory-agent co-evolution framework that adapts memory organization and agent memory use through execution feedback. Its MemGoG structure represents functional units as subgraphs of related conditions, actions, and outputs, connecting them through support, conflict, and repair relations. MACE Loop selects task-relevant units and relations within a memory budget and provides each agent with instructions or checklists for its current operation. It records the selected units, presentation formats, agent outputs, and task outcomes to update unit scores and relations for retrieval and inform subsequent presentation choices. Across eight benchmarks, MACE outperforms ten baselines with an average score of 81.11%, compared with 78.97% for the strongest baseline, SAGE.

69. On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation

论排斥型与吸引型教师:在自蒸馏中分离正确性与行为

AI 总结:本研究提出对比自蒸馏,结合正确解教师的吸引与错误解教师的排斥,以分离正确性与行为信号,提升推理性能并保持响应长度稳定。

链接:https://arxiv.org/abs/2609.21561

机构:ETH Zurich(苏黎世联邦理工学院); Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)

作者:Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub, Frederike Lübeck, Jonas Hübotter, Thomas Kleine Buening, Andreas Krause

英文摘要:On-policy self-distillation provides dense, token-level supervision by conditioning a model on privileged information and distilling the resulting teacher distribution back into the model. However, privileged information can change not only what the teacher knows, but also how it behaves, entangling correctness-relevant learning signals with unintended behavioral shifts. We study this effect in reasoning tasks by contrasting attractive self-distillation, which moves the model toward a privileged teacher, with repulsive self-distillation, which moves it away from a privileged teacher. We find that both objectives can induce strong and opposing behavioral shifts: attraction suppresses exploratory reasoning and promotes shorter, more confident responses, whereas repulsion increases response length, can trigger unintended switches into a model's latent thinking mode, and ultimately becomes unstable. Motivated by these observations, we study contrastive self-distillation, which combines attraction toward a correct-solution-conditioned teacher with repulsion from an incorrect-solution-conditioned teacher. In contrast to prior work that combines such distillation signals with a GRPO objective, we isolate the self-distillation objective and study its behavior on its own. We find that the shared behavioral shifts of the two teachers largely cancel, leaving a token-level signal that more directly reflects correctness. Across non-thinking, instruct-only, and already-thinking models, this contrastive objective improves reasoning performance while maintaining stable response lengths.

70. Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs

超越高斯世界:潜在几何对JEPA的重要性

AI 总结:本研究将JEPA的线性恢复分析从高斯推广到黎曼流形,证明球面均匀分布可替代高斯实现线性恢复,并给出更紧的近似界,实验验证几何匹配目标更优。

链接:https://arxiv.org/abs/2609.21656

作者:Léo Nicollier (CB, ATT), Enric Meinhardt-Llopis (CB), Marc Pic (ATT), Pablo Musé (CB, IFUMI), Gabriele Facciolo (CB)

英文摘要:Recent Joint-Embedding Predictive Architectures (JEPAs) prevent representation collapse by constraining learned representations to follow a prescribed target distribution, such as an isotropic Gaussian or the uniform distribution on a hypersphere. Klindt et al. (2026) showed that, under their Euclidean assumptions, matching a Gaussian target can recover Gaussian latent variables up to a linear transformation, and that the Gaussian is the unique distribution with this guarantee. We extend their analysis to latent variables supported on embedded Riemannian manifolds and derive conditions on the latent geometry and positive-pair dynamics under which alignment and exact distribution matching guarantee linear recovery. In particular, when the latent variables are uniformly distributed on a sphere and the representations are matched to the same spherical distribution, every optimal representation recovers the latent state up to an orthogonal transformation. This shows that Gaussian uniqueness is not a universal property of distribution-matched JEPAs: non-Euclidean latent geometries can admit other linearly recoverable distributions. We further derive an approximate-recovery bound that is strictly tighter for the spherical world than for the Gaussian world. Experiments on Gaussian, spherical, and toroidal latent spaces show that geometrically compatible targets yield better linear recovery when optimization succeeds, whereas mismatched targets distort the latent structure. This advantage persists in high-dimensional Clifford-torus worlds.

71. Bilevel Optimization of Topology and Hyperparameters (BOTH)

拓扑与超参数的二层优化(BOTH)

AI 总结:本文提出BOTH方法,通过自动微分对拓扑优化过程本身求超梯度,实现超参数与主优化同步调整,可扩展到数千超参数,开销仅相当于几次标准TO运行,并在应力约束和柔度问题上验证有效性。

链接:https://arxiv.org/abs/2609.21758

机构:Delft University of Technology(代尔夫特理工大学); Brown University(布朗大学)

作者:Suryanarayanan Manoj Sanu, Miguel Anibal Bessa, Alejandro Marcos Aragón

英文摘要:Topology optimization (TO) represents a significant step towards automating the design process: given a working simulation, TO can produce a viable prototype at the press of a button by differentiating the simulation and iteratively improving the design. In practice, however, TO is riddled with ``magic numbers''---hyperparameters whose tuning significantly affects the outcome. Finding the right values typically requires not only deep problem-specific knowledge but also extensive trial-and-error. While practitioners can use surrogate-assisted hyperparameter optimization as an alternative, this approach requires strictly limiting the number of hyperparameters through careful problem formulation. Here, we propose differentiating TO itself using automatic differentiation. This yields ``hypergradients'' that allow us to tune these hyperparameters in tandem with the primary optimization. We show that evaluating just one or two steps of TO is sufficiently informative and that the method scales favorably to thousands of hyperparameters at an expense comparable to only a few standard TO runs. We demonstrate this approach on stress-constrained and compliance problems, with the latter utilizing a neural parameterization of the density field.

72. RegKT: Interpretable and Robust Deep Knowledge Tracing With IRT-Regularizer

RegKT:基于IRT正则化器的可解释且稳健的深度知识追踪

AI 总结:针对深度学习知识追踪模型可解释性差和易过拟合的问题,本文提出一种基于IRT正则化的新方法,同时提升模型稳健性与可解释性,以适用于真实教育应用。

链接:https://arxiv.org/abs/2609.21791

机构:Inria-Saclay(法国国家信息与自动化研究所萨克雷中心); University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校); Telecom-Sud Paris(南巴黎电信学院)

作者:Samuel Girard, Juan D. Pinto, Jill-Jênn Vie, Amel Bouzeghoub

英文摘要:As deep learning models continue to advance, knowledge tracing models have achieved higher accuracy. However, these gains come at the cost of reduced interpretability, which is crucial for practitioners in educational settings to adopt new methodologies. Additionally, deep learning models are prone to overfitting, particularly when dealing with the small datasets that are common in educational applications. In this paper, we propose a novel regularization technique designed to enhance the robustness of deep-learning-based knowledge tracing models, while simultaneously improving their interpretability. Our method addresses both the interpretability and overfitting challenges, making it more feasible for real-world educational applications.

73. The Weight Is Over - Interactive Diffusion on Consumer GPUs

重量终结——消费级GPU上的交互式扩散

AI 总结:本文针对消费级GPU上的扩散模型推理,提出嵌入翻译器、扫描配方和交互式编辑器,实现亚秒级首令牌时间,兼顾性能、质量与模型占用。

链接:https://arxiv.org/abs/2609.21849

机构:Adobe(Adobe公司); NVIDIA(英伟达)

作者:Frieder Ganz, Maximilian Müller

英文摘要:On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a decoder, and often further postprocessing that is not as standardized as LLM inference loops are. We navigate the trade-off between performance, quality, and model footprint to reach as many client devices in the wild as possible. We make three contributions: an embedding translator that maps a small text encoder into a large encoder space to cut weight and latency; a reproducible sweep recipe for navigating the speed/quality/memory triangle in diffusion pipelines; and an interactive on-device image generation editor achieving sub-second TTFI on recent GPUs.

74. Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining

几何平均池化用于等权乘法粗粒化

AI 总结:本文提出几何平均池化(GMP),一种结合符号乘积与幅度几何均值的带符号池化算子,用于等权乘法粗粒化,在合成任务上优于平均和最大池化,但图像和分子数据上效果依赖具体设置。

链接:https://arxiv.org/abs/2609.21876

机构:University of Tennessee, Knoxville(田纳西大学诺克斯维尔分校); Rutgers University(罗格斯大学); Google(谷歌)

作者:Ang-Kun Wu, Fangdi Wen, Jingtao Zhang

英文摘要:As an alternative to the additive and extremal biases of average and max pooling, we introduce Geometric Mean Pooling (GMP), a signed pooling operator that combines the product of feature signs with the geometric mean of feature magnitudes. Motivated by local-to-global composition in quantum many-body physics, GMP retains both joint sign information and a characteristic multiplicative scale without introducing learnable pooling parameters. We show that non-overlapping hierarchical GMP preserves the corresponding global multiplicative statistic and evaluate it on synthetic sequence tasks, iterative coarse-graining, image classification, and molecular lipophilicity regression. On the synthetic tasks, GMP recovers product-based signals more accurately than average and max pooling and maintains predictive performance under the tested levels of multiplicative input noise. On image and molecular data, however, its effectiveness depends on the representation, target parameterization, and placement of local and global pooling. These results position GMP as a complementary, regime-dependent inductive bias for tasks in which equal-weight multiplicative composition is plausible, rather than as a universal replacement for standard pooling operators.

75. Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective

从自由能视角检测大型语言模型中的预训练数据

AI 总结:本文提出从自由能视角检测大型语言模型预训练数据,通过引入倾斜边界和熵校正,提出能量转移检测(ETD)方法,显著提升检测性能。

链接:https://arxiv.org/abs/2609.21888

机构:State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(中国科学技术大学认知智能全国重点实验室); Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院); College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院); iFLYTEK AI Research (Central China), iFLYTEK Co., Ltd(科大讯飞股份有限公司(华中)讯飞人工智能研究院)

作者:Chenye Ke, Zirui Liu, Qi Liu, Yan Zhuang, Jintao Zhang, Zhenya Huang, Shijin Wang

英文摘要:Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates prediction loss relative to predictive entropy. Our analysis shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member--non-member separation. We further extend the mean--variance analysis to the more general setting with a nonzero mean entropy gap. Interestingly, this entropy-adjusted score admits a Helmholtz free-energy interpretation, leading to Energy Transfer Detection (ETD), which views pretraining data detection from a macroscopic residual free-energy transfer perspective. Extensive experiments show that ETD achieves the best average detection performance, improving average AUROC by up to 3.5\% and TPR@5\%FPR by up to 5.1\%, while remaining robust across diverse settings.

76. LLMs as Feature Engineers for Text-and-Tabular Prediction

LLMs作为文本与表格预测的特征工程师

AI 总结:本文提出一个迭代框架,利用LLM自动从文本中提取可解释特征用于表格预测,通过错误反馈加速特征发现,并在多数据集上提升性能与可解释性。

链接:https://arxiv.org/abs/2609.21894

机构:Teads

作者:Merwan Barlier, Blaz Skrlj

英文摘要:We introduce an iterative framework that automates the extraction of interpretable, schema-bound categorical features from unstructured text for tabular prediction models. To navigate the feature space, a generator LLM proposes semantic definitions, a separate extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance. We optimize this search by translating explicit model errors, such as AUC ranking inversions, into natural-language feedback, steering the LLM to resolve specific predictive failures. Evaluated across three public datasets, this error-driven loop accelerates feature discovery by up to $3\times$ compared to unguided search. Empirically, the generated features demonstrate strong multi-view complementarity, strictly outperforming any subset when combined with TF-IDF and dense embeddings. Finally, the framework guarantees instance-level interpretability: the discovered features dominate SHAP importance rankings and provide a fully transparent, semantic audit trail for every prediction.

77. Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models

干预粒度至关重要:临床世界模型反事实模拟中的连贯治疗捆绑

AI 总结:本研究通过临床世界模型反事实模拟,发现干预粒度影响模型响应:完整治疗捆绑比单组件编辑更显著改变预测状态,为反事实治疗模拟提供更可靠基础。

链接:https://arxiv.org/abs/2609.21906

机构:Duke University(杜克大学)

作者:Fangzhou Wang, Yixuan Yang, Camilla Balzarotti, Rishikesan Kamaleswaran

英文摘要:Counterfactual simulation with a clinical world model means fixing a patient's history, changing the treatment, and reading off the predicted response. Doing so requires deciding what counts as one intervention. In clinical settings, interventions are documented as bundles: a co-occurrence audit of 945,707 patient-hours from MIMIC-IV shows groups of components, such as every parameter of a dialysis circuit, that never appear apart, so an edit that changes one component on its own describes an hour that never occurs in the data. We hypothesize that the granularity at which an intervention is edited changes how a world model responds, and test this with Clin-JEPA, a latent world model of patient trajectories conditioned on hourly treatment text. At 1,019 documented onsets of invasive ventilation, we keep the patient's history and other treatments fixed and compare editing one ventilator setting with editing the complete configuration recorded for a real patient with the most similar recent trajectory. The complete bundle moves the predicted next state further than any single setting, consistently across all five settings, and the difference remains after accounting for how much each edit changes the model's input. Intervention granularity therefore materially affects the response of a clinical world model: single-component edits may understate treatment sensitivity, and bundle-aware editing may offer a better-supported basis for counterfactual treatment simulation.

78. Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data

基于部分充电数据的锂离子电池剩余使用寿命联合预测与容量估计

AI 总结:提出交叉专家框架,利用部分充电数据联合预测锂离子电池剩余使用寿命与容量,在公开数据集上取得最低RUL误差。

链接:https://arxiv.org/abs/2609.21932

机构:Ton Duc Thang University(孙德胜大学); The University of Danang—University of Science and Technology(岘港大学—科技大学); AIWARE Limited Company(AIWARE有限公司)

作者:Khoa Tran, Ho-Si-Hung Nguyen, Phone Wai Yan Moe, Hung-Cuong Trinh, Thi-Hoang-Giang Tran

英文摘要:Joint remaining useful life (RUL) prediction and capacity estimation require representations of both gradual degradation and recent battery behavior. This paper presents a cross-expert framework using partial-charging measurements without measured historical full-cycle capacity as an input. The RUL Expert encodes nominal 10-min segments from ten cycles sampled within a 30-cycle history using a pretrained gated recurrent unit (GRU) encoder, a two-dimensional convolutional neural network (2D-CNN), and a temporal GRU. The Capacity Expert processes statistical descriptors of nominal 40-min segments from ten consecutive cycles using a 2D-CNN and a Transformer. A feature-wise linear modulation module uses the short-term representation to condition the long-term representation for joint prediction. Training comprises supervised autoencoder pretraining, independent expert pretraining, and fusion training with frozen experts. On two public battery-aging datasets, the reference configuration achieves mean RUL root-mean-square errors of 143.69 and 161.10 cycles and capacity errors of 12.36 and 7.28mAh, respectively. On Dataset I, fusion reduces both mean errors relative to either standalone expert. The results demonstrate a trade-off between RUL and capacity accuracy: the proposed method attains the lowest reported RUL RMSE among the compared methods on both datasets, whereas several baselines yield lower capacity errors.

79. Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction

机器学习临界热流密度模型在CTF子通道程序中方形棒束预测中的评估

AI 总结:本研究评估了基于机器学习的临界热流密度模型在CTF子通道程序中的方形棒束预测性能,发现管束训练的ML模型可有效迁移至棒束应用,局部混合LUT模型表现最佳,显著优于传统方法。

链接:https://arxiv.org/abs/2609.21995

机构:North Carolina State University(北卡罗来纳州立大学); University of Wisconsin–Madison(威斯康星大学麦迪逊分校); Oak Ridge National Laboratory(橡树岭国家实验室)

作者:Aidan Furlong, Vinicius de Melo Monteiro, Robert Salko, Juliana Pacheco Duarte, Xu Wu

英文摘要:The prediction of critical heat flux (CHF), a key safety-related quantity in nuclear thermal hydraulics, remains an important challenge due to its direct relationship with fuel performance and reactor safety. Recent studies have demonstrated that relative to traditional empirical correlations and lookup tables (LUTs), machine learning (ML) methods can substantially improve CHF prediction accuracy. Most ML-based CHF models, however, have been developed and evaluated using tube databases, leaving their applicability to reactor-relevant rod bundle geometries largely unexplored. This study evaluates ML-based CHF models deployed within the CTF subchannel code using the Electric Power Research Institute (EPRI) rod bundle CHF database. Both pure and hybrid residual correction models are considered in local and semilocal formulations. The tube-trained ML CHF models generally transferred favorably to rod bundle applications and outperformed traditional CHF methods across most geometries and operating conditions. The local hybrid LUT model produced the strongest overall performance, and the semilocal pure ML model remained highly competitive. Comparison against the Bowring correlation, W-3 correlation, and 2006 Groeneveld LUT demonstrated that substantial improvements in rod bundle CHF prediction are possible even when models are trained exclusively on tube data. These findings provide one of the first large-scale assessments of ML-based CHF models in square rod bundles within a production-level subchannel analysis environment and support their broader application in reactor thermal hydraulic analysis.

80. Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

软注意力中的弃权与噪声过滤:两个缺失的原语

AI 总结:本研究提出并验证了软注意力缺失的两种原语——弃权(不执行)与噪声过滤,发现其收益随模型规模呈相反趋势,且两者结合在10M至350M参数规模上均达到最优性能。

链接:https://arxiv.org/abs/2609.22005

机构:St. John Fisher University(圣约翰费舍尔大学)

作者:Richard Zhe Wang

英文摘要: Gating the value pathway of attention reportedly improves language model pretraining, and prior studies disagree on why. We argue and provide experimental evidence that such gates supply two different things that softmax attention lacks: abstention and noise filtering. The first is abstention, which allows an attention head to output nothing, bypassing the requirement that attention weights must sum to one. The second is noise filtering, which allows the value pathway of an attention head to suppress interference from superposed features in the residual stream. In our experiments in matched models from 10M to 350M parameters, we supply abstention through a learned per-head sink logit in the softmax and noise filtering through a gate on each value. We report three empirical findings. First, the benefit of abstention, measured as the reduction in validation loss relative to a matched baseline, declines as models grow, whereas the benefit of noise filtering increases with scale. In particular, abstention accounts for nearly all of the gain from gating at 10M and filtering for most of it at 350M. Second, the best model at every scale is the one with both primitives built in. Third, injecting controlled interference into the values a head reads confirms that the gate removes such interference, and reveals that each of the two gate forms we study has a characteristic blind spot. Supplying both primitives adds negligible parameters and remains compatible with the key-value cache.

81. COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules

COMPLEX:多参数持久性模的闭式认证嵌入

AI 总结:COMPLEX提出闭式无训练的多参数持久性模嵌入,首次提供双侧失真界,使特征保真可度量,并在多个基准上达到最先进性能。

链接:https://arxiv.org/abs/2609.22012

机构:George Washington University(乔治华盛顿大学); Montana Technological University(蒙大拿科技大学); University of Ljubljana(卢布尔雅那大学); Institute IMFM(IMFM 研究所)

作者:Sushovan Majhi, Atish Mitra, Žiga Virk, Pramita Bagchi

英文摘要:Every multiparameter persistence vectorization we know of carries a one-sided Lipschitz upper bound and nothing below it: without a lower gauge there is no sense in which the features are faithful, and no per-prediction guarantee can be built on them. This paper supplies the missing side. COMPLEX is a closed-form, training-free embedding of multiparameter modules -- slice the module along a fixed near-diagonal net, embed each slice barcode by the certified PLACE/PALACE landmark map, concatenate. Under a checkable witnessing-slice coherence condition, holding on 100% of audited pairs on Orbit5k, a single slice carries a closed-form lower gauge: separated modules stay separated in the embedding. With the standard upper bound this gives, to our knowledge, the first two-sided distortion bound for a multiparameter feature map, making faithfulness measurable. Measuring it, we find the floor tight within a small factor of realized distances yet operationally local: an RBF-SVM reaches 91% where 1-NN reaches 78% on the same features. Local per-prediction certification therefore fails for a structural reason common to every landmark embedding whose lower gauge is witnessed by one coordinate. With no learned embedding and no held-out calibration -- only a cross-validated SVM head -- COMPLEX sets the state of the art on both Orbit benchmarks (91.95% on Orbit5k, 92.98% on Orbit100k), level with or above Euler-characteristic surfaces and above transformers and graphcode. On graphs it exceeds GRIL on all four shared molecular benchmarks with one fixed configuration, including the only multiparameter method to clear COX2's majority baseline by more than three points. Closed-form selection -- of the landmark radius, the kernel (certificate-preserving), and the bifiltration set -- buys further accuracy; gradient-shaped adaptation buys none.

82. Available Guardrails: Certifying Selective Prediction across ML Systems

可用护栏:跨机器学习系统的选择性预测认证

AI 总结:本文提出通过精确二项反演和动态规划,使选择性预测的认证可用性可计算,并揭示有限数据恢复是核心挑战,从而将认证可用性作为可规划的部署资源。

链接:https://arxiv.org/abs/2609.22048

作者:Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky

英文摘要:A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every reporting unit of interest, such as a tool, policy label, or patient subgroup. The main difficulty is often not whether a granted certificate is valid, but whether finite calibration data can produce one at all. As the gate becomes safer or more fine-grained, some units may receive too little evidence to certify. We make this notion of availability computable through classical exact-binomial inversion and formulate reporting-partition selection, under a fixed group order, as a dynamic program that exposes the trade-off among safety, granularity, and served traffic. The resulting frontier reveals a large population opportunity that finite-sample estimation nearly erases: a truth-informed planner gains $0.157$ mean coverage over support balancing, whereas a naive estimator recovers only $0.005$, making recovery from finite data the central challenge. Constructing candidate partitions on one planning split and selecting among them on another recovers part of this gap, improving mean coverage over support balancing by $0.060$, with the direction reproduced in $59$ of $60$ model effects across three intent-routing datasets and two architectures. A complementary validity-preserving lever, reallocating the familywise error budget across reporting units, recovers additional coverage both with population quantities and noisy estimates. The same frontier recurs, with predictor-specific ceilings, across LLM tool-calling, content moderation, lesion classification, and recommendation. Certified availability is therefore a plannable deployment resource that determines when a safety gate can be certified, at what granularity, and over how much traffic.

83. Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise

粒子竞争与合作用于标签噪声下鲁棒图卷积网络学习

AI 总结:针对GCN对标签噪声敏感的问题,提出PCC+GCN混合框架,利用粒子竞争与合作进行标签精炼,在多个噪声场景下提升准确率并降低计算开销。

链接:https://arxiv.org/abs/2609.22053

机构:São Paulo State University - UNESP(圣保罗州立大学)

作者:Fabricio Breve

英文摘要: Graph Convolutional Networks (GCNs) are highly sensitive to label noise, since corrupted supervision can propagate through the graph and degrade learned node representations. This work proposes PCC+GCN, a hybrid framework that uses Particle Competition and Cooperation (PCC) as a graph-based label-refinement stage before GCN training. PCC identifies suspicious labeled nodes through particle domination dynamics and determines whether their labels should be preserved, removed, or reassigned before GCN training. The framework also allows the graph used by PCC to be augmented with feature-based $k$-nearest-neighbor edges, while the GCN itself is trained on the original graph structure and node features. The proposed method was evaluated on ten graph datasets from the NoisyGL benchmark under conventional Uniform, Pair, and Random label noise, as well as under instance-dependent label noise. A detailed hyperparameter analysis was also conducted on Cora, CiteSeer, and PubMed. Under conventional noise, PCC+GCN achieved the highest overall average accuracy and the best average rank among the evaluated methods, with an average gain of $1.67$ percentage points over the baseline GCN across the clean setting and all noisy scenarios. Under instance-dependent noise, PCC+GCN remained competitive with the best-performing robust methods while requiring substantially lower execution time, being the fastest robust method on eight of the ten datasets. The results indicate that PCC-based label refinement provides an effective and computationally efficient preprocessing strategy for improving GCN robustness under noisy supervision.

Python社区是高质量的Python/Django开发社区
本文地址:http://www.python88.com/topic/201319