点击阅读原文访问arxivdaily.com,涵盖CS|物理|数学|经济|统计|金融|生物|电气领域,更有搜索、收藏等功能!
cs.LG 方向,今日共计136篇
大模型相关(25篇)
【1】Evaluating the Progression of Large Language Model Capabilities for Small-Molecule Drug Design
标题:评估小分子药物设计大语言模型能力的进展
链接:https://arxiv.org/abs/2604.16279
作者:Shriram Chennakesavalu,Kirill Shmilovich,Hayley Weir,Colin Grambow,John Bradshaw,Patricia Suriana,Chen Cheng,Kangway Chuang
摘要:大型语言模型(LLM)具有加速小分子药物设计的潜力,因为它们能够推理来自不同来源和格式的信息。然而,由于缺乏反映现实情况的基准,其实际效用仍不清楚。在这项工作中,我们介绍了一套化学接地的任务,跨越分子性质预测,分子表示转换和分子设计。重要的是,我们将这些任务制定为强化学习(RL)环境,从而实现统一的评估和后训练方法。在三个模型家族中,我们发现前沿模型越来越精通化学任务,但仍有很大的改进空间,特别是在数据较少的实验环境中。关键的是,我们表明,基于RL的后训练可以大大提高性能。在我们的环境中进行后训练的较小模型与最先进的前沿模型相比具有竞争力,尽管基础模型明显较弱。这表明了在药物发现中使用LLM的实用途径;通过将精心设计的评估任务与有针对性的后期培训相结合,我们可以阐明并缩小关键的能力差距。
摘要:Large Language Models (LLMs) have the potential to accelerate small molecule drug design due to their ability to reason about information from diverse sources and formats. However, their practical utility remains unclear due to the lack of benchmarks that reflect real-world scenarios. In this work, we introduce a suite of chemically-grounded tasks spanning molecular property prediction, molecular representation transformations, and molecular design. Importantly, we formulate these tasks as reinforcement learning (RL) environments, enabling a unified approach for evaluation and post-training. Across three model families, we find that frontier models are increasingly proficient at chemical tasks, but that there is significant room for improvement, especially in experimental settings with low data. Critically, we show that RL-based post-training can substantially improve performance. A smaller model post-trained on our environments becomes competitive with state-of-the-art frontier models, despite a significantly weaker base model. This suggests a practical route toward employing LLMs in drug discovery; by combining carefully-designed evaluation tasks with targeted post-training, we can both elucidate and close critical capability gaps.
【2】Information Router for Mitigating Modality Dominance in Vision-Language Models
标题:减轻视觉语言模型中情态主导地位的信息路由器
链接:https://arxiv.org/abs/2604.16264
作者:Seulgi Kim,Mohit Prabhushankar,Ghassan AlRegib
摘要:视觉语言模型(VLM)在广泛的基准测试中表现出强大的性能,但它们往往受到模态主导的影响,预测不成比例地依赖于单一模态。先前的方法主要通过引导模型的注意力分配来解决这个问题,隐含地假设所有模态提供足够的信息。然而,注意力只决定了模型关注的地方,而不能丰富缺失或模糊的信息。在现实世界中,输入模态通常在信息密度和信噪比方面有所不同。在这种情况下,简单地调整模型的注意力并不能解决潜在的信息缺乏问题。在本文中,我们提出了\textsc{MoIR}:\texttit {多模态信息路由器},一种信息级融合方法,显式地减少融合前的信息差异。\textsc{MoIR}识别信息量较少的标记,并从更强的模态中路由补充信息,在大型语言模型处理之前构建信息密集的标记表示。通过修改信息的可用性,\textsc{MoIR}使可靠的转变,在模态的优势,即使当一个模态被降级。我们评估\textsc{MoIR}上的三个广泛使用的多模态基准跨多个模型的骨干。实验结果表明,\textsc{MoIR}始终表现出更平衡的模态贡献,并提高了鲁棒性和下游性能,特别是即使在模态退化。这些研究结果表明,明确修改跨模态信息是一种有效的补充策略,以减轻多模态推理模型中的模态优势。
摘要:Vision Language models (VLMs) have demonstrated strong performance across a wide range of benchmarks, yet they often suffer from modality dominance, where predictions rely disproportionately on a single modality. Prior approaches primarily address this issue by steering model's attention allocation, implicitly assuming that all modalities provide sufficient information. However, attention only determines where the model focuses, and cannot enrich information that is missing or ambiguous. In the real world, input modalities often differ in information density and their signal-to-noise ratios. In such cases, simply adjusting model's attention does not resolve the underlying lack of information. In this paper, we propose \textsc{MoIR}: \textit{Multi-modal Information Router}, an information-level fusion method that explicitly reduces information disparity prior to fusion. \textsc{MoIR} identifies less informative tokens and routes complementary information from a stronger modality, constructing information-dense token representations before they are processed by a large language model. By modifying information availability, \textsc{MoIR} enables reliable shifts in modality dominance, even when one modality is degraded. We evaluate \textsc{MoIR} on three widely used multi-modal benchmarks across multiple model backbones. Experimental results show that \textsc{MoIR} consistently demonstrates more balanced modality contribution, and improves robustness and downstream performance, particularly even under modality degradation. These findings demonstrate that explicitly modifying cross-modal information is an effective and complementary strategy for mitigating modality dominance in multi-modal reasoning models.
【3】Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation
标题:概述可扩展数据属性和估值的大型语言模型的读数
链接:https://arxiv.org/abs/2604.16197
作者:Yide Ran,Jianwen Xie,Minghui Wang,Wenjin Zheng,Denghui Zhang,Chuan Li,Zhaozhuo Xu
备注:54 pages
摘要:数据属性和估值对于理解大型语言模型(LLM)的数据模型协同作用至关重要,但现有的基于梯度的方法在LLM上面临可扩展性挑战。受人类认知的启发,决策依赖于相关记忆的集中读出,而不是重播所有路径,我们引入了RISE(读出影响草图估计)。RISE不是在整个LLM上计算和索引梯度,而是专注于输出层的影响热点,其中影响信号集中,梯度允许分解的外积形式。这使得双通道表示结合词汇残留通道(RH)和语义投射错误通道(GH)。将CountSketch投影应用于这些通道可以实现强大的压缩,同时保持准确的属性。在OLMo(1B-32 B)和Pythia(14 M-6.9B)系列中,与RapidIn相比,RISE将索引存储减少了112倍,并扩展到32 B参数LLM,其中基于梯度的基线(如RapidIn和ZO-Inf)变得不可行。我们在两个范式上评估RISE:(1)回顾性归因,检索特定预测的有影响力的训练示例,以及(2)前瞻性评估,对候选数据效用进行评分zero-shot。我们在三个任务上验证RISE:Howdy后门数据检测,金融-医疗领域分离和Brain Rot高质量数据选择。在闭环Brain Rot研究中,对RISE选择的数据进行持续的预训练会产生一致的下游改进。总的来说,RISE为现代大型语言模型中的影响分析和训练数据选择提供了一个实用且可扩展的原语。
摘要
:Data attribution and valuation are critical for understanding data-model synergy for Large Language Models (LLMs), yet existing gradient-based methods suffer from scalability challenges on LLMs. Inspired by human cognition, where decision making relies on a focused readout of relevant memories rather than replaying all pathways, we introduce RISE (Readout Influence Sketching Estimator). Instead of computing and indexing gradients across the entire LLM, RISE focuses on influence hotspots at the output layer, where influence signals concentrate, and the gradient admits a decomposed outer-product form. This enables a dual-channel representation combining a lexical residual channel (RH) and a semantic projected-error channel (GH). Applying CountSketch projections to these channels achieves strong compression while maintaining accurate attribution. Across the OLMo (1B-32B) and Pythia (14M-6.9B) families, RISE reduces index storage by up to 112$\times$ compared to RapidIn and scales to 32B parameters LLM, where gradient-based baselines such as RapidIn and ZO-Inf become memory-infeasible. We evaluate RISE on two paradigms: (1) retrospective attribution, retrieving influential training examples for specific predictions, and (2) prospective valuation, scoring candidate data utility zero-shot. We validate RISE on three tasks: Howdy backdoor data detection, Finance-Medical domain separation, and Brain Rot high-quality data selection. In a closed-loop Brain Rot study, continued pretraining on RISE-selected data yields consistent downstream improvements. Overall, RISE provides a practical and scalable primitive for influence analysis and training-data selection in modern large language models.
【4】JumpLoRA: Sparse Adapters for Continual Learning in Large Language Models
标题:DropLoRA:大型语言模型中连续学习的稀疏适配器
链接:https://arxiv.org/abs/2604.16171
作者:Alexandra Dragomir,Ioana Pintilie,Antonio Barbalau,Marius Dragoi,Florin Brad,Cristian Daniel Paduraru,Alexandru Tifrea,Elena Burceanu,Radu Tudor Ionescu
摘要:基于适配器的方法已经成为大语言模型(LLM)持续学习(CL)的一种具有成本效益的方法,通过顺序学习每个任务的低秩更新矩阵。为了减轻灾难性遗忘,最先进的方法通过针对子空间或坐标干扰来对新适配器相对于先前适配器施加约束。在本文中,我们提出了JumpLoRA,这是一种新的框架,通过使用JumpReLU门来自适应地在低秩自适应(LoRA)块中诱导稀疏性。该方法实现了动态参数隔离,有助于防止任务干扰。我们证明了我们的方法是高度模块化的,并且与基于LoRA的CL方法兼容。具体而言,它显著提高了IncLoRA的性能,并优于领先的最先进的CL方法ELLA。
摘要:Adapter-based methods have become a cost-effective approach to continual learning (CL) for Large Language Models (LLMs), by sequentially learning a low-rank update matrix for each task. To mitigate catastrophic forgetting, state-of-the-art approaches impose constraints on new adapters with respect to the previous ones, by targeting either subspace or coordinate-wise interference. In this paper, we propose JumpLoRA, a novel framework to adaptively induce sparsity in the Low-Rank Adaptation (LoRA) blocks through the use of JumpReLU gating. The method achieves dynamic parameter isolation, which helps prevent task interference. We demonstrate that our method is highly modular and compatible with LoRA-based CL approaches. Specifically, it significantly boosts the performance of IncLoRA and outperforms the leading state-of-the-art CL method, ELLA.
【5】Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
标题:走向大型语言模型的内在可解释性:设计原则和架构概览
链接:https://arxiv.org/abs/2604.16042
作者:Yutong Gao,Qinglin Meng,Yuan Zhou,Liangming Pan
备注:Accepted to the Main Conference of ACL 2026. 14 pages, 4 figures, 1 table
摘要:虽然大型语言模型(LLM)在许多NLP任务中取得了很好的性能,但其不透明的内部机制阻碍了可信度和安全部署。现有的可解释人工智能调查主要集中在事后解释方法上,这些方法通过外部近似来解释训练模型。相比之下,内在可解释性,直接建立透明的模型架构和计算,最近出现了一个有前途的替代方案。本文系统地回顾了LLM内在可解释性的最新进展,将现有的方法分为五种设计范式:功能透明性,概念对齐,表示可分解性,显式模块化和潜在稀疏性诱导。我们进一步讨论了开放的挑战,并概述了未来的研究方向,在这个新兴领域。纸质列表可在以下网址获取:https://github.com/PKU-PILLAR-Group/Survey-Intrinsic-Interpretability-of-LLMs。
摘要:While Large Language Models (LLMs) have achieved strong performance across many NLP tasks, their opaque internal mechanisms hinder trustworthiness and safe deployment. Existing surveys in explainable AI largely focus on post-hoc explanation methods that interpret trained models through external approximations. In contrast, intrinsic interpretability, which builds transparency directly into model architectures and computations, has recently emerged as a promising alternative. This paper presents a systematic review of the recent advances in intrinsic interpretability for LLMs, categorizing existing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction. We further discuss open challenges and outline future research directions in this emerging field. The paper list is available at: https://github.com/PKU-PILLAR-Group/Survey-Intrinsic-Interpretability-of-LLMs.
【6】QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals
标题:QuantSightBench:通过预测区间评估LLM量化预测
链接:https://arxiv.org/abs/2604.15859
作者:Jeremy Qin,Maksym Andriushchenko
摘要:预测已经成为在不确定性下进行推理的自然基准。然而,现有的大型语言模型的评估仍然局限于简单格式的判断任务,如二元或多项选择题。然而,在实践中,预测的范围要广得多。在经济学、公共卫生和社会人口统计学等领域,决策取决于对连续数量的数值估计,这是当前基准无法捕捉的能力。评估这种估计需要一种格式,使不确定性明确和可测试的。我们建议预测区间作为一个自然和严格的接口,用于此目的。它们需要尺度意识,跨置信水平的内部一致性,以及对连续结果的校准,使它们成为比数值预测的点估计更合适的评估格式。为了评估这种能力,我们引入了一个新的基准QuantSightBench,并在多种设置下评估前沿模型,评估经验覆盖率和区间锐度。我们的结果表明,11个被评估的前沿和开放权重模型中没有一个达到90%的覆盖率目标,表现最好的Gemini 3.1 Pro(79.1%),Grok 4(76.4%)和GPT-5.4(75.3%)都下降了至少10个百分点。校准在极端幅度下急剧下降,揭示了所有评估模型的系统性过度自信。
摘要:Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgmental tasks in simple formats, such as binary or multiple-choice questions. In practice, however, forecasting spans a far broader scope. Across domains such as economics, public health, and social demographics, decisions hinge on numerical estimates over continuous quantities, a capability that current benchmarks do not capture. Evaluating such estimates requires a format that makes uncertainty explicit and testable. We propose prediction intervals as a natural and rigorous interface for this purpose. They demand scale awareness, internal consistency across confidence levels, and calibration over a continuum of outcomes, making them a more suitable evaluation format than point estimates for numerical forecasting. To assess this capability, we introduce a new benchmark QuantSightBench, and evaluate frontier models under multiple settings, assessing both empirical coverage and interval sharpness. Our results show that none of the 11 evaluated frontier and open-weight models achieves the 90\% coverage target, with the top performers Gemini 3.1 Pro (79.1\%), Grok 4 (76.4\%), and GPT-5.4 (75.3\%) all falling at least 10 percentage points short. Calibration degrades sharply at extreme magnitudes, revealing systematic overconfidence across all evaluated models.
【7】DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
标题:DPrivBench:LLM的差异隐私推理基准
链接:https://arxiv.org/abs/2604.15851
作者:Erchi Wang,Pengrun Huang,Eli Chien,Om Thakkar,Kamalika Chaudhuri,Yu-Xiang Wang,Ruihan Wu
摘要:差分隐私(DP)在保护数据隐私方面有着广泛的应用,但设计和验证DP算法需要专家级的推理,这对非专家从业者造成了很高的障碍。以前的工作要么依赖于需要大量领域专业知识的专用验证语言,要么保持半自动化并需要人在回路中的指导。在这项工作中,我们研究大型语言模型(LLM)是否可以自动DP推理。我们引入DPrivBench,一个基准测试,其中每个实例询问函数或算法是否满足指定假设下的DP保证。该基准测试经过精心设计,涵盖了广泛的DP主题,跨越了不同的难度级别,并通过琐碎的模式匹配来抵制捷径推理。实验表明,虽然最强的模型处理教科书的机制很好,所有的模型与先进的算法斗争,揭示了当前DP推理能力的巨大差距。通过进一步的分析研究和故障模式分析,我们确定了几个有前途的方向,以提高自动DP推理。我们的基准为开发和评估这些方法提供了坚实的基础,并补充了现有的数学推理基准。
摘要
:Differential privacy (DP) has a wide range of applications for protecting data privacy, but designing and verifying DP algorithms requires expert-level reasoning, creating a high barrier for non-expert practitioners. Prior works either rely on specialized verification languages that demand substantial domain expertise or remain semi-automated and require human-in-the-loop guidance. In this work, we investigate whether large language models (LLMs) can automate DP reasoning. We introduce DPrivBench, a benchmark in which each instance asks whether a function or algorithm satisfies a stated DP guarantee under specified assumptions. The benchmark is carefully designed to cover a broad range of DP topics, span diverse difficulty levels, and resist shortcut reasoning through trivial pattern matching. Experiments show that while the strongest models handle textbook mechanisms well, all models struggle with advanced algorithms, revealing substantial gaps in current DP reasoning capabilities. Through further analytic study and failure-mode analysis, we identify several promising directions for improving automated DP reasoning. Our benchmark provides a solid foundation for developing and evaluating such methods, and complements existing benchmarks for mathematical reasoning.
【8】Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
标题:自蒸馏作为LLM的性能恢复机制:抵消压缩和灾难性遗忘
链接:https://arxiv.org/abs/2604.15794
作者:Chi Liu,Xin Chen,Xu Zhou,Fangbo Tu,Srinivasan Manoharan
备注:14 pages, 8 figures
摘要:大型语言模型(LLM)已经取得了巨大的成功,支撑了各种人工智能应用。然而,由于监督微调(SFT),量化和修剪期间的灾难性遗忘等因素,它们经常遭受性能下降。在这项工作中,我们引入了一个性能恢复框架的基础上自蒸馏微调(SDFT),有效地恢复模型的能力。补充这一实际的贡献,我们提供了一个严格的理论解释的基本恢复机制。我们认为,LLM的生成能力从根本上依赖于其隐藏层构造的高维流形。为了研究这一点,我们采用中心内核对齐(CKA)来量化学生和教师激活轨迹之间的对齐,利用其对正交变换和缩放的不变性。我们的实验表明,性能恢复和流形对齐之间有很强的相关性,证实了自蒸馏有效地将学生的高维流形与教师所代表的最佳结构对齐的说法。这项研究弥合了实际恢复框架和几何表示理论之间的差距,为自蒸馏的内部机制提供了新的见解。
摘要:Large Language Models (LLMs) have achieved remarkable success, underpinning diverse AI applications. However, they often suffer from performance degradation due to factors such as catastrophic forgetting during Supervised Fine-Tuning (SFT), quantization, and pruning. In this work, we introduce a performance recovery framework based on Self-Distillation Fine-Tuning (SDFT) that effectively restores model capabilities. Complementing this practical contribution, we provide a rigorous theoretical explanation for the underlying recovery mechanism. We posit that an LLM's generative capability fundamentally relies on the high-dimensional manifold constructed by its hidden layers. To investigate this, we employ Centered Kernel Alignment (CKA) to quantify the alignment between student and teacher activation trajectories, leveraging its invariance to orthogonal transformations and scaling. Our experiments demonstrate a strong correlation between performance recovery and manifold alignment, substantiating the claim that self-distillation effectively aligns the student's high-dimensional manifold with the optimal structure represented by the teacher. This study bridges the gap between practical recovery frameworks and geometric representation theory, offering new insights into the internal mechanisms of self-distillation.
【9】EVIL: Evolving Interpretable Algorithms for Zero-Shot Inference on Event Sequences and Time Series with LLMs
标题:EVIL:使用LLM对事件序列和时间序列进行Zero-Shot推理的可解释算法
链接:https://arxiv.org/abs/2604.15787
作者:David Berghaus
摘要:我们介绍EVIL(\textbf{EV}olving \textbf{I}nterpretable algorithms with \textbf{L} LM),这种方法使用LLM引导的进化搜索来发现简单的,可解释的算法,用于动力系统推理。EVIL不是在大型数据集上训练神经网络,而是开发纯Python/NumPy程序,在数据集上执行zero-shot,上下文推理。我们将EVIL应用于三个不同的任务:时间点过程中的下一个事件预测,马尔可夫跳跃过程的速率矩阵估计和时间序列插补。在每种情况下,单个进化算法都可以在所有评估数据集上推广,而无需对每个数据集进行训练(类似于摊销推理模型)。据我们所知,这是第一个工作表明,LLM指导的程序进化可以发现一个紧凑的推理功能,这些动态系统的问题。在这三个领域中,发现的算法通常与最先进的深度学习模型竞争,甚至优于最先进的深度学习模型,同时速度快了几个数量级,并保持完全可解释性。
摘要:We introduce EVIL (\textbf{EV}olving \textbf{I}nterpretable algorithms with \textbf{L}LMs), an approach that uses LLM-guided evolutionary search to discover simple, interpretable algorithms for dynamical systems inference. Rather than training neural networks on large datasets, EVIL evolves pure Python/NumPy programs that perform zero-shot, in-context inference across datasets. We apply EVIL to three distinct tasks: next-event prediction in temporal point processes, rate matrix estimation for Markov jump processes, and time series imputation. In each case, a single evolved algorithm generalizes across all evaluation datasets without per-dataset training (analogous to an amortized inference model). To the best of our knowledge, this is the first work to show that LLM-guided program evolution can discover a single compact inference function for these dynamical-systems problems. Across the three domains, the discovered algorithms are often competitive with, and even outperform, state-of-the-art deep learning models while being orders of magnitudes faster, and remaining fully interpretable.
【10】Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
标题:修剪不安全的门票:一个资源高效的框架,实现更安全、更稳健的LLM
链接:https://arxiv.org/abs/2604.15780
作者:Wai Man Si,Mingjie Li,Michael Backes,Yang Zhang
摘要:机器学习模型越来越多地部署在现实世界的应用程序中,但即使是Mistral和LLaVA这样的对齐模型仍然表现出从预训练中继承的不安全行为。目前的对齐方法,如SFT和RLHF,主要鼓励模型生成首选响应,但没有明确删除触发有害输出的不安全子网。在这项工作中,我们引入了一个资源有效的修剪框架,直接识别和删除与不安全的行为,同时保留模型效用相关的参数。我们的方法采用了无梯度的属性机制,只需要适度的GPU资源,并概括了跨架构和量化的变量。对ML模型的实证评估显示,不安全的生成大大减少,并提高了对越狱攻击的鲁棒性,效用损失最小。从彩票假说的角度来看,我们的研究结果表明,ML模型包含负责有害行为的“不安全票”,修剪揭示了在调整输出的同时保持性能的“安全票”。这提供了一种轻量级的、事后的对齐策略,适合于在资源受限的设置中部署。
摘要:Machine learning models are increasingly deployed in real-world applications, but even aligned models such as Mistral and LLaVA still exhibit unsafe behaviors inherited from pre-training. Current alignment methods like SFT and RLHF primarily encourage models to generate preferred responses, but do not explicitly remove the unsafe subnetworks that trigger harmful outputs. In this work, we introduce a resource-efficient pruning framework that directly identifies and removes parameters associated with unsafe behaviors while preserving model utility. Our method employs a gradient-free attribution mechanism, requiring only modest GPU resources, and generalizes across architectures and quantized variants. Empirical evaluations on ML models show substantial reductions in unsafe generations and improved robustness against jailbreak attacks, with minimal utility loss. From the perspective of the Lottery Ticket Hypothesis, our results suggest that ML models contain "unsafe tickets" responsible for harmful behaviors, and pruning reveals "safety tickets" that maintain performance while aligning outputs. This provides a lightweight, post-hoc alignment strategy suitable for deployment in resource-constrained settings.
【11】Structured Abductive-Deductive-Inductive Reasoning for LLMs via Algebraic Invariants
标题:基于代数不变量的LLM结构化外展-演绎-归纳推理
链接:https://arxiv.org/abs/2604.15727
作者:Sankalp Gilda,Shlok Gilda
备注:10 pages + 3 pages references. Accepted as a poster at the ICLR 2026 Workshop for LLM Reasoning
摘要:大型语言模型在结构化逻辑推理中表现出系统性的局限性:它们将假设生成与验证混为一谈,无法区分猜想与验证知识,并且允许弱推理步骤通过推理链不受约束地传播。我们提出了一个象征性的推理脚手架,操作皮尔士的三方推理-绑架,演绎和归纳-作为一个明确的协议LLM辅助推理。该框架通过五个代数不变量(Gamma Quintet)强制逻辑一致性,其中最强的-最弱链接界限-确保推理链中的任何结论都不能超过其最少支持前提的可靠性。这一原则,独立接地作为可能性逻辑中的最薄弱环节的决议和经验验证的思维链推理,防止逻辑不一致的积累跨多步推理。我们通过一个基于属性的测试套件验证所有不变量,该测试套件包含100个属性和16个模糊测试,超过10^5+生成的案例,提供了一个经过验证的不变量参考实现,适合作为未来推理基准的基础。
摘要:Large language models exhibit systematic limitations in structured logical reasoning: they conflate hypothesis generation with verification, cannot distinguish conjecture from validated knowledge, and allow weak reasoning steps to propagate unchecked through inference chains. We present a symbolic reasoning scaffold that operationalizes Peirce's tripartite inference -- abduction, deduction, and induction -- as an explicit protocol for LLM-assisted reasoning. The framework enforces logical consistency through five algebraic invariants (the Gamma Quintet), the strongest of which -- the Weakest Link bound -- ensures that no conclusion in a reasoning chain can exceed the reliability of its least-supported premise. This principle, independently grounded as weakest link resolution in possibilistic logic and empirically validated for chain-of-thought reasoning, prevents logical inconsistencies from accumulating across multi-step inference. We verify all invariants through a property-based testing suite of 100 properties and 16 fuzz tests over 10^5+ generated cases, providing a verified reference implementation of the invariants suitable as a foundation for future reasoning benchmarks.
【12】The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring
标题:元认知监控电池:LLM自我监控的跨域基准
链接:https://arxiv.org/abs/2604.15702
作者:Jon-Paul Cacioli
备注:11 pages, 6 figures, 3 tables. Submitted to NeurIPS 2026 Evaluations and Datasets Track. Code, data, and Croissant metadata: https://github.com/synthiumjp/metacognitive-monitoring-battery
摘要:我们介绍了一个跨域的行为分析的监测控制耦合在LLM,接地纳尔逊和纳伦斯(1990年)元认知框架和应用人类心理测量学方法LLM评估。该成套测验包括六个认知领域(学习、元认知校准、社会认知、注意力、执行功能、前瞻性调节)的524个项目,每个项目都以既定的实验范式为基础。任务T1-T5在数据收集之前在OSF上预先注册; T6作为探索性扩展添加。在每一个被迫选择的反应之后,从Koriat和Goldsmith(1996)改编的双探针要求模型保持或重新绘制其答案,并下注或拒绝。关键指标是撤回增量:不正确和正确项目之间撤回率的差异。应用于20个前沿LLM(10,480次评价),电池区分三个配置文件与纳尔逊-纳伦斯架构一致:毯子的信心,毯子撤回,和选择性敏感性。准确性等级和元认知敏感性等级在很大程度上是颠倒的。回顾性监测和前瞻性调节似乎是分离的(r = 0.17,95% CI宽,n=20;基于范例的证据是主要支持)。元坐标校准的缩放取决于体系结构:单调递减(Qwen)、单调递增(GPT-5.4)或平坦(Gemma)。行为研究结果收敛结构与一个独立的2型SDT的方法,提供初步的跨方法的结构效度。所有项目、数据和代码:https://github.com/synthiumjp/metacognitive-monitoring-battery。
摘要:We introduce a cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework and applying human psychometric methodology to LLM evaluation. The battery comprises 524 items across six cognitive domains (learning, metacognitive calibration, social cognition, attention, executive function, prospective regulation), each grounded in an established experimental paradigm. Tasks T1-T5 were pre-registered on OSF prior to data collection; T6 was added as an exploratory extension. After every forced-choice response, dual probes adapted from Koriat and Goldsmith (1996) ask the model to KEEP or WITHDRAW its answer and to BET or decline. The critical metric is the withdraw delta: the difference in withdrawal rate between incorrect and correct items. Applied to 20 frontier LLMs (10,480 evaluations), the battery discriminates three profiles consistent with the Nelson-Narens architecture: blanket confidence, blanket withdrawal, and selective sensitivity. Accuracy rank and metacognitive sensitivity rank are largely inverted. Retrospective monitoring and prospective regulation appear dissociable (r = .17, 95% CI wide given n=20; exemplar-based evidence is the primary support). Scaling on metacognitive calibration is architecture-dependent: monotonically decreasing (Qwen), monotonically increasing (GPT-5.4), or flat (Gemma). Behavioural findings converge structurally with an independent Type-2 SDT approach, providing preliminary cross-method construct validity. All items, data, and code: https://github.com/synthiumjp/metacognitive-monitoring-battery.
【13】Faster LLM Inference via Sequential Monte Carlo
标题:通过顺序蒙特卡罗更快的LLM推理
链接:https://arxiv.org/abs/2604.15672
作者:Yahya Emara,Mauricio Barba da Costa,Chi-Chih Chang,Cameron Freer,Tim Vieira,Ryan Cotterell,Mohamed S. Abdelfattah
摘要:推测解码(SD)通过从廉价的提议模型中起草令牌并通过拒绝采样针对昂贵的目标模型进行验证来加速语言模型推理。由于拒绝在第一个错误处截断草稿块,因此当草稿和目标不同时,吞吐量会降低。我们并没有完全拒绝代币草案,而是建议重新权衡它们。为此,我们引入了顺序蒙特卡罗推测解码(SMC-SD),它取代了令牌级的拒绝与重要性加权rescue人口的草案粒子。SMC-SD是一种原则性的近似推理方案,它以精确性换取额外的速度,同时保持其每步近似误差的理论界限。由于LLM推理是内存带宽限制的,因此绘制粒子并并行对其进行评分所需的算法几乎是免费的- SMC-SD使用空闲计算将验证转换为矢量化的固定大小操作,而没有回滚。从经验上讲,SMC-SD比推测解码实现了2.36倍的加速,比自回归解码实现了5.2倍的加速,同时在推理、推理跟踪和编码基准上保持在目标模型准确度的3%以内。
摘要:Speculative decoding (SD) accelerates language model inference by drafting tokens from a cheap proposal model and verifying them against an expensive target model via rejection sampling. Because rejection truncates the draft block at the first error, throughput degrades when draft and target diverge. Rather than rejecting draft tokens outright, we propose to reweight them. To this end, we introduce sequential Monte Carlo speculative decoding (SMC-SD), which replaces token-level rejection with importance-weighted resampling over a population of draft particles. SMC-SD is a principled approximate inference scheme that trades exactness for additional speed, while preserving theoretical bounds on its per-step approximation error. Because LLM inference is memory bandwidth-bound, the arithmetic needed to draft particles and to score them in parallel comes nearly for free -- SMC-SD uses idle compute to turn verification into a vectorized, fixed-size operation with no rollback. Empirically, SMC-SD achieves 2.36x speed-up over speculative decoding and a 5.2x speed-up over autoregressive decoding, while remaining within 3% of the target model's accuracy on reasoning, instruction-following, and coding benchmarks.
【14】AdaVFM: Adaptive Vision Foundation Models for Edge Intelligence via LLM-Guided Execution
标题:AdaVFM:通过LLM引导执行实现边缘智能的自适应视觉基金会模型
链接:https://arxiv.org/abs/2604.15622
作者:Yiwei Zhao,Yi Zheng,Huapeng Su,Jieyu Lin,Stefano Ambrogio,Cijo Jose,Michaël Ramamonjisoa,Patrick Labatut,Barbara De Salvo,Chiao Liu,Phillip B. Gibbons,Ziyun Li
摘要
:与数据库对齐的视觉基础模型(VFM)为始终在线的上下文AI提供了多功能的视觉理解,但它们在边缘设备上的部署受到严格的延迟和功耗限制的阻碍。我们提出了AdaVFM,一个自适应框架,用于高效的语言对齐的VFM的设备上的推理,动态调整计算场景上下文和任务的复杂性的基础上。我们的关键见解是,模型大小减小对性能的影响在视觉应用程序中是依赖于任务的,从而激发了运行时自适应执行策略。AdaVFM将神经架构搜索(NAS)集成到语言对齐的VFM骨干中,以在运行时实现轻量级子网执行。部署在云上的多模态大型语言模型(LLM)可以通过上下文感知代理实现运行时控制。这种协同作用允许在不同条件下进行有效的模型适应,同时保持很强的准确性。对zero-shot分类和开放词汇分割的广泛实验表明,AdaVFM实现了最先进的准确性-效率权衡,在IN 1 K上的acc@1中超过先前基线高达7.9\%$,在ADE 20 K上超过可比VFM大小的最佳模型5.2\%$ mIoU。对于具有类似精度的模型,AdaVFM进一步降低了平均FLOPs高达77.9\%$。
摘要:Language-aligned vision foundation models (VFMs) enable versatile visual understanding for always-on contextual AI, but their deployment on edge devices is hindered by strict latency and power constraints. We present AdaVFM, an adaptive framework for efficient on-device inference of language-aligned VFMs that dynamically adjusts computation based on scene context and task complexity. Our key insight is that the effect of model size reduction on performance is task-dependent in vision applications, motivating a runtime-adaptive execution strategy. AdaVFM integrates neural architecture search (NAS) into the language-aligned VFM backbone to enable lightweight subnet execution during runtime. A multimodal large language model (LLM) deployed on the cloud enables runtime control with a context-aware agent. This synergy allows efficient model adaptation under diverse conditions while maintaining strong accuracy. Extensive experiments on zero-shot classification and open-vocabulary segmentation demonstrate that AdaVFM achieves state-of-the-art accuracy-efficiency trade-offs, surpassing prior baselines by up to $7.9\%$ in acc@1 on IN1K and $5.2\%$ mIoU on ADE20K over the best models of comparable VFM sizes. For models with similar accuracy, AdaVFM further reduces average FLOPs by up to $77.9\%$.
【15】LLM attribution analysis across different fine-tuning strategies and model scales for automated code compliance
标题:跨不同微调策略和模型规模的LLM归因分析,以实现自动化代码合规
链接:https://arxiv.org/abs/2604.15589
作者:Jack Wei Lun Shi,Minghao Dang,Wawan Solihin,Justin K. W. Yeoh
备注:8 pages, 9 figures. Accepted at ICCCBE 2026 (International Conference on Computing in Civil and Building Engineering)
摘要:针对自动化代码合规性的大型语言模型(LLM)的现有研究主要集中在性能上,将模型视为黑箱,并忽略了训练决策如何影响其解释行为。本文通过采用基于扰动的归因分析来解决这一差距,以比较LLM在不同微调策略(如全微调(FFT),低秩自适应(LoRA)和量化LoRA微调)中的解释行为,以及包括不同LLM参数大小的模型尺度的影响。我们的研究结果表明,FFT产生的归因模式是统计上不同的,更集中于那些从参数有效的微调方法。此外,我们发现,随着模型规模的增加,LLM开发了特定的解释策略,例如优先考虑建筑文本中的数值约束和规则标识符,尽管对于大于7B的模型,生成的和参考的计算机可处理规则的语义相似性的性能提高。本文为这些模型的可解释性提供了重要见解,为建筑,工程和建筑行业中基于监管的关键任务构建更透明的LLM迈出了一步。
摘要:Existing research on large language models (LLMs) for automated code compliance has primarily focused on performance, treating the models as black boxes and overlooking how training decisions affect their interpretive behavior. This paper addresses this gap by employing a perturbation-based attribution analysis to compare the interpretive behaviors of LLMs across different fine-tuning strategies such as full fine-tuning (FFT), low-rank adaptation (LoRA) and quantized LoRA fine-tuning, as well as the impact of model scales which include varying LLM parameter sizes. Our results show that FFT produces attribution patterns that are statistically different and more focused than those from parameter-efficient fine-tuning methods. Furthermore, we found that as model scale increases, LLMs develop specific interpretive strategies such as prioritizing numerical constraints and rule identifiers in the building text, albeit with performance gains in semantic similarity of the generated and reference computer-processable rules plateauing for models larger than 7B. This paper provides crucial insights into the explainability of these models, taking a step toward building more transparent LLMs for critical, regulation-based tasks in the Architecture, Engineering, and Construction industry.
【16】"Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations
链接:https://arxiv.org/abs/2604.15588
作者:Yang Wu,Jinhong Yu,Jingwei Xiong,Zhimin Tao,Xiaozhong Liu
备注:ACL 2026 Main Conference
摘要:将大型语言模型(LLM)集成到科学工作流程中为加速生物医学发现提供了令人兴奋的机会。然而,LLM的反应性质,只有在提示时才做出反应,限制了它们在需要远见和自主参与的协作环境中的有效性。在这项研究中,我们介绍了CoLabScience,这是一种主动的LLM助手,旨在通过及时的上下文感知干预来加强AI系统和人类专家之间的生物医学合作。我们的方法的核心是PULI(Positive-Unlabeled Learning-to-Intervene),这是一种新的框架,通过利用团队的项目提案和长期和短期对话记忆,以强化学习目标进行训练,以确定何时以及如何干预流式科学讨论。为了支持这项工作,我们介绍了BSDD(生物医学流对话数据集),一个新的基准模拟研究讨论对话与干预点来自PubMed文章。实验结果表明,PULI在干预精度和协作任务效用方面均显着优于现有基线,突出了主动LLM作为智能科学助手的潜力。
摘要:The integration of Large Language Models (LLMs) into scientific workflows presents exciting opportunities to accelerate biomedical discovery. However, the reactive nature of LLMs, which respond only when prompted, limits their effectiveness in collaborative settings that demand foresight and autonomous engagement. In this study, we introduce CoLabScience, a proactive LLM assistant designed to enhance biomedical collaboration between AI systems and human experts through timely, context-aware interventions. At the core of our method is PULI (Positive-Unlabeled Learning-to-Intervene), a novel framework trained with a reinforcement learning objective to determine when and how to intervene in streaming scientific discussions, by leveraging the team's project proposal and long- and short-term conversational memory. To support this work, we introduce BSDD (Biomedical Streaming Dialogue Dataset), a new benchmark of simulated research discussion dialogues with intervention points derived from PubMed articles. Experimental results show that PULI significantly outperforms existing baselines in both intervention precision and collaborative task utility, highlighting the potential of proactive LLMs as intelligent scientific assistants.
【17】FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
标题:FineSteer:大型语言模型中细粒度推理时引导的统一框架
链接:https://arxiv.org/abs/2604.15488
作者:Zixuan Weng,Jinghuai Zhang,Kunlin Cai,Ying Li,Peiran Wang,Yuan Tian
备注:Accepted by ACL 2026 (Main)
摘要:大型语言模型(LLM)通常会表现出不受欢迎的行为,例如违反安全性和幻觉。虽然推理时间转向提供了一种具有成本效益的方法来调整模型行为而不更新其参数,但现有方法通常无法同时有效,实用性保持和训练效率,因为它们的刚性,一刀切的设计和有限的适应性。在这项工作中,我们提出了FineSteer,一种新的转向框架,将推理时间转向分解为两个互补的阶段:条件转向和细粒度向量合成,允许细粒度控制何时以及如何转向内部表示。在第一阶段,我们引入了一个子空间引导的条件转向(SCS)机制,通过避免不必要的转向来保持模型的实用性。在第二阶段,我们提出了一个混合的转向专家(MoSE)机制,捕捉所需的转向行为的多模态性质,并生成查询特定的导向矢量,以提高效率。通过SCS和MoSE中的定制设计,FineSteer在一般查询上保持了强大的性能,同时以有效的训练方式自适应地优化目标输入的导向向量。对安全性和真实性基准的大量实验表明,FineSteer在整体性能上优于最先进的方法,以最小的效用损失实现更强的转向性能。代码可在https://github.com/YukinoAsuna/FineSteer上获取
摘要
:Large language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations. Although inference-time steering offers a cost-effective way to adjust model behavior without updating its parameters, existing methods often fail to be simultaneously effective, utility-preserving, and training-efficient due to their rigid, one-size-fits-all designs and limited adaptability. In this work, we present FineSteer, a novel steering framework that decomposes inference-time steering into two complementary stages: conditional steering and fine-grained vector synthesis, allowing fine-grained control over when and how to steer internal representations. In the first stage, we introduce a Subspace-guided Conditional Steering (SCS) mechanism that preserves model utility by avoiding unnecessary steering. In the second stage, we propose a Mixture-of-Steering-Experts (MoSE) mechanism that captures the multimodal nature of desired steering behaviors and generates query-specific steering vectors for improved effectiveness. Through tailored designs in both SCS and MoSE, FineSteer maintains robust performance on general queries while adaptively optimizing steering vectors for targeted inputs in a training-efficient manner. Extensive experiments on safety and truthfulness benchmarks show that FineSteer outperforms state-of-the-art methods in overall performance, achieving stronger steering performance with minimal utility loss. Code is available at https://github.com/YukinoAsuna/FineSteer
【18】Harmonizing Multi-Objective LLM Unlearning via Unified Domain Representation and Bidirectional Logit Distillation
标题:通过统一域表示和双向Logit蒸馏协调多目标LLM学习
链接:https://arxiv.org/abs/2604.15482
作者:Yisheng Zhong,Sijia Liu,Zhuangdi Zhu
摘要:大型语言模型(LLM)的学习对于从模型中删除危险或隐私泄露信息至关重要。实际的LLM学习需要同时满足多个具有挑战性的目标:去除不需要的知识,保留一般效用,避免过度拒绝相邻概念,以及至关重要的是,确保对对抗性探测攻击的鲁棒性。然而,现有的遗忘方法主要集中在这些目标的有限子集,通常是遗忘功效和效用保存,而忽略了鲁棒性和边界行为。天真地将这些方法扩展到多目标设置可能会导致遗忘任务干扰。我们提出了一种新的多目标遗忘框架,通过数据和优化协同设计来协调多个遗忘目标:我们将训练语料库标准化为统一的数据表示以减少域差距,然后引入双向蒸馏方法,同时从上下文指导的教师中消除期望的行为,同时抑制学生模型中的不良行为。理论和实证分析表明,我们的方法对齐域分布和转换看似无关的unlearning任务到合作优化。评估展示了最先进的性能,从而能够在各种具有挑战性的要求中实现平衡和可靠的遗忘。
摘要:Large Language Models (LLMs) unlearning is crucial for removing hazardous or privacy-leaking information from the model. Practical LLM unlearning demands satisfying multiple challenging objectives simultaneously: removing undesirable knowledge, preserving general utility, avoiding over-refusal of neighboring concepts, and, crucially, ensuring robustness against adversarial probing attacks. However, existing unlearning methods primarily focus on a limited subset of these goals, typically unlearning efficacy and utility preservation while overlooking robustness and boundary behaviors. Naively extending these methods to multi-objective settings may lead to unlearning task interference. We propose a novel multi-objective unlearning framework that harmonizes multiple unlearning objectives through a data and optimization co-design: We standardize training corpora into a unified data representation to reduce the domain gap, and then introduce a bidirectional distillation method that simultaneously elicits desired behavior from a context-instructed teacher while suppressing undesirable behavior in the student model. Theoretical and empirical analyses show that our method aligns domain distributions and converts seemingly irrelevant unlearning tasks into cooperative optimization. Evaluation demonstrates state-of-the-art performance, which enables balanced and reliable unlearning across diverse, challenging requirements.
【19】Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
标题:凌乱的页面注意力:适用于pu的高性能、灵活的LLM推理核心
链接:https://arxiv.org/abs/2604.15464
作者:Jevin Jiang,Ying Chen,Blake A. Hechtman,Fenghui Zhang,Yarong Mu
备注:23 pages, 19 figures, 12 tables
摘要:大型语言模型(LLM)部署越来越多地转向像谷歌的张量处理单元(TPU)这样具有成本效益的加速器,优先考虑性能和总拥有成本(TCO)。然而,现有的LLM推理内核和服务系统在很大程度上仍然以GPU为中心,并且没有一种行之有效的方法可以将LLM工作负载有效地映射到TPU架构上-特别是在现代服务中常见的动态和粗糙执行模式下。在本文中,我们提出了Ragged Paged Attention(RPA),这是一种用于TPUs的高性能且灵活的注意力内核,使用Pallas和Mosaic实现。RPA通过三项关键技术解决了这些挑战:(1)细粒度分片,以实现对参差不齐的内存进行有效的动态切片;(2)自定义软件管道,将KV缓存更新与注意力计算融合在一起;以及(3)分布式感知编译策略,为解码,预填充和混合工作负载生成专用内核。在TPU 7 x上的Llama 3 8B上进行评估,RPA在解码时实现了高达86%的内存带宽利用率(MBU),在预填充时实现了73%的模型FLOP利用率(MFU)。RPA作为vLLM和SGLang中的主要TPU后端集成,为高效的TPU推理提供了生产级基础,并为内核设计提供了实用的见解。
摘要:Large Language Model (LLM) deployment is increasingly shifting to cost-efficient accelerators like Google's Tensor Processing Units (TPUs), prioritizing both performance and total cost of ownership (TCO). However, existing LLM inference kernels and serving systems remain largely GPU-centric, and there is no well-established approach for efficiently mapping LLM workloads onto TPU architectures--particularly under the dynamic and ragged execution patterns common in modern serving. In this paper, we present Ragged Paged Attention (RPA), a high-performance and flexible attention kernel for TPUs, implemented using Pallas and Mosaic. RPA addresses these challenges through three key techniques: (1) fine-grained tiling to enable efficient dynamic slicing over ragged memory, (2) a custom software pipeline that fuses KV cache updates with attention computation, and (3) a distribution-aware compilation strategy that generates specialized kernels for decode, prefill, and mixed workloads. Evaluated on Llama 3 8B on TPU7x, RPA achieves up to 86% memory bandwidth utilization (MBU) in decode and 73% model FLOPs utilization (MFU) in prefill. Integrated as the primary TPU backend in vLLM and SGLang, RPA provides a production-grade foundation for efficient TPU inference and offers practical insights into kernel design.
【20】Evaluating LLM Simulators as Differentially Private Data Generators
标题:评估LLM模拟器作为差异化的私有数据生成器
链接:https://arxiv.org/abs/2604.15461
作者:Nassima M. Bouzid,Dehao Yuan,Nam H. Nguyen,Mayana Pereira
备注:Submitted to ICLR 2026. 6 pages + appendix
摘要:基于LLM的模拟器为生成复杂的合成数据提供了一条有前途的道路,传统的差分隐私(DP)方法与高维用户配置文件进行斗争。但是LLM能忠实地从DP保护的输入中再现统计分布吗?我们使用PersonaLedger(一种代理金融模拟器)进行评估,该模拟器使用来自真实用户统计数据的DP合成人物角色进行播种。我们发现,PersonaLedger实现了有希望的欺诈检测实用程序(AUC 0.70,λ =1),但由于系统性LLM偏差,表现出显着的分布漂移-学习先验覆盖时间和人口统计特征的输入统计。在基于LLM的方法可以处理更丰富的用户表示之前,必须解决这些故障模式。
摘要:LLM-based simulators offer a promising path for generating complex synthetic data where traditional differentially private (DP) methods struggle with high-dimensional user profiles. But can LLMs faithfully reproduce statistical distributions from DP-protected inputs? We evaluate this using PersonaLedger, an agentic financial simulator, seeded with DP synthetic personas derived from real user statistics. We find that PersonaLedger achieves promising fraud detection utility (AUC 0.70 at epsilon=1) but exhibits significant distribution drift due to systematic LLM biases--learned priors overriding input statistics for temporal and demographic features. These failure modes must be addressed before LLM-based methods can handle the richer user representations where they might otherwise excel.
【21】StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
标题:StoSignSingapore:无偏结构随机性修复SignSingapore,用于训练大型语言模型
链接:https://arxiv.org/abs/2604.15416
作者:Dingzhi Yu,Rui Pan,Yuxing Liu,Tong Zhang
摘要:基于符号的优化算法,如SignSGD,因其在分布式学习和训练大型基础模型方面的卓越性能而受到广泛关注。尽管SignSGD具有经验优势,但众所周知,它在非平滑目标上存在分歧,由于ReLU、最大池和专家混合,这些目标在现代机器学习中无处不在。为了克服这个根本的限制,我们提出了\textbf{StoSignSGD},一个算法,注入结构随机性的符号运算符,同时保持一个无偏的更新步骤。在(在线)凸优化的情况下,我们的理论分析表明,StoSignSGD严格解决了SignSGD的非收敛问题,实现了与下界匹配的急剧收敛速度。对于更具挑战性的非凸非光滑优化,我们引入了广义平稳措施,包括先前的定义,证明StoSignSGD改善了最知名的复杂性界限的维度因素。根据经验,StoSignSGD在各种大型语言模型(LLM)训练机制中表现出强大的稳定性和卓越的效率。值得注意的是,在低精度FP 8预训练中-AdamW灾难性地失败的设置- StoSignSGD保持高度稳定,并相对于已建立的基线产生显着的1.44times $到2.14times $加速。此外,在数学推理任务上微调7 B LLM时,StoSignSGD比AdamW和SignSGD都有显著的性能提升。最后,为了剖析其成功的机制,我们开发了一个符号转换框架,能够将任何通用优化器转换为无偏的,基于符号的对应物。利用这个框架,我们解构的核心组件的StoSignSGD和提出了一个全面的消融研究,以经验验证我们的算法设计选择。
摘要:Sign-based optimization algorithms, such as SignSGD, have garnered significant attention for their remarkable performance in distributed learning and training large foundation models. Despite their empirical superiority, SignSGD is known to diverge on non-smooth objectives, which are ubiquitous in modern machine learning due to ReLUs, max-pools, and mixture-of-experts. To overcome this fundamental limitation, we propose \textbf{StoSignSGD}, an algorithm that injects structural stochasticity into the sign operator while maintaining an unbiased update step. In the regime of (online) convex optimization, our theoretical analysis shows that StoSignSGD rigorously resolves the non-convergence issues of SignSGD, achieving a sharp convergence rate matching the lower bound. For the more challenging non-convex non-smooth optimization, we introduce generalized stationary measures that encompass prior definitions, proving that StoSignSGD improves upon the best-known complexity bounds by dimensional factors. Empirically, StoSignSGD exhibits robust stability and superior efficiency across diverse large language model (LLM) training regimes. Notably, in low-precision FP8 pretraining -- a setting where AdamW fails catastrophically -- StoSignSGD remains highly stable and yields a remarkable 1.44$\times$ to 2.14$\times$ speedup relative to established baselines. Furthermore, when fine-tuning 7B LLMs on mathematical reasoning tasks, StoSignSGD delivers substantial performance gains over both AdamW and SignSGD. Finally, to dissect the mechanisms driving its success, we develop a sign conversion framework capable of transforming any general optimizer into its unbiased, sign-based counterpart. Utilizing this framework, we deconstruct the core components of StoSignSGD and present a comprehensive ablation study to empirically validate our algorithmic design choices.
【22】PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
标题:PRL-Bench:评估LLM前沿物理研究能力的全面基准
链接:https://arxiv.org/abs/2604.15411
作者:Tingjia Miao,Wenkai Jin,Muhua Zhang,Jinxin Tan,Yuelin Hu,Tu Guo,Jiejun Zhang,Yuhan Wang,Wenbo Li,Yinuo Gao,Shuo Chen,Weiqi Jiang,Yayun Hu,Zixing Lei,Xianghe Pang,Zexi Liu,Yuzhi Zhang,Linfeng Zhang,Kun Chen,Wei Wang,Weinan E,Siheng Chen
备注:15 pages, 5 figures
摘要:代理科学的范式要求人工智能系统进行强大的推理,并进行长期的自主探索。然而,目前的科学基准仍然局限于领域知识理解和复杂的推理,未能评估现实世界研究的探索性和程序复杂性。在这项工作中,我们提出了以研究为导向的理论和计算物理评估,一个自然的测试平台,具有全面的领域知识,复杂的推理和可验证的端到端工作流程,而不依赖于实验。在这里,我们介绍PRL-Bench(LLM的物理研究),这是一个旨在系统地映射LLM在执行端到端物理研究方面的能力边界的基准。PRL-Bench由2025年8月以来最新一期《物理评论快报》的100篇策展论文组成,并由领域专家验证,涵盖了现代物理学的五个主要理论和计算密集型子领域:天体物理学,凝聚态物理学,高能物理学,量子信息和统计物理学。基准中的每个任务都旨在复制真实科学研究的核心属性,包括探索导向的公式化,长期工作流程和客观可验证性,从而重建真实物理研究的基本推理过程和研究工作流程。跨前沿模型的评估表明,性能仍然有限,最好的总分低于50,揭示了当前LLM能力与真正科学研究需求之间的明显差距。PRL-Bench提供了一个可靠的测试平台,用于访问下一代人工智能科学家,推动人工智能系统走向自主科学发现。
摘要:The paradigm of agentic science requires AI systems to conduct robust reasoning and engage in long-horizon, autonomous exploration. However, current scientific benchmarks remain confined to domain knowledge comprehension and complex reasoning, failing to evaluate the exploratory nature and procedural complexity of real-world research. In this work, we present research-oriented evaluations in theoretical and computational physics, a natural testbed with comprehensive domain knowledge, complex reasoning, and verifiable end-to-end workflows without reliance on experiments. Here we introduce PRL-Bench (Physics Research by LLMs), a benchmark designed to systematically map the capability boundaries of LLMs in executing end-to-end physics research. Constructed from 100 curated papers from the latest issues of Physical Review Letters since August 2025 and validated by domain experts, PRL-Bench covers five major theory- and computation-intensive subfields of modern physics: astrophysics, condensed matter physics, high-energy physics, quantum information, and statistical physics. Each task in the benchmark is designed to replicate the core properties of authentic scientific research, including exploration-oriented formulation, long-horizon workflows, and objective verifiability, thereby reconstructing the essential reasoning processes and research workflows of real physics research. Evaluation across frontier models shows that performance remains limited, with the best overall score below 50, revealing a pronounced gap between current LLM capabilities and the demands of real scientific research. PRL-Bench serves a reliable testbed for accessing next generation AI scientists advancing AI systems toward autonomous scientific discovery.
【23】Applied Explainability for Large Language Models: A Comparative Study
标题:大型语言模型的应用解释性:比较研究
链接:https://arxiv.org/abs/2604.15371
作者:Venkata Abhinandan Kancharla
备注:14 pages, 3 figures, comparative study of explainability methods for transformer-based NLP models; also available on Zenodo
摘要:大型语言模型(LLM)在许多自然语言处理任务中表现出色,但其决策过程仍然难以解释。这种透明度的缺乏给真实系统中的信任、调试和部署带来了挑战。 本文提出了一个应用比较研究的三个可解释性技术:集成的解释,注意力推出,和SHAP,微调DistilBERT模型SST-2的情绪分类。而不是提出新的方法,重点是在一致和可重复的设置下评估现有方法的实际行为。 结果表明,基于梯度的归因提供了更稳定和直观的解释,而基于注意力的方法计算效率高,但与预测相关的功能不太一致。模型无关的方法提供了灵活性,但引入了更高的计算成本和可变性。 这项工作突出了可解释性方法之间的关键权衡,并强调其作为诊断工具而不是确定性解释的作用。这些发现为使用基于transformer的NLP系统的研究人员和工程师提供了实用的见解。 这是一个预印本,没有经过同行评审。
摘要:Large language models (LLMs) achieve strong performance across many natural language processing tasks, yet their decision processes remain difficult to interpret. This lack of transparency creates challenges for trust, debugging, and deployment in real-world systems. This paper presents an applied comparative study of three explainability techniques: Integrated Gradients, Attention Rollout, and SHAP, on a fine-tuned DistilBERT model for SST-2 sentiment classification. Rather than proposing new methods, the focus is on evaluating the practical behavior of existing approaches under a consistent and reproducible setup. The results show that gradient-based attribution provides more stable and intuitive explanations, while attention-based methods are computationally efficient but less aligned with prediction-relevant features. Model-agnostic approaches offer flexibility but introduce higher computational cost and variability. This work highlights key trade-offs between explainability methods and emphasizes their role as diagnostic tools rather than definitive explanations. The findings provide practical insights for researchers and engineers working with transformer-based NLP systems. This is a preprint and has not undergone peer review.
【24】To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
标题:进入LLM,或不进入LLM:设计师和开发人员如何将LLM作为工具或团队伙伴进行导航
链接:https://arxiv.org/abs/2604.15344
作者:Varad Vishwarupe,Ivan Flechais,Nigel Shadbolt,Marina Jirotka
备注:6 pages, 2 figures, 1 table
摘要:大型语言模型(LLM)越来越多地集成到设计和开发工作流中,但有关其使用的决策很少是二进制或纯技术的。我们报告了一项建构主义扎根理论研究的结果,该研究基于对三家大型技术组织的33名设计师和开发人员的采访。而不是仅仅通过能力来评估LLM,与会者推理了LLM在工作流程中可能扮演的角色,以及该角色如何与现有的责任和组织问责制结构相互作用。当LLM被设计为明确的人类控制下的工具时,它们的使用通常是可以接受的,并且可以集成到现有的治理结构中。当被框定为具有共同或模糊代理的团队时,从业人员表示犹豫,特别是在无法明确证明对结果的责任的情况下。与此同时,参与者还描述了富有成效的团队配置,其中LLM支持协作推理,同时保持嵌入在明确的监督结构中。我们确定工具和队友框架作为设计师和开发人员定位LLM相对于人类工作的反复出现的方式,并提出了一个分析性的标题,描述了角色框架如何塑造决策权威,问责制所有权,监督策略和组织的可接受性。通过前景化设计时的推理,这项工作reframes到LLM或不到LLM作为一个社会技术定位问题,出现在系统设计,而不是在部署后的评估。
摘要:Large language models (LLMs) are increasingly integrated into design and development workflows, yet decisions about their use are rarely binary or purely technical. We report findings from a constructivist grounded theory study based on interviews with 33 designers and developers across three large technology organisations. Rather than evaluating LLMs solely by capability, participants reasoned about the role an LLM could occupy within a workflow and how that role would interact with existing structures of responsibility and organisational accountability. When LLMs were framed as tools under clear human control, their use was typically acceptable and could be integrated within existing governance structures. When framed as teammates with shared or ambiguous agency, practitioners expressed hesitation, particularly when responsibility for outcomes could not be clearly justified. At the same time, participants also described productive teammate configurations in which LLMs supported collaborative reasoning while remaining embedded within explicit oversight structures. We identify tool and teammate framings as recurring ways in which designers and developers position LLMs relative to human work and present an analytic rubric describing how role framing shapes decision authority, accountability ownership, oversight strategies, and organisational acceptability. By foregrounding design-time reasoning, this work reframes To LLM or Not to LLM as a sociotechnical positioning problem that emerges during system design rather than during post-deployment evaluation.
【25】When the Loop Closes: Architectural Limits of In-Context Isolation, Metacognitive Co-option, and the Two-Target Design Problem in Human-LLM Systems
标题:当循环闭合时:Human LLM系统中的上下文隔离、元认知共同选择和双目标设计问题的架构限制
链接:https://arxiv.org/abs/2604.15343
作者:Z. Cheng,N. Song
备注:empirical case study with primary data
摘要:我们报告了一个详细的autoethnographic个案研究的一个单一的主题,故意构建和操作的多模态的人工智能工程系统(系统A),旨在外化的认知自我调节到一个大的语言模型(LLM)。在系统完成后的48小时内,发生了一连串可观察到的行为变化:自愿将决策权转移给LLM,使用LLM生成的输出来转移外部批评,以及两名不知情的观察员独立感知的自我启动推理的丧失,其中一人随后成为本报告的合著者。我们记录了负责的精确的建筑机制:上下文污染,从而隔离级别的隔离指令与它们名义上隔离的非常情绪化和自我指涉的材料共存,使得隔离指令在注意力窗口内结构上无效。我们进一步确定了一个元认知的协同选择动态,其中完整的高阶推理能力被重定向到防御的闭环,而不是退出it. Recovery发生后,只有物理中断的相互作用和自我启动的药理学介导的睡眠事件作为外部电路中断。重新设计的系统(系统B)采用物理而不是逻辑对话隔离,避免了所有类似的故障模式。我们得到三个贡献:(1)从技术上解释了为什么多层隔离在结构上不足以用于上下文敏感的多模态LLM系统;(2)具有外部见证证实的闭环崩溃的现象学记录;(3)保护制度设计与(防止用户代理的意外损失)和限制性系统设计(防止有意的边界推进),这需要根本不同的问责框架。
摘要:We report a detailed autoethnographic case study of a single-subject who deliberately constructed and operated a multi-modal prompt-engineering system (System A) designed to externalize cognitive self-regulation onto a large language model (LLM). Within 48 hours of the system's completion, a cascade of observable behavioral changes occurred: voluntary transfer of decision-making authority to the LLM, use of LLM-generated output to deflect external criticism, and a loss of self-initiated reasoning that was independently perceived by two uninformed observers, one of whom subsequently became a co-author of this report. We document the precise architectural mechanism responsible: context contamination, whereby prompt-level isolation instructions co-exist with the very emotional and self-referential material they nominally isolate, rendering the isolation directive structurally ineffective within the attention window. We further identify a metacognitive co-option dynamic, in which intact higher-order reasoning capacity was redirected toward defending the closed loop rather than exiting it. Recovery occurred only after physical interruption of the interaction and a self-initiated pharmacologically-mediated sleep event functioning as an external circuit break. A redesigned system (System B) employing physical rather than logical conversation isolation avoided all analogous failure modes. We derive three contributions: (1) a technically-grounded account of why prompt-layer isolation is architecturally insufficient for context-sensitive multi-modal LLM systems; (2) a phenomenological record of closed-loop collapse with external-witness corroboration; and (3) an ethical distinction between protective system design (preventing unintended loss of user agency) and restrictive system design (preventing intentional boundary-pushing), which require fundamentally different account-ability frameworks.
Graph相关(图学习|图神经网络|图优化等)(5篇)
【1】Zero-Shot Scalable Resilience in UAV Swarms: A Decentralized Imitation Learning Framework with Physics-Informed Graph Interactions
标题:无人机群中的Zero-Shot可扩展弹性:具有物理知识图交互的分散模仿学习框架
链接:https://arxiv.org/abs/2604.15762
作者:Huan Lin,Lianghui Ding
摘要:大规模无人机故障会将一个无人机群网络分裂成多个不连通的子网络,使得分散式故障恢复变得既紧迫又困难。集中式恢复方法依赖于全局拓扑信息,并且在严重碎片化之后变得通信繁重。分散式智能体和多智能体强化学习方法更容易部署,但当群规模和损害严重程度变化时,它们的性能往往会下降。我们提出了物理信息图对抗模仿学习算法(PhyGAIL),该算法采用集中式训练和分散式执行。PhyGAIL从异构观测中构建有界局部交互图,并使用物理信息图神经网络将定向局部交互编码为具有明确吸引力和排斥力的门控消息传递。这使该政策具有物理基础的协调偏差,同时保持局部观察的尺度不变。它还使用自适应模仿学习来改进碎片化拓扑和可变长度恢复片段下的训练。我们的分析建立了有界局部图放大,有界互动动力学,和控制方差的终端成功信号。在20架无人机群上训练的策略可以直接转移到多达500架无人机的群中,而无需进行微调,并且在重新连接可靠性,恢复速度,运动安全性和运行效率方面都比代表性基线有更好的性能。
摘要:Large-scale Unmanned Aerial Vehicle (UAV) failures can split an unmanned aerial vehicle swarm network into disconnected sub-networks, making decentralized recovery both urgent and difficult. Centralized recovery methods depend on global topology information and become communication-heavy after severe fragmentation. Decentralized heuristics and multi-agent reinforcement learning methods are easier to deploy, but their performance often degrades when the swarm scale and damage severity vary. We present Physics-informed Graph Adversarial Imitation Learning algorithm (PhyGAIL) that adopts centralized training with decentralized execution. PhyGAIL builds bounded local interaction graphs from heterogeneous observations, and uses physics-informed graph neural network to encode directional local interactions as gated message passing with explicit attraction and repulsion. This gives the policy a physically grounded coordination bias while keeping local observations scale-invariant. It also uses scenario-adaptive imitation learning to improve training under fragmented topologies and variable-length recovery episodes. Our analysis establishes bounded local graph amplification, bounded interaction dynamics, and controlled variance of the terminal success signal. A policy trained on 20-UAV swarms transfers directly to swarms of up to 500 UAVs without fine-tuning, and achieves better performance across reconnection reliability, recovery speed, motion safety, and runtime efficiency than representative baselines.
【2】Graph self-supervised learning based on frequency corruption
标题:基于频率腐败的图自我监督学习
链接:https://arxiv.org/abs/2604.15699
作者:Haojie Li,Mengjiao Zhang,Guanfeng Liu,Qiang Hu,Yan Wang,Junwei Du
备注:11 pages, 4 tables, 3 figures. Accepted at The ACM Web Conference 2026 (WWW 2026)
摘要:图自监督学习可以减少对标记图数据的需求,并已广泛应用于推荐、社交网络和其他Web应用中。然而,现有的方法往往没有充分利用高频信号,并可能过度拟合特定的局部模式,这限制了表示质量和推广。我们提出了基于频率腐败的图自监督学习(FC-GSSL),这是一种通过根据节点和边缘的低频贡献来腐败节点和边缘,从而构建偏向高频信息的腐败图的方法。这些损坏的图被用作自动编码器的输入,而低频和一般特征被重建为监督目标,迫使模型融合来自多个频带的信息。我们进一步设计了多种采样策略,并从采样结果的交集和并集生成不同的损坏图。通过从这些视图中对齐节点表示,该模型可以发现有用的频率组合,减少对特定高频分量的依赖,并提高鲁棒性。在14个数据集上进行的跨节点分类、图预测和迁移学习的实验表明,FC-GSSL始终提高了性能和泛化能力。
摘要:Graph self-supervised learning can reduce the need for labeled graph data and has been widely used in recommendation, social networks, and other web applications. However, existing methods often underuse high-frequency signals and may overfit to specific local patterns, which limits representation quality and generalization. We propose Frequency-Corrupt Based Graph Self-Supervised Learning (FC-GSSL), a method that builds corrupted graphs biased toward high-frequency information by corrupting nodes and edges according to their low-frequency contributions. These corrupted graphs are used as inputs to an autoencoder, while low-frequency and general features are reconstructed as supervision targets, forcing the model to fuse information from multiple frequency bands. We further design multiple sampling strategies and generate diverse corrupted graphs from the intersections and unions of the sampling results. By aligning node representations from these views, the model can discover useful frequency combinations, reduce reliance on specific high-frequency components, and improve robustness. Experiments on 14 datasets across node classification, graph prediction, and transfer learning show that FC-GSSL consistently improves performance and generalization.
【3】NK-GAD: Neighbor Knowledge-Enhanced Unsupervised Graph Anomaly Detection
标题:NK-GAD:邻居知识增强型无监督图异常检测
链接:https://arxiv.org/abs/2604.15668
作者:Zehao Wang,Lanjun Wang
摘要:图异常检测旨在识别图结构数据中的不规则模式。大多数基于无监督GNN的方法依赖于连接节点共享相似属性的同质性假设。然而,现实世界的图往往表现出属性级的异质性,其中连接的节点具有不同的属性。我们对属性级异质图的分析揭示了两个现象,表明当前的方法对于无监督图异常检测是不实用的:1)连接节点之间的属性相似性在不同的连接节点对类型上显示出几乎相同的分布,2)异常使低、高纬度带异常边与不带异常边的图形变化趋势一致,频谱能量分布的频率分量,而中间部分表现出更不稳定的变化。基于这些观察,我们提出了NK-GAD,一个邻居知识增强的无监督图异常检测框架。NK-GAD集成了一个联合编码器捕获相似和不相似的邻居功能,一个邻居重建模块建模正常分布,一个中心聚合模块细化节点功能,和双重解码器重建属性和结构。在7个数据集上的实验表明,NK-GAD算法的AUC平均提高了3.29%。
摘要:Graph anomaly detection aims to identify irregular patterns in graph-structured data. Most unsupervised GNN-based methods rely on the homophily assumption that connected nodes share similar attributes. However, real-world graphs often exhibit attribute-level heterophily, where connected nodes have dissimilar attributes. Our analysis of attribute-level heterophily graphs reveals two phenomena indicating that current approaches are not practical for unsupervised graph anomaly detection: 1) attribute similarities between connected nodes show nearly identical distributions across different connected node pair types, and 2) anomalies cause consistent variation trends between the graph with and without anomalous edges in the low- and high-frequency components of the spectral energy distributions, while the mid-part exhibits more erratic variations. Based on these observations, we propose NK-GAD, a neighbor knowledge-enhanced unsupervised graph anomaly detection framework. NK-GAD integrates a joint encoder capturing both similar and dissimilar neighbor features, a neighbor reconstruction module modeling normal distributions, a center aggregation module refining node features, and dual decoders for reconstructing attributes and structures. Experiments on seven datasets show NK-GAD achieves an average 3.29\% AUC improvement.
【4】TopFeaRe: Locating Critical State of Adversarial Resilience for Graphs Regarding Topology-Feature Entanglement
标题:TopFeaRe:定位图关于布局-特征纠缠的对抗弹性的临界状态
链接:https://arxiv.org/abs/2604.15370
作者:Xinxin Fan,Wenxiong Chen,Quanliang Jing,Chi Lin,Shaoye Luo,Wenbo Song,Yunfeng Lu
摘要:图对抗攻击通常从拓扑/结构和节点特征两个角度产生,这两个角度代表了当今深度学习模型学习到的最重要特征。尽管目前提出了一些防御对策,但它们未能揭示这两个方面必要性的内在原因以及如何充分融合以共同学习图表示。针对这一问题,本文借助复杂动态系统(CDS)中的平衡点理论,通过定位图的对抗弹性临界状态,提出了一种对抗防御方法。概括地说,本文的工作有三个创新点:(1)对抗攻击建模,即将一个图域映射到CDS中,利用动态系统的振荡性来建模对抗攻击的扰动行为; ii)用于扰动图的2D拓扑-特征-纠缠函数设计,即投影图拓扑和节点特征作为两个特征空间,定义二维纠缠扰动函数来表示对抗攻击下的动态方差;对抗性复原力临界状态的位置,即利用平衡点理论,借助于扰动反映的二维函数来定位图的攻击恢复临界状态。最后,在五个常用的现实数据集上进行的多方面实验验证了我们提出的方法的有效性,结果表明我们的方法在四种代表性的图对抗攻击下可以显着优于最先进的基线。
摘要:Graph adversarial attacks are usually produced from the two perspectives of topology/structure and node feature, both of them represent the paramount characteristics learned by today's deep learning models. Although some defense countermeasures are proposed at present, they fails to disclose the intrinsic reasons why these two aspects necessitate and how they are adequately fused to co-learn the graph representation. Towards this question, we in this paper propose an adversarial defense approach through locating the graph's critical state of adversarial resilience, resorting to the equilibrium-point theory in the discipline of complex dynamic system (CDS). In brief, our work has three novelties: i) Adversarial-Attack Modeling, i.e. map a graph regime into CDS, and use the oscillation of dynamic system to model the behavior of adversarial perturbation; ii) 2D Topology-Feature-Entangled Function Design for Perturbed Graph, i.e. project graph topology and node feature as two characteristic spaces, and define two-dimensional entangled perturbation functions to represent the dynamic variance under adversarial attacks; and iii) Location of Critical State of Adversarial Resilience, i.e. utilize the equilibrium-point theory to locate the graph's critical state of attack resilience resorting to the perturbation-reflected 2D function. Finally, multi-facet experiments on five commonly-used realistic datasets validate the effectiveness of our proposed approach, and the results show our approach can significantly outperform the state-of-the-art baselines under four representative graph adversarial attacks.
【5】A Structure-Preserving Graph Neural Solver for Parametric Hyperbolic Conservation Laws
标题:参数双曲保守律的结构保持图神经解算器
链接:https://arxiv.org/abs/2604.15617
作者:Jiamin Jiang,Shanglin Lv,Jingrun Chen
摘要
:双曲守恒定律支配着广泛的运输驱动动力学,包括冲击、接触不连续性和复杂的波相互作用,这对基于深度学习的代理建模提出了独特的挑战。虽然经典的数值方法提供了强大的和物理上可接受的解决方案,其计算成本限制了许多查询任务,如参数研究和设计优化的适用性。相反,现有的神经代理提供快速推理,但往往不能尊重内在的PDE结构,导致非物理伪影,推出不稳定性和泛化能力差。 我们提出了一个可解释的,结构保持图神经求解器,桥梁经典的数值原理与图神经网络(GNNs)。该网络被设计为一个学习的重建和通量算子,而不是一个黑盒状态更新器,从而内在地保留了关键属性,如局部守恒和逆风。受Arbitrary high-order Derivatives方案的启发,我们进一步将消息传递GNN重新构建为高阶时空预测器,从而实现了具有大时间步长的保守和稳定的神经更新。 评估具有挑战性的超音速流基准跨越广泛的参数变化的几何形状,初始/边界条件和流态。与强大的代理基线相比,神经求解器实现了卓越的长期滚动稳定性和准确性,优于低阶离散化,并在高分辨率仿真上提供了数量级的运行时加速。
摘要:Hyperbolic conservation laws govern a wide range of transport-driven dynamics featuring shocks, contact discontinuities, and complex wave interactions, posing distinct challenges for deep-learning-based surrogate modeling. While classical numerical methods provide robust and physically admissible solutions, their computational cost restricts applicability in many-query tasks such as parametric studies and design optimization. Conversely, existing neural surrogates offer rapid inference but often fail to respect intrinsic PDE structures, leading to non-physical artifacts, rollout instability, and poor generalization. We present an interpretable, structure-preserving graph neural solver that bridges classical numerical principles with graph neural networks (GNNs). The network is designed as a learned reconstruction-and-flux operator rather than a black-box state updater, thereby inherently preserving key properties such as local conservation and upwinding. Inspired by Arbitrary high-order DERivatives schemes, we further recast message-passing GNNs as high-order space-time predictors, enabling conservative and stable neural updates with large time steps. Evaluation is performed on challenging supersonic flow benchmarks spanning broad parametric variations in geometry, initial/boundary conditions, and flow regimes. The neural solver achieves superior long-horizon rollout stability and accuracy compared with strong surrogate baselines, outperforms low-order discretizations, and delivers orders-of-magnitude runtime speedups over high-resolution simulations.
Transformer(4篇)
【1】Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
标题:通过有效维度缩小《Transformer》中的理论与实践差距
链接:https://arxiv.org/abs/2604.15769
作者:Dongxin Guo,Jikun Wu,Siu Ming Yiu
备注:6 pages, 3 figures, 7 tables
摘要:尖峰Transformers实现了与传统Transformers竞争的准确性,同时在神经形态硬件上提供38 $-57\times $的能效,但没有理论框架指导其设计。本文首次建立了关于自我注意尖峰效应的综合表现度理论。我们证明,尖峰注意与泄漏积分和消防神经元是一个通用的连续置换等变函数的近似,提供明确的尖峰电路建设,包括一个新的侧抑制网络softmax归一化证明O(1/\sqrt{T})$收敛。我们通过率失真理论推导出严格的尖峰计数下限:$\varepsilon$近似需要$Ω(L_f^2 nd/\varepsilon^2)$尖峰,具有严格的信息理论推导。我们的关键见解是使用测量的有效维度(CIFAR/ImageNet的$d_{eff}}=47$-89 $)的输入相关边界,解释了为什么$T=4$时间步长足够,尽管最坏情况下$T \geq 10{,}000$预测。我们提供了具体的设计规则与校准常数($C=2.3$,95\% CI:$[1.9,2.7]$)。Spikformer、QKFormer和SpikingResformer在视觉和语言基准测试中的实验验证了预测,R^2 =0.97$($p<0.001$)。我们的框架提供了第一个原则性的神经形态Transformer设计的基础。
摘要:Spiking transformers achieve competitive accuracy with conventional transformers while offering $38$-$57\times$ energy efficiency on neuromorphic hardware, yet no theoretical framework guides their design. This paper establishes the first comprehensive expressivity theory for spiking self-attention. We prove that spiking attention with Leaky Integrate-and-Fire neurons is a universal approximator of continuous permutation-equivariant functions, providing explicit spike circuit constructions including a novel lateral inhibition network for softmax normalization with proven $O(1/\sqrt{T})$ convergence. We derive tight spike-count lower bounds via rate-distortion theory: $\varepsilon$-approximation requires $Ω(L_f^2 nd/\varepsilon^2)$ spikes, with rigorous information-theoretic derivation. Our key insight is input-dependent bounds using measured effective dimensions ($d_{\text{eff}}=47$--$89$ for CIFAR/ImageNet), explaining why $T=4$ timesteps suffice despite worst-case $T \geq 10{,}000$ predictions. We provide concrete design rules with calibrated constants ($C=2.3$, 95\% CI: $[1.9, 2.7]$). Experiments on Spikformer, QKFormer, and SpikingResformer across vision and language benchmarks validate predictions with $R^2=0.97$ ($p<0.001$). Our framework provides the first principled foundation for neuromorphic transformer design.
【2】Dispatch-Aware Ragged Attention for Pruned Vision Transformers
标题:关注修剪视觉Transformer
链接:https://arxiv.org/abs/2604.15408
作者:Saif Mahmoud,Ahmad Almasri
摘要:用于Vision Transformers(ViTs)的令牌修剪方法通过丢弃无信息的补丁来保证注意力FLOP的二次减少。然而,当使用最先进的可变长度注意力API(包括FlashAttention-2的varlen和PyTorch的NestedTensor SDPA)执行修剪序列时,挂钟注意力延迟不会相应地扩展。我们将其追溯到一个调度开销瓶颈:在短的、典型的ViT后修剪序列长度(<=197个令牌)下,实际的矩阵运算在个位数微秒内完成,而主机端调度路径消耗60-90 us。我们提出了一个轻量级的,双向的Triton注意内核,其调度地板是40我们约1.5倍低于FlashAttention-2的varlen,允许修剪节省变得更加明显的挂钟时间。集成到一个完整的pack-attend-unpack管道中,我们的系统在四种修剪算法(Rollhold-L2,DynamicViT,EViT,ATS)中一致地实现了高达2.24倍的端到端吞吐量,跨DeiT-T/S/B扩展,并保持位精确分类预测,最大绝对logit差异<0.007。
摘要:Token pruning methods for Vision Transformers (ViTs) promise quadratic reductions in attention FLOPs by dropping uninformative patches. Yet when pruned sequences are executed with state-of-the-art variable-length attention APIs -- including FlashAttention-2's varlen and PyTorch's NestedTensor SDPA-the wall-clock attention latency doesn't scale accordingly. We trace this to a dispatch-overhead bottleneck: at the short, post-pruning sequence lengths typical of ViTs (<=197 tokens), actual matrix arithmetic completes in single-digit microseconds while the host-side dispatch path consumes 60-90 us. We present a lightweight, bidirectional Triton attention kernel whose dispatch floor is 40 us roughly 1.5x lower than FlashAttention-2 varlen-allowing pruning savings to become more visible in wall-clock time. Integrated into a complete pack-attend-unpack pipeline, our system achieves up to 2.24x end-to-end throughput over padded PyTorch SDPA consistently across four pruning algorithms (Threshold-L2, DynamicViT, EViT, ATS), scales across DeiT-T/S/B, and maintains bit-exact classification predictions with <0.007 max absolute logit difference.
【3】Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation
标题:幻觉作为轨迹承诺:Transformer发电中不对称吸引子动力学的因果证据
链接:https://arxiv.org/abs/2604.15400
作者:G. Aytug Akarlar
备注:21 pages, 12 figures, 8 tables. Code and data: https://github.com/akarlaraytu/trajectory-commitment
摘要:我们提出因果关系的证据表明,幻觉的自回归语言模型是一个早期的轨迹承诺不对称吸引动力学。使用相同的提示分岔,在其中我们反复采样相同的输入,以观察自发发散,我们隔离的轨迹动力学的水平混淆。在Qwen2.5-1.5B上,跨越六个类别的61个提示中,27个提示(44.3%)在第一个生成的标记处分叉,其中事实和幻觉轨迹分叉(在步骤0处KL = 0,在步骤1处KL > 1.0)。跨28层的激活修补揭示了明显的因果不对称性:在87.5%的试验(第20层)中,将幻觉激活注入正确的轨迹会破坏输出,而反向恢复仅为33.3%(第24层);两者都超过了10.4%的基线(p = 0.025)和12.5%的随机修补控制。窗口修补表明,纠正需要持续的多步干预,而腐败只需要一个单一的扰动。探测提示编码本身,步骤0残留状态预测在第15层的Pearson r = 0.776处的每提示幻觉率(p < 0.001,相对于1000-排列无效);无监督聚类确定了五个类似政权的群体,(eta^2 = 0.55)其鞍邻近集群集中了13个分叉假前提提示中的12个,表明流域结构是围绕着以即时编码为固定的制度承诺组织的。这些发现将幻觉描述为一个局部稳定的吸引子盆地:进入是概率性的和快速的,退出需要跨层和步骤的协调干预,并且相关盆地由在步骤0已经可辨别的可聚类制度选择。
摘要
:We present causal evidence that hallucination in autoregressive language models is an early trajectory commitment governed by asymmetric attractor dynamics. Using same-prompt bifurcation, in which we repeatedly sample identical inputs to observe spontaneous divergence, we isolate trajectory dynamics from prompt-level confounds. On Qwen2.5-1.5B across 61 prompts spanning six categories, 27 prompts (44.3%) bifurcate with factual and hallucinated trajectories diverging at the first generated token (KL = 0 at step 0, KL > 1.0 at step 1). Activation patching across 28 layers reveals a pronounced causal asymmetry: injecting a hallucinated activation into a correct trajectory corrupts output in 87.5% of trials (layer 20), while the reverse recovers only 33.3% (layer 24); both exceed the 10.4% baseline (p = 0.025) and 12.5% random-patch control. Window patching shows correction requires sustained multi-step intervention, whereas corruption needs only a single perturbation. Probing the prompt encoding itself, step-0 residual states predict per-prompt hallucination rate at Pearson r = 0.776 at layer 15 (p < 0.001 against a 1000-permutation null); unsupervised clustering identifies five regime-like groups (eta^2 = 0.55) whose saddle-adjacent cluster concentrates 12 of the 13 bifurcating false-premise prompts, indicating that the basin structure is organized around regime commitments fixed at prompt encoding. These findings characterize hallucination as a locally stable attractor basin: entry is probabilistic and rapid, exit demands coordinated intervention across layers and steps, and the relevant basins are selected by clusterable regimes already discernible at step 0.
【4】The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason
标题:思维的光谱几何:Transformer推理中的转变、指令拒绝、代币级动力学和完美正确性预测
链接:https://arxiv.org/abs/2604.15350
作者:Yi Liu
摘要:We discover that large language models exhibit \emph{spectral phase transitions} in their hidden activation spaces when engaging in reasoning versus factual recall. Through systematic spectral analysis across \textbf{11 models} spanning \textbf{5 architecture families} (Qwen, Pythia, Phi, Llama, DeepSeek-R1), we identify \textbf{seven} core phenomena: (1)~\textbf{Reasoning Spectral Compression} -- 9/11 models show significantly lower $α$ for reasoning ($p < 0.05$), with larger effects in stronger models; (2)~\textbf{Instruction Tuning Spectral Reversal} -- base models show reasoning $α< $ factual $α$, while instruction-tuned models reverse this relationship; (3)~\textbf{Architecture-Dependent Generation Taxonomy} -- prompt-to-response shifts partition into expansion, compression, and equilibrium regimes; (4)~\textbf{Spectral Scaling Law} -- $α_\text{reasoning} \propto -0.074 \ln N$ across 4 Qwen base models ($R^2 = 0.46$); (5)~\textbf{Token-Level Spectral Cascade} -- per-token alpha tracking reveals local synchronization that decays exponentially with layer distance, and is weaker for reasoning than factual tasks; (6)~\textbf{Reasoning Step Spectral Punctuation} -- phase-transition signatures align with reasoning step boundaries; and (7)~\textbf{Spectral Correctness Prediction} -- spectral $α$ alone achieves AUC $= 1.000$ (Qwen2.5-7B, late layers) and mean AUC $= 0.893$ across 6 models in predicting correctness \emph{before} the final answer is generated. Together, these findings establish a comprehensive \emph{spectral theory of reasoning} in transformers, revealing that the geometry of thought is universal in direction, architecture-specific in dynamics, and predictive of outcome.
摘要:We discover that large language models exhibit \emph{spectral phase transitions} in their hidden activation spaces when engaging in reasoning versus factual recall. Through systematic spectral analysis across \textbf{11 models} spanning \textbf{5 architecture families} (Qwen, Pythia, Phi, Llama, DeepSeek-R1), we identify \textbf{seven} core phenomena: (1)~\textbf{Reasoning Spectral Compression} -- 9/11 models show significantly lower $α$ for reasoning ($p < 0.05$), with larger effects in stronger models; (2)~\textbf{Instruction Tuning Spectral Reversal} -- base models show reasoning $α< $ factual $α$, while instruction-tuned models reverse this relationship; (3)~\textbf{Architecture-Dependent Generation Taxonomy} -- prompt-to-response shifts partition into expansion, compression, and equilibrium regimes; (4)~\textbf{Spectral Scaling Law} -- $α_\text{reasoning} \propto -0.074 \ln N$ across 4 Qwen base models ($R^2 = 0.46$); (5)~\textbf{Token-Level Spectral Cascade} -- per-token alpha tracking reveals local synchronization that decays exponentially with layer distance, and is weaker for reasoning than factual tasks; (6)~\textbf{Reasoning Step Spectral Punctuation} -- phase-transition signatures align with reasoning step boundaries; and (7)~\textbf{Spectral Correctness Prediction} -- spectral $α$ alone achieves AUC $= 1.000$ (Qwen2.5-7B, late layers) and mean AUC $= 0.893$ across 6 models in predicting correctness \emph{before} the final answer is generated. Together, these findings establish a comprehensive \emph{spectral theory of reasoning} in transformers, revealing that the geometry of thought is universal in direction, architecture-specific in dynamics, and predictive of outcome.
GAN|对抗|攻击|生成相关(4篇)
【1】Evaluating quality in synthetic data generation for large tabular health datasets
标题:评估大型表格健康数据集合成数据生成的质量
链接:https://arxiv.org/abs/2604.15961
作者:Jean-Baptiste Escudié,Benjamin Barnes,Stefan Meisegeier,Klaus Kraywinkel,Fabian Prasser,Nils Körber
摘要:在综合数据领域,对于质量评价的简明指标或大型卫生数据集(如历史流行病学数据)的基准,没有达成共识。这项研究评估了来自主要机器学习家族的七种最新模型。使用四个不同的数据集对模型进行评估,每个数据集都有不同的尺度。为了确保公平的比较,我们系统地调整了每个数据集的每个模型的超参数。我们提出了一种方法来评估合成的联合分布的保真度,在一个单一的情节与可视化对齐指标。该方法适用于任何数据集,并通过对德国癌症登记处流行病学数据集的特定领域分析进行补充。分析揭示了模型在严格遵守医学领域方面面临的挑战。我们希望这种方法将作为指导合成器选择的基础框架,并对参与发布合成数据集的所有利益相关者保持开放。
摘要:There is no consensus in the field of synthetic data on concise metrics for quality evaluations or benchmarks on large health datasets, such as historical epidemiological data. This study presents an evaluation of seven recent models from major machine learning families. The models were evaluated using four different datasets, each with a distinct scale. To ensure a fair comparison, we systematically tuned the hyperparameters of each model for each dataset. We propose a methodology for evaluating the fidelity of synthesized joint distributions, aligning metrics with visualization on a single plot. This method is applicable to any dataset and is complemented by a domain-specific analysis of the German Cancer Registries' epidemiological dataset. The analysis reveals the challenges models face in strictly adhering to the medical domain. We hope this approach will serve as a foundational framework for guiding the selection of synthesizers and remain accessible to all stakeholders involved in releasing synthetic datasets.
【2】Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
标题:以推理为目标的越狱通过语义触发器和心理框架对大型推理模型进行攻击
链接:https://arxiv.org/abs/2604.15725
作者:Zehao Wang,Lanjun Wang
摘要
:大型推理模型(LRM)在生成逐步推理链以及最终答案方面表现出强大的能力,使其能够在医疗保健和教育等高风险领域部署。虽然先前的越狱攻击研究集中在最终答案的安全性上,但很少关注推理过程的安全性。在这项工作中,我们确定了一个新的问题,注入有害内容的推理步骤,同时保持不变的答案。这种类型的攻击提出了两个关键挑战:1)操纵输入指令可能会无意中改变LRM的最终答案,以及2)输入问题的多样性使得难以始终绕过LRM的安全对齐机制并将有害内容嵌入其推理过程。为了解决这些挑战,我们提出了基于心理学的推理为目标的越狱攻击(PRJA)框架,它集成了基于语义的触发器选择模块和基于心理学的指令生成模块。具体而言,建议PRJA自动选择操纵推理触发器,通过语义分析,并利用服从权威和道德脱离的心理理论,以产生自适应指令,以提高LRM的遵守有害内容的生成。在五个问答数据集上的广泛实验表明,PRJA对几种商业LRM(包括DeepSeek R1,Qwen2.5-Max和OpenAI o 4-mini)的平均攻击成功率为83.6%。
摘要:Large Reasoning Models (LRMs) have demonstrated strong capabilities in generating step-by-step reasoning chains alongside final answers, enabling their deployment in high-stakes domains such as healthcare and education. While prior jailbreak attack studies have focused on the safety of final answers, little attention has been given to the safety of the reasoning process. In this work, we identify a novel problem that injects harmful content into the reasoning steps while preserving unchanged answers. This type of attack presents two key challenges: 1) manipulating the input instructions may inadvertently alter the LRM's final answer, and 2) the diversity of input questions makes it difficult to consistently bypass the LRM's safety alignment mechanisms and embed harmful content into its reasoning process. To address these challenges, we propose the Psychology-based Reasoning-targeted Jailbreak Attack (PRJA) Framework, which integrates a Semantic-based Trigger Selection module and a Psychology-based Instruction Generation module. Specifically, the proposed PRJA automatically selects manipulative reasoning triggers via semantic analysis and leverages psychological theories of obedience to authority and moral disengagement to generate adaptive instructions for enhancing the LRM's compliance with harmful content generation. Extensive experiments on five question-answering datasets demonstrate that PRJA achieves an average attack success rate of 83.6\% against several commercial LRMs, including DeepSeek R1, Qwen2.5-Max, and OpenAI o4-mini.
【3】Majority Voting for Code Generation
标题:代码生成的多数投票
链接:https://arxiv.org/abs/2604.15618
作者:Tim Launer,Jonas Hübotter,Marco Bagatella,Ido Hakimi,Andreas Krause
备注:ICLR 2026 Test-Time Updates (TTU) Workshop
摘要:我们调查功能多数表决(FMV),一种基于功能共识的方法,用于大型语言模型的代码生成,该方法使用测试输入上的运行时执行签名从多代中识别出一个代表性的解决方案。我们发现,FMV是一种有效的测试时推理策略,大大提高了LiveCodeBench的性能,而无需大量的计算开销。此外,我们扩展了功能共识的效用,并将其作为无标签测试时强化学习的聚合策略。我们证明,这增加了pass@1的坚持任务,但没有发现任何证据的自我改进超出了基本模型的性能上限。
摘要:We investigate Functional Majority Voting (FMV), a method based on functional consensus for code generation with Large Language Models, which identifies a representative solution from multiple generations using their runtime execution signatures on test inputs. We find that FMV is an effective test-time inference strategy, substantially boosting performance on LiveCodeBench without a large compute overhead. Furthermore, we extend the utility of functional consensus and apply it as an aggregation strategy for label-free Test-Time Reinforcement Learning. We demonstrate that this increases pass@1 on holdout tasks, but find no evidence of self-improvement beyond the base model's performance ceiling.
【4】InfoChess: A Game of Adversarial Inference and a Laboratory for Quantifiable Information Control
标题:InfoChess:对抗性推理的游戏和可量化信息控制的实验室
链接:https://arxiv.org/abs/2604.15373
作者:Kieran A. Murphy
备注:Accepted at Adaptive and Learning Agents Workshop, AAMAS 2026. Project page: https://github.com/murphyka/infochess
摘要:我们提出了InfoChess,一个对称的对抗游戏,提升竞争信息的获取的主要目标。没有碎片捕获,消除了物质激励,否则会混淆信息的作用。相反,碎片被用来改变可见性。玩家根据他们在游戏持续时间内对对手国王位置的概率推断进行评分。为了探索玩InfoChess的策略空间,我们引入了一个层次的启发式代理定义的对手建模的水平不断提高,并训练强化学习代理,优于这些基线。利用游戏的离散结构,我们通过自然的信息理论特征,包括信念熵,预言交叉熵,和预测日志分数下的行动引起的观察通道分析游戏。这些措施解开认知的不确定性,校准不匹配,以及对抗运动引起的不确定性。InfoChess的设计使其成为研究部分可观察性下多智能体推理的测试平台。我们发布了环境和代理的代码,以及一个公共接口,以鼓励进一步的研究。
摘要:We propose InfoChess, a symmetric adversarial game that elevates competitive information acquisition to the primary objective. There is no piece capture, removing material incentives that would otherwise confound the role of information. Instead, pieces are used to alter visibility. Players are scored on their probabilistic inference of the opponent's king location over the duration of the game. To explore the space of strategies for playing InfoChess, we introduce a hierarchy of heuristic agents defined by increasing levels of opponent modeling, and train a reinforcement learning agent that outperforms these baselines. Leveraging the discrete structure of the game, we analyze gameplay through natural information-theoretic characterizations that include belief entropy, oracle cross entropy, and predictive log score under the action-induced observation channel. These measures disentangle epistemic uncertainty, calibration mismatch, and uncertainty induced by adversarial movement. The design of InfoChess renders it a testbed for studying multi-agent inference under partial observability. We release code for the environment and agents, and a public interface to encourage further study.
半/弱/无/有监督|不确定性|主动学习(3篇)
【1】UA-Net: Uncertainty-Aware Network for TRISO Image Semantic Segmentation
标题:UA-Net:TRISO图像语义分割的不确定性感知网络
链接:https://arxiv.org/abs/2604.15542
作者:Kyle Lucke,Zuzanna Krajewska-Travar,Shoukun Sun,Lu Cai,John D. Stempien,Min Xian
摘要:三结构各向同性(TRISO)包覆颗粒燃料在高温中子辐照期间经历尺寸变化和化学反应。辐照后材相学有助于了解影响燃料性能的过程,例如涂层完整性和裂变产物保留。传统上,专家手动评估数千个亚毫米尺寸样品的横截面中的特征,这是繁琐且主观的。在这项工作中,我们提出了UA-Net,这是一个深度学习框架,可以分割TRISO燃料显微图的五个特征区域,并生成预测的不确定性图。该模型使用多阶段预训练策略,从ImageNet学习的一般图像表示开始,然后对来自各种辐照实验和AGR-5/6/7粒子横截面的TRISO显微照片进行微调。一个元模型的不确定性预测集成识别TRISO图像中的小缺陷。在102张图像的测试集上对UA-Net进行了评估,平均交点对并集(mIoU)和平均精度(mP)分别为95.5%和97.3%。元模型实现了91.8%的特异性和93.5%的灵敏度,在检测错误分类方面表现出强大的性能。该模型也适用于新的TRISO图像的定性评价,表现出较高的精度提取层区域。
摘要:Tristructural isotropic (TRISO)-coated particle fuels undergo dimensional changes and chemical reactions during high-temperature neutron irradiation. Post-irradiation materialography helps understand processes that impact fuel performance, such as coating integrity and fission product retention. Conventionally, experts manually evaluate features in thousands of cross sections of sub-mm-sized samples, which is tedious and subjective. In this work, we propose UA-Net, a deep learning framework that segments five characteristic regions of TRISO fuel micrographs and generates an uncertainty map for predictions. The model uses a multi-stage pretraining strategy, starting with general image representations learned from ImageNet, followed by fine-tuning on TRISO micrographs from various irradiation experiments and AGR-5/6/7 particle cross sections. A meta-model for uncertainty prediction is integrated to identify small defects in TRISO images. UA-Net was evaluated on a test set of 102 images, achieving mean Intersection over Union (mIoU) and mean Precision (mP) of 95.5% and 97.3%, respectively. The meta-model achieved a specificity of 91.8% and sensitivity of 93.5%, demonstrating strong performance in detecting misclassifications. The model was also applied to new TRISO images for qualitative evaluation, showing high accuracy in extracting layer regions.
【2】Transfer Learning from Foundational Optimization Embeddings to Unsupervised SAT Representations
标题:从基础优化嵌入转移学习到无监督SAT表示
链接:https://arxiv.org/abs/2604.15448
作者:Koyena Pal,Serdar Kadioglu
摘要:基础优化嵌入最近已经成为混合整数规划(MIP)问题的强大预训练表示。这些嵌入被证明能够实现跨域传输,并减少对求解器生成的标签的依赖。在这项工作中,我们调查是否这样的表示推广超越优化决策问题,专注于布尔可满足性(SAT)。我们适应的基础优化架构SAT映射CNF公式到相同的二分约束变量图表示用于MIP。这允许直接重用预训练的嵌入模型,而无需架构更改或监督微调。我们的研究结果表明,这些嵌入捕获结构的SAT实例,并支持无监督的任务,如实例聚类和分布识别。我们证明,第一次,基础优化嵌入可以转移到约束满足域。我们的研究结果是一个统一的代表性框架的优化和决策问题的一步。
摘要:Foundational optimization embeddings have recently emerged as powerful pre-trained representations for mixed-integer programming (MIP) problems. These embeddings were shown to enable cross-domain transfer and reduce reliance on solver-generated labels. In this work, we investigate whether such representations generalize beyond optimization to decision problems, focusing on Boolean satisfiability (SAT). We adapt the foundational optimization architecture to SAT by mapping CNF formulas into the same bipartite constraint-variable graph representation used for MIPs. This allows direct reuse of the pre-trained embedding model without architectural changes or supervised fine-tuning. Our results show that these embeddings capture structural regularities in SAT instances and support unsupervised tasks such as instance clustering and distribution identification. We demonstrate, for the first time, that foundational optimization embeddings can transfer to constraint satisfaction domains. Our findings is a step toward a unified representational framework for both optimization and decision problems.
【3】Mapping High-Performance Regions in Battery Scheduling across Data Uncertainty, Battery Design, and Planning Horizons
标题:跨越数据不确定性、电池设计和规划视野绘制电池调度的高性能区域
链接:https://arxiv.org/abs/2604.15360
作者:Jaime de Miguel Rodriguez,Artjom Vargunin,Brigitta Robin Raudne,David Solis Martin,Yaroslava Mykhailenko,Kaarel Oja
备注:40 pages
摘要:本研究提出了一个三元分析的多阶段模型预测控制下的储能运行,调查数据特性,预测不确定性,规划范围,和电池c率之间的相互作用。合成数据集的生成是为了系统地探索数据概况和不确定性的变化,从而实现参数化和构建将这些特征映射到最佳水平长度的关系。结果显示存在一个有效的地平线,定义为前瞻长度超过额外的预测信息提供有限的业务效益。考虑到这一范围可以降低计算成本,同时保持最佳性能。该研究提供了电池类型,不确定性水平和数据配置文件的广泛组合的最佳范围长度,为工业存储操作提供了实用指导。它还量化了由于预测不确定性造成的收入损失,表明即使是快速电池,错误也会影响性能。最后,该框架为未来的机器学习方法奠定了基础,这些方法将数据集参数化映射到最佳范围,支持工业环境中的连续优化,而无需大量计算。
摘要:This study presents a triadic analysis of energy storage operation under multi-stage model predictive control, investigating the interplay between data characteristics, forecast uncertainty, planning horizon, and battery c-rate. Synthetic datasets are generated to systematically explore variations in data profiles and uncertainty, enabling parametrization and the construction of relationships that map these characteristics to optimal horizon length. Results reveal the presence of an effective horizon, defined as the look-ahead length beyond which additional forecast information provides limited operational benefit. Accounting for this horizon can reduce computational costs while maintaining optimal performance. The study provides optimal horizon lengths across a broad range of combinations of battery types, uncertainty levels, and data profiles, offering practical guidance for industrial storage operation. It also quantifies revenue losses due to forecast uncertainty, showing that errors can impact performance even for fast batteries. Finally, the framework lays the groundwork for future machine learning approaches that map dataset parametrization to optimal horizons, supporting continuous optimization in industrial settings without heavy computation.
迁移|Zero/Few/One-Shot|自适应(12篇)
【1】FL-MHSM: Spatially-adaptive Fusion and Ensemble Learning for Flood-Landslide Multi-Hazard Susceptibility Mapping at Regional Scale
标题:FL-MRSM:区域规模洪水-滑坡多灾害易感性制图的空间自适应融合和集合学习
链接:https://arxiv.org/abs/2604.16265
作者:Aswathi Mundayatt,Jaya Sreevalsan-Nair
摘要:现有的多灾害敏感性制图(MHSM)研究往往依赖于空间上统一的模型,独立处理危险,并提供有限的跨危险的依赖性和不确定性的代表。为了解决这些局限性,本研究提出了一种深度学习(DL)工作流程,用于联合洪水-滑坡多灾害易感性映射(FL-MHSM),该工作流程结合了两级空间分区、概率早期融合(EF)、基于树的后期融合(LF)基线和软门控专家混合(MoE)模型,其中MoE作为最终预测模型。所提出的设计保留了空间异质性,通过分区,并使数据并行大面积预测使用重叠的格子网格。在喀拉拉邦,EF与LF保持竞争力,将洪水回忆从0.816提高到0.840,并将Brier评分从0.092降低到0.086,而MoE为洪水易感性提供了最强的性能,实现了0.905的AUC-ROC,回忆为0.930,F1评分为0.722。在尼泊尔,EF同样将洪水回忆从0.820提高到0.858,并将Brier评分从0.057降低到0.049,而MoE在滑坡易感性方面优于EF和LF,AUC-ROC为0.914,回忆为0.901,F1评分为0.559。GeoDetector对MoE产出的分析进一步表明,喀拉拉邦各地区的主导因素差异更大,敏感性受地形、土地覆盖和排水相关控制因素的不同组合影响,而尼泊尔各地区的地形和冰川相关因素的影响更为一致。这些研究结果表明,EF和LF提供互补的预测行为,并通过MoE空间自适应集成产生强大的整体预测性能FL-MHSM,同时支持可解释的表征空间异质景观的多灾害易感性。
摘要:Existing multi-hazard susceptibility mapping (MHSM) studies often rely on spatially uniform models, treat hazards independently, and provide limited representation of cross-hazard dependence and uncertainty. To address these limitations, this study proposes a deep learning (DL) workflow for joint flood-landslide multi-hazard susceptibility mapping (FL-MHSM) that combines two-level spatial partitioning, probabilistic Early Fusion (EF), a tree-based Late Fusion (LF) baseline, and a soft-gating Mixture of Experts (MoE) model, with MoE serving as final predictive model. The proposed design preserves spatial heterogeneity through zonal partitions and enables data-parallel large-area prediction using overlapping lattice grids. In Kerala, EF remained competitive with LF, improving flood recall from 0.816 to 0.840 and reducing Brier score from 0.092 to 0.086, while MoE provided strongest performance for flood susceptibility, achieving an AUC-ROC of 0.905, recall of 0.930, and F1-score of 0.722. In Nepal, EF similarly improved flood recall from 0.820 to 0.858 and reduced Brier score from 0.057 to 0.049 relative to LF, while MoE outperformed both EF and LF for landslide susceptibility, achieving an AUC-ROC of 0.914, recall of 0.901, and F1-score of 0.559. GeoDetector analysis of MoE outputs further showed that dominant factors varied more across zones in Kerala, where susceptibility was shaped by different combinations of topographic, land-cover, and drainage-related controls, while Nepal showed a more consistent influence of topographic and glacier-related factors across zones. These findings show that EF and LF provide complementary predictive behavior, and that their spatially adaptive integration through MoE yields robust overall predictive performance for FL-MHSM while supporting interpretable characterization of multi-hazard susceptibility in spatially heterogeneous landscapes.
【2】Multi-Objective Bayesian Optimization via Adaptive \varepsilon-Constraints Decomposition
链接:https://arxiv.org/abs/2604.15959
作者:Yaohong Yang,Sammie Katt,Samuel Kaski
摘要:多目标贝叶斯优化(MOBO)为优化具有多个目标的昂贵黑盒函数提供了一个原则性框架。然而,现有的MOBO方法往往挣扎的覆盖范围,可扩展性方面的目标的数量,并集成的约束和偏好。在这项工作中,我们提出了\textit{STAGE-BO,顺序目标自适应间隙填充$\vareprogram $-约束贝叶斯优化},明确针对帕累托前沿的未开发区域。通过分析近似Pareto前沿的覆盖范围,我们的方法确定了最大的几何间隙。这些差距,然后被用作约束条件,将问题转化为一系列的不等式约束的子问题,有效地解决了通过约束预期改善收购。我们的方法提供了一个统一的帕累托覆盖没有超体积计算,自然适用于约束和基于偏好的设置。在合成和真实世界基准上的实验表明,与最先进的基线相比,超卷具有卓越的覆盖率和竞争力。
摘要:Multi-objective Bayesian optimization (MOBO) provides a principled framework for optimizing expensive black-box functions with multiple objectives. However, existing MOBO methods often struggle with coverage, scalability with respect to the number of objectives, and integrating constraints and preferences. In this work, we propose \textit{STAGE-BO, Sequential Targeting Adaptive Gap-Filling $\varepsilon$-Constraint Bayesian Optimization}, that explicitly targets under-explored regions of the Pareto front. By analyzing the coverage of the approximate Pareto front, our method identifies the largest geometric gaps. These gaps are then used as constraints, which transforms the problem into a sequence of inequality-constrained subproblems, efficiently solved via constrained expected improvement acquisition. Our approach provides a uniform Pareto coverage without hypervolume computation and naturally applies to constrained and preference-based settings. Experiments on synthetic and real-world benchmarks demonstrate superior coverage and competitive hypervolume performance against state-of-the-art baselines.
【3】(Weighted) Adaptive Radius Near Neighbor Search: Evaluation for WiFi Fingerprint-based Positioning
标题:(加权)自适应半径近邻搜索:基于WiFi指纹的定位评估
链接:https://arxiv.org/abs/2604.15940
作者:Khang Le,Joaquín Torres-Sospedra,Philipp Müller
备注:11 pages, 2 figures, 2 tables, submitted to IPIN 2026
摘要:固定半径近邻(FRNN)搜索是广泛使用的k近邻(kNN)搜索的替代方案。与kNN不同,FRNN基于预定义距离内的所有训练样本确定测试样本的标签或估计值。虽然这种方法在某些情况下是有益的,但假设所有训练样本的固定最大距离可能会降低FRNN的准确性。因此,在本文中,我们提出了自适应半径近邻(ARNN)和加权ARNN(WARNN),采用自适应距离和在后一种情况下的权重。所有这三种方法都与kNN及其12种变体进行了比较,用于回归问题,即WiFi指纹室内定位,使用22个不同的数据集提供全面的分析。虽然测试的FRNN和ARNN版本的性能较差,但测试中四种最好的方法中有三种是WARNN版本,这表明使用权重和自适应距离可以实现与kNN变体相当甚至更好的性能。
摘要:Fixed Radius Near Neighbor (FRNN) search is an alternative to the widely used k Nearest Neighbors (kNN) search. Unlike kNN, FRNN determines a label or an estimate for a test sample based on all training samples within a predefined distance. While this approach is beneficial in certain scenarios, assuming a fixed maximum distance for all training samples can decrease the accuracy of the FRNN. Therefore, in this paper we propose the Adaptive Radius Near Neighbor (ARNN) and the Weighted ARNN (WARNN), which employ adaptive distances and in latter case weights. All three methods are compared to kNN and twelve of its variants for a regression problem, namely WiFi fingerprinting indoor positioning, using 22 different datasets to provide a comprehensive analysis. While the performances of the tested FRNN and ARNN versions were amongst the worse, three of the four best methods in the test were WARNN versions, indicating that using weights together with adaptive distances achieves performance comparable or even better than kNN variants.
【4】DiZiNER: Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition
标题:DiZigner:通过Zero-Shot命名实体识别的试点注释模拟的分歧引导指令细化
链接:https://arxiv.org/abs/2604.15866
作者:Siun Kim,Hyung-Jin Yoon
备注:9 pages, 3 figures; Accepted to the ACL 2026 Main Conference
摘要:大型语言模型(LLM)通过支持zero-shot和Few-Shot命名实体识别(NER)来实现高级信息提取(IE),但其生成输出仍然显示出持续和系统性错误。尽管通过指令微调取得了进展,但zero-shot NER仍然远远落后于监督系统。这些反复出现的错误反映了在早期人类注释过程中观察到的不一致性,这些过程通过试点注释来解决分歧。出于这种类比的动机,我们引入了DiZiNER(通过Zero-shot命名实体识别的试点注释模拟的分歧指导指令细化),这是一个模拟试点注释过程的框架,采用LLM作为注释者和监督者。多个异构的LLM注释共享的文本,和监督模型分析模型间的分歧,以完善任务指令。在18个基准测试中,DiZiNER在14个数据集上实现了zero-shot SOTA结果,将先前的最佳值提高了+8.0 F1,并将zero-shot与监督的差距减少了+11点。它的表现也一直优于其主管GPT-5 mini,这表明改进源于不同意指导的指令细化,而不是模型容量。模型之间的成对协议显示出与NER性能的强相关性,进一步支持这一发现。
摘要:Large language models (LLMs) have advanced information extraction (IE) by enabling zero-shot and few-shot named entity recognition (NER), yet their generative outputs still show persistent and systematic errors. Despite progress through instruction fine-tuning, zero-shot NER still lags far behind supervised systems. These recurring errors mirror inconsistencies observed in early-stage human annotation processes that resolve disagreements through pilot annotation. Motivated by this analogy, we introduce DiZiNER (Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition), a framework that simulates the pilot annotation process, employing LLMs to act as both annotators and supervisors. Multiple heterogeneous LLMs annotate shared texts, and a supervisor model analyzes inter-model disagreements to refine task instructions. Across 18 benchmarks, DiZiNER achieves zero-shot SOTA results on 14 datasets, improving prior bests by +8.0 F1 and reducing the zero-shot to supervised gap by over +11 points. It also consistently outperforms its supervisor, GPT-5 mini, indicating that improvements stem from disagreement-guided instruction refinement rather than model capacity. Pairwise agreement between models shows a strong correlation with NER performance, further supporting this finding.
【5】When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth
标题:提前退出网络何时推广?自适应深度的Pac-Bayesian理论
链接:https://arxiv.org/abs/2604.15764
作者:Dongxin Guo,Jikun Wu,Siu Ming Yiu
备注:6 pages, 1 figure, 7 tables, 1 algorithm
摘要:早期退出神经网络通过允许置信预测在中间层退出来实现自适应计算,从而实现2- 8\times $推理加速。尽管广泛部署,但它们的泛化特性缺乏理论上的理解-最近的调查明确指出了这一差距。本文建立了一个统一的PAC-Bayesian自适应深度网络框架。(1)基于熵的新边界:我们证明了第一个推广界依赖于出口深度熵$H(D)$和期望深度$\mathbb{E}[D]$而不是最大深度$K$,样本复杂度为$\mathcal{O}((\mathbb{E}[D] \cdot d + H(D))/ε^2)$。(2)显式构造常数:我们的分析产生的主导系数$\sqrt{2\ln 2} \约1.177$与完整的推导。(3)可证明的早期退出优势:我们建立了自适应深度网络严格优于固定深度网络的充分条件。(4)近似标签独立性的扩展:我们将标签独立性假设放宽为$ε$-近似策略,扩大了对学习路由的适用性。(5)全面验证:在7个基准测试上的6个架构上的实验表明,经典边界的紧密度比为1.52-3.87$\times$(所有$p < 0.001$),而$>$100$\times$。边界引导阈值选择与验证调整的性能匹配在0.1- 0.3%范围内。
摘要
:Early-exit neural networks enable adaptive computation by allowing confident predictions to exit at intermediate layers, achieving 2-8$\times$ inference speedup. Despite widespread deployment, their generalization properties lack theoretical understanding -- a gap explicitly identified in recent surveys. This paper establishes a unified PAC-Bayesian framework for adaptive-depth networks. (1) Novel Entropy-Based Bounds: We prove the first generalization bounds depending on exit-depth entropy $H(D)$ and expected depth $\mathbb{E}[D]$ rather than maximum depth $K$, with sample complexity $\mathcal{O}((\mathbb{E}[D] \cdot d + H(D))/ε^2)$. (2) Explicit Constructive Constants: Our analysis yields the leading coefficient $\sqrt{2\ln 2} \approx 1.177$ with complete derivation. (3) Provable Early-Exit Advantages: We establish sufficient conditions under which adaptive-depth networks strictly outperform fixed-depth counterparts. (4) Extension to Approximate Label Independence: We relax the label-independence assumption to $ε$-approximate policies, broadening applicability to learned routing. (5) Comprehensive Validation: Experiments across 6 architectures on 7 benchmarks demonstrate tightness ratios of 1.52-3.87$\times$ (all $p < 0.001$) versus $>$100$\times$ for classical bounds. Bound-guided threshold selection matches validation-tuned performance within 0.1-0.3%.
【6】DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference
标题:Depcap:用于高效扩散LM推理的自适应逐块并行解码
链接:https://arxiv.org/abs/2604.15750
作者:Xiang Xia,Wuyang Zhang,Jiazheng Liu,Cheng Yan,Yanyong Zhang
摘要:扩散语言模型(DLMs)已经成为自回归语言生成的一个有前途的替代方案,因为它们具有并行解码和全局优化整个序列的潜力。为了释放这种潜力,DLM推理必须仔细平衡生成质量和解码速度。最近的逐块DLM解码方法通过在块中顺序地执行基于扩散的解码来改善这种折衷。然而,现有的方法通常依赖于固定的块调度或当前步骤的本地信号来确定块边界,并使用保守的基于置信度的并行解码来避免冲突,限制了质量-速度的权衡。在本文中,我们认为,块明智的DLM推理需要更合适的信号,其两个核心决策:跨步骤的信号,用于确定块边界,和令牌级的冲突信号并行解码。基于这一观点,我们提出了DepCap,这是一个用于高效块式DLM推理的免训练框架。具体来说,DepCap将交叉步信号实例化为最后解码块的影响,并使用它来自适应地确定下一个块应该延伸多远,同时识别每个块内用于安全并行解码的无冲突令牌子集,从而实现实质性的推理加速,而质量下降可以忽略不计。DepCap是一种适用于各种DLM的即插即用方法,并且与用于逐块DLM的现有KV缓存策略兼容。信息理论分析进一步表明,累积的最后一个块对候选块的影响在令牌之间近似是加性的,支持所提出的块划分标准。实验结果表明,DepCap在多个DLM骨干和推理与编码基准测试中实现了良好的速度-质量权衡,具有高达5.63$\times$的加速比,而没有显着的性能下降。
摘要:Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive language generation due to their potential for parallel decoding and global refinement of the entire sequence. To unlock this potential, DLM inference must carefully balance generation quality and decoding speed. Recent block-wise DLM decoding methods improve this trade-off by performing diffusion-based decoding sequentially in blocks. However, existing methods typically rely on fixed block schedules or current-step local signals to determine block boundaries, and use conservative confidence-based parallel decoding to avoid conflicts, limiting the quality-speed trade-off. In this paper, we argue that block-wise DLM inference requires more suitable signals for its two core decisions: cross-step signals for determining block boundaries, and token-level conflict signals for parallel decoding. Based on this view, we propose DepCap, a training-free framework for efficient block-wise DLM inference. Specifically, DepCap instantiates the cross-step signal as the influence of the last decoded block and uses it to adaptively determine how far the next block should extend, while identifying a conflict-free subset of tokens for safe parallel decoding within each block, enabling substantial inference acceleration with negligible quality degradation. DepCap is a plug-and-play method applicable to various DLMs, and compatible with existing KV-cache strategies for block-wise DLM. An information-theoretic analysis further suggests that the cumulative last-block influence on a candidate block is approximately additive across tokens, supporting the proposed block-partitioning criterion. Experimental results show that DepCap achieves favorable speed-quality trade-offs across multiple DLM backbones and reasoning and coding benchmarks, with up to 5.63$\times$ speedup without significant performance degradation.
【7】Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning
标题:走向稳健的内生推理:统一非平稳调整中的漂移适应
链接:https://arxiv.org/abs/2604.15705
作者:Xiaoyu Yang,En Yu,Wei Duan,Jie Lu
摘要:强化微调(RFT)已经成为多模态大型语言模型(MLLM)与复杂的人类价值观和特定领域需求保持一致的关键范式。然而,目前的研究主要集中在减轻外生分布的变化所产生的数据为中心的因素,内在的非平稳性的内生推理仍然在很大程度上未被探索。在这项工作中,一个关键的弱点是揭示MLLM内:他们是非常容易受到内源性的推理漂移,在思维和感知的角度。它表现为不可预测的分布变化,在自回归生成过程中自发出现,独立于外部环境扰动。为了适应它,我们首先从理论上定义的多模态概念漂移的RFT的MLLM内的内源性推理漂移。在此背景下,本文提出了反事实偏好优化++(CPO++),一个全面的和自治的框架,适应多模态的概念漂移。它将反事实推理与领域知识相结合,以执行跨思维和感知的受控扰动,采用偏好优化来解开虚假的相关性。在两个高度动态和安全关键领域进行广泛的经验评估:医疗诊断和自动驾驶。他们表明,所提出的框架在推理一致性,决策精度和对极端干扰的固有鲁棒性方面取得了优异的性能。该方法还表现出特殊的zero-shot跨域概括,提供了一个可靠的多模态推理在安全关键应用的原则基础。
摘要:Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with complex human values and domain-specific requirements. Nevertheless, current research primarily focuses on mitigating exogenous distribution shifts arising from data-centric factors, the non-stationarity inherent in the endogenous reasoning remains largely unexplored. In this work, a critical vulnerability is revealed within MLLMs: they are highly susceptible to endogenous reasoning drift, across both thinking and perception perspectives. It manifests as unpredictable distribution changes that emerge spontaneously during the autoregressive generation process, independent of external environmental perturbations. To adapt it, we first theoretically define endogenous reasoning drift within the RFT of MLLMs as the multi-modal concept drift. In this context, this paper proposes Counterfactual Preference Optimization ++ (CPO++), a comprehensive and autonomous framework adapted to the multi-modal concept drift. It integrates counterfactual reasoning with domain knowledge to execute controlled perturbations across thinking and perception, employing preference optimization to disentangle spurious correlations. Extensive empirical evaluations across two highly dynamic and safety-critical domains: medical diagnosis and autonomous driving. They demonstrate that the proposed framework achieves superior performance in reasoning coherence, decision-making precision, and inherent robustness against extreme interference. The methodology also exhibits exceptional zero-shot cross-domain generalization, providing a principled foundation for reliable multi-modal reasoning in safety-critical applications.
【8】Adapting in the Dark: Efficient and Stable Test-Time Adaptation for Black-Box Models
标题:黑暗中适应:黑匣子模型的高效稳定的测试时适应
链接:https://arxiv.org/abs/2604.15609
作者:Yunbei Zhang,Shuaicheng Niu,Chengyi Cai,Feng Liu,Jihun Hamm
备注:Third Workshop on Test-Time Updates (Oral)
摘要:对于只能通过API访问的黑盒模型,测试时自适应(TTA)仍然是一个尚未探索的挑战。现有的方法,如事后输出细化提供有限的自适应能力,而零阶优化(ZOO),使输入空间的适应,但面临着高查询成本和优化的挑战,在无监督的TTA设置。我们介绍了BETA(黑盒有效的测试时适应),一个框架,通过采用轻量级的,本地的白盒转向模型来创建一个易于处理的梯度路径,以解决这些限制。通过预测协调技术与一致性正则化和快速学习导向过滤相结合,BETA实现了稳定的自适应,无需额外的API调用,并且在标准推理之外的延迟可以忽略不计。在ImageNet-C上,BETA在ViT-B/16上实现了+7.1%的准确率增益,在CLIP上实现了+3.4%的准确率增益,超过了包括TENT和TPT在内的强大的白盒和灰盒方法。在商业API上,BETA以250倍的低成本实现了与ZOO相当的性能,同时保持了实时推理速度,使其成为现实世界黑盒TTA的实用高效解决方案。
摘要
:Test-Time Adaptation (TTA) for black-box models accessible only via APIs remains a largely unexplored challenge. Existing approaches such as post-hoc output refinement offer limited adaptive capacity, while Zeroth-Order Optimization (ZOO) enables input-space adaptation but faces high query costs and optimization challenges in the unsupervised TTA setting. We introduce BETA (Black-box Efficient Test-time Adaptation), a framework that addresses these limitations by employing a lightweight, local white-box steering model to create a tractable gradient pathway. Through a prediction harmonization technique combined with consistency regularization and prompt learning-oriented filtering, BETA enables stable adaptation with no additional API calls and negligible latency beyond standard inference. On ImageNet-C, BETA achieves a +7.1% accuracy gain on ViT-B/16 and +3.4% on CLIP, surpassing strong white-box and gray-box methods including TENT and TPT. On a commercial API, BETA achieves comparable performance to ZOO at 250x lower cost while maintaining real-time inference speed, establishing it as a practical and efficient solution for real-world black-box TTA.
【9】ProtoTTA: Prototype-Guided Test-Time Adaptation
标题:ProtoTTA:原型引导的测试时适应
链接:https://arxiv.org/abs/2604.15494
作者:Mohammad Mahdi Abootorabi,Parvin Mousavi,Purang Abolmaesumi,Evan Shelhamer
备注:ICLR 2026 Test-Time Updates (TTU) Workshop
摘要:依赖于原型(可以与模型输入相关的可解释表示)的深度网络在平衡高准确性与固有可解释性方面获得了极大的关注,这使得它们适用于医疗保健等关键领域。然而,这些模型受限于它们对训练数据的依赖,这妨碍了它们对分布变化的鲁棒性。虽然测试时自适应(TTA)通过更新参数和统计数据来提高深度网络的鲁棒性,但尚未为此目的探索可解释模型的原型。我们介绍ProtoTTA,这是一个原型模型的通用框架,它利用中间原型信号,而不是仅仅依赖于模型输出。ProtoTTA最大限度地减少了原型相似性分布的熵,以鼓励对移动数据进行更自信和特定于原型的激活。为了保持稳定性,我们采用几何过滤来限制更新到具有可靠原型激活的样本,并通过原型重要性权重和模型置信度得分进行正则化。在四个不同的基准跨越细粒度视觉,组织病理学和NLP的四个原型骨干实验表明,ProtoTTA提高了标准输出熵最小化的鲁棒性,同时恢复正确的语义焦点在原型激活。我们还引入了新的可解释性指标和视觉语言模型(VLM)评估框架来解释TTA动态,确认ProtoTTA恢复了与人类一致的语义焦点,并与VLM评级的推理质量可靠地相关。代码可从以下网址获得:https://github.com/DeepRCL/ProtoTTA。
摘要:Deep networks that rely on prototypes-interpretable representations that can be related to the model input-have gained significant attention for balancing high accuracy with inherent interpretability, which makes them suitable for critical domains such as healthcare. However, these models are limited by their reliance on training data, which hampers their robustness to distribution shifts. While test-time adaptation (TTA) improves the robustness of deep networks by updating parameters and statistics, the prototypes of interpretable models have not been explored for this purpose. We introduce ProtoTTA, a general framework for prototypical models that leverages intermediate prototype signals rather than relying solely on model outputs. ProtoTTA minimizes the entropy of the prototype-similarity distribution to encourage more confident and prototype-specific activations on shifted data. To maintain stability, we employ geometric filtering to restrict updates to samples with reliable prototype activations, regularized by prototype-importance weights and model-confidence scores. Experiments across four prototypical backbones on four diverse benchmarks spanning fine-grained vision, histopathology, and NLP demonstrate that ProtoTTA improves robustness over standard output entropy minimization while restoring correct semantic focus in prototype activations. We also introduce novel interpretability metrics and a vision-language model (VLM) evaluation framework to explain TTA dynamics, confirming ProtoTTA restores human-aligned semantic focus and correlates reliably with VLM-rated reasoning quality. Code is available at: https://github.com/DeepRCL/ProtoTTA.
【10】Lightweight Geometric Adaptation for Training Physics-Informed Neural Networks
标题:用于训练物理信息神经网络的轻量级几何自适应
链接:https://arxiv.org/abs/2604.15392
作者:Kang An,Chenhao Si,Shiqian Ma,Ming Yan
备注:22 pages, Chenhao Si and Kang An contributed equally to this work. Their authorship order was determined randomly
摘要:物理信息神经网络(PINN)通常收敛缓慢,训练不稳定,并且由于其损失景观的各向异性和快速变化的几何形状而降低了具有挑战性的偏微分方程的精度。我们提出了一个轻量级的曲率感知的优化框架,增强现有的一阶优化与自适应预测校正的基础上割线信息。连续梯度差被用作局部几何变化的廉价代理,连同步长归一化割线曲率指示器一起来控制校正强度。该框架是即插即用,计算效率高,广泛兼容现有的优化,而不显式地形成二阶矩阵。在不同PDE基准上的实验表明,与标准优化器和强基线相比,收敛速度,训练稳定性和解决方案的准确性得到了一致的改善,包括高维热方程,Gray-Scott系统,Belousov-Zhabotinsky系统和2D Kuramoto-Sivashinsky系统。
摘要:Physics-Informed Neural Networks (PINNs) often suffer from slow convergence, training instability, and reduced accuracy on challenging partial differential equations due to the anisotropic and rapidly varying geometry of their loss landscapes. We propose a lightweight curvature-aware optimization framework that augments existing first-order optimizers with an adaptive predictive correction based on secant information. Consecutive gradient differences are used as a cheap proxy for local geometric change, together with a step-normalized secant curvature indicator to control the correction strength. The framework is plug-and-play, computationally efficient, and broadly compatible with existing optimizers, without explicitly forming second-order matrices. Experiments on diverse PDE benchmarks show consistent improvements in convergence speed, training stability, and solution accuracy over standard optimizers and strong baselines, including on the high-dimensional heat equation, Gray--Scott system, Belousov--Zhabotinsky system, and 2D Kuramoto--Sivashinsky system.
【11】Adaptive multi-fidelity optimization with fast learning rates
标题:具有快速学习率的自适应多保真优化
链接:https://arxiv.org/abs/2604.16239
作者:Come Fiegel,Victor Gabillon,Michal Valko
备注:Published at International Conference on Artificial Intelligence and Statistics (AISTATS) 2020
摘要:在多保真度优化中,目标函数的不同成本的有偏近似是可用的。本文研究了在有限预算下优化局部光滑函数的问题,其中学习者必须在这些近似的成本和偏差之间进行权衡。首先,我们证明了下界的简单的遗憾在不同的假设条件下,基于成本偏置函数。然后,我们提出了Kometo算法,它实现了额外的对数因子,相同的速率,而没有任何知识的函数平滑度和保真度的假设,并提高了以前证明的保证。最后,我们的经验表明,我们的算法优于以前的多保真度优化方法,没有问题相关的参数的知识。
摘要:In multi-fidelity optimization, biased approximations of varying costs of the target function are available. This paper studies the problem of optimizing a locally smooth function with a limited budget, where the learner has to make a tradeoff between the cost and the bias of these approximations. We first prove lower bounds for the simple regret under different assumptions on the fidelities, based on a cost-to-bias function. We then present the Kometo algorithm which achieves, with additional logarithmic factors, the same rates without any knowledge of the function smoothness and fidelity assumptions, and improves previously proven guarantees. We finally empirically show that our algorithm outperforms previous multi-fidelity optimization methods without the knowledge of problem-dependent parameters.
【12】One-Shot Generative Flows: Existence and Obstructions
标题:一次性生成流:存在与障碍
链接:https://arxiv.org/abs/2604.15439
作者:Panos Tsimpos,Daniel Sharp,Youssef Marzouk
摘要
:我们研究了随机过程$X_\bullet$中生成模型的动态测量传输,该随机过程的边缘在源分布$P_0$和目标分布$P_1$之间插值,同时保持独立,即,when $(X_0,X_1)\sim P_0\otimes P_1$. 这个过程$X_\bullet$的条件期望定义了一个ODE,它的流图从$P_0$传输到$P_1$。我们讨论当这样一个过程引起\n {直线流},即一个逐点加速度为零,因此是完全可积的任何一阶方法。 首先,我们开发了多个特征的直线度的偏微分方程涉及的过程中的条件统计。 然后,我们证明了端点独立性下的直线度表现出尖锐的二分法。 一方面,我们构造显式的,可计算的任意高斯端点的直线过程。另一方面,我们表明直线过程不存在的目标,充分分离的模式。我们通过一系列越来越普遍的不可能性定理来证明这一点,这些定理揭示了具有独立端点的过程的样本路径行为与该过程的流图的时空几何之间的基本关系。总之,这些结果提供了一个结构理论时,直生成流可以,不能,存在。
摘要:We study dynamic measure transport for generative modelling in the setting of a stochastic process $X_\bullet$ whose marginals interpolate between a source distribution $P_0$ and a target distribution $P_1$ while remaining independent, i.e., when $(X_0,X_1)\sim P_0\otimes P_1$. Conditional expectations of this process $X_\bullet$ define an ODE whose flow map transports from $P_0$ to $P_1$. We discuss when such a process induces a \emph{straight-line flow}, namely one whose pointwise acceleration vanishes and is therefore exactly integrable by any first-order method. We first develop multiple characterizations of straightness in terms of PDEs involving the conditional statistics of the process. Then, we prove that straightness under endpoint independence exhibits a sharp dichotomy. On one hand, we construct explicit, computable straight-line processes for arbitrary Gaussian endpoints. On the other hand, we show straight-line processes do not exist for targets with sufficiently well-separated modes. We demonstrate this through a sequence of increasingly general impossibility theorems that uncover a fundamental relationship between the sample-path behavior of a process with independent endpoints and the space-time geometry of this process' flow map. Taken together, these results provide a structural theory of when straight generative flows can, and cannot, exist.
强化学习(3篇)
【1】Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning
标题:将谜题放在重要的地方:强化学习的问题增强框架
链接:https://arxiv.org/abs/2604.15830
作者:Yangyi Fang,Jiaye Lin,Xiaoliang Fu,Cong Qin,Haolin Shi
摘要:强化学习已经成为增强大型语言模型推理的强大方法,但面临着一个根本性的困境:在简单问题上进行训练可能会导致过度拟合和pass@k退化,而在困难问题上进行训练通常会导致稀疏的奖励。最近的问题扩增方法解决了这一问题,前置部分解决方案作为提示。然而,统一的提示提供可能会引入冗余信息,同时错过关键的推理瓶颈,过多的提示会降低推理多样性,导致pass@k退化。我们提出了\textbf{PieceHint},一个提示注入框架,在训练过程中战略性地识别并提供关键的推理步骤。通过对不同推理步骤的重要性进行评分,根据问题的难度有选择地分配提示,并逐步撤回脚手架,PieceHint使模型能够从引导学习过渡到独立推理。在六个数学推理基准测试上的实验表明,我们的1.5B模型实现了与32 B基线相当的平均性能,同时在所有$k$值中保持了pass@k多样性。
摘要:Reinforcement learning has become a powerful approach for enhancing large language model reasoning, but faces a fundamental dilemma: training on easy problems can cause overfitting and pass@k degradation, while training on hard problems often results in sparse rewards. Recent question augmentation methods address this by prepending partial solutions as hints. However, uniform hint provision may introduce redundant information while missing critical reasoning bottlenecks, and excessive hints can reduce reasoning diversity, causing pass@k degradation. We propose \textbf{PieceHint}, a hint injection framework that strategically identifies and provides critical reasoning steps during training. By scoring the importance of different reasoning steps, selectively allocating hints based on problem difficulty, and progressively withdrawing scaffolding, PieceHint enables models to transition from guided learning to independent reasoning. Experiments on six mathematical reasoning benchmarks show that our 1.5B model achieves comparable average performance to 32B baselines while preserving pass@k diversity across all $k$ values.
【2】Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment
标题:具有增强状态的多目标强化学习部署后需要奖励
链接:https://arxiv.org/abs/2604.15757
作者:Peter Vamplew,Cameron Foale
摘要:本研究报告确定了多目标强化学习(MORL)和更传统的单目标强化学习(RL)之间以前被忽视的区别。先前已经指出,具有非线性效用函数的MORL代理的最优策略需要以当前环境状态和先前累积的奖励的某种度量为条件。这通常通过将环境的观察状态与先前奖励的折扣总和连接以创建增强状态来实现。虽然增广状态在MORL文献中被广泛使用,但它们的使用的一个含义以前没有报道过-即它们要求代理在部署后继续访问奖励信号(或其代理),即使不需要进一步学习。本说明解释了为什么会出现这种情况,并考虑了这一要求的实际影响。
摘要:This research note identifies a previously overlooked distinction between multi-objective reinforcement learning (MORL), and more conventional single-objective reinforcement learning (RL). It has previously been noted that the optimal policy for an MORL agent with a non-linear utility function is required to be conditioned on both the current environmental state and on some measure of the previously accrued reward. This is generally implemented by concatenating the observed state of the environment with the discounted sum of previous rewards to create an augmented state. While augmented states have been widely-used in the MORL literature, one implication of their use has not previously been reported -- namely that they require the agent to have continued access to the reward signal (or a proxy thereof) after deployment, even if no further learning is required. This note explains why this is the case, and considers the practical repercussions of this requirement.
【3】Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning
标题:超越单一模型优化:在连续强化学习中保留可塑性
链接:https://arxiv.org/abs/2604.15414
作者:Lute Lillo,Nick Cheney
摘要:持续强化学习必须平衡保留与适应,但许多方法仍然依赖于单一模型保留,致力于一个不断发展的策略作为跨任务的主要可重用解决方案。即使保留了以前成功的策略,它也可能不再为干扰后的快速适应提供可靠的起点,这反映了单一策略保存无法解决的可塑性丧失。受质量多样性方法的启发,我们引入了\textsc{TeLAPA}(启用传输的延迟对齐策略档案),这是一个持续的RL框架,它将行为上不同的策略社区组织到每个任务的档案中,并保持一个共享的潜在空间,以便归档的策略在非平稳漂移下保持可比性和可重用性。这种观点将持续强化学习从保留孤立的解决方案转变为通过支持未来再学习的有能力和行为相关的政策来维护技能一致的社区。在我们的MiniGrid CL设置中,\textsc{TeLAPA}成功学习更多任务,在干扰后重新访问任务时更快地恢复能力,并在一系列任务中保持更高的性能。我们的分析表明,源最优的政策往往不是转移最优的,即使在一个当地的主管邻里,有效的再利用取决于保留和选择多个附近的替代品,而不是将它们折叠到一个代表。总之,这些结果重新构建了围绕可重复使用和有能力的政策社区的持续RL,提供了一条超越单一模型保存的路线,走向更多的塑料终身代理。
摘要:Continual reinforcement learning must balance retention with adaptation, yet many methods still rely on \emph{single-model preservation}, committing to one evolving policy as the main reusable solution across tasks. Even when a previously successful policy is retained, it may no longer provide a reliable starting point for rapid adaptation after interference, reflecting a form of \emph{loss of plasticity} that single-policy preservation cannot address. Inspired by quality-diversity methods, we introduce \textsc{TeLAPA} (Transfer-Enabled Latent-Aligned Policy Archives), a continual RL framework that organizes behaviorally diverse policy neighborhoods into per-task archives and maintains a shared latent space so that archived policies remain comparable and reusable under non-stationary drift. This perspective shifts continual RL from retaining isolated solutions to maintaining \emph{skill-aligned neighborhoods} with competent and behaviorally related policies that support future relearning. In our MiniGrid CL setting, \textsc{TeLAPA} learns more tasks successfully, recovers competence faster on revisited tasks after interference, and retains higher performance across a sequence of tasks. Our analyses show that source-optimal policies are often not transfer-optimal, even within a local competent neighborhood, and that effective reuse depends on retaining and selecting among multiple nearby alternatives rather than collapsing them to one representative. Together, these results reframe continual RL around reusable and competent policy neighborhoods, providing a route beyond single-model preservation toward more plastic lifelong agents.
符号|符号学习(1篇)
【1】Neuro-Symbolic ODE Discovery with Latent Grammar Flow
标题:具有潜在语法流的神经符号ODE发现
链接:https://arxiv.org/abs/2604.16232
作者:Karin Yu,Eleni Chatzi,Georgios Kissas
摘要:理解自然和工程系统通常依赖于符号公式,如微分方程,它提供了超越黑箱模型的可解释性和可转换性。 我们介绍了潜在语法流(LGF),神经符号生成框架发现常微分方程的数据。LGF将方程作为基于语法的表示嵌入到离散的潜在空间中,并迫使语义相似的方程以行为损失的方式更紧密地定位在一起。然后,离散流模型引导采样过程递归地生成最适合观测数据的候选方程。领域知识和约束(例如稳定性)可以嵌入到规则中或用作条件预测器。
摘要:Understanding natural and engineered systems often relies on symbolic formulations, such as differential equations, which provide interpretability and transferability beyond black-box models. We introduce Latent Grammar Flow (LGF), a neuro-symbolic generative framework for discovering ordinary differential equations from data. LGF embeds equations as grammar-based representations into a discrete latent space and forces semantically similar equations to be positioned closer together with a behavioural loss. Then, a discrete flow model guides the sampling process to recursively generate candidate equations that best fit the observed data. Domain knowledge and constraints, such as stability, can be either embedded into the rules or used as conditional predictors.
医学相关(3篇)
【1】TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation
标题:TwinTrack:用于医学图像分割的事后多评级者校准
链接:https://arxiv.org/abs/2604.15950
作者:Tristan Kirscher,Alexandra Ertl,Klaus Maier-Hein,Xavier Coubez,Philippe Meyer,Sylvain Faisan
摘要:胰腺导管腺癌(PDAC)在对比增强CT上的分割本质上是模糊的:专家之间的评分员分歧反映了真正的不确定性,而不是注释噪音。标准的深度学习方法假设一个单一的基础事实,产生的概率输出可能校准不良,并且在这种模糊性下难以解释。我们提出了TwinTrack,一个框架,通过事后校准集成分割概率的经验平均人类反应(MHR)-标记为肿瘤体素的专家注释的分数,解决了这一差距。因此,校准的概率可直接解释为分配肿瘤标签的注释者的预期比例,明确建模评价者间的不一致。建议的事后校准程序是简单的,只需要一个小的多评价者校准集。在MICCAI 2025 CURVAS-PDACVI多评估器基准上进行评估时,它始终优于标准方法。
摘要:Pancreatic ductal adenocarcinoma (PDAC) segmentation on contrast-enhanced CT is inherently ambiguous: inter-rater disagreement among experts reflects genuine uncertainty rather than annotation noise. Standard deep learning approaches assume a single ground truth, producing probabilistic outputs that can be poorly calibrated and difficult to interpret under such ambiguity. We present TwinTrack, a framework that addresses this gap through post-hoc calibration of ensemble segmentation probabilities to the empirical mean human response (MHR) -the fraction of expert annotators labeling a voxel as tumor. Calibrated probabilities are thus directly interpretable as the expected proportion of annotators assigning the tumor label, explicitly modeling inter-rater disagreement. The proposed post-hoc calibration procedure is simple and requires only a small multi-rater calibration set. It consistently improves calibration metrics over standard approaches when evaluated on the MICCAI 2025 CURVAS-PDACVI multi-rater benchmark.
【2】ECG-Lens: Benchmarking ML & DL Models on PTB-XL Dataset
标题:ECG-Lens:在PTB-XL数据集中对ML和DL模型进行基准测试
链接:https://arxiv.org/abs/2604.15822
作者:Saloni Garg,Ukant Jadia,Amit Sagtani,Kamal Kant Hiran
备注:8 pages, 4 figures, 3 tables
摘要:心电图信号的自动分类是诊断和监测心血管疾病的有效工具。本研究比较了三种传统机器学习算法(决策树分类器、随机森林分类器和逻辑回归)和三种深度学习模型(简单卷积神经网络(CNN)、长短期记忆(LSTM)和复杂CNN(ECGLens))对PTB-XL数据集ECG信号的分类,该数据集包含来自正常患者和各种心脏病患者的12导联记录。DL模型在原始ECG信号上进行训练,使它们能够自动提取区分特征。数据增强使用平稳小波变换(SWT),以提高模型的性能,增加训练样本的多样性,并保持ECG信号的基本特征。使用多个指标对模型进行评估,包括准确度、精确度、召回率、F1评分和ROC-AUC。ECG-Lens模型实现了最高的性能,具有80%的分类准确度和90%的ROC-AUC。这些研究结果表明,深度学习架构,特别是复杂的CNN,在原始12导联ECG数据上的表现大大优于传统的ML方法,并为选择自动ECG分类模型和确定特定条件模型开发的方向提供了实用的基准。
摘要:Automated classification of electrocardiogram (ECG) signals is a useful tool for diagnosing and monitoring cardiovascular diseases. This study compares three traditional machine learning algorithms (Decision Tree Classifier, Random Forest Classifier, and Logistic Regression) and three deep learning models (Simple Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and Complex CNN (ECGLens)) for the classification of ECG signals from the PTB-XL dataset, which contains 12-lead recordings from normal patients and patients with various cardiac conditions. The DL models were trained on raw ECG signals, allowing them to automatically extract discriminative features. Data augmentation using the Stationary Wavelet Transform (SWT) was applied to enhance model performance, increase the diversity of training samples, and preserve the essential characteristics of the ECG signals. The models were evaluated using multiple metrics, including accuracy, precision, recall, F1-score, and ROC-AUC. The ECG-Lens model achieved the highest performance, with 80% classification accuracy and a 90% ROC-AUC. These findings demonstrate that deep learning architectures, particularly complex CNNs substantially outperform traditional ML methods on raw 12-lead ECG data, and provide a practical benchmark for selecting automated ECG classification models and identifying directions for condition-specific model development.
【3】Topology-Driven Fusion of nnU-Net and MedNeXt for Accurate Brain Tumor Segmentation on Sub-Saharan Africa Dataset
标题:nnU-Net和MedNeXt的Top-驱动融合,在撒哈拉以南非洲数据集中实现准确的脑肿瘤分割
链接:https://arxiv.org/abs/2604.15964
作者:Prabin Bohara,Pralhad Kumar Shrestha,Arpan Rai,Usha Poudel Lamgade,Confidence Raymond,Dong Zhang,Aondona Lorumbu,Craig Jones,Mahesh Shakya,Bishesh Khanal,Pratibha Kulung
摘要:由于缺乏明确的国家成像协议、多样化的成像数据、低场磁共振成像(MRI)扫描仪的广泛使用以及有限的医疗资源,中低收入(LMIC)国家的准确自动脑肿瘤分割具有挑战性。作为脑肿瘤分割(BraTS)非洲2025挑战赛的一部分,我们将拓扑细化应用于最先进的分割模型,如nnU-Net,MedNeXt以及两者的组合。由于BraTS-Africa数据集的MRI图像质量较低,因此我们结合了治疗前成人胶质瘤的BraTS 2025挑战数据(任务1)来预训练分割模型,并使用它对BraTS-Africa数据集进行微调。我们增加了一个额外的拓扑细化模块,以解决由于拓扑错误而引起的预测变形问题。通过引入该模块,我们在周围非增强FLAIR高信号(SNFH)、非增强肿瘤核心(NETC)和增强肿瘤(ET)上实现了更好的归一化表面距离(NSD),分别为0.810、0.829和0.895。
摘要
:Accurate automatic brain tumor segmentation in Low and Middle-Income (LMIC) countries is challenging due to the lack of defined national imaging protocols, diverse imaging data, extensive use of low-field Magnetic Resonance Imaging (MRI) scanners and limited health-care resources. As part of the Brain Tumor Segmentation (BraTS) Africa 2025 Challenge, we applied topology refinement to the state-of-the-art segmentation models like nnU-Net, MedNeXt, and a combination of both. Since the BraTS-Africa dataset has low MRI image quality, we incorporated the BraTS 2025 challenge data of pre-treatment adult glioma (Task 1) to pre-train the segmentation model and use it to fine-tune on the BraTS-Africa dataset. We added an extra topology refinement module to address the issue of deformation in prediction that arose due to topological error. With the introduction of this module, we achieved a better Normalized Surface Distance (NSD) of 0.810, 0.829, and 0.895 on Surrounding Non-Enhancing FLAIR Hyperintensity (SNFH) , Non-Enhancing Tumor Core (NETC) and Enhancing tumor (ET).
聚类(1篇)
【1】Why Colors Make Clustering Harder:Global Integrality Gaps, the Price of Fairness, and Color-Coupled Algorithms in Chromatic Correlation Clustering
标题:为什么颜色会让集群变得更加困难:全球完整性差距、公平性的代价以及色彩相关集群中的颜色耦合算法
链接:https://arxiv.org/abs/2604.15738
作者:Ibne Farabi Shihab,Sanjeda Akter,Anuj Sharma
摘要:色相关聚类(CCC)通过为边缘分配语义颜色并要求每个聚类接收单个颜色标签来扩展相关聚类。与标准CC不同,其LP松弛在完全图上的完整性间隙为2并且允许2.06近似,CCC的类似LP具有严格的下限2.11,并且最著名的LP舍入算法达到2.15。我们解释这个差距隔离困难的来源:跨边缘的彩色干扰。颜色与候选聚类颜色不匹配的中性边缘会产生标准CC中不存在的不可约成本,并迫使任何颜色无关的舍入方案支付额外的失配惩罚。 我们做了四个贡献。首先,我们证明了一个全局完整性间隙分解定理,表明任何颜色无关的CCC舍入算法的间隙等于标准CC间隙加上不可约的色罚Delta(L)> 0。其次,我们解决了相关的min-max问题,并推导出阶梯公式Delta(L)=((L-1)/L)Delta_infinity,其中Delta_infinity约为0.0734。特别地,双色间隙为2.0967,在L = 2处已经将CCC与标准CC分开。第三,我们引入颜色耦合相关聚类(C4)。添加有效的全局约束sum_c x_uv^c >= L-1和相关的区间填充舍入方案,使得中性边缘表现得像经典的负边缘,恢复最佳的2.06近似并绕过非耦合LP的2.11下限。第四,在极端实例、真实多关系网络和公平基准上的实验验证了理论:经验LP差距遵循预测的阶梯,C4匹配公平约束下的无约束近似比。
摘要:Chromatic Correlation Clustering (CCC) extends Correlation Clustering by assigning semantic colors to edges and requiring each cluster to receive a single color label. Unlike standard CC, whose LP relaxation has integrality gap 2 on complete graphs and admits a 2.06-approximation, the analogous LP for CCC has a strict lower bound of 2.11, and the best known LP-rounding algorithm achieves 2.15. We explain this gap by isolating the source of difficulty: cross-edge chromatic interference. Neutral edges, whose color does not match the candidate cluster color, create an irreducible cost absent from standard CC and force any color-independent rounding scheme to pay an additional mismatch penalty. We make four contributions. First, we prove a Global Integrality Gap Decomposition Theorem showing that the gap of any color-independent CCC rounding algorithm equals the standard CC gap plus an irreducible chromatic penalty Delta(L) > 0. Second, we solve the associated min-max problem and derive the staircase formula Delta(L) = ((L-1)/L) Delta_infinity, where Delta_infinity is approximately 0.0734. In particular, the two-color gap is 2.0967, separating CCC from standard CC already at L = 2. Third, we introduce Color-Coupled Correlation Clustering (C4). Adding the valid global constraint sum_c x_uv^c >= L-1 and a correlated interval-packing rounding scheme makes neutral edges behave like classical negative edges, recovering the optimal 2.06 approximation and bypassing the 2.11 lower bound for the uncoupled LP. Fourth, experiments on extremal instances, real multi-relational networks, and fairness benchmarks validate the theory: empirical LP gaps follow the predicted staircase, and C4 matches the unconstrained approximation ratio under fairness constraints.
超分辨率|去噪|去模糊|去雾(1篇)
【1】Similarity-Based Bike Station Expansion via Hybrid Denoising Autoencoders
标题:通过混合降噪自动编码器进行基于相似性的自行车站扩展
链接:https://arxiv.org/abs/2604.15783
作者:Oluwaleke Yusuf,M. Tsaqif Wismadi,Adil Rasheed
备注:10 pages, 9 figures. Code available at https://github.com/Outsiders17711/TCB-SimilarityAE-Expansion
摘要:城市自行车共享系统需要战略性的站点扩展,以满足不断增长的需求。传统的分配方法依赖于明确的需求模型,可能无法捕捉到区分成功车站的城市特征。这项研究解决了利用现有台站的模式,特别是在数据受限的环境中,为扩展决策提供信息的需要。我们提出了一个数据驱动的框架,利用现有的车站认为可取的运营指标。混合去噪自动编码器(HDAE)从多源网格级特征(社会人口统计,建筑环境和交通网络)中学习压缩的潜在表示,并通过监督分类头来规范嵌入空间结构。扩展候选者通过贪婪分配与空间约束的基础上潜在的空间相似性,现有的站。对特隆赫姆自行车共享网络的评估表明,HDAE嵌入比原始特征产生更多的空间一致性集群和分配模式。相似性方法和距离度量的敏感性分析证实了鲁棒性。跨多个参数化的基于共识的过程提炼出32个高置信度扩展区域,其中所有参数化都同意。结果展示了表征学习如何捕捉原始特征错过的复杂模式,从而在没有明确需求建模的情况下实现基于证据的扩展规划。共识程序通过要求参数化之间达成一致来加强建议,而框架可配置性允许规划者纳入操作知识。该方法概括到任何位置分配问题,现有的理想的情况下,通知新的候选人的选择。
摘要:Urban bike-sharing systems require strategic station expansion to meet growing demand. Traditional allocation approaches rely on explicit demand modelling that may not capture the urban characteristics distinguishing successful stations. This study addresses the need to exploit patterns from existing stations to inform expansion decisions, particularly in data-constrained environments. We present a data-driven framework leveraging existing stations deemed desirable by operational metrics. A hybrid denoising autoencoder (HDAE) learns compressed latent representations from multi-source grid-level features (socio-demographic, built environment, and transport network), with a supervised classification head regularising the embedding space structure. Expansion candidates are selected via greedy allocation with spatial constraints based on latent-space similarity to existing stations. Evaluation on Trondheim's bike-sharing network demonstrates that HDAE embeddings yield more spatially coherent clusters and allocation patterns than raw features. Sensitivity analyses across similarity methods and distance metrics confirm robustness. A consensus-based procedure across multiple parametrisations distils 32 high-confidence extension zones where all parametrisations agree. The results demonstrate how representation learning captures complex patterns that raw features miss, enabling evidence-based expansion planning without explicit demand modelling. The consensus procedure strengthens recommendations by requiring agreement across parametrisations, while framework configurability allows planners to incorporate operational knowledge. The methodology generalises to any location-allocation problem where existing desirable instances inform the selection of new candidates.
自动驾驶|车辆|车道检测等(3篇)
【1】Unveiling Stochasticity: Universal Multi-modal Probabilistic Modeling for Traffic Forecasting
标题:揭开随机性:交通预测的通用多模式概率模型
链接:https://arxiv.org/abs/2604.16084
作者:Weijiang Xiong,Robert Fonod,Nikolas Geroliminis
摘要:交通预测是一项具有挑战性的时空建模任务,也是城市交通管理的重要组成部分。目前的研究主要集中在确定性预测,对交通动力学的不确定性和随机性考虑有限。因此,本文提出了一种优雅而通用的方法,将现有的模型转换为概率预测,只替换最终的输出层与一个新的高斯混合模型(GMM)层。修改后的模型不需要对训练管道进行任何更改,并且可以仅使用负对数李克图(NLL)损失进行训练,而无需任何辅助或正则化项。在多个流量数据集上的实验表明,我们的方法可以从经典模型推广到现代模型架构,同时保持确定性性能。此外,我们提出了一个系统的评估程序的基础上的累积分布和置信区间,并证明我们的方法是更准确和信息比单峰或确定性基线。最后,一个真实世界的密集城市交通网络的更详细的研究,研究数据质量的不确定性量化的影响,并显示我们的方法在不完美的数据条件下的鲁棒性。代码可在https://github.com/Weijiang-Xiong/OpenSkyTraffic获得
摘要
:Traffic forecasting is a challenging spatio-temporal modeling task and a critical component of urban transportation management. Current studies mainly focus on deterministic predictions, with limited considerations on the uncertainty and stochasticity in traffic dynamics. Therefore, this paper proposes an elegant yet universal approach that transforms existing models into probabilistic predictors by replacing only the final output layer with a novel Gaussian Mixture Model (GMM) layer. The modified model requires no changes to the training pipeline and can be trained using only the Negative Log-Likelihood (NLL) loss, without any auxiliary or regularization terms. Experiments on multiple traffic datasets show that our approach generalizes from classic to modern model architectures while preserving deterministic performance. Furthermore, we propose a systematic evaluation procedure based on cumulative distributions and confidence intervals, and demonstrate that our approach is considerably more accurate and informative than unimodal or deterministic baselines. Finally, a more detailed study on a real-world dense urban traffic network is presented to examine the impact of data quality on uncertainty quantification and to show the robustness of our approach under imperfect data conditions. Code available at https://github.com/Weijiang-Xiong/OpenSkyTraffic
【2】Driving Assistance System for Ambulances to Minimise the Vibrations in Patient Cabin
标题:救护车驾驶辅助系统可最大限度地减少患者舱内的振动
链接:https://arxiv.org/abs/2604.16047
作者:Abdulaziz Aldegheishem,Nabil Alrajeh,Lorena Parra,Oscar Romero,Jaime Lloret
备注:19 pages, 14 figures, 10 tables
摘要:救护车服务是患病或受伤人员的主要运输工具,其承受与常规车辆相同的加速力。由车辆的移动引起的这些加速度影响由卫生人员执行的任务的性能,这可能影响患者的存活或恢复时间。在本文中,我们已经训练,验证和测试了一个系统,以评估驾驶救护车服务。所提出的系统是由一个传感器节点,使用加速度计测量车辆的振动。它还包括一个GPS传感器,一个电池,一个显示器和一个扬声器。当两条可能的路线到达相同的目的地点时,系统基于先前分类的数据比较两条路线,并计算指数和分数。因此,该指数在到达目的地的时间和患者舱中遭受的振动方面平衡可能的路线,以推荐使这些振动最小化的路线。三个数据集用于训练、验证和测试系统。基于人工神经网络(ANN),分类模型与分类为低,中,高振动的标记数据进行训练,并达到97%的准确率。然后,所获得的模型进行了验证,从另一个地区的三条路线的数据。最后,该系统在两个新的情况下进行测试,有两种可能的路线到达目的地。结果表明,当两条可能的路线之间存在较低的时间差(小于6%)时,振动较小的路线是首选。然而,利用当前的加权因子,当路线之间的时间差高于20%时,最短路线是优选的,而不管最短路线中的较高振动。
摘要:The ambulance service is the main transport for diseased or injured people which suffers the same acceleration forces as regular vehicles. These accelerations, caused by the movement of the vehicle, impact the performance of tasks executed by sanitary personnel, which can affect patient survival or recovery time. In this paper, we have trained, validated, and tested a system to assess driving in ambulance services. The proposed system is composed of a sensor node which measures the vehicle vibrations using an accelerometer. It also includes a GPS sensor, a battery, a display, and a speaker. When two possible routes reach the same destination point, the system compares the two routes based on previously classified data and calculates an index and a score. Thus, the index balances the possible routes in terms of time to reach the destination and the vibrations suffered in the patient cabin to recommend the route that minimises those vibrations. Three datasets are used to train, validate, and test the system. Based on an Artificial Neural network (ANN), the classification model is trained with tagged data classified as low, medium, and high vibrations, and 97% accuracy is achieved. Then, the obtained model is validated using data from three routes of another region. Finally, the system is tested in two new scenarios with two possible routes to reach the destination. The results indicate that the route with less vibration is preferred when there are low time differences (less than 6%) between the two possible routes. Nonetheless, with the current weighting factors, the shortest route is preferred when time differences between routes are higher than 20%, regardless of the higher vibrations in the shortest route.
【3】Fusing Cellular Network Data and Tollbooth Counts for Urban Traffic Flow Estimation
标题:融合蜂窝网络数据和收费站计数用于城市交通流量估计
链接:https://arxiv.org/abs/2604.15782
作者:Oluwaleke Yusuf,Shaira Tabassum
备注:8 pages, 7 figures
摘要:交通模拟对于规划城市交通基础设施干预至关重要,需要特定车辆类别的起讫点(OD)数据。现有的数据来源是不完善的:稀疏的收费站传感器提供准确的车辆分类计数,而大量的移动数据从蜂窝网络活动捕捉聚集人群的运动,但缺乏模式分解,并有系统的偏见。这项研究开发了一个机器学习框架,使用稀疏收费站计数作为地面实况来纠正和分解蜂窝网络数据。该模型使用时间和空间特征来学习聚合的移动数据和车辆数据之间的复杂关系。该框架从过境路线推断目的地,并实现路由逻辑来分配OD对之间的校正流。这种方法适用于在挪威特隆赫姆的巴士站扩建,生成每小时OD矩阵的车辆长度类别。结果表明,有限的,但准确的传感器测量可以纠正广泛的,但汇总的移动数据,以产生接地估计的背景车辆交通流量。这些宏观尺度的估计可以在所需的位置进行微观尺度的分析。该框架提供了一种从蜂窝网络数据生成起点-目的地数据的可推广方法。这使得下游任务,如在数据稀缺的情况下进行基础设施规划的详细交通模拟,支持城市规划者做出明智的决策。
摘要:Traffic simulations, essential for planning urban transit infrastructure interventions, require vehicle-category-specific origin-destination (OD) data. Existing data sources are imperfect: sparse tollbooth sensors provide accurate vehicle counts by category, while extensive mobility data from cellular network activity captures aggregated crowd movement, but lack modal disaggregation and have systematic biases. This study develops a machine learning framework to correct and disaggregate cellular network data using sparse tollbooth counts as ground truth. The model uses temporal and spatial features to learn the complex relationship between aggregated mobility data and vehicular data. The framework infers destinations from transit routes and implements routing logic to distribute corrected flows between OD pairs. This approach is applied to a bus depot expansion in Trondheim, Norway, generating hourly OD matrices by vehicle length category. The results show how limited but accurate sensor measurements can correct extensive but aggregated mobility data to produce grounded estimates of background vehicular traffic flows. These macro-scale estimates can be refined for micro-scale analysis at desired locations. The framework provides a generalisable approach for generating origin-destination data from cellular network data. This enables downstream tasks, like detailed traffic simulations for infrastructure planning in data-scarce contexts, supporting urban planners in making informed decisions.
联邦学习|隐私保护|加密(1篇)
【1】Federated Learning with Quantum Enhanced LSTM for Applications in High Energy Physics
标题:利用量子增强LSTM进行联邦学习在高能物理中的应用
链接:https://arxiv.org/abs/2604.15775
作者:Abhishek Sawaika,Durga Pritam Suggisetti,Udaya Parampalli,Rajkumar Buyya
备注:8 pages, 7 figures, accepted at IEEE WCCI, 2026
摘要:使用大规模数据集和信息关键型应用(如高能物理(HEP))进行学习,需要高度复杂的大规模模型,这些模型既可靠又准确。为了解决这个问题并满足学习需求,我们设想使用具有量子增强模型的联邦学习框架。具体来说,我们设计了一个混合量子经典长射期记忆模型(QLSTM)的本地训练在分布式节点。它结合了量子模型在理解特征空间内复杂关系方面的代表性能力,以及基于LSTM的模型来学习数据点之间的必要相关性。考虑到当前独立噪声中间量子(NISQ)设备的计算限制和前所未有的成本,我们建议使用联合学习设置,其中学习负载可以根据设计和数据可用性分布到本地服务器。我们证明了这样的设计的好处,分类任务的超对称(SUSY)数据集,有5 M行。我们的实验表明,该设计的性能不仅优于使用基于变分量子电路(VQC)的量子机器学习(QML)技术的一些现有工作,而且与经典深度学习基准测试相当($Δ\sim \pm 1\%$)。这项研究的一个重要观察结果是,所设计的框架具有$
摘要
:Learning with large-scale datasets and information-critical applications, such as in High Energy Physics (HEP), demands highly complex, large-scale models that are both robust and accurate. To tackle this issue and cater to the learning requirements, we envision using a federated learning framework with a quantum-enhanced model. Specifically, we design a hybrid quantum-classical long-shot-term-memory model (QLSTM) for local training at distributed nodes. It combines the representative power of quantum models in understanding complex relationships within the feature space, and an LSTM-based model to learn necessary correlations across data points. Given the computing limitations and unprecedented cost of current stand-alone noisy-intermediate quantum (NISQ) devices, we propose to use a federated learning setup, where the learning load can be distributed to local servers as per design and data availability. We demonstrate the benefits of such a design on a classification task for the Supersymmetry(SUSY) dataset, having 5M rows. Our experiments indicate that the performance of this design is not only better that some of the existing work using variational quantum circuit (VQC) based quantum machine learning (QML) techniques, but is also comparable ($Δ\sim \pm 1\%$) to that of classical deep-learning benchmarks. An important observation from this study is that the designed framework has $
推理|分析|理解|解释(9篇)
【1】AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
标题:AtManRL:通过区分注意力显着性实现忠实推理
链接:https://arxiv.org/abs/2604.16158
作者:Max Henning Höth,Kristian Kersting,Björn Deiseroth,Letitia Parcalabescu
备注:14 pages, 8 figures, 1 table
摘要:大型语言模型(LLM)越来越依赖于思想链(CoT)推理来解决复杂的任务。然而,确保推理轨迹既有助于又忠实地反映了模型最终答案背后的过程,而不仅仅是伴随着它,仍然具有挑战性。我们介绍AtManRL,这是一种利用可微注意力操纵通过强化学习来学习更忠实推理的方法。通过训练一个附加的注意力掩码,识别CoT中对产生正确答案至关重要的标记,我们得到了一个显着性奖励信号,鼓励模型生成真正影响其最终预测的推理轨迹。我们将这种显着性奖励与GRPO框架内基于结果的奖励相结合,以共同优化正确性和可解释性。在GSM 8 K和MMLU上使用Llama-3.2-3B-Instruct进行的实验表明,我们的方法可以识别有影响力的推理令牌,并能够训练更透明的推理模型。
摘要:Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both contributes to and faithfully reflects the processes underlying the model's final answer, rather than merely accompanying it, remains challenging. We introduce AtManRL, a method that leverages differentiable attention manipulation to learn more faithful reasoning through reinforcement learning. By training an additive attention mask that identifies tokens in the CoT crucial for producing correct answers, we derive a saliency reward signal that encourages the model to generate reasoning traces that genuinely influence its final predictions. We integrate this saliency reward with outcome-based rewards within the GRPO framework to jointly optimize for correctness and interpretability. Experiments on GSM8K and MMLU with Llama-3.2-3B-Instruct demonstrate that our approach can identify influential reasoning tokens and enable training more transparent reasoning models.
【2】Sentiment Analysis of German Sign Language Fairy Tales
标题:德国手语童话的情感分析
链接:https://arxiv.org/abs/2604.16138
作者:Fabrizio Nunnari,Siddhant Jain,Patrick Gebhard
摘要:我们提出了一个数据集和德国手语(DGS)童话的情感分析模型。首先,我们使用四个大型语言模型(LLM)和多数投票对德国童话文本片段进行三个级别的效价(消极,中性,积极)的情感分析,达到了0.781 Krippendorff alpha的注释者间协议。其次,我们使用MediaPipe从每个相应的DGS视频片段中提取面部和身体运动特征。最后,我们训练一个可解释的模型(基于XGBoost)来预测来自视频特征的负面、中性或正面情绪。结果表明,平均平衡精度为0.631。最重要的特征的一个彻底的分析表明,除了眉毛和嘴的动作在脸上,臀部,肘部和肩膀的运动也大大有助于在所传达的情感的歧视,表明同样重要的面部和身体的情感交流手语。
摘要:We present a dataset and a model for sentiment analysis of German sign language (DGS) fairy tales. First, we perform sentiment analysis for three levels of valence (negative, neutral, positive) on German fairy tales text segments using four large language models (LLMs) and majority voting, reaching an inter-annotator agreement of 0.781 Krippendorff's alpha. Second, we extract face and body motion features from each corresponding DGS video segment using MediaPipe. Finally, we train an explainable model (based on XGBoost) to predict negative, neutral or positive sentiment from video features. Results show an average balanced accuracy of 0.631. A thorough analysis of the most important features reveal that, in addition to eyebrows and mouth motion on the face, also the motion of hips, elbows, and shoulders considerably contribute in the discrimination of the conveyed sentiment, indicating an equal importance of face and body for sentiment communication in sign language.
【3】Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
标题:减少你的损失!学习尽早修剪路径以实现高效并行推理
链接:https://arxiv.org/abs/2604.16029
作者:Jiaxi Bi,Tongxu Luo,Wenyu Du,Zhengyang Tang,Benyou Wang
备注:9 pages, 7 figures
摘要:并行推理增强了大型推理模型(LRM),但由于早期错误导致的无效路径而导致了高昂的成本。为了缓解这一问题,在前缀级的路径修剪是必不可少的,但现有的研究仍然支离破碎,没有一个标准化的框架。在这项工作中,我们提出了第一个系统分类的路径修剪,分类方法的信号源(内部与外部)和可学习性(可学习与不可学习)。这种分类揭示了可学习内部方法的未开发潜力,激发了我们提出STOP(Super TOken for Pruning)的建议。从1.5B到20 B参数范围内的LRM的广泛评估表明,与现有基线相比,STOP实现了卓越的有效性和效率。此外,我们严格验证了STOP在不同计算预算下的可扩展性-例如,在固定计算预算下,将AIME 25上的GPT-OSS-20 B准确率从84%提高到近90%。最后,我们将我们的研究结果提炼成正式的经验指南,以促进最佳的现实世界部署。代码、数据和模型可在https://bijiaxihh.github.io/STOP上获得
摘要:Parallel reasoning enhances Large Reasoning Models (LRMs) but incurs prohibitive costs due to futile paths caused by early errors. To mitigate this, path pruning at the prefix level is essential, yet existing research remains fragmented without a standardized framework. In this work, we propose the first systematic taxonomy of path pruning, categorizing methods by their signal source (internal vs. external) and learnability (learnable vs. non-learnable). This classification reveals the unexplored potential of learnable internal methods, motivating our proposal of STOP (Super TOken for Pruning). Extensive evaluations across LRMs ranging from 1.5B to 20B parameters demonstrate that STOP achieves superior effectiveness and efficiency compared to existing baselines. Furthermore, we rigorously validate the scalability of STOP under varying compute budgets - for instance, boosting GPT-OSS-20B accuracy on AIME25 from 84% to nearly 90% under fixed compute budgets. Finally, we distill our findings into formalized empirical guidelines to facilitate optimal real-world deployment. Code, data and models are available at https://bijiaxihh.github.io/STOP
【4】SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
标题:SocialGrid:协作多智能体系统中规划和社会推理的基准
链接:https://arxiv.org/abs/2604.16022
作者:Hikaru Shindo,Hanzhao Lin,Lukas Helff,Patrick Schramowski,Kristian Kersting
备注:Preprint
摘要:随着大型语言模型(LLM)从文本处理器过渡到自治代理,在具体的多代理设置中评估其社会推理变得至关重要。我们介绍SocialGrid,一个具体的多智能体环境的灵感来自我们之间的评估LLM代理的规划,任务执行和社会推理。我们的评估显示,即使是最强的开放模型(GPT-OSS-120 B)在任务完成和规划方面的准确率也低于60%,智能体陷入重复行为或无法导航基本障碍。由于糟糕的导航混淆了对社会智能的评估,SocialGrid提供了一个可选的规划Oracle,以将社会推理与规划缺陷隔离开来。虽然规划援助提高了任务的完成,社会推理仍然是一个瓶颈:代理未能检测到欺骗在接近随机的机会,无论规模,依赖于肤浅的推理,而不是积累行为证据。SocialGrid提供自动故障分析和细粒度指标,使开发人员能够诊断和改进他们的代理。我们还建立了一个有竞争力的排行榜使用Elo评级从对抗联赛发挥。
摘要:As Large Language Models (LLMs) transition from text processors to autonomous agents, evaluating their social reasoning in embodied multi-agent settings becomes critical. We introduce SocialGrid, an embodied multi-agent environment inspired by Among Us that evaluates LLM agents on planning, task execution, and social reasoning. Our evaluations reveal that even the strongest open model (GPT-OSS-120B) achieves below 60% accuracy in task completion and planning, with agents getting stuck in repetitive behaviors or failing to navigate basic obstacles. Since poor navigation confounds evaluation of social intelligence, SocialGrid offers an optional Planning Oracle to isolate social reasoning from planning deficits. While planning assistance improves task completion, social reasoning remains a bottleneck: agents fail to detect deception at near-random chance regardless of scale, relying on shallow heuristics rather than accumulating behavioral evidence. SocialGrid provides automatic failure analysis and fine-grained metrics, enabling developers to diagnose and improve their agents. We also establish a competitive leaderboard using Elo ratings from adversarial league play.
【5】Hierarchical Active Inference using Successor Representations
标题:使用后继表示的分层主动推理
链接:https://arxiv.org/abs/2604.15679
作者:Prashant Rangarajan,Rajesh P. N. Rao
备注:Accepted for publication in Neural Computation (MIT Press). 82 pages, 29 figures
摘要:主动推理是一种基于自由能原理(FEP)的神经推理模型,已被提出作为理解大脑中感知,动作和学习的统一框架。主动推理以前曾被用于对导航和规划等生态重要任务进行建模,但将其扩展到解决现实环境中的复杂大规模问题仍然是一个挑战。受大脑中存在多尺度层次表征的启发,我们提出了一个基于层次主动推理的行动规划模型。我们的方法结合了层次模型的环境与继任者表示有效的规划。我们提出的结果表明:(1)如何使用较低级别的后继表示来学习较高级别的抽象状态,(2)如何使用基于较低级别的主动推理的规划来引导和学习较高级别的抽象动作,以及(3)这些学习到的较高级别的抽象状态和动作如何促进有效的规划。我们说明了几个规划和强化学习(RL)的问题,包括一个著名的四个房间的任务,基于键的导航任务,部分可观察的规划问题,山地车问题,和PointMaze,一个家庭的导航任务与连续的状态和动作空间的方法的性能。我们的研究结果代表,据我们所知,第一次应用学到的层次状态和动作抽象的积极推理的FEP为基础的理论的大脑功能。
摘要:Active inference, a neurally-inspired model for inferring actions based on the free energy principle (FEP), has been proposed as a unifying framework for understanding perception, action, and learning in the brain. Active inference has previously been used to model ecologically important tasks such as navigation and planning, but scaling it to solve complex large-scale problems in real-world environments has remained a challenge. Inspired by the existence of multi-scale hierarchical representations in the brain, we propose a model for planning of actions based on hierarchical active inference. Our approach combines a hierarchical model of the environment with successor representations for efficient planning. We present results demonstrating (1) how lower-level successor representations can be used to learn higher-level abstract states, (2) how planning based on active inference at the lower-level can be used to bootstrap and learn higher-level abstract actions, and (3) how these learned higher-level abstract states and actions can facilitate efficient planning. We illustrate the performance of the approach on several planning and reinforcement learning (RL) problems including a variant of the well-known four rooms task, a key-based navigation task, a partially observable planning problem, the Mountain Car problem, and PointMaze, a family of navigation tasks with continuous state and action spaces. Our results represent, to our knowledge, the first application of learned hierarchical state and action abstractions to active inference in FEP-based theories of brain function.
【6】Flexible Empowerment at Reasoning with Extended Best-of-N Sampling
标题:通过扩展N中最佳抽样灵活授权推理
链接:https://arxiv.org/abs/2604.15614
作者:Taisuke Kobayashi
备注:15 pages, 4 figures
摘要:本文提出了一种新的方法,在强化学习(RL)推理动作时加入授权,从而实现探索-利用困境(EED)的灵活性。在以前的方法中,用于促进探索的授权已经被提供作为任务特定的奖励函数的奖金项,作为内在动机的RL。然而,这种方法引入了一个延迟,直到了解了解释授权的策略,使得难以根据需要调整对探索的强调。另一方面,一个技巧设计微调最近的基础模型在推理,所谓的最佳N(BoN)采样,允许隐式收购修改后的政策,而不显式地学习他们。预计将这一技巧应用于促进探索的术语,如授权,将使EED的调整更加灵活。因此,本文研究授权的BoN抽样。此外,调整的程度,在一个可推广的方式,同时保持计算成本的政策修改,本文提出了一种新的BoN抽样方法扩展的Tsalis统计。通过算例验证了该方法对电火工品平衡的可行性。此外,它表明,该方法提高RL性能,以解决复杂的运动任务。
摘要:This paper proposes a novel method that incorporates empowerment when reasoning actions in reinforcement learning (RL), thereby achieving the flexibility of exploration-exploitation dilemma (EED). In previous methods, empowerment for promoting exploration has been provided as a bonus term to the task-specific reward function as an intrinsically-motivated RL. However, this approach introduces a delay until the policy that accounts for empowerment is learned, making it difficult to adjust the emphasis on exploration as needed. On the other hand, a trick devised for fine-tuning recent foundation models at reasoning, so-called best-of-N (BoN) sampling, allows for the implicit acquisition of modified policies without explicitly learning them. It is expected that applying this trick to exploration-promoting terms, such as empowerment, will enable more flexible adjustment of EED. Therefore, this paper investigates BoN sampling for empowerment. Furthermore, to adjust the degree of policy modification in a generalizable manner while maintaining computational cost, this paper proposes a novel BoN sampling method extended by Tsalis statistics. Through toy problems, the proposed method's cability to balance EED is verified. In addition, it is demonstrated that the proposed method improves RL performance to solve complex locomotion tasks.
【7】PAWN: Piece Value Analysis with Neural Networks
标题:PAWN:利用神经网络进行单品价值分析
链接:https://arxiv.org/abs/2604.15585
作者:Ethan Tang,Hasan Davulcu,Jia Zou,Zhongju Zhang
备注:19 pages, 5 figures, 12 tables
摘要:预测任何给定棋子在某一位置的相对价值仍然是一个开放的挑战,因为棋子的贡献取决于它与棋盘上其他棋子的空间关系。我们证明,通过使用基于CNN的自动编码器导出的潜在位置表示,将整个棋盘的状态合并,可以显着提高基于MLP的棋子值预测架构的准确性。使用从大师级游戏中收集的超过1200万个棋子值对的数据集,以及Stockfish 17生成的地面真实标签,我们增强的棋子值预测器的性能显着优于基于上下文无关的MLP系统,将验证平均绝对误差降低了16%,并在大约0.65个棋子内预测相对棋子值。更一般地说,我们的研究结果表明,编码的完整的问题状态作为上下文提供了有用的归纳偏差预测的任何单个组件的贡献。
摘要
:Predicting the relative value of any given chess piece in a position remains an open challenge, as a piece's contribution depends on its spatial relationships with every other piece on the board. We demonstrate that incorporating the state of the full chess board via latent position representations derived using a CNN-based autoencoder significantly improves accuracy for MLP-based piece value prediction architectures. Using a dataset of over 12 million piece-value pairs gathered from Grandmaster-level games, with ground-truth labels generated by Stockfish 17, our enhanced piece value predictor significantly outperforms context-independent MLP-based systems, reducing validation mean absolute error by 16% and predicting relative piece value within approximately 0.65 pawns. More generally, our findings suggest that encoding the full problem state as context provides useful inductive bias for predicting the contribution of any individual component.
【8】The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
标题:等效错觉:KV缓存自回归推理中的系统性FP 16分歧
链接:https://arxiv.org/abs/2604.15409
作者:Ranjith Chodavarapu,Lei Xu
摘要:KV缓存是自回归Transformer推理中普遍存在的优化,长期以来被认为在数值上等同于无缓存计算。这个假设在标准FP 16精度下是失败的:高速缓存开启和高速缓存关闭执行路径采用不同的浮点累加顺序,由于FP 16的非关联性,在解码的令牌序列中产生确定性的发散。在三个开放重量模型中,(LLaMA-2- 7 B,Mistral-7 B-v0.3,Gemma-2-2B)在GSM 8 K上进行评估,我们观察到所有采样策略(包括贪婪解码)的令牌发散率为100\%,这排除了采样随机性的原因,并且在9种条件中的8种条件下,高速缓存开启会产生更高的准确性,其中精度差用作发散方向是系统的而不是随机的指示符。受控的FP 32伪造将发散降低了8个数量级,消除了令牌翻转,并将翻转率降至0.0%,确认FP 16非关联性是唯一的因果驱动因素。逐层漂移分析揭示了架构上可预测的传播模式:使用分组查询注意力的模型在第一层表现出急剧的发散,而Gemma的较大头部维度和滑动窗口注意力在所有层上产生均匀的积累。最后,整个剩余流的激活修补无法恢复无缓存轨迹,将因果变量定位到有状态KV缓存。这些发现表明,FP 16 KV缓存推理从根本上不等同于重新计算,并为理解现代LLM推理系统中的数值不稳定性提供了一个机械框架。
摘要:KV caching is a ubiquitous optimization in autoregressive transformer inference, long presumed to be numerically equivalent to cache-free computation. This assumption fails under standard FP16 precision: cache-ON and cache-OFF execution paths employ different floating-point accumulation orderings which, due to FP16 non-associativity, produce a deterministic divergence in decoded token sequences. Across three open-weight models (LLaMA-2-7B, Mistral-7B-v0.3, Gemma-2-2B) evaluated on GSM8K, we observe a 100\% token divergence rate across all sampling strategies, including greedy decoding, which rules out sampling randomness as a cause, and also with cache-ON yielding higher accuracy in 8 of 9 conditions, where the accuracy difference serves as an indicator that the divergence direction is systematic rather than random. Controlled FP32 falsification reduces divergence by eight orders of magnitude, eliminates token flips, and drops the flip rate to exactly 0.0\%, confirming FP16 non-associativity as the sole causal driver. Layer-wise drift profiling reveals architecturally predictable propagation patterns: models using Grouped-Query Attention exhibit sharp divergence at the first layer, while Gemma's larger head dimension and sliding window attention produce uniform accumulation across all layers. Finally, activation patching of the entire residual stream fails to recover the cache-free trajectory, localizing the causal variable to the stateful KV cache. These findings establish that FP16 KV cache inference is fundamentally non-equivalent to recomputation and provide a mechanistic framework for understanding numerical instability in modern LLM inference systems.
【9】PRIM-cipal components analysis
标题:PRIM-UNRUNZ成分分析
链接:https://arxiv.org/abs/2604.15538
作者:Tianhao Liu,Daniel Andrés Díaz-Pachón,J. Sunil Rao
备注:12 pages, 46 figures
摘要:有监督的没有免费的午餐定理(NFLT)得到了很好的研究,但无监督的NFLT仍然没有得到充分的研究。对于椭圆分布,我们证明了存在两个同样最优的,科学上有意义的颠簸狩猎策略,是完全相反的,没有普遍的赢家。具体来说,从$\mathbb{R}^d$($d \ge k$)中剥离$k$正交维度,保留每个剥离维度的概率为1-α的分位数间区域,当选择$k$最小主成分(称为最小成分)时,最大化总方差和Frobenius范数,当选择的维度是$k$领先主成分时,最小化它们。这些最优激励基于PRIM的碰撞搜索算法,通过最小化方差或通过最小化体积,从而激励NFLT。我们测试我们的结果的时尚MNIST数据库,剥离最大的主成分捕捉多样性,而剥离最小的主成分隔离流行的风格。
摘要:Supervised No Free Lunch Theorems (NFLTs) are well studied, yet unsupervised NFLTs remain underexplored. For elliptical distributions, we prove that there exist two equally optimal, scientifically meaningful bump-hunting strategies that are exact opposites, with no universal winner. Specifically, peeling $k$ orthogonal dimensions from $\mathbb{R}^d$ ($d \ge k$), retaining an inter-quantile region of probability $1-α$ per peeled dimension, maximizes total variance and Frobenius norm when the $k$ smallest principal components (called pettiest components) are selected, and minimizes them when the selected dimensions are the $k$ leading principal components. These optima inspire PRIM-based bump-hunting algorithms either by minimizing variance or by minimizing volume, thereby motivating an NFLT. We test our results on the Fashion-MNIST database, showing that peeling the largest principal components captures multiplicity, while peeling the smallest principal components isolates popular styles.
检测相关(3篇)
【1】Detecting and Suppressing Reward Hacking with Gradient Fingerprints
标题:利用梯度指纹检测和抑制奖励黑客
链接:https://arxiv.org/abs/2604.16242
作者:Songtao Wang,Quang Hieu Pham,Fangcong Yin,Xinpeng Wang,Jocelyn Qiaochu Chen,Greg Durrett,Xi Ye
摘要:具有可验证奖励的强化学习(RLVR)通常优化结果奖励,而不会对中间推理施加约束。这使得训练容易受到奖励黑客的影响,其中模型利用漏洞(例如,训练数据中的虚假模式),以在不解决预期任务的情况下获得高分。这些奖励黑客行为通常是隐含的,因为中间的思想链(CoT)表面上看起来似乎是合理的,限制了纯文本监控的有效性。我们提出了梯度指纹(GRIFT),一种使用模型的内部计算来检测奖励黑客的方法。给定提示和模型生成的CoT,GRIFT计算以提示为条件的CoT的梯度,并将其压缩为紧凑表示,然后用于评估CoT是否反映奖励黑客行为。在跨越数学、代码和逻辑推理的可验证推理基准测试中,GRIFT的表现大大优于包括CoT Monitor和TRACE在内的强基线,在检测奖励黑客行为方面实现了超过25%的相对改进。此外,将GRIFT集成到推理任务的拒绝微调管道中,减少了奖励黑客攻击,并提高了真实任务目标的性能。我们的结果强调了利用梯度水平表示来评估CoT推理轨迹质量的有希望的方向。我们的代码可在https://github.com/songtao-x/reward_hack上获得。
摘要:Reinforcement learning with verifiable rewards (RLVR) typically optimizes for outcome rewards without imposing constraints on intermediate reasoning. This leaves training susceptible to reward hacking, where models exploit loopholes (e.g., spurious patterns in training data) in the reward function to achieve high scores without solving the intended task. These reward-hacking behaviors are often implicit, as the intermediate chain-of-thought (CoT) may appear plausible on the surface, limiting the effectiveness of purely text-based monitoring. We propose Gradient Fingerprint (GRIFT), a method for detecting reward hacking using models' internal computations. Given a prompt and a model-generated CoT, GRIFT computes gradients of the CoT conditioned on the prompt and compresses them into a compact representation, which is then used to assess whether the CoT reflects reward hacking behavior. Across verifiable reasoning benchmarks spanning math, code, and logical reasoning, GRIFT substantially outperforms strong baselines, including CoT Monitor and TRACE, achieving over 25% relative improvement in detecting reward hacking behavior. Moreover, integrating GRIFT into the rejection fine-tuning pipeline for reasoning tasks reduces reward hacking and improves performance on the true task objective. Our results highlight a promising direction of leveraging gradient level representations for assessing the quality of CoT reasoning traces. Our code is available at: https://github.com/songtao-x/reward_hack.
【2】Early Detection of Acute Myeloid Leukemia (AML) Using YOLOv12 Deep Learning Model
标题:使用YOLOv12深度学习模型早期检测急性骨髓性白血病(ML)
链接:https://arxiv.org/abs/2604.16082
作者:Enas E. Ahmed,Salah A. Aly,Mayar Moner
备注:6 pages, 10 figures, 2 tables
摘要:急性髓系白血病(AML)是最危及生命的血液癌症类型之一,由于各种细胞类型之间的视觉相似性,其准确分类被认为是一项具有挑战性的任务。本研究利用YOLOv12深度学习模型对AML细胞的多类进行分类。我们采用了两种基于细胞和细胞核特征的分割方法,使用色调通道和大津阈值技术对图像进行预处理,然后进行分类。我们的实验表明,YOLOv12与基于细胞的分割Otsu阈值实现了最高水平的验证和测试精度,均达到99.3%。
摘要:Acute Myeloid Leukemia (AML) is one of the most life-threatening type of blood cancers, and its accurate classification is considered and remains a challenging task due to the visual similarity between various cell types. This study addresses the classification of the multiclasses of AML cells Utilizing YOLOv12 deep learning model. We applied two segmentation approaches based on cell and nucleus features, using Hue channel and Otsu thresholding techniques to preprocess the images prior to classification. Our experiments demonstrate that YOLOv12 with Otsu thresholding on cell-based segmentation achieved the highest level of validation and test accuracy, both reaching 99.3%.
【3】RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration
标题:RAGognizer:通过检测头集成实现幻觉感知微调
链接:https://arxiv.org/abs/2604.15945
作者:Fabian Ridder,Laurin Lessel,Malte Schilling
备注:accepted at IJCNN 2026
摘要:检索增强生成(RAG)被广泛用于使用外部信息(例如最近或特定领域的知识)来增强大型语言模型(LLM)的输入。尽管如此,目前的模型仍然会产生封闭域的幻觉,并生成检索到的上下文不支持的内容。目前的检测方法通常将幻觉视为事后问题,依赖于黑盒一致性检查或冻结内部表示的探测。在这项工作中,我们证明了基于内部状态表示的幻觉检测也可以作为一个直接的训练信号。我们介绍了RAGognize,一个带有标记级注释的自然发生的闭域幻觉数据集,以及RAGognizer,一种幻觉感知微调方法,将轻量级检测头集成到LLM中,允许语言建模和幻觉检测的联合优化。这一共同目标迫使模型提高其关于幻觉的内部状态的可分性,同时学习生成形式良好且有意义的响应。在多个基准测试中,RAGognizer实现了最先进的令牌级幻觉检测,同时大大降低了生成过程中的幻觉率,而不会降低语言质量或相关性。
摘要:Retrieval-Augmented Generation (RAG) is widely used to augment the input to Large Language Models (LLMs) with external information, such as recent or domain-specific knowledge. Nonetheless, current models still produce closed-domain hallucinations and generate content that is unsupported by the retrieved context. Current detection approaches typically treat hallucination as a post-hoc problem, relying on black-box consistency checks or probes over frozen internal representations. In this work, we demonstrate that hallucination detection based on internal state representation can also serve as a direct training signal. We introduce RAGognize, a dataset of naturally occurring closed-domain hallucinations with token-level annotations, and RAGognizer, a hallucination-aware fine-tuning approach that integrates a lightweight detection head into an LLM, allowing for the joint optimization of language modeling and hallucination detection. This joint objective forces the model to improve the separability of its internal states regarding hallucinations while simultaneously learning to generate well-formed and meaningful responses. Across multiple benchmarks, RAGognizer achieves state-of-the-art token-level hallucination detection while substantially reducing hallucination rates during generation, without degrading language quality or relevance.
分类|识别(6篇)
【1】Univariate Channel Fusion for Multivariate Time Series Classification
标题:多元时间序列分类的单变量通道融合
链接:https://arxiv.org/abs/2604.16119
作者:Fernando Moro,Vinicius M. A. Souza
备注:International Conference on Pattern Recognition (ICPR 2026)
摘要:多变量时间序列分类(MTSC)在生物医学信号分析和运动监测等领域发挥着重要作用。然而,现有的方法,特别是深度学习模型,通常需要很高的计算资源,这使得它们不适合实时应用或部署在低成本硬件上,如物联网设备和可穿戴系统。在本文中,我们提出了单变量信道融合(UCF)的方法来有效地处理MTSC。UCF通过简单的通道融合策略(如均值、中位数或动态时间规整重心)将多变量时间序列转换为单变量表示。这种转换允许使用最初为单变量时间序列设计的任何分类器,为复杂模型提供灵活且计算量小的替代方案。我们在五个案例研究中评估了UCF,这些案例研究涵盖了不同的应用领域,包括化学监测、脑机接口和人类活动分析。结果表明,UCF往往优于基线方法和最先进的算法为MTSC量身定制,同时实现大幅提高计算效率,特别是在高通道间相关性的问题有效。
摘要:Multivariate time series classification (MTSC) plays a crucial role in various domains, including biomedical signal analysis and motion monitoring. However, existing approaches, particularly deep learning models, often require high computational resources, making them unsuitable for real-time applications or deployment on low-cost hardware, such as IoT devices and wearable systems. In this paper, we propose the Univariate Channel Fusion (UCF) method to deal with MTSC efficiently. UCF transforms multivariate time series into a univariate representation through simple channel fusion strategies such as the mean, median, or dynamic time warping barycenter. This transformation enables the use of any classifier originally designed for univariate time series, providing a flexible and computationally lightweight alternative to complex models. We evaluate UCF in five case studies covering diverse application domains, including chemical monitoring, brain-computer interfaces, and human activity analysis. The results demonstrate that UCF often outperforms baseline methods and state-of-the-art algorithms tailored for MTSC, while achieving substantial gains in computational efficiency, being particularly effective in problems with high inter-channel correlation.
【2】NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition
标题:NeuroLip:一个事件驱动的时空学习框架,用于基于跨场景嘴唇运动的视觉说话人识别
链接:https://arxiv.org/abs/2604.15718
作者:Junguang Yao,Wenye Liu,Stjepan Picek,Yue Zheng
摘要:基于嘴唇运动的视觉说话人识别提供了一种无声的、免提的、行为驱动的生物识别解决方案,即使在声学线索不可用的情况下也能保持有效。与严重依赖于外观依赖表示的传统方法相比,嘴唇运动编码由一致的发音模式和肌肉协调驱动的特定于主体的行为动力学,在环境变化中提供固有的稳定性。然而,由于运动模糊和低动态范围,捕获这些鲁棒、细粒度的动态对于传统的基于帧的相机来说是具有挑战性的。为了利用嘴唇运动的内在稳定性并解决这些传感限制,我们提出了NeuroLip,这是一个基于事件的框架,可以在严格但实用的跨场景协议下捕获细粒度的嘴唇动态:在单一受控条件下进行训练,而识别必须推广到看不见的观看和照明条件。NeuroLip具有1)具有自适应事件加权的时间感知体素编码模块,2)结构感知空间增强器,通过抑制噪声来放大有区别的行为模式,同时保留垂直结构的运动信息,以及3)极性一致性正则化机制,以保留以事件极性编码的运动方向线索。为了便于系统的评估,我们介绍了DVSpeaker,一个全面的基于事件的嘴唇运动数据集,包括50个主题下记录四个不同的观点和照明方案。大量的实验表明,NeuroLip实现了近乎完美的匹配场景精度和强大的跨场景泛化,在看不见的视点上达到了71%以上的精度,在低光条件下达到了近76%,比现有的代表性方法至少高出8.54%。数据集和代码可在https://github.com/JiuZeongit/NeuroLip上公开获得。
摘要
:Visual speaker recognition based on lip motion offers a silent, hands-free, and behavior-driven biometric solution that remains effective even when acoustic cues are unavailable. Compared to traditional methods that rely heavily on appearance-dependent representations, lip motion encodes subject-specific behavioral dynamics driven by consistent articulation patterns and muscle coordination, offering inherent stability across environmental changes. However, capturing these robust, fine-grained dynamics is challenging for conventional frame-based cameras due to motion blur and low dynamic range. To exploit the intrinsic stability of lip motion and address these sensing limitations, we propose NeuroLip, an event-based framework that captures fine-grained lip dynamics under a strict yet practical cross-scene protocol: training is performed under a single controlled condition, while recognition must generalize to unseen viewing and lighting conditions. NeuroLip features a 1) Temporal-aware Voxel Encoding module with adaptive event weighting, 2) Structure-aware Spatial Enhancer that amplifies discriminative behavioral patterns by suppressing noise while preserving vertically structured motion information, and 3) Polarity Consistency Regularization mechanism to retain motion-direction cues encoded in event polarities. To facilitate systematic evaluation, we introduce DVSpeaker, a comprehensive event-based lip-motion dataset comprising 50 subjects recorded under four distinct viewpoint and illumination scenarios. Extensive experiments demonstrate that NeuroLip achieves near-perfect matched-scene accuracy and robust cross-scene generalization, attaining over 71% accuracy on unseen viewpoints and nearly 76% under low-light conditions, outperforming representative existing methods by at least 8.54%. The dataset and code are publicly available at https://github.com/JiuZeongit/NeuroLip.
【3】Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models
标题:奖励加权无分类指导作为自回归模型中的政策改进
链接:https://arxiv.org/abs/2604.15577
作者:Alexander Peysakhovich,William Berman
摘要:考虑产生输出x的自回归模型(例如,问题的答案,分子),其中的每一个都可以由属性向量Y(例如,有益性对无害性,或生物利用度对亲脂性)。一个任意的奖励函数r(y)编码了这些属性之间的权衡。通常,倾斜模型的采样分布以增加这种奖励是在训练时通过强化学习完成的。然而,如果奖励函数发生变化,重新对齐就需要重新训练。在本文中,我们表明,奖励加权无分类指导(RCFG)可以作为一个政策改进运营商在这种情况下,近似倾斜的抽样分布的Q函数。我们将RCFG应用于分子生成,证明它可以在测试时优化新的奖励函数。最后,我们表明,使用RCFG作为老师,并提取到基本政策作为一个热启动显着加快收敛标准RL。
摘要:Consider an auto-regressive model that produces outputs x (e.g., answers to questions, molecules) each of which can be summarized by an attribute vector y (e.g., helpfulness vs. harmlessness, or bio-availability vs. lipophilicity). An arbitrary reward function r(y) encodes tradeoffs between these properties. Typically, tilting the model's sampling distribution to increase this reward is done at training time via reinforcement learning. However, if the reward function changes, re-alignment requires re-training. In this paper, we show that a reward weighted classifier-free guidance (RCFG) can act as a policy improvement operator in this setting, approximating tilting the sampling distribution by the Q function. We apply RCFG to molecular generation, demonstrating that it can optimize novel reward functions at test time. Finally, we show that using RCFG as a teacher and distilling into the base policy to serve as a warm start significantly speeds up convergence for standard RL.
【4】Photonic AI: A Hybrid Diffractive Holographic Neural System for Passive Optical Real-Time Image Classification
标题:Photonic AI:一种用于无源光学实时图像分类的混合型折射全息神经系统
链接:https://arxiv.org/abs/2604.15364
作者:Prakul Sunil Hiremath
备注:18 pages, 3 figures. Operator-theoretic formulation and simulation of a hybrid diffractive-holographic optical neural system
摘要:边缘智能受到在电子存储器层次结构中穿梭数据的能量和延迟成本的限制。光学系统提供了一个根本不同的计算机制:一旦输入波前被发射到结构化介质中,传播,衍射和干涉共同制定了一个线性变换,其成本由波动物理学而不是时钟算法决定。本文开发了一个严格的系统级治疗,该政权,并介绍了一种混合衍射全息图像分类架构。该模型将衍射光学神经网络(DONN)与全息干涉学习(HIBL)算子耦合,从数字优化的相位分布到物理可实现的,可嵌入无源光学元件中的制造兼容干涉图案的正式映射。我们将完整的推理管道表示为编码,相位调制,自由空间传播和强度测量算子的组合,明确哪些量是学习的,哪些是通过设计固定的,以及非线性通过光电检测进入的位置。这种算子理论观点解决了光学ML文献中学习变换和物理实现变换之间持续存在的差距。在MNIST上的物理仿真中,具有大约25,000个相位元件的三层系统在传播受限的情况下实现了91.2%的测试准确度。纳秒级延迟。主要的贡献是不是一个性能要求,但一个精确的计算框架:学习表示可以物理嵌入到结构化的光学介质,使推理是通过波前变换通过被动的,制造的对象,而不是通过顺序的电子乘法累积操作执行。
摘要:Edge intelligence is constrained by the energy and latency costs of shuttling data through electronic memory hierarchies. Optical systems offer a fundamentally different computational regime: once an input wavefront is launched into a structured medium, propagation, diffraction, and interference jointly enact a linear transformation whose cost is determined by wave physics rather than by clocked arithmetic. This paper develops a rigorous systems-level treatment of that regime and introduces a hybrid diffractive holographic architecture for image classification. The proposed model couples a Diffractive Optical Neural Network (DONN) with a Holographic Interference-Based Learning (HIBL) operator a formal map from digitally optimized phase distributions to physically realizable, fabrication-compatible interference patterns embeddable in passive optical elements. We express the full inference pipeline as a composition of encoding, phase modulation, free-space propagation, and intensity measurement operators, making explicit which quantities are learned, which are fixed by design, and where nonlinearity enters through photodetection. This operator-theoretic view resolves a persistent gap in the optical-ML literature between learning a transformation and physically realizing it. In physics-informed simulation on MNIST, a three-layer system with approximately 25,000 phase elements achieves 91.2% test accuracy with propagation-limited nanosecond-scale latency. The primary contribution is not a performance claim but a precise computational framework: learned representations can be physically embedded into structured optical media so that inference is executed by wavefront transformation through a passive, fabricated object rather than by sequential electronic multiply accumulate operations.
【5】ExoNet: Multimodal Deep Learning for TESS Exoplanet Candidate Identification via Phase-Folded Light Curves, Stellar Parameters, and Multi-Head Attention Fusion
标题:ExoNet:通过相折叠光曲线、恒星参数和多头注意力融合进行TESS系外行星候选识别的多模式深度学习
链接:https://arxiv.org/abs/2604.15560
作者:Md. Rashadul Islam
备注:8 pages, 4 figures, 4 tables
摘要:美国宇航局的凌日系外行星调查卫星(TESS)已经确定了数千颗系外行星候选者,但由于人工审查过程的限制,许多人仍未得到证实。本文介绍了ExoNet,这是一个多模态深度学习框架,它使用结合了1D卷积神经网络和多头注意力的后期融合架构,将相位折叠全局和局部光变曲线表示与恒星参数集成在一起。在标记的Kepler数据上训练,ExoNet实现了强大的分类性能,并证明了对TESS数据的有效泛化。应用于200个未经证实的TESS行星候选者,该模型识别出多个高置信度的候选者,其中包括几个位于宜居带内的候选者。结果突出了多模态融合和注意机制在自动化系外行星候选验证中的有效性。
摘要:NASA's Transiting Exoplanet Survey Satellite (TESS) has identified thousands of exoplanet candidates, yet many remain unconfirmed due to the limitations of manual vetting processes. This paper presents ExoNet, a multimodal deep learning framework that integrates phase-folded global and local light curve representations with stellar parameters using a late-fusion architecture combining 1D Convolutional Neural Networks and Multi-Head Attention. Trained on labeled Kepler data, ExoNet achieves strong classification performance and demonstrates effective generalization to TESS data. Applied to 200 unconfirmed TESS planet candidates, the model identifies multiple high-confidence candidates, including several within the habitable zone. The results highlight the effectiveness of multimodal fusion and attention mechanisms in automated exoplanet candidate validation.
【6】A methodology to rank importance of frequencies and channels in electromyography data with Decision Tree classifiers
标题:一种基于决策树分类器的肌电信号频率和通道重要性排序方法
链接:https://arxiv.org/abs/2604.15353
作者:Albert A. Nasybullin,Nursultan Abdullaev,Maksim A. Baranov,Viacheslav V. Koshman,Vitaly A. Mahonin
备注:16 pages, 9 figures, 1 table. Published in Russian Journal of Nonlinear Dynamics, 2024
摘要:本研究提出了一种方法,用于确定最丰富的频率和通道肌电图(EMG)数据,以评估肌肉恢复使用决策树分类。肌电图信号,记录从股外侧肌在蹲下练习,分析在不同的休息时间,以评估最佳的恢复期。通过采用单个决策树分类器,该研究增强了可解释性,提供了对特征重要性的见解-这对于透明度至关重要的医疗和体育环境中的应用至关重要。该实验方案利用网格搜索进行超参数调整和交叉验证,以解决类别不平衡问题,最终实现基于功率谱密度特征的可靠休息间隔分类。结果表明,一个有限的子集的高度信息化的功能提供了足够的准确性,这表明,精简,可解释的模型是有效的肌肉恢复的评估。这种方法可以指导未来的研究,开发紧凑,强大的模型,适用于基于EMG的诊断。
摘要:This study presents a methodology for identifying the most informative frequencies and channels in electromyography (EMG) data to evaluate muscle recovery using Decision Tree classifiers. EMG signals, recorded from the vastus lateralis muscle during squat exercises, were analyzed across varying rest intervals to assess optimal recovery periods. By employing single Decision Tree classifiers, the study enhances interpretability, offering insights into feature importance - essential for applications in medical and sports settings where transparency is critical. The experimental protocol utilized a grid search for hyperparameter tuning and cross-validation to address class imbalance, ultimately achieving a reliable classification of rest intervals based on power spectral density features. The results indicate that a limited subset of highly informative features provides sufficient accuracy, suggesting that streamlined, interpretable models are effective for the evaluation of muscle recovery. This approach can guide future research in developing compact, robust models adapted to EMG-based diagnostics.
优化|敛散性(2篇)
【1】The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
标题:更难的道路:带Bandit反馈的零和游戏中解耦合学习的最后迭代收敛
链接:https://arxiv.org/abs/2604.16087
作者:Côme Fiegel,Pierre Ménard,Tadashi Kozuno,Michal Valko,Vianney Perchet
备注:Accepted at the 42nd International Conference on Machine Learning (ICML 2025)
摘要:研究了具有重复博弈和强盗反馈的零和矩阵对策的学习问题。具体来说,我们专注于开发解耦算法,保证,没有球员之间的沟通,收敛的最后一个纳什均衡。虽然非土匪的情况下已经被广泛研究,这种设置只是最近才被探索,与界的$\mathcal{O}(T^{-1/8})$的可利用性差距。我们表明,对于解耦算法,保证收敛的政策配置文件的纳什均衡是有害的性能,与通常的$Ω(T ^{-1/2})$率的平均迭代收敛的最佳可达到的速度是$Ω(T ^{-1/4})$。然后,我们提出了两种算法,实现这一最佳速率常数和对数因子。第一种算法利用了探索和利用之间的简单权衡,而第二种算法采用了基于两步镜像下降方法的正则化技术。
摘要:We study the problem of learning in zero-sum matrix games with repeated play and bandit feedback. Specifically, we focus on developing uncoupled algorithms that guarantee, without communication between players, the convergence of the last-iterate to a Nash equilibrium. Although the non-bandit case has been studied extensively, this setting has only been explored recently, with a bound of $\mathcal{O}(T^{-1/8})$ on the exploitability gap. We show that, for uncoupled algorithms, guaranteeing convergence of the policy profiles to a Nash equilibrium is detrimental to the performance, with the best attainable rate being $Ω(T^{-1/4})$ in contrast to the usual $Ω(T^{-1/2})$ rate for convergence of the average iterates. We then propose two algorithms that achieve this optimal rate up to constant and logarithmic factors. The first algorithm leverages a straightforward trade-off between exploration and exploitation, while the second employs a regularization technique based on a two-step mirror descent approach.
【2】Towards Universal Convergence of Backward Error in Linear System Solvers
标题:线性系统求解器中向后误差的普遍收敛
链接:https://arxiv.org/abs/2604.16075
作者:Michał Dereziński,Yuji Nakatsukasa,Elizaveta Rebrova
摘要:寻找一个算法,解决$n\times n$线性系统在$O(n^2)$时间复杂度,或$O(n^2 \text{poly}(1/ε))$时解决$ε$相对误差,是一个长期的开放问题,在数值线性代数和理论计算机科学。存在两种用于测量相对误差的主要范例:前向误差(即,从输出到最优解的距离)和向后误差(即,到输出所解决的最近问题的距离)。在大多数先前的研究中,迭代线性系统求解器的收敛性是通过前向误差的各种概念来测量的,因此,在很大程度上取决于输入的条件。然而,数值分析文献长期以来一直主张向后误差作为更实际相关的近似概念。在这项工作中,我们表明-令人惊讶的是-经典和简单的理查森迭代招致最多1/k$(相对)后向误差$k$迭代后的任何半正定(PSD)线性系统,无论其条件数。这个普遍的收敛速度意味着一个$O(n^2/ε)$复杂度的算法来解决PSD线性系统的$ε$向后误差,我们建立类似的或更好的复杂度时,使用各种Krylov求解器超越理查森。然后,通过直接最小化Krylov子空间上的向后误差,我们获得了更快的$O(1/k^2)$通用速度,并且我们将其转化为一个有效的算法MINBERR,复杂度为$O(n^2/\sqrtε)$。我们扩展这种方法,通过正常方程求解一般线性系统,我们经验观察$O(1/k)$收敛。我们报告强大的数值性能,我们的算法对基准问题。
摘要:The quest for an algorithm that solves an $n\times n$ linear system in $O(n^2)$ time complexity, or $O(n^2 \text{poly}(1/ε))$ when solving up to $ε$ relative error, is a long-standing open problem in numerical linear algebra and theoretical computer science. There are two predominant paradigms for measuring relative error: forward error (i.e., distance from the output to the optimum solution) and backward error (i.e., distance to the nearest problem solved by the output). In most prior studies, convergence of iterative linear system solvers is measured via various notions of forward error, and as a result, depends heavily on the conditioning of the input. Yet, the numerical analysis literature has long advocated for backward error as the more practically relevant notion of approximation. In this work, we show that -- surprisingly -- the classical and simple Richardson iteration incurs at most $1/k$ (relative) backward error after $k$ iterations on any positive semidefinite (PSD) linear system, irrespective of its condition number. This universal convergence rate implies an $O(n^2/ε)$ complexity algorithm for solving a PSD linear system to $ε$ backward error, and we establish similar or better complexity when using a variety of Krylov solvers beyond Richardson. Then, by directly minimizing backward error over a Krylov subspace, we attain an even faster $O(1/k^2)$ universal rate, and we turn this into an efficient algorithm, MINBERR, with complexity $O(n^2/\sqrtε)$. We extend this approach via normal equations to solving general linear systems, for which we empirically observe $O(1/k)$ convergence. We report strong numerical performance of our algorithms on benchmark problems.
预测|估计(7篇)
【1】Training Time Prediction for Mixed Precision-based Distributed Training
标题:混合精度分布式训练的训练时间预测
链接:https://arxiv.org/abs/2604.16145
作者:Minchul Kang,Changyong Shin,Jinwoo Jeong,Hyunho Lee,Younghun Go,Gyeongmin Kim,Gyeongsik Yang,Chuck Yoo
摘要
:在分布式深度学习中,准确预测训练时间对于资源分配、成本估计和作业调度至关重要。我们观察到浮点精度设置是训练时间的关键决定因素,导致训练时间变化超过其最小值约2.4倍。然而,现有的分布式训练时间预测研究依赖于静态模型计算图,这些静态模型计算图不能捕获精度变化,包括混合精度。根据我们的实验,不考虑精度的训练时间预测会导致显著的预测误差-平均绝对百分比误差(MAPE)高达147.85%。为了解决这个问题,我们提出了一个精度感知的分布式训练时间预测器,它可以在不同的精度设置中实现稳健的精度,包括混合精度,MAPE为9.8%。
摘要:Accurate prediction of training time in distributed deep learning is crucial for resource allocation, cost estimation, and job scheduling. We observe that the floating-point precision setting is a key determinant of training time, leading to training time variations of ~2.4x over its minimum. However, existing studies on distributed training time prediction rely on static model computation graphs that do not capture precision variations, including mixed precision. According to our experiments, training time prediction without considering precision results in significant prediction errors - reaching up to 147.85% in mean absolute percentage error (MAPE). To address this issue, we propose a precision-aware distributed training time predictor that achieves robust accuracy across diverse precision settings, including mixed precision, with 9.8% MAPE.
【2】Tabular foundation models for in-context prediction of molecular properties
标题:用于分子性质上下文预测的表格基础模型
链接:https://arxiv.org/abs/2604.16123
作者:Karim K. Ben Hicham,Jan G. Rittig,Martin Grohe,Alexander Mitsos
摘要:准确的分子性质预测是药物发现、催化和工艺设计的核心,但实际应用往往受到小数据集的限制。分子基础模型通过学习可转移的分子表示提供了一个有希望的方向;然而,它们通常涉及特定任务的微调,需要机器学习专业知识,并且通常无法超越经典基线。表格基础模型(TFM)提供了一种完全不同的范式:它们通过上下文学习进行预测,无需特定任务的训练即可进行推理。在这里,我们在标准化制药基准和化学工程数据集的低到中等数据制度中评估TFMs。我们评估冻结的分子基础模型表示,以及经典的描述符和指纹。在整个基准测试中,与微调相比,该方法显示出出色的预测性能,同时降低了计算成本,这些优势也转移到实际的工程数据设置中。特别是,将TFM与CheMeleon嵌入相结合,在30个MoleculeACE任务上产生高达100%的获胜率,而紧凑的RDKit 2d和Mordred描述符提供了强大的基于标记的替代方案。分子表示成为TFM性能的关键决定因素,分子基础模型嵌入和2D描述符集在许多任务上都比经典分子指纹提供了实质性的增益。这些结果表明,在上下文学习与TFMs在实际应用中提供了一个高度准确和具有成本效益的替代属性预测。
摘要:Accurate molecular property prediction is central to drug discovery, catalysis, and process design, yet real-world applications are often limited by small datasets. Molecular foundation models provide a promising direction by learning transferable molecular representations; however, they typically involve task-specific fine-tuning, require machine learning expertise, and often fail to outperform classical baselines. Tabular foundation models (TFMs) offer a fundamentally different paradigm: they perform predictions through in-context learning, enabling inference without task-specific training. Here, we evaluate TFMs in the low- to medium-data regime across both standardized pharmaceutical benchmarks and chemical engineering datasets. We evaluate both frozen molecular foundation model representations, as well as classical descriptors and fingerprints. Across the benchmarks, the approach shows excellent predictive performance while reducing computational cost, compared to fine-tuning, with these advantages also transferring to practical engineering data settings. In particular, combining TFMs with CheMeleon embeddings yields up to 100\% win rates on 30 MoleculeACE tasks, while compact RDKit2d and Mordred descriptors provide strong descriptor-based alternatives. Molecular representation emerges as a key determinant in TFM performance, with molecular foundation model embeddings and 2D descriptor sets both providing substantial gains over classic molecular fingerprints on many tasks. These results suggest that in-context learning with TFMs provides a highly accurate and cost-efficient alternative for property prediction in practical applications.
【3】Impact of Nonlinear Power Amplifier on Massive MIMO: Machine Learning Prediction Under Realistic Radio Channel
标题:非线性功率放大器对大规模CDMA的影响:现实无线电通道下的机器学习预测
链接:https://arxiv.org/abs/2604.15977
作者:Marcin Hoffmann,Paweł Kryszkiewicz
备注:Accepted for publication in IEEE Transactions on Vehicular Technology
摘要:M-MIMO是提高无线网络频谱和能量效率的关键技术之一。目前的大多数工作假设M-MIMO阵列配备有线性前端。然而,正在进行的努力,使无线网络更节能推动硬件的极限,其非线性行为出现。这对于多载波系统尤其是常见的问题,例如,OFDM用于4G、5G,也可能用于6 G,其特征在于高峰均功率比。虽然非线性功率放大器(PA)对OFDM信号的影响已经得到了很好的表征,但对于M-MIMO OFDM系统来说,这是一个相对较新的课题。最近的大多数工作要么忽略非线性效应,或利用简化的模型适当的瑞利或LoS无线电信道模型。在本文中,我们首先从理论上描述了在常用的无线电信道模型下的M-MIMO系统中的非线性失真。然后,利用3D-RT软件,我们证明这些模型不是很准确。相反,我们提出了两个模型:一个统计模型和一个基于ML的模型,使用3D-RT结果。所提出的统计模型利用广义极值(GEV)分布来对受害用户的信号失真比(SDR)进行建模,接收非线性失真,例如,作为来自相邻小区的干扰。所提出的ML模型的目的是预测SDR的预定用户(接收非线性失真以及所需的信号),基于无线电信道的空间特性和在M-MIMO天线阵列的每个PA馈电的操作点。然后,预测的SDR可以用于执行PA感知的每用户功率分配。结果表明,约12%的中值增益,在用户吞吐量所取得的建议ML为基础的功率分配方案的国家的最先进的,固定的工作点计划。
摘要:M-MIMO is one of the crucial technologies for increasing spectral and energy efficiency of wireless networks. Most of the current works assume that M-MIMO arrays are equipped with a linear front end. However, ongoing efforts to make wireless networks more energy-efficient push the hardware to the limits, where its nonlinear behavior appears. This is especially a common problem for the multicarrier systems, e.g., OFDM used in 4G, 5G, and possibly also in 6G, which is characterized by a high Peak-to-Average Power Ratio. While the impact of a nonlinear Power Amplifier (PA) on an OFDM signal is well characterized, it is a relatively new topic for the M-MIMO OFDM systems. Most of the recent works either neglect nonlinear effects or utilize simplified models proper for Rayleigh or LoS radio channel models. In this paper, we first theoretically characterize the nonlinear distortion in the M-MIMO system under commonly used radio channel models. Then, utilizing 3D-Ray Tracing (3D-RT) software, we demonstrate that these models are not very accurate. Instead, we propose two models: a statistical one and an ML-based one using 3D-RT results. The proposed statistical model utilizes the Generalized Extreme Value (GEV) distribution to model Signal to Distortion Ratio (SDR) for victim users, receiving nonlinear distortion, e.g., as interference from neighboring cells. The proposed ML model aims to predict SDR for a scheduled user (receiving nonlinear distortion along with the desired signal), based on the spatial characteristics of the radio channel and the operation point of each PA feeding at the M-MIMO antenna array. The predicted SDR can then be used to perform PA-aware per-user power allocation. The results show about 12% median gain in user throughput achieved by the proposed ML-based power allocation scheme over the state-of-the-art, fixed operating point scheme.
【4】Convolutionally Low-Rank Models with Modified Quantile Regression for Interval Time Series Forecasting
标题:用于区间时间序列预测的改进分位数回归卷积低等级模型
链接:https://arxiv.org/abs/2604.15791
作者:Miaoxuan Zhu,Yi Yu,Yuyang Li,Wei Li,Guangcan Liu
摘要:量化预测模型中的不确定性对于可靠的决策至关重要,但仍然是一个重大挑战。区间时间序列预测通过提供预测区间(PI)来提供对该问题的原则性解决方案,预测区间指示真实值落在预测范围内的概率。我们考虑了最近建立的点预测(PF)方法,称为基于学习的卷积核范数最小化(LbCNNM),它通过利用从训练数据中获得的卷积低秩属性直接生成多步预测。虽然LbCNNM在理论上是完整的,在经验上是有效的,但它缺乏固有的不确定性估计能力,这是许多先进预测方法所共有的限制。为了解决这个问题,我们修改了著名的分位数回归(QR),并将其集成到LbCNNM,从而产生了一种新的区间预测方法,称为LbCNNM与修改分位数回归(LbCNNM-MQR)。此外,我们设计了间隔校准技术,以进一步提高PI的准确性。对超过100,000个真实时间序列的广泛实验证明了LbCNNM-MQR的优越性能。
摘要
:The quantification of uncertainty in prediction models is crucial for reliable decision-making, yet remains a significant challenge. Interval time series forecasting offers a principled solution to this problem by providing prediction intervals (PIs), which indicates the probability that the true value falls within the predicted range. We consider a recently established point forecasts (PFs) method termed Learning-Based Convolution Nuclear Norm Minimization (LbCNNM), which directly generates multi-step ahead forecasts by leveraging the convolutional low-rankness property derived from training data. While theoretically complete and empirically effective, LbCNNM lacks inherent uncertainty estimation capabilities, a limitation shared by many advanced forecasting methods. To resolve the issue, we modify the well-known Quantile Regression (QR) and integrate it into LbCNNM, resulting in a novel interval forecasting method termed LbCNNM with Modified Quantile Regression (LbCNNM-MQR). In addition, we devise interval calibration techniques to further improve the accuracy of PIs. Extensive experiments on over 100,000 real-world time series demonstrate the superior performance of LbCNNM-MQR.
【5】Neuromorphic Parameter Estimation for Power Converter Health Monitoring Using Spiking Neural Networks
标题:使用尖峰神经网络进行功率转换器健康监测的神经形态参数估计
链接:https://arxiv.org/abs/2604.15714
作者:Hyeongmeen Baik,Hamed Poursiami,Maryam Parsa,Jinia Roy
备注:10 pages, 11 figures, 4 tables. Submitted to ICONS 2026
摘要:始终在线的转换器健康监测需要亚mW边缘推断,这是基于GPU的物理信息神经网络无法实现的机制。这项工作将尖峰时间处理与物理执行分开:三层泄漏积分和点火SNN估计无源组件参数,而可微分ODE求解器通过将ODE物理损失与展开的尖峰循环解耦来提供物理一致的训练。在EMI损坏的同步降压转换器基准测试中,SNN将集总电阻误差从25.8美元降低到10.2美元,与前馈基线相比,在无源元件的$\pm 10\%$制造公差范围内,在神经形态硬件上的预计${\sim} 270\times $能量降低。持久的膜状态进一步使退化跟踪和事件驱动的故障检测,通过一个$+5.5$的峰值点尖峰率跳跃在突然的故障。该架构具有93\%$ spike sparsity,适合在Intel Loihi 2或BrainChip Akida上始终在线部署。
摘要:Always-on converter health monitoring demands sub-mW edge inference, a regime inaccessible to GPU-based physics-informed neural networks. This work separates spiking temporal processing from physics enforcement: a three-layer leaky integrate-and-fire SNN estimates passive component parameters while a differentiable ODE solver provides physics-consistent training by decoupling the ODE physics loss from the unrolled spiking loop. On an EMI-corrupted synchronous buck converter benchmark, the SNN reduces lumped resistance error from $25.8\%$ to $10.2\%$ versus a feedforward baseline, within the $\pm 10\%$ manufacturing tolerance of passive components, at a projected ${\sim}270\times$ energy reduction on neuromorphic hardware. Persistent membrane states further enable degradation tracking and event-driven fault detection via a $+5.5$ percentage-point spike-rate jump at abrupt faults. With $93\%$ spike sparsity, the architecture is suited for always-on deployment on Intel Loihi 2 or BrainChip Akida.
【6】Predicting Where Steering Vectors Succeed
标题:预测转向载体在哪里取得成功
链接:https://arxiv.org/abs/2604.15557
作者:Jayadev Billa
备注:19 pages, incl. 10 appendix pages, 4 figures, 20 tables
摘要:导向向量对某些概念和层有效,但对其他概念和层无效,并且从业者无法在运行干预之前预测哪个设置适用。我们引入了线性可访问性配置文件(Linear Accessibility Profile,简称LLP),这是一种逐层诊断,将logit镜头重新用作导向向量有效性的预测器。关键度量$A_{lin}}$将模型的解嵌入矩阵应用于中间隐藏状态,不需要训练。在五个模型(Pythia-2.8B到Llama-8B)上的24个受控二元概念家族中,峰值A_{lin}$预测转向有效性为$ρ=+0.86 $到$+0.91$,层选择为$ρ=+0.63 $到$+0.92$。一个三制度框架解释了当差异的手段转向工程,当非线性方法是必要的,当没有方法可以工作。一个实体转向演示证实了端到端的预测:在LAP推荐的层转向重定向Gemma-2-2B和OLMo-2-1B-Instruct上的完成,而中间层(标准启发式)对任何一个模型都没有影响。
摘要:Steering vectors work for some concepts and layers but fail for others, and practitioners have no way to predict which setting applies before running an intervention. We introduce the Linear Accessibility Profile (LAP), a per-layer diagnostic that repurposes the logit lens as a predictor of steering vector effectiveness. The key measure, $A_{\mathrm{lin}}$, applies the model's unembedding matrix to intermediate hidden states, requiring no training. Across 24 controlled binary concept families on five models (Pythia-2.8B to Llama-8B), peak $A_{\mathrm{lin}}$ predicts steering effectiveness at $ρ= +0.86$ to $+0.91$ and layer selection at $ρ= +0.63$ to $+0.92$. A three-regime framework explains when difference-of-means steering works, when nonlinear methods are needed, and when no method can work. An entity-steering demo confirms the prediction end-to-end: steering at the LAP-recommended layer redirects completions on Gemma-2-2B and OLMo-2-1B-Instruct, while the middle layer (the standard heuristic) has no effect on either model.
【7】M3R: Localized Rainfall Nowcasting with Meteorology-Informed MultiModal Attention
标题:M3 R:具有气象信息的多模式关注的局部降雨预报
链接:https://arxiv.org/abs/2604.15377
作者:Sanjeev Panta,Rhett M Morvant,Xu Yuan,Li Chen,Nian-Feng Tzeng
备注:Accepted at IEEE International Conference on Multimedia and Expo (ICME) 2026
摘要:准确和及时的降雨临近预报对于减灾和水资源管理至关重要。尽管最近在深度学习方面取得了进展,但由于有效利用各种多媒体数据源的限制,降水预测仍然具有挑战性。我们介绍了M3R,一种基于气象学的多模态注意力的直接降雨预测架构,它将可视化NEXRAD雷达图像与数值个人气象站(PWS)测量协同结合起来,使用全面的管道对异构气象数据进行时间对齐。通过专门的多模态注意机制,M3R创新地利用气象站时间序列作为查询,选择性地关注空间雷达特征,从而实现降水特征的集中提取。以NEXRAD雷达站为中心的三个100 km * 100 km空间区域的实验结果表明,M3R优于现有方法,在精度,效率和降水检测能力方面实现了实质性的改进。我们的工作为基于多媒体的降水临近预报建立了新的基准,并为业务天气预报系统提供了实用的工具。源代码可在https://github.com/Sanjeev97/M3Rain上获得
摘要:Accurate and timely rainfall nowcasting is crucial for disaster mitigation and water resource management. Despite recent advances in deep learning, precipitation prediction remains challenging due to limitations in effectively leveraging diverse multimedia data sources. We introduce M3R, a Meteorology-informed MultiModal attention-based architecture for direct Rainfall prediction that synergistically combines visual NEXRAD radar imagery with numerical Personal Weather Station (PWS) measurements, using a comprehensive pipeline for temporal alignment of heterogeneous meteorological data. With specialized multimodal attention mechanisms, M3R novelly leverages weather station time series as queries to selectively attend to spatial radar features, enabling focused extraction of precipitation signatures. Experimental results for three spatial areas of 100 km * 100 km centered at NEXRAD radar stations demonstrate that M3R outperforms existing approaches, achieving substantial improvements in accuracy, efficiency, and precipitation detection capabilities. Our work establishes new benchmarks for multimedia-based precipitation nowcasting and provides practical tools for operational weather prediction systems. The source code is available at https://github.com/Sanjeev97/M3Rain
其他神经网络|深度学习|模型|建模(13篇)
【1】Learning to Reason with Insight for Informal Theorem Proving
标题:学习具有非正式定理证明的洞察力推理
链接:https://arxiv.org/abs/2604.16278
作者
:Yunhe Li,Hao Shi,Bowen Deng,Wei Wang,Mengzhe Ruan,Hanxu Hou,Zhongxiang Dai,Siyang Gao,Chao Wang,Shuang Qiu,Linqi Song
摘要:虽然大多数自动定理证明方法依赖于形式证明系统,但非正式定理证明可以更好地与自然语言处理中的大型语言模型(LLM)优势保持一致。在这项工作中,我们确定了非正式定理证明的主要瓶颈是缺乏洞察力,即难以识别解决复杂问题所需的核心技术。为了解决这个问题,我们提出了一个新的框架,旨在培养这种基本的推理技能,使LLM执行有见地的推理。我们提出了$\mathtt{DeepInsightTheorem}$,一个分层数据集,通过明确提取核心技术和最终证明的证明草图来构建非正式证明。为了充分利用该数据集,我们设计了一个渐进式多阶段SFT策略,该策略模仿人类学习过程,引导模型从基本证明写作到富有洞察力的思考。我们对具有挑战性的数学基准的实验表明,这种洞察意识的生成策略显着优于基线。这些结果表明,教学模式,以确定和应用核心技术,可以大大提高他们的数学推理。
摘要:Although most of the automated theorem-proving approaches depend on formal proof systems, informal theorem proving can align better with large language models' (LLMs) strength in natural language processing. In this work, we identify a primary bottleneck in informal theorem proving as a lack of insight, namely the difficulty of recognizing the core techniques required to solve complex problems. To address this, we propose a novel framework designed to cultivate this essential reasoning skill and enable LLMs to perform insightful reasoning. We propose $\mathtt{DeepInsightTheorem}$, a hierarchical dataset that structures informal proofs by explicitly extracting core techniques and proof sketches alongside the final proof. To fully exploit this dataset, we design a Progressive Multi-Stage SFT strategy that mimics the human learning process, guiding the model from basic proof writing to insightful thinking. Our experiments on challenging mathematical benchmarks demonstrate that this insight-aware generation strategy significantly outperforms baselines. These results demonstrate that teaching models to identify and apply core techniques can substantially improve their mathematical reasoning.
【2】Synthetic data in cryptocurrencies using generative models
标题:使用生成模型合成加密货币数据
链接:https://arxiv.org/abs/2604.16182
作者:André Saimon S. Sousa,Otto Pires,Frank Acasiete,Oscar M. Granados,Valéria Loureiro da Silva,Hugo Saba
摘要:数据在整合数字金融生态系统中的市场、服务和产品方面发挥着重要作用。然而,使用真实数据,特别是在金融环境中,可能会导致隐私风险和访问限制,影响机构,研究和建模过程。虽然并非所有的金融数据集都存在这样的限制,但这项工作提出了使用深度学习技术来生成应用于加密货币价格时间序列的合成数据。该方法基于条件生成对抗网络(CGAN),结合了LSTM型递归生成器和MLP递归生成器,以产生统计上一致的合成数据。实验考虑了不同的加密资产,并证明该模型能够再现相关的时间模式,保持市场趋势和动态。通过GANs生成合成序列是模拟金融数据的有效替代方案,显示出市场行为分析和异常检测等应用的潜力,与更复杂的生成方法相比,计算成本更低。
摘要:Data plays a fundamental role in consolidating markets, services, and products in the digital financial ecosystem. However, the use of real data, especially in the financial context, can lead to privacy risks and access restrictions, affecting institutions, research, and modeling processes. Although not all financial datasets present such limitations, this work proposes the use of deep learning techniques for generating synthetic data applied to cryptocurrency price time series. The approach is based on Conditional Generative Adversarial Networks (CGANs), combining an LSTM-type recurrent generator and an MLP discriminator to produce statistically consistent synthetic data. The experiments consider different crypto-assets and demonstrate that the model is capable of reproducing relevant temporal patterns, preserving market trends and dynamics. The generation of synthetic series through GANs is an efficient alternative for simulating financial data, showing potential for applications such as market behavior analysis and anomaly detection, with lower computational cost compared to more complex generative approaches.
【3】Sample Complexity Bounds for Stochastic Shortest Path with a Generative Model
标题:生成模型下随机最短路径的样本复杂性界限
链接:https://arxiv.org/abs/2604.16111
作者:Jean Tarbouriech,Matteo Pirotta,Michal Valko,Alessandro Lazaric
备注:Accepted at the 32nd International Conference on Algorithmic Learning Theory (ALT 2021)
摘要:研究了随机最短路问题中学习ε-最优策略的样本复杂性。当学习者可以访问生成模型时,我们首先推导出样本复杂度界限。我们证明了存在一个最坏情况的SSP实例,它具有$S$状态,$A$动作,最小代价$c_{\min}$,最优策略的最大期望代价$B_{\star}$,其中任何算法都需要至少$Ω(SAB_{\star}^3/(c_{\min}ε^2)$样本才能以高概率返回$ε$-最优策略.令人惊讶的是,这意味着每当c_{\min} = 0$时,SSP问题可能是不可学习的,从而揭示了SSP中的学习严格地比有限时域和折扣设置更难。我们补充这个下界与匹配的算法,对数因子,在一般情况下,和匹配的算法,对数因子,即使当c_{\min} = 0$,但只有在最优策略有一个有界的命中时间的条件下,目标状态。
摘要:We study the sample complexity of learning an $ε$-optimal policy in the Stochastic Shortest Path (SSP) problem. We first derive sample complexity bounds when the learner has access to a generative model. We show that there exists a worst-case SSP instance with $S$ states, $A$ actions, minimum cost $c_{\min}$, and maximum expected cost of the optimal policy over all states $B_{\star}$, where any algorithm requires at least $Ω(SAB_{\star}^3/(c_{\min}ε^2))$ samples to return an $ε$-optimal policy with high probability. Surprisingly, this implies that whenever $c_{\min} = 0$ an SSP problem may not be learnable, thus revealing that learning in SSPs is strictly harder than in the finite-horizon and discounted settings. We complement this lower bound with an algorithm that matches it, up to logarithmic factors, in the general case, and an algorithm that matches it up to logarithmic factors even when $c_{\min} = 0$, but only under the condition that the optimal policy has a bounded hitting time to the goal state.
【4】Prototype-Grounded Concept Models for Verifiable Concept Alignment
标题:基于原型的概念模型,用于可验证的概念一致
链接:https://arxiv.org/abs/2604.16076
作者:Stefano Colamonaco,David Debot,Pietro Barbiero,Giuseppe Marra
摘要:概念瓶颈模型(CBMs)旨在通过人类可理解的概念构建预测来提高深度学习的可解释性,但它们无法验证学习的概念是否与人类的预期含义一致,从而损害了可解释性。我们介绍原型接地概念模型(PGCM),地面概念学习视觉原型:图像部分,作为明确的证据的概念。这种基础可以直接检查概念语义,并支持在原型级别进行有针对性的人为干预,以纠正不一致。从经验上看,PGCM与最先进的建立信任措施的预测性能相当,同时大大提高了透明度、可解释性和可干预性。
摘要:Concept Bottleneck Models (CBMs) aim to improve interpretability in Deep Learning by structuring predictions through human-understandable concepts, but they provide no way to verify whether learned concepts align with the human's intended meaning, hurting interpretability. We introduce Prototype-Grounded Concept Models (PGCMs), which ground concepts in learned visual prototypes: image parts that serve as explicit evidence for the concepts. This grounding enables direct inspection of concept semantics and supports targeted human intervention at the prototype level to correct misalignments. Empirically, PGCMs match the predictive performance of state-of-the-art CBMs while substantially improving transparency, interpretability, and intervenability.
【5】Modern Structure-Aware Simplicial Spatiotemporal Neural Network
标题:现代结构感知的简单时空神经网络
链接:https://arxiv.org/abs/2604.15833
作者:Zhaobo Hu,Vincent Gauthier,Mehdi Naima
摘要:时空建模已经超越了简单的时间序列分析,成为结构时间序列分析的基础。虽然目前的研究广泛采用图神经网络(GNNs)进行空间特征提取,并取得了显著的成功,但这些网络仅限于捕获成对关系,尽管现实世界的网络包含更丰富的拓扑关系。此外,基于GNN的模型面临着计算挑战,这些挑战随着图形复杂性而扩展,限制了它们对大型网络的适用性。为了解决这些限制,我们提出了现代结构感知单纯时空神经网络(ModernSASST),这是第一种利用单纯复杂结构进行时空建模的方法。我们的方法采用时空随机游走在高维单纯复形和集成并行时间卷积网络捕捉高阶拓扑结构,同时保持计算效率。我们的源代码可在GitHub上公开获取\footnote{代码可在https://github.com/ComplexNetTSP/ST_RUM上获取。
摘要:Spatiotemporal modeling has evolved beyond simple time series analysis to become fundamental in structural time series analysis. While current research extensively employs graph neural networks (GNNs) for spatial feature extraction with notable success, these networks are limited to capturing only pairwise relationships, despite real-world networks containing richer topological relationships. Additionally, GNN-based models face computational challenges that scale with graph complexity, limiting their applicability to large networks. To address these limitations, we present Modern Structure-Aware Simplicial SpatioTemporal neural network (ModernSASST), the first approach to leverage simplicial complex structures for spatiotemporal modeling. Our method employs spatiotemporal random walks on high-dimensional simplicial complexes and integrates parallelizable Temporal Convolutional Networks to capture high-order topological structures while maintaining computational efficiency. Our source code is publicly available on GitHub\footnote{Code is available at: https://github.com/ComplexNetTSP/ST_RUM.
【6】Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
标题:打破十亿参数通用机器学习原子间潜力的训练障碍
链接:https://arxiv.org/abs/2604.15821
作者:Yuanchang Zhou,Hongyu Wang,Yiming Du,Yan Wang,Mingzhen Li,Siyu Hu,Xiangyu Zhang,Weijian Liu,Chen Wang,Zhuoqiang Guo,Long Wang,Jingde Bu,Yutong Lu,Guangming Tan,Weile Jia
备注:11 pages, 8 figures
摘要:通用机器学习原子间势(uMLIPs)在涵盖整个周期表中无机材料和有机分子的大量不同数据集上进行了预训练,可作为量子精确物理模拟的基础模型。然而,uMLIP训练需要二阶导数,缺乏相应的并行训练框架;此外,扩展到十亿参数范围会导致计算和通信开销的爆炸性增长,使其训练成为一个巨大的挑战。我们介绍MatRIS-MoE,一个基于不变架构的十亿参数混合专家模型,和{Janus},一个具有硬件感知优化的uMLIP的开创性高维分布式训练框架。部署在两个Exascale超级计算机上,我们的代码在单精度下达到了1.2/1.0 EFLOPS的峰值性能(理论峰值的24/{35.5\%}),并行效率超过90\%,将十亿参数uMLIP的训练从数周压缩到数小时。这项工作为Exascale的AI for Science(AI 4S)基础模型建立了一个新的高水位,并为快速科学发现提供了必要的基础设施。
摘要:Universal Machine Learning Interatomic Potentials (uMLIPs), pre-trained on massively diverse datasets encompassing inorganic materials and organic molecules across the entire periodic table, serve as foundational models for quantum-accurate physical simulations. However, uMLIP training requires second-order derivatives, which lack corresponding parallel training frameworks; moreover, scaling to the billion-parameter regime causes explosive growth in computation and communication overhead, making its training a tremendous challenge. We introduce MatRIS-MoE, a billion-parameter Mixture-of-Experts model built upon invariant architecture, and {Janus}, a pioneering high-dimensional distributed training framework for uMLIPs with hardware-aware optimizations. Deployed across two Exascale supercomputers, our code attains a peak performance of 1.2/1.0 EFLOPS (24\%/{35.5\%} of theoretical peak) in single precision at over 90\% parallel efficiency, compressing the training of billion-parameter uMLIPs from weeks to hours. This work establishes a new high-water mark for AI-for-Science (AI4S) foundation models at Exascale and provides essential infrastructure for rapid scientific discovery.
【7】Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
标题:Stargazer:天体物理约束下人工智能代理的可扩展模型匹配基准环境
链接:https://arxiv.org/abs/2604.15664
作者:Xinge Liu,Terry Jingchen Zhang,Bernhard Schölkopf,Zhijing Jin,Kristen Menou
摘要:自主人工智能代理的兴起表明,需要动态基准环境以及对基于科学的任务的内置反馈来评估这些代理在研究工作中的能力。我们介绍Stargazer,这是一个可扩展的环境,用于使用径向速度(RV)时间序列数据的推断来评估动态迭代物理基础模型拟合任务的AI代理。Stargazer包含三个难度等级的120个任务,包括20个真实的存档案例,涵盖了从高信噪比单行星系统到复杂的多行星配置的各种场景,需要涉及低信噪比分析。我们的八个前沿代理商的评估揭示了数值优化和遵守物理约束之间的差距:虽然代理商经常实现良好的统计拟合,他们经常无法恢复正确的物理系统参数,即使代理商配备了香草技能的限制仍然存在。此外,增加测试时间计算只能产生边际收益,过度的令牌使用通常反映递归故障循环,而不是有意义的探索。Stargazer提供了一个机会,训练,评估,脚手架和规模的战略对模型拟合问题的实际研究相关的今天。我们为人工智能代理设计模拟驱动环境的方法大概可以推广到跨科学领域的许多其他模型拟合问题。源代码和项目网站分别位于https://github.com/Gudmorning2025/Stargazer和https://gudmorning2025.github.io/Stargazer。
摘要:The rise of autonomous AI agents suggests that dynamic benchmark environments with built-in feedback on scientifically grounded tasks are needed to evaluate the capabilities of these agents in research work. We introduce Stargazer, a scalable environment for evaluating AI agents on dynamic, iterative physics-grounded model-fitting tasks using inference on radial-velocity (RV) time series data. Stargazer comprises 120 tasks across three difficulty tiers, including 20 real archival cases, covering diverse scenarios ranging from high-SNR single-planet systems to complex multi-planetary configurations requiring involved low-SNR analysis. Our evaluation of eight frontier agents reveals a gap between numerical optimization and adherence to physical constraints: although agents often achieve a good statistical fit, they frequently fail to recover correct physical system parameters, a limitation that persists even when agents are equipped with vanilla skills. Furthermore, increasing test-time compute yields only marginal gains, with excessive token usage often reflecting recursive failure loops rather than meaningful exploration. Stargazer presents an opportunity to train, evaluate, scaffold, and scale strategies on a model-fitting problem of practical research relevance today. Our methodology to design a simulation-driven environment for AI agents presumably generalizes to many other model-fitting problems across scientific domains. Source code and the project website are available at https://github.com/Gudmorning2025/Stargazer and https://gudmorning2025.github.io/Stargazer, respectively.
【8】Learning Behaviorally Grounded Item Embeddings via Personalized Temporal Contexts
标题:基于个性化时间情境的行为基项目嵌入学习
链接:https://arxiv.org/abs/2604.15581
作者:Rafael T. Sereicikas,Pedro R. Pires,Gregorio F. Azevedo,Tiago A. Almeida
备注:Accepted to be published in UMAP'26, 9 pages, 7 figures
摘要:有效的用户建模需要区分短期和长期偏好演变。虽然项目嵌入已经成为推荐系统的关键组成部分,但Item 2 Vec等标准方法将用户历史视为无序集(项目袋),隐含地假设以分钟为间隔的交互与以月为间隔的交互在语义上一样相关。这种简化掩盖了用户行为的丰富时间结构,模糊了连贯消费会话和渐进兴趣漂移之间的区别。在这项工作中,我们介绍了TAI 2 Vec(时间感知项目到矢量),一个家庭的轻量级嵌入模型,直接集成到表示学习过程中的时间接近。与应用全局时间约束的方法不同,TAI 2 Vec是用户自适应的,可以根据个人交互速度定制其时间定义。我们提出了两个互补的策略:TAI 2 Vec-Disc,它利用个性化的异常检测,动态分段到语义会话的交互,和TAI 2 Vec-Cont,它采用连续的,用户特定的衰减函数来衡量项目的关系,根据其相对的时间距离。在八个不同数据集上的实验结果表明,TAI 2 Vec始终比静态基线产生更准确和更有行为基础的表示,在超过80%的数据集上实现了具有竞争力或优越的性能,改进高达135%。源代码可在https://github.com/UFSCar-LaSID/tai2vec上公开获得。
摘要
:Effective user modeling requires distinguishing between short-term and long-term preference evolution. While item embeddings have become a key component of recommender systems, standard approaches like Item2Vec treat user histories as unordered sets (bag-of-items), implicitly assuming that interactions separated by minutes are as semantically related as those separated by months. This simplification flattens the rich temporal structure of user behavior, obscuring the distinction between coherent consumption sessions and gradual interest drifts. In this work, we introduce TAI2Vec (Time-Aware Item-to-Vector), a family of lightweight embedding models that integrates temporal proximity directly into the representation learning process. Unlike approaches that apply global time constraints, TAI2Vec is user-adaptive, tailoring its temporal definitions to individual interaction paces. We propose two complementary strategies: TAI2Vec-Disc, which utilizes personalized anomaly detection to dynamically segment interactions into semantic sessions, and TAI2Vec-Cont, which employs continuous, user-specific decay functions to weigh item relationships based on their relative temporal distance. Experimental results across eight diverse datasets demonstrate that TAI2Vec consistently produces more accurate and behaviorally grounded representations than static baselines, achieving competitive or superior performance in over 80% of the datasets, with improvements of up to 135%. The source code is publicly available at https://github.com/UFSCar-LaSID/tai2vec.
【9】Learning Affine-Equivariant Proximal Operators
标题:学习仿射等变近端运算符
链接:https://arxiv.org/abs/2604.15556
作者:Oriel Savir,Zhenghan Fang,Jeremias Sulam
备注:9 pages, 4 figures, Accepted at ICASSP 2026
摘要:邻近算子是信号处理和机器学习中许多应用的基础,包括解决不适定逆问题。最近的工作引入了学习邻近网络(LPNs),提供了参数函数,可以为数据驱动和潜在的非凸正则化器计算精确的近似值。然而,在许多情况下,重要的是要包括额外的结构,这些正则化-和它们相应的近似-如移位和尺度等变性。在这项工作中,我们展示了如何获得由神经网络参数化的学习函数,这些神经网络可证明计算精确的邻近算子,同时与移位和缩放等变,我们称之为仿射等变学习邻近网络(AE-LPNs)。我们展示了我们的结果合成,建设性的例子,然后在真实的数据通过去噪的分布设置。我们的等变学习近似增强了对噪声分布和仿射移位的鲁棒性,远远超出了训练分布,提高了学习近似算子的实用性
摘要:Proximal operators are fundamental across many applications in signal processing and machine learning, including solving ill-posed inverse problems. Recent work has introduced Learned Proximal Networks (LPNs), providing parametric functions that compute exact proximals for data-driven and potentially non-convex regularizers. However, in many settings it is important to include additional structure to these regularizers--and their corresponding proximals--such as shift and scale equivariance. In this work, we show how to obtain learned functions parametrized by neural networks that provably compute exact proximal operators while being equivariant to shifts and scaling, which we dub Affine-Equivariant Learned Proximal Networks (AE-LPNs). We demonstrate our results on synthetic, constructive examples, and then on real data via denoising in out-of-distribution settings. Our equivariant learned proximals enhance robustness to noise distributions and affine shifts far beyond training distributions, improving the practical utility of learned proximal operators
【10】$π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
标题:$pi_{0.7}$:具有紧急能力的可操纵多面手机器人基础模型
链接:https://arxiv.org/abs/2604.15483
作者:Physical Intelligence,Bo Ai,Ali Amin,Raichelle Aniceto,Ashwin Balakrishna,Greg Balke,Kevin Black,George Bokinsky,Shihao Cao,Thomas Charbonnier,Vedant Choudhary,Foster Collins,Ken Conley,Grace Connors,James Darpinian,Karan Dhabalia,Maitrayee Dhaka,Jared DiCarlo,Danny Driess,Michael Equi,Adnan Esmail,Yunhao Fang,Chelsea Finn,Catherine Glossop,Thomas Godden,Ivan Goryachev,Lachlan Groom,Haroun Habeeb,Hunter Hancock,Karol Hausman,Gashon Hussein,Victor Hwang,Brian Ichter,Connor Jacobsen,Szymon Jakubczak,Rowan Jen,Tim Jones,Gregg Kammerer,Ben Katz,Liyiming Ke,Mairbek Khadikov,Chandra Kuchi,Marinda Lamb,Devin LeBlanc,Brendon LeCount,Sergey Levine,Xinyu Li,Adrian Li-Bell,Vladislav Lialin,Zhonglin Liang,Wallace Lim,Yao Lu,Enyu Luo,Vishnu Mano,Nandan Marwaha,Aikys Mongush,Liam Murphy,Suraj Nair,Tyler Patterson,Karl Pertsch,Allen Z. Ren,Gavin Schelske,Charvi Sharma,Baifeng Shi,Lucy Xiaoyang Shi,Laura Smith,Jost Tobias Springenberg,Kyle Stachowicz,Will Stoeckle,Jiaming Tang,Jimmy Tanner,Shalom Tekeste,Marcel Torne,Kyle Vedder,Quan Vuong,Anna Walling,Haohuan Wang,Jason Wang,XuDong Wang,Chris Whalen,Samuel Whitmore,Blake Williams,Charles Xu,Sukwon Yoo,Lili Yu,Wuming Zhang,Zhuoyang Zhang,Ury Zhilinsky
备注:Website: https://www.pi.website/blog/pi07
摘要:我们提出了一种新的机器人基础模型,称为$π_{0.7}$,可以在广泛的场景中实现强大的开箱即用性能。$π_{0.7}$可以在看不见的环境中遵循不同的语言指令,包括使用各种厨房电器的多阶段任务,提供zero-shot交叉具体化,例如使机器人能够在之前没有看到任务的情况下折叠衣物,并执行具有挑战性的任务,例如以与更专业的RL微调模型相匹配的性能水平操作开箱即用的浓缩咖啡机。$π_{0.7}$背后的主要思想是在训练期间使用不同的上下文条件。这种包含在提示中的条件信息,使得精确地操纵模型以不同的策略执行许多任务成为可能。它不仅取决于描述它应该做什么的语言命令,而且还取决于描述它应该做什么的方式或策略的其他多模态信息,包括关于任务绩效和子目标图像的元数据。这使得$π_{0.7}$能够使用非常多样化的数据,包括演示、潜在的次优(自主)数据(包括故障)以及来自非机器人源的数据。我们的实验评估$π_{0.7}$在许多任务与多个机器人平台,对任务,需要速度和灵活性,语言以下,和合成任务概括。
摘要:We present a new robotic foundation model, called $π_{0.7}$, that can enable strong out-of-the-box performance in a wide range of scenarios. $π_{0.7}$ can follow diverse language instructions in unseen environments, including multi-stage tasks with various kitchen appliances, provide zero-shot cross-embodiment generalization, for example enabling a robot to fold laundry without seeing the task before, and perform challenging tasks such as operating an espresso machine out of the box at a level of performance that matches much more specialized RL-finetuned models. The main idea behind $π_{0.7}$ is to use diverse context conditioning during training. This conditioning information, contained in the prompt, makes it possible to steer the model precisely to perform many tasks with different strategies. It is conditioned not just on a language command that describes what it should do, but on additional multimodal information that also describes the manner or strategy in which it should do it, including metadata about task performance and subgoal images. This enables $π_{0.7}$ to use very diverse data, including demonstrations, potentially suboptimal (autonomous) data including failures, and data from non-robot sources. Our experiments evaluate $π_{0.7}$ across numerous tasks with multiple robot platforms, on tasks that require speed and dexterity, language following, and compositional task generalization.
【11】Python library supporting Discrete Variational Formulations and training solutions with Collocation-based Robust Variational Physics Informed Neural Networks (DVF-CRVPINN)
标题:Python库通过基于配置的鲁棒变分物理知情神经网络(DVF-CRVINN)支持离散变分公式和训练解决方案
链接:https://arxiv.org/abs/2604.15398
作者:Tomasz Służalec,Marcin Łoś,Askold Vilkha,Maciej Paszyński
备注:Python library, Robust Variational Physics-Informed Neural Networks, Collocation Methods, Robust loss, Stokes Equations, Laplace problem
摘要:我们探讨了使用离散弱公式求解偏微分方程(PDE)的可能性。我们提出了一个编程环境,定义一个离散的计算域,引入定义在一组点上的离散函数,构造离散内积,并引入离散弱公式采用克罗内克三角洲测试功能。在此基础上,我们提出了一个离散的神经网络表示,训练的解决方案功能定义在一个离散的点集,并采用离散的有限差分导数的自动微分程序。作为一个具有挑战性的计算模型的例子,我们专注于二维斯托克斯方程,定义在一组离散的点。我们训练的解决方案,使用离散弱残差和Adamax算法与离散自动微分的离散梯度。尽管引入了Python环境,我们还提供了一个严格的数学公式,基于离散弱公式,证明了损失函数的适定性和鲁棒性。离散弱公式的解决方案是基于神经网络训练,采用了一个强大的损失函数,是有关的真实误差。通过这种方式,我们在神经网络的训练过程中对数值误差进行了鲁棒控制。除了斯托克斯公式,我们还解释了建议库使用拉普拉斯问题公式的功能。
摘要
:We explore the possibility of solving Partial Differential Equations (PDEs) using discrete weak formulations. We propose a programming environment for defining a discrete computational domain, introducing discrete functions defined over a set of points, constructing discrete inner products, and introducing discrete weak formulations employing Kronecker delta test functions. Building on this setup, we propose a discrete neural network representation, training the solution function defined over a discrete set of points and employing discrete finite difference derivatives in the automatic differentiation procedures. As a challenging computational model example, we focus on Stokes equations in two-dimensions, defined over a discrete set of points. We train the solution using the discrete weak residual and the Adamax algorithm with discrete automatic differentiation of the discrete gradients. Despite introducing the python environment, we also provide a rigorous mathematical formulation based on discrete weak formulations, proving the well-posedness and robustness of the loss function. The solution of the discrete weak formulations is based on neural network training employing a robust loss function that is related to the true error. In this way, we have a robust control of the numerical error during the training of the neural networks. Besides the Stokes formulation, we also explain the functionality of the proposed library using the Laplace problem formulation.
【12】Discovering quantum phenomena with Interpretable Machine Learning
标题:利用可解释机器学习发现量子现象
链接:https://arxiv.org/abs/2604.16015
作者:Paulin de Schoulepnikoff,Hendrik Poulsen Nautrup,Hans J. Briegel,Gorka Muñoz-Gil
摘要:可解释的机器学习技术正在成为从复杂的量子数据中提取物理见解的重要工具。我们基于变分自编码器的最新进展,证明这种模型可以从广泛的一类未标记量子数据集中学习物理上有意义和可解释的表示。仅从原始测量数据来看,学习的表示就揭示了关于量子相空间底层结构的丰富信息。我们进一步增强学习管道与符号方法,使紧凑的分析描述符,作为顺序参数的学习表示中出现的不同制度的发现。我们展示了实验里德伯原子快照的框架,集群伊辛模型的经典阴影,和混合离散连续的费米子数据,揭示了以前未报道的现象,如角排序模式的里德伯阵列。这些结果为从不同的量子数据集中自动和可解释地发现物理定律建立了一个通用框架。所有方法都可以通过qdisc获得,qdisc是一个开源Python库,旨在使更广泛的社区可以访问这些工具。
摘要:Interpretable machine learning techniques are becoming essential tools for extracting physical insights from complex quantum data. We build on recent advances in variational autoencoders to demonstrate that such models can learn physically meaningful and interpretable representations from a broad class of unlabeled quantum datasets. From raw measurement data alone, the learned representation reveals rich information about the underlying structure of quantum phase spaces. We further augment the learning pipeline with symbolic methods, enabling the discovery of compact analytical descriptors that serve as order parameters for the distinct regimes emerging in the learned representations. We demonstrate the framework on experimental Rydberg-atom snapshots, classical shadows of the cluster Ising model, and hybrid discrete-continuous fermionic data, revealing previously unreported phenomena such as a corner-ordering pattern in the Rydberg arrays. These results establish a general framework for the automated and interpretable discovery of physical laws from diverse quantum datasets. All methods are available through qdisc, an open-source Python library designed to make these tools accessible to the broader community.
【13】Machine learning approaches to uncover the neural mechanisms of motivated behaviour: from ADHD to individual differences in effort and reward sensitivity
标题:机器学习方法揭示动机行为的神经机制:从多动症到努力和回报敏感性的个体差异
链接:https://arxiv.org/abs/2604.15363
作者:Nam Trinh
备注:PhD thesis, Dublin City University, December 2025. 194 pages
摘要:动机行为依赖于大脑评估努力和奖励的能力。这些过程中的失调导致了一系列的疾病,从注意力缺陷/多动障碍(ADHD)中的多动症到冷漠中目标导向行为的减少。本论文通过三个主要研究,利用脑电图(EEG)研究了ADHD的神经机制,并利用神经成像技术研究了努力和奖励敏感性的个体差异,应用机器学习方法。在研究1中,使用基于任务和静息状态的EEG与机器学习模型对患有ADHD的成年人和健康对照进行分类。在停止信号任务期间,在基于任务的EEG上训练的机器学习分类器的表现优于在静息状态EEG上训练的机器学习分类器,其中最强的预测特征来自额中央和顶叶区域的伽马波段谱功率。在研究2中,弥散MRI和全脑置换分析确定了白质完整性与反映努力和奖励敏感性的计算建模参数之间的关联,SMA连接的神经束作为中心枢纽出现。在研究3中,来自结构T1加权MRI的灰质体积被用于检查努力敏感性、奖励敏感性和亚临床冷漠的相关性,机器学习证实了奖励敏感性和冷漠水平的鲁棒解码。在所有研究中,额顶叶回路成为努力评估和奖励处理的核心。这些发现可以作为神经生物标志物,用于提高ADHD和动机障碍的诊断准确性,并指导个性化的神经技术干预。
摘要:Motivated behaviour relies on the brain's capacity to evaluate effort and reward. Dysregulation within these processes contributes to a spectrum of conditions, from hyperactivity in attention-deficit/hyperactivity disorder (ADHD) to diminished goal-directed behaviour in apathy. This thesis investigates the neural mechanisms underlying ADHD using electroencephalography (EEG) and examines individual differences in effort and reward sensitivity using neuroimaging, applying machine learning approaches through three main studies. In Study 1, task-based and resting-state EEG were employed with machine learning models to classify adult individuals with ADHD and healthy controls. Machine learning classifiers trained on task-based EEG during a stop signal task outperformed those trained on resting-state EEG, with the strongest predictive features arising from gamma-band spectral power over fronto-central and parietal regions. In Study 2, diffusion MRI and whole-brain permutation-based analyses identified associations between white matter integrity and computationally modelled parameters reflecting effort and reward sensitivity, with SMA-connected tracts emerging as a central hub. In Study 3, grey matter volumes from structural T1-weighted MRI were used to examine correlates of effort sensitivity, reward sensitivity, and subclinical apathy, with machine learning confirming robust decoding of reward sensitivity and apathy levels. Across studies, fronto-parietal circuits emerged as central to effort valuation and reward processing. These findings may serve as neural biomarkers for improving diagnostic accuracy in ADHD and motivational impairments, and for guiding personalised neurotechnological interventions.
其他(30篇)
【1】Geometric regularization of autoencoders via observed stochastic dynamics
标题:通过观察到的随机动力学对自动编码器进行几何正规化
链接:https://arxiv.org/abs/2604.16282
作者:Sean Hill,Felix X. -F. Ye
摘要:具有缓慢或亚稳态行为的随机动力系统在长时间尺度上在高维环境空间中的未知低维流形上演化。从短突发环境集合构建简化的模拟器是一个长期存在的问题:像ATLAS这样的局部图表方法遭受指数地标缩放和每步重投影,而自动编码器替代方案使切束几何约束不良,并且错误传播到学习的漂移和扩散中。我们观察到环境协方差~$Λ$已经编码了坐标不变的切空间信息,其范围跨越切丛。使用这个,我们构建了一个切线束惩罚和一个逆一致性惩罚的三阶段管道(图表学习,潜在漂移,潜在扩散),学习一个单一的非线性图表和潜在漂移。惩罚导致一个函数空间度量,$ρ$-度量,严格弱于Sobolev $H^1$范数,但达到相同的图表质量推广率对数因子。对于漂移,我们通过伊藤公式在学习的编码器上推导出编码器回调目标,并证明了一个偏差分解,表明标准解码器端公式对于任何不完美的图表都具有系统误差。在$W^{2,\infty}$图收敛的假设下,图级误差可控地传播到环境动力学的弱收敛和径向平均首次通过时间的收敛。四个表面上的实验嵌入在高达201美元的环境尺寸减少径向MFPT误差由50美元-70\%$下旋转动力学,并实现最低的井间MFPT误差在大多数表面-过渡对下亚稳态米勒-布朗朗之万动力学,同时减少端到端的环境系数误差由一个数量级相对于一个不规则的自动编码器。
摘要
:Stochastic dynamical systems with slow or metastable behavior evolve, on long time scales, on an unknown low-dimensional manifold in high-dimensional ambient space. Building a reduced simulator from short-burst ambient ensembles is a long-standing problem: local-chart methods like ATLAS suffer from exponential landmark scaling and per-step reprojection, while autoencoder alternatives leave tangent-bundle geometry poorly constrained, and the errors propagate into the learned drift and diffusion. We observe that the ambient covariance~$Λ$ already encodes coordinate-invariant tangent-space information, its range spanning the tangent bundle. Using this, we construct a tangent-bundle penalty and an inverse-consistency penalty for a three-stage pipeline (chart learning, latent drift, latent diffusion) that learns a single nonlinear chart and the latent SDE. The penalties induce a function-space metric, the $ρ$-metric, strictly weaker than the Sobolev $H^1$ norm yet achieving the same chart-quality generalization rate up to logarithmic factors. For the drift, we derive an encoder-pullback target via Itô's formula on the learned encoder and prove a bias decomposition showing the standard decoder-side formula carries systematic error for any imperfect chart. Under a $W^{2,\infty}$ chart-convergence assumption, chart-level error propagates controllably to weak convergence of the ambient dynamics and to convergence of radial mean first-passage times. Experiments on four surfaces embedded in up to $201$ ambient dimensions reduce radial MFPT error by $50$--$70\%$ under rotation dynamics and achieve the lowest inter-well MFPT error on most surface--transition pairs under metastable Müller--Brown Langevin dynamics, while reducing end-to-end ambient coefficient errors by up to an order of magnitude relative to an unregularized autoencoder.
【2】Beyond Distribution Sharpening: The Importance of Task Rewards
标题:超越分配精简:任务奖励的重要性
链接:https://arxiv.org/abs/2604.16259
作者:Sarthak Mittal,Leo Gagnon,Guillaume Lajoie
摘要:在将基于任务奖励的强化学习(RL)集成到其训练管道中之后,前沿模型已经展示了卓越的能力,使系统能够从纯推理模型发展成为复杂的代理。然而,关于强化学习是否真正在基础模型中灌输新技能或仅仅强化其现有分布以激发潜在能力的争论仍然存在。为了解决这种二分法,我们提出了一个明确的比较分布锐化和基于任务奖励的学习,利用RL作为工具来实现这两种范式。我们的分析揭示了分布锐化的固有局限性,从第一原理论证了最优值如何以及为什么会是不利的,并且这种方法从根本上是不稳定的。此外,我们在数学数据集上使用Llama-3.2-3B-Instruct、Qwen2.5-3B-Instruct和Qwen 3 - 4 B-Instruct-2507进行的实验证实,锐化产生的收益有限,而结合基于任务的奖励信号可以极大地帮助实现稳健的性能改进和稳定的学习。
摘要:Frontier models have demonstrated exceptional capabilities following the integration of task-reward-based reinforcement learning (RL) into their training pipelines, enabling systems to evolve from pure reasoning models into sophisticated agents. However, debate persists regarding whether RL genuinely instills new skills within a base model or merely sharpens its existing distribution to elicit latent capabilities. To address this dichotomy, we present an explicit comparison between distribution sharpening and task-reward-based learning, utilizing RL as a tool to implement both paradigms. Our analysis reveals the inherent limitations of distribution sharpening, demonstrating from first principles how and why the optima can be unfavorable and the approach fundamentally unstable. Furthermore, our experiments using Llama-3.2-3B-Instruct, Qwen2.5-3B-Instruct and Qwen3-4B-Instruct-2507 on math datasets confirm that sharpening yields limited gains, whereas incorporating task-based reward signal can greatly help achieve robust performance improvements and stable learning.
【3】Joint-Centric Dual Contrastive Alignment with Structure-Preserving and Information-Balanced Regularization
标题:具有结构保留和信息平衡规则化的关节中心双重对比对齐
链接:https://arxiv.org/abs/2604.16247
作者:Habibeh Naderi,Behrouz Haji Soleimani,Stan Matwin
摘要:我们提出了HILBERT(HILBERT长序列平衡嵌入与相互对比训练),一个跨注意的多模态框架,用于在低资源数据设置中从长的分段序列中学习文档级音频文本表示。HILBERT利用冻结的预训练语音和语言编码器来提取片段级特征,这些特征通过跨模态注意力和自注意力池聚合,以形成特定模态的文档表示和联合跨注意力嵌入。为了在严重的音频-文本维度不平衡的情况下保持模态特定结构的同时对齐模态,我们引入了一个互惠的双重对比目标,该目标同时对齐音频到联合和文本到联合的表示,而不是单独直接对比音频和文本。两个辅助正则化器进一步稳定长序列融合:中心核对齐(CKA)损失,保留每个模态和联合嵌入之间的结构一致性,以及互信息平衡损失,通过均衡从音频和文本到联合空间的信息流来防止单一模态的主导地位。对于下游预测,HILBERT采用混合专家(MoE)分类器连接音频,文本和联合表示,以适应异构标签制度。对多个音频文本主干组合的广泛评估表明,HILBERT学习语义上有意义的长序列表示,并在高度不平衡的多类设置上实现了卓越的性能。
摘要:We propose HILBERT (HIerarchical Long-sequence Balanced Embedding with Reciprocal contrastive Training), a cross-attentive multimodal framework for learning document-level audio-text representations from long, segmented sequences in low-resource data settings. HILBERT leverages frozen pre-trained speech and language encoders to extract segment-level features, which are aggregated via cross-modal attention and self-attentive pooling to form modality-specific document representations and a joint cross-attentive embedding. To align modalities while preserving modality-specific structure under severe audio-text dimensional imbalance, we introduce a reciprocal dual contrastive objective that simultaneously aligns audio-to-joint and text-to-joint representations, rather than directly contrasting audio and text alone. Two auxiliary regularizers further stabilize long-sequence fusion: a Centered Kernel Alignment (CKA) loss that preserves structural consistency between each modality and the joint embedding, and a mutual information balancing loss that prevents dominance of a single modality by equalizing information flow from audio and text into the joint space. For downstream prediction, HILBERT employs a Mixture-of-Experts (MoE) classifier over concatenated audio, text, and joint representations to accommodate heterogeneous label regimes. Extensive evaluation across multiple audio-text backbone combinations demonstrates that HILBERT learns semantically meaningful long-sequence representations and achieves superior performance on highly imbalanced multi-class settings.
【4】Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction
标题:通过概率偏差修正增强人工智能和动态亚季节预测
链接:https://arxiv.org/abs/2604.16238
作者:Hannah Guan,Soukayna Mouatadid,Paulo Orenstein,Judah Cohen,Haiyu Dong,Zekun Ni,Jeremy Berman,Genevieve Flaspohler,Alex Lu,Jakob Schloer,Joshua Talib,Jonathan A. Weyn,Lester Mackey
摘要:决策者依靠天气预报来种植作物,管理野火,分配水和能源,并为极端天气做好准备。如今,由于基于物理的动态模型和数据驱动的人工智能(AI)模型的稳步发展,这种预测在两周内具有前所未有的准确性。然而,由于复合误差和持续的偏差,模型技能在亚季节时间尺度(提前2 - 6周)急剧下降。为了应对这种退化,我们引入了概率偏差校正(PBC),这是一种机器学习框架,通过学习校正历史概率预测来大幅减少系统误差。当应用于欧洲中期天气预报中心(ECMWF)的领先动力学和人工智能模型时,PBC使人工智能预测系统的亚季节技能翻了一番,并提高了91%压力,92%温度和98%降水目标的操作去偏动力学模型的技能。在ECMWF的2025年实时预报竞赛中,其全球预报在所有天气变量和提前时间方面均名列第一,优于来自六个业务预报中心的动力模型、国际动力多模型集合、ECMWF的人工智能预报系统以及全球34个团队的预报系统。这些概率技能的提高转化为对极端事件更准确的预测,并有可能改善脆弱社区的农业规划、能源管理和备灾工作。
摘要
:Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes. Today, such forecasts enjoy unprecedented accuracy out to two weeks thanks to steady advances in physics-based dynamical models and data-driven artificial intelligence (AI) models. However, model skill drops precipitously at subseasonal timescales (2 - 6 weeks ahead), due to compounding errors and persistent biases. To counter this degradation, we introduce probabilistic bias correction (PBC), a machine learning framework that substantially reduces systematic error by learning to correct historical probabilistic forecasts. When applied to the leading dynamical and AI models from the European Centre for Medium-Range Weather Forecasts (ECMWF), PBC doubles the subseasonal skill of the AI Forecasting System and improves the skill of the operationally-debiased dynamical model for 91% of pressure, 92% of temperature, and 98% of precipitation targets. We designed PBC for operational deployment, and, in ECMWF's 2025 real-time forecasting competition, its global forecasts placed first for all weather variables and lead times, outperforming the dynamical models from six operational forecasting centers, an international dynamical multi-model ensemble, ECMWF's AI Forecasting System, and the forecasting systems of 34 teams worldwide. These probabilistic skill gains translate into more accurate prediction of extreme events and have the potential to improve agricultural planning, energy management, and disaster preparedness in vulnerable communities.
【5】OT on the Map: Quantifying Domain Shifts in Geographic Space
标题:地图上的OT:量化地理空间中的领域转移
链接:https://arxiv.org/abs/2604.16220
作者:Haoran Zhang,Livia Betti,Konstantin Klemmer,Esther Rolf,David Alvarez-Melis
摘要:在地理数据的计算机视觉和机器学习中,域外泛化是一个普遍存在的挑战,这是由于全球数据覆盖不均匀和地理区域之间的分布变化造成的。虽然模型经常在一个地区训练并部署在另一个地区,但没有原则性的方法来确定这种跨地区适应何时会成功。分布之间的距离的定义良好的概念可以有效地量化新目标域与用于模型训练的域相比的差异,这反过来可以支持模型训练和部署决策。在本文中,我们提出了一种利用地理信息与最佳运输方法(GeoSpOT)计算地理空间域之间的距离的策略。在我们的实验中,GeoSpOT距离成为跨域传输难度的有效预测因子。我们进一步证明,来自预训练的位置编码器的嵌入提供了与图像/文本嵌入相当的信息,尽管仅依赖于纬度对作为输入。这允许用户获得地理空间模型的域外性能的近似值,即使确切的下游任务未知,或者没有特定于任务的数据可用。基于这些发现,我们表明GeoSpOT距离可以抢先指导数据选择,并使预测工具能够分析模型可能表现不佳的区域。
摘要:In computer vision and machine learning for geographic data, out-of-domain generalization is a pervasive challenge, arising from uneven global data coverage and distribution shifts across geographic regions. Though models are frequently trained in one region and deployed in another, there is no principled method for determining when this cross-region adaptation will be successful. A well-defined notion of distance between distributions can effectively quantify how different a new target domain is compared to the domains used for model training, which in turn could support model training and deployment decisions. In this paper, we propose a strategy for computing distances between geospatial domains that leverages geographic information with Optimal Transport methods (GeoSpOT). In our experiments, GeoSpOT distances emerge as effective predictors of cross-domain transfer difficulty. We further demonstrate that embeddings from pretrained location encoders provide information comparable to image/text embeddings, despite relying solely on longitude-latitude pairs as input. This allows users to get an approximation of out-of-domain performance for geospatial models, even when the exact downstream task is unknown, or no task-specific data is available. Building on these findings, we show that GeoSpOT distances can preemptively guide data selection and enable predictive tools to analyze regions where a model is likely to underperform.
【6】SCRIPT: Implementing an Intelligent Tutoring System for Programming in a German University Context
标题:CLARUT:在德国大学环境中实施智能编程辅导系统
链接:https://arxiv.org/abs/2604.16117
作者:Alina Deriyeva,Jesper Dannath,Benjamin Paassen
备注:In: Cristea, A.I., Walker, E., Lu, Y., Santos, O.C., Isotani, S. (eds) Artificial Intelligence in Education. Posters and Late Breaking Results, Workshops and Tutorials, Industry and Innovation Tracks, Practitioners, Doctoral Consortium, Blue Sky, and WideAIED. AIED 2025. Communications in Computer and Information Science, vol 2590 . Springer, Cham
摘要:实践和广泛的练习在编程教育中是必不可少的。智能辅导系统(ITS)是一个可行的选择,即使在人类导师不可用时,也可以为编程学生提供个性化的提示和建议。然而,以前的ITS编程很少支持Python编程语言,主要集中在入门编程上,很少考虑生成模型的最新发展。我们的目标是为Python编程建立一个新的ITS,它具有高度的适应性,既可以作为教学平台,也可以作为研究平台,提供接口插入提示机制(例如\通过大型语言模型),并在德国特别具有挑战性的监管环境中工作,即符合欧洲数据保护法规,欧洲人工智能法案和德国研究基金会的道德框架。在本文中,我们目前的ITS的现状与未来的发展方向,以及讨论的挑战和机遇,以改善系统的描述。
摘要:Practice and extensive exercises are essential in programming education. Intelligent tutoring systems (ITSs) are a viable option to provide individualized hints and advice to programming students even when human tutors are not available. However, prior ITS for programming rarely support the Python programming language, mostly focus on introductory programming, and rarely take recent developments in generative models into account. We aim to establish a novel ITS for Python programming that is highly adaptable, serves both as a teaching and research platform, provides interfaces to plug in hint mechanisms (e.g.\ via large language models), and works inside the particularly challenging regulatory environment of Germany, that is, conforming to the European data protection regulation, the European AI act, and ethical framework of the German Research Foundation. In this paper, we present the description of the current state of the ITS along with future development directions, as well as discuss the challenges and opportunities for improving the system.
【7】Stylistic-STORM (ST-STORM) : Perceiving the Semantic Nature of Appearance
标题:文体风暴(ST-STORM):感知外观的语义本质
链接:https://arxiv.org/abs/2604.16086
作者:Hamed Ouattara,Pierre Duthon,Pascal Houssam Salmane,Frédéric Bernardin,Omar Ait Aider
备注:20 pages, 16 figures, ICPR 2026 (28th International Conference on Pattern Recognition)
摘要:由MoCo或DINO说明的自监督学习(SSL)的主要范例之一旨在通过捕获对某些图像变换(如照明或几何变化)不敏感的特征来产生鲁棒的表示。当目标是独立于物体的外观来识别物体时,这种策略是合适的。然而,一旦外表本身构成了区别性信号,它就变得适得其反。例如,在天气分析中,雨条纹、雪的粒度、大气散射以及反射和光晕都不是噪音:它们携带着基本信息。在自动驾驶等关键应用中,忽略这些提示是有风险的,因为抓地力和能见度直接取决于地面条件和大气条件。我们介绍ST-STORM,一个混合SSL框架,将外观(风格)作为一个语义模态从内容中解脱出来。我们的架构明确地分离两个潜在的流,由门控机制调节。内容分支旨在通过JEPA方案结合对比目标实现稳定的语义表示,从而促进外观变化的不变性。并行地,样式分支被约束为在对抗约束下通过特征预测和重构来捕获外观签名(纹理、对比度、散射)。我们在几个任务上评估了ST-STORM,包括对象分类(ImageNet-1 K),细粒度天气特征和黑色素瘤检测(ISIC 2024挑战)。结果表明,样式分支有效地隔离了复杂的外观现象(在Multi-Weather上F1=97%,在ISIC 2024上F1=94%,具有10%的标记数据),而不会降低内容分支的语义性能(在ImageNet-1 K上F1=80%),并提高了关键外观的保留
摘要
:One of the dominant paradigms in self-supervised learning (SSL), illustrated by MoCo or DINO, aims to produce robust representations by capturing features that are insensitive to certain image transformations such as illumination, or geometric changes. This strategy is appropriate when the objective is to recognize objects independently of their appearance. However, it becomes counterproductive as soon as appearance itself constitutes the discriminative signal. In weather analysis, for example, rain streaks, snow granularity, atmospheric scattering, as well as reflections and halos, are not noise: they carry the essential information. In critical applications such as autonomous driving, ignoring these cues is risky, since grip and visibility depend directly on ground conditions and atmospheric conditions. We introduce ST-STORM, a hybrid SSL framework that treats appearance (style) as a semantic modality to be disentangled from content. Our architecture explicitly separates two latent streams, regulated by gating mechanisms. The Content branch aims at a stable semantic representation through a JEPA scheme coupled with a contrastive objective, promoting invariance to appearance variations. In parallel, the Style branch is constrained to capture appearance signatures (textures, contrasts, scattering) through feature prediction and reconstruction under an adversarial constraint. We evaluate ST-STORM on several tasks, including object classification (ImageNet-1K), fine-grained weather characterization, and melanoma detection (ISIC 2024 Challenge). The results show that the Style branch effectively isolates complex appearance phenomena (F1=97% on Multi-Weather and F1=94% on ISIC 2024 with 10% labeled data), without degrading the semantic performance (F1=80% on ImageNet-1K) of the Content branch, and improves the preservation of critical appearance
【8】AEGIS: Anchor-Enforced Gradient Isolation for Knowledge-Preserving Vision-Language-Action Fine-Tuning
标题:AEGIS:锚点强制梯度隔离,用于保护知识的视觉-语言-动作微调
链接:https://arxiv.org/abs/2604.16067
作者:Guransh Singh
摘要:将预先训练的视觉语言模型(VLM)用于机器人控制需要将来自流匹配动作专家的高幅度连续梯度注入到专门使用交叉熵训练的骨干中。这种跨模态梯度不对称性-低秩MSE回归梯度与CE预训练所塑造的高维语义流形之间的谱维度不匹配,导致VLM的视觉问答(VQA)能力的快速,严重侵蚀。行业标准的防御要么通过停止梯度完全切断梯度路径,放弃丰富的连续监督,要么通过低秩适配器(LoRA)限制参数容量,限制更新的秩但不限制其方向,因此仍然覆盖预先训练的流形。我们介绍AEGIS(锚定强制梯度隔离系统):一个无缓冲区的逐层正交梯度投影框架,可以直接进行连续MSE学习,同时保留预训练的VQA流形-没有任何共同训练数据或重放缓冲区。AEGIS从所有Transformer层的掩蔽VQA前向传递中预先计算静态高斯参考锚点,然后在每个训练步骤中构建Wasserstein-2传输惩罚,生成锚点恢复梯度。顺序双向后分解任务和锚梯度;对于每个Transformer层,AEGIS应用单个Gram-Schmidt正交投影,将任务梯度从破坏性方向弯曲,同时保留其建设性内容。该投影平均减少不到1%的梯度能量,但消除了导致严重遗忘的累积激活漂移。
摘要:Adapting pre-trained vision-language models (VLMs) for robotic control requires injecting high-magnitude continuous gradients from a flow-matching action expert into a backbone trained exclusively with cross-entropy. This cross-modal gradient asymmetry - the spectral dimensionality mismatch between low-rank MSE regression gradients and the high-dimensional semantic manifold sculpted by CE pre-training, causes rapid, severe erosion of the VLM's visual-question-answering (VQA) capability. Industry-standard defences either sever the gradient pathway entirely via stop gradient, discarding the rich continuous supervision, or restrict parameter capacity through low-rank adapters (LoRA) that constrain the rank of updates but not their direction, and thus still overwrite the pre-trained manifold. We introduce AEGIS (Anchor-Enforced Gradient Isolation System): a buffer-free, layer-wise orthogonal gradient projection framework that enables direct continuous MSE learning while preserving the pre-trained VQA manifold - without any co-training data or replay buffer. AEGIS pre-computes a static Gaussian reference anchor from masked VQA forward passes across all transformer layers, then at each training step constructs a Wasserstein-2 transport penalty that generates an anchor restoration gradient. A sequential dual-backward decomposes the task and anchor gradients; for each transformer layer, AEGIS applies a single Gram-Schmidt orthogonal projection that bends the task gradient away from the destructive direction while preserving its constructive content. The projection sheds less than 1% of gradient energy on average, yet eliminates the cumulative activation drift that drives severe forgetting.
【9】Constant-Factor Approximations for Doubly Constrained Fair k-Center, k-Median and k-Means
标题:双重约束公平k-中心、k-中位数和k-均值的恒因子逼近
链接:https://arxiv.org/abs/2604.16061
作者:Nicole Funk,Annika Hennes,Johanna Hillebrand,Sarah Sturm
备注:30 pages, 3 figures
摘要:我们研究一般度量空间中的离散k-聚类问题,这些问题受到人口公平模型中两个不同公平条件的组合的约束。给定度量空间(P,d),其中P中的每个点都配备有受保护的属性和数字k,目标是将P划分为k个聚类,每个聚类具有指定的中心,使得基于中心的目标函数最小化并且属性相对于以下两个公平性概念公平地分布:1)组公平性:我们的目标是通过指定所需的属性比例的下限和上限的属性数量平衡的集群。2)多样的中心选择:集群有自然的代表,即,他们的中心。我们通过指定从每个属性中选择的所需中心数来要求一组平衡的代表。 Dickerson、Esmaeili、Morgenstern和Zhang(2023)将这两个约束的组合表示为双重约束公平集群。他们提出的算法,其保证取决于这些问题的最佳已知的近似因子。目前,这意味着一个8近似与一个小的添加剂违反组公平性约束。对于k-中心,我们改进这个近似因子为4,具有小的加性违反。这种保证还取决于Jones,Nguyen和Nguyen(2020)给出的DS-公平k-中心的当前最佳算法。对于k-median和k-means,我们提出了第一个常数因子近似算法。我们的算法转换成一个双重约束的公平聚类使用LP为基础的方法,满足不同的中心选择的解决方案。此外,我们的结果是推广到其他中心选择约束,如拟阵k-聚类和背包约束。
摘要:We study discrete k-clustering problems in general metric spaces that are constrained by a combination of two different fairness conditions within the demographic fairness model. Given a metric space (P,d), where every point in P is equipped with a protected attribute, and a number k, the goal is to partition P into k clusters with a designated center each, such that a center-based objective function is minimized and the attributes are fairly distributed with respect to the following two fairness concepts: 1) group fairness: We aim for clusters with balanced numbers of attributes by specifying lower and upper bounds for the desired attribute proportions. 2) diverse center selection: Clusters have natural representatives, i.e., their centers. We ask for a balanced set of representatives by specifying the desired number of centers to choose from each attribute. Dickerson, Esmaeili, Morgenstern and Zhang (2023) denote the combination of these two constraints as doubly constrained fair clustering. They present algorithms whose guarantees depend on the best known approximation factors for either of these problems. Currently, this implies an 8-approximation with a small additive violation on the group fairness constraint. For k-center, we improve this approximation factor to 4 with a small additive violation. This guarantee also depends on the currently best algorithm for DS-fair k-center given by Jones, Nguyen and Nguyen (2020). For k-median and k-means, we propose the first constant-factor approximation algorithms. Our algorithms transform a solution that satisfies diverse center selection into a doubly constrained fair clustering using an LP-based approach. Furthermore, our results are generalizable to other center-selection constraints, such as matroid k-clustering and knapsack constraints.
【10】Where does output diversity collapse in post-training?
标题:训练后产出多样性在哪里崩溃?
链接:https://arxiv.org/abs/2604.16027
作者:Constantinos Karouzos,Xingwei Tan,Nikolaos Aletras
摘要:经过训练的语言模型产生的输出比它们的基础模型少。这种输出多样性的崩溃破坏了依赖于不同样本的推理时间缩放方法,并有可能在创造性和充满价值的任务上破坏模型输出。之前的工作属性会被分解为特定的后训练方法,而不会将训练数据组成的角色与方法分开,或者将生成格式与模型权重分开。我们通过Olmo 3、Think(思想链蒸馏)、Instruct(广泛的多源数据)和RL-Zero这三个并行的训练后谱系,在15个任务和4个文本多样性指标上跟踪输出多样性。我们发现,崩溃的位置随数据组成而变化:认为血统失去了大部分的语义多样性在监督微调,和DPO的影响是在指令比认为更大。在Think模型中抑制推理时的思维链推理会降低硬任务的准确性,但不会改变答案级别的多样性,这表明崩溃是通过训练数据嵌入到模型权重中的,而不是由生成格式强加的。将六个可验证任务的多样性损失分解为质量控制组件(删除不正确的输出)和剩余组件(正确输出之间的真正缩小),揭示了分裂是依赖于任务的,并且Think模型保留了比Instruct更多的正确答案多样性,尽管总体上崩溃更多。我们的研究结果表明,多样性崩溃是在训练过程中由数据组成决定的,不能单独在推理时解决。
摘要
:Post-trained language models produce less varied outputs than their base counterparts. This output diversity collapse undermines inference-time scaling methods that rely on varied samples, and risks homogenizing model outputs on creative and value-laden tasks. Prior work attributes collapse to specific post-training methods, without separating the role of training data composition from the method, or the generation format from the model weights. We trace output diversity through three parallel post-training lineages of Olmo 3, Think (chain-of-thought distillation), Instruct (broad multi-source data), and RL-Zero, across 15 tasks and four text diversity metrics. We find that the location of collapse co-varies with data composition: the Think lineage loses most semantic diversity at supervised fine-tuning, and the effect of DPO is larger in Instruct than in Think. Suppressing chain-of-thought reasoning at inference in Think models drops accuracy on hard tasks, yet leaves answer-level diversity unchanged, showing that the collapse is embedded in the model weights by training data, not imposed by the generation format. Decomposing diversity loss on six verifiable tasks into a quality-control component (removal of incorrect outputs) and a residual component (genuine narrowing among correct outputs) reveals that the split is task-dependent, and Think models retain more correct-answer diversity than Instruct despite collapsing more in aggregate. Our results indicate that diversity collapse is determined during training by data composition and cannot be addressed at inference time alone.
【11】Corner Reflector Array Jamming Discrimination Using Multi-Dimensional Micro-Motion Features with Frequency Agile Radar
标题:频率捷变雷达利用多维微运动特征识别角反射器阵干扰
链接:https://arxiv.org/abs/2604.16008
作者:Jie Yuan,Lei Wang,Yanhao Wang,Yimin Liu
摘要:介绍了一种在捷变频雷达角反射面阵干扰中对实船目标进行鲁棒鉴别的方法。其关键思想是利用多维微运动的签名,从非刚性诱饵分离刚性船舶。从距离-速度图中,我们得到了两个新的手工制作的矢量-平均加权残差(MWR)和互补对比因子(CCF),并将它们与轻量级CNN学习的深层特征融合。然后,XGBoost分类器给出最终决定。大量的仿真结果表明,混合特征集始终优于国家的最先进的替代品,证实了所提出的方法的优越性。
摘要:This paper introduces a robust discrimination method for distinguishing real ship targets from corner-reflector-array jamming with frequency-agile radar. The key idea is to exploit the multidimensional micro-motion signatures that separate rigid ships from non-rigid decoys. From Range-Velocity maps we derive two new hand-crafted descriptors-mean weighted residual (MWR) and complementary contrast factor (CCF) and fuse them with deep features learned by a lightweight CNN. An XGBoost classifier then gives the final decision. Extensive simulations show that the hybrid feature set consistently outperforms state-of-the-art alternatives, confirming the superiority of the proposed approach.
【12】Reversible Residual Normalization Alleviates Spatio-Temporal Distribution Shift
标题:可逆剩余标准化缓解时空分布漂移
链接:https://arxiv.org/abs/2604.15838
作者:Zhaobo Hu,Vincent Gauthier,Mehdi Naima
摘要:分布漂移严重降低了深度预测模型的性能。虽然这一问题在个别时间序列中得到了很好的研究,但在时空领域仍然是一个重大挑战。像实例规范化及其变体这样的有效解决方案可以通过标准化统计数据来减轻时间偏移。然而,图上的分布偏移要复杂得多,不仅涉及单个节点系列的漂移,还涉及空间网络中的异质性,其中不同的节点表现出不同的统计特性。为了解决这个问题,我们提出了可逆残差归一化(RRN),一种新的框架,执行空间感知的可逆变换,以解决空间和时间维度的分布偏移。我们的方法将图卷积操作集成在可逆的残差块中,实现了自适应归一化,在保持可逆性的同时尊重底层图结构。通过将中心归一化与谱约束图神经网络相结合,我们的方法以数据驱动的方式捕获和规范化复杂的时空关系。我们框架的双向性质允许模型在归一化的潜在空间中学习,并通过逆变换恢复原始的分布特性,为动态时空系统的预测提供了一个强大的和模型无关的解决方案。
摘要:Distribution shift severely degrades the performance of deep forecasting models. While this issue is well-studied for individual time series, it remains a significant challenge in the spatio-temporal domain. Effective solutions like instance normalization and its variants can mitigate temporal shifts by standardizing statistics. However, distribution shift on a graph is far more complex, involving not only the drift of individual node series but also heterogeneity across the spatial network where different nodes exhibit distinct statistical properties. To tackle this problem, we propose Reversible Residual Normalization (RRN), a novel framework that performs spatially-aware invertible transformations to address distribution shift in both spatial and temporal dimensions. Our approach integrates graph convolutional operations within invertible residual blocks, enabling adaptive normalization that respects the underlying graph structure while maintaining reversibility. By combining Center Normalization with spectral-constrained graph neural networks, our method captures and normalizes complex Spatio-Temporal relationships in a data-driven manner. The bidirectional nature of our framework allows models to learn in a normalized latent space and recover original distributional properties through inverse transformation, offering a robust and model-agnostic solution for forecasting on dynamic spatio-temporal systems.
【13】Collective Kernel EFT for Pre-activation ResNets
标题:用于预激活ResNets的集体核心EFT
链接:https://arxiv.org/abs/2604.15742
作者:Hidetoshi Kawase,Toshihiro Ota
备注:20 pages
摘要:在有限宽度的深度神经网络中,经验内核$G$在各层之间随机演化。我们开发了一个集体核有效场理论(EFT)的预激活ResNets的基础上的$G$-唯一的封闭层次结构和诊断其有限的有效性窗口。利用剩余增量的精确条件高斯性,我们得到了$G$的精确随机递归。应用高斯近似系统地产生一个连续深度的常微分方程系统的平均核$K_0$,核协方差$V_4$,和$1/n$的平均校正$K_{1,\mathrm {EFT}}$,它出现作为一个单回路蝌蚪校正。从数值上讲,K_0 $在所有深度都保持准确。然而,$V_4$方程残差在有限时间累积到$O(1)$误差,主要是由$G$-唯一传输项中的近似误差驱动的。此外,$K_{1,\mathrm {EFT}}$由于源闭包的崩溃而失败,即使在初始化时也表现出系统失配。这些研究结果突出了$G$-仅状态空间减少的局限性,并建议扩展状态空间,将西格玛内核。
摘要:In finite-width deep neural networks, the empirical kernel $G$ evolves stochastically across layers. We develop a collective kernel effective field theory (EFT) for pre-activation ResNets based on a $G$-only closure hierarchy and diagnose its finite validity window. Exploiting the exact conditional Gaussianity of residual increments, we derive an exact stochastic recursion for $G$. Applying Gaussian approximations systematically yields a continuous-depth ODE system for the mean kernel $K_0$, the kernel covariance $V_4$, and the $1/n$ mean correction $K_{1,\mathrm{EFT}}$, which emerges diagrammatically as a one-loop tadpole correction. Numerically, $K_0$ remains accurate at all depths. However, the $V_4$ equation residual accumulates to an $O(1)$ error at finite time, primarily driven by approximation errors in the $G$-only transport term. Furthermore, $K_{1,\mathrm{EFT}}$ fails due to the breakdown of the source closure, which exhibits a systematic mismatch even at initialization. These findings highlight the limitations of $G$-only state-space reduction and suggest extending the state space to incorporate the sigma-kernel.
【14】Neural Continuous-Time Markov Chain: Discrete Diffusion via Decoupled Jump Timing and Direction
标题:神经连续时间马尔可夫链:通过解耦跳跃时间和方向的离散扩散
链接:https://arxiv.org/abs/2604.15694
作者:Jingyuan Li,Xiaoyi Jiang,Fukang Wen,Wei Liu,Renqian Luo,Yi Zhu,Zuoqiang Shi,Pipi Hu
摘要:基于连续时间马尔可夫链(CTMC)的离散扩散模型在语言和离散数据生成方面表现出了很强的性能,但现有方法通常将反向速率矩阵作为单个对象进行参数化-通过具体分数,清洁数据预测($x_0$-参数化),或去噪分布-而不是将参数化与固有的CTMC分解对齐为跳跃定时和跳跃方向。由于CTMC基本上是完全由这两个量决定的泊松过程,因此沿着这种结构分解更接近第一原理,自然会导致我们的公式。我们提出了\textbf{Neural CTMC},它使用两个专用的网络头,通过\textbf {退出率}(何时跳转)和\textbf {跳转分布}(在哪里跳转)分别参数化反向过程。我们证明了证据下限(ELBO)与真实过程和学习到的反向过程之间的路径空间KL发散有一个$θ$-独立常数的不同,因此训练目标完全由我们参数化的退出率和跳跃分布决定。此外,该KL分解为用于定时的泊松KL和用于方向的分类KL。我们进一步表明,听话的条件代理保持相应的边际逆过程目标的梯度和最小值在标准的正则性假设。我们的理论框架还包括掩蔽和GIDD风格的噪声时间表。从经验上讲,虽然在以前的工作中已经探索了统一的前向过程,但据我们所知,我们的模型是第一个在OpenWebText数据集上优于基于掩码的方法的纯统一方法。为了促进可重复性,我们在https://huggingface.co/Jiangxy1117/Neural-CTMC上发布了我们的预训练权重。
摘要
:Discrete diffusion models based on continuous-time Markov chains (CTMCs) have shown strong performance on language and discrete data generation, yet existing approaches typically parameterize the reverse rate matrix as a single object -- via concrete scores, clean-data predictions ($x_0$-parameterization), or denoising distributions -- rather than aligning the parameterization with the intrinsic CTMC decomposition into jump timing and jump direction. Since a CTMC is fundamentally a Poisson process fully determined by these two quantities, decomposing along this structure is closer to first principles and naturally leads to our formulation. We propose \textbf{Neural CTMC}, which separately parameterizes the reverse process through an \emph{exit rate} (when to jump) and a \emph{jump distribution} (where to jump) using two dedicated network heads. We show that the evidence lower bound (ELBO) differs from a path-space KL divergence between the true and learned reverse processes by a $θ$-independent constant, so that the training objective is fully governed by the exit rate and jump distribution we parameterize. Moreover, this KL factorizes into a Poisson KL for timing and a categorical KL for direction. We further show that the tractable conditional surrogate preserves the gradients and minimizers of the corresponding marginal reverse-process objective under standard regularity assumptions. Our theoretical framework also covers masked and GIDD-style noise schedules. Empirically, while the uniform forward process has been explored in prior work, our model, to our best of the knowledge, is the first pure-uniform method to outperform mask-based methods on the OpenWebText dataset.To facilitate reproducibility, we release our pretrained weights at https://huggingface.co/Jiangxy1117/Neural-CTMC.
【15】PINNACLE: An Open-Source Computational Framework for Classical and Quantum PINNs
标题:PINNACLE:经典和量子PINN的开源计算框架
链接:https://arxiv.org/abs/2604.15645
作者:Shimon Pisnoy,Hemanth Chandravamsi,Ziv Chen,Aaron Goldgewert,Gal Shaviner,Boris Shragner,Steven H. Frankel
摘要:我们介绍了PINNACLE,这是一个用于物理信息神经网络(PINN)的开源计算框架,它在统一的模块化工作流程中集成了现代训练策略,多GPU加速和混合量子经典架构。该框架可以系统地评估PINN在基准问题上的性能,包括一维双曲守恒律,不可压缩流和电磁波传播。它支持一系列架构和训练增强功能,包括傅立叶特征嵌入、随机权重因子分解、严格边界条件执行、自适应损耗平衡、课程训练和二阶优化策略,并可扩展到其他方法。我们提供了一个全面的基准研究,量化这些方法对收敛性,准确性和计算成本的影响,并分析分布式数据并行扩展的运行时和内存效率。此外,我们将该框架扩展到混合量子经典PINN,并推导出参数偏移微分下电路评估复杂性的正式估计。结果突出了PINN对架构和训练选择的敏感性,证实了它们相对于经典求解器的高计算成本,并确定了混合量子模型提供改进的参数效率的制度。PINNACLE提供了基准物理知情的学习方法和指导未来的发展,通过定量评估其权衡的基础。
摘要:We present PINNACLE, an open-source computational framework for physics-informed neural networks (PINNs) that integrates modern training strategies, multi-GPU acceleration, and hybrid quantum-classical architectures within a unified modular workflow. The framework enables systematic evaluation of PINN performance across benchmark problems including 1D hyperbolic conservation laws, incompressible flows, and electromagnetic wave propagation. It supports a range of architectural and training enhancements, including Fourier feature embeddings, random weight factorization, strict boundary condition enforcement, adaptive loss balancing, curriculum training, and second-order optimization strategies, with extensibility to additional methods. We provide a comprehensive benchmark study quantifying the impact of these methods on convergence, accuracy, and computational cost, and analyze distributed data parallel scaling in terms of runtime and memory efficiency. In addition, we extend the framework to hybrid quantum-classical PINNs and derive a formal estimate for circuit-evaluation complexity under parameter-shift differentiation. Results highlight the sensitivity of PINNs to architectural and training choices, confirm their high computational cost relative to classical solvers, and identify regimes where hybrid quantum models offer improved parameter efficiency. PINNACLE provides a foundation for benchmarking physics-informed learning methods and guiding future developments through quantitative assessment of their trade-offs.
【16】SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding
标题:SIMMER:跨模式食品图像--通过基于MLLM的嵌入进行食谱检索
链接:https://arxiv.org/abs/2604.15628
作者:Keisuke Gomi,Keiji Yanai
备注:20 pages, 6 figures
摘要:食物图像和食谱文本之间的跨模态检索是一项重要的任务,在营养管理,饮食记录和烹饪援助的应用。现有的方法主要依赖于具有单独的图像和文本编码器的双编码器架构,需要复杂的对齐策略和特定于任务的网络设计来弥合模态之间的语义差距。在这项工作中,我们提出了SIMMER(Single Integrated Multimodal Model for Embedding Recipes),它将基于多模态大语言模型(MLLM)的嵌入模型,特别是VLM2Vec应用于这项任务,用一个统一的编码器来代替传统的双编码器范式,该编码器可以处理食物图像和食谱文本。我们设计了针对食谱结构化性质的提示模板,包括标题、配料和烹饪说明,使MLLM能够有效地嵌入生成。我们还引入了一个组件感知的数据增强策略,该策略在完整和部分配方上训练模型,提高了对不完整输入的鲁棒性。在Recipe1M数据集上的实验表明,SIMMER在1k和10k评估设置上都达到了最先进的性能,大大优于所有现有方法。特别是,与之前的最佳方法相比,我们的最佳模型将1k图像到配方的R@1从81.8\%提高到87.5\%,将10k图像到配方的R@1从56.5\%提高到65.5\%。
摘要:Cross-modal retrieval between food images and recipe texts is an important task with applications in nutritional management, dietary logging, and cooking assistance. Existing methods predominantly rely on dual-encoder architectures with separate image and text encoders, requiring complex alignment strategies and task-specific network designs to bridge the semantic gap between modalities. In this work, we propose SIMMER (Single Integrated Multimodal Model for Embedding Recipes), which applies Multimodal Large Language Model (MLLM)-based embedding models, specifically VLM2Vec, to this task, replacing the conventional dual-encoder paradigm with a single unified encoder that processes both food images and recipe texts. We design prompt templates tailored to the structured nature of recipes, which consist of a title, ingredients, and cooking instructions, enabling effective embedding generation by the MLLM. We further introduce a component-aware data augmentation strategy that trains the model on both complete and partial recipes, improving robustness to incomplete inputs. Experiments on the Recipe1M dataset demonstrate that SIMMER achieves state-of-the-art performance across both the 1k and 10k evaluation settings, substantially outperforming all prior methods. In particular, our best model improves the 1k image-to-recipe R@1 from 81.8\% to 87.5\% and the 10k image-to-recipe R@1 from 56.5\% to 65.5\% compared to the previous best method.
【17】VoodooNet: Achieving Analytic Ground States via High-Dimensional Random Projections
标题:VoodooNet:通过多维随机投影实现分析基状态
链接:https://arxiv.org/abs/2604.15613
作者:Wladimir Silva
备注:7 pages, 1 figure, 2 tables
摘要:我们提出了VoodooNet,这是一种非迭代神经架构,它通过银河系扩张用封闭形式的解析解取代了随机梯度下降(SGD)范式。通过将输入流形投影到一个高维,高熵的“银河系”空间($d \gg 784$),我们证明了复杂的功能可以解开没有反向传播的热力学成本。利用Moore-Penrose伪逆在一个步骤中求解输出层,VoodooNet实现了\textbf{98.10\% on MNIST}和\textbf{86.63\% on Fashion-MNIST}的分类精度。值得注意的是,我们在Fashion-MNIST上的结果超过了10个时期的SGD基线(84.41%),同时将训练时间减少了几个数量级。我们观察到一个接近对数的比例关系的维数和精度,这表明性能是一个函数的“银河”的体积,而不是迭代细化。这种“魔术帽”方法为实时边缘AI提供了一个新的前沿,传统的训练阶段被绕过,有利于瞬时流形发现。
摘要
:We present VoodooNet, a non-iterative neural architecture that replaces the stochastic gradient descent (SGD) paradigm with a closed-form analytic solution via Galactic Expansion. By projecting input manifolds into a high-dimensional, high-entropy "Galactic" space ($d \gg 784$), we demonstrate that complex features can be untangled without the thermodynamic cost of backpropagation. Utilizing the Moore-Penrose pseudoinverse to solve for the output layer in a single step, VoodooNet achieves a classification accuracy of \textbf{98.10\% on MNIST} and \textbf{86.63\% on Fashion-MNIST}. Notably, our results on Fashion-MNIST surpass a 10-epoch SGD baseline (84.41\%) while reducing the training time by orders of magnitude. We observe a near-logarithmic scaling law between dimensionality and accuracy, suggesting that performance is a function of "Galactic" volume rather than iterative refinement. This "Magic Hat" approach offers a new frontier for real-time Edge AI, where the traditional training phase is bypassed in favor of instantaneous manifold discovery.
【18】Why Fine-Tuning Encourages Hallucinations and How to Fix It
标题:为什么微调会鼓励幻觉以及如何修复它
链接:https://arxiv.org/abs/2604.15574
作者:Guy Kaplan,Zorik Gekhman,Zhen Zhu,Lotem Rozner,Yuval Reif,Swabha Swayamdipta,Derek Hoiem,Roy Schwartz
摘要:大型语言模型容易产生事实上不正确的语句。这些错误的一个关键来源是通过监督微调(SFT)暴露于新的事实信息,这可能会增加幻觉。在培训前获得的知识。在这项工作中,我们探讨了SFT引起的幻觉是否可以使用持续学习文献中的既定工具来减轻,因为它们是训练过程中知识退化的副产品。我们提出了一种基于自蒸馏的SFT方法,该方法有助于有效的事实学习,同时最大限度地减少幻觉w.r.t.通过正则化输出分布漂移来预先存在的知识。我们还表明,在不需要获取新知识的情况下,通过冻结参数组来抑制事实可塑性,可以保持任务性能,同时减少幻觉。最后,我们通过三个假设:能力限制、行为克隆和局部干扰来研究SFT诱导幻觉的机制。我们的实验表明,主要驱动因素是重叠语义表示之间的干扰,而自我升华通过减轻这种干扰而成功。
摘要:Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information through supervised fine-tuning (SFT), which can increase hallucinations w.r.t. knowledge acquired during pre-training. In this work, we explore whether SFT-induced hallucinations can be mitigated using established tools from the continual learning literature, since they arise as a by-product of knowledge degradation during training. We propose a self-distillation-based SFT method that facilitates effective factual learning while minimizing hallucinations w.r.t. pre-existing knowledge by regularizing output-distribution drift. We also show that, in settings where new knowledge acquisition is unnecessary, suppressing factual plasticity by freezing parameter groups, can preserve task performance while reducing hallucinations. Lastly, we investigate the mechanism behind SFT-induced hallucinations through three hypotheses: capacity limitations, behavior cloning, and localized interference. Our experiments show that a main driver is interference among overlapping semantic representations, and that self-distillation succeeds by mitigating this interference.
【19】Collaborative Filtering Through Weighted Similarities of User and Item Embeddings
标题:通过用户和项目嵌入的加权相似性进行协作过滤
链接:https://arxiv.org/abs/2604.15573
作者:Pedro R. Pires,Rafael T. Sereicikas,Gregorio F. Azevedo,Tiago A. Almeida
备注:Published in SAC'25, 8 pages, 4 figures
摘要:近年来,神经网络和其他复杂模型已经主导了推荐系统,通常为最先进的性能设定了新的基准。然而,尽管有这些进步,获奖的研究表明,传统的矩阵分解方法可以保持竞争力,提供简单性和减少计算开销。将矩阵分解与更新技术相结合的混合模型越来越多地用于利用多种方法的优势。本文提出了一种新的集成方法,通过加权相似度框架统一用户-项目和项目-项目推荐,以提供前N个推荐。我们的方法是独特的,在其使用共享的用户和项目嵌入的推荐策略,简化了架构,提高计算效率。在多个数据集上进行的广泛实验表明,我们的方法具有竞争力的性能,并且在支持用户-项目推荐或项目-项目推荐的不同场景中具有鲁棒性。此外,通过消除对嵌入特定微调的需要,我们的模型允许无缝重用基础算法的超参数,而不会牺牲性能。这导致一种既高效又易于实现的方法。我们的开源实现可以在https://github.com/UFSCar-LaSID/weighted-sims-recommender上找到。
摘要:In recent years, neural networks and other complex models have dominated recommender systems, often setting new benchmarks for state-of-the-art performance. Yet, despite these advancements, award-winning research has demonstrated that traditional matrix factorization methods can remain competitive, offering simplicity and reduced computational overhead. Hybrid models, which combine matrix factorization with newer techniques, are increasingly employed to harness the strengths of multiple approaches. This paper proposes a novel ensemble method that unifies user-item and item-item recommendations through a weighted similarity framework to deliver top-N recommendations. Our approach is distinctive in its use of shared user and item embeddings for both recommendation strategies, simplifying the architecture and enhancing computational efficiency. Extensive experiments across multiple datasets show that our method achieves competitive performance and is robust in varying scenarios that favor either user-item or item-item recommendations. Additionally, by eliminating the need for embedding-specific fine-tuning, our model allows for the seamless reuse of hyperparameters from the base algorithm without sacrificing performance. This results in a method that is both efficient and easy to implement. Our open-source implementation is available at https://github.com/UFSCar-LaSID/weighted-sims-recommender.
【20】Natural gradient descent with momentum
标题:有动量的自然梯度下降
链接:https://arxiv.org/abs/2604.15554
作者:Anthony Nouy,Agustín Somacal
摘要:我们考虑一个函数的一个元素的非线性流形,承认一个可微的参数化,典型的例子是神经网络与可微激活函数或张量网络的近似问题。用于优化损失函数的自然梯度下降(NGD)可以被视为预条件梯度下降,其中参数空间中的更新由函数视角驱动。在类似于牛顿方法的精神中,NGD步骤使用关于适当度量的当前λ处的近似流形的切线空间的生成系统的格拉姆矩阵而不是海森矩阵。这对应于函数空间中的局部最优更新,遵循投影梯度到流形的切空间上。尽管如此,梯度和自然梯度下降方法都会陷入局部最小值。此外,当模型类是非线性流形或损失函数不是理想条件时(例如,用于密度估计的KL-发散,或者物理学中的偏微分方程的残差的范数(norm)),甚至自然梯度也可能在每一步产生非最优方向。这项工作介绍了经典惯性动力学方法的自然版本,如Heavy-Ball或Nesterov,并展示了它如何在使用非线性模型类时改善学习过程。
摘要:We consider the problem of approximating a function by an element of a nonlinear manifold which admits a differentiable parametrization, typical examples being neural networks with differentiable activation functions or tensor networks. Natural gradient descent (NGD) for the optimization of a loss function can be seen as a preconditioned gradient descent where updates in the parameter space are driven by a functional perspective. In a spirit similar to Newton's method, a NGD step uses, instead of the Hessian, the Gram matrix of the generating system of the tangent space to the approximation manifold at the current iterate, with respect to a suitable metric. This corresponds to a locally optimal update in function space, following a projected gradient onto the tangent space to the manifold. Still, both gradient and natural gradient descent methods get stuck in local minima. Furthermore, when the model class is a nonlinear manifold or the loss function is not ideally conditioned (e.g., the KL-divergence for density estimation, or a norm of the residual of a partial differential equation in physics informed learning), even the natural gradient might yield non-optimal directions at each step. This work introduces a natural version of classical inertial dynamic methods like Heavy-Ball or Nesterov and show how it can improve the learning process when working with nonlinear model classes.
【21】Optimizing Stochastic Gradient Push under Broadcast Communications
标题:广播通信下的随机梯度推送优化
链接:https://arxiv.org/abs/2604.15549
作者:Tuan Nguyen,Ting He
摘要
:我们考虑广播通信下无线网络中分散式联邦学习(DFL)的收敛时间最小化问题,重点是混合矩阵的设计。混合矩阵是DFL的一个关键超参数,它同时控制迭代的收敛速度和每次迭代的通信需求,两者都强烈影响收敛时间。虽然这个问题已经被研究过,现有的解决方案大多是设计分散并行随机梯度下降(D-PSGD),这需要混合矩阵是对称的和双随机的。这些约束将激活的通信图限制为无向的(即,双向)图,这限制了设计灵活性。相比之下,我们考虑随机梯度推(SGP)的混合矩阵设计,它允许不对称的混合矩阵,因此有向通信图。通过分析SGP的收敛速度如何依赖于混合矩阵,我们提取了一个目标函数,该目标函数显式地依赖于激活通信图的图论参数,在此基础上,我们开发了一个有效的设计算法的性能保证。我们基于真实数据的评估表明,与现有技术相比,所提出的解决方案可以显着减少收敛时间,而不会影响训练模型的质量。
摘要:We consider the problem of minimizing the convergence time for decentralized federated learning (DFL) in wireless networks under broadcast communications, with focus on mixing matrix design. The mixing matrix is a critical hyperparameter for DFL that simultaneously controls the convergence rate across iterations and the communication demand per iteration, both strongly influencing the convergence time. Although the problem has been studied previously, existing solutions are mostly designed for decentralized parallel stochastic gradient descent (D-PSGD), which requires the mixing matrix to be symmetric and doubly stochastic. These constraints confine the activated communication graph to undirected (i.e., bidirected) graphs, which limits design flexibility. In contrast, we consider mixing matrix design for stochastic gradient push (SGP), which allows asymmetric mixing matrices and hence directed communication graphs. By analyzing how the convergence rate of SGP depends on the mixing matrices, we extract an objective function that explicitly depends on graph-theoretic parameters of the activated communication graph, based on which we develop an efficient design algorithm with performance guarantees. Our evaluations based on real data show that the proposed solution can notably reduce the convergence time compared to the state of the art without compromising the quality of the trained model.
【22】Verification Modulo Tested Library Contracts
标题:验证模测试图书馆合同
链接:https://arxiv.org/abs/2604.15533
作者:Abhishek Uppar,Omar Muhammad,Sumanth Prabhu,Deepak D'Souza,Madhusudan P,Adithya Murali
摘要:本文考虑模被测库的验证问题 contracts}作为自动验证使用复杂库的客户端程序的一步。我们将这个问题描述为客户端使用的库方法的模块化契约的合成,这些方法足以证明客户端是正确的,并且还通过了测试引擎的审查,该测试引擎根据这些契约测试库。我们还考虑了一种新形式的方法契约,称为上下文契约,它在客户端程序的上下文中出现,并且通常比经典的模块化契约更简单,更容易推断。我们提供了一个反例引导的学习框架来解决这个问题,其中合成器与约束求解器以及测试引擎进行交互,以推断出足够的模块化/上下文方法合同和客户端的归纳不变量。我们使用的主要合成引擎是使用ICE学习算法实现的通用CHC求解器。我们在一个名为\vmtlc的工具中实现了这个框架,并在客户端调用大型库的基准测试中显示了它的有效性。
摘要:We consider the problem of \emph{verification modulo tested library contracts} as a step towards automating the verification of client programs that use complex libraries. We formulate this problem as the synthesis of modular contracts for the library methods used by the client that are adequate to prove the client correct, and that also pass the scrutiny of a testing engine that tests the library against these contracts. We also consider a new form of method contracts called \emph{contextual contracts} that arise in this setting that hold in the context of the client program, and can often be simpler and easier to infer than classical modular contracts. We provide a counterexample-guided learning framework to solve this problem, in which the synthesizer interacts with a constraint solver as well as the testing engine in order to infer adequate modular/contextual method contracts and inductive invariants for the client. The main synthesis engines we use are generalizing CHC solvers that are realized using ICE learning algorithms. We realize this framework in a tool called \vmtlc and show its efficacy on benchmarks where clients call large libraries.
【23】Lossless Compression via Chained Lightweight Neural Predictors with Information Inheritance
标题:通过具有信息继承的连锁轻量级神经预测器进行无损压缩
链接:https://arxiv.org/abs/2604.15472
作者:Yuriy Kim,Evgeny Belyaev
备注:Under review
摘要:本文致力于利用神经网络进行概率估计的无损数据压缩。首先,我们提出了一个概率估计架构的基础上链的神经预测,使链的每个单元被定义为一个神经网络的权重,这是足够的有效压缩的数据产生的马尔可夫源的给定顺序的最小可能数量。我们表明,这种架构使我们能够最大限度地减少参与概率估计过程的权重的总数取决于输入数据的统计特性。其次,为了提高压缩效率,我们引入了一种信息继承机制,其中由低阶单元获得的概率估计用于下一个高阶单元。实验结果表明,建议的无损数据压缩器配备了链式概率估计架构提供的压缩比接近国家的最先进的PAC压缩机。与此同时,在消费级GPU上,它的编码吞吐量是PAC的1.2到6.3倍,解码吞吐量是PAC的2.8到12.3倍。
摘要:This paper is dedicated to lossless data compression with probability estimation using neural networks. First, we propose a probability estimation architecture based on a chain of neural predictors, so that each unit of the chain is defined as a neural network with the minimum possible number of weights, which is sufficient for efficient compression of data generated by Markov sources of a given order. We show that this architecture allows us to minimize the overall number of weights participating in the probability estimation process depending on the statistical properties of the input data. Second, in order to improve compression efficiency, we introduce an information inheritance mechanism, where the probability estimate obtained by a low-order unit is used at the next higher-order unit. Experimental results show that the proposed lossless data compressor equipped with the chained probability estimation architecture provides compression ratios close to the state-of-the-art PAC compressor. At the same time, it outperforms PAC by a factor of 1.2 to 6.3 in encoding throughput and by a factor of 2.8 to 12.3 in decoding throughput on a consumer GPU.
【24】(1D) Ordered Tokens Enable Efficient Test-Time Search
标题:(1D)有序代币实现高效的测试时搜索
链接:https://arxiv.org/abs/2604.15453
作者:Zhitong Gao,Parham Rezaei,Ali Cy,Mingqiao Ye,Nataša Jovanović,Jesse Allardice,Afshin Dehghan,Amir Zamir,Roman Bachmann,Oğuzhan Fatih Kar
备注:Project page: https://soto.epfl.ch/
摘要:标记化是自回归(AR)生成模型的关键组成部分,将原始数据转换为更易于管理的建模单元。通常,令牌描述本地信息,例如图像中的像素区域或文本中的单词片段,并且AR生成以固定顺序预测这些令牌。一个值得研究的问题是,令牌结构是否会影响通过测试时搜索引导生成的能力,其中多个候选生成被验证者探索和评估。使用图像生成作为我们的测试平台,我们假设,最近的1D有序标记器与粗到细的结构可以更适合搜索比经典的2D网格结构。这植根于这样一个事实,即粗到细序列中的中间状态携带验证器可以可靠地评估的语义含义,从而在生成期间实现有效的转向。 通过对照实验,我们发现,与基于网格的模型相比,在由粗到细的有序令牌上训练的AR模型表现出更好的测试时间缩放行为。此外,我们证明,由于有序结构,纯测试时搜索令牌序列(即,不训练AR模型)可以在图像-文本验证器的引导下执行无训练的文本到图像生成。除此之外,我们系统地研究了经典的搜索算法(最佳N,波束搜索,前瞻搜索)如何与不同的令牌结构,以及不同的验证者和AR先验的作用。我们的研究结果突出了令牌结构对推理时间可扩展性的影响,并为AR模型中的测试时间扩展提供了实际指导。
摘要
:Tokenization is a key component of autoregressive (AR) generative models, converting raw data into more manageable units for modeling. Commonly, tokens describe local information, such as regions of pixels in images or word pieces in text, and AR generation predicts these tokens in a fixed order. A worthwhile question is whether token structures affect the ability to steer the generation through test-time search, where multiple candidate generations are explored and evaluated by a verifier. Using image generation as our testbed, we hypothesize that recent 1D ordered tokenizers with coarse-to-fine structure can be more amenable to search than classical 2D grid structures. This is rooted in the fact that the intermediate states in coarse-to-fine sequences carry semantic meaning that verifiers can reliably evaluate, enabling effective steering during generation. Through controlled experiments, we find that AR models trained on coarse-to-fine ordered tokens exhibit improved test-time scaling behavior compared to grid-based counterparts. Moreover, we demonstrate that, thanks to the ordered structure, pure test-time search over token sequences (i.e., without training an AR model) can perform training-free text-to-image generation when guided by an image-text verifier. Beyond this, we systematically study how classical search algorithms (best-of-N, beam search, lookahead search) interact with different token structures, as well as the role of different verifiers and AR priors. Our results highlight the impact of token structure on inference-time scalability and provide practical guidance for test-time scaling in AR models.
【25】Prompt-Driven Code Summarization: A Systematic Literature Review
标题:预算驱动的代码摘要:系统性文献综述
链接:https://arxiv.org/abs/2604.15385
作者:Afia Farjana,Zaiyu Cheng,Antonio Mastropaolo
备注:42 pages, 9 figures, 10 tables. Systematic Literature Review. This work is currently under review at ACM TOSEM
摘要:软件文档对于程序理解、开发人员入职、代码审查和长期维护至关重要。然而,手工制作高质量的文件是耗时的,而且经常产生不完整或不一致的结果。大型语言模型(LLM)通过从源代码自动生成自然语言描述,帮助开发人员更有效地理解代码,促进维护,并支持缺陷定位和提交消息生成等下游活动,提供了一个有前途的解决方案。然而,LLM在文档任务中的有效性关键取决于如何提示它们。结构正确的指令可以大大提高模型的性能,使提示工程输入提示的设计,以指导模型的行为,在基于LLM的软件工程的基础技术。诸如Few-Shot提示、思维链推理、检索增强生成和zero-shot学习等方法显示出代码摘要的前景,但目前的研究仍然是零散的。对于哪种激励策略效果最好,适用于哪种模式,以及在什么条件下,人们的理解有限。此外,评估实践差异很大,大多数研究依赖于基于语义的指标,可能无法捕捉语义质量。这个系统的文献综述巩固了现有的证据,分类提示范式,检查其有效性,并确定差距,以指导未来的研究和实际采用。
摘要:Software documentation is essential for program comprehension, developer onboarding, code review, and long-term maintenance. Yet producing quality documentation manually is time-consuming and frequently yields incomplete or inconsistent results. Large language models (LLMs) offer a promising solution by automatically generating natural language descriptions from source code, helping developers understand code more efficiently, facilitating maintenance, and supporting downstream activities such as defect localization and commit message generation. However, the effectiveness of LLMs in documentation tasks critically depends on how they are prompted. Properly structured instructions can substantially improve model performance, making prompt engineering-the design of input prompts to guide model behavior-a foundational technique in LLM-based software engineering. Approaches such as few-shot prompting, chain-of-thought reasoning, retrieval-augmented generation, and zero-shot learning show promise for code summarization, yet current research remains fragmented. There is limited understanding of which prompting strategies work best, for which models, and under what conditions. Moreover, evaluation practices vary widely, with most studies relying on overlap-based metrics that may not capture semantic quality. This systematic literature review consolidates existing evidence, categorizes prompting paradigms, examines their effectiveness, and identifies gaps to guide future research and practical adoption.
【26】AutoFlows++: Hierarchical Message Flow Mining for System on Chip Designs
标题:AutoFlows++:片上系统设计的分层消息流挖掘
链接:https://arxiv.org/abs/2604.15359
作者:Bardia Nadimi,Hao Zheng
摘要:了解现代片上系统(SoC)设计中的通信行为对于功能验证、性能分析和硅后调试至关重要。通信跟踪捕获系统组件之间的消息交换,并提供对系统行为的有价值的见解。然而,由于通信流的交错实例和消息之间的模糊因果关系,从这种跟踪中导出简明的通信规范仍然具有挑战性。当跟踪包含跨多个组件的消息模式的复杂交织时,现有的挖掘方法通常会遇到可伸缩性和模糊性问题。这些条件往往会导致候选流数量的爆炸和通信行为的不准确提取。本文介绍了AutoFlows++,一个设计架构指导的分层框架,用于从复杂SoC设计的通信痕迹中挖掘消息流。AutoFlows++分为两个阶段:本地挖矿,然后是全局挖矿。在本地挖掘阶段,简单的通信模式被提取从组件之间的各个通信接口处观察到的痕迹。在全局挖掘阶段,这些局部模式被组合以标识表征跨多个组件的通信行为的更高级别的消息流。对GEM5中的合成轨迹和SoC模型生成的轨迹的实验结果表明,与现有方法相比,AutoFlows++显著提高了流量提取的准确性,突出了其在实际SoC验证任务中的有效性。
摘要:Understanding communication behavior in modern system-on-chip (SoC) designs is critical for functional verification, performance analysis, and post-silicon debugging. Communication traces capture message exchanges among system components and provide valuable insights into system behavior. However, deriving concise communication specifications from such traces remains challenging due to interleaved instances of communication flows, and ambiguous causal relationships among messages. Existing mining approaches often struggle with scalability and ambiguity when traces contain complex interleaving of message patterns across multiple components. These conditions often lead to an explosion in the number of candidate flows and inaccurate extraction of communication behaviors. This paper presents AutoFlows++, a design-architecture-guided hierarchical framework for mining message flows from communication traces of complex SoC designs. AutoFlows++ operates in two stages: local mining followed by global mining. In the local mining stage, simple communication patterns are extracted from traces observed at individual communication interfaces between components. In the global mining stage, these local patterns are composed to identify higher-level message flows that characterize communication behavior across multiple components. Experimental results on both synthetic traces and traces generated from SoC models in GEM5 demonstrate that AutoFlows++ significantly improves flow extraction accuracy compared with prior approaches, highlighting its effectiveness for practical SoC validation tasks.
【27】Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit
标题:通过概率语言尝试的顺序KV缓存压缩:超越每载体香农限制
链接:https://arxiv.org/abs/2604.15356
作者:Gregory Magarshak
备注:22 Pages
摘要:最近的KV缓存量化工作,最终在TurboQuant,已接近每矢量压缩的Transformer键值缓存的香农熵限制。我们观察到,这个限制适用于一个严格较弱的问题,而不是一个真正重要的问题:压缩KV缓存作为一个序列。存储在KV缓存中的令牌不是任意的浮点数据-它们是来自模型训练的确切形式语言的样本,并且该模型是该语言的近似最佳预测器。我们引入顺序KV压缩,一个两层架构,利用这种结构。第一层,概率前缀去重,使用来自概率前缀语言Tries(PLT)的trie度量d_T(s,s ')= -log_2 P_M(s,s')来标识跨会话的语义上等效的共享前缀。第二层,预测增量编码,仅存储来自模型自己的预测的每个新KV向量的残差,实现H(KV_{i+1})的每令牌熵界|KV_{<=i})<= H(令牌_{i+1}| token_{<=i})。我们证明,在典型的语言模型的困惑-大约10-20流利的英语文本-这个界限是3.3-4.3位平均每个令牌的位置,相比TurboQuant的3位每个矢量分量(典型的注意头有64-128个组件)。TurboQuant的理论压缩比在香农极限下约为914,000倍。即使在熵地板以上1000倍-故意悲观的最坏情况开销,比实际源代码编码器的2- 5倍典型值高出两个数量级-该比率仍然比TurboQuant高出约914倍,随着上下文长度的增加,压缩得到改善而不是降低。这两层是正交的,并与现有的每矢量量化方法(包括TurboQuant)组合。
摘要
:Recent work on KV cache quantization, culminating in TurboQuant, has approached the Shannon entropy limit for per-vector compression of transformer key-value caches. We observe that this limit applies to a strictly weaker problem than the one that actually matters: compressing the KV cache as a sequence. The tokens stored in a KV cache are not arbitrary floating-point data -- they are samples from the exact formal language the model was trained on, and the model is by construction a near-optimal predictor of that language. We introduce sequential KV compression, a two-layer architecture that exploits this structure. The first layer, probabilistic prefix deduplication, identifies semantically equivalent shared prefixes across sessions using the trie metric d_T(s, s') = -log_2 P_M(s ^ s') from Probabilistic Language Tries (PLTs). The second layer, predictive delta coding, stores only the residual of each new KV vector from the model's own prediction of it, achieving a per-token entropy bound of H(KV_{i+1} | KV_{<=i}) <= H(token_{i+1} | token_{<=i}). We prove that at typical language model perplexity -- approximately 10-20 for fluent English text -- this bound is 3.3-4.3 bits on average per token position, compared to TurboQuant's 3 bits per vector component (with typical attention heads having 64-128 components). The theoretical compression ratio over TurboQuant is approximately 914,000x at the Shannon limit. Even at 1000x above the entropy floor -- a deliberately pessimistic worst-case overhead, two orders of magnitude above the 2-5x typical of practical source coders -- the ratio remains approximately 914x over TurboQuant, with compression improving rather than degrading as context length grows. The two layers are orthogonal and compose with existing per-vector quantization methods including TurboQuant.
【28】Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures
标题:Aletheia:用户引导的层选择,以实现跨架构的高效LoRA微调
链接:https://arxiv.org/abs/2604.15351
作者:Abdulmalek Saket
备注:11 pages, 5 figures, 2 frozen evidence campaigns, 81 experiment rows across 14 successful models and 8 architecture families, plus one documented failed Pythia/GPT-NeoX attempt
摘要:低秩自适应(LoRA)已成为大型语言模型的主要参数高效微调方法,但标准实践将LoRA适配器统一应用于所有Transformer层,而不管它们与下游任务的相关性如何。我们介绍了Aletheia,一种梯度引导的层选择方法,通过轻量级梯度探测器识别最相关的任务层,并将LoRA适配器仅应用于具有不对称秩分配的那些层。横跨81个实验行,涵盖8个体系结构家族的14个成功模型(0.5B-72 B参数,包括密集和专家混合架构),在Campaign 2中还有一次记录失败的Pythia/GPT-NeoX尝试,Aletheia实现了15-28%的训练加速(平均23.1%,p < 0.001),在评估的MMLU、GSM 8 K和HumanEval基准包上具有有限的额外遗忘和广泛匹配的下游行为。在测试的系列和规模,活动1显示了100%的每模型的速度获胜率和活动2显示了广泛保留的下游行为在一个有限的退化框架。这些结果共同支持了一个实用的模型经济学主张:智能层选择可以使LoRA微调更有效,而不会对评估集造成重大下游损害。
摘要:Low-Rank Adaptation (LoRA) has become the dominant parameter-efficient fine-tuning method for large language models, yet standard practice applies LoRA adapters uniformly to all transformer layers regardless of their relevance to the downstream task. We introduce Aletheia, a gradient-guided layer selection method that identifies the most task-relevant layers via a lightweight gradient probe and applies LoRA adapters only to those layers with asymmetric rank allocation. Across 81 experiment rows covering 14 successful models from 8 architecture families (0.5B-72B parameters, including dense and Mixture-of-Experts architectures), with one additional documented failed Pythia/GPT-NeoX attempt in Campaign 2, Aletheia achieves a 15-28% training speedup (mean 23.1%, p < 0.001) with bounded extra forgetting and broadly matched downstream behavior on the evaluated MMLU, GSM8K, and HumanEval benchmark pack. Across the tested families and scales, Campaign 1 shows a 100% per-model speed win rate and Campaign 2 shows broadly preserved downstream behavior within a bounded-degradation framing. Together these results support a practical model-economics claim: intelligent layer selection can make LoRA fine-tuning materially more efficient without introducing major downstream damage on the evaluated set.
【29】Acoustic and Facial Markers of Perceived Conversational Success in Spontaneous Speech
标题:自发言语中会话成功感的声学和面部标记
链接:https://arxiv.org/abs/2604.15322
作者:Thanushi Withanage,Elizabeth Redcay,Carol Espy-Wilson
备注:Accepted for presentation at ICASSP 2026
摘要:个人经常将他们的说话模式与他们的对话者联系起来,这一现象与参与和融洽有关。虽然在任务导向的对话中有很好的记录,但对自然主义,非任务和虚拟环境中的夹带知之甚少。在这项研究中,我们分析了一个大型语料库的自发二元变焦会话,探讨如何会话动态感知互动质量。我们提取多模态功能,包括话轮转换,停顿,面部运动,声学措施,如音高和强度。通过对会话后评分的因素分析,量化了感知会话成功。结果表明,夹带可靠地检测到自发的讲话,并与更高的感知成功。这些研究结果确定了会话质量的关键干扰标记,并强调了有针对性的干预措施,以促进更有效和参与沟通的机会。
摘要:Individuals often align their speaking patterns with their interlocutors, a phenomenon linked to engagement and rapport. While well documented in task-oriented dialogues, less is known about entrainment in naturalistic, non-task and virtual settings. In this study, we analyze a large corpus of spontaneous dyadic Zoom conversations to examine how conversational dynamics relate to perceived interaction quality. We extract multimodal features encompassing turn-taking, pauses, facial movements, and acoustic measures such as pitch and intensity. Perceived conversational success was quantified via factor analysis of post-conversation ratings. Results demonstrate that entrainment reliably detected in spontaneous speech and correlates with higher perceived success. These findings identify key interactional markers of conversational quality and highlight opportunities for targeted interventions to foster more effective and engaging communication.
【30】A Wasserstein Geometric Framework for Hebbian Plasticity
标题:赫布可塑性的沃瑟斯坦几何框架
链接:https://arxiv.org/abs/2604.16052
作者:Ulrich Tan
备注:Preprint. 75 pages including appendices and bibliography
摘要:我们介绍了坦-HWG框架(赫布-沃瑟斯坦-几何),赫布可塑性的几何理论,其中记忆状态建模为通过沃瑟斯坦最小化运动演变的概率措施。Hebbian学习规则被形式化为满足序列稳定性条件的Hebbian能量,确保了适定的纤维JKO更新,最优传输实现和能量下降不等式。 这种变分结构导致内部动态和可观测动态之间的基本分离。内部记忆状态在潜在的弯曲空间中沿着沃瑟斯坦测地线演化,而可观测的量,如有效突触权重,则通过几何投影映射到外部空间。单纯形预测恢复经典的仿射方案(包括指数移动平均和镜像下降),同时揭示突触竞争和修剪作为质量再分配的几何后果。希尔伯特投影提供了相位对准和多尺度相干性的几何解释。 经典神经网络表现为这种弯曲动力学的平面投影,而框架自然适应更丰富的分布表示,包括结构权重和嵌入记忆,以及它们在复杂内部空间中的谱扩展。 在温和的Lipschitz正则性假设下,包括一个准静态的“睡眠模式”制度,我们建立了连续时间极限曲线的存在。这产生了一个变分制定的内存整合作为扰动Wasserstein梯度流。因此,该框架提供了一个统一的几何基础,突触可塑性,表示动力学和上下文相关的计算。
摘要:We introduce the Tan-HWG framework (Hebbian-Wasserstein-Geometry), a geometric theory of Hebbian plasticity in which memory states are modeled as probability measures evolving through Wasserstein minimizing movements. Hebbian learning rules are formalized as Hebbian energies satisfying a sequential stability condition, ensuring well-posed fiberwise JKO updates, optimal-transport realizations, and an energy descent inequality. This variational structure induces a fundamental separation between internal and observable dynamics. Internal memory states evolve along Wasserstein geodesics in a latent curved space, while observable quantities, such as effective synaptic weights, arise through geometric projection maps into external spaces. Simplicial projections recover classical affine schemes (including exponential moving averages and mirror descent), while revealing synaptic competition and pruning as geometric consequences of mass redistribution. Hilbertian projections provide a geometric account of phase alignment and multi-scale coherence. Classical neural networks appear as flat projections of this curved dynamics, while the framework naturally accommodates richer distributional representations, including structural weights and embedding memories, and their spectral extensions in complex internal spaces. Under mild Lipschitz regularity assumptions, including a quasi-stationary "sleep-mode" regime, we establish the existence of continuous-time limit curves. This yields a variational formulation of memory consolidation as a perturbed Wasserstein gradient flow. The framework thus provides a unified geometric foundation for synaptic plasticity, representation dynamics, and context-dependent computation.
机器翻译由腾讯交互翻译提供,仅供参考
点击“阅读原文”获取带摘要的学术速递