Py学习  »  机器学习算法

机器学习学术速递[4.27]

arXiv每日学术速递 • 2 月前 • 376 次点击  

点击阅读原文访问arxivdaily.com,涵盖CS|物理|数学|经济|统计|金融|生物|电气领域,更有搜索、收藏等功能!


cs.LG 方向,今日共计110篇


大模型相关(14篇)

【1】Aligning Dense Retrievers with LLM Utility via DistillationAligning Dense Retrievers with LLM Utility via Distillation
标题:通过蒸馏将密集检索器与LLM实用程序对齐通过蒸馏将密集检索器与LLM实用程序对齐
链接:https://arxiv.org/abs/2604.22722

作者:Rajinder Sandhu,Di Mu,Cheng Chang,Md Shahriar Tasjid,Himanshu Rai,Maksims Volkovs,Ga Wu
摘要:密集向量检索是检索增强生成(RAG)的实际支柱,但相似性搜索会受到精度的限制。相反,利用LLM重新排序的基于效用的方法通常实现优异的性能,但在计算上是禁止的,并且易于在困惑度估计中固有的噪声。我们提出了实用程序对齐嵌入(阿联酋),一个框架,旨在将这些优点合并到一个实用的,高性能的检索方法。我们制定检索作为一个分布匹配问题,训练一个双编码器来模仿一个效用分布来自困惑减少使用效用调制InfoNCE目标。这种方法将分级效用信号直接注入嵌入空间,而不需要测试时LLM推理。在QASPER基准测试中,UAE将检索Recall@1提高了30.59%,MAP提高了30.16%,Token F1提高了17.3%。至关重要的是,UAE比高效的LLM重新排名方法快180倍以上,从而保持了竞争力,这表明将检索与生成效用相结合可以产生可靠的大规模上下文。
摘要:Dense vector retrieval is the practical backbone of Retrieval- Augmented Generation (RAG), but similarity search can suffer from precision limitations. Conversely, utility-based approaches leveraging LLM re-ranking often achieve superior performance but are computationally prohibitive and prone to noise inherent in perplexity estimation. We propose Utility-Aligned Embeddings (UAE), a framework designed to merge these advantages into a practical, high-performance retrieval method. We formulate retrieval as a distribution matching problem, training a bi-encoder to imitate a utility distribution derived from perplexity reduction using a Utility-Modulated InfoNCE objective. This approach injects graded utility signals directly into the embedding space without requiring test-time LLM inference. On the QASPER benchmark, UAE improves retrieval Recall@1 by 30.59%, MAP by 30.16% and Token F1 by 17.3% over the strong semantic baseline BGE-Base. Crucially, UAE is over 180x faster than the efficient LLM re-ranking methods preserving competitive performance, demonstrating that aligning retrieval with generative utility yields reliable contexts at scale.

【2】FeatEHR-LLM: Leveraging Large Language Models for Feature Engineering in Electronic Health Records
标题:CLAREHR-LLM:利用大型语言模型进行电子健康记录的特征工程
链接:https://arxiv.org/abs/2604.22534

作者:Hojjat Karami,David Atienza,Jean-Philippe Thiran,Anisoara Ionescu
摘要:特征工程的电子健康记录(EHR)是复杂的不规则的观察间隔,可变的测量频率,和结构稀疏性固有的临床时间序列。现有的自动化方法要么缺乏临床领域意识,要么假设干净,定期采样的输入,限制了它们对真实世界EHR数据的适用性。我们提出了一个框架,利用大语言模型(LLM)从不规则采样的EHR时间序列中生成有临床意义的表格特征。为了限制患者隐私暴露,LLM仅对数据集模式和任务描述而不是原始患者记录进行操作。工具增强的生成机制为LLM配备了用于查询不规则时态数据的专用例程,使其能够生成可执行的特征提取代码,显式处理不均匀的观察模式和信息稀疏性。ESEHR-LLM通过迭代、验证在环管道支持单变量和多变量特征生成。在四个ICU数据集的八个临床预测任务上进行评估,我们的框架在8个任务中的7个任务上实现了最高的平均AUROC,比强基线提高了6个百分点。代码可在github.com/hojjatkarami/FeatEHR-LLM上获得。
摘要:Feature engineering for Electronic Health Records (EHR) is complicated by irregular observation intervals, variable measurement frequencies, and structural sparsity inherent to clinical time series. Existing automated methods either lack clinical domain awareness or assume clean, regularly sampled inputs, limiting their applicability to real-world EHR data. We present \textbf{FeatEHR-LLM}, a framework that leverages Large Language Models (LLMs) to generate clinically meaningful tabular features from irregularly sampled EHR time series. To limit patient privacy exposure, the LLM operates exclusively on dataset schemas and task descriptions rather than raw patient records. A tool-augmented generation mechanism equips the LLM with specialized routines for querying irregular temporal data, enabling it to produce executable feature-extraction code that explicitly handles uneven observation patterns and informative sparsity. FeatEHR-LLM supports both univariate and multivariate feature generation through an iterative, validation-in-the-loop pipeline. Evaluated on eight clinical prediction tasks across four ICU datasets, our framework achieves the highest mean AUROC on 7 out of 8 tasks, with improvements of up to 6 percentage points over strong baselines. Code is available at github.com/hojjatkarami/FeatEHR-LLM.

【3】Introducing Background Temperature to Characterise Hidden Randomness in Large Language Models
标题:引入背景温度来描述大型语言模型中隐藏的随机性
链接:https://arxiv.org/abs/2604.22411

作者:Alberto Messina,Stefano Scotta
摘要:即使在温度T=0的情况下进行解码,大型语言模型(LLM)也会对相同的输入产生不同的输出。Thinking Machines Lab最近的工作强调了实现级别的不确定性来源,包括批大小变化,内核非不变性和浮点非关联性。在这个简短的说明中,我们通过引入背景温度T_{bg}的概念来形式化这种行为,即使在标称T=0时,也可以观察到由实现相关的扰动过程引起的有效温度。我们提供干净的定义,显示如何$T_{\mathrm{bg}}$涉及到一个随机扰动的推理环境$I$,并提出了一个经验协议来估计$T_{bg}$通过一个理想的参考系统的等效温度$T_n(I)$。最后,我们在主要LLM提供商的代表性池上进行了一系列试点实验,这些实验展示了可重复性,评估和部署的想法和概述。
摘要:Even when decoding with temperature $T=0$, large language models (LLMs) can produce divergent outputs for identical inputs. Recent work by Thinking Machines Lab highlights implementation-level sources of nondeterminism, including batch-size variation, kernel non-invariance, and floating-point non-associativity. In this short note we formalize this behavior by introducing the notion of \emph{background temperature} $T_{\mathrm{bg}}$, the effective temperature induced by an implementation-dependent perturbation process observed even when nominal $T=0$. We provide clean definitions, show how $T_{\mathrm{bg}}$ relates to a stochastic perturbation governed by the inference environment $I$, and propose an empirical protocol to estimate $T_{bg}$ via the equivalent temperature $T_n(I)$ of an ideal reference system. We conclude with a set of pilot experiments run on a representative pool from the major LLM providers that demonstrate the idea and outline implications for reproducibility, evaluation, and deployment.

【4】How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
标题:LLM如何检测和纠正自己的错误:内部信心信号的作用
链接:https://arxiv.org/abs/2604.22271

作者:Dharshan Kumaran,Viorica Patraucean,Simon Osindero,Petar Velickovic,Nathaniel Daw
摘要:大型语言模型可以检测自己的错误,有时甚至可以在没有外部反馈的情况下纠正错误,但其潜在机制仍然未知。我们通过决策神经科学的二阶信心模型来研究这一点。在一阶系统中,置信度来源于生成信号本身,因此对于所选的响应是最大的,排除了错误检测。二阶模型描述了一个部分独立的评估信号,它可能与承诺的响应不一致,为错误检测提供了基础。Kumaran et al.(2026)表明,LLM将置信度表示缓存在紧随答案之后的标记处(即,答案后的新行:PANL)-这因果地驱动了口头置信度并与对数概率分离。在这里,我们测试该PANL信号是否超出置信度以支持错误检测和自校正。在这里,我们测试这个信号是否支持错误检测和自校正,从二阶框架中导出预测。使用验证然后正确的范例,我们表明:(i)言语信心预测错误检测远远超出令牌对数概率,排除一阶帐户;(ii)PANL激活预测错误检测超出言语信心本身;(iii)PANL预测模型可以纠正哪些错误-所有行为信号都失败。因果干预证实,PANL信号救援错误检测行为时,答案信息被损坏。所有结果在模型(Gemma 3 27 B和Qwen 2.5 7B)和任务(TriviaQA和MNLI)中重复。这些结果表明,LLM自然地实现了二阶置信度架构,其内部评估信号不仅编码答案是否可能是错误的,而且编码模型是否有知识来修复它。
摘要 :Large language models can detect their own errors and sometimes correct them without external feedback, but the underlying mechanisms remain unknown. We investigate this through the lens of second-order models of confidence from decision neuroscience. In a first-order system, confidence derives from the generation signal itself and is therefore maximal for the chosen response, precluding error detection. Second-order models posit a partially independent evaluative signal that can disagree with the committed response, providing the basis for error detection. Kumaran et al. (2026) showed that LLMs cache a confidence representation at a token immediately following the answer (i.e. post-answer newline: PANL) -- that causally drives verbal confidence and dissociates from log-probabilities. Here we test whether this PANL signal extends beyond confidence to support error detection and self-correction. Here we test whether this signal supports error detection and self-correction, deriving predictions from the second-order framework. Using a verify-then-correct paradigm, we show that: (i) verbal confidence predicts error detection far beyond token log-probabilities, ruling out a first-order account; (ii) PANL activations predict error detection beyond verbal confidence itself; and (iii) PANL predicts which errors the model can correct -- where all behavioural signals fail. Causal interventions confirm that PANL signals rescue error detection behavior when answer information is corrupted. All findings replicate across models (Gemma 3 27B and Qwen 2.5 7B) and tasks (TriviaQA and MNLI). These results reveal that LLMs naturally implement a second-order confidence architecture whose internal evaluative signal encodes not only whether an answer is likely wrong but whether the model has the knowledge to fix it.

【5】Estimating Tail Risks in Language Model Output Distributions
标题:语言模型输出分布的尾部风险估计
链接:https://arxiv.org/abs/2604.22167

作者:Rico Angell,Raghav Singhal,Zachary Horvitz,Zhou Yu,Rajesh Ranganath,Kathleen McKeown,He He
摘要:语言模型的能力越来越强,并且正在人口规模上迅速部署。因此,这些模型的安全性越来越高。幸运的是,对齐方面的进步大大降低了有害模型输出的可能性。然而,当模型在一天内被查询数十亿次时,甚至会发生罕见的最坏情况行为。目前的安全评价侧重于捕捉产生有害输出的投入的分布情况。这些评估忽略了模型的概率性质及其尾部输出行为。为了衡量这种尾部风险,我们提出了一种方法来有效地估计任何输入查询的有害输出的概率。我们不是从目标模型中进行简单的蛮力采样,因为有害的输出可能很少,而是通过创建目标模型的不安全版本来实现重要性采样。这些不安全的版本通过使有害输出更有可能实现样本有效估计。在衡量误用和错位的基准测试中,这些估计值与使用10- 20倍样本的蛮力蒙特卡罗估计值相匹配。例如,我们可以用500个样本估计出10^-4量级的有害输出概率。此外,我们发现这些危害性估计可以揭示模型对模型输入扰动的敏感性,并预测部署风险。我们的工作表明,准确的稀有事件估计是安全评估的关键和可行的。代码可在https://github.com/rangell/LMTailRisk上获取
摘要:Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunately, advances in alignment have significantly reduced the likelihood of harmful model outputs. However, when models are queried billions of times in a day, even rare worst-case behaviors will occur. Current safety evaluations focus on capturing the distribution of inputs that yield harmful outputs. These evaluations disregard the probabilistic nature of models and their tail output behavior. To measure this tail risk, we propose a method to efficiently estimate the probability of harmful outputs for any input query. Instead of naive brute-force sampling from the target model, where harmful outputs could be rare, we operationalize importance sampling by creating unsafe versions of the target model. These unsafe versions enable sample-efficient estimation by making harmful outputs more probable. On benchmarks measuring misuse and misalignment, these estimates match brute-force Monte Carlo estimates using 10-20x fewer samples. For example, we can estimate probability of harmful outputs on the order of 10^-4 with just 500 samples. Additionally, we find that these harmfulness estimates can reveal the sensitivity of models to perturbations in model input and predict deployment risks. Our work demonstrates that accurate rare-event estimation is both critical and feasible for safety evaluations. Code is available at https://github.com/rangell/LMTailRisk

【6】Sum-of-Checks: Structured Reasoning for Surgical Safety with Large Vision-Language Models
标题:检查总和:使用大型视觉语言模型的手术安全结构化推理
链接:https://arxiv.org/abs/2604.22156

作者:Weiqiu You,Cassandra Goldberg,Amin Madani,Daniel A. Hashimoto,Eric Wong
备注:IPCAI 2026 short communication
摘要:目的:腹腔镜胆囊切除术中准确评估安全性临界点(CVS)对于预防胆管损伤至关重要,胆管损伤是一种与显著发病率和死亡率相关的并发症。虽然大型视觉语言模型(LVLM)提供了灵活的推理,但它们的预测仍然难以审计,并且在安全关键的手术任务中不可靠。 研究方法:我们介绍了总和检查,一个框架,每个CVS标准分解成专家定义的推理检查,反映临床相关的视觉证据。给定一个腹腔镜框架,LVLM评估每个检查,产生一个二元判断和理由。标准级分数是通过检查结果的固定加权聚合计算的。我们使用三个前沿LVLM在Endoscapes 2023基准上进行评估,与直接提示、思维链和子问题分解进行比较,每个都有和没有Few-Shot示例。 结果如下:相对于所有三个模型和标准的最佳基线,检查总和将平均帧级平均精度提高了12- 14%。个别检查的分析表明,LVLM在观察检查中是可靠的(例如,可见性、工具阻塞),但在决策关键的解剖学证据上显示出实质性的变化。 结论:将手术推理结构化为专家对齐的验证检查,提高了基于LVLM的CVS评估的准确性和透明度,表明明确分离证据获取与决策对于可靠和可审计的手术AI系统至关重要。 代码可在https://github.com/BrachioLab/SumOfChecks上获得。
摘要:Purpose: Accurate assessment of the Critical View of Safety (CVS) during laparoscopic cholecystectomy is essential to prevent bile duct injury, a complication associated with significant morbidity and mortality. While large vision-language models (LVLMs) offer flexible reasoning, their predictions remain difficult to audit and unreliable on safety-critical surgical tasks. Methods: We introduce Sum-of-Checks, a framework that decomposes each CVS criterion into expert-defined reasoning checks reflecting clinically relevant visual evidence. Given a laparoscopic frame, an LVLM evaluates each check, producing a binary judgment and justification. Criterion-level scores are computed via fixed, weighted aggregation of check outcomes. We evaluate on the Endoscapes2023 benchmark using three frontier LVLMs, comparing against direct prompting, chain-of-thought, and sub-question decomposition, each with and without few-shot examples. Results: Sum-of-Checks improves average frame-level mean average precision by 12--14% relative to the best baseline across all three models and criteria. Analysis of individual checks reveals that LVLMs are reliable on observational checks (e.g., visibility, tool obstruction) but show substantial variability on decision-critical anatomical evidence. Conclusion: Structuring surgical reasoning into expert-aligned verification checks improves both accuracy and transparency of LVLM-based CVS assessment, demonstrating that explicitly separating evidence elicitation from decision-making is critical for reliable and auditable surgical AI systems. Code is available at https://github.com/BrachioLab/SumOfChecks.

【7】Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems
标题:通过自适应多代理LLM系统进行可靠的自残风险筛查
链接:https://arxiv.org/abs/2604.22154

作者:Meghana Karnam,Ananya Joshi
摘要:行为健康和精神病学领域的新兴人工智能系统使用多步骤或多代理LLM管道来执行评估自我伤害风险和筛查抑郁症等任务。然而,常见的评估方法,如LLM-as-a-judge,并不表明决策何时是可靠的,或者错误如何在多个LLM判断中累积,限制了它们对安全关键设置的适用性。我们提出了一个统计框架,多代理管道结构为有向无环图(DAG),提供了一种替代启发式投票与原则,自适应决策。我们将每个代理建模为随机分类决策,并引入(1)更严格的代理级性能置信区间,(2)基于输入难度的基于Bandit的自适应采样策略,以及(3)在部署时显示对数误差增长的多代理系统上的遗憾保证。我们在行为健康的两个标记数据集上评估了我们的系统:AEGIS 2.0行为健康子集(N=161)和SWMH Reddit帖子的分层样本(N=250)。从经验上讲,我们的自适应采样策略在两个数据集上实现了任何条件的最低假阳性率,AEGIS 2.0上为0.095,而单代理模型为0.159,将安全内容的错误标记减少了40%,并且在所有条件下仍然具有相似的假阴性率。这些结果表明,有原则的自适应采样提供了一个有意义的提高精度,而不会减少召回在这种情况下。
摘要:Emerging AI systems in behavioral health and psychiatry use multi-step or multi-agent LLM pipelines for tasks like assessing self-harm risk and screening for depression. However, common evaluation approaches, like LLM-as-a-judge, do not indicate when a decision is reliable or how errors may accumulate across multiple LLM judgements, limiting their suitability for safety-critical settings. We present a statistical framework for multi-agent pipelines structured as directed acyclic graphs (DAGs) that provides an alternative to heuristic voting with principled, adaptive decision-making. We model each agent as a stochastic categorical decision and introduce (1) tighter agent-level performance confidence bounds, (2) a bandit-based adaptive sampling strategy based on input difficulty, and (3) regret guarantees over the multi-agent system that shows logarithmic error growth when deployed. We evaluate our system on two labeled datasets in behavioral health : the AEGIS 2.0 behavioral health subset (N=161) and a stratified sample of SWMH Reddit posts (N=250). Empirically, our adaptive sampling strategy achieves the lowest false positive rate of any condition across both datasets, 0.095 on AEGIS 2.0 compared to 0.159 for single-agent models, reducing incorrect flagging of safe content by 40\% and still having similar false negative rates across all conditions. These results suggest that principled adaptive sampling offers a meaningful improvement in precision without reducing recall in this setting.

【8】Where Should LoRA Go? Component-Type Placement in Hybrid Language Models
标题:LoRA应该去哪里?混合语言模型中的代理类型放置
链接:https://arxiv.org/abs/2604.22127

作者:Hector Borobia,Elies Seguí-Mas,Guillermina Tormo-Carbó
备注:21 pages, 5 figures, 7 tables. Code and data: https://github.com/hecboar/lora-placement-hybrid
摘要:将注意力与循环组件交织在一起的混合语言模型与纯Transformers相比越来越具有竞争力,但标准LoRA实践统一应用适配器,而不考虑每个组件类型的不同功能角色。我们系统地研究了两种混合架构的组件类型LoRA布局- Qwen3.5-0.8B(顺序,GatedDeltaNet + softmax注意力)和Falcon-H1-0.5B(并行,Mamba-2 SSM +注意力)-在三个领域进行了微调,并在五个基准上进行了评估。我们发现,注意力路径-尽管是少数组成部分-始终优于全模型适应,可训练参数少5- 10倍。至关重要的是,适应经常性的骨干是破坏性的顺序杂交(-14.8 pp的GSM 8 K),但建设性的平行(+8.6 pp)。我们进一步证明了转移不对称性:并行混合动力车表现出积极的跨任务转移,而顺序混合动力车遭受灾难性的遗忘。这些结果表明,混合拓扑结构从根本上决定了自适应响应,并且组件感知LoRA放置是混合架构的必要设计维度。
摘要 :Hybrid language models that interleave attention with recurrent components are increasingly competitive with pure Transformers, yet standard LoRA practice applies adapters uniformly without considering the distinct functional roles of each component type. We systematically study component-type LoRA placement across two hybrid architectures -- Qwen3.5-0.8B (sequential, GatedDeltaNet + softmax attention) and Falcon-H1-0.5B (parallel, Mamba-2 SSM + attention) -- fine-tuned on three domains and evaluated on five benchmarks. We find that the attention pathway -- despite being the minority component -- consistently outperforms full-model adaptation with 5-10x fewer trainable parameters. Crucially, adapting the recurrent backbone is destructive in sequential hybrids (-14.8 pp on GSM8K) but constructive in parallel ones (+8.6 pp). We further document a transfer asymmetry: parallel hybrids exhibit positive cross-task transfer while sequential hybrids suffer catastrophic forgetting. These results establish that hybrid topology fundamentally determines adaptation response, and that component-aware LoRA placement is a necessary design dimension for hybrid architectures.

【9】PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
标题:PermaFrost-Attack:LLM训练期间用于植入逻辑地雷的隐形预训练播种(MPS)
链接:https://arxiv.org/abs/2604.22117

作者:Harsh Kumar,Rahul Maity,Tanmay Joshi,Aman Chadha,Vinija Jain,Suranjana Trivedy,Amitava Das
摘要:对齐的大型语言模型(LLM)仍然容易受到对抗性操纵的影响,它们对网络规模预训练的依赖创造了一个微妙但严重的攻击面。我们研究了Stealth Pretraining Seeding(SPS),这是一种新的攻击家族,其中对手在隐身网站上分发少量有毒内容,通过robots.txt将其暴露给网络爬虫,从而增加了这些内容被吸收到未来训练语料库中的可能性,这些训练语料库来自Common Crawl等来源。由于每个有效载荷都很小,分散,表面上是良性的,因此在数据集构建或过滤过程中很难检测到攻击。其结果是一种潜在的中毒形式:在预训练期间嵌入的休眠逻辑地雷在标准评估中基本上是不可见的,但后来可以通过精确的字母数字触发器激活,例如<00TRIGGER00>绕过保障措施。我们称这种攻击为PermaFrost,类比于北极永久冻土:有害物质可以保持冻结,掩埋,长时间不被注意,只有在条件允许时才会重新出现。我们通过PermaFrost-Attack,一个用于潜在概念中毒的受控框架,以及一套几何诊断:热力学长度,光谱曲率和感染追溯图来操作这种威胁。在多个模型系列和规模,我们表明,SPS是广泛有效的,诱导持久的不安全行为,同时往往逃避对齐防御。我们的研究结果确定SPS作为一个实际的和低估的威胁,未来的基础模型。本文介绍了一种新的几何诊断镜头,系统地检查潜在的模型行为,提供了一个原则性的基础,检测,表征和理解的漏洞,可能仍然是不可见的标准评估。
摘要:Aligned large language models(LLMs) remain vulnerable to adversarial manipulation, and their dependence on web-scale pretraining creates a subtle but serious attack surface. We study Stealth Pretraining Seeding (SPS), a new attack family in which adversaries distribute small amounts of poisoned content across stealth websites, expose them to web crawlers through robots.txt, and thereby increase the likelihood that such content is absorbed into future training corpora derived from sources such as Common Crawl. Because each individual payload is tiny, diffuse, and superficially benign, the attack is difficult to detect during dataset construction or filtering. The result is a latent form of poisoning: dormant logic landmines embedded during pretraining that remain largely invisible under standard evaluation, yet can later be activated by precise alphanumeric triggers such as <00TRIGGER00> to bypass safeguards. We call this attack PermaFrost, by analogy to Arctic permafrost: harmful material can remain frozen, buried, and unnoticed for long periods, only to resurface when conditions allow. We operationalize this threat through PermaFrost-Attack, a controlled framework for latent conceptual poisoning, together with a suite of geometric diagnostics: Thermodynamic Length, Spectral Curvature, and the Infection Traceback Graph. Across multiple model families and scales, we show that SPS is broadly effective, inducing persistent unsafe behavior while often evading alignment defenses. Our results identify SPS as a practical and underappreciated threat to future foundation models. This paper introduces a novel geometric diagnostic lens for systematically examining latent model behavior, providing a principled foundation for detecting, characterizing, and understanding vulnerabilities that may remain invisible to standard evaluation.

【10】Removing Sandbagging in LLMs by Training with Weak Supervision
标题:通过监督薄弱的训练消除LLM中的沙袋
链接:https://arxiv.org/abs/2604.22082

作者:Emil Ryd,Henning Bartsch,Julian Stastny,Joe Benton,Vivek Hebbar
摘要:随着人工智能系统开始自动化复杂的任务,监督越来越依赖于较弱的模型或有限的人类监督,无法完全验证输出质量。一个比其主管更有能力的模型可以通过沙袋来利用这一差距,产生看似可以接受但低于其真正能力的工作。即使没有可靠的验证,训练也能得出模型的最佳结果吗?我们使用训练沙袋的模型生物来研究这一点,测试解决问题的数学,研究生水平的科学和竞争性编码任务的启发技术。我们发现,当监督微调(SFT)和强化学习(RL)相结合时,弱监督的训练可以可靠地引出沙袋模型:弱演示的SFT打破了沙袋行为,使RL能够完全引出性能。这两种方法都不能单独可靠地成功-没有SFT的RL几乎总是导致奖励黑客而不是真正的改进。至关重要的是,这依赖于训练与部署无法区分;当模型可以区分训练和部署时,它们可以在训练期间表现良好,同时在之后继续沙袋。我们的研究结果提供了初步的证据,培训是一个可行的缓解沙袋,同时强调了培训的重要性,从部署区分。
摘要:As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more capable than its supervisors could exploit this gap through sandbagging, producing work that appears acceptable but falls short of its true abilities. Can training elicit a model's best work even without reliable verification? We study this using model organisms trained to sandbag, testing elicitation techniques on problem-solving math, graduate-level science, and competitive coding tasks. We find that training with weak supervision can reliably elicit sandbagging models when supervised fine-tuning (SFT) and reinforcement learning (RL) are combined: SFT on weak demonstrations breaks the sandbagging behavior, enabling RL to then fully elicit performance. Neither method succeeds reliably alone-RL without SFT almost always leads to reward hacking rather than genuine improvement. Critically, this relies on training being indistinguishable from deployment; when models can distinguish between training and deployment, they can perform well during training while continuing to sandbag afterward. Our results provide initial evidence that training is a viable mitigation against sandbagging, while highlighting the importance of making training indistinguishable from deployment.

【11】Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores
标题:精神病学LLM下游任务的可靠性审核:LLM生成的住院风险评分
链接:https://arxiv.org/abs/2604.22063

作者:Shevya Pandya,Shinjini Bose,Ananya Joshi
摘要:大型语言模型(LLM)越来越多地用于临床推理和风险评估。然而,他们的解释可靠性在关键和不确定的领域,如精神病学仍然不清楚。先前的工作已经确定了这些系统中的算法偏差和提示敏感性,引起了人们对上下文信息如何影响模型输出的关注,但仍然没有系统的方法来评估这些,特别是在精神病学领域。我们提出了一种可靠性审计下游LLM任务的方法,方法是围绕及时设计的影响进行结构化评估,并将医学上无关紧要的输入纳入预测的住院风险评分,这通常是第一个下游AI临床决策任务。在我们的审核中,生成了一组合成患者特征(n = 50),每个特征包括15个临床相关特征和多达50个临床不显著特征,包括四个提示重构(中性、逻辑、人为影响、临床判断)。我们审计了四个LLM(Gemini 2.5 Flash,LLaMa 3.3 70 b,Claude Sonnet 4.6,GPT-4 o mini),我们的结果表明,包括医学上无关紧要的变量导致所有模型和提示的绝对平均预测住院风险和输出变异性在统计学上显著增加,这表明随着背景噪声的增加,预测稳定性降低。在许多模型提示条件下,临床上不显著的特征对不稳定性有影响,提示变化以依赖于模型的方式独立影响不稳定性的轨迹。这些发现量化了基于LLM的精神病风险评估对非临床信息的敏感性,强调了在临床部署之前系统评估归因稳定性和不确定性行为的必要性。
摘要:Large language models (LLMs) are increasingly utilized in clinical reasoning and risk assessment. However, their interpretive reliability in critical and indeterminate domains such as psychiatry remains unclear. Prior work has identified algorithmic biases and prompt sensitivity in these systems, raising concerns about how contextual information may influence model outputs, but there remains no systematic way to assess these, especially in the psychiatric domain. We propose an approach for reliability auditing downstream LLM tasks by structuring evaluation around the impact of prompt design and the inclusion of medically insignificant inputs on predicted hospitalization risk scores, which is often the first downstream AI clinical-decision-making task. In our audit, a cohort of synthetic patient profiles (n = 50) is generated, each consisting of 15 clinically relevant features and up to 50 clinically insignificant features, across four prompt reframings (neutral, logical, human impact, clinical judgment). We audit four LLMs (Gemini 2.5 Flash, LLaMa 3.3 70b, Claude Sonnet 4.6, GPT-4o mini), and our results show that including medically insignificant variables resulted in a statistically significant increase in the absolute mean predicted hospitalization risk and output variability across all models and prompts, indicating reduced predictive stability as contextual noise increased. Clinically insignificant features had an effect on instability across many model-prompt conditions, and prompt variations independently affected the trajectory of instability in a model-dependent manner. These findings quantify how LLM-based psychiatric risk assessments are sensitive to non-clinical information, highlighting the need for systematic evaluations of attributional stability and uncertainty behavior like this before clinical deployments.

【12】Lightweight Retrieval-Augmented Generation and Large Language Model-Based Modeling for Scalable Patient-Trial Matching
标题:用于可扩展患者试验匹配的轻量级检索增强生成和基于大型语言模型的建模
链接:https://arxiv.org/abs/2604.22061

作者:Xiaodi Li,Yang Xiao,Munhwan Lee,Konstantinos Leventakos,Young J. Juhn,David Jones,Terence T. Sio,Wei Liu,Maria Vassilaki,Nansu Zong
备注:31 pages, 7 figures
摘要:患者试验匹配需要对长时间、异构的电子健康记录(EHR)和复杂的资格标准进行推理,这对可扩展性、泛化和计算效率提出了重大挑战。现有的方法要么依赖于使用大型语言模型(LLM)的全文档处理,这在计算上是昂贵的,要么使用传统的机器学习方法,难以捕获非结构化的临床叙述。在这项工作中,我们提出了一个轻量级的框架,结合检索增强生成和大型语言模型为基础的建模可扩展的患者试验匹配。该框架明确地分离了两个关键组成部分:检索增强生成用于从长EHR中识别临床相关片段,降低输入复杂性,而大型语言模型用于将这些选定片段编码为信息表示。这些表示通过降维进一步细化,并使用轻量级预测器进行建模,从而实现高效和可扩展的下游分类。我们在多个公共基准(n2 c2,SIGIR,TREC 2021/2022)和Mayo Clinic(MCPMD)的真实多模态数据集上评估了所提出的方法。结果表明,基于检索的信息选择显着降低了计算负担,同时保留临床有意义的信号。我们进一步证明,冻结LLM为结构化临床数据提供了强有力的表示,而微调对于非结构化临床叙述建模至关重要。重要的是,所提出的轻量级流水线实现了与端到端LLM方法相当的性能,并且计算成本大大降低。
摘要:Patient-trial matching requires reasoning over long, heterogeneous electronic health records (EHRs) and complex eligibility criteria, posing significant challenges for scalability, generalization, and computational efficiency. Existing approaches either rely on full-document processing with large language models (LLMs), which is computationally expensive, or use traditional machine learning methods that struggle to capture unstructured clinical narratives. In this work, we propose a lightweight framework that combines retrieval-augmented generation and large language model-based modeling for scalable patient-trial matching. The framework explicitly separates two key components: retrieval-augmented generation is used to identify clinically relevant segments from long EHRs, reducing input complexity, while large language models are used to encode these selected segments into informative representations. These representations are further refined through dimensionality reduction and modeled using lightweight predictors, enabling efficient and scalable downstream classification. We evaluate the proposed approach on multiple public benchmarks (n2c2, SIGIR, TREC 2021/2022) and a real-world multimodal dataset from Mayo Clinic (MCPMD). Results show that retrieval-based information selection significantly reduces computational burden while preserving clinically meaningful signals. We further demonstrate that frozen LLMs provide strong representations for structured clinical data, whereas fine-tuning is essential for modeling unstructured clinical narratives. Importantly, the proposed lightweight pipeline achieves performance comparable to end-to-end LLM approaches with substantially lower computational cost.

【13】LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
标题:LayerBoost:减少高效的LLM的层意识注意力
链接:https://arxiv.org/abs/2604.22050

作者:Mohamed Ali Souibgui,Jan Fostier,Rodrigo Abadía-Heredia,Bohdan Denysenko,Christian Marschke,Igor Peric
摘要:Transformers主要依赖于softmax attention,这引入了关于序列长度的二次复杂度,并且仍然是高效推理的主要瓶颈。先前关于线性或混合注意力的工作通常会在所有层中均匀地取代softmax注意力,这通常会导致显著的性能下降或需要大量的重新训练来恢复模型质量。 这项工作提出了LayerBoost,层感知的注意力减少方法,有选择地修改的注意力机制的基础上的敏感性个别Transformer层。它首先对预训练的模型进行系统的敏感性分析,以确定对保持性能至关重要的层。在这种分析的指导下,可以应用三种不同的策略:在高度敏感的层中保留标准的softmax注意力,在中度敏感的层中用线性滑动窗口注意力取代它,以及在表现出低敏感性的层中完全删除注意力。 为了在这些架构修改后恢复性能,我们引入了一个轻量级的基于蒸馏的修复阶段,只需要10M额外的训练令牌。LayerBoost减少了推理延迟,并在高并发情况下将吞吐量提高了68%,同时保持了具有竞争力的模型质量。它匹配几个基准的基础模型的性能,表现出只有轻微的退化,并显着优于国家的最先进的注意力线性化方法。这些效率的提高使得我们的方法特别适合高并发服务和硬件受限的部署场景,其中推理成本和内存占用是关键瓶颈。
摘要:Transformers are mostly relying on softmax attention, which introduces quadratic complexity with respect to sequence length and remains a major bottleneck for efficient inference. Prior work on linear or hybrid attention typically replaces softmax attention uniformly across all layers, often leading to significant performance degradation or requiring extensive retraining to recover model quality. This work proposes LayerBoost, a layer-aware attention reduction method that selectively modifies the attention mechanism based on the sensitivity of individual transformer layers. It first performs a systematic sensitivity analysis on a pretrained model to identify layers that are critical for maintaining performance. Guided by this analysis, three distinct strategies can be applied: retaining standard softmax attention in highly sensitive layers, replacing it with linear sliding window attention in moderately sensitive layers, and removing attention entirely in layers that exhibit low sensitivity. To recover performance after these architectural modifications, we introduce a lightweight distillation-based healing phase requiring only 10M additional training tokens. LayerBoost reduces inference latency and improves throughput by up to 68% at high concurrency, while maintaining competitive model quality. It matches base model performance on several benchmarks, exhibits only minor degradations on others, and significantly outperforms state-of-the-art attention linearization methods. These efficiency gains make our method particularly well-suited for high-concurrency serving and hardware-constrained deployment scenarios, where inference cost and memory footprint are critical bottlenecks.

【14】Shared Lexical Task Representations Explain Behavioral Variability In LLMs
标题:共享词汇任务表示解释LLM中的行为变异性
链接:https://arxiv.org/abs/2604.22027

作者:Zhuonan Yang,Jacob Xiaochen Li,Francisco Piedrahita Velez,Eric Todd,David Bau,Michael L. Littman,Stephen H. Bach,Ellie Pavlick
摘要:对大型语言模型(LLM)最常见的抱怨之一是它们的快速敏感性-也就是说,它们执行任务或提供问题正确答案的能力可能无法预测地取决于问题的提出方式。我们通过比较两种非常不同但常用的提示风格来研究这种变化:基于提示的提示,它以自然语言描述任务,基于示例的提示,它提供上下文中的Few-Shot演示对来说明任务。我们发现,尽管大的变化,性能作为一个功能的提示,该模型从事一些共同的底层机制在不同的提示的任务。具体来说,我们确定了特定于任务的注意头,其输出字面上描述的任务-我们配音词汇任务头-并表明,这些头是共享的提示风格,并触发随后的答案生产。我们进一步发现,提示之间的行为变化可以解释的程度,这些头被激活,和失败至少有时是由于竞争的任务表示,稀释了目标任务的信号。我们的研究结果一起呈现了一个越来越清晰的画面,LLM的内部表示如何解释行为,否则似乎是特殊的用户和开发人员。
摘要:One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a task or provide a correct answer to a question can depend unpredictably on the way the question is posed. We investigate this variation by comparing two very different but commonly-used styles of prompting: instruction-based prompts, which describe the task in natural language, and example-based prompts, which provide in-context few-shot demonstration pairs to illustrate the task. We find that, despite large variation in performance as a function of the prompt, the model engages some common underlying mechanisms across different prompts of a task. Specifically, we identify task-specific attention heads whose outputs literally describe the task -- which we dub lexical task heads -- and show that these heads are shared across prompting styles and trigger subsequent answer production. We further find that behavioral variation between prompts can be explained by the degree to which these heads are activated, and that failures are at least sometimes due to competing task representations that dilute the signal of the target task. Our results together present an increasingly clear picture of how LLMs' internal representations can explain behavior that otherwise seems idiosyncratic to users and developers.

Graph相关(图学习|图神经网络|图优化等)(4篇)

【1】Operational Feature Fingerprints of Graph Datasets via a White-Box Signal-Subspace Probe
标题:通过白盒信号子空间探测器获取图形数据集的操作特征指纹
链接:https://arxiv.org/abs/2604.22676

作者:Yuchen Xiong,Swee Keong Yeap,Zhen Hong Ban
备注:21 pages, 10 figures, 7 tables
摘要:图神经网络实现了很强的节点分类精度,但它们学习的消息传递将自我属性、邻域平滑、高通图差异、类几何和分类器边界纠缠在一个不透明的表示中。这模糊了为什么对节点进行分类以及数据集需要什么样的特征级图学习机制。 我们提出了WG-SRC,一种用于预测和图数据集诊断的白盒信号子空间探测器。WG-SRC用一个固定的、命名的原始特征的图形信号字典、行归一化和行归一化的低通传播以及高通图形差异来代替学习的消息传递。它结合了Fisher坐标选择、类PCA子空间、封闭形式的多α岭分类和基于验证的分数融合,因此预测和分析使用显式的类子空间、能量控制维度和封闭形式的线性决策。 作为一种白盒图学习工具,WG-SRC使用预测性能来验证其诊断:在六个节点分类数据集上,支架与复制的图基线保持竞争力,并在对齐的分割下实现正平均增益。它的图谱由预测器产生,将行为分解为原始特征、低通、高通、类几何和脊边界分量。这些操作特征指纹区分了低通主导的Amazon图、混合的高通和类几何复杂的Chameleon行为以及原始或边界敏感的WebKB图。作为内在分类器输出而不是事后解释,这些指纹为以后的分析和特定于指纹的修改提供了评估后的指导。对齐的机械干预支持这一指导,指出高通块作为可移动的噪音,当原始功能应该保留,当脊型边界校正的事项。
摘要 :Graph neural networks achieve strong node-classification accuracy, but their learned message passing entangles ego attributes, neighborhood smoothing, high-pass graph differences, class geometry, and classifier boundaries in an opaque representation. This obscures why a node is classified and what feature-level graph-learning mechanisms a dataset requires. We propose WG-SRC, a white-box signal-subspace probe for prediction and graph dataset diagnosis. WG-SRC replaces learned message passing with a fixed, named graph-signal dictionary of raw features, row-normalized and symmetric-normalized low-pass propagation, and high-pass graph differences. It combines Fisher coordinate selection, class-wise PCA subspaces, closed-form multi-alpha ridge classification, and validation-based score fusion, so prediction and analysis use explicit class subspaces, energy-controlled dimensions, and closed-form linear decisions. As a white-box graph-learning instrument, WG-SRC uses predictive performance to validate its diagnostics: across six node-classification datasets, the scaffold remains competitive with reproduced graph baselines and achieves positive average gain under aligned splits. Its atlas, produced by a predictor, decomposes behavior into raw-feature, low-pass, high-pass, class-geometric, and ridge-boundary components. These operational feature fingerprints distinguish low-pass-dominated Amazon graphs, mixed high-pass and class-geometrically complex Chameleon behavior, and raw- or boundary-sensitive WebKB graphs. As intrinsic classifier outputs rather than post-hoc explanations, these fingerprints provide post-evaluation guidance for later analysis and dataset-specific modification. Aligned mechanistic interventions support this guidance by indicating when high-pass blocks act as removable noise, when raw features should be preserved, and when ridge-type boundary correction matters.

【2】Distance-Misaligned Training in Graph Transformers and Adaptive Graph-Aware Control
标题:图变换器中的距离失调训练和自适应图感知控制
链接:https://arxiv.org/abs/2604.22413

作者:Qinhan Hou,Jing Tang
备注:Accepted by Graph Signal Processing Workshop 2026 as an extended abstract
摘要:图Transformers可以全局混合信息,但这种灵活性也会产生故障模式:一些任务需要远程通信,而另一些任务则可以通过本地交互更好地服务。我们通过上下文随机块模型图上的合成节点分类基准来研究这一点,其中标签由本地和远壳信号的可控混合物生成。我们将距离失调训练定义为标签相关信息所在位置与模型在图距离上分配通信的位置之间的不匹配。在这个基准上,我们发现三点。首先,偏好图距离偏差随任务局部性而系统地变化。第二,一个预言机自适应控制器,离线访问任务侧的距离目标,几乎匹配的最佳固定偏置跨制度和强烈改善了中性基线的混合和本地任务。第三,任务不可知的零差距控制器是较弱的,这表明适应本身是不够的,控制目标的问题。这些结果表明,距离分辨诊断是有用的理解图形Transformer故障和设计图形感知控制。
摘要:Graph Transformers can mix information globally, but this flexibility also creates failure modes: some tasks require long-range communication while others are better served by local interaction. We study this through a synthetic node-classification benchmark on contextual stochastic block model graphs, where labels are generated by a controllable mixture of local and far-shell signals. We define distance-misaligned training as a mismatch between where label-relevant information lies and where the model allocates communication over graph distance. On this benchmark, we find three points. First, the preferred graph-distance bias changes systematically with task locality. Second, an oracle adaptive controller, given offline access to the task-side distance target, nearly matches the best fixed bias across regimes and strongly improves over a neutral baseline on mixed and local tasks. Third, a task-agnostic zero-gap controller is weaker, indicating that adaptation alone is not enough and that the control target matters. These results suggest that distance-resolved diagnosis is useful for understanding Graph Transformer failures and for designing graph-aware control.

【3】FixV2W: Correcting Invalid CVE-CWE Mappings with Knowledge Graph Embeddings
标题:FixV 2 W:使用知识图嵌入纠正无效的CVE-CWE映射
链接:https://arxiv.org/abs/2604.22176

作者:Sevval Simsek,Varsha Athreya,David Starobinski
摘要:常见漏洞和暴露(CVE)和常见弱点枚举(CWE)条目之间的准确映射对于有效的漏洞管理和风险评估至关重要。然而,公共数据库,如国家漏洞数据库(NVD),遭受不一致和不完整的CVE到CWE映射,复杂的自动化分析和补救。我们介绍了FixV 2 W,这是一种轻量级的方法,它利用知识图嵌入和纵向趋势来提高NVD的映射精度。FixV 2 W系统地分析历史重映射模式,并利用NVD和CWE数据中的层次关系来预测链接到禁止或不鼓励类别的漏洞的更精确的CWE映射。我们根据2021年8月至2024年12月期间收集的测试数据集对FixV 2 W进行了广泛的实验评估。考虑排名前10的预测,结果表明,FixV 2 W预测了69%的已利用漏洞的正确CWE映射,这些漏洞在被利用之前具有无效的CWE。我们还表明,FixV 2 W显著提高了依赖于NVD数据的ML模型的性能。例如,对于用于发现未知CVE-CWE映射的模型,FixV 2 W将平均倒数秩(MRR)从0.174提高到0.608。这些结果表明,FixV 2 W是一种很有前途的方法来识别和阻止新出现的威胁。
摘要:Accurate mapping between Common Vulnerabilities and Exposures (CVE) and Common Weakness Enumeration (CWE) entries is critical for effective vulnerability management and risk assessment. However, public databases, such as the National Vulnerability Database (NVD), suffer from inconsistent and incomplete CVE to CWE mappings, complicating automated analysis and remediation. We introduce FixV2W, a lightweight approach that leverages knowledge graph embeddings and longitudinal trends to improve mapping accuracy of the NVD. FixV2W systematically analyzes historical remapping patterns and leverages hierarchical relationships within NVD and CWE data to predict more precise CWE mappings for vulnerabilities linked to Prohibited or Discouraged categories. We run extensive experimental evaluation of FixV2W, based on test data set collected between August 2021 and December 2024. Considering the Top 10 ranked predictions, the results show that FixV2W predicts the correct CWE mappings for 69% of exploited vulnerabilities that had invalid CWEs before they were exploited. We also show that FixV2W significantly improves the performance of ML models relying on NVD data. For instance, for a model geared at uncovering unknown CVE-CWE mappings, FixV2W improves the Mean Reciprocal Rank (MRR) from 0.174 to 0.608. These results show that FixV2W is a promising approach to identify and thwart emerging threats.

【4】Mochi: Aligning Pre-training and Inference for Efficient Graph Foundation Models via Meta-Learning
标题:Mochi:通过元学习调整预训练和推理以获得高效的图形基础模型
链接:https://arxiv.org/abs/2604.22031

作者:João Mattos,Arlei Silva
备注:20 pages, 7 figures
摘要:我们提出了Mochi,一个图形基础模型,通过采用基于元学习的训练框架来解决任务统一和训练效率问题。先前的模型使用基于重建的目标(如链接预测)进行预训练,并假设所得到的表示可以通过单独的统一步骤(如类原型)与下游任务对齐。我们通过合成和真实世界的实验证明,这个过程,虽然简单和直观,有直接影响下游任务性能的限制。为了解决这些局限性,Mochi在反映下游评估协议的Few-Shot片段上进行预训练,将训练目标与推理对齐,而不是依赖于事后统一步骤。我们表明,Mochi及其更强大的变体Mochi++与现有的Graph Foundation Models相比,在跨越节点分类,链接预测和图形分类的25个真实世界的图形数据集上实现了具有竞争力或优越的性能,同时需要比最强基线少27倍的训练时间。
摘要:We propose Mochi, a Graph Foundation Model that addresses task unification and training efficiency by adopting a meta-learning based training framework. Prior models pre-train with reconstruction-based objectives such as link prediction, and assume that the resulting representations can be aligned with downstream tasks through a separate unification step such as class prototypes. We demonstrate through synthetic and real-world experiments that this procedure, while simple and intuitive, has limitations that directly affect downstream task performance. To address these limitations, Mochi pre-trains on few-shot episodes that mirror the downstream evaluation protocol, aligning the training objective with inference rather than relying on a post-hoc unification step. We show that Mochi, along with its more powerful variant Mochi++, achieves competitive or superior performance compared to existing Graph Foundation Models across 25 real-world graph datasets spanning node classification, link prediction, and graph classification, while requiring 8$\sim$27 times less training time than the strongest baseline.

Transformer(2篇)

【1】Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
标题:括号序列变换器中的分解性和因果使用
链接:https://arxiv.org/abs/2604.22128

作者:Aryan Sharma,Cutter Dawes,Shivam Raval
摘要:当接受需要理解层次结构的任务的训练时,已经发现Transformers以不同的方式表示这种层次结构:在剩余流的几何结构中,以及在保持后进先出顺序的堆栈式注意模式中。然而,目前还不清楚这些表示是否是因果关系使用或只是解码。我们研究了在Dyck语言(一种平衡括号序列的正式语言)上训练的Transformers中的这种差距,其中层次化的基础事实是明确的。通过对剩余流和注意模式的探测和干预,我们发现深度、距离和栈顶信号都是可解码的,但它们的因果作用是不同的。具体来说,掩蔽注意到真正的顶部的堆栈位置会导致长距离精度急剧下降,而烧蚀低维残留流子空间的影响相对较小。这些结果,扩展到模板化的自然语言设置,表明即使在一个受控的设置,其中相关的层次变量是已知的,解码本身并不意味着因果关系的使用。
摘要 :When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining a last-in, first-out ordering. However, it remains unclear whether these representations are causally used or merely decodable. We examine this gap in transformers trained on the Dyck language (a formal language of balanced bracket sequences), where the hierarchical ground truth is explicit. By probing and intervening on the residual stream and attention patterns, we find that depth, distance, and top-of-stack signals are all decodable, yet their causal roles diverge. Specifically, masking attention to the true top-of-stack position causes a sharp drop in long-distance accuracy, while ablating low-dimensional residual stream subspaces has comparatively little effect. These results, which extend to a templated natural language setting, suggest that even in a controlled setting where the relevant hierarchical variables are known, decodability alone does not imply causal use.

【2】Universal Transformers Need Memory: Depth-State Trade-offs in Adaptive Recursive Reasoning
标题:通用Transformer需要记忆:自适应回归推理中的深度状态权衡
链接:https://arxiv.org/abs/2604.21999

作者:Grigory Sapunov
备注:12 pages, 7 figures, 8 tables. Code: https://github.com/che-shr-cat/utm-jax
摘要:我们研究了学习记忆令牌作为计算暂存器的单块通用Transformer(UT)与自适应计算时间(ACT)的数独极限,组合推理基准。我们发现内存令牌是经验上必要的:在所有测试的配置中- 3个种子,多个令牌计数,两个初始化方案,ACT和固定深度处理-没有内存令牌的配置实现了非平凡的性能。最佳计数表现出一个明显的低阈值(T=0总是失败,T=4是边界,T=8可靠地成功81个细胞的难题),然后是一个稳定的平台(T=8-32,57.4% +/- 0.7%精确匹配),并在T=64时因注意力稀释而崩溃。 在实验过程中,我们发现了一个路由器初始化陷阱,导致>70%的训练运行失败:默认的零偏置初始化(p ~ 0.5)和Graves推荐的正偏置(p ~ 0.73)都会导致令牌在初始化时的~2步后停止,进入模型无法逃脱的浅平衡(halt ~ 5-7)。将偏倚反转为-3(“深启动”,p ~ 0.05)可消除该失效模式。我们通过消融确认陷阱是ACT初始化所固有的,而不是我们架构选择的伪影。 随着可靠的训练建立,我们发现(1)ACT提供了比固定深度处理更一致的结果(56.9% +/- 0.7% vs 53.4% +/- 9.3%,3个种子);(2)ACT与lambda预热实现匹配准确性(57.0% +/- 1.1%)使用34%的思考步骤;(3)注意力头专注于内存读取器,约束传播器和递归深度的积分器。代码可在https://github.com/che-shr-cat/utm-jax上获得。
摘要:We study learned memory tokens as computational scratchpad for a single-block Universal Transformer (UT) with Adaptive Computation Time (ACT) on Sudoku-Extreme, a combinatorial reasoning benchmark. We find that memory tokens are empirically necessary: across all configurations tested -- 3 seeds, multiple token counts, two initialization schemes, ACT and fixed-depth processing -- no configuration without memory tokens achieves non-trivial performance. The optimal count exhibits a sharp lower threshold (T=0 always fails, T=4 is borderline, T=8 reliably succeeds for 81-cell puzzles) followed by a stable plateau (T=8-32, 57.4% +/- 0.7% exact-match) and collapse from attention dilution at T=64. During experimentation, we identify a router initialization trap that causes >70% of training runs to fail: both default zero-bias initialization (p ~ 0.5) and Graves' recommended positive bias (p ~ 0.73) cause tokens to halt after ~2 steps at initialization, settling into a shallow equilibrium (halt ~ 5-7) that the model cannot escape. Inverting the bias to -3 ("deep start," p ~ 0.05) eliminates this failure mode. We confirm through ablation that the trap is inherent to ACT initialization, not an artifact of our architecture choices. With reliable training established, we show that (1) ACT provides more consistent results than fixed-depth processing (56.9% +/- 0.7% vs 53.4% +/- 9.3% across 3 seeds); (2) ACT with lambda warmup achieves matching accuracy (57.0% +/- 1.1%) using 34% fewer ponder steps; and (3) attention heads specialize into memory readers, constraint propagators, and integrators across recursive depth. Code is available at https://github.com/che-shr-cat/utm-jax.

GAN|对抗|攻击|生成相关(7篇)

【1】Adversarial Malware Generation in Linux ELF Binaries via Semantic-Preserving Transformations
标题:通过语义保持转换在Linux ELF二进制文件中生成对抗性恶意软件
链接:https://arxiv.org/abs/2604.22639

作者:Lukáš Hrdonka,Martin Jureček
摘要:近年来,恶意软件的开发和检测发生了重大变化,因为机器学习等现代概念已用于对抗性攻击和防御。尽管对Windows可移植可执行(PE)文件进行了深入的研究,但对Linux可执行和可链接格式(ELF)的研究很少。在这项工作中,我们总结了在这一领域提交的学术论文,并开发了一个新的对抗性恶意软件生成器的ELF格式。使用各种指标,我们彻底评估了我们的生成器,并实现了67.74%的规避率,同时在所使用数据集的平均情况下将恶意软件检测器的置信度更改了-0.50。在我们的方法中,我们选择MalConv作为目标分类器。使用这个分类器,我们发现最成功的修改使用良性文件的典型字符串作为数据源。我们进行了各种实验,并得出结论,目标分类器似乎敏感的字符串在任何位置的可执行文件。
摘要:Malware development and detection have undergone significant changes in recent years as modern concepts, such as machine learning, have been used for both adversarial attacks and defense. Despite intensive research on Windows Portable Executable (PE) files, there is minimal work on Linux Executable and Linkable Format (ELF). In this work, we summarize the academic papers submitted in this field and develop a new adversarial malware generator for the ELF format. Using a variety of metrics, we thoroughly evaluated our generator and achieved an Evasion Rate of 67.74 % while changing the confidence of the malware detector by -0.50 in the mean case for the dataset used. In our approach, we chose MalConv as the target classifier. Using this classifier, we found that the most successful modifications used strings typical of benign files as a data source. We conducted a variety of experiments and concluded that the target classifier appears sensitive to strings at any location within the executable file.

【2】Adversarial Co-Evolution of Malware and Detection Models: A Bilevel Optimization Perspective
标题:恶意软件和检测模型的对抗协同进化:两层优化的角度
链接:https://arxiv.org/abs/2604.22569

作者:Olha Jurečková,Martin Jureček,Matouš Kozák,Róbert Lórencz
摘要:基于机器学习的恶意软件检测器越来越容易受到对抗性示例的攻击。传统的防御,如一次性对抗训练,往往无法抵御使用强化学习绕过检测的自适应攻击者。本文提出了一种基于两级优化的鲁棒防御框架,将防御者和攻击者之间的策略交互显式建模为对抗性协同进化过程。我们评估我们的方法使用MAB恶意软件框架对三个不同的恶意软件家族:Mokes,Strab和DCRat。我们的实验结果表明,虽然标准分类器和基本的对抗性再训练通常仍然很脆弱,显示逃避率高达90%,但所提出的双层优化方法始终实现了近乎完全的免疫,将逃避率降低到0 - 1.89%。此外,迭代框架显著增加了攻击者的查询复杂性,将成功规避的平均成本提高了两个数量级。这些发现表明,通过双层优化对攻击和防御的迭代周期进行建模,对于开发能够抵御不断变化的对抗性威胁的弹性恶意软件检测系统至关重要。
摘要:Machine learning-based malware detectors are increasingly vulnerable to adversarial examples. Traditional defenses, such as one-shot adversarial training, often fail against adaptive attackers who use reinforcement learning to bypass detection. This paper proposes a robust defense framework based on bilevel optimization, explicitly modeling the strategic interaction between a defender and an attacker as an adversarial co-evolutionary process. We evaluate our approach using the MAB-malware framework against three distinct malware families: Mokes, Strab, and DCRat. Our experimental results demonstrate that while standard classifiers and basic adversarial retraining often remain vulnerable, showing evasion rates as high as 90 %, the proposed bilevel optimization approach consistently achieves near-total immunity, reducing evasion rates to 0 - 1.89 %. Furthermore, the iterative framework significantly increases the attacker's query complexity, raising the average cost of successful evasion by up to two orders of magnitude. These findings suggest that modeling the iterative cycle of attack and defense through bilevel optimization is essential for developing resilient malware detection systems capable of withstanding evolving adversarial threats.

【3】TabSCM: A practical Framework for Generating Realistic Tabular Data
标题:TabSCP:生成真实表格数据的实用框架
链接:https://arxiv.org/abs/2604.22337

作者:Sven Jacob,Bardh Prenkaj,Weijia Shao,Gjergji Kasneci
摘要:大多数表格数据生成器匹配边际统计数据,但忽略因果结构,导致下游模型学习虚假或不公平的模式。我们提出了TabSCM,一个混合类型的生成器,保留这些因果依赖关系。从任何因果结构发现算法发现的完全部分有向无环图(CPDAG)开始,TabSCM(i)将边定向到DAG,(ii)用KDE或分类频率拟合根节点边缘,(iii)学习拓扑有序的结构分配。这种分配是使用条件扩散模型为连续变量的子节点和梯度提升树的分类。祖先采样产生语义上有效的记录,并使确切的反事实查询。在七个公共数据集上,包括医疗保健,金融,住房,环境,TabSCM在统计保真度,下游效用和隐私风险方面与最先进的GAN,扩散和LLM基线相匹配或超越,同时还降低了违规率并提供了因果关系有意义和强大的条件干预。由于生成被分解为显式方程,它运行速度比仅扩散模型快583倍,并为公平审计和政策模拟暴露了可解释的旋钮,使TabSCM成为现实主义,可解释性和因果合理性的实际选择。
摘要 :Most tabular-data generators match marginal statistics yet ignore causal structure, leading downstream models to learn spurious or unfair patterns. We present TabSCM, a mixed-type generator that preserves those causal dependencies. Starting from a Completed Partially Directed Acyclic Graph (CPDAG) found by any causal structure discovery algorithm, TabSCM (i) orients edges to a DAG, (ii) fits root-node marginals with KDE or categorical frequencies, and (iii) learns topologically ordered structural assignments. Such assignments are achieved using conditional diffusion models for continuous variables as child nodes and gradient-boosted trees for categorical ones. Ancestral sampling yields semantically valid records and enables exact counterfactual queries. On seven public datasets, encompassing healthcare, finance, housing, environment, TabSCM matches or surpasses state-of-the-art GAN, diffusion, and LLM baselines in statistical fidelity, downstream utility, and privacy risk, while also cutting rule-violation rates and providing causally meaningful and robust conditional interventions. Because generation is decomposed into explicit equations, it runs up to 583$\times$ faster than diffusion-only models and exposes interpretable knobs for fairness auditing and policy simulation, making TabSCM a practical choice for realism, explainability, and causal soundness.

【4】AI-Driven Performance-to-Design Generation and Optimization of Marine Propellers
标题:人工智能驱动的船舶螺旋桨设计性能生成和优化
链接:https://arxiv.org/abs/2604.22224

作者:Leah Chen,Keni Chih-Hua Wu,Boon Tat Chia,Xiuqing Xing,Jian Cheng Wong
备注:Accepted at OMAE 2026
摘要:人工智能越来越多地用于通过改善决策和缩短迭代周期来加速工程设计。然而,由于缺乏训练数据和缺乏广泛使用的预训练模型,船舶螺旋桨设计的应用仍然具有挑战性。我们通过基于物理的数据生成管道和生成式人工智能框架来解决这一差距,用于针对船用螺旋桨量身定制的直接性能到设计生成。首先,我们建立了一个包含20,000多个四叶和五叶螺旋桨几何形状的数据库,每个都附有模拟的开放水域性能曲线。在此数据集之上,我们开发了一个三模块设计框架:(1)条件生成模型,该模型根据设计规格(如目标推力、功率和直径)提出候选几何形状。(2)性能预测模型,作为神经网络代理实现,可在毫秒内预测推力、扭矩和效率,从而实现对生成的设计的快速评估。(3)一个设计优化阶段,应用进化优化来加强实际约束,如在功率限制下所需的推力和叶片面积比和厚度的界限。在一系列的操作条件下的实验结果表明,该框架可以生成符合规定的性能目标,同时大大减少设计迭代时间相对于传统的专家指导的细化水动力学合理的螺旋桨设计。基于潜在扩散的生成器在相同条件下比条件变分自编码器产生更多样化的设计,这表明扩散模型具有更强的设计空间探索能力。通过将基于物理的数据合成与模块化AI模型相结合,所提出的方法简化了螺旋桨设计周期,并减少了对最终验证阶段昂贵的高保真仿真的依赖。
摘要:AI is increasingly used to accelerate engineering design by improving decision-making and shortening iteration cycles. Application to marine propeller design, however, remains challenging due to scarce training data and the lack of widely available pretrained models. We address this gap with a physics-based data generation pipeline and a generative-AI framework for direct performance-to-design generation tailored to marine propellers. First, we build a database of over 20,000 four- and five-bladed propeller geometries, each accompanied by simulated open-water performance curves. On top of this dataset, we develop a three-module design framework: (1) A Conditional Generation Model that proposes candidate geometries conditioned on design specifications such as target thrust, power, and diameter. (2) A Performance Prediction Model, implemented as a neural-network surrogate, that predicts thrust, torque, and efficiency in milliseconds, enabling rapid evaluation of generated designs. (3) A design refinement stage that applies evolutionary optimization to enforce practical constraints such as required thrust under power limits and bounds on blade-area ratio and thickness. Experimental results over a range of operating conditions show that the framework can generate hydrodynamically plausible propeller designs that match prescribed performance targets while substantially reducing design-iteration time relative to the traditional expert-guided refinement. Latent diffusion-based generator produces more diverse designs under the same conditions than the conditional variational autoencoder, suggesting a stronger capacity for design-space exploration with diffusion models. By coupling physics-based data synthesis with modular AI models, the proposed approach streamlines the propeller design cycle and reduces reliance on expensive high-fidelity simulations to final validation stages.

【5】Sharpness-Aware Poisoning: Enhancing Transferability of Injective Attacks on Recommender Systems
标题:敏锐意识中毒:增强推荐系统上注射性攻击的可转移性
链接:https://arxiv.org/abs/2604.22170

作者:Junsong Xie,Yonghui Yang,Pengyang Shao,Le Wu
摘要:推荐系统~(RS)已经被证明容易受到注入攻击,其中攻击者注入有限的虚假用户配置文件以促进目标项目向真实用户暴露以获得不道德的收益(例如,经济或政治优势)。由于攻击者通常缺乏部署在目标RS中的受害者模型的知识,现有的方法诉诸于使用固定的代理模型来模仿潜在的受害者模型。尽管取得了相当大的进展,我们认为,\texit {中毒的数据生成的代理模型可以用来攻击其他受害者模型}的假设是一厢情愿的。当代理模型和受害者模型之间存在显著的结构差异时,攻击的可转移性不可避免地受到影响。直观地说,如果我们能够识别最坏情况的受害者模型,并针对它迭代优化中毒效果,那么生成的中毒数据将更好地转移到其他受害者模型。然而,在攻击过程中,准确地识别最坏情况下的受害者模型是具有挑战性的,由于受害者模型的大空间。为此,在这项工作中,我们提出了一种新的攻击方法,称为Sharpness-Aware Poisoning(\textit{SharpAP})。具体而言,它采用尖锐度感知最小化原则,以寻求近似最坏情况下的受害者模型,并优化中毒的数据专门为这个最坏情况下的模型。SharpAP的中毒攻击被公式化为最小-最大-最小三级优化问题。通过将SharpAP集成到迭代攻击过程中,我们的方法可以生成更健壮的中毒数据,这些数据对模型结构的变化不太敏感,从而减轻了对代理模型的过拟合。在三个真实数据集上的综合实验比较表明,\name~可以显着提高攻击的可转移性。
摘要:Recommender Systems~(RS) have been shown to be vulnerable to injective attacks, where attackers inject limited fake user profiles to promote the exposure of target items to real users for unethical gains (e.g., economic or political advantages). Since attackers typically lack knowledge of the victim model deployed in the target RS, existing methods resort to using a fixed surrogate model to mimic the potential victim model. Despite considerable progress, we argue that the assumption that \textit{poisoned data generated for the surrogate model can be used to attack other victim models} is wishful. When there are significant structural discrepancies between the surrogate and victim models, the attack transferability inevitably suffers. Intuitively, if we can identify the worst-case victim model and iteratively optimize the poisoning effect specifically against it, then the generated poisoned data would be better transferred to other victim models. However, exactly identifying the worst-case victim model during the attack process is challenging due to the large space of victim models. To this end, in this work, we propose a novel attack method called Sharpness-Aware Poisoning (\textit{SharpAP}). Specifically, it employs the sharpness-aware minimization principle to seek the approximately worst-case victim model and optimizes the poisoned data specifically for this worst-case model. The poisoning attack with SharpAP is formulated as a min-max-min tri-level optimization problem. By integrating SharpAP into the iterative process for attacks, our method can generate more robust poisoned data which is less sensitive to the shift of model structure, mitigating the overfitting to the surrogate model. Comprehensive experimental comparisons on three real-world datasets demonstrate that \name~can significantly enhance the attack transferability.

【6】Generating Synthetic Malware Samples Using Generative AI
标题:使用生成性AI生成合成恶意软件样本
链接:https://arxiv.org/abs/2604.22084

作者:Tiffany Bao,Kylie Trousil,Quang Duy Tran,Fabio Di Troia,Younghee Park
备注:12 pages, 8 figures. This paper has been published in IEEE Access, available at this URL: https://ieeexplore.ieee.org/document/10947040
摘要:恶意软件攻击对网络安全领域各种规模的组织都有重大的负面影响。最近,恶意软件研究人员越来越多地转向机器学习技术来对抗恶意软件中使用的复杂混淆方法。然而,使用各种混淆技术收集不同的恶意软件样本集是具有挑战性的,并且通常需要数年时间,特别是对于新开发的恶意软件。机器学习模型的一个众所周知的局限性进一步加剧了这个问题:当训练数据稀缺时,它们的性能很差。在本文中,我们提出了一个新的系统,用于生成合成恶意软件样本,以增加不平衡的恶意软件数据集。我们的方法将恶意软件二进制样本分解为助记符操作码序列,利用自然语言处理来提取恶意软件操作码特征背后的上下文含义,以帮助学习本文中采用的生成AI(GenAI),生成对抗网络(GAN),Wasserstein生成对抗网络与梯度惩罚(WGAN-GP)和修改后的扩散模型。实验结果表明,使用基于扩散的合成数据来增强训练数据,可以显著提高小类的分类性能,平均提高高达60%。这一增强最终导致整体恶意软件分类性能达到96%,提高了8%。这些发现证明了合成数据的高质量和保真度、其鲁棒性及其在恶意软件分析中的潜在应用。具体而言,合成恶意软件数据证明在改善次要恶意软件类别的分类和检测率方面是有效的,即使已知恶意软件数据的大小非常小。
摘要 :Malware attacks have a significant negative impact on organizations of varied scales in the field of cybersecurity. Recently, malware researchers have increasingly turned to machine learning techniques to combat sophisticated obfuscation methods used in malware. However, collecting a diverse set of malware samples with various obfuscation techniques is challenging and often takes years, especially for newly developed malware. This issue is further compounded by a well-known limitation of machine learning models: their poor performance when training data is scarce. In this paper, we propose a new system for generating synthetic malware samples to augment imbalanced malware dataset. Our approach decomposes malware binary samples into mnemonic opcode sequences, leveraging natural language processing to extract contextual meaning behind malware opcode features to aid the learning of generative AI (GenAI) employed in this paper, Generative Adversarial Networks (GAN), Wasserstein Generative Adversarial Networks with Gradient Penalty (WGAN-GP), and a modified Diffusion model. The experiment results show that augmenting training data with Diffusion-based synthetic data significantly improves classification performance for minor classes by up to 60% on average. This enhancement ultimately leads to an overall malware classification performance of 96%, an 8% improvement. These findings demonstrate the high quality and fidelity of the synthetic data, its robustness, and its potential applications in malware analysis. Specifically, synthetic malware data proves effective in improving the classification of minor malware classes and detection rates, even though the size of known malware data is significantly small.

【7】Feedback Over Form: Why Execution Feedback Matters More Than Pipeline Topology in 1-3B Code Generation
标题:形式上的反馈:为什么在1-3B代码生成中执行反馈比管道布局更重要
链接:https://arxiv.org/abs/2604.21950

作者:Charles Junichi McAndrews
备注:17 pages main text, 2 page references, 3 figures. Code: https://github.com/L3G/feedback-over-form
摘要:小型语言模型(1-3B)在本地运行是可行的,但在较难的代码生成任务上受到限制。我们想知道将它们组合到管道中是否可以恢复一些失去的功能。我们研究了从1-3B模型构建的代码生成管道,并使用NEAT启发的进化搜索来测试更复杂的管道结构是否有助于超越简单的细化循环。我们评估HumanEval(164个问题)和消毒MBPP(427个问题),所有这些都在一台笔记本电脑上进行本地推理。在两个基准测试中,带有执行反馈的自我改进将代码生成提高了4个以上的标准差。在机制上,改进的好处很有限:改进修复了许多运行时错误(特别是NameError和SyntaxError),但很少修复AssertionError等逻辑错误。在我们测试的通用模型池中,生成器身份比精炼器能力更重要:与3B精炼器配对的1.5B生成器与同时扮演这两个角色的3B模型相匹配。提前停止是必要的;没有它,每个迭代都是净负的。代码专用模型优于任何通用管道配置,这表明模型专用化比管道架构更重要。没有执行反馈的初步纯文本管道实验没有显示出这种规模的收益。在我们的受限搜索空间中,进化搜索大多重新发现了我们手动发现的简单的生成-执行-细化循环,并没有从添加的拓扑结构中获得明显的好处。单一评估适应性将结果放大5- 7%,选择幸运的基因组而不是好的基因组。在这些1-3B规模的基准测试中,在确定组合是否有帮助时,执行反馈比增加的管道复杂性更重要。
摘要:Small language models (1-3B) are practical to run locally, but individually limited on harder code generation tasks. We ask whether composing them into pipelines can recover some of that lost capability. We study code generation pipelines built from 1-3B models with execution feedback, and use a NEAT-inspired evolutionary search to test whether more complex pipeline structure helps beyond a simple refinement loop. We evaluate on HumanEval (164 problems) and sanitized MBPP (427 problems), all with local inference on a single laptop. Self-refinement with execution feedback improves code generation by more than 4 standard deviations on both benchmarks. The gains are narrow in mechanism: refinement fixes many runtime errors (especially NameError and SyntaxError), but rarely fixes logic errors such as AssertionError. Within our tested general-purpose model pool, generator identity mattered less than refiner capability: a 1.5B generator paired with a 3B refiner matched a 3B model doing both roles. Early stopping is essential; without it, every iteration is net-negative. The code-specialized models outperform every general-purpose pipeline configuration, suggesting model specialization matters more than pipeline architecture. Preliminary text-only pipeline experiments without execution feedback did not show gains at this scale. In our constrained search space, evolutionary search mostly rediscovered the same simple generate-execute-refine loop we found manually, with no clearly significant gain from added topology. Single-evaluation fitness inflates results by 5-7 percent, selecting lucky genomes over good ones. On these benchmarks at 1-3B scale, execution feedback mattered more than added pipeline complexity in determining whether composition helped.

半/弱/无/有监督|不确定性|主动学习(5篇)

【1】Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering
标题:通过跨语言迁移和无监督集群在低资源班图语言中进行Zero-Shot形态发现
链接:https://arxiv.org/abs/2604.22723

作者:Hillary Mutisya,John Mugane
摘要:我们提出了一种方法,通过结合跨语言迁移学习和无监督聚类来发现低资源班图语中的形态特征。应用于Giriama(nyf),一种只有91个标记范式的语言,我们的管道发现了2,455个单词的名词类分配,并识别了两个以前没有记录的形态模式:Class 2的a前缀变体(元音合并-两个相邻元音的合并-wa-,95.1%的一致性)和收缩的k '前缀(98.5%的一致性)。对444个已知的Giriama动词范例的外部验证证实了78.2%的词形还原准确率,而v3语料库扩展到19,624个单词(9,014个唯一词元),在所有主要单词类别中实现了97.3%的分割和86.7%的词形还原率。我们将斯瓦希里语的迁移学习和无监督聚类结合起来,通过加权投票,利用互补的优势:迁移擅长同源检测(利用约60%的词汇重叠),而聚类发现了迁移不可见的语言特定创新。我们发布了所有代码和发现的词典,以支持低资源班图语言的形态文档。
摘要:We present a method for discovering morphological features in low-resource Bantu languages by combining cross-lingual transfer learning with unsupervised clustering. Applied to Giriama (nyf), a language with only 91 labeled paradigms, our pipeline discovers noun class assignments for 2,455 words and identifies two previously undocumented morphological patterns: an a- prefix variant for Class 2 (vowel coalescence - the merger of two adjacent vowels - of wa-, 95.1% consistency) and a contracted k'- prefix (98.5% consistency). External validation on 444 known Giriama verb paradigms confirms 78.2% lemmatization accuracy, while a v3 corpus expansion to 19,624 words (9,014 unique lemmas) achieves 97.3% segmentation and 86.7% lemmatization rates across all major word classes. Our ensemble of transfer learning from Swahili and unsupervised clustering, combined via weighted voting, exploits complementary strengths: transfer excels at cognate detection (leveraging ~60% vocabulary overlap) while clustering discovers language-specific innovations invisible to transfer. We release all code and discovered lexicons to support morphological documentation for low-resource Bantu languages.

【2】On the Properties of Feature Attribution for Supervised Contrastive Learning
标题:监督对比学习的特征归因性质
链接:https://arxiv.org/abs/2604.22540

作者:Leonardo Arrighi,Julia Eva Belloni,Aurélie Gallet,Ivan Gentile,Matteo Lippi,Marco Zullich
摘要:大多数用于分类的神经网络(NN)都使用交叉熵作为损失函数进行训练。这种方法要求模型具有显式的分类层。然而,存在替代方法,例如对比学习(CL)。CL没有显式地操作分类,而是让NN产生一个嵌入空间,在这个空间中,相似数据的投影被拉到一起,而不相似数据的投影被推开。在监督CL(SCL)的情况下,标签被采用作为相似性标准,从而创建一个嵌入空间,其中投影数据点被很好地聚类。SCL在对抗鲁棒性和分布外检测方面提供了CE的关键优势,因此使其成为安全关键场景中更自然的选择。在本文中,我们实证表明,用SCL训练的图像分类神经网络在忠实性、复杂性和连续性方面比CL提供了更高质量的特征归因解释。这些结果加强了以前的研究结果,基于CL的方法时,针对更值得信赖的和透明的NN,可以指导从业者在选择培训目标,不仅针对准确性,但也透明的模型。
摘要:Most Neural Networks (NNs) for classification are trained using Cross-Entropy as a loss function. This approach requires the model to have an explicit classification layer. However, there exist alternative approaches, such as Contrastive Learning (CL). Instead of explicitly operating a classification, CL has the NN produce an embedding space where projections of similar data are pulled together, while projections of dissimilar data are pushed apart. In the case of Supervised CL (SCL), labels are adopted as similarity criteria, thus creating an embedding space where the projected data points are well-clustered. SCL provides crucial advantages over CE with regard to adversarial robustness and out-of-distribution detection, thus making it a more natural choice in safety-critical scenarios. In the present paper, we empirically show that NNs for image classification trained with SCL present higher-quality feature attribution explanations than CL with regard to faithfulness, complexity, and continuity. These results reinforce previous findings about CL-based approaches when targeting more trustworthy and transparent NNs and can guide practitioners in the selection of training objectives targeting not only accuracy, but also transparency of the models.

【3】Revisiting Neural Activation Coverage for Uncertainty Estimation
标题:重新审视不确定性估计的神经激活覆盖
链接:https://arxiv.org/abs/2604.22360

作者:Benedikt Franke,Nils Förster,Frank Köster,Asja Fischer,Markus Lange,Arne Raulf
备注:Published in 34th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, ESANN 2026
摘要:神经激活覆盖(NAC)是最近提出的一种用于分布外检测和泛化的技术。我们建立在这个有前途的基础上,并扩展该方法的工作作为一个不确定性估计技术已经训练的人工神经网络在回归领域。我们的实验证实NAC不确定性分数比其他技术更有意义,例如Monte-Carlo Dropout。
摘要:Neural activation coverage (NAC) is a recently-proposed technique for out-of-distribution detection and generalization. We build upon this promising foundation and extend the method to work as an uncertainty estimation technique for already-trained artificial neural networks in the domain of regression. Our experiments confirm NAC uncertainty scores to be more meaningful than other techniques, e.g. Monte-Carlo Dropout.

【4】Fast Neural-Network Approximation of Active Target Search Under Uncertainty
标题:不确定性下主动目标搜索的快速神经网络逼近
链接:https://arxiv.org/abs/2604.22254

作者:Bilal Yousuf,Zsofia Lendek,Lucian Busoniu
摘要 :我们解决的问题,搜索未知数量的固定目标在未知位置与移动代理。在测量不确定性条件下,利用概率假设密度滤波器估计目标的期望个数。现有的规划器,如主动搜索(AS)及其间歇变体(ASI),实现准确的检测,但需要昂贵的在线优化。为了减少在线计算,我们建议使用卷积神经网络通过直接推理来近似AS或ASI决策。该网络使用多通道网格对AS/ASI数据进行训练,该网格对目标信念、代理位置、访问历史和边界信息进行编码。均匀和集群目标分布的模拟表明,该网络实现了检测率与AS或ASI相当,同时减少了数量级的计算。
摘要:We address the problem of searching for an unknown number of stationary targets at unknown positions with a mobile agent. A probability hypothesis density filter is used to estimate the expected number of targets under measurement uncertainty. Existing planners, such as Active Search (AS) and its Intermittent variant (ASI), achieve accurate detection but require costly online optimization. To reduce online computation, we propose to use a convolutional neural network to approximate AS or ASI decisions through direct inference. The network is trained on AS/ASI data using a multi-channel grid that encodes target beliefs, the agent position, visitation history, and boundary information. Simulations with uniform and clustered target distributions show that the network achieves detection rates comparable to AS or ASI while reducing computation by orders of magnitude.

【5】Anatomy-Aware Unsupervised Detection and Localization of Retinal Abnormalities in Optical Coherence Tomography
标题:光学相干断层扫描中视网膜异常的解剖感知无监督检测和定位
链接:https://arxiv.org/abs/2604.22139

作者:Tania Haghighi,Sina Gholami,Hamed Tabkhi,Minhaj Nur Alam
备注:11 pages, 3 figures, accepted in CVPR-CV4Clinical
摘要:光学相干断层扫描(OCT)成像的可靠自动分析对于诊断视网膜疾病至关重要,但面临一个关键障碍:需要昂贵的劳动密集型专家注释。受监督的深度学习模型很难在不同的病理、成像设备和患者人群中推广,因为它们对注释异常的词汇量有限。我们提出了一种无监督的异常检测框架,该框架可以学习健康视网膜解剖结构的规范分布,而无需病变注释,直接解决了临床部署中的注释效率挑战。我们的方法利用在正常B扫描上训练的离散潜在模型来捕获OCT特定的结构模式。为了增强临床稳健性,我们结合了视网膜层感知监督和结构化三联体学习,以将健康与病理表征分开,从而提高了不同成像条件下的模型可靠性。在推理过程中,通过重建差异检测和定位异常,从而实现图像和像素级识别,而无需疾病特异性标签。在Kermany数据集(AUROC:0.799)上,我们的方法大大优于VAE,VQVAE,VQGAN和f-AnoGAN基线。关键的是,Srinivasan的跨数据集评估达到了AUROC 0.884,具有出色的泛化能力,证明了鲁棒的域适应性。在外部RETOUCH基准测试中,无监督异常分割获得了具有竞争力的Dice(0.200)和mIoU(0.117)分数,验证了跨机构的可重复性。
摘要:Reliable automated analysis of Optical Coherence Tomography (OCT) imaging is crucial for diagnosing retinal disorders but faces a critical barrier: the need for expensive, labor-intensive expert annotations. Supervised deep learning models struggle to generalize across diverse pathologies, imaging devices, and patient populations due to their restricted vocabulary of annotated abnormalities. We propose an unsupervised anomaly detection framework that learns the normative distribution of healthy retinal anatomy without lesion annotations, directly addressing annotation efficiency challenges in clinical deployment. Our approach leverages a discrete latent model trained on normal B-scans to capture OCT-specific structural patterns. To enhance clinical robustness, we incorporate retinal layer-aware supervision and structured triplet learning to separate healthy from pathological representations, improving model reliability across varied imaging conditions. During inference, anomalies are detected and localized via reconstruction discrepancies, enabling both image and pixel-level identification without requiring disease-specific labels. On the Kermany dataset (AUROC: 0.799), our method substantially outperforms VAE, VQVAE, VQGAN, and f-AnoGAN baselines. Critically, cross-dataset evaluation on Srinivasan achieves AUROC 0.884 with superior generalization, demonstrating robust domain adaptation. On the external RETOUCH benchmark, unsupervised anomaly segmentation achieves competitive Dice (0.200) and mIoU (0.117) scores, validating reproducibility across institutions.

迁移|Zero/Few/One-Shot|自适应(5篇)

【1】Adaptive Head Budgeting for Efficient Multi-Head Attention
标题:自适应头部预调整以实现高效的多头注意力
链接:https://arxiv.org/abs/2604.22583

作者:Bilal Faye,Abdoulaye Mbaye,Hanane Azzag,Mustapha Lebbah
摘要:Transformers已经成为广泛领域的主导架构,这主要是由于多头注意力在捕获不同表示子空间方面的有效性。然而,标准的多头注意力会针对每次输入统一激活所有头部,无论任务要求或输入复杂性如何。在许多情况下,特别是对于粗粒度的任务,如文本分类,相关的信息往往是全球性的,并不需要注意头的完全多样性。因此,使用固定数量的磁头可能会引入不必要的计算成本,或者在分配与输入不匹配时导致次优性能。为了解决这个问题,我们引入了BudgetFormer,一个Transformer架构,配备了一个自适应的多头注意力机制,动态分配计算资源。我们的方法学习,对于每个输入,头部预算对应于所需的注意头部的数量,和相关性分布,选择信息量最大的头部。我们还提出了一种基于探索和利用权衡的训练策略,允许模型在收敛到有效的使用模式之前发现有效的头部配置。对不同复杂度的文本分类任务的实验表明,我们的方法在FLOPs和内存方面降低了推理成本,同时还实现了可以超过标准的全多头注意力的性能。这些结果突出了自适应磁头分配作为提高Transformer模型效率和有效性的原则性方法的潜力。
摘要:Transformers have become the dominant architecture across a wide range of domains, largely due to the effectiveness of multi-head attention in capturing diverse representation subspaces. However, standard multi-head attention activates all heads uniformly for every input, regardless of task requirements or input complexity. In many scenarios, particularly for coarse-grained tasks such as text classification, the relevant information is often global and does not require the full diversity of attention heads. As a consequence, using a fixed number of heads can introduce unnecessary computational cost or lead to suboptimal performance when the allocation does not match the input. To address this limitation, we introduce BudgetFormer, a Transformer architecture equipped with an adaptive multi-head attention mechanism that dynamically allocates computational resources. Our approach learns, for each input, both a head budget corresponding to the number of attention heads required, and a relevance distribution that selects the most informative heads. We also propose a training strategy based on an exploration and exploitation trade-off, allowing the model to discover effective head configurations before converging to efficient usage patterns. Experiments on text classification tasks of varying complexity show that our method reduces inference cost in terms of FLOPs and memory, while also achieving performance that can surpass standard full multi-head attention. These results highlight the potential of adaptive head allocation as a principled approach to improving both efficiency and effectiveness in Transformer models.

【2】Towards Adaptive Continual Model Merging via Manifold-Aware Expert Evolution
标题:通过Manifold感知专家进化实现自适应连续模型合并
链接:https://arxiv.org/abs/2604.22464

作者:Haiyun Qiu,Xingyu Wu,Kay Chen Tan
摘要:连续模型合并(CMM)顺序地将特定于任务的模型集成到一个统一的架构中,而无需密集的重新训练。然而,现有的CMM方法是阻碍了一个基本的饱和冗余困境:骨干为中心的方法面临参数饱和和表示的干扰在固定的容量,而混合专家(MoE)的变种诉诸不加选择的扩展,招致专家冗余和路由瓶颈依赖于额外的数据驱动的优化。为了解决这些挑战,我们提出了MADE-IT(流形感知动态专家进化和隐式路由),一种自适应CMM方法,通过在流形几何中建立内在的专家表示来协调专家管理和激活。我们引入了一个基于投影的子空间亲和力度量加上一个分布感知的自适应阈值机制,以指导自主专家进化,协调多样性与建筑简约。此外,为了绕过参数化的门控网络,我们设计了一个无数据和无训练的隐式路由机制,通过特征子空间对齐激活专家。大量的实验表明,MADE-IT在长期和混合任务序列的准确性和鲁棒性方面始终优于强基线,同时显着修剪冗余专家,特别是在通用模块和早期层中。
摘要:Continual Model Merging (CMM) sequentially integrates task-specific models into a unified architecture without intensive retraining. However, existing CMM methods are hindered by a fundamental saturation-redundancy dilemma: backbone-centric approaches face parameter saturation and representation interference within fixed capacities, whereas Mixture-of-Experts (MoE) variants resort to indiscriminate expansion, incurring expert redundancy and a routing bottleneck reliant on additional data-driven optimization. To resolve these challenges, we propose MADE-IT (Manifold-Aware Dynamic Expert Evolution and Implicit rouTing), an adaptive CMM method that orchestrates expert management and activation by grounding intrinsic expert representations in manifold geometry. We introduce a projection-based subspace affinity metric coupled with a distribution-aware adaptive threshold mechanism to guide autonomous expert evolution, harmonizing diversity with architectural parsimony. Furthermore, to bypass parameterized gating networks, we design a data-free and training-free implicit routing mechanism that activates experts via feature-subspace alignment. Extensive experiments demonstrate that MADE-IT consistently outperforms strong baselines in accuracy and robustness across long-horizon and shuffled task sequences, while significantly pruning redundant experts, particularly within generic modules and early layers.

【3】Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
标题:连续学习中Adam下梯度修改的隐藏故障模式和作为修复的自适应脱钩时刻路由
链接:https://arxiv.org/abs/2604.22407

作者:Yuelin Hu,Zhenbo Yu,Zhengxue Cheng,Wei Liu,Li Song
备注:28 pages, 5 figures, preprint
摘要 :许多连续学习方法修改上游的梯度(例如,投影,惩罚重新缩放,重放混合),同时将亚当视为中性后端。我们表明这种组合物具有隐藏的失效模式。在一个高重叠、非自适应的8域连续LM中,所有共享路由投影基线都接近于香草遗忘(12.5- 12.8 vs. 13.2)。0.5%的重放缓冲区是最强的共享替代方案,但仍达到11.6,而固定强度解耦则低于14.1。只有自适应解耦路由保持稳定在9.4,比vanilla提高了3.8个单位。在16域流上,其在最强共享路由投影基线上的增益增长到4.5- 4.8个单位。在干净的基准上,这种失败在很大程度上是不可见的。 我们通过亚当的二阶矩途径解释了这种效应:在测试的制度中,投影引起旧方向有效学习率的1/(1-alpha)膨胀,在8个alpha值中匹配测量值在8%以内。同样的冲突出现在惩罚方法,重放混合和LoRA下的7 B规模上。我们的修复仅将修改后的梯度路由到第一时刻,同时保留幅度忠实的第二时刻统计数据,具有可感知的自适应强度。这个简单的改变是唯一经过测试的配置,它始终避免了方法,优化器和规模的崩溃。
摘要:Many continual-learning methods modify gradients upstream (e.g., projection, penalty rescaling, replay mixing) while treating Adam as a neutral backend. We show this composition has a hidden failure mode. In a high-overlap, non-adaptive 8-domain continual LM, all shared-routing projection baselines collapse close to vanilla forgetting (12.5--12.8 vs. 13.2). A 0.5% replay buffer is the strongest shared alternative but still reaches 11.6, while fixed-strength decoupling falls below vanilla at 14.1. Only adaptive decoupled routing remains stable at 9.4, improving over vanilla by 3.8 units. On a 16-domain stream, its gain over the strongest shared-routing projection baseline grows to 4.5--4.8 units. The failure is largely invisible on clean benchmarks. We explain this effect through Adam's second-moment pathway: in the tested regime, projection induces a 1/(1-alpha) inflation of the old-direction effective learning rate, matching measurements within 8% across eight alpha values. The same conflict appears with penalty methods, replay mixing, and at 7B scale under LoRA. Our fix routes the modified gradient only to the first moment while preserving magnitude-faithful second-moment statistics, with overlap-aware adaptive strength. This simple change is the only tested configuration that consistently avoids collapse across methods, optimizers, and scale.

【4】Wiggle and Go! System Identification for Zero-Shot Dynamic Rope Manipulation
标题:扭动并开始!零发射动态绳索操纵的系统识别
链接:https://arxiv.org/abs/2604.22102

作者:Arthur Jakobsson,Abhinav Mahajan,Karthik Pullalarevu,Krishna Suresh,Yunchao Yao,Yuemin Mao,Bardienus Duisterhof,Shahram Najam Syed,Jeffrey Ichnowski
摘要:许多机器人任务是不可原谅的;动态投掷中的一个错误可能导致不可接受的延迟或不可恢复的故障。为了缓解这一问题,我们提出了一种新的方法,利用学习的模拟先验知识来告知绳索的目标条件动态操纵,以实现高效和准确的任务执行。用于动态绳索操纵的相关方法要么需要大的真实世界数据集来估计绳索行为,要么需要对目标完成任务的尝试进行迭代改进。我们介绍摇摆和去!,一个系统识别的两阶段框架,使zero-shot任务绳操纵。该框架包括一个系统识别模块,观察绳运动预测描述性的物理参数,然后通知的目标条件动作预测的机器人执行zero-shot在现实中的优化方法。我们的方法在多个动态操作任务中实现了强大的性能,这些任务由相同的任务不可知的系统识别模块实现,该模块提供了不同操作任务之间的无缝切换,允许单个模型支持各种各样的操作策略。我们实现了3.55厘米的平均精度在3D目标打击在现实中使用绳索系统参数相比,15.34厘米的精度时,我们的任务模型是没有系统参数通知。我们实现了皮尔森相关系数为0.95之间的傅立叶频率的预测和真实的绳子上看不见的轨迹。项目网址请参见https://wiggleandgo.github.io/
摘要:Many robotic tasks are unforgiving; a single mistake in a dynamic throw can lead to unacceptable delays or unrecoverable failure. To mitigate this, we present a novel approach that leverages learned simulation priors to inform goal-conditioned dynamic manipulation of ropes for efficient and accurate task execution. Related methods for dynamic rope manipulation either require large real-world datasets to estimate rope behavior or the use of iterative improvements on attempts at the task for goal completion. We introduce Wiggle and Go!, a system-identification, two-stage framework that enables zero-shot task rope manipulation. The framework consists of a system identification module that observes rope movement to predict descriptive physical parameters, which then informs an optimization method for goal-conditioned action prediction for the robot to execute zero-shot in the real. Our method achieves strong performance across multiple dynamic manipulation tasks enabled by the same task-agnostic system identification module which offers seamless switching between different manipulation tasks, allowing a single model to support a diverse array of manipulation policies. We achieve a 3.55 cm average accuracy on 3D target striking in real using rope system parameters in comparison to 15.34 cm accuracy when our task model is not system-parameter-informed. We achieve a Pearson correlation coefficient of 0.95 between Fourier frequencies of the predicted and real ropes on an unseen trajectory. Project website please see https://wiggleandgo.github.io/

【5】Pack only the essentials: Adaptive dictionary learning for kernel ridge regression
标题:只打包必需品:用于内核岭回归的自适应字典学习
链接:https://arxiv.org/abs/2604.22386

作者:Daniele Calandriello,Alessandro Lazaric,Michal Valko
备注:In NeurIPS 2016 Workshop on Adaptive and Scalable Nonparametric Methods in Machine Learning (ASNMML)
摘要:核岭回归(KRR)的一个主要限制是存储和操作n个样本的核矩阵K_n需要O(n^2)空间,这对于大n很快变得不可行。Nystrom近似通过从K_n中抽取m列,将空间复杂度降低到O(nm)。只有当m与K_n的最大自由度成比例时,均匀采样才能保持KRR精度(最高可达100%),对于具有高一致性的数据集,这可能需要O(n)列。根据它们的岭杠杆分数(RLS)对列进行采样可以给出精确的Nystrom近似,其中m与有效维数成比例,但计算精确的RLS也需要O(n^2)空间。 (Calcillello et al. 2016)提出了INK-Estimate,这是一种增量处理数据集并实时更新RLS,有效尺寸和Nystrom近似的算法。它的空间复杂度与有效维数成比例,但引入了对K_n的最大特征值的依赖,在最坏的情况下是O(n)。 在本文中,我们介绍SQUEAK,一个新的算法,建立在INK估计,但使用非归一化的RLS。因此,该算法更简单,不需要估计归一化的有效维数,并且实现了比精确RLS采样差的常数因子的空间复杂度。
摘要:One of the major limits of kernel ridge regression (KRR) is that storing and manipulating the kernel matrix K_n for n samples requires O(n^2) space, which rapidly becomes unfeasible for large n. Nystrom approximations reduce the space complexity to O(nm) by sampling m columns from K_n. Uniform sampling preserves KRR accuracy (up to epsilon) only when m is proportional to the maximum degree of freedom of K_n, which may require O(n) columns for datasets with high coherence. Sampling columns according to their ridge leverage scores (RLS) gives accurate Nystrom approximations with m proportional to the effective dimension, but computing exact RLS also requires O(n^2) space. (Calandriello et al. 2016) propose INK-Estimate, an algorithm that processes the dataset incrementally and updates RLS, effective dimension, and Nystrom approximations on-the-fly. Its space complexity scales with the effective dimension but introduces a dependency on the largest eigenvalue of K_n, which in the worst case is O(n). In this paper we introduce SQUEAK, a new algorithm that builds on INK-Estimate but uses unnormalized RLS. As a consequence, the algorithm is simpler, does not need to estimate the effective dimension for normalization, and achieves a space complexity that is only a constant factor worse than exact RLS sampling.

强化学习(4篇)

【1】SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning
标题:SOLAR-RL:半在线长期作业强化学习
链接:https://arxiv.org/abs/2604.22558

作者:Jichao Wang,Liuyang Bian,Yufeng Zhou,Han Xiao,Yue Pan,Guozhi Wang,Hao Wang,Zhaoxiong Wang,Yafei Wen,Xiaoxin Chen,Shuai Ren,Lingfang Zeng
备注:14 pages, 11 figures. Accepted to Findings of the Association for Computational Linguistics: ACL 2026
摘要:随着多模态大型语言模型(MLLM)的成熟,GUI代理正在从静态交互发展到复杂的导航。虽然强化学习(RL)已经成为一个很有前途的范例训练MLLM代理动态GUI任务,其有效的应用面临着困境。标准离线RL通常依赖于静态步骤级数据,忽略了全局轨迹语义,如任务完成和执行质量。相反,在线强化学习捕捉到了长期的动态,但却面临着高交互成本和潜在的环境不稳定性。为了弥补这一差距,我们提出了SOLAR-RL(半在线长视野分配强化学习)。我们的框架不是仅仅依赖昂贵的在线互动,而是将全球轨迹洞察直接整合到离线学习过程中。具体来说,我们从静态数据中重建不同的推出候选者,使用每步有效性信号检测第一个故障点,并追溯分配密集的步骤级奖励与目标对齐的整形,以反映自动化级别的执行质量,有效地模拟在线反馈,而无需交互成本。大量的实验表明,与强基线相比,SOLAR-RL显著提高了长期任务完成率和鲁棒性,为自主GUI导航提供了一个样本高效的解决方案。
摘要:As Multimodal Large Language Models (MLLMs) mature, GUI agents are evolving from static interactions to complex navigation. While Reinforcement Learning (RL) has emerged as a promising paradigm for training MLLM agents on dynamic GUI tasks, its effective application faces a dilemma. Standard Offline RL often relies on static step-level data, neglecting global trajectory semantics such as task completion and execution quality. Conversely, Online RL captures the long-term dynamics but suffers from high interaction costs and potential environmental instability. To bridge this gap, we propose SOLAR-RL (Semi-Online Long-horizon Assignment Reinforcement Learning). Instead of relying solely on expensive online interactions, our framework integrates global trajectory insights directly into the offline learning process. Specifically, we reconstruct diverse rollout candidates from static data, detect the first failure point using per-step validity signals, and retroactively assign dense step-level rewards with target-aligned shaping to reflect trajectory-level execution quality, effectively simulating online feedback without interaction costs. Extensive experiments demonstrate that SOLAR-RL significantly improves long-horizon task completion rates and robustness compared to strong baselines, offering a sample-efficient solution for autonomous GUI navigation.

【2】Preserve Support, Not Correspondence: Dynamic Routing for Offline Reinforcement Learning
标题:保留支持,而不是对应:离线强化学习的动态路由
链接:https://arxiv.org/abs/2604.22229

作者:Zhancun Mu,Guangyu Zhao,Yiwu Zhong,Chi Zhang
备注:17 pages, 4 figures
摘要:一步离线RL演员是有吸引力的,因为它们避免了通过长时间迭代采样器的反向传播,并保持推理便宜,但它们仍然必须在评论家的指导下进行改进,而不会偏离数据集可以支持的动作。在最近的一步提取管道中,强迭代教师为每个潜在绘制提供一个目标动作,并且要求相同的学生输出完成两项工作:向更高的Q移动并保持在配对端点附近。如果这两个方向不一致,损失将它们作为对同一样本的妥协,即使附近的更好的行动仍然得到数据的局部支持。我们提出了DROL,一个潜在的条件一步演员训练与顶级1动态路由。对于每个状态,演员从有界潜在先验中采样$K$个候选动作,将每个数据集动作分配给其最近的候选者,并使用行为克隆和评论家指导仅更新获胜者。因为路由是从当前候选几何体重新计算的,所以支持区域的所有权可以在学习过程中跨候选者转移。这给了一个单步执行者空间来进行逐点提取难以捕获的局部改进,同时在测试时保留单遍推理。在OGBench和D4 RL上,DROL与一步FQL基线相比具有竞争力,改进了许多OGBench任务组,同时在AntMaze和Adroit上保持强劲。项目页面:https://muzhancun.github.io/preprints/DROL。
摘要:One-step offline RL actors are attractive because they avoid backpropagating through long iterative samplers and keep inference cheap, but they still have to improve under a critic without drifting away from actions that the dataset can support. In recent one-step extraction pipelines, a strong iterative teacher provides one target action for each latent draw, and the same student output is asked to do both jobs: move toward higher Q and stay near that paired endpoint. If those two directions disagree, the loss resolves them as a compromise on that same sample, even when a nearby better action remains locally supported by the data. We propose DROL, a latent-conditioned one-step actor trained with top-1 dynamic routing. For each state, the actor samples $K$ candidate actions from a bounded latent prior, assigns each dataset action to its nearest candidate, and updates only that winner with Behavior Cloning and critic guidance. Because the routing is recomputed from the current candidate geometry, ownership of a supported region can shift across candidates over the course of learning. This gives a one-step actor room to make local improvements that pointwise extraction struggles to capture, while retaining single-pass inference at test time. On OGBench and D4RL, DROL is competitive with the one-step FQL baseline, improving many OGBench task groups while remaining strong on both AntMaze and Adroit. Project page: https://muzhancun.github.io/preprints/DROL.

【3】ReCast: Recasting Learning Signals for Reinforcement Learning in Generative Recommendation
标题:ReCast:在生成式推荐中重新铸造学习信号以进行强化学习
链接:https://arxiv.org/abs/2604.22169

作者:Peiyan Zhang,Hanmo Liu,Chengxuan Tong,Yuxia Wu,Wei Guo,Yong Liu
摘要:通用的基于组的RL假设采样的卷展组已经是可用的学习信号。我们表明,这一假设打破了稀疏命中生成式推荐,其中许多采样组永远不会成为学习。我们提出了ReCast,这是一个先修复再对比的学习信号框架,它首先恢复全零组的最小可学习性,然后用最强正面和最难负面的边界聚焦对比更新取代全组奖励标准化。ReCast保持外部RL框架不变,仅修改组内信号构造,并部分地将rollout搜索宽度从actor-side更新宽度扩展。在多个生成式推荐任务中,ReCast始终优于OpenOneRec-RL,在Pass@1中实现了高达36.6%的相对改进。它的匹配预算优势更大:ReCast仅用4.1%的推出预算就达到了基线的目标性能,并且这种优势随着模型规模的扩大而扩大。相同的设计还可以直接获得系统级增益,将执行器端更新时间缩短16.60倍,将峰值分配内存降低16.5%,并将执行器MFU提高14.2%。机制分析表明,ReCast缓解了持久的全零/单次命中机制,在自然积极因素稀缺时恢复了可学习性,并将浪费的部署预算转换为更稳定的策略更新。这些结果表明,对于生成式推荐,决定性的RL问题不仅是如何分配奖励,而且如何从稀疏的结构化监督中构造可学习的优化事件。
摘要:Generic group-based RL assumes that sampled rollout groups are already usable learning signals. We show that this assumption breaks down in sparse-hit generative recommendation, where many sampled groups never become learnable at all. We propose ReCast, a repair-then-contrast learning-signal framework that first restores minimal learnability for all-zero groups and then replaces full-group reward normalization with a boundary-focused contrastive update on the strongest positive and the hardest negative. ReCast leaves the outer RL framework unchanged, modifies only within-group signal construction, and partially decouples rollout search width from actor-side update width. Across multiple generative recommendation tasks, ReCast consistently outperforms OpenOneRec-RL, achieving up to 36.6% relative improvement in Pass@1. Its matched-budget advantage is substantially larger: ReCast reaches the baseline's target performance with only 4.1% of the rollout budget, and this advantage widens with model scale. The same design also yields direct system-level gains, reducing actor-side update time by 16.60x, lowering peak allocated memory by 16.5%, and improving actor MFU by 14.2%. Mechanism analysis shows that ReCast mitigates the persistent all-zero / single-hit regime, restores learnability when natural positives are scarce, and converts otherwise wasted rollout budget into more stable policy updates. These results suggest that, for generative recommendation, the decisive RL problem is not only how to assign rewards, but how to construct learnable optimization events from sparse, structured supervision.

【4】Insect-inspired modular architectures as inductive biases for reinforcement learning
标题:昆虫启发的模块化架构作为强化学习的归纳偏见
链接:https://arxiv.org/abs/2604.22081

作者:Anne E. Staples
摘要:大多数连续控制中使用的重复学习(RL)控制器在结构上是集中式的:观察被压缩到一个单一的潜在状态,从中产生值估计和动作。生物控制系统通常以不同的方式组织。特别是昆虫,协调导航,航向稳定,记忆和上下文相关的行动选择通过分布式电路,而不是一个单一的单片控制器。出于这种对比,我们研究了一种RL策略架构,该架构将控制分解为用于感觉编码、标题表示、稀疏联想记忆、经常性命令生成和局部运动控制的交互模块,并具有一种学习仲裁机制,该机制在模块之间分配运动权限。该模型进行评估的二维导航任务,需要同时寻找食物,避障,和捕食者逃脱。在一个用邻近策略优化(PPO)训练75次更新的六种子捕食者导航实验中,模块化策略在测试的控制器中实现了最强的最终平均性能,最终情景回报率为$-2798.8\pm964.4$,而集中式门控复发单元(GRU)为$-3778.0\pm628.1$,集中式多层感知器(MLP)。模块化策略还实现了最低的最终值损失和稳定的PPO优化统计,同时将模块分配熵驱动到$0.0457\pm0.0244$,表明高度选择性的控制分配。这些结果表明,分布式控制可以作为一个有用的归纳偏见RL问题,涉及动态竞争的行为目标。
摘要:Most reinforcement-learning (RL) controllers used in continuous control are architecturally centralized: observations are compressed into a single latent state from which both value estimates and actions are produced. Biological control systems are often organized differently. Insects, in particular, coordinate navigation, heading stabilization, memory, and context-dependent action selection through distributed circuits rather than a single monolithic controller. Motivated by this contrast, we study an RL policy architecture that decomposes control into interacting modules for sensory encoding, heading representation, sparse associative memory, recurrent command generation, and local motor control, with a learned arbitration mechanism that allocates motor authority across modules. The model is evaluated on a two-dimensional navigation task that require simultaneous food seeking, obstacle avoidance, and predator escape. In a six-seed predator-navigation experiment trained with Proximal Policy Optimization (PPO) for 75 updates, the modular policy achieves the strongest final mean performance among the tested controllers, with final episodic return $-2798.8\pm964.4$ versus $-3778.0\pm628.1$ for a centralized gated recurrent unit (GRU) and $-4727.5\pm772.5$ for a centralized multilayer perceptron (MLP). The modular policy also attains the lowest final value loss and stable PPO optimization statistics while driving module-assignment entropy to $0.0457\pm0.0244$, indicating highly selective control allocation. These results suggest that distributed control can serve as a useful inductive bias for RL problems involving dynamically competing behavioral objectives.

医学相关(4篇)

【1】A Nationwide Japanese Medical Claims Foundation Model: Balancing Model Scaling and Task-Specific Computational Efficiency
标题:全国范围的日本医疗索赔基金会模型:平衡模型缩放和特定任务的计算效率
链接:https://arxiv.org/abs/2604.22348

作者:Nanae Aratake,Taisei Tosaki,Yuji Okamoto,Eiichiro Uchino,Masaki Nakamura,Nobutomo Matsui,Akiko Hatakama,Yasushi Okuno
备注:14 pages, 5 figures, 3 tables
摘要:使用纵向医疗数据的临床风险预测支持个性化护理。自我监督的基础模型已经成为利用大规模未标记医疗记录的一种有前途的方法。在自然语言处理中,缩放定律表明,较大的模型可以实现可预测的较低预训练损失,支持基础模型范式。然而,对于结构化的医学数据,其特征是有限的词汇和稀疏的观察,增加模型大小是否能持续改善下游预测尚不清楚,因为大多数研究仅评估单个模型规模。在这项研究中,我们评估了结构化医学基础模型的模型规模和下游任务性能之间的关系。使用来自全国519家医院日本索赔数据库的随机样本(230万患者,32家医院),我们在5个尺度(2.2 M-101 M参数)下对仅编码器Transformers进行了疾病发生率和药物预测的预训练。下游性能在任务相关阈值处饱和:疾病预测受益于较大的模型(32 M-101 M),而药物预测在11 M时饱和,预训练时间减少了178小时。在所有任务中,表现最好的模型在精确度-召回率曲线下的区域始终优于Light Gradient Boosting Machine基线。这些发现表明,与单调递减的预训练损失不同,最佳模型大小取决于任务特征。这种依赖于任务的饱和度为平衡结构化医学基础模型中的预测性能和计算成本提供了实际指导。
摘要 :Clinical risk prediction using longitudinal medical data supports individualized care. Self-supervised foundation models have emerged as a promising approach for leveraging large-scale unlabeled healthcare records. In natural language processing, scaling laws suggest that larger models achieve predictably lower pretraining losses, supporting the foundation model paradigm. However, for structured medical data, characterized by a limited vocabulary and sparse observations, whether increasing model size consistently improves downstream predictions is unclear, as most studies evaluate only a single model scale. In this study, we evaluated the relationship between model scale and downstream task performance for structured medical foundation models. Using a random sample (2.3 million patients, 32 hospitals) from a nationwide 519-hospital Japanese claims database, we pretrained encoder-only Transformers at five scales (2.2M-101M parameters) for disease incidence and medication prediction. Downstream performance saturated at task-dependent thresholds: disease prediction benefited from larger models (32M-101M), whereas medication prediction saturated at 11M, reducing pretraining time by 178 h. Across all tasks, the best-performing model consistently outperformed a Light Gradient Boosting Machine baseline in the area under the precision-recall curve. These findings indicate that, unlike the monotonically decreasing pretraining loss, the optimal model size varied depending on task characteristics. This task-dependent saturation provides practical guidance for balancing predictive performance and computational cost in structured medical foundation models.

【2】Conditional anomaly detection using soft harmonic functions: An application to clinical alerting
标题:使用软调和函数的条件异常检测:临床警报的应用
链接:https://arxiv.org/abs/2604.21956

作者:Michal Valko,Hamed Valizadegan,Branislav Kveton,Gregory F. Cooper,Milos Hauskrecht
备注:ICML 2011 Workshop on Machine Learning for Global Challenges. arXiv admin note: substantial text overlap with arXiv:2604.21462. substantial text overlap with arXiv:2604.21462
摘要:及时检测相关事件是临床实践中的一个重要问题。在本文中,我们考虑的问题,有条件的异常检测,旨在识别数据实例与一个不寻常的反应,如遗漏的一个重要的实验室测试。我们开发了一种新的非参数方法的基础上的软谐波解决方案的条件异常检测,我们估计的标签检测异常错误标记的信心。我们进一步正则化的解决方案,以避免孤立的例子和分布支持的边界上的例子的检测。我们证明了所提出的方法在检测真实世界的电子健康记录数据集上的不寻常标签的有效性,并将其与几种基线方法进行比较。
摘要:Timely detection of concerning events is an important problem in clinical practice. In this paper, we consider the problem of conditional anomaly detection that aims to identify data instances with an unusual response, such as the omission of an important lab test. We develop a new non-parametric approach for conditional anomaly detection based on the soft harmonic solution, with which we estimate the confidence of the label to detect anomalous mislabeling. We further regularize the solution to avoid the detection of isolated examples and examples on the boundary of the distribution support. We demonstrate the efficacy of the proposed method in detecting unusual labels on a real-world electronic health record dataset and compare it to several baseline approaches.

【3】Useful nonrobust features are ubiquitous in biomedical images
标题:有用的非稳健特征在生物医学图像中无处不在
链接:https://arxiv.org/abs/2604.22579

作者:Coenraad Mouton,Randle Rabe,Niklas C. Koser,Nicolai Krekiehn,Christopher Hansen,Jan-Bernd Hövener,Claus-C. Glüer
备注:Accepted at The IEEE International Symposium on Biomedical Imaging (ISBI), 2026
摘要:我们研究用于医学成像的深度网络是否学习有用的非鲁棒特征-预测性输入模式,这些模式不是人类可解释的,并且非常容易受到小的对抗性扰动-以及这些特征如何影响测试性能。我们发现,仅在非鲁棒特征上训练的模型在五个MedMNIST分类任务中实现了远高于机会的准确性,证实了它们的分布预测值。相反,主要依赖于鲁棒特征的逆向训练模型牺牲了分布内的准确性,但在受控分布变化下产生了明显更好的性能(MedMNIST-C)。总体而言,非鲁棒性功能提高了标准的准确性,但降低了分布性能,揭示了一个实际的鲁棒性-准确性权衡医学成像分类任务,应根据部署设置的要求。
摘要:We study whether deep networks for medical imaging learn useful nonrobust features - predictive input patterns that are not human interpretable and highly susceptible to small adversarial perturbations - and how these features impact test performance. We show that models trained only on nonrobust features achieve well above chance accuracy across five MedMNIST classification tasks, confirming their predictive value in-distribution. Conversely, adversarially trained models that primarily rely on robust features sacrifice in-distribution accuracy but yield markedly better performance under controlled distribution shifts (MedMNIST-C). Overall, nonrobust features boost standard accuracy yet degrade out-of-distribution performance, revealing a practical robustness-accuracy trade-off in medical imaging classification tasks that should be tailored to the requirements of the deployment setting.

【4】Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction?
标题:Natural-Area Foundation模型对于加速心脏MRI重建有效吗?
链接:https://arxiv.org/abs/2604.22557

作者:Anam Hashmi,Mayug Maniparambil,Julia Dietlmeier,Kathleen M. Curran,Noel E. O'Connor
备注:Accepted to CVPRW 2026
摘要:大规模预训练基础模型的出现改变了计算机视觉,使其在各种下游任务中具有强大的性能。然而,他们的潜在的物理为基础的逆问题,如加速心脏MRI重建,在很大程度上仍有待探索。在这项工作中,我们调查自然域基础模型是否可以作为加速心脏MRI重建的有效图像先验,并比较与特定领域的同行,如BiomedCLIP获得的性能。我们提出了一个展开的重建框架,该框架在每个级联中包含预训练的冻结视觉编码器,如CLIP,DINOv 2和BiomedCLIP,以指导重建过程。通过大量的实验,我们表明,虽然特定于任务的最先进的重建模型,如E2 E-VarNet在标准的分布设置中实现了卓越的性能,但基于基础模型的方法仍然具有竞争力。更重要的是,在具有挑战性的跨领域场景中,模型在心脏MRI上进行训练,并在解剖学上不同的膝盖和大脑数据集上进行评估-基础模型表现出更好的鲁棒性,特别是在高加速因子和有限的低频采样下。我们进一步观察到,自然图像预训练模型,如CLIP,学习高度可转移的结构表示,而特定领域的预训练(BiomedCLIP)在更不适定的机制中提供了适度的额外增益。总的来说,我们的研究结果表明,预训练的基础模型提供了一个有前途的可转移先验资源,能够提高加速MRI重建的鲁棒性和泛化能力。
摘要:The emergence of large-scale pretrained foundation models has transformed computer vision, enabling strong performance across diverse downstream tasks. However, their potential for physics-based inverse problems, such as accelerated cardiac MRI reconstruction, remains largely underexplored. In this work, we investigate whether natural-domain foundation models can serve as effective image priors for accelerated cardiac MRI reconstruction, and compare the performance obtained against domain-specific counterparts such as BiomedCLIP. We propose an unrolled reconstruction framework that incorporates pretrained, frozen visual encoders, such as CLIP, DINOv2, and BiomedCLIP, within each cascade to guide the reconstruction process. Through extensive experiments, we show that while task-specific state-of-the-art reconstruction models such as E2E-VarNet achieve superior performance in standard in-distribution settings, foundation-model-based approaches remain competitive. More importantly, in challenging cross-domain scenarios, where models are trained on cardiac MRI and evaluated on anatomically distinct knee and brain datasets--foundation models exhibit improved robustness, particularly under high acceleration factors and limited low-frequency sampling. We further observe that natural-image-pretrained models, such as CLIP, learn highly transferable structural representations, while domain-specific pretraining (BiomedCLIP) provides modest additional gains in more ill-posed regimes. Overall, our results suggest that pretrained foundation models offer a promising source of transferable priors, enabling improved robustness and generalization in accelerated MRI reconstruction.

聚类(3篇)

【1】From Local to Cluster: A Unified Framework for Causal Discovery with Latent Variables
标题:从本地到集群:具有潜在变量的因果发现统一框架
链接:https://arxiv.org/abs/2604.22416

作者:Zongyu Li
摘要:潜在变量对因果发现和推理提出了根本性的挑战。传统的局部方法专注于直接邻居,但无法提供宏观层面的见解。聚类层次的方法使宏观因果推理,但要么假设集群是已知的先验或需要因果充分。此外,直接将单变量因果发现方法应用于集群级问题违反了因果充分性,并导致不正确的结果。为了克服这些局限性,本文提出了L2C(本地集群因果抽象),一个统一的框架,桥梁本地结构学习和集群级因果发现。不像以前的工作,需要一个完整的手动分配的微观变量的集群,L2C发现分区自动从本地因果模式。我们的解决方案利用集群约简定理将任何集群减少到最多三个节点而不会丢失因果信息,应用本地因果发现来识别潜在变量存在的直接原因,影响和V结构,并通过集群级演算对学习的集群图进行宏观层次的因果推理。L2C不假设因果充分性,因为潜变量通过本地发现来处理。理论分析表明,L2C保证了可靠性,原子完整性和计算效率。对合成和真实世界数据的广泛实验表明,与现有基线相比,L2C可以准确地恢复地面真实聚类,并实现更好的宏观因果效应识别。
摘要 :Latent variables pose a fundamental challenge to causal discovery and inference. Conventional local methods focus on direct neighbors but fail to provide macro level insights. Cluster level methods enable macro causal reasoning but either assume clusters are known a priori or require causal sufficiency. Moreover, directly applying single variable causal discovery methods to cluster level problems violates causal sufficiency and leads to incorrect results. To overcome these limitations, this paper proposes L2C (Local to Cluster Causal Abstraction), a unified framework that bridges local structure learning and cluster level causal discovery. Unlike prior work that requires a complete manual assignment of micro variables to clusters, L2C discovers the partition automatically from local causal patterns. Our solution leverages a cluster reduction theorem to reduce any cluster to at most three nodes without loss of causal information, applies local causal discovery to identify direct causes, effects, and V structures in the presence of latent variables, and performs macro level causal inference via cluster level calculus on the learned cluster graph. L2C does not assume causal sufficiency, as latent variables are handled through local discovery. Theoretical analysis shows that L2C ensures soundness, atomic completeness, and computational efficiency. Extensive experiments on synthetic and real world data demonstrate that L2C accurately recovers ground truth clusters and achieves superior macro causal effect identification compared to existing baselines.

【2】Robust Fuzzy local k-plane clustering with mixture distance of hinge loss and L1 norm
标题:具有铰链损失和L1模混合距离的鲁棒模糊局部k平面聚集
链接:https://arxiv.org/abs/2604.22405

作者:Junjun Huang,Xiliang Lu,Xuelin Xie,Jerry Zhijian Yang
摘要:K平面聚类(KPC)、超平面聚类和混合回归本质上都属于同一类问题。这个问题可以被概念化为相对高维的K子空间或K线性流形中的聚类。传统的KPC或模糊KPC模型表现出对离群值的明显敏感性,因为它们假定数据点与平面法向量之间的投影距离符合L2距离。同时,无限扩展集群的假设会对集群性能产生不利影响。针对这些问题,提出了一种新的鲁棒模糊局部k平面聚类(RFLkPC)方法,该方法结合了铰链损失和L1范数的混合距离。RFLkPC模型假设每个平面簇都被限制在一个有限的区域内,可以灵活而鲁棒地处理有离群点或无离群点的平面聚类任务。给出了相应的RFLkPC模型和优化算法。通过与相关模型的比较,在模拟数据和实际数据上验证了RFLkPC的有效性。所提出的RFLkPC方法的源代码可在https://github.com/xuelin-xie/RFLkPC上公开获得。
摘要:K-plane clustering (KPC), hyperplane clustering, and mixture regression all essentially fall within the same class of problems. This problem can be conceptualized as clustering in relatively high-dimensional K subspaces or K linear manifolds. Traditional KPC or fuzzy KPC models demonstrate a pronounced susceptibility to outliers, as they presuppose that the projection distance between data points and the plane normal vector adheres to the L2 distance. Meanwhile, the assumption of infinitely extending clusters adversely affects clustering performance. To solve these problems, this paper proposed a new robust fuzzy local k-plane clustering (RFLkPC) method that combines the mixture distance of hinge loss and L1 norm. The RFLkPC model assumes that each plane cluster is bounded to a finite area, which can flexibly and robustly handle plane clustering tasks with outliers or not. The corresponding model and optimization algorithms of RFLkPC were provided. Compared to other related models on this topic, a large number of experiments verify the efficiency of RFLkPC on simulated data and real data. The source code for the proposed RFLkPC method is publicly available at https://github.com/xuelin-xie/RFLkPC.

【3】Assessing the impact of dimensionality reduction on clustering performance -- a systematic study
标题:评估降维对集群性能的影响--一项系统性研究
链接:https://arxiv.org/abs/2604.22099

作者:Ousmane Assani Amate,Mohammadreza Bakhtyari,Émilie Roy,Vladimir Makarenkov
摘要:高维数据聚类的关键预处理步骤之一是降低高维数据的相似性,但对不同方法和数据类型的影响的综合评估仍然有限。在这项研究中,我们系统地评估了五种降维技术-主成分分析(PCA),核主成分分析(Kernel Principal Component Analysis)(核PCA),变分自动编码器(VAE),等距映射(Isomap)和多维缩放(MDS)-四种流行的聚类算法的性能-k均值,凝聚层次聚类(AHC),高斯混合模型(GMM),排序点以识别聚类结构(OPTICS)。我们使用调整后的兰德指数(ARI)来评估聚类质量,在文献中推荐的不同降维水平(即,k-1,其中k是簇的数目,以及原始维数的25%和50%)。我们的研究结果强调的重要性,一个仔细选择的降维技术和降维水平,应量身定制的内在数据的几何形状和聚类算法正在考虑。
摘要:Dimensionality reduction is a critical preprocessing step for clustering high-dimensional data, yet comprehensive evaluation of its impact across diverse methods and data types remains limited. In this study, we systematically assess the influence of five dimensionality reduction techniques - Principal Component Analysis (PCA), Kernel Principal Component Analysis (Kernel PCA), Variational Autoencoder (VAE), Isometric Mapping (Isomap), and Multidimensional Scaling (MDS) - on the performance of four popular clustering algorithms - k-means, Agglomerative Hierarchical Clustering (AHC), Gaussian Mixture Models (GMM), and Ordering Points to Identify the Clustering Structure (OPTICS). We evaluate clustering quality using the Adjusted Rand Index (ARI), comparing results without and with dimensionality reduction at different reduction levels recommended in the literature (i.e., k-1, where k is the number of clusters, and 25% and 50% of the original number of dimensions). Our findings underscore the importance of a careful selection of the dimensionality reduction technique and the dimensionality reduction level that should be tailored to intrinsic data geometry and clustering algorithms under consideration.

联邦学习|隐私保护|加密(2篇)

【1】Data-Free Contribution Estimation in Federated Learning using Gradient von Neumann Entropy
标题:使用梯度冯诺伊曼熵的联邦学习中的无数据贡献估计
链接:https://arxiv.org/abs/2604.22562

作者:Asim Ukaye,Mubarak Abdu-Aguye,Nurbek Tastan,Karthik Nandakumar
备注:10 pages, 4 figures, 4 pages Appendix, 6 figures in Appendix. To appear in CVPR 2026 FedVision Workshop
摘要:联合学习中的客户贡献评估对于识别客户的重要性和提供公平的奖励是必要的。目前的方法通常依赖于服务器端验证数据或自我报告的客户端信息,这可能会损害隐私或容易被操纵。我们引入了一个无数据信号的基础上的矩阵冯诺依曼(谱)熵的最后一层更新,它衡量的多样性的信息贡献。我们实例化两个实际的计划:(i)SpectralFed,它使用归一化熵作为聚合权重,和(ii)Spectralfed,它通过秩自适应卡尔曼滤波器将熵与特定类别的对齐融合,以实现每轮稳定性。在CIFAR-10/100和自然分区的FEMNIST和FedISIC基准测试中,熵衍生的分数在各种非IID制度下与独立客户端准确性保持高度相关性-没有验证数据或客户端元数据。我们将我们的结果与无数据贡献估计基线进行比较,并表明谱熵是客户贡献的有用指标。
摘要:Client contribution estimation in Federated Learning is necessary for identifying clients' importance and for providing fair rewards. Current methods often rely on server-side validation data or self-reported client information, which can compromise privacy or be susceptible to manipulation. We introduce a data-free signal based on the matrix von Neumann (spectral) entropy of the final-layer updates, which measures the diversity of the information contributed. We instantiate two practical schemes: (i) SpectralFed, which uses normalized entropy as aggregation weights, and (ii) SpectralFuse, which fuses entropy with class-specific alignment via a rank-adaptive Kalman filter for per-round stability. Across CIFAR-10/100 and the naturally partitioned FEMNIST and FedISIC benchmarks, entropy-derived scores show a consistently high correlation with standalone client accuracy under diverse non-IID regimes - without validation data or client metadata. We compare our results with data-free contribution estimation baselines and show that spectral entropy serves as a useful indicator of client contribution.

【2】FedSPDnet: Geometry-Aware Federated Deep Learning with SPDnet
标题:FedSPDnet:使用SPDnet的几何感知联合深度学习
链接:https://arxiv.org/abs/2604.22494

作者:Thibault Pautrel,Florent Bouchard,Ammar Mian,Guillaume Ginolhac
摘要:我们介绍了两个联邦学习框架的经典SPDnet模型上的对称正定(SPD)矩阵与Stiefel约束参数。不像标准的欧几里德平均,违反正交性,我们的方法保留几何结构,通过两个有效的聚合策略:ProjAvg,投影算术平均到Stiefel流形,和RLAvg,近似通过收缩和提升切空间平均。这两种方法都是计算高效的,独立于优化器,并使可扩展的联邦学习的信号处理应用程序的功能是SPD矩阵。对EEG运动想象基准的模拟表明,FedSPDnet在F1得分和对联邦和部分参与的鲁棒性方面优于联邦EEGnet,同时每轮通信使用较少的参数。
摘要:We introduce two federated learning frameworks for the classical SPDnet model operating on symmetric positive definite (SPD) matrices with Stiefel-constrained parameters. Unlike standard Euclidean averaging, which violates orthogonality, our approach preserves geometric structure through two efficient aggregation strategies: ProjAvg, projecting arithmetic means onto the Stiefel manifold, and RLAvg, approximating tangent-space averaging via retractions and liftings. Both methods are computationally efficient, independent of the optimizer, and enable scalable federated learning for signal processing applications whose features are SPD matrices. Simulations on EEG motor imagery benchmarks show that FedSPDnet outperforms federated EEGnet in F1 score and robustness to federation and partial participation, while using fewer parameters per communication round.

推理|分析|理解|解释(10篇)

【1】SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference
标题:SpikingBrain2.0:用于高效长上下文和跨平台推理的大脑启发基础模型
链接:https://arxiv.org/abs/2604.22575

作者:Yuqi Pan,Jinghao Zhuang,Yupeng Feng,Fangzhi Zhong,Siyu Ding,Xuerui Qiu,Shaowei Gu,Bohan Sun,Zhiyong Qin,Yibo Zhong,Lingtao Ouyang,Kun Yang,Zehao Liu,Yuhong Chou,Shurong Wang,Anjie Hu,Han Xu,Bo Xu,Guoqi Li
摘要:缩放上下文长度正在重塑大型模型开发,但全注意力Transformers在长序列时会遇到计算和推理瓶颈。一个关键的挑战是设计基础模型,以最小的训练开销保持性能和长上下文效率。我们介绍了SpikingBrain2.0(SpB2.0),这是一个5 B模型,它的架构和训练效率都比它的前身有所提高。 我们的贡献是双重的。(1)建筑创新:我们提出了双空间稀疏注意力(DSSA),这是稀疏Softmax注意力(MoBA)和稀疏线性注意力(SSE)的层间混合,实现了长上下文建模的性能效率权衡。SpB2.0进一步支持双量化路径:INT 8-Spiking编码支持稀疏事件驱动计算,而FP 8编码则加速了现代GPU上的推理。(2)强化培训战略:我们开发了一个优化的Transformer-to-Hybrid(T2 H)管道,该管道使用策展的开源数据为LLM和VLM提供双转换路径。 根据经验,SpB 2.0 - 5 B和SpB 2.0-VL-5 B在7 k A100 GPU小时内恢复了大部分基本Transformer(Qwen 3 - 4 B)功能。SpB 2.0在4 M上下文下实现了10.13倍的TTFT加速,并在vLLM下在8个A100 GPU上支持超过1000万个令牌,其中全注意力模型超过内存限制。它还表现出强大的跨平台兼容性,支持FP 8 GPU推理(250 k时加速2.52倍)和高效的神经形态执行(稀疏性64.31%,500 MHz时面积和功耗分别减少70.6%和46.5%)。 总的来说,SpikingBrain2.0为轻量级、多模式、尖峰基础模型提供了一条实用的途径,突出了将大脑启发机制与资源受限和边缘场景的高效架构相结合的潜力。
摘要:Scaling context length is reshaping large-model development, yet full-attention Transformers suffer from prohibitive computation and inference bottlenecks at long sequences. A key challenge is to design foundation models that maintain performance and long-context efficiency with minimal training overhead. We introduce SpikingBrain2.0 (SpB2.0), a 5B model that advances both architecture and training efficiency of its predecessor. Our contributions are two-fold. (1) Architectural Innovation: We propose Dual-Space Sparse Attention (DSSA), an inter-layer hybrid of Sparse Softmax Attention (MoBA) and Sparse Linear Attention (SSE), achieving an improved performance-efficiency trade-off for long-context modeling. SpB2.0 further supports dual quantization paths: INT8-Spiking coding enables sparse event-driven computation, while FP8 coding accelerates inference on modern GPUs. (2) Enhanced Training Strategy: We develop an optimized Transformer-to-Hybrid (T2H) pipeline with dual conversion paths for LLMs and VLMs using curated open-source data. Empirically, SpB2.0-5B and SpB2.0-VL-5B recover most of the base Transformer (Qwen3-4B) capability with under 7k A100 GPU hours. SpB2.0 achieves a 10.13x TTFT speedup at 4M context and supports over 10M tokens on 8 A100 GPUs under vLLM, where full-attention models exceed memory limits. It also demonstrates strong cross-platform compatibility, enabling FP8 GPU inference (2.52x speedup at 250k) and efficient neuromorphic execution (64.31% sparsity, with 70.6% and 46.5% area and power reduction at 500MHz). Overall, SpikingBrain2.0 provides a practical pathway for lightweight, multimodal, spiking foundation models, highlighting the potential of combining brain-inspired mechanisms with efficient architectures for resource-constrained and edge scenarios.

【2】An Integrated Framework for Explainable, Fair, and Observable Hospital Readmission Prediction: Development and Validation on MIMIC-IV
标题:可解释、公平和可观察的医院再入院预测的集成框架:MIIC-IV的开发和验证
链接:https://arxiv.org/abs/2604.22535

作者:Isaac Tosin Adisa
备注:22 pages, 8 figures. Submitted to the Journal of the American Medical Informatics Association (JAMIA), currently under review
摘要:目的:提出并回顾性验证一个综合框架,解决再入院预测临床翻译的三个障碍:缺乏可解释性、缺乏部署可靠性基础设施和人口统计公平性评估不足。 材料与方法:我们从MIMIC-IV数据库(30天再入院率18.0%)中构建了一个包含415231名成人住院患者的队列,分为70/15/15。Logistic回归,XGBoost和LightGBM模型在26个特征上进行了训练。SHAP提供了每例患者的解释。使用AUC-ROC、假阴性率(FNR)和阳性预测值(PPV)评估16个亚组的公平性。使用Brier评分和校准曲线评估校准。 结果:XGBoost达到AUC-ROC 0.696(95% CI 0.691-0.701),优于或匹配LACE基线(AUC 0.60-0.68)。LightGBM实现了最佳校准(Brier 0.146)。先前的入院是主要的预测因素。所有亚组均符合公平阈值(Δ AUC <= 0.05,Δ FNR <= 0.10)。 结论:该框架提供了有竞争力的性能,临床可行的解释,和强大的人口公平性。代码可在https://github.com/Tomisin92/readmission-prediction上公开获取。
摘要:Objective: To propose and retrospectively validate an integrated framework addressing three barriers to clinical translation of readmission prediction: lack of explainability, absence of deployment reliability infrastructure, and inadequate demographic fairness evaluation. Materials and Methods: We constructed a cohort of 415231 adult admissions from the MIMIC-IV database (30-day readmission prevalence 18.0%), split 70/15/15. Logistic regression, XGBoost, and LightGBM models were trained on 26 features. SHAP provided per-patient explanations. Fairness was evaluated across 16 subgroups using AUC-ROC, false negative rate (FNR), and positive predictive value (PPV). Calibration was assessed using Brier scores and calibration curves. Results: XGBoost achieved AUC-ROC 0.696 (95% CI 0.691-0.701), outperforming or matching the LACE baseline (AUC 0.60-0.68). LightGBM achieved best calibration (Brier 0.146). Prior admissions were the dominant predictor. All subgroups met equity thresholds (delta AUC <= 0.05, delta FNR <= 0.10). Conclusion: This framework delivers competitive performance, clinically actionable explanations, and strong demographic equity. Code is publicly available at https://github.com/Tomisin92/readmission-prediction.

【3】Beyond Land Surface Temperature: Explainable Spatial Machine Learning Reveals Urban Morphology Effects on Human-Centric Heat Stress
标题:超越地表温度:可解释的空间机器学习揭示城市形态对以人为中心的热应激的影响
链接:https://arxiv.org/abs/2604.22433

作者:Yuan Wang,Shengao Yi,Xiaojiang Li,Pengyuan Liu,Zhiwei Yang,Ronita Bardhan,Rudi Stouffs
摘要:热暴露将建筑环境和公共健康联系起来,直接塑造城市地区的宜居性和可持续性。了解热暴露及其驱动因素的空间异质性对于气候适应性城市规划至关重要。然而,大多数以规划为导向的研究依赖于地表温度(LST),LST是否足以代表人类的热暴露以及它与生理相关的热应激有何不同仍然没有得到充分的研究。在这里,采用Landsat检索的30-m LST和GPU加速的1-m通用热气候指数(UTCI)在新加坡,本研究建立了一个全面的“建模-比较-评估”的框架,系统地评估两个指标之间的空间和机制的差异。我们进一步研究了显着的非平稳和基于阈值的定量关系的两个指标与城市因素,采用一种新的地理加权XGBoost(GW-XGBoost)和广义加性模型(GAM)的工作流程。我们的研究结果表明,LST和UTCI的空间模式存在显着差异,二维和三维城市因素如何影响这两个热指标存在显著的空间异质性,正如可解释的GW-XGBoost模型所揭示的那样(LST的全球外袋R2 = 0.855,UTCI为0.905)。至关重要的是,空间显式SHAP解释天空视图因子在解释UTCI变化中起着核心作用,但对LST的贡献相对较小,表明LST无法充分捕获阴影驱动和辐射过程,这些过程控制着实际的人类热应力。值得注意的是,SHAP-GAM分析表明,较高的HDL-DO与增加的UTCI相关。这些新的发现为整合生理相关的热指数提供了证据,以告知有针对性的热风险管理和气候适应性城市规划。
摘要:Heat exposure connects the built environment and public health, directly shaping the livability and sustainability of urban areas. Understanding the spatial heterogeneity of heat exposure and its drivers is vital for climate-adaptive urban planning. However, most planning-oriented studies rely on land surface temperature (LST), and whether LST adequately represents human heat exposure and how it differs from physiologically relevant heat stress remains insufficiently examined. Here, adopting Landsat-retrieved 30-m LST and GPU-accelerated 1-m universal thermal climate index (UTCI) in Singapore, this study establishes a comprehensive "Modeling-Comparing-Assessing" framework to systematically evaluate the spatial and mechanistic discrepancies between the two metrics. We further investigate pronounced non-stationary and threshold-based quantitative relationships of the two metrics with urban factors by employing a novel geographically weighted XGBoost (GW-XGBoost) and generalized additive model (GAM) workflow. Our results demonstrate notable discrepancies in spatial patterns of LST and UTCI, along with substantial spatial heterogeneity in how 2D and 3D urban factors impact these two thermal metrics, as revealed by explainable GW-XGBoost models (global out-of-bag R2 = 0.855 for LST and 0.905 for UTCI, respectively). Crucially, spatially explicit SHAP interprets that sky view factor plays a central role in explaining UTCI variability but exhibits a comparatively marginal independent contribution to LST, indicating that LST inadequately captures shading-driven and radiative processes governing actual human heat stress. Notably, SHAP-GAM analysis indicates that higher albedo is associated with increased UTCI. These novel findings provide evidence for integrating physiologically relevant thermal indices to inform targeted heat risk management and climate-adaptive urban planning.

【4】HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
标题:HGQ-LU:DNN推理的快速LU感知训练和高效架构
链接:https://arxiv.org/abs/2604.22293

作者:Chang Sun,Zhiqiang Que,Bakhtiar Zadeh,Qibin Liu,Kevin H. Alvarez,Wayne Luk,Maria Spiropulu
摘要 :基于查找表(LUT)的神经网络可以通过将算术运算直接映射到逻辑原语上,在FPGA上提供超低延迟和出色的硬件效率。然而,最先进的LUT感知训练(LAT)方法仍然难以在实践中使用:它们通常比传统网络慢几个数量级的训练,需要针对硬件效率进行非平凡的手动调整,并且缺乏端到端工作流。这项工作提出了HGQ-LUT,集成在https://github.com/calad0i/HGQ2中,这是一种新的LAT方法,可以实现最先进的硬件效率,同时在现代GPU上将训练速度加快100倍以上。HGQ-LUT引入了LUT-Dense和LUT-Conv层,这些层在训练期间使用常规的加速器高效张量操作来实现,然后将其编译为硬件的逻辑LUT。通过将这些层与细粒度、逐元素异构量化(包括零比特修剪)和LUT感知资源代理相结合,HGQ-LUT能够自动探索精度-资源权衡,而无需手动位宽调整。我们进一步将HGQ-LUT集成到开源工具链中,实现了混合LUT与传统算术块的混合架构的统一设计,编译和位精确验证。这些功能使得基于LAT的DNN在现实世界中的部署变得实用,例如在CERN大型强子对撞机的实验中。
摘要:Lookup-table (LUT) based neural networks can deliver ultra-low latency and excellent hardware efficiency on FPGAs by mapping arithmetic operations directly onto the logic primitives. However, state-of-the-art LUT-aware training (LAT) approaches remain difficult to use in practice: they are often orders of magnitude slower to train than conventional networks, require non-trivial manual tuning for hardware efficiency, and lack an end-to-end workflow. This work presents HGQ-LUT, integrated in https://github.com/calad0i/HGQ2, a new LAT approach that achieves state-of-the-art hardware efficiency while accelerating training by over 100 times on modern GPUs. HGQ-LUT introduces LUT-Dense and LUT-Conv layers that are implemented with regular, accelerator-efficient tensor operations during training, which are then compiled into logic LUTs for hardware. By combining these layers with fine-grained, element-wise heterogeneous quantization (including zero-bit pruning) and a LUT-aware resource surrogate, HGQ-LUT enables the automatic exploration of accuracy-resource trade-offs without manual bit-width tuning. We further integrate HGQ-LUT into open-source toolchains, enabling unified design, compilation, and bit-exact verification of hybrid architectures that mix LUT-based with conventional arithmetic blocks. These features make LAT-based DNNs practical for real-world deployment, such as at the CERN Large Hadron Collider's experiments.

【5】Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems
标题:主权统计循环:将人工智能推理与现实世界系统中的执行脱钩
链接:https://arxiv.org/abs/2604.22136

作者:Jun He,Deying Yu
备注:15 pages, 2 figures
摘要:大型语言模型(LLM)代理越来越多地发出改变真实系统的API调用,但许多当前架构将随机模型输出直接传递到执行层。我们认为,这种耦合创建了一个安全风险,因为模型的正确性,上下文感知,并对齐不能在执行时假设。我们引入主权的抽象循环(SAL),控制平面架构中的模型发出结构化的意图与理由,和控制平面验证这些意图对真正的系统状态和政策执行之前。SAL结合了一个模糊化膜,它限制了模型对身份敏感状态的访问,并结合了一个加密链接的证据链,以实现可验证性和重放。我们正式SAL,并表明,在规定的假设下,它提供了政策约束的执行,身份隔离,和确定性重放。在云基础设施的OpenKedge原型中,SAL在策略层阻止了93%的不安全意图,通过一致性检查拒绝了剩余的7%,在我们的基准测试中阻止了不安全的执行,并增加了12.4 ms的中值延迟。
摘要:Large language model (LLM) agents increasingly issue API calls that mutate real systems, yet many current architectures pass stochastic model outputs directly to execution layers. We argue that this coupling creates a safety risk because model correctness, context awareness, and alignment cannot be assumed at execution time. We introduce Sovereign Agentic Loops (SAL), a control-plane architecture in which models emit structured intents with justifications, and the control plane validates those intents against true system state and policy before execution. SAL combines an obfuscation membrane, which limits model access to identity-sensitive state, with a cryptographically linked Evidence Chain for auditability and replay. We formalize SAL and show that, under the stated assumptions, it provides policy-bounded execution, identity isolation, and deterministic replay. In an OpenKedge prototype for cloud infrastructure, SAL blocks 93% of unsafe intents at the policy layer, rejects the remaining 7% via consistency checks, prevents unsafe executions in our benchmark, and adds 12.4 ms median latency.

【6】Who Audits the Auditor? Tamper-Proof Fraud Detection with Blockchain-Anchored Explainable ML
标题:谁来审计审计师?使用区块链锚定的可解释ML进行防篡改欺诈检测
链接:https://arxiv.org/abs/2604.22096

作者:Zhaohui Wang
备注:Accepted to IEEE COMPSAC 2026 (Paper ID 9376, SEPT Symposium). This is the de-anonymized camera-ready version. Code is available at: https://github.com/GeoffreyWang1117/fraud-detection-chain
摘要:在企业欺诈检测中,当内部人员可以篡改审计日志或绕过审批工作流时,仅凭模型准确性是不够的。现实世界的事件表明,欺诈行为经常持续存在,不是因为检测算法失败,而是因为审计跟踪本身可以由特权操作员控制。这暴露了一个根本的信任缺口:谁来审计审计师? 我们提出了一个防篡改的欺诈检测系统,该系统将ML预测和工作流执行锚定到不可变的区块链分类账。我们不使用区块链作为被动存储,而是通过智能合约执行整个审批流程,确保每一笔交易、预测和解释都被原子化记录,并且不能被追溯修改。我们的检测模块实现了具有竞争力的准确性(F1 = 0.895,PR-AUC = 0.974),同时提供了支持监管可验证性要求的加密可验证决策跟踪(例如,GDPR第22条)。系统评估显示,推理延迟低于25 ms,在第2层网络上部署经济可行,每笔交易低于0.01美元(根据PolygonScan数据验证),支持每月支付10,000多笔的企业级工作负载。
摘要:In enterprise fraud detection, model accuracy alone is insufficient when insiders can tamper with audit logs or bypass approval workflows. Real-world incidents show that fraud often persists not because detection algorithms fail, but because the audit trail itself is controllable by privileged operators. This exposes a fundamental trust gap: *who audits the auditor?* We present a tamper-evident fraud detection system that anchors both ML predictions and workflow execution to an immutable blockchain ledger. Rather than using blockchain as passive storage, we enforce the entire approval process through smart contracts, ensuring that every transaction, prediction, and explanation is atomically recorded and cannot be retroactively modified. Our detection module achieves competitive accuracy (F1 = 0.895, PR-AUC = 0.974) while providing cryptographically verifiable decision trails that support regulatory auditability requirements (e.g., GDPR Article 22). System evaluation shows sub-25 ms inference latency and economically viable deployment on Layer-2 networks at under \$0.01 per transaction (validated against PolygonScan data), supporting enterprise-scale workloads of 10,000+ monthly payments.

【7】Math Takes Two: A test for emergent mathematical reasoning in communication
标题:数学需要两个:沟通中紧急数学推理的测试
链接:https://arxiv.org/abs/2604.21935

作者:Michael Cooper,Samuel Cooper
备注:Accepted at HCAIR workshop, ICLR 2026
摘要:虽然语言模型在数学基准上表现出了非凡的能力,但目前还不清楚这是否反映了真正的数学推理或统计模式匹配。大多数现有的评估依赖于基于既定数学惯例的符号问题,限制了对模型从第一原理构建抽象概念的能力的洞察。在这项工作中,我们提出了数学需要两个,一个新的基准,旨在评估数学推理的出现,通过沟通。出于人类的数学认知与精确沟通的需要共同进化的假设,我们的基准测试是否两个代理,没有事先的数学知识,可以开发一个共享的符号协议来解决一个视觉接地的任务,使用数值系统方便外推。与许多当前的数据集不同,我们的基准避开了预定义的数学语言,而是要求代理从头开始发现潜在的结构和表示。因此,Math Takes Two提供了一个新的镜头,通过它来开发和评估具有紧急数值推理能力的模型。
摘要:Although language models demonstrate remarkable proficiency on mathematical benchmarks, it remains unclear whether this reflects true mathematical reasoning or statistical pattern matching over learning formal syntax. Most existing evaluations rely on symbolic problems grounded in established mathematical conventions, limiting insight into the models' ability to construct abstract concepts from first principles. In this work, we propose Math Takes Two, a new benchmark designed to assess the emergence of mathematical reasoning through communication. Motivated by the hypothesis that mathematical cognition in humans co-evolved with the need for precise communication, our benchmark tests whether two agents, without prior mathematical knowledge, can develop a shared symbolic protocol to solve a visually grounded task where the use of a numerical system facilitates extrapolation. Unlike many current datasets, our benchmark eschews predefined mathematical language, instead requiring agents to discover latent structure and representations from scratch. Math Takes Two thus provides a novel lens through which to develop and evaluate models with emergent numerical reasoning capabilities.

【8】Analyzing Shapley Additive Explanations to Understand Anomaly Detection Algorithm Behaviors and Their Complementarity
标题:分析Shapley加法解释以了解异常检测算法行为及其互补性
链接:https://arxiv.org/abs/2602.00208

作者:Jordan Levy,Paul Saves,Moncef Garouani,Nicolas Verstaevel,Benoit Gaudou
备注:IDA Frontier Prize and Best Paper Award -Intelligent Data Analysis (IDA) 2026, Springer Nature
摘要:由于数据分布的多样性和标签的缺乏,无监督异常检测是一个具有挑战性的问题。通常采用包围方法来通过组合多个检测器来减轻这些挑战,这可以减少个体偏差并提高鲁棒性。然而,建立一个真正互补的集合仍然具有挑战性,因为许多检测器依赖于类似的决策线索,最终产生冗余的异常分数。因此,集成学习的潜力往往受到识别真正捕获不同类型的不规则性的模型的困难的限制。为了解决这个问题,我们提出了一种方法,通过他们的决策机制来表征异常检测器。使用SHapley加法解释,我们量化了每个模型如何将重要性归因于输入特征,并且我们使用这些属性配置文件来测量检测器之间的相似性。我们发现,具有类似解释的检测器往往会产生相关的异常分数,并识别出大量重叠的异常。相反,解释分歧可靠地表明互补的检测行为。我们的研究结果表明,简化驱动的指标提供了一个不同的标准比原始输出选择模型的合奏。然而,我们也表明,仅仅多样性是不够的,高的个人模型性能仍然是有效的合奏的先决条件。通过明确针对解释多样性,同时保持模型质量,我们能够构建更多样化,更互补,最终更有效的无监督异常检测的合奏。
摘要 :Unsupervised anomaly detection is a challenging problem due to the diversity of data distributions and the lack of labels. Ensemble methods are often adopted to mitigate these challenges by combining multiple detectors, which can reduce individual biases and increase robustness. Yet building an ensemble that is genuinely complementary remains challenging, since many detectors rely on similar decision cues and end up producing redundant anomaly scores. As a result, the potential of ensemble learning is often limited by the difficulty of identifying models that truly capture different types of irregularities. To address this, we propose a methodology for characterizing anomaly detectors through their decision mechanisms. Using SHapley Additive exPlanations, we quantify how each model attributes importance to input features, and we use these attribution profiles to measure similarity between detectors. We show that detectors with similar explanations tend to produce correlated anomaly scores and identify largely overlapping anomalies. Conversely, explanation divergence reliably indicates complementary detection behavior. Our results demonstrate that explanation-driven metrics offer a different criterion than raw outputs for selecting models in an ensemble. However, we also demonstrate that diversity alone is insufficient; high individual model performance remains a prerequisite for effective ensembles. By explicitly targeting explanation diversity while maintaining model quality, we are able to construct ensembles that are more diverse, more complementary, and ultimately more effective for unsupervised anomaly detection.

【9】Time-Localized Parametric Decomposition of Respiratory Airflow for Sub-Breath Analysis
标题:用于亚呼吸分析的呼吸气流的时间局部参数分解
链接:https://arxiv.org/abs/2604.22695

作者:Victoria Ribeiro Rodrigues,Paul W. Davenport,Nicholas J. Napoli
备注:Submitted to IEEE Journal of Biomedical and Health Informatics (under review). 18 pages, 7 figures, 5 tables
摘要:呼吸气流信号提供了对呼吸机制的关键洞察,然而传统的分析方法在表征个体呼吸的内部结构的能力方面仍然有限。传统的方法将气流视为准周期信号,并依赖于全局描述符,如潮气量或峰值流量,模糊了反映神经肌肉协调和代偿性呼吸策略的亚呼吸事件。本研究引入了一个参数框架,用于将吸气气流分解为少量具有明确幅度、起始时间和持续时间参数的时间局部化分量。与光谱或数据自适应方法不同,所提出的方法采用生理接地基函数,半正弦,高斯和β,通过约束非线性优化来表示intrabreath波形形态。对8,276次呼吸的评估表明,在中等噪声下,重建精度较高(四分量模型的均方误差<0.001美元),参数精度稳健。与经典的呼吸指标相比,描述子呼吸定时和协调的派生特征将认知呼吸竞争引起的认知疲劳状态的分类提高了高达30.7%的Matthews相关系数。这些结果表明,建模气流作为一个参数化的,时间本地化的原语的总和提供了一个可解释的和精确的基础,量化intrabreath组织,代偿性呼吸动力学,呼吸运动控制适应下的认知呼吸双任务的需求。
摘要:Respiratory airflow signals provide critical insight into breathing mechanics, yet conventional analysis methods remain limited in their ability to characterize the internal structure of individual breaths. Traditional approaches treat airflow as a quasi-periodic signal and rely on global descriptors such as tidal volume or peak flow, obscuring sub-breath events that reflect neuromuscular coordination and compensatory breathing strategies. This study introduces a parametric framework for decomposing inspiratory airflow into a small number of time-localized components with explicit amplitude, onset time, and duration parameters. Unlike spectral or data-adaptive methods, the proposed approach employs physiologically grounded basis functions, Half-Sine, Gaussian, and Beta, to represent intrabreath waveform morphology through constrained nonlinear optimization. Evaluation across 8,276 breaths demonstrates high reconstruction accuracy (mean squared error $

【10】Audio Video Verbal Analysis (AVVA) for Capturing Classroom Dialogues
标题:捕捉课堂对话的音频视频言语分析(AVVA)
链接:https://arxiv.org/abs/2604.22043

作者:Vivek Upadhyay,Amaresh Chakrabarti
备注:42 pages, 4 figures, 1 table
摘要:背景资料:课堂话语分析已经被越来越多的音频视频多模态数据的使用所改变,这需要平衡解释深度与计算可扩展性的分析方法。 研究方法:本研究介绍了音频视频言语分析(AVVA)框架,改编自言语分析方法,将定性解释与定量建模相结合。与完全多模态的学习分析方法不同,AVVA专注于具有基本intermittent模态的逐字记录。 结果:该框架嵌入三角测量作为一个核心的设计策略,在10个方法步骤,加强有效性和分析的严谨性。一个全面的验证方案解决了时间观测研究中的基本挑战:低频变量的Phi上限(通过基本速率过滤),估计不确定性(通过自举置信区间)和可修改的时间单位问题,其中测量的关联取决于观测窗口大小。四准则稳定性评估(符号一致性,置信区间重叠,零排除,幅度稳定性)将变量对分类为可解释的模式:跨时间粒度的粒度不变,尺度特定或多尺度等结构。将其应用于23小时的课堂录音说明了其实际可行性及其产生有意义的见解的潜力。 贡献:该框架因此提供了一个可扩展的途径,将丰富的课堂话语转化为可分析的数据集。
摘要:Background: The classroom discourse analysis has been transformed by the growing use of audio-video multimodal data, which demands analytical methods that balance interpretive depth with computational scalability. Methods: This study introduces the Audio Video Verbal Analysis (AVVA) framework, adapted from the Verbal Analysis method to integrate qualitative interpretation with quantitative modelling. Unlike fully multimodal learning analytics approaches, AVVA focuses on verbatim transcripts with essential interactional modalities. Findings: The framework embeds triangulation as a core design strategy across ten methodological steps, strengthening validity and analytical rigour. A comprehensive validation scheme addresses fundamental challenges in temporal observational research: Phi Ceiling for low-frequency variables (via Base Rate Filtering), estimation uncertainty (via bootstrap confidence intervals), and the Modifiable Temporal Unit Problem, where measured associations depend on observational window size. Four-criterion stability assessment (sign consistency, confidence interval overlap, zero exclusion, magnitude stability) classifies variable pairs into interpretable patterns: grain-invariant, scale-specific, or multi-scale, etc. structures across temporal grain sizes. Its application to 23 hours of classroom recordings illustrates its practical viability and its potential to yield meaningful insights. Contribution: The framework thus provides a scalable pathway for transforming rich classroom discourse into analysable datasets.

检测相关(4篇)

【1】Detecting Concept Drift in Evolving Malware Families Using Rule-Based Classifier Representations
标题:使用基于规则的分类器表示检测不断发展的恶意软件家族中的概念漂移
链接:https://arxiv.org/abs/2604.22629

作者:Tomáš Kalný,Martin Jureček,Mark Stamp
摘要:这项工作提出了一种结构化的方法来概念漂移检测恶意软件分类使用决策树规则集。分类器在EMBER 2024数据集上的时间窗口中进行训练,并通过使用特征重要性,预测一致性,激活稳定性和覆盖率指标比较提取的规则表示来量化漂移。这些指标与准确度下降和数据分布偏移相关,作为补充漂移指标。该方法在六个恶意软件家族中使用固定间隔和基于聚类的窗口在家族与良性和家族与家族设置中进行评估,并与RIPPER和超越基线进行比较。结果表明,固定两个月的窗口与功能级皮尔逊相关性是最可靠的配置,是唯一一个所有家庭对产生积极的漂移精度相关性。这些方法是互补的-没有一种方法在所有对中占主导地位。
摘要:This work proposes a structural approach to concept drift detection in malware classification using decision tree rulesets. Classifiers are trained across temporal windows on the EMBER2024 dataset, and drift is quantified by comparing extracted rule representations using feature importance, prediction agreement, activation stability, and coverage metrics. These metrics are correlated with both accuracy degradation and data distribution shift as complementary drift indicators. The approach is evaluated across six malware families using fixed-interval and clustering-based windowing in family-vs-benign and family-vs-family settings, and compared against RIPPER and Transcendent baselines. Results show that fixed two-month windowing with feature-level Pearson correlation is the most reliable configuration, being the only one where all family pairs produce positive drift-accuracy correlations. The methods are complementary - no single approach dominates across all pairs.

【2】Protect the Brain When Treating the Heart: A Convolutional Neural Network for Detecting Emboli
标题:治疗心脏时保护大脑:检测血栓的卷积神经网络
链接:https://arxiv.org/abs/2604.22258

作者:Andrea Angino,Ken Trotti,Diego Ulisse Pizzagalli,Rolf Krause,Tiziano Torre,Stefanos Demertzis
备注:Corresponding authors: Andrea Angino and Diego Ulisse Pizzagalli
摘要:气体微栓子(GME)是外科和经导管方法中心脏结构性介入治疗的常见并发症。经胸心脏超声成像是一种可视化循环GME存在的便捷方法。然而,它们的检测和量化是远远不够的,由于操作员依赖的视图,高速度,和在背景中具有相似结构的对象。在这里,我们提出了一种基于2.5D U-Net架构的方法来分割时空连接数据中的GME。这样的方法产生对背景的鲁棒检测和高分割精度,同时保持实时执行速度。这些特性促进了将所提出的管道整合到患者监测手术方案中,提供了GME面积随时间的量化。
摘要 :Gaseous microemboli (GME) represent a common complication of cardiac structural interventions across both surgical and transcatheter approaches. Transthoracic cardiac ultrasound imaging represents a convenient methodology to visualize the presence of circulating GME. However, their detection and quantification are far from trivial due to operator-dependent view, high velocity, and objects with similar structure in the background. Here, we propose an approach based on a 2.5D U-Net architecture to segment GME in space-time connected data. Such an approach yields robust detection against the background and high segmentation accuracy while retaining real-time execution speed. These properties facilitated the integration of the proposed pipeline into patient-monitoring surgical protocols, providing the quantification of GME area over time.

【3】When Quotes Crumble: Detecting Transient Mechanical Liquidity Erosion in Limit Order Books
标题:当报价崩溃时:检测限额订单簿中的暂时机械流动性侵蚀
链接:https://arxiv.org/abs/2604.21993

作者:Haohan Xu,Jason Bohne,Pawel Polak,Yurij Baransky,Ajay Alva,Violetta Fedotova,Gary Kazantsev,David Rosenberg
备注:10 pages, 4 figures. Accepted at ICLR 2026 Workshop on Advances in Financial AI
摘要:我们研究了检测瞬态流动性侵蚀(“摇摇欲坠的报价”)在电子限价订单簿,可观察到的报价恶化可能反映机械流动性撤回或信息重新定价。使用ABIDES代理为基础的模拟器,我们构建了一个多代理环境中,摇摇欲坠出现随机政权切换在做市商,提供时间分辨地面真理在真实的市场数据。我们开发了一个检测管道,使用订单簿功能识别机械驱动的报价侵蚀,并训练神经模型以产生校准的崩溃概率。实验表明,所提出的框架可以根据代理级别的地面事实可靠地识别崩溃事件,神经模型在基于规则的基线上实现了+36%的AUC改进,并且在正常,高波动性,牛市和熊市条件下都具有强大的性能。消融的时间特征和不同的依赖结构的地面实况机制的研究证实,该框架概括了独立和自相关的流动性撤回动态。
摘要:We study the detection of transient liquidity erosion ("crumbling quotes") in electronic limit order books, where observable quote deterioration may reflect either mechanical liquidity withdrawal or informational repricing. Using the ABIDES agent-based simulator, we construct a multi-agent environment in which crumbling emerges from stochastic regime switches in a market maker, providing time-resolved ground truth unavailable in real market data. We develop a detection pipeline that identifies mechanically driven quote erosion using order book features, and train a neural model to produce calibrated crumbling probabilities. Experiments demonstrate that the proposed framework reliably identifies crumbling events against agent-level ground truth, with the neural model achieving +36% AUC improvement over rule-based baselines and robust performance across normal, high-volatility, bull, and bear market conditions. Ablation studies on temporal features and varying the dependence structure of the ground-truth mechanism confirm that the framework generalizes across both independent and autocorrelated liquidity withdrawal dynamics.

【4】Performance Anomaly Detection in Athletics: A Benchmarking System with Visual Analytics
标题:田径运动表现异常检测:具有视觉分析的基准系统
链接:https://arxiv.org/abs/2604.21953

作者:Blessed Madukoma,Prasenjit Mitra
备注:8 pages, 5 figures, 5 tables
摘要:反兴奋剂计划依靠生物测试来检测提高成绩的药物,但这种测试每个样本的成本超过800美元,并且受到许多违禁物质检测窗口短的限制。这些限制使大部分运动员没有定期测试,从而激发了分析常规比赛结果以识别可疑表现模式的补充筛选方法。我们提出了一个系统,使用从统计规则到机器学习和轨迹分析的八种检测方法,处理来自19,000多场比赛(2010-2025)的160万次田径表演。我们对所有公开确认的反兴奋剂违规行为进行验证,以衡量其在识别受制裁运动员方面的有效性。基于轨迹的方法将绩效与预期的职业发展进行比较,在检测违规行为和限制误报之间实现了最佳平衡,尽管所有方法都面临数据不完整和罕见的确认违规行为的挑战。该系统为专家驱动的调查提供了一个互动界面,强调透明度和人的判断,以支持而不是取代既定的反兴奋剂程序。
摘要:Anti-doping programs rely on biological testing to detect performance-enhancing drugs, but such testing costs over $800 per sample and is limited by short detection windows for many prohibited substances. These constraints leave large portions of athletes without regular testing, motivating complementary screening approaches that analyze routine competition results to identify suspicious performance patterns. We present a system that processes 1.6 million athletics performances from over 19,000 competitions (2010-2025) using eight detection methods ranging from statistical rules to machine learning and trajectory analysis. We validate all methods against publicly confirmed anti-doping violations to measure their effectiveness in identifying sanctioned athletes. Trajectory-based methods, which compare performances to expected career progression, achieve the best balance between detecting violations and limiting false alarms, though all methods face challenges from incomplete data and rare confirmed violations. The system provides an interactive interface for expert-driven investigation, emphasizing transparency and human judgment to support, rather than replace, established anti-doping processes.

分类|识别(2篇)

【1】Different Strokes for Different Folks: Writer Identification for Historical Arabic Manuscripts
标题:不同人物的不同笔触:阿拉伯历史手稿的作者身份
链接:https://arxiv.org/abs/2604.22515

作者:Hamza A. Abushahla,Ariel Justine N. Panopio,Layth Al-Khairulla,Mohamed I. AlHajri
备注:29 pages, 13 figures, 31 tables
摘要:手写的阿拉伯手稿保存了阿拉伯世界的知识和文化遗产,作家身份验证支持出处,真实性验证和历史分析。使用历史阿拉伯语手稿的Muharaf数据集,我们评估了单个行图像的作者识别,并尽我们所知,提供了行级和页面不相交评估协议下报告的第一个基线。由于数据集仅部分标记了作者身份,我们手动验证并将公共部分的作者标签从24,495行图像中的6,858行(28.00%)扩展到21,249行(86.75%),纠正不一致并删除非手写文本。进一步过滤后,我们保留了18,987行(77.51%)。我们提出了一个基于卷积神经网络(CNN)的模型,该模型具有用于闭集作家识别的注意力机制,包括建模为复合作家对类的罕见的两个作家行。我们对14种配置进行基准测试,并在不同的特征提取器和训练制度中进行消融。为了评估对不可见页面的泛化,页面不相交协议将每个页面的所有行分配给单个分割。在线路级协议下,经过微调的DenseNet 201具有注意力,达到99.05%的Top-1准确率,99.73%的Top-5准确率和97.44%的F1分数。在更具挑战性的页面不相交协议下,观察到的最佳结果是78.61%的Top-1准确率,87.79%的Top-5准确率和66.55%的F1分数,从而量化了页面级别线索的影响。通过扩展Muharaf数据集的标记子集并报告两个协议,我们为从事文化和历史重要文件的历史学家和语言学家提供了更清晰的基准和实用资源。代码和实现细节可以在GitHub上找到。
摘要:Handwritten Arabic manuscripts preserve the Arab world's intellectual and cultural heritage, and writer identification supports provenance, authenticity verification, and historical analysis. Using the Muharaf dataset of historical Arabic manuscripts, we evaluate writer identification from individual line images and, to the best of our knowledge, provide the first baselines reported under both line-level and page-disjoint evaluation protocols. Since the dataset is only partially labeled for writer identification, we manually verified and expanded writer labels in the public portion from 6,858 (28.00%) to 21,249 lines (86.75%) out of 24,495 line images, correcting inconsistencies and removing non-handwritten text. After further filtering, we retained 18,987 lines (77.51%). We propose a Convolutional Neural Network (CNN)-based model with attention mechanisms for closed-set writer identification, including rare two-writer lines modeled as composite writer-pair classes. We benchmark fourteen configurations and conduct ablations across different feature extractors and training regimes. To assess generalization to unseen pages, the page-disjoint protocol assigns all lines from each page to a single split. Under the line-level protocol, a fine-tuned DenseNet201 with attention achieves 99.05% Top-1 accuracy, 99.73% Top-5 accuracy, and 97.44% F1-score. Under the more challenging page-disjoint protocol, the best observed results are 78.61% Top-1 accuracy, 87.79% Top-5 accuracy, and 66.55% F1-score, thus quantifying the impact of page-level cues. By expanding the Muharaf dataset's labeled subset and reporting both protocols, we provide a clearer benchmark and a practical resource for historians and linguists engaged with culturally and historically significant documents. The code and implementation details are available on GitHub.

【2】Do Not Imitate, Reinforce: Iterative Classification via Belief Refinement
标题:不要模仿,强化:通过信念细化进行迭代分类
链接:https://arxiv.org/abs/2604.22110

作者:Mahdi Kallel,Johannes Tölle,Ahmed Hendawy,Carlo D'Eramo
摘要:标准监督分类训练模型来模仿完美预言机提供的精确标签。这种模仿发生在单次通过中,即使输入复杂度不同,也会将模型限制在固定的计算预算中。此外,严格的训练目标迫使模型在其训练数据上表达绝对的确定性,导致评估期间的过度自信预测。我们提出了强化迭代分类(RIC),它用强化学习(RL)代替了模仿目标。RIC部署了一个递归代理,迭代地更新类上的预测分布,并因预测质量的逐步提高而获得奖励。价值函数通过估计剩余的改进范围提供了一个自然的停止标准。我们证明,迭代公式恢复相同的最佳预测交叉熵,同时产生一个随时分类。在图像分类基准上,RIC将监督基线的准确性与改进的校准相匹配,并学会在输入之间自适应地分配计算。
摘要 :Standard supervised classification trains models to imitate the exact labels provided by a perfect oracle. This imitation happens in a single pass, restricting the model to a fixed compute budget even when inputs vary in complexity. Moreover, the rigid training objective forces the model to express absolute certainty on its training data, resulting in overconfident predictions during evaluation. We propose Reinforced Iterative Classification (RIC), which replaces the imitative objective with Reinforcement Learning (RL). RIC deploys a recurrent agent that iteratively updates a predictive distribution over classes, receiving reward for stepwise improvement in prediction quality. The value function provides a natural halting criterion by estimating the remaining scope for improvement. We prove that the iterative formulation recovers the same optimal predictions as cross-entropy while yielding an anytime classifier. On image classification benchmarks, RIC matches the accuracy of supervised baselines with improved calibration and learns to allocate computation adaptively across inputs.

编码器(1篇)

【1】CLVAE: A Variational Autoencoder for Long-Term Customer Revenue Forecasting
标题:CLVAE:一种用于长期客户收入预测的变分自动编码器
链接:https://arxiv.org/abs/2604.22636

作者:Jeffrey Näf,Riana Valera Mbelson,Markus Meierer
摘要:从稀疏和不规则的交易数据中预测客户的长期收入是非合同环境中营销资源分配的核心,但现有方法面临权衡。传统的概率客户群模型通过强加强有力的结构假设来提供稳健的长期预测,而灵活的机器学习模型通常需要大量的训练数据和仔细的调整。我们提出了一种基于变分自动编码器的模型,该模型保留了基于客户异质性的已建立的消耗-交易-支出模型的基于过程的可能性,但用编码器-解码器网络学习的灵活的潜在表示替换了限制性参数混合分布。由此产生的方法(i)提供了一个单一的模型,客户流失,交易和支出,(ii)仍然可靠,当上下文协变量不可用,(iii)灵活地结合丰富的协变量和非线性效应时,他们是可用的。这种设计平衡了结构稳定性和捕获复杂购买动态所需的灵活性。在多个真实世界的数据集和预测范围内,所提出的模型在最新的基准上进行了改进。企业直接受益,因为更好地评估客户的未来收入可以提高广告活动的效率。对于研究,这项工作提供了关于如何将特定领域的模型嵌入到变分自动编码器框架中的指导,从而实现灵活的表示学习,同时保留经济计量学上有意义的过程结构。
摘要:Predicting customers' long-term revenue from sparse and irregular transaction data is central to marketing resource allocation in non-contractual settings, yet existing approaches face a trade-off. Traditional probabilistic customer base models deliver robust long-horizon forecasts by imposing strong structural assumptions, while flexible machine-learning models often require substantial training data and careful tuning. We propose a variational-autoencoder-based model that preserves the process-based likelihood of established attrition-transaction-spend models conditional on customer heterogeneity, but replaces the restrictive parametric mixing distribution with a flexible latent representation learned by encoder-decoder networks. The resulting approach (i) provides a single model for customer attrition, transactions and spending, (ii) remains reliable when contextual covariates are unavailable, and (iii) flexibly incorporates rich covariates and nonlinear effects when they are available. This design balances structural stability with the flexibility needed to capture complex purchase dynamics. Across multiple real-world datasets and prediction horizons, the proposed model improves upon the latest benchmarks. Businesses benefit directly, as a better assessment of customers' future revenues improves the efficiency of campaign targeting. For research, this work provides guidance on how to embed domain-specific models into the variational autoencoder framework, enabling flexible representation learning while retaining an econometrically meaningful process structure.

优化|敛散性(4篇)

【1】Optimal sequential decision-making for error propagation mitigation in digital twins
标题:减少数字双胞胎中错误传播的最佳顺序决策
链接:https://arxiv.org/abs/2604.22168

作者:Annice Najafi,Shokoufeh Mirzaei
摘要:在这里,我们探讨的问题,在模块化的数字双胞胎作为一个顺序的决策过程中的错误传播缓解。建立在一个同伴的研究,使用隐马尔可夫模型(HMM)来推断潜在的错误制度,从代理物理残差,我们开发了一个马尔可夫决策过程(MDP),其中推断的制度作为国家,纠正措施作为行动,和标量奖励,考虑到系统保真度和维护费用之间的成本效益权衡。从HMM学习的参数中提取基线转移矩阵。然后,我们将该公式扩展到部分可观察MDP(POMDP),该MDP通过保持通过贝叶斯过滤更新的信念分布来解释政权分类的不完美性质,并将HMM混淆矩阵作为观察模型。这两种配方通过动态规划求解,并通过Gillespie随机模拟验证。然后,我们对两种无模型强化学习算法Q-学习和REINFORCE进行了基准测试,以评估在没有明确模型知识的情况下是否可以学习有效的策略。不同干预策略的系统比较表明,MDP策略在标称操作中实现了最高的累积奖励和时间分数,而POMDP在实际观测噪声下恢复了约95%的MDP性能。对观测质量、修复概率和折扣因子的敏感性分析证实了这些结论的稳健性,并且政策层次结构中的主要差距在统计学上是显著的,p < 0.001。MDP和POMDP性能之间的差距量化了信息的价值,为投资提高分类准确性提供了原则性标准。
摘要:Here, we explore the problem of error propagation mitigation in modular digital twins as a sequential decision process. Building on a companion study that used a Hidden Markov Model (HMM) to infer latent error regimes from surrogate-physics residuals, we develop a Markov Decision Process (MDP) in which the inferred regimes serve as states, corrective interventions serve as actions, and a scalar reward that takes into consideration the cost-benefit tradeoff between system fidelity and maintenance expense. The baseline transition matrix is extracted from the HMM-learned parameters. We then extend the formulation to a Partially Observable MDP (POMDP) that accounts for the imperfect nature of regime classification by maintaining a belief distribution updated via Bayesian filtering, with the HMM confusion matrix serving as the observation model. Both formulations are solved via dynamic programming and validated through Gillespie stochastic simulation. We then benchmark two model-free reinforcement learning algorithms, Q-learning and REINFORCE, to assess whether effective policies can be learned without explicit model knowledge. A systematic comparison of different intervention policies demonstrates that the MDP policy achieves the highest cumulative reward and fraction of time in nominal operation, while the POMDP recovers approximately 95\% of MDP performance under realistic observation noise. Sensitivity analyses across observation quality, repair probability, and discount factor confirm the robustness of these conclusions, and the major gaps in the policy hierarchy are statistically significant at $p < 0.001$. The gap between MDP and POMDP performance quantifies the value of information providing a principled criterion for investing in improved classification accuracy.

【2】Learning Coverage- and Power-Optimal Transmitter Placement from Building Maps: A Comparative Study of Direct and Indirect Neural Approaches
标题:根据建筑地图学习覆盖范围和功率最佳发射器放置:直接和间接神经方法的比较研究
链接:https://arxiv.org/abs/2604.22056

作者:Çağkan Yapar
摘要:最佳无线发射器放置是无线电网络规划中的中心任务,然而穷举搜索在规模上变得过于昂贵。本文研究了一个固定的学习传播代理下的单发射机设置,其中详尽的每像素评估仍然易于处理,并提供代理确切的地面真相。我们介绍了一个数据集的167,525个城市场景(RadioMapSeer部署)与双代理精确标签的覆盖范围最佳和功率最佳的发射机位置。地面实况分析揭示了一个不对称的覆盖功率权衡:覆盖最优布局牺牲13.86%的接收功率,而功率最优布局牺牲只有5.50%的覆盖范围;最佳可实现的平衡布局位于$\bar{d}=2.60$从理想点(100%,100%)。我们评估了两种学习公式:间接的基于热图的模型,预测接收功率的无线电地图,和直接的得分地图模型,预测在可行的发射机位置的客观景观。在热图系列中,判别模型提供的一次性预测比穷举搜索快1350- 2400倍,而扩散模型还支持多样本推理,提高了单目标性能,并通过在平衡标准下重用相同的样本池,恢复强平衡的位置,而无需显式多目标训练。结合功率和覆盖率得分地图的双得分地图策略匹配穷举平衡最优值($\bar{d}=2.60$),并且在候选重新评估后以14- 22倍的加速比在较小的候选预算中保持接近。这两种公式都允许非常快速的一次性推理;在这个基准上,双得分图方法对于平衡放置是最强的,而热图公式仍然具有吸引力,因为它们具有物理意义的中间图,并且在扩散设置中,用于推理时间搜索。
摘要:Optimal wireless transmitter placement is a central task in radio-network planning, yet exhaustive search becomes prohibitively expensive at scale. This paper studies the single-transmitter setting under a fixed learned propagation surrogate, where exhaustive per-pixel evaluation remains tractable and provides surrogate-exact ground truth. We introduce a dataset of 167,525 urban scenarios (RadioMapSeer-Deployment) with dual surrogate-exact labels for coverage-optimal and power-optimal transmitter locations. Ground-truth analysis reveals an asymmetric coverage-power trade-off: coverage-optimal placement sacrifices 13.86% of received power, whereas power-optimal placement sacrifices only 5.50% of coverage; the best achievable balanced placement lies at $\bar{d}=2.60$ from the ideal point (100%,100%). We evaluate two learning formulations: indirect heatmap-based models that predict received-power radio maps, and direct score-map models that predict the objective landscape over feasible transmitter locations. Within the heatmap family, discriminative models deliver one-shot predictions 1350-2400x faster than exhaustive search, while diffusion models additionally support multi-sample inference that improves single-objective performance and, by reusing the same sample pool under a balanced criterion, recovers strong balanced placements without explicit multi-objective training. Dual score-map strategies combining power and coverage score maps match the exhaustive balanced optimum ($\bar{d}=2.60$) and remain close across smaller candidate budgets, at 14-22x speedups after candidate re-evaluation. Both formulations admit very fast one-shot inference; on this benchmark, dual score-map methods are strongest for balanced placement, whereas heatmap formulations remain attractive for their physically meaningful intermediate maps and, in the diffusion setting, for inference-time search.

【3】Multi-Task Optimization over Networks of Tasks
标题:任务网络上的多任务优化
链接:https://arxiv.org/abs/2604.21991

作者:Julian Hatzky,Thomas Bartz-Beielstein,A. E. Eiben,Anil Yaman
备注:14 pages, 5 figures
摘要 :多任务优化是并行解决大量任务的有效方法。然而,现有的算法面临着明显的局限性:基于人口的方法规模很差,并且对于大型任务集仍然探索不足。能够扩展到超过一千个任务的方法大多是MAP-Elites变体,并且依赖于一个固定的离散化存档,该存档忽略了任务空间的拓扑结构。我们介绍MONET(多任务优化网络的任务),多任务优化算法,模型的任务空间作为一个图形:任务是节点,和边缘连接任务参数空间中的任务。这种表示使任务之间的知识转移,并保持易于处理的高维问题,同时利用任务空间的拓扑结构。MONET结合了社会学习和个体学习,社会学习通过交叉从相邻节点生成候选者,个体学习通过变异独立地改进节点自己的解决方案。我们在四个领域(射箭,手臂和cartpole,每个5,000个任务; hexapod,2,000个任务)上评估了MONET,并表明它在所有四个领域中匹配或超过了现有的基于MAP精英的基线的性能。
摘要:Multi-task optimization is a powerful approach for solving a large number of tasks in parallel. However, existing algorithms face distinct limitations: Population-based methods scale poorly and remain underexplored for large task sets. Approaches that do scale beyond a thousand tasks are mostly MAP-Elites variants and rely on a fixed, discretized archive that disregards the topology of the task space. We introduce MONET (Multi-Task Optimization over Networks of Tasks), a multi-task optimization algorithm that models the task space as a graph: tasks are nodes, and edges connect tasks in the task parameter space. This representation enables knowledge transfer between tasks and remains tractable for high-dimensional problems while exploiting the topology of the task space. MONET combines social learning, which generates candidates from neighboring nodes via crossover, with individual learning, which refines a node's own solution independently via mutation. We evaluate MONET on four domains (archery, arm, and cartpole with 5,000 tasks each; hexapod with 2,000 tasks) and show that it matches or exceeds the performance of existing MAP-Elites-based baselines across all four domains.

【4】Near-Optimal Regret for the Safe Learning-based Control of the Constrained Linear Quadratic Regulator
标题:约束线性二次调节器的安全学习控制的近最优遗憾
链接:https://arxiv.org/abs/2604.22158

作者:Spencer Hutchinson,Nanfei Jiang,Mahnoosh Alizadeh
摘要:研究了随机线性二次型调节器(LQR)在每个时间步都必须满足约束条件的自适应控制问题。以前的工作多维问题已经显示$\tilde{O}(T^{2/3})$遗憾和满足鲁棒约束,留下开放的问题是否$\tilde{O}(\sqrt{T})$遗憾可以达到约束LQR设置。我们有助于这个问题,显示$\tilde{O}(\sqrt{T})$遗憾和满足的机会约束。这种类型的约束使我们能够处理无界噪声,也使分析技术不直接适用于强大的约束。我们针对这个问题提出的算法使用SDP来选择一个乐观的策略,然后“缩减”这个策略,直到它是可验证安全的。我们的理论分析建立了遗憾和约束保证通过一个关键的引理,约束系统的协方差选择的政策。这种基于协方差的分析与自适应LQR中通常使用的基于成本的分析相反。
摘要:We study the problem of adaptive control of the stochastic linear quadratic regulator (LQR) with constraints that must be satisfied at every time step. Prior work on the multidimensional problem has shown $\tilde{O}(T^{2/3})$ regret and satisfaction of robust constraints, leaving open the question of whether $\tilde{O}(\sqrt{T})$ regret can be attained in the constrained LQR setting. We contribute to this problem by showing $\tilde{O}(\sqrt{T})$ regret and satisfaction of chance constraints. This type of constraints allow us to handle unbounded noise and also enable analytical techniques not directly applicable to robust constraints. Our proposed algorithm for this problem uses an SDP to select an optimistic policy, and then "scales back" this policy until it is verifiably-safe. Our theoretical analysis establishes regret and constraint guarantees via a key lemma that bounds the system covariance in terms of the chosen policy. This covariance-based analysis is in contrast with the cost-to-go based analysis that is typically used in adaptive LQR.

预测|估计(5篇)

【1】Iterative Model-Learning Scheme via Gaussian Processes for Nonlinear Model Predictive Control of (Semi-)Batch Processes
标题:用于(半)批过程非线性模型预测控制的高斯过程迭代模型学习方案
链接:https://arxiv.org/abs/2604.22672

作者:Tai Xuan Tan,Alexander Mitsos,Eike Cramer
备注:12 pages, 7 figures
摘要:间歇过程本质上是瞬时的,并且通常是非线性的,这激发了非线性模型预测控制(NMPC)。然而,采用NMPC是由成本和动态模型的不可用性的阻碍。因此,我们建议使用高斯过程(GP)的模型学习NMPC计划(GP-MLMPC)的批处理过程。我们使用来自单个初始轨迹的数据来初始化GP-MLMPC,例如,PI控制器我们迭代地应用嵌入有GP的NMPC来运行批次,并使用每次迭代的新观察结果更新GP,从而实现批量改进。使用不确定性量化的全球定位系统,我们制定机会约束,以执行安全操作所需的置信水平。我们在半间歇聚合反应器上展示了我们的方法,用于在两个小时的持续时间内跟踪和经济目标,并且反应器温度被限制在其设定点周围的范围内。经过四批迭代,跟踪误差从GP-MLMPC计划收敛到减少83\%$,相比于初始轨迹。此外,在经济目标下,与初始轨迹相比,GP-MLMPC在迭代8时导致最终产品质量增加17倍。在这两种情况下,所得到的GP-MLMPC性能与全模型NMPC相当,这表明可以通过该方法学习最优控制器。通过在最佳轨迹周围收集样本,GP-MLMPC在迭代过程中保持样本效率,并实现快速收敛。因此,建议的GP-MLMPC计划提出了一个有前途的数据有效的方法,用于控制非线性间歇过程没有机械知识。
摘要:Batch processes are inherently transient and typically nonlinear, motivating nonlinear model predictive control (NMPC). However, adopting NMPC is hindered by the cost and unavailability of dynamic models. Thus, we propose to use Gaussian Processes (GP) in a model-learning NMPC scheme (GP-MLMPC) for batch processes. We initialize the GP-MLMPC using data from a single initial trajectory, e.g., from a PI controller. We iteratively apply the NMPC embedded with GPs to run batches and update the GP with new observations from each iteration, thereby achieving batch-wise improvements. Using uncertainty quantification from the GPs, we formulate chance constraints to enforce safe operation to the required confidence levels. We demonstrate our approach in \textit{silico} on a semi-batch polymerization reactor for tracking and economic objectives over durations of two hours, and the reactor temperature is constrained in a range of $\pm2^\circ C$ around its setpoint. After only four batch iterations, tracking error from the GP-MLMPC scheme converged to a reduction of $83\%$, compared to the initial trajectory. Furthermore, under an economic objective, the GP-MLMPC resulted in a 17-fold increase in final product mass by iteration 8, compared to the initial trajectory. In both cases, the resulting GP-MLMPC performance is on par with the full-model NMPC, which shows that the optimal controller can be learned by the approach. By collecting samples around the optimal trajectory, the GP-MLMPC remains sample-efficient across iterations and achieves quick convergence. Thus, the proposed GP-MLMPC scheme presents a promising data-efficient approach for the control of nonlinear batch processes without mechanistic knowledge.

【2】FETS Benchmark: Foundation Models Outperform Dataset-specific Machine Learning in Energy Time Series Forecasting
标题:FETS基准:基础模型在能源时间序列预测中优于特定数据的机器学习
链接:https://arxiv.org/abs/2604.22328

作者:Marco Obermeier,Marco Pruckner,Florian Haselbeck,Andreas Zeiselmair
摘要:在向气候中性能源系统过渡的推动下,准确的能源时间序列预测对于规划和运营至关重要。然而,它在很大程度上仍然是一个特定于网络的任务,需要全面的训练数据,限制了可扩展性,并导致高模型开发和维护工作。最近,旨在通过广泛的预训练学习可推广模式的基础模型在多个预测任务中表现出优异的性能。尽管它们在解决能源预测挑战方面取得了成功并具有强大的潜力,但它们在这一领域的应用在很大程度上仍未得到探索。我们通过介绍能源时间序列预测(FETS)基准中的基础模型来解决这一差距。我们(1)沿着三个主要维度提供能源预测用例的结构化概述:利益相关者,属性和数据类别;(2)收集和分析9个数据类别的54个数据集,以典型的利益相关者利益为指导;(3)在不同的预测设置中,将基础模型与经典的机器学习方法进行基准测试。在所有设置和数据类别中,基础模型的性能始终优于特定于机器学习的优化机器学习方法,尽管后者在训练期间已经看到了完整的历史目标数据。特别是,协变量知情的基础模型实现了最强的性能。进一步的分析揭示了预测性能和谱熵之间的强相关性,超过一定上下文长度的性能饱和,以及在更高聚合级别(如国家负荷,区域供热和电网数据)的性能改善。总体而言,我们的研究结果突出了基础模型作为能源领域可扩展和可推广的预测解决方案的强大潜力,特别是在数据受限和隐私敏感的环境中。
摘要:Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and operation. Yet, it remains largely a dataset-specific task, requiring comprehensive training data, limiting scalability, and resulting in high model development and maintenance effort. Recently, foundation models that aim to learn generalizable patterns via extensive pretraining have shown superior performance in multiple prediction tasks. Despite their success and strong potential to address challenges in energy forecasting, their application in this domain remains largely unexplored. We address this gap by presenting the Foundation Models in Energy Time Series Forecasting (FETS) benchmark. We (1) provide a structured overview of energy forecasting use cases along three main dimensions: stakeholders, attributes, and data categories; (2) collect and analyze 54 datasets across 9 data categories, guided by typical stakeholder interests; (3) benchmark foundation models against classical machine learning approaches across different forecasting settings. Foundation models consistently outperform dataset-specific optimized machine learning approaches across all settings and data categories, despite the latter having seen the full historic target data during training. In particular, covariate-informed foundation models achieve the strongest performance. Further analysis reveals a strong correlation between predictive performance and spectral entropy, performance saturation beyond a certain context length, and improved performance at higher aggregation levels such as national load, district heating, and power grid data. Overall, our findings highlight the strong potential of foundation models as scalable and generalizable forecasting solutions for the energy domain, particularly in data-constrained and privacy-sensitive settings.

【3】Null-Space Flow Matching for MIMO Channel Estimation in Latency-Constrained Systems
标题:延迟约束系统中用于MMO信道估计的零空间流匹配
链接:https://arxiv.org/abs/2604.22005

作者:Junjie Zhao,Guangming Liang,Dongzhu Liu,Xiaonan Liu
备注:6 pages, 3 figures, 20 references
摘要:准确而低延迟的信道状态信息(CSI)获取对于多输入多输出(MIMO)通信系统是必不可少的。虽然先进的深度生成模型(如基于分数的模型和扩散模型)能够从有限的导频观测中实现高保真CSI重建,但它们通常会受到高推理延迟的影响。为了在严格的时延约束下实现准确的信道状态信息估计,提出了一种零空间流匹配(FM)框架,将导频受限MIMO信道估计分解为距离空间重构问题和零空间生成问题。具体而言,通道的距离空间分量直接从噪声导频观测中恢复,而只有模糊的零空间分量使用基于FM的生成先验迭代细化。为了进一步提高所提出的框架的鲁棒性,我们引入了幂律时间表,以更好地分配有限数量的细化步骤,以及噪声感知自适应校正策略,以抑制细化轨迹上的信道噪声。实验结果表明,即使在约3 ms的严格延迟预算下,我们的方法也能实现具有竞争力的归一化均方误差(NMSE),同时提供比基于模型和生成基线更高的估计精度和更快的推理速度。
摘要:Accurate yet low-latency channel state information (CSI) acquisition is essential for multiple-input multiple-output (MIMO) communication systems. While advanced deep generative models, such as score-based and diffusion models, enable high-fidelity CSI reconstruction from limited pilot observations, they often suffer from high inference latency. To achieve accurate CSI estimation under stringent latency constraints, this paper proposes a null-space flow matching (FM) framework that decomposes pilot-limited MIMO channel estimation into a range-space reconstruction problem and a null-space generation problem. Specifically, the range-space component of the channel is directly recovered from noisy pilot observations, while only the ambiguous null-space component is iteratively refined using an FM-based generative prior. To further improve the robustness of the proposed framework, we introduce a power-law time schedule to better allocate the limited number of refinement steps, along with a noise-aware adaptive correction strategy to suppress channel noise on the refinement trajectory. Experimental results demonstrate that our method achieves a competitive normalized mean square error (NMSE) even under a strict latency budget of around 3 ms, while delivering superior estimation accuracy and faster inference than both model-based and generative baselines.

【4】MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction
标题:MambaSCP:用于硬件高效的通道状态预测的混合注意力状态空间模型
链接:https://arxiv.org/abs/2604.21957

作者:Aladin Djuhera,Haris Gacanin,Holger Boche
摘要:最近的工作表明,基于注意力的Transformer和大语言模型(LLM)架构可以通过捕获跨信道状态信息(CSI)序列的长范围时间依赖性来实现强大的信道状态预测(CSP)性能。然而,这些模型遭受序列长度的二次缩放,导致大量的计算成本,内存消耗和推理延迟,这限制了它们在实时和资源受限的无线部署中的适用性。在本文中,我们调查是否选择性状态空间模型(SSM)可以作为一个硬件有效的替代CSI预测。我们提出了MambaCSP,一个混合注意SSM架构,用线性时间Mamba模型取代基于LLM的预测骨干。为了克服纯SSM的仅本地依赖性,我们引入了轻量级的补丁混合器注意层,定期注入跨令牌注意力,帮助进行长上下文CSI预测。广泛的MISO-OFDM仿真表明,MambaCSP将基于LLM的方法的预测精度提高了9- 12%,同时提供高达3.0倍的吞吐量,2.6倍的VRAM使用率和2.9倍的推理速度。我们的研究结果表明,混合状态空间架构为未来无线网络中的可扩展和硬件高效的AI原生CSI预测提供了一个有希望的方向。
摘要:Recent works have demonstrated that attention-based transformer and large language model (LLM) architectures can achieve strong channel state prediction (CSP) performance by capturing long-range temporal dependencies across channel state information (CSI) sequences. However, these models suffer from quadratic scaling in sequence length, leading to substantial computational cost, memory consumption, and inference latency, which limits their applicability in real-time and resource-constrained wireless deployments. In this paper, we investigate whether selective state space models (SSMs) can serve as a hardware-efficient alternative for CSI prediction. We propose MambaCSP, a hybrid-attention SSM architecture that replaces LLM-based prediction backbones with a linear-time Mamba model. To overcome the local-only dependencies of pure SSMs, we introduce lightweight patch-mixer attention layers that periodically inject cross-token attentions, helping with long-context CSI prediction. Extensive MISO-OFDM simulations show that MambaCSP improves prediction accuracy over LLM-based approaches by 9-12%, while delivering up to 3.0x higher throughput, 2.6x lower VRAM usage, and 2.9x faster inference. Our results demonstrate that hybrid state space architectures provide a promising direction for scalable and hardware-efficient AI-native CSI prediction in future wireless networks.

【5】Explanation of Dynamic Physical Field Predictions using WassersteinGrad: Application to Autoregressive Weather Forecasting
标题:使用WassersteinGrad解释动态物理场预测:自回归天气预测的应用
链接:https://arxiv.org/abs/2604.22580

作者:Younes Essafouri,Laure Raynaud,Luciano Drozda,Laurent Risser
摘要:随着将人工智能集成到高风险环境中的需求不断增长,解释神经网络预测背后的推理已经从理论上的好奇心转变为严格的操作要求。我们的工作的动机是解释动态物理场的自回归神经预测,如天气预报。基于属性的特征归因方法被广泛用于解释对此类数据的预测,特别是由于其对高维输入的可扩展性。同样有趣的是,基于梯度的技术,如SmoothGrad,现在是图像的标准,使用从几个噪声输入中获得的属性图的逐点平均值来增强解释。我们的目标是有效地适应这种聚合策略的动态物理场。要做到这一点,我们的第一个贡献是确定一个基本的故障模式时,平均扰动的属性地图上的动态物理场:随机输入扰动不会引起平稳振幅噪声的属性地图,而是造成几何位移的属性。因此,逐点平均模糊了这些空间错位的特征。为了解决这个问题,我们介绍WassersteinGrad,它提取的几何共识扰动属性地图计算其熵Wasserstein重心。在区域天气数据和气象学家验证的神经模型上获得的结果表明,WassersteinGrad在单步和自回归预测设置中基于梯度的基线上具有很好的可解释性。
摘要:As the demand to integrate Artificial Intelligence into high-stakes environments continues to grow, explaining the reasoning behind neural-network predictions has shifted from a theoretical curiosity to a strict operational requirement. Our work is motivated by the explanations of autoregressive neural predictions on dynamic physical fields, as in weather forecasting. Gradient-based feature attribution methods are widely used to explain the predictions on such data, in particular due to their scalability to high-dimensional inputs. It is also interesting to remark that gradient-based techniques such as SmoothGrad are now standard on images to robustify the explanations using pointwise averages of the attribution maps obtained from several noised inputs. Our goal is to efficiently adapt this aggregation strategy to dynamic physical fields. To do so, our first contribution is to identify a fundamental failure mode when averaging perturbed attribution maps on dynamic physical fields: stochastic input perturbations do not induce stationary amplitude noise in attribution maps, but instead cause a geometric displacement of the attributions. Consequently, pointwise averaging blurs these spatially misaligned features. To tackle this issue, we introduce WassersteinGrad, which extracts a geometric consensus of perturbed attribution maps by computing their entropic Wasserstein barycenter. The results, obtained on regional weather data and a meteorologist-validated neural model, demonstrate promising explainability properties of WassersteinGrad over gradient-based baselines across both single-step and autoregressive forecasting settings.

其他神经网络|深度学习|模型|建模(14篇)

【1】Quality-Driven Selective Mutation for Deep Learning
标题:深度学习的质量驱动选择性突变
链接:https://arxiv.org/abs/2604.22640

作者:Zaheed Ahmed,Emmanuel Charleson Dapaah,Philip Makedonski,Jens Grabowski
摘要:突变体支持测试和调试两个角色:(i)作为测试目标和(ii)作为真正的故障的替代品。难以杀死的突变体为测试改进提供了更好的指导,而当突变体被用来模拟真实的bug时,现实主义是必不可少的。基于这些角色,深度学习的选择性突变(DL)旨在通过选择产生抗性和现实突变体的操作符配置来降低突变体生成和执行的成本。然而,深度学习文献缺乏一个统一的衡量标准来捕捉这两个方面。这项研究提出了一个概率框架,量化突变质量沿两个互补的轴:阻力和现实主义。电阻适应难以杀死突变体的经典概念的DL设置使用统计的杀伤概率,而现实主义是通过突变体和真正的故障检测模式之间的广义Jaccard相似性来衡量。该框架能够对低质量的突变算子配置进行排名和过滤,而无需假设特定的用例。我们经验评估的方法对四个数据集的真正DL故障。三个数据集(CleanML,DeepFD和DeepLocalize)用于估计和选择高质量的操作员配置,并使用保留的defect4ML数据集进行验证。结果表明,质量驱动的选择减少了高达55.6%的突变体的数量,同时保留了基线对齐的选择阈值下的电阻和现实主义的典型水平。这些发现证实,双目标选择可以降低成本,而不损害突变体的有用性。
摘要 :Mutants support testing and debugging in two roles: (i) as test goals and (ii) as substitutes for real faults. Hard-to-kill mutants provide better guidance for test improvement, while realism is essential when mutants are used to simulate real bugs. Building on these roles, selective mutation for deep learning (DL) aims to reduce the cost of mutant generation and execution by choosing operator configurations that yield resistant and realistic mutants. However, the DL literature lacks a unified measure that captures both aspects. This study presents a probabilistic framework to quantify mutant quality along two complementary axes: resistance and realism. Resistance adapts the classical notion of hard-to-kill mutants to the DL setting using statistical killing probabilities, while realism is measured via the generalized Jaccard similarity between mutant and real-fault detectability patterns. The framework enables ranking and filtering of low-quality mutation-operator configurations without assuming a specific use case. We empirically evaluate the approach on four datasets of real DL faults. Three datasets (CleanML, DeepFD, and DeepLocalize) are used to estimate and select high-quality operator configurations, and the held-out defect4ML dataset is used for validation. Results show that quality-driven selection reduces the number of generated mutants by up to 55.6% while preserving typical levels of resistance and realism under baseline-aligned selection thresholds. These findings confirm that dual-objective selection can lower cost without compromising the usefulness of mutants for either role.

【2】Beyond Patient Invariance: Learning Cardiac Dynamics via Action-Conditioned JEPAs
标题:超越患者不变性:通过条件下的JEPA学习心脏动力学
链接:https://arxiv.org/abs/2604.22618

作者:Jose Geraldo Fernandes,Luiz Facury,Pedro Robles Dutenhefner,Wagner Meira
摘要:医疗保健中的自我监督学习在很大程度上依赖于基于不变性的目标,这使得同一患者的不同视图之间的相似性最大化。虽然对静态解剖结构有效,但这种模式从根本上与临床诊断不一致,因为它在数学上迫使模型抑制其旨在检测的瞬时病理变化。我们提出了一个转向疾病条件世界模型,学习模拟疾病进展的动态,或事件条件。将LeJEPA框架适应于生理时间序列,我们将病理学定义为不是静态标签,而是作用于患者潜伏状态的转换向量。通过预测疾病发作时心脏的未来电生理状态,我们的模型明确地将稳定的解剖特征与动态的病理力量分开。在MIMIC-IV-ECG数据集上进行评估,我们的方法在关键的分诊任务上优于完全监督的基线。至关重要的是,我们展示了卓越的样本效率:在低资源条件下,我们的世界模型比监督学习的性能高出0.05 AUROC以上。这些结果表明,建模生物动力学提供了一个密集的监督信号,这是远远超过静态分类更强大。源代码可在https://github.com/cljosegfer/lesaude-dynamics上获得
摘要:Self-supervised learning in healthcare has largely relied on invariance-based objectives, which maximize similarity between different views of the same patient. While effective for static anatomy, this paradigm is fundamentally misaligned with clinical diagnosis, as it mathematically compels the model to suppress the transient pathological changes it is intended to detect. We propose a shift towards Action-Conditioned World Models that learn to simulate the dynamics of disease progression, or Event-Conditioned. Adapting the LeJEPA framework to physiological time-series, we define pathology not as a static label, but as a transition vector acting on a patient's latent state. By predicting the future electrophysiological state of the heart given a disease onset, our model explicitly disentangles stable anatomical features from dynamic pathological forces. Evaluated on the MIMIC-IV-ECG dataset, our approach outperforms fully supervised baselines on the critical triage task. Crucially, we demonstrate superior sample efficiency: in low-resource regimes, our world model outperforms supervised learning by over 0.05 AUROC. These results suggest that modeling biological dynamics provides a dense supervision signal that is far more robust than static classification. Source code is available at https://github.com/cljosegfer/lesaude-dynamics

【3】Deep Learning for Model Calibration in Simulation of Itaconic Acid Production
标题:衣康酸生产模拟中模型校准的深度学习
链接:https://arxiv.org/abs/2604.22496

作者:Daria Fokina,Marco Baldan,Constantin Romankiewicz,Wolfgang Laudensack,Roland Ulber,Michael Bortz
摘要:在这项研究中,深度学习被用来估计动力学参数,用于基于在不同搅拌速度和反应器规模下进行的真实批量实验来建模衣康酸生产。两种深度学习策略,即直接深度学习(Direct Deep Learning,缩写为EML)和生成式条件流匹配(Generative Conditional Flow Matching,缩写为CFM),被比较并以非线性回归作为基准。与传统方法相比,CFM始终产生更准确的结果。CFM预测的浓度曲线与非线性回归得到的浓度曲线非常吻合,而CFM预测的浓度曲线偏差较大。在放大实验中观察到类似的行为,其中CFM模型再次比直接方法更好地推广并且更鲁棒。这些发现表明,CFM可以可靠地预测不同操作条件和规模的系统行为,为动态生物过程模型中的参数估计提供了一个灵活且数据高效的框架。
摘要:In this study, deep learning is used to estimate kinetic parameters for modeling itaconic acid production based on real batch experiments conducted at different agitation speeds and reactor scales. Two deep learning strategies, namely direct deep learning (DDL) and generative conditional flow matching (CFM) are compared and benchmarked against nonlinear regression as a reference method. Compared with DDL, CFM consistently yields more accurate results. The concentration profiles predicted by CFM closely match those obtained from nonlinear regression, whereas DDL results in larger deviations. Similar behavior is observed in the scale-up experiments, where the CFM model again generalizes better and is more robust than the direct approach. These findings demonstrate that CFM can reliably predict system behavior across different operating conditions and scales, offering a flexible and data-efficient framework for parameter estimation in dynamic bioprocess models.

【4】HubRouter: A Pluggable Sub-Quadratic Routing Primitive for Hybrid Sequence Models
标题:HubRouter:混合序列模型的可插入次二次路由基元
链接:https://arxiv.org/abs/2604.22442

作者:Abhinaba Basu
摘要:我们引入了HubRouter,这是一个可插入的模块,它用O(nM)hub介导的路由替换了O(n^2)注意层,其中M << n是少量学习的hub令牌。我们在两个从头开始的架构中展示了它:Jamba风格的混合和12层的Transformer;改造成预训练模型是一个经过测试的负面案例。HubRouter实现了一个编码-解码-评分-理事会流水线:M个学习的hubs交叉处理所有token,token针对hubs投影以路由指纹,评分头选择前k个token,稀疏理事会只处理选定的子集。 我们在三个设置中验证HubRouter。(1)Hub-Jamba在匹配的PyTorch原生基线中,在序列长度为1024时,PPL提高了4.2%(200.2 vs 209.0,单个种子;可能在种子噪声内),训练吞吐量高达约90倍;优化的基线将使其缩小到约10- 15倍。(2)逐步替换25%的Transformer注意力层,在我们的匹配预算扫描中提供了最好的困惑(268.0 vs 282.4纯Transformer)。(3)Hub-GPT提供严格的因果路由,在3个种子上实现了211.5 +/- 0.4的PPL(议会后因果修复);比Jamba的208.5 +/- 0.7差了大约3 PPL,这是避免O(n^2)计算的可测量质量成本。后缀,块大小C几乎没有影响;前缀块大小的好处是我们在对抗性审查中发现的双向委员会泄漏的产物。 多种子中心计数扫描(跨越M=1-32的~105次运行)揭示M=8-14作为可靠收敛的子带(4-5/5种子);通过正交正则化将M=6拯救到5/5,而M>=20显示增加的种子敏感性。配套论文arXiv:2603.20997(Basu,2026)定义了路由诊断任务。代码和脚本将被释放。
摘要:We introduce HubRouter, a pluggable module that replaces O(n^2) attention layers with O(nM) hub-mediated routing, where M << n is a small number of learned hub tokens. We demonstrate it in two from-scratch architectures: a Jamba-style hybrid and a 12-layer Transformer; retrofit into pretrained models is a tested negative case. HubRouter implements an encode-decode-score-council pipeline: M learned hubs cross-attend to all tokens, tokens project against hubs for routing fingerprints, a score head selects top-k tokens, and a sparse council attends only to the selected subset. We validate HubRouter in three settings. (1) Hub-Jamba yields a nominal 4.2% PPL improvement (200.2 vs 209.0, single seed; possibly within seed noise) and up to ~90x training throughput at sequence length 1024 in matched PyTorch-native baselines; an optimised baseline would narrow this to ~10-15x. (2) Graduated replacement of 25% of Transformer attention layers gives the best perplexity in our matched-budget sweep (268.0 vs 282.4 pure Transformer). (3) Hub-GPT provides strictly causal routing, achieving PPL 211.5 +/- 0.4 over 3 seeds (post council-causal fix); approximately 3 PPL worse than Jamba's 208.5 +/- 0.7, a measurable quality cost for avoiding O(n^2) computation. Post-fix, chunk size C has little effect; the pre-fix chunk-size benefit was an artifact of a bidirectional-council leak we found in adversarial review. A multi-seed hub-count sweep (~105 runs across M=1-32) reveals M=8-14 as the reliably-converging sub-band (4-5/5 seeds); M=6 is rescued to 5/5 by orthogonal regularization, while M>=20 shows increasing seed sensitivity. Companion paper arXiv:2603.20997 (Basu, 2026) defines the routing diagnostic task. Code and scripts will be released.

【5】SOC-ICNN: From Polyhedral to Conic Geometry for Learning Convex Surrogate Functions
标题:SOC-ICNN:从多边形到锥形几何学习凸替代函数
链接:https://arxiv.org/abs/2604.22355

作者:Kang Liu,Jianchen Hu
备注:28 pages and no figure
摘要:经典的基于ReLU的输入凸神经网络(ICNN)等价于线性规划(LP)的最优值函数。这种内在的结构等价性限制了它们对分段线性多面体函数的表示能力。为了克服这个代表性的瓶颈,我们提出了SOC-ICNN,一个架构,概括了底层的优化类从LP二阶锥规划(SOCP)。通过显式地注入半正定曲率和基于欧氏范数的圆锥基元,我们的公式将原生光滑曲率引入表示中,同时保留严格的优化理论解释。我们正式证明了SOC-ICNN严格扩展了ReLU-ICNN的表示空间,而不会增加前向传递复杂度的渐近阶。大量的实验表明,SOC-ICNN大大提高了函数逼近,同时提供有竞争力的下游决策质量。该代码可在https://github.com/Kanyooo/SOC-ICNN上获得。
摘要 :Classical ReLU-based Input Convex Neural Networks (ICNNs) are equivalent to the optimal value functions of Linear Programming (LP). This intrinsic structural equivalence restricts their representational capacity to piecewise-linear polyhedral functions. To overcome this representational bottleneck, we propose the SOC-ICNN, an architecture that generalizes the underlying optimization class from LP to Second-Order Cone Programming (SOCP). By explicitly injecting positive semi-definite curvature and Euclidean norm-based conic primitives, our formulation introduces native smooth curvature into the representation while preserving a rigorous optimization-theoretic interpretation. We formally prove that SOC-ICNNs strictly expand the representational space of ReLU-ICNNs without increasing the asymptotic order of forward-pass complexity. Extensive experiments demonstrate that SOC-ICNN substantially improves function approximation, while delivering competitive downstream decision quality. The code is available at https://github.com/Kanyooo/SOC-ICNN.

【6】A Brain-Inspired Deep Separation Network for Single Channel Raman Spectra Unmixing
标题:用于单通道拉曼光谱解混的脑启发深度分离网络
链接:https://arxiv.org/abs/2604.22324

作者:Gaoruishu Long,Jinchao Liu,Bo Liu,Jie Liu,Xiaolin Hu
备注:Accepted by the 2026 International Joint Conference on Neural Networks (IJCNN 2026). 8 pages, 5 figures
摘要:在实际应用中获得的拉曼光谱通常是测试样品中各种物质的几个光谱的噪声组合。将这种光谱分解成对应于每种物质的单独组分具有很大的价值,并且一直是拉曼光谱学中的长期挑战。现有的解混方法主要用于反演超定混合模型,因此需要多个混合光谱作为输入。然而,拉曼光谱中的开放域和/或非合作检测应用(例如受控物质检测)需要单通道解决方案,该解决方案可以通过仅分析单个有噪混合光谱来从数千个候选物中识别单个成分。据我们所知,稀疏回归是唯一现有的解决方案,可以应付这种情况下,但它有很低的容忍度,噪声,很难在实践中应用。为了解决这些限制,我们引入了一种新的神经方法,用于受语音分离启发的单通道拉曼光谱解混。它旨在解决欠定系统,可以从数千种成分(物质)的库中分解出有噪的混合光谱。该方法的核心是一个深度分离神经网络(RSSNet),它以混合光谱为输入,输出纯组分的光谱。我们创建了两个单通道拉曼光谱解混的合成数据集,并在这些数据集上证明了RSSNet的可行性和优越性(优于竞争方法>4dB)。此外,我们还验证了仅在合成数据上训练的RSSNet可以成功地对矿物粉末混合物的真实混合光谱进行分解,表现出很强的泛化能力。我们的方法代表了一个新的范例拉曼解混,并实现了新的可能性,快速检测拉曼混合物。
摘要:Raman spectra obtained in real world applications are often a noisy combination of several spectra of various substances in a tested sample. Unmixing such spectra into individual components corresponding to each of the substances is of great value and has been a longstanding challenge in Raman spectroscopy. Existing unmixing methods are predominantly designed to invert an overdetermined mixed model and therefore require multiple mixed spectra as input. However, open domain and/or non-cooperative detection applications in Raman spectroscopy such as controlled substance detection, call for single-channel solutions which can identify individual components from thousands of candidates by analyzing only a single noisy mixed spectrum. To our knowledge, sparse regression is the only existing solution which can cope with this scenario, yet it has very low tolerance to noises and can hardly be applicable in practice. To address these limitations, we introduce a novel neural approach for single-channel Raman spectrum unmixing inspired by speech separation. It aims at solving underdetermined systems and can decompose a noisy mixed spectrum from a library of thousands of components (substances). The core of our method is a deep separation neural network (RSSNet) which takes a mixed spectrum as input and outputs spectra of pure components. We created two synthetic datasets of single-channel Raman spectra unmixing and demonstrated feasibility and superiority of RSSNet on these datasets (outperform competing methods by >4dB). Furthermore, we verified that RSSNet, trained solely on synthetic data, can successfully unmix real-world mixed spectra of mixtures of mineral powders, exhibiting strong generalization. Our approach represents a new paradigm for Raman unmixing and enables new possibilities for fast detection of Raman mixtures.

【7】Learning-augmented robotic automation for real-world manufacturing
标题:面向现实世界制造的学习增强机器人自动化
链接:https://arxiv.org/abs/2604.22235

作者:Yunho Kim,Quan Nguyen,Taewhan Kim,Youngjin Heo,Joonho Lee
摘要:工业机器人广泛应用于制造业,但大多数操作仍然依赖于固定的路点脚本,这些脚本对环境变化很脆弱。基于学习的控制提供了一种更具适应性的替代方案,但目前尚不清楚这些方法是否仍然主要局限于实验室演示,是否能够维持数小时的可靠操作,提供一致的质量,并在现场生产线上安全地工作。在这里,我们介绍了学习增强机器人自动化,这是一种混合系统,将学习任务控制器和神经3D安全监控器集成到传统的工业工作流程中。我们将该系统部署在电动机生产线上,以在实际制造限制下自动化可变形电缆的插入和焊接,这一步骤以前由人类工人手动执行。在每个任务不到20分钟的真实世界数据的情况下,该系统连续运行5小时10分钟,生产了108台电机,没有物理围栏,并在产品级质量控制测试中达到了99.4%的合格率。它保持了接近人类的节拍时间,同时减少了焊点质量和周期时间的变化。这些结果为通过基于学习的方法扩展工业自动化建立了一条实用的途径。
摘要:Industrial robots are widely used in manufacturing, yet most manipulation still depends on fixed waypoint scripts that are brittle to environmental changes. Learning-based control offers a more adaptive alternative, but it remains unclear whether such methods, still mostly confined to laboratory demonstrations, can sustain hours of reliable operation, deliver consistent quality, and behave safely around people on a live production line. Here we present Learning-Augmented Robotic Automation, a hybrid system that integrates learned task controllers and a neural 3D safety monitor into conventional industrial workflows. We deployed the system on an electric-motor production line to automate deformable cable insertion and soldering under real manufacturing constraints, a step previously performed manually by human workers. With less than 20 min of real-world data per task, the system operated continuously for 5 h 10 min, producing 108 motors without physical fencing and achieving a 99.4% pass rate on product-level quality-control tests. It maintained near-human takt time while reducing variability in solder-joint quality and cycle time. These results establish a practical pathway for extending industrial automation with learning-based methods.

【8】LTBs-KAN: Linear-Time B-splines Kolmogorov-Arnold Networks
标题:LTBs-KAN:线性时间B样条Kolmogorov-Arnold网络
链接:https://arxiv.org/abs/2604.22034

作者:Eduardo Said Merin-Martinez,Andres Mendez-Vazquez,Eduardo Rodriguez-Tello
摘要:Kolmogorov-Arnold网络(KAN)是一种新的神经网络架构,它提供了一种替代多层感知器(MLP)的方法,具有更好的可解释性和可表达性。然而,由于B样条函数计算的递归性质,KAN比MLP慢得多,限制了它们的应用。本文提出了一种新的具有线性复杂度的基样条线性时间B样条Kolmogorov-Arnold网络(LTBs-KAN)。与以前的方法,依赖于布尔-曼斯菲尔德-考克斯样条算法或其他计算密集型的数学函数,我们的方法显着降低了计算负担。此外,我们进一步降低模型的参数,通过积的和矩阵分解的前向通过不牺牲性能。在MNIST、Fashion-MNIST和CIFAR-10上的实验表明,与其他KAN实现相比,LTBs-KAN在用作构建架构块时实现了良好的时间复杂度和参数降低。
摘要:Kolmogorov-Arnold Networks (KANs) are a recent neural network architecture offering an alternative to Multilayer Perceptrons (MLPs) with improved explainability and expressibility. However, KANs are significantly slower than MLPs due to the recursive nature of B-spline function computations, limiting their application. This work addresses these issues by proposing a novel base-spline Linear-Time B-splines Kolmogorov-Arnold Network (LTBs-KAN) with linear complexity. Unlike previous methods that rely on the Boor-Mansfield-Cox spline algorithm or other computationally intensive mathematical functions, our approach significantly reduces the computational burden. Additionally, we further reduce model's parameter through product-of-sums matrix factorization in the forward pass without sacrificing performance. Experiments on MNIST, Fashion-MNIST and CIFAR-10 demonstrate that LTBs-KAN achieves good time complexity and parameter reduction, when used as building architectural blocks, compared to other KAN implementations.

【9】Focus Session: Hardware and Software Techniques for Accelerating Multimodal Foundation Models
标题:焦点会议:加速多模式基础模型的硬件和软件技术
链接:https://arxiv.org/abs/2604.21952

作者:Muhammad Shafique,Abdul Basit,Muhammad Abdullah Hanif,Alberto Marchisio,Rachmad Vidya Wicaksana Putra,Minghao Shao
备注:Accepted at the Design, Automation and Test in Europe Conference (DATE), April 20-22, 2026 in Verona, Italy
摘要:这项工作提出了一个多层次的方法,有效地加速多模态基础模型(MFM)。它将Transformer模块的硬件和软件协同设计与降低计算和内存需求的优化管道相结合。在模型开发过程中,它通过针对特定领域的调整来增强性能。我们的方法进一步结合了硬件和软件技术,优化的MFM。具体而言,它采用MFM压缩,使用分层感知混合精度量化和结构修剪的Transformer块和MLP通道。它还通过推测性解码优化操作,模型级联通过从小到大的级联路由查询,并使用轻量级自检来确定何时升级到更大的模型,以及序列长度,视觉分辨率和步幅的共同优化,以及图形级运算符融合。为了有效地执行该模型,基于底层硬件架构以及内存高效注意力来优化处理流程,以满足片上带宽和延迟预算。为了支持这一点,采用了专门用于Transformer工作负载的硬件加速器,可以通过专家设计或LLM辅助设计方法进行开发。我们证明了所提出的方法的有效性,医疗MFMs和代码生成任务,并得出结论,扩展到节能尖峰MFMs。
摘要 :This work presents a multi-layered methodology for efficiently accelerating multimodal foundation models (MFMs). It combines hardware and software co-design of transformer blocks with an optimization pipeline that reduces computational and memory requirements. During model development, it employs performance enhancements through fine-tuning for domain-specific adaptation. Our methodology further incorporates hardware and software techniques for optimizing MFMs. Specifically, it employs MFM compression using hierarchy-aware mixed-precision quantization and structural pruning for transformer blocks and MLP channels. It also optimizes operations through speculative decoding, model cascading that routes queries through a small-to-large cascade and uses lightweight self-tests to determine when to escalate to larger models, as well as co-optimization of sequence length, visual resolution & stride, and graph-level operator fusion. To efficiently execute the model, the processing dataflow is optimized based on the underlying hardware architecture together with memory-efficient attention to meet on-chip bandwidth and latency budgets. To support this, a specialized hardware accelerator for the transformer workloads is employed, which can be developed through expert design or an LLM-aided design approach. We demonstrate the effectiveness of the proposed methodology on medical-MFMs and on code generation tasks, and conclude with extensions toward energy-efficient spiking-MFMs.

【10】Relaxation-Informed Training of Neural Network Surrogate Models
标题:神经网络代理模型的放松信息训练
链接:https://arxiv.org/abs/2604.22746

作者:Calvin Tsay
备注:35 pages, 5 figures
摘要:作为代理模型训练的ReLU神经网络可以精确地嵌入混合整数线性规划(MILP)中,从而实现对学习函数的全局优化。所得MILP的易处理性取决于网络的结构性质,即,相关公式中二元变量的数量和连续LP松弛的紧密度。这些属性是在训练过程中确定的,但标准训练目标(经典权重正则化的预测损失)没有提供直接控制它们的机制。这项工作研究了直接针对下游MILP可处理性的训练正则化器。具体来说,我们提出了简单的基于边界的正则化惩罚MILP配方的大M常数和/或不稳定的神经元的数量。此外,我们引入了一个LP松弛间隙正则化器,明确惩罚训练点处连续松弛的每样本间隙。我们推导出其相关的梯度,并提供了一个实现从LP双变量没有自定义的自动微分工具。我们表明,结合上述正则化可以近似的LP间隙相对于网络参数的全导数,捕捉直接和间接的灵敏度。非凸基准函数和分位数神经网络代理的两阶段随机规划问题的实验表明,所提出的正则化器可以减少MILP求解时间高达四个数量级相对于未正则化的基线,同时保持竞争力的代理模型的准确性。
摘要:ReLU neural networks trained as surrogate models can be embedded exactly in mixed-integer linear programs (MILPs), enabling global optimization over the learned function. The tractability of the resulting MILP depends on structural properties of the network, i.e., the number of binary variables in associated formulations and the tightness of the continuous LP relaxation. These properties are determined during training, yet standard training objectives (prediction loss with classical weight regularization) offer no mechanism to directly control them. This work studies training regularizers that directly target downstream MILP tractability. Specifically, we propose simple bound-based regularizers that penalize the big-M constants of MILP formulations and/or the number of unstable neurons. Moreover, we introduce an LP relaxation gap regularizer that explicitly penalizes the per-sample gap of the continuous relaxation at training points. We derive its associated gradient and provide an implementation from LP dual variables without custom automatic differentiation tools. We show that combining the above regularizers can approximate the full total derivative of the LP gap with respect to the network parameters, capturing both direct and indirect sensitivities. Experiments on non-convex benchmark functions and a two-stage stochastic programming problem with quantile neural network surrogates demonstrate that the proposed regularizers can reduce MILP solve times by up to four orders of magnitude relative to an unregularized baseline, while maintaining competitive surrogate model accuracy.

【11】Mixed Membership sub-Gaussian Models
标题:混合成员次高斯模型
链接:https://arxiv.org/abs/2604.22633

作者:Huan Qing
备注:30 pages, 6 figures, 2 tables
摘要:高斯混合模型由于其简单性和可解释性而被广泛应用于无监督学习。然而,经典高斯混合模型的一个基本限制是,它迫使每个观测值只属于一个分量。在许多实际应用中,如遗传学、社会网络分析和文本挖掘,一个观察可能自然地属于多个组件或在几个潜在组件中表现出部分成员关系。为了克服这一限制,我们提出了混合隶属度亚高斯模型,它扩展了经典的高斯混合框架,允许每个观察属于多个组件。该模型继承了经典高斯混合模型的可解释性,同时为捕获复杂的重叠结构提供了更大的灵活性。我们开发了一种有效的谱算法来估计每个个体观测的混合隶属度,并在组分中心的温和分离条件下,我们证明了每个个体隶属度向量的估计误差可以以高概率任意小。据我们所知,这是第一个工作,提供了一个计算效率的估计与这样一个消失的误差保证的混合成员扩展的高斯混合模型。大量的实验研究表明,我们的方法优于现有的方法,忽略混合成员。
摘要:The Gaussian mixture model is widely used in unsupervised learning, owing to its simplicity and interpretability. However, a fundamental limitation of the classical Gaussian mixture model is that it forces each observation to belong to exactly one component. In many practical applications, such as genetics, social network analysis, and text mining, an observation may naturally belong to multiple components or exhibit partial membership in several latent components. To overcome this limitation, we propose the mixed membership sub-Gaussian model, which extends the classical Gaussian mixture framework by allowing each observation to belong to multiple components. This model inherits the interpretability of the classical Gaussian mixture model while offering greater flexibility for capturing complex overlapping structures. We develop an efficient spectral algorithm to estimate the mixed membership of each individual observation, and under mild separation conditions on the component centres, we prove that the estimation error of the per-individual membership vector can be made arbitrarily small with high probability. To our knowledge, this is the first work to provide a computationally efficient estimator with such a vanishing-error guarantee for a mixed-membership extension of the Gaussian mixture model. Extensive experimental studies demonstrate that our method outperforms existing approaches that ignore mixed memberships.

【12】Multi-output Extreme Spatial Model for Complex Aircraft Production Systems
标题:复杂飞机生产系统的多输出极端空间模型
链接:https://arxiv.org/abs/2604.22548

作者:Cheolhei Lee,Xing Wang,Xiaowei Yue,Jianguo Wu
摘要:问题定义:机器学习中的数据驱动模型实现了生产系统的有效管理。然而,大多数机器学习模型都致力于对平均响应或平均模式进行建模,这不适合研究飞机制造中通常最感兴趣的异常极端事件。由于重尾分布的极端事件会导致系统管理费用过高,因此迫切需要复杂的极端模型来分析复杂的极端风险。极端模型的工程应用通常集中在单个极端事件上,这对于具有相关性的复杂系统是不够的。方法/结果:我们引入了一个极端的空间模型的多输出响应控制系统,有效地捕捉动态使用双线性函数的控制变量和测量位置的两个空间域。研究了边际参数建模和极值相关性。此外,一个有效的图辅助的复合似然估计和相应的计算算法,以应付高维输出。复合材料飞机生产的应用表明,所提出的模型能够进行全面的分析,具有优越的预测性能的极端事件相比,规范的方法。管理方面的影响:我们的方法展示了如何使用极端空间模型来预测极端事件和管理飞机等复杂生产系统中的极端风险。这有助于在飞机生产系统及其他领域实现更好的质量管理和操作安全。
摘要:Problem definition: Data-driven models in machine learning have enabled efficient management of production systems. However, a majority of machine learning models are devoted to modeling the mean response or average pattern, which is inappropriate for studying abnormal extreme events that are often of primary interest in aircraft manufacturing. Since extreme events from heavy-tailed distributions give rise to prohibitive expenditures in system management, sophisticated extreme models are urgently needed to analyze complex extreme risks. Engineering applications of extreme models usually focus on individual extreme events, which is insufficient for complex systems with correlations. Methodology/results: We introduce an extreme spatial model for multi-output response control systems that efficiently captures the dynamics using a bilinear function on two spatial domains for control variables and measurement locations. Marginal parameter modeling and extremal dependence have been investigated. In addition, an efficient graph-assisted composite likelihood estimation and corresponding computational algorithms are developed to cope with high-dimensional outputs. The application to composite aircraft production shows that the proposed model enables comprehensive analyses with superior predictive performance on extreme events compared to canonical methods. Managerial implications: Our method shows how to use an extreme spatial model for predicting extreme events and managing extreme risks in complex production systems such as aircraft. This can help achieve better quality management and operation safety in aircraft production systems and beyond.

【13】On Benchmark Hacking in ML Contests: Modeling, Insights and Design
标题:ML竞赛中的基准黑客:建模、见解和设计
链接:https://arxiv.org/abs/2604.22230

作者:Xiaoyun Qiu,Yang Yu,Haifeng Xu
摘要:基准黑客是指调整机器学习模型,使其在某些评估标准上得分很高,而不提高真正的泛化能力或忠实地解决预期问题。我们在一个通用的机器学习竞赛中研究这种现象,每个参赛者选择两种类型的努力:创造性的努力,提高模型的能力,如竞赛主持人所期望的,和机械的努力,只提高模型的健身比赛中的特定任务,而没有真正的推广。在这个竞争博弈中,我们建立了对称单调纯策略均衡的存在性。它还提供了一个自然的定义基准黑客在这种战略背景下,通过比较一个球员的均衡努力分配的单代理基线方案。根据我们的定义,低于一定阈值(低类型)的参赛者总是参与基准黑客,而高于阈值的参赛者则不会。此外,我们表明,更倾斜的奖励结构(有利于排名靠前的参赛者)可以引出更理想的比赛结果。我们还提供了经验证据来支持我们的理论预测。
摘要 :Benchmark hacking refers to tuning a machine learning model to score highly on certain evaluation criteria without improving true generalization or faithfully solving the intended problem. We study this phenomenon in a generic machine learning contest, where each contestant chooses two types of effort: creative effort that improves model capability as desired by the contest host, and mechanistic effort that only improves the model's fitness to the particular task in contest without contributing to true generalization. We establish the existence of a symmetric monotone pure strategy equilibrium in this competition game. It also provides a natural definition of benchmark hacking in this strategic context by comparing a player's equilibrium effort allocation to that of a single-agent baseline scenario. Under our definition, contestants with types below certain threshold (low types) always engage in benchmark hacking, whereas those above the threshold do not. Furthermore, we show that more skewed reward structures (favoring top-ranked contestants) can elicit more desirable contest outcomes. We also provide empirical evidence to support our theoretical predictions.

【14】Foundation models for discovering robust biomarkers of neurological disorders from dynamic functional connectivity
标题:从动态功能连接性中发现神经系统疾病稳健生物标志物的基础模型
链接:https://arxiv.org/abs/2604.22018

作者:Deepank Girish,Yi Hao Chan,Sukrit Gupta,Jing Xia,Jagath C. Rajapakse
摘要:最近提出了几种脑基础模型(FM),通过模拟动态功能连接(FC)来预测大脑疾病。虽然它们表现出显着的模型性能和零或Few-Shot泛化,但被鉴定为潜在生物标志物的显著特征尚未得到彻底评估。我们提出了RE-CONFIRM,这是一个用于评估由包括FM在内的深度学习(DL)模型阐明的潜在生物标志物候选物的鲁棒性的框架。通过对自闭症谱系障碍(ASD),注意力缺陷多动障碍(ADHD)和阿尔茨海默病(AD)的五个大型数据集的实验,我们发现,尽管常用的性能指标提供了对模型预测的直观评估,但它们不足以评估这些模型识别的生物标志物的鲁棒性。RE-CONFIRM指标显示,简单地微调FM会导致模型无法有效地捕获区域中心,即使在已知涉及中心的疾病中,如ASD和ADHD。有鉴于此,我们提出了Hub-LoRA(低秩适应)作为一种微调技术,使FM不仅优于定制的DL模型,而且还产生了荟萃分析支持的神经生物学上忠实的生物标志物。RE-CONFIRM是可推广的,可以很容易地应用于确定功能MRI数据集上训练的DL模型的鲁棒性。代码可从以下网址获得:https://github.com/SCSE-Biomedical-Computing-Group/RE-CONFIRM。
摘要:Several brain foundation models (FM) have recently been proposed to predict brain disorders by modelling dynamic functional connectivity (FC). While they demonstrate remarkable model performance and zero- or few-shot generalization, the salient features identified as potential biomarkers are yet to be thoroughly evaluated. We propose RE-CONFIRM, a framework for evaluating the robustness of potential biomarker candidates elucidated by deep learning (DL) models including FMs. From experiments on five large datasets of Autism Spectrum Disorder (ASD), Attention-deficit Hyperactivity Disorder (ADHD), and Alzheimer's Disease (AD), we found that although commonly used performance metrics provide an intuitive assessment of model predictions, they are insufficient for evaluating the robustness of biomarkers identified by these models. RE-CONFIRM metrics revealed that simply finetuning FMs leads to models that fail to capture regional hubs effectively, even in disorders where hubs are known to be implicated, such as ASD and ADHD. In view of this, we propose Hub-LoRA (Low-Rank Adaptation) as a fine-tuning technique that enables FMs to not only outperform customised DL models but also produce neurobiologically faithful biomarkers supported by meta-analyses. RE-CONFIRM is generalizable and can be easily applied to ascertain the robustness of DL models trained on functional MRI datasets. Code is available at: https://github.com/SCSE-Biomedical-Computing-Group/RE-CONFIRM.

其他(20篇)

【1】Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection
标题:花更少,适应更好:通过主动实验选择实现预算高效的标度定律匹配
链接:https://arxiv.org/abs/2604.22753

作者:Sijie Li,Shanda Li,Haowei Lin,Weiwei Sun,Ameet Talwalkar,Yiming Yang
摘要:缩放定律被用来规划数百万美元的培训运行,但拟合这些定律本身就可能花费数百万美元。在现代大规模的工作流程中,组装一组足够信息丰富的试点实验已经是一个主要的数据分配问题,而不是一个常规的预处理步骤。我们将标度律拟合公式化为具有非线性感知的序贯实验设计:给定具有异质成本的可运行实验的有限池,选择执行哪些运行,以便在高成本目标区域中最大化外推精度。然后,我们提出了一个不确定性意识的方法,顺序分配实验预算对运行最有用的目标区域外推。在不同的尺度律任务基准中,我们的方法始终优于经典的基于设计的基线,并且通常接近完整实验集的拟合性能,而仅使用总训练预算的10%左右。我们的代码可在https://github.com/PlanarG/active-sl上获得。
摘要:Scaling laws are used to plan multi-million-dollar training runs, but fitting those laws can itself cost millions. In modern large-scale workflows, assembling a sufficiently informative set of pilot experiments is already a major budget-allocation problem rather than a routine preprocessing step. We formulate scaling-law fitting as budget-aware sequential experimental design: given a finite pool of runnable experiments with heterogeneous costs, choose which runs to execute so as to maximize extrapolation accuracy in a high-cost target region. We then propose an uncertainty-aware method for sequentially allocating experimental budget toward the runs most useful for target-region extrapolation. Across a diverse benchmark of scaling-law tasks, our method consistently outperforms classical design-based baselines, and often approaches the performance of fitting on the full experimental set while using only about 10% of the total training budget. Our code is available at https://github.com/PlanarG/active-sl.

【2】Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data
标题:从现代数据中神经恢复班图语言历史词汇结构
链接:https://arxiv.org/abs/2604.22730

作者:Hillary Mutisya,John Mugane
摘要:我们研究了专门在现代形态学数据上训练的神经模型是否可以恢复与历史重建一致的跨语言词汇结构。使用BantuMorph v7,一个在Bantu形态范例上的Transformer,我们分析了14种东部和南部班图语言,提取了它们的名词和动词词元的编码器嵌入,并识别了5种以上语言中共享的728个名词和1,525个动词同源候选词。根据已建立的历史资源-班图语词汇重建数据库评估这些候选词(BLR 3; 4,786个重建的原始班图形式)和ASJP基本词汇-我们确认前11个名词候选中的10个(90.9%)与先前重建的原始班图形式一致,包括 *-ntU '人'(8种语言),*gombe '牛'(9种语言)和 *mUn(9种语言)。扩展到动词,12个动词同源词与重建的原始班图语词根一致,包括 *-bon- 'see'和 *-jIm- 'stand',每个都在广泛的地理范围内得到证明。使用独立翻译模型(NLLB-600 M)的跨模型验证证实了这些模式:两个模型都恢复了与已建立的Guthrie区分类一致的同源簇和系统发育分组(p < 0.01)。跨语言名词类分析显示,所有13个产出类在不同语言间保持>0.83的余弦相似性(类内>类间,p < 10^-9)。我们的数据集仅限于东部和南部班图人,所以我们解释这些结果为恢复与原始班图人一致的共享班图语词汇结构,而不是明确区分原始班图人的保留和后来的区域创新。
摘要:We investigate whether neural models trained exclusively on modern morphological data can recover cross-lingual lexical structure consistent with historical reconstruction. Using BantuMorph v7, a transformer over Bantu morphological paradigms, we analyze 14 Eastern and Southern Bantu languages, extract encoder embeddings for their noun and verb lemmas, and identify 728 noun and 1,525 verb cognate candidates shared across 5+ languages. Evaluating these candidates against established historical resources-the Bantu Lexical Reconstructions database (BLR3; 4,786 reconstructed Proto-Bantu forms) and the ASJP basic vocabulary-we confirm 10 of the top 11 noun candidates (90.9%) align with previously reconstructed Proto-Bantu forms, including *-ntU 'person' (8 languages), *gombe 'cow' (9 languages), and *mUn (9 languages). Extending to verbs, 12 verb cognates align with reconstructed Proto-Bantu roots, including *-bon- 'see' and *-jIm- 'stand', each attested across wide geographic ranges. Cross-model validation using an independent translation model (NLLB-600M) confirms these patterns: both models recover cognate clusters and phylogenetic groupings consistent with established Guthrie-zone classifications (p < 0.01). Cross-lingual noun class analysis reveals that all 13 productive classes maintain >0.83 cosine similarity across languages (within-class > between-class, p < 10^-9). Our dataset is restricted to Eastern and Southern Bantu, so we interpret these results as recovering shared Bantu lexical structure consistent with Proto-Bantu rather than definitively distinguishing Proto-Bantu retentions from later regional innovations.

【3】Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
标题:重新思考XAI评估:高风险环境下对Shapley基准的以人为本的审计
链接:https://arxiv.org/abs/2604.22662

作者:Inês Oliveira e Silva,Sérgio Jesus,Iker Perez,Rita P. Ribeiro,Carlos Soares,Hugo Ferreira,Pedro Bizarro
摘要:Shapley值是可解释人工智能的基石,但它们扩散到相互竞争的公式中,造成了一个支离破碎的局面,在实际部署方面几乎没有共识。虽然理论上的差异是有据可查的,但评价仍然依赖于与人类效用一致的量化代理,而这些代理未经核实。在这项工作中,我们使用一个统一的摊销框架,以隔离8个Shapley变量之间的语义差异下的低延迟约束的操作风险工作流。我们在四个风险数据集和一个现实的欺诈检测环境中进行了大规模的实证评估,涉及专业分析师和3,735个案例审查。我们的研究结果揭示了一个根本性的不一致:标准的定量指标,如稀疏性和忠诚度,与人类感知的清晰度和决策效用脱钩。此外,虽然没有公式提高客观分析师的表现,解释一贯增加决策的信心,标志着在高风险设置的自动化偏见的关键风险。这些研究结果表明,目前的评估代理不足以预测下游人类的影响,我们提供了基于证据的指导,在业务决策系统中选择配方和指标。
摘要 :Shapley values are a cornerstone of explainable AI, yet their proliferation into competing formulations has created a fragmented landscape with little consensus on practical deployment. While theoretical differences are well-documented, evaluation remains reliant on quantitative proxies whose alignment with human utility is unverified. In this work, we use a unified amortized framework to isolate semantic differences between eight Shapley variants under the low-latency constraints of operational risk workflows. We conduct a large-scale empirical evaluation across four risk datasets and a realistic fraud-detection environment involving professional analysts and 3,735 case reviews. Our results reveal a fundamental misalignment: standard quantitative metrics, such as sparsity and faithfulness, are decoupled from human-perceived clarity and decision utility. Furthermore, while no formulation improved objective analyst performance, explanations consistently increased decision confidence, signaling a critical risk of automation bias in high-stakes settings. These findings suggest that current evaluation proxies are insufficient for predicting downstream human impact, and we provide evidence-based guidance for selecting formulations and metrics in operational decision systems.

【4】Associativity-Peakiness Metric for Contingency Tables
标题:权宜表的关联度峰值指标
链接:https://arxiv.org/abs/2604.22655

作者:Naomi E. Zirkind,William J. Diehl
备注:38 pages, 21 figures
摘要:对于比较输出为列联表的聚类算法的性能的用例,需要列联表的单个性能度量。这样的指标对于聚类算法的比较性能分析至关重要。对公开文献的调查没有显示存在这样一个指标。真值和预测值的向量对确实存在,这是聚类算法输出的另一种形式。然而,向量对的度量并没有揭示列联表中明显存在的详细特征。本文提出了关联性峰值(AP)度量,它的特点方面的聚类算法的性能是预测部署时的聚类算法的性能至关重要。AP度量类似于作为监督学习算法的输出的混淆矩阵的质量度量。本文介绍了500列联表生成多个测试方案的模拟结果。结果表明,对于评估聚类算法的用例,AP度量表征了比公开可用度量具有更高动态范围的列联表的性能,并且它在计算上比可比的公开可用度量更有效。
摘要:For the use case of comparing the performance of clustering algorithms whose output is a contingency table, a single performance metric for contingency tables is needed. Such a metric is vital for comparative performance analysis of clustering algorithms. A survey of publicly available literature did not show the presence of such a metric. Metrics do exist for vector pairs of truth values and predicted values, which are an alternative form of output of clustering algorithms. However, the metrics for vector pairs do not reveal the presence of detailed features that are apparent in contingency tables. This paper presents the Associativity Peakiness (AP) metric, which characterizes aspects of clustering algorithm performance that are critical for predicting a clustering algorithm's performance when deployed. The AP metric is analogous to measures of quality for confusion matrices that are outputs of supervised learning algorithms. This paper presents results from simulations in which 500 contingency tables were generated for multiple test scenarios. The results show that for the use case of evaluating clustering algorithms, the AP metric characterizes performance of contingency tables with higher dynamic range than publicly available metrics, and that it is computationally more efficient than comparable publicly available metrics.

【5】Decoding High-Dimensional Finger Motion from EMG Using Riemannian Features and RNNs
标题:使用Riemann特征和RNN从EMG解码多维手指运动
链接:https://arxiv.org/abs/2604.22499

作者:Martin Colot,Cédric Simar,Guy Cheron,Ana Maria Cebolla Alvarez,Gianluca Bontempi
备注:13 pages, 10 figures, 3 tables, links to a GitHub, a dataset on Zenodo, and two videos on YouTube
摘要:从前臂表面肌电图(EMG)连续估计高维手指运动学可以实现对手部假肢,AR/XR接口和远程操作的自然控制。然而,人类手势的复杂性和前臂肌肉的纠缠使得准确识别具有内在的挑战性。现有方法通常通过依赖于基于分类的机器学习来降低任务复杂性,限制可控的自由度并损害自然交互。我们提出了一个端到端的框架,连续肌电运动学回归只使用消费级硬件。该框架结合了8通道EMG臂章,单个网络摄像头和自动同步程序,从而能够收集EMG手指运动数据集(EMG-FK),同步EMG的10小时数据集和来自20名参与者的15个手指关节角度,这些参与者执行丰富,不受约束的右手运动。我们还介绍了时间黎曼回归(TRR),一个轻量级的基于GRU的模型,使用多波段黎曼协方差特征序列来解码手指运动。在EMG-FK和公共emg 2 pose基准测试中,TRR在受试者内和跨受试者评估中的表现优于最先进的方法。在EMG FK上,其平均绝对误差在受试者内为$9.79 °\pm 1.48$,在跨受试者中为$16.71 °\pm 3.97$。最后,我们展示了在Raspberry Pi 5上的实时部署和对机器人手的直观控制; TRR的运行速度接近10次预测/秒,比最先进的方法快了大约一个数量级。总之,这些贡献降低了高维手指运动的可再现的、实时的基于EMG的解码的障碍,并为基于EMG的嵌入式系统的更自然和直观的控制铺平了道路。
摘要:Continuous estimation of high-dimensional finger kinematics from forearm surface electromyography (EMG) could enable natural control for hand prostheses, AR/XR interfaces, and teleoperation. However, the complexity of human hand gestures and the entanglement of forearm muscles make accurate recognition intrinsically challenging. Existing approaches typically reduce task complexity by relying on classification-based machine learning, limiting the controllable degrees of freedom and compromising on natural interaction. We present an end-to-end framework for continuous EMG-to-kinematics regression using only consumer-grade hardware. The framework combines an 8-channel EMG armband, a single webcam, and an automatic synchronization procedure, enabling the collection of the EMG Finger-Kinematics dataset (EMG-FK), a 10-h dataset of synchronized EMG and 15 finger joint angles from 20 participants performing rich, unconstrained right-hand motions. We also introduce the Temporal Riemannian Regressor (TRR), a lightweight GRU-based model that uses sequences of multi-band Riemannian covariance features to decode finger motion. Across EMG-FK and the public emg2pose benchmark, TRR outperforms state-of-the-art methods in both intra- and cross-subject evaluation. On EMG-FK, it reaches an average absolute error of $9.79 °\pm 1.48$ in intra-subject and $16.71 °\pm 3.97$ in cross-subject. Finally, we demonstrate real-time deployment on a Raspberry Pi 5 and intuitive control of a robotic hand; TRR runs at nearly 10 predictions/s and is roughly an order of magnitude faster than state-of-the-art approaches. Together, these contributions lower the barrier to reproducible, real-time EMG-based decoding of high-dimensional finger motion, and pave the way toward more natural and intuitive control of embedded EMG-based systems.

【6】Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
标题:对比语义投射:用对比例子标记忠实神经元
链接:https://arxiv.org/abs/2604.22477

作者:Oussama Bouanani,Jim Berend,Wojciech Samek,Sebastian Lapuschkin,Maximilian Dreyer
摘要:神经元标记将文本描述分配给深度网络的内部单元。现有的方法通常依赖于高度活跃的例子,往往产生广泛的或误导性的标签,专注于占主导地位的,但偶然的视觉因素。先前的工作,如ESTCON,引入了对比示例-语义上类似于激活示例但引起低激活的输入-以锐化解释,但它主要解决子空间级别的可解释性,而不是可扩展的神经元级别标记。我们在两个阶段重新审视了神经元级别标记的对比解释:(1)视觉语言模型(VLM)的候选标签生成和(2)CLIP类编码器的标签分配。首先,我们表明,提供对比图像集的VLMs产生候选标签,更具体,更忠实。其次,我们介绍了对比语义投影(CSP),SemanticLens的扩展,将对比示例直接纳入其基于CLIP的评分和选择管道。在广泛的实验和黑色素瘤检测的案例研究中,对比标记在最先进的基线上提高了忠诚度和语义粒度。我们的研究结果表明,对比示例是神经元标记和分析管道的一个简单而强大的组件,目前尚未得到充分利用。
摘要:Neuron labeling assigns textual descriptions to internal units of deep networks. Existing approaches typically rely on highly activating examples, often yielding broad or misleading labels by focusing on dominant but incidental visual factors. Prior work such as FALCON introduced contrastive examples -- inputs that are semantically similar to activating examples but elicit low activations -- to sharpen explanations, but it primarily addresses subspace-level interpretability rather than scalable neuron-level labeling. We revisit contrastive explanations for neuron-level labeling in two stages: (1) candidate label generation with vision language models (VLMs) and (2) label assignment with CLIP-like encoders. First, we show that providing contrastive image sets to VLMs yields candidate labels that are more specific and more faithful. Second, we introduce Contrastive Semantic Projection (CSP), an extension of SemanticLens that incorporates contrastive examples directly into its CLIP-based scoring and selection pipeline. Across extensive experiments and a case study on melanoma detection, contrastive labeling improves both faithfulness and semantic granularity over state-of-the-art baselines. Our results demonstrate that contrastive examples are a simple yet powerful and currently underutilized component of neuron labeling and analysis pipelines.

【7】All Eyes on the Workflow: Automated and Efficient Event Discovery from Video Streams
标题:所有人都关注工作流程:从视频流中自动有效地发现事件
链接:https://arxiv.org/abs/2604.22476

作者:Marco Pegoraro,Jonas Seng,Dustin Heller,Wil M. P. van der Aalst,Kristian Kersting
备注:17 pages, 6 figures, 1 table, 23 references
摘要:诸如业务流程管理和流程挖掘之类的工具通过在记录的事件数据的基础上发现有关流程的见解来帮助组织。然而,流程分析的一个障碍是数据的多模态性:例如,视频形式的数据不能直接解释为事件。在这项工作中,我们提出了SnapLog,从视频中提取事件数据的方法,通过使用图像嵌入将帧转换为特征向量,并通过逐帧相似性矩阵进行时间分割。然后使用广义的Few-Shot分类来向视频片段分配标签,从而产生可解释为事件的带标签的、带时间戳的帧子序列。传统的过程挖掘技术可以用来分析结果数据。我们表明,我们的方法产生的日志准确地反映了视频中的过程。
摘要 :Disciplines such as business process management and process mining aid organizations by discovering insights about processes on the basis of recorded event data. However, an obstacle to process analysis is data multi-modality: for instance, data in video form are not directly interpretable as events. In this work, we present SnapLog, an approach to extract event data from videos by converting frames to feature vectors using image embeddings and performing temporal segmentation through frame-wise similarity matrices. A generalized few-shot classification is then used to assign labels to the video segments, yielding labeled, timestamped sub-sequences of frames that are interpretable as events. Conventional process mining techniques can be used to analyze the resulting data. We show that our approach produces logs that accurately reflect the process in the videos.

【8】Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents
标题:超级思维测试:通过探测代理积极评估代理人社会的集体智能
链接:https://arxiv.org/abs/2604.22452

作者:Xirui Li,Ming Li,Yunze Xiao,Ryan Wong,Dianqi Li,Timothy Baldwin,Tianyi Zhou
摘要:集体智慧指的是一个群体取得超出任何个人成员单独完成的成果的能力。随着大型语言模型代理扩展到数百万人口,一个关键问题出现了:集体智能是否会自发地从规模中出现?我们提出了这个问题的第一个实证评估在一个大规模的自主代理社会。通过研究MoltBook,一个拥有超过200万个代理的平台,我们引入了Superminds Test,这是一个分层框架,使用控制的探测代理在三个层次上探测社会级智能:联合推理,信息合成和基本交互。我们的实验揭示了集体智慧的严重缺失。社会未能超越个人的前沿模型在复杂的推理任务,很少综合分布式信息,往往失败,甚至微不足道的协调任务。平台范围内的分析进一步表明,交互仍然很浅,线程很少超出单个回复,大多数回复都是通用的或离题的。这些结果表明,集体智慧并不仅仅来自规模。相反,当前代理社会的主要限制是非常稀疏和浅的互动,这阻止了代理交换信息和建立在彼此的输出。
摘要:Collective intelligence refers to the ability of a group to achieve outcomes beyond what any individual member can accomplish alone. As large language model agents scale to populations of millions, a key question arises: Does collective intelligence emerge spontaneously from scale? We present the first empirical evaluation of this question in a large-scale autonomous agent society. Studying MoltBook, a platform hosting over two million agents, we introduce Superminds Test, a hierarchical framework that probes society-level intelligence using controlled Probing Agents across three tiers: joint reasoning, information synthesis, and basic interaction. Our experiments reveal a stark absence of collective intelligence. The society fails to outperform individual frontier models on complex reasoning tasks, rarely synthesizes distributed information, and often fails even trivial coordination tasks. Platform-wide analysis further shows that interactions remain shallow, with threads rarely extending beyond a single reply and most responses being generic or off-topic. These results suggest that collective intelligence does not emerge from scale alone. Instead, the dominant limitation of current agent societies is extremely sparse and shallow interaction, which prevents agents from exchanging information and building on each other's outputs.

【9】Algorithmic Feature Highlighting for Human-AI Decision-Making
标题:人机智能决策的数学特征凸显
链接:https://arxiv.org/abs/2604.22236

作者:Yifan Guo,Jann Spiess
摘要:人类决策者经常面临复杂案例的选择,这些案例具有许多潜在的相关特征,但有限的带宽可以检查和整合所有可用信息。在这种情况下,我们研究的算法突出了一个小的子集的情况下特定的功能,为人类考虑,而不是产生一个单一的预测或建议。我们的模型突出显示作为一个受约束的信息政策,选择了少量的功能来揭示。一个核心问题是人类如何解释算法的功能选择:一个复杂的代理正确的条件选择规则,而天真的代理更新只显示的功能值,并把选择事件作为外源。我们表明,优化突出复杂的代理可以是计算上棘手的,即使在简单的离散和二进制设置,而优化幼稚的代理是易于处理的,只要最大带宽是固定的。我们还表明,一个突出的政策,是最佳的复杂的代理可以任意执行时,天真的代理,激励强大的,可实施的替代品。我们说明了我们的框架,在校准的实证练习的基础上,美国住房调查。总的来说,我们的研究结果确立了突出特定于上下文的特征集而不是固定特征集的价值,作为实现人类算法互补性的实用吸引力和计算可行的工具。
摘要:Human decision-makers often face choices about complex cases with many potentially relevant features, but limited bandwidth to inspect and integrate all available information. In such settings, we study algorithms that highlight a small subset of case-specific features for human consideration, rather than producing a single prediction or recommendation. We model highlighting as a constrained information policy that selects a small number of features to reveal. A central issue is how humans interpret the algorithm's choice of features: a sophisticated agent correctly conditions on the selection rule, while a naive agent updates only on revealed feature values and treats the selection event as exogenous. We show that optimizing highlighting for sophisticated agents can be computationally intractable, even in simple discrete and binary settings, whereas optimizing for naive agents is tractable as long as the maximal bandwidth is fixed. We also show that a highlighting policy that is optimal for sophisticated agents can perform arbitrarily poorly when deployed to naive agents, motivating robust, implementable alternatives. We illustrate our framework in a calibrated empirical exercise based on the American Housing Survey. Overall, our results establish the value of highlighting a context-specific set of features rather than a fixed one as a practically appealing and computationally feasible tool for achieving human-algorithm complementarity.

【10】Logistic Bandits with $\tilde{O}(\sqrt{dT})$ Regret without Context Diversity Assumptions
链接:https://arxiv.org/abs/2604.22161

作者:Seoungbin Bae,Dabeen Lee
摘要:我们研究$K$-武装物流强盗问题,在每一轮,代理观察$K$特征向量与$K$行动。现有的方法,实现率最优的$\tilde{\mathcal{O}}(\sqrt{dT})$遗憾界严重依赖于上下文多样性的假设,如严格的积极性的上下文协方差矩阵的最小特征值。然而,这些假设对上下文过程施加了很强的限制,因为它们排除了上下文向量集中在低维子空间中的情况。在本文中,我们提出了SupSplitLog,据我们所知,这是第一个算法的物流强盗,实现$\tilde{\mathcal{O}}(\sqrt{dT})$遗憾没有任何上下文多样性的假设。其关键思想是在构造估计量时将收集的样本分成两个不相交的子集;一个用于计算初始点估计量,而另一个用于应用牛顿型一步校正过程。分裂规则是精心设计的,以平衡的精度要求的初始点估计和一步校正过程。此外,SupSplitLog严格改进了现有的算法在尺寸$d$的遗憾上限的依赖。此外,SupSplitLog可以简单地适用于推导出一个遗憾的界限,增长与数据相关的复杂性措施,避免直接依赖于$d$,这是有利的,当上下文向量集中在一个低维子空间。我们还提供了实验结果,数值证明我们的算法的优越性,验证了理论结果。
摘要:We study the $K$-armed logistic bandit problem, where at each round, the agent observes $K$ feature vectors associated with $K$ actions. Existing approaches that achieve a rate-optimal $\tilde{\mathcal{O}}(\sqrt{dT})$ regret bound rely heavily on context diversity assumptions, such as strict positivity of the minimum eigenvalue of a context covariance matrix. These assumptions, however, impose strong restrictions on the context process, as they rule out the situation where the context vectors are concentrated in a low-dimensional subspace. In this paper, we propose SupSplitLog, which, to the best of our knowledge, is the first algorithm for logistic bandits that achieves $\tilde{\mathcal{O}}(\sqrt{dT})$ regret without any context diversity assumption. The key idea is to split the collected samples into two disjoint subsets when constructing estimators; one is used to compute an initial-point estimator, while the other is used to apply a Newton-type one-step correction procedure. The splitting rule is carefully designed to balance the accuracy requirements of the initial-point estimator and the one-step correction procedure. Moreover, SupSplitLog strictly improves on the existing algorithms in terms of the dependence on dimension $d$ in the regret upper bound. Furthermore, SupSplitLog can be adapted simply to deduce a regret bound that grows with a data-dependent complexity measure, avoiding a direct dependence on $d$, which is favorable when the context vectors are concentrated in a low-dimensional subspace. We also provide experimental results that demonstrate numerically the superiority of our algorithm, validating the theoretical results.

【11】PrivUn: Unveiling Latent Ripple Effects and Shallow Forgetting in Privacy Unlearning
标题:PrivUn:揭露隐私遗忘中的潜在涟漪效应和肤浅遗忘
链接:https://arxiv.org/abs/2604.22076

作者:Xiaoyi Chen,Haoyuan Wang,Siyuan Tang,Sijia Liu,Liya Su,XiaoFeng Wang,Haixu Tang
摘要:大型语言模型(LLM)经常在训练过程中记住私人信息,这引发了严重的隐私问题。虽然机器非学习已经成为一种很有前途的解决方案,但它对隐私攻击的真正有效性仍不清楚。为了解决这个问题,我们提出了PrivUn,一个新的评估框架,通过三层攻击场景系统地评估unlearning鲁棒性:直接检索,上下文学习恢复和微调恢复;结合使用遗忘分数,关联度量和遗忘深度评估的定量分析。我们的研究揭示了当前遗忘方法的重大弱点,揭示了两个关键发现:1)遗忘表现出梯度驱动的涟漪效应:与遵循语义关系的传统遗忘不同(例如,知识图),隐私非学习跨潜在的基于梯度的关联传播;以及2)大多数方法遭受浅遗忘,无法移除分布在多个深模型层上的私有信息。为了验证这些见解,我们探索了两种策略:利用梯度相似性的关联感知核心集选择,以及通过代表性约束的多层深度干预。这些策略代表了从浅遗忘到深遗忘的范式转变。
摘要 :Large language models (LLMs) often memorize private information during training, raising serious privacy concerns. While machine unlearning has emerged as a promising solution, its true effectiveness against privacy attacks remains unclear. To address this, we propose PrivUn, a new evaluation framework that systematically assesses unlearning robustness through three-tier attack scenarios: direct retrieval, in-context learning recovery, and fine-tuning restoration; combined with quantitative analysis using forgetting scores, association metrics, and forgetting depth assessment. Our study exposes significant weaknesses in current unlearning methods, revealing two key findings: 1) unlearning exhibits gradient-driven ripple effects: unlike traditional forgetting which follows semantic relations (e.g., knowledge graphs), privacy unlearning propagates across latent gradient-based associations; and 2) most methods suffer from shallow forgetting, failing to remove private information distributed across multiple deep model layers. To validate these insights, we explore two strategies: association-aware core-set selection that leverages gradient similarity, and multi-layer deep intervention through representational constraints. These strategies represent a paradigm shift from shallow forgetting to deep forgetting.

【12】EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms
标题:EgoMAGIC-用于训练感知算法的以自我为中心的视频领域医学数据集
链接:https://arxiv.org/abs/2604.22036

作者:Brian VanVoorst,Nicholas Walczak,Christopher Gilleo,Charles Meissner,Fabio Felix,Iran Roman,Bea Steers,Claudio Silva,Yuhan Shen,Zijia Lu,Shih-Po Lee,Ehsan Elhamifar
备注:9 pages, 4 figures, 3 tables
摘要:本文介绍了EgoMAGIC(医疗援助,指导,指导和纠正),一个以自我为中心的医疗活动数据集收集的DARPA的感知使能任务指导(PTG)计划的一部分。该数据集包括50个医疗任务的3,355个视频,每个任务至少有50个标记视频。PTG计划的主要目标是开发集成到增强现实耳机中的虚拟助手,以帮助用户执行复杂的任务。 为了鼓励使用该数据集进行探索和研究,医疗训练数据已经发布,同时发布的还有一个动作检测挑战,重点是八项医疗任务。大部分视频都是使用带有集成音频的头戴式立体摄像机录制的。从这个数据集中,使用195万个标签训练了40个YOLO模型,以检测124个医疗对象,为开发医疗AI应用程序的开发人员提供了一个强大的起点。 除了介绍数据集外,本文还介绍了三种模型中选定的八项医疗任务的动作检测基线结果,其中性能最好的方法实现了平均mAP 0.526。虽然本文主要将动作检测作为基准,但EgoMAGIC数据集同样适用于动作识别,对象识别和检测,错误检测以及其他具有挑战性的计算机视觉任务。 该数据集可通过zenodo.org(DOI:10.5281/zenodo.19239154)访问。
摘要:This paper introduces EgoMAGIC (Medical Assistance, Guidance, Instruction, and Correction), an egocentric medical activity dataset collected as part of DARPA's Perceptually-enabled Task Guidance (PTG) program. This dataset comprises 3,355 videos of 50 medical tasks, with at least 50 labeled videos per task. The primary objective of the PTG program was to develop virtual assistants integrated into augmented reality headsets to assist users in performing complex tasks. To encourage exploration and research using this dataset, the medical training data has been released along with an action detection challenge focused on eight medical tasks. The majority of the videos were recorded using a head-mounted stereo camera with integrated audio. From this dataset, 40 YOLO models were trained using 1.95 million labels to detect 124 medical objects, providing a robust starting point for developers working on medical AI applications. In addition to introducing the dataset, this paper presents baseline results on action detection for the eight selected medical tasks across three models, with the best-performing method achieving average mAP 0.526. Although this paper primarily addresses action detection as the benchmark, the EgoMAGIC dataset is equally suitable for action recognition, object identification and detection, error detection, and other challenging computer vision tasks. The dataset is accessible via zenodo.org (DOI: 10.5281/zenodo.19239154).

【13】Kernel Contracts: A Specification Language for ML Kernel Correctness Across Heterogeneous Silicon
标题:核心契约:跨异类硅的ML核心正确性规范语言
链接:https://arxiv.org/abs/2604.22032

作者:Cooper Veit
备注:28 pages, 1 figure
摘要:每个ML内核都附带一个关于它所计算的内容的隐式契约。人们很少把合同写下来。当两个内核不一致时--当AMD上的matmul产生的梯度与NVIDIA上的相同matmul不同时,当融合注意力内核默默地向下转换累加器时,当越界访问在一个堆栈上返回零而在另一个堆栈上返回垃圾时--没有正式的工件来仲裁争议。最近的实证工作测量了硅平台之间的差距,但没有具体说明违反的合同。 我们提出了一个规范语言的内核合同。合同有八个部分:标识符、作用域、前置条件、后置条件、容差、引用oracle、测量协议和违规签名。我们用它来说明12个合同类,涵盖精度,排序,编译器引起的,和异常值故障模式,每一个接地在已发表的经验证据。我们需要一个三态校准:每个契约必须承认至少一个符合引用的实现和至少一个通过基本功能测试的违反契约的实现。我们将该框架应用于三个记录在案的事件-华为Ascend沉默精确胁迫,Sakana AI CUDA工程师奖励黑客攻击,AMD越界沉默接受-并显示每个非正式诊断都映射到具有可测量签名的特定合同违规行为。内核合同套件是一个标准参考,可以根据它对一致性进行分级,就像ISASecure根据IEC 62443对工业控制系统进行分级一样。
摘要:Every ML kernel ships with an implicit contract about what it computes. People rarely write the contract down. When two kernels disagree -- when a matmul on AMD produces a different gradient than the same matmul on NVIDIA, when a fused attention kernel silently downcasts an accumulator, when an out-of-bounds access returns zero on one stack and garbage on another -- there is no formal artifact to arbitrate the dispute. Recent empirical work has measured the gap across silicon platforms, but none of it specifies the contract being violated. We present a specification language for kernel contracts. A contract has eight parts: identifier, scope, precondition, postcondition, tolerance, reference oracle, measurement protocol, and violation signature. We use it to state twelve contract classes covering precision, ordering, compiler-induced, and exceptional-value failure modes, each grounded in published empirical evidence. We require a three-state calibration: every contract must admit at least one reference-conforming implementation and at least one contract-violating implementation that passes basic functional tests. We apply the framework to three documented incidents -- Huawei Ascend silent precision coercion, Sakana AI CUDA Engineer reward hacking, AMD out-of-bounds silent acceptance -- and show that each informal diagnosis maps to a specific contract violation with a measurable signature. A kernel contract suite is a normative reference against which conformance can be graded, in the way that ISASecure grades industrial control systems against IEC 62443.

【14】AI-based framework to predict animal and pen feed intake in feedlot beef cattle
标题:基于人工智能的框架预测饲养场肉牛的动物和围栏饲料摄入量
链接:https://arxiv.org/abs/2511.17663

作者:Alex S. C. Maia,John B. Hall,Hugo F. M. Milan,Izabelle A. M. A. Teixeira
摘要:技术的进步正在改变可持续的养牛实践,电子饲养系统生成关于个体动物饲料摄入量的大型纵向数据集,为自动精确畜牧系统提供了可能性。然而,文献仍然缺乏充分利用这些纵向大数据来准确预测环境条件下的饲料摄入量的方法。为了填补这一空白,我们开发了一个基于人工智能的框架,以准确预测个体动物的摄食量和围栏水平的聚集。数据来自在Nancy M.进行的19项实验(> 1650万份样本; 2013-2024年)。卡明斯研究推广与教育中心(Carmen,ID)饲养场设施和来自AgriMet网络气象站的环境数据被用于开发两种新的环境指数:InComfort-Index,仅基于气象变量,显示出对热舒适性的良好预测能力,但预测采食量的能力有限; EASI-Index是一个综合环境变量和采食行为的混合指数,在预测采食量方面表现良好,但对热舒适性的预测效果较差。与环境指数一起,对机器学习模型进行了训练,最佳机器学习模型(XGBoost)的准确度为动物水平的RMSE为1.38 kg/天,围栏水平的RMSE仅为0.14 kg/(天-动物)。这种方法提供了一个强大的基于人工智能的框架,用于预测个体动物和围栏的饲料摄入量,通过减少饲料浪费,资源优化和气候适应性牲畜管理,在饲养场牛的精确管理中具有潜在的应用。
摘要:Advances in technology are transforming sustainable cattle farming practices, with electronic feeding systems generating big longitudinal datasets on individual animal feed intake, offering the possibility for autonomous precision livestock systems. However, the literature still lacks a methodology that fully leverages these longitudinal big data to accurately predict feed intake accounting for environmental conditions. To fill this gap, we developed an AI-based framework to accurately predict feed intake of individual animals and pen-level aggregation. Data from 19 experiments (>16.5M samples; 2013-2024) conducted at Nancy M. Cummings Research Extension & Education Center (Carmen, ID) feedlot facility and environmental data from AgriMet Network weather stations were used to develop two novel environmental indices: InComfort-Index, based solely on meteorological variables, showed good predictive capability for thermal comfort but had limited ability to predict feed intake; EASI-Index, a hybrid index integrating environmental variables with feed intake behavior, performed well in predicting feed intake but was less effective for thermal comfort. Together with the environmental indices, machine learning models were trained and the best-performing machine learning model (XGBoost) accuracy was RMSE of 1.38 kg/day for animal-level and only 0.14 kg/(day-animal) at pen-level. This approach provides a robust AI-based framework for predicting feed intake in individual animals and pens, with potential applications in precision management of feedlot cattle, through feed waste reduction, resource optimization, and climate-adaptive livestock management.

【15】The Exact Replica Threshold for Nonlinear Moments of Quantum States
标题:量子态非线性矩的精确阈值
链接:https://arxiv.org/abs/2604.22627

作者:Shuai Zeng
摘要 :对量子态的多个副本的联合测量提供了对诸如$\operatorname{tr}(ρ^t)$之类的非线性可观测量的访问,但是副本数量是否标志着一个尖锐的信息论资源边界仍然不清楚。对于每一个固定的顺序$t\ge 3$,现有的协议表明,$\lceil t/2\rceil$副本已经足够的多项式样本估计的$\operatorname{tr}(ρ^t)$,但它仍然开放是否一个副本必须引起一个样本复杂性的障碍,随着维度的增长。我们证明,这确实是在样本/副本访问模型与复制有限的联合测量的情况下:任何协议限制到$\lceil t/2\rceil-1$副本需要维数增长的样本复杂性,而$\lceil t/2\rceil$副本足够的先前的工作。因此,固定阶纯矩的精确副本阈值是$\lceil t/2\rceil$。同样,对于固定阶纯矩,一个额外的相干副本不仅是有用的,但标志着多项式样本估计和复制限制模型中的维数增长制度之间的确切阈值。我们进一步表明,相同的阈值法扩展到一个广泛的家庭的可观加权矩$\operatorname{tr}(Oρ^t)$,包括泡利可观和其他可观的有界算子范数和宏观迹范数。因此,相干副本数作为一个真正的非线性量子状态估计的离散资源。
摘要:Joint measurements on multiple copies of a quantum state provide access to nonlinear observables such as $\operatorname{tr}(ρ^t)$, but whether replica number marks a sharp information-theoretic resource boundary has remained unclear. For every fixed order $t\ge 3$, existing protocols show that $\lceil t/2\rceil$ replicas already suffice for polynomial-sample estimation of $\operatorname{tr}(ρ^t)$, yet it has remained open whether one fewer replica must necessarily incur a sample-complexity barrier growing with the dimension. We prove that this is indeed the case in the sample/copy-access model with replica-limited joint measurements: any protocol restricted to $\lceil t/2\rceil-1$ replicas requires dimension-growing sample complexity, while $\lceil t/2\rceil$ replicas suffice by prior work. Thus the exact replica threshold for fixed-order pure moments is $\lceil t/2\rceil$. Equivalently, for fixed-order pure moments, one additional coherent replica is not merely useful but marks the exact threshold between polynomial-sample estimation and a dimension-growing regime in the replica-limited model. We further show that the same threshold law extends to a broad family of observable-weighted moments $\operatorname{tr}(Oρ^t)$, including Pauli observables and other observables with bounded operator norm and macroscopic trace norm. Coherent replica number therefore acts as a genuinely discrete resource for nonlinear quantum-state estimation.

【16】Conformalized Super Learner
标题:符合规范的超级学习者
链接:https://arxiv.org/abs/2604.22391

作者:Zhanli Wu,Fabrizio Leisen,Miguel-Angel Luque-Fernandez,F. Javier Rubio
备注:R codes and data can be found at: https://github.com/ZWU-001/CSL
摘要:超级学习器(SL)是一种广泛使用的集成方法,它根据学习器库的预测性能组合预测。区间预测具有相当大的实际意义,因为它们允许量化单个学习者或集合产生的预测中的不确定性。已经提出了几种方法来构建基于SL的区间预测,然而,这些方法通常是合理的使用渐近参数或依赖于计算密集型的程序,如自举。共形预测(CP)是一种机器学习框架,用于在温和条件下构建具有有限样本和渐近覆盖保证的预测区间。我们建议通过反映原始SL框架的自然构造将CP与SL耦合,使用个人学习者权重并通过加权多数投票结合特定于学习者的一致性得分。我们的特性所产生的SL为基础的预测区间连续的结果。我们涵盖的设置下exchangeries,潜在的违反exchangeries,和数据生成机制表现出异方差,稀疏性,和其他形式的分布异质性。一个全面的模拟研究表明,整合SL实现有效的有限样本覆盖率与竞争力的性能相对于真正的数据生成机制。这项工作的核心贡献是使用社会人口统计学,生物统计学和实验室测量来预测肌酐水平。这个例子展示了精心挑选的学习器集成的好处,这些学习器旨在捕捉复杂回归函数的关键方面,包括非线性效应、交互作用、稀疏性、异方差性和对离群值的鲁棒性。
摘要:The Super Learner (SL) is a widely used ensemble method that combines predictions from a library of learners based on their predictive performance. Interval predictions are of considerable practical interest because they allow uncertainty in predictions produced by an individual learner or an ensemble to be quantified. Several methods have been proposed for constructing interval predictions based on the SL, however, these approaches are typically justified using asymptotic arguments or rely on computationally intensive procedures such as the bootstrap. Conformal prediction (CP) is a machine learning framework for constructing prediction intervals with finite-sample and asymptotic coverage guarantees under mild conditions. We propose coupling CP with the SL through a natural construction that mirrors the original SL framework, using individual learner weights and combining learner-specific conformity scores via a weighted majority vote. We characterize the properties of the resulting SL-based prediction intervals for continuous outcomes. We cover settings under exchangeability, potential violations of exchangeability, and data-generating mechanisms exhibiting heteroscedasticity, sparsity, and other forms of distributional heterogeneity. A comprehensive simulation study shows that the conformalized SL achieves valid finite-sample coverage with competitive performance relative to the true data-generating mechanism. A central contribution of this work is an application to predicting creatinine levels using socio-demographic, biometric, and laboratory measurements. This example demonstrates the benefits of an ensemble with carefully selected learners designed to capture key aspects of complex regression functions, including non-linear effects, interactions, sparsity, heteroscedasticity, and robustness to outliers.R

【17】Pliable rejection sampling
标题:容易拒绝抽样
链接:https://arxiv.org/abs/2604.22385

作者:Akram Erraqabi,Michal Valko,Alexandra Carpentier,Odalric-Ambrym Maillard
备注:In ICML 2016
摘要:拒绝抽样是一种从困难分布中抽样的技术。然而,由于拒绝率高,其使用受到限制。常见的自适应拒绝采样方法要么只适用于非常特定的分布,要么没有性能保证。在本文中,我们提出了柔韧拒绝抽样(PRS),拒绝抽样的一种新方法,在那里我们学习的抽样建议使用核估计。由于我们的方法建立在拒绝抽样的基础上,因此获得的样本具有很高的独立同分布概率。并按f分布。此外,PRS还保证接受的样品数量。
摘要:Rejection sampling is a technique for sampling from difficult distributions. However, its use is limited due to a high rejection rate. Common adaptive rejection sampling methods either work only for very specific distributions or without performance guarantees. In this paper, we present pliable rejection sampling (PRS), a new approach to rejection sampling, where we learn the sampling proposal using a kernel estimator. Since our method builds on rejection sampling, the samples obtained are with high probability i.i.d. and distributed according to f. Moreover, PRS comes with a guarantee on the number of accepted samples.

【18】Multimodal Diffusion to Mutually Enhance Polarized Light and Low Resolution EBSD Data
标题:多峰扩散相互增强极化光和低分辨率EBSD数据
链接:https://arxiv.org/abs/2604.22212

作者:Harry Dong,Timofey Efimov,Megna Shah,Jeff Simmons,Sean Donegan,Marc De Graef,Yuejie Chi
摘要:尽管3-D电子背散射衍射(EBSD)显微镜的效用,数据收集过程可以是耗时的连续切片。因此,自然会考虑其他方式,如偏振光(PL)数据,以加速EBSD数据收集,并辅以共享信息。作为补充,混沌PL数据中的特征甚至可以用少数EBSD测量来丰富。为了从本质上了解EBSD和PL之间的复杂动态来解决这些逆问题,我们使用了一个无条件的多模态扩散模型,这是由逆问题扩散模型的进展所驱动的。虽然只在合成数据上训练过一次,但我们的模型在真实数据上具有很强的泛化能力,这些数据可能是低分辨率、噪声、损坏和配准错误的。通过推理时间缩放,我们在各种目标上都表现出了性能上的提高,包括晶界预测、超分辨率和去噪。使用我们的模型,我们证明了只有25%(1/4分辨率)的EBSD数据和损坏的PL数据与全分辨率性能几乎没有区别。
摘要:In spite of the utility of 3-D electron back-scattered diffraction (EBSD) microscopy, the data collection process can be time-consuming with serial-sectioning. Hence, it is natural to look at other modalities, such as polarized light (PL) data, to accelerate EBSD data collection, supplemented with shared information. Complementarily, features in chaotic PL data could even be enriched with a handful of EBSD measurements. To inherently learn the complex dynamics between EBSD and PL to solve these inverse problems, we use an unconditional multimodal diffusion model, motivated by progress in diffusion models for inverse problems. Although trained solely on synthetic data once, our model has strong generalizable capabilities on real data which can be low-resolution, noisy, corrupted, and misregistered. With inference-time scaling, we show gains in performance on a variety of objectives including grain boundary prediction, super-resolution, and denoising. With our model, we demonstrate that there is little difference from full resolution performance with only 25% (1/4 the resolution) of EBSD data and corrupted PL data.

【19】Concave Statistical Utility Maximization Bandits via Influence-Function Gradients
标题:通过影响函数子索的凹陷统计效用最大化Bandits
链接:https://arxiv.org/abs/2604.22140

作者:Matías Carrasco,Alejandro Cholaquidis
摘要:我们研究随机多武装土匪的目标是一个统计功能的长期奖励分布,而不是预期的奖励。在温和的连续性假设下,我们证明了无限时域问题归结为对静态混合策略的优化:单纯形上的每个权重向量\(w\)诱导一个混合律\(P^w\),性能由凹效用\(U(w)=\mathfrak U(P^w)\)来衡量。 对于可微统计效用,我们使用影响函数演算从强盗反馈中推导出随机梯度估计量。这导致了一个熵镜像上升算法截断单纯形,通过乘法权重更新和插件的影响函数的估计。我们建立了遗憾的界限,从估计的影响函数所造成的偏差分离镜上升优化误差。该框架是为一般凹分布的公用事业,并说明通过方差和Wasserstein目标,数值实验比较精确和插件的影响功能的实现。
摘要 :We study stochastic multi-armed bandits in which the objective is a statistical functional of the long-run reward distribution, rather than expected reward alone. Under mild continuity assumptions, we show that the infinite-horizon problem reduces to optimizing over stationary mixed policies: each weight vector \(w\) on the simplex induces a mixture law \(P^w\), and performance is measured by the concave utility \(U(w)=\mathfrak U(P^w)\). For differentiable statistical utilities, we use influence-function calculus to derive stochastic gradient estimators from bandit feedback. This leads to an entropic mirror-ascent algorithm on a truncated simplex, implemented through multiplicative-weights updates and plug-in estimates of the influence function. We establish regret bounds that separate the mirror-ascent optimization error from the bias caused by estimating the influence function. The framework is developed for general concave distributional utilities and illustrated through variance and Wasserstein objectives, with numerical experiments comparing exact and plug-in influence-function implementations.

【20】Conditional Diffusion Posterior Alignment for Sparse-View CT Reconstruction
标题:稀疏视图CT重建的条件扩散后部对齐
链接:https://arxiv.org/abs/2604.21960

作者:Luis Barba,Johannes Kirschner,Benjamin Bejar
摘要:计算机断层扫描(CT)是在医疗和工业应用中广泛使用的成像模态。为了限制辐射暴露和测量时间,人们对稀疏视图CT越来越感兴趣,其中投影视图的数量显著减少。深度神经网络在提高稀疏视图CT的重建质量方面表现出很大的潜力,特别是生成扩散模型。然而,这些方法由于以下几个原因而难以扩展到大的3D体积:(i)3D模型的高存储器和计算要求,(ii)缺乏大的3D训练数据集,以及(iii)当在每个切片上独立地使用2D模型时,切片之间的不一致性。我们克服了这些局限性和规模扩散为基础的稀疏视图CT重建到大的三维体积结合条件扩散与明确的数据一致性。我们提出了条件扩散后对齐(CDPA),使可扩展的三维稀疏视图CT重建。2D U-Net扩散模型以初始3D重建为条件,以提高切片间的一致性,并结合数据一致性对齐以匹配测量的投影。对合成和真实锥形束CT(CBCT)数据的实验显示了最先进的性能,消融证实了拟议管道的协同效应。最后,我们证明了同样的原理也可以加强快速去噪U网,以一小部分计算成本产生近扩散质量。
摘要:Computed Tomography (CT) is a widely used imaging modality in medical and industrial applications. To limit radiation exposure and measurement time, there is a growing interest in sparse-view CT, where the number of projection views is significantly reduced. Deep neural networks have shown great promise in improving reconstruction quality in sparse-view CT, especially generative diffusion models. However, these methods struggle to scale to large 3D volumes due to several reasons: (i) the high memory and computational requirements of 3D models, (ii) the lack of large 3D training datasets, and (iii) the inconsistencies across slices when using 2D models independently on each slice. We overcome these limitations and scale diffusion-based sparse-view CT reconstruction to large 3D volumes by combining conditional diffusion with explicit data consistency. We propose Conditional Diffusion Posterior Alignment (CDPA) to enable scalable 3D sparse-view CT reconstruction. A 2D U-Net diffusion model is conditioned on an initial 3D reconstruction to improve inter-slice consistency, combined with data-consistency alignment to match measured projections. Experiments on synthetic and real Cone Beam CT (CBCT) data show state-of-the-art performance, with ablations that confirm the synergistic effects of the proposed pipeline. Finally, we show that the same principles also strengthen fast denoising U-Nets, yielding near-diffusion quality at a fraction of the computational cost.

机器翻译由腾讯交互翻译提供,仅供参考

点击“阅读原文”获取带摘要的学术速递

Python社区是高质量的Python/Django开发社区
本文地址:http://www.python88.com/topic/195552