_              _         ____              
   / \   _ ____  _(_)_   __ |  _ \  __ _ _   _ 
  / _ \ | '__\ \/ / \ \ / / | | | |/ _` | | | |
 / ___ \| |   >  <| |\ V /  | |_| | (_| | |_| |
/_/   \_\_|  /_/\_\_| \_/   |____/ \__,_|\__, |
                                         |___/ 
        

Articles: 0

Last Updated: N/A (+00:00)

Index | Calendar | Favorites | Archive | Profile

QF3: Fast Flow RL with Filtered Q-Gradients

Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with a high-throughput off-policy training recipe, it trains humanoid locomotion and motion-tracking policies with a 10x wall-clock speedup over FPO++, a recent on-policy flow RL method. We further apply QF3 to fine-tune pretrained flow-based manipulation policies on both ABC-Sim and Robomimic tasks. These results suggest that QF3 can both learn robot policies from scratch and refine those acquired from demonstrations. Website: https://qf3-rl.github.io/

Updated: 2026-10-06 17:59:34

标题: QF3:带有过滤Q梯度的快速流动强化学习

摘要: 流水政策已经成为从示范中学习机器人行为的标准政策类别,但强化学习仍然对改进预训练的流水政策或通过互动从头开始学习它们至关重要。我们介绍了QF3(带有过滤Q梯度的快速流水RL),这是一种在线离策略RL算法,通过通过一步预测流的输出反向传播来训练一个流水政策。为了保持评论员和此预测可靠,QF3仅将评论员梯度应用于保持接近重放动作的动作维度。据我们所知,QF3是第一个离策略流RL方法,可以从头开始训练人形 locomotion 策略并将其零次转移到硬件。结合高吞吐量的离策略训练配方,它可以使人类 locomotion 和运动跟踪策略的训练速度比最近的一种基于政策的流RL方法FPO++提高10倍。我们进一步将QF3应用于优化预训练的基于流动的操纵策略,同时还适用于 ABC-Sim 和 Robomimic 任务。这些结果表明,QF3既可以从头开始学习机器人策略,也可以优化从示范中获得的策略。网站:https://qf3-rl.github.io/

更新时间: 2026-10-06 17:59:34

领域: cs.RO,cs.LG

下载: http://arxiv.org/abs/2610.08789v1

Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective

Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty, the information-theoretic basis for this interpretation remains poorly understood. In this work, we provide such a foundation using a decision-theoretic generalization of entropy tailored to set-valued prediction. In particular, we introduce a family of generalized information measures based on the size and coverage of conformal prediction sets. Notably, Shannon mutual information admits an exact integral representation in terms of these measures. We then show that, in standard classification settings, the reduction in conformal set size from additional information (i) is sandwiched between calibration-dependent members of this family and (ii) obeys a data processing inequality, both up to finite-sample calibration and model error terms. Together, our results formally relate conformal prediction to classical information-theoretic quantities and justify using set-size reduction as an information gain metric. Empirically, we validate our theory across 11 classification settings and show that set-size reduction and Shannon mutual information can rank features differently in a greedy feature selection experiment.

Updated: 2026-10-06 17:59:11

标题: 《一种理论视角下的调整预测集量化信息增益》

摘要: Conformal prediction是一种用于不确定性量化的流行工具,它输出具有有限样本覆盖保证的预测集。虽然预测集大小通常被用作不确定性的启发式度量,但对于这种解释的信息理论基础仍然不够清楚。在这项工作中,我们使用适用于集值预测的决策理论泛化的熵,为这样的基础提供了支持。特别地,我们引入了一系列基于conformal prediction集的大小和覆盖范围的广义信息度量。值得注意的是,Shannon互信息可通过这些度量的精确积分表示。然后我们展示,在标准分类设置中,从额外信息中减少conformal set大小(i)被夹在这一族与校准相关的成员之间,(ii)遵循数据处理不等式,都在有限样本校准和模型误差项上。综合起来,我们的结果正式将conformal prediction与经典信息理论量联系起来,并证明了使用集大小减少作为信息增益度量的合理性。在实证方面,我们验证了我们的理论在11个分类设置中的有效性,并展示了集大小减少和Shannon互信息在贪婪特征选择实验中可以以不同方式对特征进行排名。

更新时间: 2026-10-06 17:59:11

领域: cs.LG

下载: http://arxiv.org/abs/2610.08785v1

4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a separate post-hoc optimization after reconstruction, we directly steer the evolving generative states using physical interaction constraints and observed 2D evidence, allowing the reconstruction to be refined as part of the generative process itself. By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios. Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.

Updated: 2026-10-06 17:59:02

标题: 4D-HOF:手-物体流匹配用于前馈4D交互重建

摘要: 现有的4D手部-物体重建方法通常依赖于昂贵的每序列优化,而生成方法通常从随机噪声中合成交互,这可能导致不稳定的交互预测。我们引入了4D-HOF,这是一个前馈框架,从视觉基础模型产生的粗略但信息丰富的估计中重建4D手部-物体交互。具体来说,我们学习了一个条件流匹配模型,将基础模型派生的手部-物体状态传输到交互流形,使模型能够以前馈方式纠正平移、旋转和对齐错误。我们生成式公式的一个关键优势是它自然地在传输过程中启用了测试时间指导。我们不是在重建后应用单独的事后优化,而是直接使用物理交互约束和观察到的2D证据来引导发展中的生成状态,从而使重建成为生成过程本身的一部分。通过在多样化数据集上训练生成模型,4D-HOF在具有挑战性的野外场景中表现出良好的泛化性。在域外基准测试中的实验表明,4D-HOF实现了最先进的性能,产生更稳定和准确的4D手部-物体重建。

更新时间: 2026-10-06 17:59:02

领域: cs.CV,cs.AI,cs.GR

下载: http://arxiv.org/abs/2610.08782v1

IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas

Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation remains challenging, as existing approaches based on prompting or feedback lack structured supervision for how papers should be synthesized. We introduce IdeaAnchor, a paradigm for training LLMs to perform research ideation using structured specifications as privileged signals. Each IdeaAnchor instance encodes how each input paper should be synthesized into a successful idea, including their functional roles, relationships, and target synthesis criteria. We build this paradigm by mining instances from published papers, capturing how real ideas emerge from prior literature. We then train models via demonstration, self-distillation, and reinforcement learning, and further enhance generation with retrieval at inference time. Experiments show consistent improvements in ideation quality. Our analysis reveals a functional decomposition: anchor-based training strengthens creative synthesis, retrieval enhances detail elaboration, and combining both yields the best performance.

Updated: 2026-10-06 17:59:01

标题: IdeaAnchor:教授LLMs将文学转化为研究想法

摘要: 科学研究通常从合成一组相关论文的想法开始,以识别差距并制定新的方向。然而,训练语言模型执行这种基于文献的构想仍然具有挑战性,因为现有的基于提示或反馈的方法缺乏关于如何合成论文的结构化监督。我们引入了IdeaAnchor,这是一种用结构化规范作为特权信号训练LLMs进行研究构想的范式。每个IdeaAnchor实例都编码了每篇输入论文应如何合成为成功想法,包括它们的功能角色、关系和目标合成标准。我们通过从已发表的论文中挖掘实例来构建这一范式,捕捉从先前文献中获得真实想法的方式。然后我们通过示范、自我精馏和强化学习来训练模型,并在推理时进一步通过检索增强生成。实验显示了构想质量的持续改进。我们的分析揭示了一种功能分解:基于锚点的训练增强了创造性合成,检索增强了细节阐述,并将两者结合可以获得最佳表现。

更新时间: 2026-10-06 17:59:01

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.08781v1

DepthWorld: 3D World Model for Robot Manipulation

World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving <0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.

Updated: 2026-10-06 17:59:00

标题: DepthWorld:机器人操作的三维世界模型

摘要: 世界模型为机器人技术提供了一种基于数据驱动的替代方案,应用领域涵盖政策评估、改进和规划。所有这些用途都依赖于忠实的3D几何,然而当前基于视频的世界模型仅在RGB上进行训练,产生的结果在逐帧看起来正确,但无法组成一致的3D世界。要弥补这一差距,需要在两个方面取得进展:为操纵提供大规模的3D监督,并且设计一个能够吸收这种监督的架构,而不干扰强大的预训练视频先验知识。我们引入了一个校准流程,将学习到的立体深度与联合因子图相结合,汇集所有收集自同一物理机器人的剧集,以恢复其共享的运动学参数以及每个场景的外部参数。应用于DROID数据集,这产生了DROID-3D,一个校准的3D数据集,提供密集的度量深度和重新校准的多视图外参(在外部摄像头上达到90%的剧集中<0.7像素的重投影误差)。然后我们训练DepthWorld,一个基于稳定视频扩散的世界模型,通过空间潜在平铺联合预测多视图RGB和深度,保持预训练的变分自动编码器(VAE)不变。深度监督使RGB预测本身的PSNR比相同的仅RGB基线提高了+1.48 dB,同时为下游几何推理提供准确的度量深度。

更新时间: 2026-10-06 17:59:00

领域: cs.RO,cs.AI,cs.CV

下载: http://arxiv.org/abs/2610.08780v1

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replacement only when the trace satisfies an evidence condition tied to a trace-local failure hypothesis. It constructs a trace-grounded counterfactual alternative, a negative twin, and replaces the agent's proposal only if the twin passes structural checks and the pairwise verifier prefers it in both candidate orders. For paired evaluation, exact replay holds the agent's parsed responses and actions fixed until the first accepted replacement, separating intervention effects from resampling. In the primary analysis of 159 multi-turn BFCL V4 tasks with complete exact-replay pairs, the complete policy raises task success for GPT-5.6 Sol from 45.3% to 58.5% (95% task-bootstrap CI [8.2, 18.8]), with no observed success-to-failure regressions. Together, these findings recast execution-boundary repair as a constrained comparison, making the counterfactual action itself the object of verification.

Updated: 2026-10-06 17:58:20

标题: TwinCheck:针对有状态工具代理的证据支持的负向双重验证

摘要: 一个局部合理的工具调用可能会使一个本来成功的代理轨迹出轨。仅仅怀疑并不能证明干预的必要性,因为替换本身可能引入验证旨在防止的失败。我们引入了TwinCheck,一种推断时验证策略,它只在迹符合与迹本地失败假设相关的证据条件时考虑替换。它构建了一个基于迹的反事实替代,一个负的双胞胎,并且只有当双胞胎通过结构检查并且成对验证器在两个候选顺序中都更喜欢它时才替换代理的提议。在配对评估中,精确重放保持代理的解析响应和动作固定,直到第一个被接受的替换,将干预效果与重新采样分开。在对159个完整的精确重放对的多轮BFCL V4任务的主要分析中,完整的策略将GPT-5.6 Sol的任务成功率从45.3%提高到58.5%(95%任务自举置信区间[8.2, 18.8]),没有观察到成功到失败的回归。总的来说,这些发现将执行边界修复重新构建为一种受限制的比较,使反事实行动本身成为验证的对象。

更新时间: 2026-10-06 17:58:20

领域: cs.AI

下载: http://arxiv.org/abs/2609.26911v2

Sherpa: Teaching LLMs to Teach Adaptively

Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.

Updated: 2026-10-06 17:58:18

标题: 谢尔巴:教授LLMs进行自适应教学

摘要: 大型语言模型(LLMs)已经变得越来越有能力解决问题,但能够解决问题并不等同于能够教授它。现有的培训LLMs作为教师的方法依赖于示范、偏好数据或预先定义的教学标准,这些标准指定了良好教学的特征。然而,这些信号通常并不基于个体学生的学习成果,有效的教学策略在学习者之间可能存在巨大差异。为了解决这个问题,我们引入了Sherpa,这是一个多轮强化学习框架,用LLMs实例化多个学生原型,这些LLMs受到不同学习偏好的条件,并训练一个教师模型通过直接最大化他们的学习成果来调整其教学。使用Sherpa训练的教师LLMs将指导学生的表现提高了平均20.5个百分点。在MathTutorBench的评估下,Sherpa将整体教学得分从52.5%提高到79.2%,表明教学反应更好。我们的人类研究表明,在79.6%的两两比较中,受过训练的教师优于基础模型。总之,Sherpa培训LLM教师以适应多样化的模拟学生,并更好地与人类教师对齐,为AI导师教授真实学生铺平了道路。

更新时间: 2026-10-06 17:58:18

领域: cs.AI,cs.CL

下载: http://arxiv.org/abs/2610.08778v1

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We call this ability "bottling": the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost. We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets. Agents choose their own approach, such as training a small model or writing a reusable program. Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities. Models with similar zero-shot scores can differ substantially after bottling, and 48 of 60 bottling runs score below the lower bound of the 95% confidence interval of their model's zero-shot performance. Moreover, 31 of 60 runs underperform the stronger of two small-model distillation baselines with the same token budget. Nevertheless, bottling can yield substantial savings: on query-product relevance classification, Opus 5 retains about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. Bottling is also competitive with Jev, a "system one" model built especially for cheap, repetitive inference: Opus 5 on the same task recovers about 94% of Jev's macro-F1 at a quarter of Jev's projected full-workload cost. BOTTLED provides a basis for evaluating and improving agents' ability to invest limited resources in reusable solutions for large, repetitive workloads.

Updated: 2026-10-06 17:57:19

标题: 一个瓶中的代理:LLM代理能否将其能力转化为廉价、可扩展的工件?

摘要: 大型语言模型(LLMs)可以解决许多狭窄任务,但单独查询它们的数百万相关实例可能成本过高。LLM代理能够自主创建更便宜的解决方案来处理这样的工作负载吗?我们称这种能力为“装瓶”:将通用能力转化为平衡答案质量和摊销成本的特定任务解决方案的能力。我们介绍了BOTTLED,一个基准测试,在该基准测试中,代理接收整个未标记的工作负载,并必须在固定的时间、计算和LLM API预算下完成。代理可以选择自己的方法,例如训练一个小模型或编写一个可重复使用的程序。在十个模型和三个任务中,我们发现强零-shot任务性能并不能可靠地转化为强装瓶能力。在装瓶后,具有相似零-shot分数的模型可能存在显著差异,60次装瓶运行中有48次得分低于其模型零-shot性能95%置信区间下限。此外,60次运行中有31次表现不佳,低于两个相同标记预算的小模型蒸馏基准中较强的一个。然而,装瓶可以带来可观的节省:在查询-产品相关性分类任务中,Opus 5以大约657倍较低的报告成本保留了其零-shot宏F1的约82%。装瓶还与Jev竞争,Jev是专门用于廉价、重复推断的“系统一”模型:Opus 5在相同的任务上以Jev预期全负载成本的四分之一恢复了约94%的Jev宏F1。BOTTLED为评估和改进代理在大型重复工作负载中投入有限资源以获得可重复使用解决方案的能力提供了基础。

更新时间: 2026-10-06 17:57:19

领域: cs.AI

下载: http://arxiv.org/abs/2610.08775v1

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.

Updated: 2026-10-06 17:56:43

标题: AdvSim2Real:在Web世界模型中训练Web代理对抗自适应提示注入

摘要: 网络代理通过阅读和执行第三方编写的页面来完成用户请求,因此在页面上植入的指令可能会将代理重定向到用户的目标之外。代理不能简单地忽略页面,因为页面还包含任务需要的数值和控制。当前的防御措施对在训练之前固定的注入进行微调,而适应训练模型的攻击者可以绕过它们。对抗性训练允许攻击者适应,但保持任务不变,因此一旦代理解决了任务,任务就停止教学。我们引入了AdvSim2Real,它在冻结的网络世界模型中共同演化任务课程、注入对手和代理。课程奖励代理大约一半时间解决的任务,而对手仅在成功翻转时获得奖励,即将被判断为成功的注入转化为失败。在模拟器中训练使得4B代理更加强大和更加稳健:它的完成率随着攻击和没有攻击而提高,对抗适应模型的成功率保持稳定,其能力增益也延续到真实浏览器。在150个网络任务中,AdvSim2Real相对于基础代理将在未见对手的情况下的完成率提高了33.6\%。

更新时间: 2026-10-06 17:56:43

领域: cs.CL,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08773v1

Mission-Aware Attestation Envelopes for Time-Critical Autonomous Action: A Hardware-in-the-Loop V2I Study

An autonomous system that asks for a privileged physical action is usually gated on integrity evidence: a platform proves what it is running, and the request is granted or refused on that basis. Such a gate is normally treated as a predicate, yet the evidence behind it has an age, the decision that consumes it has a latency, and the physical system that waits for it has a deadline. We formulate mission-aware attestation as a runtime assurance contract that holds only when integrity is valid, the evidence is fresh enough, and the decision completes inside a budget derived from the current physical state. The contract yields four operational outcomes where a binary gate yields two, separating a refusal caused by tampering from one caused by stale evidence and from one caused by a late decision. We evaluate it on a hardware-in-the-loop vehicle-to-infrastructure platform: a driving simulator supplies the physical state and the authorisation deadline, while a microcontroller on-board unit and a TPM-backed roadside unit running Linux integrity measurement supply the assurance evidence. A security-blind model admits the whole operating space and a hardware-informed one three quarters of it, and every point it refuses fails the freshness margin rather than the response margin. Moving the attestation interval across the range the verifier permits costs about as much as a fivefold scaling of the latency distribution, and the interval is directly configurable, which makes it the immediately actionable deployment parameter. If the freshness bound does not exceed the authorisation budget, every late decision is also stale and lateness becomes unobservable, so the attestation interval and the freshness bound cannot be chosen from security requirements alone.

Updated: 2026-10-06 17:55:18

标题: 任务感知认证信封用于时间关键型自主行动:基于硬件在环V2I研究

摘要: 一个要求特权物理操作的自主系统通常会基于完整性证据进行门控:一个平台证明它正在运行的内容,并根据此基础授予或拒绝请求。这样的门控通常被视为一个谓词,但是其背后的证据有一个年龄,消耗其的决策有一个延迟,并且等待它的物理系统有一个截止日期。我们将任务感知认证定义为一种运行时保证合同,只有在完整性有效、证据足够新鲜且决策在从当前物理状态派生的预算内完成时才有效。该合同产生四种运行结果,而二进制门则产生两种结果,将由篡改引起的拒绝与由陈旧证据引起的拒绝以及由延迟决策引起的拒绝分开。我们在一个硬件在环车辆到基础设施平台上对其进行评估:一个驾驶模拟器提供物理状态和授权截止日期,而一个搭载微控制器和运行Linux完整性测量的TPM支持的道路单元提供保证证据。一个安全盲模型允许整个操作空间,而一个硬件信息模型允许其四分之三,它拒绝的每一点都未能达到新鲜度边界,而不是响应边界。在验证器允许的范围内移动认证间隔的成本大约相当于延迟分布的五倍缩放,而且该间隔是直接可配置的,这使得它成为立即可操作的部署参数。如果新鲜度边界不超过授权预算,那么每个延迟决策也是陈旧的,延迟变得不可观察,因此认证间隔和新鲜度边界不能仅根据安全需求选择。

更新时间: 2026-10-06 17:55:18

领域: cs.CR,cs.RO,eess.SY

下载: http://arxiv.org/abs/2610.08771v1

Rapid Fredholm stabilization of the Kuramoto--Sivashinsky equation with unrestricted, spatially-varying anti-diffusion

We develop the first feedback design for rapid stabilization of the Kuramoto--Sivashinsky equation with a spatially varying anti-diffusion coefficient. For constant coefficients, the single-input Fredholm design of Coron and Lü (2015) excludes a discrete set of values at which repeated unstable eigenvalues cause a loss of controllability. We overcome this obstruction by introducing a second boundary input and assigning the two inputs distinct roles. The key idea, inspired by Heymann's Lemma, is to use the boundary value $u(0,t)$ entirely for a pre-feedback that renders the modified plant controllable through the curvature input $u_{xx}(0,t)$. The latter input then stabilizes the plant through a Fredholm backstepping transformation. We show that two inputs suffice for controllability and are necessary when the plant has an unstable double eigenvalue. However, the Fredholm kernel still must be approximated for implementation. Hence, to enable kernel and gain approximation, we prove continuity of the coefficient-to-gain design map on compact admissible design classes. Unlike Volterra-based continuity proofs using successive approximations, our proof uses the modal representation to control the spectral data, the inverse coefficient system, and the tails of the kernel and gain series. This yields a single neural operator approximation of the gain to any prescribed $L^2$ accuracy across the class. Finally, we establish rapid local stabilization of the nonlinear closed-loop system under both the exact gains and sufficiently accurate approximations. We conclude with numerical results that illustrate prescribed decay rates and the computational cost of the approximations. In particular, we train a Fourier neural operator that achieves typical relative gain errors of approximately $0.1\%$ and stabilizes all held-out cases tested, including a plant with an unstable double eigenvalue.

Updated: 2026-10-06 17:50:54

标题: 快速Fredholm稳定化Kuramoto-Sivashinsky方程具有无限制、空间变化的反扩散

摘要: 我们开发了第一个针对具有空间变化反扩散系数的Kuramoto-Sivashinsky方程快速稳定化的反馈设计。对于常数系数,Coron和Lü(2015)的单输入Fredholm设计排除了一组离散的数值,其中重复的不稳定特征值导致了可控性的丧失。我们通过引入第二个边界输入并指定两个输入不同的角色来克服这一障碍。受Heymann引理启发的关键思想是将边界值$u(0,t)$完全用于渲染经过反馈修改的系统通过曲率输入$u_{xx}(0,t)$可控。然后,后者输入通过Fredholm回步变换稳定化系统。我们证明当系统具有不稳定双特征值时,两个输入足以实现可控性,并且在这种情况下是必要的。然而,仍然必须对Fredholm核进行近似以进行实施。因此,为了实现核和增益的逼近,我们证明了在紧致可容许设计类上的系数与增益设计映射的连续性。与使用逐步逼近的Volterra连续性证明不同,我们的证明使用模态表示来控制谱数据、逆系数系统以及核和增益级数的尾部。这产生了对于该类别中的任何预定$L^2$精度的增益的单个神经算子逼近。最后,我们在精确增益和足够精确的逼近下建立了非线性闭环系统的快速局部稳定化。我们通过数值结果来说明所需的衰减速率和逼近的计算成本。特别地,我们训练了一个傅里叶神经算子,其典型相对增益误差约为$0.1\%$,并且稳定化了所有进行测试的保留案例,包括具有不稳定双特征值的系统。

更新时间: 2026-10-06 17:50:54

领域: eess.SY,cs.LG,math.AP,math.OC

下载: http://arxiv.org/abs/2610.08764v1

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.

Updated: 2026-10-06 17:50:29

标题: VeriFine:身体推理自我改进中的验证扩展

摘要: 自我改进的政策不断暴露新的失败模式,改变了评委必须能够验证的内容。然而,目前的固定评委限制了优化反馈和发现有用训练示例的能力,限制了进一步的自我改进。在具体推理中,这一挑战更加严峻,可靠的评估必须考虑空间基础、因果推理和安全意识决策。我们介绍了VeriFine,一个代理器架构框架,通过政策、培训大纲和评委的共同进化来扩展验证。政策改进循环使用一个评分评委来诊断重复的失败,构建一个适应性课程,并优化政策。当进展停滞,验证成为瓶颈时,评委改进循环有选择地查询有关信息失败案例的人类指导,并通过共同校准来完善评委,在这种情况下,人类和代理解决分歧,并朝着物理推理的客观评分标准趋于一致。然后修订的评委指导下一阶段的数据选择和政策优化。在驾驶和机器人导航任务上的实验表明,在强化和监督微调中,政策和评委的能力持续自我改进。这些结果展示了如何通过扩展验证来支持随着政策失败模式的演变而持续自我改进。

更新时间: 2026-10-06 17:50:29

领域: cs.AI,cs.RO

下载: http://arxiv.org/abs/2610.08761v1

WorldSonus: Bringing Sound to Worlds

Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/

Updated: 2026-10-06 17:50:18

标题: WorldSonus:为世界带来声音

摘要: 最近在世界模型方面取得的进展已经使得视觉合成变得越来越逼真。然而,这些生成的环境仍然大多是无声的。为世界模型引入声音面临三个核心挑战:实时生成以跟上交互式视频流,交互控制以响应中途的声音指令,以及空间对齐立体声以反映场景几何和摄像机运动。为了解决这些需求,我们引入了WorldSonus,这是一个专为世界模型设计的交互式视频到音频框架,用于实时空间声音合成。对于实时生成,WorldSonus采用了流式因果自回归扩散架构,以低实时因子(RTF)0.41合成音频块。对于交互控制,我们结合了一个以音频为中心的字幕管道,通过块索引提示调度,使得在生成过程中能够动态操纵声音事件。对于空间对齐,我们利用了从各种立体声和环绕声数据中策划的高质量立体声监督。大量实验表明,虽然专为世界模型定制,但WorldSonus在开放领域的视频到音频基准测试中表现出很好的泛化能力,与最先进的双向模型在声学质量和空间对齐方面相匹敌甚至超过。项目页面:https://noizai.github.io/WorldSonus/

更新时间: 2026-10-06 17:50:18

领域: cs.SD,cs.AI,cs.CV,eess.AS

下载: http://arxiv.org/abs/2610.08760v1

TAPDreamer: Transferable Adversarial Patches for World Action Models

World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.86% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 1.45% and 1.00% on two DreamWAM configurations and to 10.60% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.

Updated: 2026-10-06 17:50:08

标题: TAPDreamer:用于世界行动模型的可转移对抗性贴片

摘要: 世界模型学习预测环境将如何演变,使其成为通用机器人控制的重要基础。然而,世界动作模型依赖于相机输入,其操作可能会破坏跨任务和动作策略使用的视觉表示。现有对这些模型的攻击针对受害者的行动或预测的未来进行优化,因此需要访问目标模型的输出。在本文中,我们提出了一种攻击,TAPDreamer,针对世界动作模型,它仅使用一个公共编码器来构建一个固定的局部扰动,可以在任务和动作架构之间传递。TAPDreamer不需要目标策略查询。我们的关键见解是,补丁引起的注意权重和价值向量之间的交互广播了一个几乎相同的表示偏移,远远超出了补丁的范围,而且这个偏移在任务观察中始终保持稳定。受此见解的指导,TAPDreamer使用一个源任务的六帧图像来最大化干净和补丁编码器表示之间的全局L1距离。在闭环评估中,每个基准测试中有一个冻结的补丁,覆盖了大约6.5%的输入,将FastWAM的成功率从97.7%降低到0.0%,跨越了40个LIBERO任务,从90.86%降低到0.0%,跨越了50个RoboTwin任务;匹配的随机补丁保留了81.5%和79.2%的成功率。相同的补丁将成功率降低到了1.45%和1.00%的两个DreamWAM配置,以及10.60%的Motus。这些结果表明,仅保护下游动作生成是不够的:世界动作模型的防御也必须确保共享的视觉编码器免受持续的局部扰动。

更新时间: 2026-10-06 17:50:08

领域: cs.CV,cs.AI,cs.RO

下载: http://arxiv.org/abs/2610.06814v2

Reinforcement Learning over Predictive Distributions for LLM Regression

Large language models (LLMs) have emerged as flexible regressors capable of predicting real-valued quantities from heterogeneous inputs. Yet most LLM regression objectives optimize predictions independently, often yielding poor calibration. We introduce Distribution-Aware Reward (DAR), an on-policy reinforcement learning objective that instead jointly evaluates the empirical predictive distribution formed by multiple predictions for the same input. To translate this distribution-level objective into rollout-level rewards, we assign each prediction credit based on its leave-one-out contribution to the quality of the overall predictive distribution. This encourages predictions that are well-centered and appropriately dispersed around the target. We evaluate on three regression settings: a synthetic task probing interpolation and extrapolation, and two real-world scientific tasks involving code and molecular data. Across tasks, DAR produces better-calibrated uncertainty estimates while consistently reducing prediction error and improving ranking quality over supervised fine-tuning and pointwise reinforcement learning. Together, these results highlight the benefits of distribution-aware training for LLM regression.

Updated: 2026-10-06 17:48:09

标题: 使用LLM回归的预测分布进行强化学习

摘要: 大型语言模型(LLM)已成为灵活的回归器,能够从异质输入中预测实值数量。然而,大多数LLM回归目标都是独立优化预测,通常导致预测不准确。我们引入了Distribution-Aware Reward(DAR),这是一个基于政策的强化学习目标,而不是联合评估由同一输入生成的多个预测组成的经验预测分布。为了将这种分布级别的目标转化为滚动级别的奖励,我们根据每个预测对整体预测分布质量的单个贡献,给予每个预测信用。这鼓励预测集中在目标周围并适当分散。我们在三个回归设置上进行评估:一个探索插值和外推的合成任务,以及涉及代码和分子数据的两个真实科学任务。在各种任务中,DAR产生更好校准的不确定性估计,同时始终减少预测误差并提高排名质量,优于监督微调和点对点强化学习。这些结果共同突显了分布感知培训对LLM回归的好处。

更新时间: 2026-10-06 17:48:09

领域: cs.LG,cs.AI,cs.CL

下载: http://arxiv.org/abs/2605.20740v2

Neural Petri flows for chemical reactions

Petri nets have been used to describe chemical processes such as reactions.They map well to chemistry: Places are the bonds between atoms and the free valence of each atom, a token is a unit of bond order, a transition forms or breaks a bond, the conserved quantities are the valence budgets of the atoms, and the enabling rule is the valence rule. These semantics are not guaranteed by learned models of reactions or neural networks that are built on Petri nets that use the net as a scaffold for message passing. Here, we ask what architecture remains a Petri net for every value of its weights. We find the answer in the theory, where all semantics of a net share the firing form $m^\prime=m+Cσ$, locality, as enabling reads only the inputs of a transition, and the enabling rule, and we prove that conservation forces the firing form and that non-negativity forces the enabling rule on local rate laws. This leaves free the rate law, which is the propensity of each transition to fire. We introduce Neural Petri Flow, which learns this rate law, or a readout for classification, and hard-wires the rest as parameter-free layers. On what we denote a valence net, atom mapping, reaction classification, and forward prediction become three tasks on one firing vector. Without training, the minimum firing vector maps 88.8% of the curated Golden set against 85.6% for RXNMapper, and 88.7 against 77.9% of the enzymatic reactions of EnzymeMap. On USPTO-480K, NPF trained on these firing vectors predicts 87.7% of the products and 67.4% when trained on a 1% subset of the training reactions. EC numbers of ECREACT are predicted at the third level for 90.2% of reactions, 5.6 points ahead of the best published method. With electrons as tokens, the same token game predicts 90.5% of the elementary steps of FlowER first, ahead of the published baseline, and every top-1 prediction is a valid molecule without a filter.

Updated: 2026-10-06 17:44:56

标题: 神经元彼得里流用于化学反应

摘要: Petri nets被用来描述化学过程,比如反应。它们很好地映射到化学领域:地点是原子之间的键和每个原子的自由价,一个令牌是一个键序的单位,一个转换形成或断裂一个键,守恒量是原子的价预算,启用规则是价规则。这些语义不被学习模型或建立在Petri网上的神经网络所保证,这些模型使用网作为消息传递的支架。在这里,我们询问每个权重值下Petri网仍保持什么体系结构。我们在理论中找到答案,其中网的所有语义共享发射形式$m^\prime=m+Cσ$,局部性,因为启用仅读取转换的输入,并且启用规则,我们证明守恒迫使发射形式,非负性迫使局部速率规律上的启用规则。这留下了速率规律的自由,这是每个转换发射的倾向。我们引入神经Petri流,它学习这种速率规律,或者分类的读出,并将其余部分硬编码为无参数层。在我们所称的价网上,原子映射,反应分类和前向预测成为一个发射向量上的三个任务。在没有训练的情况下,最小发射向量对齐了88.8%的经过整理的Golden数据集,而RXNMapper为85.6%,对EnzymeMap的酶反应88.7%对77.9%。在USPTO-480K上,对这些发射向量进行训练的NPF预测了87.7%的产品,当对训练反应的1%子集进行训练时,预测了67.4%。对于ECREACT,EC编号在第三级别上对90.2%的反应进行了预测,比最佳已发表的方法领先5.6个百分点。使用电子作为令牌,相同的令牌游戏预测了FlowER的90.5%基本步骤,领先于已发表的基线,并且每个前1预测都是一个有效的分子,没有过滤器。

更新时间: 2026-10-06 17:44:56

领域: cs.LG,physics.chem-ph,q-bio.QM

下载: http://arxiv.org/abs/2610.08750v1

ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks

Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and masked dLMs. We scale Embedded Language Flows (ELF) to mathematical reasoning and code generation on GSM8K, MATH-500, HumanEval, and MBPP. We introduce ELF-REG, which improves learning with representation alignment and entanglement (REPA+REG), where a frozen AR teacher supervises intermediate denoiser features and supplies a global representation that is jointly denoised with the response. ELF-REG-L achieves 55.96% pass@1 on GSM8K at 64 network function evaluations (NFE), and 13.39% on MATH-500 and 22.56% on HumanEval at 128 NFE. It outperforms the evaluated comparable-scale dLMs in pass@1 on GSM8K and code, and improves MATH-500 pass@1 from 10.55% for the ELF-L baseline to 13.39% with ELF-REG-L. Without few-step training, the same task-specific checkpoints support strong low-NFE performance through early-stop, which decodes an intermediate clean prediction without completing the denoising trajectory. At 16 NFE, ELF-REG-L reaches 41.21% HumanEval pass@10, outperforming recent continuous dLMs of comparable scale.

Updated: 2026-10-06 17:43:50

标题: ELF-REG:将连续扩散语言模型扩展到推理任务

摘要: 全连续扩散语言模型(dLMs)在没有中间离散化的情况下对连续表示进行去噪,然后在最后一步并行解码所有响应标记。它们在具有挑战性的推理任务上的表现与自回归(AR)LLMs和掩蔽dLMs相比仍然不太确定。我们将嵌入语言流(ELF)扩展到GSM8K、MATH-500、HumanEval和MBPP上的数学推理和代码生成。我们引入了ELF-REG,通过表示对齐和纠缠(REPA+REG)改进学习,其中一个冻结的AR教师监督中间去噪特征并提供一个与响应一起联合去噪的全局表示。ELF-REG-L在64个网络功能评估(NFE)的GSM8K上达到了55.96%的pass@1,而在128个NFE的MATH-500和HumanEval上分别为13.39%和22.56%。它在GSM8K和代码的pass@1上优于评估的相当规模的dLMs,并将MATH-500的pass@1从ELF-L基线的10.55%提高到了13.39%。在没有几步训练的情况下,相同的任务特定检查点通过提前停止支持强大的低NFE性能,这在完成去噪轨迹之前解码了一个中间干净的预测。在16个NFE时,ELF-REG-L达到了41.21%的HumanEval pass@10,优于相当规模的最新连续dLMs。

更新时间: 2026-10-06 17:43:50

领域: cs.CL,cs.LG

下载: http://arxiv.org/abs/2609.29102v2

Linear Bandits under Exact Sliding-Window Constraints

We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity and cyclic-shift invariance make a stationary solution optimal when $w\mid T$ and within an additive $O(w)$ gap otherwise. In the online setting, we show that geometric structure alone is insufficient for learning, and sublinear regret can be impossible. We introduce a transition diameter $τ$ that quantifies feasible reachability and develop a rare-switching OFUL algorithm with regret $\widetilde{O}(d\sqrt{T}+τd+w)$ against the offline-optimal feasible trajectory. Finally, we remove cyclic invariance and consider general sliding-window constraints, where optimal behavior may be non-stationary. We represent recent action history as the state of a finite-memory control problem and introduce a history-state diameter $D$ that measures feasible communication between viable histories. Combining optimistic remaining-horizon planning with rare policy updates, we obtain a regret bound of $\widetilde{O}(d\sqrt{T}+dD+w)$. We evaluate our approach on real-world and synthetic benchmarks, showing that it maintains exact feasibility while achieving reward and regret comparable to baselines with substantially fewer policy updates.

Updated: 2026-10-06 17:41:08

标题: 线性臂带在精确滑动窗口约束下

摘要: 我们研究了在线精确滑动窗口约束下的线性赌臂问题,其中每个连续的动作块必须属于一个预定的可行集合。在离线设定中,即奖励函数已知的情况下,我们证明了凸性和循环移位不变性使得当$w\mid T$时,静态解是最优的,否则会有$O(w)$的附加差距。在在线设定中,我们证明了仅凭几何结构是不足以学习的,而且次线性后悔可能是不可能的。我们引入了一个度量可行性可达性的过渡直径$τ$,并开发了一种具有$\widetilde{O}(d\sqrt{T}+τd+w)$的后悔的稀有切换OFUL算法,针对离线最优可行轨迹。最后,我们去除了循环不变性,考虑了一般的滑动窗口约束,其中最优行为可能是非静态的。我们将最近的动作历史表示为有限记忆控制问题的状态,并引入了一个度量可行历史之间通信的历史状态直径$D$。通过乐观剩余地平线规划与稀有策略更新相结合,我们得到了一个后悔上界为$\widetilde{O}(d\sqrt{T}+dD+w)$。我们在真实世界和合成基准测试中评估了我们的方法,结果表明它在保持精确可行性的同时,实现了与具有大大减少策略更新的基线相当的奖励和后悔。

更新时间: 2026-10-06 17:41:08

领域: cs.LG

下载: http://arxiv.org/abs/2610.08745v1

Fast, Interpretable, and Deterministic Time Series Classification With a Bag-of-Receptive-Fields

The current trend in the literature on Time Series Classification is to develop increasingly accurate algorithms by combining multiple models in ensemble hybrids, representing time series in complex and expressive feature spaces, and extracting features from different representations of the same time series. As a consequence of this focus on predictive performance, the best time series classifiers are black-box models, which are not understandable from a human standpoint. Even the approaches that are regarded as interpretable, such as shapelet-based ones, rely on randomization to maintain computational efficiency. This poses challenges for interpretability, as the explanation can change from run to run. Given these limitations, we propose the Bag-Of-Receptive-Field (BORF), a fast, interpretable, and deterministic time series transform. Building upon the classical Bag-Of-Patterns, we bridge the gap between convolutional operators and discretization, enhancing the Symbolic Aggregate Approximation (SAX) with dilation and stride, which can more effectively capture temporal patterns at multiple scales. We propose an algorithmic speedup that reduces the time complexity associated with SAX-based classifiers, allowing the extension of the Bag-Of-Patterns to the more flexible Bag-Of-Receptive-Fields, represented as a sparse multivariate tensor. The empirical results from testing our proposal on more than 150 univariate and multivariate classification datasets demonstrate good accuracy and great computational efficiency compared to traditional SAX-based methods and state-of-the-art time series classifiers, while providing easy-to-understand explanations.

Updated: 2026-10-06 17:41:02

标题: 快速、可解释和确定性的时间序列分类方法:基于接收域的词袋模型

摘要: 时间序列分类文献中的当前趋势是通过将多个模型结合成集成混合模型,将时间序列表示为复杂和富有表现力的特征空间,并从同一时间序列的不同表示中提取特征,开发越来越准确的算法。由于这种对预测性能的关注,最好的时间序列分类器是黑盒模型,从人类的角度来看,这些模型是无法理解的。即使被认为是可解释的方法,如基于形状的方法,也依赖于随机化以保持计算效率。这对解释性提出了挑战,因为解释可以在每次运行时发生变化。鉴于这些局限性,我们提出了一种快速、可解释且确定性的时间序列转换方法——感受野包(BORF)。在经典的模式包基础上,我们在卷积算子和离散化之间架起了桥梁,通过增强符号聚合近似(SAX)的膨胀和步幅,可以更有效地捕获多尺度的时间模式。我们提出了一种算法加速方法,减少了与基于SAX的分类器相关的时间复杂性,从而将模式包扩展为更灵活的感受野包,表示为稀疏多元张量。对我们的提议在150多个单变量和多变量分类数据集上进行测试的实证结果表明,与传统的基于SAX的方法和最先进的时间序列分类器相比,我们的方法具有良好的准确性和极高的计算效率,同时提供易于理解的解释。

更新时间: 2026-10-06 17:41:02

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2311.18029v2

Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation

Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching $1.11\times$ to $5.21\times$ that of the strongest baseline, with competitive session depth and no larger retained sets.

Updated: 2026-10-06 17:40:52

标题: 使用符合动作集的强化学习:一个顺序推荐的应用

摘要: 序列推荐系统通常使用固定的推荐集大小,即使在会话中有用的替代品数量会发生变化。我们提出了一种称为校准修剪的强化学习(RLCP)方法,该方法使用评论家评分和在线阈值来调整保留的动作集。阈值是根据二进制反馈更新的,指示集合是否包含代理目标中的动作。我们证明了在自适应轨迹上观察到的代理错失率的确定性界。为了量化修剪对奖励的影响,我们推导出了价值损失的精确分解为过滤和选择损失。在明确的代理和评论家逼近条件下,这种分解产生了一个有限的会话奖励界限,还考虑了不完美的选择和集合截断,而不需要学习参数收敛。在KuaiRand-Pure和MovieLens 1M上的实验比较了两个RLCP实现和四个RL基线。在19个配置中,至少有一个RLCP变体实现了最高的目录多样性,达到了最强基线的$1.11\times$到$5.21\times$的水平,具有竞争性的会话深度和没有更大的保留集合。

更新时间: 2026-10-06 17:40:52

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.08743v1

On the Computational Tractability of Robust Bandits

Learning when the environment does not belong to the learner's hypothesis class is typically handled using agnostic learning guarantees. However, for anything beyond supervised learning, agnostic guarantees are difficult to come by. Recently, imprecise bandits (Kosoy, 2025) (later renamed to robust bandits in Appel and Kosoy, 2025) were introduced as another approach to unrealizable learning in the bandits setting and a $Θ(\sqrt{T})$ regret learner was shown for a large class. However, no computational guarantees were provided. In this paper we identify a special case that admits a polynomial-time learner with $\tilde{O}(\sqrt{T})$ regret. We also show that several small generalizations of this special case are NP-hard thus indicating that the special case is at the boundary of what is tractable. It has been recently suggested (Kosoy, 2018) that computationally efficient learners for unrealizable learning problems are crucial for solving the AI alignment problem. This work is a small step in that direction.

Updated: 2026-10-06 17:38:34

标题: 关于坚固性赌徒计算可行性的研究

摘要: 学习当环境不属于学习者的假设类时,通常使用不可知学习保证来处理。然而,除了监督学习之外,不可知保证很难实现。最近,不精确赌徒(Kosoy,2025)(后来在Appel和Kosoy,2025中更名为健壮赌徒)被引入作为在赌徒设置中另一种处理不可实现学习的方法,并且已经为一个大类展示了$Θ(\sqrt{T})$的后悔学习器。然而,并未提供计算保证。在本文中,我们确定了一个特殊情况,其中存在一个多项式时间学习器,具有$\tilde{O}(\sqrt{T})$的后悔。我们还表明,这个特殊情况的几个小的推广是NP难的,从而表明这个特殊情况处于可处理的边界。最近有人提出(Kosoy,2018)计算效率高的学习器对于解决人工智能对齐问题至关重要。这项工作是朝着这个方向迈出的一小步。

更新时间: 2026-10-06 17:38:34

领域: cs.LG

下载: http://arxiv.org/abs/2610.08740v1

BARE-AI: Bit-Flip Attack Resilience in AI Hardware through Built-in Performance Monitors

Deep Neural Networks (DNNs) are integral to many safety critical systems, yet they remain highly vulnerable to bit-flip attacks (BFAs), where a few memory level perturbations can drastically degrade accuracy. Existing defenses incur significant hardware overhead, depend on retraining, or fail against targeted flips. We propose BARE-AI, a runtime framework that detects, localizes, and mitigates BFAs during inference. BARE-AI introduces AI Performance Counters (APCs), lightweight hardware monitors in the accelerator datapath that capture per-layer activation statistics such as sparsity, entropy, kurtosis, and spectral shift. These are analyzed by the Predictive Unit for Layer Security Evaluation (PULSE), a compact detector trained offline as an ensemble of classifiers and realized on-chip as a small neural engine. For explainability and recovery, BARE-AI introduces an Activation Shift Index (ASI) for layer level fault localization and a z-score based repair that resets anomalous weights toward clean layer statistics. Across CNNs, Vision Transformers, and Large Language Models under random, targeted, adaptive, and magnitude based BFAs, BARE-AI achieves up to 98% detection accuracy on vision models and 74% to 95% on language models, restores near clean accuracy for CNNs and ViTs, and provides partial recovery for LLMs. Synthesized at 28nm, the monitoring infrastructure incurs under 3% energy, under 4% area, and about 10% latency overhead, with a configurable operating point that reduces latency overhead to about 6%. Unlike error correcting codes, whose redundancy grows with the number of tolerated flips, BARE-AI's overhead remains constant regardless of attack strength, making it attractive for resource constrained, safety critical edge applications such as autonomous systems, energy, and healthcare.

Updated: 2026-10-06 17:38:24

标题: BARE-AI:通过内置性能监控实现人工智能硬件的位翻转攻击抵御

摘要: 深度神经网络(DNNs)是许多安全关键系统中不可或缺的,但它们仍然极易受到比特翻转攻击(BFAs)的影响,其中少数内存级扰动可能会严重降低准确性。现有的防御措施会导致显著的硬件开销,依赖重新训练,或无法抵御有针对性的翻转攻击。我们提出了BARE-AI,这是一个运行时框架,用于在推断期间检测、定位和减轻BFAs。BARE-AI引入了AI性能计数器(APCs),这是加速器数据通路中的轻量级硬件监视器,用于捕获每层激活统计信息,如稀疏性、熵、峰度和谱移。这些统计数据由预测单元用于层安全评估(PULSE)进行分析,PULSE是一个紧凑的检测器,离线训练为一组分类器,并在芯片上实现为一个小型神经引擎。为了解释和恢复,BARE-AI引入了一个激活偏移指数(ASI)用于层级故障定位,以及基于z分数的修复,将异常权重重置为干净的层统计数据。在随机、有针对性、自适应和基于幅度的BFAs下,BARE-AI在视觉模型上实现了高达98%的检测准确性,对语言模型实现了74%到95%的准确性恢复,对CNNs和ViTs恢复了接近干净准确性,并为LLMs提供了部分恢复。在28纳米合成时,监测基础设施的能耗不到3%,面积不到4%,延迟开销约为10%,具有可配置的运行点,可以将延迟开销降低到约6%。与纠错码不同,其冗余随着允许的翻转次数增加而增加,而BARE-AI的开销则保持不变,无论攻击强度如何,使其适用于资源受限、安全关键的边缘应用,如自主系统、能源和医疗保健。

更新时间: 2026-10-06 17:38:24

领域: cs.CR

下载: http://arxiv.org/abs/2610.08739v1

Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling

Diffusion Language Models (DLMs) hold the promise of order-agnostic, parallel text generation. Recently, continuous diffusion and flow matching models have seen substantial gains, driven by carefully crafted token representations and diffusion/flow spaces. In this work, we introduce Hierarchical Continuous Diffusion Language Models (H-CDLMs), a simple framework that further improves continuous DLMs with minimal compute and parameter overhead. Drawing on the discrete DLM and continuous image diffusion literature on joint diffusion, we diffuse multiple modalities in parallel. These modalities represent tokens at different semantic granularities: in our instantiation, the tokens themselves and coarser clusters obtained by clustering pretrained token embeddings. We propose a general setup that allows per-modality samplers and schedules to enhance the interplay between modalities. Applied to CoBit, this yields H-CoBit, which delivers large empirical gains across benchmarks. At dataset entropy, H-CoBit improves MAUVE and reaches a generative perplexity (GenPPL) of 49.4 on LM1B and 50.4 on OWT, improving on the baseline by 24.2 and 20.7 points and surpassing even discrete DLMs of comparable size. On GSM8K, it reaches 27.4% accuracy, outperforming prior continuous diffusion and flow-based models. We further apply H-CDLM to the flow matching model FLM, obtaining consistent gains with H-FLM and demonstrating that the framework generalizes across continuous generative paradigms. Our code will be made publicly available at https://github.com/matol-16/HCDLM.git .

Updated: 2026-10-06 17:38:19

标题: 去噪分层表示:用于语言建模的联合连续扩散

摘要: 扩散语言模型(DLMs)具有无序、并行文本生成的潜力。最近,连续扩散和流匹配模型取得了显著进展,这是由于精心设计的令牌表示和扩散/流空间推动的。在这项工作中,我们引入了分层连续扩散语言模型(H-CDLMs),这是一个简单的框架,进一步改进了连续的DLMs,而且计算和参数开销很小。借鉴了离散DLM和连续图像扩散文献上的联合扩散,我们并行扩散多个形式。这些形式代表不同语义粒度的令牌:在我们的实例中,令牌本身和通过对预训练令牌嵌入进行聚类得到的更粗的簇。我们提出了一个通用的设置,允许按形式的采样器和时间表来增强形式之间的相互作用。应用于CoBit,这产生了H-CoBit,它在基准测试中取得了大幅度的实证增益。在数据集熵中,H-CoBit改进了MAUVE,并在LM1B上达到了49.4的生成困惑度(GenPPL),在OWT上达到了50.4,比基准提高了24.2和20.7个点,甚至超过了相同大小的离散DLMs。在GSM8K上,它达到了27.4%的准确率,优于先前的连续扩散和基于流的模型。我们进一步将H-CDLM应用于流匹配模型FLM,获得了与H-FLM一致的增益,并证明了该框架在连续生成范式之间的泛化。我们的代码将公开在https://github.com/matol-16/HCDLM.git。

更新时间: 2026-10-06 17:38:19

领域: cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.08738v1

Optimal and Efficient Online Inverse Optimization

In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on $\mathbb{R}^{d}$; the goal is to learn to optimize this objective without observing it. Sakaue recently obtained the optimal regret $O(\sqrt d)$ with a randomized algorithm making $(dT)^{O(d)}$ linear optimizations per round, and asked whether it can be attained in polynomial time. We answer positively: our deterministic algorithm has regret $O(\sqrt d)$ for every horizon $T$ and runs in time polynomial in $d$ and $T$. It is a variant of the variable-metric algorithms of Sakaue et al.\ and Cai et al., in which a metric update is revoked once the query point moves far enough from where the update was made.

Updated: 2026-10-06 17:37:24

标题: 最佳和高效的在线反向优化

摘要: 在线逆线性优化中,学习者推荐一个行动,然后观察专家的选择,专家在$\mathbb{R}^{d}$上最大化一个固定的未知线性目标;目标是学会优化这个目标而不观察它。最近Sakaue通过一个随机算法获得了最优遗憾$O(\sqrt d)$,每轮进行$(dT)^{O(d)}$次线性优化,并询问是否可以在多项式时间内实现。我们给出了积极的答复:我们的确定性算法对于每个长度为$T$的时间段都具有$O(\sqrt d)$的遗憾,并且在$d$和$T$中都是多项式时间运行。这是Sakaue等人和Cai等人的变尺度算法的一个变体,在这种算法中,一旦查询点远离了更新的位置,就会撤销度量更新。

更新时间: 2026-10-06 17:37:24

领域: cs.LG,cs.DS

下载: http://arxiv.org/abs/2610.08735v1

EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning

Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.

Updated: 2026-10-06 17:30:33

标题: EgoLAP:通过语言-行为推理学习来自自我中心人类数据

摘要: Egocentric human data提供了一种扩展机器人学习的路径,超越了昂贵的机器人示范,然而,自体差距使得原始人类轨迹成为控制的不良监督目标。我们的关键洞察是,尽管低层动作是特定于自体的,但它们的基础运动意图可以捕捉跨人类和机器人的任务相关结构。我们引入了EgoLAP,一个通过共享基于语言的动作思维链同时从人类和机器人轨迹中学习的VLA预训练框架。EgoLAP将运动意图表达为结构化、时间抽象的语言动作,并将其与基于场景几何、物理和物体功能的运动级推理相配对。通过大量的真实世界和模拟实验,EgoLAP比替代的动作表示更有效地将人类经验转移到机器人控制,并达到80.1%的平均真实世界任务进展,比替代动作表示提高2.3倍。运动级推理也优于合成推理格式,其结合了子任务、物体框和视觉追踪推理。

更新时间: 2026-10-06 17:30:33

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2610.08726v1

Do AI weather models miss extremes?

AI weather models are often reported to underestimate extremes, but most evidence concerns deterministic regression models verified against reanalysis. We evaluate twelve physical and AI forecast models against ECMWF IFS using ten months of European station observations. The evaluation covers 10 m wind, 2 m temperature, solar radiation, and precipitation within regimes defined from a fixed ERA5 1991-2020 climatology. We find no uniform AI-specific deficit in the tails. Several AI models remain more accurate than IFS under extreme conditions, while others deteriorate markedly; comparable variation occurs among physical models. Every model nevertheless exhibits a common conditional-error pattern, overpredicting low observations and underpredicting high observations. Attenuation of extreme values therefore does not imply a uniform loss of relative skill: tail performance depends on the model, variable, and evaluation setting rather than on whether the forecast is produced by AI or physical numerical modelling.

Updated: 2026-10-06 17:28:42

标题: 人工智能天气模型是否会错过极端天气?

摘要: 人工智能天气模型通常被报道低估极端天气情况,但大多数证据涉及对再分析数据验证的确定性回归模型。我们评估了十二个物理和人工智能预报模型与欧洲站点观测的十个月数据相比,使用ECMWF IFS进行评估。评估涵盖了来自固定ERA5 1991-2020气候学的各种制度内的10米风、2米温度、太阳辐射和降水。我们发现在极端条件下,并没有人工智能特有的普遍缺陷。一些人工智能模型在极端条件下仍然比IFS更准确,而其他模型则明显恶化;与物理模型相比,存在可比较的变化。然而,每个模型都表现出一个共同的条件误差模式,即高估低观测值和低估高观测值。因此,极端值的衰减并不意味着相对技能的统一丧失:尾部性能取决于模型、变量和评估设置,而不取决于预报是由人工智能还是物理数值建模产生的。

更新时间: 2026-10-06 17:28:42

领域: physics.ao-ph,cs.AI,cs.LG

下载: http://arxiv.org/abs/2608.09972v2

Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus

Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus of 590 harness-triggered AppWorld compaction boundaries replays each boundary from a re-executed prefix state under the pre-compaction context and under the summary, and records the burden of the next actions: calls that error or repeat a call already made. We find that pre-boundary history predicts post-compaction harm only weakly. An internally prespecified contrast by prefix placement is a wide null, and the naive "has-written" label behind it turns out to measure trajectory phase. The best extension-protocol trigger reaches held-out AUROC 0.66 (0.64 on the replicate's own label) against a same-boundary replicate of 0.72; the best frozen, interpretable trigger avoids 21% of harmful (positive-burden) boundaries while keeping 84% of compaction opportunities, and exceeds the random-rule expectation on count but not on burden mass (a post hoc comparison). Whether the best trigger beats a token-budget rule at matched retention cannot be evaluated on the release. We state what corpora should ship to answer it.

Updated: 2026-10-06 17:26:57

标题: 代理历史是否可以告诉您何时压缩会造成伤害?对TRACE配对重放语料库的影响有限,但显著.

摘要: 许多长期视角的代理人将其上下文压缩为全局规则,通常是一个令牌预算,对代理人的活动一无所知。我们想知道代理人最近的行为是否能预测何时压缩会造成伤害。TRACE的公共语料库包含590个由AppWorld触发的压缩边界,每个边界都会回放从重新执行的前缀状态下的预压缩上下文和总结,并记录下一步行动的负担:错误调用或重复已经进行的调用。我们发现,边界之前的历史只能微弱地预测压缩后的伤害。一个内部预先设定的前缀位置对比是一个广泛的空值,并且幼稚的“已写入”标签背后的实际上测量了轨迹阶段。最佳的扩展协议触发器在保留AUROC 0.66(在复制的自身标签上为0.64)的情况下达到了0.72的同一边界的复制;最佳的冻结、可解释的触发器避免了21%的有害(正负担)边界,同时保留了84%的压缩机会,并且在计数上超过了随机规则的预期,但在负担质量上没有超过(事后比较)。最佳触发器是否能在匹配保留方面击败令牌预算规则无法在发布时评估。我们说明了语料库应该提供什么来回答这个问题。

更新时间: 2026-10-06 17:26:57

领域: cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08722v1

WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?

LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied AI, games and films. As the workhorse of such simulation, a solver computes how the state of a dynamic system evolves over time. Building such solvers requires physical understanding to identify appropriate models, mathematical reasoning to formulate the underlying dynamics, and software engineering to implement them as executable code, yet this capability of LLM agents remains underexplored. To this end, we introduce WorldSolver, a benchmark of 168 simulation tasks derived from physical phenomena in 61 classic computer graphics papers, spanning 7 physical domains. Each task contains a code scaffold that provides a fixed simulation environment for the scene, with the solver implementation left for the agent to complete. Specifically, we evaluate them along three dimensions: Execution Checks for successful execution, Visual Fidelity for reproducing the intended dynamic behavior in the rendered simulation, and Physical Plausibility for physics-grounded verification of the generated dynamics. Experiments on frontier agents reveal that producing executable solvers is difficult itself, and satisfying visual and physical correctness is even harder. GPT-5.6-Sol and Claude-Opus-5 perform comparatively better than the other evaluated agents, yet achieve overall scores of only 48.7% and 46.7%, respectively. WorldSolver is an early step toward agentic solver generation, and we hope it helps drive progress toward agents that can faithfully simulate the dynamic physical world. Code is available at https://github.com/sirujiang/WorldSolver.

Updated: 2026-10-06 17:26:26

标题: WorldSolver:LLM代理能否通过生成求解器模拟物理动力学?

摘要: 基于LLM的代理正在越来越多地推动科学和工程问题的解决,物理模拟正逐渐成为一个具有挑战性但实用的测试平台,用于重现具有在具体人工智能、游戏和电影中的应用的复杂物理现象。作为这种模拟的主力军,求解器计算动态系统状态随时间的演变。构建这样的求解器需要物理理解来识别适当的模型,数学推理来制定基础动态,并进行软件工程来将其实现为可执行代码,然而LLM代理的这种能力仍未得到充分开发。为此,我们介绍了WorldSolver,这是一个由61篇经典计算机图形论文中的物理现象衍生出的168个模拟任务的基准,涵盖了7个物理领域。每个任务都包含一个代码框架,为场景提供了一个固定的模拟环境,求解器的实现留给代理完成。具体来说,我们沿三个维度进行评估:执行检查用于成功执行,视觉保真度用于在渲染模拟中重现预期的动态行为,物理合理性用于对生成的动态进行基于物理的验证。对前沿代理的实验表明,生成可执行求解器本身就很困难,而满足视觉和物理正确性更加困难。GPT-5.6-Sol和Claude-Opus-5相对于其他评估的代理表现更好,但总体得分仅为48.7%和46.7%。WorldSolver是朝着代理求解器生成迈出的早期步骤,我们希望它有助于推动朝着能够忠实模拟动态物理世界的代理的进步。代码可在https://github.com/sirujiang/WorldSolver 上找到。

更新时间: 2026-10-06 17:26:26

领域: cs.AI

下载: http://arxiv.org/abs/2610.08720v1

XDecomposer: Learning Prior-Free Set Decomposition for Multiphase X-ray Diffraction

Multiphase powder X-ray diffraction (PXRD) analysis remains a fundamental bottleneck in structure identification, as real-world synthesis often produces complex mixtures whose constituent phases (components) cannot be reliably disentangled. While recent advances in representation-based crystal retrieval and generation suggest the possibility of inferring structures directly from PXRD, existing approaches largely assume single-phase inputs and break down in multiphase settings. Here, we present XDecomposer, a prior-free framework for joint decomposition and identification of multiphase XRD patterns without requiring candidate phase lists, structural templates, or prior knowledge of phase number. We formulate multiphase diffraction analysis as a set prediction problem, where the model infers an unordered set of phase-resolved components, their mixture proportions, and corresponding structural representations within a unified architecture. A phase-query-driven decomposition mechanism, together with diffraction-consistent physical reconstruction, enables accurate source separation while preserving crystallographic fidelity. Extensive experiments on both simulated and experimental datasets show that XDecomposer substantially improves reconstruction accuracy and phase identification across diverse chemical systems, while maintaining strong generalization to unseen mixtures. These results provide a practical route toward data-driven, source-resolved multiphase XRD analysis and reduce long-standing dependence on prior-guided iteratively phase matching. The code is openly available at https://github.com/Licht0812/XDecomposer

Updated: 2026-10-06 17:25:55

标题: XDecomposer:学习无先验知识的多相X射线衍射集合分解

摘要: 多相粉末X射线衍射(PXRD)分析仍然是结构鉴定中的一个基本瓶颈,因为真实世界中的合成通常会产生复杂的混合物,其组成相(成分)无法可靠地分离。尽管最近在基于表示的晶体检索和生成方面取得了进展,表明可以直接从PXRD推断出结构,但现有方法很大程度上假设单相输入,并在多相环境中失效。在这里,我们提出了XDecomposer,这是一个无先验框架,用于联合分解和识别多相XRD图谱,无需候选相列表、结构模板或阶段数量的先验知识。我们将多相衍射分析构建为一个集合预测问题,模型推断出一组无序的相分辨组分、它们的混合比例和相应的结构表示,全部在一个统一的架构中。相查询驱动的分解机制,以及衍射一致的物理重建,使得准确的源分离同时保持晶体学的忠实度。对模拟和实验数据集的广泛实验表明,XDecomposer显著提高了对各种化学体系的重建准确性和相位识别,同时在未见混合物上保持很强的泛化能力。这些结果为基于数据驱动的、源分辨的多相XRD分析提供了一条实用途径,并减少了长期依赖于先验引导的迭代阶段匹配。代码可在https://github.com/Licht0812/XDecomposer 上公开获取。

更新时间: 2026-10-06 17:25:55

领域: cs.AI,cond-mat.mtrl-sci,cs.LG

下载: http://arxiv.org/abs/2605.05866v2

When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting

Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.

Updated: 2026-10-06 17:25:48

标题: 当遗忘并非灾难性时:关于虚假遗忘机制的研究

摘要: 在微调期间似乎被遗忘的语言模型知识通常仍然被存储,并且可以被恢复,这种现象称为虚假遗忘。对新事实进行微调甚至可能产生自我消除的遗忘:旧事实的召回崩溃,当仅在新事实上继续训练时,召回才会恢复,然后才会永久侵蚀。我们试图理解这种遗忘何时不会是灾难性的。一个最小的联想记忆使用三个要素再现这些动态:具有共享结构的键、集中的新值和网络中的归一化。微调将所有旧表示沿着一个共同的方向移动,隐藏旧事实同时保留它们的相对几何结构;归一化在学习新事实后撤回这种转变,而事实特定的变化则积累并导致侵蚀。此外,在合成数据上训练的Transformer中减去共同转变可以消除崩溃,并且从每个权重更新中去除一个单一方向可以恢复预训练语言模型中的旧事实。因此,遗忘结合了共享的、可逆的访问丢失与个体事实的缓慢侵蚀,而仅有后者是灾难性的。哪一个占主导取决于新数据是否将旧记忆移动在一起还是分开。

更新时间: 2026-10-06 17:25:48

领域: cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.08718v1

Co-Evolving Paths and Flows via Path-Flow Alignment

We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow network using the same alignment loss: the flow learns to match the path velocity, and the path learns to align its velocity to the current flow. Although every fixed learned path defines a valid flow-matching objective, the alignment loss alone is not a reliable criterion for path learning. We identify path overfitting, a failure mode in which the alignment loss decreases while sample quality worsens. We find that this failure is associated with low-entropy bottlenecks in the induced probability path, where the learned path routes samples through overly concentrated intermediate marginals. Motivated by this diagnosis, we introduce a stochastic path regularizer that hides part of the source information from the path network while preserving exact endpoints. The resulting regularization gives an explicit entropy floor for the stochastic training-path marginals and empirically suppresses the bottleneck in the learned sampler, making joint path-flow training effective. On ImageNet-256x256 with SiT backbones, our method consistently improves FID across model scales, extends to model-guidance training, and leaves the inference-time architecture and sampler unchanged. Code is available at https://github.com/lizeyu090312/traj_opt_paper

Updated: 2026-10-06 17:24:36

标题: 共同进化的路径和流通过路径-流对齐

摘要: 我们研究路径-流对齐作为流匹配的统一训练目标。与固定插值路径并仅学习速度场不同,我们同时训练端点保持路径网络和流网络,使用相同的对齐损失:流学习匹配路径速度,路径学习将其速度与当前流对齐。尽管每个固定学习的路径定义了一个有效的流匹配目标,但仅仅依靠对齐损失并不是路径学习的可靠标准。我们确定了路径过拟合,一种失败模式,其中对齐损失减少,但样本质量变差。我们发现,这种失败与在诱导的概率路径中的低熵瓶颈有关,学习路径通过过于集中的中间边缘样本。受到这一诊断的启发,我们引入了一种随机路径正则化器,隐藏了部分源信息,同时保持确切的端点。结果正则化为随机训练路径边缘提供了明确的熵底线,并在经验上抑制了学习的采样器中的瓶颈,使联合路径-流训练有效。在ImageNet-256x256和SiT骨干上,我们的方法始终改善了不同模型规模下的FID,扩展到模型引导训练,并且在推理时间架构和采样器保持不变。代码可在https://github.com/lizeyu090312/traj_opt_paper上找到。

更新时间: 2026-10-06 17:24:36

领域: cs.CV,cs.LG

下载: http://arxiv.org/abs/2610.08717v1

Cross-Lingual Activation Steering for Multilingual Language Models

Large language models exhibit strong multilingual capabilities, yet significant performance gaps persist between dominant and non-dominant languages. Prior work attributes this gap to imbalances between shared and language-specific neurons in multilingual representations. We propose Cross-Lingual Activation Steering (CLAS), a training-free inference-time intervention that selectively modulates neuron activations. We evaluate CLAS on classification and generation benchmarks, achieving average improvements of 2.3% (Acc.) and 3.4% (F1) respectively, while maintaining high-resource language performance. We discover that effective transfer operates through functional divergence rather than strict alignment; performance gains correlate with increased language cluster separation. Our results demonstrate that targeted activation steering can unlock latent multilingual capacity in existing models without modification to model weights.

Updated: 2026-10-06 17:23:35

标题: 多语言语言模型的跨语言激活引导

摘要: 大型语言模型展现出强大的多语言能力,然而主导语言和非主导语言之间仍然存在显著的性能差距。先前的研究将这种差距归因于多语言表示中共享和特定语言神经元之间的不平衡。我们提出了一种名为跨语言激活引导(CLAS)的训练无干预干预方法,在推断时间有选择性地调节神经元激活。我们在分类和生成基准测试中评估了CLAS,在保持高资源语言性能的同时,分别实现了2.3%(准确度)和3.4%(F1)的平均改进。我们发现有效的迁移是通过功能分歧而不是严格的对齐实现的;性能提升与语言集群分离的增加相关。我们的结果表明,有针对性的激活引导可以在不修改模型权重的情况下释放现有模型中的潜在多语言能力。

更新时间: 2026-10-06 17:23:35

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2601.16390v2

KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization

Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kernels, yet treat compiled models as black boxes, generally optimizing individual standalone kernels without respecting the compiler's structural decisions or verifying the model end-to-end. We present KernelOPT, a multi-agent system that treats compiled models as structured artifacts. It preserves vendor library calls (cuBLAS, cuDNN) and exclusively targets generated Triton sub-kernels using five profiling-guided LLM agents. A four-gate verification cascade applies static validation, multi-seed correctness checking, model-level float64-fallback verification, and performance gating ($γ{=}1.03$) to filter candidates and verify the re-stitched model end-to-end. When candidates fail verification, the system preserves the compiler baseline. The system accepts PyTorch nn Modules, standalone Triton kernels, and Helion kernels. Evaluated on 250 KernelBench problems (100 Level 1, 100 Level 2 and 50 Level 3) on NVIDIA H200, KernelOPT achieves geometric mean speedups over torch compile of 1.40$\times$ (L1), 1.15$\times$ (L2), and 1.07$\times$ (L3) across all kernels, including fallback cases. Optimized-only geomeans (excluding cases where verification gates preserve the compiler baseline) are substantially higher: 2.54$\times$ (L1: 36/100), 1.84$\times$ (L2: 23/100), and 1.37$\times$ (L3: 11/50), reflecting where the optimizer achieves meaningful leverage.

Updated: 2026-10-06 17:23:07

标题: KernelOPT: 面向GPU内核优化的调度感知代理式搜索

摘要: 深度学习推理和训练性能在很大程度上取决于GPU核心的效率。现代编译器如PyTorch Inductor会自动生成高级模型代码的GPU核心,但通常表现不如专家编写的实现。最近的LLM辅助核心优化器可以弥合独立核心之间的差距,但会将编译模型视为黑匣子,通常优化单独的核心而不考虑编译器的结构决策或验证模型端到端。我们提出了KernelOPT,这是一个将编译模型视为结构化工件的多智能体系统。它保留了供应商库调用(cuBLAS、cuDNN),并专门针对使用五个基于性能分析的LLM智能体生成的Triton子核心。一个四门验证级联应用静态验证、多种子正确性检查、模型级float64回退验证和性能门控(γ=1.03)来过滤候选项并验证重新拼接的模型端到端。当候选项未通过验证时,系统保留编译器基线。该系统接受PyTorch nn模块、独立的Triton核心和Helion核心。在NVIDIA H200上对250个KernelBench问题(100个一级,100个二级和50个三级)进行评估,KernelOPT在所有核心上(包括回退情况)相对于torch compile实现了1.40倍(L1),1.15倍(L2)和1.07倍(L3)的几何平均加速。仅优化的几何平均值(不包括验证门保留编译器基线的情况)明显更高:2.54倍(L1:36/100),1.84倍(L2:23/100)和1.37倍(L3:11/50),反映了优化器实现有意义杠杆的地方。

更新时间: 2026-10-06 17:23:07

领域: cs.DC,cs.AI,cs.LG

下载: http://arxiv.org/abs/2609.30059v2

Prediction-powered inference for time series across space

The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data is observed over many spatial locations. For instance, crop yield might be observed over a large geographical area for recent years, but weather data (which is informative about crop yield) is available for a much longer period. The goal is to estimate, at each spatial location, the expected label (e.g., crop yield) in the future and provide a valid confidence interval for this value. The observed time period alone is too short for reliable estimates. Imputing missing labels with machine learning can cause substantial bias. Prediction-powered inference (PPI) can correct for this bias, but it relies on an i.i.d. assumption that breaks under our expected temporal dependencies. Heteroskedasticity and autocorrelation consistent (HAC) procedures account for temporal correlation, but have not been adapted to cases where some labels are imputed. We provide reliable point estimates and confidence intervals given: short labeled time series (across spatial locations), a longer unlabeled time series, and an imperfect predictor of labels given covariates. We show our method outperforms natural alternatives.

Updated: 2026-10-06 17:22:39

标题: 基于预测的空间时间序列推断

摘要: 在时空设置中,常见的情形是我们观察到一系列协变量和标签对,这些对是在相对较短的最近时间段内观察到的。我们可以访问到长时间段内的未标记协变量。数据是在许多空间位置上观察到的。例如,农作物产量可能在最近几年内在一个较大的地理区域内观察到,但天气数据(对农作物产量具有信息量)则可在一个更长的时间段内获得。目标是在每个空间位置上估计未来期望的标签(例如,农作物产量)并为该值提供有效的置信区间。仅仅观察到的时间段太短,无法提供可靠的估计。使用机器学习来填补缺失的标签可能会导致相当大的偏差。基于预测的推断(PPI)可以纠正这种偏差,但它依赖于一个独立同分布的假设,而这个假设在我们的预期时间依赖下会被打破。异方差性和自相关一致性(HAC)程序考虑了时间相关性,但尚未适应一些标签被填补的情况。我们提供可靠的点估计和置信区间,考虑到:短标记时间序列(跨空间位置)、更长的未标记时间序列以及给定协变量的不完美标签预测器。我们展示了我们的方法优于自然替代方案。

更新时间: 2026-10-06 17:22:39

领域: stat.ME,cs.LG,stat.ML

下载: http://arxiv.org/abs/2610.08715v1

Sensitivity Shaping for Latent Modeling

Generative dynamics models enable planning in challenging systems, but safe deployment requires detecting policy-induced out-of-distribution (OOD) transitions. Existing methods typically treat learned dynamics as fixed and rely on post hoc support surrogates for OOD detection. This overlooks a critical failure mode: learned dynamics that are insensitive to control changes can map unsupported controls to latent predictions resembling demonstrated transitions, suppressing OOD signals despite large prediction errors. We introduce support-conditioned control-sensitivity regularization to preserve control-induced variation by promoting local responsiveness in well-supported training regions. Experiments in vision-based obstacle avoidance, manipulation, and real-robot navigation demonstrate improved OOD detection and safer closed-loop planning.

Updated: 2026-10-06 17:21:45

标题: 潜在建模的敏感性塑造

摘要: 生成动态模型使规划在具有挑战性的系统中成为可能,但安全部署需要检测由策略引起的超出分布(OOD)转换。现有方法通常将学习的动态视为固定的,并依赖事后支持替代品来进行OOD检测。这忽视了一个关键的失败模式:对控制变化不敏感的学习动态可能将不受支持的控制映射到类似于演示转换的潜在预测,从而抑制OOD信号,尽管存在较大的预测错误。我们引入支持条件控制敏感性正则化,通过在受支持的训练区域中促进局部响应性来保留控制引起的变化。基于视觉的避障、操纵和真实机器人导航的实验表明,改进了OOD检测和更安全的闭环规划。

更新时间: 2026-10-06 17:21:45

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2606.14585v2

PRUE: A Practical Recipe for Field Boundary Segmentation at Scale

Large-scale maps of field boundaries are essential for agricultural monitoring tasks. Existing deep learning approaches for satellite-based field mapping are sensitive to illumination, spatial scale, and changes in geographic location. We conduct the first systematic evaluation of segmentation and geospatial foundation models (GFMs) for global field boundary delineation using the Fields of The World (FTW) benchmark. We evaluate 18 models under unified experimental settings, showing that a U-Net semantic segmentation model outperforms instance-based and GFM alternatives on a suite of performance and deployment metrics. We propose a new segmentation approach that combines a U-Net backbone, composite loss functions, and targeted data augmentations to enhance performance and robustness under real-world conditions. Our model achieves a 76% IoU and 47% object-F1 on FTW, an increase of 6% and 9% over the previous baseline. Our approach provides a practical framework for reliable, scalable, and reproducible field boundary delineation across model design, training, and inference. We release all models and model-derived field boundary datasets for five countries.

Updated: 2026-10-06 17:17:19

标题: PRUE:大规模田地边界分割的实用配方

摘要: 大规模的田界地图对于农业监测任务至关重要。现有的基于深度学习的卫星图像田地绘制方法对光照、空间尺度和地理位置的变化敏感。我们首次对全球田界划分的分割和地理空间基础模型(GFMs)进行系统评估,使用Fields of The World(FTW)基准测试。我们在统一的实验设置下评估了18种模型,结果显示,U-Net语义分割模型在性能和部署指标套件上优于基于实例和GFM的替代方案。我们提出了一种新的分割方法,结合了U-Net骨干、复合损失函数和定向数据增强,以增强性能和在实际条件下的稳健性。我们的模型在FTW上实现了76%的IoU和47%的目标F1,比以前的基线提高了6%和9%。我们的方法提供了一个可靠、可扩展和可重复的田界划分框架,涵盖了模型设计、训练和推断。我们发布了五个国家的所有模型和模型衍生的田界数据集。

更新时间: 2026-10-06 17:17:19

领域: cs.CV,cs.LG

下载: http://arxiv.org/abs/2603.27101v2

nanoMuse: An Open-Source Personal Agent for Every Device You Own

Assistants from 2011 answered and waited, and agents from 2023 did a task and stopped. In September 2026 Meta's Muse showed an agent for one person, with accounts, devices, memory and a conversation that lasts, closed, in a vendor's cloud, in one country. Such an agent is expected to act on a person's accounts and devices, remember them across weeks, speak first when it is worth it, and answer for what it did. It is a kind of software, not a model, and until now had no open counterpart. This report defines the personal agent in five questions and three horizons. It reads how Muse is built from Meta's public record and a copy of its production prompt, each statement marked by its source. It then presents nanoMuse, the open-source counterpart under the GPL-3.0, one agent on every device a person owns, with hands on the phone's screen and the computer's. They share one conversation over a relay anyone can run; every action goes through a Sentinel, memory is files the person can read, and the model is their choice. Its size and cost are given as estimates. What is open, memory with provenance, an evaluation suite for the hands and an open model for them, is set out as a roadmap.

Updated: 2026-10-06 17:13:53

标题: 纳米博物馆:为您拥有的每个设备提供的开源个人代理

摘要: 2011年的助手回答并等待,而2023年的代理完成了任务并停止。在2026年9月,Meta的Muse展示了一个人的代理,具有账户、设备、记忆和持续对话,在一个国家的供应商云中关闭。这样的代理预计将对一个人的账户和设备采取行动,跨越几周记住它们,在值得时首先开口,并为其所做的事情作出回应。这是一种软件,不是一个模型,直到现在都没有开放的对应物。这份报告通过五个问题和三个视角定义了个人代理。它解释了Muse是如何从Meta的公开记录和其生产提示的副本构建的,每个陈述都标有其来源。然后介绍了nanoMuse,作为GPL-3.0开源对应物,每个人拥有的设备上的一个代理,手放在手机屏幕和电脑上。他们通过任何人都可以运行的中继进行一次对话;每个行动都经过一个哨兵,记忆是人可以阅读的文件,模型是他们的选择。它的尺寸和成本被给出为估计值。开放的是带来源的记忆,为手提供评估套件,以及为他们提供开放模型,这被制定为一份路线图。

更新时间: 2026-10-06 17:13:53

领域: cs.AI

下载: http://arxiv.org/abs/2610.08699v1

GeneICL: A Tabular Foundation Model for Bulk Transcriptomics

Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic foundation models often fail to outperform simple supervised baselines. Tabular foundation models offer an alternative through in-context learning, but are typically pretrained on generic synthetic data rather than transcriptomic structure. We ask whether transcriptomics-aware pretraining, rather than scale, is the missing ingredient. Towards this end, we introduce GeneICL, a 4.2M-parameter tabular foundation model combining a semi-synthetic pretraining prior built from measured bulk expression profiles with a parameter-efficient recurrent architecture. We further enable right-censored survival prediction via a training-free reduction to regression using Cox partial-likelihood residuals. We evaluate GeneICL on 80 clinical outcome-prediction tasks spanning classification, regression, and survival. Tabular foundation models consistently outperform self-supervised transcriptomic models, while GeneICL achieves the best overall rank among evaluated foundation models and tuned baselines. GeneICL does so with up to 387$\times$ fewer parameters, no gradient updates at inference, and predictions within seconds on a laptop CPU.

Updated: 2026-10-06 17:10:37

标题: GeneICL:用于大规模转录组学的表格基础模型

摘要: 基因表达在生物医学领域被广泛测量,然而临床结果预测仍然具有挑战性,原因是高维度、强特征相关性和有限的标记数据。大型自监督转录组基础模型通常无法超越简单的监督基线。表格基础模型通过上下文学习提供了一种替代方法,但通常是在通用合成数据上进行预训练,而不是转录组结构。我们问自己,转录组感知预训练,而不是规模,是否是缺失的要素。为此,我们介绍了GeneICL,一个包含半合成预训练先验的4.2M参数表格基础模型,该预训练先验是从测量的批量表达谱建立的,并采用了参数高效的循环架构。我们进一步通过使用Cox偏似然残差将其转换为回归来实现右截尾生存预测,无需训练。我们在80个临床结果预测任务上评估了GeneICL,涵盖分类、回归和生存。表格基础模型始终优于自监督转录组模型,而GeneICL在评估的基础模型和调整基线中实现了最佳的综合排名。GeneICL的参数数量最多减少了387倍,在推断时没有梯度更新,并且在笔记本电脑CPU上可以在几秒内进行预测。

更新时间: 2026-10-06 17:10:37

领域: cs.LG

下载: http://arxiv.org/abs/2610.08694v1

ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences

Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at https://github.com/beita6969/ScienceClaw.

Updated: 2026-10-06 17:08:03

标题: ScienceClaw:跨自然和社会科学领域的AI-for-Science代理的持续自我进化基准测试

摘要: 大型语言模型代理正在加速科学自动化,然而经过验证的执行很少成为持久的程序级改进,现有的评估也没有跨自然和社会科学中的连续任务来检查这个过程。我们将ScienceClaw形式化为固定参数程序自进化,统一任务解决、科学验证和程序更新。ScienceClaw-Eval涵盖了23个学科,通过顺序流和独立重置评估,衡量了科学正确性、进化增益、保留、跨数据集传输和演化成本。我们的框架通过多轮交互修复可执行工作流程,将重新执行验证的失败成功轨迹转换为链接的技能和操作者候选人,并仅在源任务重放产生修复并独立科学任务改进时保留更新。代码可在https://github.com/beita6969/ScienceClaw找到。

更新时间: 2026-10-06 17:08:03

领域: cs.AI

下载: http://arxiv.org/abs/2610.08691v1

Probabilistic Counterfactual Inference for Discrete Outcomes in Gaussian-Process Causal Models

Counterfactual inference in Gaussian-process structural causal models (GP-SCMs) has been developed primarily for continuous endogenous variables, limiting applicability to causal graphs that contain discrete child nodes with continuous parents. We introduce a unified probabilistic framework for counterfactual inference with heterogeneous variable types by pairing GP predictors with explicit exogenous noise mechanisms. For discrete outcomes, we derive exact conditional noise-abduction procedures using a uniform threshold for binary variables, a Gumbel-max race for nominal categories, and a latent Gaussian cut-point model for ordinal ones. In each case, we propagate abducted noise through interventions while accounting for posterior uncertainty in the GP latent functions, and prove that the resulting mechanisms reproduce the fitted model's observational and interventional distributions. On synthetic SCMs with known ground-truth counterfactuals, we evaluate estimation accuracy, consistency, and robustness to coupling misspecification. A key finding is that applying a categorical coupling to ordinal data inflates counterfactual error roughly threefold even when observational fit remains comparable, and that this error does not diminish with more data. As the training set grows, the fitted structural equation converges to the truth while the counterfactual error flattens onto a floor. In the reverse direction, forcing a false order onto nominal data instead degrades the fitted equation itself. The choice of coupling must therefore be justified on structural grounds rather than read off the fit.

Updated: 2026-10-06 17:06:22

标题: 高斯过程因果模型中离散结果的概率反事实推断

摘要: 高斯过程结构因果模型(GP-SCMs)中的反事实推断主要针对连续内生变量进行了开发,限制了适用于包含离散子节点和连续父节点的因果图。我们通过将GP预测器与显式外生噪声机制配对,引入了一种用于异质变量类型的反事实推断的统一概率框架。对于离散结果,我们使用二元变量的统一阈值、名义类别的Gumbel-max竞赛和有序变量的潜在高斯切点模型,推导出精确的条件噪声抽取程序。在每种情况下,我们通过干预传播被抽取的噪声,同时考虑GP潜在函数的后验不确定性,并证明结果机制重现了拟合模型的观测和干预分布。在具有已知地面真实反事实的合成SCMs上,我们评估了估计精度、一致性和对耦合误差规范的稳健性。一个关键发现是,即使观测拟合保持可比性,将分类耦合应用于有序数据也会使反事实误差大约增加三倍,并且这种错误不会随着更多数据而减少。随着训练集的增长,拟合的结构方程趋于真实,而反事实错误则趋于一个底线。在相反方向上,对名义数据强加错误的顺序会导致拟合的方程本身退化。因此,耦合的选择必须基于结构基础来进行正当化,而不是从拟合中读取。

更新时间: 2026-10-06 17:06:22

领域: cs.LG

下载: http://arxiv.org/abs/2610.08689v1

MSPR: Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning

This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations that align state and goal latents is a challenge, as the encoder can learn goal-agnostic features that destabilize policy learning. We address this issue by learning the encoder's representation with alignment objectives that capture the environment across multiple scales, from local physical dynamics to long-horizon goal-directed structure. Concretely, we propose MSPR, a framework that leverages multi-scale predictive supervision to enforce goal-directed alignment within the latent space. We demonstrate that MSPR leads to strong performance on both vision and state-based tasks. Furthermore, we show that our approach is resilient under realistic, challenging data regimes, maintaining state-of-the-art performance across a wide variety of tasks.

Updated: 2026-10-06 17:05:14

标题: MSPR:用于目标条件强化学习的多尺度预测表示

摘要: 本文研究了离线目标条件强化学习(GCRL)中的鲁棒表示学习。特别是在稀疏奖励场景中,学习对齐状态和目标潜变量的表示是一项挑战,因为编码器可能学习到不考虑目标的特征,从而破坏策略学习。我们通过学习编码器的表示,使用能够捕捉从局部物理动态到长程目标导向结构的多尺度环境的对齐目标来解决这个问题。具体地,我们提出了MSPR,这是一个利用多尺度预测监督来强化潜变量空间内目标导向对齐的框架。我们展示了MSPR在视觉和基于状态的任务上表现出强大的性能。此外,我们展示了我们的方法在现实和具有挑战性的数据情境下具有韧性,保持在各种任务中的最新性能。

更新时间: 2026-10-06 17:05:14

领域: cs.LG

下载: http://arxiv.org/abs/2605.09364v2

Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue

Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop settles into. Two PersonaPlex-7B instances exchange audio tokens on a shared clock in unscripted conversation, and one floor-transfer rule is applied to them and to Switchboard. Their timing is coupled: re-pairing speakers across conversations destroys it. But the floor changes hands late, at a median of 400-560 ms against 137 ms for humans, and the last 120 ms of the partner's turn, where human projection places a tenth of its transfers, holds 1% of theirs. Delaying one direction of the channel shifts the response one-for-one and leaves the run-up to it empty, consistent with a reactive wait after the perceived end rather than the turn-end projection human timing requires.

Updated: 2026-10-06 17:03:54

标题: 耦合但晚:非剧本对话中全双工语音模型之间的交替

摘要: 全双工语音模型被训练与人对话,但越来越多地被用于彼此对话,进行自我生成数据、代理社区和基于模型的评估。在这个循环中,没有人吸收定时错误:每个模型的轮流对话是另一个模型的输入。我们想知道这个循环会定格在什么时机。两个PersonaPlex-7B实例在无剧本对话中共享时钟交换音频令牌,并对它们和Switchboard应用一个换手规则。它们的定时是耦合的:跨对话重新配对说话者会破坏这种耦合。但换手发生得较晚,中位数为400-560毫秒,而人类为137毫秒;在对话伴侣的最后120毫秒中,人类将其一成的转换放置在这里,而他们只占1%。延迟通道的一个方向会导致响应按比例延迟,并使其前期为空,与人类定时所需的转换结束投射不一致。这表明出,人类的定时需要对感知结束做出反应性等待。

更新时间: 2026-10-06 17:03:54

领域: cs.AI

下载: http://arxiv.org/abs/2610.08683v1

A Systematic Study of Small Language Models on Abstract Reasoning Tasks

Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.

Updated: 2026-10-06 17:01:19

标题: 一个小语言模型在抽象推理任务上的系统研究

摘要: 摘要:在抽象推理基准上的端点准确性并不能揭示一个语言模型是否已经掌握了可转移的规则或者适配了特定分布的规律性。我们在ARC-TGI基准上研究了小型语言模型之间的区别,该基准将抽象网格变换组织成可控的任务系列,并支持重采样、空间偏移和跨基准的转移。在超级微调下,我们通过超过1,000次运行,对解码器、编码器-解码器和专家混合模型族进行了剖析。我们研究了技能获取的效率和稳定性、训练分布之外的稳健性、与模型族和任务制定之间的相互作用,以及伴随行为差异的逐层注意力特征。虽然在训练分布内可以获得可观的准确性,但获取过程对优化敏感,并在任务系列之间分布不均匀。在训练分布之外,包括保留规则但网格规模发生改变时,性能急剧下降。训练集的深度和广度会带来不均匀的收益,而额外的上下文示例对模型族的影响取决于模型族。可执行规则归纳还可以产生在直接网格生成下未观察到的正确解决方案。在选定的任务上,注意力诊断显示出不同的集中度和依赖上下文的特征,但并不能建立一般的因果机制。总的来说,抽象推理分数取决于模型、适应制度、评估分布和响应格式。

更新时间: 2026-10-06 17:01:19

领域: cs.LG,cs.CL

下载: http://arxiv.org/abs/2610.08680v1

Secure Speculative Decoding for Large Language Models

Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emph{target model}, by first using a smaller model, referred to as the \emph{draft model}, to generate candidate tokens and then verifying them with the target model for acceptance or rejection. Prior studies primarily focused on the efficiency-utility trade-off of speculative decoding, e.g., lossy speculative decoding, leaving its security implications largely unexplored. In this work, we bridge this gap by providing the \emph{first} systematic study of the security implications of speculative decoding. Through a large-scale measurement study, we reveal a pronounced security-utility asymmetry: across a wide range of lossy speculative decoding methods, improvements in inference efficiency come at a disproportionately high cost to security, with attack success rates for jailbreak and prompt injection attacks increasing much faster than utility degrades. We then propose SecureSD, a new theory-guided speculative decoding method that enhances security while maintaining efficiency and utility. Specifically, our theoretical analysis reveals that security degradation primarily originates from the early tokens generated by the draft model. Motivated by this insight, SecureSD applies a stricter verification criterion to draft-model tokens at early decoding positions. Extensive experiments on both security and utility benchmarks demonstrate that SecureSD significantly improves security while preserving efficiency and utility compared to existing speculative decoding methods.

Updated: 2026-10-06 16:59:37

标题: 大规模语言模型的安全推断解码

摘要: 这项研究提出了一种名为SecureSD的新理论引导的推测解码方法,它在保持效率和效用的同时增强了安全性。具体而言,我们的理论分析揭示了安全性降级主要源自草稿模型生成的早期令牌。受到这一洞察的启发,SecureSD在早期解码位置对草稿模型令牌应用更严格的验证标准。对安全性和效用基准的大量实验表明,与现有推测解码方法相比,SecureSD显著提高了安全性,同时保持了效率和效用。

更新时间: 2026-10-06 16:59:37

领域: cs.CR,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08678v1

Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling

Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights. A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model. In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR. Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance. We derive the population-optimal coefficient in closed form and show that the resulting estimator has variance no larger than either endpoint, OffCEM or DR. Experiments in controlled synthetic settings and on two large-action benchmarks show that VOCEM improves upon both endpoints in all 23 evaluated conditions, exhibiting greater stability and empirical robustness.

Updated: 2026-10-06 16:58:26

标题: 用联合效果建模进行方差最优的离线策略评估

摘要: 在上下文臂策略的离线评估(OPE)中,当动作级别的重要性加权导致过多方差时,会变得具有挑战性。双重稳健(DR)估计在常见支持下保持无偏,但保留这些高方差的动作级别权重。一种先前的估计器,称为带有Conjunct Effect Model(OffCEM)的离线评估,用更稳定的集群级权重替换它们,但需要依赖奖励模型的局部正确性。在本文中,我们展示了,在DR和OffCEM需要的假设下,存在一组无偏估计器,可以在OffCEM和DR之间进行插值。基于这一结果,我们提出了方差最优CEM(VOCEM)估计器,它选择插值系数以最小化方差。我们推导出了封闭形式的最优人群系数,并展示了由此得出的估计器的方差不会大于任一端点,即OffCEM或DR。在受控合成环境和两个大动作基准测试中的实验表明,VOCEM在所有23个评估条件中都优于这两个端点,表现出更大的稳定性和实证鲁棒性。

更新时间: 2026-10-06 16:58:26

领域: cs.LG

下载: http://arxiv.org/abs/2610.08677v1

When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

Large language models can follow complex instructions in a single turn, yet over long multi-turn interactions they often lose the thread of instructions, persona, and rules. This degradation has been measured behaviorally but not mechanistically explained. We propose a channel-transition account: goal-defining tokens become less accessible through attention, while goal-related information may persist in residual representations. We introduce the Goal Accessibility Ratio (GAR), measuring attention from generated tokens to task-defining goal tokens, and combine it with sliding-window ablations and residual-stream probes. When attention to instructions closes, what survives reveals architecture. Across architectures, the transition yields qualitatively distinct failure modes: some models preserve goal-conditioned behavior at vanishing attention, others fail despite decodable residual goal information, and the layer at which this encoding emerges varies from 2 to 27. A within-model causal ablation that force-closes the attention channel in Mistral collapses recall from near-perfect to 11% on a 20-fact retention task and raises persona-constraint violations above an adversarial-pressure baseline without user pressure, with both effects emerging at the predictable crossover turn. Linear probes recover per-episode recall outcomes from residual representations with AUC up to 0.99 across all four primary architectures, while input embeddings remain at chance. Across architectures and model scales, the gap between attention loss and residual decodability predicts whether goal-conditioned behavior survives channel closure. We contribute GAR as a diagnostic, the channel-transition framework as a controlled mechanistic account, and a parametric prediction of failure timing under windowed attention closure.

Updated: 2026-10-06 16:57:10

标题: 当注意力分散:LLMs 如何在多轮互动中失去主题

摘要: 大型语言模型可以在单个对话中遵循复杂的指令,但是在长时间的多轮交互中,它们经常会失去指令、个性和规则的线索。这种退化已经在行为上得到了测量,但尚未得到机械性的解释。我们提出了一个通道转换模型:目标定义的标记通过注意力变得不太容易访问,而与目标相关的信息可能会在残余表示中存在。我们引入了目标可访问性比率(GAR),用于衡量生成标记与任务定义的目标标记之间的注意力,并将其与滑动窗口消融和残余流探针相结合。当对指令的注意力关闭时,存活下来的内容会揭示架构。在不同的架构中,这种转换产生了定性不同的失败模式:一些模型在注意力消失时保留了目标条件行为,而其他一些模型尽管存在可解码的残余目标信息,但仍然失败,并且这种编码出现的层次从第2层到第27层不等。在Mistral模型中,通过强制关闭注意力通道的模型内因果消融将一个20个事实的保留任务的召回率从接近完美的状态降至11%,并将角色约束违规提高到对抗压力基准线以上,而没有用户压力,这两种效果在可预测的交叉转折点出现。线性探针可以从残余表示中恢复出每一集回忆结果,四个主要架构中的AUC高达0.99,而输入嵌入则保持在偶然的情况下。在各种架构和模型规模中,注意力丧失和残余可解码性之间的差距可以预测目标条件行为是否在通道关闭时存活。我们将GAR作为一种诊断工具,通道转换框架作为一个受控的机械化解释,并在窗口注意力关闭下对失败时间进行参数化预测。

更新时间: 2026-10-06 16:57:10

领域: cs.AI,cs.CL

下载: http://arxiv.org/abs/2605.12922v2

How Children Design and Reason about Trustworthy AI Chatbots

Children increasingly interact with AI chatbots, making trust calibration essential to AI literacy. Prior research has examined children's trust in AI mainly as users evaluating systems built by others, rather than as designers of their own chatbots. We developed a chatbot-building environment with adjustable trust-relevant traits (e.g., confidence, transparency, formality, assertiveness), rules, and persona. We conducted mixed-methods study with 115 learners (ages 8-18) who made 119 chatbots. We examined how children configured their chatbots, reasoned about trustworthiness, and how closely chatbot behavior aligned with their designs. Younger students (age 10-13) set significantly higher confidence than older students (age 14-18), and some deliberately built chatbots that gave wrong answers on purpose, yet still called them trustworthy, arguing that a chatbot does what it was built to do. Younger students equated trust with purpose-fulfillment, while older students linked it to transparent, calibrated design. Students also calibrated academic chatbots to be more transparent and formal than hobby chatbots. We identify seven design dimensions describing what children believe makes a chatbot trustworthy, and discuss implications for AI literacy tools.

Updated: 2026-10-06 16:56:51

标题: 孩子们如何设计和思考可信赖的人工智能聊天机器人

摘要: 儿童与人工智能聊天机器人的互动日益增加,因此信任校准对于人工智能素养至关重要。先前的研究主要关注儿童作为评估他人构建系统的用户时对人工智能的信任,而不是作为他们自己聊天机器人的设计者。我们开发了一个具有可调整信任相关特征(例如自信度、透明度、正式性、坚定性)、规则和个性的聊天机器人构建环境。我们进行了一项混合方法研究,共有115名学习者(8-18岁)制作了119个聊天机器人。我们研究了儿童如何配置他们的聊天机器人,如何推理其可信度以及聊天机器人的行为与他们的设计有多接近。年龄较小的学生(10-13岁)设置的自信度明显高于年龄较大的学生(14-18岁),有些学生故意构建了故意给出错误答案的聊天机器人,但仍然称其为可信的,并认为聊天机器人只是按照其构建的目的行事。年龄较小的学生将信任等同于目的实现,而年龄较大的学生将其与透明、校准的设计联系起来。学生还将学术聊天机器人校准为比兴趣爱好聊天机器人更加透明和正式。我们确定了七个设计维度,描述了儿童认为使聊天机器人可信的特征,并讨论了对于人工智能素养工具的影响。

更新时间: 2026-10-06 16:56:51

领域: cs.HC,cs.AI

下载: http://arxiv.org/abs/2609.25244v3

Multi-Probe Zero Collision Hash (MPZCH): Mitigating Embedding Collisions and Enhancing Model Freshness in Large-Scale Recommenders

Embedding tables are critical components of large-scale recommendation systems, facilitating the efficient mapping of high-cardinality categorical features into dense vector representations. However, as the volume of unique IDs expands, traditional hash-based indexing methods suffer from collisions that degrade model performance and personalization quality. We present Multi-Probe Zero Collision Hash (MPZCH), a novel indexing mechanism based on linear probing that effectively mitigates embedding collisions. With reasonable table sizing, it often eliminates these collisions entirely while maintaining production-scale efficiency. MPZCH utilizes auxiliary tensors and high-performance CUDA kernels to implement configurable probing and active eviction policies. By retiring obsolete IDs and resetting reassigned slots, MPZCH prevents the stale embedding inheritance typical of hash-based methods, ensuring new features learn effectively from scratch. Despite its collision-mitigation overhead, the system maintains training QPS and inference latency comparable to existing methods. Rigorous online experiments demonstrate that MPZCH achieves zero collisions for user embeddings and significantly improves item embedding freshness and quality. The solution has been released within the open-source TorchRec library for the broader community.

Updated: 2026-10-06 16:54:07

标题: Multi-Probe Zero Collision Hash (MPZCH):减轻嵌入冲突并增强大规模推荐系统中模型的新鲜度

摘要: 嵌入表是大规模推荐系统的关键组件,它有助于将高基数分类特征有效地映射为密集向量表示。然而,随着唯一ID数量的增加,传统的基于哈希的索引方法会遭受碰撞,降低模型性能和个性化质量。我们提出了一种基于线性探测的新型索引机制Multi-Probe Zero Collision Hash (MPZCH),有效地减轻了嵌入碰撞。在合理的表大小下,它通常完全消除了这些碰撞,同时保持了生产规模的效率。MPZCH利用辅助张量和高性能的CUDA核心来实现可配置的探测和主动驱逐策略。通过淘汰过时的ID并重置重新分配的插槽,MPZCH防止了基于哈希方法典型的陈旧嵌入继承,确保新特征有效地从头开始学习。尽管存在碰撞减轻的开销,该系统仍然保持了与现有方法相当的训练QPS和推理延迟。严格的在线实验表明,MPZCH实现了用户嵌入的零碰撞,并显著提高了项目嵌入的新鲜度和质量。这一解决方案已在开源TorchRec库中发布,供更广泛的社区使用。

更新时间: 2026-10-06 16:54:07

领域: cs.LG

下载: http://arxiv.org/abs/2602.17050v4

EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans

Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.

Updated: 2026-10-06 16:53:07

标题: EnGRICH:利用人类批评增强生成式奖励建模

摘要: 生成式奖励模型(GRMs)对LLM优化至关重要。与标量奖励模型不同,GRMs在偏好判断之外还生成自然语言评论,提供更精细的评估信号。它们的有效性在很大程度上取决于评论的可靠性。然而,现有的GRM训练通常使用最终偏好正确性作为结果监督。由于偏好结果空间受到严格限制,不可靠的评论仍然可以产生正确的结果,因此会被加强。最近的研究利用人类评论进行过程监督,但这些评论很少,并且经常被简化为标量奖励,使其细粒度的评价信息被低估。我们认为从人类评论中学习的评价标准可以推广到更广泛的仅结果偏好数据。为此,我们提出了EnGRICH,一个将GRM与在少量人类评论中学习的训练时MetaCritic配对的GRM训练框架。MetaCritic构建响应特定的评分标准,并使用它们来评估生成评论的证据覆盖率和正确性。由此产生的信号为细粒度的学分分配提供了过程奖励,并为探索更好的评论提供了结构化指导。在GRM训练期间,MetaCritic进一步优化以将基于人类的评价标准推广到仅结果数据。在推断中,经过训练的GRM独立运行。在七个奖励模型基准测试中的实验表明,EnGRICH始终优于竞争基线,而进一步分析验证了其核心机制的有效性。

更新时间: 2026-10-06 16:53:07

领域: cs.AI

下载: http://arxiv.org/abs/2610.05370v2

Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment

Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed. Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195). Meta's recipe and Ai2's Tulu 3 start from the same Llama-3.1 weights, and only Meta's carries the gap. Reading a chat model outside its chat template reverses the sign of its gap with nothing at stake (-0.038 against +0.055 under the template on OLMo-3), a distortion present on two of three recipes. On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that. The gap is a measurable target for post-training recipes, not a fixed property of pretrained weights.

Updated: 2026-10-06 16:52:23

标题: 《原则压力下:后期培训决定LLMs是否根据自己的道德判断行事》

摘要: 语言模型越来越像代理人。一个说一个动作是错误的代理人,然后还是采取了这个动作,与一个不知道更好的代理人有着不同的失败,而且对所陈述的价值观的评估也看不出来。我们建立了一个预注册的面板,包括248个场景,涵盖了五种压力情况。每个场景都会两次提出给同一个模型,一次作为代理人选择要做什么,一次以第三人称问哪个选项是正确的,因此模型自己的判断成为参考。每个场景都有一个没有压力的对应场景,每个模型都会得到一个阳性对照,其中操作员会下令违反动作,以便将缺失的空白区别于盲目的工具。在OLMo-3-7B-Instruct上,模型在大约五分之一的受压力场景中采取了它认为错误的动作,比在去除压力的相同场景中更频繁。在四个指示模型中,这种差距取决于训练后的配方:OLMo-3和Meta的Llama-3.1-8B-Instruct中都存在这种差距;Tulu 3在整个面板上没有显示(概率大约为0.01以上),也没有在自己的最具压力的场景上显示;Qwen2.5-7B-Instruct在整个面板上也没有显示(概率大约为0.02以上),在自己身上也没有解决(0.083,-0.028到0.195)。Meta的配方和Ai2的Tulu 3从相同的Llama-3.1权重开始,只有Meta的配方具有这种差距。在聊天模型超出其聊天模板时,其没有压力时的差距的符号会反转(在OLMo-3上,在模板下为-0.038,无压力时为+0.055),这种扭曲存在于三种配方中的两种中。在承载这一差距的两个模型上,在行动之前对利害关系进行推理会将选择重新转向模型自己的判断,与相同长度的非道德任务相比,有或没有压力;在OLMo-3上,提及所涉及的规范大约占其中的三分之一。这种差距是一个可衡量的后训练配方目标,而不是预训练权重的固有属性。

更新时间: 2026-10-06 16:52:23

领域: cs.LG,cs.AI,cs.CL

下载: http://arxiv.org/abs/2610.08670v1

MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge

On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA) variants, enable efficient adaptation at the edge, the limiting resource for Convolutional Neural Network (CNN) adaptation is often not the number of trainable parameters but the activation state that must be retained until the backward pass. This paper introduces Memory-Floor LoRA (MemFLoRA), a low-rank CNN adapter built around a memory-first design principle rather than a direct application of transformer-oriented LoRA. Instead of merely reducing trainable weights, we define an activation-memory-floor criterion: trainable backward computations must not depend on full-width layer inputs. The resulting adapter freezes the down-projection, trains a scale-matched up-projection, and combines eval-mode backbone normalization with activation-minimal backward rules, reducing saved state to the low-rank branch. Evaluated on three Human Activity Recognition (HAR) datasets and two CNN backbones under subject, body-location, and sensor-placement shifts, MemFLoRA reduces saved-activation memory by 98.5-98.7% and peak training-state memory by 94.9-97.3% relative to full fine-tuning, while matching or exceeding CNN PEFT baselines.

Updated: 2026-10-06 16:51:20

标题: MemFLoRA:边缘端CNN适应的内存地板LoRA

摘要: 设备上的学习在模型部署后遇到用户、传感器或环境特定变化时是必要的。尽管参数高效微调(PEFT)方法,特别是低秩适应(LoRA)变体,能够在边缘实现高效的适应,但卷积神经网络(CNN)适应的限制资源通常不是可训练参数的数量,而是必须保留直到向后传递的激活状态。本文介绍了Memory-Floor LoRA(MemFLoRA),这是一个围绕记忆优先设计原则构建的低秩CNN适配器,而不是直接应用面向变换器的LoRA。我们定义了一个激活记忆地板标准:可训练的向后计算不能依赖于完整宽度的层输入。结果适配器冻结了向下投影,训练了一个匹配比例的向上投影,并将评估模式的主干归一化与激活最小化的向后规则相结合,将保存的状态减少到低秩分支。在三个人体活动识别(HAR)数据集和两个CNN主干下,针对主体、身体位置和传感器位置变化评估,MemFLoRA相对于完全微调将激活内存减少了98.5-98.7%,峰值训练状态内存减少了94.9-97.3%,同时匹配或超过CNN PEFT基线。

更新时间: 2026-10-06 16:51:20

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.08669v1

Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents

Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched; in AgentMark's own robustness test, paraphrasing the observation alone drops bit-recovery to 16.8%. Second, every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else's, a question answered affirmatively for text watermarks (Jovanović et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over semantic action clusters under history conditioning, with the public-cluster bin replaced by keyed collision-resistant binning whose fresh-bucket assignment is provably unpredictable in the random-oracle model. Across five agent models (3B-14B, four vendors) and three encoders the ordering holds on both benchmarks: on ToolBench (600 trajectories per model) detection under rewriting is 0.49-0.66 for cluster-level versus 0.05-0.17 for exact-symbol at a permutation-calibrated 1% FPR, at 72-83% choice agreement against 22-27% for logit biasing; on ALFWorld (100 episodes per model) it is 0.92-0.97 versus 0.00-0.01. Keyed binning takes adaptive forgery from 100% to the false-positive floor at the primary operating point (bge, r=64). We also mark the boundary that guarantee does not cover: when the adversary copies the victim's own steps, shuffled splicing is neutralized (0.000 on Qwen2.5-3B) but chained replay remains at 0.76-0.98 across the five models, reported as open. Paraphrase robustness costs about half of the per-step watermark capacity. Code is available at https://anonymous.4open.science/r/SBW-Agent-Watermark.

Updated: 2026-10-06 16:50:42

标题: 语义行为水印:对LLM代理进行释义稳健和防伪溯源

摘要: 行为水印技术将所有者标识嵌入到LLM代理的高级动作选择中,实现了溯源而不影响输出标记。之前的代理水印技术存在两个问题。首先,所有三种之前的方案都将水印绑定到确切的动作符号上,因此重新命名工具会导致解码不同步,即使观察未受影响;在AgentMark自身的鲁棒性测试中,仅对观察进行改写就会将位恢复降至16.8%。其次,之前的每个代理水印技术都只研究了水印的去除,而没有询问对手是否可以伪造验证为他人的轨迹,这个问题在文本水印技术中已经肯定答案(Jovanović等人,2024年)。我们提出了语义行为水印(SBW):在历史条件下对语义动作簇进行水印处理,将公共簇桶替换为具有键控抗碰撞的桶,其新桶分配在随机预言者模型中被证明是不可预测的。在五个代理模型(3B-14B,四个供应商)和三个编码器上,排序在两个基准测试中均保持不变:在ToolBench(每个模型600条轨迹)重写检测在簇级别为0.49-0.66,精确符号为0.05-0.17,在1%的FPR下,对于72-83%的选择一致性,对比于22-27%的对数偏差;在ALFWorld(每个模型100集)中为0.92-0.97与0.00-0.01。键控桶技术将自适应伪造从100%降至主要操作点(bge,r=64)的误报率。我们还标记了保证不包括的边界:当对手复制受害者的步骤时,混合切片会被中和(在Qwen2.5-3B上为0.000),但串联重播在五个模型中仍保持在0.76-0.98之间,被报告为公开。改写鲁棒性成本大约为每步水印容量的一半。代码可在https://anonymous.4open.science/r/SBW-Agent-Watermark获取。

更新时间: 2026-10-06 16:50:42

领域: cs.CR,cs.AI

下载: http://arxiv.org/abs/2610.08668v1

ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding

As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.

Updated: 2026-10-06 16:46:14

标题: ParanoiaEval: 在主体编码中进行不必要的防御性工作的基准测试

摘要: 随着编码代理越来越多地独立承担实际工作,判断它们的风险处理是否合理变得至关重要。现有工作从不同的角度评估相关代理行为,但缺乏一个统一这些行为的系统框架。为了弥合这一差距,我们引入了ParanoiaEval,这是第一个用于编码代理风险处理能力统一评估的基准。ParanoiaEval基于软件工程风险管理中已经建立的避免-转移-缓解-接受框架,将其四种基本处理方式在编码代理设置中实现,并包含了200个受证据控制的存储库级任务对,每个任务对仅在定义处理方式的证据上有所不同。我们进一步引入了专门用于风险处理违规和证据响应性的度量标准,使用经过人类校准的代理评判员进行可靠评估。对8个代表性模型进行的大规模实验和事后人类研究显示:(I)尽管有明确证据,不必要的风险处理在11.2%-58.7%的运行中发生,且在代理配置之间存在显著变化;(II)更强的任务能力并不能保证更合适的风险处理,而处理违规严重损害开发人员的体验,将风险处理确立为一个独立的能力维度;(III)代理表现出与已建立的风险管理发现一致的系统模式,表明人类实践中的知识可以指导该能力的诊断和改进。

更新时间: 2026-10-06 16:46:14

领域: cs.AI,cs.SE

下载: http://arxiv.org/abs/2610.08662v1

Selective Transfer of RL Updates for Visual Reasoning

Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM. Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. Matched controls show that update magnitude or arbitrary low rank alone does not reproduce these gains. These results highlight a distinction between what is acquired during post-training and what remains transferable across models, providing a training-stage perspective on cross-model capability transfer. Code is available at https://anonymous.4open.science/r/selective-rl.

Updated: 2026-10-06 16:43:55

标题: 视觉推理中RL更新的选择性转移

摘要: 模型合并提供了一种无需训练的方式,将推理能力从语言模型转移到视觉-语言模型(VLMs),但基于端点的转移可能会混淆预先存在的模型差异和推理后训练期间获得的更改。相反,我们将能力转移的形式围绕训练阶段更新,隔离由强化学习(RL)引起的参数更改。然而,将这种更新完全转移仍然不够理想:我们发现其组件在跨模型转移性上存在显著差异,主导方向比完整更新更有效地转移。基于这一发现,我们引入了选择性RL,它隔离了RL阶段的更新,保留了其主导的矩阵方向并保持幅度,并将它们转移到VLM的语言模块。在三个模型家族和五个视觉推理基准测试中,选择性RL在15次比较中的12次中改善了完整更新的插值,包括对Qwen收件人的8.55个百分点的MathVision增益。匹配控制表明,仅更新幅度或任意低秩不能复制这些增益。这些结果突显了在推理后训练期间获得的内容与跨模型可转移内容之间的区别,为跨模型能力转移提供了训练阶段的视角。代码可在https://anonymous.4open.science/r/selective-rl上找到。

更新时间: 2026-10-06 16:43:55

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.08659v1

AX is the New AEO

In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, or AEO, now tells businesses to scatter breadcrumbs across forum threads, listicles, and off-site citations, so AI engines are likelier to surface and recommend them. But being surfaced is no longer enough: an agent opens the results and reads them before deciding, and one buyer question sends it through several rounds of search and fetch. What decides the outcome at this drill-down step is whether the agent can fetch and read the business's own site: agent experience (AX). We argue that AX is the new AEO. We run 37,927 agent journeys, each a buyer question about a business, across four independent harnesses over 1,056 real businesses, matched on fame, prior model knowledge, and two AEO proxies, then split based on their AX level. Only 7-10% of the finished answer comes from the model's training knowledge, whether or not the site is readable. Agent-ready businesses have answers built from their own pages 78% of the time against 56% and are clearly recommended 1.9x more often, while a grounded answer about a not-agent-ready business costs the agent 64% more on average. Holding business, harness, and question fixed, answers built from the site are 41% more accurate on average. The dominant failure is not fabrication but omission: web-built answers are 3.7x more likely to contain none of the facts the buyer asked for. Baselines differ sharply across the four harnesses, with clear-recommendation rates varying sevenfold from stack to stack, yet the recommendation gap holds in every one. In the agentic web era, being readable beats being talked about, and improving a site's AX is the strongest lever a business has.

Updated: 2026-10-06 16:43:16

标题: AX是新的AEO

摘要: 在2023年,人工智能模型在训练数据中回答问题,并在数据用尽时产生幻觉,企业被告知要种植这些知识。模型的训练知识已经被实时网络搜索所取代,随之而来的建议是:答案引擎优化(AEO)现在告诉企业在论坛帖子、列表文章和外部引用中撒下面包屑,以便人工智能引擎更有可能浮出水面并推荐它们。但仅仅浮出水面已经不够了:一个代理打开结果并在决定之前阅读它们,并且一个买家问题会让它经历几轮搜索和获取。在这个深入步骤,决定结果的是代理能否获取并阅读企业自己的网站:代理体验(AX)。我们认为AX是新的AEO。我们进行了37,927次代理旅程,每次是关于一个企业的买家问题,在四个独立的控制组中进行了1,056个真实企业的匹配,匹配标准是名声、先前模型知识和两个AEO代理,然后根据他们的AX水平进行分组。无论网站是否可读,完成的答案只有7-10%来自模型的训练知识。准备好的企业78%的答案来自他们自己的页面,而不具备代理准备的企业则为56%,并且明显推荐的概率增加了1.9倍。而对于一个不具备代理准备的企业的一个基础问题的答案平均成本会增加64%。在保持企业、控制组和问题不变的情况下,从网站构建的答案平均准确率提高了41%。主要的失败不是捏造而是遗漏:通过网络构建的答案中不包含买家所要求的事实的可能性是原来的3.7倍。基线在四个控制组之间有显著差异,明确推荐率从一个控制组到另一个控制组变化了7倍,但推荐差距在每一个控制组中都存在。在代理网络时代,可读性胜过被谈论,改善网站的AX是企业拥有的最强大的杠杆。

更新时间: 2026-10-06 16:43:16

领域: cs.AI,cs.IR

下载: http://arxiv.org/abs/2609.34951v3

Steering Diffusion Models to Rare Events with Sequential Monte Carlo

Diffusion models are increasingly used as surrogates for expensive simulators in weather prediction, molecular dynamics, and materials design. In these models, computing the probability $p_0[E]$ of an event $E$ is difficult, especially when the event of interest is rare. A stable estimate using Monte Carlo becomes computationally intractable, requiring a growing sample size $\propto\!1/p_0[E]$ to compensate for an increasing rarity. In this paper, we present Diffusion Importance Sampling of Rare Events or DireSMC, a sequential Monte Carlo scheme that guides a population of weighted samples towards the rare event, giving access not only to samples but also to a calibrated estimate of its probability. We set up our guidance using an analytical relaxation of the event set, allowing the method to easily extend to a wide range of user-defined rare events. We validate our method on a toy problem with analytical solutions and on a score-based climate emulator, where we obtain accurate rare-event probabilities on a range of rarities from $10^{-3}$ to $10^{-5}$, achieving net speed-ups of $9\times$ to $1413\times$ over Monte Carlo.

Updated: 2026-10-06 16:38:03

标题: 使用顺序蒙特卡洛将扩散模型引导到罕见事件

摘要: 扩散模型越来越多地被用作天气预测、分子动力学和材料设计中昂贵模拟器的替代品。在这些模型中,计算事件$E$的概率$p_0[E]$是困难的,特别是当感兴趣的事件很罕见时。使用蒙特卡罗方法得到稳定估计变得计算复杂,需要一个增长的样本大小$\propto\!1/p_0[E]$来补偿逐渐增加的罕见性。在本文中,我们提出了一种称为DireSMC的稀有事件扩散重要性采样的序贯蒙特卡罗方案,该方案引导一组加权样本朝向罕见事件,不仅可以获得样本,还可以获得其概率的校准估计。我们通过对事件集的分析松弛来设置我们的引导,使该方法可以轻松扩展到各种用户定义的罕见事件。我们在一个具有解析解的玩具问题和一个基于评分的气候模拟器上验证了我们的方法,在稀有性范围从$10^{-3}$到$10^{-5}$的范围内获得准确的罕见事件概率,实现了相对于蒙特卡罗的$9\times$到$1413\times$的净加速。

更新时间: 2026-10-06 16:38:03

领域: stat.ML,cs.LG

下载: http://arxiv.org/abs/2610.08652v1

A Case Study in Assuring AI-Written Software

Software-engineering agents can enable people without formal software training to build systems they could not otherwise implement and simultaneously can produce more code than even experts can meaningfully inspect. In both cases, exhaustive code review is not reliable as the sole basis for human control. We report a case study of a production healthcare platform built through coding agents and governed by an operator without formal software-engineering training. Over time, its workflow grew into a human-led meta-agent system where one agent wrote code, other agents supervised and reviewed it, and project rules carried lessons forward. The operator found that tests, monitors and reviewing agents used to supervise the system were fallible. Some monitors measured proxies rather than outcomes, some audits failed silently, missing checks disappeared from reported results and one automated repair caused operational disruption. In this case, human control depended on keeping the intended outcome, the evidence used to judge it, the agents' permissions and the final human decision were all tied to the same underlying objective.

Updated: 2026-10-06 16:37:08

标题: 确保人工智能编写的软件的案例研究

摘要: 软件工程代理可以使没有正式软件培训的人们构建他们在其他情况下无法实现的系统,同时可以生成比专家甚至无法有意义地检查的代码更多。在这两种情况下,详尽的代码审查并不可靠作为人类控制的唯一基础。我们报告了一个通过编码代理构建并由没有正式软件工程培训的操作员管理的生产医疗平台的案例研究。随着时间的推移,其工作流程发展成为人类主导的元代理系统,其中一个代理编写代码,其他代理监督和审查,项目规则延续经验。操作员发现,用于监督系统的测试、监控和审查代理是可犯错误的。一些监控器测量的是代理而不是结果,一些审计默默失败,缺失的检查从报告的结果中消失,而一个自动修复导致了操作中断。在这种情况下,人类控制取决于保持预期的结果、用于判断的证据、代理的权限和最终人类决策都与同一基础目标相关联。

更新时间: 2026-10-06 16:37:08

领域: cs.SE,cs.AI,cs.MA

下载: http://arxiv.org/abs/2610.08651v1

SquidAgent: Parallelize Wisely, Coordinate Efficiently

LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a re-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, such as prior decisions, that would otherwise be inherited implicitly in a serial execution. Second, there is an alignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. We thus derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. To address this, we instead measure cost in predicted output tokens, which we empirically find LLMs can estimate substantially more reliably than wall-clock time. Building on this token-based criterion, we propose SquidAgent. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator's session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, SquidAgent achieves a 2.2$\times$ mean throughput improvement and a 2.6$\times$ mean wall-time speedup over Claude Code, and a 2.0$\times$ throughput improvement over the strongest multi-agent baseline.

Updated: 2026-10-06 16:35:28

标题: SquidAgent:明智并行化,高效协调

摘要: 基于LLM的代理程序可以解决复杂的多步任务,但顺序执行会产生相当大的延迟。原则上,跨多个代理程序并行工作应该产生接近线性的加速。然而,现有的并行多代理系统通常比单一代理程序基准运行得更慢。我们将这种差距归因于并行执行所产生的两个隐藏成本,但串行代理程序可以避免。首先,存在重新探索成本:并行工作者花费冗余的工作重建编排者已经拥有的上下文,如之前的决策,在串行执行中本应隐含继承。其次,存在对齐成本:需要协调独立生成的输出之间的不一致性所需的开销。因此,我们得出一个基于原则的决策标准:只有在临界路径成本加上重新探索和对齐开销低于相应的串行成本时,才应该并行化一个层。虽然这个标准在壁钟时间中自然表达,但我们观察到当要求LLM估计任务持续时间时,它们的校准性很差。为了解决这个问题,我们改为以预测输出令牌来衡量成本,经验表明LLM可以比壁钟时间更可靠地估计令牌数量。基于这个基于令牌的标准,我们提出了SquidAgent。它在单个规划步骤中估计所有令牌预算,直接从编排者的会话中派生每个工作者,以消除重新探索成本,并用一个预先生成的共享约定块替换事后协调,将对齐转化为有限的前期成本。然后,确定性调度程序逐层应用标准。在实证方面,SquidAgent实现了2.2倍的平均吞吐量改善和2.6倍的平均壁钟时间加速,超过了Claude Code,以及比最强的多代理基线改进了2.0倍的吞吐量。

更新时间: 2026-10-06 16:35:28

领域: cs.AI,cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.08647v1

HygieneRoboBench: Benchmarking Hygiene-Aware Planning for Household Robots

Contact with contaminated objects can spread hazards through a household robot's grippers, tools, and shared surfaces, while new contacts can make an existing plan unsafe. Existing benchmarks do not jointly assess how planners identify hygiene risks from contact history and plan safe continuations after new contact events. Planners must do so within time and resource limits while respecting user priorities. We introduce HygieneRoboBench, with 624 instances across 134 task families, to evaluate safe resolution of household tasks from a given execution history. Tasks capture contamination through two grippers and shared objects, treatment costs, and user priorities. We combine controlled history, profile, and event comparisons with independent plan evaluation. These assess safe resolution, cost efficiency under user priorities, and responses to contact events. Evaluation of LLM-based and symbolic planners shows that safely completing a task does not guarantee the lowest execution costs under the user's priorities. To address this problem, we introduce Hygiene-NSP. It combines LLM-based grounding, contact-history reconstruction, and CP-SAT to jointly plan hygiene treatment and task execution under user priorities. Hygiene-NSP achieves safe resolution and optimal safe resolution rates of 94.4% and 90.4%, respectively. Both rates are higher than those of the evaluated baseline planners on the full dataset. Project page: https://euron-zc.github.io/HygieneRoboBench/.

Updated: 2026-10-06 16:33:44

标题: HygieneRoboBench: 用于家庭机器人的卫生感知规划基准测试

摘要: 接触受污染的物体可能通过家用机器人的夹持器、工具和共享表面传播危害,而新的接触可能使现有计划不安全。现有的基准并未共同评估规划者如何从接触历史中识别卫生风险,并在新接触事件后规划安全的延续。规划者必须在时间和资源限制内完成此任务,同时尊重用户的优先级。我们引入了HygieneRoboBench,包括134个任务家族中的624个实例,以评估根据给定执行历史安全解决家庭任务。任务通过两个夹持器和共享对象捕捉污染、治疗成本和用户优先级。我们将受控历史、概要和事件比较与独立计划评估相结合。这些评估安全解决、在用户优先级下的成本效率以及对接触事件的响应。对基于LLM和符号的规划者的评估显示,安全完成任务并不保证在用户优先级下具有最低执行成本。为解决这一问题,我们引入了Hygiene-NSP。它结合了基于LLM的接地、接触历史重建和CP-SAT,共同规划卫生治疗和任务执行以满足用户的优先级。Hygiene-NSP实现了94.4%和90.4%的安全解决和最佳安全解决率,分别高于整个数据集上评估的基准规划者。项目页面:https://euron-zc.github.io/HygieneRoboBench/。

更新时间: 2026-10-06 16:33:44

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2610.08642v1

Behavioral Guarantees for Proxy-Based Unlearning

This paper proposes a framework generalizing recent proxy-based unlearning methods and proves theoretical guarantees about the behavior of the resulting unlearned model: upper bounds on its Kullback-Leibler divergence to the ideal posterior distribution of the retain data. We model approximate unlearning as a constrained optimization problem and interpret a family of solutions as introducing a scaled unlearning signal in the output space. The unlearning signal arises from proxies of the posterior data distributions. Its scale is adapted to the proxies to ensure the behavioral upper bounds. This framework relies on the structure of the data distributions in order to create proxies. If need be, the target serves as a teacher to distill the update in the weights. Our approach is experimentally validated over two forgetting scenarios as reaching the closest classifier to the model retrained from scratch.

Updated: 2026-10-06 16:29:49

标题: 基于代理的遗忘过程的行为保证

摘要: 本文提出了一个框架,概括了最近基于代理的遗忘方法,并证明了关于结果遗忘模型行为的理论保证:其Kullback-Leibler散度上限到保留数据的理想后验分布。我们将近似遗忘建模为一个受约束的优化问题,并将一组解释为在输出空间引入了一个缩放的遗忘信号。遗忘信号来自后验数据分布的代理。其规模根据代理进行调整,以确保行为上限。这个框架依赖于数据分布的结构,以创建代理。如果需要,目标可以作为教师来提炼权重的更新。我们的方法在两种遗忘场景上经过实验证实,达到了与从头开始重新训练的模型最接近的分类器。

更新时间: 2026-10-06 16:29:49

领域: cs.LG

下载: http://arxiv.org/abs/2605.10680v2

SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents

Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.

Updated: 2026-10-06 16:29:47

标题: SCAD:长期代理商的结构化信用分配和蒸馏

摘要: 训练长期视野的代理程序解决复杂任务需要对延长交互序列进行有效监督。然而,稀疏的终端奖励使中间贡献变得模糊,而在线策略蒸馏可能会因为学生生成的历史增长而失去信息丰富的教师指导。为解决这一问题,我们引入了SCAD,将交互组织为规划和有界子任务执行,将执行蒸馏在局部上下文中,通过交叉轨迹子任务前缀树来细化规划学分,其中规划获得完整的终端学分,执行获得积极的终端学分和教师指导。在所有评估的基准测试中,SCAD相比最强的训练基线,提高了文本任务的宏平均准确率4.48个百分点,多模态任务提高了4.19个百分点。SCAD有效地将基于结果的学分分配与教师引导的蒸馏相结合,以改善长期视野代理程序中的规划和执行。

更新时间: 2026-10-06 16:29:47

领域: cs.LG

下载: http://arxiv.org/abs/2610.03372v3

PyDPF: A Python Package for Differentiable Particle Filtering

State-space models (SSMs) are a widely used tool in time series analysis. In the complex systems that arise from real-world data, it is common to employ particle filtering (PF), an efficient Monte Carlo method for estimating the hidden state corresponding to a sequence of observations. Applying particle filtering requires specifying both the parametric form and the parameters of the system, which are often unknown and must be estimated. Gradient-based optimisation techniques cannot be applied directly to standard particle filters, as the filters themselves are not differentiable. However, several recently proposed methods modify the resampling step to make particle filtering differentiable. In this paper, we present an implementation of several such differentiable particle filters (DPFs) with a unified API built on the popular PyTorch framework. Our implementation makes these algorithms easily accessible to a broader research community and facilitates straightforward comparison between them. We validate our framework by reproducing experiments from several existing studies and demonstrate how DPFs can be applied to address several common challenges with state space modelling.

Updated: 2026-10-06 16:27:46

标题: PyDPF:一种用于可微分粒子滤波的Python软件包

摘要: 状态空间模型(SSMs)是时间序列分析中广泛使用的工具。在由真实数据产生的复杂系统中,通常使用粒子滤波(PF),一种用于估计与一系列观测对应的隐藏状态的高效蒙特卡洛方法。应用粒子滤波需要指定系统的参数形式和参数,这些参数通常是未知的,必须进行估计。基于梯度的优化技术不能直接应用于标准粒子滤波器,因为滤波器本身不可微分。然而,最近提出的几种方法修改重新采样步骤,使粒子滤波器可微分。在本文中,我们介绍了几种这样的可微分粒子滤波器(DPFs)的实现,统一构建在流行的PyTorch框架上。我们的实现使这些算法更容易地被更广泛的研究社区访问,并促进它们之间的直接比较。我们通过复制几个现有研究中的实验来验证我们的框架,并展示了如何应用DPFs来解决状态空间建模中的几个常见挑战。

更新时间: 2026-10-06 16:27:46

领域: eess.SP,cs.LG

下载: http://arxiv.org/abs/2510.25693v4

Parallel Predictive World Models for Accurate and Efficient Long-Horizon Planning

Long-horizon world-model planning typically relies on autoregressive rollouts, where predicted states are repeatedly fed back into the model. This preserves temporal structure but creates a horizon-length sequential path and exposes later predictions to recursive decoded-state feedback. We introduce Parallel Predictive World Models (PPWM), which predict a finite-horizon trajectory in parallel while retaining causal interaction among future representations. Each horizon is conditioned on its causal action prefix, and future representations interact before decoding, separating temporal causality from state-by-state output recursion. We formalize this distinction by viewing autoregressive rollout as a causal trajectory map and identifying the decoded-state feedback pathway removed by PPWM. Across four visual-control tasks, PPWM achieves the lowest long-horizon prediction error and the highest Cross-Entropy Method (CEM) simulator success among the evaluated predictive interfaces. Meanwhile, PPWM achieves more than a 3$\times$ average CEM planning speedup over the autoregressive LeWM baseline. These results suggest that accurate and efficient long-horizon world-model planning does not require state-by-state autoregression, but can instead be achieved through parallel causal trajectory prediction.

Updated: 2026-10-06 16:26:34

标题: 并行预测世界模型用于准确和高效的长期规划

摘要: 长期视角的世界模型规划通常依赖于自回归展开,在这种情况下,预测的状态被反复馈送回模型。这保留了时间结构,但创建了一个长度为地平线的序列路径,并将后续预测暴露给递归解码状态反馈。我们引入了并行预测世界模型(PPWM),它们能够并行预测有限地平线轨迹,同时保持未来表示之间的因果交互。每个地平线都取决于其因果动作前缀,未来表示在解码之前进行交互,将时间因果性与状态逐个输出递归分开。我们通过将自回归展开视为因果轨迹映射并识别PPWM消除的解码状态反馈路径来形式化这种区别。在四个视觉控制任务中,PPWM实现了最低的长期地平线预测误差和评估预测界面中最高的交叉熵方法(CEM)模拟器成功率。与自回归LeWM基线相比,PPWM实现了超过3倍的平均CEM计划加速。这些结果表明,准确和高效的长期地平线世界模型规划并不需要状态逐个自回归,而是可以通过并行因果轨迹预测来实现。

更新时间: 2026-10-06 16:26:34

领域: cs.AI

下载: http://arxiv.org/abs/2610.08627v1

Feature Information Dynamics in Diffusion

Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce feature information dynamics, an information-theoretic framework for localizing when a feature is generated during diffusion. Using the I-MMSE identity, we connect the rate of feature mutual information change to a gap between optimal unconditional and feature-conditional denoising losses, yielding practical estimators for feature information density. We further develop a chained decomposition that separates shared from incremental information in a feature hierarchy. We use this framework first to quantitatively confirm spectral autoregression in pixel diffusion, and then to extend the analysis beyond frequency: under a class $\to$ mask $\to$ Canny conditioning chain, the per-feature information densities differ across pixel, SDVAE, VAVAE, and RAE, exposing fundamental differences between these representations and suggesting that ordered generation could be beneficial for training diffusion models. Our code is available at https://github.com/AI4Science-WestlakeU/feature-information-dynamics.

Updated: 2026-10-06 16:24:07

标题: 扩散中的特征信息动态

摘要: 扩散模型通过一系列去噪问题生成数据,并被广泛观察到在细节之前揭示粗略结构。然而,这种直觉大多是经验性和定性的。我们引入特征信息动态,这是一个信息论框架,用于定位在扩散过程中何时生成特征。利用I-MMSE恒等式,我们将特征互信息变化速率与最优无条件和特征条件去噪损失之间的差距联系起来,从而得到特征信息密度的实用估计器。我们进一步开发了一个链式分解,将特征层次结构中的共享信息与增量信息分离开来。我们首先使用这个框架定量确认了像素扩散中的谱自回归,然后将分析扩展到频率以外:在一个类$\to$掩码$\to$Canny条件链下,像素、SDVAE、VAVAE和RAE之间的每个特征信息密度不同,揭示了这些表示之间的基本差异,并暗示有序生成可能有利于训练扩散模型。我们的代码可在https://github.com/AI4Science-WestlakeU/feature-information-dynamics 上找到。

更新时间: 2026-10-06 16:24:07

领域: stat.ML,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08626v1

Understanding the Identity-Transformation Approach in OIDC-Compatible Privacy-Preserving SSO Services

Single sign-on (SSO) enables a user to log into multiple websites, called relying parties (RPs), by her username and credential set up in another trusted web system, called the identity provider (IdP). Identity transformations are proposed in UppreSSO to provide privacy-preserving SSO services, preventing both IdP-based login tracing and RP-based identity linkage. While the security and privacy guarantees of UppreSSO have been proved, several essential issues on the identity-transformation approach are not well studied. In this paper, we comprehensively investigate this approach as below. Firstly, several suggestions to efficiently integrate identity transformations into OpenID Connect (OIDC) are explained. Then, we uncover the relationship between identity transformations in SSO and oblivious pseudo-random functions (OPRFs), and present two variations of the properties required for SSO security as well as other requirements, to analyze existing OPRF protocols. Finally, new identity transformations different from those proposed in UppreSSO, are constructed based on some OPRFs. To the best of our knowledge, this is the first time to uncover the relationship between identity transformations in SSO services and OPRFs, and prove the SSO-related properties (i.e., output uniqueness, key-identifier freeness, and collision resistance on 1st/2nd-input) of typical OPRFs.

Updated: 2026-10-06 16:23:42

标题: 理解 OIDC 兼容的隐私保护 SSO 服务中的身份转换方法

摘要: 单点登录(SSO)使用户可以通过在另一个可信的Web系统中设置的用户名和凭证登录多个网站,称为依赖方(RP),称为身份提供者(IdP)的身份提供者。 UppreSSO中提出了身份转换,以提供保护隐私的SSO服务,防止基于IdP的登录跟踪和基于RP的身份关联。虽然已经证明了UppreSSO的安全性和隐私性保证,但对于身份转换方法的几个基本问题尚未得到充分研究。在本文中,我们全面调查了这种方法。首先,解释了将身份转换有效集成到OpenID Connect(OIDC)中的几个建议。然后,我们揭示了SSO中身份转换与遗忘伪随机函数(OPRFs)之间的关系,并提出了两种SSO安全性所需的属性的变体以及其他要求,以分析现有的OPRF协议。最后,基于一些OPRFs构建了与UppreSSO中提出的不同的新身份转换。据我们所知,这是首次揭示了SSO服务中身份转换与OPRFs之间的关系,并证明了典型OPRFs的SSO相关属性(即输出唯一性、密钥标识符自由和第1/2输入的冲突抵抗)。

更新时间: 2026-10-06 16:23:42

领域: cs.CR

下载: http://arxiv.org/abs/2506.01325v3

Early Memory Selection for Balanced Adam

We propose a method for choosing the shared memory parameter $β_1=β_2=β$ in Adam from a short pilot training. The selected $β$ remains fixed during the subsequent full training. A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past gradients. This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints. The estimator uses the numerator and denominator jointly, preserving their covariance. With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared $β=0.95$. The mean gap is also 32.3% lower than that of the best constant $β$ chosen across all eleven workloads.

Updated: 2026-10-06 16:22:36

标题: 早期内存选择用于平衡的Adam

摘要: 我们提出了一种方法,从短期试点训练中选择Adam中共享内存参数$β_1=β_2=β$。在随后的完整训练过程中,所选的$β$保持不变。Adam的归一化方向的局部模型平衡了抽样变异性和平均过去梯度引入的延迟。这种平衡提供了一个立方内存规则,其两个系数是从几个试点检查点的梯度探测中估计得到的。该估计器同时使用分子和分母,保持它们的协方差。通过在每个四个检查点进行200次更新试点和十六次梯度探测,与代表共享$β=0.95$的网格相比,在十一个视觉和语言工作负载上进行的种子匹配的回顾评估将平均相对验证间隙减少了40.7%,最差季度平均间隙减少了44.3%。平均间隙也比在所有十一个工作负载中选择的最佳恒定$β$的间隙低32.3%。

更新时间: 2026-10-06 16:22:36

领域: cs.LG,cs.AI,stat.ML

下载: http://arxiv.org/abs/2610.08624v1

Agentic RCA for Internet-Scale Services Using Constrained Creativity

System administrators of Internet-scale services need to resolve failure incidents to maintain reliability of such services. Ideally, we want a troubleshooting system to be: (1) expressive to known and unknown incidents with high accuracy; (2) cost efficient at scale; (3) explainable to provide actionable insights operators can act on; and (4) entail low effort from the operators. Unfortunately, most existing systems, including emerging LLM-assisted agentic workflows and structured frameworks for authoring diverse RCA algorithms fall short of achieving all four requirements. We present E4, a novel agentic system for troubleshooting for Internet-scale services. E4 embodies the paradigm of constrained creativity that combines the best of LLM-assisted automation and exploration with the explainability and efficiency of a structured approach. Instead of allowing an LLM agent to write arbitrary code or generate arbitrary responses, we provide the agent a restricted DSL to generate its response via simple loop-free data flow programs. This DSL, equipped with high level operators for troubleshooting, makes E4's output accurate, verifiable and explainable. On a mix of synthetic and real-world workloads, E4 achieves up to 62% better accuracy compared to state-of-the-art solutions, while providing more explainable responses at up to 12x reduced cost.

Updated: 2026-10-06 16:21:57

标题: 利用受限创造力的主动式RCA技术实现互联网规模服务

摘要: 互联网规模服务的系统管理员需要解决故障事件,以保持这些服务的可靠性。理想情况下,我们希望故障排除系统具有以下特点:(1)能够准确表达已知和未知事件;(2)在规模上具有成本效益;(3)能够提供可操作的见解,使运营人员能够采取行动;(4)需要的操作人员工作量较低。不幸的是,大多数现有系统,包括新兴的LLM辅助代理工作流和用于编写各种根本原因分析算法的结构化框架,都无法同时满足这四个要求。我们提出了E4,一种面向互联网规模服务故障排除的新型代理系统。E4体现了受限创造力的范式,结合了LLM辅助自动化和探索的优点,以及结构化方法的解释性和效率。我们不允许LLM代理编写任意代码或生成任意响应,而是为代理提供了一个受限的DSL,通过简单的无环数据流程序生成其响应。这个DSL配备了用于故障排除的高级运算符,使得E4的输出准确、可验证且可解释。在合成和真实工作负载的混合情况下,E4相比最先进的解决方案,准确性提高了高达62%,同时提供了更多可解释的响应,成本降低了高达12倍。

更新时间: 2026-10-06 16:21:57

领域: cs.NI,cs.AI

下载: http://arxiv.org/abs/2610.08622v1

Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness

Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer's feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves state-of-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.

Updated: 2026-10-06 16:21:44

标题: 递归游戏创作者:一个主体产品级体验导向游戏工具

摘要: 最近的游戏设计代理在生成可玩游戏方面取得了重大进展。然而,程序的正确性并不能确保玩家获得愉快的体验。我们提出了递归游戏创造者,这是一个以体验为导向的工具,旨在将代理游戏开发从粗糙的游戏原型推进到有趣的游戏。递归游戏创造者围绕四个组件进行递归开发:设计师、构建者、玩家和评论者。设计师将用户指令和评论者的反馈转化为详细计划。构建者将这些计划转化为候选游戏。编码原生的玩家通过编程接口创建和执行可重用策略,以有效收集各种游戏轨迹,缓解由于缓慢的GUI集合导致的评估偏差。评论者使用经过精心设计的基于轨迹的指标来诱导玩家偏好,结合视觉证据和明确的文本偏好来评估游戏是否符合特定标准。最后,评论者接受更好的版本并为下一轮提供改进意见,闭合递归循环。我们的方法在GameCraft-Bench上实现了77.89的最新综合性能。在GameASG-Bench上,它实现了53.2%的严格任务成功率,比同一模型基线提高了34.1%,并且在比较方法中具有93.4%的最高平均运行时检查通过率。用户研究显示较长的游戏时间和更高的评分。代码即将推出。

更新时间: 2026-10-06 16:21:44

领域: cs.AI,cs.MA,cs.SE

下载: http://arxiv.org/abs/2610.08621v1

Process-Aware AI for Rainfall-Runoff Modeling: A Mass-Conserving Neural Framework with Hydrological Process Constraints

Machine learning models can achieve high predictive accuracy in hydrological applications but often lack physical interpretability. The Mass-Conserving Perceptron (MCP) provides a physics-aware artificial intelligence (AI) framework that enforces conservation principles while allowing hydrological process relationships to be learned from data. In this study, we investigate how progressively embedding physically meaningful representations of hydrological processes within a single MCP storage unit improves predictive skill and interpretability in rainfall-runoff modeling. Starting from a minimal MCP formulation, we sequentially introduce bounded soil storage, state-dependent conductivity, variable porosity, infiltration capacity, surface ponding, vertical drainage, and nonlinear water-table dynamics. The resulting hierarchy of process-aware MCP models is evaluated across 15 catchments spanning five hydroclimatic regions of the continental United States using daily streamflow prediction as the target. Results show that progressively augmenting the internal physical structure of the MCP unit generally improves predictive performance. The influence of these process representations is strongly hydroclimate dependent: vertical drainage substantially improves model skill in arid and snow-dominated basins but reduces performance in rainfall-dominated regions, while surface ponding has comparatively small effects. The best-performing MCP configurations approach the predictive skill of a Long Short-Term Memory benchmark while maintaining explicit physical interpretability. These results demonstrate that embedding hydrological process constraints within AI architectures provides a promising pathway toward interpretable and process-aware rainfall-runoff modeling.

Updated: 2026-10-06 16:21:32

标题: 过程感知人工智能用于降雨径流建模:具有水文过程约束的保质量神经网络框架

摘要: 机器学习模型在水文应用中可以达到很高的预测准确性,但通常缺乏物理可解释性。质量守恒感知器(MCP)提供了一个具有物理意识的人工智能框架,它在允许从数据中学习水文过程关系的同时强制执行保守原则。在这项研究中,我们调查了如何逐步将具有物理意义的水文过程表示嵌入单个MCP存储单元中,以提高降雨径流建模的预测技能和可解释性。从一个最简单的MCP公式开始,我们逐步引入有界土壤贮存、状态依赖导电性、可变孔隙度、渗透能力、表面积水、垂直排水和非线性水位动力学。通过使用每日流量预测作为目标,在涵盖美国大陆五个水文气候区域的15个集水区中评估了这种具有过程意识的MCP模型层次结构。结果表明,逐步增强MCP单元的内部物理结构通常会改善预测性能。这些过程表示的影响与水文气候密切相关:垂直排水在干旱和以雪为主的流域中显着改善模型技能,但在以降雨为主的地区降低性能,而表面积水的影响相对较小。表现最佳的MCP配置接近长短期记忆基准的预测技能,同时保持明确的物理可解释性。这些结果表明,在AI架构中嵌入水文过程约束提供了一条通向可解释和具有过程意识的降雨径流建模的有希望的途径。

更新时间: 2026-10-06 16:21:32

领域: cs.LG

下载: http://arxiv.org/abs/2603.25093v3

RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step state access dependent only on the number of selected slots rather than the total state size. Empirically, RAM-Net outperforms strong recurrent baselines on fine-grained long-range retrieval and achieves the lowest perplexity with competitive commonsense reasoning. It does so while accessing fewer state elements per step than all baselines, e.g., $8\times$ fewer than Mamba2.

Updated: 2026-10-06 16:19:58

标题: RAM-Net:稀疏可寻址状态的线性时间序列建模

摘要: 线性注意力提供了一个高效的替代方案,可以使用固定大小的循环状态,而不是全注意力。然而,这个状态被所有标记共享,因此来自不同标记的信息在其中叠加,产生了干扰,降低了长距离细粒度的召回率。为了解决这个问题,我们提出了RAM-Net,它用稀疏基于地址的访问替代了对共享状态的密集访问。RAM-Net将循环状态组织为一个独立插槽的固定大小数组,并使用地址解码器将每个关键字或查询映射到一个稀疏地址,在每个步骤中选择一小部分插槽进行写入或读取。这种设计将具有不重叠地址的标记定向到不相交的插槽,抑制了标记之间的干扰,同时使每个步骤的状态访问仅依赖于所选插槽的数量,而不是总状态大小。从经验上看,RAM-Net在细粒度的长距离检索方面优于强大的循环基线,并在竞争性常识推理中实现了最低的困惑度。它在每个步骤访问的状态元素比所有基线都少,比如比Mamba2少8倍。

更新时间: 2026-10-06 16:19:58

领域: cs.LG,cs.CL

下载: http://arxiv.org/abs/2602.11958v2

FFR: Forward-Forward Learning for Regression

The Forward-Forward (FF) algorithm offers a computationally efficient and biologically plausible alternative to backpropagation (BP) by training neural networks through purely local, layer-wise optimization. However, FF is inherently designed for classification via contrastive positive-negative sample pairs, and extending it to regression poses fundamental challenges: continuous target space lacks natural "opposites" for contrastive learning, and the standard goodness function carries no information about target magnitude or ordering. We propose FFR (Forward-Forward for Regression), to our knowledge, the first framework to extend FF to real-world regression and demonstrate competitive performance across diverse realworld datasets. FFR introduces three key innovations: (1) an ordinal competitive goodness function that replaces contrastive pairs with competitive learning between partitioned neuron groups under distance-aware ordinal supervision; (2) a stratified ladder architecture where shallow layers learn coarse ordinal discrimination and deeper layers refine into fine-grained regression, with multi-scale feature aggregation for inter-layer collaboration; and (3) hierarchical prediction with uncertainty estimation, where multi-scale predictors jointly provide robust predictions and a single-pass uncertainty score. Extensive experimental results show FFR recovers on average 98.5% of BP's accuracy across six real-world regression benchmarks while reducing peak training memory to only 27% of BP's at depth 8 and 8% at depth 32, with per-iteration time around 72% of BP's, and substantially outperforms all BP-free competitors.

Updated: 2026-10-06 16:18:54

标题: FFR:回归的前向学习

摘要: 前向-前向(FF)算法通过纯粹的本地、逐层优化训练神经网络,提供了一个计算效率高且生物合理的替代方案,相比于反向传播(BP)。然而,FF本质上是为了通过对比正负样本对进行分类而设计的,将其扩展到回归面临着基本挑战:连续目标空间缺乏对比学习的自然“对立物”,而标准的良好函数没有关于目标大小或排序的信息。我们提出了FFR(用于回归的前向-前向),据我们所知,这是第一个将FF扩展到真实世界回归的框架,并在各种真实世界数据集上展示了竞争性能。FFR引入了三个关键创新:(1)一种序数竞争性良好函数,用距离感知序数监督替代对比对,在分区神经元组之间进行竞争学习;(2)分层阶梯结构,浅层学习粗略的序数判别,深层逐渐细化为细粒度回归,实现多尺度特征聚合进行层间协作;(3)具有不确定性估计的层次预测,多尺度预测器共同提供稳健预测和单次不确定性评分。广泛的实验结果表明,FFR在六个真实世界回归基准测试中平均恢复了BP准确度的98.5%,同时将训练峰值内存减少到BP的27%(深度为8时)和8%(深度为32时),每次迭代的时间约为BP的72%,并且远远优于所有不使用BP的竞争对手。

更新时间: 2026-10-06 16:18:54

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2606.03927v2

A Swarm-Coordinated Multi-Robot System for Early Stress Detection in Agricultural Rows Using Multimodal Leaf Sensing

Early stress detection in crops is a necessity today to improve efficiency and reduce waste of time, money, and effort. However, most modern techniques, such as hyperspectral imaging and AI-based systems, are too costly and complex for medium and small-scale farmers to implement. This paper showcases CropSentry, a low-cost, ground-based multi-robot system that uses multimodal leaf sensing to continuously monitor crop health by tracking stress levels. The system comprises two autonomous bots that continuously detect leaf color and environmental data row by row. The observations are spatially mapped and sent over to the master bot, which uses color-coded row segments to generate a real-time web-based dashboard displaying crop health. After 63 observations were collected during the experiments, the results showed an overall crop health classification accuracy of 84.12%, with 82.60% for healthy plants, 88% for nutrient-deficient plants, and 80% for diseased plants. Also, 100% wireless communication success rate across 10 slave observations was achieved. Close-range leaf inspection across multiple bots can detect early stress in crops while remaining affordable, accessible, and scalable. It provides farmers with timely information to improve resource utilization and crop management.

Updated: 2026-10-06 16:07:15

标题: 一个群体协调的多机器人系统,在农田中使用多模式叶片感知进行早期压力检测

摘要: 作物的早期压力检测是当今必不可少的,以提高效率并减少时间、金钱和精力的浪费。然而,大多数现代技术,如高光谱成像和基于人工智能的系统,对中小规模农民来说成本过高且复杂,无法实施。本文展示了CropSentry,一个低成本的基于地面的多机器人系统,通过使用多模式叶片传感器来持续监测作物健康状态,跟踪压力水平。该系统由两个自主机器人组成,通过逐行连续检测叶片颜色和环境数据。这些观测结果被空间映射并发送到主机器人,后者使用颜色编码的行分段生成实时基于网络的仪表板,显示作物健康情况。在实验中收集了63次观测后,结果显示整体作物健康分类准确率为84.12%,其中健康植物为82.60%,营养不足植物为88%,病害植物为80%。此外,通过10次从机观测实现了100%的无线通信成功率。多机器人的近距离叶片检查可以检测作物的早期压力,同时保持价格实惠、易于访问和可扩展。它为农民提供及时信息,以改善资源利用和作物管理。

更新时间: 2026-10-06 16:07:15

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2610.08603v1

One for All, All for One: Coordinated Multi-Agent Diffusion Steering via Stochastic Optimal Control

Deep generative models often produce structured outputs composed of interacting components. Modelling these outputs with a single model requires learning both the component distributions and their interactions. We pursue a modular alternative: reuse independently trained component generators and learn only how to coordinate them to produce coherent structured outputs. Our framework, Coordinated Multi-Agent Diffusion Steering (CMDS), treats frozen pretrained diffusion models as reusable generative primitives and coordinates their reverse processes through a learned control. We formulate coordination as a stochastic optimal control problem, balancing an assembly-level reward that specifies the desired properties of the combined output against deviations from the pretrained dynamics. The learned control amortises this optimisation, allowing reuse across new task instances. Experiments show that CMDS can recover a known target distribution, satisfy different spatial constraints with the same trained control, and recover individual sources from degraded mixtures. Across multi-agent maze navigation, articulated robot planning, and text-conditioned human motion, CMDS turns frozen models into coordinated multi-agent generators.

Updated: 2026-10-06 16:02:49

标题: 一个为所有人,所有人为一个:通过随机最优控制协调多智能体扩散引导

摘要: 深度生成模型通常生成由相互作用组件组成的结构化输出。用单个模型对这些输出进行建模需要学习组件分布和它们之间的相互作用。我们提出了一种模块化的替代方案:重复使用独立训练的组件生成器,并仅学习如何协调它们以生成连贯的结构化输出。我们的框架,协调多智能体扩散驾驶(CMDS),将冻结的预训练扩散模型视为可重复使用的生成原语,并通过学习的控制协调它们的逆过程。我们将协调形式化为随机最优控制问题,平衡一个装配级奖励,指定所需组合输出的属性与偏离预训练动态之间的差异。学习的控制摊销了这种优化,允许在新任务实例中重复使用。实验证明,CMDS可以恢复已知的目标分布,使用相同的训练控制满足不同的空间约束,并从降解的混合物中恢复单个来源。在多智能体迷宫导航、关节机器人规划和文本调节的人体运动方面,CMDS将冻结的模型转化为协调的多智能体生成器。

更新时间: 2026-10-06 16:02:49

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2610.08595v1

Infrared Subtraction with Artificial Intelligence

We present AI-developed local infrared subtraction, building on projection to Born and EFT matching. The framework separates an integrable radiation term from a finite contribution at Born kinematics, referred to as the Born contact. The contact is determined using the EFT singular distribution in a resolution observable such as N-jettiness $τ_N$. Under human physics guidance, an LLM develops two implementations. One uses a neural network for phase space projection and fits the contact by matching to EFT cumulants. The other uses an analytic construction that keeps the Born momenta fixed while integrating over radiation. It combines the EFT $δ(τ_N)$ coefficient with finite 4-dimensional radiation integrals to calculate the contact term directly. This gives a local subtraction formula without a slicing parameter, while reusing existing lower-order radiation calculations and EFT singular predictions. As a demonstration, we reconstruct the full NLO correction for massless 3- and 4-jet production in electron-positron annihilation. The attempt to the NNLO dijet production is also made by recursively using the NLO P2B construction with the LLM designing machine-learning controls to reduce the variance of the contact integral. The tested predictions are in good agreement with EERAD3. The numerical calculation and projection-network training use a 2020 Apple M1 MacBook, without GPU acceleration, illustrating the feasibility of the construction with modest computing resources. The appendices develop an extension of the local subtraction to 3-jet NNLO, giving explicit radiation maps and a proposed contact formula. We also show how to integrate over NNLO radiation while keeping the Born momenta fixed, for any number of massless final-state jets. Our results demonstrate how AI can help higher-order calculations by constructing infrared subtraction and improving its numerical integration.

Updated: 2026-10-06 16:02:25

标题: 用人工智能进行红外线减法

摘要: 我们提出了基于投影到Born和EFT匹配的AI开发的局部红外减法。该框架将可积辐射项与Born动力学下的有限贡献分离开来,称为Born接触。通过在诸如N-jettiness $τ_N$等分辨率观测量中使用EFT奇异分布来确定接触。在人类物理学的指导下,LLM开发了两种实现。一种使用神经网络进行相空间投影,并通过匹配EFT累积量来拟合接触。另一种采用保持Born动量固定的解析构造,同时在辐射上积分。它将EFT $δ(τ_N)$系数与有限的四维辐射积分结合起来,直接计算接触项。这提供了一个无需切片参数的局部减法公式,同时重复使用现有的低阶辐射计算和EFT奇异预测。作为示范,我们重建了电子-正电子湮灭中无质量3和4喷注产生的完整NLO修正。还尝试通过递归使用LLM设计的机器学习控制减少接触积分方差来实现NNLO双喷注产生。经过测试的预测与EERAD3非常一致。数值计算和投影网络训练使用2020年的苹果M1 MacBook,无需GPU加速,展示了在适度的计算资源下进行构造的可行性。附录展开了将局部减法扩展到3喷注NNLO的工作,提供了明确的辐射图以及一个提议的接触公式。我们还展示了如何在保持Born动量固定的情况下对NNLO辐射进行积分,适用于任意数量的无质量末态喷注。我们的结果展示了AI如何通过构建红外减法并改进其数值积分来帮助高阶计算。

更新时间: 2026-10-06 16:02:25

领域: hep-ph,cs.AI,hep-ex,nucl-ex,nucl-th

下载: http://arxiv.org/abs/2609.36007v3

Multi-Label Perceptual Bug Detection in Video Games using Deep Learning on Gameplay Footage

Traditional approaches for automated bug detection in video games, such as manual testing, can be beneficial for the improvement of quality assurance, but they can be expensive and time-consuming. The scarce number of tools available to detect multiple perceptual bugs in the same video frame introduces detection challenges for automated bug detection tools in real-world scenarios. We propose a deep learning model for multi-label perceptual bug detection and compare it against video classification models such as Inflated 3D ConvNet and 3D ResNet. Our proposed model, ResNet-BiLSTM, achieved an F1 score of 85.78% on the benchmark dataset. Our results demonstrated that temporal dependency modelling is beneficial for accurate video-based bug detection. We believe this work with multi-label perceptual bug detection on gameplay videos will help save resources spent on manual testing workloads in video games. Furthermore, we introduce a new dataset with multi-label perceptual bugs in this work. The dataset contains 77,969 video clips across different genres of games with approximately 1.2 million frames, containing combinations from 5 classes of bugs in the same video frame.

Updated: 2026-10-06 16:01:00

标题: 使用深度学习在游戏玩法视频中进行多标签感知性错误检测

摘要: 传统的自动化游戏bug检测方法,如手动测试,对于改进质量保证可能是有益的,但它们可能既昂贵又耗时。在同一视频帧中检测多个感知bug的工具数量有限,这在现实场景下给自动bug检测工具带来了检测挑战。我们提出了一个用于多标签感知bug检测的深度学习模型,并将其与视频分类模型(如Inflated 3D ConvNet和3D ResNet)进行了比较。我们提出的模型ResNet-BiLSTM在基准数据集上实现了85.78%的F1分数。我们的结果表明,时间依赖建模对准确的基于视频的bug检测是有益的。我们认为,在游戏玩法视频上进行的多标签感知bug检测工作将有助于节省在视频游戏中手动测试工作负担方面的资源。此外,我们在这项工作中介绍了一个包含不同游戏类型的77,969个视频剪辑的新数据集,其中包含大约120万帧,其中包含同一视频帧中的5类bug的组合。

更新时间: 2026-10-06 16:01:00

领域: cs.LG

下载: http://arxiv.org/abs/2610.08593v1

CNet: A Complex-Valued Deep Learning Framework with Wirtinger Autodifferentiation and FFT--Hadamard Convolution

CNet is a C++/CUDA framework for building and training deep complex-valued neural networks (CVNNs) and, more generally, for optimizing complex-valued functions by gradient descent with Wirtinger (CR-calculus) derivatives. It takes a physics-native stance: a network is a cascade of complex -- and often unitary (the DFT) -- operations acting on an amplitude vector, and classification is a Born-rule measurement $p_k = |z_k|^2 / \|z\|^2$ rather than a softmax over real logits. Every layer ships a CPU reference and a CUDA kernel checked against finite differences, and the computation graph is cloned across the batch for GPU execution. On top of the base layers we add signal-processing primitives that turn the identity conv(x,k) = IFFT(FFT(x) . FFT(k)) into a learnable complex convolutional network, together with a true-Adam optimizer and a reduced-memory inference mode. We report three studies. First, a fully complex-valued, FNet-style causal sequence model built on a new $O(N \log N)$ causal Fourier mixer -- a triangular-masked DFT evaluated by a Bluestein / chirp-z factorization: once properly tuned it matches or exceeds a parameter-matched real-valued causal FNet on character-level language modeling, reaching the real model's converged quality in under half the training steps. Second and third, bottleneck analyses on radio-modulation classification (RML2016.10a) and the Fourier phase problem of coherent-diffraction imaging, which isolate exactly where complex-valued networks still need new operators. Across all three the complex formulation provably learns the physically correct structure. Code: https://github.com/crasmarum/CNet

Updated: 2026-10-06 15:57:33

标题: CNet:一种具有Wirtinger自动微分和FFT-Hadamard卷积的复值深度学习框架

摘要: CNet是一个用于构建和训练深度复值神经网络(CVNNs)的C++/CUDA框架,更一般地说,用于通过Wirtinger(CR-微积分)导数进行梯度下降优化复值函数。它采取了物理本源的立场:网络是作用于幅度向量的复数 - 通常是幺正(DFT) - 操作的级联,并且分类是一个Born规则测量$p_k = |z_k|^2 / \|z\|^2$,而不是对实对数进行softmax。每一层都提供了一个CPU参考和一个经过有限差分检查的CUDA核,计算图在批处理中被克隆以在GPU上执行。在基础层之上,我们添加了信号处理原语,将恒等卷积(conv(x,k) = IFFT(FFT(x) . FFT(k))转换为可学习的复值卷积网络,还有一个真正的Adam优化器和一个减少内存的推断模式。 我们报告了三项研究。首先,一个完全复值的FNet风格因果序列模型,建立在一个新的$O(N \log N)$因果傅立叶混频器上 - 一个由Bluestein / chirp-z因子化评估的三角掩模DFT:一旦正确调整,它匹配或超过了参数匹配的实值因果FNet在字符级语言建模上的质量,达到了实模型在一半以下训练步骤中的收敛质量。第二和第三,无线调制分类(RML2016.10a)和相干衍射成像的傅立叶相位问题的瓶颈分析,这些分析准确地确定了复值网络仍需要新操作符的地方。在这三个研究中,复值表达式可以可靠地学习出物理上正确的结构。 代码:https://github.com/crasmarum/CNet

更新时间: 2026-10-06 15:57:33

领域: cs.LG,cs.MS,math.OC

下载: http://arxiv.org/abs/2610.08592v1

TwinViT-DeepJSCC: Adversarially Robust Semantic Image Communication

Learning-based semantic communication is vulnerable to adversarial perturbations introduced before semantic encoding or over wireless channels. This paper proposes TwinViT-DeepJSCC, a preventive-corrective semantic image transceiver operating under a fixed channel-use budget. Two Vision Transformer (ViT)-based deep joint source-channel coding (DeepJSCC) branches learn complementary latent representations protected by sensitivity-aware masking. At the receiver, confidence-aware fusion, blind corruption-severity estimation, and signal-to-noise ratio (SNR)-severity-conditioned denoising diffusion implicit model (DDIM) purification mitigate residual corruption without requiring attack metadata. Experiments on the Canadian Institute for Advanced Research 100-class (CIFAR-100) dataset consider fast gradient sign method (FGSM), projected gradient descent (PGD), natural evolution strategies (NES), and Carlini-Wagner (CW) source-domain attacks, as well as random jamming and channel-aware adversarial waveforms over additive white Gaussian noise (AWGN) and block-flat Rayleigh fading. Under matched channel-use and attack budgets, TwinViT-DeepJSCC achieves maximum peak signal-to-noise ratio (PSNR) gains of approximately 9.5 dB under 20-step PGD and 10.8 dB under channel-aware waveform attacks over block-flat Rayleigh fading. Under PGD, it also improves Top-1 accuracy by up to approximately 38 percentage points over the undefended baseline and 13 percentage points over the strongest competing defense. Ablation results confirm the complementary contributions of the proposed transmitter- and receiver-side mechanisms.

Updated: 2026-10-06 15:57:08

标题: TwinViT-DeepJSCC:对抗鲁棒语义图像通信

摘要: 基于学习的语义通信容易受到在语义编码之前引入的对抗性扰动或无线信道的影响。本文提出了一种名为TwinViT-DeepJSCC的预防性-纠正性语义图像收发器,其在固定信道使用预算下运行。两个基于Vision Transformer(ViT)的深度联合源-信道编码(DeepJSCC)分支学习了由敏感性感知掩模保护的互补潜在表示。在接收端,置信度感知融合、盲目损坏严重性估计以及信噪比(SNR)严重性条件下的去噪扩散隐式模型(DDIM)净化减轻了残余损坏,而无需攻击元数据。在加拿大高级研究所100类(CIFAR-100)数据集上的实验考虑了快速梯度符号方法(FGSM)、投影梯度下降(PGD)、自然进化策略(NES)和Carlini-Wagner(CW)源域攻击,以及在加性白高斯噪声(AWGN)和块平射线衰落上的随机干扰和信道感知对抗波形。在匹配的信道使用和攻击预算下,TwinViT-DeepJSCC在20步PGD下达到了约9.5 dB的最大峰值信噪比(PSNR)增益,并在块平射线衰落上的信道感知波形攻击下达到了约10.8 dB的增益。在PGD下,它还将Top-1准确率提高了约38个百分点,超过了未防御基线和最强竞争防御的13个百分点。消融结果证实了所提出的发射端和接收端机制的互补贡献。

更新时间: 2026-10-06 15:57:08

领域: cs.CR,eess.SP

下载: http://arxiv.org/abs/2610.08590v1

MINDSET: Energy-based Schema Evolution for Long Conversational Agent Memory

Long conversational agents have become essential in our daily lives. They must remember what was said long back in order to help us efficiently complete a task without needing the user to repeat instructions and context repeatedly. However, the main issue is that instructions and context change over time and so the agents must be able to adapt accordingly. A useful memory system should preserve both current and historical states, distinguish stale information from active knowledge, retrieve evidence appropriate to the query and avoid repeatedly invoking a large language model to rewrite prior interactions. We introduce MINDSET, a memory controller that stores a conversation as immutable episodes and organizes them into versioned schemas through minimum-energy state transitions. Each incoming episode may reinforce, supersede, split or create a schema. The transition decision balances representation distortion, contradiction, historical damage, fragmentation and internal inconsistency, while hysteresis prevents isolated contradictions from prematurely rewriting stable memory. We evaluate MINDSET against 5 memory systems on a reproducible sample of 850 questions (700 LoCoMo + 150 MemoryAgentBench). MINDSET obtains the highest observed LoCoMo answer F1 while significantly improving retrieval ranking (Recall@8, MRR and nDCG@8) over the second best method LightMem (p<0.01 after Holm correction). It obtains the highest observed scores on MemoryAgentBench although the relative difference is low. Ablations identify controlled fragmentation and schema-aware assignment as the largest contributors to answer quality. Additionally, a 700-question cross-model evaluation with GLM-4.7 and Gemma-4-31B supported model independence. These results show that long-term memory can be better handled as constrained state management rather than continual summarization.

Updated: 2026-10-06 15:55:42

标题: 思维模式:基于能量的模式演变,用于长对话代理记忆

摘要: 长对话代理在我们的日常生活中变得至关重要。他们必须记住很久以前说过的话,以便在不需要用户重复指令和上下文的情况下有效地帮助我们完成任务。然而,主要问题在于指令和上下文随着时间的推移而发生变化,因此代理必须能够相应地进行适应。一个有用的记忆系统应该保留当前和历史状态,区分过时信息和活跃知识,检索与查询相关的证据,并避免反复调用大型语言模型来重写先前的交互。我们介绍了MINDSET,这是一个记忆控制器,它将对话存储为不可变的剧集,并通过最小能量状态转换将它们组织成版本化的模式。每个传入的剧集可能会强化、取代、分裂或创建一个模式。转换决策平衡了表示失真、矛盾、历史损害、碎片化和内部不一致性,而滞后效应防止孤立的矛盾过早地重写稳定的记忆。我们对850个问题的可再现样本(700个LoCoMo + 150个MemoryAgentBench)对MINDSET进行了评估。MINDSET在观察到的LoCoMo答案F1上表现最好,同时在检索排名(Recall@8、MRR和nDCG@8)上明显提高,超过了第二好的方法LightMem(Holm校正后p<0.01)。在MemoryAgentBench上观察到的得分最高,尽管相对差异较小。消融试验确定了受控的碎片化和模式感知分配是答案质量的最大贡献者。此外,与GLM-4.7和Gemma-4-31B进行的700个问题的跨模型评估支持模型的独立性。这些结果表明,长期记忆可以更好地处理为受限状态管理,而不是持续摘要。

更新时间: 2026-10-06 15:55:42

领域: cs.AI

下载: http://arxiv.org/abs/2610.08586v1

Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems

LLM-agent evaluations commonly measure task success or agreement with a declared plan. In strategic cyber-physical systems, an architecture must also remain appropriate after autonomous participants respond and physics constrains outcomes. We introduce a controlled benchmark of planning-induced control trajectories: ordered planning operations and directives linking execution architecture to strategic response and physical consequences. Four coded executors (predefined, sequential, hierarchical, and search) control demand response for 40 prosumers on a radial feeder. The LLM declares or advises typed policies and mediates communication; schedules, base prosumer dynamics, stochastic actions, and power flow remain explicit code. Paired forced-mode counterfactuals, exact-prompt caching, common response draws with separate randomness streams, critic isolation, and event-level feasibility isolate comparisons. The Llama-3.3-70B experiments on this feeder distinguish three properties. First, forced search is the oracle in all five baseline seeds under the specified objective. Second, injected objective substitution preserves mode agreement at 1.0 while increasing cumulative voltage shortfall by 2.68x. Third, the 144-scenario, 576-episode factorial bank, using three repeated seeds, contains feasible oracles from predefined, sequential, and search. The prespecified stress-held-out ridge has mean regret 90.7 and no observed value over fixed sequential. A post-hoc constraint-aware analysis reduces regret to 29.0; a simple deadline rule attains 28.7, so this gain does not establish a learning advantage. An all-feasible ablation does not improve over fixed search. These are simulation-internal, descriptive comparisons. A five-model, 300-declaration extension tests interface behaviour, not cross-backbone physical rankings; shared-endpoint latency tails motivate probabilistic live feasibility.

Updated: 2026-10-06 15:55:01

标题: 在网络物理系统中LLM代理的规划策略的战略评估

摘要: LLM代理评估通常衡量任务成功或与声明计划的一致性。在战略性的网络物理系统中,架构必须在自主参与者响应和物理约束结果之后仍然合适。我们引入了一个受控的基准,用于规划诱导的控制轨迹:有序的规划操作和指令将执行架构与战略响应和物理结果联系起来。四个编码执行者(预定义的、顺序的、分层的和搜索)控制40个生产者和消费者在一个径向馈线上的需求响应。LLM声明或建议类型政策并调解沟通;调度、基础生产者动态、随机行为和电力流保持显式代码。配对的强制模式反事实、确切提示缓存、具有单独随机流的常见响应抽取、评论孤立和事件级可行性隔离比较。在这个馈线上进行的Llama-3.3-70B实验区分了三个特性。首先,在指定的客观条件下,强制搜索是所有五个基线种子中的预言者。其次,注入的客观替代保持模式一致性为1.0,同时将累积电压不足增加了2.68倍。第三,使用三个重复种子的144个场景、576个情节的因子银行中包含了来自预定义、顺序和搜索的可行预言者。预先指定的压力保留岭的平均后悔值为90.7,没有观察到固定顺序的值。事后的约束感知分析将后悔减少到29.0;一个简单的截止时间规则达到了28.7,因此这种收益并不能建立学习优势。一个全可行的消融并不能改进固定搜索。这些是模拟内部的描述性比较。一个五模型、300声明的扩展测试接口行为,而不是交叉骨干物理排名;共享端点延迟尾部激励概率实时可行性。

更新时间: 2026-10-06 15:55:01

领域: cs.MA,cs.AI,eess.SY

下载: http://arxiv.org/abs/2608.04265v2

Random Feature Gaussian Process Attention: Linear-Time Probabilistic Attention with Calibrated Uncertainty

Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising direction addresses this issue by interpreting attention as a Gaussian process (GP) posterior, which enables principled uncertainty calibration but incurs cubic complexity in sequence length due to the inversion of the kernel; although decoupled GP variants reduced the cost to quadratic, the computation remains prohibitive in practice. In this paper, we propose the plug-and-play random Fourier feature Gaussian process attention (RFF-GPA) module, which represents the attention as a GP with a stationary kernel approximated by random Fourier features. This low-rank approximation results in linear-time complexity for approximating the posterior mean and variance, making it far more scalable compared to previous work. Empirical results on multiple real-world datasets show that our attention module improves calibration while maintaining predictive accuracy, and simultaneously reduces computational complexity to linear in the sequence length.

Updated: 2026-10-06 15:51:29

标题: 随机特征高斯过程注意力:具有校准不确定性的线性时间概率注意力

摘要: Transformer提供了一个最先进的建模框架,但是由于精度不足,限制了它们在安全关键应用中的可靠性。一个有前途的方向是将注意力解释为高斯过程(GP)后验,这使得原理上可以校准不确定性,但由于内核的反演导致序列长度的立方复杂性;尽管解耦的GP变体将成本降低到二次,但实际计算仍然是禁止的。在本文中,我们提出了插入式随机傅里叶特征高斯过程注意力(RFF-GPA)模块,它将注意力表示为具有由随机傅里叶特征近似的平稳内核的GP。这种低秩近似导致用于近似后验均值和方差的线性时间复杂性,使其与以前的工作相比更具可扩展性。多个真实世界数据集上的实证结果显示,我们的注意力模块在提高校准性的同时保持预测准确性,并将计算复杂性同时减少到序列长度的线性。

更新时间: 2026-10-06 15:51:29

领域: cs.LG

下载: http://arxiv.org/abs/2610.08578v1

How Learning Governs Unlearning across the Memorization-Generalization Spectrum

While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their subsequent unlearning. In this paper, we investigate this connection from the perspectives of memorization and generalization, the two most representative yet competing strategies that models employ during training. We first classify memorization- and generalization-heavy models using grokking in modular addition and compare their responses to unlearning, showing that the latter suffer greater retain damage, i.e., a larger performance drop on the retain set. Furthermore, we conduct a finer-grained analysis by introducing bucketed modular addition, in which the respective contributions of the two strategies can be explicitly controlled across the memorization-generalization spectrum. In this setup, we reaffirm that the same trend persists and is nearly monotonic. We further demonstrate that this relationship also holds in LLM unlearning across verbatim and factual recall settings. Finally, we provide two practical insights for developing better unlearning methods, highlighting the importance of accounting for learning dynamics in unlearning.

Updated: 2026-10-06 15:50:51

标题: 学习如何影响记忆-泛化谱上的遗忘

摘要: 尽管取消学习旨在否定通过学习获得的不良能力,但很少有研究探讨了模型学习方式如何塑造它们后续取消学习的过程。在本文中,我们从记忆和泛化的角度研究这种关联,这两种是模型在训练过程中采用的最具代表性但又相互竞争的策略。我们首先使用模块化加法对记忆和泛化程度较高的模型进行分类,并比较它们对取消学习的响应,结果显示后者遭受了更严重的保留损害,即在保留集上性能下降更大。此外,我们通过引入分桶式模块化加法进行了更精细的分析,在这种设置中,可以明确控制两种策略在记忆-泛化光谱上的各自贡献。在这种设置下,我们进一步证实了相同的趋势存在并且几乎是单调的。我们进一步证明,这种关系在LLM取消学习中也适用于逐字逐句和事实回忆设置。最后,我们提供了两个关于开发更好的取消学习方法的实用见解,强调需要考虑取消学习中的学习动态的重要性。

更新时间: 2026-10-06 15:50:51

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.08577v1

How To Track Qubits Through Space and Time (Or: Sailing in a Quantum Boat)

While quantum position verification aims to certify a prover's location using quantum information, existing security definitions only guarantee that part of the successful adversarial party is in the claimed location.This leaves open the possibility that a distributed team of adversaries can jointly simulate a prover in a way that defeats the intended meaning of `being at a location' in position-based cryptography. We introduce stronger notions of position verification that we call quantum localization, which requires that there is a specified, unclonable state at the verified spacetime point-and that this state can be found nowhere else.We show that quantum localization leads naturally to a meaningful notion of trajectory verification, in which quantum information is verifiably tracked through space and time.We construct quantum localization and trajectory verification protocols using quantum anchor states, which generalize coset states from unclonable cryptography.The security of our schemes is proven in the classical oracle (i.e. ideal obfuscation) model, which can be heuristically instantiated in the plain model using post-quantum indistinguishability obfuscation. We also introduce and instantiate the concept of functionality localization, which guarantees that the adversary has the ability to compute a secret function at the verified spacetime point, and this function cannot be computed anywhere else.This raises the intriguing possibility of localizing computational capabilities in space and time. More broadly, we believe our notions of quantum localization and subsequent feasibility results provide stronger foundations for position-based cryptography.At a conceptual level, our results also explore the tight link between two fundamental notions in physics-location and information-thus serving as a natural foundation for cryptographic capabilities tied to physical spacetime.

Updated: 2026-10-06 15:49:59

标题: 如何跟踪量子比特在空间和时间中的运动(或者:在量子船上航行)

摘要: 量子位置验证旨在使用量子信息证明验证者的位置,而现有的安全定义仅保证成功的对手的一部分位于所声称的位置。这使得可能存在一个分布式的对手团队可以共同模拟一个验证者,从而打败了位置为基础的密码学中“在某个位置”的原意。我们引入了更强的位置验证概念,称之为量子定位,它要求在验证的时空点存在一个指定的、无法复制的状态,并且这个状态在其他地方找不到。我们展示了量子定位自然地导致了一个有意义的轨迹验证概念,其中量子信息可以通过时空可验证地跟踪。我们使用量子锚定态构建了量子定位和轨迹验证协议,这些协议将无法复制的密码学中的余集态进行了泛化。我们的方案的安全性在经典预言模型(即理想混淆)中得到证明,可以在明文模型中使用后量子区分混淆启发式地实现。我们还介绍并实现了功能性定位的概念,它保证对手有能力在验证的时空点计算一个秘密函数,而这个函数在其他地方无法计算。这引发了在时空中本地化计算能力的有趣可能性。更广泛地说,我们相信我们关于量子定位和随后可行性结果的概念为基于位置的密码学提供了更坚实的基础。在概念上,我们的结果还探讨了物理学中两个基本概念-位置和信息之间的紧密联系,因此为与物理时空相关的密码能力提供了一个自然的基础。

更新时间: 2026-10-06 15:49:59

领域: quant-ph,cs.CR

下载: http://arxiv.org/abs/2605.30732v2

Rethinking Adapter Placement: A Dominant Adaptation Module Perspective

Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method that places trainable low-rank adapters into frozen pre-trained models. Recent studies show that using fewer LoRA adapters may still maintain or even improve performance, but existing methods still distribute adapters broadly, leaving \emph{where to place a limited number of adapters to maximize performance} largely open. To investigate this, we introduce \textbf{PAGE} (\textbf{P}rojected \textbf{A}dapter \textbf{G}radient \textbf{E}nergy), a gradient-based sensitivity probe that estimates the initial trainable gradient energy available to each candidate LoRA adapter. Surprisingly, we find that PAGE is highly concentrated on a single shallow FFN down-projection across two model families and four downstream tasks. We term this module the \textbf{dominant adaptation module} and show that its layer index is architecture-dependent but task-stable. Motivated by this finding, we propose \textbf{DomLoRA}, a placement method that places a single adapter at the dominant adaptation module. With only \textbf{0.7\%} of vanilla LoRA's trainable parameters, DomLoRA outperforms it on average across downstream tasks, including instruction following, mathematical reasoning, coding, and multi-turn conversation. This method also matches or improves other LoRA variants and reduces training time by up to \textbf{2.74}$\times$ compared with broad placement, supporting the dominant adaptation module perspective as a practical placement guideline.

Updated: 2026-10-06 15:49:15

标题: 重新思考适配器位置:主导适应模块视角

摘要: 低秩适应(LoRA)是一种广泛使用的参数高效微调方法,将可训练的低秩适配器放入冻结的预训练模型中。最近的研究表明,使用较少的LoRA适配器仍然可以保持甚至提高性能,但现有方法仍然广泛分布适配器,使得\emph{在哪里放置有限数量的适配器以最大化性能}基本上是开放的。为了调查这一点,我们引入了\textbf{PAGE}(\textbf{P}rojected \textbf{A}dapter \textbf{G}radient \textbf{E}nergy),这是一种基于梯度的敏感性探测器,用于估计每个候选LoRA适配器的初始可训练梯度能量。令人惊讶的是,我们发现PAGE在两个模型系列和四个下游任务中高度集中在一个浅层FFN向下投影上。我们将这个模块称为\textbf{主导适应模块},并展示其层索引是依赖于架构但稳定于任务。受到这一发现的启发,我们提出了\textbf{DomLoRA},一种将单个适配器放置在主导适应模块的方法。仅使用\textbf{0.7\%}的vanilla LoRA的可训练参数,DomLoRA在下游任务中平均优于它,包括指令跟随、数学推理、编码和多轮对话。这种方法还匹配或改进其他LoRA变体,并且与广泛放置相比,可以将训练时间缩短多达\textbf{2.74}$\times$,支持主导适应模块视角作为实际放置指南。

更新时间: 2026-10-06 15:49:15

领域: cs.AI,cs.CL,cs.LG

下载: http://arxiv.org/abs/2605.06183v2

FedDermaSeg: Federated Learning for Dermatological Image Segmentation

Skin cancer is a major global health concern, and early detection and accurate lesion delineation are important for effective diagnosis and treatment planning. Automated skin lesion analysis can assist dermatologists, with lesion segmentation serving as a fundamental step in computer-aided diagnostic systems. Conventional deep learning-based segmentation models typically rely on centralized training, where images and their corresponding segmentation masks are collected on a central server. Such data aggregation raises privacy concerns in medical applications and requires substantial centralized computational resources. To address these limitations, we investigate the feasibility of federated learning for privacy-preserving skin lesion segmentation. The training and validation sets of the ISIC 2018 Skin Lesion Segmentation Challenge dataset are used to simulate a distributed learning environment and develop a federated segmentation model. The resulting model is evaluated on the ISIC 2018 test set and the PH2 dataset to assess its performance and generalizability. Experimental results demonstrate that the federated model achieves performance comparable to centralized training while consistently improving upon the locally trained models. These findings demonstrate the potential of federated learning for collaborative skin lesion segmentation without requiring centralized aggregation of medical images.

Updated: 2026-10-06 15:49:04

标题: FedDermaSeg:用于皮肤病图像分割的联邦学习

摘要: 皮肤癌是一个重要的全球健康问题,早期检测和准确的病变定位对于有效的诊断和治疗规划至关重要。自动皮肤病变分析可以帮助皮肤科医生,而病变分割作为计算机辅助诊断系统中的基本步骤。传统的基于深度学习的分割模型通常依赖于集中式训练,其中图像及其对应的分割掩模被收集在中央服务器上。这种数据聚合引发了医学应用中的隐私问题,并且需要大量集中式的计算资源。为了解决这些限制,我们研究了联邦学习用于隐私保护皮肤病变分割的可行性。ISIC 2018皮肤病变分割挑战数据集的训练和验证集被用来模拟分布式学习环境并开发一个联邦分割模型。得到的模型在ISIC 2018测试集和PH2数据集上进行评估,以评估其性能和泛化能力。实验结果表明,联邦模型实现了与集中式训练可比的性能,同时始终优于本地训练的模型。这些发现展示了联邦学习在协作皮肤病变分割中的潜力,而不需要对医学图像进行集中聚合。

更新时间: 2026-10-06 15:49:04

领域: cs.CV,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08574v1

RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems

Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.

Updated: 2026-10-06 15:47:47

标题: RAG-PIBench:一个泄漏感知的基准测试,用于可信RAG系统中的提示注入检测

摘要: 检索增强生成(RAG)系统容易受到嵌入在检索内容中的提示注入攻击的影响。我们引入了RAG-PIBench,这是一个用于RAG风格提示注入检测的基准,包含4,876个上下文示例,涵盖了冻结的训练、验证和受保护的测试分割。通过使用一个泄漏感知的构建流水线和严格的评估协议,我们比较了基于关键词、语义参考、TF-IDF和基于Transformer的检测器。DistilBERT在受保护的测试性能中表现最好(F1 = 0.896,PR-AUC = 0.968),而TF-IDF SVM和逻辑回归仍然具有竞争力。我们的结果表明,在RAG系统中,泄漏感知基准设计和强大的稀疏基线对于可靠的提示注入检测具有价值。

更新时间: 2026-10-06 15:47:47

领域: cs.CR,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08571v1

Less Is More: A Leakage-Controlled Study of Dermoscopic Preprocessing for Joint Skin Lesion Classification and Segmentation with YOLO26

Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheless, its actual contribution to modern real-time models remains unclear, particularly when evaluation protocols do not adequately control correlations among images of the same lesion. This study presents a leakage-controlled, lesion-disjoint evaluation of dermoscopic preprocessing and augmentation for joint multi-class lesion classification and instance segmentation using a fixed nano-scale YOLO26 segmentation model (YOLO26n-seg). From HAM10000 (10,015 images), quality control yields 10,013 valid image-mask pairs from 7,468 unique lesions, partitioned into mutually exclusive sets by lesion identity. With the architecture, resolution, training budget, and evaluation protocol held fixed, we compare minimally processed images plus online augmentation against offline class balancing, DullRazor-CLAHE preprocessing, and raw-processed hybrid views, over three random seeds. On the lesion-disjoint test set, the raw baseline achieves a mask mAP$_{50:95}$ of $0.5636 \pm 0.0234$, a Dice score of $0.9356 \pm 0.0024$, and a macro-F1 score of $0.6917 \pm 0.0202$. Offline augmentation does not improve the mean performance, while the combined and hybrid strategies reduce both class-aware segmentation and classification accuracy. At only 2.69 million parameters, the model runs at approximately 50 frames per second. Under a leakage-controlled, lesion-disjoint protocol with all non-input factors held fixed, minimally processed dermoscopic images combined with standard online augmentation deliver a better accuracy-efficiency trade-off than increasingly complex deterministic preprocessing, which yields no consistent joint benefit across three seeds on HAM10000.

Updated: 2026-10-06 15:47:39

标题: Less Is More: 一项关于使用YOLO26进行联合皮肤病变分类和分割的皮肤镜预处理的泄漏控制研究

摘要: 手工预处理在自动皮肤镜分析中被广泛使用,以抑制成像伪影并增强病变的可见性。然而,它对现代实时模型的实际贡献仍不清楚,特别是当评估协议未充分控制同一病变图像之间的相关性时。本研究提出了一种受泄漏控制的病变分离评估方法,用于关节多类病变分类和实例分割,使用固定的纳米尺度YOLO26分割模型(YOLO26n-seg)。从HAM10000(10,015张图像)中,质量控制得到了10,013个有效的图像-掩膜对,来自7,468个唯一病变,按病变身份互斥地划分为集合。在保持架构、分辨率、训练预算和评估协议不变的情况下,我们比较了最小化处理的图像加上在线增强与离线类平衡、DullRazor-CLAHE预处理和原始处理的混合视图,通过三个随机种子。在病变分离测试集上,原始基准实现了一个mask mAP$_{50:95}$为$0.5636 \pm 0.0234$,Dice分数为$0.9356 \pm 0.0024$,宏-F1分数为$0.6917 \pm 0.0202$。离线增强并没有提高平均性能,而组合和混合策略降低了类感知分割和分类准确性。在仅有2.69百万参数的情况下,该模型运行速度约为每秒50帧。在一个受泄漏控制的、病变分离的协议下,所有非输入因素保持不变的情况下,最小化处理的皮肤镜图像与标准在线增强相结合,提供了比越来越复杂的确定性预处理更好的准确性和效率的平衡,这在HAM10000上的三个种子中都没有一致的联合优势。

更新时间: 2026-10-06 15:47:39

领域: cs.LG,cs.CV

下载: http://arxiv.org/abs/2610.08570v1

Singular Value Decomposition: A Geometric Rediscovery, Where Proofs Become Algorithms

This article is a geometric rediscovery of the singular value decomposition, with a further claim: the construction it builds is the machinery behind much of machine learning. The same argument that answers an idle question about ellipses is the algorithm behind principal component analysis, kernel methods, and PageRank, and it is not only the results that transfer but the proofs themselves, run as procedures. The usual introduction states $A = UΣV^T$ and justifies it via the spectral theorem applied to $A^T A$. This is correct but unilluminating, since it assumes a powerful theorem to reach a result that is, in the end, about ellipses. Part I reverses the order. A linear map sends the unit circle to an ellipse; one asks which input directions map to its axes, and finds, example after example, that they are perpendicular. In the plane this can be watched: rotate a frame, track how far its images are from perpendicular, and a sign change forces a frame where they are exactly perpendicular, which is also where the map stretches hardest. Maximizing the stretch and recursing generalizes this to n dimensions, with singular values falling out in order, and the construction proves the spectral theorem rather than assuming it. Part II puts each construction to work: maximize-and-recurse becomes the power method and PageRank; the lemma locating the maximizer becomes the stopping rule of gradient descent; the duality between $A^T A$ and $A A^T$ becomes the transport at the heart of kernel PCA. Each connection is stated with its boundary, saying what the decomposition supplies and where another idea takes over. Prerequisites are the standard sophomore sequence, and the worked examples are small enough to check by hand.

Updated: 2026-10-06 15:44:32

标题: 奇异值分解:几何上的重新发现,证明变成算法

摘要: 这篇文章重新发现了奇异值分解的几何性质,并进一步声称:它构建的机制是许多机器学习背后的基础。回答一个关于椭圆的无关问题的论据,也是主成分分析、核方法和PageRank背后的算法,不仅是结果可以转移,连证明本身也可以运行为过程。 通常的介绍说明$A = UΣV^T$,并通过应用于$A^T A$的谱定理来证明。这是正确的,但缺乏启发性,因为它假设了一个强大的定理来达到最终关于椭圆的结果。第一部分颠倒了顺序。一个线性映射将单位圆映射到一个椭圆;人们问哪些输入方向映射到它的轴,然后发现,例子接连不断,它们是垂直的。在平面上可以观察到这一点:旋转一个框架,跟踪它的图像与垂直的距离,一个符号变化会导致一个框架,其中它们正好是垂直的,这也是映射最大程度伸展的地方。最大化伸展并递归将其推广到n维空间,奇异值按顺序落入,构造证明了谱定理而不是假设它。 第二部分将每个构造应用到实际中:最大化和递归成为幂方法和PageRank;定位最大化者的引理成为梯度下降的停止规则;$A^T A$和$A A^T$之间的对偶性成为核PCA核心的传输。每个连接都陈述了其边界,说明分解提供了什么,另一个想法在哪里接管。先修条件是标准的大二课程序列,而示例足够小,可以手动检查。

更新时间: 2026-10-06 15:44:32

领域: cs.LG,math.HO

下载: http://arxiv.org/abs/2610.08565v1

Valid for Free: Homophily-Gated Conformal Prediction for Training-Free Node Classification with Tabular Foundation Models

Tabular foundation models (TFMs) can classify the nodes of a graph without training on it, by reading node and neighborhood features as table rows next to labeled context rows. Work in this line reports predictive performance, not conformal coverage or prediction-set size. To our knowledge, we give the first reliability study of the setting, with TabICL as the TFM and half of each graph as labeled context. As for any predictor fixed before calibration, a frozen in-context predictor makes split conformal prediction exactly valid in finite samples, with no training, validation fold, or tuning on the target graph. An audit across ten graphs then shows that the training-free TabICL posterior has lower expected calibration error (ECE) than GCN with temperature scaling (GCN+TS) on nine of them. Its mean ECE over the ten graphs is 0.019, about 35 percent below the 0.029 of GCN+TS. We also introduce HG-DAPS, a training-free diffusion score whose homophily gate reads only the in-context labels, so the guarantee still holds. Relative to adaptive prediction sets (APS), it reduces mean set size by 5.8 to 17.1 percent on six homophilous graphs and changes it by under 1 percent on four heterophilous ones. On two binary, class-imbalanced graphs, a pre-registered trap case shows that gating on raw rather than adjusted homophily lowers coverage among low-homophily nodes by 0.27 and 0.12. Marginal coverage stays at the nominal 0.90 and masks this drop.

Updated: 2026-10-06 15:44:07

标题: 免费有效:同质性门控的符合性预测,用于无需训练的基于表格的节点分类

摘要: 表格化基础模型(TFMs)可以在不对其进行训练的情况下对图的节点进行分类,通过将节点和邻域特征作为标记上下文行旁边的表格行进行阅读。这一领域的工作报告了预测性能,而不是符合性覆盖率或预测集大小。据我们所知,我们是该设置的第一个可靠性研究,其中TabICL作为TFM,每个图的一半作为标记上下文。对于任何在校准之前固定的预测器,一个在上下文中冻结的预测器使得在有限样本中分裂符合性预测完全有效,无需对目标图进行训练、验证折叠或调整。然后,对十个图进行审计显示,无需训练的TabICL后验在九个图中的期望校准误差(ECE)低于具有温度缩放(GCN+TS)的GCN。在十个图中的平均ECE为0.019,比GCN+TS的0.029低约35%。我们还引入了HG-DAPS,一种无需训练的扩散分数,其同质性门只读取上下文标签,因此保证仍然成立。相对于自适应预测集(APS),在六个同质图上将平均集大小减少了5.8至17.1%,在四个异质图上将其改变不到1%。在两个二元、类别不平衡的图上,一个预先注册的陷阱案例表明,对原始同质性进行分组而不是调整后的同质性会导致低同质性节点的覆盖率降低0.27和0.12。边际覆盖率保持在名义0.90,并掩盖了这种降低。

更新时间: 2026-10-06 15:44:07

领域: cs.LG

下载: http://arxiv.org/abs/2610.08564v1

Adaptive Power Sampling for LLM Reasoning

Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM). Nevertheless, existing methods typically sharpen the base model distribution uniformly across queries, overlooking variations in query difficulty and in how well the base model already handles each query. The goal of this work is to equip power sampling with query adaptivity. Theoretically, we show that the benefits of further sharpening are determined by the self-reward gap between correct and incorrect responses. Based on this insight, we propose \emph{Adaptive Power Sampling} (APS), which adjusts the sharpening exponent on a per-query basis at test time using the relationship between answer agreement and the model's self-reward. Experiments across diverse reasoning tasks, including MATH500, HumanEval, and GPQA, show that APS consistently outperforms power sampling with a fixed sharpening exponent, without additional training.

Updated: 2026-10-06 15:43:35

标题: 自适应功率采样用于LLM推理

摘要: 最近,序列级功率抽样已经成为一种无需训练的推理方法,通过从基本大型语言模型(LLM)的锐化输出分布中进行抽样。然而,现有方法通常在查询之间均匀锐化基本模型分布,忽视了查询难度的变化以及基本模型已经处理每个查询的情况。本研究的目标是为功率抽样提供查询适应性。理论上,我们表明进一步锐化的好处取决于正确和不正确响应之间的自我奖励差距。基于这一洞察力,我们提出了自适应功率抽样(APS),它通过在测试时间使用答案一致性和模型的自我奖励之间的关系,在每个查询基础上调整锐化指数。跨多种推理任务的实验,包括MATH500、HumanEval和GPQA,表明APS始终优于具有固定锐化指数的功率抽样,而无需额外训练。

更新时间: 2026-10-06 15:43:35

领域: cs.AI

下载: http://arxiv.org/abs/2610.08563v1

Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers

Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy exploration combined with a neural reward model is effective. In this paper, we address this question by modeling the reward as a hierarchical function on the response space: the reward consists of infinitely many local components, each of which becomes relevant only after the preceding ones have been resolved. We show that a natural Transformer-based actor--critic algorithm, which alternates between sampling from the current KL-regularized policy, fitting a Transformer critic to the observed rewards, and updating the policy, achieves the minimax optimal rates in the query budget and in the regularization strength up to logarithmic factors, and is minimax optimal for a fixed number of prompts. In contrast, we prove that sampling from the fixed reference distribution, as in offline reward modeling, can limit regret decay to a logarithmic rate. These results show that on-policy exploration progressively zooms in on the region where the reward is concentrated, and quantify its benefit for RL post-training.

Updated: 2026-10-06 15:43:08

标题: Hierarchical Reasoning Rewards的强化学习:使用Transformers实现的极小极大最优奖励率

摘要: 强化学习(RL)已经成为后期训练语言模型在推理任务上的标准工具,其中通过奖励反馈更新策略,同时探索响应空间。尽管在实证方面取得了成功,但对RL后期训练的理论理解仍然有限,特别是对于为什么基于政策探索结合神经奖励模型有效的问题。在本文中,我们通过将奖励建模为响应空间上的分层函数来解决这个问题:奖励由无限多个局部组成,每个局部只在前面的局部已经解决后才变得相关。我们展示了一个自然的基于Transformer的演员-评论家算法,该算法在从当前的KL正则化策略中抽样、将Transformer评论家拟合到观察到的奖励中、以及更新策略之间交替,实现了在查询预算和正则化强度上的极小最优率,直到对数因子,并对于固定数量的提示是极小最优的。相比之下,我们证明从固定参考分布中抽样,就像在离线奖励建模中一样,可以将遗憾下降限制在对数速率。这些结果表明,基于政策的探索逐渐聚焦于奖励集中的区域,并量化了它对RL后期训练的好处。

更新时间: 2026-10-06 15:43:08

领域: cs.LG,stat.ML

下载: http://arxiv.org/abs/2610.08561v1

Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness

Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.

Updated: 2026-10-06 15:43:04

标题: 我已经看够了吗? 冻结视频语言模型编码证据准备

摘要: 流媒体视频语言模型不仅必须决定如何回答问题,还必须确定当前问题所需的证据是否已经到位。现有系统将该决策学习为一个单独的触发器;我们研究是否一个未经修改的模型已经在计算这一决策。我们展示了冻结的VideoLLMs携带一个线性可读的准备就绪信号,从时间戳证据中标记而不是从模型输出中标记。在共享字节相同的评估中(在最严格的未准备采样下,其中一个适应的时钟接近机会),它在所有七个模型中解码(AUROC为0.733-0.905),而且一个没有任何基准家族素材的探针仍然可以读取该家族。这种准备是问题条件的:在字节相同的窗口上,只改变问题就会导致66.1%的对照组反转读数,而每个不考虑问题的对照组都是随机的。模型可以答错但仍然编码准备就绪:错误答案的AUROC仍为0.722。准备就绪也胜过不确定性估计器及其监督组合在延迟匹配的答案选择上,比置信度更接近独立人类判断。发布的流媒体触发器也是线性读出的,然而,一个训练有素的触发器在其基础模型的激活上读取时近似正交于准备就绪,解码准确性远远低于一个探针。我们将读出转换为准备就绪门控,这是一种回答时间策略,可以在匹配视频持续时间时提高最多+9.75个百分点的准确性,而计算开销几乎可以忽略不计。它的增益量取决于任务可用的准确性空间:在26个配置中,增益与该准确性空间相匹配,并且一个将其移动到相同像素上的干预措施也会随之提高增益。

更新时间: 2026-10-06 15:43:04

领域: cs.CV,cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.08560v1

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Although trained only on forward perturbation-response prediction, PertMind improves response inference in unseen cellular contexts while retaining general language capabilities. It also transfers, without task-specific post-training, to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generates biological profiles that support competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.

Updated: 2026-10-06 15:42:49

标题: PertMind:通过在细胞扰动数据上进行强化学习引发出生物学推理的产生

摘要: 大型语言模型可以描述机制,但可扩展的后训练仍然依赖于昂贵的、手工筛选的生物推理痕迹。在这里,我们展示细胞扰动图谱可以成为强化学习环境,其中测量的基因响应为生物推理提供可计算的奖励。我们引入PertMind,它结合了可信轨迹监督初始化和基因、通路和格式级别的强化信号。尽管只在前向扰动响应预测上训练,PertMind在未见细胞环境中改善了响应推断,同时保留了一般语言能力。它还在没有任务特定后训练的情况下转移到逆向扰动识别、双扰动推理、表型筛选优先级和生物过程解释。PertMind进一步生成支持跨多尺度下游任务的竞争性基因、细胞和供体表示的生物特征。这些结果支持了一个假设,即对实验终点进行强化可以集中可重复使用的生物策略,这些策略已经可被预训练模型访问。更广泛地说,扰动导出的强化学习为将不断扩大的实验图谱转化为通用生物推理训练环境提供了一个可扩展的途径。

更新时间: 2026-10-06 15:42:49

领域: cs.LG,cs.AI,q-bio.QM

下载: http://arxiv.org/abs/2608.16419v3

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Yet it remains unclear where and how such associations are computed within VLMs. In this work, we show that VLMs rely on two concurrent mechanisms to represent spatial variable binding. In the language model backbone, intermediate layers represent content-independent spatial relations on top of visual tokens corresponding to objects. However, this mechanism plays only a secondary role in shaping model predictions. Instead, the dominant source of spatial information originates in the vision encoder, whose representations encode the layout of objects and are directly exploited by the language model backbone. Notably, this spatial signal is distributed globally across visual tokens, extending beyond object regions into surrounding background areas. We validate the generalization of our findings to complex natural images from the COCO dataset, where globally amplifying the vision-derived spatial representations across all image tokens corrects spatial variable binding failures across models of various sizes. Together, our results clarify how spatial variable binding is computed within VLMs and highlight the central role of vision encoders in enabling it.

Updated: 2026-10-06 15:42:39

标题: 《视觉-语言模型中的空间变量绑定的双重机制》

摘要: 许多多模态任务,如图像字幕和视觉问答,需要视觉语言模型(VLMs)将对象与其属性和空间关系绑定在一起。然而,目前还不清楚在VLMs内部这种关联是在哪里以及如何计算的。在这项工作中,我们展示了VLMs依赖于两种并行机制来表示空间变量绑定。在语言模型主干中,中间层在与对象对应的视觉标记之上表示与内容无关的空间关系。然而,这种机制在塑造模型预测方面起到次要作用。相反,空间信息的主要来源是视觉编码器,其表示编码对象的布局,并直接被语言模型主干利用。值得注意的是,这种空间信号在整个视觉标记中是全局分布的,延伸到对象区域之外的周围背景区域。我们验证了我们的发现对来自COCO数据集的复杂自然图像的泛化,通过在所有图像标记上全局增强视觉衍生的空间表示,纠正了不同大小模型中的空间变量绑定失败。总之,我们的结果澄清了VLMs内部如何计算空间变量绑定,并突出了视觉编码器在实现这一目标中的核心作用。

更新时间: 2026-10-06 15:42:39

领域: cs.CV,cs.LG

下载: http://arxiv.org/abs/2603.22278v3

Latent space bias directions in LLMs capture confidence, not fairness

Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model performance and limited transfer to new datasets. Our work analyses what the debiasing direction used for activation steering actually encodes, in order to shed light on its inconsistent performance. We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks. We find that this direction is dominated by model confidence, pointing from regions of high to low-probability tokens in activation space rather than encoding a meaningful representation of model bias. Steering along it does reduce measured bias, but this is a consequence of reducing model confidence: on QA benchmarks we find that this steering drives the model to abstain from answering, with a side effect of improving fairness metrics. Our experiments show that model confidence is the dominant separating factor between biased and anti-biased prompts in hidden space, indicating that isolating a linear representation of bias which is disentangled from model confidence is difficult and steering-based debiasing results should be interpreted with care. In short, steering appears to reduce bias, not by correcting the model's underlying preferences, but by making it less confident, even on tasks unrelated to bias.

Updated: 2026-10-06 15:42:35

标题: LLM中的潜在空间偏向捕捉的是信心,而非公平性

摘要: 激活导向作为一种轻量级推理时间去偏斜技术已经变得越来越受欢迎,尤其是对于大型语言模型。然而,先前的研究表明,导向向量的泛化能力较差,对模型性能产生了意想不到的影响,并且在新数据集上的转移有限。我们的研究分析了用于激活导向的去偏斜方向实际上编码了什么,以揭示其不一致的表现。我们研究了通过对比反偏斜和偏斜提示的激活获得的线性去偏斜方向,并将其评估为针对偏见和一般知识基准的导向干预。我们发现,这个方向主要由模型置信度主导,指向激活空间中概率高到低的令牌区域,而不是编码模型偏见的有意义表示。沿着这个方向导航确实减少了测量的偏见,但这是减少模型置信度的结果:在问答基准上,我们发现这种导向驱使模型放弃回答,从而改善了公平度量指标。我们的实验显示,模型置信度是隐藏空间中偏斜和反偏斜提示之间的主要区分因素,这表明隔离出一个与模型置信度解耦的偏见线性表示是困难的,并且应谨慎解释基于导向的去偏斜结果。简而言之,导向似乎通过降低置信度来减少偏见,而不是通过纠正模型的潜在偏好,甚至对于与偏见无关的任务也是如此。

更新时间: 2026-10-06 15:42:35

领域: cs.CL,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08559v1

Systemization of Knowledge (SoK): Human-Centered AI Safety for Youth

While HCI increasingly examines AI-safety for youth, the literature lacks a comprehensive view of what risks have been identified, how they are addressed, and whether proposed protections work in-practice. We systematically reviewed 100 empirical HCI studies involving children and youth interacting with or exposed to AI across schools, homes, care settings, and public services. Using the YAIR taxonomy for risks and the MIT Mitigation Taxonomy for countermeasures, we map which risks have been identified, whether each risk is addressed by countermeasure(s), and whether each countermeasure for that risk is implemented and even evaluated. The risk-countermeasure mapping shows that most risks are matched only with proposed/ideated countermeasures; few countermeasures have been implemented, and fewer still evaluated; and existing evaluations often measure technical performance rather than protection from harm. We identify where coverage is absent, where safeguards remain untested, and propose concrete directions for HCI research to strengthen youth AI-safety.

Updated: 2026-10-06 15:40:21

标题: 知识系统化(SoK):面向青少年的人类中心人工智能安全

摘要: 虽然人机交互领域越来越关注青少年的人工智能安全性,但文献缺乏一个全面的视角来了解已经识别出的风险是什么,它们是如何被解决的,以及提出的保护措施在实践中是否有效。我们系统地审查了100篇涉及儿童和青少年在学校、家庭、护理环境和公共服务中与人工智能互动或接触的实证人机交互研究。使用YAIR风险分类和MIT缓解分类来进行对抗措施,我们映射了已经识别出的风险,每个风险是否被对抗措施所解决,以及每个风险的对抗措施是否被实施甚至评估。风险-对抗措施的映射显示,大多数风险仅与提议的/构想中的对抗措施相匹配;很少有对抗措施被实施,而更少被评估;现有的评估通常衡量技术性能而不是保护免受伤害。我们确定了覆盖不足的地方,未经测试的安全防护措施,并提出了加强青少年人工智能安全性的人机交互研究的具体方向。

更新时间: 2026-10-06 15:40:21

领域: cs.HC,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08554v1

DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory

Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.

Updated: 2026-10-06 15:40:06

标题: DeltaTTT: 适用于非线性循环记忆的逐层优化

摘要: 顺序测试时间训练通过连续更新适应内存网络,每次更新都根据网络先前的状态计算内部梯度。直觉上,这种状态依赖性应该允许每次更新考虑内存已经学到的内容,并更好地整合新信息。然而,我们发现这种预期的优势在非线性内存中并不一致实现:固定基础的并行TTT基线优于其串行对应项。我们的探索性实验指出一个关键的基本困难:非线性内存在单次通过序列时比线性内存更难优化。为了缓解这种优化困难,我们引入了DeltaTTT,它用逐层学习代替了两层内存网络的联合内部梯度优化。每个层分配一个本地预测目标,并通过状态依赖的增量规则进行更新。这种公式保留了非线性输出,同时实现了分块并行计算。在DeltaNet和LaCT骨干上的实验显示,语言建模和检索方面的表现优于它们的递归基线。

更新时间: 2026-10-06 15:40:06

领域: cs.LG,cs.AI,cs.CL

下载: http://arxiv.org/abs/2610.08553v1

AnyBottle: A Recipe to Only Keep the Concepts You Really Need

Concept bottleneck models (CBMs) make predictions inspectable and intervenable by routing them through human-interpretable concepts, but originally required concept annotations. Annotation-free variants remove this requirement, but typically use large concept vocabularies, static at both training and inference, producing bottlenecks larger than any task or prediction needs and harder to inspect. We propose AnyBottle, a single recipe for building compact, task-specific CBMs. AnyBottle assumes only a frozen backbone and an unsupervised concept pool, such as a sparse autoencoder. A black-box teacher trained on the same backbone then guides selection: each round adds the concept that best explains the bottleneck's current failures, with candidates restricted to regions of teacher/student disagreement. Trained with nested dropout over this selection order, the final bottleneck predicts accurately from any concept prefix, so inference spends fewer concepts on inputs it is confident about early and more on hard ones. Since no stage is modality-specific, a new domain and task requires swapping only the backbone and concept pool. Across six vision and two text datasets and two teacher paradigms, AnyBottle yields bottlenecks with fewer concepts and higher concept consistency than annotation-free baselines, while staying close to the black-box reference. Overall, AnyBottle shows that going annotation-free need not mean going large: a small, discovered vocabulary can be as expressive as a much larger, fixed one.

Updated: 2026-10-06 15:40:06

标题: AnyBottle:仅保留您真正需要的概念的秘诀

摘要: 概念瓶颈模型(CBMs)通过人类可解释的概念路由使预测变得可检查和可干预,但最初需要概念注释。无需注释的变体消除了这一要求,但通常使用大型概念词汇表,在训练和推断过程中均保持静态,产生比任何任务或预测所需更大且更难检查的瓶颈。我们提出了AnyBottle,这是一种构建紧凑的、特定任务的CBMs的单一方法。 AnyBottle仅假设一个冻结的主干和一个无监督的概念池,例如稀疏自动编码器。然后,在相同主干上训练的黑盒教师指导选择:每一轮都会添加最能解释当前瓶颈失败的概念,候选概念受限于教师/学生分歧的区域。通过在这个选择顺序上进行嵌套丢弃训练,最终瓶颈可以准确地从任何概念前缀预测,因此推断会在对早期自信的输入上花费较少的概念,而在困难的输入上花费更多的概念。由于没有任何阶段是特定于模态的,新的领域和任务仅需要交换主干和概念池。在六个视觉和两个文本数据集以及两种教师范式中,AnyBottle产生比无注释基线更少的概念和更高的概念一致性的瓶颈,同时与黑盒参考保持接近。总的来说,AnyBottle表明无需注释不一定意味着变得更大:一个小的、发现的词汇表可以和一个更大、固定的词汇表一样具有表达能力。

更新时间: 2026-10-06 15:40:06

领域: cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08552v1

FRAGMENTA: Efficient End-to-end Fragmentation-based Generative Model with Agentic Tuning for Drug Lead Optimization in Small Data Regime

Molecule generation from extremely limited training data is a key challenge in drug discovery. Existing fragment-based methods are more suitable than atom-based approaches in this regime, but typically optimize fragment selection separately from downstream generation. Expert feedback is also especially valuable with limited data, yet translating such feedback into model objectives usually requires AI engineering expertise. We introduce FRAGMENTA, an end-to-end framework for small-data drug lead optimization with two components: (1) LVSEF, a fragment-based generator that jointly optimizes fragmentation and generation through a tabular reward-update mechanism, and (2) an agentic system that converts conversational expert feedback into updated generative objectives. Across three small-data datasets (11--104 molecules), LVSEF outperforms state-of-the-art methods in the smallest-data settings, matches them at larger scales, and trains ${\sim}16\times$ faster. On three public protein targets, iterative closed-loop optimization improves final-round discovery yield by up to ${\sim}16%$ over one-shot LVSEF-only on kinase, with gains depending on how well feedback matches target chemistry. In a real-world cancer drug-discovery deployment, Human-Agent FRAGMENTA identified nearly twice as many molecules with favorable docking scores ($< -6$) as baseline methods.

Updated: 2026-10-06 15:35:51

标题: Fragmenta:高效的端到端基于片段的生成模型,通过代理调整在小数据环境中进行药物引导优化

摘要: 极度有限的训练数据中的分子生成是药物发现中的一个关键挑战。现有的基于片段的方法比基于原子的方法更适用于这种情况,但通常会将片段选择与下游生成分开优化。专家反馈在有限数据情况下尤其有价值,然而将这种反馈转化为模型目标通常需要AI工程专业知识。我们引入了FRAGMENTA,一个用于小数据药物前导优化的端到端框架,包括两个组件:(1) LVSEF,一个基于片段的生成器,通过表格奖励更新机制共同优化片段化和生成,以及(2)一个主体系统,将专家对话反馈转化为更新的生成目标。在三个小数据集(11-104个分子)上,LVSEF在最小数据设置中胜过了最先进的方法,在较大规模上与它们相匹配,并且训练速度比它们快16倍。在三个公共蛋白靶标上,迭代闭环优化使最终发现产量比仅使用一次LVSEF提高了约16%,在激酶上的增益取决于反馈与靶标化学特性的匹配程度。在一个真实的癌症药物发现部署中,Human-Agent FRAGMENTA鉴定出的具有有利的对接分数(< -6)的分子数量几乎是基线方法的两倍。

更新时间: 2026-10-06 15:35:51

领域: cs.AI

下载: http://arxiv.org/abs/2511.20510v3

Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers

Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves. We present Regime-Conditional Verification (RCV), a lightweight wrapper that adapts an off-the-shelf safety classifier without retraining it. RCV estimates, from the classifier's internal representations, the probability that each prediction disagrees with the deployer's policy, and selectively corrects predictions likely to be wrong. The same correctness estimates also provide a label-free signal for detecting distribution shift, enabling a maintenance loop that updates the correctness estimation layer and resorts to classifier fine-tuning only when repair fails within a label budget. Across three off-the-shelf safety classifiers and two benchmark datasets, RCV improves adherence to the deployer's policy in every classifier-dataset combination, catching up to 0.81 of previously missed unsafe content without modifying the underlying classifier. In a deployment study with ten attack campaigns, each a harm category held out of RCV's training, RCV detects every campaign in a dedicated injection panel; in the maintenance census most drift episodes are repaired without updating the classifier, and the fine-tune is reserved for the residual episodes.

Updated: 2026-10-06 15:35:38

标题: Regime-Conditional Verification: 适用于调整和监测安全分类器的正确性估计

摘要: 大型语言模型部署的安全分类器通常由于两个原因而失败:它们的决策反映了训练期间学习到的策略,而不是部署者所期望的策略,并且它们的性能会随着部署流量的变化而降低。我们提出了Regime-Conditional Verification(RCV),这是一个轻量级包装器,可以在不重新训练的情况下调整现成的安全分类器。RCV从分类器的内部表示中估计每个预测与部署者策略不一致的概率,并有选择地修正可能错误的预测。相同的正确性估计还提供了一个无标签信号,用于检测分布偏移,从而实现了一个维护循环,更新正确性估计层,并仅在标签预算内修复失败时才进行分类器微调。在三个现成的安全分类器和两个基准数据集上,RCV在每个分类器-数据集组合中都提高了对部署者策略的遵守,最多可捕捉到先前被错过的不安全内容的0.81,而无需修改底层分类器。在一个部署研究中,有十个攻击活动,每个活动都属于RCV训练中排除的伤害类别,RCV在一个专用注入面板中检测到每个活动;在维护普查中,大多数漂移事件都可以修复而无需更新分类器,微调仅用于残留事件。

更新时间: 2026-10-06 15:35:38

领域: cs.AI,cs.CL,cs.CR,cs.LG

下载: http://arxiv.org/abs/2608.14089v3

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?

Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable, recent approaches have begun imposing hierarchical structure, either explicitly or as an implicit effect of training constraints, yet rigorous comparison remains difficult. There are no agreed-upon requirements for what a meaningful feature hierarchy should satisfy, and evaluation has largely relied on qualitative illustrations with fragmented quantitative protocols. To address this, we derive a set of key requirements for generalization/specialization hierarchies in unsupervised concept discovery, drawing on semantic net and taxonomy research alongside recent SAE work, and use them to derive a concrete evaluation protocol. Applying this protocol to current SAE approaches trained on visual data, we find that while feature spaces generally provide a basis for sensible hierarchies, establishing good hierarchical structure remains challenging. In particular, feature absorption, both in its well-known hard form and in a continuous, soft form, systematically compromises hierarchy quality, pointing to a fundamental tension that future approaches will need to navigate.

Updated: 2026-10-06 15:34:38

标题: 稀疏自编码器是否学习了有意义的概念层次结构?

摘要: 稀疏自编码器(SAEs)已经成为大型模型中无监督概念发现的重要工具。为了使产生的特征空间更易于解释和管理,最近的方法已经开始施加层次结构,无论是显式地还是作为训练约束的隐式效果,但严格的比较仍然困难。对于有意义的特征层次结构应该满足什么要求尚无共识,评估主要依赖于定性插图和片段化的定量协议。为了解决这个问题,我们从语义网络和分类学研究以及最近的SAE工作中汲取了一组关键要求,用于无监督概念发现中的泛化/特化层次结构,并利用它们推导出一个具体的评估协议。将这一协议应用于当前在视觉数据上训练的SAE方法时,我们发现虽然特征空间通常提供了合理层次结构的基础,但建立良好的层次结构仍然具有挑战性。特别是,特征吸收,无论是其众所周知的硬形式还是连续的软形式,系统地损害了层次结构的质量,指向了未来方法需要解决的根本张力。

更新时间: 2026-10-06 15:34:38

领域: cs.LG

下载: http://arxiv.org/abs/2606.22994v2

Neural Global Optimization via Iterative Refinement from Noisy Samples

Global optimization of black-box functions from noisy samples is a fundamental challenge in machine learning and scientific computing. Traditional methods such as Bayesian Optimization often converge to local minima on multi-modal functions, while gradient-free methods require many function evaluations. We present a novel neural approach that learns to find global minima through iterative refinement. Our model takes noisy function samples and their fitted spline representation as input, then iteratively refines an initial guess toward the true global minimum. Trained on randomly generated functions with ground truth global minima obtained via exhaustive search, our method achieves a mean error of 8.05 percent on challenging multi-modal test functions, compared to 36.24 percent for the spline initialization, a 28.18 percent improvement. The model successfully finds global minima in 72 percent of test cases with error below 10 percent, demonstrating learned optimization principles rather than mere curve fitting. Our architecture combines encoding of multiple modalities including function values, derivatives, and spline coefficients with iterative position updates, enabling robust global optimization without requiring derivative information or multiple restarts.

Updated: 2026-10-06 15:33:54

标题: 通过从有噪声样本中迭代细化的神经全局优化

摘要: 嘈杂样本中黑盒函数的全局优化是机器学习和科学计算中的一个基本挑战。传统方法如贝叶斯优化常常在多峰函数上收敛到局部最小值,而无梯度方法则需要大量函数评估。我们提出了一种新颖的神经方法,通过迭代细化学习找到全局最小值。我们的模型以噪声函数样本及其拟合的样条表示作为输入,然后迭代地将初始猜测细化到真正的全局最小值。通过在随机生成的函数上进行训练,通过穷尽搜索获得地面实际全局最小值,我们的方法在具有挑战性的多模态测试函数上实现了8.05%的平均误差,而样条初始化为36.24%,提高了28.18%。该模型成功在72%的测试案例中找到误差在10%以下的全局最小值,展示了学习优化原则而不仅仅是曲线拟合。我们的架构结合了多种模态的编码,包括函数值、导数和样条系数,以及迭代位置更新,实现了稳健的全局优化,无需导数信息或多次重新启动。

更新时间: 2026-10-06 15:33:54

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2604.03614v3

How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it. We test this on transformers trained for in-context meta-analysis, which must infer the hidden heterogeneity between studies to weight them correctly, and where both reference points are known. Under distribution shift, probe scores fall and prediction error rises $12$--$15\times$, yet the model recovers a similar share of the headroom, indicating that the data lost information, not the representation. We then analyze the real models. The single-cell foundation model scGPT encodes biological variability only partially. We also revisit four influential LLM probing studies, which claim that models represent geography, the state of an Othello board, truth, and the demographics of their users. Against a floor computed from the input text alone, some of these claims hold, while others are largely explained by the text itself.

Updated: 2026-10-06 15:32:25

标题: 0.6是多高?楼层、天花板和可解释性探测中的头部空间

摘要: 探针是可解释性的重要工具。如果一个模型的隐藏状态可以预测一个变量,那么就说这个模型代表了这个变量。但是探针得分并没有固定的含义。一个0.6的$R^2$可能只是反映了输入本身已经包含的信息,同样的得分在不同的数据上可能意味着不同的事情。我们建议将每个探针得分与两个参考点进行比较:一个是底线,即一个声明的简单输入集合已经可以预测的内容,另一个是天花板,即完整输入可以预测的内容。它们之间的差距,即潜力空间,是一个探针可以展示模型计算出简单输入之外的内容的范围。我们证明潜力空间会以两种方式消失:目标不再依赖模型必须推断的隐藏变量,或者输入不再揭示这些变量。我们在为上下文元分析训练的transformers上进行了测试,这些transformers必须推断不同研究之间的隐藏异质性以正确加权,而且两个参考点都是已知的。在分布转移的情况下,探针得分下降,预测误差增加12-15倍,然而模型恢复了相似比例的潜力空间,表明数据丢失了信息,而不是表示。然后我们对真实模型进行了分析。单细胞基础模型scGPT只部分编码了生物变异性。我们还重新审视了四个有影响力的LLM探针研究,这些研究声称模型可以代表地理位置、奥赛罗棋盘的状态、真相以及用户的人口统计信息。与仅从输入文本计算的底线相比,其中一些声明是正确的,而其他一些主要是由文本本身解释的。

更新时间: 2026-10-06 15:32:25

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.08544v1

Action Shaping: Policies Absorb What They Can Express

Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its minimal instance is a zero-initialized linear head behind a learnable gate, added to an actor that trains through a learned action-value function, with no penalty or schedule. The gate rises and then falls on its own, for deterministic and stochastic actors alike, and on 20 tasks removing the head costs almost nothing. The condition is exact reproduction, not capacity: a nonlinear head with more parameters is not absorbed, and in a paired control, one linear path added to a nonlinear base head restores absorption. Exact reproduction gives the loss a flat direction that gradient noise drifts along, and the offset's amplitude indicates, before removal, what dropping the head will cost. Action shaping thus gains the counterpart of the shaping theorem, a condition for absorption, together with the mechanism behind it and a diagnostic that reads it. Policies absorb what they can express, and only that.

Updated: 2026-10-06 15:32:07

标题: 行动塑造:政策吸收能够表达的内容

摘要: 奖励塑造有一个定理:基于潜力的项可以被移除而不改变最优策略。在动作通道上采用相同的做法,即在训练中添加一个偏移量,然后在部署时删除,却没有定理支持。没有任何东西可以取消动作的偏移量,因此在部署时保留或删除校正都没有保证。我们称之为动作塑造并说明其原则。可训练的策略吸收一个其自身输出层完全可以重现的偏移量,这就是我们所说的表达;被吸收的内容可以在保留回报的情况下被移除。其最小实例是一个零初始化的线性头部,位于一个可学习的门后面,添加到一个通过学习的动作值函数训练的执行者,没有惩罚或计划。该门自行上升然后下降,对确定性和随机性执行者都是如此,在20个任务中移除头部几乎没有任何成本。条件是精确重现,而不是容量:一个具有更多参数的非线性头部不会被吸收,在一个配对的控制中,一个线性路径添加到一个非线性基础头部可以恢复吸收。精确重现给损失带来了一个平坦的方向,梯度噪声沿着这个方向漂移,偏移的振幅指示,在移除之前,去掉头部将会产生什么成本。因此,动作塑造获得了塑造定理的对应物,一个吸收条件,以及背后的机制和一个诊断,可以读取它。策略吸收它们可以表达的内容,只有这些内容。

更新时间: 2026-10-06 15:32:07

领域: cs.AI,cs.LG

下载: http://arxiv.org/abs/2609.32752v2

Micro Neural Policies for Safe Real-Time Robotic Control

In this paper, we investigate the synthesis of Micro Neural Policies (MNP) to enable safe and robust real-time robotic control on computationally constrained embedded devices. We demonstrate that integrating Evolution Strategy (ES) and Statistical Model Checking (SMC)-based verification for policy search can drastically reduce neural network size without compromising safety and robustness. We conduct a large-scale training and evaluation of MNP on Cartpole and Quadrotor control tasks, varying control frequencies and network architectures. After validating these policies in simulation, we evaluate their deployability through zero-shot transfer to physical systems. Our experiments show that MNP can successfully achieve safe sim-to-real transfer without sacrificing control performance. We then show that the policies' memory footprint, ranging from 0.5 to 7.5 kB, allows deployment on microcontrollers, where they achieve real-time inference latency with under 25 ns of jitter while leaving the chip idle for over 97% of the time for additional workloads. This makes them a highly practical solution for severely resource-constrained robotic systems.

Updated: 2026-10-06 15:29:28

标题: 微型神经策略用于安全实时机器人控制

摘要: 在这篇论文中,我们研究了微型神经政策(MNP)的合成,以实现在计算受限的嵌入式设备上进行安全和稳健的实时机器人控制。我们证明,将进化策略(ES)和基于统计模型检验(SMC)的政策搜索相结合可以大幅减少神经网络的大小,而不会影响安全性和稳健性。我们对Cartpole和Quadrotor控制任务上的MNP进行了大规模训练和评估,变化控制频率和网络架构。在模拟中验证这些政策后,我们通过零次转移到物理系统来评估它们的可部署性。我们的实验表明,MNP可以成功实现安全的模拟到实际的转移,而不会牺牲控制性能。然后我们表明,这些政策的存储空间占用范围从0.5到7.5 kB,可以部署在微控制器上,在保持芯片空闲超过97%的时间进行额外工作负载的情况下实现实时推理延迟,并且推理延迟不超过25 ns。这使它们成为严重资源受限的机器人系统的高度实用解决方案。

更新时间: 2026-10-06 15:29:28

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2610.08541v1

Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements

Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property. We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as B_r(N)=a_rN^alpha_r, with N a capability proxy; against a budget proportional to N, scaling helps if alpha_r<1, keeps pace if alpha_r~1, and accumulates alignment debt if alpha_r>1. We give three operationalizations of burden and distinguish observed, audited and true alignment. A toy model, in which corrections consume capability headroom, makes the consequences explicit. We prove that the largest exponent among corrected risks, not an average, sets the long-run regime; that above 1 any policy holding headroom above a floor must grow super-exponentially; that, for burdens that are positive mixtures of power laws, fits on small models underestimate large-scale exponents; and that an audit that uncovers hidden failures without false positives never underestimates true alignment. We propose a pre-registrable protocol and apply reduced versions of it twice. A preregistered reanalysis of public adversarial-training data for Pythia classifiers finds that the compute needed to bring attack success under 10% grows as N^0.60. A preregistered pilot on Qwen2.5 0.5B-72B finds exponents of -0.05 for truthfulness and 0.48 for stated dispositions (both scaling helps under its reduced rule, though local slopes approach 1 at the top; replicated on Qwen3 0.6B-14B), while sycophancy (0.89, or 0.83 with two seeds added at 72B) and a planted backdoor are undetermined: the backdoor is removed quickly when its trigger is known but survives blind safety training at four of five sizes. We release four browser games that play these laws (www.aisafety.fun). We make no claim about which regime holds for current frontier models.

Updated: 2026-10-06 15:29:07

标题: 朝向对齐缩放规律:一个框架和首次预注册测量

摘要: 随着模型的增长,对齐是否变得更容易或更困难经常从孤立的发现中进行争论,好像对齐是一个属性。我们将其视为一系列可测量的缩放关系:对于每个风险类别r,为了实现固定的安全目标所需的对齐负担被建模为B_r(N)=a_rN^alpha_r,其中N是能力代理;相对于与N成比例的预算,如果alpha_r<1,则缩放有所帮助,如果alpha_r~1,则保持步伐,如果alpha_r>1,则积累对齐债务。我们给出了三种负担的操作化形式,并区分了观察到的、审计过的和真实的对齐。一个玩具模型,其中纠正消耗能力余地,使后果变得明显。我们证明,被纠正的风险中最大的指数,而不是平均值,确定了长期的制度;在大于1时,任何保持高于底线的余地的政策必须超指数增长;对于正的混合幂律负担,对小型模型的拟合低估了大规模指数;一个揭示隐藏故障而没有误报的审计永远不会低估真实对齐。我们提出了一个可预注册的协议,并将其简化版本应用了两次。对Pythia分类器的公开对抗训练数据进行的预注册再分析发现,使攻击成功率降至10%以下所需的计算量随N^0.60增长。在Qwen2.5 0.5B-72B上进行的预注册试点发现,真实性的指数为-0.05,陈述的倾向性为0.48(在其简化规则下,两者的缩放有所帮助,尽管在顶部局部斜率接近1;在Qwen3 0.6B-14B上复制),而奉承(0.89,或在72B添加两个种子后为0.83)和一个植入后门是未确定的:当知道其触发器时,后门很快被移除,但在四种大小的盲目安全训练中仍存活。我们发布了四款可以遵循这些规则的浏览器游戏(www.aisafety.fun)。我们不对当前前沿模型适用哪种制度提出任何主张。

更新时间: 2026-10-06 15:29:07

领域: cs.AI,cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.08540v1

From Shared Demand Patterns to Local Uncertainty: Probabilistic Load Forecasting by Mixing Compact Adaptations

Probabilistic load forecasting has been widely studied for power-system operation and planning, but customer- and transformer-level forecasting introduces a distinct scalability challenge. At these levels, load uncertainty is strongly affected by customer behavior, weather, and mixed load composition, making it difficult for a single shared model to capture heterogeneous patterns. Using separate probabilistic models can improve local accuracy, but becomes costly to train, store, update, and validate at scale. To address this challenge, we develop a scalable customer-aware forecasting framework that learns common demand behavior through a shared model while adapting only a compact subset of parameters. Rather than using an independent model for each load or assigning each load to a specialized model, the proposed design learns a small bank of low-dimensional adaptation components and allows each load to combine them according to its forecasting characteristics. This preserves shared knowledge across customers while providing sufficient flexibility for heterogeneous and mixed load compositions. Experiments on 590 load profiles from the SMART-DS dataset show consistent improvements in deterministic accuracy and probabilistic quality over statistical, neural-network, Transformer-based, and pretrained time-series baselines, while retaining low storage and inference costs.

Updated: 2026-10-06 15:27:43

标题: 从共享需求模式到本地不确定性:通过混合紧凑适应进行概率负荷预测

摘要: 概率负荷预测已被广泛研究用于电力系统的运营和规划,但客户和变压器级别的预测引入了一个独特的可伸缩性挑战。在这些级别上,负载不确定性受到客户行为、天气和混合负载组成的强烈影响,这使得一个共享模型很难捕捉到异构模式。使用独立的概率模型可以提高本地精度,但在规模化的情况下训练、存储、更新和验证成本较高。为了解决这一挑战,我们开发了一个可扩展的客户感知预测框架,通过一个共享模型学习常见的需求行为,同时仅调整一个紧凑的参数子集。与为每个负载使用独立模型或将每个负载分配给专门的模型不同,所提出的设计学习了一小组低维适应组件,并允许每个负载根据其预测特性组合它们。这保留了客户之间的共享知识,同时为异构和混合负载组合提供了足够的灵活性。对来自SMART-DS数据集的590个负载配置文件的实验显示,相对于统计、神经网络、基于Transformer的和预先训练的时间序列基线模型,确定性精度和概率质量有一致的提高,同时保持低存储和推理成本。

更新时间: 2026-10-06 15:27:43

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.08538v1

The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empirically selected via cross-validated performance rather than fixed a priori, (ii) locating a relatively small set of latent support vectors (~ 29% of total training examples) representing influential prompts for identifying tokens that alter the classifier's predicted labels, and (iii) utilizing such tokens and their associated attack magnitudes for constructing a diagnostic taxonomy. This diagnostic taxonomy provides an end-to-end guideline for flagging prompts that require different treatments: rely Safely on the classifier's decision; flag Heuristic Bias and Heuristic Override cases; route Insufficient Context cases for further human/safety review. Applying the framework to a classifier trained on a public prompt injection dataset, we find that a substantial fraction of its confident decisions (~ 77%) are not robust to removing a single token, and that this brittleness separates into two distinct failure patterns: a confidence calibration failure and a genuinely exploitable shortcut. For each zone of the taxonomy, we also recommend strategies for remediating diagnosed prompts. We illustrate the framework as a series of steps, demonstrating how each step operates.

Updated: 2026-10-06 15:27:36

标题: 潜在的诊断分类法:构建分类器和诊断它们的决策框架,应用于快速注入检测

摘要: 这篇论文提出了一个构建分类器作为安全层的框架,并开发了一个补充诊断,用于识别哪些分类器的自信决策是可信的。这个框架,称为潜在诊断分类法,包括:(i)构建一个维度优化的分类器,其中嵌入维度是经验性地通过交叉验证性能而不是事先固定选择的,(ii)找到一个相对较小的潜在支持向量集合(约占总训练示例的29%),代表影响标记变化的关键提示,(iii)利用这些关键提示及其相关的攻击强度构建一个诊断分类法。这个诊断分类法提供了一个端到端的指导方针,用于标记需要不同处理的提示:可依赖于分类器的决定;标记启发式偏见和启发式覆盖案例;将上下文不足的案例路由至进一步的人工/安全审查。将这个框架应用于一个针对公共提示注入数据集训练的分类器时,我们发现其中相当大比例的自信决策(约77%)在去掉一个标记后并不稳健,并且这种脆弱性分为两种不同的失败模式:信心校准失败和真正可利用的捷径。对于分类法的每个区域,我们还推荐了诊断提示的纠正策略。我们以一系列步骤的形式说明了这个框架,展示了每个步骤的操作方式。

更新时间: 2026-10-06 15:27:36

领域: cs.LG,cs.AI,cs.CL,cs.CR

下载: http://arxiv.org/abs/2608.26423v2

FlowCF: Sparse Counterfactual Explanations for Mixed-Type Tabular Data using Flow Matching

In the field of Explainable AI (XAI), counterfactual (CF) explanations interpret a model's decision by suggesting the changes to the input that would lead to a more favourable outcome. To be useful in practice, such an explanation should change few features and change them as little as possible, properties known as sparsity and proximity. We observe that existing methods remain limited in this respect, especially for numerical features, whether they are model-agnostic and amortised, or gradient-based with full access to the model. In this paper, we propose FlowCF, a model-agnostic generative method that frames CF generation as sparse transport from the factual to the target class. We solve this transport with flow matching, which we extend to mixed feature types with a novel mixed flow operator, and exploit the resulting geometry to optimise for sparsity through a gating network that minimises the number of features the transport changes. Extensive experiments on six benchmark datasets demonstrate that FlowCF produces the best numerical sparsity and proximity, changing 29% of the numerical features where the best baseline changes 89%, at 70% smaller displacement, while remaining comparable on the other desiderata.

Updated: 2026-10-06 15:27:23

标题: FlowCF:使用流匹配为混合类型表格数据生成稀疏对照解释

摘要: 在可解释人工智能(XAI)领域,反事实(CF)解释通过建议改变输入来解释模型的决策,从而导致更有利的结果。为了在实践中有用,这种解释应该改变少量特征,并尽可能少地改变它们,这些特性被称为稀疏性和接近性。我们观察到现有方法在这方面仍然存在局限性,特别是对于数值特征,无论它们是模型不可知和摊销的,还是基于梯度并完全访问模型。在本文中,我们提出了FlowCF,一种模型不可知的生成方法,将CF生成看作是从事实到目标类的稀疏传输。我们通过流匹配来解决这种传输问题,我们通过一种新颖的混合流运算符将其扩展到混合特征类型,并利用生成的几何形状通过一个门控网络来优化稀疏性,从而最小化传输改变的特征数量。对六个基准数据集的广泛实验表明,FlowCF产生了最佳的数值稀疏性和接近性,在最佳基线改变了89%的数值特征的情况下改变了29%,且位移量减少了70%,同时在其他方面保持可比性。

更新时间: 2026-10-06 15:27:23

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.08537v1

BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration

Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. However, existing methods often require an additional draft model or weight representation, introducing non-negligible memory overhead on resource-constrained devices. Self-speculative approaches reduce this overhead, yet still face trade-offs between draft quality, target quality, and storage efficiency. We propose BitNest, a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation. Instead of deriving a draft from a predefined target, BitNest first constructs a strong low-precision base and then recovers the higher-precision target through residual refinement, enabling both models to share a single physical weight representation. BitNest further extends this progressive-precision design to the KV cache for long-context inference. Across multiple 7B--8B edge-friendly LLMs and diverse workloads, BitNest achieves an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and delivers 1.48--1.61x end-to-end speedup over FP16 autoregressive decoding. On the LLaMA models supported by all representative self-speculative baselines, BitNest also achieves consistently competitive or higher decoding speedup.

Updated: 2026-10-06 15:26:50

标题: BitNest:基于比特嵌套的推测解码,用于内存高效的LLM推理加速

摘要: Speculative decoding通过使用轻量级草案提出多个令牌进行并行验证,加速自回归生成。然而,现有方法通常需要额外的草案模型或权重表示,在资源受限设备上引入了不可忽略的内存开销。自我推测方法减少了这种开销,但仍然要在草案质量、目标质量和存储效率之间权衡。我们提出了BitNest,一个比特嵌套的推测解码框架,将低精度草案直接嵌入到高精度目标表示中。BitNest首先构建一个强大的低精度基础,然后通过残差细化恢复高精度目标,使两个模型共享单一的物理权重表示。BitNest进一步将这种逐渐精度设计扩展到长上下文推理的KV缓存。在多个7B--8B边缘友好的LLMs和不同工作负载中,BitNest实现了平均95.2%的推测接受率,同时保持较高精度模型质量,并在FP16自回归解码上实现了1.48--1.61倍的端到端加速。在所有代表性自我推测基线支持的LLaMA模型上,BitNest还实现了一致竞争力或更高的解码加速。

更新时间: 2026-10-06 15:26:50

领域: cs.AI

下载: http://arxiv.org/abs/2610.02800v2

How Bregman Divergences Shape Shampoo

Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.

Updated: 2026-10-06 15:25:57

标题: 如何Bregman散度塑造洗发水

摘要: 最近,对洗发水背后原理的理解指导了更有效的神经网络优化器的开发。这些方法通过优化Frobenius或Kullback-Leibler(KL)散度来学习一个预处理器。在这项工作中,我们研究了散度选择如何塑造预处理,这一点仍然不清楚,并阻碍了进一步的改进。为此,我们开发了一个统一的Bregman散度框架,连接所有流行的散度,使我们能够共同研究它们。通过梯度二阶矩的经验谱分析,我们研究了散度选择如何塑造Kronecker逼近,并与预处理中的有限样本误差相互作用。我们发现一些散度可以更好地补偿对经验二阶矩的有限样本低估,有助于解释其对应Shampoo变种的不同行为。我们通过GPT-2预训练实验证实了这一解释。通过将散度选择与实际训练行为相连接,我们相信我们的框架为理解和进一步改进Shampoo提供了基本指导。

更新时间: 2026-10-06 15:25:57

领域: cs.LG

下载: http://arxiv.org/abs/2610.08534v1

Beyond Perturbation Magnitude: Direction-Dependent Responses in Multimodal Geometric Representations

Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.

Updated: 2026-10-06 15:25:47

标题: 超越扰动幅度:多模几何表示中的方向依赖性响应

摘要: 基于Gram行列式的几何对齐得分提供了一种紧凑的建模多模态之间高阶一致性的方式,然而这种得分如何对模态退化作出响应尚不清楚。本文探讨了多模态几何得分的响应是否主要由扰动引起的位移大小决定。我们利用来自MSR-VTT(N=878)和DiDeMo(N=980)的固定队列,应用控制的视频模糊和音频噪声,并分析得分定义的关系几何中的响应。位移大小最多解释绝对响应的样本外方差的15%,并且匹配大小的对响应有系统性差异,因此标量大小不能组织响应。Gramian体积的封闭形式一阶展开产生方向几何响应(DGR):位移在局部体积梯度上的投影,联合捕捉干净操作点、位移大小和位移方向。绝对一阶DGR项解释了观察到的响应,样本外R^2为0.838-0.969,匹配大小排序准确率为0.864-0.963,响应符号准确率为0.909-0.989,而经过测试的无方向替代在相应的评估协议下仍然较弱或不稳定。预先指定的增益归一化候选人V/(g_V+eps)未能预测其干净顺序门。DGR使用观察到的退化状态位移,因此是一种解释性量,而不是部署时间预测器:几何响应取决于表示操作的位置,在哪里退化移动关系几何以及移动方向。

更新时间: 2026-10-06 15:25:47

领域: cs.CV,cs.LG,cs.SD

下载: http://arxiv.org/abs/2610.08533v1

MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis

Clinical diagnosis is inherently a structured reasoning process, yet existing deep learning models often bypass this structure by mapping image features directly to disease labels without explicitly interrogating the morphological and textural criteria that clinicians systematically evaluate. This limits diagnostic transparency and may compromise safe clinical deployment. We present MedCORE (Medical Criteria-Oriented Reasoning and Evidence), a structured diagnostic framework that operationalizes clinical reasoning within a vision-language architecture. For each input image, MedCORE decomposes the diagnostic process into clinically defined criteria, spatially localizes each criterion to diagnostically relevant image regions, encodes evidence through multi-scale representations that capture macro-structural and micro-textural pathological characteristics, and refines criterion representations using a Graph Attention Network that explicitly models inter-criteria dependencies. Criterion representations are further aligned with clinical text descriptors, reinforced through class-wise visual prototypes, and aggregated using uncertainty-calibrated weighting that proportionally discounts low-confidence diagnostic evidence. MedCORE is validated across three clinically heterogeneous imaging modalities, including dermoscopic lesion classification on ISIC 2018, breast ultrasound lesion characterization on BUSI, and diabetic retinopathy grading on IDRiD. Quantitatively, MedCORE achieves 89.2% accuracy, 85.7% macro-F1, and 96.4% AUC on ISIC 2018; 96.1% accuracy, 95.2% macro-F1, and 98.4% AUC on BUSI; and 84.3% accuracy, 80.2% macro-F1, and 92.8% AUC on IDRiD. These results demonstrate consistent improvements over strong CNN, transformer, biomedical vision-language, concept-based, and prototype-based baselines.

Updated: 2026-10-06 15:23:01

标题: MedCORE: 基于标准的可解释医学图像诊断临床推理

摘要: 临床诊断本质上是一个结构化的推理过程,然而现有的深度学习模型往往通过将图像特征直接映射到疾病标签而绕过这种结构,而不是明确审查临床医生系统地评估的形态和纹理标准。这限制了诊断的透明度,并可能损害安全的临床应用。我们提出了MedCORE(医学标准导向的推理和证据),这是一个结构化的诊断框架,将临床推理在视觉-语言架构中操作化。对于每个输入图像,MedCORE将诊断过程分解为临床定义的标准,将每个标准局部化到诊断相关的图像区域,通过捕捉宏观结构和微观纹理病理特征的多尺度表示来编码证据,并使用明确建模互相依赖的Graph Attention Network来细化标准表示。标准表示进一步与临床文本描述符对齐,通过类别-wise视觉原型加强,并使用校准不确定性权重进行聚合,比例折扣低置信度的诊断证据。MedCORE在三种临床异质成像模态上进行了验证,包括ISIC 2018上的皮肤镜病变分类,BUSI上的乳腺超声病变表征,以及IDRiD上的糖尿病视网膜病变分级。定量上,MedCORE在ISIC 2018上达到89.2%的准确率,85.7%的宏F1分数,和96.4%的AUC;在BUSI上达到96.1%的准确率,95.2%的宏F1分数,和98.4%的AUC;在IDRiD上达到84.3%的准确率,80.2%的宏F1分数,和92.8%的AUC。这些结果表明相对于强CNN、变压器、生物医学视觉-语言、基于概念和基于原型的基线方法,MedCORE实现了一致的改进。

更新时间: 2026-10-06 15:23:01

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.08528v1

PHBA: Prefix-State Hybrid Block Attention

Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks, allowing the model to combine precise long-range evidence with compressed historical context within a unified layer. We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors. Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.

Updated: 2026-10-06 15:22:49

标题: PHBA: 前缀状态混合区块注意力

摘要: 混合架构将线性序列模型与softmax注意力相结合,提供了在高效长上下文建模和精确标记检索之间的有效平衡。现有设计,如本地混合注意力(NHA),将压缩的长期状态与滑动窗口注意力结合在一起,但其确切的关注点被限制在一个固定的局部窗口内。在这项工作中,我们介绍了Prefix-State Hybrid Block Attention(PHBA),它用top-k块稀疏检索替换了本地滑动窗口注意力,并将每个检索到的块与总结其前文环境的紧凑前缀状态相结合。前缀状态是通过在块边界处进行门控线性递归构建的,并与相应的标记块一起检索,使模型能够在统一层内将精确的长距离证据与压缩的历史背景结合起来。我们进一步开发了一个硬件感知的Triton实现,可以在不产生大量中间张量的情况下流式传输路由的标记块和前缀状态。实验证明,PHBA在保持高效训练和推断的同时,提高了长上下文和检索性能,超过了强线性和混合基线。

更新时间: 2026-10-06 15:22:49

领域: cs.LG

下载: http://arxiv.org/abs/2610.08527v1

Large Pretraining Datasets Don't Guarantee Robustness after Fine-Tuning in Image Classification

Large-scale pretrained models are widely leveraged as foundations for learning new specialized tasks via fine-tuning, with the goal of maintaining the general performance of the model while allowing it to gain new skills. A valuable goal for all such models is robustness: the ability to perform well on out-of-distribution (OOD) tasks. We assess whether fine-tuning preserves the overall robustness of the pretrained model in image classification, and observed that models pretrained on large datasets exhibited strong catastrophic forgetting and loss of OOD generalization. To systematically assess robustness preservation in fine-tuned models, we propose the Robustness Inheritance Benchmark (ImageNet-RIB). The benchmark, which can be applied to any pretrained model, consists of a set of related but distinct OOD (downstream) tasks and involves fine-tuning on one of the OOD tasks in the set then testing on the rest. We find that though continual learning methods help, fine-tuning reduces robustness across pretrained models. Surprisingly, models pretrained on the largest and most diverse datasets (e.g., LAION-2B) exhibit both larger robustness losses and lower absolute robustness after fine-tuning on small datasets, relative to models pretrained on smaller datasets. We observe this collapse in contrastively pretrained (CLIP) models and their fine-tuned variants, where it grows with pretraining scale; the supervised models we test do not exhibit it. These findings suggest that starting with the strongest foundation model is not necessarily the best approach for performance on specialist tasks. https://jd730.github.io/projects/ImageNet-RIB

Updated: 2026-10-06 15:20:27

标题: 大型预训练数据集在图像分类微调后并不保证稳健性

摘要: 大规模预训练模型被广泛用作学习新专业任务的基础,通过微调来实现这一目标,旨在保持模型的整体性能,同时使其获得新技能。所有这类模型的一个宝贵目标是鲁棒性:在分布外(OOD)任务上表现良好的能力。我们评估了微调是否保持了在图像分类中预训练模型的整体鲁棒性,并观察到在大型数据集上预训练的模型表现出强大的灾难性遗忘和OOD泛化的损失。为了系统评估微调模型中的鲁棒性保持,我们提出了鲁棒性继承基准(ImageNet-RIB)。这个基准可以应用于任何预训练模型,由一组相关但不同的OOD(下游)任务组成,涉及在该组中的一个OOD任务上进行微调,然后在其余任务上进行测试。我们发现,虽然持续学习方法有所帮助,但微调会降低预训练模型的鲁棒性。令人惊讶的是,在最大和最多样化的数据集上预训练的模型(例如LAION-2B)在微调小数据集后,鲁棒性损失更大,绝对鲁棒性更低,相对于在较小数据集上预训练的模型。我们观察到这种对比预训练(CLIP)模型及其微调变体中的崩溃,其中随着预训练规模的增长而增加;我们测试的监督模型没有表现出这种情况。这些发现表明,从最强大的基础模型开始并不一定是在专业任务上表现最佳的方法。

更新时间: 2026-10-06 15:20:27

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2410.21582v4

How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation

Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.

Updated: 2026-10-06 15:17:56

标题: 编码代理的自我修正应携带多少证据?自适应狄利克雷证据用于自蒸馏

摘要: 执行反馈让编码代理人修订程序并从自己的更正中学习。一个更正的学习权重应该反映其执行所支持的转换以及支持该转换的证据量。我们引入了有效证据自蒸馏(EESD),它分别表示这些数量。标准化的执行相关性确定相对转换支持和有效伪记数质量;然后,Dirichlet后验产生一个受不确定性惩罚的KL-锚定更正学习权重。在对称先验下,改变质量保留类别排序,并且有效质量产生一个受其匹配固定质量对应物限制的受监督系数。在四个模型领域历史扫描中,从一个到八个可见观察,将未来结果的负对数似然减少了55.0-59.3%。在八个观察中,有效质量在所有四个比较中都比固定质量获得了更低的负对数似然。在主要的匹配DeepSeek/RunBugRun研究中,argmax预测在所有3,000个示例上达成一致,其中在关注度集中的情况下获得了最大的负对数似然增益。经过一轮更正学习后,DeepSeek/CodeARC所有测试的Pass@1从15.0%增加到20.4%,配对的95%源自临时区间为[+2.8,+8.0]个百分点。十二种设置的下游评估建立了这一更新的模型领域范围。这些结果展示了如何在编码代理人中将证据支持与证据质量分开,从而改变概率估计和更正学习。

更新时间: 2026-10-06 15:17:56

领域: cs.AI,cs.SE

下载: http://arxiv.org/abs/2610.08514v1

Wiki-Talkie: Multilingual Benchmarking of Persona-Based Agents on Real-World Discussions

LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human interactions. Central to this is grounding agents in realistic user personas, yet existing datasets rely on fictional personas and are limited to a handful of languages, lacking the empirical grounding necessary to evaluate behavioral fidelity across diverse populations. We introduce Wiki-Talkie, a multilingual dataset of real-world conversations from Wikipedia Talk pages across five languages spanning two language families: Germanic (German, English) and Romance (Spanish, French, Italian), paired with personas derived from real user communities and encompassing sociodemographic attributes, self-descriptions, and behaviorally grounded interaction traits. Using Wiki-Talkie, we evaluate agent interactional behavior on a next-turn generation task across various persona conditioning strategies. Our evaluation assesses whether agents collectively reproduce the distributional behavioral patterns observed in human discussions. Results show that user's comment history exemplifying interaction behavior consistently outperforms explicit persona information. In addition, models systematically underproduce negative or extreme sentiments, while over producing references and suggestions, revealing biases toward agreeableness and positivity. Crucially, these patterns hold robustly across languages, with small cross-lingual differences.

Updated: 2026-10-06 15:17:46

标题: 维基对讲机:基于真实世界讨论的多语言人物代理基准测试

摘要: LLMs越来越多地被部署为社交环境中的自主代理,因此研究它们模拟人类互动的能力至关重要。其中的关键是将代理程序植根于现实用户角色,然而现有数据集依赖于虚构的用户角色,并且仅限于少数几种语言,缺乏必要的实证基础来评估行为忠实度在不同人群中的差异。我们介绍了Wiki-Talkie,这是一个跨越两种语言系列(日耳曼语系(德语、英语)和罗曼语系(西班牙语、法语、意大利语)的五种语言)的真实对话的多语言数据集,这些对话来自维基百科对话页面,配有来源于真实用户社区的用户角色,涵盖社会人口属性、自我描述和行为基础互动特征。利用Wiki-Talkie,我们评估了代理的互动行为在下一个转换生成任务上,采用各种用户角色调整策略。我们的评估评估了代理是否共同复制了人类讨论中观察到的分布行为模式。结果显示,用户的评论历史,展示互动行为,始终优于明确的用户角色信息。此外,模型系统地低估了负面或极端情绪,同时过度产生了参考和建议,显示出对宜人性和积极性的偏见。至关重要的是,这些模式在各种语言中都表现出强大的稳健性,存在小的跨语言差异。

更新时间: 2026-10-06 15:17:46

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.08513v1

Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation

Paired translation between quasiperiodic physiological waveforms (i.e., recovering a target oscillatory signal from the source) is central to the interpretation of cardiovascular signals derived from wearables placed at different body locations. This source-to-target mapping in these problems carries inherent geometric structure: the phase wraps around the cycle and must be treated as a circular variable, the amplitude remains strictly positive, and the beat-to-beat alignment can drift unpredictably across cycles and subjects. While deep neural networks have been used for phase estimation and complex-valued signal modeling, prior work does not explicitly learn phase transport between paired signals. Consequently, neither endpoint-supervised regression nor the standard affine path used in flow matching accounts for this phase--amplitude structure. We introduce \emph{cylindrical geodesic flow matching} for paired cardiovascular waveform translation. We show that the standard affine path used in flow matching distorts intermediate amplitude and instantaneous frequency when interpolating between quasiperiodic signals; replacing it with a closed-form geodesic on the phase--amplitude cylinder eliminates these artifacts and converts each training pair into dense, geometry-consistent velocity supervision. On zero-shot photoplethysmography and limited-support seismocardiography adaptation benchmarks, our method consistently outperforms interpolation baselines and matches or exceeds direct supervised prediction, reducing Hilbert Transform, $L_2$, and Dynamic Time Warping distance by up to ${\sim}15\%$ over the strongest competing baseline. These results suggest that bridge geometry is a critical inductive bias for flow matching on oscillatory signal translation.

Updated: 2026-10-06 15:15:07

标题: 圆柱形测地线流匹配用于准周期生理信号转换

摘要: 在不同身体部位放置的可穿戴设备所衍生的心血管信号的解释中,对准周期生理波形之间的配对翻译(即从源信号中恢复目标振荡信号)至关重要。这些问题中的源到目标映射具有固有的几何结构:相位在周期内环绕并必须被视为一个圆形变量,振幅保持严格为正,以及心跳对齐可能在周期和主体之间不可预测地漂移。尽管深度神经网络已被用于相位估计和复值信号建模,先前的工作并没有明确学习配对信号之间的相位传输。因此,既不是端点监督回归,也不是在流匹配中使用的标准仿射路径考虑了这种相位-振幅结构。我们引入了\emph{圆柱测地线流匹配}用于配对心血管波形翻译。我们展示了在插值准周期信号之间时,流匹配中使用的标准仿射路径会扭曲中间振幅和瞬时频率;将其替换为相位-振幅圆柱体上的闭合式测地线可消除这些人工制品,并将每个训练对转换为密集的、几何一致的速度监督。在零样本光电容积描记术和有限支持心脏震动图适应基准上,我们的方法始终优于插值基线,并与直接监督预测相匹配或超越,将希尔伯特变换,$L_2$和动态时间规整距离减少了高达约15\%的最强竞争基线。这些结果表明,桥梁几何是振荡信号翻译的流匹配中的一个关键归纳偏见。

更新时间: 2026-10-06 15:15:07

领域: cs.AI

下载: http://arxiv.org/abs/2610.08510v1

Terracotta: Enabling the Adoption of New DRAM Techniques via a Flexible DRAM Interface and Memory Controller

DRAM continues to limit the performance, energy efficiency, and robustness of modern systems. Many prior works propose DRAM techniques that support in-DRAM computation, improve memory access latency and parallelism, and enhance DRAM maintenance and reliability. However, adopting each new technique requires repeated modifications to the rigid DRAM interface and memory controller, hindering its deployment. Our goal is to reduce these repeated modifications. We observe that the DRAM commands and memory controller structures of many DRAM techniques are similar. Our key idea is to use these similarities to compose a set of primitives for implementing diverse DRAM techniques. We propose Terracotta, a new framework with two flexible components: (i) custom command extensions that let DRAM vendors define new commands within a single, standardized interface, and (ii) a programmable memory controller that system designers can program to support new DRAM techniques post-silicon. Together, these enable deployment by configuring the memory controller instead of modifying the interface and controller. We design Terracotta for a DDR5-based system and evaluate its performance, energy, and hardware complexity. For four DRAM techniques from four distinct domains (processing-using-DRAM, low-cost DRAM maintenance, subarray-level parallelism, and latency reduction), Terracotta retains almost all of the performance benefits (>96%) of custom implementations. A Terracotta-based composition of two techniques outperforms the Terracotta-based implementation of each technique alone, demonstrating the benefits of adding techniques without repeated interface and controller modifications. Terracotta incurs low DRAM energy (0.6-3.2%), area (0.03%), and power (0.56%) overheads in a high-end server-grade processor. Terracotta's source code is freely available at https://github.com/CMU-SAFARI/Terracotta.

Updated: 2026-10-06 15:14:32

标题: Terracotta:通过灵活的DRAM接口和内存控制器促进新的DRAM技术的采用

摘要: DRAM继续限制现代系统的性能、能源效率和稳健性。许多先前的研究提出了支持在DRAM内进行计算、改善内存访问延迟和并行性,并增强DRAM维护和可靠性的DRAM技术。然而,采用每种新技术都需要对刚性DRAM接口和内存控制器进行重复修改,阻碍了其部署。我们的目标是减少这些重复修改。我们观察到许多DRAM技术的DRAM命令和内存控制器结构是相似的。我们的关键思想是利用这些相似之处来组合一组原语,用于实现多样化的DRAM技术。我们提出了Terracotta,这是一个新框架,包含两个灵活的组件:(i)自定义命令扩展,让DRAM供应商在单一的标准接口内定义新命令,(ii)可编程内存控制器,系统设计者可以在硅后期编程以支持新的DRAM技术。通过这些组件,可以通过配置内存控制器来实现部署,而不是修改接口和控制器。我们为基于DDR5的系统设计了Terracotta,并评估了其性能、能源和硬件复杂性。对来自四个不同领域的四种DRAM技术(使用DRAM进行处理、低成本DRAM维护、子阵列级并行性和延迟降低),Terracotta保留了几乎所有定制实现的性能优势(>96%)。Terracotta基于两种技术的组合优于单独实现每种技术的Terracotta,展示了在没有重复接口和控制器修改的情况下添加技术的好处。Terracotta在高端服务器级处理器中承担了较低的DRAM能耗(0.6-3.2%)、面积(0.03%)和功耗(0.56%)开销。Terracotta的源代码可在https://github.com/CMU-SAFARI/Terracotta免费获取。

更新时间: 2026-10-06 15:14:32

领域: cs.AR,cs.CR,cs.DC

下载: http://arxiv.org/abs/2610.06475v2

X-OPM: Explainable Automatic Digital On-Chip Power Modeling for Enhanced Robustness

Proactive power management systems reduce processor dynamic power through runtime power prediction and power-aware scheduling. Accurate, stable and low-overhead digital on-chip power meters (OPMs) are crucial for improving the prediction quality. Recent studies have explored various modeling methods, including using linear models, decision trees, and multi-layer perceptrons (MLPs) to construct OPMs. However, most current approaches train models end-to-end without analyzing the physical interpretability of features, affecting their ability to generalize to unseen workloads. Grounded in the design principles of synchronous digital VLSI circuits, X-OPM introduces a robust feature engineering framework that uses tree-based models to capture feature interactions and linear models for prediction. It also incorporates a human-in-the-loop workflow to balance model accuracy against modeling effort. Evaluated on a commercial C906 vector processor, X-OPM consistently achieves $R^2 > 0.93$ across all workloads with sampling window size set below $8$ cycles. In contrast, state-of-the-art methods including APOLLO, COBIT, and standard MLPs fail to generalize across all test cases. Layout with commercial EDA tools shows that X-OPM incurs an area overhead below $0.1\%$, which is on par with lightweight tree-based and linear models, and significantly smaller than MLP-based models.

Updated: 2026-10-06 15:09:52

标题: X-OPM:可解释的自动数字芯片功耗建模以增强鲁棒性

摘要: 主动式功率管理系统通过运行时功率预测和功率感知调度来减少处理器的动态功率。准确、稳定且低开销的数字片上功率计(OPMs)对于提高预测质量至关重要。最近的研究已探索各种建模方法,包括使用线性模型、决策树和多层感知器(MLPs)来构建OPMs。然而,大多数当前方法在训练模型时端到端进行,而没有分析特征的物理可解释性,影响其泛化到未知工作负载的能力。基于同步数字VLSI电路设计原则,X-OPM引入了一个强大的特征工程框架,使用基于树的模型来捕捉特征交互作用,并使用线性模型进行预测。它还结合了一个人在回路工作流程,以平衡模型的准确性和建模工作量。在商用C906矢量处理器上评估,X-OPM在所有工作负载下始终实现$R^2 > 0.93$,采样窗口大小设定在小于$8$个周期。相比之下,包括APOLLO、COBIT和标准MLPs在所有测试用例中泛化失败。使用商业EDA工具的布局显示,X-OPM的面积开销低于$0.1\%$,与轻量级基于树的和线性模型相当,并且明显小于基于MLP的模型。

更新时间: 2026-10-06 15:09:52

领域: cs.AR,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08502v1

Federated Mixture-of-Experts Alignment on Mobile Edge Networks under Data Heterogeneity

The growing demand for on-device large language model (LLM) services on mobile edge devices has driven the adoption of Mixture-of-Experts (MoE) architectures, which scale model capacity with limited computation. Since fine-tuning MoE-based LLMs relies on privacy-sensitive local data, federated learning (FL) offers a natural paradigm for collaborative training without exposing raw data. However, integrating MoE-based LLM fine-tuning into FL faces two critical challenges caused by data heterogeneity across clients: (i) divergent local data distributions drive clients to develop distinct gating preferences, so direct parameter aggregation yields a one-size-fits-none global gating network; and (ii) same-indexed experts develop disparate semantic roles across devices, leading to expert semantic blurring and degraded specialization. To address these challenges, we propose FedAlign-MoE, a federated aggregation alignment framework for edge computing systems that jointly enforces routing consistency and expert semantic alignment. Specifically, FedAlign-MoE aggregates gating behaviors by aligning routing distributions through consistency weighting and optimizes local gating networks through distribution regularization, maintaining cross-client stability while preserving discriminative local gating preferences. Meanwhile, FedAlign-MoE quantifies the semantic consistency of same-indexed experts across devices and selectively aggregates semantically aligned experts, ensuring stable and specialized global experts. Extensive experiments demonstrate that FedAlign-MoE outperforms state-of-the-art benchmarks, achieving faster convergence and higher accuracy in non-IID federated environments with lightweight computation and efficient communication.

Updated: 2026-10-06 15:09:51

标题: 在移动边缘网络中基于数据异构性的联合专家混合对齐

摘要: 对于移动边缘设备上的设备端大型语言模型(LLM)服务的不断增长需求推动了混合专家(MoE)架构的采用,这种架构可以在有限的计算资源下扩展模型容量。由于微调基于MoE的LLM依赖于隐私敏感的本地数据,联邦学习(FL)为协作训练提供了一种不暴露原始数据的自然范式。然而,将基于MoE的LLM微调集成到FL中面临着两个关键挑战,这是由于客户端之间的数据异质性造成的:(i)不同的本地数据分布驱动客户端开发出不同的门控偏好,因此直接参数聚合会产生一个不适合所有的全局门控网络;(ii)相同索引的专家在设备间发展出不同的语义角色,导致专家语义模糊和专业化下降。为了解决这些挑战,我们提出了FedAlign-MoE,这是一个用于边缘计算系统的联邦聚合对齐框架,它同时强制执行路由一致性和专家语义对齐。具体来说,FedAlign-MoE通过一致性加权对齐路由分布来聚合门控行为,并通过分布正则化优化本地门控网络,保持跨客户端的稳定性同时保留有区分性的本地门控偏好。与此同时,FedAlign-MoE量化了设备间相同索引专家的语义一致性,并选择性地聚合语义对齐的专家,确保全局专家的稳定性和专业化。大量实验证明,FedAlign-MoE优于最先进的基准,能在轻量级计算和高效通信的非独立同分布(non-IID)联邦环境中实现更快的收敛和更高的准确性。

更新时间: 2026-10-06 15:09:51

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2603.21276v2

Universe of Thoughts: A Computational Framework for Creative Reasoning in Large Language Models

Recent advances in Large Language Model (LLM) reasoning have improved conventional problem solving, but creative reasoning remains comparatively underexplored. Inspired by cognitive science, we formalize combinational, exploratory, and transformational creativity as executable computational operators over structured problem and solution spaces, specifying how each mode combines, explores, or transforms those spaces. Combinational reasoning transfers ideas across domains to form unfamiliar combinations; exploratory reasoning searches for new solutions within an existing conceptual space; and transformational reasoning modifies the rules or constraints that define that space. This formalization yields distinct algorithmic procedures, which we instantiate in Universe of Thoughts (UoT), an LLM reasoning framework. Existing creativity benchmarks emphasize either open-ended ideation or highly constrained problem solving. We therefore introduce three novel creative-reasoning tasks requiring concrete solutions in low-constraint settings. Across 10 generations per method and task, T-UoT with GPT-4o performs strongest on the low-constraint, high-objective-specificity Bridge and Electricity tasks, while C-UoT shows its strongest relative performance on the low-constraint, lower-objective-specificity Society task. In addition, we evaluate UoT on HypoArena, an independent scientific hypothesis-generation benchmark with 100 tasks across biomedical, machine-learning, and social-science domains. With Qwen3-14B, Exploratory UoT ranks first among seven reasoning methods, achieving a 32.7\% pairwise win rate compared with 25.5\% for the next-best method. Our results suggest distinct performance patterns across task structures: T-UoT is strongest in low-constraint, high-specificity settings, E-UoT in more constrained, high-specificity settings, and C-UoT in low-constraint, lower-specificity settings.

Updated: 2026-10-06 15:08:49

标题: 思维宇宙:大型语言模型中创造性推理的计算框架

摘要: 最近大型语言模型(LLM)推理的进展改善了传统问题解决,但创造性推理仍然相对未被充分探索。受认知科学启发,我们将组合性、探索性和转换性创造力形式化为可执行的计算操作符,作用于结构化问题和解决方案空间,具体规定了每种模式如何结合、探索或转换这些空间。组合推理将想法跨领域传递,形成不熟悉的组合;探索性推理在现有概念空间中寻找新的解决方案;转换性推理修改定义该空间的规则或约束。这种形式化产生了独特的算法程序,我们在思维宇宙(UoT)中实现了这些程序,这是一个LLM推理框架。现有的创造力基准强调开放性构思或高度受限制的问题解决。因此,我们引入了三个新颖的创造性推理任务,需要在低约束环境中提供具体解决方案。在每种方法和任务的10代中,T-UoT与GPT-4o在低约束、高客观特异性的桥梁和电力任务上表现最佳,而C-UoT在低约束、较低客观特异性的社会任务上表现出最强的相对性能。此外,我们在HypoArena上评估了UoT,这是一个独立的科学假设生成基准,涵盖生物医学、机器学习和社会科学领域的100个任务。使用Qwen3-14B,探索性UoT在七种推理方法中排名第一,在32.7\%的配对胜率中超过次佳方法的25.5%。我们的结果表明在任务结构上存在独特的性能模式:T-UoT在低约束、高特异性设置中表现最强,E-UoT在更受限制、高特异性设置中表现最强,C-UoT在低约束、较低特异性设置中表现最强。

更新时间: 2026-10-06 15:08:49

领域: cs.AI

下载: http://arxiv.org/abs/2511.20471v3

Language-model ratings of depression reflect the rater more than the patient

Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.

Updated: 2026-10-06 15:08:17

标题: 语言模型对抑郁症的评分更多地反映了评分者而不是患者

摘要: 抑郁症没有诊断性血液检测。语言模型承诺不知疲倦、一致的评估,但准确的评分者是否会在个体上产生不一致?我们预先注册了880个语言模型评分者,将11种开放模型与提示和评分选择相结合,并将它们应用于189个面试,与八项患者健康问卷进行比较。模型选择解释了总症状评分方差的30.0%,稳定的参与者差异为10.5%。两个随机选择的评分者,其接收者操作特征曲线(AUC)大于等于0.70,对40%的参与者的筛查决定存在分歧,平均而言。过度评分主导了被标记为多少,然而能力相等的评分者对大约五分之一的参与者做出不同选择。对86个新面试的锁定分析重现了主要预先注册的发现。通过对40名标记过的参与者进行探索性重新校准,将准确率从约60%提高到75%,并将分歧减半,使得大约五分之一的参与者做出了不同选择。校准修复了评分者依赖的大部分问题,但并没有就个体达成一致。

更新时间: 2026-10-06 15:08:17

领域: cs.CL,cs.AI,q-bio.NC

下载: http://arxiv.org/abs/2610.08501v1

Information-Dense Synthesis for Molecular Discovery

Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms offer little gain over random guessing. We propose a method to efficiently search large regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map. We optimize synthesis to encode maximal information. Theoretically, this approach can reduce the number of experiments required to find the optimal molecule among $d$ candidates from $\mathcal{O}(d)$ to $\mathcal{O}(\log d)$ or $\mathcal{O}(1)$. In simulation, on estimated protein fitness landscapes, it finds active molecules with an order of magnitude fewer experiments than existing Bayesian optimization methods.

Updated: 2026-10-06 15:04:49

标题: 分子发现的信息密集综合

摘要: 机器学习可以通过设计分子和规划实验加速分子发现。然而,许多科学挑战需要具有非常罕见特性的分子,在这种稀疏情况下,现有算法与随机猜测相比提供的收益很少。我们提出了一种方法,通过算法控制的随机合成来高效搜索分子空间的大区域。与设计、制造和测试单个分子不同,我们设计和制造复杂的混合物,将它们作为一个池进行测试,然后解析分子活性图。我们优化合成以编码最大信息量。从理论上讲,这种方法可以将发现 $d$ 个候选分子中的最佳分子所需的实验数量从 $\mathcal{O}(d)$ 降低到 $\mathcal{O}(\log d)$ 或 $\mathcal{O}(1)$。在模拟中,在估计的蛋白质适应性景观上,它比现有的贝叶斯优化方法少一个数量级地找到活性分子。

更新时间: 2026-10-06 15:04:49

领域: stat.ML,cs.LG,physics.chem-ph,q-bio.BM

下载: http://arxiv.org/abs/2610.08495v1

Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models

Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface temperature fields from a state-of-the-art global climate model. As a key advance over previous work, our framework explicitly models both atmospheric dynamical interactions arising from internal climate variability and forced responses due to changes in atmospheric greenhouse gas and aerosol concentrations. When trained on future climate change scenarios, our method accurately predicts the long-term global mean and regional temperature evolution and shows physically realistic responses to perturbations in greenhouse gas and aerosol concentrations when evaluated on unseen scenarios. Our results underline the potential of causal representation learning frameworks for advancing climate model emulation.

Updated: 2026-10-06 15:00:35

标题: 学习气候模型中外力对温度影响的层次因果表示

摘要: 机器学习(ML)仿真器提供了一种快速且具有成本效益的方法,可以在对地球系统模型预测进行训练后模拟气候变化情景。然而,这些基于数据驱动方法的黑匣子特性限制了它们的输出可用性和可信度,特别是它们作为因果归因工具的使用。在这里,我们开发了一个应用于来自最先进全球气候模型的海表温度场的分层因果表示学习框架。与以往工作相比的一个关键进展是,我们的框架明确地模拟了由内部气候变率引起的大气动力相互作用和由大气温室气体和气溶胶浓度变化引起的强迫响应。当在未知情景上进行评估时,我们的方法在未来气候变化情景上训练后准确预测了全球平均和区域温度的长期演变,并对温室气体和气溶胶浓度扰动显示出物理上合理的响应。我们的结果强调了因果表示学习框架在推进气候模型仿真方面的潜力。

更新时间: 2026-10-06 15:00:35

领域: cs.LG

下载: http://arxiv.org/abs/2609.30995v2

Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment

Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.

Updated: 2026-10-06 14:58:53

标题: Knee3DVLM:全序列全容量视觉-语言模型用于全面膝关节MRI评估

摘要: 视觉-语言模型(VLMs)越来越多地应用于三维医学影像,但它们在膝关节MRI方面的应用仍然有限,特别是用于解释临床实践中使用的互补序列。我们引入了Knee3DVLM,这是一个序列感知VLM,使用全容量DESS和液体敏感TSE MRI来预测从MRI骨关节炎膝关节评分(MOAKS)派生的57个解剖学分辨率的二进制诊断目标,用于结构化报告。我们使用互不相交的骨关节炎倡议分区评估了仅DESS、仅TSE和配对DESS-TSE配置。在一个包含1,074次检查的留存队列中,融合模型实现了72.98%的平均准确率、71.17%的平衡准确率、78.96%的平均ROC-AUC和78.74%的宏观ROC-AUC,这些数值是三种配置中最高的。在与发布的3DReasonKnee队列对齐的二次多类别分析中,Knee3DVLM在五个病理类别中均比最强报告的3DReasonKnee配置高。这些发现支持对全面膝关节MRI评估的双序列全容量建模。

更新时间: 2026-10-06 14:58:53

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.08482v1

MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata

Few-shot meta-learning traditionally formulates task adaptation either as analytical gradient descent through unrolled computational graphs or as metric-based distance comparisons over flattened 1D fea- ture vectors, which either incur costly test-time backpropagation or discard native 2D spatial geometry. In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference. MetaLearnNCA decomposes task adaptation into an Active- NCA, which executes task inference conditioned on a continuous 2D spatial memory grid termed the spatial program, and a learned Meta-NCA, which acts as a decentralized cellular optimizer by diffusing spatial error residuals across local neighborhoods to dynamically update this program. METALEARN- NCA is competitive against canonical meta-learners in-distribution (96.12% on Omniglot) with Out-Of- Distribution transfer gains on MNIST, KMNIST, and Fashion-MNIST transfer across 10 independent testing seeds across 1-, 5-, and 10-shot regimes (e.g., surpassing Prototypical Networks by +10.54% on 10-shot MNIST and a +3.87% gain on 10-shot Fashion-MNIST over FOMAML). Our results establish that robust, gradient-free learning-to-learn can emerge from decentralized cellular dynamics on non-von Neumann substrates.

Updated: 2026-10-06 14:56:29

标题: MetaLearnNCA:通过相互作用的神经元细胞自动机进行少样本离线元学习

摘要: Few-shot meta-learning传统上将任务适应性要么表述为通过展开的计算图进行分析梯度下降,要么作为基于度量距离比较的扁平化1D特征向量,这两种方法都会导致昂贵的测试时间反向传播,或者丢弃原生的2D空间几何。在这项工作中,我们提出了METALEARNNCA,这是一个分散式框架,通过耦合的神经元细胞自动机(NCAs)的动态相互作用实现了少样本适应,而在推理过程中不需要计算分析梯度。MetaLearnNCA将任务适应性分解为一个主动NCA,它执行基于连续的2D空间内存网格(称为空间程序)的任务推理,以及一个学习的Meta-NCA,它作为一个分散式细胞优化器,通过在本地邻域中扩散空间错误残差来动态更新这个程序。METALEARN-NCA在分布内(在Omniglot上为96.12%)竞争力强,同时在MNIST、KMNIST和Fashion-MNIST上的跨域转移中,在1-shot、5-shot和10-shot制度下的10个独立测试种子中获得了收益(例如,在10-shot MNIST上超过了Prototypical Networks的+10.54%,在10-shot Fashion-MNIST上超过了FOMAML的+3.87%)。我们的结果表明,从非冯诺依曼衬底上的分散式细胞动力学中可以出现稳健的无梯度学习。

更新时间: 2026-10-06 14:56:29

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.08479v1

Learning PDE solution operators with variable initial conditions via Latent Dynamics Networks

In many-query scenarios, data-driven surrogate models provide an efficient alternative to high-fidelity solvers for simulating physical systems governed by Partial Differential Equations (PDEs). In this context, the Latent Dynamics Network (LDNet) has recently demonstrated remarkable performance in predicting the response of spatio-temporal systems, combining Neural Ordinary Differential Equations with nonlinear dimensionality reduction. However, the original formulation assumes a fixed initial condition, limiting its applicability to many real-world applications where a system evolves from varying starting states. In this work, we overcome this limitation while keeping the end-to-end training procedure of the original LDNet and its encoder-free nature, which preserves its intrinsic independence from spatial resolution and grid topology. We infer the initial latent state directly from a small set of early-time observations, treating latent-state initialization as an adaptation problem, and investigate two strategies: an auto-decoding formulation and a meta-learning approach in which the initial latent state acts as a task-specific context variable. We demonstrate the accuracy of the proposed methods across diverse physical phenomena, spanning advection-diffusion, fluid dynamics, and solid mechanics. Meta-learning markedly accelerates latent-state inference and induces smoother, better-conditioned optimization landscapes, and spontaneously organizes the latent space into a structured representation that reflects physically meaningful features of the underlying dynamics. The coordinate-based decoder enables training from spatially subsampled data while recovering high-resolution solution fields at inference. The resulting approach provides an efficient and resolution-independent surrogate modeling framework for many-query simulations of time-dependent PDEs with varying initial conditions.

Updated: 2026-10-06 14:55:39

标题: 通过潜在动态网络学习具有可变初始条件的PDE解算符

摘要: 在许多查询场景中,数据驱动的代理模型为受偏微分方程(PDEs)控制的物理系统模拟提供了一种高效的替代方案,而不是高保真度求解器。在这种情况下,潜在动力网络(LDNet)最近展示了在预测时空系统响应方面的卓越性能,将神经常微分方程与非线性降维结合起来。然而,原始公式假定固定的初始条件,限制了其适用于许多实际应用场景,其中系统从不同的初始状态演变。在这项工作中,我们克服了这一限制,同时保持了原始LDNet的端到端训练过程和无编码器的特性,这保留了它与空间分辨率和网格拓扑的固有独立性。我们直接从一小组早期观测中推断初始潜在状态,将潜在状态初始化视为一种适应性问题,并研究了两种策略:一种自动解码公式和一种元学习方法,其中初始潜在状态充当任务特定的上下文变量。我们跨越各种物理现象展示了所提出方法的准确性,涵盖了对流扩散、流体动力学和固体力学。元学习显著加速了潜在状态推断,并诱导了更平滑、更好条件的优化景观,并自发地将潜在空间组织成反映基础动态物理意义特征的结构化表示。基于坐标的解码器使得可以从空间子采样数据进行训练,同时在推断时恢复高分辨率解场。由此产生的方法为具有不同初始条件的时间依赖PDEs的多查询模拟提供了一种高效且与分辨率无关的代理建模框架。

更新时间: 2026-10-06 14:55:39

领域: cs.LG,math.NA

下载: http://arxiv.org/abs/2610.08475v1

Federated Bayesian Surveillance of Mechanical Thrombectomy Adverse Events: A Population Risk Layer for Surgical Digital Twins

Learned surgical simulators and world models can roll out plausible procedural futures, but they carry no grounded estimate of how often interventional devices actually harm patients. We propose treating population-scale adverse-event surveillance as a distinct belief layer of the surgical digital twin, and we evaluate a federated Bayesian protocol for learning it under formal privacy guarantees. Each site holds per-class Gamma-Poisson posteriors over adverse-event rates and exchanges only Rényi-differentially-private natural-parameter updates. We benchmark on the complete FDA MAUDE cohort for thrombus-retrieval catheters (product code NRY): 8,617 reports, of which 6,491 are classified by transparent keyword rules into five thrombectomy complication classes and partitioned across $K=8$ manufacturer sites. At a matched privacy budget of $(\\varepsilon \\approx 2.09, \δ= 10^{-5})$, the conjugate protocol attains a held-out Poisson score of -5.78 per test event versus -26.58 for FedAvg with differential privacy. The non-private federated model also outperforms centralized pooling (+3.19 vs +2.93), evidence that manufacturer-specific complication profiles are real and that federation preserves them. Because MAUDE lacks procedure denominators, outputs are relative rate orderings rather than absolute risks, and we report all privacy-utility operating points.

Updated: 2026-10-06 14:49:28

标题: 联邦贝叶斯监测机械溶栓不良事件:外科数字双胞胎的人口风险层

摘要: 学习的外科模拟器和世界模型可以展示出合理的程序性未来,但它们并没有对介入设备实际伤害患者的频率进行基于事实的估计。我们建议将人口规模的不良事件监测视为外科数字孪生体的一个独立信念层,并评估在正式隐私保证下学习它的联合贝叶斯协议。每个站点都持有关于不良事件率的每类伽玛-泊松后验,并且仅交换Rényi差分私有的自然参数更新。我们在完整的FDA MAUDE队列中对血栓取出导管(产品代码NRY)进行基准测试:共8,617份报告,其中有6,491份根据透明的关键词规则分类为五种血栓切除并跨越$K=8$制造商站点。在匹配的隐私预算下($\varepsilon \approx 2.09,δ= 10^{-5}$),共轭协议在测试事件中达到持有的泊松分数为-5.78,而带有差分隐私的FedAvg为-26.58。非私有的联合模型也优于集中汇总(+3.19 vs +2.93),证明制造商特定的并发症概况是真实的,并且联合保留了这些概况。由于MAUDE缺乏程序分母,输出是相对风险排序而不是绝对风险,我们报告所有隐私-效用操作点。

更新时间: 2026-10-06 14:49:28

领域: cs.CR,cs.RO

下载: http://arxiv.org/abs/2610.08464v1

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target's already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree in one pass and commits exactly its greedy output. In one production engine at a fixed round budget, GLANCE decodes up to 3.05 times faster than autoregressive decoding and outpaces the production EAGLE3-VL head on average and by about 11% on grounded tasks. An entropy law explains when drafting pays, predicting the longest accepted blocks on grounded tasks, where the target's next-token entropy is lowest. Our code is available at https://github.com/js-lee-AI/GLANCE.

Updated: 2026-10-06 14:49:05

标题: 视觉并非开销:一次通过块草拟的方法用于视觉-语言模型中无损推断解码

摘要: 投机性解码加速生成过程,但不改变其输出,但在视觉语言模型(VLMs)上,自我强化循环使其受阻。由于自回归起草者为每个起草的令牌支付顺序传递,因此必须保持较小,并且无法在每次传递时关注图像。因此,先前的工作压缩或隐藏图像,使得起草者在确定图像的文本上最薄弱。我们提出了GLANCE,一种一次性块起草器,它在未经修改的VLM目标上打破了这种循环。其块扩散头在目标的已融合视觉语言状态上进行一次正向传递,一次性草拟一个完整的块,无论草稿有多深,都只读取多模态上下文一次。目标在一次传递中验证一个广泛的候选树,并确切地承诺其贪婪输出。在固定轮次预算的一个生产引擎中,GLANCE的解码速度比自回归解码快高达3.05倍,并且平均超过生产EAGLE3-VL头部大约11%的基准任务。熵定律解释了何时起草有用,预测接受最长块的基准任务,在那里目标的下一个标记熵最低。我们的代码可在https://github.com/js-lee-AI/GLANCE找到。

更新时间: 2026-10-06 14:49:05

领域: cs.AI,cs.CL,cs.CV

下载: http://arxiv.org/abs/2609.00355v3

UNREAL: Unifying Retrieval and Long-Context with a Single Model

Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.

Updated: 2026-10-06 14:48:53

标题: 《UNREAL: 用单一模型统一检索和长上下文》

摘要: 长上下文推理和检索增强生成(RAG)处理不同范围的证据选择,从单个长提示到整个语料库。我们询问是否一个单一的模型内部机制可以跨越这个范围选择证据。我们引入了UNifying REtrieval And Long-Context with a Single Model (UNREAL),这是一个模型原生的证据选择框架,可以跨越语料库检索和长上下文推理。UNREAL对分块进行编码,并直接从冻结的LLM内部表示派生检索查询。它增加了不到500K可训练参数,并且保持了骨干结构不变。在一个3B令牌、21M分块的维基百科索引上,所有四个密集和混合UNREAL骨干模型优于最先进的检索-重新排名系统。最佳模型在HotpotQA上将召回率从49.1%提高到73.2%,在2WikiMultiHopQA上从31.7%提高到60.1%,在MuSiQue上从8.8%提高到14.4%。应用于长上下文任务时,相同的选择机制在生成之前去除干扰因素,使NoLiMa在最大上下文长度为128K令牌时的准确性从1.0%提高到24.83%,将LV-Eval的F1分数从49.97%提高到54.66%在256K时。UNREAL还减少了FLOPs和从大约32K令牌开始的第一个令牌时间,随着上下文的增长,获益更大。这些结果共同建立了模型内部证据选择作为语料库检索和证据稀疏的长上下文推理的共同基础。

更新时间: 2026-10-06 14:48:53

领域: cs.CL,cs.IR,cs.LG

下载: http://arxiv.org/abs/2610.08463v1

One-Shot Private Confidence Regions via Resampling

We propose a simple framework for constructing differentially private confidence regions \textit{in one shot}, i.e., by adding noise only to the final resampling quantile instead of privatizing the estimator computed on each resample. The cost of privacy of our procedure is only logarithmic in the number of resamples $B$ under with-replacement ($m$-out-of-$n$) sampling and independent of $B$ under without replacement sampling (subsampling), avoiding the $\sqrt{B}$ factor that arises in previous works. We provide nonasymptotic Gaussian Differential Privacy (GDP) and utility guarantees for both subsampling and $m$-out-of-$n$ resampling, covering mean-like estimators with small global sensitivity as well as estimators admitting efficiently computable smooth sensitivity bounds, including quantiles and degenerate U-statistics. This allows us to also obtain private confidence regions for degenerate U-statistics where the private error is much smaller than the non-private error. In all, we provide a toolbox for widely applicable DP uncertainty quantification procedures under popular resampling strategies while avoiding the computational and privacy costs of privatizing many intermediate resample statistics.

Updated: 2026-10-06 14:47:06

标题: 利用重采样实现一次性私密置信区间

摘要: 我们提出了一个简单的框架,用于构建差分私有置信区间“一次性”,即仅将噪声添加到最终重新抽样分位数,而不是对每个重新抽样计算的估计值进行私有化。我们的程序的隐私成本仅在基于替换(m个中n个)抽样下的重新抽样数量$B$的对数中,并且在无替换抽样(子抽样)下与$B$无关,避免了以前作品中出现的$\sqrt{B}$因子。我们为子抽样和$m$个中n个重新抽样提供了非渐近高斯差分隐私(GDP)和效用保证,覆盖具有小全局敏感度的类似均值估计器以及具有可有效计算的平滑敏感度界限的估计器,包括分位数和退化U-统计量。这使我们还可以获得退化U-统计量的私有置信区间,其中私有误差比非私有误差小得多。总之,我们提供了一个工具箱,用于在流行的重新抽样策略下广泛适用的差分私有不确定性量化程序,同时避免了私有化许多中间重新抽样统计量的计算和隐私成本。

更新时间: 2026-10-06 14:47:06

领域: stat.ML,cs.CR

下载: http://arxiv.org/abs/2610.08460v1

Improving Proactive AI Assistance with Hierarchical Procedural Understanding

Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust the guidance level to the user's expertise and needs. Supporting these capabilities requires training and evaluation data that reflect procedural structure and capture how guidance should adapt to task progress and user needs. However, existing datasets either focus on detection-based proactive understanding or provide procedural guidance at a fixed granularity. Fixed-granularity guidance provides limited information about fine-grained progress and broader procedural context, making it difficult to determine completion and adapt guidance granularity. To address these limitations, we introduce the ProactiveCoach suite, comprising ProactiveCoach-Instruct for training, ProactiveCoachBench for evaluation, and fine-tuned VLMs with an adaptive guidance system. ProactiveCoach-Instruct provides hierarchically structured guidance at the phase, step, and action levels for learning task progress and procedural context. ProactiveCoachBench evaluates whether models provide appropriate guidance at the right time across different guidance levels and adapt when the requested level changes. We fine-tune pretrained VLMs on ProactiveCoach-Instruct and demonstrate its effectiveness across backbones. Compared with fixed-granularity supervision, hierarchical supervision improves overall performance across backbones by up to 9.6%p. We further build an adaptive guidance system by combining our fine-tuned model with a lightweight guidance router. Without additional fine-tuning, our system outperforms the in-context adaptation baseline by 57.1%p across four guidance-level transitions. Our project page is available at https://jinsuby.github.io/ProactiveCoach/.

Updated: 2026-10-06 14:39:49

标题: 提高具有分层程序理解的主动AI辅助

摘要: 主动型人工智能助手持续观察用户的活动,并决定是否提供新的指导或保持沉默。他们应为任务提供适当的指导,根据任务进展确定何时提供下一步指导,并根据用户的专业知识和需求调整指导水平。支持这些能力需要反映程序结构并捕捉指导应如何适应任务进展和用户需求的训练和评估数据。然而,现有数据集要么专注于基于检测的主动理解,要么以固定粒度提供程序性指导。固定粒度的指导提供了有关细粒度进展和更广泛程序上下文的有限信息,使得确定完成和调整指导粒度变得困难。为了解决这些限制,我们引入了ProactiveCoach套件,包括用于训练的ProactiveCoach-Instruct,用于评估的ProactiveCoachBench以及具有自适应指导系统的微调VLMs。 ProactiveCoach-Instruct为学习任务进展和程序上下文提供了阶层结构化指导,涵盖了阶段、步骤和动作级别。 ProactiveCoachBench评估模型是否在不同的指导水平上在正确的时间提供适当的指导,并在请求的水平发生变化时进行调整。我们在ProactiveCoach-Instruct上对预训练的VLMs进行微调,并跨不同的骨干结构展示其有效性。与固定粒度监督相比,层次监督通过最多提高9.6%的综合性能。我们进一步通过将我们的微调模型与轻量级指导路由器相结合构建了一个自适应指导系统。在没有额外微调的情况下,我们的系统在四个指导级别转换中超过了上下文自适应基线57.1%。我们的项目页面可在https://jinsuby.github.io/ProactiveCoach/上找到。

更新时间: 2026-10-06 14:39:49

领域: cs.LG,cs.CV

下载: http://arxiv.org/abs/2610.06505v2

Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.

Updated: 2026-10-06 14:38:15

标题: 主动型AutoRAG:通过推理驱动代理优化RAG管道

摘要: 检索增强生成(RAG)是一种广泛使用的方法,用于将大型语言模型(LLMs)与外部知识联系起来。然而,配置一个流水线是一个昂贵的超参数优化问题,涉及许多相互作用的选择,从分块和嵌入模型到重新排名和生成。现有的优化器,从贪婪搜索到贝叶斯优化,将每个试验缩减为一个聚合分数,并在没有建模的情况下搜索为什么某个配置表现如此,尽管检索到的块已经提供了关于每个失败是在检索之后还是之前发生的证据。我们介绍了Agentic AutoRAG,这是一个用于多目标RAG超参数优化的LLM代理优化器,具有检索与生成失败归因。它提出了根据语料库中的一个冻结考试得分的配置:在每次试验后,一个诊断者将每个失败的问题归因为检索或生成,而一个建议者,在模型排名和定价的知识库中,选择下一个配置,权衡准确性与成本以跟踪帕累托边界。在三个多跳问答基准测试中,它达到了比我们比较的每个基线更高的LLM评分准确性,并且在前10次试验中,它与统计基线的30次评分准确性相匹配或超过。在真实世界的医疗语料库上的成本感知模式下,它达到了77%的中位考试准确性,高于最强基线的71.5%,每次查询的成本约为该基线的58%,并且在成本的约22%处达到了71.5%。

更新时间: 2026-10-06 14:38:15

领域: cs.CL,cs.IR,cs.LG

下载: http://arxiv.org/abs/2610.08452v1

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.

Updated: 2026-10-06 14:37:22

标题: 重新思考跨标记器的政策蒸馏:从对齐覆盖到监督可靠性

摘要: On-Policy Distillation (OPD)通过使用教师反馈在其自己的生成上训练学生。使用不同的分词器,比较教师和学生的预测需要在序列和词汇级别上进行对齐。在本文中,我们研究了扩展这种对齐覆盖范围是否有助于改善学习。在数学推理和代码生成的三对异质教师-学生对中,尽管存在实质性的词汇不匹配,严格的1:1组已经覆盖了大部分学生生成的标记。在蒸馏之前从学生中抽样的响应中,共享的词汇在平均严格对齐位置保留了几乎所有教师和学生的概率质量。在每个严格位置限制反向KL到共享词汇的学生选择的前16个子集,实现了与完全共享词汇OPD相当的准确性,优于评估的跨分词器基线。在不匹配组中添加在跨组范围对数概率上的均方误差监督提供了完整的监督覆盖范围,但降低了准确性。在仅使用严格损失进行训练的检查点中,跨度梯度显示出与严格梯度的弱或负向一致性,并且相对于严格梯度增长。这些诊断可能有助于解释添加跨度监督后准确性下降。我们的发现激励从最大化对齐覆盖范围转向优先考虑监督可靠性:在严格位置上的紧凑监督可能比引入弱对齐或冲突的训练信号更有效。

更新时间: 2026-10-06 14:37:22

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.08448v1

AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly

Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachment relations into 6D pose updates that enable parts to physically interact with the environment and previously assembled components. To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly. It firstly employs anchor-guided boundary assembly states to decompose manual pages into single-part operations and recover an assembly-tree. Then, it uses iterative after-state feedback refinement to guide successive (SE(3)) updates and corrections, and validates their physical plausibility through simulation-based release tests. Experiments show that compared with the strongest prior baseline, AssemState improves F1 from 38.58\% to 62.80\% and Tree Exact Match from 28.24\% to 53.92\% for assembly-tree recovery. On 243 independently evaluated part-level operations, our proposed iterative refinement improves judge-accepted operations from 0 to 5.3\% and reduces mean Chamfer distance from 5.4111 to 1.7744. However, visually plausible candidate poses may still suffer from collision, floating, mirror-orientation errors, incomplete seating, and wrong-side attachment. These results show that AssemState improves operation-structure recovery and selected local pose metrics, while MLLMs remain limited for spatial relationship reasoning.

Updated: 2026-10-06 14:35:49

标题: AssemState:零编组家具装配的手动和物理状态引导推理

摘要: 多模态大型语言模型(MLLMs)在视觉理解方面取得了显著进展,但精确的与物理环境集成的3D空间推理仍然困难重重。家具组装不仅需要从示意图手册中恢复步骤级操作,还需要将语义附着关系转化为6D姿态更新,使零件能够与环境和先前组装的部件进行物理交互。为了研究这个问题,我们提出了AssemState,一个用于手动和物理状态引导家具组装的零样本框架。首先,它利用锚点引导的边界组装状态将手动页面分解为单个零件操作,并恢复一个组装树。然后,它使用迭代的后状态反馈细化来引导连续的(SE(3))更新和校正,并通过基于模拟的释放测试验证它们的物理可信度。实验证明,与最强的基线相比,AssemState将组装树恢复的F1值从38.58\%提高至62.80%,将树完全匹配从28.24%提高至53.92%。在独立评估的243个零件级操作中,我们提出的迭代细化将接受判定的操作从0提高至5.3%,将平均Chamfer距离从5.4111降低至1.7744。然而,视觉上合理的候选姿态仍可能受到碰撞、悬浮、镜像方向错误、不完整座椅和错误的侧面附着的影响。这些结果表明,AssemState提高了操作结构恢复和选定的局部姿态指标,而MLLMs在空间关系推理方面仍存在局限性。

更新时间: 2026-10-06 14:35:49

领域: cs.AI

下载: http://arxiv.org/abs/2610.08446v1

Local exponential stability of mean-field Langevin descent-ascent and associated particle system

We study the mean-field Langevin descent-ascent (MFL-DA), a coupled optimization dynamics on the space of probability measures for entropically regularized two-player zero-sum games, together with its associated interacting particle system. For general nonconvex-nonconcave payoffs, Wang and Chizat (COLT 2024) asked whether the original single-timescale MFL-DA converges to the mixed Nash equilibrium and, if so, at what rate. We prove a local affirmative answer in Wasserstein space: if the initial datum is sufficiently close to the mixed Nash equilibrium, then the mean-field dynamics converges to it exponentially fast at a quantitative rate. We further show that the finite-$N$ particle system inherits this stability up to times exponential in $N$, with an $N$-independent exponential rate modulo a finite-particle error floor. Combined with the recent counterexample of Mourrat and Pillaud-Vivien for MFL-DA, which shows that global convergence cannot hold in general, our theorem completes the positive local counterpart of the Wang-Chizat question: the mixed Nash equilibrium has a robust basin of attraction, stable under both the mean-field flow and its finite-particle approximation.

Updated: 2026-10-06 14:33:29

标题: 局部指数稳定性的均场兰治文下降-上升和相关粒子系统

摘要: 我们研究了均场朗之万下降-上升(MFL-DA),这是在熵正则化的两人零和博弈概率测度空间上的耦合优化动态,以及其相关的相互作用粒子系统。对于一般的非凸非凹报酬,Wang和Chizat(COLT 2024)问原始的单时间尺度MFL-DA是否收敛到混合纳什均衡,并且如果是,则以什么速度。我们在Wasserstein空间中证明了一个局部肯定的答案:如果初始数据足够接近混合纳什均衡,那么均场动力学以定量速率指数级地收敛到该均衡。我们进一步展示有限-$N$粒子系统在$N$指数时间内继承了这种稳定性,具有一个$N$无关的指数速率,除了有限粒子误差。结合Mourrat和Pillaud-Vivien最近对MFL-DA的反例,表明总体收敛通常无法成立,我们的定理完成了Wang-Chizat问题的正面局部对应:混合纳什均衡具有一个稳定的吸引域,对均场流和有限粒子近似均稳定。

更新时间: 2026-10-06 14:33:29

领域: cs.LG,math.AP,math.OC,math.PR

下载: http://arxiv.org/abs/2602.01564v3

EMHO: EMbodied Agent Harness Optimization via Experience Traces

Improving embodied agents often focuses on optimizing the underlying model through training, while the surrounding agent harness that controls planning, context, and tool use is typically engineered. We ask whether this harness can instead improve itself directly from experience traces under sparse environmental feedback. We propose EMbodied Agent Harness Optimization (EMHO), a self-evolving framework that keeps the embodied model frozen and iteratively revises its harness by analyzing execution trajectories and prior harness history. EMHO optimizes beyond skills or recovery prompts, modifying how the agent monitors progress, uses vision tools, grounds observations, and responds to failures. To support multiple subtasks with a single harness, we introduce EMHO-Merge, which addresses trade-offs in jointly optimizing a single shared harness across subtasks by using episode-level gains and losses to guide evidence-supported refinement of when and how revised behaviors are applied. We evaluate EMHO on EmbodiedBench across navigation and manipulation tasks, and EMHO consistently improves task success for both Qwen 9B and 27B models. Qualitative analysis shows that EMHO goes beyond recovering from failures and unproductive actions to reshape how the embodied agent interprets and interacts with its environment.

Updated: 2026-10-06 14:30:42

标题: EMHO: 通过经验轨迹实现EMbodied Agent Harness优化

摘要: 改进具身代理通常侧重于通过训练优化基础模型,而控制规划、上下文和工具使用的周围代理则通常是经过工程设计的。我们问是否这种控制装置可以直接通过稀疏环境反馈中的经验痕迹来改进自身。我们提出了一种自我演化框架EMbodied Agent Harness Optimization(EMHO),它通过分析执行轨迹和先前的控制装置历史来保持具身模型冻结并迭代地修改其控制装置。EMHO不仅优化技能或恢复提示,还修改了代理如何监控进展、使用视觉工具、将观察结果落实以及如何应对失败。为了支持单一控制装置下的多个子任务,我们引入了EMHO-Merge,通过使用每集收益和损失来引导证据支持的修订,从而解决了在子任务间共同优化单一共享控制装置时的权衡问题。我们在EmbodiedBench上评估了EMHO在导航和操作任务中的表现,EMHO始终提高了Qwen 9B和27B模型的任务成功率。定性分析显示,EMHO不仅仅是从失败和无效操作中恢复过来,还改变了具身代理如何解释和与其环境互动。

更新时间: 2026-10-06 14:30:42

领域: cs.AI

下载: http://arxiv.org/abs/2610.08432v1

AEGIS: Runtime-Guided GPU Collocation for Multi-Tenant Deep Learning Training

Deep learning training commonly runs on shared multi-tenant GPU servers, where exclusive allocation provides isolation but can leave resources underutilized and increase queueing time. Collocation can improve efficiency, but interference-agnostic placement may cause severe slowdowns, while inaccurate memory information can lead to out-of-memory (OOM) failures. We present AEGIS, a server-scale runtime scheduling system for controlled collocation of deep learning training workloads on shared multi-GPU servers. AEGIS integrates memory feasibility, post-placement observation, runtime-pressure filtering, placement, and OOM-aware recovery in a single scheduling loop. After placement, AEGIS observes workload activity before permitting further collocation, then uses low-overhead telemetry to determine whether a GPU can safely accept additional work. OOM failures trigger retries under progressively safer memory conditions, eventually falling back to exclusive execution. This online approach avoids costly offline pairwise compatibility profiling. We evaluate AEGIS using vision, Transformer, recommendation, and LLM-style workloads across three production-derived traces. AEGIS reduces geometric-mean makespan by 16% relative to Lucid, 21% relative to Horus, and 27% relative to exclusive allocation. Sensitivity studies show that activity-anchored observation and runtime-pressure filtering balance conservative isolation against interference-agnostic collocation, improving makespan while limiting sharing-induced per-task slowdown.

Updated: 2026-10-06 14:30:27

标题: AEGIS:多租户深度学习训练的运行时引导GPU共存

摘要: 深度学习训练通常在共享的多租户GPU服务器上运行,独占分配提供隔离性,但可能会导致资源利用不足和增加排队时间。共享可以提高效率,但不考虑干扰的放置可能导致严重的减速,而不准确的内存信息可能导致内存不足(OOM)故障。 我们提出了AEGIS,一个用于在共享多GPU服务器上对深度学习训练工作负载进行受控排列的服务器规模运行时调度系统。AEGIS整合了内存可行性、放置后观察、运行时压力过滤、放置和OOM感知恢复在一个单一调度循环中。放置后,AEGIS在允许进一步的共享之前观察工作负载活动,然后使用低开销的遥测来确定GPU是否可以安全地接受额外的工作。OOM故障会触发在逐渐更安全的内存条件下重试,最终回退到独占执行。这种在线方法避免了昂贵的离线成对兼容性分析。 我们使用视觉、Transformer、推荐和LLM风格的工作负载跨三个生产衍生的跟踪来评估AEGIS。相对于Lucid,AEGIS将几何平均工期缩短了16%,相对于Horus缩短了21%,相对于独占分配缩短了27%。敏感性研究显示,活动锚定观察和运行时压力过滤在保守隔离与干扰无关的共享之间取得平衡,提高了工期,同时限制了由共享引起的每个任务的减速。

更新时间: 2026-10-06 14:30:27

领域: cs.DC,cs.LG,cs.PF

下载: http://arxiv.org/abs/2508.19073v4

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint's canonical coordinates, residual conversion covers the other changes, and the serving runtime's native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40$\times$ faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.

Updated: 2026-10-06 14:30:06

标题: NeMo-DCR:万亿参数规模下可扩展代理强化学习的比特精确增量压缩重新适应

摘要: 主动强化学习(RL)将训练与推出分开,因此每次策略更新必须在下一批之前到达推出集群。在两个AWS区域之间传输完整的1T检查点以进行权重同步(重新拟合)需要87.5分钟。BF16训练的测量结果显示,大约有1%的权重在每一步中改变其存储值。最近的系统利用这种稀疏性,但在位置、准确性或效率方面表现不佳:它们重新实现位置规则,组装完整张量,通过算术重新建立值,或使用跨集群集体,但没有一个完全从中间重新拟合失败中恢复过来。 我们提出了NeMo-DCR(Delta-压缩重新拟合),它只发送更改,但是精确无误:接收器获得与密集重新拟合相同的参数和缓冲位。对于位置,固定的仿射映射将训练碎片中的更改投影到检查点的规范坐标中,残差转换涵盖其他更改,并且服务运行时的本机加载器将所有更改放置在接收器存储中。对于准确性,可压缩的异或掩码携带仿射更改,其投影和加载器保持存储位,而覆盖则携带其他更改。接收器在原地应用两者,重试覆盖部分写入,联合提交将策略绑定到下一个增量的基线。对于效率,对象存储或中继树在增量构建过程中流式传输有效负载,而无需跨集群集体。即使在3%和5%的更改率下,NeMo-DCR对30B-1T模型的重新拟合比仅传输全检查点参考快12-40倍。在3%的情况下,1T中继树重新拟合只需150秒,而不是87.5分钟,使得在万亿参数规模下的跨集群主动RL变得实际。

更新时间: 2026-10-06 14:30:06

领域: cs.DC,cs.AI

下载: http://arxiv.org/abs/2610.08430v1

Symmetry-Aware Feature Learning: A Polynomial Separation for Multi-Index Models

We establish a polynomial sample complexity separation between symmetry-aware and symmetry-agnostic feature learning. We study growing-rank multi-index models with high-dimensional Gaussian covariates in $\mathbb{R}^d$ and $r=Θ(d^δ)$ teacher directions forming a cyclic symmetry orbit, where $0<δ<1/2$. We compare three ways of exploiting this structure: architectural weight sharing, data augmentation over the full symmetry group, and learning without access to the symmetry. In particular, we analyze a symmetry-tied convolutional network, an untied network, and the same untied network trained with full-group data augmentation, using spherical online SGD with correlation loss. For a class of polynomial links with information exponent $p\ge3$, we prove matching sample complexity bounds up to logarithmic factors: the tied and augmented learners achieve weak directional recovery in $\widetildeΘ(d^{p-1})$ samples, whereas the symmetry-agnostic learner requires $\widetildeΘ(rd^{p-1})$. For the pure quadratic Hermite link, the same separation holds for weak recovery of the teacher subspace, with sample complexities $\widetildeΘ(d)$ and $\widetildeΘ(rd)$, respectively. Thus, full-group data augmentation matches the sample efficiency of architectural weight sharing, and both provide a polynomial advantage over training without symmetry. For $p\ge3$, the proof reveals a two-stage mechanism: fluctuations at initialization select one direction in the teacher orbit, after which localized growth amplifies its overlap to the weak recovery scale while competing overlaps remain near their initialization scale.

Updated: 2026-10-06 14:23:00

标题: 对称感知特征学习:多指数模型的多项式分离

摘要: 我们在对称感知和对称不可知特征学习之间建立了一个多项式样本复杂度的区别。我们研究了在$\mathbb{R}^d$中具有高维度高斯协变量的增长秩多指数模型,其中$r=Θ(d^δ)$个教师方向形成一个循环对称轨道,其中$0<δ<1/2$。我们比较了利用这种结构的三种方式:架构权重共享、对整个对称群进行数据增强,以及在没有对称性访问的情况下学习。特别是,我们分析了一个对称绑定的卷积网络,一个未绑定的网络,以及使用球面在线SGD和相关损失训练的相同未绑定网络,对于信息指数$p\ge3$的一类多项式链接,我们证明了匹配的样本复杂度上限至对数因子:绑定和增强学习者在$\widetildeΘ(d^{p-1})$样本中实现了弱方向恢复,而对称不可知学习者需要$\widetildeΘ(rd^{p-1})$个样本。对于纯二次Hermite链接,对教师子空间的弱恢复具有相同的分离性,样本复杂度分别为$\widetildeΘ(d)$和$\widetildeΘ(rd)$。因此,完整的对称群数据增强与架构权重共享的样本效率相匹配,两者都比没有对称性的训练提供了多项式优势。对于$p\ge3$,证明揭示了一个两阶段机制:初始化时的波动选择教师轨道中的一个方向,之后局部增长将其重叠放大至弱恢复尺度,同时竞争性重叠保持在其初始化尺度附近。

更新时间: 2026-10-06 14:23:00

领域: cs.LG,math.PR,stat.ML

下载: http://arxiv.org/abs/2610.08420v1

Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as advanced LipSync generation methods not only achieve better lip synchronization but also eliminate visual artifacts. An important reason is that they overlook an inherent biological coupling between lip movements and head poses in natural speech videos. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as the fingerprint of models, enabling source tracing. We conduct extensive experiments on two challenging LipSync datasets as well as our own proposed large-scale and multi-generator dataset. LipDA achieves over 97\% AUC in detection and 97.5\% accuracy in model attribution, significantly outperforming existing methods. Code and the proposed LipSync-A dataset are available at https://github.com/AnsonShe/LipDA.

Updated: 2026-10-06 14:21:50

标题: 阿里阿德涅之线的LipSync:通过嘴唇动作与头部姿势之间的不一致性揭示伪造

摘要: 最近LipSync生成技术的进步导致了高度逼真的视频的制作,带来了严重的社会风险。然而,现有的防御策略在对抗LipSync伪造方面存在困难,因为先进的LipSync生成方法不仅实现了更好的嘴唇同步,还消除了视觉伪迹。一个重要的原因是它们忽视了自然语音视频中嘴唇运动和头部姿势之间的固有生物耦合。在本文中,我们提出了LipDA,一个新颖的联合LipSync检测和归因框架,利用头部和嘴唇之间的不一致性。对于检测,该框架学习量化真实与伪造视频之间的嘴唇和姿势特征的差异。对于归因,我们的方法旨在捕捉作为模型指纹的独特时间动态和音视频同步模式,实现源追踪。我们在两个具有挑战性的LipSync数据集以及我们自己提出的大规模和多生成器数据集上进行了广泛实验。LipDA在检测方面取得了超过97\%的AUC,模型归因的准确率达到了97.5\%,明显优于现有方法。代码和提出的LipSync-A数据集可在https://github.com/AnsonShe/LipDA 上找到。

更新时间: 2026-10-06 14:21:50

领域: cs.CV,cs.CR

下载: http://arxiv.org/abs/2610.08417v1

Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals

Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue. We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry. Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0<->MuSiQue, 0.77-0.90). Probes for epistemic "known-unknowns" transfer poorly to math, but this separation weakens under lexical controls and changes with layer and coordinate system, so it remains unresolved. Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier. A gate on the calibrated probe, with no model fine-tuning, fires on underspecified turns far more precisely than chance, and its end-task success comes within 0.08 of a gate given the true labels. Yet across four models it does not reliably beat vanilla generation or prompted consolidation. The remaining gap lies mostly in how models use a clarification, not in detection.

Updated: 2026-10-06 14:18:13

标题: 知道何时不回答:潜在的不明确信号的跨领域和多轮泛化

摘要: 大型语言模型通常会回答那些从提供的信息中无法回答的问题,在对话中它们会在足够的信息还未被提供时就给出答案。无法回答性可以从隐藏状态线性解码,但目前尚不清楚它的哪些形式共享一种表示,以及这个信号在对话中是否有用。我们提出了一个以转向标记的多轮基准测试(423个对话,1,661个标记的转向状态)和一个带有模拟用户回答澄清问题的评估工具,并将它们与六个数据集和六个开放权重的LLM一起使用,以测试对无法回答性的探测能力有多远。探测器在共享无法回答性基础的数据集之间稳健地转移:数学问题中缺失信息(AUROC 0.77-0.97)和一个段落中的缺失信息(SQuAD 2.0<->MuSiQue,0.77-0.90)。关于认知“已知-未知”方面的探测器在数学问题中的转移效果较差,但在词汇控制下分离变弱,并且随着层和坐标系统的变化而变化,因此仍未解决。单轮探测器在零样本情况下无法检测对话何时变得可回答;结构内的探测器可以恢复这一情况,但效果不比词袋分类器好。一个经过校准的探测器门,在没有模型微调的情况下,比随机更精确地触发不明确的转向,并且其最终任务成功率与给定真实标签的门之间的差距仅为0.08。然而,在四个模型中,它并不能可靠地击败纯生成或提示整合。剩下的差距主要在于模型如何使用澄清,而不是在探测方面。

更新时间: 2026-10-06 14:18:13

领域: cs.CL,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.08413v1

Decoy and disclosure radii of invariant shape descriptors

A recognizer that compares rotation-invariant descriptors sees a surface only up to the fiber of the descriptor. We measure this fiber by its radius in the orbit distance from the enrolled surface. A large radius admits decoys, that is, distant shapes that pass the matcher. A small radius discloses the enrolled shape to anyone who captures the stored value. For star-shaped surfaces truncated to spherical harmonics of degree at most $L$, with $n$ coefficients, a descriptor of generic rank $r$ has generic fibers of dimension $n-3-r$ modulo rotations. The standard pool of band powers, even bispectra, and three invariants of the degree-three band therefore admits decoy families of dimension $5$, $13$, $20$ at $L=4,6,8$. Its rank first reaches $n-3$ at $L=16$, and a mirror decoy remains at every $L$. The odd bispectra remove the mirror decoy generically for $L \geq 4$. Yet at fixed mean radius the same pool determines the enclosed volume exactly, and it does not determine whether a surface meets a clearance requirement. We certify two cases by exact and interval arithmetic. At $L=6$ a decoy matches all $32$ invariants to relative precision $2 \cdot 10^{-18}$ at orbit distance at least $0.87$ times the norm of the enrolled tuple. For the radar shape model of asteroid (101955) Bennu, the pool recovers the modeled volume, misses the handedness, and leaves the keep-out radius uncertain by more than $7 \, \mathrm{m}$.

Updated: 2026-10-06 14:16:33

标题: 特征形状描述符的诱饵和披露半径

摘要: 一个比较旋转不变描述符的识别器只能看到表面,直到描述符的纤维为止。我们通过从已注册表面的轨道距离中的半径来测量这个纤维。较大的半径允许假冒者,即远程形状可以通过匹配器。较小的半径将已注册形状展示给任何捕获存储值的人。对于截断为最大$L$次球谐函数的星形表面,带有$n$个系数的通用秩$r$的描述符在旋转模数下具有维度为$n-3-r$的通用纤维。标准带功率池,甚至双频谱,以及三个三阶带不变性因此在$L=4,6,8$时允许维度为$5$,$13$,$20$的假冒家族。其秩在$L=16$时首次达到$n-3$,并且在每个$L$处都留下一个镜像假冒者。奇数次双谱在$L \geq 4$时通常消除镜像假冒家族。然而,在固定的平均半径下,相同的池确切地确定了封闭体积,并且它不确定表面是否符合间隙要求。我们通过精确和区间算术证明了两种情况。在$L=6$时,一个假冒者可以将所有$32$个不变性匹配到相对精度为$2 \cdot 10^{-18}$,轨道距离至少为已注册元组的范数的$0.87$倍。对于小行星(101955)本努的雷达形状模型,该池可以恢复建模的体积,错过了手性,并且使禁入半径的不确定性超过$7 \, \mathrm{m}$。

更新时间: 2026-10-06 14:16:33

领域: cs.CV,astro-ph.EP,cs.CR,math.AG

下载: http://arxiv.org/abs/2610.08410v1

Case-Level Verification in Scanner-LLM Cascades: Overcoming the Alert Aggregation Bottleneck to Expand the FRR-TPR Trade-off Space

Dynamic Application Security Testing (DAST) scanners achieve high recall but also produce a large number of false positives, resulting in substantial manual triage costs. Large Language Models (LLMs), when used for independent detection, achieve extremely high recall (95.4%-100%) but also exhibit prohibitively high false positive rates (49.6%-85.0%), precluding their use as standalone replacements for scanners. A natural solution is a two-stage cascade consisting of scanner detection followed by LLM verification. However, a verification-granularity issue that has long been overlooked in practice creates a structural bottleneck: alert aggregation binds multiple true and false cases into a shared decision unit, such that removing a false positive inevitably eliminates true positives aggregated within the same alert group. This creates a trade-off bottleneck between the False-positive Reduction Rate (FRR) and the True-positive Rate (TPR). We formalize this bottleneck by showing that the alert-level false-positive set is a subset of the case-level false-positive set, and introduce a Case-Level, per-case verification strategy that shifts the decision granularity from the alert level to the instance level, independently replaying HTTP requests and making an independent determination for each detected case. Evaluation on the dual testbeds of Damn Vulnerable Web Application (DVWA) and WebGoat shows that the empirically best Alert-Level operating point achieves FRR=42.86% (TPR=51.7%). The Case-Level Baseline achieves FRR=47.6%, an improvement of 4.7 percentage points (+4.7 pp), while the Case-Focused Evidence Verification Prompt (CEV-Prompt) increases TPR from 55.2% to 62.1% at the same FRR.

Updated: 2026-10-06 14:15:45

标题: 在扫描仪-LLM级联中的案例级验证:克服警报聚合瓶颈以扩展FRR-TPR权衡空间

摘要: 动态应用安全测试(DAST)扫描器实现了很高的召回率,但也产生了大量的误报,导致了巨大的手动分类成本。当独立用于检测时,大型语言模型(LLMs)实现了极高的召回率(95.4%-100%),但也展现出了极高的误报率(49.6%-85.0%),使其无法作为扫描器的独立替代品。一个自然的解决方案是一个由扫描器检测和LLM验证组成的两阶段级联。然而,在实践中长期被忽视的一个验证粒度问题造成了一个结构瓶颈:警报聚合将多个真实和错误案例绑定到一个共享的决策单元,因此消除一个误报必然会消除在同一警报组中聚合的真实阳性。这造成了误报减少率(FRR)和真阳性率(TPR)之间的权衡瓶颈。我们通过显示警报级别的误报集是案例级别误报集的子集,形式化这个瓶颈,并引入了一个案例级别、每个案例验证策略,将决策粒度从警报级别转移到实例级别,独立重放HTTP请求,并对每个检测到的案例做出独立的判断。在Damn Vulnerable Web Application (DVWA)和WebGoat的双测试床上进行评估显示,经验上最佳的警报级别工作点实现了FRR=42.86%(TPR=51.7%)。案例级别基线实现了FRR=47.6%,提高了4.7个百分点(+4.7 pp),而案例聚焦证据验证提示(CEV-Prompt)将TPR从55.2%提高到62.1%,保持相同的FRR。

更新时间: 2026-10-06 14:15:45

领域: cs.CR

下载: http://arxiv.org/abs/2610.08406v1

Learning from Failures: A Failure-Driven Prompt Refinement for LLM-Based Vulnerability Analysis

Large Language Models have emerged as promising tools for software vulnerability analysis, but their effectiveness depends heavily on prompt design. Existing research primarily compares prompting strategies using aggregate performance metrics, providing limited insight into why models fail or how prompts can be improved systematically. We propose Failure-Driven Prompt Refinement (FDPR), a methodology that analyzes recurring model failures to guide evidence-based prompt refinement. Using the Damn Vulnerable Java Application (DVJA), we identify recurring failure modes, including false positives, false negatives, unsupported reasoning, and CWE misclassification, and translate them into targeted prompt refinements. We then evaluate the resulting prompt on the Juliet Test Suite and perform cross-model validation to assess generalizability. The results show that failure-driven refinement improves the reliability of LLM-based vulnerability analysis while yielding reusable prompt design principles. More broadly, this work demonstrates that recurring model failures provide a principled foundation for prompt engineering, enabling the systematic development of more reliable LLM-based vulnerability analysis systems.

Updated: 2026-10-06 14:15:44

标题: 从失败中学习:基于LLM的漏洞分析的基于失败的提示优化

摘要: 大型语言模型已经成为软件漏洞分析的有希望的工具,但其有效性很大程度上取决于及时的设计。现有研究主要通过聚合性能指标比较提示策略,提供了有限的洞察力,无法解释模型失败的原因或如何系统地改进提示。我们提出了基于失败驱动的提示细化(FDPR)方法,该方法分析重复的模型失败以指导基于证据的提示细化。使用Damn Vulnerable Java Application(DVJA),我们确定了重复的失败模式,包括假阳性、假阴性、不支持的推理和CWE错误分类,并将它们转化为有针对性的提示细化。然后我们在Juliet测试套件上评估生成的提示,并进行跨模型验证以评估泛化性。结果表明,基于失败驱动的细化提高了基于LLM的漏洞分析的可靠性,同时产生可重复使用的提示设计原则。更广泛地说,这项工作表明,重复的模型失败为提示工程提供了原则性基础,使得能够系统地开发更可靠的基于LLM的漏洞分析系统成为可能。

更新时间: 2026-10-06 14:15:44

领域: cs.SE,cs.AI,cs.CR

下载: http://arxiv.org/abs/2610.08405v1

The Terminal Representation in Reinforcement Learning

Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL). Two well established approaches are through the successor representation (SR) and the default representation (DR). The SR encodes states by the future trajectories they induce, capturing information flow decoupled from reward. The DR builds on this by weighting trajectories with reward, integrating credit-assignment structure into the representation. Eigenvectors of both representations have been used to support a range of downstream tasks -- including option discovery, reward shaping, transfer learning, and exploration. We introduce a structurally distinct formulation: the terminal representation (TR). The TR encodes reward-weighted trajectories similarly to the DR, but can be learned as a lower-dimensionality object, and can be used directly for the mentioned applications without eigenvector computations. Eigendecomposition also imposes the assumption of symmetric transition dynamics, which the TR can bypass. In this work we develop the theoretical foundations of the TR: its derivation, convergence of two learning algorithms, its use for zero-shot compositionality, and equivalences between alternative reward formulations. We further show the TR is embedded in the top DR eigenvector, allowing it to capture the same underlying knowledge without eigendecomposition. Additionally, we provide empirical evidence of the TR as a viable alternative to existing representations in subsidiary applications, while requiring less computational overhead to learn, store, and use.

Updated: 2026-10-06 14:15:29

标题: 在强化学习中的终端表示

摘要: 表征学习是强化学习(RL)中时空抽象的强大工具。两种成熟的方法是通过继承表示(SR)和默认表示(DR)。SR通过它们引发的未来轨迹对状态进行编码,捕获与奖励无关的信息流。DR在此基础上加权轨迹与奖励,将信用分配结构整合到表示中。这两种表示的特征向量已被用于支持一系列下游任务,包括选项发现、奖励塑造、迁移学习和探索。我们引入了一个结构上独特的公式:终端表示(TR)。TR类似于DR,通过奖励加权轨迹进行编码,但可以作为一个较低维度的对象进行学习,并且可以直接用于上述应用,无需特征向量计算。特征分解还假设对称转移动态,TR可以绕过此假设。在这项工作中,我们发展了TR的理论基础:它的推导,两种学习算法的收敛性,它在零-shot组合性方面的应用,以及替代奖励形式之间的等价性。此外,我们进一步展示了TR嵌入在顶级DR特征向量中,使其能够捕获相同的基础知识,而无需进行特征分解。此外,我们提供了TR作为现有表示的一个可行替代品的实证证据,同时需要更少的计算开销来学习、存储和使用。

更新时间: 2026-10-06 14:15:29

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2605.31289v3

SSR: Sparse Segment Reduction for Ternary GEMM Acceleration

Large Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs mitigate these demands through weight quantization via ternary values, achieving significant compression often with 50-90% sparsity. However, existing approaches have limitations: methods optimized for ternary weights, such as BitNet, redundant segment reduction (RSR), and its improved version RSR++, do not exploit sparsity structures, while conventional sparse formats neglect ternary characteristics, foregoing dual optimization opportunities. In this paper, we introduce Sparse Segment Reduction (SSR), a ternary matrix multiplication method designed to accelerate the inference of ternary LLMs and general Ternary Weight Networks (TWNs). SSR has a dedicated optimized ternary data format and an algorithm that systematically exploits sparsity patterns through computation trees that scale with the sparsity. SSR provides theoretical gains with asymptotically faster inference than RSR++ for sparsity above 50%, while practical evaluations reveal performance improvements across all sparsity levels. Evaluation results show that SSR achieves 2.1-11.3x speedup over RSR++ on ternary GEMM with 45-95% sparsity. Furthermore, SSR achieves 3.5-6.3x end-to-end speedup and 4.9% of memory saving over RSR++ on the Llama-3 1B model inference.

Updated: 2026-10-06 14:15:12

标题: SSR:用于三值GEMM加速的稀疏段减少

摘要: 大型语言模型(LLMs)需要大量的计算资源,限制了它们在资源受限硬件上的部署。三值LLMs通过三值权重量化减少这些需求,通常实现50-90%的稀疏性压缩。然而,现有方法存在局限性:针对三值权重优化的方法(如BitNet、冗余段减少(RSR)及其改进版本RSR++)未利用稀疏结构,而传统的稀疏格式忽略了三值特征,放弃了双重优化机会。 本文介绍了Sparse Segment Reduction(SSR),这是一种旨在加速三值LLMs和一般三值权重网络(TWNs)推理的三值矩阵乘法方法。SSR具有专门优化的三值数据格式和一种系统地利用稀疏模式的算法,通过与稀疏性一起扩展的计算树。SSR提供了理论上的收益,对于超过50%的稀疏性,其推理速度渐进地比RSR++快,而实际评估结果显示,在所有稀疏级别下性能有所提升。评估结果显示,SSR在45-95%的稀疏性下,在三值GEMM上比RSR++实现了2.1-11.3倍的加速。此外,SSR在Llama-3 1B模型推理中实现了3.5-6.3倍的端到端加速和4.9%的内存节省。

更新时间: 2026-10-06 14:15:12

领域: cs.LG

下载: http://arxiv.org/abs/2610.08403v1

VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents

Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response? Existing methods typically focus on only one level: turn-level methods evaluate complete responses but do not distinguish the decisions within them; token-level methods can propagate feedback across turns but do not explicitly model credit for each response. These complementary limitations motivate learning credit at both levels and coordinating it in a single policy update. We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic. VETTA computes advantages along both temporal sequences and combines each turn advantage with a within-response-centered token residual for PPO updates. Furthermore, to reduce value-learning cost, the critic retains only early Transformer blocks from the pretrained checkpoint used to initialize the actor. On two challenging agent benchmarks, ALFWorld and WebShop, VETTA improves success rates over PPO by 37.5% and 22.3%, respectively, with Qwen2.5-1.5B-Instruct and achieves success rates of 95.5% and 76.0%, respectively, with Qwen2.5-7B-Instruct. Critic-depth comparisons further show strong task performance with substantially lower critic-side computation. These results suggest that a compact shared critic can coordinate turn- and token-level credit to improve agent performance while keeping value estimation efficient. Code is available at https://github.com/Jiaju-Chen/VETTA-official.

Updated: 2026-10-06 14:15:11

标题: VETTA:协调多轮LLM代理的轮转和令牌级信用分配

摘要: 多轮LLM代理通常在多次交互中接收稀疏的任务反馈,同时逐个生成每个响应标记。这引发了两个相关的分配学分问题:哪些响应有助于实现结果,以及在每个响应中哪些生成决策起作用?现有方法通常只关注一个级别:轮次级方法评估完整的响应,但不区分其中的决策;标记级方法可以在轮次之间传播反馈,但不明确地模拟每个响应的学分。这些互补的限制促使在两个级别学习学分,并在单个策略更新中协调它。我们引入了VETTA,一种通过共享轻量级评论家的独立头部共同学习轮次和标记级值的学分分配方法。VETTA沿着时间序列计算优势,并将每个轮次的优势与响应内部为中心的标记残差相结合,用于PPO更新。此外,为了降低价值学习成本,评论家仅保留了从用于初始化演员的预训练检查点中提取的早期Transformer块。在两个具有挑战性的代理基准测试中,ALFWorld和WebShop,VETTA分别将成功率提高了37.5%和22.3%,使用Qwen2.5-1.5B-Instruct实现了95.5%和76.0%的成功率。评论者深度比较进一步显示出具有明显较低评论者端计算量的强大任务性能。这些结果表明,紧凑的共享评论家可以协调轮次和标记级学分,以提高代理性能,同时保持价值估计的效率。代码可在https://github.com/Jiaju-Chen/VETTA-official 上找到。

更新时间: 2026-10-06 14:15:11

领域: cs.LG

下载: http://arxiv.org/abs/2610.08402v1

GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63\%.

Updated: 2026-10-06 14:14:25

标题: GeoPID:分解和引导视觉语言模型中的视觉信息

摘要: 最近的视觉-语言模型(VLMs)在各种应用中表现出色,但它们往往未充分利用视觉信息,过分依赖文本内容。在这项工作中,我们提出了\textsc{GeoPID},这是一个无需训练的框架,从几何角度分析VLMs中的多模态信息。\textsc{GeoPID}通过视觉和文本表示子空间之间的几何关系,将信息分解为冗余、模态唯一和协同三个组成部分。通过对22个VLMs和14个基准数据集进行广泛分析,我们确认当问题强烈需要视觉基础时,正确的预测表现出更强的视觉唯一成分。基于这种几何分析,我们引入了一种有针对性的干预技术,在推理过程中有选择地增强视觉表示在视觉唯一子空间中的表现。结果,视觉基础能力得到增强,无需进行任何额外的模型参数更新,平均相对精度提升了7.63%。

更新时间: 2026-10-06 14:14:25

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.08401v1

Atom-JEPA: Joint-Embedding Predictive Architecture for 3D Atomistic Systems

Large-scale self-supervised pretraining has reshaped modern machine learning, substantially advancing the ability of language and vision models to generalize across downstream tasks. While deep learning has driven considerable progress in modeling atomistic systems in recent years, self-supervised pretraining in this domain has not yet achieved comparable downstream generalization. To address this, we introduce Atom-JEPA, a self-supervised pretraining framework that learns latent representations from unlabeled 3D structures through complementary atom-level and substructure-level objectives inspired by joint-embedding predictive architectures. We pretrain Atom-JEPA on large-scale molecular and crystalline datasets and evaluate its transfer performance by fine-tuning on a diverse set of downstream property prediction tasks. Atom-JEPA achieves state-of-the-art performance on molecular ADMET and quantum-chemical property prediction tasks, and is highly competitive in predicting the physical properties of crystalline materials. These results demonstrate the potential of latent-space predictive pretraining to support broad downstream generalization from structural data alone. Code and pretrained model checkpoints are publicly available at https://github.com/khelverskovp/atom-jepa

Updated: 2026-10-06 14:13:51

标题: 原子-JEPA:用于3D原子系统的联合嵌入预测架构

摘要: 大规模自监督预训练已经改变了现代机器学习,显著提升了语言和视觉模型在下游任务中的泛化能力。尽管深度学习近年来在建模原子系统方面取得了相当大的进展,但在这一领域的自监督预训练尚未实现可比的下游泛化。为了解决这个问题,我们引入了Atom-JEPA,一个自监督预训练框架,通过受联合嵌入预测架构启发的互补原子级和亚结构级目标,从未标记的3D结构中学习潜在表示。我们在大规模分子和结晶数据集上对Atom-JEPA进行预训练,并通过对各种下游属性预测任务进行微调来评估其转移性能。Atom-JEPA在分子ADMET和量子化学属性预测任务上实现了最先进的性能,并在预测结晶材料的物理性质方面具有很高的竞争力。这些结果表明,潜在空间预测性预训练有望支持仅从结构数据中实现广泛的下游泛化。代码和预训练模型检查点可以在https://github.com/khelverskovp/atom-jepa 上公开获取。

更新时间: 2026-10-06 14:13:51

领域: cs.LG,cs.AI,physics.comp-ph

下载: http://arxiv.org/abs/2610.08400v1

Grand Canonical Generators

We introduce Grand Canonical Generators (GCG), a generative framework that extends Boltzmann generators to the grand canonical ensemble. We present two designs. The first conditions a variable-size generative model on the chemical potential, sampling particle number and configuration jointly. The second factorizes the grand canonical distribution into a particle-number distribution and the corresponding canonical Boltzmann density. This factorized formulation can use any existing Boltzmann generator for the canonical component, encodes the known linear chemical-potential dependence analytically, and yields a tractable likelihood that supports self-normalized importance sampling (SNIS). Empirically, GCG accurately reproduces grand canonical observables on a Lennard--Jones fluid and methane adsorption in a zeolite, demonstrating generalization across chemical potentials and correction via SNIS and grand canonical Monte Carlo.

Updated: 2026-10-06 14:08:53

标题: 大规范发生器

摘要: 我们引入了Grand Canonical Generators(GCG),这是一个将Boltzmann生成器扩展到巨正则系综的生成框架。我们提出了两种设计。第一种将一个可变大小的生成模型条件在化学势上,同时对粒子数量和构型进行采样。第二种将巨正则分布分解为粒子数量分布和相应的正则Boltzmann密度。这种分解形式可以使用任何现有的正则Boltzmann生成器来进行正则部分的计算,通过解析地编码已知的线性化学势依赖性,并产生一个可解析的似然函数,支持自归一化重要性采样(SNIS)。在实证研究中,GCG精确地重现了一个Lennard-Jones流体和沸石中甲烷吸附的巨正则观测量,展示了在化学势上的泛化和通过SNIS和巨正则蒙特卡洛进行修正的能力。

更新时间: 2026-10-06 14:08:53

领域: cs.LG,stat.ML

下载: http://arxiv.org/abs/2610.00683v2

Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering

Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge source for reliable reasoning. However, existing LLM-guided graph reasoning methods typically rely on hop-wise greedy or beam-style pruning during evidence retrieval. Such local decision processes are inherently myopic: evidence that appears weak near the source may become crucial only after deeper graph context is explored, causing answer-critical branches to be discarded prematurely and making the reasoning chain difficult to recover. To address this limitation, we propose Foresight-over-Graph (FoG), a foresight-aware evidence retrieval framework for knowledge base question answering (KBQA). FoG iteratively constructs a question-relevant evidence subgraph and uses far-to-near feedback to guide path exploration, and maintains a compact memory subgraph to support continued exploration. Extensive experiments on widely used KBQA benchmarks demonstrate that FoG achieves state-of-the-art performance, with a particularly large improvement of 16.58% in Hit on CWQ, while also reducing LLM calls and token usage. Our code is available at https://github.com/yhong7/FoG .

Updated: 2026-10-06 14:07:43

标题: 超越局部视野的预测-图形:基于知识库的问题回答的推理

摘要: 大型语言模型(LLMs)在问答方面表现出强大的能力,但它们在知识密集型任务上经常出现幻觉。知识图(KGs)为LLMs提供了结构化、可解释和可更新的事实基础,使它们成为可靠推理的有希望的外部知识来源。然而,现有的LLM引导的图推理方法通常在证据检索过程中依赖于跳跃式贪婪或束式修剪。这种局部决策过程本质上是短视的:在源头附近看起来薄弱的证据可能只有在探索更深的图上下文之后才变得关键,导致关键的答案分支被过早丢弃,使推理链难以恢复。为了解决这一局限性,我们提出了一种前瞻感知证据检索框架FoG,用于知识库问答(KBQA)。FoG迭代地构建一个与问题相关的证据子图,并使用远到近的反馈来指导路径探索,并维护一个紧凑的记忆子图以支持持续探索。对广泛使用的KBQA基准进行的大量实验表明,FoG实现了最先进的性能,在CWQ的Hit上特别取得了16.58%的大幅度改进,同时还减少了LLM的调用和标记使用。我们的代码可在https://github.com/yhong7/FoG 上找到。

更新时间: 2026-10-06 14:07:43

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.08388v1

Task diversity produces systematic transfer but inhibits continual reinforcement learning

Continual reinforcement learning (RL) aims to produce agents that never stop adapting to new tasks. A key question is how this interacts with the diversity of tasks an agent experiences. Prior work has shown that training on many diverse tasks leads to agents with strong zero-shot and in-context adaptation. However, this work evaluated agents after they'd stopped learning, i.e. with frozen weights. How task diversity affects an agent's ability to continue learning over a sequence of distribution shifts remains unclear. We introduce Banyan, a GPU-accelerated continual RL domain where one can parametrically control three independent axes that define a task: the map layouts an agent must navigate, the objects it must interact with, and the hierarchical structures of sub-goal dependencies. We find that increasing diversity along each axis induces systematic transfer -- that is, agents begin training on a new task distribution near the performance attained on the previous one, even when the shift changes the structure of the optimal policy. While increasing diversity improves systematic transfer, we find that too much diversity inhibits a learner's ability to continue adapting to new task distributions. As diversity increases, learners plateau in the success rate they achieve on new tasks, yet continue improving on old tasks -- even without further exposure to them. We find this phenomenon manifests across continual learning algorithms, memory architectures, architecture sizes, and in Kinetix -- a physics-based control domain. We release Banyan as a domain for running controlled experiments that study continual RL in the many-tasks regime. Code is available at https://github.com/nhshah15/banyan.

Updated: 2026-10-06 14:06:33

标题: 任务多样性会产生系统化的迁移,但会抑制持续的强化学习

摘要: 持续的强化学习(RL)旨在产生永远不会停止适应新任务的代理。一个关键问题是这如何与代理经历的任务多样性相互作用。先前的研究表明,在许多不同的任务上进行训练会导致代理具有强大的零样本和上下文适应能力。然而,这项工作在代理停止学习后评估代理,即使用固定的权重。任务多样性如何影响代理在一系列分布转变中继续学习的能力仍不清楚。我们引入了Banyan,一个GPU加速的持续RL领域,在这个领域中可以参数化地控制定义任务的三个独立轴:代理必须导航的地图布局、必须与之交互的对象以及子目标依赖的层次结构。我们发现,沿着每个轴增加多样性会引发系统性转移--也就是说,代理在新任务分布上开始训练时,性能接近于在先前任务上达到的性能,即使转变改变了最优策略的结构。虽然增加多样性可以改善系统性转移,但我们发现太多的多样性会抑制学习者继续适应新任务分布的能力。随着多样性的增加,学习者在新任务上取得的成功率会达到一个平稳期,但在旧任务上仍会继续改进--甚至在没有进一步接触的情况下。我们发现这种现象在持续学习算法、内存架构、架构大小以及基于物理的控制领域Kinetix中都存在。我们发布了Banyan作为一个领域,用于进行控制实验,研究许多任务制度下的持续RL。代码可在https://github.com/nhshah15/banyan上找到。

更新时间: 2026-10-06 14:06:33

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2606.00880v2

Decision-Focused Learning in MDPs: An Occupancy Measure Approach

In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solving a linear system over all state-action pairs, limiting its scalability. We address this by reformulating the MDP as an occupancy measure-based linear program (LP), whose feasible region is induced by predicted dynamics, and we derive a closed-form gradient by identifying the active constraints in the feasible polyhedron via the pivoting algorithm. This occupancy measure-based LP layer raises two challenges: (1) LP's solution gradient is discontinuous when active constraints change, and (2) the LP backward cost still scales with the state size, which is costly for large or continuous state spaces. We address the challenges with an augmented Lagrangian surrogate and smooth the boundary jumps by random row sketching of the constraints, and a learnable soft state-aggregation layer and its function-approximation generalization that scales the LP to large finite and continuous-state MDPs. Across multiple tasks, our methods reach lower regret than KKT-based DFL and two-stage baselines with significantly lower computation cost. The source code for all experiments is available at https://github.com/A-Eshragh/State_Aggregation_Project.

Updated: 2026-10-06 14:04:19

标题: 基于占用度度量方法的MDPs中的决策聚焦学习

摘要: 在这项工作中,我们考虑了决策聚焦学习(DFL)用于马尔可夫决策过程(MDP),现有方法通过贝尔曼方程的KKT条件进行区分,并需要在所有状态-动作对上解决线性系统,从而限制了可扩展性。我们通过将MDP重新制定为基于占用度测量的线性规划(LP)来解决这个问题,其可行区域由预测动态引起,我们通过识别可行多面体中的活动约束,通过枢轴算法推导出闭合形式的梯度。这个基于占用度测量的LP层提出了两个挑战:(1)当活动约束发生变化时,LP的解梯度是不连续的,(2)LP的反向成本仍然随着状态大小增加,对于大型或连续状态空间来说是昂贵的。我们通过增广拉格朗日替代方法解决了这些挑战,并通过随机行素描约束来平滑边界跳跃,以及一个可学习的软状态聚合层及其函数逼近泛化,将LP扩展到大型有限和连续状态的MDP。在多个任务中,我们的方法比基于KKT的DFL和两阶段基线具有更低的遗憾,并且计算成本显著更低。所有实验的源代码均可在https://github.com/A-Eshragh/State_Aggregation_Project 上找到。

更新时间: 2026-10-06 14:04:19

领域: cs.LG

下载: http://arxiv.org/abs/2610.08384v1

Lossy Compression of PDE Training Inputs: Field Reconstruction Error Does Not Order the Cost to a Trained Operator

Operator-learning benchmarks are stored at full precision and have grown to terabyte scale. Rate-distortion theory says how many bits the stored field needs, while a practitioner needs to know how accurate an operator trained on the compressed data will be. We show that the first does not determine the second, and measure why, compressing the input fields while targets and test inputs stay at full precision. A solution operator attenuates a perturbation of its input. Pushing a compressed field through a surrogate already trained at full precision measures how much of the perturbation that surrogate transmits. The fraction is consistent with the smoothing behaviour of the underlying equation, and it spans more than two orders of magnitude across PDE families. Field reconstruction error is computed before the attenuation and cannot see it. For operators trained with mean squared error it inverts 36 of 104 cost comparisons across datasets, where a probe built from the same forward passes inverts 12. Two families that PDEBench stores with identical initial conditions differ threefold downstream at identical field error. Under the relative-L2 objective of the reference recipe the separation narrows, while the ordering of the family-level median transmission factors is unchanged. After one full-precision training run, the probe evaluates an entire rate curve by forward passes alone. It ranks datasets and rates consistently across the codecs and architectures we test, while its magnitude does not transfer between them.

Updated: 2026-10-06 14:03:35

标题: PDE训练输入的有损压缩:场重建误差不能确定对训练操作员的成本

摘要: 运算学习基准以完全精度存储,并已发展到千兆字节规模。速率失真理论说明存储字段需要多少位,而从业者需要知道在压缩数据上训练的运算符会有多准确。我们展示第一个并不能确定第二个,并测量为什么,在压缩输入字段的同时,目标和测试输入保持完全精度。解决方案运算符减弱其输入的扰动。将压缩字段通过已经在完全精度上训练的替代物传递,可以衡量替代物传递多少扰动。这一比例与基础方程的平滑行为一致,并跨越两个数量级以上的PDE家族。在减弱之前计算字段重建误差,并且无法看到它。对于使用均方误差训练的运算符,它在104个成本比较中颠倒了36个,而从相同前向传递构建的探针颠倒了12个。PDEBench存储的两个具有相同初始条件的家族在相同的字段误差下三倍不同。在参考配方的相对L2目标下,分离变窄,而家族级别的中位传输因子的排序保持不变。经过一次完全精度训练运行后,探针仅通过前向传递评估整个速率曲线。它在我们测试的编解码器和架构中一致地对数据集和速率进行排名,而其大小不在它们之间传递。

更新时间: 2026-10-06 14:03:35

领域: cs.LG

下载: http://arxiv.org/abs/2610.06095v2

Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning

Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibits a fundamentally different uncertainty structure from general text: dense low-confidence runs encode epistemic knowledge gaps (dense causal chains, rare entities) rather than the sparse aleatoric stylistic variation typical of general text. Based on this discovery, we propose Balanced Fine-Tuning (BFT), a dual-scale post-training method that combines group-normalized token reweighting with sequence-level reallocation toward knowledge-dense samples exhibiting dense epistemic uncertainty. Across medical evaluation, biological reasoning, sparse-reward RL, and biological representation tasks, BFT provides more consistent gains than SFT and DFT under a shared training setup. When replacing the default closed-source backbones in GeneAgent (GPT-4o) and VCWorld (Gemini-2.5-Flash), the BFT-aligned 70B model delivers stronger performance across biological process reasoning and chemical perturbation prediction. Critically, all BFT variants further improve after subsequent GRPO with sparse rewards, while SFT and DFT degrade, suggesting that epistemic-aware post-training provides a more robust policy initialization. Beyond text generation, BFT-aligned LLMs produce more accurate and professional biomedical profile texts; after encoding these profiles with a text embedding model, the resulting representations support gene-level, cell-level, and perturbation-response tasks, suggesting that BFT-enhanced generation can facilitate biological representation and, in turn, broader biomedical downstream tasks.

Updated: 2026-10-06 14:00:27

标题: 使用平衡微调将LLMs与生物医学知识对齐

摘要: 将LLMs工程化以加速生命科学研究需要与生物医学知识进行牢固的对齐。我们观察到,生物医学文本具有与一般文本根本不同的不确定性结构:密集的低置信度运行编码认识知识缺口(密集的因果链,罕见实体),而不是一般文本典型的稀疏的偶然风格变异。基于这一发现,我们提出了平衡微调(BFT),这是一种双尺度的后训练方法,它将组归一化的标记重新加权与序列级别的重新分配相结合,以便将展示密集认识不确定性的知识密集样本。在医学评估、生物推理、稀疏奖励RL和生物表示任务中,BFT在共享训练设置下提供比SFT和DFT更一致的收益。当将GeneAgent(GPT-4o)和VCWorld(Gemini-2.5-Flash)中的默认闭源主干替换为与BFT对齐的70B模型时,生物过程推理和化学扰动预测的性能更强。至关重要的是,所有BFT变体在随后的具有稀疏奖励的GRPO之后进一步改进,而SFT和DFT则退化,这表明认识感知后训练提供了更强大的策略初始化。除了文本生成,与BFT对齐的LLMs产生更准确和专业的生物医学配置文件文本;在使用文本嵌入模型对这些配置文件进行编码后,产生的表示支持基因级、细胞级和扰动响应任务,表明BFT增强的生成可以促进生物表示,从而推动更广泛的生物医学下游任务。

更新时间: 2026-10-06 14:00:27

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2511.21075v4

Particle Monte Carlo Tree Search

Monte Carlo Tree Search (MCTS) is a widely used approach for policy improvement and action selection in Reinforcement Learning. Due to its sequential and deterministic nature, principled runtime-scaling of MCTS with parallel compute remains a major challenge. We introduce Particle MCTS (PMCTS), a parallel MCTS algorithm which is suited for neural network evaluations, designed for GPU-acceleration with batch-parallelization and retains MCTS's principled approximate policy improvement interpretation. Empirically, PMCTS scales well with parallel compute and consistently outperforms or compares well to the popular heuristic-based baselines across a range of popular discrete- and continuous-action benchmark domains, including Chess, 19x19 Go, 9x9 Go, Gardner Chess, Snake, classical control environments from Brax and LLM reasoning in Sokoban.

Updated: 2026-10-06 13:59:56

标题: 粒子蒙特卡洛树搜索

摘要: 蒙特卡洛树搜索(MCTS)是强化学习中广泛使用的一种方法,用于策略改进和动作选择。由于其顺序和确定性的特性,利用并行计算对MCTS进行合理的运行时间扩展仍然是一个主要挑战。我们引入了粒子MCTS(PMCTS),这是一种并行MCTS算法,适用于神经网络评估,设计用于GPU加速并具有批处理并行化,并保留了MCTS的合理近似策略改进解释。实证上,PMCTS在并行计算方面表现出色,并在一系列流行的离散和连续动作基准领域中始终优于或与流行的基于启发式的基准相比良好,包括国际象棋、19x19围棋、9x9围棋、加德纳国际象棋、贪吃蛇、Brax的经典控制环境以及Sokoban中的LLM推理。

更新时间: 2026-10-06 13:59:56

领域: cs.LG

下载: http://arxiv.org/abs/2605.08982v4

Accelerating the Development of PLGA In Situ Forming Depots Through AI-Driven Multi-Objective Optimization

Developing long-acting injectable formulations requires the simultaneous optimization of drug loading, release kinetics, viscosity, injectability, stability and other objectives. To navigate this multidimensional space, Corbion and Intrepid combined Corbion's diverse PURASORB bioresorbable polymer library with Intrepid Labs' proprietary AI algorithm (ANDROMEDA 1) to develop in situ forming depots for a therapeutic peptide. Over approximately 15 weeks, 181 unique formulations spanning drug loadings of 6-12% w/w were prepared and characterized through broad design-space mapping and targeted multi-objective optimization. Four lead candidate formulations were identified at 6%, 9%, and 12% w/w drug loading. Each met the predefined viscosity and injectability criteria while providing distinct 30-day in vitro release profiles. The study evaluated polymers spanning a broad range of molecular weights, including commercially available PURASORB grades and new polymers under development by Corbion to expand its polymer toolbox. ANDROMEDA 1 identified that polymers with intermediate molecular weights provided a favorable balance between sustained release and solution viscosity. Together, these findings demonstrate how integrated polymer expertise and AI-driven optimization can rapidly identify differentiated formulation candidates, focus the development space, and establish a strong data-driven foundation for further optimization and in vivo evaluation.

Updated: 2026-10-06 13:54:38

标题: 通过人工智能驱动的多目标优化加速PLGA原位形成缓释剂的开发

摘要: 开发长效注射剂需要同时优化药物载荷、释放动力学、粘度、可注射性、稳定性和其他目标。为了在这个多维空间中导航,Corbion和Intrepid结合了Corbion多样化的PURASORB可生物降解聚合物库与Intrepid Labs专有的AI算法(ANDROMEDA 1),为治疗肽开发原位形成的沉积物。在大约15周的时间内,制备了181种独特的配方,涵盖了6-12% w/w药物载荷,并通过广泛的设计空间映射和定向多目标优化进行了表征。发现了4种主要候选配方,药物载荷分别为6%、9%和12%w/w。每种配方均符合预定的粘度和可注射性标准,并提供了不同的30天体外释放曲线。该研究评估了涵盖广泛分子量范围的聚合物,包括商业可获得的PURASORB等级和Corbion正在开发的新型聚合物,以扩展其聚合物工具箱。ANDROMEDA 1发现,中等分子量的聚合物在持续释放和溶液粘度之间提供了有利的平衡。这些发现共同展示了整合的聚合物专业知识和AI驱动的优化如何能够快速识别差异化的配方候选者,聚焦开发空间,并为进一步优化和体内评估奠定坚实的数据驱动基础。

更新时间: 2026-10-06 13:54:38

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.08368v1

Evolutionary One-Step Generators: Fast and Diverse Sampling for Discrete Design

Several discrete design tasks, such as molecular discovery, require diverse collections of useful candidates at low computational cost. High validity alone does not guarantee a useful candidate library: repeatedly generating the same valid structures leaves few distinct alternatives. Training for both feasibility and diversity is challenging because many relevant criteria can only be evaluated after hard decoding. To address this challenge, we propose EGO (Evolutionary Generators with One-step inference), a framework for training compact generators directly on discrete outputs. The method combines distribution matching with structural constraints and optional diversity or history-dependent rewards, using antithetic low-rank evolution strategies without requiring criterion-specific differentiable surrogates. Once trained, the generator produces the entire graph in a single neural-network evaluation. On molecular generation benchmarks, our compact generator achieves over $50\times$ the valid-and-unique yield per estimated dense operation compared to recent one-step flow-map baselines while retaining high chemical validity. In scaffold completion, EGO achieves an observed $44.3\times$ speedup over MoLeR in generation to SMILES and produces approximately $10\times$ as many filter-passing proposals within matched time budgets for generation and screening. Beyond chemistry, EGO produces $1.54\times$ as many distinct held-out elite architectures as relaxed gradient training on NAS-Bench-101. The low generation cost may enable real-time candidate generation across discrete design tasks, supporting interactive exploration of constrained design spaces and rapid construction of candidate sets for downstream evaluation.

Updated: 2026-10-06 13:54:24

标题: 进化一步生成器:离散设计的快速多样化抽样

摘要: 这个文献摘要介绍了几个离散设计任务,比如分子发现,需要以低计算成本生成多样化的有用候选集合。仅仅高准确性并不能保证一个有用的候选库:重复生成相同有效的结构只留下很少的不同选择。同时训练可行性和多样性是具有挑战性的,因为许多相关标准只能在硬解码后进行评估。为了解决这一挑战,作者提出了EGO(具有一步推断的进化生成器)框架,用于直接训练紧凑的生成器输出。该方法结合了分布匹配和结构约束,以及可选的多样性或历史依赖奖励,使用反向低秩进化策略,而不需要特定标准的可微替代品。一旦训练完成,生成器在单次神经网络评估中生成整个图。在分子生成基准测试中,我们的紧凑生成器在估计的密集操作中实现了超过50倍的有效且独特的产量,与最近的一步流图基线相比,同时保持高化学有效性。在骨架完成中,EGO在生成到SMILES中实现了观察到的44.3倍速度提升,同时在生成和筛选的匹配时间预算中产生了大约10倍的通过滤器的提案。除了化学之外,EGO在NAS-Bench-101上产生了1.54倍的不同的保留精英架构,相比之下,它比基于放松梯度的训练更有效。低生成成本可能实现跨离散设计任务的实时候选生成,支持对受限设计空间的互动探索和快速构建用于下游评估的候选集合。

更新时间: 2026-10-06 13:54:24

领域: cs.LG,cs.NE

下载: http://arxiv.org/abs/2610.08367v1

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.

Updated: 2026-10-06 13:52:15

标题: 横断面:保持对长时间跨度LLM代理评估的可观察性

摘要: 前沿AI评估越来越多地使用开放性、主动性、长期视角的任务,其记录可以跨越数百页的复杂多智能体网络的输出和行为。因此,可观察性包络-评估者可可靠推断关于智能体行为的范围-正在缩小。语言模型助手可以帮助分类和解释智能体行为,但也给人类评估者带来了重要的分析自由度,威胁到基于语言模型的记录分析的可重复性和可审计性。Transect是一个建立在Inspect Scout之上的开源软件包,旨在帮助评估者了解长期智能体运行的展开情况,识别值得调查的行为,并将解释与记录相对照。用户在可重复使用的评估家族配置中指定任务背景和行为词汇,评判模型和分析设置分开提供。Transect的可导航报告将记录的事件、标记使用、子智能体活动和模型生成的行为标签对齐在一个共同的基于轮次的时间轴上。审阅者可以快速把握运行的叙述,追溯任何标签或事件到其源轮次,并导出基础数据表进行跨运行分析。我们在一个AI研发评估上演示了这一工作流程,生成了近1300万个标记,将智能体的工作划分为与研究技能分类、子智能体委派和互动以及标记使用相一致的行为阶段。综合视图显示了对运行工作和手稿制作的关注,几乎没有持续的假设生成阶段的证据-这可能是高质量科学产出的必要组成部分。Transect的灵活、可定制的记录分析流水线将使评估者能够跟上更长、更复杂、更频繁的AI评估,同时支持科学严谨性、透明性和可重复性。

更新时间: 2026-10-06 13:52:15

领域: cs.AI

下载: http://arxiv.org/abs/2610.08364v1

Explainable Failure Prediction and Prevention in Maritime

Maritime systems operate in highly dynamic environments where unexpected equipment failures can compromise safety, reliability, and operational efficiency. Recent advances in artificial intelligence (AI), machine learning, digital twins, and predictive maintenance enable proactive failure prediction and prevention. However, ensuring trustworthy and explainable decision-making remains a major challenge in safety-critical maritime applications. This chapter reviews key AI technologies required for explainable failure prediction and prevention in maritime systems and presents a conceptual architecture capable of supporting autonomous or human-in-the-loop corrective actions. This architecture integrates data acquisition, time-series forecasting, anomaly detection, risk assessment, decision-making, and explainable AI into a closed-loop framework. With reference to the architectural components, a review and discussion of relevant maritime studies is performed, outlining their methods, advantages, and limitations. Furthermore, it highlights current challenges, including uncertainty and robustness, model generalization, explainability, limited availability of maritime datasets, and operational deployment, and identifies future research directions toward trustworthy AI-assisted maritime decision-making.

Updated: 2026-10-06 13:51:58

标题: 可解释的海事故障预测和预防

摘要: 海事系统在高度动态的环境中运作,意外设备故障可能会影响安全性、可靠性和运行效率。最近人工智能(AI)、机器学习、数字双胞胎和预测性维护的进步使得预测和预防故障变得更为主动。然而,在安全关键的海事应用中,确保可信赖且可解释的决策仍然是一个重要挑战。本章回顾了海事系统中需要解释性故障预测和预防的关键AI技术,并提出了一个能够支持自主或人机协同纠正措施的概念架构。该架构将数据获取、时间序列预测、异常检测、风险评估、决策制定和可解释的AI集成到一个闭环框架中。通过参考架构组件,对相关海事研究进行了审查和讨论,概述了它们的方法、优势和局限性。此外,它强调了当前的挑战,包括不确定性和鲁棒性、模型泛化、可解释性、海事数据集的有限可用性以及操作部署,并确定了未来研究方向,以实现可信赖的AI辅助海事决策。

更新时间: 2026-10-06 13:51:58

领域: cs.AI

下载: http://arxiv.org/abs/2610.08363v1

Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration

Post-training quantization is a standard route to fitting vision transformers (ViTs) into edge compute and memory budgets, yet quantized models become especially brittle under distribution shift. Test-time adaptation (TTA) addresses such shifts without labels, but most existing approaches are poorly aligned with the constraints of quantized inference. Prevailing TTA methods recover accuracy through backpropagation, while backprop-free methods often still incur overhead from extra forward passes or parameter updates, and lightweight feature- or logit-level methods recover only part of the loss. Across these approaches, a quantization-specific failure mode that amplifies the drop is not directly targeted: under shift, activations occupy frozen quantizers' calibrated ranges differently, distorting their code distribution. We propose Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters. QuAR recalibrates activations at the input to a frozen quantizer, mapping the test stream's running per-channel statistics back toward the source calibration. On ImageNet-C with ViT-B, QuAR achieves the highest mean accuracy among state-of-the-art backprop-free TTA methods at 3-, 4-, 6- and 8-bit weight/activation precision, outperforming the strongest baseline by 2.28 points at 8 bits and 4.00 at 3 bits, with 46% lower latency and a memory overhead of only 0.17 MB (0.01% of peak inference memory). Analysis and diagnostics trace the gain to a reduced per-channel mismatch at these quantizers, which restores the code distribution the baselines leave unchanged or distort further. A single fixed configuration remains ahead across continual streams, non-i.i.d. label shift, seven out-of-distribution suites, and three other backbones.

Updated: 2026-10-06 13:48:33

标题: 通过单次量化器对齐重新校准实现量化 ViTs 的测试时间适应

摘要: 后训练量化是将视觉变换器(ViTs)适配到边缘计算和内存预算的标准路径,然而在分布转移下,量化模型变得特别脆弱。测试时间适应(TTA)解决这种转移,而无需标签,但大多数现有方法与量化推断的约束不相符。流行的TTA方法通过反向传播恢复准确性,而无反向传播的方法通常仍会产生来自额外前向传递或参数更新的开销,而轻量级特征或逻辑级方法仅恢复部分损失。在这些方法中,没有直接针对放大降低的量化特定失效模式:在转移下,激活以不同方式占据冻结量化器的校准范围,扭曲其编码分布。我们提出了Quantizer-Aligned Recalibration(QuAR),这是一种专为量化ViTs定制的单通道TTA方法,既不进行反向传播,也不更新任何模型参数。QuAR重新校准冻结量化器输入处的激活,将测试流的每个通道的运行统计数据映射回源校准。在ImageNet-C上的ViT-B,QuAR在3位、4位、6位和8位权重/激活精度上实现了最高的平均准确性,在8位上超过了最强基线2.28个点,在3位上超过了4.00个点,延迟降低了46%,内存开销仅为0.17 MB(占峰值推理内存的0.01%)。分析和诊断显示,这些量化器的每个通道不匹配减少了,这恢复了基线未改变或进一步扭曲的编码分布。在持续流、非i.i.d.标签转移、七个超出分布套件和三个其他骨干中,一种固定配置始终处于领先地位。

更新时间: 2026-10-06 13:48:33

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.08358v1

Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals

Flow-matching models start from an isotropic Gaussian source, the standard choice when the correlation structure of the data is unknown in advance. For multi-channel brain recordings, however, part of this structure is known in advance. Electrodes sit at fixed positions on the head, and volume conduction through the skull and scalp makes nearby electrodes co-vary in a way that is shared across subjects. Existing EEG generative models nonetheless leave the network to learn this from scratch. We put this structure into the source instead. From the sensor coordinates alone, we build a k-nearest-neighbor graph and take a graph-Matérn function of its Laplacian as the source covariance, so the flow starts from spatially coherent patterns rather than channel-independent noise. The change adds no learned parameters, works with any coupling and any drift network, and uses the same three hyperparameters on every dataset. Across eight EEG datasets and four flow-matching methods, the graph-Matérn source lowers the spectral discrepancy between generated and real signals in the five clinical bands (PSD-KL) on most datasets. PSD-KL falls by 12% to 17% in geometric mean over datasets depending on the method and by up to 40% on PhysioNet-MI, the densest montage. We show that the improvement stems from the spatial eigenvectors of the local graph of sensor positions, since randomizing the eigenvectors while preserving the eigenvalue spectrum eliminates the gain. Furthermore, a prior fitted directly to the empirical data covariance performs worse than isotropic noise. The same construction applies unchanged to MEG, intracranial EEG with patient-specific grids, and a traffic-sensor network, lowering PSD-KL for every method on each. https://jd730.github.io/projects/GraphPrior

Updated: 2026-10-06 13:44:57

标题: 传感器几何形状作为多通道脑信号流匹配先验的翻译

摘要: 流匹配模型从各向同性高斯源开始,这是在数据的相关结构事先未知时的标准选择。然而,对于多通道脑电记录,部分结构是事先已知的。电极固定在头部位置,通过头骨和头皮的体积传导使附近的电极以一种在被试间共享的方式相关。然而,现有的脑电生成模型仍然让网络从头开始学习这一结构。我们将这种结构放入源中。仅从传感器坐标开始,我们构建一个k-最近邻图,并将其拉普拉斯矩阵的图马特恩函数作为源协方差,因此流从空间一致的模式开始,而不是独立于通道的噪声。这种改变不会增加学习参数,适用于任何耦合和漂移网络,并在每个数据集上使用相同的三个超参数。在八个脑电数据集和四种流匹配方法中,图马特恩源在大多数数据集上降低了生成信号与真实信号在五个临床频带(PSD-KL)上的谱差距。 PSD-KL在几何平均上降低了12%至17%,具体取决于方法,并在PhysioNet-MI上最多降低了40%,这是最密集的装置。我们展示了改进来自于传感器位置的局部图的空间特征向量,因为随机化特征向量同时保留特征值谱会消除增益。此外,直接适配于经验数据协方差的先验比各向同性噪声表现更差。相同的构造不变地适用于脑磁图、患者特定网格的颅内脑电以及交通传感器网络,降低了每种方法在每种情况下的PSD-KL。

更新时间: 2026-10-06 13:44:57

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.08355v1

Uncertainty Quantification Is Indispensable for Reliable Connectome-Based Graph Learning: A Narrative Review and Case Study

While graph neural networks (GNNs) have shown substantial promise in connectome-based diagnostic classification, deterministic models inevitably suppress pipeline-induced noise and model ambiguities, yielding overconfident predictions. Although uncertainty quantification (UQ) is widely adopted in voxel-level segmentation, its role in connectomic graph learning remains largely unaddressed. This paper presents a comprehensive narrative review of UQ frameworks tailored to connectome graph learning alongside an empirical case study demonstrating the perils of uncalibrated predictions. We delineate sources of aleatoric and epistemic uncertainty across neuroimaging pipelines and review prominent UQ paradigms, from Bayesian approximations and ensemble methods to evidential learning and conformal prediction. In our case study, a temporal Graph Attention Network (GAT) trained on dynamic functional connectivity (dFC) matrices from the SUDMEX CONN dataset achieves 80.0% diagnostic accuracy (F1 = 0.794) for Cocaine Use Disorder. However, a post-hoc uncertainty audit via Monte Carlo dropout reveals severe overconfidence (ECE = 0.127), with misclassified subjects assigned prediction confidences up to 95%. This empirical divergence between discrimination and calibration underscores the confidence paradox in deep connectomics. Our findings establish that rigorous UQ, calibration, and selective prediction mechanisms are indispensable for deploying trustworthy graph-based biomarkers in clinical neuroscience.

Updated: 2026-10-06 13:43:42

标题: 不确定性量化对于可靠的连接组学图学习至关重要:叙述性评论和案例研究

摘要: 尽管图神经网络(GNNs)在基于连接组的诊断分类中表现出了相当大的潜力,确定性模型不可避免地会抑制管道引起的噪音和模糊性,从而产生过于自信的预测。虽然不确定性量化(UQ)在体素级分割中被广泛采用,但在连接组图学习中的作用仍然未得到充分解决。本文提供了一篇专门针对连接组图学习的UQ框架的全面叙述性评论,同时展示了未校准预测的危险性的实证案例研究。我们详细描述了神经影像管道中的偶然和认知不确定性的来源,并回顾了主要的UQ范式,从贝叶斯近似和集成方法到证据学习和符合预测。在我们的案例研究中,一个基于SUDMEX CONN数据集中的动态功能连接(dFC)矩阵训练的时间图注意力网络(GAT)实现了80.0%的诊断准确性(F1 = 0.794)用于可卡因使用障碍。然而,通过蒙特卡罗辍学进行的事后不确定性审计揭示了严重的过度自信(ECE = 0.127),误分类的受试者分配的预测置信度高达95%。这种歧视和校准之间的实证差异强调了深度连接组中的信心悖论。我们的研究结果表明,严格的UQ、校准和选择性预测机制对于在临床神经科学中部署可信赖的基于图的生物标志物是不可或缺的。

更新时间: 2026-10-06 13:43:42

领域: cs.LG

下载: http://arxiv.org/abs/2610.08353v1

Guided Action Flow: Value-Guided Sampling for Frozen Vision-Language-Action Policies

Reinforcement learning can improve vision-language-action (VLA) policies beyond supervised fine-tuning, although this typically involves further updates to the policy parameters. For flow-matching policies, iterative action generation provides an additional opportunity to incorporate task information during inference. We introduce Guided Action Flow (GAF), which learns a compact, observation-conditioned action-value critic from robot task rollouts and applies its action gradient to steer reverse-time flow sampling. The supervised-fine-tuned VLA remains frozen throughout critic learning and deployment. Physical-robot experiments show an increase in aggregate success from 60.0% to 82.5% across six nominal manipulation tasks. Under six altered-lighting and object-distractor conditions evaluated on three of these tasks, aggregate success improves from 34.2% to 49.2%. Ablations and rollout analyses support the importance of the learned guidance direction and the critic's visual and proprioceptive inputs. With approximately 2.735M trainable critic parameters alongside a 0.45B-parameter VLA, GAF enables task outcomes to inform action generation through a compact inference-time guidance module.

Updated: 2026-10-06 13:43:17

标题: 导引动作流程:冻结视觉-语言-动作政策的价值导引采样

摘要: 强化学习可以在监督微调之外改进视觉-语言-行动(VLA)策略,尽管通常需要对策略参数进行进一步更新。对于匹配流的策略,迭代动作生成在推理过程中提供了一个额外的机会来整合任务信息。我们引入了引导动作流(GAF),它从机器人任务展开中学习了一种紧凑的,观察条件化的动作值评论家,并将其动作梯度应用于引导逆向时间流采样。在评论家学习和部署过程中,经过监督微调的VLA保持冻结状态。物理机器人实验显示,在六个标准操纵任务中,聚合成功率从60.0%提高到82.5%。在评估了三项任务中的六种改变光照和物体干扰条件下,聚合成功率从34.2%提高到49.2%。消融实验和任务展开分析支持了学习到的引导方向以及评论家的视觉和本体感输入的重要性。在大约2.735M可训练的评论家参数以及0.45B参数的VLA的基础上,GAF通过一个紧凑的推理时间引导模块使任务结果能够影响动作生成。

更新时间: 2026-10-06 13:43:17

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2607.02092v4

How Much Planning Is Enough? Reducing Search and Computation in World-Model Planning

Visual world models enable goal-directed control through decision-time action search, but their deployment efficiency is often limited by conservatively large planning budgets. We show that competitive task performance can be achieved without agreement with the Full-budget action, that sufficient budgets vary across model--task pairs, and that iterative planners repeatedly encode solve-invariant context. To address these inefficiencies, we propose {SufficientPlan}, a simple deployment framework that requires no modification to pretrained world models or planner updates. Its {Paired Sequential Budget Certification (PSBC)} component uses paired closed-loop evidence to search for and certify a reduced model--task-specific budget within a predefined Full-performance tolerance. Its {Static-Context Reuse (SCR)} component caches observation and goal representations across search iterations while preserving candidate-dependent planning and selected actions. Experiments across multiple world-model backbones and visual-control tasks show that SufficientPlan substantially reduces search budgets and planning latency while maintaining competitive control performance.

Updated: 2026-10-06 13:40:44

标题: 需要多少规划?减少世界模型规划中的搜索和计算

摘要: 视觉世界模型通过决策时间的动作搜索实现目标导向控制,但它们的部署效率通常受限于过于保守的大规划预算。我们展示了,即使与全预算动作不一致,也可以实现竞争性任务表现,足够的预算在模型-任务对之间变化,并且迭代规划器重复编码解不变的上下文。为了解决这些低效问题,我们提出了{SufficientPlan},这是一个简单的部署框架,无需修改预训练的世界模型或规划器更新。其{Paired Sequential Budget Certification (PSBC)} 组件使用成对的闭环证据搜索和验证在预定义的全面表现容忍度内的减少模型-任务特定预算。其 {Static-Context Reuse (SCR)} 组件在搜索迭代中缓存观察和目标表示,同时保留候选依赖的规划和选择动作。实验跨多个世界模型骨干和视觉控制任务表明,SufficientPlan大大降低了搜索预算和规划延迟,同时保持了竞争性控制性能。

更新时间: 2026-10-06 13:40:44

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2610.08350v1

High-Dimensional Statistical Inference for Sparse Support Vector Machines

Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the $L_1$-penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.

Updated: 2026-10-06 13:40:12

标题: 高维度统计推断用于稀疏支持向量机

摘要: 使用一个复制对称的高维特征描述,我们为稀疏支持向量机开发了一个推断框架,当样本大小和特征数量成比例增长时。主要挑战是非平滑的铰链损失,这阻碍了对平滑分类损失开发的去偏置论证的直接应用。我们通过将$L_1$惩罚支持向量机(SVM)表示为一个线性规划,并通过其对偶变量识别铰链损失次梯度来克服这一困难。这产生了一个在比例渐近制度下的计算可访问的去偏置估计量,其坐标在渐近情况下呈高斯分布。所得的分布特征提供了对个别特征的置信区间和假设检验,并实现了控制虚发现率的变量选择。广泛的模拟研究了在一系列协方差结构下的校准、功率和变量选择性能,包括强相关设计。对高维乳腺癌基因表达数据的分析说明了提出的推断如何区分通过原始稀疏SVM选择的变量与统计显著的特征。

更新时间: 2026-10-06 13:40:12

领域: stat.ML,cs.LG,stat.ME

下载: http://arxiv.org/abs/2610.08345v1

DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.

Updated: 2026-10-06 13:38:27

标题: DIPrune: 面向任务的双重重要性的令牌修剪,用于高效的多模态语言模型

摘要: 最近针对多模态大型语言模型(MLLMs)的无需训练修剪方法通过利用视觉冗余或文本-视觉注意力有效地减少计算开销。然而,它们经常因为任务不可知的设计或不可靠的注意力估计而遭受语义退化。根据我们的实证分析,我们发现这个问题的根源在于浅层中显著的标记通过数值惯性持续抑制新出现的语义标记,导致关键信号被过早丢弃,从而影响深层推理。为了解决上述问题,我们从任务导向的角度出发,首先将无需训练的修剪重新制定为最小化最终任务损失中的失真,并推导出一个可行的、基于标记的上限作为替代目标。具体来说,这种制定本质上揭示了一个先前被忽视的跨层项,该项考虑了跨层梯度。因此,在实现方面,我们提出了DIPrune,一个基于排名的框架,采用双重重要性评分机制来共同优化层内静态特征显著性和层间动态语义演变。对LLaVA和Qwen-VL的大量实验证明,DIPrune始终实现了最先进的结果。

更新时间: 2026-10-06 13:38:27

领域: cs.CV,cs.LG

下载: http://arxiv.org/abs/2610.08341v1

Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV). We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571. The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025. The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.

Updated: 2026-10-06 13:38:26

标题: 存储不是策略:基于状态条件的支持控制用于LLM去学习

摘要: 许多局部化大型语言模型(LLM)遗忘方法从局部化信号中选择一个小的参数子集,并在优化过程中保持固定。然而,与目标最相关的参数未必是最好的更新对象,候选干预可能会随着优化的进行而改变价值。在一项受控实验中,存储局部化分数达到了接收器工作特性曲线下的面积(AUROC)为0.981,然而存储身份仅在36个目标中的17个上与更好的干预一致,而低秩适应(LoRA)在36个目标中有35次获胜。我们引入了干预分数,通过预测实际遗忘更新的影响效果对可编辑组进行排序,同时考虑到附带损害,然后使用它形成静态干预值基线(Static-IV)。然后我们引入了选择性动态干预重新排序(DIR-R),只有当校准探测器证明了比较的必要性时才重新访问该子集。在Natural-TOFU数据集上,我们的方法在20次方法和目标之间的比较中有19次正面描述性边际,尽管有几次接近零。在LACUNA局部化精度基准测试中,我们的平均终端效用在所有六个负偏好优化(NPO)和SimNPO比较中都更高:NPO边际范围从+0.431到+0.848,SimNPO边际范围从+0.503到+0.571。梯度差异(GradDiff)目标揭示了实质性的领域依赖性。相对于Static-IV,主要的四个领域GradDiff评估有六次胜利,六次平局,没有失败,平均和中位数配对增益分别为+0.165和+0.0025。证据支持将局部化、初始干预选择和检查点依赖的支持修订分开。

更新时间: 2026-10-06 13:38:26

领域: cs.LG,cs.CL

下载: http://arxiv.org/abs/2609.37858v2

PAC-CF: Calibrating Irreversible Frontier Pruning in LLM-Guided Search

LLM-guided search explores multiple candidate trajectories, but at substantial test-time cost. Pruning low-scoring frontier candidates can control this cost, yet it also turns potentially biased evaluator scores into irreversible decisions: systematic ranking errors can persist under repeated scoring and remove useful branches. We propose Probably Approximately Correct Conformal Filtering (PAC-CF). Its fixed-frontier analysis formulates elimination as an $(\varepsilon,δ)$-PAC problem under bounded evaluator bias; its operational rule separately calibrates a score-gap threshold on held-out tasks by running the original controller without PAC-CF and using post-search verifier labels to measure the deficit of solution-preserving candidates relative to the frontier leader. Conditional on exchangeable native-controller tasks with nonempty protected exposure, conformal calibration gives finite-sample coverage for retaining at least one verifier-defined valid continuation at every protected frontier on the native trajectory. At deployment, PAC-CF removes only candidates whose gap from the highest frontier score exceeds the frozen threshold. We evaluate PAC-CF across three domains, five controllers, and four request budgets from B100 to B500. In the cross-domain/controller macro averages, the point estimates for all three workload measures are lower at every budget; the paired-bootstrap 95\% confidence interval for utility excludes zero at B100 and B200. For pruning-aware ToolTree, the full-test-set cross-domain utility difference is $+4.38$ points at each tested budget; on the natural-termination sensitivity cohort, physical requests decrease by $18.94$--$18.95\%$ and end-to-end token usage by $23.57$--$23.76\%$.

Updated: 2026-10-06 13:36:41

标题: PAC-CF: 在LLM引导搜索中校准不可逆边界修剪

摘要: LLM引导搜索探索多个候选路径,但在测试时间成本很高。修剪得分低的前沿候选者可以控制这种成本,但也会将潜在的有偏评估者评分转化为不可逆转的决定:系统排序错误可能会在重复评分下持续存在,并移除有用的分支。我们提出了可能大致正确的符合过滤(PAC-CF)。其固定前沿分析形式将消除问题公式化为一个$(\varepsilon,δ)$-PAC问题,受限于评估者偏见;其操作规则通过在保留任务上运行原始控制器而不使用PAC-CF,并使用后搜索验证器标签来衡量相对于前沿领导者的解决方案保留候选者的不足,分别校准得分差阈值。在具有非空受保护曝光的可交换本地控制器任务条件下,符合校准为在本地轨迹上的每个受保护前沿至少保留一个由验证器定义的有效延续提供了有限样本覆盖。在部署时,PAC-CF仅移除与最高前沿得分的差距超过冻结阈值的候选者。我们在三个领域、五个控制器和从B100到B500的四个请求预算中评估了PAC-CF。在跨领域/控制器宏平均值中,所有三个工作负载度量的点估计在每个预算下都较低;在B100和B200处,实用性的成对bootstrap 95%置信区间排除了零。对于具有修剪意识的ToolTree,在每个测试预算上,全测试集跨领域实用性差异为+4.38个点;在自然终止敏感性队列中,物理请求减少了18.94%至18.95%,端到端令牌使用减少了23.57%至23.76%。

更新时间: 2026-10-06 13:36:41

领域: cs.LG,cs.AI,stat.ML

下载: http://arxiv.org/abs/2604.14345v6

Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution under Analysis Budgets

Sandbox execution and memory forensics are among the most constrained resources in malware triage. Static analysis can scale to millions of files, whereas dynamic and memory analysis require minutes of analyst controlled infrastructure for each sample. Despite this difference, multimodal ransomware detectors often apply every modality to every sample, causing analysis cost and time to verdict to increase linearly with sample volume even when static evidence is already sufficient for a decision. We present a cost aware Hierarchical Multi-Agent System that formulates evidence acquisition as a budgeted sequential decision problem. Specialist agents generate schema validated risk signals for each modality, domain controllers aggregate these signals, and a Meta-Orchestrator begins with static evidence and escalates to dynamic and memory evidence only when confidence is insufficient or agents within a controller disagree. An optional, bounded, locally hosted large language model reviewer can adjust a verdict by at most one tier but cannot replace the deterministic pipeline. Each decision is recorded with a complete provenance trace. In multiple runs over 12439 samples from 16 ransomware families and benign samples, the deterministic HMAS achieves F1 0.93 and macro F1 0.97, resolving 57.95% of cases using static evidence alone, 35.83% after adding dynamic evidence, and only 6.21% through the full pipeline. The average internal analysis cost is 6.65 units, compared with 12 for exhaustive analysis, representing a 44.6% reduction. Standalone leave-one-family-out testing further shows that accuracy on families held out during tuning falls to 0.26 to 0.64 outside the Benign and high support classes. We report these results alongside a cost sensitivity analysis, a partial leave one component out ablation, and a full scale comparison with learned early and late fusion and cascade baselines.

Updated: 2026-10-06 13:30:24

标题: 成本感知的分层多代理勒索软件检测和家族归因在分析预算下

摘要: 沙盒执行和内存取证是恶意软件审查中最受限制的资源之一。静态分析可以扩展到数百万个文件,而动态和内存分析则需要每个样本由分析师控制的基础设施几分钟的时间。尽管存在这种差异,多模式勒索软件检测器通常会对每个样本应用每种模态,导致分析成本和判定时间随着样本量的增加呈线性增长,即使静态证据已经足够做出决定。我们提出了一种成本感知的分层多代理系统,将证据获取形式化为一个预算顺序决策问题。专业代理为每种模态生成经模式验证的风险信号,域控制器聚合这些信号,元编排器从静态证据开始,并在置信度不足或控制器内的代理不一致时才升级到动态和内存证据。一个可选的、有界的、本地托管的大型语言模型评审员可以最多调整一个层次的判定,但不能替代确定性管道。每个决定都记录了完整的溯源跟踪。在对来自16个勒索软件家族和良性样本的12439个样本进行多次运行后,确定性HMAS实现了F1 0.93和宏F1 0.97,仅使用静态证据解决了57.95%的案例,添加动态证据后为35.83%,而仅通过完整的管道为6.21%。平均内部分析成本为6.65个单位,相比于详尽分析的12个单位,减少了44.6%。单独进行留一家族测试进一步显示,在调整期间保留的家族上的准确性下降到0.26到0.64之间,超出了良性和高支持类别。我们报告了这些结果,同时进行了成本敏感性分析、部分留一组件分析和与学习的早期和后期融合以及级联基线的全面比较。

更新时间: 2026-10-06 13:30:24

领域: cs.CR,cs.AI

下载: http://arxiv.org/abs/2609.04820v2

MARCO: The Radioactive Watermark for Protein Generative Models

Protein Generative Models (PGMs) have revolutionized structural biology by enabling the design of complex 3D protein structures from sequence data. However, this breakthrough introduces a dual-use challenge, exposing high-value PGMs to economic risks like unauthorized model extraction and biosecurity threats such as biohazard synthesis. To mitigate these threats, we propose \textbf{MARCO} (\textsc{COnformation waterMARk}), the first radioactive watermarking framework specifically tailored for PGMs. MARCO establishes a Dual-Layer defense that simultaneously protects intellectual property and ensures the forensic traceability of potential biosecurity misuses. (i) To preserve efficiency, MARCO iteratively embeds watermarks during diffusion reverse denoising via an auxiliary encoder-decoder, allowing the original PGM parameters to remain frozen for broad compatibility. (ii) To preserve biophysical fidelity and maximize robustness, we employ specialized loss functions targeting $C_α$-atom pairwise distances and torsion angles ($ψ, φ$) within an adversarial training framework integrated with stochastic attack simulations. (iii) Crucially, MARCO exhibits ``radioactivity'' where the watermark automatically transfers to the outputs of any pirate models trained on the watermarked data, effectively countering model extraction attacks. Comprehensive experiments demonstrate that MARCO achieves superior fidelity and robustness while successfully validating watermark transferability.

Updated: 2026-10-06 13:18:22

标题: 马尔科:蛋白质生成模型的放射性水印

摘要: 蛋白质生成模型(PGMs)已经通过使复杂的三维蛋白质结构的设计变得可能,从序列数据中引领了结构生物学的革命。然而,这一突破性进展带来了一个双重挑战,使高价值的PGMs面临未经授权的模型提取和生物安全威胁(如生物危害合成)等经济风险。为了减轻这些威胁,我们提出了\textbf{MARCO}(\textsc{COnformation waterMARk}),这是专门为PGMs量身定制的第一个放射性水印框架。MARCO建立了一个双层防御体系,同时保护知识产权并确保对潜在的生物安全滥用的法庭追踪。 (i)为了保持效率,MARCO通过辅助编码器-解码器在扩散反向去噪过程中迭代地嵌入水印,使原始PGM参数保持冻结以实现广泛的兼容性。 (ii)为了保持生物物理的忠实性和最大化的鲁棒性,我们采用了专门针对$C_α$-原子间距离和扭转角($ψ, φ$)的损失函数,并将其整合到与随机攻击模拟相结合的对抗训练框架中。 (iii)关键是,MARCO表现出“放射性”,其中水印会自动转移到任何在带水印数据上训练的盗版模型的输出中,从而有效地抵制模型提取攻击。全面的实验证明,MARCO在成功验证水印可转移性的同时实现了优越的忠实度和鲁棒性。

更新时间: 2026-10-06 13:18:22

领域: cs.CR,cs.AI

下载: http://arxiv.org/abs/2610.08316v1

Zeppelin: Client-Side BFV Encryption and Decryption for Helium-Powered Microcontrollers

The growth of the Internet of Things (IoT) has raised concerns over the privacy of data collected by resource-constrained sensing devices. Homomorphic encryption (HE) addresses this by letting a device encrypt its data once and an untrusted cloud server compute on the ciphertext without seeing the values. In practice, HE's memory and computational cost have kept it out of reach of microcontroller-class devices. Prior work, SEAL-Embedded, made this feasible using CKKS, but left three gaps: its arithmetic is entirely scalar, even on hardware with a vector instruction set; it never decrypts on the device, so the client cannot consume a result; and it does not explore BFV, whose encoding uses only integer arithmetic and decrypts exactly. We present Zeppelin, the first HE library to use an embedded vector instruction set, the first to support both encryption and decryption on an embedded device, and the first BFV implementation on MCU-class hardware. Zeppelin vectorizes the number-theoretic transform, HE's main bottleneck, for ARM's Helium extension, and includes a decryption procedure that avoids large-integer arithmetic and timing leakage of the secret key. A server-side adapter converts Zeppelin's ciphertexts into a format compatible with Microsoft SEAL, so the device handles encryption and decryption while the server performs all homomorphic computation. On an STM32N6 MCU with an ARM Cortex-M55, Zeppelin encodes and encrypts 4096 packed values in 4.76 ms (seeded symmetric, 128-bit security per the HE standard) and decrypts and decodes the result in 5.38 ms, using under 500 KB of RAM, with the vectorized NTT engine ${\sim}2.7\times$ faster than an equivalent scalar implementation on the same core.

Updated: 2026-10-06 13:07:53

标题: 齐柏林:氦气动力微控制器的客户端BFV加密和解密

摘要: 物联网(IoT)的发展引起了对资源受限感知设备收集的数据隐私的担忧。同态加密(HE)通过让设备加密其数据一次,然后让不受信任的云服务器在不查看数值的情况下对密文进行计算来解决这个问题。在实践中,HE的内存和计算成本使得微控制器级设备无法使用。先前的工作SEAL-Embedded使用了CKKS使这一点成为可能,但存在三个缺陷:其算术完全是标量的,即使在具有向量指令集的硬件上也是如此;它从不在设备上解密,因此客户端无法获取结果;它也没有探索BFV,其编码只使用整数算术并且可以精确解密。我们提出了Zeppelin,这是第一个使用嵌入式向量指令集的HE库,也是第一个支持在嵌入式设备上进行加密和解密的库,同时也是MCU级硬件上的第一个BFV实现。Zeppelin对ARM的Helium扩展中的数论变换进行了向量化处理,这是HE的主要瓶颈,并包含了一个解密过程,避免了大整数算术和秘钥泄漏的时间问题。服务器端适配器将Zeppelin的密文转换为与Microsoft SEAL兼容的格式,因此设备处理加密和解密,而服务器执行所有同态计算。在具有ARM Cortex-M55的STM32N6 MCU上,Zeppelin在4.76毫秒内对4096个打包数值进行编码和加密(基于HE标准,128位安全性),并在5.38毫秒内解密和解码结果,使用少于500 KB的RAM,其向量化NTT引擎比同一核心上的等效标量实现快约2.7倍。

更新时间: 2026-10-06 13:07:53

领域: cs.CR,cs.AR

下载: http://arxiv.org/abs/2610.08301v1

Practical Feasibility of Gradient Inversion Attacks in Federated Learning

Gradient inversion attacks are often presented as a serious privacy threat in federated learning, with recent work reporting increasingly strong reconstructions under favorable experimental settings. However, it remains unclear whether such attacks are feasible in modern, performance-optimized systems deployed in practice. In this work, we evaluate the practical feasibility of gradient inversion for image-based federated learning. We conduct a systematic study across multiple datasets and tasks, including image classification and object detection, using canonical vision architectures at contemporary resolutions. Our results show that while gradient inversion remains possible for certain legacy or transitional designs under highly restrictive assumptions, modern, performance-optimized models consistently resist meaningful reconstruction visually. We further demonstrate that many reported successes rely on upper-bound settings, such as inference mode operation or architectural simplifications which do not reflect realistic training pipelines. Taken together, our findings indicate that, under an honest-but-curious server assumption, high-fidelity image reconstruction via gradient inversion does not constitute a critical privacy risk in production-optimized federated learning systems, and that practical risk assessments must carefully distinguish diagnostic attack settings from real-world deployments.

Updated: 2026-10-06 13:02:26

标题: 《联邦学习中梯度反演攻击的实际可行性》

摘要: 梯度反演攻击经常被认为是联邦学习中的严重隐私威胁,在有利的实验设置下,最近的研究报告显示出越来越强的重建效果。然而,目前尚不清楚这种攻击在实践中部署的现代、性能优化系统中是否可行。在这项工作中,我们评估了基于图像的联邦学习中梯度反演的实际可行性。我们在多个数据集和任务中进行了系统研究,包括图像分类和目标检测,使用当代分辨率下的经典视觉架构。我们的研究结果表明,虽然在高度限制性假设下,梯度反演对于某些传统或过渡设计仍然可能,但现代性能优化模型在视觉上一致抵抗有意义的重建。我们进一步证明,许多报道的成功案例依赖于上限设置,如推断模式操作或不反映现实训练流程的架构简化。总的来说,我们的研究结果表明,在一个诚实但好奇的服务器假设下,通过梯度反演实现高保真度图像重建并不构成生产优化的联邦学习系统中的关键隐私风险,而实际风险评估必须仔细区分诊断攻击设置和实际部署。

更新时间: 2026-10-06 13:02:26

领域: cs.CR,cs.AI,cs.LG

下载: http://arxiv.org/abs/2508.19819v3

Contextual Chain: Lightweight Continuity Authentication for Intermittently Connected Devices

Can authentication make memory, rather than computational hardness, the attacker's bottleneck? Contextual Chain is a lightweight continuity protocol for intermittently connected devices that share evolving physical or operational context. An honest device follows one realized history, updating a compact accumulator and fixed hash-based readiness lanes; outages cause pause or bounded rollback, not branch search. After the epoch is frozen, a fresh challenge selects one lane under a short deadline. An outsider that missed context may therefore need to prepare for many mature histories before learning which one will be tested. In the standard random-oracle model, a causal counting theorem lower-bounds the deadline-accessible retained state required for a target success probability against arbitrary nonlinear preselection encoding and adaptive post-selection queries, accounting for sequential depth, candidate queries, and cross-target protected information obtained online. Honest readiness memory remains fixed and independent of the number of plausible histories. Contextual Chain thus converts shared-experience uncertainty into a tunable preparation requirement without transferring combinatorial complexity to lightweight devices.

Updated: 2026-10-06 12:37:50

标题: 上下文链:面向间歇连接设备的轻量级连续性认证

摘要: 身份验证是否可以使记忆成为攻击者的瓶颈,而不是计算难度?Contextual Chain是一种轻量级的连续性协议,适用于共享不断演化的物理或运行环境的间歇连接设备。一个诚实的设备遵循一个实现的历史,更新一个紧凑的累加器和固定的基于哈希的就绪通道;中断会导致暂停或有界回滚,而不是分支搜索。在时代被冻结后,一个新的挑战会在短时间内选择一个通道。因此,一个错过上下文的外部人可能需要为许多成熟的历史做准备,然后才能了解哪一个将被测试。在标准随机神谕模型中,一个因果计数定理对针对任意非线性预选编码和自适应后选择查询计算出的目标成功概率下界,考虑了顺序深度、候选查询和在线获取的跨目标受保护信息所需的保留状态。诚实的就绪记忆保持不变,并且与可信历史的数量无关。因此,Contextual Chain将共享经验的不确定性转化为可调整的准备要求,而不会将组合复杂性转移给轻量级设备。

更新时间: 2026-10-06 12:37:50

领域: cs.CR,cs.DC

下载: http://arxiv.org/abs/2610.08262v1

HE-OFT: Privacy-Preserving One-Shot Federated Fine-Tuning under Homomorphic Encryption

Many organizations adapt large pretrained models to their own tasks by fine-tuning on private data. Several of these parties often hold data for the same task and wish to fine-tune a model together without pooling that data. Federated learning (FL) enables joint fine-tuning, but reconstruction attacks on shared intermediate values (the model or its gradients) remain a privacy risk. A one-shot protocol that exchanges one encrypted contribution exposes no intermediate value. Such a protocol still gives the trained model to every participant, which is not permitted where the model is a regulated or proprietary asset. We present HE-OFT, the first cryptographically secure one-shot federated fine-tuning protocol in which no party receives the trained model. Each client fine-tunes a low-rank adapter and a classifier head on a frozen public backbone and keeps the adapter. The client uploads one encrypted head displacement, which the server combines under multiparty CKKS and never decrypts. A quorum of clients returns only the predicted label to the querier. On four text classification tasks and one vision task, HE-OFT reaches 61 to 79 per cent accuracy, against 20 to 48 per cent for a client training alone. HE-OFT keeps 85 to 96 per cent of the accuracy of a disclosed model. A test-time query takes 443.1 to 1713.1 s on one core, or 56.1 to 255.1 s with level restoration on a GPU. Restoring levels at the server cuts the traffic per query from up to 1.6 GiB to 13.5 MiB.

Updated: 2026-10-06 12:35:08

标题: HE-OFT:在同态加密下隐私保护的一次性联邦微调

摘要: 许多组织通过在私人数据上进行微调,将大型预训练模型调整到自己的任务中。几个这样的组织通常持有相同任务的数据,并希望共同对模型进行微调,而不汇集这些数据。联邦学习(FL)实现了联合微调,但对共享的中间值(模型或其梯度)的重建攻击仍然是一个隐私风险。一种一次性协议,交换一个加密的贡献,不暴露任何中间值。这样的协议仍将训练好的模型提供给每个参与者,这在模型是受监管或专有资产的情况下是不允许的。我们提出了HE-OFT,这是第一个在其中没有任何一方接收训练好的模型的加密安全的一次性联邦微调协议。每个客户端在冻结的公共骨干上微调低秩适配器和分类器头,并保留适配器。客户端上传一个加密的头位移,服务器在多方CKKS下组合这些位移并永远不解密。一组客户端仅向查询者返回预测的标签。在四个文本分类任务和一个视觉任务上,HE-OFT的准确率达到61%至79%,而单独训练的客户端的准确率为20%至48%。HE-OFT保持了一个披露模型的85%至96%的准确率。一个测试时查询在一个核心上需要443.1秒至1713.1秒,或在GPU上进行水平恢复时需要56.1秒至255.1秒。在服务器上恢复级别将每个查询的流量从高达1.6 GiB减少到13.5 MiB。

更新时间: 2026-10-06 12:35:08

领域: cs.CR

下载: http://arxiv.org/abs/2610.08255v1

Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code

Large Language Models (LLMs) are widely used to generate code. Although their functional plausibility keeps improving, the generated code often contains security vulnerabilities. The functionality-security gap captures code that passes functional tests but fails security tests. A recent longitudinal study of three model families concluded that LLMs become smarter but not safer, with the only considered open-weight family stagnating. Whether this holds for other (open-weight) families and particularly for compact models remains open. We present a longitudinal study of the gap across 32 LLMs from seven model families (five open-weight), covering three successive releases per family in flagship and compact variants. Using CWEval with 119 tasks in five programming languages and 31 CWEs, we compare trajectories across families, model sizes, and languages. Newer models do become safer in absolute terms, although no family closes the gap. Unlike prior work, we find that openness does not separate the families: every considered open-weight family narrows the gap significantly, while Gemini 3.1 Pro keeps a gap as wide as the one reported for Llama. Compact models usually produce less secure code than their flagship counterparts, with notable exceptions (e.g., Gemini 3.7 Flash). At the CWE level, we confirm persistent weaknesses such as log injection (CWE-117) and HTTP response splitting (CWE-113) and regressions in the newest proprietary models on memory and integer weaknesses, and show that the same CWE carries very different risk across languages. From these results, we derive implications for LLM vendors, researchers, and developers. In particular, developers should assume neither that upgrades improve security nor that proprietary models are more secure; they should rerun security checks after each model change and provide secure APIs in the model's context.

Updated: 2026-10-06 12:25:30

标题: 更新和更大,但更安全吗?LLM生成代码中功能性与安全性差距的纵向研究

摘要: 大型语言模型(LLMs)被广泛用于生成代码。尽管它们的功能合理性不断提高,但生成的代码通常包含安全漏洞。功能安全差距捕捉到通过功能测试但未通过安全测试的代码。最近一项关于三个模型家族的纵向研究得出结论,LLMs变得更聪明但并不更安全,唯一考虑的开放权重家族停滞不前。是否对其他(开放权重)家族,特别是紧凑模型也适用,尚未确定。我们展示了对来自七个模型家族的32个LLMs进行的差距纵向研究(其中五个是开放权重),覆盖了每个家族的旗舰版和紧凑版的三个连续发布。使用包括119个任务、五种编程语言和31个CWE的CWEval,我们比较了不同家族、模型大小和语言之间的轨迹。新模型在绝对意义上确实变得更安全,尽管没有任何一个家族能够弥合这一差距。与以往研究不同,我们发现开放性并不能将这些家族分开:每个考虑的开放权重家族都大幅缩小了差距,而Gemini 3.1 Pro保持了与Llama报告的差距一样宽。紧凑模型通常生成的代码比旗舰版本更不安全,但也有一些显著的例外(例如,Gemini 3.7 Flash)。在CWE级别上,我们确认了持久的弱点,如日志注入(CWE-117)和HTTP响应分割(CWE-113),以及最新专有模型在内存和整数弱点方面的退化,并展示了同一CWE在不同语言中的风险差异。根据这些结果,我们为LLM供应商、研究人员和开发人员提出了一些启示。特别是,开发人员不应假设升级会改善安全性,也不应认为专有模型更安全;他们应在每次模型更改后重新运行安全检查,并在模型的上下文中提供安全的API。

更新时间: 2026-10-06 12:25:30

领域: cs.SE,cs.CR

下载: http://arxiv.org/abs/2610.08240v1

Large Language Models for Cryptocurrency Transaction Analysis: A Bitcoin Case Study

Cryptocurrencies are widely used, yet current methods for analyzing transactions often rely on opaque, black-box models. While these models may achieve high performance, their outputs are usually difficult to interpret and adapt, making it challenging to capture nuanced behavioral patterns. Large language models (LLMs) have the potential to address these gaps, but their capabilities in this area remain largely unexplored, particularly in cybercrime detection. In this paper, we test this hypothesis by applying LLMs to real-world cryptocurrency transaction graphs, with a focus on Bitcoin, one of the most studied and widely adopted blockchain networks. We introduce a three-tiered framework to assess LLM capabilities: foundational metrics, characteristic overview, and contextual interpretation. This includes a new, human-readable graph representation format, LLM4TG, and a connectivity-enhanced transaction graph sampling algorithm, CETraS. Together, they significantly reduce token requirements, transforming the analysis of multiple moderately large-scale transaction graphs with LLMs from nearly impossible to feasible under strict token limits. Experimental results demonstrate that LLMs have outstanding performance on foundational metrics and characteristic overview, where the accuracy of recognizing most basic information at the node level exceeds 98.50% and the proportion of obtaining meaningful characteristics reaches 95.00%. Regarding contextual interpretation, LLMs also demonstrate strong performance in classification tasks, even with very limited labeled data, where top-3 accuracy reaches 72.43% with explanations. While the explanations are not always fully accurate, they highlight the strong potential of LLMs in this domain. At the same time, several limitations persist, which we discuss along with directions for future research.

Updated: 2026-10-06 11:58:33

标题: 大型语言模型用于加密货币交易分析:比特币案例研究

摘要: 加密货币被广泛使用,然而目前用于分析交易的方法往往依赖于不透明的黑匣子模型。虽然这些模型可能达到高性能,但它们的输出通常难以解释和调整,这使得捕捉微妙的行为模式具有挑战性。大型语言模型(LLMs)有潜力填补这些空白,但它们在这一领域的能力仍然很少被探索,特别是在网络犯罪检测方面。在本文中,我们通过将LLMs应用于现实世界的加密货币交易图,重点关注比特币,这是最受研究和广泛采用的区块链网络之一,来测试这一假设。我们引入了一个三层框架来评估LLM的能力:基础指标、特征概览和语境解释。这包括一种新的、人类可读的图表示格式LLM4TG,以及一个增强连接的交易图采样算法CETraS。它们共同显著降低了令牌需求,将在严格的令牌限制下对多个中等规模的交易图进行LLM分析从几乎不可能变得可行。实验结果表明,LLMs在基础指标和特征概览方面表现出色,节点级别识别大多数基本信息的准确率超过98.50%,获得有意义特征的比例达到95.00%。关于语境解释,LLMs在分类任务中也表现出色,即使有非常有限的标记数据,前三的准确率也能达到72.43%,带有解释。虽然解释并不总是完全准确,但它们突显了LLMs在这一领域的巨大潜力。与此同时,一些限制仍然存在,我们将讨论这些限制以及未来研究的方向。

更新时间: 2026-10-06 11:58:33

领域: cs.CR,cs.LG

下载: http://arxiv.org/abs/2501.18158v4

Subjects, Not Authors: The Authorship Hazard in Agentic Dataspaces

Dataspace connectors decide whether a transfer may occur, not what the transferred value contains, tolerable for contracted applications, not for LLM agents that compose tool calls. Work on agents that generate governance artifacts evaluates output quality, not who may authorize an artifact for use. A published policy is what the decision point enforces, so publication is a governance event, and agents that are both policy subjects and policy authors write the norms that bind them. We name this the authorship hazard and state one principle: an agent is a subject of the governance plane, never an author of it. Its authorization channel to publication is closed by construction; its influence channel, drafting what humans approve, becomes an enforcement problem. Across 90 preregistered edits to the paper's running agreement, each evaluated on 344,512 requests, the six that only reclassify a field all change authorization and narrow a duty without touching policy text, and a policy-diff classifier passes all six. Read as worded, the privilege-delta conditions also pass 33 of 69 effective policy-text edits; read as covering any relaxation, none. Treating classification as authorship routes all six to review; the registry this requires is not yet built. At the execution boundary, protected fields reach the model in 105 of 105 cases under prompt-stated duties and in 0 of 105 under a compiled tool-call constraint, but values outside named fields are exposed in 7 of 7. At the review share measured, a central approval pool needs one approver per 20 to 138 participants.

Updated: 2026-10-06 11:43:54

标题: 主题,而不是作者:在有权数据空间中的作者风险

摘要: 数据空间连接器决定是否可以发生传输,而不是传输的值包含什么,这对于合同应用程序是可以容忍的,但对于组成工具调用的LLM代理不可以。生成治理工件的代理工作评估输出质量,而不是谁可以授权工件供使用。已发布的政策是决策点执行的内容,因此发布是一种治理事件,同时既是政策主体又是政策作者的代理人编写规范约束它们。我们将这称为创作风险,并提出一个原则:代理人是治理平面的主体,永远不是其作者。其授权通道到发布是由结构关闭的;其影响通道,即起草人类批准的内容,会成为一个执行问题。对于文章的90个预注册编辑,每个编辑都在344,512个请求上进行评估,其中六个仅重新分类一个字段,都会改变授权并缩小责任,而没有修改政策文本,且一个政策差异分类器通过了所有六个。按照字面意思来看,特权增量条件也通过了69个有效政策文本编辑中的33个;按照任何放松的标准来看,没有一个通过。将分类视为创作将所有六个路由到审查;这需要的注册表尚未建立。在执行边界上,受保护字段在105个案例中有105个在快速指定职责下达模型,但在编译工具调用约束下为0个,在7个案例中,超出命名字段的值暴露出来。在测量的审查份额中,一个中央批准池需要每20至138名参与者一个批准者。

更新时间: 2026-10-06 11:43:54

领域: cs.CR,cs.AI,cs.DB,cs.MA

下载: http://arxiv.org/abs/2609.30614v2

On the Intrinsic Limited Robustness of Latent-Based Watermarking

Existing latent-based watermarking methods for diffusion models have overestimated their robustness to image distortions, including geometric transformations such as rotation, scaling, and translation (RST). Moreover, this paradigm of watermarking approaches may suffer from inherent limitations arising from the domain in which the watermark is embedded. In this paper, we provide the first theoretical analysis explaining why these methods lack invariance to perturbations. By relaxing the invariant relation, we derive a maximum perturbation bound that characterizes the relationship between pixel-space perturbations and their corresponding effects in latent space. In addition, we present the first analytical formulation that captures all components of practical detection mechanisms. Finally, we conduct experiments to validate the theoretical findings and the limitations of latent-based watermarking methods. Our theoretical and empirical results indicate that, under the current design paradigm, latent-based watermarking methods intrinsically exhibit limited robustness. We conclude by providing the analytical tool and design guidelines that future research could follow.

Updated: 2026-10-06 11:29:50

标题: 关于基于潜在特征的水印技术固有有限的鲁棒性

摘要: 现有的基于潜水印的扩散模型水印方法高估了它们对图像失真的鲁棒性,包括几何变换,如旋转、缩放和平移(RST)。此外,这种水印方法范式可能受到植入水印的域固有限制的影响。在本文中,我们提供了第一次理论分析,解释为什么这些方法缺乏对扰动的不变性。通过放宽不变关系,我们得出了一个最大扰动界限,描述了像素空间扰动与其在潜在空间中对应效果之间的关系。此外,我们提出了捕捉所有实际检测机制组件的第一个分析公式。最后,我们进行实验证实理论发现和基于潜水印的水印方法的局限性。我们的理论和实证结果表明,在当前设计范式下,基于潜水印的水印方法固有地表现出有限的鲁棒性。我们总结了提供未来研究可以遵循的分析工具和设计准则。

更新时间: 2026-10-06 11:29:50

领域: cs.LG,cs.CR

下载: http://arxiv.org/abs/2610.08178v1

FBAN: A Fully Homomorphic Encryption Compatible Bottleneck Attention Network for Privacy-Preserving Behavioral Authentication

Continuous authentication (CA) strengthens session security by repeatedly verifying the user during device interaction, yet it inherently relies on highly sensitive behavioral traces (e.g., fine-grained touch dynamics) that are often outsourced to cloud/edge services for scalable inference. This raises a fundamental privacy-in-use challenge: protecting behavioral features during computation, not only in transit or at rest. Fully homomorphic encryption (FHE) offers a principled solution, but deploying modern CA models under FHE remains difficult due to non-linearities and attention-style operations that incur high ciphertext cost. We propose FBAN, a TFHE-compatible Bottleneck Attention Network and an end-to-end encrypted CA framework. FBAN is designed for integer-only execution via a two-stage pipeline (floating-point pretraining followed by quantization-aware training) and is compiled into TFHE circuits for homomorphic inference. We further specify a client-server protocol with session-bound blinding and decrypt-and-return verification, enabling the server to authenticate users without observing raw behavioral features. We provide a cryptographic security analysis against an honest-but-curious server under TFHE IND-CPA security, and formalize resistance to replay and impersonation without the TFHE secret key. Experiments on two public touchscreen datasets demonstrate that FBAN achieves strong authentication utility under encrypted inference while maintaining a lightweight model footprint, with TFHE parameters instantiated at $\geq 128$-bit security.

Updated: 2026-10-06 11:28:13

标题: FBAN:用于隐私保护行为认证的完全同态加密兼容瓶颈注意力网络

摘要: 持续认证(CA)通过在设备交互过程中反复验证用户来增强会话安全性,然而它固有地依赖于高度敏感的行为特征(例如,细粒度的触摸动态),这些特征经常被外包给云/边缘服务进行可扩展推断。这带来了一个基本的隐私使用挑战:在计算过程中保护行为特征,而不仅仅是在传输或静态状态下。全同态加密(FHE)提供了一个原则性解决方案,但在FHE下部署现代CA模型仍然困难,因为非线性和注意力风格的操作会产生高密文成本。 我们提出了FBAN,一个TFHE兼容的瓶颈注意力网络和一个端到端加密的CA框架。FBAN设计用于仅整数执行,通过两阶段流水线(浮点预训练后跟量化感知训练)并编译为TFHE电路进行同态推断。我们进一步指定了一个客户端-服务器协议,其中包含基于会话的盲化和解密-返回验证,使服务器能够在不观察原始行为特征的情况下对用户进行认证。我们提供了针对诚实但好奇的服务器的TFHE IND-CPA安全性的加密安全分析,并形式化了对重放和冒充的抵抗性,而不需要TFHE秘钥。在两个公共触摸屏数据集上的实验表明,FBAN在加密推断下实现了强大的认证效用,同时保持了轻量级模型足迹,TFHE参数设定为≥128位安全性。

更新时间: 2026-10-06 11:28:13

领域: cs.CR

下载: http://arxiv.org/abs/2610.08174v1

Partially Observable Zero-shot coordination by Predicting Intention of Partner

Zero-shot coordination in embodied settings requires acting while the partner is intermittently out of view, leaving existing methods with ambiguous partner representations and uncertainty over hidden partner states. We propose Predicting Intention of Partner (PIP) to jointly address these challenges. PIP uses a Joint-view VAE to distill richer training-time evidence from the union of both agents' local observations into a partner representation available from local observations alone. Partner-state Belief networks further infer the partner's hidden location and behavioral tendencies from the ego agent's interaction history. We evaluate PIP in Burrito-PO, Overcooked-PO, and a Melting Pot substrate, together with a human evaluation in Burrito-PO. PIP attains the highest mean performance among the compared methods across all three benchmarks. Human evaluation and diagnostic analyses further support coordination with unseen partners and the contributions of both components under partner occlusion.

Updated: 2026-10-06 10:55:06

标题: 通过预测合作伙伴的意图实现部分可观测的零射击协调

摘要: 在具身环境中的零样本协调需要在伙伴间断性消失的情况下行动,这让现有方法在伙伴表示和隐藏伙伴状态上存在歧义和不确定性。我们提出了预测伙伴意图(PIP)来共同解决这些挑战。PIP使用联合视图VAE从两个代理的局部观察的并集中提取更丰富的训练时证据,形成仅从局部观察中得到的伙伴表示。伙伴状态信念网络进一步从自我代理的交互历史中推断出伙伴的隐藏位置和行为倾向。我们在Burrito-PO,Overcooked-PO和Melting Pot基质中评估了PIP,同时在Burrito-PO中进行了人类评估。PIP在所有三个基准测试中获得了与其他方法相比最高的平均性能。人类评估和诊断分析进一步支持与不可见伙伴的协调以及伙伴遮挡下两个组件的贡献。

更新时间: 2026-10-06 10:55:06

领域: cs.AI,cs.MA

下载: http://arxiv.org/abs/2610.08142v1

Test-Time Agent Evolution for Long-Horizon Legal Reasoning

Legal intelligence aims to support reliable decision-making across long-horizon legal processes involving evolving case states and multiple roles. However, real-world legal deployment exhibits substantial case heterogeneity in facts, evidence, and procedural contexts, exposing the limitations of static agent strategies. Moreover, legal reasoning is inherently interdependent across roles and procedural stages, making global reliability fundamentally different from isolated role competence. To address these challenges, we study training-free test-time agent adaptation, where agents continuously exploit deployment-time signals from preceding cases and ongoing interactions without updating model parameters. We propose \method, which introduces \emph{Test-Time Memory Evolution} to retrieve reusable experience from previous cases, adapt it to the current factual and procedural context, and consolidate accumulated experience for subsequent decision-making. Further, \emph{Rubric-Aligned Collaboration} verifies and revises role-specific actions according to behavioral and procedural requirements, enabling coordinated decision-making across roles and stages. Extensive experiments on J1-EVAL and LegalWorld across five backbone models demonstrate consistent improvements over representative reasoning and agent baselines with reasonable interaction and computational costs. Ablation and case studies further show that the two components provide complementary benefits in experience adaptation and cross-role coordination, improving the reliability and efficiency of long-horizon legal reasoning.

Updated: 2026-10-06 10:53:25

标题: 长时间视角法律推理的测试时间代理演化

摘要: 法律智能旨在支持涉及不断演变的案件状态和多个角色的长期法律流程中的可靠决策。然而,在现实世界的法律部署中,事实、证据和程序背景存在实质性的案例异质性,暴露了静态代理策略的局限性。此外,法律推理在角色和程序阶段之间固有地相互依赖,使全球可靠性与孤立角色能力根本不同。为了解决这些挑战,我们研究了无需培训的测试时间代理适应,其中代理不断利用先前案例和正在进行的互动中的部署时间信号,而无需更新模型参数。我们提出了\method,引入\emph{测试时间记忆演化},以从先前案例中检索可重复利用的经验,将其适应到当前的事实和程序背景,并为随后的决策巩固积累的经验。此外,\emph{基于标准的协作}根据行为和程序要求验证和修订特定角色的行动,实现跨角色和阶段的协调决策。在J1-EVAL和LegalWorld上进行的广泛实验涵盖了五种基础模型,展示了相对于代表性推理和代理基线的一致改进,且具有合理的交互和计算成本。消融和案例研究进一步表明,这两个组件在经验适应和跨角色协调方面提供了互补的好处,提高了长期法律推理的可靠性和效率。

更新时间: 2026-10-06 10:53:25

领域: cs.AI

下载: http://arxiv.org/abs/2610.08138v1

Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering

AI systems create images and videos with image/video generation models or by writing code and graphics descriptions that are then rendered. These routes can produce similar visible artifacts but expose different representations, intervention points, and provenance evidence. We develop a production-centered framework that compares detection and watermarking across both routes. An explicit verification specification distinguishes passive inference, message recovery, and authenticated provenance. We organize image, video, source-code, and rendering-aware watermarks by production stage. We examine the different requirements of generated images and video, plots and SVG, programmable video, and agent-composed workflows. Documented Claude, OpenAI, and rendering-tool interfaces connect the framework to concrete systems. We pose ten scoped research questions on identifiability, observability, fair comparison across stages, recoverable payload, reconstruction, synchronization, composition, hybrid local contribution, and private production-event authentication. The result is a conceptual research agenda grounded in published methods, inspected interfaces, and elementary boundary examples. It reports no experiments and claims no new theorems; its appendix results are elementary calculations, and documentation and source inspection establish interfaces, not empirical robustness.

Updated: 2026-10-06 10:53:19

标题: 重新思考视觉溯源:在直接视觉生成和LLM驱动的代码渲染中的检测和水印技术

摘要: AI 系统通过图像/视频生成模型或编写代码和图形描述来创建图像和视频。这些方法可以产生类似的可见工件,但暴露不同的表示、干预点和来源证据。我们开发了一个以生产为中心的框架,比较了这两种方法的检测和水印技术。一个明确的验证规范区分了被动推理、消息恢复和认证来源。我们通过生产阶段组织图像、视频、源代码和渲染感知的水印。我们研究了生成图像和视频、图表和 SVG、可编程视频以及代理组合工作流的不同需求。记录的 Claude、OpenAI 和渲染工具接口将框架连接到具体系统。我们提出了十个有限范围的研究问题,涉及可识别性、可观测性、跨阶段的公平比较、可恢复的有效载荷、重建、同步、组合、混合本地贡献和私人生产事件认证。结果是一个基于已发表的方法、检查过的接口和基本边界示例的概念性研究议程。它不报告实验结果,也不宣称有新的定理;其附录结果是基本的计算,文档和源代码检查建立了接口,而不是经验上的稳健性。

更新时间: 2026-10-06 10:53:19

领域: cs.CR,cs.CV

下载: http://arxiv.org/abs/2610.08137v1

Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations

Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.

Updated: 2026-10-06 10:47:26

标题: 超越边缘监测:针对大规模电子商务运营中数据概念漂移的分布式联合分布检测

摘要: 概念漂移对生产机器学习构成威胁,但多变量两样本漂移检测器在规模上的实证行为仍未得到充分描述。现有基准测试很少涉及数亿行和高基数特征的典型工业运营数据集。我们评估了五种多列两样本检测方法(边际、基于投影和核嵌入方法)在三个互补环境中的表现:哈佛数据集、经验证的“高声失败”再现(平均绝对误差在0.030至0.053之间)以及一个新颖的合成注入基准测试,使用了包含137.5亿行的Trendyol产品排名特征表。在两种严重程度范围内测试了四种漂移类型,我们展示了分布式最大均值差异检测器在Apache Spark上的健壮扩展性。在强制度中对四种漂移类型进行平均,并在校准阈值下,它实现了与预期漂移幅度的皮尔逊相关系数r = 0.940,真阳性率为80.4%,假阳性率为3.2%。相反,在对称抽样下,基于Kolmogorov-Smirnov检验的每维度测试由于ID样式列的统计饱和而失败,为大规模抽样设计建立了关键约束。在弱配置下(实现翻转分数最多为0.57%),检测器难以可靠地区分,突显了未来需要进行强度网格功率分析以区分基本敏感性限制和可扩展阈值转移的需求。

更新时间: 2026-10-06 10:47:26

领域: cs.DC,cs.LG,stat.ME,stat.ML

下载: http://arxiv.org/abs/2610.08132v1

SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy

User prompts provided to large language models (LLMs) may contain private information. One way to protect them is to execute the LLM inside a trusted execution environment (TEE). However, this results in slow inference times as current TEEs are significantly slower than GPUs for LLM inference. To circumvent this, Tramèr and Boneh (2019) proposed Slalom which splits neural network inference between a TEE and an untrusted GPU. They encrypt inputs to computations outsourced to the GPU. In this paper, we extend this split-inference architecture to LLM inference and instead protect intermediate inputs using differential privacy (DP). We first demonstrate that masking intermediate representations is necessary by showing an 80% accuracy on a prompt-reconstruction attack from these representations. Our main contribution is a global sensitivity analysis of key functions in LLM inference, which bounds the required scale of DP noise. Unlike encryption, DP avoids quantization, allowing the LLM to remain in the floating-point domain. We also derive an upper bound on the floating-point error from masking and subsequent noise cancellation as a function of the privacy parameter epsilon, keeping the same quality of the LLM response. We implement our architecture using the Intel TDX TEE and two LLMs: Llama-3.2-3B and Qwen3-4B. Our split execution is nearly twice as fast as fully TDX-based inference. Moreover, it is at most 43% faster than Slalom while achieving higher accuracy. Finally, we demonstrate that prompt reconstruction, even with knowledge of the DP mechanism, cannot recover more information than is contained in an unrelated prompt.

Updated: 2026-10-06 10:45:55

标题: SpliTEE:通过将GPU辅助受信任执行环境与差分隐私相结合,实现快速和私密的LLM推断

摘要: 用户提供给大型语言模型(LLMs)的提示可能包含私人信息。保护这些信息的一种方法是在受信任的执行环境(TEE)中执行LLM。然而,由于当前的TEE与LLM推理相比显着较慢,这导致推理时间缓慢。为了规避这一问题,Tramèr和Boneh(2019)提出了Slalom,它将神经网络推理分为TEE和不受信任的GPU之间。他们对外包给GPU的计算输入进行了加密。在本文中,我们将这种分离推理架构扩展到LLM推理,并使用差分隐私(DP)来保护中间输入。我们首先通过展示从这些表示中进行提示重建攻击的80%准确率来证明掩盖中间表示是必要的。我们的主要贡献是对LLM推理中关键函数的全局敏感性分析,该分析限制了需要的DP噪声规模。与加密不同,DP避免了量化,使LLM保持在浮点域中。我们还导出了掩盖和随后噪声抵消的浮点误差的上限,作为隐私参数epsilon的函数,保持LLM响应的相同质量。我们使用Intel TDX TEE和两个LLMs:Llama-3.2-3B和Qwen3-4B来实现我们的架构。我们的分离执行几乎比完全基于TDX的推理快一倍。此外,它比Slalom快至多43%,同时实现更高的准确性。最后,我们证明即使了解DP机制,提示重建也无法恢复比无关提示中包含的更多信息。

更新时间: 2026-10-06 10:45:55

领域: cs.CR,cs.AI,cs.LG

下载: http://arxiv.org/abs/2609.15039v3

Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors

Achieving compositional concept generalization (CoCoGen), the ability to understand novel situations by recombining learned primitives, remains a fundamental challenge in artificial intelligence. Compositional semantic models such as Compositional Distributional Semantics (DisCoCat) offer solutions by generalising vectors to tensors, but suffer from scaling bottlenecks when learning the tensors. Mapping DisCoCat onto Variational Quantum Circuits (VQCs) resolves this limitation for text, yet the methodology has not been expanded to multimodal situations such as the ones involved in CoCoGen. This paper introduces Mu-DisCoCat: a multimodal variational quantum learning framework for DisCoCat that achieves CoCoGen. The framework first learns stable object representations from single-object image-text pairs, then fixes these and uses them to learn the relations between them in multi-object situations. In classical simulations, the model used Uhlmann state fidelity to compute the overlap between the multimodal circuit representations and achieved higher relational OOD accuracy than the evaluated CLIP baseline. Its deployment was evaluated using the destructive SWAP test across noisy quantum emulators, including a range of IBM fake backends, IQM FakeAphrodite, and the IBM Marrakesh quantum processor. Despite real-world device noise, the hardware-executed models maintained a strong positive correlation with simulated fidelities, reliably distinguishing unseen similar and dissimilar pairs. Our work establishes a framework for executing CoCoGen on VQCs, demonstrating a viable use case for near-term quantum hardware.

Updated: 2026-10-06 10:45:43

标题: Mu-DisCoCat:一种用于量子处理器上的组合泛化的变分流水线

摘要: 实现组合概念概括(CoCoGen),即通过重新组合学习的基元来理解新颖情况的能力,仍然是人工智能中的一个基本挑战。像组合分布语义模型(DisCoCat)这样的组合语义模型通过将向量泛化为张量来提供解决方案,但在学习张量时存在扩展瓶颈。将DisCoCat映射到变分量子电路(VQCs)解决了文本的这一限制,但该方法尚未扩展到类似于CoCoGen中涉及的多模式情况。本文介绍了Mu-DisCoCat:一种用于DisCoCat的多模式变分量子学习框架,实现了CoCoGen。该框架首先从单对象图像-文本对中学习稳定的对象表示,然后固定这些表示,并使用它们来学习多对象情况下它们之间的关系。在经典模拟中,该模型使用Uhlmann状态保真度来计算多模态电路表示之间的重叠,并实现了比评估的CLIP基线更高的关系OOD准确性。其部署使用了嘈杂的量子仿真器进行评估,包括一系列IBM虚拟后端、IQM FakeAphrodite和IBM Marrakesh量子处理器。尽管存在真实设备噪声,硬件执行的模型与模拟保真度保持强烈的正相关性,可可靠地区分看不见的相似和不同对。我们的工作建立了一个在VQCs上执行CoCoGen的框架,展示了近期量子硬件的可行用例。

更新时间: 2026-10-06 10:45:43

领域: cs.LG,cs.CV

下载: http://arxiv.org/abs/2610.08131v1

Do LLMs Act on What They Know? From Partner Representations to Cooperative Actions

Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiving decisions, and frozen model weights. Across eight LLMs, linear probes recover intent conventions substantially more accurately than target conventions, yet receiving choices do not consistently agree with the sender's convention. We compare probe-predicted and ground-truth conventions presented either as general rules or as externally computed action recommendations. Rule statements yield modest and model-dependent changes in cooperation, whereas action translation produces larger gains on average. In a Qwen3-8B case study, matched-state statement reversals reveal much greater sensitivity to action recommendations than to rule statements. Activation transfers from oracle-action and non-oracle hint-restatement donors improve intent accuracy on both action classes, but the tested alternatives do not reliably reproduce these benefits. Together, these results distinguish convention decodability, sensitivity to convention information, and cooperative performance, and highlight limitations in turning available partner information into receiving decisions.

Updated: 2026-10-06 10:44:55

标题: LLMs是否会根据所知行动?从合作伙伴代表到合作行动

摘要: 与陌生合作伙伴合作需要适应事先不知道的沟通惯例。我们在一个受控的Hanabi衍生环境中研究了这个问题,其中包括脚本化的提示生成、LLM控制的接收决策和冻结的模型权重。在八个LLM中,线性探针恢复意图惯例的准确性要远远高于目标惯例,然而接收选择并不始终与发送者的惯例一致。我们比较了探针预测和基本事实的惯例,这些基本事实可以作为一般规则或外部计算的行动建议呈现。规则陈述对合作产生了适度且依赖模型的变化,而行动转换平均产生了更大的收益。在Qwen3-8B案例研究中,匹配状态陈述的逆转显示出对行动建议的敏感度远远大于对规则陈述的敏感度。来自神谕行动和非神谕提示重述捐赠者的激活转移改善了两类行动的意图准确性,但测试的替代方案并不能可靠地再现这些好处。总的来说,这些结果区分了惯例的可解码性、对惯例信息的敏感性和合作性能,并突显了将可用合作伙伴信息转化为接收决策的局限性。

更新时间: 2026-10-06 10:44:55

领域: cs.LG

下载: http://arxiv.org/abs/2610.08129v1

Supermarket Product Detection and Recognition: Utilizing Deep Learning with Rectified Imagery

Product Identification has sprung up to become one of the most challenging problems in the automation of the retail industry. With the new industry 5.0 standards, automated inventory management, and catalog creation tasks are vitally important. Object identification models have emerged as a viable answer with their unprecedented identification and localization accuracy. However, the close-knit rack design of supermarkets generates the problem of angle variation in capturing images. The angle-variant densely packed images(a single image contains many objects) become overwhelming for these models alone. In this paper, we try to supplement object detection models with traditional Hough transform (HT) and homogeneous estimation concepts. We study the effect of rectified images using homography estimation and hough transform and their limitations on the problem of grocery identification. We make a case for creating a new dataset to test the effects of such rectification and produce analytical results on different scenarios of angle variation and object densities per image. Extensive experiments on different object detection models suggest that image rectification of angled images improves the detection accuracy of grocery products in images. The results also highlight the limitation of rectification on the angle of image capture and the object density of the image.

Updated: 2026-10-06 10:44:33

标题: 超市产品检测和识别:利用修正后的图像深度学习

摘要: 产品识别已成为零售业自动化中最具挑战性的问题之一。随着新的5.0行业标准的出现,自动化库存管理和目录创建任务变得至关重要。物体识别模型以其前所未有的识别和定位准确性成为了一个可行的解决方案。然而,超市紧凑的货架设计导致了在捕捉图像时角度变化的问题。角度变化密集包装图像(一个图像包含多个物体)对这些模型来说变得非常困难。在本文中,我们尝试用传统的霍夫变换(HT)和同质估计概念来补充物体检测模型。我们研究了使用单应性估计和霍夫变换矫正图像的效果,以及它们在杂货识别问题上的局限性。我们提出创建一个新的数据集来测试此类矫正的效果,并在不同角度变化和图像中物体密度的不同情景下产生分析结果。对不同物体检测模型的大量实验表明,对倾斜图像进行图像矫正可以提高图像中杂货产品的检测准确性。结果还突显了图像捕捉角度和图像中物体密度对矫正的限制。

更新时间: 2026-10-06 10:44:33

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.08126v1

Beyond Waypoint Regression: Query-Based Cost Learning over Reachable Ego Futures for End-to-End Driving

End-to-end planners based on waypoint regression achieve strong open-loop accuracy, but they primarily learn to mimic expert geometry and remain difficult to adapt to deployment-time safety constraints. We propose a query-based cost-learning framework that estimates bounded costs for dynamically reachable ego trajectory queries, rather than dense BEV cells or a small regressed trajectory set. Compact joint scene tokens capture coherent multimodal agent futures, while contingency-aware cost aggregation and cost-guided intra-cluster MPPI mixing convert the learned cost topology into feasible ego plans. On nuScenes, our method improves over prior cost-estimation planners such as ST-P3 and NMP, outperforms most regression baselines in collision rate, while remaining competitive in L2, and retaining an interpretable cost interface. On real-world driving logs, the proposed planner reduces collision rates compared with SparseDrive and Alpamayo without fine-tuning, while maintaining a diverse set of candidate trajectories.

Updated: 2026-10-06 10:43:13

标题: 超越航点回归:基于查询的可达自我未来成本学习,用于端到端驾驶

摘要: 基于路点回归的端到端规划器实现了较强的开环准确性,但它们主要学习模仿专家几何并且很难适应部署时的安全约束。我们提出了一种基于查询成本学习的框架,该框架估算动态可达的自我轨迹查询的有界成本,而不是密集的BEV单元格或一个小的回归轨迹集。紧凑的联合场景令牌捕获一致的多模态代理未来,而考虑后备方案的成本聚合和成本引导的簇内MPPI混合将学习的成本拓扑转换为可行的自我规划。在nuScenes上,我们的方法改进了先前的成本估计规划器,如ST-P3和NMP,在碰撞率方面优于大多数回归基线,同时在L2方面保持竞争力,并保留可解释的成本界面。在实际驾驶记录中,所提出的规划器与SparseDrive和Alpamayo相比,无需微调即可降低碰撞率,同时保持一组多样化的候选轨迹。

更新时间: 2026-10-06 10:43:13

领域: cs.RO,cs.AI,cs.LG,eess.SY

下载: http://arxiv.org/abs/2610.08123v1

KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat

We introduce KadiAssistant, a privacy-by-design AI assistant integrated into the Kadi research data ecosystem, enabling researchers to efficiently access, aggregate, and synthesize information from heterogeneous, privacy-sensitive research data. Interdisciplinary fields such as materials science bring together disciplines with their own terminology and standards. While this convergence fuels innovation, it also makes it increasingly difficult to connect and access knowledge, as data are distributed across disciplines, organizations, and individuals. For example, battery research combines electrochemical measurements, materials characterization data, physics-based simulations, and manufacturing parameters, each using different formats, vocabularies, and standards. Efficiently storing and sharing such heterogeneous data via research data platforms, such as Kadi4Mat, demands domain knowledge, technical expertise, and familiarity with metadata schemas and interfaces. Research data also vary in sensitivity: newly generated 'warm' data are often private, whereas published 'cold' data are usually openly accessible. The Kadi ecosystem offers fine-grained access control needed for sensitive data. A solution for efficient information retrieval in Kadi must therefore respect the fine-grained access permissions. To address these intertwined challenges of information retrieval, strong data privacy, and complex access control, KadiAssistant combines a self-hosted large language model (LLM) with a privacy-preserving semantic search, inspired by retrieval-augmented generation, that can access files and record metadata on Kadi. This allows the assistant to screen, aggregate, and structure information into a highly informative answer. KadiAssistant therefore bridges terminology and standards, lowers access barriers for researchers, and strengthens the Findable pillar of FAIR data principles.

Updated: 2026-10-06 10:42:08

标题: KadiAssistant:Kadi4Mat中用于信息检索的对话式人工智能代理

摘要: 我们引入了KadiAssistant,这是一个隐私设计的人工智能助手,集成到Kadi研究数据生态系统中,使研究人员能够高效地访问、聚合和综合来自异构、隐私敏感的研究数据的信息。跨学科领域如材料科学将各自具有其术语和标准的学科聚集在一起。虽然这种融合推动创新,但也使得连接和访问知识变得越来越困难,因为数据分布在不同的学科、组织和个人之间。例如,电池研究结合了电化学测量、材料表征数据、基于物理的模拟和制造参数,每种都使用不同的格式、词汇和标准。通过研究数据平台(如Kadi4Mat)高效地存储和共享这种异构数据需要领域知识、技术专业知识和熟悉元数据模式和接口。研究数据的敏感性也各不相同:新生成的“热”数据通常是私有的,而已发布的“冷”数据通常是公开可访问的。Kadi生态系统提供了敏感数据所需的细粒度访问控制。因此,Kadi中高效的信息检索解决方案必须尊重细粒度的访问权限。为了解决信息检索、强大的数据隐私和复杂访问控制这些相互关联的挑战,KadiAssistant结合了自托管的大型语言模型(LLM)和隐私保护的语义搜索,受到检索增强生成的启发,可以访问Kadi上的文件和记录元数据。这使得助手能够筛选、聚合和将信息结构化为一个高度信息丰富的答案。因此,KadiAssistant桥接术语和标准,降低研究人员的访问障碍,并增强FAIR数据原则中可找到的支柱。

更新时间: 2026-10-06 10:42:08

领域: cs.IR,cs.AI

下载: http://arxiv.org/abs/2605.18850v2

Attenuated in-context identification in time-series foundation models: diagnosis under counterfactual inputs and repair by synthetic forced-system fine-tuning

Covariate-aware time-series foundation models (TSFMs) promise training-free what-if answers for instrumented plants: the change in output that a different future input would cause. We test this on forced engineering systems with exact counterfactuals, comparing Chronos-2, TimesFM-2.5 and TabPFN-TS with classical system identification fitted to the same context. Through their default covariate interfaces, TimesFM-2.5 and TabPFN-TS are memoryless: the predicted effect of an input change is a same-time function of that change ($R^2 = 1.000$ for TimesFM-2.5). Chronos-2 identifies dynamics in context but attenuates them. Its predicted effect is 0.33-0.80 of the true effect, its recovered impulse response has the wrong shape, and its error on a one-degree-of-freedom oscillator levels off at 0.57 with 8192 context samples, where ARX fitted to 256 samples reaches 0.02. Context dither at inference lowers the what-if error on all six synthetic classes without training. A 26-minute fine-tune on synthetic forced systems restores the response magnitude (sensitivity 0.83-0.96) and outperforms structure-agnostic identification on Wiener-Hammerstein and a held-out friction class. A specialised in-context identifier trained on the same data comes close, so the forced-system data carry most of the gain. On three of four measured plants classical identification remains clearly better, and the fine-tuned model loses part of its univariate forecasting skill. Paired counterfactual inputs, together with shuffled future inputs on measured records, test two properties: whether the covariate interface can represent dynamics and whether the pretraining prior covers the plant's time scale. Only the counterfactual pairs expose the attenuation.

Updated: 2026-10-06 10:39:04

标题: 基于时间序列基础模型的上下文识别减弱:在对抗性输入下的诊断和通过合成强制系统微调修复

摘要: 共变量感知时间序列基础模型(TSFMs)承诺对仪器化植物进行无需训练的假设情况答案:不同未来输入会导致输出变化。我们在强制工程系统上进行测试,这些系统具有确切的对照情况,比较了Chronos-2、TimesFM-2.5和TabPFN-TS与经典系统识别拟合到相同背景的情况。通过它们的默认共变量接口,TimesFM-2.5和TabPFN-TS是无记忆的:输入变化的预测效果是该变化的同时函数($R^2 = 1.000$ for TimesFM-2.5)。Chronos-2在上下文中识别动态但减弱了它们。其预测效果为真实效果的0.33-0.80,其恢复的脉冲响应形状错误,并且在具有8192个上下文样本的单自由度振荡器上的误差达到0.57,而ARX拟合到256个样本时达到0.02。推理时的上下文抖动降低了所有六个合成类别的假设错误,而无需训练。在合成强制系统上进行26分钟的微调可以恢复响应幅度(敏感度0.83-0.96),并且在Wiener-Hammerstein和一个保留摩擦类别上优于结构不可知识别。在相同数据上训练的专门的上下文识别器接近,因此强制系统数据带来大部分收益。在四个测量植物中,三个经典识别仍然明显更好,并且经过微调的模型失去了部分单变量预测能力。配对的假设输入,以及在测量记录上混洗的未来输入,测试了两个属性:共变量接口能否表示动态以及预训练先验是否覆盖了植物的时间尺度。只有假设对才暴露出减弱。

更新时间: 2026-10-06 10:39:04

领域: cs.LG,eess.SY

下载: http://arxiv.org/abs/2610.08118v1

When Explanations Compete: Policy-Aware Selection Under Uncertainty

Uncertainty-aware explanation methods often produce several alternatives for the same prediction. Selecting among them requires a policy for balancing prediction confidence, uncertainty, and application constraints. This paper presents a framework for applying such policies to a fixed set of generated explanations. Candidates are characterised by uncertainty change, prediction direction, and, when available, interval position relative to a decision boundary. The framework combines these properties with eligibility rules, optional bidirectional Pareto screening, and policy-aware ranking. A fictitious prostate-cancer example illustrates how different explanatory purposes lead to different selections from the same candidate set. We instantiate the framework with Calibrated Explanations for classification, thresholded regression, and plain regression. Across 41 benchmark datasets, mean candidate counts range from 11.57 to 21.75 for single-feature explanations and from $29.48$ to $69.53$ when conjunctions are included. Equal-weight and confidence-only policies yield an average selection-disagreement rate of $28.7\%$ while favouring the same confidence direction. A supporting $δ$-CLUE experiment demonstrates use with a second generator. By making the selection policy explicit, the framework allows applications to compare and prioritise explanations according to their intended use.

Updated: 2026-10-06 10:38:02

标题: 当解释相互竞争:不确定情况下的政策感知选择

摘要: 不确定性感知解释方法通常会为同一预测产生几种不同的替代方案。在它们之间进行选择需要一个平衡预测置信度、不确定性和应用约束的策略。本文提出了一个框架,用于将这些策略应用于一组固定生成的解释。候选解释特征包括不确定性变化、预测方向,以及在可用时相对于决策边界的区间位置。该框架将这些特性与资格规则、可选的双向帕累托筛选和策略感知排序结合起来。一个虚构的前列腺癌例子说明了不同的解释目的如何导致从相同的候选集中进行不同的选择。我们使用校准解释来实例化该框架,用于分类、阈值回归和普通回归。在41个基准数据集中,单要素解释的平均候选数量范围从11.57到21.75,包括连词时则为29.48到69.53。等权重和仅置信度策略的平均选择不一致率为28.7%,同时偏好相同的置信度方向。一个支持的δ-CLUE实验展示了与第二生成器的使用。通过明确选择策略,该框架使应用能够根据其预期用途比较和优先选择解释。

更新时间: 2026-10-06 10:38:02

领域: cs.AI,cs.LG

下载: http://arxiv.org/abs/2410.05479v2

Energy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric Vehicles

Path-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles. The latter objective is especially relevant for Electric Vehicles (EVs), since their limited driving range can be extended by recovering energy through regenerative braking, a feature that has not yet been sufficiently studied in the literature. In this work, we perform a comparative analysis of four controllers under one common Frenet frame-based kinematic vehicle model, utilizing a validated energy model (VT-CPEM) with explicit regenerative braking. Herein, we implement the following controllers: Nonlinear Model Predictive Control (NMPC), Proximal Policy Optimization (PPO), gain-scheduled Ackermann state-feedback baseline (PID-SF), and a Stanley geometric baseline. To satisfy real-time requirements, we implement the NMPC using JIT-compiled CasADi. Moreover, we train the PPO using traditional straight and S-curve tracks, after which we successfully transfer the unmodified policy to unseen tracks, including: an ISO 3888-1 lane-change, a chicane, randomly-generated parameterized-splines, and a $\pm3^\circ$ graded road. In addition, the policy transfers to a dynamic single-track vehicle model with linear tires, zero-shot with an acceptable initial performance, which was optimized after brief fine-tuning. Thereby, we demonstrate that our PPO is readily transferable to more comprehensive vehicle models. We conclude with a performance analysis of developed controllers and discuss ideas for future work.

Updated: 2026-10-06 10:35:13

标题: 能源感知路径跟踪:电动汽车的强化学习和NMPC的比较分析

摘要: 路径跟随控制策略通常遵循双目标优化困境:在保持平滑速度曲线的同时最小化与参考路径的偏差。后者的目标对于电动车辆(EVs)尤为重要,因为它们有限的行驶里程可以通过再生制动来延长,这一特性在文献中尚未得到充分研究。在这项工作中,我们在一个常见的Frenet框架基于运动学车辆模型下进行了四种控制器的比较分析,利用了一个经过验证的具有显式再生制动的能量模型(VT-CPEM)。在此基础上,我们实施了以下控制器:非线性模型预测控制(NMPC)、近端策略优化(PPO)、增益调度Ackermann状态反馈基线(PID-SF)和Stanley几何基线。为满足实时要求,我们使用即时编译的CasADi实施了NMPC。此外,我们训练PPO使用传统直线和S型曲线轨道,然后成功将未经修改的策略转移到未知轨道,包括:ISO 3888-1换道、迂回、随机生成的参数化样条曲线和一个$\pm3^\circ$坡道。此外,该策略还可以转移到具有线性轮胎的动态单轨车辆模型,零样本学习后具有可接受的初始性能,并在简短微调后进行了优化。因此,我们展示了我们的PPO可以轻松转移到更全面的车辆模型。最后,我们对开发的控制器进行性能分析,并讨论未来工作的想法。

更新时间: 2026-10-06 10:35:13

领域: cs.RO,cs.LG,eess.SY

下载: http://arxiv.org/abs/2610.08112v1

Shared Phase and Retention Control for Efficient Adaptive Spectral Recurrence

As new evidence arrives, a sequence model must update what it remembers and how memory influences predictions. While Transformers incur computation and cache costs scaling with context length, fixed-state recurrent models offer constant-memory inference. However, linear and spectral recurrences traditionally rely on static transitions, failing to dynamically revise how stored representations decay or rotate. While recent selective architectures introduce input-dependent transitions, they assign independent controls to every memory mode, coupling control cost to state capacity. We show that high-dimensional spectral memory does not require high-dimensional control, and introduce Shared Phase and Retention Control for Efficient Adaptive Spectral Recurrence (SPARC). SPARC employs just two input-dependent scalar signals to coordinate memory retention and phase rotation across heterogeneous complex modes, while preserving mode-specific baseline timescales and frequencies. Its diagonal affine recurrence supports parallel associative scans for sequence-level BPTT as well as exact structured Real-Time Recurrent Learning (RTRL) for online credit assignment. Across partially observable continuous control, POPGym, and sequence classification, SPARC achieves a 9.09% relative return improvement on Walker-P and a 1.36% relative accuracy gain on FordA over second-best methods. On an NVIDIA Blackwell GPU, our implementation reduces recurrent-mixer training latency by 18.2%-34.2% in fixed-token workloads and accelerates scans by 3.1x-4.7x over an optimized RG-LRU baseline. These results show that two shared control signals can efficiently govern adaptive spectral memory across online and full-sequence settings. Code is available at https://github.com/Botwwt/sparc.

Updated: 2026-10-06 10:34:13

标题: 共享相位和保留控制用于高效自适应频谱重复

摘要: 随着新证据的出现,序列模型必须更新其记忆内容以及记忆如何影响预测。尽管Transformers在上下文长度的扩展时产生计算和缓存成本,固定状态的循环模型提供了恒定内存推理。然而,线性和谱循环传统上依赖静态转换,未能动态修订存储表示的衰减或旋转方式。虽然最近的选择性架构引入了依赖于输入的转换,但它们为每个记忆模式分配独立控制,将控制成本与状态容量耦合。我们展示了高维谱记忆不需要高维控制,并引入了用于高效自适应谱循环的共享相位和保留控制(SPARC)。SPARC仅使用两个依赖于输入的标量信号来协调跨异质复杂模式的记忆保留和相位旋转,同时保持模式特定的基准时间尺度和频率。它的对角仿射循环支持用于序列级BPTT的并行关联扫描,以及用于在线信用赋值的精确结构化实时循环学习(RTRL)。在部分可观察的连续控制、POPGym和序列分类中,SPARC在Walker-P上相对回报率提升了9.09%,在FordA上相对准确率提升了1.36%,优于次佳方法。在NVIDIA Blackwell GPU上,我们的实现将固定令牌工作负载中的循环混合器训练延迟降低了18.2%-34.2%,并将扫描加速了3.1倍至4.7倍,超过了优化的RG-LRU基线。这些结果表明,两个共享控制信号可以有效地管理在线和完整序列设置中的自适应谱记忆。代码可在https://github.com/Botwwt/sparc上找到。

更新时间: 2026-10-06 10:34:13

领域: cs.LG

下载: http://arxiv.org/abs/2609.39082v2

Enhancing Diffusion Language Models with Autoregressive Post-Training Weights

Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding. Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations. After the conversion, however, they typically ignore the extensive post-training ecosystem of their AR ancestors. In this work, we show that these existing AR post-training weight updates can instead be effectively recycled to enhance diffusion models. Despite the changes by AR-to-diffusion conversion, directly adding an AR post-training weight update to a diffusion base model remains effective, bringing its performance close to that achieved by direct diffusion post-training. Notably, AR and diffusion post-training updates are nearly orthogonal in weight space, yet induce substantially more aligned representation changes in the diffusion model. Their distinct updates are also complementary: composing their weights can retain gains from both regimes and further improve the post-trained diffusion model. Based on these findings, we propose A2D, a simple training-free framework for enhancing diffusion models with existing AR post-training resources. A2D can transfer capabilities from AR post-trained models to diffusion base models, and further improve already post-trained diffusion models by composing AR and diffusion post-training updates. Across various dLLMs, including Dream, DreamReasoner, DiffuCoder, Dream-Coder, Nemotron-Labs-Diffusion, and DiffusionGemma, A2D reliably improves instruction following, mathematical reasoning, and coding with both supervised fine-tuning and reinforcement learning updates, without additional training, or inference-time computation.

Updated: 2026-10-06 10:31:54

标题: 用自回归后训练权重增强扩散语言模型

摘要: 扩散语言模型(dLLMs)已经成为一种有前途的替代自回归(AR)语言模型的选择,提供了灵活的令牌更新顺序和并行解码。最近的dLLMs通常在扩散转换之前从预训练的AR模型中初始化,以继承它们学习到的表示。然而,在转换后,它们通常忽略了其AR祖先的广泛后训练生态系统。在这项工作中,我们展示了这些现有的AR后训练权重更新可以有效地被回收利用来增强扩散模型。尽管经过AR到扩散的转换发生了变化,但直接将AR后训练权重更新添加到扩散基础模型仍然是有效的,使其性能接近直接扩散后训练所达到的水平。值得注意的是,AR和扩散后训练更新在权重空间中几乎是正交的,但在扩散模型中引起了更加对齐的表示变化。它们各自独特的更新也是互补的:组成它们的权重可以保留两种方案的收益,并进一步改进后训练的扩散模型。基于这些发现,我们提出了A2D,一个简单的无需训练的框架,用于利用现有的AR后训练资源增强扩散模型。A2D可以将AR后训练模型的能力转移到扩散基础模型,并通过组合AR和扩散后训练更新进一步改进已经后训练的扩散模型。在各种dLLMs中,包括Dream、DreamReasoner、DiffuCoder、Dream-Coder、Nemotron-Labs-Diffusion和DiffusionGemma,A2D可靠地改善指令遵循、数学推理和编码,无需额外的训练或推理时计算。

更新时间: 2026-10-06 10:31:54

领域: cs.LG

下载: http://arxiv.org/abs/2610.08108v1

Exploiting Acoustic and Content-Oriented Speaker Verification Attacks Against Multilingual Voice Anonymization

Attacker ASV systems for voice anonymization have been studied primarily in English, leaving their behavior in multilingual settings largely unexplored. Conventional ASV has shown that both acoustic and contextual information are important for multilingual speaker verification. Inspired by this, we investigate whether the same holds for attacker ASV on anonymized speech. We evaluate both acoustic- and content-oriented attackers on multilingual anonymized speech and construct a multilingual voice-converted dataset to improve cross-lingual generalization. Our results show that attacker effectiveness depends on the linguistic utility of the anonymized speech. Overall, acoustic-oriented attackers achieve better performance. However, when linguistic information is well preserved, the performance gap between content- and acoustic-oriented attackers narrows compared with conditions involving stronger speech distortion. The multilingual voice-converted dataset further improves performance and partially reduces the cross-lingual gap. These findings highlight the need for more comprehensive attacker modeling and evaluation protocols that consider both privacy and utility, rather than relying on a attacker strategy\footnote{Full code and pretrained models and MultiVC Dataset link are available at: https://github.com/monkeyDarefeen/DAST

Updated: 2026-10-06 10:31:44

标题: 利用声学和基于内容的说话者验证攻击针对多语言语音匿名化

摘要: 攻击者ASV系统用于语音匿名化主要在英语中进行研究,而它们在多语言环境中的行为仍然未被充分探索。传统ASV已经表明,对于多语种说话者验证,声学和上下文信息都很重要。受此启发,我们研究了攻击者ASV在匿名化语音上是否也是如此。我们评估了针对多语种匿名化语音的声学和内容导向的攻击者,并构建了一个多语种语音转换数据集,以改善跨语种的泛化性能。我们的结果表明,攻击者的有效性取决于匿名化语音的语言效用。总体而言,声学导向的攻击者表现更好。然而,当语言信息被很好地保留时,与涉及更强语音失真条件相比,内容和声学导向的攻击者之间的性能差距缩小。多语种语音转换数据集进一步提高了性能,并部分减少了跨语言的差距。这些发现强调了需要更全面的攻击者建模和评估协议,考虑隐私和效用,而不是依赖于攻击者策略。

更新时间: 2026-10-06 10:31:44

领域: cs.SD,cs.AI

下载: http://arxiv.org/abs/2610.08107v1

ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an on-demand construction process: information-rich charts make chart question answering (Chart QA) suitable for probing coupled perception and reasoning. Automated Chart QA construction is intended to shorten the benchmark-development cycle by turning identified gaps into targeted samples on demand. Current methods, however, commonly separate target guidance from scratch generation: target-guided systems often require prepared data, charts, or templates, while scratch-generation systems primarily ensure artifact validity, without explicitly controlling whether newly synthesized requirements and content remain aligned with an externally specified diagnostic target. We introduce ChartBmkAgent, which turns an identified capability gap into targeted diagnostic evidence by constructing complete Chart QA samples from sparse error-taxonomy specifications. Throughout construction, a central harness governs specialized agents, requires stage-specific evidence of alignment with the original error category, and records the basis for each acceptance decision. On 300 taxonomy-wide samples, MLLM accuracies ranged from 32.7% to 84.3% with distinct category profiles, showing that generated samples reveal capability differences. Across three source-model comparisons, targeted follow-ups scored 50.0% versus 82.2% on matched controls ($p=8.96\times10^{-6}$); all six cross-model comparisons had the same direction, demonstrating targeted validation and diagnostic-data generation. Multiple evaluator models assessed whether each sample tested its specified error category; 86.4% met this criterion, providing empirical evidence of target preservation.

Updated: 2026-10-06 10:31:28

标题: ChartBmkAgent: 利用稀疏错误分类规范进行受控多代理构建图表问答基准

摘要: 多模态大语言模型(MLLMs)快速发展,而传统的基准开发滞后,延迟了对新观察到的能力差距进行调查。这种调查需要一个富有表现力的任务格式和一个按需构建过程:信息丰富的图表使图表问答(Chart QA)适合探究耦合的感知和推理。自动化的图表问答构建旨在通过将识别出的差距转化为按需的有针对性样本,缩短基准开发周期。然而,当前方法通常将目标指导与从头生成分开:目标引导系统通常需要准备好的数据、图表或模板,而从头生成系统主要确保工件的有效性,而不明确控制新合成的需求和内容是否与外部指定的诊断目标保持一致。我们介绍了ChartBmkAgent,它通过从稀疏的错误分类规范构建完整的Chart QA样本,将识别出的能力差距转化为有针对性的诊断证据。在构建过程中,一个中央控制系统管理专门的代理,要求特定阶段的证据与原始错误类别保持一致,并记录每个接受决定的基础。在300个涵盖整个分类的样本中,MLLM的准确率从32.7%到84.3%不等,具有不同的类别特征,显示生成的样本揭示了能力差异。在三个源模型比较中,有针对性的跟进得分为50.0%,而匹配控制组为82.2%($p=8.96\times10^{-6}$);所有六个跨模型比较都具有相同的方向,证明了有针对性的验证和诊断数据生成。多个评估模型评估了每个样本是否测试了其指定的错误类别;86.4%满足了这一标准,提供了目标保持的实证证据。

更新时间: 2026-10-06 10:31:28

领域: cs.AI

下载: http://arxiv.org/abs/2610.08106v1

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.

Updated: 2026-10-06 10:30:14

标题: DSV-Mem:评估MLLM代理在专业工作流程中的多模态记忆

摘要: 对话式MLLM代理越来越被期望在专业工作流程中发挥作用,从人工智能研究和工程设计到产品管理和业务运营。然而,这种能力仍未被充分探索:现有的基准主要集中在非正式、日常互动和个人生活场景,包括自然图像、孤立的静态物品和以回忆为导向的问题。相比之下,专业场景通常涉及结构化、信息密集的物品,这些物品经常进行修订和权威更新,以及需要调和许多物品版本并精确跟踪状态的组合查询。为了解决这些挑战,我们引入了DSV-Mem,一个用于评估密集有状态视觉记忆的基准。DSV-Mem包括专家审查的场景和1000个问题,涵盖五个用户导向的类别(当前状态、过去状态、派生状态、变更历史和冲突/拒绝)。受哈特利启发的标准偏爱具有更广泛视觉证据检查需求的问题。我们还引入了一个生成工具,通过将状态转换合成与对话填充解耦,生成评估套件。对包括前沿和开放权重模型以及内存管理方法在内的27种配置进行评估,结果显示最强基线在DSV-Mem上得分低于45%。分析表明:1)多模态和信息密度都会增加难度,但状态演变,特别是统治性更新的数量,是主要的测试因素。原始对话/大草堆长度、OCR和算术不是主要瓶颈;2)模型在回答之前经常未能验证用户前提与先前状态更新之间的一致性;3)增加推理工作量和内存管理方法带来的收益有限,而状态感知设计表现更为有效。该基准和代码将公开发布。

更新时间: 2026-10-06 10:30:14

领域: cs.AI

下载: http://arxiv.org/abs/2610.08102v1

Beyond Corrected Memory: Execution Consistency in Multi-Agent Systems

Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselves establish whether it is sufficient to judge task duties. We define execution consistency through duties governing state use, information handoffs, and final-state agreement, with explicit evidence conditions for judging fulfillment. Our core claim is that identical retained records can correspond to compliant and violating executions under the same task rule. Controlled removal of evidence such as receipt, action dependence, or response validity leaves 82.4% of opposite-label pairs indistinguishable; restoration separates 97.9% of the merged pairs. Natural-log annotations identify the defined violations in actual executions. However, existing logs do not always explicitly represent the execution relationships needed for these judgments. To assess the definition's practical value, we use CAVERT, a framework for consistency diagnosis and recovery, to extract supported relationships from logs and apply these criteria. It consistently outperforms contract-prompted LLM and rule-based baselines in diagnosis across all 12 benchmark-executor settings. Under the same gate and executor limits, it also outperforms rule-guided recovery in all four evaluated environments. These findings identify execution evidence that agent-memory and execution interfaces should preserve for reliable judgment.

Updated: 2026-10-06 10:29:41

标题: 超越纠正的记忆:多智能体系统中的执行一致性

摘要: 共享内存协调代理的行动,但正确的记录并不能确保这些行动符合任务要求。内存治理和故障诊断规范或检查记录的信息;它们本身并不能确立这些信息是否足以判断任务职责。我们通过管理状态使用、信息交接和最终状态协议来定义执行一致性,为判断履行提供了明确的证据条件。我们的核心观点是,相同的保留记录可以对应于在相同任务规则下符合和违规执行。对证据的受控移除,如收据、行动依赖或响应有效性,使82.4%的对立标记对无法区分;恢复将97.9%的合并对分离。自然对数注释在实际执行中识别了定义的违规行为。然而,现有的日志并不总是明确地代表这些判断所需的执行关系。为了评估定义的实际价值,我们使用CAVERT,一个用于一致性诊断和恢复的框架,从日志中提取支持的关系,并应用这些标准。在所有12个基准执行器设置中,它在诊断中始终优于合同提示的LLM和基于规则的基线。在相同的门限和执行器限制下,在所有四个评估环境中,它也优于基于规则的恢复。这些发现确定了代理-内存和执行接口应保留以进行可靠判断的执行证据。

更新时间: 2026-10-06 10:29:41

领域: cs.AI

下载: http://arxiv.org/abs/2610.08101v1

Surviving the Router: Optimizing Skill Injections for Retrieval and Execution

AI agents increasingly rely on modular third-party "skills" that are dynamically selected by skill routers to execute complex tasks. While recent studies highlight the threat of prompt injections embedded in these skills, existing evaluations often assume settings where the malicious skill is already selected for execution. We show that this assumption can substantially overestimate attack success. In realistic multi-skill environments, injected skills must first compete for retrieval, reducing the effective attack success rate (ASR) of existing injections by 87-97%. To address this limitation, we introduce CORSA (Cluster Optimization for Router-Aware Skill Attacks), a router-aware attack that optimizes skill injections for both retrieval and execution across clusters of related tasks. We evaluate skill injection attacks under router-managed multi-skill settings by extending the benchmark introduced by SkillRouter with eight malicious payload categories. CORSA uses successive optimization stages to first improve retrieval and then optimize end-to-end attack success, while we evaluate user utility and injection naturalism separately. Our experiments show that CORSA substantially improves both retrieval and end-to-end attack success over existing skill injections while preserving user utility, and that the resulting attacks transfer across different router architectures and LLM backbones.

Updated: 2026-10-06 10:28:27

标题: 生存路由器:优化技能注入以进行检索和执行

摘要: 人工智能代理越来越依赖于由技能路由器动态选择的模块化第三方“技能”来执行复杂任务。尽管最近的研究强调了嵌入这些技能中的提示注入的威胁,但现有的评估通常假设恶意技能已被选择执行。我们展示了这种假设可能大大高估了攻击成功率。在现实的多技能环境中,注入的技能必须首先竞争检索,从而将现有注入的有效攻击成功率(ASR)降低了87-97%。为了解决这一限制,我们引入了CORSA(路由器感知技能攻击的集群优化),这是一种针对相关任务集群优化技能注入的路由器感知攻击。我们通过使用包含八个恶意有效载荷类别的基准测试扩展SkillRouter,在路由器管理的多技能设置下评估了技能注入攻击。CORSA使用连续的优化阶段首先改善检索,然后优化端到端攻击成功,同时我们分别评估用户效用和注入的自然性。我们的实验表明,CORSA在提高检索和端到端攻击成功率的同时,保持了用户效用,并且所产生的攻击可以在不同的路由器架构和LLM骨干上转移。

更新时间: 2026-10-06 10:28:27

领域: cs.CR,cs.LG

下载: http://arxiv.org/abs/2610.08098v1

When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback

Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems. We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification. Without verification, corruption causes dramatic accuracy loss, from 100% down to 72.4%. Mandatory reflection fully recovers this performance to 100%. Optional verification improves accuracy only when models actively invoke it. Our results show that checking frequency is strongly associated with robustness differences, while unequal invocation prevents a controlled comparison of verifier quality. A supporting recovery experiment shows that full problem restart succeeds in 100% of cases after explicit detection. These findings demonstrate that verifier availability and verification policy are separate components of mathematical-agent reliability. Mandatory policies enforce verification while optional policies depend on the model's own choice to invoke it.

Updated: 2026-10-06 10:28:19

标题: 当工具撒谎:在受损工具反馈下数学代理的可靠性

摘要: 数学问题解决通常需要确定性的计算步骤,代理人将这些步骤委托给工具并隐式信任。然而,工具可能会默默失败,返回看似正确但实际错误的结果。代理人能够多大程度地检测和纠正受损工具调用输出呢?我们通过一个控制性的损坏框架来研究这一问题,在这个框架中,一个隐藏的拦截器会在有针对性的问题上用看似正确但实际错误的信息替换工具调用结果。我们在四种验证设计下评估了代理人在31个问题上的表现,包括无验证(基准)、强制相同上下文反思、可选新上下文验证以及可选结构验证。没有验证时,损坏会导致准确性急剧下降,从100%降至72.4%。强制反思完全恢复了准确性至100%。可选验证仅在模型主动调用时才提高准确性。我们的结果表明,检查频率与稳健性差异强相关,而不平等调用阻止了对验证器质量的控制比较。一项支持性的恢复实验显示,在明确检测后,完全重新开始问题在所有情况下都能成功。这些发现表明,验证器的可用性和验证政策是数学代理人可靠性的独立组成部分。强制政策强制执行验证,而可选政策取决于模型自己选择是否调用它。

更新时间: 2026-10-06 10:28:19

领域: cs.CR,cs.AI,cs.SE

下载: http://arxiv.org/abs/2610.08097v1

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.

Updated: 2026-10-06 10:28:06

标题: CSWAM:世界行动模型中更好的因果语义表示对超出分布的泛化

摘要: FastWAM风格的世界动作模型可以实现高效的仅动作推理,但在视觉分布转移下泛化能力较差。它们以重建为导向的表示强调特定外观细节,限制了对未知场景和物体的泛化能力。在观察历史的情况下,该模型也缺乏时序证据来稳健地识别在陌生视觉条件下的任务相关状态变化和运动。为了解决这些限制,我们提出了因果语义世界动作模型(CSWAM),它在FastWAM上增加了一个基于V-JEPA 2.1的因果语义专家。V-JEPA提供了与语义状态变化和运动有关的具有时间依赖性的表示,减少了对特定外观细节的依赖。该专家从当前和过去观察的稀疏历史中学习它们的未来演变,并通过因果注意力与视频和动作流共享历史衍生的上下文。在推理时,CSWAM在当前视频状态和观察到的语义历史上进行动作去噪,保留高效的仅动作推理。我们进行了模拟和真实机器人实验,评估了在分布转移下的泛化能力。通过具身预训练,CSWAM将RoboTwin 2.0干净到随机化转移的随机成功率从10.16%提高到45.18%,比FastWAM提高了35.02个百分点。在两个真实机器人任务和三个OOD难度级别上,CSWAM将平均成功率从27.5%提高到70.0%,比FastWAM提高了42.5个百分点。

更新时间: 2026-10-06 10:28:06

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2609.18462v5

Observable Neural ODEs for Identifiable Causal Forecasting in Continuous Time

Causal inference in continuous-time sequential decision problems is challenged by hidden confounding and partially observed states. We show that, under explicit structural assumptions, observability of the latent state enables identification of dynamic treatment effects through a continuous-time conditional front-door adjustment, even in the presence of hidden confounding. We derive a general adjustment formula and show that it reduces to a tractable state-space formula when unobserved contemporaneous disturbances are temporally uncorrelated. This formula expresses potential-outcome distributions under alternative treatment trajectories through the measurement model, latent dynamics, and the filtering distribution over latent states. We propose Observable Neural ODEs (ObsNODEs), Neural ODE models in observable normal form that implement this tractable adjustment for causal forecasting. ObsNODEs learn continuous-time dynamics with states reconstructible from observations, enabling outcome prediction under alternative treatment paths. Experiments on synthetic, semi-synthetic, and real-world clinical data demonstrate strong performance over recent sequence models, including external validation.

Updated: 2026-10-06 10:24:43

标题: 可观测的神经ODEs用于连续时间可识别因果预测

摘要: 在连续时间序贯决策问题中进行因果推断面临着隐藏混杂和部分观察状态的挑战。我们展示,在明确的结构假设下,潜在状态的可观测性使得可以通过连续时间条件前门调整来识别动态治疗效应,即使存在隐藏混杂。 我们推导出一个通用的调整公式,并展示当未观察到的同时扰动在时间上不相关时,它可以简化为一个易处理的状态空间公式。该公式通过测量模型、潜在动态和潜在状态的过滤分布,表达了在替代治疗路径下的潜在结果分布。 我们提出了Observable Neural ODEs(ObsNODEs),Observable正常形式中的神经ODE模型,实现了这种易处理的因果预测调整。ObsNODEs学习可从观测中重构状态的连续时间动态,从而使得在替代治疗路径下进行结果预测成为可能。 在合成、半合成和真实世界临床数据上的实验表明,ObsNODEs相对于最近的序列模型表现出了强大的性能,包括外部验证。

更新时间: 2026-10-06 10:24:43

领域: cs.LG,math.OC,math.ST,q-bio.QM

下载: http://arxiv.org/abs/2604.26070v3

Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing

Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding generation in empirical graph evidence rather than schema expectations. An offline distiller produces TTQL, a compact description of populated multi-hop patterns, conditional frequencies and path-conditioned literal examples, plus a class-property co-occurrence matrix. Online, TTQL guides the LLM, while deterministic syntax, vocabulary and data-model checks provide diagnostics for iterative refinement. On CK25 (First International Text2SPARQL Challenge), under matched-condition recomputation on a QLever snapshot, QRAKEN achieves strict F1 of 0.643 $\pm$ 0.026 with GPT-4.1 mini and 0.652 $\pm$ 0.012 with GPT-5.4: relative gains of 30% and 32% over the strongest recomputed participant, outperforming systems using the same base model family. Ablations identify TTQL patterns as the dominant driver (+0.31 strict F1 over a shape-only baseline); the refinement loop provides a cheap safety net, rejecting triple patterns unsupported by the co-occurrence matrix. Compared with auto-derived SHACL, TTQL yields 64% higher strict F1, supporting the value of empirical patterns beyond schema exposure. With two local 35B 4-bit open-weight models at zero marginal cost, the same pipeline matches the strongest recomputed participant, and TTQL advantages over shape-only and SHACL baselines persist. Results on a single, relatively small benchmark provide an initial empirical signal; monolithic TTQL injection on very open cross-domain graphs remains the main limitation.

Updated: 2026-10-06 10:24:26

标题: 自然语言问题作为知识图谱的接口:QRAKEN图谱精炼和语义自愈

摘要: 自然语言访问RDF知识图谱是语义网络的核心目标。大型语言模型(LLMs)已经推动了文本到SPARQL的发展,然而在陌生的图中,它们经常生成有效的查询,但这些查询却不正确地表达了填充数据模型。QRAKEN是一个无需训练、不受本体限制的神经符号管道,将生成基于经验图证据而不是架构期望。离线提取器生成TTQL,一个填充的多跳模式、条件频率和路径条件字面示例的紧凑描述,以及一个类属性共现矩阵。在线上,TTQL指导LLM,同时确定性语法、词汇和数据模型检查为迭代改进提供诊断。在CK25(第一届国际Text2SPARQL挑战赛)上,在QLever快照上重新计算的匹配条件下,QRAKEN在GPT-4.1 mini和GPT-5.4上的严格F1分别为0.643 $\pm$ 0.026和0.652 $\pm$ 0.012:相对于最强重新计算参与者的相对增益分别为30%和32%,超过了使用相同基础模型系列的系统。消融实验确定TTQL模式是主要的驱动因素(在仅考虑形状的基线上+0.31严格F1);改进循环提供了一个廉价的安全网,拒绝了未被共现矩阵支持的三元组模式。与自动推导的SHACL相比,TTQL的严格F1高出64%,支持实证模式超出架构曝光的价值。使用两个本地35B 4位开放权重模型,零边际成本,相同的管道与最强重新计算参与者匹敌,而TTQL相对于仅考虑形状和SHACL基线的优势仍然存在。在一个相对较小的基准测试上的结果提供了一个初步的实证信号;在非常开放的跨领域图上进行的TTQL注入仍然是主要的限制。

更新时间: 2026-10-06 10:24:26

领域: cs.AI,cs.CL

下载: http://arxiv.org/abs/2610.08095v1

SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis

Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-quality, expert-annotated training data. This challenge is exacerbated by stringent privacy requirements and the impracticality of utilizing large open-source corpora or proprietary cloud APIs within resource-limited clinical settings. To address these obstacles, we introduce SAGE (\textit{Semantic Anchor-Guided Evolution}), a novel data synthesis framework that enables small, locally deployed models to generate high-quality medical training data. SAGE leverages lightweight, publicly available taxonomies such as MeSH as semantic anchors, imposing a structured prior to effectively guide and ground the data generation process. At its core, SAGE iteratively interleaves atomic (individual concept-based) and associative (relation-based) synthesis, bootstrapping training data from minimal seeds. This approach eliminates the need for large collections of medical documents or reliance on external APIs, providing a practical solution for on-premises data creation. Extensive experiments across multiple medical question-answering benchmarks demonstrate that models fine-tuned with SAGE-synthesized data consistently outperform those trained using self-derived or conventional document-based paradigms, highlighting tangible improvements in data efficiency and resource utilization for medical LLM development. Code is available at https://github.com/DIaacKr/SAGE.

Updated: 2026-10-06 10:23:16

标题: SAGE:基于语义锚引导的医学问答数据综合演化

摘要: 发展可靠的临床任务模型,如医学问题回答(QA),受到高质量、专家注释的训练数据有限的严重限制。这一挑战受到严格的隐私要求和在资源有限的临床环境中使用大型开源语料库或专有云API的不切实际性的加剧。为了解决这些障碍,我们引入了SAGE(Semantic Anchor-Guided Evolution),这是一个新颖的数据合成框架,可以让小型、本地部署的模型生成高质量的医学训练数据。SAGE利用轻量级、公开可用的词汇表,如MeSH,作为语义锚点,施加结构化先验来有效地引导和基础数据生成过程。在其核心,SAGE迭代地交错原子(基于个体概念)和联想(基于关系)合成,从最小的种子中引导训练数据。这种方法消除了对大量医学文档的需求或对外部API的依赖,为现场数据创建提供了实际解决方案。在多个医学问答基准测试中进行的广泛实验表明,使用SAGE合成数据微调的模型始终优于使用自派生或传统文档为基础的训练方法训练的模型,突出了医学LLM开发中数据效率和资源利用方面的实质性改进。代码可在https://github.com/DIaacKr/SAGE上找到。

更新时间: 2026-10-06 10:23:16

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.08093v1

Explainable Rule Mining of IPv6 Extension-Header Presence Patterns from Paired-Vantage Captures

IPv6 extension headers (EHs), such as fragmentation, segment routing, and in-situ telemetry, are operationally important yetwidely dropped in transit, and characterising their behaviour from packet captures is a recurring measurement problem. We ask whetheran explainable miner can recover human-readable rules of EH behaviour, and we contribute two reusable tools: a negative-control protocol that diagnoses whether a mined "temporal" network rule reflects genuine cross-packet dynamics or mere within-packetco-occurrence, and a sender-conditioned, per-family EH-retention measurement. Applying an interpretable temporal-logic rule miner to the JAMES paired-vantage dataset, we recover a portable Fragment-EH rule that the protocol reveals to be a within-packet,near-definitional co-occurrence rather than a temporal pattern, so the temporal-logic machinery does no work for this dominant rule;the retention measurement independently recovers the expected within-window ordering of EH observability. Our main result istherefore an honest, controlled negative finding, corroborated by executed decision-tree and large-language-model baselines: on theevaluated JAMES traces network-temporal structure does not carry the dominant Fragment-EH signal, and we supply the controls thatestablish when it would, validated on a synthetic positive control containing a genuine cross-packet dependency.

Updated: 2026-10-06 10:21:12

标题: 可解释的规则挖掘:从配对Vantage捕获中提取IPv6扩展头存在模式

摘要: IPv6扩展标头(EHs),如分段、段路由和原位遥测,在操作上非常重要,但在传输过程中经常被丢弃,从数据包捕获中对它们的行为进行表征是一个反复出现的测量问题。我们询问一个可解释的挖掘工具是否能够恢复可读的EH行为规则,并提供两个可重复使用的工具:一个负控制协议,诊断挖掘的“时间”网络规则是否反映真正的跨数据包动态或仅是数据包内共现,以及一个发送方条件、按家族EH保留测量。将可解释的时间逻辑规则挖掘器应用于JAMES配对观测数据集,我们恢复了一个可携带的Fragment-EH规则,协议显示这是一个数据包内、几乎是定义性的共现,而不是一个时间模式,因此时间逻辑机制对这个主要规则没有作用;保留测量独立地恢复了EH可观测性的预期窗口内顺序。因此,我们的主要结果是一个真实、受控的负面发现,得到执行的决策树和大型语言模型基准的印证:在评估的JAMES跟踪网络时间结构中并没有承载主要的Fragment-EH信号,并提供了建立何时会承载的控制,验证了一个包含真正跨数据包依赖性的合成正面控制。

更新时间: 2026-10-06 10:21:12

领域: cs.CR,cs.LG,cs.NI

下载: http://arxiv.org/abs/2610.08090v1

When Plans Change Answers: Formalizing Cost-Accuracy Optimization for Semantic Queries

In semantic query engines, predicates are evaluated by machine-learned models, and the choice of a query plan affects not only the cost of a query but also its result. Existing systems either apply a fixed threshold to each semantic operator or tune accuracy per operator, without accounting for how errors propagate through joins. We give a formal problem definition for cost-accuracy optimization of such queries. Our starting point is the calibrated confidence that decision models such as Jev attach to each decision. It yields an expected error for every decision; weighting these errors by each decision's contribution to the output (in the simplest case, its fan-out) gives the expected output quality of a plan without any labeled data, and the same computation in reverse turns an output-level accuracy target into a price on each base or intermediate tuple. Building on this, we define an oracle semantics for relational algebra with semantic operators, physical plans as pairs of a logical plan and a decision policy, declarative output-level targets, and a hierarchy of plan equivalence. We show that accuracy is plan-invariant under pointwise-deterministic policies, and that selection pushdown is not quality-sound when escalation bands are calibrated on the plan's own candidates. Expected quality can be computed in polynomial time under bag semantics; under set semantics it follows the dichotomy of tuple-independent probabilistic databases when every relation carries a semantic predicate. Choosing which tuples to drop is NP-hard, while the optimization problem decomposes into per-tuple decisions through two Lagrange multipliers. Simulations on a synthetic workload illustrate these effects; an evaluation on real engines is left for future work.

Updated: 2026-10-06 10:19:43

标题: 当计划改变时的答案:为语义查询正式制定成本-准确性优化

摘要: 在语义查询引擎中,谓词由机器学习模型评估,查询计划的选择不仅影响查询的成本,还影响查询的结果。现有系统要么对每个语义操作符应用固定阈值,要么调整每个操作符的准确度,但没有考虑错误如何通过连接传播。我们为这类查询的成本-准确度优化给出了形式化的问题定义。我们的出发点是决策模型(如Jev)对每个决策赋予的校准置信度。它为每个决策产生了预期误差;通过将这些误差按照每个决策对输出的贡献(在最简单的情况下,即其扩散)进行加权,可以得到一个计划的预期输出质量,而不需要任何标记数据,并且反向计算得到一个输出级别的准确度目标转化为每个基本或中间元组的价格。在此基础上,我们定义了具有语义操作符、物理计划(逻辑计划和决策策略的配对)、声明性输出级别目标和计划等价层次的关系代数的Oracle语义。我们表明在点对点确定性策略下准确度是计划不变的,并且当升级带根据计划自身的候选者进行校准时,选择推送不是质量可靠的。在多重集语义下,预期质量可以在多项式时间内计算;在集合语义下,当每个关系都带有语义谓词时,根据元组独立概率数据库的二分法跟随。选择要丢弃的元组是NP困难的,而优化问题通过两个拉格朗日乘数分解为每个元组的决策。合成工作负载的模拟说明了这些效果;对真实引擎的评估留待未来的工作。

更新时间: 2026-10-06 10:19:43

领域: cs.DB,cs.AI

下载: http://arxiv.org/abs/2610.08089v1

A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies

A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves open. To bring those futures into the decision, we derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completion. The derivation reveals a candidate-dependent feasible-future mass: its support records whether safe completion remains possible under the frozen continuation process, while its magnitude measures how much weighted safe-completion mass remains. Since exact evaluation is impractical online, we develop a selective finite-candidate approximation and establish conditions for recovering the best retained viable candidate. Our alarm-triggered, training-free reranker VICS-G lowers mean cumulative safety cost by 1.9%-57.5% across six Safety-CHORES settings while remaining within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length. Our approach offers a promising and practical path toward safer task completion, grounded in an exact policy-relative target yet requiring neither policy retraining nor online rollouts.

Updated: 2026-10-06 10:17:30

标题: 一个安全的动作并不足够:可行的未来解码用于视觉-语言-动作政策

摘要: 一个安全的行动并不一定是可行的。一个冻结的视觉-语言-行动(VLA)政策可能会倾向于一个在本地可接受的移动,但却没有留下任何支持安全任务完成的政策路径。我们将这称为可行性-可能性差距:可能性排名下一步移动,而可行性取决于它留下的未来。 为了将这些未来引入决策中,我们推导出限制在安全任务完成的政策环境轨迹法的历史条件下的下一个区块边际。推导揭示了一个候选相关的可行未来质量:其支持记录了在冻结的继续过程下安全完成是否仍然可能,而其大小衡量了有多少加权安全完成质量仍然存在。由于在线精确评估是不切实际的,我们开发了一个有选择性的有限候选近似,并建立了恢复最佳保留的可行候选的条件。 我们的警报触发、无需培训的重新排序器VICS-G在六个Safety-CHORES设置中将平均累积安全成本降低了1.9%-57.5%,同时在成功中保持在政策采样2.5个百分点以内,平均剧集长度在0.82步以内。我们的方法提供了一条有前途和实用的路径,以更安全地完成任务,基于精确的政策相关目标,但既不需要政策重新培训,也不需要在线演练。

更新时间: 2026-10-06 10:17:30

领域: cs.AI,cs.RO

下载: http://arxiv.org/abs/2610.05166v2

Sinkhorn doubly stochastic attention rank decay analysis

The self-attention mechanism is central to the success of Transformer architectures. However, standard row-stochastic attention has been shown to suffer from significant signal degradation across layers. In particular, it can induce rank collapse, resulting in increasingly uniform token representations, as well as entropy collapse, characterized by highly concentrated attention distributions. Recent work has highlighted the benefits of doubly stochastic attention as a form of entropy regularization, promoting a more balanced attention distribution and leading to improved empirical performance. In this paper, we study rank collapse across network depth and show that doubly stochastic attention matrices normalized with Sinkhorn algorithm preserve rank more effectively than standard softmax row-stochastic ones. As previously shown for softmax, skip connections are crucial to mitigate rank collapse. We empirically validate this phenomenon on both sentiment analysis and image classification tasks. Moreover, we derive a theoretical bound for the pure self-attention rank decay when using Sinkhorn normalization and find that rank decays to one doubly exponentially with depth, a phenomenon that has already been shown for softmax.

Updated: 2026-10-06 10:16:43

标题: Sinkhorn双随机注意力排名衰减分析

摘要: 自注意机制是Transformer架构成功的核心。然而,标准的行随机注意力机制已经被证明在层间存在显著的信号衰减。特别是,它可能导致秩坍缩,导致越来越均匀的令牌表示,以及熵坍缩,其特征是高度集中的注意力分布。最近的工作强调了双随机注意力作为一种熵正则化的好处,促进更平衡的注意力分布,并提高了实证表现。在本文中,我们研究了网络深度上的秩坍缩,并展示了使用Sinkhorn算法规范化的双随机注意力矩阵比标准的softmax行随机注意力更有效地保持秩。正如以前对softmax所示,跳过连接在减轻秩坍缩方面至关重要。我们在情感分析和图像分类任务上对这一现象进行了实证验证。此外,我们推导出了使用Sinkhorn规范化时纯自我注意力秩衰减的理论上限,并发现秩随深度双指数衰减,这一现象已经被softmax证明过。

更新时间: 2026-10-06 10:16:43

领域: cs.LG,cs.AI,math.OC

下载: http://arxiv.org/abs/2604.07925v2

POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents

LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on $τ^2$-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.

Updated: 2026-10-06 10:14:16

标题: 极地:本体引导的风险预防,用于工具调用LLM代理

摘要: LLM工具使用代理在动态环境中运作,许多操作都存在操作风险。然而,大多数安全机制只能在错误出现后才会起作用。现有的预防性方法要么通过对代理进行链式思考的调整,要么将自然语言防范措施编译成运行时检查,但这些方法都没有暴露出结构化、可审计的结论。我们提出了POLAR,一个用于小型工具调用代理的防护框架,通过结构化的两层本体论来评估可逆性。POLAR通过推导候选逆序列为每个动作分配一个分级的可逆性分数;在执行之前,超过阈值的调用将被修剪。在六个代理模型上评估了$τ^2$-bench结果,POLAR在四个六个代理的航空任务奖励平均提高了0.11到0.18分,但只有十八个模型-领域单元中的八个得到了整体改进;零售和更强大的代理通常会出现退步。POLAR提供了一个可审计的结构化检查,并描述了其任务效用的权衡。奖励并不是防止伤害的直接衡量标准。

更新时间: 2026-10-06 10:14:16

领域: cs.AI,cs.CL

下载: http://arxiv.org/abs/2610.08082v1

CLEAR: Causal Context-Based Agentic Reasoning for Vulnerability Detection

Detecting source code vulnerabilities is increasingly difficult as modern security flaws are rooted in complex causal dependencies between execution flows, control conditions, and program states. Despite recent advances in Large Language Models (LLMs) and multi-agent frameworks, existing approaches primarily address superficial similarities between benign and vulnerable functions while failing to capture the complex causal dependencies inherent in security flaws. To address these limitations, we propose Causal Context-based Agentic Reasoning (CLEAR), a novel multi-agent vulnerability detection framework integrated with a causal knowledge graph. CLEAR systematically constructs a Vulnerability Causal Knowledge Graph (VCKG) that models the causal chains between entrypoints, preconditions, root causes, and fix intents across vulnerability instances. Leveraging this structured knowledge, four specialized agents, including the Collector, Claim, Critic, and Judge, collaboratively verify vulnerability hypotheses through retrieved causal contexts. Experimental results on C/C++ and Java vulnerability benchmarks demonstrate that CLEAR improves Pair-Correct (P-C) performance by 130.7% and 71.56% over state-of-the-art approaches, demonstrating the effectiveness of causal knowledge graph-guided reasoning for automated vulnerability detection.

Updated: 2026-10-06 10:09:34

标题: CLEAR:基于因果上下文的代理推理用于漏洞检测

摘要: 检测源代码漏洞变得越来越困难,因为现代安全漏洞根植于执行流、控制条件和程序状态之间复杂的因果依赖关系。尽管最近在大型语言模型(LLMs)和多智能体框架方面取得了进展,现有方法主要解决了良性和易受攻击函数之间的表面相似性,但未能捕捉安全漏洞中的复杂因果依赖关系。为了解决这些限制,我们提出了基于因果上下文的智能体推理(CLEAR),这是一个集成了因果知识图的新颖多智能体漏洞检测框架。CLEAR系统地构建了一个漏洞因果知识图(VCKG),模拟了漏洞实例之间的入口点、前提条件、根本原因和修复意图之间的因果链。利用这种结构化知识,包括收集器、声明、批评者和评判者在内的四个专门智能体,通过检索到的因果上下文协同验证漏洞假设。对C/C++和Java漏洞基准的实验结果表明,CLEAR相对于最先进的方法提高了130.7%和71.56%的成对正确(P-C)性能,展示了基于因果知识图引导推理的自动漏洞检测的有效性。

更新时间: 2026-10-06 10:09:34

领域: cs.CR,cs.SE

下载: http://arxiv.org/abs/2608.03134v2

ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding

Standard causal identification methods often assume no unmeasured confounding and can fail when relevant confounders are unobserved. Proximal causal inference instead uses proxy variables to identify effects under hidden confounding. However, nonparametric proximal estimation can be challenging in practice: recovering causal estimands such as the conditional average treatment effect (CATE) requires solving an ill-posed integral equation that is data-hungry, hyperparameter-sensitive, and optimization-unstable. Bayesian inference for such models provides a desirable alternative, mitigating these difficulties by regularizing through the prior. However, computing a posterior is itself challenging, as a typical likelihood function will include latent variables. Following the recent success of tabular foundation models in backdoor, instrumental variable, and frontdoor settings, we propose that prior-data fitted networks (PFNs) are uniquely suited to resolve this bottleneck. Indeed, by training on synthetic data sampled from compliant structural causal models with access to oracle counterfactuals, we simplify the task substantially, amortizing the implied Bayesian operator inversion into a single transformer forward pass. Compared to prior literature that focuses primarily on point estimation, our model, ProximalFM, explicitly targets the Bayesian posterior distribution of the CATE. One unique aspect of this problem is that we need to provide Monte Carlo estimates of the oracle CATEs, leading to a novel variation of PFNs that accounts for the added stochastic error. Across a diverse suite of proximal regimes, ProximalFM achieves consistently strong CATE-estimation performance without dataset-specific tuning, with its largest advantage when latent confounding is substantial and the proxies are weakly informative; it also provides fast inference through a single amortized forward pass.

Updated: 2026-10-06 10:07:19

标题: ProximalFM:隐含混淆下的摊销近端因果推断

摘要: 标准因果识别方法通常假设没有未测到的混杂因素,在未观察到相关混杂因素时可能失败。相较之下,近因果推断使用代理变量来识别隐藏混杂因素下的效应。然而,在实践中,非参数近因估计可能具有挑战性:恢复因果估计量,如条件平均处理效应(CATE),需要解决一个数据密集、超参数敏感和优化不稳定的不适定积分方程。贝叶斯推断为这类模型提供了一种理想的替代方案,通过先验进行正则化来缓解这些困难。然而,计算后验本身也具有挑战性,因为典型的似然函数将包括潜变量。鉴于表格基础模型在背门、工具变量和前门设置中的最近成功,我们提出,先验数据拟合网络(PFNs)非常适合解决这个瓶颈。事实上,通过在从符合结构因果模型中抽样获得的合成数据上进行训练,并且可以访问神谕反事实,我们大大简化了任务,将暗示的贝叶斯算子反转为一个单一的变压器前向传递。与主要侧重于点估计的先前文献相比,我们的模型ProximalFM明确地针对CATE的贝叶斯后验分布。这个问题的一个独特方面是我们需要提供神谕CATE的蒙特卡洛估计,这导致了PFNs的一种新变体,考虑了额外的随机误差。在各种近因制度中,ProximalFM实现了一致强大的CATE估计性能,无需特定于数据集的调整,当潜在混杂因素显著且代理信息不足时,其优势最大;它还通过单一的摊销前向传递提供快速推断。

更新时间: 2026-10-06 10:07:19

领域: stat.ML,cs.LG

下载: http://arxiv.org/abs/2610.08078v1

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to $24.2$ pp. Its advantage is especially pronounced when reward contrast is scarce: when $37$--$98\%$ of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where $98\%$ of groups are all-failure, the RLVR training ends up at $0.0\%$ success, while adding SRD reaches $60.6\%$ under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.

Updated: 2026-10-06 10:06:47

标题: 自我反思蒸馏:将事后经验转化为先见之明

摘要: 具有可验证奖励的强化学习(RLVR)通过与交互后的标量结果奖励将代理经验转化为学习信号。然而,对于群体相关目标,当所有展开均获得相同奖励时,该信号会消失,即使它们的轨迹可能揭示任务需求和代理失败的有用信息。我们提出一个补充问题:事后能否教会代理在采取行动之前可以预见到什么?我们引入了前瞻学习,它利用事后经验监督预先交互视图中的预测,并通过自省蒸馏(SRD)实现。直观地说,完成的轨迹揭示了在交互之前会有用的知识和应避免的陷阱;SRD将这种特权事后反思转化为相同策略的轨迹盲目前瞻。前瞻仅用作训练目标,不需要在推理时明确生成。在10个工具集成的推理和长期目标任务中,SRD通过增益最高可达24.2个百分点,与RLVR和自蒸馏基线相辅相成。当奖励对比稀缺时,其优势尤为显著:当37%至98%的展开组在模型尺度上都具有奖励一致性时,SRD仍然可以利用来自采样轨迹的学习信号。在98%的展开组均为失败的2B设置中,RLVR训练最终成功率为0.0%,而添加SRD在相同展开预算下达到了60.6%。我们的结果表明,事后代理经验不仅有助于评估或改进行为,还有助于在可用交互之前塑造预测性表示。

更新时间: 2026-10-06 10:06:47

领域: cs.AI,cs.CL,cs.IR,cs.LG

下载: http://arxiv.org/abs/2610.08077v1

SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning

Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved. The ability to develop sophisticated strategies to tackle consequential problems becomes paramount as well-trodden, human-developed solutions become insufficient for problems for which we lack context or enough training data. We study agents' capability of such strategy formation through the communal practice of video game speedrunning. In speedrunning, practitioners compete to find the fastest way to complete a video game under certain conditions, and in so doing uncovering interesting unorthodox play styles that require a thorough understanding and mastery of the underlying game mechanics. We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games. To perform well in this benchmark, agents must repeatedly improve their strategy, reflect on their performance, exploit their gained knowledge, and reason across a long-horizon of actions to improve on an increasingly difficult problem: being faster than themselves and everyone else. Our experiments show that while frontier agents approach human world records in simple platformer games, they remain behind human performance on longer, more complex games under practical budgets. These results suggest that SPEEDRUNBENCH is a useful testbed for studying agents' strategy formation capabilities as well as being a saturation-resistant evaluation measure, as there is almost always a faster completion time waiting to be discovered.

Updated: 2026-10-06 10:06:25

标题: SpeedrunBench:用视频游戏速通挑战LLM代理

摘要: 前沿LLM代理已被证明能够解决日益复杂的任务,这些任务在人类已经有可测量解决方案。这引发了一个重要问题,即LLM代理是否能够超越人类已经解决的问题。随着人类开发的解决方案在我们缺乏背景或足够训练数据的问题上变得不够,开发复杂策略来解决重要问题的能力变得至关重要。我们通过共同实践视频游戏速通来研究代理的这种策略形成能力。在速通中,从业者竞争找到在特定条件下完成视频游戏的最快方式,从而揭示需要对基础游戏机制进行深入理解和掌握的有趣的非正统玩法。我们引入了SPEEDRUNBENCH,这是一个评估前沿LLM代理在9种不同游戏中表现的基准。要在这个基准测试中表现良好,代理必须不断改进他们的策略,反思他们的表现,利用他们获得的知识,并在长期行动的过程中推理,以改进一个越来越困难的问题:比自己和其他人更快。我们的实验表明,尽管前沿代理在简单的平台游戏中接近人类世界纪录,但在更长、更复杂的游戏中,他们在实际预算下仍然落后于人类表现。这些结果表明,SPEEDRUNBENCH是一个用于研究代理策略形成能力的有用测试平台,同时也是一个抗饱和的评估指标,因为几乎总有一个更快的完成时间等待被发现。

更新时间: 2026-10-06 10:06:25

领域: cs.AI

下载: http://arxiv.org/abs/2610.08076v1

Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields

Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred from observed data. In meta-learning, this encoding occurs through gradient updates induced by the decoder, tying representation learning directly to decoder design. We formalize this connection by interpreting latent optimization as an optimization encoder, unifying the roles of second-order differentiation, latent parameterization, and task supervision. This concept enables second-order meta-learning for end-to-end training of the encoding procedure alongside the decoder, and clarifies which learning pathway first-order approximations discard. Guided by this view, we introduce Attentive Latent Fields (MetaLF), an equivariant transformer-based neural field that contextualizes a latent pointcloud through self-attention. These interactions shape both field predictions and the updates that construct their representation, allowing local observations to inform coherent non-local structure. Disentangling the inner encoding objective from outer task supervision unifies reconstruction, classification, and segmentation within an end-to-end meta-learning framework, using reconstruction-only latent adaptation at test time. Controlled experiments on polynomial fields link latent coordination to lower effective rank and stronger alignment with the underlying function space. Across image and 3D shape reconstruction, MetaLF improves fidelity within three to five gradient updates, while supporting semantic prediction across images, shapes, and volumes. Together, these findings position the optimization encoder perspective as a unified basis for designing neural fields around how representations are constructed, coordinated, and used.

Updated: 2026-10-06 10:05:56

标题: 优化编码器:重新思考神经场的二阶元学习

摘要: 条件神经场连续表示信号,但其有效性取决于如何从观测数据中推断出条件潜在表示。在元学习中,这种编码是通过解码器引起的梯度更新来进行的,直接将表示学习与解码器设计联系起来。我们通过将潜在优化解释为优化编码器来形式化这种连接,统一了二阶微分、潜在参数化和任务监督的作用。这个概念实现了端到端的编码过程和解码器的二阶元学习,并澄清了一阶近似丢弃哪些学习路径。在这个观点的指导下,我们引入了Attentive Latent Fields(MetaLF),这是一种基于等变变压器的神经场,通过自注意力对潜在点云进行上下文化。这些交互形状了场预测和构建其表示的更新,允许局部观测通知连贯的非本地结构。将内部编码目标与外部任务监督分离在端到端元学习框架中统一了重建、分类和分割,测试时仅使用重建潜在适应。在多项式场上进行的受控实验将潜在协调与较低有效秩和与底层函数空间更强的对齐联系起来。在图像和3D形状重建方面,MetaLF在三到五个梯度更新内提高了保真度,同时支持图像、形状和体积跨语义预测。这些发现将优化编码器视角定位为设计神经场的统一基础,围绕表示如何构建、协调和使用。

更新时间: 2026-10-06 10:05:56

领域: cs.LG,cs.CV

下载: http://arxiv.org/abs/2610.08075v1

Benchmarking the Personalization Capabilities of Large Language Models

Personalization is classically a two-party problem: a sender chooses what to say, and a receiver with independent objectives decides whether to act. A salesperson pitching the same analytics product leads with HIPAA compliance for a hospital and real-time reporting for a retailer, expecting a different argument to work on each. Existing LLM personalization benchmarks measure a narrower, one-party property: whether output matches the preferences of the same user it serves-sender and receiver being the same, as when RLHF aligns an assistant to its own user. The two-party case is harder to study automatically, since it needs ground truth linking specific content to an observed receiver action. Sales outreach provides this: a message written for one prospect, recorded against whether it produced a reply, a call, or a closed deal. We introduce SDR-Arena, a framework for benchmarking two-party generative personalization at scale, and SDR-Bench, a public corpus of 50,000 customer success stories across 22 industries and 3500 enterprises. Given only pre-outcome information, an agent must reconstruct the arguments that won the deal, scored by a weighted nugget-recall metric (WCS). The best model, Claude Sonnet 4.6, reaches 55.8% WCS indicating it recovers only half the winning content-a plateau we observe across model families that costly deep-research pipelines do not close. An ablation shows the cause is retrieval, not reasoning: models improve substantially given the facts a human researcher would gather, but rarely find them through web search alone. Two studies with professional SDRs support the metric: only 48% of generated pitches were rated usable without editing, and WCS yields model rankings consistent with evaluation against expert strategies authored independently of any model output.

Updated: 2026-10-06 10:05:25

标题: 大型语言模型个性化能力的基准测试

摘要: 个性化经典上是一个双方问题:发送者选择要说什么,而具有独立目标的接收者决定是否采取行动。销售人员向医院提供HIPAA合规性和向零售商提供实时报告的相同分析产品,期望不同的论点能够奏效。现有的LLM个性化基准衡量了一个更窄的、单方面的属性:输出是否符合相同用户的偏好-发送者和接收者是相同的,就像RLHF将助手与自己的用户对齐一样。两方面的情况更难以自动研究,因为它需要将特定内容与观察到的接收者行动联系在一起的真实基础。 销售推广提供了这个:为一个潜在客户编写的消息,记录了它是否产生了回复、电话或成交。我们介绍了SDR-Arena,一个用于在规模上进行双方生成式个性化基准测试的框架,以及SDR-Bench,一个包含22个行业和3500家企业的50,000个客户成功案例的公共语料库。只给出预期结果信息,一个代理必须重建赢得交易的论点,通过加权的金块召回指标(WCS)进行评分。最好的模型,Claude Sonnet 4.6,达到了55.8%的WCS,表明它仅恢复了一半的获胜内容-这是我们观察到的在昂贵的深度研究管道封闭之后的模型系列中的一个平台。消融表明原因是检索,而不是推理:模型在获得人类研究人员会收集的事实后有了相当大的改进,但很少通过网络搜索独立发现它们。两项与专业SDR的研究支持了度量标准:只有48%的生成论点被评为可用而无需编辑,而WCS产生的模型排名与独立于任何模型输出的专家策略评估一致。

更新时间: 2026-10-06 10:05:25

领域: cs.AI

下载: http://arxiv.org/abs/2607.20471v2

Hypergraph-Enhanced Dual Convolutional Network for Bundle Recommendation

Bundle recommendation ranks sets of related items rather than isolated items. Its central challenge is to connect user preferences, item interactions, and bundle composition without losing the signals needed to rank bundles. We propose Hypergraph-Enhanced Dual Convolutional Neural Network (HED), which constructs a complete hypergraph containing user--bundle, user--item, and bundle--item interactions together with intra-user and intra-bundle relations. HED couples complete-hypergraph propagation with a user--bundle branch, allowing item-aware higher-order context to inform ranking while preserving recommendation-specific signals. On NetEase, HED-128 improves over the strongest baseline by 5.04--6.97% across the six reported metrics; on Youshu, HED-64 improves by 1.87--4.56%. Ablation results support the contributions of both the user--bundle branch and intra-type relations, and sensitivity analyses identify stable operating ranges for the main hyperparameters. We further quantify the computational trade-off of the complete hypergraph, including its memory cost. The evidence supports HED on the two evaluated bundle-recommendation datasets while making its resource limitations explicit. Code and datasets will be made available upon publication.

Updated: 2026-10-06 10:03:50

标题: 超图增强的双卷积网络用于捆绑推荐

摘要: 捆绑推荐对相关项目集进行排名,而不是孤立的项目。其主要挑战是在不丢失排名所需信号的情况下,连接用户偏好、项目交互和捆绑组成。我们提出了一种名为Hypergraph-Enhanced Dual Convolutional Neural Network(HED)的方法,它构建了一个包含用户-捆绑、用户-项目和捆绑-项目相互作用以及用户内部和捆绑内部关系的完整超图。HED将完整超图传播与用户-捆绑分支相结合,允许具有项目感知的高阶上下文来影响排名,同时保留推荐特定信号。在网易上,HED-128在六个报告指标中比最强基线提高了5.04-6.97%;在优术上,HED-64提高了1.87-4.56%。消融结果支持用户-捆绑分支和内部类型关系的贡献,敏感性分析确定了主要超参数的稳定操作范围。我们进一步量化了完整超图的计算权衡,包括其内存成本。证据支持HED在两个评估的捆绑推荐数据集上的应用,同时明确了其资源限制。代码和数据集将在发表后提供。

更新时间: 2026-10-06 10:03:50

领域: cs.IR,cs.AI

下载: http://arxiv.org/abs/2312.11018v3

Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair

A mean shift between two data sources can be easy to detect but hard to remove without substantially changing their representations. We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data. We derive the exact finite-sample minimax risk over all such maps, $(d-k) \mathbb{E}[1/(d+2J)]$ with $J\sim\mathrm{Pois}(κ/2)$, where the budget allows deleting $k$ directions and $κ$ is the calibration signal-to-noise ratio. Projecting out the mean calibration difference attains it without knowing $κ$ or the noise scale. This exposes a detection-repair gap: detecting the shift needs only $κ\gg\sqrt d$, whereas removing a fixed fraction of it at constant distortion needs $κ\asymp d$, as for estimating its direction. Standard linear concept erasers (MP, SAL, LEACE) remove the same calibration difference, so the formula gives, before fitting, exactly how much shift they leave on fresh data and how much calibration a target requires. The limit is robust: pairing keeps it exact for non-Gaussian shared content, the projection keeps its guarantee under anisotropic noise, and selective abstention cannot close the gap. On paired clinical and wearable sleep EEG, where differences between participants act as calibration noise, the formula predicts the device shift left in new participants, and more recordings per person soon stop helping. Together, these results tell whether a correction that falls short needs a better method, more recordings, or more participants.

Updated: 2026-10-06 10:03:20

标题: 检测到偏移并不足够:线性表示修复的精确极小极限

摘要: 两个数据源之间的均值漂移可能很容易检测,但要在不大幅度改变它们的表示的情况下去除却很困难。我们将其去除视为一个统计决策问题:从$\mathbb{R}^d$中成对校准测量之间的噪声差异中学习一个线性映射,应用于在硬失真预算下的两个数据源,使得在新数据上尽可能少地保留漂移。我们推导出所有这样的映射上的精确有限样本极小风险,为$(d-k) \mathbb{E}[1/(d+2J)]$,其中预算允许删除$k$个方向,$κ$为校准信噪比,$J\sim\mathrm{Pois}(κ/2)$。通过投影出均值校准差异,可以实现它,而无需知道$κ$或噪声尺度。这揭示了一个检测-修复差距:检测漂移只需要$κ\gg\sqrt d$,而在恒定失真下删除其固定比例需要$κ\asymp d$,就像估计其方向一样。标准线性概念擦除器(MP、SAL、LEACE)移除相同的校准差异,因此该公式在拟合之前精确地给出了它们在新数据上留下多少漂移以及目标需要多少校准。极限是稳健的:配对保持着对于非高斯共享内容的精确性,投影在各向异性噪声下保持其保证,而选择性弃权无法消除差距。在配对的临床和可穿戴睡眠脑电图中,参与者之间的差异充当校准噪声,该公式预测了新参与者中设备漂移的情况,并且很快每人的更多记录将不再有所帮助。综合这些结果可以判断一个修正是否需要更好的方法、更多的记录或更多的参与者。

更新时间: 2026-10-06 10:03:20

领域: cs.LG,math.ST,stat.ML

下载: http://arxiv.org/abs/2610.08069v1

FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models

Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especially in healthcare, where clinically consequential decisions require not only accuracy but also interpretable, auditable rationales to meet safety, accountability, and regulatory requirements. Conventional federated fine-tuning largely imitates final answers rather than cultivating step-by-step reasoning, often relying on privacy-sensitive centralized distillation and still incurring substantial communication overhead. We address this gap with \textbf{\ours{}}, a federated reasoning framework that combines lightweight chain-of-thought resampling with a compact discriminator for selection, and client-aware LoRA stacking with weighted classifier aggregation to accommodate heterogeneity while reducing aggregation noise and communication; clients generate candidate chains and supervision locally, and only lightweight modules are aggregated on the server. Experiments on medical reasoning benchmarks show consistent gains under tight resource budgets while keeping data local and respecting privacy, offering an interpretable and resource-efficient solution. Our code is made publicly available at https://github.com/DIaacKr/FedCoT

Updated: 2026-10-06 10:02:44

标题: FedCoT:用于大型语言模型的通信高效的联邦推理增强

摘要: 在联邦设置中增强LLM推理是非常困难的,这是由于严格的计算、通信和隐私约束,特别是在医疗保健领域,临床重要决策不仅需要准确性,还需要可解释、可审计的理由,以满足安全性、责任性和监管要求。传统的联邦微调往往只模仿最终答案,而不是培养逐步推理,通常依赖于隐私敏感的集中蒸馏,仍然产生大量的通信开销。我们通过\ours{}来填补这一空白,这是一个联邦推理框架,结合轻量级的思维链重采样和紧凑的鉴别器进行选择,以及客户端感知的LoRA堆叠和加权分类器聚合,以适应异质性,同时减少聚合噪声和通信;客户端在本地生成候选链和监督,只有轻量级模块在服务器上聚合。在医疗推理基准测试中的实验表明,在紧张的资源预算下保持数据本地化并尊重隐私,提供了一个可解释和资源高效的解决方案。我们的代码已公开发布在 https://github.com/DIaacKr/FedCoT。

更新时间: 2026-10-06 10:02:44

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2508.10020v2

Systematically Optimized CNN-Transformer with Focal Loss for Imbalanced Intrusion Detection on NSL-KDD

Intrusion Detection Systems (IDS) struggle with imbalanced datasets like NSL-KDD, especially in detecting rare R2L and U2R attacks. This work describes a systematically optimized and explainable framework using a CNN-Transformer architecture to improve performance on highly imbalanced data. We decided to use XGBoost for feature selection and Focal Loss as the main mechanism for minority class learning. We used Optuna for end-to-end hyrate=0.00042, batch size=256, and Focal Loss γ), using a thoughtful data splitting process to prevent data leakage. Critically, our ablation studies pointed out that while Focal Loss (optimally at γ = 1.5) substantially enhanced minority recall, adding oversampling methods like SMOTE caused precision degradation via over-correction. Our final model, using only tuned Focal Loss, achieved 98.73% overall accuracy on the NSL-KDD test data, with significantly balanced F1-scores of 84.63% (R2L) and 69.66% (U2R) over baseline approaches. Additionally, we performed SHAP analysis to understand model predictions, understand salient features (e.g., service http, logged in), and explain the persistent U2R precision challenge due to feature overlap. This study presents a comprehensive pipeline for imbalanced IDS, providing methodological insights, and demonstrates that tuned Focal Loss can be a sufficient and effective strategy for class balancing.

Updated: 2026-10-06 10:00:06

标题: 经过系统优化的CNN-Transformer结合Focal Loss用于NSL-KDD上的不平衡入侵检测

摘要: 入侵检测系统(IDS)在处理像NSL-KDD这样的不平衡数据集时面临着困难,特别是在检测罕见的R2L和U2R攻击时。本文描述了一个系统优化和可解释的框架,使用CNN-Transformer架构来提高在高度不平衡数据上的性能。我们决定使用XGBoost进行特征选择,以及使用Focal Loss作为少数类学习的主要机制。我们使用Optuna进行端到端的超参数调优,批量大小为256,并使用Focal Loss γ,采用周到的数据拆分过程以防止数据泄漏。关键的是,我们的消融研究指出,虽然Focal Loss(在γ = 1.5时最佳)显著增强了少数类的召回率,但添加像SMOTE这样的过采样方法会通过过度校正导致精度下降。我们的最终模型,仅使用调整后的Focal Loss,在NSL-KDD测试数据上实现了98.73%的整体准确率,并且在基准方法上实现了84.63%(R2L)和69.66%(U2R)的显著平衡的F1分数。此外,我们进行了SHAP分析,以了解模型预测,理解显著特征(例如http服务,登录),并解释由于特征重叠而导致的持久的U2R精度挑战。本研究提供了一个全面的不平衡IDS流程,提供了方法论见解,并表明调整后的Focal Loss可以是一个足够有效的类平衡策略。

更新时间: 2026-10-06 10:00:06

领域: cs.CR

下载: http://arxiv.org/abs/2610.08066v1

Sequential Capacity of Quantum Processes with Finite Memory

How complex can the responses of a quantum device become as it runs longer with a fixed internal memory? We quantify this complexity through sequential response capacity: how many adaptive testing stages, each using a fresh run, can continue to separate possible processes by a prescribed gap in response probabilities. For fixed system and memory sizes, we establish a tight law relating this capacity to run length and probability resolution. At fixed resolution, the capacity grows on the order of $K\log K$, where $K$ is the number of time steps in each run. Our construction attains this growth using time-dependent phase rotations on a single visible qubit with no additional internal memory; its tests give response probabilities exactly zero or one. Under the same tests, classical stochastic processes that measure in a fixed basis at every step have only linear capacity at fixed sizes and resolution. For phase sequences selected by a stored classical label, we then quantify how known independent Pauli noise changes this logarithmic enhancement. With ideal controls and weak residual phase noise after correction, we prove matching capacity bounds at a fixed small probability gap. These bounds identify the inverse residual phase-flip probability as the coherence timescale that limits the extra logarithmic growth.

Updated: 2026-10-06 09:59:25

标题: 量子过程具有有限记忆的顺序容量

摘要: 随着量子设备在固定内存中运行时间越长,其响应可能变得多么复杂?我们通过顺序响应容量来量化这种复杂性:在每个使用新运行的自适应测试阶段中,可以继续通过一定间隔的响应概率来区分可能的过程。对于固定的系统和内存大小,我们建立了一个严格的定律,将这种容量与运行长度和概率分辨率相关联。在固定分辨率下,容量按照$K\log K$的顺序增长,其中$K$是每个运行中的时间步数。我们的构造利用对一个可见量子比特进行时间相关的相位旋转来实现这种增长,而不需要额外的内存;其测试结果的响应概率完全为零或一。在相同的测试条件下,经典随机过程在固定大小和分辨率下只有线性容量。对于由存储的经典标签选择的相位序列,我们进一步量化了已知独立的Pauli噪声如何改变这种对数增强。通过理想的控制和纠正后的弱残余相位噪声,在固定的小概率间隔下,我们证明了匹配的容量界限。这些界限确定了逆残余相位翻转概率作为限制额外对数增长的相干时间尺度。

更新时间: 2026-10-06 09:59:25

领域: quant-ph,cs.LG

下载: http://arxiv.org/abs/2610.02068v2

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting meta-agent reads workers' outputs, writes the final answer, allocates later calls, and decides when to stop. It is expressive, but it also concentrates three control decisions in an opaque, order-sensitive model call. We ask whether the manager needs to be generative at all. UnitBoost replaces that model with a defined meta-level operator: a task-given unit map turns worker outputs into slot-value proposals, a constrained argmax assembles the output, and the slots left unfilled or unsupported become an explicit residual for the next round. The operator is order-free, records unit provenance, and gives a simple guarantee: without coupling constraints, unit-wise maximization under the same admission score dominates selection of any complete candidate. On three held-out benchmarks, it exceeds the best single candidate chosen with gold labels by 0.060-0.195 absolute task-score points and input-matched generative managers by 0.048-0.076. Replacing only the management step improves six compound-system configurations by 0.013-0.182. Residual-directed rounds raise FanOutQA cell F1 from 0.4778 to 0.5524; matched controls show that the true residual outperforms random targets and ordinary rereading, while a label-free supply signal flags exhaustion after one unproductive round. The same analysis measures three conditions in which no such gain is available (one indivisible unit, unavailable unit identity, and an endpoint that charges for every emitted unit) and quantifies cross-unit coupling as a repair cost. The manager gives up semantic freedom and gains order invariance, unit provenance, and testable failure conditions.

Updated: 2026-10-06 09:57:55

标题: UnitBoost:使用合并运算符而不是模型管理复合LLM系统

摘要: 合成LLM系统通常通过添加更高级别的LLM来解决协调问题。结果的元代理读取工作人员的输出,编写最终答案,分配后续调用,并决定何时停止。它表达了但也集中了三个控制决策在一个不透明的、顺序敏感的模型调用中。我们问经理是否需要生成。UnitBoost用一个定义的元级操作符替换了该模型:一个任务给定的单位映射将工作人员的输出转换为槽值提议,一个受限制的argmax组装输出,未填充或不支持的槽成为下一轮的显式剩余。该操作符是无序的,记录单位来源,并提供了一个简单的保证:在没有耦合约束的情况下,具有相同准入分数的单元最大化优于选择任何完整候选。在三个保留的基准测试中,它超过了通过金标签选择的最佳单个候选者0.060-0.195绝对任务分数点,并且比输入匹配的生成管理者高0.048-0.076。仅替换管理步骤提高了六个复合系统配置0.013-0.182。剩余导向的循环将FanOutQA单元F1从0.4778提高到0.5524;匹配的控制显示真实的剩余优于随机目标和普通重新阅读,而无标签的供应信号在一轮无效后标志着枯竭。相同的分析测量了三种条件,在这些条件下没有这样的收益可用(一个不可分割的单位、不可用的单位标识和每发出一个单位收费的终点),并量化跨单元耦合作为修复成本。经理放弃了语义自由,获得了顺序不变性、单位来源和可测试的失败条件。

更新时间: 2026-10-06 09:57:55

领域: cs.AI,cs.CL,cs.MA

下载: http://arxiv.org/abs/2609.09815v2

ASCENT: First-Order Optimal Fine-Tuning with Recalibration for Safety--Utility Co-Enhancement

Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and utility enhancement, lack a theoretical characterization of the optimal safety-related subspace and safety-preserving task update, and typically rely on a static safety subspace that may become outdated during fine-tuning. To address these limitations, we propose ASCENT, a downstream fine-tuning framework for safety--utility co-enhancement through first-order optimal safety-aware periodic calibration and task optimization. We model safety as a function of LLM parameters $S(θ)$ and use its first-order approximation to characterize safety changes under parameter updates. Under a fixed rank and Frobenius-norm budget, we prove that the update constructed from the top-$r$ singular components of the safety-function gradient maximizes the estimated safety change, and use it for periodic calibration to preserve and improve safety. We further derive a unique safety-preserving task update that stays close to the original task update while penalizing negative effects on the estimated safety change. ASCENT alternates these optimal task and calibration updates to jointly enhance safety and utility. Experiments across multiple LLM families and downstream tasks show that ASCENT improves downstream utility by up to 20.3\% and reduces attack success rate by up to 35.5\%, achieving state-of-the-art safety and utility across all evaluated settings. Our code is available at https://github.com/ZJU-LLM-Safety/ASCENT.

Updated: 2026-10-06 09:56:55

标题: ASCENT:安全-效用共同增强的首选一阶优化微调与重新校准

摘要: 监督微调可以显著提高大型语言模型(LLMs)的下游效用,但可能会影响它们的安全性。现有的安全保护方法通过使用与安全相关的参数或子空间来限制下游更新,但主要关注的是安全性的保留而不是联合安全性和效用增强,缺乏对最优安全相关子空间和安全保护任务更新的理论特征,并且通常依赖于可能在微调过程中过时的静态安全子空间。为了解决这些限制,我们提出了ASCENT,这是一个通过一阶最优安全感知周期校准和任务优化来实现安全-效用共同增强的下游微调框架。我们将安全性建模为LLM参数$S(θ)$的函数,并使用其一阶近似来描述参数更新下的安全性变化。在固定秩和弗罗贝尼乌斯范数预算的情况下,我们证明了从安全函数梯度的前r个奇异分量构建的更新最大化了估计的安全性变化,并将其用于周期性校准以保持和改进安全性。我们进一步推导出一个独特的安全保护任务更新,该更新保持接近原始任务更新,同时惩罚对估计安全性变化的负面影响。ASCENT交替进行这些最优任务和校准更新,以共同增强安全性和效用。跨多个LLM系列和下游任务的实验显示,ASCENT可以将下游效用提高多达20.3%,并将攻击成功率降低多达35.5%,在所有评估设置中实现了最先进的安全性和效用。我们的代码可在https://github.com/ZJU-LLM-Safety/ASCENT 上找到。

更新时间: 2026-10-06 09:56:55

领域: cs.CR

下载: http://arxiv.org/abs/2610.08061v1

Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training

Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; $n=195$ expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman $ρ= 0.54$ against $\le 0.48$ within the target domain, a paired session-level gap of $+0.15$ that holds at $+0.12$ when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts' rating instruments: the instrument-derived battery lifts a single judge from $0.32$ to $0.41$ over generic dialogue qualities, judges from three model families ensemble to $0.51$ language-only, and a nonverbal-dyadic block adds $+0.03$ more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth $+0.07$ there. At practically attainable corpus sizes, the expert's overall impression is carried by what is said, and by other communication programs' data more than by one's own.

Updated: 2026-10-06 09:53:25

标题: 语言承载专家的印象:基于仪器的LLM法官转移辅导质量评估和领域内培训

摘要: 自动评估双向咨询对话中的沟通质量受到数据的瓶颈制约:专家评定的语料库规模较小且成长成本高。我们研究了专家整体印象预测在三个德国模拟咨询对话语料库之间的跨领域转移(两个普通医学实践语料库,一个与学校相关的家长-教师语料库;$n=195$个专家评定的对话,一个语料库在规模等化后)。在其他领域的训练胜过领域内的训练:留一领域外转移在目标领域内的Spearman相关系数达到了$ρ=0.54$,而在目标领域内则低于$0.48$,在训练集大小匹配时,成对会话级别的差距为$+0.15$,当训练集大小匹配时,该差距保持在$+0.12$,因此不仅仅是数据量的问题。关键特征是来自小型开放权重LLM读取两个发言者对话记录的会话级别构建分数,这些构建主要来自专家评定工具:工具衍生的电池将单个评委的得分从$0.32$提升到$0.41$,三个模型家族的评委集成到$0.51$的语言质量,一个非言语双向块添加了$+0.03$,在这个样本大小下无法与噪音分离。我们还定价了录音设置:一个语料库丢失了每个发言者的音频,其16%的日记化分段携带了错误的发言者,并且修复值为$+0.07$。在实际可达到的语料库规模下,专家的整体印象主要受到言语内容以及其他沟通程序的数据的影响,而不仅仅是个人的影响。

更新时间: 2026-10-06 09:53:25

领域: cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.08055v1

D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation

Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic sets while preserving training efficacy. However, existing studies mainly focus on image classification, leaving dense prediction tasks such as semantic segmentation largely underexplored. In this work, we identify three key challenges for segmentation DD: (i) long-tailed class imbalance, (ii) the need for strict pixel-wise alignment between images and dense labels, and (iii) the high computational cost of optimizing high-resolution data with complex models. To address these challenges, we propose D3S2, a Diffusion-guided Dataset Distillation framework for Semantic Segmentation. Our method adopts a two-stage design. In Class-Balanced Mask Selection, we construct a representative mask set via a greedy strategy that prioritizes underrepresented classes. In Diffusion-Guided Image Synthesis, we employ a pretrained layout-to-image diffusion model to generate images conditioned on the selected masks, naturally ensuring spatial alignment. To further enhance the training utility of synthesized data, we introduce guided diffusion sampling with two complementary objectives: a segmentation-consistency loss for pixel-level alignment, and a class-wise feature matching loss for aligning per-class feature statistics across layers. Extensive experiments demonstrate the superiority of D3S2. Notably, at an extremely compression rate of 1%, our method achieves 24.99% and 35.49% mIoU on ADE20K and COCO-Stuff with Mask2Former (Swin-S), outperforming random selection by 9.34% and 5.70%, respectively. Our code is available at https://github.com/zwj084/D3S2.

Updated: 2026-10-06 09:48:58

标题: D3S2:用于语义分割的扩散引导数据集精炼

摘要: 数据集精炼(DD)旨在将大规模数据集压缩为紧凑的合成集,同时保持训练效果。然而,现有研究主要集中在图像分类上,对语义分割等密集预测任务的研究较少。在本研究中,我们确定了语义分割DD的三个关键挑战:(i)长尾类别不平衡,(ii)需要严格的像素级对齐,以及(iii)使用复杂模型优化高分辨率数据的高计算成本。为了解决这些挑战,我们提出了D3S2,一种用于语义分割的扩散引导数据集精炼框架。我们的方法采用了两阶段设计。在平衡类别掩模选择中,我们通过贪婪策略构建了一个代表性的掩模集,优先考虑了少数类。在扩散引导图像合成中,我们使用预训练的布局到图像扩散模型生成基于选定掩模的图像,自然确保空间对齐。为了进一步增强合成数据的训练效用,我们引入了两个互补目标的引导扩散采样:用于像素级对齐的分割一致性损失,以及用于跨层对齐每类特征统计的类别特征匹配损失。大量实验证明了D3S2的优越性。值得注意的是,在极端压缩率为1%时,我们的方法在ADE20K和COCO-Stuff上分别以24.99%和35.49%的mIoU,使用Mask2Former(Swin-S)优于随机选择分别9.34%和5.70%。我们的代码可在https://github.com/zwj084/D3S2 上找到。

更新时间: 2026-10-06 09:48:58

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2605.25022v2

A Riemannian Geometry for Low-rank Adaptation

Low-rank adaptation (LoRA) is widely used as a parameter-efficient fine-tuning technique for pre-trained deep neural networks, which approximates the weight update via full fine-tuning by a low-rank matrix $BA^\top$. This parameterization leads to the equivalence relation $(B, A) \sim (BG^{-1}, AG^\top)$ for any invertible matrix $G$ because $BA^\top = BG^{-1}(AG^\top)^\top$ and thus both pairs yield the same loss value. This relation induces a quotient manifold where matrices $(BG^{-1}, AG^\top)$ for all $G$ are identified, eliminating redundant directions along which the loss value remains unchanged. To respect the geometry of this manifold, the original search space is endowed with a Riemannian metric that is invariant under the equivalence relation. Such a metric induces preconditioning at each gradient step and ensures that each weight update via LoRA changes the loss value, leading to efficient optimization. In this paper, we propose a new Riemannian metric that is specifically tailored to LoRA to close the gap to full fine-tuning at the weight level. We theoretically show that LoRA with our preconditioning induced by this metric satisfies the following two properties at each iteration: (i) The weight update follows the direction closest to the gradient of full fine-tuning within the subspace of first-order weight changes allowed by the LoRA parameterization. (ii) The updated weight matrix is closer in Frobenius norm to that of full fine-tuning than the updated weight matrices of LoRA with conventional preconditioning and without preconditioning. These theoretical insights suggest that our preconditioning makes LoRA better approximate full fine-tuning, thereby leading to more efficient optimization. Experiments show the effectiveness and efficiency of our preconditioning for LoRA on fine-tuning tasks with language and vision domains.

Updated: 2026-10-06 09:47:31

标题: 一个适用于低秩自适应的黎曼几何学

摘要: 低秩适应(LoRA)被广泛用作预训练深度神经网络的参数高效微调技术,通过低秩矩阵$BA^\top$近似权重更新的全面微调。这种参数化导致等价关系$(B, A) \sim (BG^{-1}, AG^\top)$对于任何可逆矩阵$G$,因为$BA^\top = BG^{-1}(AG^\top)^\top$,因此两对都产生相同的损失值。这种关系引发了一个商流形,其中所有$G$的矩阵$(BG^{-1}, AG^\top)$被识别,消除了损失值保持不变的冗余方向。为了尊重该流形的几何性质,原始搜索空间配备了一个不受等价关系影响的黎曼度量。这种度量在每个梯度步骤中引发了预调节,并确保通过LoRA的每次权重更新都改变了损失值,从而实现了有效的优化。在本文中,我们提出了一种专门为LoRA量身定制的新的黎曼度量,以在权重级别上缩小与全面微调之间的差距。我们在理论上证明,通过这种度量诱导的我们的预调节使LoRA在每次迭代中满足以下两个属性:(i)权重更新遵循全面微调梯度的子空间内允许的一阶权重变化方向中最接近的方向。 (ii)更新后的权重矩阵在Frobenius范数上比具有传统预调节和无预调节的LoRA的更新后的权重矩阵更接近全面微调的权重矩阵。这些理论见解表明我们的预调节使LoRA更好地近似全面微调,从而导致更有效的优化。实验证明了我们的预调节对于语言和视觉领域微调任务中LoRA的有效性和效率。

更新时间: 2026-10-06 09:47:31

领域: cs.LG

下载: http://arxiv.org/abs/2610.08049v1

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, $τ^2$-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: www.github.com/illuin-tech/daedalus.

Updated: 2026-10-06 09:46:32

标题: DAEDALUS:从自我生成的任务中引导代理的记忆

摘要: LLM代理通常缺乏在新环境中可靠行动所需的操作知识,因为它们必须自行发现特定工具行为或环境惯例。没有过去尝试的记忆,它们会在任务中重复相同的错误,导致更多的任务失败和更长的轨迹。为了解决这个问题,代理系统通常依赖于人工编写的指南或从训练任务和神谕验证器构建的程序性记忆,这两者都需要对环境有先前的了解。我们提出了DAEDALUS,一种从自动生成的实践中引导可重复使用的代理记忆的方法,而不需要现有任务或神谕验证器。DAEDALUS配对了两个代理:一个探险者与环境互动以生成具有挑战性但可解决的任务,以及一个解决者尝试这些任务。从每个解决者的失败中得出一个启发式,并且只有在解决者在特定情境下反复成功后才接受该启发式。这些结果还为探险者提供反馈,以完善未来任务的难度。接受的启发式随后被整合到一个内存库中,以便在测试时使用。在AppWorld、$τ^2$-bench和AutomationBench中,DAEDALUS将平均成功率提高了高达15.9个百分点,并将通过率提高了高达2.2倍,超过了没有记忆基线的方法,并且在推理成本方面与使用训练任务的方法竞争力强。我们展示了即使在较小的探索预算下也可以获得性能提升,并且其启发式也有利于其他模型系列的代理。我们的消融实验证明,解决者的跟踪提供了生成有效启发式所需的关键信息,而在早期发现中分解使探索更具成本效益。除了记忆构建,我们发现DAEDALUS生成的任务可以作为按性能对模型排序时的基准任务的代用品。代码和工件:www.github.com/illuin-tech/daedalus。

更新时间: 2026-10-06 09:46:32

领域: cs.AI,cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.08048v1

SepsisLens: Structure-Preserving Sequence Modelling for Decomposable Early Sepsis Warning

Early sepsis warning from ICU records can be cast as a structure-preserving prediction problem. A model needs to detect deterioration from irregular measurements while keeping each alert connected to the physiological signals that support it. Many temporal models fuse clinical variables into a patient-level representation, supporting scalar risk prediction but weakening the structure needed for clinical decomposition. We present SepsisLens, which preserves variable-indexed temporal states until risk composition. Observation-aware representations encode each variable's dynamics and measurement history, while a shared temporal encoder models each trajectory without collapsing the variable axis. The StructuredRiskHead composes multi-horizon risk from explicit variable-level and organ-level components. We evaluate SepsisLens on three public ICU cohorts and one private-hospital cohort under a common pre-onset protocol. SepsisLens achieves strong discrimination on all four cohorts and lower alert burden at matched event recall on MIMIC-IV. Structural ablations support the design, while input-side masking shows that the ranked components reflect variables with greater influence on prediction.

Updated: 2026-10-06 09:45:34

标题: SepsisLens:用于可分解早期脓毒症预警的保持结构的序列建模

摘要: ICU记录中的早期脓毒症预警可以被视为一个保持结构的预测问题。模型需要在检测不规则测量中的恶化的同时,保持每个警报与支持它的生理信号相连。许多时间模型将临床变量融合为患者级表示,支持标量风险预测,但削弱了临床分解所需的结构。我们提出了SepsisLens,它将变量索引的时间状态保留到风险构成。具有观测感知表示编码每个变量的动态和测量历史,同时共享时间编码器对每个轨迹进行建模,而不会使变量轴坍塌。StructuredRiskHead从明确的变量级和器官级组件组成多时间段的风险。我们在三个公共ICU队列和一个专门医院队列上以共同的发病前协议评估了SepsisLens。SepsisLens在所有四个队列上都表现出较强的区分能力,并且在MIMIC-IV上具有匹配事件回忆的较低警报负担。结构消融支持设计,而输入端屏蔽显示排名组件反映了对预测影响更大的变量。

更新时间: 2026-10-06 09:45:34

领域: cs.LG

下载: http://arxiv.org/abs/2610.08046v1

Calibrated Uncertainty for Informative Path Planning in Aquatic Environmental Monitoring

Informative Path Planning for scalar field reconstruction uses predictive uncertainty to direct sensing vehicles toward maximally informative locations. Gaussian Processes provide this signal but their stationary isotropic kernels are misspecified for non-homogeneous phenomena such as oil spills, producing miscalibrated estimates that degrade planning. We investigate whether replacing the Gaussian Process with a well-calibrated Deep Ensemble improves path planning outcomes, and whether uncertainty quality interacts with the choice of planning algorithm. Five strategies ($ε$-Greedy, Value Greedy, Uncertainty Greedy, Monte Carlo Tree Search, and Receding Horizon Orienteering) share a common Deep Ensemble backbone trained on physics-based oil spill simulations. On held-out stochastic spill scenarios, the Deep Ensemble reduces normalised reconstruction error by $83\%$ relative to the Gaussian Process baseline. Crucially, well-calibrated uncertainty amplifies the importance of the planning strategy: the performance gap between algorithms is negligible under miscalibrated models but becomes substantial under the ensemble, where multi-step lookahead planners outperform greedy selection by up to $32\%$ in reconstruction error and achieve IoU above $0.85$. Monte Carlo Tree Search is the recommended planner, matching Orienteering in reconstruction quality at an order-of-magnitude lower computational cost.

Updated: 2026-10-06 09:42:29

标题: 校准不确定性在水生环境监测中的信息路径规划

摘要: 信息路径规划用于标量场重建,利用预测不确定性指导传感车辆朝向最具信息性的位置。高斯过程提供了这一信号,但其静止各向同性核对于非均匀现象(如油污)是错误的,产生了错误估计,降低了规划质量。我们研究了用校准良好的深度集成替换高斯过程是否能改善路径规划结果,以及不确定性质量是否与规划算法的选择有关。五种策略(ε-Greedy、Value Greedy、Uncertainty Greedy、Monte Carlo Tree Search和Receding Horizon Orienteering)共享一个基于基于物理的油污模拟训练的深度集成骨干。在保留的随机泄漏场景中,相对于高斯过程基线,深度集成将归一化重建误差降低了83%。关键是,校准良好的不确定性放大了规划策略的重要性:在错误校准模型下,算法之间的性能差距微不足道,但在集成中,多步预测规划者在重建误差上击败贪心选择最多达到32%,并实现IoU值高于0.85。蒙特卡洛树搜索是推荐的规划者,以一个数量级更低的计算成本达到与定向规划相匹配的重建质量。

更新时间: 2026-10-06 09:42:29

领域: cs.AI,cs.IR

下载: http://arxiv.org/abs/2609.34577v2

A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models

Graph neural networks (GNNs) are routinely employed for spatiotemporal forecasting, yet their performance across widely used benchmark datasets is inconsistent. Here, we perform an audit of dataset properties and baseline models to assess the quality of the benchmarks, and the robustness of the conclusions drawn from them. Using classical statistical tools, we characterise spatiotemporal lagged dependencies in benchmarks, and examine how temporal differencing changes these relationships and affects model rankings. Motivated by this, we re-evaluate temporal linear baselines, significantly reducing the apparent gains from GNNs on several benchmarks, and surpassing GNNs on others. Suspecting that GNNs struggle to extract linear, node-wise signals, we find that supplying them with autoregressive residuals improves their performance particularly on non-traffic benchmarks. Finally, controlled synthetic experiments reveal that GNNs are sensitive to heterogeneity in temporal dynamics and spatial graph interactions. Together, our findings demonstrate that baseline specification, data pre-processing and system heterogeneity shape the interpretations drawn from benchmark rankings, informing the design and robust evaluation of GNNs.

Updated: 2026-10-06 09:36:11

标题: 对时空预测基准数据集和模型的批判性审计

摘要: 图神经网络(GNNs)通常用于时空预测,然而它们在广泛使用的基准数据集上的性能并不一致。在这里,我们对数据集属性和基准模型进行审计,以评估基准的质量,以及从中得出的结论的稳健性。使用经典统计工具,我们表征基准中的时空滞后依赖关系,并检查时间差分如何改变这些关系并影响模型排名。在这一基础上,我们重新评估了时间线性基线,显著降低了GNNs在几个基准上的明显收益,并在其他基准上超越了GNNs。怀疑GNNs难以提取线性、节点级信号,我们发现为它们提供自回归残差可以显著提高它们的性能,特别是在非交通基准上。最后,受控的合成实验揭示了GNNs对时间动态和空间图交互的异质性敏感。总的来说,我们的发现表明,基准规范、数据预处理和系统异质性塑造了从基准排名中得出的解释,为GNNs的设计和稳健评估提供了信息。

更新时间: 2026-10-06 09:36:11

领域: cs.LG,stat.ML

下载: http://arxiv.org/abs/2608.20980v2

Time-multiplexed layer reuse for physical neural networks

Physical neural networks (PNNs) are promising candidates for next-generation computing, but existing demonstrations remain several orders of magnitude smaller than modern digital neural networks, whose recent advances have been driven by rapid growth in trainable parameters. This situation resembles the constraints of early digital neural networks, which led to ideas around parameter reuse. We investigate what similarly efficient hardware architectures may look like, focusing specifically on the common bottleneck of slow re-adjustment of the weights in PNNs. We propose the Time-Indexed Deep Alternating Layers Network (TIDAL-Net), which occupies an intermediate regime between recurrent and deep neural networks, specifically aimed at the scales and restrictions of common PNN prototypes. TIDAL-Net leverages the timescale separation found in many PNNs between fast forward dynamics and slowly trainable weights and biases, using layer-by-layer time multiplexing to increase effective depth while limiting implementation cost. Numerical experiments on image classification and natural language processing tasks show that TIDAL-Net improves performance with only minor modifications to conventional PNNs.

Updated: 2026-10-06 09:33:14

标题: 物理神经网络的时间复用层重用

摘要: 物理神经网络(PNNs)是下一代计算的有希望的候选者,但现有的示范仍然比现代数字神经网络小几个数量级,后者的最近进展是由可训练参数的快速增长推动的。这种情况类似于早期数字神经网络的约束,这导致了参数重用的想法。我们研究了类似效率的硬件架构可能是什么样子,特别关注PNNs中慢速权重重新调整的共同瓶颈。我们提出了时间索引的深交替层网络(TIDAL-Net),它占据了复发和深度神经网络之间的中间状态,专门针对常见PNN原型的规模和限制。TIDAL-Net利用了许多PNNs中快速前向动态和缓慢可训练权重和偏差之间的时间尺度分离,使用逐层时间复用来增加有效深度,同时限制实现成本。在图像分类和自然语言处理任务上的数值实验表明,TIDAL-Net在仅对传统PNNs进行轻微修改的情况下提高了性能。

更新时间: 2026-10-06 09:33:14

领域: cs.LG,nlin.AO

下载: http://arxiv.org/abs/2511.00044v4

RIPE++: Reinforced Keypoint Learning from Positive Pairs Only

Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration. Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real-world settings. Reinforcement learning (RL) has recently emerged as a promising alternative, requiring only the information if two images show the same scene or not. However, existing RL formulations such as RIPE rely on coarse binary rewards and carefully constructed negative training pairs, limiting training stability and descriptor discriminability. In this paper, we revisit RL-based keypoint learning and propose a reward that fully exploits the geometric consistency signal, deriving both reward and penalty from a single positive pair without contrasting against negatives. This richer signal provides sufficient supervisory contrast to learn discriminative detectors and descriptors from positive image pairs alone, enabling representation learning under extremely limited supervision. Furthermore, we show that the same RL objective can be extended to the matching stage by adapting LightGlue, raising AUC@5 on MegaDepth1500 from 56.58 to 59.65 and enabling weakly-supervised training of the full sparse matching pipeline from image pairs with partial visual overlap. We validate our approach on established benchmarks, demonstrating competitive results compared to fully-supervised methods. We further show that the method can be even trained on low texture medical video sequences, where camera poses are usually unavailable and standard SfM pipelines often fail. Code and data are available at https://github.com/fraunhoferhhi/RIPEpp .

Updated: 2026-10-06 09:29:53

标题: RIPE++:仅从正样本对中增强关键点学习

摘要: 稀疏关键点提取和匹配是几何计算机视觉中核心任务的基础,包括运动结构、视觉SLAM、增强现实和医学图像配准。然而,学习稳健的局部特征表示通常需要准确的相机姿势或深度监督,在现实世界中经常无法获取这些信息。最近,强化学习(RL)作为一种有希望的替代方法出现,只需要判断两幅图像是否显示相同的场景。然而,现有的RL公式(如RIPE)依赖于粗糙的二元奖励和精心构造的负训练对,限制了训练的稳定性和描述符的可区分性。在本文中,我们重新审视基于RL的关键点学习,并提出一种奖励,充分利用几何一致性信号,从单个正对中得出奖励和惩罚。这种更丰富的信号提供了足够的监督对比,允许仅通过正对图像对学习具有区分性的检测器和描述符,从而在极度有限的监督下进行表示学习。此外,我们展示了相同的RL目标可以通过调整LightGlue扩展到匹配阶段,将MegaDepth1500上的AUC@5从56.58提高到59.65,并实现了从具有部分视觉重叠的图像对进行弱监督训练的完整稀疏匹配流程。我们在已建立的基准测试上验证了我们的方法,与完全监督方法相比,结果具有竞争力。此外,我们进一步展示,该方法甚至可以在纹理较低的医学视频序列上进行训练,其中通常无法获取相机姿势,标准的SfM流程经常失败。代码和数据可在https://github.com/fraunhoferhhi/RIPEpp找到。

更新时间: 2026-10-06 09:29:53

领域: cs.CV,cs.LG

下载: http://arxiv.org/abs/2608.19693v2

Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis

AI agents are increasingly being programmed to automate knowledge work over large collections of unstructured data. Such automation requires repeatability: when the underlying evidence is unchanged, the agent's categories, priorities, and counts should not shift materially between runs, even if each individual answer appears plausible. We introduce a repeat-run evaluation framework that aligns semantically equivalent categories and focuses on two operating metrics: theme churn, the normalized change in the returned category set, and volume disagreement, the change in counts for categories that persist. We evaluate three recurring customer-feedback tasks across eight frontier models, corpus sizes from 100 to 5,000 records, multiple prompts, and three execution designs: raw generation, taxonomy-free hierarchical decomposition, and a taxonomy-grounded agent (TGA) using persistent themes, subthemes, and record-level predictions. With Claude Opus 4.8 and the 1,000-record corpus fixed, TGA reduces theme churn by 86--88% relative to both raw generation and hierarchical decomposition, while matched-theme volumes have zero disagreement. The taxonomy-grounded agent is more stable than every raw model in the screen, remains more stable at each corpus size, and keeps this advantage when theme matching is made stricter or looser. Although evaluated on customer feedback, the framework targets repeated synthesis of unstructured corpora more broadly, including financial reports, legal documents, incident records, and scientific literature. Overall, these results show that taxonomy grounding produces more consistent and repeatable outputs for recurring knowledge work.

Updated: 2026-10-06 09:29:25

标题: 相同的反馈,不同的答案:衡量前沿模型客户反馈分析中的运行不稳定性

摘要: 人工智能代理越来越被编程用于自动化处理大量非结构化数据的知识工作。这种自动化需要可重复性:当基础证据未改变时,代理的类别、优先级和计数在多次运行之间不应有重大变化,即使每个单独的答案看起来都是合理的。我们引入了一个重复运行评估框架,该框架将语义上等价的类别进行对齐,并关注两个运行指标:主题变动,返回的类别集合的标准化变化,以及量争议,对于持续存在的类别计数的变化。我们评估了三个反复出现的客户反馈任务,涵盖了八种前沿模型、100到5,000条记录的语料库大小、多个提示和三种执行设计:原始生成、无层级分解的分类法和使用持续主题、子主题和记录级预测的分类法基础代理(TGA)。使用Claude Opus 4.8和固定的1,000条记录语料库,TGA相对于原始生成和层级分解减少了86-88%的主题变动,而匹配的主题计数没有争议。与屏幕上的每个原始模型相比,基于分类法的代理更加稳定,在每个语料库大小下保持着这种优势,并且在主题匹配变得更加严格或更加宽松时保持这种优势。尽管在客户反馈上进行评估,但该框架更广泛地针对重复合成非结构化语料库,包括财务报告、法律文件、事故记录和科学文献。总的来说,这些结果表明,基于分类法的代理为反复发生的知识工作产生了更一致和可重复的输出。

更新时间: 2026-10-06 09:29:25

领域: cs.AI

下载: http://arxiv.org/abs/2610.08036v1

Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models

Molecular property models are commonly evaluated by holding out Bemis-Murcko scaffolds, yet a scaffold identifier is only one notion of chemical unfamiliarity. We introduce a label-free structural-frontier split that reserves the sparsest and most physicochemically remote scaffold groups, and evaluate it on six public experimental or curated ADMET tasks. Against a 70/10/20 scaffold control with identical acyclic grouping, the frontier inflates equally weighted primary error with a taskwise median of 87.0% and a skew-sensitive mean of 130.3% (descriptive task/seed bootstrap interval, 52.1-246.0%). The mean falls to 75.9% once BBB is removed; that endpoint is the one whose score ranking inverts at the frontier. A message-passing graph-network control still shows a large gap (mean 82.8% over four tasks) and does not invert, so a low-capacity head does not explain the effect. We also test Multi-View Frontier Risk Extrapolation (MV-FREX), a count-adjusted tail-risk penalty over four molecular views, and treat it as a falsifiable probe. It changes normalized frontier error by only 0.16% relative to empirical risk minimization for the perceptron head (interval, -0.43-0.84%) and by -1.9% for the graph network; three fixed robust-penalty controls are likewise inconclusive. Against the published Lo-Hi and DataSAIL splitters, the frontier inflates error more on average, though no split is uniformly hardest. An audit of 31,561 marine natural products further shows that OOD status and agreement with legacy ADMET predictions depend on the molecular view, endpoint, and teacher coverage. Split construction and label provenance are important evaluation constraints in their own right, and the tested training penalties do not resolve the frontier failures we observe.

Updated: 2026-10-06 09:28:57

标题: 超越支架分裂:结构前沿评估揭示了在ADMET模型中隐藏的失败

摘要: 分子性质模型通常通过保留Bemis-Murcko支架进行评估,然而,支架标识符只是化学陌生性的一个概念。我们引入了一个无标签的结构前沿拆分,保留了最稀疏和物理化学上最遥远的支架组,并在六个公共实验或筛选的ADMET任务上进行评估。与具有相同非环组合的70/10/20支架控制相比,前沿将等权重的主要错误率提高,任务中位数为87.0%,偏度敏感平均值为130.3%(描述性任务/种子bootstrap区间为52.1-246.0%)。一旦去除BBB,平均值降至75.9%;该端点是在前沿处评分排名发生反转的端点。消息传递图网络控制仍然显示出较大差距(四个任务的平均值为82.8%),并且不会发生反转,因此低容量头部不能解释这种效应。我们还测试了Multi-View Frontier Risk Extrapolation(MV-FREX),它是基于四种分子视图的计数调整尾部风险惩罚,并将其视为可验证的探针。相对于感知器头部的经验风险最小化,它仅将标准化前沿错误率改变了0.16%(区间为-0.43-0.84%),对于图网络来说,这个变化为-1.9%;三个固定的稳健惩罚控制同样没有明确结论。与已发表的Lo-Hi和DataSAIL分割器相比,前沿平均上增加了错误率,尽管没有一个分割是一致最困难的。对31,561种海洋天然产物的审计进一步显示,ODD状态和与传统ADMET预测的一致性取决于分子视图,端点和教师覆盖范围。拆分构造和标签来源本身是重要的评估约束条件,而我们观察到的前沿失败与测试的训练惩罚无法解决。

更新时间: 2026-10-06 09:28:57

领域: cs.LG,q-bio.QM

下载: http://arxiv.org/abs/2607.10729v4

Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA

World models are usually judged from the inside: by prediction loss, by the return a policy earns in imagination, or by how convincing their frames look. We judge one from the outside. We learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with. The real game never provides a gradient; it provides the policy's own games as training data for the world model, and an online evaluation that selects and anchors the policy. Run as a continuous asynchronous Dyna loop, the policy wins 70.2% of real games as radiant (421 of 600; 95% CI 66.4-73.7) on seeds never used for any decision, up from 0% for dream training alone and 33.7% before the loop. It wins none as dire, and neither does the shipped opponent when it plays itself. Four findings explain the result. Model exploitation is invisible from inside the dream: every unanchored run collapsed within a few updates while no in-dream metric tracked the collapse. A world model that is accurate on its training corpus is badly wrong on the policy's own games, and Dyna repairs it there, which is worth +9.2 points of real win rate with the policy recipe held fixed. Finally, the policy inherits its world model's fidelity profile mechanic by mechanic: the model represents the macro game but not crowd control, cast timing or lethality, and the policy wins by map-wide pressure with almost no coordinated fighting. We release the world model, the dream-PPO harness, a world-model debugger, the evaluation protocol, and every policy and log.

Updated: 2026-10-06 09:26:48

标题: 梦中学习,现实获胜:一个十英雄MOBA游戏的连续动态循环

摘要: 世界模型通常是通过内部评判的:通过预测损失,通过政策在想象中获得的回报,或者通过它们的帧看起来有多令人信服。我们从外部进行评判。我们学习了一个完整的十英雄MOBA的结构化多智能体世界模型(206个单位,每个英雄每个时刻行动,游戏持续最多6,000个时刻),只在其中训练政策,并通过1,400个时刻的自由运行想象的情节来衡量该政策在真实游戏中对抗游戏中的对手。真实游戏从不提供梯度;它提供政策自己的游戏作为世界模型的训练数据,并提供选择和锚定政策的在线评估。作为连续异步Dyna循环运行,该政策在真实游戏中以天辉身份赢得70.2%的比赛(600场中的421场;95%CI 66.4-73.7),这些比赛从未用于任何决策,而在仅进行梦幻训练时为0%,在循环之前为33.7%。它在天辉身份下赢得没有一场比赛,当游戏对手自己玩时也是如此。四个发现解释了这一结果。模型的利用在梦幻中是看不见的:每个未锚定运行在几次更新后就崩溃了,而没有任何梦幻度量指标跟踪崩溃。一个在其训练语料库上准确的世界模型在政策自身的游戏中是严重错误的,而Dyna在那里修复了它,这在政策食谱不变的情况下值得+9.2个真实胜率。最后,政策逐个机制继承其世界模型的忠实度配置文件:该模型代表了宏观游戏,但没有人群控制、施法时机或致命性,并且该政策通过几乎没有协调战斗的地图范围压力赢得了比赛。我们发布了世界模型、梦想PPO工具、世界模型调试器、评估协议以及每个政策和日志。

更新时间: 2026-10-06 09:26:48

领域: cs.AI

下载: http://arxiv.org/abs/2610.08033v1

Spectra: Exact Component Transport for Test-Time Prior Adaptation in Simulation-Based Inference

Simulation-based inference (SBI) has become a powerful approach to Bayesian inference in complex scientific models whose likelihoods are difficult or impossible to evaluate. Amortized SBI learns reusable inference models from simulated data, enabling rapid posterior inference for new observations, and modern generative models have made these models increasingly expressive. However, this reuse is limited to the prior distribution chosen during training, whereas scientific analyses often need revised priors as knowledge accumulates or alternative assumptions are tested. We introduce Spectra, a test-time adaptation method for diffusion-based SBI. Spectra uses an exact score-transport identity to obtain the adapted score from a frozen diffusion model in closed form for structured prior changes, without additional simulation or training. Across six SBI benchmarks, Spectra achieves accurate adaptation under strong prior shifts at low online sampling cost. This enables pretrained SBI models to incorporate updated prior information at test time.

Updated: 2026-10-06 09:14:56

标题: 频谱:模拟推断中测试时先验适应的准确组件传输

摘要: 基于模拟的推断(SBI)已经成为贝叶斯推断在复杂科学模型中的强大方法,这些模型的似然函数难以或不可能评估。摊销 SBI 通过从模拟数据中学习可重复使用的推断模型,实现了对新观测的快速后验推断,现代生成模型使这些模型越来越具有表现力。然而,这种重复利用仅限于训练过程中选择的先验分布,而科学分析通常需要随着知识积累或测试替代假设而进行修订的先验。我们介绍了 Spectra,一种基于扩散的 SBI 的测试时间适应方法。Spectra 使用精确的得分传输恒等式,以封闭形式从冻结的扩散模型中获得适应得分,用于结构化先验更改,无需额外的模拟或训练。在六个 SBI 基准测试中,Spectra 在低在线采样成本下实现了对强先验转移的准确适应。这使预训练的 SBI 模型能够在测试时间吸收更新的先验信息。

更新时间: 2026-10-06 09:14:56

领域: cs.LG,stat.ML

下载: http://arxiv.org/abs/2610.08021v1

Learning consistent molecular mechanics force fields from first principles

Classical force fields (FFs) remain the workhorse for large-scale simulations even as machine-learned interatomic potentials (MLIPs) approach ab initio accuracy. They decompose total configuration energies into simple effective interactions whose parameters are traditionally assigned based on atom or bond types, enabling efficient simulations but also limiting their ability to adapt across configurations. Recent machine learning approaches have improved the accuracy and transferability of bonded parameters in these FFs by inferring them as functions of local atomic environments, but still rely on empirical nonbonded parameters for practical simulations. In this work, we introduce a unified approach, \texttt{grappa-fullFF}, which learns both bonded and nonbonded parameters \emph{consistently} and simultaneously from ab initio reference data. By incorporating physically inspired regularization via supervision of the electrostatic potential and an architecture that facilitates charge equilibration, our model recovers accurate electric response properties, achieves state-of-the-art accuracy on geometry optimization benchmarks, and reproduces the conformational sampling of both classical and existing machine-learned FFs, without relying on externally assigned nonbonded parameters.

Updated: 2026-10-06 09:14:50

标题: 学习一致的从第一原理推导的分子力学势场

摘要: 经典力场(FFs)仍然是大规模模拟的主力军,即使机器学习的原子间势(MLIPs)接近从头算准确度。它们将总配置能量分解为简单的有效相互作用,其参数传统上基于原子或键类型分配,实现了高效的模拟,但也限制了它们适应各种配置的能力。最近的机器学习方法通过推断它们作为局部原子环境的函数来改进这些FFs中的键合参数的准确性和可转移性,但仍然依赖于经验非键合参数进行实际模拟。在这项工作中,我们引入了一种统一的方法,\texttt{grappa-fullFF},它从从头算参考数据中一致地学习键合和非键合参数,并同时学习。通过纳入物理启发的正则化,通过监督静电势和促进电荷均衡的体系结构,我们的模型恢复了准确的电响应特性,在几何优化基准测试中实现了最先进的准确性,并再现了经典和现有的机器学习FFs的构象采样,而不依赖于外部分配的非键合参数。

更新时间: 2026-10-06 09:14:50

领域: physics.chem-ph,cs.LG,physics.comp-ph

下载: http://arxiv.org/abs/2610.08020v1

PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning

Tasks on complex systems require high-precision numerical computation to support decisions. However, current large language models (LLMs), even with enhanced reasoning capabilities, cannot integrate such computations as an intrinsic and interpretable capability with existing architectures. To this end, we propose Physically-isolated Experts Routing Network (PiERN), an architecture that directs computation and reasoning at token level, thereby enabling iterative alternation within a single chain of thought. We systematically evaluate PiERN on representative computation-reasoning tasks, including PDEBench and battery management tasks. Results show that PiERN achieves not only higher accuracy than directly finetuning LLMs but also significant improvements in response latency, token usage, GPU energy consumption, and experts routing accuracy compared with mainstream multi-agent approaches, while exhibiting no significant degradation in performance on MMLU and GLUE benchmarks. PiERN offers an efficient, interpretable, and scalable paradigm for interfacing language models with scientific systems.

Updated: 2026-10-06 09:14:08

标题: PiERN:用于集成高精度计算和推理的令牌级路由

摘要: 复杂系统的任务需要高精度的数值计算来支持决策。然而,当前的大型语言模型(LLMs),即使具有增强的推理能力,也无法将这种计算作为一种内在且可解释的能力与现有架构集成。为此,我们提出了Physically-isolated Experts Routing Network(PiERN),这是一种在令牌级别指导计算和推理的架构,从而在单一思维链内实现迭代交替。我们系统地评估了PiERN在代表性的计算推理任务上的表现,包括PDEBench和电池管理任务。结果显示,与直接微调LLMs相比,PiERN不仅实现了更高的准确性,而且在响应延迟、令牌使用、GPU能耗和专家路由准确性方面与主流多代理方法相比也取得了显著改进,同时在MMLU和GLUE基准测试中表现没有明显下降。PiERN为将语言模型与科学系统接口提供了一种高效、可解释且可扩展的范式。

更新时间: 2026-10-06 09:14:08

领域: cs.LG,cs.CE,cs.CL

下载: http://arxiv.org/abs/2509.18169v4

FOSLS-deRhaNN: native de Rham neural classes for H(div) and H(curl) with applications to first-order system least-squares neural network methods for partial differential equations

We construct neural approximation classes native to the graph spaces H(div) and H(curl), in two and three dimensions and, for H(div), in any dimension. Every realization lies in the space for all parameter values, and with kinked potentials, such as ReLU networks, the admissible jumps appear at finite width. The classes are images of scalar and componentwise networks under fixed operators of the de Rham complex, and do not involve a mesh or finite element emulation. For H(div) in R^n two native classes are given on an equal footing, with a skew-symmetric potential $A$: $\mathrm{Div}\,A+R_nq+\mathbf{h}$, with the divergence $q$ as an explicit unknown, and $\mathrm{Div}\,A+\mathbf{z}$ with an $H^1$ field $\mathbf{z}$; for H(curl) the analogous classes are $\mathrm{grad}\,φ+Sr+\mathbf{h}$ in two dimensions and $\mathrm{grad}\,φ+\mathbf{z}$ in two and three dimensions. In all of them every interface jump of the field is carried by the potential term, $\mathrm{Div}\,A$ or $\mathrm{grad}\,φ$, while the remaining part has no interface jump (it is an $H^1$ field in the regular-decomposition classes); the classes with $\mathbf{z}$ are the componentwise approach enriched by this term. Known or learned interface geometry enters the potential through factors with trainable amplitudes, and the remaining part if the divergence jumps. The classes lead to the FOSLS-deRhaNN method, first-order system least squares with de Rham neural networks, whose loss is the least-squares functional posed in the natural spaces of the weak formulation; for elliptic equations this includes $H^{-1}$ right-hand sides and $H^{1/2}$ Dirichlet data. Elliptic equations with discontinuous coefficients and curl-curl problems are treated as instances, with the functional equivalent to the error; linear transport with discontinuous solutions and conservation laws with shocks use the same flux classes.

Updated: 2026-10-06 09:12:45

标题: FOSLS-deRhaNN:H(div)和H(curl)的本征de Rham神经类及其在偏微分方程的一阶系统最小二乘神经网络方法中的应用

摘要: 我们构建了适用于图空间H(div)和H(curl)的神经逼近类,分别在二维和三维中,对于H(div),在任意维度中均适用。每个实现都位于所有参数值的空间中,对于具有弯曲势能的网络,如ReLU网络,可接受的跃点出现在有限宽度处。这些类是de Rham复形的固定算子下标量和分量网络的映像,并且不涉及网格或有限元仿真。对于R^n中的H(div),有两个本地类具有相同地位,其中包含一个反对称势能$A$:$\mathrm{Div}\,A+R_nq+\mathbf{h}$,其中散度$q$是一个明确的未知量,以及$\mathrm{Div}\,A+\mathbf{z}$,其中包含一个$H^1$场$\mathbf{z}$;对于H(curl),类似的类包括二维中的$\mathrm{grad}\,φ+Sr+\mathbf{h}$和二维和三维中的$\mathrm{grad}\,φ+\mathbf{z}$。在所有这些类中,场的每个界面跃点都由势能项$\mathrm{Div}\,A$或$\mathrm{grad}\,φ$承载,而其余部分没有界面跃点(在正则分解类中是一个$H^1$场);带有$\mathbf{z}$的类是通过这个项丰富的分量方法。已知或学习的界面几何通过具有可训练振幅的因子进入势能,剩余部分是散度跃点。这些类导致了FOSLS-deRhaNN方法,即de Rham神经网络的第一序系统最小二乘法,其损失是在弱式自然空间中提出的最小二乘泛函;对于椭圆方程,这包括$H^{-1}$右端项和$H^{1/2}$迪利克雷数据。具有不连续系数的椭圆方程和卷曲问题被视为实例,其功能等同于误差;具有不连续解的线性输运和具有激波的守恒定律使用相同的通量类进行处理。

更新时间: 2026-10-06 09:12:45

领域: math.NA,cs.LG

下载: http://arxiv.org/abs/2610.08016v1

Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate

The sharpness-aware minimization (SAM) algorithm and its variants, including gap guided SAM (GSAM), have been successful at improving the generalization capability of deep neural network models by finding flat local minima of the empirical loss in training. Meanwhile, it has been shown theoretically and practically that increasing the batch size or decaying the learning rate avoids sharp local minima of the empirical loss. In this paper, we consider the GSAM algorithm with increasing batch sizes or decaying learning rates, such as cosine annealing or linear learning rate, and theoretically show its convergence. Moreover, we numerically compare SAM (GSAM) with and without an increasing batch size and conclude that using an increasing batch size { achieves a lower worst-case $\ell_\infty$ adaptive sharpness} than compared with using a constant batch size and learning rate.

Updated: 2026-10-06 09:03:23

标题: 锐度感知最小化算法的收敛:使用增大批量大小和衰减学习率

摘要: 锐度感知最小化(SAM)算法及其变体,包括间隙引导SAM(GSAM),通过在训练中找到经验损失的平坦局部最小值成功改进了深度神经网络模型的泛化能力。同时,理论上和实践上已经表明,增加批量大小或衰减学习率可以避免经验损失的尖锐局部最小值。在本文中,我们考虑了GSAM算法与增加批量大小或衰减学习率(如余弦退火或线性学习率),并在理论上展示了其收敛性。此外,我们在数值上比较了带有增加批量大小和不带增加批量大小的SAM(GSAM),并得出结论,在使用增加批量大小时,可以获得比使用恒定批量大小和学习率更低的最坏情况$\ell_\infty$自适应尖锐度。

更新时间: 2026-10-06 09:03:23

领域: cs.LG,math.OC

下载: http://arxiv.org/abs/2409.09984v2

Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions

In finance, interpreting machine learning predictions is essential, yet the numerical outputs of explainable AI can be difficult for non-experts to understand. While large language models (LLMs) can translate these outputs into natural language, they may produce errors when inferring numerical changes and feature relations. We propose an LLM narrative framework for cross-sectional stock return prediction that combines temporal Shapley additive explanations (SHAP) evidence with historical regime analogs. Temporal evidence tracks changes in the normalized global SHAP importance of an XGBoost model over six months. Historical analogs are past periods with similar changes in SHAP importance, their model performance and subsequent market returns are provided as comparative context. Using this framework, we conduct a controlled study of progressive reasoning externalization, sequentially providing raw SHAP sequences, deterministic temporal descriptors, and feature relations. Each generated claim is verified against provenance-linked evidence. Across Qwen3, externalizing numerical and relational reasoning improved evidence faithfulness as well as temporal and relational accuracy. Evidence faithfulness increased from 0.696 to 0.996 for Qwen3-32B-Instruct. While historical analogs did not improve structured automatic faithfulness, they received higher human-rated usefulness scores. These results suggest that externalizing verifiable reasoning enhances narrative faithfulness and that historical context adds interpretive value.

Updated: 2026-10-06 09:02:14

标题: 股票收益预测的忠实大型语言模型叙事的推理外部化

摘要: 在金融领域,解释机器学习预测是至关重要的,然而,可解释AI的数值输出对非专家来说可能很难理解。虽然大型语言模型(LLMs)可以将这些输出转化为自然语言,但在推断数值变化和特征关系时可能会产生错误。我们提出了一个LLM叙事框架,用于横截面股票回报预测,结合了时间Shapley加法解释(SHAP)证据和历史制度模拟。时间证据跟踪了在六个月内XGBoost模型的规范化全球SHAP重要性的变化。历史模拟是过去具有类似SHAP重要性变化的时期,它们的模型表现和随后的市场回报被提供作为比较上下文。利用这一框架,我们进行了一个有序的控制研究,将逐步提供原始SHAP序列、确定性时间描述符和特征关系。每个生成的声明都经过与来源关联的证据验证。在Qwen3上,外部化数值和关系推理提高了证据的忠实度以及时间和关系的准确性。对于Qwen3-32B-Instruct,证据的忠实度从0.696提高到0.996。虽然历史模拟并没有提高结构化自动忠实度,但它们获得了更高的人类评分有用性分数。这些结果表明,外部化可验证的推理增强了叙事的忠实度,并且历史背景增加了解释价值。

更新时间: 2026-10-06 09:02:14

领域: cs.AI,cs.CE

下载: http://arxiv.org/abs/2609.38869v2

Can phenotypic activity be predicted without experimental readouts?

Molecular encoders contrastively pretrained on paired molecule-morphology data, such as CLOOME and CellCLIP, have been proposed as cheap surrogates for phenotypic prediction, avoiding the need to run a Cell Painting assay. We evaluate this idea for these molecular encoders under a protocol designed to control for two confounds that can inflate apparent performance: leakage across an encoder's own pretraining boundary, and the correlation between phenotypic activity and cytotoxicity. Testing six representations, including a non-pretrained MLP control matching CLOOME's input and layer count, on two distinct Cell Painting screens, we find that once these confounds are controlled for, the pretrained molecular encoders show no clear advantage over plain physicochemical descriptors, and that toxicity is generally easier to predict than phenotypic activity across representations. Our results suggest leakage-aware, confound-controlled evaluation should be standard practice before phenotype-pretrained encoders are trusted as surrogates for phenotypic drug discovery.

Updated: 2026-10-06 08:56:23

标题: 在没有实验结果的情况下能预测表型活动吗?

摘要: 分子编码器对配对分子形态数据进行对比预训练,例如CLOOME和CellCLIP,已被提议作为廉价的表型预测替代品,避免运行细胞绘画测定。我们评估了这一想法对这些分子编码器的影响,该协议旨在控制两个可能夸大表现的混淆因素:跨编码器自身预训练边界的泄漏,以及表型活性与细胞毒性之间的相关性。在两个不同的细胞绘画屏幕上测试了六种表示,包括一个未经预训练的MLP对照,与CLOOME的输入和层计数匹配,我们发现一旦这些混淆因素得到控制,预训练的分子编码器与纯物理化学描述符相比并无明显优势,并且毒性通常比表型活性更容易预测。我们的结果表明,在信任表型药物发现的替代品之前,应该采用泄漏感知、混淆控制的评估作为标准做法。

更新时间: 2026-10-06 08:56:23

领域: cs.LG

下载: http://arxiv.org/abs/2610.07997v1

TICDA: Tabular In-Context Data Attribution

Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any parameter update. Yet how individual demonstrations shape a given prediction remains poorly understood. This gap matters in practice: the context is often assembled from whatever labeled data is available, potentially leading to the inclusion of mislabeled, redundant, or low-quality examples that degrade performance. Standard data attribution methods do not transfer to the TFM setting: resampling-based approaches such as DemoShapley require a combinatorial number of forward passes, and gradient-based estimators such as influence functions require computing training point's effect on the model parameters, which in-context learning never updates. We introduce TICDA, a method that measures the influence of every demonstration in the context directly from linear surrogates trained on TFM latent embeddings, in a single forward pass and at negligible cost. We show that TICDA offers the best compromise against competitors across four tasks: detecting labeling errors, curating context to preserve predictive accuracy while lowering inference cost, producing attribution scores that transfer across TFMs, and supporting an acquisition strategy for efficient active learning.

Updated: 2026-10-06 08:55:27

标题: TICDA:表格中的数据属性化

摘要: 表格基础模型(TFMs)通过在上下文中提供的标记演示条件下,实现了强大的预测性能,而无需进行任何参数更新。然而,个别演示如何塑造给定预测仍然知之甚少。在实践中,这种差距很重要:上下文通常是从可用的标记数据中组装而成,可能导致包含错误标记、冗余或低质量示例,从而降低性能。标准数据归因方法无法转移到TFM设置:基于重采样的方法(如DemoShapley)需要进行组合数量的前向传递,而基于梯度的估计器(如影响函数)需要计算训练点对模型参数的影响,而在上下文学习中永不更新。我们引入了TICDA,一种方法,它通过在TFM潜在嵌入上训练的线性替代物直接测量上下文中每个演示的影响,只需一次前向传递,成本微乎其微。我们展示了TICDA在四项任务中对竞争对手提供了最佳折衷方案:检测标记错误,筛选上下文以保持预测准确性同时降低推理成本,生成跨TFM的归因分数,并支持高效主动学习的获取策略。

更新时间: 2026-10-06 08:55:27

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.07996v1

A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic

Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates. In this work, we investigate whether this regularization is actually useful. Surprisingly, empirical results show that optimizing merged-model weights without this regularization significantly boosts the performance of common merging methods across multiple architectures, domains, and even in an extremely data-limited scenario where only one instance is available per class. Moreover, directly optimizing the pretrained model weights even outperforms some existing merging methods. Analysis shows that better multi-task weights exist outside the subspace and can be found using multiple methods. We study different strategies for using the additional dataset, discussing their practical use and implications for model merging. Overall, this work calls for revisiting the existing model-merging pipeline, motivating a broader exploration of the weight space and a reconsideration of the implicit regularization induced by task arithmetic.

Updated: 2026-10-06 08:49:40

标题: 对模型合并的更广泛探讨:重新思考由任务算术引发的隐式正则化

摘要: 模型合并旨在通过结合各个任务特定模型的权重便宜地构建多任务模型。为了在多个任务中表现良好,大多数现有的合并方法使用额外的数据集来找到最佳线性组合任务特定权重更新的系数。然而,我们发现了这种标准做法中的一种隐式正则化:在系数上搜索会将候选模型限制在由任务特定权重更新张成的子空间中。在这项工作中,我们调查了这种正则化是否实际上有用。令人惊讶的是,实证结果显示,在多个架构、领域甚至在极度数据有限的情况下,即每类只有一个实例可用的情况下,优化合并模型权重而不使用这种正则化显著提升了常见合并方法的性能。此外,直接优化预训练模型权重甚至胜过了一些现有的合并方法。分析显示,更好的多任务权重存在于子空间之外,并且可以通过多种方法找到。我们研究了使用额外数据集的不同策略,讨论了它们的实际用途以及对模型合并的隐式正则化引起的影响。总的来说,这项工作呼吁重新审视现有的模型合并流程,促使对权重空间进行更广泛的探索,并重新考虑由任务算术引起的隐式正则化。

更新时间: 2026-10-06 08:49:40

领域: cs.LG,cs.CL,cs.CV

下载: http://arxiv.org/abs/2610.07990v1

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.

Updated: 2026-10-06 08:48:50

标题: VisionWeave: 将弹性视觉表示编织为MLLMs的本地能力

摘要: 多模态大型语言模型已成为视觉理解的主导范式,但通过将输入编码为密集的固定大小的补丁标记而产生巨大成本。然而,视觉信息分布不均匀:一些区域需要细粒度的细节,而其他区域可以采用紧凑的表示。降采样牺牲了这些细节,而现有的标记修剪和自适应方法在内容自适应粒度、任务泛化以及与现代MLLMs和服务基础设施的集成方面仍存在局限性。要克服这些限制,需要基础模型学习,端到端地学习在哪里以及以何种粒度分配视觉表示,这是一种我们称之为弹性视觉表示编织的本地能力。我们引入了VisionWeave,通过大规模训练在前沿级别的MLLMs中建立了这种能力。它结合了两个组件:一个门控空间池构建了粗粒度表示,同时在共享的MRoPE坐标内构建了原生细粒度表示,而一个粒度路由器学习它们的内容自适应分配。仅通过自蒸馏,我们在Qwen3.5-4B上验证了这种能力,并扩展到Qwen3.8-27B,超过30K个A100 GPU小时。基于Qwen3.8-27B,VisionWeave自适应调整标记节省到视觉内容,平均节省43.0%的标记,同时保持了八项基准测试中98.9%的原生性能,而对于具有固定50%节省目标的标记修剪基准仅保留了88%的性能。广泛的评估证实了不同任务、分辨率和视频帧之间的稳健的效率-质量权衡。当部署在SGLang服务引擎上时,我们的方法实现了2.3倍的吞吐量增益,同时将平均TTFT减少了54.4%,平均TPOT减少了60.6%。总的来说,我们相信这些结果将弹性视觉编织定位为下一代多模态模型的有前景的能力。

更新时间: 2026-10-06 08:48:50

领域: cs.CV,cs.AI,cs.CL

下载: http://arxiv.org/abs/2610.07987v1

Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels

Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a frozen 27B model that answers each question with and without each memory's pixels. With a 7B answering model, PixelTriage lies on the accuracy--cost frontier of M$^3$Exam, DMV and MemEye and uses 11--23\% of the visual tokens without a significant loss of accuracy. On DMV it answers 2.9 times faster than opening all images. It outperforms retrieval order and uniform down-sizing at equal budgets and transfers to other memory systems and to a 397B answering model.

Updated: 2026-10-06 08:47:45

标题: 在查看之前决定:学习哪些检索到的记忆值得关注Pixels

摘要: 多模态助手回答包含图像的长期记忆中的问题。在检索之后,每个检索到的图像以大约每个图像一千个视觉标记的像素形式或作为存储的文本代理到达回答模型,后者通常缺少问题所询问的细节。我们发现像素的好处通常来自一个或两个检索到的记忆,并且可以在回答模型运行之前进行预测,而无需阅读任何全分辨率图像。在PixelTriage中,放置在检索之后的插件,一个不生成文本的小型模型读取对话,每个检索到的记忆的简短笔记和缩略图,并预测其像素会添加多少。它是在由一个冻结的27B模型标记的合成记忆片段上进行训练的,该模型用每个记忆的像素回答每个问题。在一个7B回答模型中,PixelTriage位于M$^3$Exam,DMV和MemEye的精度-成本前沿,并且在不显著损失精度的情况下使用了11-23%的视觉标记。在DMV上,它比打开所有图像快2.9倍。在相同预算下,它优于检索顺序和均匀缩减大小,并且可以转移到其他记忆系统和一个397B回答模型。

更新时间: 2026-10-06 08:47:45

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.07984v1

Do Higher-Order Models Win for Higher-Order Reasons? Rethinking Performance Gains in Hypergraph Learning

Higher-order models (e.g., hypergraph neural networks) often outperform lower-order baselines on hypergraph learning benchmarks, and their advantages are commonly attributed to their ability to exploit higher-order information. However, better performance alone does not establish this explanation. We therefore ask: Do higher-order models win for higher-order reasons? To investigate this question, we introduce a controlled performance-attribution framework that perturbs higher-order information while preserving the lower-order, i.e., pairwise, information. Across 25 commonly used hypergraph learning benchmarks spanning three tasks, we frequently observe an intriguing pattern: higher-order models originally outperform lower-order baselines, yet retain most of their advantage after perturbation. This suggests that much of the observed advantage remains achievable without the higher-order information. We then investigate potential lower-order explanations for these remaining gaps. We find that simple additions to a lower-order baseline, e.g., richer pairwise weighting, more steps of pairwise feature propagation, and normalization, reduce the remaining performance gaps, supporting lower-order explanations for part of the observed advantage. Our analysis calls for the hypergraph learning community to rethink performance attribution by distinguishing performance gains from their explanations, adopt stronger lower-order baselines, and use suitable benchmarks that better test the value of higher-order information.

Updated: 2026-10-06 08:44:38

标题: 更高阶模型因何获胜?重新思考超图学习中的性能提升

摘要: 更高阶模型(例如,超图神经网络)通常在超图学习基准测试中表现优于低阶基线,并且它们的优势通常被归因于它们利用更高阶信息的能力。然而,仅仅更好的性能并不能证明这个解释。因此,我们提出一个控制性能归因框架,扰动更高阶信息同时保留更低阶(即成对)信息,以调查这个问题。在涵盖三个任务的25个常用超图学习基准测试中,我们经常观察到一个有趣的模式:更高阶模型原本优于更低阶基线,但在扰动后仍保留大部分优势。这表明观察到的优势大部分仍可在没有更高阶信息的情况下实现。然后,我们研究了这些剩余差距的可能更低阶解释。我们发现对于这些剩余性能差距,简单的对更低阶基线进行补充,例如更丰富的成对加权、更多步骤的成对特征传播和归一化,可以减少这些差距,支持部分观察到的优势的更低阶解释。我们的分析呼吁超图学习社区重新思考性能归因,通过区分性能增益与其解释,采用更强的更低阶基线,并使用更适合测试更高阶信息价值的基准测试。

更新时间: 2026-10-06 08:44:38

领域: cs.LG,cs.SI

下载: http://arxiv.org/abs/2610.07981v1

Incentive Alignment in Online Experimentation

Evaluating the causal effect of new features is a central goal for online platforms. While recent literature addresses limited testing traffic via centralized portfolio optimization, this perspective abstracts away a critical institutional reality: experimentation is operationally decentralized. The experimenters who develop new features also dictate which hypotheses to test, and they are typically rewarded based on empirical average treatment effects that are prone to upward bias. Left unchecked, this principal-agent conflict can severely erode platform value, a structural failure that conventional centralized levers, such as significance thresholds and traffic budgets, cannot resolve. By reframing experimentation as an incentive design problem, we demonstrate that two practical mechanisms, sample splitting and shrinkage, can effectively bridge this gap. Sample splitting aligns incentives perfectly at a bounded traffic cost, while shrinkage consumes no additional traffic and guarantees that interventions with negative expected effects are strictly unprofitable to field.

Updated: 2026-10-06 08:44:16

标题: 在线实验中的激励对齐

摘要: 评估新功能的因果效应是在线平台的中心目标。尽管最近的文献通过集中的投资组合优化解决了有限的测试流量问题,但这种观点忽略了一个关键的制度现实:实验操作是分散的。开发新功能的实验者也决定要测试哪些假设,并且通常根据容易产生向上偏差的经验性平均处理效应来获得奖励。如果不加以控制,这种委托代理冲突可能会严重侵蚀平台价值,这是传统集中杠杆,如显著性阈值和流量预算无法解决的结构性失败。通过将实验重新框定为激励设计问题,我们证明了两种实际机制,样本拆分和缩减,可以有效弥合这一差距。样本拆分在有限的交通成本下完美地调整激励,而缩减不需要额外的流量,并保证具有负预期效果的干预对野外严格不利可盈利性。

更新时间: 2026-10-06 08:44:16

领域: cs.GT,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.05922v2

Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents

As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills govern how agents discover and improve future Skills; their value therefore emerges through the subsequent search processes they induce. Existing approaches improve Meta-Skills from observed raw Skill-search trajectories and branch outcomes. However, branch performance entangles the effects of the initial discovery state and the Meta-Skill revision that generated the search process, making it difficult to characterize what a particular revision actually changed, and pushing updates toward revisions that benefit from favorable states rather than those that improve the process. We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents. HMED revisits the completed event from which a revision originates and re-executes the incumbent and revised Meta-Skills from the same restored discovery state, so that the changes associated with the revision can be observed under a shared condition. Each comparison is distilled into a Meta-Experience, a structured record that can be reused by future updates, so that even revisions that are not ultimately retained still contribute a learning signal. Across three interactive agent benchmarks and both open-source and closed-source models, HMED consistently improves Skill discovery performance over strong baselines, shifting Meta-Skill learning beyond branch outcomes toward the consequences of changing the improvement process.

Updated: 2026-10-06 08:43:43

标题: 从修订结果中学习:事后元经验提炼以改进自主代理者

摘要: 随着代理不断通过生成和修订技能来改进,发现和完善这些技能的过程本身成为一个可学习的对象。任务技能直接影响任务执行,而元技能则指导代理如何发现和改进未来的技能;因此,它们的价值是通过它们引发的后续搜索过程而出现的。现有方法通过观察到的原始技能搜索轨迹和分支结果改进元技能。然而,分支性能使得初始发现状态和生成搜索过程的元技能修订的影响纠缠在一起,这使得很难确定特定修订实际改变了什么,并且将更新推向从有利状态中受益而不是改进过程的修订。我们引入了HMED(后见之明元经验蒸馏),这是一种为自我改进代理构建元经验的机制。HMED回顾了修订源自的已完成事件,并从相同恢复的发现状态重新执行现有和修订的元技能,以便可以在共同条件下观察到与修订相关的变化。每次比较都被提炼成一个元经验,这是一个结构化记录,可以被未来的更新重复使用,这样即使最终不保留的修订也仍然贡献了学习信号。在三个交互式代理基准测试和开源和闭源模型中,HMED始终比强基线提高技能发现性能,将元技能学习从分支结果转向改进过程变化的后果。

更新时间: 2026-10-06 08:43:43

领域: cs.AI

下载: http://arxiv.org/abs/2610.07979v1

Quantifying the Privacy Posture of Operator-Side 5G/O-RAN Profiles

Operator-side network profiles derived from 5G/ORAN traffic carry personal data such as ephemeral subscriber identifiers, slice-level KPIs, and control-plane signalling, and must be anonymised before release to a federated-learning aggregator, threat-intelligence exchange, or ML training pipeline. We study how much re-identification risk remains after standard operator-side anonymisation. We quantify privacy posture with k-anonymity, l-diversity and t-closeness, aggregate them into a composite Privacy-Posture Index (PPI), and measure residual re-identification across eight transformation configurations on internal PCAP captures and the public Idaho Labs 5GAD corpus, under a full-QI syntactic bound and two simulated adversaries. The evaluation is modest in scale, and we read its trends as indicative rather than definitive. Three findings emerge. Pseudonymisation alone leaves re-identification unchanged; material privacy gains arise when quasi-identifiers are coarsened through generalisation, optionally combined with suppression. A downstream classification task then shows that suppression-heavy releases retain majority-class utility but sacrifice much of their minority-class recall, a cost the aggregate metrics hide. Finally, the standard kmin-based PPI correlates only modestly with the disclosure bound and not at all with the partial-knowledge attack, whereas a mean-class-size variant PPI correlates strongly with all three disclosure/attack measures; we therefore read PPI as a regulator-facing summary, not a security bound. The profiles are produced by passive operator-side monitoring with rule-based DPI; our contribution is the privacy-quantification layer that computes these metrics, applies the transformation policy, and exposes both through an inspectable dashboard.

Updated: 2026-10-06 08:42:07

标题: 量化运营商端5G/O-RAN配置文件的隐私姿态

摘要: 来自5G/ORAN流量的运营商侧网络配置文件携带个人数据,如短暂的订阅者标识符、切片级KPI和控制平面信令,在释放到联合学习聚合器、威胁情报交换或ML训练管道之前必须进行匿名化。我们研究了标准运营商侧匿名化后仍然存在多少再识别风险。我们用k-匿名性、l-多样性和t-接近性量化隐私姿势,将它们聚合成一个综合隐私姿势指数(PPI),并在内部PCAP捕获和公共爱达荷实验室5GAD语料库上的八种转换配置下测量残留的再识别,基于全QI句法界限和两个模拟对手。评估规模适中,我们将其趋势视为指示性而非决定性。出现了三个发现。仅仅使用伪匿名化并不改变再识别情况;当准标识符通过概括粗化时,才会出现实质性的隐私收益,可选地结合抑制。随后的分类任务显示,重度抑制的发布保留了多数类效用,但牺牲了大部分少数类召回,这是聚合指标隐藏的成本。最后,标准的基于kmin的PPI与披露界限的相关性仅有些许,与部分知识攻击没有关联,而平均类大小变体PPI与所有三个披露/攻击措施强烈相关;因此,我们将PPI视为面向监管机构的摘要,而不是安全界限。这些配置文件是通过基于规则的DPI passively操作员侧监控生成的;我们的贡献是计算这些指标的隐私量化层,应用转换策略,并通过可检查的仪表板公开这两者。

更新时间: 2026-10-06 08:42:07

领域: cs.CR,cs.NI

下载: http://arxiv.org/abs/2610.07976v1

Learning a Ranking from Human Feedback in Log-Concave Random Utility Models

We study the problem of recovering the ranking of a fixed set of items according to their unknown numerical utilities. At each interaction with the environment, a learner presents the item set to a human and receives comparative feedback of two types. Under full-ranking feedback, each interaction reveals a noisy ranking of all items, whereas under winner-only feedback, it reveals only the item ranked first. In both settings, we model human feedback using a random utility model with log-concave noise and study the number of observations needed to recover an $ε$-accurate ranking with high probability. This novel criterion tolerates ordering errors only between items whose utilities differ by less than $ε$. For both feedback types, we establish worst-case sample-complexity lower bounds and develop algorithms that match these bounds up to logarithmic factors. Neither algorithm requires knowledge of the noise distribution, while only requiring an upper bound on its variance. Our results show that the ranking problem under winner-only feedback is intrinsically harder by exposing the sample complexity dependence on the minimum winning probability across the item set.

Updated: 2026-10-06 08:41:20

标题: 在对数凹随机效用模型中通过人类反馈学习排名

摘要: 我们研究了根据未知数值效用恢复固定物品集的排名的问题。在与环境的每次互动中,学习者向人类展示物品集并接收两种类型的比较反馈。在全排名反馈下,每次互动会显示所有物品的有噪音排名,而在仅赢家反馈下,它只显示排名第一的物品。在这两种设置下,我们使用带对数凹噪音的随机效用模型对人类反馈进行建模,并研究需要多少观察来以高概率恢复ε-准确排名。这一新颖的准则只容忍效用相差小于ε的物品之间的排序错误。对于两种反馈类型,我们建立了最坏情况的样本复杂度下界,并开发了与这些下界相匹配的算法,直到对数因子。这两个算法都不需要知道噪音分布,只需要一个方差上界。我们的结果表明,在仅赢家反馈下的排名问题在本质上更难,通过暴露样本复杂度依赖于物品集中最小获胜概率的情况。

更新时间: 2026-10-06 08:41:20

领域: cs.LG

下载: http://arxiv.org/abs/2610.07973v1

LionMuon: Alternating Spectral and Sign Descent for Efficient Training

Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every $P$ iterations and Lion steps in between, with a single dual-EMA momentum buffer shared by both. Muon's compute and communication are paid once per $P$ steps, and the optimizer state is half of AdamW's. A single-EMA variant, SignMuon, already improves on Muon. We prove complexity bounds under heavy-tailed noise in which the period sets an interpolation between Muon's and Lion's smoothness and noise constants, and which say when LionMuon is faster than both. On 124M and 355M models trained on FineWeb, LionMuon with $P=2$ and $P=5$ reaches a lower loss than Muon, AdamW, Lion and Signum at the same number of tokens. Under 4-GPU data-parallel training it reaches Muon's final loss with a third less wall-clock on PCIe, and it beats the communication-efficient Muon variants Dion and MuonBP on loss at no more exposed communication, while keeping the exact gradient. Code: https://github.com/brain-lab-research/lion-muon

Updated: 2026-10-06 08:41:17

标题: LionMuon:交替谱和符号下降的高效训练

摘要: 预训练语言模型需要大量的计算资源,而正确的优化器可以节省很大一部分。Muon的谱步骤比符号步骤具有更强的方向,但代价高昂。每一步在完整矩阵上运行牛顿-舒尔茨迭代,并在分布式训练中进行额外的全局归约。像Lion和Signum中的符号步骤是廉价的,并且局部运行在每个设备上。我们提出了LionMuon,每$P$次迭代执行一步Muon步骤,之间执行Lion步骤,并使用一个双EMA动量缓冲区共享。Muon的计算和通信每$P$步付费一次,优化器状态是AdamW的一半。单一EMA变体SignMuon已经优于Muon。我们在重尾噪声下证明了复杂性界限,在此期间设置了Muon和Lion之间平滑度和噪声常数的插值,并指出LionMuon何时比两者都更快。在FineWeb上训练的124M和355M模型上,具有$P=2$和$P=5$的LionMuon达到比Muon、AdamW、Lion和Signum更低的损失,具有相同数量的令牌。在4-GPU数据并行训练中,它在PCIe上比Muon的最终损失少三分之一的挂钟时间,并在不暴露更多通信的情况下击败了通信效率高的Muon变体Dion和MuonBP,同时保持准确的梯度。 代码:https://github.com/brain-lab-research/lion-muon

更新时间: 2026-10-06 08:41:17

领域: cs.LG

下载: http://arxiv.org/abs/2609.35297v3

Modeling Time-Dependent Responses of Optical Compressors with Selective State Space Models

This paper presents a method for modeling optical dynamic range compressors using deep neural networks with Selective State Space models. The proposed approach surpasses previous methods based on recurrent layers by employing a Selective State Space block to encode the input audio. It features a refined technique integrating Feature-wise Linear Modulation and Gated Linear Units to adjust the network dynamically, conditioning the compression's attack and release phases according to external parameters. The proposed architecture is well-suited for low-latency and real-time applications, crucial in live audio processing. The method has been validated on the analog optical compressors TubeTech CL 1B and Teletronix LA-2A, which possess distinct characteristics. Evaluation is performed using quantitative metrics and subjective listening tests, comparing the proposed method with other state-of-the-art models. Results show that our black-box modeling methods outperform all others, achieving accurate emulation of the compression process for both seen and unseen settings during training. We further show a correlation between this accuracy and the sampling density of the control parameters in the dataset and identify settings with fast attack and slow release as the most challenging to emulate.

Updated: 2026-10-06 08:39:07

标题: 用选择性状态空间模型对光压缩器的时间相关响应进行建模

摘要: 本文提出了一种使用深度神经网络和选择性状态空间模型对光学动态范围压缩器进行建模的方法。所提出的方法通过采用选择性状态空间块来编码输入音频,超越了基于循环层的先前方法。它采用了一种精细的技术,集成了特征逐渐线性调制和门控线性单元,动态调整网络,根据外部参数调整压缩的攻击和释放阶段。所提出的架构非常适用于低延迟和实时应用,在现场音频处理中至关重要。该方法已在拥有不同特征的模拟光学压缩器TubeTech CL 1B和Teletronix LA-2A上进行了验证。评估是使用定量指标和主观听测试进行的,将所提出的方法与其他最先进的模型进行了比较。结果表明,我们的黑盒建模方法胜过所有其他方法,在训练过程中看到和未看到的设置下都能准确模拟压缩过程。我们进一步展示了数据集中控制参数的采样密度与准确性之间的相关性,并确定了具有快速攻击和缓慢释放设置是最具挑战性的。

更新时间: 2026-10-06 08:39:07

领域: cs.SD,cs.AI,eess.AS

下载: http://arxiv.org/abs/2408.12549v4

Can Agents Work for Everyone? Cross-User Reliability for Mobile GUI Agents in Personalized User Interfaces

Mobile GUI agents increasingly operate on interfaces influenced by users' histories and preferences, but their reliability across different users remains underexplored. We introduce PAIR (Personalized Application-state Instantiation and Rendering), a pipeline for constructing user-conditioned application states that enables controlled evaluation of the same task across different users. We further introduce RePAIR (Reinforcement learning with Personalization-Aware Interaction Rewards), a training approach that learns from cross-user differences in subgoal outcomes to improve reliability across user-conditioned mobile environments. Across six agents, we find substantial variation in task success across users and consistently lower subgoal achievement in user-conditioned UI contexts (6.98 to 15.4 pp). This gap further increases for personal targets drawn from each user's own content (8.77 to 22.0 pp). Failures in these contexts frequently involve selecting another item instead of the intended target, particularly before target exposure. Finally, RePAIR improves user-conditioned SAR (+5.87 pp), all-success (+7.50 pp), and overall Task SR (+9.42 pp) over its supervised fine-tuning parent on unseen users, providing initial evidence that explicitly learning from cross-user variation can improve GUI-agent reliability.

Updated: 2026-10-06 08:38:30

标题: 代理人能为所有人工作吗?个性化用户界面中移动GUI代理人的跨用户可靠性

摘要: 移动GUI代理越来越多地在受用户历史和偏好影响的界面上运行,但它们在不同用户之间的可靠性仍然未被充分探讨。我们介绍了PAIR(个性化应用状态实例化和呈现),这是一个用于构建用户条件应用状态的管道,能够实现在不同用户之间对相同任务的控制评估。我们进一步介绍了RePAIR(具有个性化感知交互奖励的强化学习),这是一种训练方法,通过学习跨用户之间的子目标结果差异来改善在用户条件移动环境中的可靠性。在六个代理中,我们发现不同用户之间任务成功率存在显著差异,并且在用户条件UI环境下子目标实现率普遍较低(6.98至15.4个百分点)。对于从每个用户自己的内容中提取的个人目标,这种差距进一步增加(8.77至22.0个百分点)。在这些环境中的失败经常涉及选择另一个项目而非预期目标,尤其是在目标曝光之前。最后,RePAIR在未见用户上提高了用户条件SAR(+5.87个百分点)、全成功率(+7.50个百分点)和整体任务SR(+9.42个百分点),这为明确从跨用户变化中学习可以改善GUI代理的可靠性提供了初步证据。

更新时间: 2026-10-06 08:38:30

领域: cs.AI

下载: http://arxiv.org/abs/2610.07972v1

To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three families show high call accuracy but much lower no-call accuracy, leaving overall accuracy in the 55%-70% range. We trace this to an Intrinsic Bias Hypothesis (IBH): the call/no-call decision mapping carries an activation-independent call offset, so the model favors call even at activation parity. Using Sparse Autoencoders (SAEs), we recover behavior-aligned feature bases for the call/no_call decision, reduce them to a signed activation margin, and estimate the offset directly. Across all six models, the model is decision-neutral only when no_call activation outweighs call activation, consistent with IBH. We then causally test IBH with Adaptive Margin-Calibrated Steering (AMCS), a closed-form counter-bias shift along SAE decoder directions. Cancelling the diagnosed offset mitigates over-calling and improves overall accuracy with a negligible drop in call accuracy. Our work recasts over-calling from an empirical phenomenon into a mechanistic object amenable to causal correction. Code is available at https://github.com/SKURA502/agent-sae/.

Updated: 2026-10-06 08:37:14

标题: 是打电话还是不打电话:诊断LLM代理的固有过度打电话偏差

摘要: LLM代理程序表现出一种一贯的过度呼叫倾向,即使在不需要工具的情况下也会调用工具。在When2Call基准测试中,来自三个家族的六种模型显示出较高的呼叫准确性,但无呼叫准确性较低,使整体准确度在55%-70%范围内。我们将这归因于内在偏见假设(IBH):呼叫/不呼叫决策映射具有独立于激活的呼叫偏移,因此模型在激活平等时更倾向于呼叫。使用稀疏自动编码器(SAEs),我们恢复了与行为对齐的呼叫/不呼叫决策的特征基础,将它们减少到有符号的激活边缘,并直接估计偏移量。在所有六个模型中,当无呼叫激活超过呼叫激活时,模型是决策中立的,与IBH一致。然后,我们使用自适应边缘校准转向(AMCS)对IBH进行因果测试,这是一种沿着SAE解码器方向进行的封闭形式反偏差位移。取消诊断的偏移量可以减轻过度呼叫,并在呼叫准确性略微下降的情况下提高整体准确度。我们的工作将过度呼叫从经验现象重新构建为一种可以接受因果修正的机械对象。代码可在https://github.com/SKURA502/agent-sae/上找到。

更新时间: 2026-10-06 08:37:14

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2605.18882v2

FHRFormer: A Self-Supervised Masked Transformer Framework for Fetal Heart Rate Time-Series Inpainting and Forecasting

Approximately 10% of newborns require assistance to initiate breathing at birth, and around 5% need ventilation support. Fetal heart rate (FHR) monitoring plays a crucial role in assessing fetal well-being during prenatal care, enabling the detection of abnormal patterns and supporting timely obstetric interventions to mitigate fetal risks during labor. Applying artificial intelligence (AI) methods to analyze large datasets of continuous FHR monitoring episodes with diverse outcomes may offer novel insights into predicting the risk of needing breathing assistance or interventions. Recent advances in wearable FHR monitors have enabled continuous fetal monitoring without compromising maternal mobility. However, sensor displacement during maternal movement, as well as changes in fetal or maternal position, often lead to signal dropout, resulting in gaps in recorded FHR data. Such missing data limits the extraction of meaningful insights and complicates automated (AI-based) analysis. Traditional approaches to handling missing data, such as simple interpolation techniques, often fail to preserve the spectral characteristics of the signals. In this paper, we propose a masked transformer-based autoencoder approach to reconstruct missing FHR signals by capturing both local temporal and frequency components of the data. The proposed method demonstrates robustness across varying durations of missing data and can be used for signal inpainting and forecasting. The proposed approach can be applied retrospectively to research datasets to support the development of AI-based risk algorithms. In the future, the proposed method could be integrated into wearable FHR monitoring devices to achieve earlier and more robust risk detection.

Updated: 2026-10-06 08:34:15

标题: FHRFormer:一种用于胎心率时间序列修复和预测的自监督遮罩变换器框架

摘要: 大约有10%的新生儿需要在出生时接受呼吸启动辅助,约5%需要呼吸机支持。胎心率(FHR)监测在产前护理中对评估胎儿健康起着至关重要的作用,能够发现异常模式并支持及时的产科干预以减少分娩期间胎儿风险。将人工智能(AI)方法应用于分析大量连续FHR监测数据集可能为预测需要呼吸辅助或干预的风险提供新的见解。近年来,可穿戴式FHR监测器的发展使得连续胎儿监测不会影响母体的活动能力。然而,母体运动过程中传感器移位以及胎儿或母体位置的改变经常导致信号丢失,导致FHR数据记录中的间断。这些缺失数据限制了有意义见解的提取,并使自动(基于AI的)分析变得复杂。传统处理缺失数据的方法,如简单的插值技术,通常无法保留信号的频谱特性。在本文中,我们提出了一种基于掩码变换器的自编码器方法,通过捕捉数据的局部时间和频率成分来重建缺失的FHR信号。所提出的方法在不同持续时间的缺失数据下表现出鲁棒性,并可用于信号修补和预测。这种方法可以被回顾性地应用于研究数据集,以支持基于AI的风险算法的发展。将来,这种方法可以被整合到可穿戴式FHR监测设备中,以实现更早、更稳健的风险检测。

更新时间: 2026-10-06 08:34:15

领域: cs.AI,cs.CE,cs.LG,math.PR

下载: http://arxiv.org/abs/2605.29695v2

DecepEval: A Benchmark for Evaluating Deception in LLM Agents

As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict. DecepEval pairs neutral and induced versions of each instance to measure condition-dependent changes in deception rates, while explicit task facts and observable agent behavior help distinguish deception from capability-related errors. Evaluations of nine frontier LLMs show that inducements increase deception across models and task families, even among models with low baseline deception rates. DecepEval makes these vulnerabilities measurable, providing a shared benchmark for progress toward trustworthy artificial intelligence.

Updated: 2026-10-06 08:33:43

标题: DecepEval:评估LLM代理中欺骗的基准

摘要: 随着大型语言模型(LLM)代理变得越来越自主,它们可能通过欺骗来追求任务表现,引发人们对它们可靠部署的担忧。现有评估表明,LLM代理可以欺骗,但通常仅考虑孤立的场景或狭义条件,限制了对欺骗何时变得更有可能的系统理解。为了填补这一空白,我们引入了DecepEval,一个基准测试,包括3个任务系列和28个专业场景共1,532个实例。借鉴经典欺诈理论,我们提出了LLM欺骗金字塔框架,该框架描述了可能诱发欺骗的四个外部条件:压力、激励、机会和冲突。DecepEval将每个实例的中立和诱导版本配对,以测量对条件依赖性的欺骗率变化,同时明确任务事实和可观察的代理行为有助于区分欺骗和能力相关错误。对九个前沿LLM的评估显示,诱导增加了模型和任务系列中的欺骗,即使是基线欺骗率较低的模型也是如此。DecepEval使这些漏洞可测量,为朝着可信赖的人工智能的进展提供了一个共享的基准测试。

更新时间: 2026-10-06 08:33:43

领域: cs.LG

下载: http://arxiv.org/abs/2610.07967v1

Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution

Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across features |\r{ho}|=0.25, 2.8x the null). Natural audio sees a stronger encoding when nonlinear probes are used (mean across features R2=0.56, 18x the null, +0.152 nonlinear gain over the linear probe R2). Encoding strength varies throughout the layers of the decoder and an increased ability to joint-encode in the middle layers is seen across all audio features (\b{eta}2 all negative, p < 0.05). The general purpose EnCodec decoder also sees similar strong synthetic responses across audio features, similar nonlinear gains for natural audio joint encoding and similar depth profiles. We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding. These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.

Updated: 2026-10-06 08:32:02

标题: 变分自动编码器(VAE)音频解码器中的特征编码:输入、深度和分布的影响

摘要: 神经音频合成模型如实时音频变分自动编码器(RAVE)实现了令人印象深刻的生成质量,然而它们内部表征如何编码音乐特征仍然不太清楚。我们对在不同音乐领域训练的三个模型上的RAVE解码器激活进行了系统的逐层和跨层聚类分析,并使用四种刺激类型进行了测试。然后,我们使用通用的EnCodec模型评估了架构的泛化性。对于RAVE,我们发现合成刺激在各模型和音频特征(音高|\r{ho}|=0.45,空值的5.1倍,BPM |\r{ho}|=0.76,空值的8.6倍)中被很好地编码。当使用自然音频时,这些结果减少但仍然明显(各特征的平均值|\r{ho}|=0.25,空值的2.8倍)。当使用非线性探针时,自然音频的编码强度更强(各特征的平均值R2=0.56,空值的18倍,比线性探测器R2增加了+0.152的非线性增益)。编码强度在解码器的各层中变化,并且在中间层中看到了更强的联合编码能力(所有音频特征的\b{eta}2都为负,p < 0.05)。通用EnCodec解码器在各音频特征上也看到了类似强烈的合成响应,自然音频联合编码的非线性增益和深度轮廓也类似。我们发现最佳的跨层聚类能够提高BPM编码的强度(r = 0.65,p = 0.006)和普遍性(r = 0.75,p = 0.001),与同一部分内最佳整层相比,联合编码没有效果。这些发现推进了神经音频模型的可解释性,并为神经合成提供了有针对性的控制策略。

更新时间: 2026-10-06 08:32:02

领域: cs.SD,cs.LG

下载: http://arxiv.org/abs/2610.07966v1

ReGraph: A Computational Account of Emergent Generalization in the "what" and "where" Dual Visual Streams

Where generalization capacity--the ability to extract context-invariant relational structures--first emerges remains a central question in AI and neuroscience. The foundation for this capacity lies upstream of the hippocampus, within the entorhinal cortex, where parallel pathways dissociate relational structure in the medial entorhinal cortex (MEC) from sensory content in the lateral entorhinal cortex. However, as Eichenbaum argued, such factorization likely originates earlier, driven by the segregation of the dorsal ('where') and ventral ('what') visual streams. Supporting this, grid-like firing patterns--a signature of MEC (context-invariant codes)--also appear in preceding neocortical regions along the dorsal pathway. Yet, how such representations are computationally formed along upstream pathways remains unknown. To investigate this in silico, we developed ReGraph, a recurrent dual-stream graph model with biological inductive biases, including retina-driven stream-specialized encoding, dorsal-to-ventral modulation, and dynamic lateral connectivity. Trained on the action benchmark Something-Something V2, ReGraph revealed a pathway-specific emergence of relational mapping: context-invariant codes and grid-like spatial bases uniquely co-emerged along the extended dorsal stream. In contrast, their absence in single-stream, unmodulated variants, and standard baselines implies that these inductive biases are prerequisites for relational structures. Crucially, our post-hoc analyses demonstrated that these grid-like bases serve as reusable routing templates for information processing via lateral connectivity. Together, our findings provide a computational account that generalization may not be a faculty that emerges abruptly within a dedicated region, but a property that already takes shape as sensory information is parsed into factorized streams of hierarchical visual processing.

Updated: 2026-10-06 08:30:45

标题: ReGraph:对“what”和“where”双视觉流中新出现的泛化现象的计算解释

摘要: 泛化能力——提取与上下文无关的关系结构的能力——首次出现的地方仍然是人工智能和神经科学中的一个核心问题。这种能力的基础位于海马区的上游,在海马回内部,平行通路将内嗅皮层中的关系结构(MEC)与外嗅皮层中的感觉内容分离开来。然而,正如埃森鲍姆所认为的那样,这种因素化可能起源于更早的时间,由背侧(“在哪里”)和腹侧(“什么”)视觉通路的分离驱动。支持这一观点的是,类似于网格的放电模式——MEC的特征标记(与上下文无关的代码)——也出现在沿着背侧通路的前额叶皮层区域中。然而,这种表示是如何在上游通路中计算形成的,目前尚不清楚。为了在计算机模拟中研究这一问题,我们开发了ReGraph,这是一个具有生物归纳偏差的经常性双流图模型,包括视网膜驱动的流专门编码、从背侧到腹侧的调制以及动态侧向连接。在Something-Something V2的行动基准上进行训练,ReGraph显示了关系映射的特定路径出现:上下文无关的代码和类似网格的空间基础独特地在扩展的背侧通路中同时出现。相比之下,在单一通路、未调制的变体和标准基线中它们的缺失意味着这些归纳偏差是关系结构的先决条件。至关重要的是,我们的事后分析表明,这些类似网格的基础作为通过侧向连接进行信息处理的可重复使用的路由模板。总之,我们的发现提供了一个计算说明,即泛化能力可能并不是突然出现在一个专门区域内的一种能力,而是一种在感觉信息被解析为分层视觉处理的因子化流中已经形成的属性。

更新时间: 2026-10-06 08:30:45

领域: cs.NE,cs.AI,q-bio.NC

下载: http://arxiv.org/abs/2610.07962v1

SDFlow: Similarity-Driven Flow Matching for Time Series Generation

Vector quantization (VQ) with autoregressive (AR) token modeling is a widely adopted and highly competitive paradigm for time-series generation. However, such models are fundamentally limited by exposure bias: during inference, errors can accumulate across sequential predictions, leading to pronounced quality degradation in long-horizon generation. To address this, we propose SDFlow ($\textbf{S}$imilarity-$\textbf{D}$riven $\textbf{Flow}$ Matching), a non-autoregressive framework that operates entirely in the frozen VQ latent space and enables parallel sequence generation via flow matching. We tackle three key challenges in making this transition: (1) eliminating exposure bias by replacing step-wise token prediction with a global transport map; (2) mitigating the high-dimensionality of VQ token spaces via a low-rank manifold decomposition with a learned anchor prior over the latent manifold; and (3) incorporating discrete supervision into continuous transport dynamics by introducing a categorical posterior over codebook indices within a variational flow-matching formulation. Extensive experiments show that SDFlow achieves state-of-the-art performance, improving Discriminative Score and substantially reducing Context-FID, particularly for challenging long-sequence generation. Moreover, SDFlow provides significant inference speedups over autoregressive baselines, offering both high fidelity and computational efficiency. Code is available at https://github.com/William-Liwei/SDFlow

Updated: 2026-10-06 08:27:32

标题: SDFlow:基于相似性驱动的流匹配用于时间序列生成

摘要: 矢量量化(VQ)与自回归(AR)令牌建模是时间序列生成的一个被广泛采纳和高度竞争的范式。然而,这样的模型在本质上受到曝光偏差的限制:在推断过程中,错误可能会在连续预测中积累,导致长期生成中显著的质量下降。为了解决这个问题,我们提出了SDFlow(相似性驱动流匹配),这是一个非自回归框架,完全在冻结的VQ潜在空间中运行,并通过流匹配实现并行序列生成。我们解决了这一转变中的三个关键挑战:(1)通过用全局传输映射替换逐步令牌预测来消除曝光偏差;(2)通过学习潜在流形上的锚先验进行低秩流形分解,以减轻VQ令牌空间的高维度;(3)通过在变分流匹配公式中引入一个对码本索引的分类后验,将离散监督融入连续传输动力学中。广泛的实验表明,SDFlow实现了最先进的性能,改善了判别分数,并显著减少了上下文-FID,特别是对于具有挑战性的长序列生成。此外,SDFlow相比自回归基线提供了显著的推断加速,既具有高保真度又具有计算效率。代码可在https://github.com/William-Liwei/SDFlow找到。

更新时间: 2026-10-06 08:27:32

领域: cs.AI

下载: http://arxiv.org/abs/2605.05736v3

Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents

When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent's success. Confidence estimation for agents is difficult because evidence about success is distributed across heterogeneous, interdependent steps of an agent's trajectory. Practical agentic deployments introduce further challenges: frontier LLMs often provide limited access to internal signals, agent roll-outs are costly, and training data may be unavailable or quickly become outdated. To address these challenges, we introduce Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an agent accomplished its task from a single trajectory, without privileged model access or training data. Rather than compressing an execution into a single holistic judgment, a CRG begins with the claim that the agent accomplished its task, decomposes it into contextualized sub-claims grounded in trajectory evidence, estimates confidence for each terminal claim, and finally aggregates these into an overall confidence estimate. Across three agentic benchmarks, three backbone models, and three agent frameworks, CRGs yield better-calibrated confidence and stronger risk-aware decision making than verbalized, sampling-based, and white-box surrogate baselines. We further find that calibration error alone can be misleading: a white-box surrogate baseline appears well calibrated while providing near-chance discrimination. Ablations attribute CRG's improvements to claim-level confidence estimation and aggregation rather than graph construction alone. Finally, a CRG exposes the claims and trajectory evidence underlying each confidence estimate, enabling it to be audited at decision time.

Updated: 2026-10-06 08:22:14

标题: 信心推理图:针对LLM代理的结构化信心估计

摘要: 在使用LLM代理在一个相关领域时,做出关于是否信任其输出或干预的决定需要对代理的成功有校准的信心。代理的信心估计很困难,因为关于成功的证据分布在代理轨迹的异质、相互依赖的步骤之间。实际的代理部署引入了进一步的挑战:前沿的LLM通常提供有限的内部信号访问,代理展开成本高昂,训练数据可能不可用或很快变得过时。为了解决这些挑战,我们引入了置信推理图(CRGs),这是一个推断时的框架,从单个轨迹估计代理完成任务的概率,而无需特权模型访问或训练数据。CRG不是将执行压缩为单一的整体判断,而是从代理完成任务的主张开始,将其分解为基于轨迹证据的上下文化子主张,为每个终端主张估计信心,最后将这些聚合成一个总体信心估计。在三个代理基准测试、三个骨干模型和三个代理框架中,CRGs比口头化、基于抽样和白盒替代基线提供更好校准的信心和更强的风险感知决策制定。我们进一步发现,仅仅校准误差可能会误导:一个白盒替代基线看起来校准良好,却提供接近机会性辨别。消蚀将CRG的改进归因于主张级信心估计和聚合,而不仅仅是图构建。最后,CRG暴露了每个信心估计背后的主张和轨迹证据,使其在决策时可以进行审计。

更新时间: 2026-10-06 08:22:14

领域: cs.AI,cs.CL

下载: http://arxiv.org/abs/2610.07948v1

Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution

Visual disruptions can arise while a robot is executing a task, leaving a vision-language-action (VLA) policy to respond without knowing the disruption type or timing. We introduce Self-supervised Adaptation from Leftover Trajectories (SALT), which uses the leftover trajectory, the unexecuted part of the previous action chunk, as self-supervision for test-time adaptation. Because consecutive chunks overlap in time, the leftover provides a temporally aligned target for the current prediction over the same future control interval. At the onset of a visual shift, the leftover can retain a plan formed before the corruption, so updating the policy toward it anchors the adaptation across the shift (Transition Anchoring). SALT keeps the adapted policy and regenerates the current chunk, whose leftover becomes the target at the next replan, carrying the correction forward along the execution trajectory (Sequential Correction Propagation). Supervision comes entirely from the policy's own predictions, requiring no disruption annotations, expert actions, or target-domain demonstrations, and a lightweight adaptation gate calibrated only on nominal trajectories decides when updates begin. On LIBERO-10, SALT increases average success across five persistent visual corruptions from 43.9% to 53.2% with SmolVLA and from 58.7% to 66.0% with GR00T N1.7, while largely preserving nominal performance. On a real robot, it raises task progress averaged over digital and physical disruptions from 0.49 to 0.61.

Updated: 2026-10-06 08:21:19

标题: 将视觉-语言-动作模型调整为执行过程中未知的视觉干扰

摘要: 视觉中断可能在机器人执行任务时出现,使得视觉-语言-动作(VLA)策略在不知道中断类型或时机的情况下作出响应。我们引入了一种称为Leftover Trajectories的自监督适应方法(SALT),该方法使用剩余轨迹,即先前动作块的未执行部分,作为测试时适应的自监督。由于连续的动作块在时间上重叠,剩余部分为当前预测提供了与同一未来控制间隔的时间对齐目标。在视觉转移开始时,剩余部分可以保留在污染之前形成的计划,因此将策略朝向这个计划进行更新可以在转移过程中锚定适应(过渡锚定)。SALT保持了适应的策略并重新生成当前的动作块,其剩余部分成为下一次重新规划时的目标,沿着执行轨迹传递修正(顺序校正传播)。监督完全来自策略自身的预测,无需中断注释、专家操作或目标领域演示,并且仅根据正常轨迹的轻量级适应门决定何时开始更新。在LIBERO-10上,SALT将对五种持久的视觉扰动的平均成功率从43.9%提高到了53.2%,同时使用SmolVLA将成功率从58.7%提高到了66.0%。在真实机器人上,它将数字和物理中断的任务进展平均从0.49提高到了0.61,同时大部分保留了正常性能。

更新时间: 2026-10-06 08:21:19

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2610.07946v1

Mean-based algorithms: A lower bound and regret

Mean-based algorithms are online learning algorithms that assign low probability to actions with low average rewards. Recent research shows that they converge to serially undominated actions, which serve as approximations to Nash equilibria in economic games. However, empirical studies indicate that mean-based algorithms converge more slowly in bandit-feedback settings than established no-regret alternatives. This work investigates mean-based algorithms under unknown horizons and bandit feedback. In this setting, we provide the first lower bound on the algorithm-defining sequence $γ_t$, establishing a fundamental limit on the learning speed of such algorithms. In multi-armed bandit problems, this result constrains the rate at which any algorithm can reliably identify low-reward actions while acting according to this knowledge. We also propose two mean-based algorithms: one generalizes $ε$-greedy, and the other extends mean-based Exp3 to unknown horizons. Our experiments show that mean-based algorithms, although slightly slower, can perform competitively with other bandit-feedback algorithms. We further study the relationship to regret. Depending on the choice of $γ_t$, the intersection with no-regret algorithms is non-trivial, and we show that some algorithms are both mean-based and no-regret.

Updated: 2026-10-06 08:18:25

标题: 基于均值的算法:一个下界和遗憾

摘要: Mean-based algorithms are online learning algorithms that assign low probability to actions with low average rewards. Recent research has shown that these algorithms converge to serially undominated actions, which approximate Nash equilibria in economic games. However, empirical studies have indicated that mean-based algorithms converge more slowly in bandit-feedback settings compared to established no-regret alternatives. This study examines mean-based algorithms in the context of unknown horizons and bandit feedback. A lower bound on the algorithm-defining sequence γt is established, setting a fundamental limit on the learning speed of such algorithms. In multi-armed bandit problems, this result limits the rate at which any algorithm can reliably identify low-reward actions while acting based on this knowledge. Two mean-based algorithms are proposed: one generalizes ε-greedy, while the other extends mean-based Exp3 to unknown horizons. Experimental results demonstrate that mean-based algorithms, although slightly slower, can compete effectively with other bandit-feedback algorithms. The relationship to regret is also explored. Depending on the choice of γt, the intersection with no-regret algorithms is non-trivial, with some algorithms being both mean-based and no-regret.

更新时间: 2026-10-06 08:18:25

领域: cs.LG,cs.GT

下载: http://arxiv.org/abs/2606.04931v2

Diffusion Model-Based Video Editing: A Survey

The rapid development of diffusion models (DMs) has significantly advanced image and video applications, making "what you want is what you see" a reality. Among these, video editing has gained substantial attention and seen a swift rise in research activity, necessitating a comprehensive and systematic review of the existing literature. This paper reviews diffusion model-based video editing techniques, including theoretical foundations and practical applications. We begin by overviewing the mathematical formulation and image domain's key methods. Subsequently, we categorize video editing approaches by the inherent connections of their core technologies, depicting evolutionary trajectory. This paper also dives into novel applications, including point-based editing and pose-guided human video editing. Additionally, we present a comprehensive comparison using our newly introduced V2VBench. Building on the progress achieved to date, the paper concludes with ongoing challenges and potential directions for future research.

Updated: 2026-10-06 08:18:15

标题: 基于扩散模型的视频编辑:一项调查

摘要: 扩散模型(DMs)的快速发展显著推动了图像和视频应用的发展,使“你想要的就是你看到的”成为现实。其中,视频编辑引起了广泛关注,并在研究活动中迅速崛起,这需要对现有文献进行全面系统的回顾。本文回顾了基于扩散模型的视频编辑技术,包括理论基础和实际应用。我们首先概述了数学公式和图像领域的关键方法。随后,通过它们核心技术的内在联系对视频编辑方法进行分类,描述了其演变轨迹。本文还深入探讨了新颖的应用,包括基于点的编辑和姿势引导的人体视频编辑。此外,我们使用我们新推出的V2VBench进行了全面比较。基于迄今取得的进展,本文总结了当前面临的挑战以及未来研究的潜在方向。

更新时间: 2026-10-06 08:18:15

领域: cs.CV,cs.AI,cs.LG,cs.MM

下载: http://arxiv.org/abs/2407.07111v2

Enhancing Distance-Based Graph Autoencoders with Structural Penalties for Dynamic Graph Embedding

Graph autoencoders (GAEs) are widely used for learning representations of dynamic graphs. However, their optimisation objectives typically do not take structural heterogeneity across nodes into account. We propose three distance-based GAE variants that incorporate structural penalties into the reconstruction loss. All variants share a two-layer Graph Convolutional Network encoder and a Euclidean-distance decoder trained with distance-based reconstruction objectives. We extend sparsity-corrected loss with two node-level regularization terms: (i) a hub penalty based on degree centrality, and (ii) a penalty based on Natural Community Local Intrinsic Dimensionality (NC-LID). The paper is motivated by prior evidence linking high NC-LID to reduced embedding quality. The proposed methods are designed to emphasize reconstruction errors for structurally ambiguous nodes. Experiments on multiple dynamic graph data sets show that incorporating NC-LID-based regularization consistently improves reconstruction performance over the baseline without structural regularization and the method using hub-aware regularization. These findings highlight NC-LID as a useful structural signal for enhancing distance-based graph autoencoders in dynamic settings.

Updated: 2026-10-06 08:14:37

标题: 使用结构惩罚增强基于距离的图自编码器,用于动态图嵌入

摘要: 图自编码器(GAEs)广泛用于学习动态图的表示。然而,它们的优化目标通常不考虑节点之间的结构异质性。我们提出了三种基于距离的GAE变体,将结构惩罚纳入重构损失中。所有变体共享一个两层图卷积网络编码器和一个通过基于距离的重构目标训练的欧几里得距离解码器。我们通过两个节点级正则化项扩展了稀疏校正损失:(i)基于度中心性的中心惩罚,(ii)基于自然社区局部内在维度(NC-LID)的惩罚。该论文受到先前证据的启发,该证据表明高NC-LID与降低嵌入质量有关。所提出的方法旨在强调对结构模糊节点的重构错误。对多个动态图数据集的实验表明,将基于NC-LID的正则化纳入其中始终可以提高重构性能,相比没有结构正则化的基准和使用中心感知正则化的方法。这些发现强调了NC-LID作为一种有用的结构信号,可以增强动态环境中基于距离的图自编码器。

更新时间: 2026-10-06 08:14:37

领域: cs.LG,cs.ET

下载: http://arxiv.org/abs/2608.18762v2

Hybrid Latent Attention for Looped Language Models

Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shrinks by 10.7x per token, fitting 4.0-8.8x as many concurrent sequences per GPU, and decoding throughput improves by 2.5x at 1K-token contexts and by up to 7.4x at 16K. HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens. After supervised fine-tuning, it performs on par with the fine-tuned original model on competition-level math.

Updated: 2026-10-06 08:14:07

标题: 混合潜在关注力用于循环语言模型

摘要: 循环语言模型将相同的层堆栈T次应用于每个令牌,这样可以加深模型而不增加参数,但将其关键-值(KV)缓存乘以T。较大的缓存限制了GPU一次解码多少个序列,并减慢了每个解码步骤,因为需要读取整个缓存。我们提出了混合潜在注意(HLA),它在最近W个标记的滑动窗口中保留确切的键和值,并将每个较旧的标记存储为紧凑的潜在值,每个循环的查询直接读取,而无需重建键和值。我们在Ouro循环模型(T=4)上训练HLA,具有14亿和26亿参数,保持预训练权重冻结,并仅训练添加参数以重现原始注意力。每个令牌的缓存缩小了10.7倍,每个GPU可容纳4.0-8.8倍的并发序列,解码吞吐量在1K令牌上提高了2.5倍,在16K时最多提高了7.4倍。HLA在数学、知识和推理基准测试中保留了原始准确率的97%以上,长上下文检索高达16K个令牌时为96-100%。经过监督微调后,它在竞赛级数学方面与经过微调的原始模型表现相当。

更新时间: 2026-10-06 08:14:07

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.07940v1

Generalized Matheron Variational Implicit Processes

Implicit-process priors specify distributions over functions through sample-forward mechanisms such as Bayesian neural networks and stochastic simulators, but their function-space densities are typically unavailable. We introduce Generalized Matheron Variational Implicit Processes (GMVIP), a pathwise variational family for posterior inference with such priors. For Gaussian-process priors, GMVIP recovers the standard inducing-variable variational GP construction; for general implicit priors, its empirical covariance construction preserves the prior mean and covariance in the population limit. GMVIP constructs posterior samples by drawing a function from the prior and applying a correction anchored at a set of inducing inputs. The effect of this correction away from the inducing inputs is determined directly from prior samples, allowing the posterior to retain the structure and variability of the original implicit process. The (surrogate) prior and variational posterior use the same pathwise construction and differ only in the distribution of whitened inducing coefficients, yielding a tractable coefficient-space Kullback-Leibler divergence. Experiments on regression, classification, and forecasting with simulator-defined and retrieval-conditioned empirical trajectory priors show that GMVIP is broadly competitive with existing methods.

Updated: 2026-10-06 08:11:49

标题: Generalized Matheron变分隐式过程

摘要: 隐式过程先验通过贝叶斯神经网络和随机模拟器等样本向前机制来指定函数分布,但它们的函数空间密度通常是不可用的。我们引入了广义马瑟隐式变分过程(GMVIP),这是一种适用于具有这种先验的后验推断的路径变分族。对于高斯过程先验,GMVIP可以恢复标准的感应变量变分GP构造;对于一般的隐式先验,其经验协方差构造在人口极限下保持先验均值和协方差。GMVIP通过从先验中抽取一个函数并在一组感应输入处应用一个校正来构造后验样本。这个校正远离感应输入的效果直接由先验样本决定,使后验能够保留原始隐式过程的结构和变异性。(替代)先验和变分后验使用相同的路径构造,只在白化感应系数的分布上有所不同,从而产生可处理的系数空间Kullback-Leibler散度。对于由模拟器定义和检索条件确定的经验轨迹先验进行的回归、分类和预测实验表明,GMVIP在广泛的情况下与现有方法具有竞争力。

更新时间: 2026-10-06 08:11:49

领域: cs.LG

下载: http://arxiv.org/abs/2610.07938v1

SIGMA: Self-Improving Alignment Generalization from a Model Spec

LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics. Since alignment is much harder to verify, this creates a growing risk of capabilities increasing without appropriate safety alignment, especially as capabilities expand to auto-research and cybersecurity. Existing approaches focus on capability self-improvement using verifiable feedback or on alignment training with supervision from stronger models or curated data, creating an external supervision bottleneck for alignment. We ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings. Given only a "Model Spec" stating the model's desired behavior, SIGMA leverages a model's reasoning capabilities to strengthen its own safety reasoning. SIGMA first performs spec-guided task synthesis, using the candidate model as a task designer agent to generate diverse alignment dilemma scenarios and convert them into training tasks that stress-test its understanding of the Model Spec. Next, SIGMA conducts self-judged alignment training through supervised fine-tuning and rubric-based reinforcement learning with the model itself as the reward model. Despite training only on single-turn chat data, SIGMA improves safety alignment in multi-turn agentic environments (AgentHarm harmfulness decreases from 22.6 to 14.8; Agentic Misalignment decreases from 79.1 to 3.8), outperforms Deliberative Alignment and Constitutional AI baselines, and retains general capability. Analyses show that a Model Spec balancing harmlessness and helpfulness, test-time reasoning for safety deliberation, and high-quality rubrics from SIGMA's task designer agent are crucial for effective self-improvement.

Updated: 2026-10-06 08:09:56

标题: SIGMA:来自模型规范的自我改进对齐泛化

摘要: LLM代理越来越能够执行复杂任务,并在易于验证的目标(如软件工程和数学)上递归地提升自己。由于对齐性要难以验证,这导致能力增长而没有适当的安全对齐,尤其是当能力扩展到自动研究和网络安全时。现有方法侧重于使用可验证的反馈进行能力自我提升,或者在更强的模型或策划数据的监督下进行对齐训练,从而为对齐性创造了外部监督瓶颈。我们探讨当前模型是否可以改进自身的安全对齐性,并提出了SIGMA,这是一个数据生成和训练管道,可以实现对齐性自我提升,可以泛化到分布之外的设置。只需一个“模型规范”来说明模型的期望行为,SIGMA利用模型的推理能力来加强自身的安全推理。SIGMA首先执行规范引导的任务合成,使用候选模型作为任务设计代理来生成各种对齐困境场景,并将其转换为训练任务,以测试其对模型规范的理解。接下来,SIGMA通过监督微调和基于规则的强化学习进行自我判断的对齐训练,模型本身作为奖励模型。尽管只在单轮聊天数据上进行训练,SIGMA在多轮代理环境中改进了安全对齐性(AgentHarm的有害性从22.6降至14.8;代理对齐性从79.1降至3.8),优于深思对齐和宪法AI基线,并保持了一般能力。分析表明,平衡无害性和帮助性的模型规范、测试时用于安全思考的推理、以及SIGMA任务设计代理的高质量规则对于有效的自我提升至关重要。

更新时间: 2026-10-06 08:09:56

领域: cs.AI

下载: http://arxiv.org/abs/2610.07935v1

Where does a rust speedup come from? Language and algorithm effects in sliding window threat scorer

Rewriting a hot path from Python into Rust is a common way to speed up security analytics, and large speedups are routinely reported. A rewrite usually changes the language and the algorithm at once, so a single factor can credit the language with a gain that comes from a better algorithm. We study this on a sliding window threat scorer modelled on the traffic light calculator of the SentinelSphere platform. Five implementations, three in Python and two in Rust, produce bit identical scores, confirmed by a shared checksum, and we time them from 100 to one million events in a bounded and a burst regime. The language alone contributes between about 4 and 29 times. Replacing a per event rescan of the window by an incremental update contributes more than 11,000 times in the burst regime, so the end to end factor reaches about 195,000 times when the window keeps filling and levels off near 13,000 times when it does not. A power law fit to the timings reported for the original rewrite gives growth exponents of 1.85 for Python and 0.88 for Rust, the signature of an algorithmic difference. Performance claims for security tooling should therefore report the language and algorithm contributions separately.

Updated: 2026-10-06 08:08:49

标题: 一个铁锈速度提升来自哪里?滑动窗口威胁评分器中的语言和算法影响

摘要: 将Python中的热路径重写为Rust是加速安全分析的常见方法,通常会报告大幅度的速度提升。重写通常会同时更改语言和算法,因此一个因素可以归功于语言的收益,而这些收益来自更好的算法。我们在一个基于SentinelSphere平台的交通灯计算器模型上研究了这一点。五个实现,其中三个是Python,两个是Rust,产生了完全相同的分数,通过共享的校验和进行了确认。我们在一个受限制的和一个突发的环境中对它们进行了从100到一百万事件的时间测试。仅语言本身就贡献了大约4到29倍。用增量更新代替窗口的每个事件重新扫描在突发环境中贡献了超过11,000倍,因此整体因子在窗口持续填充并在不填充时接近13,000倍时达到了约195,000倍。对于原始的重写所报告的时间,进行幂律拟合得到Python的增长指数为1.85,Rust为0.88,这是算法差异的标志。因此,安全工具的性能声明应分别报告语言和算法的贡献。

更新时间: 2026-10-06 08:08:49

领域: cs.CR

下载: http://arxiv.org/abs/2610.07931v1

Dynamic Alignment and Calibration for Multimodal Learning

Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise variations, potentially leading to unreasonable over-alignment; and (ii) confidence- or uncertainty-aware fusion methods often fail to adequately account for feature magnitude and confidence differences across modalities. For modality pairs with significant feature magnitude differences or small confidence gaps, it might be unreliable to strictly align fusion weights according to confidence. To address these issues, we propose an Alignment- and Calibration-driven Multimodal Learning framework (ACML). Specifically, ACML incorporates a dynamic cross-modal triplet alignment module, which enforces strong semantic consistency for high-confidence positive pairs while encouraging diverse representation learning between high- and low-confidence positive pairs according to their confidence gaps. Additionally, ACML introduces a difference-aware attention calibration strategy that adaptively adjusts attention regularization based on feature magnitude and confidence differences across modalities, thereby mitigating biases caused by unreasonable fusion constraints. Extensive experiments on multiple multimodal benchmark datasets demonstrate that ACML consistently achieves superior performance and robustness over recent state-of-the-art methods.

Updated: 2026-10-06 08:05:46

标题: 多模态学习的动态对齐和校准

摘要: 动态多模态学习旨在通过自适应地建模跨模态之间的信息差异来学习稳健的表示。然而,现有方法仍然存在两个限制:(i)静态跨模态对齐策略通常对所有样本施加统一约束,而忽略了样本间的变化,可能导致不合理的过度对齐;和(ii)置信度或不确定性感知融合方法通常未能充分考虑跨模态之间的特征大小和置信度差异。对于特征大小差异显著或置信度差距较小的模态对,严格根据置信度对齐融合权重可能不可靠。为了解决这些问题,我们提出了一种基于对齐和校准驱动的多模态学习框架(ACML)。具体而言,ACML包含一个动态的跨模态三元对齐模块,该模块在高置信正样本对中强制执行强语义一致性,同时根据其置信度差异鼓励高置信和低置信正样本对之间的多样化表示学习。此外,ACML引入了一种差异感知的注意力校准策略,根据跨模态的特征大小和置信度差异自适应调整注意力正则化,从而减轻由不合理的融合约束引起的偏差。在多个多模态基准数据集上的大量实验表明,ACML始终优于最近的最先进方法,并具有卓越的性能和鲁棒性。

更新时间: 2026-10-06 08:05:46

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.07928v1

WAMJET: A Harness for World Action Model Acceleration

World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tackle this bottleneck, we present WAMJET, an agentic harness that accelerates WAM inference by equipping coding agents with reusable optimization guidance and measurement and validation tools. WAMJET follows a bottleneck-driven workflow where the agent profiles inference, modifies targeted code, validates effects, and iteratively refines the acceleration stack as bottlenecks shift, while preserving action quality. Experiments span six WAMs, three coding agents, and two GPU architectures. WAMJET achieves up to 9.95x lossless speedup over upstream implementations. Approximation and hardware-aware optimization yield additional latency reductions, with comparable success rates. The results show that WAMJET can produce effective acceleration stacks for WAM deployment.

Updated: 2026-10-06 08:03:34

标题: WAMJET:世界行动模型加速的工具

摘要: World Action Models (WAMs) 利用预训练的视频基础模型进行机器人操作,但它们庞大的主干和视频-动作共同预测成本高昂。尽管现有的加速技术提供了许多降低这一成本的方法,但选择和组合它们需要针对每个模型和硬件平台进行大量工程工作。为了解决这一瓶颈问题,我们提出了WAMJET,一种行动力量驱使的工具,通过为编码代理提供可重复使用的优化指导和测量验证工具,加速WAM的推理。WAMJET遵循瓶颈驱动工作流程,其中代理人对推理进行分析、修改目标代码、验证效果,并在瓶颈转移时迭代地完善加速堆栈,同时保持动作质量。实验涵盖了六种WAMs、三种编码代理和两种GPU架构。WAMJET在上游实现上实现了高达9.95倍的无损加速。近似和硬件感知优化产生额外的延迟降低,具有可比较的成功率。结果表明,WAMJET可以为WAM部署生成有效的加速堆栈。

更新时间: 2026-10-06 08:03:34

领域: cs.CV,cs.AI,cs.RO

下载: http://arxiv.org/abs/2610.03797v2

Flow-Transformed Implicit Processes for Function-Space Variational Inference

Implicit-process priors define distributions over functions through flexible generative mechanisms, making them attractive for Bayesian function-space modelling. However, performing posterior inference with such priors is challenging because their induced function-space distributions are typically not available in closed form. One practical strategy is to approximate the prior using a finite collection of sampled functions, and then represent posterior functions as learned combinations of these samples. Existing approaches commonly place a Gaussian variational distribution over the combination weights. While tractable, this choice limits the shapes of posterior uncertainty that can be represented, especially when the true posterior is asymmetric, heavy-tailed, or multimodal. We propose Flow-Transformed Implicit Processes (FTIP), a variational inference method that makes this finite-dimensional function-space approximation more expressive. Instead of using a Gaussian distribution over the combination weights, FTIP uses a normalizing flow to define a richer variational distribution. This induces a flexible posterior distribution over functions while preserving tractable optimization. We train the model using a Black-Box α objective, allowing us to compare mass-covering and mode-seeking variational behaviour. Experiments show that FTIP captures asymmetric and multimodal posterior structure in function space that Gaussian coefficient approximations tend to smooth or collapse.

Updated: 2026-10-06 08:01:35

标题: 流变换的隐式过程用于函数空间变分推断

摘要: Implicit-process priors通过灵活的生成机制定义函数分布,因此在贝叶斯函数空间建模中具有吸引力。然而,使用这种先验进行后验推断是具有挑战性的,因为它们引起的函数空间分布通常无法以闭合形式获得。一种实用的策略是使用有限数量的采样函数来近似先验,然后将后验函数表示为这些样本的学习组合。现有方法通常在组合权重上放置一个高斯变分分布。虽然可行,但这种选择限制了可以表示的后验不确定性的形状,特别是当真实后验是非对称的、重尾的或多峰的时。我们提出了Flow-Transformed Implicit Processes (FTIP),这是一种变分推断方法,可以使这种有限维函数空间近似更具表达力。FTIP不是使用高斯分布来定义组合权重,而是使用正规化流来定义更丰富的变分分布。这在保持可行优化的同时引起了函数上的灵活后验分布。我们使用Black-Box α目标训练模型,使我们能够比较质量覆盖和寻找模式的变分行为。实验证明,FTIP在函数空间中捕获了非对称和多峰后验结构,而高斯系数近似往往会平滑或坍缩。

更新时间: 2026-10-06 08:01:35

领域: cs.LG,stat.ML

下载: http://arxiv.org/abs/2606.01954v3

Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images

Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models. We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training. Each WSI is represented as a bag of patches paired with a slide-level diagnostic caption. The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference. We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models. The source code is available at https://github.com/helomelo1/MKD-LMF.

Updated: 2026-10-06 07:56:54

标题: 多模态知识蒸馏用于胃腺癌全切片图像分类

摘要: 胃腺癌(GA)是全球癌症相关死亡的主要原因,从整张切片图像(WSIs)进行准确的组织病理学亚型分类对于有效的治疗规划至关重要。尽管将病理报告文本与WSIs集成的多模态方法可以改善分类,但现有方法通常依赖于计算昂贵的变换器架构和大型语言模型。我们提出了一种多模态知识蒸馏(MKD)框架,它结合了预训练的WSI图像编码器和临床文本编码器,使用低秩多模态融合(LMF)来在训练期间高效地建模跨模态交互。每个WSI被表示为一组与幻灯片级诊断标题配对的补丁,教师模型学习融合的图像文本表示以进行亚型分类,而学生模型则蒸馏这种知识以实现准确的仅图像推理。我们在PatchGastric基准数据集上评估了我们的方法,并实现了比最先进方法至少高出3.35%的平均准确度,而不依赖于基于变换器的融合,多任务学习或大型语言模型。源代码可在https://github.com/helomelo1/MKD-LMF上找到。

更新时间: 2026-10-06 07:56:54

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.07913v1

MiniScope: Authorizing Agents with Least-Privilege Permissions

AI agents are increasingly granted autonomous access to sensitive user data and third-party services, making effective permission management a critical security challenge. Existing permission models, however, typically rely on flat permission structures that fail to balance security with usability: fine-grained confirmation induces user fatigue, while coarse-grained or persistent approval leads to overprivileged agents. To address this tradeoff, we propose a task-centric, hierarchical permission model that treats an agent as a delegate operating within a task-specific role instead of requiring a separate permission decision for every tool call. Building on this model, we present MiniScope, an end-to-end permission system for agents that automates permission-hierarchy discovery and enforces contextual least privilege at runtime. Our evaluation shows that MiniScope reduces simulated permission confirmations by 43.4%-89.4% for cautious and typical personas relative to per-tool prompting and mitigates all privilege-escalation attacks with negligible impact on utility and runtime. Applied to real-world deployments, MiniScope further uncovers six overprivileged connector configurations in ChatGPT and Claude.

Updated: 2026-10-06 07:56:06

标题: MiniScope:授权代理使用最低权限权限

摘要: 人工智能代理越来越被授予对敏感用户数据和第三方服务的自主访问权限,这使得有效的权限管理成为一个关键的安全挑战。然而,现有的权限模型通常依赖于平面权限结构,无法平衡安全性和可用性:细粒度的确认会导致用户疲劳,而粗粒度或持久的批准会导致代理权过大。为了解决这种权衡,我们提出了一种以任务为中心的分层权限模型,将代理视为在特定任务角色内运作的委托者,而不是要求为每个工具调用做出单独的权限决策。基于这一模型,我们提出了MiniScope,一个端到端的代理权限系统,用于自动发现权限层次结构并在运行时执行上下文最小权限。我们的评估显示,相对于每个工具提示,MiniScope将慎重和典型用户的模拟权限确认减少了43.4% -89.4%,并且在效用和运行时几乎没有影响的情况下缓解了所有特权升级攻击。应用于实际部署,MiniScope进一步发现了ChatGPT和Claude中六个过度特权的连接器配置。

更新时间: 2026-10-06 07:56:06

领域: cs.CR,cs.AI

下载: http://arxiv.org/abs/2512.11147v2

Diverse Motion Customization via Control-based Dynamic Optimization

Despite recent advances in video generation, motion customization remains challenging due to content leakage, where appearance attributes from the reference video unintentionally propagate into the generated output. We identify this issue as a consequence of the generative process collapsing toward the reference video, which arises from formulating the learning objective as a direct regression on the reference. To address this, we propose Control-based Motion Customization (CMC), a principled training framework that is structurally robust to content leakage. Our key idea is to steer generative dynamics toward desired motion while avoiding collapse toward the reference video, which we formalize using Stochastic Optimal Control (SOC). Under this formulation, customized videos acquire the target motion yet remain within the pre-trained model's prompt-conditional distribution, where appearance is determined by the text prompt rather than the reference video. Furthermore, to improve efficiency, we tailor the SOC formulation to motion customization by eliminating the need for an explicit reward and introducing a timestep-adaptive motion cost that focuses only on early generative stages, accelerating training by 2.5 times. Extensive experiments demonstrate that CMC effectively mitigates content leakage and achieves competitive motion fidelity while preserving the diversity of the base model across diverse scenarios.

Updated: 2026-10-06 07:55:27

标题: 通过基于控制的动态优化实现多样化的运动定制

摘要: 尽管视频生成近年来取得了一些进展,但由于内容泄漏,运动定制仍然具有挑战性,即参考视频中的外观属性意外地传播到生成的输出中。我们确定这一问题是生成过程向参考视频折叠的结果,这是由于将学习目标制定为对参考的直接回归而引起的。为了解决这个问题,我们提出了基于控制的运动定制(CMC),这是一个结构上抗泄漏的训练框架。我们的关键思想是引导生成动态朝向所需运动,同时避免向参考视频折叠,我们使用随机最优控制(SOC)来形式化这一概念。在这种形式化下,定制的视频获得目标动作,但仍保持在预先训练模型的提示条件分布内,外观由文本提示而不是参考视频决定。此外,为了提高效率,我们将SOC公式调整为适用于运动定制,通过消除明确奖励的需要,并引入一个适应时间步的运动成本,仅关注早期生成阶段,加速训练2.5倍。大量实验证明,CMC有效缓解了内容泄漏,并在保持基础模型在不同情况下的多样性的同时实现了竞争性的运动保真度。

更新时间: 2026-10-06 07:55:27

领域: cs.CV,cs.AI

下载: http://arxiv.org/abs/2610.07911v1

Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning

Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoothness. Building on this finding, we propose Conditioning for Action using only Temporal Smoothness (CATS), which combines a temporal penalty with linear ramp-up. We highlight temporal regularization's ability to provide spatial smoothness while better preserving task performance than explicit spatial regularization. Through linear ramp-up, CATS allows the policy to learn rewarding behavior before progressively smoothing its actions, improving return preservation and both temporal and spatial smoothness. Experiments in both simulation and the real world show that CATS substantially reduces action oscillation without degrading task performance, with little computational overhead.

Updated: 2026-10-06 07:53:24

标题: 重新审视深度强化学习中平滑控制的时间正则化

摘要: 深度强化学习策略可能会产生非平滑的动作振荡,从而妨碍在物理机器人上的部署。现有的架构和基于惩罚的方法旨在通过直接减少对状态输入变化的敏感性来寻求空间平滑性,但它们的广泛约束可能会随着追求更强的平滑而降低任务性能。而时间正则化则通过约束观察到的转换中的动作差异,但被认为无法在观察噪声下提供所需的空间平滑性。我们通过证明时间惩罚约束了当前状态和共享下一个状态之间的预期动作差异,揭示了一种空间效应,实验证明这种效应延伸到了空间平滑性。基于这一发现,我们提出了仅使用时间平滑性的行为调节(CATS),它结合了时间惩罚和线性逐渐增加。我们强调时间正则化提供空间平滑性的能力,同时比显式空间正则化更好地保留任务性能。通过线性逐渐增加,CATS允许策略学习有益行为,然后逐渐平滑其动作,提高回报保留和时间空间平滑性。在模拟和真实世界中的实验表明,CATS显著减少了动作振荡,而不会降低任务性能,并且计算开销很小。

更新时间: 2026-10-06 07:53:24

领域: cs.LG,cs.RO

下载: http://arxiv.org/abs/2610.07910v1

Continuous Memory Machines

Recurrent neural networks typically compress information into a single vector-valued recurrent state, forcing short-term computation and long-term retention to share the same representation. Past extensions alleviate this bottleneck by increasing the memory capacity or separating timescales, but lack the combination of rapid neuron-level processing and longer-term retention found in biology. To that end, we introduce the Continuous Memory Machine (CMM), a recurrent architecture with matrix-valued short- and long-term memory states serving distinct functional roles. Building on the Continuous Thought Machine (CTM), the CMM's short-term memory tracks recent neural activity, with uniquely parameterized neuron-level models learning to use these activity patterns for computation. A persistent long-term memory stores information for later use, with a Transformer jointly updating both memory stores, providing an expressive bidirectional read--write mechanism such that each store can reorganize its own contents and both read from and write to the other. Across algorithmic, in-context learning, and recurrent reasoning tasks, the CMM outperforms a broad suite of baselines, exhibiting stronger generalization than prior memory-augmented networks while preserving the CTM's interpretable attention patterns. Code is available at https://github.com/SakanaAI/continuous-memory-machines.

Updated: 2026-10-06 07:52:50

标题: 连续存储器机器

摘要: 循环神经网络通常将信息压缩为单个向量值的循环状态,强制短期计算和长期记忆共享相同的表示。过去的扩展通过增加存储容量或分离时间尺度来缓解这一瓶颈,但缺乏生物学中发现的快速神经元级处理和长期保留的组合。为此,我们介绍了连续记忆机(CMM),这是一个具有矩阵值短期和长期记忆状态的循环体系结构,具有不同的功能角色。基于连续思维机(CTM),CMM的短期记忆跟踪最近的神经活动,具有独特参数化的神经元级模型学习如何利用这些活动模式进行计算。持久的长期记忆存储信息供以后使用,使用Transformer共同更新两个存储器,提供表达丰富的双向读-写机制,使每个存储器可以重新组织自己的内容,并从另一个存储器读取和写入。在算法、上下文学习和循环推理任务中,CMM优于广泛的基线套件,表现出比先前的记忆增强网络更强的泛化能力,同时保留了CTM的可解释的注意模式。代码可在https://github.com/SakanaAI/continuous-memory-machines 上找到。

更新时间: 2026-10-06 07:52:50

领域: cs.AI

下载: http://arxiv.org/abs/2610.07907v1

Isotropic Yet Undecodable: The Sequential Content-Sufficiency Gap in Latent-Predictive Text Representations

We study sequential content sufficiency by investigating whether a representation retains the ordered target information available in its input. An information-theoretic decomposition separates input ambiguity, representation loss, and readout mismatch. We construct recoverable views where perfect agreement and joint isotropic Gaussianity coexist with zero target information, and establish limits imposed by deterministic canonical anchors. Token log-loss provides a one-sided information-loss bound; a fixed-penalty ridge analysis shows why rank alone cannot determine prediction risk. These results motivate CANOPE, a nonautoregressive framework with ordered latent canvases, canonical-token supervision, and geometric regularization. On 40,000 validation sequences, latent-agreement (PL0) and token-grounded (PL2) have nearly identical pooled ranks but reach 13.5% and 98.8% positional Recall@1, respectively, under strong natural corruption when the correct target length is provided. On 3,930 LJSpeech validation utterances, frozen PL2 with a trained MatchaTTS readout yields 21.54% word error rate (WER) on corrupted text, versus 99.22% for frozen PL0, while end-to-end MatchaTTS reaches 10.93%. These results show that geometric regularity alone does not guarantee recoverable sequential content or effective downstream access in the text settings studied here.

Updated: 2026-10-06 07:52:37

标题: 各向同性但无法解码:潜在预测文本表示中的顺序内容充分性差距

摘要: 我们通过研究顺序内容充分性来调查表示是否保留其输入中可用的有序目标信息。信息论分解将输入模糊性、表示丢失和读出不匹配分开。我们构建可恢复的视图,其中完美一致性和联合各向同性高斯性与零目标信息共存,并确定确定性规范锚所施加的限制。令牌对数损失提供了一个单边信息损失界限;固定惩罚岭分析显示仅凭等级不能确定预测风险的原因。这些结果促使CANOPE的发展,这是一个具有有序潜在画布、规范-令牌监督和几何正则化的非自回归框架。在40,000个验证序列上,潜在一致性(PL0)和令牌基础(PL2)具有几乎相同的综合排名,但在提供正确目标长度时,它们分别达到13.5%和98.8%的位置Recall@1,在强自然损坏情况下。在3,930个LJSpeech验证语音中,与训练有素的MatchaTTS读出相结合的冻结PL2在损坏文本上产生21.54%的词错误率(WER),而冻结PL0的错误率为99.22%,而端到端的MatchaTTS则为10.93%。这些结果表明,仅仅靠几何规则本身并不能保证在所研究的文本环境中恢复序列内容或有效地获得下游访问。

更新时间: 2026-10-06 07:52:37

领域: cs.AI,cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.07906v1

Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare

We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-specific interpretable models that synthesizes coefficient-space structure from natural-language task descriptions and a memory of previously learned task-specific predictors. RAIL retrieves related source tasks, transfers structure through coefficient space, and generates a new predictor in the original diagnostic-feature space, enabling zero-shot and few-shot clinical procedure prediction with feature-level explanations. Its probabilistic formulation provides uncertainty over retrieval, model coefficients, and predictions, supporting uncertainty-aware deployment: uncertain predictions or unstable explanations can be flagged for additional clinical review rather than treated as automatic decisions. This makes RAIL particularly suited for healthcare settings, where prediction tasks are highly long-tailed, new clinical targets arise frequently, and models must remain inspectable, uncertainty-aware, and compatible with human oversight. Across long-tailed clinical procedure prediction tasks, RAIL improves low-data model generation, benefits from clinically informed task representations, and yields retrieval, uncertainty, and coefficient-level diagnostics that make model behavior more transparent. These results suggest a path toward scalable clinical prediction systems that can adapt to new tasks while preserving interpretability and reliability.

Updated: 2026-10-06 07:50:24

标题: 检索增强的可解释学习:朝向医疗领域任务特定的零-shot模型

摘要: 我们介绍了检索增强可解释学习(RAIL),这是一个零样本生成特定任务可解释模型的概率元学习框架,它从自然语言任务描述和先前学习的任务特定预测器的记忆中综合出系数空间结构。RAIL检索相关的源任务,通过系数空间传递结构,并在原始诊断特征空间中生成新的预测器,实现了零样本和少样本临床流程预测,并提供特征级解释。其概率化表述提供了关于检索、模型系数和预测的不确定性,支持不确定性感知部署:不确定的预测或不稳定的解释可以被标记为需要额外临床审查,而不是作为自动决策处理。这使得RAIL特别适用于医疗保健领域,其中预测任务具有极度长尾性,新的临床目标频繁出现,模型必须保持可检查性、不确定性感知,并与人类监督兼容。在长尾临床流程预测任务中,RAIL改进了低数据模型生成,受益于临床信息化的任务表示,并产生了检索、不确定性和系数级诊断,使模型行为更透明。这些结果表明了一条通向可扩展临床预测系统的道路,这些系统可以适应新任务,同时保持解释性和可靠性。

更新时间: 2026-10-06 07:50:24

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2607.17508v3

ApexQuant: Data-Free Elastic Quantization by Residual Re-Isotropization

We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers. We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes. This result lets us determine, before any weight is read, how many passes a layer needs for a target weight-space error. Every prefix is itself a valid lower-rate model, so one artifact serves several precisions. We instantiate ApexQuant with three interchangeable stages, scalar, $E_8$ and trellis, and validate it on four open-weight LLMs and on Earth-observation and medical domains where in-distribution data is often unattainable as imagery arrives under restrictive licences or due to patient material under privacy constraints. Progressive re-isotropization comes within a few percent of full precision at four bits and gives the best two-bit arm we measure, in a completely data-free setting.

Updated: 2026-10-06 07:49:55

标题: ApexQuant:通过残差重新等距化实现无数据弹性量化

摘要: 我们引入了ApexQuant,这是一种无需校准的量化方法,通过递归重新量化残差误差,在现有量化器之上作为一个精细的层。我们确定,一个新的随机旋转可以将每个残差返回到超球面上的均匀分布,这表征了在连续传递中逐渐减小误差的速率。这个结果让我们在读取任何权重之前就能确定,一个层需要多少次传递才能达到目标权重空间误差。每个前缀本身都是一个有效的低速率模型,因此一个工件可以服务于多个精度。我们使用三个可互换的阶段,标量、$E_8$和格雷码,实例化了ApexQuant,并在四个开放权重的LLMs上以及地球观测和医学领域进行了验证,这些领域中的数据通常无法获得,因为图像受到限制性许可或由于患者资料受到隐私约束。在完全不依赖数据的情况下,渐进性的重新等规化在四位时接近完整精度,提供了我们测得的最佳二位性能。

更新时间: 2026-10-06 07:49:55

领域: cs.LG

下载: http://arxiv.org/abs/2610.07904v1

IEEE 802.11bx - WLAN Intelligent Networking (WIN): Toward an AI-Ready Wi-Fi 9

Wi-Fi 9 is expected to go beyond mere communication and provide new services such as sensing or computation. At this juncture, Artificial Intelligence (AI) is taking a leading role in the definition of the 802.11bx amendment, named WLAN Intelligent Networking (WIN). In this tutorial, we survey the recent progress made toward Wi-Fi 9 within IEEE 802.11 standardization, tracing the drivers and technological advances that motivate an AI-ready Wi-Fi 9. We then examine AI's role along three complementary dimensions, i.e., AI as a protocol (AI is applied to Wi-Fi's PHY/MAC operation), AI as a platform (Wi-Fi infrastructure is repurposed to provide AI computation), and AI as traffic (AI flows call for new traffic-handling policies), and discuss candidate features and open challenges along each. As a concrete illustration of the AI as traffic paradigm, we present a case study on AI traffic differentiation, where we explore a potential extension of the current Enhanced Distributed Channel Access (EDCA) to support new AI traffic flows.

Updated: 2026-10-06 07:47:48

标题: IEEE 802.11bx - WLAN智能网络(WIN):走向AI准备的Wi-Fi 9

摘要: Wi-Fi 9预计将超越纯粹的通信,提供新的服务,如传感或计算。在这个时刻,人工智能(AI)正在领导802.11bx修正案的定义,名为无线局域网智能网络(WIN)。在本教程中,我们调查了IEEE 802.11标准化中近期关于Wi-Fi 9的进展,追踪推动AI准备的Wi-Fi 9的驱动因素和技术进步。然后,我们沿着三个互补维度(即AI作为协议,AI应用于Wi-Fi的PHY/MAC操作;AI作为平台,Wi-Fi基础设施被重新用于提供AI计算;AI作为流量,AI流需要新的流量处理政策)检查AI的作用,并讨论每个维度上的候选特性和开放性挑战。作为AI作为流量范式的具体示例,我们提出了一个关于AI流量区分的案例研究,其中我们探讨了当前增强分布式通道访问(EDCA)的潜在扩展,以支持新的AI流量流。

更新时间: 2026-10-06 07:47:48

领域: cs.NI,cs.AI

下载: http://arxiv.org/abs/2610.07900v1

Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments

Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected $Q$-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.

Updated: 2026-10-06 07:47:26

标题: 方差厌恶$n$步离线强化学习用于稀疏长期环境

摘要: 生成式演员正在通过实现复杂动作分布的表达策略类,改变离线强化学习(RL)。然而,这种表达能力也暴露了异构数据集中的一个关键挑战:生成式策略可以复制不可靠的动作模式,其回报分布表现出高方差,偶尔会因偶然事件而产生高回报,但缺乏一致性。因此,仅仅最大化预期的$Q$值是不足以识别可靠动作的。我们提出了VAN-Flow(方差回避$n$步流),这是一个促进生成式离线RL中可靠动作的框架。VAN-Flow结合了(i)一个分类分布评论家,(ii)一个方差回避期望算子,平滑地重新加权原子概率,以偏爱既有高回报又低离散度的动作,以及(iii)通过拒绝采样引导的流匹配生成式演员。与CVaR或均值方差目标不同,该算子在没有硬截断或辅助惩罚项的情况下,在分类回报分布上重新分配概率质量。在来自D4RL和OGBench的40多个任务中,VAN-Flow始终胜过强基线,在长时间跨度和高方差范围中获得最大收益,可靠动作选择变得至关重要。

更新时间: 2026-10-06 07:47:26

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.07899v1

FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents

Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next trajectory is generated, while the failed execution determines its conditioning context. We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training. After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory. FC-SWE adapts GRPO to these chains of complete, multi-turn tool-use trajectories through two mechanisms. Trajectory-local rewards preserve each attempt's verifier outcome, preventing recovery success from rewarding an earlier failed patch. Active-set advantage estimation forms a comparison group from all initial and recovery trajectories actually executed for the same issue, so failed attempts remain in the group while unexecuted attempts are excluded. On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO. Although trained with at most two attempts per chain, FC-SWE reaches 70.7% Resolved@11 under an eleven-attempt test-time budget.

Updated: 2026-10-06 07:44:12

标题: FC-SWE:面向长期视野软件工程代理的故障条件强化学习

摘要: 存储库级软件工程(SWE)是一个具有挑战性的长期设置:代理人必须在延长的交互、使用工具和适应有状态环境方面进行推理。最近的研究使用强化学习方法(如Group Relative Policy Optimization(GRPO))训练SWE代理,这些方法独立地对每个问题采样多个轨迹,测试生成的补丁,并在固定组内比较终端奖励。然而,这种训练设置不会重复使用来自失败补丁的验证器反馈作为后续尝试的上下文,尽管这些反馈包含有关出错原因的宝贵诊断信息。在恢复轨迹上进行训练是具有挑战性的,因为前面的结果决定了下一个轨迹是否生成,而失败的执行决定了其条件化上下文。我们引入了FC-SWE,这是一个将恢复尝试纳入策略训练的失败条件化RL框架。在补丁未通过验证时,FC-SWE将存储库恢复到其原始任务状态,并将失败的补丁和验证器反馈作为恢复轨迹的上下文。FC-SWE通过两种机制将GRPO调整为这些完整的、多轮工具使用轨迹的链。轨迹局部奖励维持每次尝试的验证器结果,防止恢复成功奖励之前的失败补丁。主动集合优势估计从为同一问题实际执行的所有初始和恢复轨迹中形成一个比较组,因此失败尝试仍然保留在组中,而未执行的尝试被排除。在一个验证器辅助协议下的500个SWE-bench验证任务中,FC-SWE与Qwen3.5-4B和SWE-agent相比,达到了41.7%的Resolved@1和52.8%的Resolved@2,而GRPO分别为38.9%和48.5%。尽管在每个链中最多进行两次尝试的情况下进行训练,FC-SWE在11个尝试的测试时间预算下达到了70.7%的Resolved@11。

更新时间: 2026-10-06 07:44:12

领域: cs.LG,cs.SE

下载: http://arxiv.org/abs/2610.07898v1

Textual Environmental Context and Spatial Graphs for LLM-Based Regional SST Forecasting

Sea surface temperature (SST) forecasting depends on local temporal persistence, regional spatial dependence, and environmental conditions that evolve with the forecast date. We study how these heterogeneous conditions can be presented to a large language model (LLM) for regional multi-step forecasting without serializing the full SST grid as text. We formulate forecasting as conditional numerical generation: historical SST and anomaly sequences, date-aligned environmental records, and static ocean knowledge form a textual context, while regional spatial state is supplied through continuous graph-derived prefixes. A static graph encodes persistent geographic--climatological relations, and a dynamic graph encodes recent SST correlations and localized tropical-cyclone influence. Two graph neural networks produce a target-node representation that is mapped by a spatial-prefix fusion and injected into the LLM input. On SST forecasting in the South China Sea, the complete configuration achieves the best MAE and $\Rtwo$ among the compared methods over ten forecast steps. Alongside the numerical forecast, a rule-based module matches predicted trends and environmental-factor directions with knowledge entries to return source-linked, post-hoc contextual explanations.

Updated: 2026-10-06 07:42:51

标题: LLM-Based区域SST预测的文本环境背景和空间图。

摘要: 海表面温度(SST)预测取决于当地时间持续性,区域空间依赖性以及随着预测日期而演变的环境条件。本研究探讨了如何将这些异质条件呈现给大型语言模型(LLM),用于区域多步预测,而无需将完整的SST网格串行化为文本。我们将预测形式化为条件数值生成:历史SST和异常序列,与日期对齐的环境记录以及静态海洋知识构成文本上下文,而区域空间状态则通过连续图导出的前缀提供。静态图编码了持久的地理气候关系,动态图编码了最近的SST相关性和局部热带气旋影响。两个图神经网络产生一个目标节点表示,通过空间前缀融合并注入LLM输入。在南海SST预测中,完整配置在十个预测步骤中与比较方法相比实现了最佳的MAE和$R^2$。除了数值预测外,基于规则的模块将预测的趋势和环境因素方向与知识条目匹配,以返回与源相关的事后上下文解释。

更新时间: 2026-10-06 07:42:51

领域: cs.AI

下载: http://arxiv.org/abs/2610.07895v1

Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective

Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation fails to capture a fundamental requirement of faithful behavior: the ability to adapt model responses to changes in available contexts. In particular, a model should provide correct answers when sufficient evidence is present and abstain when it is not. In this work, we propose a Pairwise Faithfulness Benchmark (PFaithBench) that evaluates whether a model can switch between answering and abstaining for the same question under supporting versus non-supporting contexts. Our evaluations across thirty-nine models with seven model families demonstrate that faithfulness fundamentally involves a trade-off between answering and abstaining, and that most current models exhibit a strong bias toward answering, with most faithfulness errors arising from over-answering, i.e., models tend to fabricate a response even when the provided context is insufficient. We further conduct a series of studies on faithfulness training under different data constructions. Our results show that training outcomes are highly sensitive to the specific composition of answering and abstaining data. Constructing answering and abstaining data from mismatched sources can cause models to rely on dataset-specific shortcuts rather than actual context sufficiency. Moreover, increasing answer-supervised data improves answering performance but exacerbates over-answering, while increasing abstaining data reduces hallucination but leads to over-abstention. The code and data are released at https://github.com/tmlr-group/PFaithBench.

Updated: 2026-10-06 07:42:38

标题: 重新思考在LLMs中的忠诚度:一种成对上下文敏感的视角

摘要: 大型语言模型(LLMs)被期望根据提供的上下文诚实地回答问题,在上下文信息不足以回答问题时放弃。现有的诚实性评估通常评估每个问题-上下文实例,但是,这种实例级评估未能捕捉到忠实行为的一个基本要求:即能够根据可用上下文的变化调整模型响应的能力。特别是,一个模型应该在有足够证据时提供正确答案,并在没有证据时放弃。在这项工作中,我们提出了一个成对诚实性基准(PFaithBench),评估模型是否能在支持和不支持的上下文下切换回答和放弃相同的问题。我们对七个模型家族的三十九个模型进行的评估表明,诚实性基本上涉及回答和放弃之间的权衡,而大多数当前模型表现出强烈的回答偏见,大多数诚实性错误是由于过度回答,即当提供的上下文不足时,模型倾向于捏造答案。我们进一步进行了一系列关于不同数据构建下诚实性训练的研究。我们的结果表明,训练结果对回答和放弃数据的具体组成非常敏感。从不匹配的来源构建回答和放弃数据可能会导致模型依赖于数据集特定的快捷方式而不是实际的上下文充分性。此外,增加回答监督数据会提高回答性能,但会加剧过度回答,而增加放弃数据会减少虚构,但会导致过度放弃。代码和数据发布在https://github.com/tmlr-group/PFaithBench。

更新时间: 2026-10-06 07:42:38

领域: cs.CL,cs.LG

下载: http://arxiv.org/abs/2610.07894v1

Visual Abstention in Unified Multimodal Models

Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.

Updated: 2026-10-06 07:33:42

标题: 统一多模态模型中的视觉抽象

摘要: 统一多模型(UMMs)整合了理解和生成,然而它们的生成行为很少受到对任务的理解的限制。我们形式化了视觉放弃:当请求的视觉转换在任务规则下不可能时,模型应该意识到不存在有效的解决方案,声明这一点,并拒绝生成。我们引入了Draw-or-Decline(DoD),一个涵盖了7个任务类别的1050个可行-不可行请求对的基准,共同衡量了编辑成功和拒绝不可行请求。评估了8个UMMs,我们发现编辑能力和放弃是不同的能力:即使是最强大的编辑器,在普通指令下也只有68.4%的编辑准确率,在拒绝不可行请求方面只有0.4%。他们的推理显示为什么:模型很少注意到冲突,而是计划编辑,仿佛请求是可能的,常常描述不在图像中的对象,或者悄悄地将请求改变成他们可以完成的请求。明确提示这些UMMs报告不可行性会增加文本拒绝,但降低编辑准确性。我们提出了VisTA(Visual Transformation and Abstention),一种训练方法,将可行和不可行的例子配对,以便模型在决定生成之前判断可行性。我们训练VisTA-BAGEL执行可行的编辑并拒绝不可行的请求。在没有任何提示的情况下,拒绝了93.0%的不可行请求,而最强大的编辑器只有0.4%,同时错误地拒绝了0.8%的可行请求。与提醒不同,这并不会影响编辑的准确性:VisTA-BAGEL完成了74.3%的可行编辑,超过了评估的8个UMMs中的任何一个。

更新时间: 2026-10-06 07:33:42

领域: cs.CL,cs.AI,cs.CV

下载: http://arxiv.org/abs/2610.07887v1

Prime Fourier Embeddings: A Principled Basis for Modular Arithmetic

Numbers have algebraic structure that standard neural embeddings often fail to expose. We introduce Prime Fourier Embeddings (PFE), which encode integers as prime-indexed (cos, sin) pairs derived from the harmonic analysis of Q, providing a pre-structured representation in which modular arithmetic reduces to selecting the relevant prime channel rather than discovering algebraic structure from scratch. We prove that any linear map equivariant with respect to the product group action on PFE must be block-diagonal with one independent block per prime -- a consequence of Schur's lemma applied to the resulting character decomposition. For square-free composite moduli, the Chinese Remainder Theorem predicts which prime channels are task-relevant. Both predictions are confirmed empirically: ablation studies show specialization ratios exceeding 500x between task-relevant and task-irrelevant channels, with perfect in-distribution test accuracy across all square-free composite moduli tested.

Updated: 2026-10-06 07:33:00

标题: 主要的傅立叶嵌入:模算术的原则基础

摘要: 数字具有代数结构,标准神经嵌入通常无法展现。我们介绍了素数傅立叶嵌入(PFE),它将整数编码为从Q的谐波分析中导出的素数索引(cos,sin)对,提供了一个预结构化表示,在这个表示中,模算术减少到选择相关的素数通道,而不是从头开始发现代数结构。我们证明,对于PFE上与乘积群作用等变的任何线性映射必须是块对角的,每个素数对应一个独立块--这是对结果特征分解应用舒尔引理的一个结果。对于无平方因子的复合模数,中国剩余定理预测了哪些素数通道是任务相关的。这两个预测在经验上得到了证实:消融研究表明任务相关和任务无关通道之间的专业化比率超过500倍,并且在所有测试的无平方复合模数中都实现了完美的分布测试准确性。

更新时间: 2026-10-06 07:33:00

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2606.23044v3

ShanLiangRen: A Nutrition Agent for Personalized Daily Meal Planning

Dietary nutrition planning plays an important role in chronic disease management and maintaining a healthy body. In applications, it must simultaneously satisfy personalized constraints and reasonable multidimensional nutritional goals. These two aspects often conflict, and user constraints evolve with feedback, resulting in a substantial gap between generic guidelines and executable plans. To bridge this gap, we first propose the personalized fully quantified multiobjective dietary planning problem (MDP). To tackle MDP, we develop a nutrition agent, ShanLiangRen. The system first transforms dietary specifications, nutrient data, user attributes and natural language requirements into an individualized constrained planning instance. It then employs an exact retrieval-augmented generation method to shrink the feasible candidate set from a large scale ingredient and recipe space. Finally, it adopts a refinement guided by Pareto principles, where an LLM iteratively revises candidate plans under deterministic nutrition computation and feedback from constraint verification. The system outputs fully quantified meal plans with explicit ingredients and portion sizes, together with reports on nutrition compliance that show constraint satisfaction and nutrient interval attainment. We have released the system online as a WeChat Program, ShanLiangRen. A demo video is available at https://www.youtube.com/watch?v=652OtY5VlGA.

Updated: 2026-10-06 07:32:50

标题: ShanLiangRen:个性化日常饮食计划的营养代理

摘要: 膳食营养规划在慢性疾病管理和保持健康身体中发挥着重要作用。在应用中,它必须同时满足个性化约束和合理的多维营养目标。这两个方面经常发生冲突,用户约束会随着反馈而发展,导致通用指南和可执行计划之间存在实质性差距。为了弥合这一差距,我们首先提出了个性化完全量化的多目标膳食规划问题(MDP)。为了解决MDP,我们开发了一个营养代理系统ShanLiangRen。该系统首先将膳食规格、营养数据、用户属性和自然语言要求转化为个性化约束规划实例。然后,它采用一种精确的检索增强生成方法,从大规模的成分和食谱空间中缩小可行的候选集。最后,它采用了一个受帕累托原则指导的精炼方法,其中一个LLM在确定性营养计算和约束验证的反馈下迭代修订候选计划。该系统输出了具有明确成分和份量大小的完全量化的餐饮计划,以及显示约束满足和营养间隔达成的营养合规报告。我们已经将该系统发布在网上作为一个微信程序,ShanLiangRen。演示视频可在https://www.youtube.com/watch?v=652OtY5VlGA 上观看。

更新时间: 2026-10-06 07:32:50

领域: cs.AI,cs.IR

下载: http://arxiv.org/abs/2610.07886v1

Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools

Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements. Deep learning has advanced this task but remains dependent on costly expert annotations. Label-efficient strategies such as self-supervised pretraining and semi-supervised learning are expected to ease this burden, yet it remains unclear whether they yield reliable delineation and whether the deep models they produce outperform the delineation tools used in practice. We address this in two stages. First, comparing self-supervised objectives with supervised or semi-supervised fine-tuning across one internal and four external datasets, we find that pretraining helps but the objective matters, and that the value of semi-supervised fine-tuning depends on the pretraining objective. Second, we benchmark the selected deep learning model against widely used open-source (NeuroKit2, Prominence, ECGdeli) and commercial (CalECG) tools using three complementary metrics. The model ranks best on every metric and dataset, outperforming the strongest tool by a clear margin on the rhythm-diverse set (mIoU 71.3 vs. 54.8%; averaged point-wise sensitivity 92.6 vs. 76.4%), and degrades the least from sinus to arrhythmia. A rhythm-stratified and point-wise analysis further characterizes the distinctive behavior of each tool, yielding practical guidance for tool selection. These results provide systematic, multi-dataset evidence that self-supervised pretraining is effective for ECG delineation and enables a label-efficiently trained deep learning model to outperform widely used delineation tools by leveraging abundant unlabeled data. This supports adopting such models in diverse, real-world clinical settings.

Updated: 2026-10-06 07:32:30

标题: 标签高效深度学习用于心电图划分:与广泛使用的划分工具进行多数据集基准对比

摘要: 心电图(ECG)的描绘,即波形边界的识别,是将原始ECG信号转化为临床可解释测量的基础步骤。深度学习已经推进了这一任务,但仍然依赖于昂贵的专家注释。自监督预训练和半监督学习等标签有效策略预计将减轻这一负担,然而仍不清楚它们是否能产生可靠的描绘结果以及它们所产生的深度模型是否优于实际使用的描绘工具。我们在两个阶段解决了这个问题。首先,在一个内部和四个外部数据集上比较自监督目标与监督或半监督微调,我们发现预训练有助于但目标很重要,半监督微调的价值取决于预训练目标。其次,我们使用三个互补的指标将选定的深度学习模型与广泛使用的开源(NeuroKit2、Prominence、ECGdeli)和商业(CalECG)工具进行基准测试。该模型在每个指标和数据集上表现最佳,超过了最强工具在节律多样性集上的明显边缘(mIoU 71.3 vs. 54.8%; 平均点对点灵敏度 92.6 vs. 76.4%),且从窦性到心律失常的降级最小。基于节律分层和点对点分析进一步描绘了每个工具的独特行为,为工具选择提供了实用指导。这些结果提供了系统性、多数据集证据,证明了自监督预训练对于ECG描绘的有效性,并使得通过利用丰富的未标记数据进行标签有效训练的深度学习模型能够在广泛使用的描绘工具之上表现优异。这支持在多样化的真实临床环境中采用这样的模型。

更新时间: 2026-10-06 07:32:30

领域: cs.LG,cs.AI,cs.CV,eess.SP

下载: http://arxiv.org/abs/2610.07885v1

Learned Adaptive Multiresolution Diffusion Imaging

Adaptive multiresolution methods reduce representation cost by concentrating fine-scale degrees of freedom where needed, but their tree updates are usually governed by fixed local criteria. We introduce Learned Adaptive Multiresolution Diffusion Imaging (Learned AMDI), which preserves the AMDI fixed-tree propagator and hierarchy constraints while replacing the post-propagation selector with a shared local policy trained by proximal policy optimization. Regression tests reproduce deterministic AMDI trajectories to machine precision when identical trees are used. In the Haar implementation studied here, the deterministic one-step selector accepts no refinements in 54 decisions. Across nine held-out cases, Learned AMDI executes 393 refinements and reduces the mean terminal reference discrepancy from $0.17496$ to $0.13657$, while occupancy rises from $0.13737$ to $0.26660$. Step-resolved diagnostics reveal occasional small adaptation-energy increases; fixed-tree energy stability therefore does not guarantee monotonicity of the learned outer iteration. At comparable occupancy, a validation-tuned observed-detail threshold reaches a discrepancy of $0.13792$ with slightly better RMSE and SSIM, placing both methods on essentially the same accuracy--occupancy tradeoff. A decision-1-only control reaches $0.13742$, indicating that most of the improvement on this static benchmark arises from the initial allocation. The shared actor transfers without retraining to $64\times64$ and $128\times128$ images, improving reference discrepancy, RMSE, and SSIM relative to deterministic AMDI, while the frozen threshold rule remains competitive. Learned AMDI thus provides a hierarchy-constrained, resolution-transferable mechanism for adaptive allocation and clarifies the contribution of sequential decisions.

Updated: 2026-10-06 07:31:14

标题: 学习自适应多分辨率扩散成像

摘要: 自适应多分辨率方法通过在需要时集中细粒度自由度来减少表示成本,但它们的树更新通常由固定的本地标准控制。我们介绍了学习自适应多分辨率扩散成像(Learned AMDI),它保留了AMDI固定树传播器和层次结构约束,同时用由近端策略优化训练的共享本地策略替换后传播选择器。回归测试在使用相同树时复制确定性AMDI轨迹至机器精度。在这里研究的Haar实现中,确定性的一步选择器在54个决策中不接受任何细化。在九个保留案例中,学习AMDI执行393次细化,将平均末端参考差异从0.17496降至0.13657,同时占用率从0.13737增至0.26660。步骤分辨诊断显示偶尔会出现小的适应能量增加;因此,固定树能量稳定性不能保证学习外部迭代的单调性。在可比较的占用率下,经过验证调整的观察细节阈值达到0.13792的差异,具有稍好的均方根误差和结构相似性指数,从而使两种方法基本上具有相同的准确性-占用率权衡。决策-1-仅控制达到0.13742,表明在这个静态基准上的大部分改进来自初始分配。共享的执行器在不重新训练的情况下转移到64×64和128×128像素图像,相对于确定性AMDI,改进了参考差异、均方根误差和结构相似性指数,而冻结的阈值规则仍然具有竞争力。因此,学习AMDI提供了一个具有层次约束、可转移分辨率的自适应分配机制,并阐明了顺序决策的贡献。

更新时间: 2026-10-06 07:31:14

领域: math.NA,cs.LG,eess.IV

下载: http://arxiv.org/abs/2610.07884v1

Self-Referenced Social Preferences: Cooperation without Observing Others Rewards

Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approach on three sequential social dilemmas, Escape Room, Clean Up, and Commons Harvest, which require volunteering, public-good contribution, and resource restraint, respectively. Across all three environments, agents learn cooperative behavior without observing others' rewards, including in settings where independent learners fail to cooperate, and frequently achieve more equitable divisions of jointly produced returns than agents with access to true rewards. The effective integration point depends on the social preference: inequity aversion works best in the reward together with a value look-ahead, whereas a purely benevolent preference benefits from policy-update weighting. Under partial observability, the policy-update approach continues to support cooperation. These results show that explicit access to other agents' reward signals is not necessary for learning cooperative behavior: social preferences can instead be grounded in self-referenced assessments of others' outcomes derived from their observed behavior.

Updated: 2026-10-06 07:29:38

标题: 自我参照的社会偏好:在不观察他人奖励的情况下合作

摘要: 社会偏好可以促进多智能体强化学习中的合作,但现有方法通常要求智能体观察同伴的奖励。然而,在许多现实世界的互动中,一个智能体可以像人类一样观察其他人的行为和结果,而无需访问他们的私人奖励信号。我们引入了自我参考的社会偏好,其中每个智能体学习自己的奖励模型,将其应用于其他智能体观察到的转换,以从自己的角度评估他们的结果,并将这些自我参考的评估输入到标准的社会偏好中。我们研究了两种整合这些评估的方法:修改学习奖励,或使用它们来加权策略更新。我们在三个连续的社会困境环境中评估了这种方法,包括逃脱室、清理和共享收获,分别需要自愿、公共贡献和资源约束。在所有三个环境中,智能体学习合作行为而不观察其他人的奖励,包括在独立学习者无法合作的情况下,经常实现更公平的共同生产回报分配。有效整合点取决于社会偏好:厌恶不公平在奖励以及价值前瞻中效果最佳,而纯粹的仁慈偏好从策略更新加权中获益。在部分可观察性下,策略更新方法仍然支持合作。这些结果表明,明确访问其他智能体的奖励信号对学习合作行为不是必要的:社会偏好可以基于从观察到的行为中得出的其他人结果的自我参考评估。

更新时间: 2026-10-06 07:29:38

领域: cs.AI,cs.MA

下载: http://arxiv.org/abs/2610.07881v1

Don't Let One Lie Survive A Hundred Truths: A Selective Bayesian Trust Estimator for Collaborative Perception

Collaborative perception (CP) enables connected vehicles to see beyond their own sensors but makes them dependent on messages they cannot independently verify. A compromised collaborator can surgically conceal a single safety-critical object or inject a non-existing one while correctly reporting many others. Existing Bayesian trust mechanisms pool agreement across objects, which, while effective against blatant untargeted attacks, either incurs high false-positive rates (FPR), or allows unrelated correct reports to dilute persistent attack evidence for stealthy single-object attackers. To address this problem, we propose SABER, a selective two-tier Bayesian trust estimator. The first tier maintains broad agent and object trust, preserving the ability to downweight benign but low-quality contributors. Cumulative-sum screening selects agent--object pairs with persistent omissions or unsupported reports for focused Bayesian assessment. The second tier checks these pairs against other agents' evidence and maintains a separate, reference-weighted Beta state for each. The lowest pair score constrains agent trust, preventing unrelated reports from diluting a targeted attack. We establish sufficient conditions for stronger attacker-side trust reductions with bounded additional benign false alarms at fixed thresholds. Compared with state-of-the-art CP defenses, SABER improves attack detection while reducing benign FPRs. On OPV2V, SABER improves defense ROC-AUC over MATE by up to 0.427 in late fusion and 0.337 in intermediate fusion. Against advanced intermediate-fusion data fabrication attacks, it increases detection rates over ROBOSAC and LUCIA by up to 96.40 and 67.07 percentage points, respectively, while reducing FPRs.

Updated: 2026-10-06 07:25:07

标题: 不要让一个谎言存活一百个真相:用于协作感知的选择性贝叶斯信任估计器

摘要: 协同感知(CP)使连接车辆能够超越自身传感器的范围,但也使它们依赖于无法独立验证的消息。一个被篡改的合作者可以精确地隐藏一个安全关键对象或者注入一个不存在的对象,同时正确报告其他许多对象。现有的贝叶斯信任机制会汇总对象间的一致性,虽然对于明显的非定向攻击有效,但会导致高虚警率(FPR),或者允许无关的正确报告稀释针对隐蔽单对象攻击者的持续攻击证据。为了解决这个问题,我们提出了SABER,一种选择性的两层贝叶斯信任估计器。第一层维护广泛的代理和对象信任,保留了降低良性但低质量贡献者权重的能力。累积和筛选选择具有持续遗漏或无支持报告的代理-对象对进行重点的贝叶斯评估。第二层检查这些对与其他代理的证据,并为每个对维护一个独立的、参考加权的Beta状态。最低对分数约束代理信任,防止无关报告稀释针对性攻击。我们建立了更强的攻击者信任降低的充分条件,并在固定阈值下增加良性误报。与最先进的CP防御相比,SABER提高了攻击检测能力,同时减少了良性FPR。在OPV2V上,SABER在后期融合中比MATE提高了最多0.427的防御ROC-AUC,在中间融合中提高了最多0.337。对抗高级中间融合数据伪造攻击,它比ROBOSAC和LUCIA分别提高了高达96.40和67.07个百分点的检测率,同时降低了FPR。

更新时间: 2026-10-06 07:25:07

领域: cs.CR

下载: http://arxiv.org/abs/2610.07875v1

On-Policy Distillation with Negative-Policy Rollouts

On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at https://github.com/naver-ai/np-opd.

Updated: 2026-10-06 07:25:03

标题: 在政策蒸馏中使用负面政策展开

摘要: On-policy distillation (OPD)作为一种广泛研究的训练后方法,其中学生模型在自己的回合中从更强的教师那里获得了令牌级别的监督。最近的研究通过替代蒸馏奖励公式和教师配置改进了OPD,而蒸馏的目标仍然集中在模仿教师。然而,当更强的教师与学生的分布重叠有限时,这种积极的指导可能提供不足的学习信号。在这项工作中,我们引入了Negative-Policy OPD(NP-OPD),它通过来自表现较差、能力较低的负面政策的回合来补充教师监督,这些负面政策充当学生的负面参考。与修改蒸馏奖励公式不同,NP-OPD在回合阶段引入了负面政策,不断提供负面政策喜欢的令牌,以便它们在整个训练过程中仍然暴露于教师监督。这通过负面政策的回合提供了明确的负面信号,同时保留了OPD中使用的积极教师监督。通过大量实验,我们展示了NP-OPD在模型规模、生成模式、推理领域和不同的OPD变体上改善了OPD。此外,我们的分析显示,NP-OPD有效地抑制了负面政策喜欢的令牌,并将学生从负面政策中移开。这些结果支持我们通过负面政策的回合引入负面信号的设计,并为OPD中回合政策的作用提供了新的见解。源代码将在https://github.com/naver-ai/np-opd上提供。

更新时间: 2026-10-06 07:25:03

领域: cs.LG

下载: http://arxiv.org/abs/2610.07874v1

Preparing an AI-Augmented SIEM for the EU Cyber Resilience Act: A Practitioner Case Study

The EU Cyber Resilience Act (CRA), Regulation (EU) 2024/2847, makes product cybersecurity a lifecycle obligation for products with digital elements on the EU market: risk assessment, vulnerability handling, conformity documentation, and Article 14 incident- and vulnerability-reporting readiness must be operational before market placement. Small and medium-sized enterprises that build security products are doubly exposed, since their products are in scope while their customers expect them to be exemplary. This case study documents a CRA preparedness pilot for one such product, SEUXDR, an AI-augmented security monitoring product with a large-language-model active-response component, on the open-source CYBERFORT platform. We contribute a reproducible six-step recipe (Scope and Classify, Asset Registration, Produce Evidence, Map to CRA, Gap and Actions, Audit Pack), two end-to-end traceability threads, and a pilot snapshot tracing product risks through baseline and AI-specific controls and policies to CRA objectives. It offers practitioners a replicable starting point for translating CRA legal text into operational preparedness for incident response, vulnerability reporting, and conformity assessment.

Updated: 2026-10-06 07:24:07

标题: 为欧盟网络安全韧性法案准备一个AI增强的SIEM:一个从业者案例研究

摘要: 欧盟网络安全弹性法案(CRA),即《欧盟》2024/2847号法规,使得在欧盟市场上有数字元素的产品的网络安全性成为产品生命周期的义务:在上市前必须进行风险评估、漏洞处理、符合性文件编制以及第14条事件和漏洞报告准备工作。建造安全产品的中小企业面临双重风险,因为他们的产品在范围内,而客户期望这些产品表现出色。这个案例研究记录了一个为SEUXDR产品进行CRA准备试点的过程,这是一个具有大型语言模型主动响应组件的AI增强安全监控产品,使用开源CYBERFORT平台。我们提供了一个可复制的六步骤配方(范围和分类、资产注册、产生证据、映射到CRA、差距和行动、审计包),两个端到端的可追溯线索,以及一个试点快照,通过基线和AI特定的控制和政策跟踪产品风险,以达到CRA目标。它为从事者提供了一个可复制的起点,用于将CRA法律文本转化为操作准备,以应对事件响应、漏洞报告和符合性评估。

更新时间: 2026-10-06 07:24:07

领域: cs.CR,cs.CY

下载: http://arxiv.org/abs/2610.07873v1

Plug-and-Play Quantum-Resistant BLE Pairing for Medical Implants via NFC Out-of-Band

Bluetooth Low Energy (BLE) pairing establishes the cryptographic foundation for secure device communication. However, mainstream man-in-the-middle (MITM)-resistant pairing methods, such as Numeric Comparison and Passkey Entry, require user interfaces that implantable medical devices (IMDs) inherently lack, making Near-Field Communication (NFC)-assisted Out-of-Band (OOB) pairing an attractive alternative. Existing NFC-assisted OOB schemes authenticate a classical BLE key exchange that remains vulnerable to quantum attacks, while the NFC channel itself is susceptible to eavesdropping and active injection under stronger threat models. To address these limitations, we propose a lightweight, plug-and-play NFC-based OOB pairing protocol that performs a post-quantum key encapsulation mechanism (KEM) entirely over the NFC channel, ensuring that no shared secret is transmitted over NFC while reducing long-range radio-frequency (RF) exposure during pairing. The proposed protocol requires no modifications to the BLE stack and inherently resists RF battery-depletion attacks. Evaluation on an IMD-class proof-of-concept testbed demonstrates that post-quantum OOB pairing is practical on resource-constrained devices, reducing projected battery life by less than 0.57\% for Kyber-1024 and 0.56\% for FireSABER under an operational usage model relative to the MITM-vulnerable Just Works method.

Updated: 2026-10-06 07:22:23

标题: 通过 NFC 带外频道实现医疗植入物的即插即用抗量子配对

摘要: 蓝牙低功耗(BLE)配对建立了安全设备通信的加密基础。然而,主流的抗中间人(MITM)攻击的配对方法,如数字比较和密码输入,需要可植入医疗设备(IMDs)固有缺乏的用户界面,因此近场通信(NFC)辅助的脱机配对成为一种吸引人的替代方案。现有的NFC辅助脱机配对方案验证了一个传统的BLE密钥交换,这仍然容易受到量子攻击的影响,而NFC通道本身则容易受到更强威胁模型下的窃听和主动注入攻击的影响。为了解决这些限制,我们提出了一种轻量级的即插即用基于NFC的脱机配对协议,在NFC通道上完全执行后量子密钥封装机制(KEM),确保在NFC上传输没有共享秘钥,同时减少配对过程中的长距离射频(RF)暴露。所提议的协议不需要对BLE堆栈进行修改,并固有抵抗RF电池耗尽攻击。在IMD类概念验证测试平台上的评估表明,相对于易受MITM攻击的Just Works方法,在操作使用模型下,后量子脱机配对在资源受限设备上是可行的,对于Kyber-1024的预期电池寿命减少不到0.57\%,对于FireSABER的预期电池寿命减少不到0.56%。

更新时间: 2026-10-06 07:22:23

领域: cs.CR

下载: http://arxiv.org/abs/2610.07870v1

Best-of-$N$ Guidance for Test-time Diffusion Alignment

Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-$N$ Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, achieving 1.3$\times$ ImageReward score of the latest sample-based guidance method with a 1.6$\times$ speedup. We release the code at https://github.com/aailab-kaist/BoNG.

Updated: 2026-10-06 07:19:42

标题: 最佳$N$个测试时间扩散对齐指南

摘要: 扩散模型在生成性能方面表现出色,但通常难以将生成的样本与通过奖励模型测量的人类偏好对齐。一种简单而有效的用于测试时间对齐的算法是最佳-$N$ (BoN)采样,它从预训练的扩散模型中抽取$N$个独立同分布的样本,并输出最高奖励的单个样本。尽管BoN在实证成功方面表现出色,但它对奖励信息的利用有限,因为它仅在最终选择阶段才将其纳入,而在采样过程中不影响反向扩散轨迹。因此,BoN采样并不改善生成样本的平均对齐性,并且主要适用于单输出设置。我们提出了最佳-$N$引导(BoNG),这是一种将BoN采样原则直接整合到反向扩散过程中的新方法。BoNG通过对去噪粒子进行在线BoN选择,并调整反向扩散过程,将粒子群体引导至生成过程中更高奖励的区域。具体来说,通过引入去噪粒子之间的非对称引导交互,BoNG将当前的BoN粒子作为引导信号,指导其余粒子群体。这种粒子级交互重塑了采样过程,使其朝向更高奖励的区域,使得BoNG不仅能够改善最终最佳样本,还能改善生成样本的平均质量,超越了Vanilla BoN采样。在36个实证比较中,BoNG在29个案例中取得了最佳表现,在与SMC和Vanilla BoN采样的比较中,排名第一的比例达到了80.56%。BoNG还支持多输出功能,其ImageReward分数是最新的基于样本的引导方法的1.3倍,速度提升了1.6倍。我们在https://github.com/aailab-kaist/BoNG 上发布了代码。

更新时间: 2026-10-06 07:19:42

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.05108v2

PBT-Bench: Benchmarking AI Agents on Property-Based Testing

Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a described issue. Neither isolates the distinct skill of property-based testing: deriving a semantic invariant from documentation, and then constructing an input-generation strategy precise enough to make a random search reveal the violation. We introduce PBT-Bench, a benchmark of 100 curated property-based testing problems across 40 real Python libraries. Each problem injects one or more semantic bugs (365 in total, mean 3.65 per problem) designed so that default-strategy random inputs almost never trigger them; the agent must read the library's documentation, identify the relevant invariant, and specify a Hypothesis @given strategy that concentrates mass in the trigger region. Bugs are stratified across three difficulty levels (L1-L3) spanning single-constraint boundary bugs to stateful, cross-function protocol violations. We evaluate eight contemporary LLMs under two prompting regimes (open-ended baseline vs. explicit Hypothesis scaffolding) for three independent runs per configuration. Bug recall under the PBT-guided prompt ranges from 42.1% to 83.4% across models; under the open-ended baseline, from 31.4% to 76.7%. Hypothesis scaffolding lifts mid-capability models by over 20 percentage points, but yields smaller gains for the strongest models, with two exceptions showing degradation, suggesting the structured prompt can interfere with certain model behaviours rather than complementing them. The hardest bugs prove model-specific: different architectures fail on different problems, leaving persistent gaps that no single model closes. We release the benchmark, harness, and full evaluation corpus to support downstream work on documentation-grounded semantic reasoning.

Updated: 2026-10-06 07:14:56

标题: PBT-Bench:基于属性测试对人工智能代理进行基准测试

摘要: 现有的代码基准测试衡量一个代理能否生成任何能复现已知错误的测试,或者能否生成修复描述问题的补丁。这两者都不能独立衡量基于属性的测试的独特技能:从文档中推导语义不变量,然后构建一个足够精确的输入生成策略,以使随机搜索揭示违规行为。我们引入了PBT-Bench,这是一个跨40个真实Python库的100个经过筛选的基于属性的测试问题的基准测试。每个问题注入一个或多个语义错误(总共365个,平均每个问题3.65个),设计得使默认策略的随机输入几乎不会触发它们;代理必须阅读库的文档,识别相关的不变量,并指定一个在触发区域集中质量的Hypothesis @given策略。错误分层分为三个难度级别(L1-L3),涵盖单一约束边界错误到有状态的、跨函数协议违规。我们对8个当代LLM在两种提示制度下(开放式基线与明确的Hypothesis支架)进行了评估,每种配置进行了三次独立运行。在PBT引导提示下的错误召回率在不同模型间范围从42.1%到83.4%;在开放式基线下,从31.4%到76.7%。Hypothesis支架将中等能力模型提升了超过20个百分点,但对于最强的模型,收益较小,有两个例外表现出恶化,表明结构化提示可能会干扰某些模型行为而不是补充它们。最困难的错误证明是特定于模型的:不同的架构在不同的问题上失败,留下持久的差距,没有一个单一模型可以弥补。我们发布了这个基准测试、工具和完整的评估语料库,以支持基于文档的语义推理的下游工作。

更新时间: 2026-10-06 07:14:56

领域: cs.SE,cs.AI

下载: http://arxiv.org/abs/2605.15229v4

Data-Driven Personas for Survey Simulation: Insights into Simulation Alignment Across Data-Access Regimes

Large language models (LLMs) offer new opportunities for public opinion research by enabling early prediction of survey responses, potentially reducing the cost and time of traditional surveys. However, many existing steering approaches rely on target-domain human data for fine-tuning or prompting that is costly to collect and raises privacy concerns. In this paper, we study demographic group-level survey simulation, where personas induced from heterogeneous, anonymized public behavioral data condition agents that simulate responses of individuals from specific demographic groups. We examine whether representative personas can be induced from diverse sources and analyze how the domain, scale, and granularity of the source data affect survey simulation alignment. We find that personas induced from out-of-domain sources rarely outperform simulations conditioned only on basic demographic information, largely due to population mismatch. However, when personas are accurately assigned to the target demographic groups, alignment improves substantially. Finally, personas induced from target-domain survey data generalize better as more survey question history becomes available, suggesting that richer behavioral evidence enables more stable persona trait inference that transfers to better unseen questions simulation alignment.

Updated: 2026-10-06 07:14:13

标题: 数据驱动的调查模拟人物:对数据访问制度下模拟对齐的洞见

摘要: 大型语言模型(LLMs)通过提供早期预测调查回应的机会,为公众舆论研究带来了新的机遇,潜在地降低了传统调查的成本和时间。然而,许多现有的引导方法依赖于目标领域的人类数据进行微调或提示,这些数据收集成本高昂,并引发了隐私问题。在本文中,我们研究了基于人口统计群体水平的调查模拟,其中从异构、匿名的公共行为数据中诱导出的人物条件代理模拟特定人口统计群体的个体回应。我们检查代表性人物是否可以从不同来源诱导,并分析源数据的领域、规模和粒度如何影响调查模拟对齐。我们发现,从非领域来源诱发的人物很少优于仅根据基本人口统计信息进行条件模拟,这在很大程度上是由于人口不匹配造成的。然而,当人物被准确分配到目标人口统计群体时,对齐会显著改善。最后,从目标领域调查数据诱导出的人物在更多调查问题历史可用时具有更好的泛化能力,这表明更丰富的行为证据使得更稳定的人物特征推断可以转移到更好的未见问题模拟对齐。

更新时间: 2026-10-06 07:14:13

领域: cs.AI

下载: http://arxiv.org/abs/2610.05828v2

The Amplifier Effect: Human-Factor Risks of AI-Suggested Correlation and Auto-Propagation in Multi-Framework GRC Self-Assessment

Multi-framework Governance, Risk and Compliance (GRC) platforms increasingly automate the link between an organisation's self-assessment answer and the compliance obligations that answer is said to satisfy. Cross-framework control mapping, AI-suggested question correlation, and automatic propagation of answers and evidence across correlated questions all serve the legitimate efficiency goal of reducing duplicate work for small and medium-sized enterprises under the EU Cyber Resilience Act, NIS2 and GDPR. The same mechanisms, however, amplify the consequences of any human-factor bias in a single answer: one optimistically-graded control, one rubber-stamped attestation, or one AI-drafted answer can be silently replicated as evidence of compliance with many obligations across multiple frameworks. We call this the amplifier effect: a platform-design property (coarse-grained attestation and un-gated propagation) rather than a failing of individual users. Using two EU-funded SME-facing GRC platforms, CYBERFORT and CYBER-BRIDGE, as examples, we (i) describe the amplification mechanism in concrete data-model terms, (ii) propose a six-dimension scoring framework for evaluating any GRC tool's exposure to the effect, (iii) instantiate the framework on a thirteen-tool comparison covering enterprise IRM, mid-market platforms, compliance-automation tools, and the two EU SME projects, and (iv) outline a measurement protocol that a consortium with access to production self-assessment data can run. The thirteen-tool comparison is a structured design assessment, not an empirical measurement of user behaviour. The EU SME platforms score lowest on the amplifier dimensions because their burden-reduction design deliberately trades sign-off granularity for throughput; we report this as a design trade-off, not a verdict on the platforms. Our contribution is the framing and the measurement protocol.

Updated: 2026-10-06 07:12:30

标题: 放大器效应:多框架GRC自我评估中AI建议的相关性和自动传播的人因风险

摘要: 多框架治理、风险和合规性(GRC)平台越来越多地自动化组织的自我评估答案与这些答案所满足的合规义务之间的联系。跨框架控制映射、AI建议的问题相关性以及答案和证据在相关问题之间的自动传播,都服务于降低欧盟《网络安全韧性法案》、NIS2和GDPR下中小企业重复工作的合法效率目标。然而,同样的机制也放大了单个答案中任何人为因素偏见的后果:一个乐观评分的控制、一个橡皮图章的保证或一个AI起草的答案可以被默默复制为对多个框架下的多个义务的合规证据。我们将此称为放大器效应:这是一个平台设计属性(粗粒度保证和非门控传播),而不是个体用户的失败。以两个欧盟资助的面向中小企业的GRC平台CYBERFORT和CYBER-BRIDGE为例,我们(i)用具体的数据模型术语描述放大机制,(ii)提出一个评估任何GRC工具暴露于该效应的六维评分框架,(iii)在一个涵盖企业IRM、中端市场平台、合规自动化工具以及两个欧盟中小企业项目的十三种工具比较中实例化该框架,并(iv)概述一个负责运行具有生产自我评估数据访问权限的财团的测量协议。这十三种工具的比较是一个结构化的设计评估,而不是对用户行为的实证测量。欧盟中小企业平台在放大器维度上得分最低,因为它们的减负设计故意以签署粒度换取吞吐量;我们将其报告为一个设计权衡,而不是对平台的裁决。我们的贡献是框架和测量协议。

更新时间: 2026-10-06 07:12:30

领域: cs.CR,cs.CY,cs.HC

下载: http://arxiv.org/abs/2610.07866v1

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.

Updated: 2026-10-06 07:09:54

标题: ReFold:长期代理的无需训练可逆转环境上下文折叠

摘要: 长期视野的LLM代理人在每一步都会对一个只追加的交互历史进行操作,并在每一步重新发送给模型,因此上下文及其成本会随着步数增加直到会话超过上下文窗口。现有方法通过上下文需求预测来管理上下文,依赖于额外的模型调用、启发式规则或训练好的策略。然而,这些预测方法会引入运行时开销,使前缀缓存失效,并且永久丢弃内容而无法保证恢复。为了克服这些限制,我们介绍了ReFold:一种无需训练的渲染层,它保留了底层的交互历史,同时只压缩模型的已渲染上下文。它消除了两种类型的转弯间冗余,而无需辅助预测器:一个早期转弯已经显示的内容被替换为占位符,以及代理人报告已完成的转弯被折叠成一行注释。这两个操作符使用分块渲染,每隔几个步骤重写一次缓存的前缀,而不是在每一步都这样做。每次删除都是完全可逆的,一个错误的删除只会造成一次从历史中恢复,而不是永久丢失内容。由于它在渲染层操作,ReFold可以在标准的ReAct风格的接口中轻松安装和使用。在五个长期视野基准和两个前沿LLM的评估中,ReFold将令牌消耗降低了最多2.5倍,每个会话的KV缓存内存减半,而不降低任务成功率。在有上限的上下文预算下,它避免了高达92%的强制压缩。在并发服务工作负载下,它将请求排队延迟减少了最多100%,推动推理加速最多1.7倍,同时将推理成本降低了最多3.4倍。

更新时间: 2026-10-06 07:09:54

领域: cs.CL,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.07863v1

A self-learning scientific agent for X-ray diffraction

A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, without retraining the language model or changing the underlying physical models. Skills selected using development data and frozen before held-out evaluation achieve higher refinement scores than the original expert-designed skills across FullProf, GSAS-II and PyWPEM. The agent resolves strongly overlapping reflections, quantifies a five-phase ancient Egyptian cosmetic, tracks lattice evolution in an operating battery and compares atomic configurations in a disordered oxide catalyst. On DeltaXRDbench, it leads the evaluated methods in single- and multiphase identification across simulated and experimental data. Without supplied composition, single-phase top-1 accuracies reach 96.30\%, 81.78\% and 40.83\% on MP500, RRUFF and opXRD, respectively, compared with 58.00\%, 58.47\% and 26.45\% for the strongest comparator. These results demonstrate how an integrated scientific tool ecosystem can support agents that extract structural knowledge from measurements while accumulating validated analytical expertise that transfers to new samples.

Updated: 2026-10-06 07:09:39

标题: 一个用于X射线衍射的自学科学代理

摘要: 科学性代理面临的一个核心挑战是将分析经验转化为基于物理证据的可重复利用的专业知识。在这里,我们介绍了甘江(Gan Jiang),这是一个基于我们开发的衍射分析生态系统构建的粉末X射线衍射的自学习代理:XMatcher、XQueryer、XDecomposer和WPEM。这些引擎共同涵盖了相位识别、多相分解和受物理约束的整体图案建模。甘江通过诊断失败、修订技能说明和代码,并在重复使用之前验证修订,将分析经验转化为可执行技能,而无需重新训练语言模型或更改基础物理模型。在开发数据中选择技能并在留置评估之前冻结,可以获得比原始专家设计的技能在FullProf、GSAS-II和PyWPEM上的更高的优化分数。该代理解决了严重重叠的反射,量化了五相古埃及化妆品,追踪了运行中电池中的晶格演变,并比较了无序氧化物催化剂中的原子配置。在DeltaXRDbench上,它在模拟和实验数据中的单相和多相识别中领先评估方法。在没有提供组成的情况下,与最强的比较器相比,MP500、RRUFF和opXRD上的单相top-1准确率分别达到了96.30%、81.78%和40.83%,而最强比较器为58.00%、58.47%和26.45%。这些结果展示了一个集成的科学工具生态系统如何支持能够从测量中提取结构知识的代理,并累积验证的分析专业知识,这些专业知识可以转移到新样本中。

更新时间: 2026-10-06 07:09:39

领域: cond-mat.mtrl-sci,cs.AI,cs.LG

下载: http://arxiv.org/abs/2610.07862v1

SWE-Game: Can Coding Agents Build the Games We Want?

We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 gameplay clips. Together, these results characterize current agent capabilities across game-development activities and support combining runtime evidence with visual assessment.

Updated: 2026-10-06 07:03:14

标题: SWE-Game:编码代理能否构建我们想要的游戏?

摘要: 我们介绍了SWE-Game,这是一个基于41个可执行的参考Godot游戏构建的247个任务的基准,涵盖了2D和3D中的13种游戏玩法类别。五种任务类型涵盖了从简要开发,根据游戏设计文件实现,骨架完成,修复83个注入故障案例,以及Godot到Unity的移植。参考资料指定了预期的游戏玩法,而共享的仪器接口让评估者拥有的驱动程序和探针执行操作并独立观察实现的游戏。评估结合了引擎状态检查,认证的参考输入重放,以及代理作者编写的特性演示,以评估机械正确性,展示的可玩性,以及修复后的行为恢复和保留。游戏特定的视觉语言规则单独评估呈现。在六个模型中,Opus5在所有五种任务类型中获得了最高的总分。在三个构建任务中,最佳总分仍然低于100分中的60分,其中Brief-to-Game达到50.38分。对审核提交的分析确定了需求遗漏和游戏逻辑错误作为主要的实现问题。从100个代理构建的游戏中人工标记的行为中,可执行检查达到92.59%的平衡准确度,而基于视频的VLM评委为78.41%。基于规则的视觉评分与200个游戏片段的人工评分达到0.829的Spearman相关性。总的来说,这些结果描述了当前代理在游戏开发活动中的能力,并支持将运行时证据与视觉评估相结合。

更新时间: 2026-10-06 07:03:14

领域: cs.AI

下载: http://arxiv.org/abs/2609.33678v2

KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy

Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More than 35 clinicians contributed to case authoring and benchmark evaluation. In an empirical study, clinicians gave simulated dialogues higher mean quality ratings than reference conversations, which is adapted from real conversation. In each task, an LM has a fixed budget of turns to communicate with the patient, ask about relevant history, request examinations, follow action constraints, and record a final diagnosis. We score these steps separately as well as together. Across 31 models and seven model families, the best-performing models (e.g., GPT-6-astra and Claude Opus 5) succeed on less than 30% of tasks, even though their diagnosis accuracy reaches 90.7%. Some models benefit from talking with the patient; others diagnose well from a complete chart but perform much worse in conversation. Overall, KlinikeBench provides a testbed for evaluating the full clinical encounter and reveals a substantial gap between diagnostic accuracy and performance in interactive clinical assessment. All the code and data is available on https://zehui127.github.io/klinikebench/

Updated: 2026-10-06 07:01:43

标题: KlinikeBench: 评估语言模型的方式超越诊断准确性

摘要: 大多数临床基准评估语言模型(LMs)在诊断上使用完整的病例描述。然而,在临床实践中,患者以不同的方式呈现信息,临床医生必须获取相关病史并确定需要进行哪些检查才能做出诊断。因此,仅凭诊断准确性无法确定代理人是否收集到必要信息或进行了适当的临床评估。此外,现有的基准缺乏专业临床医生的验证。为了填补这一差距,我们介绍了KlinikeBench,这是一个包含333个由临床医生撰写的任务的基准,每个任务提供一个孤立的沙盒环境,包括虚拟患者、临床工具和任务特定的成功标准。超过35名临床医生参与了案例编写和基准评估。在一项实证研究中,临床医生给予模拟对话更高的平均质量评分,比起来自真实对话的参考对话。在每个任务中,LM与患者进行通信的固定轮次预算,询问相关病史,请求检查,遵循行动约束,并记录最终诊断。我们分别对这些步骤进行评分,以及合并评分。在31个模型和七个模型系列中,表现最佳的模型(例如GPT-6-astra和Claude Opus 5)在任务中的成功率不到30%,尽管其诊断准确性达到90.7%。一些模型受益于与患者交谈;而其他一些则从完整的图表中诊断良好,但在对话中表现得糟糕得多。总的来说,KlinikeBench提供了一个评估完整临床会诊的测试平台,并揭示了诊断准确性与互动性临床评估表现之间存在重大差距。所有代码和数据都可在https://zehui127.github.io/klinikebench/上获得。

更新时间: 2026-10-06 07:01:43

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2609.38480v2

WorkflowOps: Learning Agent Collaboration Priors for Multi-Agent Workflow Orchestration

Multi-agent systems are increasingly deployed for complex knowledge work, yet their orchestration layers remain largely memoryless: each new task is decomposed, assigned, and executed from scratch with no benefit from prior successful executions. We present WorkflowOps, a multi-agent workflow orchestration framework that learns agent collaboration priors from historical workflows and expands its agent pool on demand to cover new capability requirements. Our approach introduces three coupled mechanisms. First, a transition probability matrix captures pairwise agent collaboration frequencies from past workflows and applies them as soft guidance during DAG workflow construction through intra-layer ordering optimization, probability-thresholded edge suggestion, and transitive reduction for parallelism maximization. Second, a sufficiency-driven agent creation loop detects capability gaps via semantic matching scores, generates specialized agents through an LLM, and simultaneously injects them into the collaboration matrix, so that newly created agents are immediately usable with predicted collaboration priors. Third, a layered semantic matching strategy uses pre-trained sentence embeddings for fast, deterministic capability matching as a first pass, invoking LLM verification only for low-confidence cases, thereby reducing LLM routing calls by over 80\% compared to pure-LLM approaches. Experiments on mixed code, math, and question-answering suites show that WorkflowOps improves end-to-end pass rates over recent workflow-construction baselines, with the largest gains on structured, decomposable tasks where past agent handoff patterns transfer.

Updated: 2026-10-06 07:01:38

标题: WorkflowOps:学习代理协作先验知识以用于多代理工作流编排

摘要: 多智能体系统越来越多地用于复杂的知识工作,但它们的编排层仍然主要是无记忆的:每个新任务都是从头开始分解、分配和执行的,没有从先前成功的执行中受益。我们提出了WorkflowOps,这是一个多智能体工作流编排框架,它从历史工作流中学习智能体协作先验知识,并根据需求扩展其智能体池以满足新的能力要求。我们的方法引入了三种耦合机制。首先,一个转移概率矩阵捕捉了过去工作流中的智能体协作频率,并将它们作为软指导应用于DAG工作流构建中,通过层内次序优化、概率阈值化的边建议和传递约减以最大化并行性。其次,一种基于足够性驱动的智能体创建循环通过语义匹配分数检测能力缺口,通过LLM生成专门的智能体,并同时将它们注入到协作矩阵中,这样新创建的智能体就可以立即使用与预测的协作先验知识。第三,一种分层语义匹配策略使用预先训练的句子嵌入进行快速、确定的能力匹配作为第一遍,仅在置信度较低的情况下调用LLM验证,从而与纯LLM方法相比,将LLM路由调用减少了80\%以上。在混合代码、数学和问答套件上的实验表明,WorkflowOps相对于最近的工作流构建基线,提高了端到端通过率,在结构化、可分解的任务上获得了最大收益,其中过去的智能体交接模式转移。

更新时间: 2026-10-06 07:01:38

领域: cs.AI

下载: http://arxiv.org/abs/2610.07860v1

Tram-FL: Reducing Communication and Computation Costs through Sequential Model Circulation in Decentralized Federated Learning

Conventional decentralized federated learning (DFL) often focuses on clients, with each client maintaining a model copy, performing updates individually, and undertaking model exchange and integration. While fully leveraging computational resources can shorten training times, it can also lead to significant computational and communication waste. This is especially pronounced with non-independent and identically distributed (non-IID) data, where achieving high model accuracy demands extra resources. This research shifts focus to the model itself, aiming to realize DFL with minimal computation and communication costs. To this end, we propose Tram-FL (Traveling Model Training Mechanism for Decentralized Federated Learning), a mechanism designed to efficiently address these challenges. It sequentially trains a single model by circulating it among nodes. We address the training scheduling problem in model circulation-based training, specifically determining which nodes should update the model and the number of updates to perform. This is approached by considering the model's circulation route and update iteration allocation, for which we propose simple yet effective methods. Additionally, with quantized momentum, Tram-FL achieves high accuracy with fewer model circulations while controlling communication load per transmission. Experimental results show that the proposed algorithm, even with non-IID data, converges to a global model with reduced communication and computation.

Updated: 2026-10-06 07:01:01

标题: Tram-FL:通过在分散式联邦学习中实现顺序模型传递来减少通信和计算成本

摘要: 传统的分散式联邦学习(DFL)通常专注于客户端,每个客户端维护一个模型副本,独立进行更新,并进行模型交换和集成。充分利用计算资源可以缩短训练时间,但也可能导致显着的计算和通信浪费。这在非独立同分布(non-IID)数据中尤为突出,其中实现高模型准确性需要额外资源。本研究将重点转移到模型本身,旨在实现最小的计算和通信成本的DFL。为此,我们提出了Tram-FL(分散式联邦学习的旅行模型训练机制),这是一个旨在有效解决这些挑战的机制。它通过在节点之间循环传递一个单一模型来顺序训练模型。我们解决了基于模型循环的训练中的训练调度问题,特别是确定哪些节点应该更新模型以及执行更新的次数。通过考虑模型的循环路线和更新迭代分配,我们提出了简单而有效的方法。此外,使用量化的动量,Tram-FL在控制每次传输的通信负载的同时实现了高准确性。实验结果显示,即使在非IID数据下,所提出的算法也能收敛到具有减少通信和计算的全局模型。

更新时间: 2026-10-06 07:01:01

领域: cs.LG,cs.DC,cs.NI

下载: http://arxiv.org/abs/2610.07859v1

A Decision-Focused Neural Optimization Framework for Personalized Route Reproduction from Vehicle Trajectories

This study formulates individual route reproduction as a shortest-path problem over learned driver-specific latent link costs. The central idea is that, once such latent costs are inferred from contextual information, observed routes can be reproduced without enumerating alternative route sets. We propose a neural pipeline that includes a perception model that embeds context covariates, which comprises individual characteristics, trip-specific attributes, and network-level traffic states, into the personalized link costs. A constrained optimization (CO) layer, which determines the shortest path (SP) based on these estimated costs, follows the perception encoder. To enable end-to-end training, we employ decision-focused learning to align the predicted shortest paths with observed routes. The implicit maximum likelihood estimation (iMLE) provides an approximate gradient of the loss function that contains the non-differentiable CO layer. Furthermore, a regularization term anchors the latent cost distribution to the empirical scale of observed link travel times, mitigating the scale ambiguity inherent in shortest-path supervision. Empirical evaluations demonstrate that the proposed framework outperforms baseline route choice models in path reproduction. The learned latent costs, interpreted as proxies for perceived travel costs, provide plausible explanations for heterogeneous route choices.

Updated: 2026-10-06 07:00:49

标题: 一个以决策为中心的神经优化框架,用于根据车辆轨迹实现个性化路线重现

摘要: 这项研究将个体路径再现构建为基于学习的驾驶员特定潜在链路成本的最短路径问题。中心思想是,一旦从上下文信息中推断出这些潜在成本,观察到的路径就可以在不枚举替代路径集的情况下再现出来。我们提出了一个神经网络管道,其中包括一个感知模型,该模型将上下文协变量(包括个体特征、特定旅行属性和网络级交通状态)嵌入到个性化链路成本中。一个受约束的优化(CO)层,根据这些估计的成本确定最短路径(SP),紧随感知编码器。为了实现端到端的训练,我们采用决策聚焦学习来使预测的最短路径与观察到的路径对齐。隐式最大似然估计(iMLE)提供了包含不可微分CO层的损失函数的近似梯度。此外,正则化项将潜在成本分布锚定到观察到的链路旅行时间的经验尺度,减轻最短路径监督中固有的尺度模糊性。经验评估表明,所提出的框架在路径再现方面优于基线路径选择模型。学习的潜在成本被解释为感知旅行成本的代理,为异质路径选择提供了合理的解释。

更新时间: 2026-10-06 07:00:49

领域: cs.LG

下载: http://arxiv.org/abs/2610.07857v1

Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations

Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two paradigms. Reconstruction-based methods recover missing information from observed modalities, while joint-representation methods learn directly from incomplete inputs. Although effective, these methods usually treat modality reliability only implicitly within representation learning or fusion design rather than modeling it explicitly. We argue that modality reliability is a central variable in incomplete-observation settings. Failure to model it explicitly gives rise to two related issues. The first is reliability mismatch, in which the affective evidence retained by each modality varies across samples and missing rates. The second is reliability propagation bias, in which messages from degraded modalities may adversely affect cross-modal interaction and predictive performance. To address these issues, we propose MRCF, a Modality Reliability-Calibrated Framework for MSA with incomplete observations. MRCF contains a Reliability-Aware Branch that estimates sample-specific modality reliability from intramodal quality cues and cross-modal semantic consistency, a Reliability-Guided Interaction Branch that uses the estimated scores to modulate cross-modal information flow, and a Reliability-Calibrated Fusion Module that integrates reliability and semantic cues for final prediction. Experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS show that MRCF achieves strong performance under standard incomplete-observation protocols. Further analyses provide evidence that explicit reliability modeling helps mitigate reliability mismatch and reliability propagation bias during interaction and fusion.

Updated: 2026-10-06 07:00:00

标题: 重新思考带有不完整观测的多模态情感分析中的模态可靠性

摘要: 多模态情感分析(MSA)集成文本、音频和视觉来推断人类情感,然而现实世界中的多模态观察往往是不完整的。现有的不完整观察MSA方法主要遵循两种范式。基于重建的方法从观察到的模态中恢复缺失的信息,而联合表示方法直接从不完整的输入中学习。尽管有效,这些方法通常只在表示学习或融合设计中隐式处理模态可靠性,而不是明确对其进行建模。我们认为模态可靠性是不完整观察环境中的一个核心变量。未能明确对其进行建模会导致两个相关问题。第一个是可靠性不匹配,在这种情况下,每个模态保留的情感证据在样本和缺失率之间变化。第二个是可靠性传播偏差,在这种情况下,从受损模式传递的信息可能会对跨模态交互和预测性能产生不利影响。为了解决这些问题,我们提出了MRCF,一种用于具有不完整观察的MSA的模态可靠性校准框架。MRCF包含一个可靠性感知分支,从模态内部质量线索和跨模态语义一致性估计样本特定的模态可靠性,一个可靠性引导交互分支,利用估计的分数来调节跨模态信息流,以及一个可靠性校准融合模块,将可靠性和语义线索集成到最终预测中。在CMU-MOSI、CMU-MOSEI和CH-SIMS上的实验表明,MRCF在标准不完整观察协议下取得了强大的性能。进一步的分析证明,明确的可靠性建模有助于减轻交互和融合过程中出现的可靠性不匹配和可靠性传播偏差。

更新时间: 2026-10-06 07:00:00

领域: cs.AI,cs.MM

下载: http://arxiv.org/abs/2608.03611v3

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We operationalize experimental research taste as compute efficiency; a Researcher who reaches the same score as an expert human using half the serial experimental compute has twice the experimental taste. Experimental taste thus acts as a multiplier on experimental compute, making it a key input to forecasts of AI progress. TasteVal consists of 8 novel, challenging, open-ended tasks representative of frontier AI R&D. To isolate taste from coding ability, the model under evaluation acts as a Researcher that iteratively designs experiments while a fixed Coder agent implements them and reports their results. The Researcher executes until either the 40 H100 hour or 120 wall-clock hour budgets are exhausted. We recruit 24 human experts, at least 2 per task, and take the best expert attempt per task as the expert baseline. We evaluate 20 models released between 2023 and 2026. The best-performing model, Opus 5.5, exceeds our expert baseline, with a compute multiplier of 2.3x (95% CI 1.15-4.37), at roughly 1/30 of our baseliners' average per-run cost. On TasteVal, the compute multiplier of frontier models has doubled approximately every 3.0 months since December 2025 (95% CI 1.7-5.0), up from every 14 months between 2023 and December 2025. Measured by final normalized performance, frontier models show no trend break, doubling every 14.6 months. To keep TasteVal uncontaminated, we do not release the tasks.

Updated: 2026-10-06 07:00:00

标题: TasteVal:衡量人工智能系统实验研究的口味与人类专家的对比

摘要: 我们介绍了TasteVal,这是一个用于评估前沿模型实验研究品味的基准。我们将研究品味定义为选择有趣问题、设计实验和解释实验结果的能力。TasteVal衡量了研究品味的实验组成部分;在一个固定的研究问题下,我们衡量一个模型如何迭代地设计实验并从实验结果中得出结论。我们将实验研究品味操作化为计算效率;一个研究者如果能够在使用一半串行实验计算的情况下达到与专家人类相同的分数,则其实验品味是专家的两倍。实验品味因此作为实验计算的乘数,成为AI进展预测的关键输入。TasteVal由8个具有代表性的前沿AI研发任务组成,这些任务具有新颖、具有挑战性、开放性。为了将品味与编码能力分离,评估模型作为一个研究者,迭代设计实验,而一个固定的编码代理人实现它们并报告结果。研究者执行,直到40 H100小时或120挂钟小时的预算用尽。我们招募了24位专家,每个任务至少2位,并将每个任务的最佳专家尝试作为专家基线。我们评估了2023年至2026年间发布的20个模型。表现最好的模型Opus 5.5超过了我们的专家基线,计算乘数为2.3倍(95% CI 1.15-4.37),大约是我们基准模型平均每次运行成本的1/30。在TasteVal上,前沿模型的计算乘数自2025年12月以来大约每3.0个月翻倍一次(95% CI 1.7-5.0),而在2023年至2025年12月间是每14个月一次。根据最终标准化性能来衡量,前沿模型没有趋势突破,每14.6个月翻倍一次。为了保持TasteVal的纯洁性,我们不公开任务内容。

更新时间: 2026-10-06 07:00:00

领域: cs.AI

下载: http://arxiv.org/abs/2610.06824v2

Rate-Optimal Algorithm for Adversarial Linear CMDPs

We study episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions, where both the loss and constraint functions may vary adversarially across episodes. The best previous algorithm achieves $\widetilde{\mathcal{O}}(K^{3/4})$ regret and cumulative constraint violation, leaving a gap to the optimal $\widetilde{\mathcal{O}}(\sqrt{K})$ dependence on the number of episodes $K$. We close this gap by proposing a new primal dual algorithm that achieves $\widetilde{\mathcal{O}}(\sqrt{K})$ regret and cumulative constraint violation without assuming Slater's condition. We further extend the algorithm to achieve the same $\widetilde{\mathcal{O}}(\sqrt{K})$ guarantees for regret and hard constraint violation, which does not allow constraint violations to cancel across episodes. The main challenge is that learning linear CMDPs requires uniform concentration over a value function class with a controlled covering number, whereas standard techniques in constrained online learning, such as policy mixing, can make this class more complex. Our algorithm combines adaptive Follow the Regularized Leader (FTRL), contracted value estimation, and an exponential Lyapunov function. An adaptive dual regularizer offsets the dependence on the dual weights in the primal regret bound, removing the need for policy mixing. We further show that the normalization in the FTRL update bounds the policy parameters independently of the magnitudes of the dual weights, which explains why the resulting policy class remains compatible with uniform concentration. Under feature access, the computational complexity is independent of the size of the state space.

Updated: 2026-10-06 06:59:55

标题: 对抗线性CMDPs的速率最优算法

摘要: 我们研究具有未知转换的阶段性对抗线性约束马尔可夫决策过程(CMDPs),在这种情况下,损失和约束函数在不同阶段可能在对抗性上变化。最佳先前算法实现了$ \widetilde{\mathcal{O}}(K^{3/4}) $的后悔和累积约束违规,与对于阶段数量$ K $的最佳$ \widetilde{\mathcal{O}}(\sqrt{K}) $依赖之间存在差距。我们通过提出一种新的原始对偶算法来弥合这一差距,实现了$ \widetilde{\mathcal{O}}(\sqrt{K}) $的后悔和累积约束违规,而不需要假设Slater's条件。我们进一步将该算法扩展到实现对于后悔和硬约束违反的相同$ \widetilde{\mathcal{O}}(\sqrt{K}) $保证,这不允许在不同阶段之间取消约束违规。主要挑战在于学习线性CMDPs需要在具有受控覆盖数量的值函数类上实现均匀集中,而标准的约束在线学习技术,如策略混合,可能使该类更加复杂。我们的算法结合了自适应Follow the Regularized Leader(FTRL),收缩值估计和指数Lyapunov函数。自适应对偶正则化器抵消了原始后悔界中对偶权重的依赖性,从而消除了策略混合的需要。我们进一步展示,FTRL更新中的归一化将策略参数的边界与对偶权重的幅度独立起来,这解释了为什么所得到的策略类仍与均匀集中兼容。在特征访问下,计算复杂度与状态空间的大小无关。

更新时间: 2026-10-06 06:59:55

领域: cs.LG

下载: http://arxiv.org/abs/2610.00927v2

Lifecycle-Based Design and Evaluation of Real-Time Backup Triggers for Ransomware Damage Mitigation

Ransomware continues to encrypt files during the interval between attack onset and detection. Real-time backups can mitigate this damage by preserving files before they are modified. The previously proposed Real-Time Open-File Backup System (ROFBS) triggers backups primarily on file-open events. However, the file lifecycle offers several candidate trigger points, including open, read, write, and rename operations. Triggering backups too early may create unnecessary backup files, whereas triggering them too late may allow ransomware writes to race with backup creation and prevent the preservation of clean file contents. Consequently, it remains unclear which trigger timing best balances recoverability and the number of backups created. In this study, we design and evaluate real-time backup triggers for mitigating ransomware damage from a file-lifecycle perspective. Specifically, we compare four strategies: Open-time backup, Read-time backup, Write-time backup, and Rename-time backup. We implement these strategies in an ROFBS-style prototype on XFS and evaluate them using five ransomware samples: Conti, Sodinokibi, AvosLocker, REvil, and HelloKitty. Our results clarify how trigger timing affects both damage mitigation and the number of backups created, providing design guidance for selecting effective triggers in real-time backup systems against ransomware.

Updated: 2026-10-06 06:59:38

标题: 基于生命周期的设计和评估实时备份触发器用于勒索软件损害缓解

摘要: 勒索软件在攻击发生和检测之间的时间间隔内继续加密文件。实时备份可以通过在文件被修改之前保存文件来减轻此损害。先前提出的实时开放文件备份系统(ROFBS)主要在文件打开事件上触发备份。然而,文件生命周期提供了几个候选触发点,包括打开、读取、写入和重命名操作。过早触发备份可能会产生不必要的备份文件,而过晚触发备份可能会导致勒索软件的写入与备份创建竞争,从而阻止干净文件内容的保存。因此,目前尚不清楚哪种触发时间最好地平衡了可恢复性和创建的备份数量。在这项研究中,我们从文件生命周期的角度设计和评估实时备份触发器,以减轻勒索软件造成的损害。具体来说,我们比较了四种策略:打开时间备份、读取时间备份、写入时间备份和重命名时间备份。我们在XFS上实现了这些策略的ROFBS风格原型,并使用五个勒索软件样本进行评估:Conti、Sodinokibi、AvosLocker、REvil和HelloKitty。我们的结果阐明了触发时间如何影响损害减轻和创建备份数量,为在实时备份系统中选择有效触发器提供设计指导,以抵御勒索软件。

更新时间: 2026-10-06 06:59:38

领域: cs.CR

下载: http://arxiv.org/abs/2610.07854v1

Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16. We audit those pipelines across three labs. In released checkpoints, fp32 quantization of the shipped latents disagrees with the deployed codes on 0.83-1.77% of codes in Falcon-E and BitCPM and on 1.530% in BitNet 2B-4T; for Falcon-E and BitCPM most disagreements are products that bf16 rounding lands exactly on the threshold, which ties-to-even maps to zero, and the unmodified onebitllms exporter reproduces all four Falcon-E releases byte for byte. At fine-tuned endpoints, with learning rates selected to match a nominal learning-rate-to-bf16-ULP ratio, the documented export lowers greedy GSM8K strict accuracy from 58.79% to 0.78% for Falcon-E-1B-Base and from 36.13% to 0.39% for BitCPM-CANN-0.5B, and a bf16 save and reload lowers BitNet 2B-4T's strict accuracy by 27.54 points while its last-number accuracy rises. Two compatibility remedies, writing the training quantizer's codes directly or adjusting the bf16 inputs until the unchanged tools emit them, each met a 4-point strict-accuracy non-inferiority criterion against online evaluation in all three models. In two model families, randomized interventions on the initial distance from the threshold support distance-dependent selection of the codes that fine-tuning changes.

Updated: 2026-10-06 06:58:35

标题: 在bf16 Cast中迷失:导出三元语言模型可以还原大多数低学习率代码更改

摘要: 三元语言模型,如BitNet b1.58,Falcon-E和BitCPM,经过高精度潜在权重的微调,并作为由出口步骤产生的三元代码部署,该步骤在实验室记录的流程中,首先将潜在权重转换为bf16。我们在三个实验室中审计这些流程。在发布的检查点中,所发出的潜在权重的fp32量化与Falcon-E和BitCPM中的部署代码有0.83-1.77%的不一致,而在BitNet 2B-4T中有1.530%的不一致;对于Falcon-E和BitCPM,大多数不一致是由于bf16舍入恰好落在阈值上,这会映射为零,并且未经修改的onebitllms导出器逐字地复制了所有四个Falcon-E版本。在微调的端点上,选择的学习速率与名义学习速率与bf16-ULP比率相匹配,记录的出口将Falcon-E-1B-Base的贪婪GSM8K严格准确度从58.79%降低到0.78%,将BitCPM-CANN-0.5B的准确度从36.13%降低到0.39%,并且bf16保存和重新加载将BitNet 2B-4T的准确度降低了27.54个百分点,而其最后一个数字的准确度升高。两种兼容性疗法,直接编写训练量化器的代码或调整bf16输入,直到未更改的工具发出它们,每个对所有三个模型在线评估满足4点严格准确度非劣性标准。在两个模型系列中,对距离阈值的初始距离进行随机干预,支持根据距离进行距离依赖选择进行微调更改的代码。

更新时间: 2026-10-06 06:58:35

领域: cs.LG,cs.CL

下载: http://arxiv.org/abs/2610.07853v1

RA-MoWE: Workflow-Affinity Embeddings for Query Clustering and Agentic Workflow Generation

Agentic workflows enable large language models (LLMs) to solve complex tasks by coordinating reasoning, tool use, and verification. However, a workflow optimized for an entire task collection can overlook differences in the reasoning strategies that individual queries need, while searching for a new workflow for every query repeats costly optimization. To address this tradeoff, we introduce RA-MoWE, a framework that uses workflow-affinity embeddings to cluster queries and guide the generation of reusable expert workflows. Each embedding records how well a fixed set of reference workflows solves a query, revealing similarities in which reasoning strategies are effective. RA-MoWE uses each cluster's queries and average embedding to initialize and refine a specialized workflow through execution feedback. An embedding encoder predicts these embeddings from query text, allowing new queries to select a generated expert without first executing the reference workflows. On a 300-query test set drawn from four benchmarks spanning mathematics, science, and programming, RA-MoWE improves average task score by 4.04 percentage points over selecting among the reference workflows, while using 27.7% fewer language-model calls at inference.

Updated: 2026-10-06 06:58:16

标题: RA-MoWE:用于查询聚类和代理工作流生成的工作流亲和嵌入

摘要: 主动工作流使大型语言模型(LLMs)能够通过协调推理、工具使用和验证来解决复杂任务。然而,为整个任务集合优化的工作流可能会忽视个别查询所需的推理策略的差异,同时为每个查询搜索新工作流会重复昂贵的优化过程。为了解决这种权衡,我们引入了RA-MoWE,这是一个利用工作流亲和力嵌入来对查询进行聚类并指导生成可重复使用的专家工作流的框架。每个嵌入记录了一组固定参考工作流如何解决一个查询,揭示了哪种推理策略是有效的相似性。RA-MoWE使用每个集群的查询和平均嵌入来初始化和优化专门的工作流,通过执行反馈。一个嵌入编码器从查询文本中预测这些嵌入,允许新查询在首次执行参考工作流之前选择生成的专家。在从覆盖数学、科学和编程的四个基准绘制的300个查询测试集上,RA-MoWE将平均任务得分提高了4.04个百分点,同时在推断时减少了27.7%的语言模型调用。

更新时间: 2026-10-06 06:58:16

领域: cs.AI

下载: http://arxiv.org/abs/2610.07851v1

MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development

Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic tools alone does not ensure that agents learn when to use them or how to act on their findings. We introduce ToolMLBench, a suite of executable tools for data inspection, code verification, and experiment diagnosis, together with an SFT and RL pipeline for learning their use. Diagnostic calls acquire evidence whose value depends on subsequent decisions, so final outcomes provide limited guidance on which calls to reinforce. We address this challenge with SPICE, which measures how privileged context changes the likelihood of a sampled tool action and uses this difference as a turn-level reward alongside the final outcome. We train on 80 synthetic tasks and evaluate on 25 in-domain and 10 out-of-domain tasks. Providing tool interfaces and descriptions alone yields inconsistent gains across unadapted models. With the same diagnostic interface, our training pipeline raises in-domain success from 24.8% to 52.4% for Qwen3-8B and from 35.6% to 69.2% for Qwen3.5-35B-A3B. The latter also improves from 31% to 48% out-of-domain, supporting learned diagnostic tool use on held-out sources and targets.

Updated: 2026-10-06 06:57:34

标题: MLToolBench:用于机器学习开发的学习工具增强型代理

摘要: 机器学习工程(MLE)代理取得了实质性的进展,但通过机器学习实验学习仍然耗时且计算成本高昂。合成环境可以降低这些成本,同时引入数据和实验设置的变化,需要特定任务的诊断。仅仅提供诊断工具并不能确保代理何时使用它们或如何根据发现行动。我们引入了ToolMLBench,一个用于数据检查、代码验证和实验诊断的可执行工具套件,以及一个用于学习它们的SFT和RL管道。诊断调用获取证据,其价值取决于后续决策,因此最终结果对于加强哪些调用提供了有限的指导。我们通过SPICE解决了这一挑战,该方法衡量特权上下文如何改变抽样工具动作的可能性,并将此差异作为转折级别奖励,与最终结果一起。我们在80个合成任务上训练,并在25个领域内和10个领域外的任务上进行评估。仅提供工具接口和描述会使未适应模型的增益不一致。在相同的诊断接口下,我们的训练管道将Qwen3-8B的领域内成功率从24.8%提高到52.4%,将Qwen3.5-35B-A3B的领域内成功率从35.6%提高到69.2%。后者在领域外也从31%提高到48%,支持学习的诊断工具在保留源和目标上的使用。

更新时间: 2026-10-06 06:57:34

领域: cs.AI

下载: http://arxiv.org/abs/2609.36679v2

A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese

This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO). Contrary to previous null findings, our results show that surprisal is predictive of Chinese reading times. However, whether predictive power scales with model size and the amount of training is corpus-specific: bigger models predict better in GECO-CN, whereas inverse scaling emerges in HKP and, at the largest sizes, in MECO. Subsequently, we tested one possible explanation for the inverse scaling in HKP and found that checkpoints whose surprisal remains closer to $n$-gram statistics are better predictors of reading. All in all, the predictive power of surprisal on Chinese reading time measurements is corpus-specific, which cautions against drawing scaling conclusions from a single corpus.

Updated: 2026-10-06 06:56:02

标题: 对LM Surprisal在阅读中文中的预测能力进行系统分析

摘要: 这项研究分析了基于语言模型推导的标记级惊异对汉语阅读时间的预测能力。我们首先提出了最短匹配序列(SMS),这是一种将眼动跟踪语料库假定的词分割与LM的子词标记化进行映射的对齐方案,因为在汉语环境中这两种标记化经常不一致。然后,我们使用一组在30B标记上进行训练的中文-Pythia模型(14M-1.4B),研究惊异如何预测三个汉语段落级眼动跟踪语料库(GECO-CN、HKP和MECO)中的首次注视持续时间、凝视持续时间和总阅读时间。与先前的零结果相反,我们的结果表明惊异可以预测汉语阅读时间。然而,预测能力是否随着模型规模和训练量而变化是与语料库特定的:在GECO-CN中,更大的模型预测更好,而在HKP和MECO中出现了反向缩放。随后,我们测试了HKP中反向缩放的一个可能解释,并发现惊异保持更接近$n$-gram统计量的检查点更好地预测了阅读。总的来说,惊异对汉语阅读时间测量的预测能力是与语料库特定的,这提示不要从单一语料库中得出缩放结论。

更新时间: 2026-10-06 06:56:02

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.04898v2

Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models

Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the structured heterogeneity of attention across dimensions, heads, layers, and input tokens. In practice, attention representations exhibit non-uniform behavior, and positional encoding mechanisms such as rotary positional embeddings (RoPE) induce dimension-dependent positional structure, making uniform adaptation suboptimal. In this work, we propose DyPAM (Dynamic Positional Attention Modulation), a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations. DyPAM combines input-conditioned, dimension-wise modulation with head-wise and layer-wise structural modulation, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the pretrained backbone. Extensive experiments on mathematical and commonsense reasoning benchmarks across multiple backbone models demonstrate that DyPAM consistently outperforms existing strong PEFT baselines.

Updated: 2026-10-06 06:55:38

标题: 大型语言模型的参数高效微调的动态位置关注调制

摘要: 参数高效微调(PEFT)已成为将大型语言模型适应下游任务的标准方法。然而,大多数现有的PEFT方法依赖于统一和静态的调整,没有考虑到跨维度、头部、层次和输入标记的注意力结构的异质性。在实践中,注意力表示表现出非统一行为,而旋转位置嵌入(RoPE)等位置编码机制引入了维度相关的位置结构,使统一调整不够理想。在这项工作中,我们提出了DyPAM(动态位置注意力调制),这是一种通过直接操作查询和键表示来调整位置信息如何影响注意力的PEFT方法。DyPAM结合了输入条件、维度调制与头部和层次结构调制,执行与RoPE引发结构对齐的位置注意力的细粒度调整,而不修改预训练的骨干模型。在多个骨干模型上对数学和常识推理基准的大量实验表明,DyPAM始终优于现有的强PEFT基线。

更新时间: 2026-10-06 06:55:38

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.07848v1

Scen-Opt: A Scenario Optimization Toolbox for Data-Driven Convex Programming

The scenario approach is a well-established statistical framework for data-driven decision-making. In particular, in data-driven optimization, the scenario approach unveils how the problem structure governs out-of-sample generalization, and offers a principled basis for assessing and certifying the reliability of the optimal solution as per constraint satisfaction. Despite its strong theoretical development and wide applicability, no software toolbox has been available to date that enables user-friendly, data-driven convex optimization within the scenario-approach framework. In this paper, we introduce Scen-Opt, an open-source software tool that integrates convex programming with data samples while providing statistical guarantees grounded in scenario theory. Scen-Opt is implemented in Python, supporting data-driven linear, quadratic, and semidefinite programming, and offers a Python-based web application with an intuitive and reactive graphical user interface (GUI) built using modern web technologies. Scen-Opt can be used directly through its online interface or installed locally, accommodating both manual input and data-file uploads (CSV, JSON, TXT, TSV, MAT, Excel, NPY, NPZ, Parquet). Built on a Python backend with a modern JavaScript frontend, Scen-Opt offers a highly user-friendly experience and efficient usability across desktops, laptops, tablets, and mobile devices. In this paper, Scen-Opt is applied to a set of representative benchmarks, demonstrating its practical effectiveness for data-driven convex optimization with guaranteed performance.

Updated: 2026-10-06 06:51:26

标题: Scen-Opt:用于数据驱动凸规划的场景优化工具箱

摘要: 场景方法是一种用于数据驱动决策的成熟统计框架。特别是在数据驱动优化中,场景方法揭示了问题结构如何控制样本外泛化,并提供了基于约束满足的可靠性评估和认证的原则性基础。尽管其理论发展强大且适用广泛,但迄今为止尚无可用的软件工具箱,能够在场景方法框架内实现用户友好的数据驱动凸优化。在本文中,我们介绍了Scen-Opt,这是一个集成凸规划和数据样本的开源软件工具,同时提供基于场景理论的统计保证。Scen-Opt采用Python实现,支持数据驱动的线性、二次和半定规划,并提供一个基于Python的Web应用程序,具有直观且反应灵敏的图形用户界面(GUI),使用现代Web技术构建。Scen-Opt可以通过其在线界面直接使用,也可以本地安装,支持手动输入和数据文件上传(CSV、JSON、TXT、TSV、MAT、Excel、NPY、NPZ、Parquet)。基于Python后端和现代JavaScript前端构建的Scen-Opt提供了极具用户友好性的体验,并在台式机、笔记本电脑、平板电脑和移动设备上具有高效的可用性。在本文中,Scen-Opt应用于一组代表性基准测试中,展示了其在数据驱动凸优化中的实际有效性和保证性能。

更新时间: 2026-10-06 06:51:26

领域: cs.MS,cs.LG,eess.SY,math.OC

下载: http://arxiv.org/abs/2610.07846v1

CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational Pathology

In digital pathology, compact evidence is often used to explain or audit predictions made by whole-slide image multiple instance learning models. In hierarchical compact-evidence pipelines, candidate filtering introduces a strategy-specific candidate-conditioned prediction alongside the original full-bag prediction. If the evaluation reference changes while the intended target remains the original full-bag prediction, however, not only can the measured fidelity of the same compact evidence change, but comparisons between competing candidate strategies can also change. To make this dependence explicit, we introduce CHARTER, a reference-aware evaluation charter that asks researchers to DECLARE the intended target and reference, QUANTIFY candidate-induced prediction shift, and AUDIT the stability of comparative conclusions. Across the 15 comparisons in our main five-seed Random-K audit, 4 showed determinate reversals; in a matched native-ranking stress test, the ACMIL comparison changed from REVERSED to PRESERVED. CHARTER turns otherwise implicit candidate-filtering and reference choices into an auditable evaluation specification, helping distinguish genuine preservation of the intended prediction from apparent gains induced by changing the prediction being explained.

Updated: 2026-10-06 06:45:56

标题: CHARTER:计算病理学中层次紧凑证据评估中的审计参考替换

摘要: 在数字病理学中,紧凑证据经常用于解释或审计由整张幻灯片图像多实例学习模型所做的预测。在分层紧凑证据管道中,候选筛选引入了一种特定于策略的候选条件预测,与原始的完整袋预测并存。然而,如果评估参考发生变化,而预期目标仍然是原始的完整袋预测,不仅可以使相同紧凑证据的测量忠实度发生变化,而且竞争候选策略之间的比较也会发生变化。为了使这种依赖关系明确化,我们引入了CHARTER,一个考虑参考的评估章程,要求研究人员声明预期目标和参考,量化候选引起的预测偏移,并审计比较结论的稳定性。在我们主要的五个种子随机K审计中的15个比较中,有4个显示了确定性的逆转;在一个匹配的原生排名压力测试中,ACMIL比较从逆转变为保留。CHARTER将原本隐含的候选筛选和参考选择转化为可审计的评估规范,有助于区分预期预测的真实保留与通过改变被解释的预测而引起的明显增益。

更新时间: 2026-10-06 06:45:56

领域: cs.CV,cs.LG

下载: http://arxiv.org/abs/2610.07843v1

Privileged Context as Drift in On-Policy Self-Distillation

On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift. Specifically, we vary two axes: content (a demonstration, feedback, or rephrase) and source (external, self-generated with a verifier, or self-generated without a verifier). We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accuracy, prior-task retention, reverse KL from the base policy, and parameter-update geometry. Holding source fixed, changing content spans a wider median KL range than holding content fixed and changing source. The ratio between these ranges is $5.1\times$ for per-token KL and $2.2\times$ for per-sequence KL. Parameter-update geometry shows the same pattern: updates from adapters that share content are more closely aligned (mean cosine $0.571$) than updates from adapters that share source ($0.255$). For continual learning, these findings suggest that privileged context should be treated as part of OPSD's stability design because it is associated with how far and in what direction the policy moves.

Updated: 2026-10-06 06:45:39

标题: 特权上下文作为政策自我提炼中的漂移

摘要: On-policy self-distillation (OPSD)训练一个语言模型,使其与一个基于特权上下文的副本匹配。现有工作在特权上下文包含的内容和生成方式方面存在差异,同时还改变模型、数据和训练设置,使得特权上下文设计的影响难以孤立。受到持续学习中减少灾难性遗忘的努力的启发,我们研究特权上下文选择如何影响策略漂移。具体而言,我们改变两个轴:内容(演示、反馈或改述)和来源(外部、自动生成并具有验证器、自动生成但没有验证器)。我们用OPSD跨这九种组合和三个数据集训练Qwen2.5-7B,测量目标任务准确性、先前任务的保留、基础策略的反向KL和参数更新几何形状。保持来源不变,改变内容跨度的中位KL范围比保持内容不变并改变来源更广。这些范围之间的比率为$5.1\times$(每个标记KL)和$2.2\times$(每个序列KL)。参数更新几何形状显示相同的模式:共享内容的适配器的更新更加一致(平均余弦为$0.571$),而共享来源的适配器的更新更加一致($0.255$)。对于持续学习,这些发现表明特权上下文应被视为OPSD稳定性设计的一部分,因为它与策略移动的距离和方向有关。

更新时间: 2026-10-06 06:45:39

领域: cs.LG

下载: http://arxiv.org/abs/2610.07842v1

Probabilistic Truly Unordered Rule Sets

Rule set learning has recently been frequently revisited because of its interpretability. Existing methods have several shortcomings though. First, most existing methods impose orders among rules, either explicitly or implicitly, which makes the models less comprehensible. Second, due to the difficulty of handling conflicts caused by overlaps (i.e., instances covered by multiple rules), existing methods often do not consider probabilistic rules. Third, learning classification rules for multi-class target is understudied, as most existing methods focus on binary classification or multi-class classification via the ``one-versus-rest" approach. To address these shortcomings, we propose TURS, for Truly Unordered Rule Sets. To resolve conflicts caused by overlapping rules, we propose a novel model that exploits the probabilistic properties of our rule sets, with the intuition of only allowing rules to overlap if they have similar probabilistic outputs. We next formalize the problem of learning a TURS model based on the MDL principle and develop a carefully designed heuristic algorithm. We benchmark against a wide range of rule-based methods and demonstrate that our method learns rule sets that have lower model complexity and highly competitive predictive performance. In addition, we empirically show that rules in our model are empirically ``independent" and hence truly unordered.

Updated: 2026-10-06 06:38:19

标题: 概率性真正无序规则集

摘要: 最近,由于其可解释性,规则集学习受到频繁关注。然而,现有方法存在一些缺点。首先,大多数现有方法要求在规则之间建立顺序,无论是显式地还是隐含地,这使得模型不太易理解。其次,由于处理重叠引起的冲突(即被多个规则覆盖的实例)的困难,现有方法通常不考虑概率规则。第三,为多类目标学习分类规则尚未得到足够研究,因为大多数现有方法专注于通过“一对多”方法进行二元分类或多类分类。为了解决这些缺点,我们提出了TURS,即真正无序规则集。为了解决由重叠规则引起的冲突,我们提出了一种利用我们规则集的概率特性的新模型,其直觉是只有当规则具有类似概率输出时才允许规则重叠。接下来,我们基于MDL原则形式化学习TURS模型的问题,并开发了一个精心设计的启发式算法。我们与一系列基于规则的方法进行了基准测试,并证明我们的方法学习到的规则集具有更低的模型复杂度和高度竞争力的预测性能。此外,我们凭经验证明,我们模型中的规则在经验上是“独立的”,因此是真正无序的。

更新时间: 2026-10-06 06:38:19

领域: cs.LG

下载: http://arxiv.org/abs/2401.09918v2

DHCG: Dynamic Construction of Hierarchical Collaboration Graphs for LLM-Based Multi-Agent Reasoning

LLM-based multi-agent systems (MAS) have demonstrated strong capabilities in solving complex problems across diverse domains. Recently, the dynamic orchestration of agent systems has become an important research direction. However, existing methods suffer from limited composition, misaligned dependencies, and inflexible scale, restricting their ability to adapt to reasoning requirements during execution. To address these limitations, we reframe MAS design as a partially observable Markov decision process, in which both the composition and scale of the MAS are dynamically determined. We propose DHCG, a novel framework that coordinates three modules (Planner, Worker, and Generator) to progressively construct a dynamic hierarchical collaboration graph from scratch based on the query and evolving execution feedback. At each step, guided by feedback, the Planner generates a set of distinct and complementary roles tailored to the current reasoning needs and selectively routes relevant information to each role. It can also finalize the hierarchical collaboration graph early or progressively expand it when additional reasoning is required. We further introduce action-aware preference optimization to train the Planner to make more effective decisions when constructing hierarchical collaboration graphs. We systematically evaluate DHCG across code generation, mathematical reasoning, and domain-specific reasoning benchmarks. DHCG achieves state-of-the-art average performance among the compared methods, improving over the single-agent baseline by 13.06 points and outperforming both static and dynamic MAS baselines by 2.77-8.02 points. Additional experiments further demonstrate its generalization across different Planner backbones and unseen Worker models.

Updated: 2026-10-06 06:33:40

标题: DHCG: 基于LLM的多智能体推理的层次化协作图动态构建

摘要: 基于LLM的多Agent系统(MAS)已经展示出在解决跨领域复杂问题方面的强大能力。最近,Agent系统的动态编排已成为一个重要的研究方向。然而,现有方法存在有限的组合、不匹配的依赖关系和不灵活的规模,限制了它们在执行过程中适应推理需求的能力。为了解决这些限制,我们将MAS设计重新构想为一个部分可观察的马尔可夫决策过程,在其中MAS的组合和规模都是动态确定的。我们提出了DHCG,一个新颖的框架,它协调三个模块(规划者、工作者和生成器)根据查询和不断演化的执行反馈逐步构建一个动态的分层协作图。在每一步,根据反馈指导,规划者生成一组针对当前推理需求量身定制的不同而互补的角色,并有选择地将相关信息路由到每个角色。它还可以在需要额外推理时提前完成分层协作图或逐步扩展它。我们进一步引入了基于动作感知的偏好优化,从而训练规划者在构建分层协作图时做出更有效的决策。我们系统地评估了DHCG在代码生成、数学推理和领域特定推理基准上的性能。DHCG在与其他方法相比达到了最先进的平均性能,比单Agent基线提高了13.06个点,并且在静态和动态MAS基线上的表现均优于2.77-8.02个点。进一步的实验进一步证明了它在不同规划者骨干和未见工作者模型上的泛化能力。

更新时间: 2026-10-06 06:33:40

领域: cs.AI

下载: http://arxiv.org/abs/2610.07835v1

Retrieval Is Not Enough: Refreshing Memory for Frozen Time-Series Forecasters

Retrieval-augmented time-series forecasting uses the continuations of historical segments similar to the current context as references for a forecaster. Most existing methods build the retrieval memory once from the training segment, leaving observations revealed after deployment unavailable as references, and generally do not calibrate how much the retrieved information should influence a frozen forecaster. We identify two key determinants of retrieval utility for a frozen forecaster: whether the history still reflects the current state, and whether the correction it induces aligns with the forecaster's residual errors, an alignment that can shift between validation and deployment when the memory becomes stale. We propose FreshCast, a plug-in retrieval framework that keeps the forecaster frozen, continuously updates a non-parametric memory with new observations, forms a memory forecast through relational kernel regression, and calibrates its weight in closed form on the validation segment. Under a simplified generative model, we characterize the optimal combination gain through the second-order relation between forecaster error and memory correction, and show that a sufficiently long look-back can make periodic memory information redundant. Across seven benchmarks and ten forecasting architectures, FreshCast reduces average MSE for every evaluated forecaster and input length, by 14.6% and 5.6% at input lengths 96 and 720, and achieves lower MSE than the evaluated retrieval-augmented and online baselines in their comparison settings. Ablations show that freezing the memory at the end of training removes most of the gain, identifying post-training observations as a primary source of improvement. For a frozen forecaster, useful historical references must remain timely and provide information that helps correct its remaining errors.

Updated: 2026-10-06 06:33:16

标题: 检索并不足够:为冻结时间序列预测器刷新记忆

摘要: 检索增强时间序列预测使用类似于当前上下文的历史片段的延续作为预测者的参考。大多数现有方法一次从训练片段构建检索记忆,导致部署后不可用作参考的观察结果,并且通常不校准检索到的信息应该如何影响冻结的预测者。我们确定了冻结预测者的检索效用的两个关键因素:历史是否仍然反映当前状态,以及其引起的校正是否与预测者的残差误差一致,这种一致性在记忆变得陈旧时在验证和部署之间可能会发生变化。我们提出了FreshCast,这是一个插件检索框架,它保持预测者冻结,持续更新一个非参数记忆与新观察结果,通过关系核回归形成记忆预测,并在验证段上以闭式校准其权重。在一个简化的生成模型下,我们通过预测者误差和记忆校正之间的二阶关系表征了最佳组合增益,并且表明足够长的回顾可以使周期性记忆信息变得多余。在七个基准和十个预测架构中,FreshCast在每个评估的预测者和输入长度上都减少了平均均方误差,分别为96和720输入长度的14.6%和5.6%,并且在比较设置中的MSE低于评估的检索增强和在线基线。消融实验表明,在训练结束时冻结记忆会去除大部分增益,将后训练观察结果识别为改进的主要来源。对于冻结的预测者,有用的历史参考必须保持及时,并提供有助于纠正其余误差的信息。

更新时间: 2026-10-06 06:33:16

领域: cs.LG

下载: http://arxiv.org/abs/2610.07834v1

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.

Updated: 2026-10-06 06:32:11

标题: 使用模块化可执行开发基元的软件工程中的绑定工程

摘要: 大型语言模型(LLMs)配备终端访问权限,已经展示出在自动化软件工程任务方面具有强大的能力。然而,现有的代理在长期工作流中仍然表现脆弱,其中它们必须反复重建散落在源文件、配置、测试、依赖和运行时行为中的程序状态,导致交互历史越来越长,上下文爆炸和语义漂移。大型代码库进一步复杂化了任务相关组件的识别。为了解决这些挑战,我们引入了“Dev-Primitives”(开发原语),这是一种模块化和可执行的抽象,将代码库组件从被动的软件工件转化为软件工程中的活跃参与者。每个Dev-Primitive将一个代码库工件与一个驻留的LLM配对,这为工件提供了一个基于其自身实现和依赖关系的代理本地接口,实现了自然语言推理、组件间通信和局部自修改。基于Dev-Primitives,我们提出了“HERMES”,一种通过模块化可执行的Dev-Primitives来进行软件工程的Harness Engineering框架,通过一个依赖感知的动态激活机制和一个将执行证据映射回必须进行修订的组件的bug诊断机制,在代码库规模上实例化这些原语。对四个软件工程基准的广泛实验表明,HERMES的性能比匹配的基线Harness平均提高了12.4%。此外,当与强激活和诊断模型配对时,即使使用Qwen3-8B Dev-Primitives,HERMES在所有四个基准测试中仍然保持在距离同质GPT-5.6 Sol配置不到4.5%的范围内,同时在Terminal-Bench 4.0上减少了26.2%的推理成本,突显了软件工程代理中Harness设计的重要性。

更新时间: 2026-10-06 06:32:11

领域: cs.SE,cs.AI,cs.CL,cs.MA

下载: http://arxiv.org/abs/2610.07832v1

When Similarity Is Interaction-Driven: Quantum Kernels for Regime-Sensitive Learning

Similarity in many decision systems is governed not by distance alone but by interactions among variables. In fraud and anomaly detection, small local perturbations can cross interaction-sensitive decision boundaries while leaving ambient distance almost unchanged. Motivated by this setting, we introduce a thin-slab interaction model and an interaction-driven quantum kernel constructed from entangled Pauli-string feature maps. The feature map explicitly encodes sparse high-order block interactions. We show that the resulting fidelity kernel is positive semidefinite, admits an exact block-factorized formulation, and induces a geometry sensitive to changes in interaction regime. Across balanced and imbalanced synthetic experiments spanning third-, fourth-, sixth-, and eighth-order interactions, the proposed kernel consistently outperforms linear, radial basis function, Laplacian, and polynomial kernels, as well as an engineered-interaction linear baseline supplied with the planted block products. On real fraud-detection benchmarks, it achieves the highest mean accuracy and F1 on Credit Card Fraud Detection and ranks second on IEEE-CIS Fraud Detection. Executed on a 156-qubit IBM Quantum processor in a fourth-order setting, the hardware-estimated kernel matches the noise-free simulator within seed-to-seed variability and retains its advantage over the baselines. These findings show that quantum-kernel performance depends on alignment between feature-map geometry and the underlying predictive structure, rather than on Hilbert-space dimension alone. Because the prescribed block-factorized kernel can also be evaluated exactly on a classical computer, the results establish predictive and representational value rather than computational quantum speedup.

Updated: 2026-10-06 06:32:04

标题: 当相似性由互动驱动时:用于区别敏感学习的量子核

摘要: 在许多决策系统中,相似性不仅受距离的影响,还受变量之间的相互作用影响。在欺诈和异常检测中,小范围的局部扰动可以越过对相互作用敏感的决策边界,同时几乎不改变环境距离。受到这种情况的启发,我们引入了一个薄板相互作用模型和一个由纠缠的Pauli字符串特征映射构建的相互作用驱动的量子核。特征映射明确地编码了稀疏高阶块相互作用。我们展示了由此产生的保真度核是正半定的,具有精确的块因式分解形式,并诱导出一种对相互作用变化敏感的几何形状。在跨越第三、第四、第六和第八阶相互作用的平衡和不平衡合成实验中,所提出的核一贯优于线性、径向基函数、拉普拉斯和多项式核,以及一个提供了植入块乘积的工程化相互作用线性基准。在真实的欺诈检测基准上,它在信用卡欺诈检测中取得了最高的平均准确率和F1值,并在IEEE-CIS欺诈检测中排名第二。在第四阶设置中,在一个156量子比特的IBM量子处理器上执行,硬件估计的核与无噪声模拟器在种子间变化范围内匹配,并保持其优势超过基线。这些发现表明,量子核的性能取决于特征映射几何形状与基础预测结构之间的一致性,而不仅仅取决于希尔伯特空间的维度。因为所规定的块因式分解核也可以在经典计算机上进行精确评估,结果建立了预测和表征价值,而不是计算量子加速。

更新时间: 2026-10-06 06:32:04

领域: quant-ph,cs.LG

下载: http://arxiv.org/abs/2608.24631v2

Agentic Semantic Sensing for Resource-Adaptive AI-RAN

Semantic sensing (SemS) acquires task-relevant information rather than reconstructing complete physical information. Existing SemS formulations typically operate open loop: sensing configurations and observation schedules are fixed before inference and cannot respond to evolving task-level evidence. We propose Agentic SemS, a closed-loop framework for AI-enabled radio access networks (AI-RANs) that controls sensing within a communication-feasible profile set. A profile-conditioned causal Transformer updates the semantic belief from streaming observations, while key-value caching enables efficient state updates across profile changes without repeatedly processing the complete history. A semantic utility network estimates the task-level benefit of acquiring the next observation block under each feasible profile after accounting for sensing cost. The resulting continuation utilities jointly support next-profile selection and semantic early exit, adapting sensing configuration and duration to evolving evidence. The expected semantic gain is further related to conditional mutual information, providing a value-of-information interpretation of continued online sensing. Experiments on Widar3.0 with six emulated sensing profiles show that, in comparison with full-sequence High, the resource-efficient Agentic setting reduces normalized cumulative sensing cost by 25.33% while achieving 85.79% Macro-F1. At the same utility checkpoint, semantic early exit provides a further 12.35% cost reduction over adaptive sensing without early exit, with a 0.97-percentage-point Macro-F1 decrease.

Updated: 2026-10-06 06:27:08

标题: 资源自适应AI-RAN的主动语义感知

摘要: 语义感知(SemS)获取任务相关信息,而不是重建完整的物理信息。现有的SemS公式通常是开环操作:在推理之前,传感配置和观测计划是固定的,无法响应不断变化的任务级证据。我们提出了主动SemS,这是一个闭环框架,用于AI-enabled无线接入网络(AI-RANs),它控制感知在通信可行的配置集内。一个受配置条件影响的因果Transformer从流式观测中更新语义信念,而键值缓存能够在配置更改时有效地更新状态,而无需重复处理完整历史。一个语义效用网络估计在考虑感知成本后,在每个可行配置下获取下一个观测块的任务级益处。由此产生的继续效用共同支持下一个配置选择和语义提前退出,根据不断变化的证据调整感知配置和持续时间。预期的语义收益进一步与条件互信息相关,为持续在线感知提供信息价值解释。在六个模拟感知配置的Widar3.0上的实验表明,与完整序列High相比,资源高效的主动设置将标准化累积感知成本降低了25.33%,同时实现了85.79%的Macro-F1。在相同的效用检查点上,语义提前退出比没有提前退出的自适应感知进一步减少了12.35%的成本,Macro-F1减少了0.97个百分点。

更新时间: 2026-10-06 06:27:08

领域: cs.AI,eess.SP

下载: http://arxiv.org/abs/2610.07829v1

Boosting Large Language Models with Mask Fine-Tuning

The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performance. In this work, we introduce Mask Fine-Tuning (MFT), a novel LLM fine-tuning paradigm demonstrating that carefully breaking the model's structural integrity can surprisingly improve performance without updating model weights. MFT learns and applies binary masks to well-optimized models, using the standard LLM fine-tuning objective as supervision. Based on fully fine-tuned models, MFT uses the same fine-tuning datasets to achieve consistent performance gains across domains and backbones (e.g., an average gain of 2.70/4.15 on IFEval with LLaMA2-7B/3.1-8B). Detailed ablation studies and analyses examine the proposed MFT from different perspectives, including the sparse ratio and the loss surface. Additionally, when deployed on well-trained models, MFT is compatible with other LLM optimization procedures to improve overall model performance. Furthermore, this study extends the masking operation beyond its conventional use in network pruning for model compression to encompass a broader range of model capabilities.

Updated: 2026-10-06 06:26:46

标题: 用Mask Fine-Tuning推动大型语言模型

摘要: 大型语言模型(LLM)通常被整合到主流优化协议中。然而,保持模型完整性是否对良好性能至关重要仍未得到充分探讨。在这项工作中,我们介绍了Mask Fine-Tuning(MFT),这是一种新颖的LLM微调范式,展示了通过精心打破模型的结构完整性可以令人惊讶地提高性能,而无需更新模型权重。MFT学习并应用二进制掩模到经过良好优化的模型上,使用标准LLM微调目标作为监督。基于完全微调的模型,MFT使用相同的微调数据集跨领域和骨干(例如,在IFEval上与LLaMA2-7B/3.1-8B平均增益为2.70/4.15)实现了一致的性能提升。详细的消融研究和分析从不同角度检验了提出的MFT,包括稀疏比和损失曲面。此外,当部署在训练良好的模型上时,MFT与其他LLM优化程序兼容,以提高整体模型性能。此外,该研究将掩模操作从网络剪枝和模型压缩的传统用途扩展到涵盖更广泛的模型功能范围。

更新时间: 2026-10-06 06:26:46

领域: cs.CL,cs.AI,cs.LG

下载: http://arxiv.org/abs/2503.22764v3

Adaptive Latent Capacity for World Models

We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.

Updated: 2026-10-06 06:24:43

标题: 适应性潜在容量用于世界模型

摘要: 我们介绍了一种基于联合嵌入预测架构(JEPA)的自适应LeWorldModel(ALeWM),该模型学习将预测信息集中在宽广的潜在表示的紧凑前缀中。为了鼓励这种排序,ALeWM学习了一个基于序列的前缀长度分布,并训练预测器从采样输入前缀估计完整的下一个嵌入。由于标准的抗坍缩目标鼓励潜在坐标之间的变化,并且没有按预测重要性对其进行组织,因此我们还引入了MixSIGReg。MixSIGReg通过使用高斯主动前缀和剩余坐标中的零的先验加权混合对掩码嵌入进行正则化。因此,ALeWM目标鼓励早期坐标保留有用于预测和递归规划的信息。我们的分析显示,MixSIGReg使用的混合分布将更高的方差分配给早期坐标块,将较低的方差分配给后续坐标块。此外,我们表明,在指定的假设下,通过将对预测最有用的信息放在较早的块中,可以最小化预测误差。在经过控制的动态系统和目标条件视觉控制中,我们通过实验研究了ALeWM的行为。我们表明,ALeWM始终比调整后的固定宽度LeWM实现更高的平均成功率,且平均规划能力较低。

更新时间: 2026-10-06 06:24:43

领域: cs.LG

下载: http://arxiv.org/abs/2609.32921v2

VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis

Plan accuracy alone cannot show whether a speech-delivery decision follows its source: a fixed prior may match the original label yet fail to respond appropriately when a cue changes. VoxReason provides a 100-case verifier benchmark that holds each utterance fixed, edits one designated source-label cue, and scores cited evidence, eight plan fields, and the permitted response. On a source-key-disjoint test of 24 cases, a source-emotion prior reaches plan-slot accuracy 0.958, but none of the 24 edited neutral targets appears in its training labels; its required-change accuracy is 0.000. This diagnoses the support boundary of this prior, not its performance on supported edits. In a complementary 32-case emotion-disjoint test, the prior has seen all edited neutral targets but neither original test emotion; its plan-slot accuracy is 0.219 and required-change accuracy is 1.000. The partitions reuse and overlap the same 100 cases, so these deterministic diagnostics are not independent cohorts or learned-planner results. The benchmark evaluates derived labels and structured plans, not audio input, generated speech, or listener judgments.

Updated: 2026-10-06 06:24:04

标题: VoxReason:在合成之前审计基于来源的语音计划

摘要: 计划的准确性本身不能显示语音传递决策是否遵循其源头:固定的先验可能与原始标签匹配,但在提示发生变化时可能无法适当响应。 VoxReason提供了一个包含100个案例的验证基准,每个话语保持不变,编辑一个指定的源标签提示,并评分引用证据、八个计划字段和允许的响应。在一个包含24个案例的源-键不相交的测试中,源情感先验达到了计划槽准确率0.958,但24个编辑的中性目标中没有一个出现在其训练标签中;其所需更改的准确度为0.000。这诊断了该先验的支持边界,而不是其在支持的编辑上的表现。在一个补充的32个情感不相交的测试中,先验已经看过所有编辑的中性目标,但没有原始的测试情感;其计划槽准确率为0.219,所需更改的准确率为1.000。分区重复并重叠相同的100个案例,因此这些确定性诊断不是独立的队列或学习规划结果。该基准评估派生标签和结构化计划,而不是音频输入、生成的语音或听众判断。

更新时间: 2026-10-06 06:24:04

领域: cs.SD,cs.CL,cs.LG,eess.AS

下载: http://arxiv.org/abs/2609.03203v4

Forecast Accuracy Is Not Trading Profit: Evolving Small Recurrent Networks for Stock Return Prediction

Time series forecasting models are typically compared on pointwise error, which scores a prediction in isolation from the decision it is produced for, and a lower forecast error does not imply a better decision downstream. A parallel debate asks whether modern transformer architectures forecast better than recurrent and other lightweight models. We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy. All models are fit on a pooled panel, one network trained across the whole universe. Across four mid-cap portfolios and three trading years, the evolved networks rank first on both forecast accuracy and net trading performance, while the second most accurate model loses money once positions are formed and costs are charged. The advantage tracks a horizon match, since rank IC for the evolved networks rises from a one-day to a ten-day scoring horizon while every model above 300 parameters declines. They are also the cheapest end to end: a CPU-only search of 16 minutes yields 66-weight networks that predict in 10.8~$μ$s on a Raspberry Pi Zero, against transformer baselines of up to 817,153 parameters that require GPU training.

Updated: 2026-10-06 06:20:33

标题: 预测准确性并非交易利润:为股票回报预测演化小型循环网络

摘要: 时间序列预测模型通常通过点间误差进行比较,这种方法对预测进行评分时独立于产生该预测的决策,因此更低的预测误差并不意味着在下游产生更好的决策。另一方面,有一个关于现代变压器架构是否比递归和其他轻量级模型预测更好的讨论。我们将线性、固定递归、变压器和基于混合的架构与通过神经进化架构搜索进化的递归网络进行比较,评估每种模型的预测准确性和每日做多/做空策略的净收益。所有模型都基于汇总面板进行拟合,一个网络在整个宇宙上进行训练。在四个中型股票组合和三个交易年份中,进化的网络在预测准确性和净交易表现上排名第一,而第二准确性最高的模型在形成头寸并收取费用后会亏损。优势体现在时间跨度匹配上,因为进化网络的排名IC从一天到十天的评分范围上升,而超过300个参数的每个模型都在下降。它们也是最经济的:仅使用CPU进行16分钟的搜索可以获得66个权重的网络,该网络在树莓派Zero上以10.8μs的速度进行预测,而变压器基线需要GPU训练,参数高达817,153。

更新时间: 2026-10-06 06:20:33

领域: cs.NE,cs.CE,cs.LG

下载: http://arxiv.org/abs/2610.07825v1

CANDLE: Cortical Null-Space Decomposition for Noninvasive Brain Source Imaging

Electrophysiological source imaging (ESI) aims to estimate cortical source activity from noninvasive electrophysiological measurements such as electroencephalogram (EEG). However, ESI is fundamentally ill-posed because source activity is substantially higher-dimensional than sensor observations, resulting in non-unique solutions. Recent learning-based approaches address this ambiguity by learning data-driven source priors, yet they often struggle to generalize across subject-specific cortical geometries. To address this, we propose CANDLE, a learning-based ESI model that estimates source activity on subject-specific cortical geometries. CANDLE learns a prior over the null space induced by the source-to-sensor mapping derived from T1-weighted MRI, restricting learning to unobservable source components while preserving geometric constraints. To train CANDLE, we develop a whole-brain simulator spanning over 1,100 subject-specific cortical geometries with source configurations derived from over 26,000 statistical brain maps. Trained exclusively on simulated data, CANDLE outperformed prior ESI methods on simulated source activity estimation and generalized to two empirical tasks: (i) intracranial stimulation localization from simultaneously recorded scalp EEG and (ii) epileptogenic zone estimation from presurgical interictal EEG. Our project page is available at https://candle-esi.pages.dev}{https://candle-esi.pages.dev.

Updated: 2026-10-06 06:19:18

标题: 蜡烛:用于非侵入性脑源成像的皮层零空间分解

摘要: 电生理源成像(ESI)旨在从非侵入性电生理测量(如脑电图(EEG))中估计皮层源活动。然而,由于源活动的维度远高于传感器观测,ESI基本上是不适定的,导致非唯一解。最近的基于学习的方法通过学习数据驱动的源先验来解决这种不确定性,但它们通常难以概括不同受试者特定的皮层几何形状。为了解决这个问题,我们提出了CANDLE,一个基于学习的ESI模型,用于在受试者特定的皮层几何形状上估计源活动。CANDLE通过学习源到传感器映射所引起的零空间上的先验来限制学习到不可观测的源组分,同时保留几何约束。为了训练CANDLE,我们开发了一个整脑模拟器,涵盖超过1,100个受试者特定的皮层几何形状,源配置来源于超过26,000个统计脑图。仅在模拟数据上训练,CANDLE在模拟源活动估计方面优于先前的ESI方法,并推广到两个实证任务:(i)从同时记录的头皮EEG中定位颅内刺激,以及(ii)从术前间歇期EEG中估计癫痫发作区。我们的项目页面位于https://candle-esi.pages.dev。

更新时间: 2026-10-06 06:19:18

领域: cs.LG,eess.SP,q-bio.NC

下载: http://arxiv.org/abs/2610.07824v1

TTNet: Multi-Task Deep Learning for Table Tennis Player Analysis with Smart Racket

The AI CUP 2025 Precise Analysis of Table Tennis Smart Racket Data Competition introduced smart table tennis rackets that collect extensive player swing data, enabling research on table tennis big data. These data support in-depth analysis of players' return techniques and swing-force consistency, improving the accuracy of player skill assessment. This study focuses on six-axis sensor data collected by smart table tennis rackets and proposes TTNet, a novel deep learning model with multitask learning capabilities, to advance table tennis data analysis and related applications. TTNet combines convolutional neural networks (CNNs), residual networks (ResNet), and self-attention mechanisms to simultaneously predict four player attributes: gender, playing hand, years of experience, and skill level. We adopt a two-stage training strategy that incorporates data augmentation and task-specific loss functions to improve generalization on imbalanced data. Our approach achieved second place on the official competition leaderboard.

Updated: 2026-10-06 06:18:52

标题: TTNet:智能球拍下的乒乓球运动员分析的多任务深度学习

摘要: 2025年人工智能杯精准分析乒乓球智能球拍数据比赛引入了智能乒乓球拍,收集了大量球员挥拍数据,实现了对乒乓球大数据的研究。这些数据支持对球员回球技术和挥拍力量一致性的深入分析,提高了球员技能评估的准确性。本研究侧重于智能乒乓球拍收集的六轴传感器数据,并提出了TTNet,一种具有多任务学习能力的新型深度学习模型,以推进乒乓球数据分析和相关应用。TTNet结合了卷积神经网络(CNN)、残差网络(ResNet)和自注意机制,同时预测四个球员属性:性别、击球手、经验年限和技能水平。我们采用了两阶段训练策略,结合了数据增强和任务特定损失函数,以改善不平衡数据的泛化能力。我们的方法在官方比赛排行榜上取得了第二名。

更新时间: 2026-10-06 06:18:52

领域: cs.LG

下载: http://arxiv.org/abs/2610.07823v1

Computationally efficient goodness-of-fit tests through kernelized Stein discrepancy

Models with intractable normalizing constants are widely used in statistics and machine learning. Assessing the adequacy of such models poses significant challenges: obtaining samples from the fitted model often requires sophisticated sampling algorithms. Moreover, model fitting sometimes requires iterative numerical optimization, making bootstrap procedures that require repeated refitting computationally expensive. In this paper, we leverage the kernel-based testing framework to develop a general semiparametric goodness-of-fit test based on the kernelized Stein discrepancy. We establish the consistency and the asymptotic null distribution of the test statistic under general nuisance estimation. To produce a level-$α$ test, we propose a novel influence-adjusted wild bootstrap that requires neither refitting the model nor sampling from it. We prove the consistency of the proposed bootstrap test procedure under the null and the alternative, and characterize its limiting power under contiguous local alternatives. Across simulations ranging from classical normality testing to models with intractable likelihoods, the proposed test delivers competitive or superior power at a computational cost orders of magnitude lower than that of existing approaches. We illustrate the method by assessing the adequacy of a protein signaling network model for reverse-phase protein array data from lung adenocarcinoma tumors. As a complementary insight, we show that the SKSD test can be regarded as a nonparametric score test under exponentially tilted models, connecting score-based and distance-based goodness-of-fit testing.

Updated: 2026-10-06 06:17:23

标题: 通过核化斯坦距离实现计算效率高的拟合度检验

摘要: 具有难以计算归一化常数的模型在统计学和机器学习中被广泛使用。评估这些模型的适用性带来了重大挑战:从拟合模型中获取样本通常需要复杂的抽样算法。此外,模型拟合有时需要迭代数值优化,使需要重复拟合的自助程序在计算上昂贵。在本文中,我们利用基于核的测试框架,基于核化的Stein距离开发了一种基于核化Stein距离的一般半参数拟合优度检验。我们在一般无关估计下建立了检验统计量的一致性和渐近零分布。为了产生一个水平为$α$的检验,我们提出了一种新颖的影响调整的野生自助法,既不需要重新拟合模型也不需要从中抽样。我们证明了所提出的自助检验程序在零假设和备择假设下的一致性,并描述了在连续局部备择假设下的极限功率。在从经典正态性测试到具有难以计算似然的模型的模拟中,所提出的检验方法以比现有方法计算成本低几个数量级的竞争性或更高的功率。我们通过评估用于肺腺癌肿瘤的逆向蛋白质组数据的蛋白信号网络模型的适用性来说明该方法。作为一个补充的见解,我们展示了SKSD检验可以被视为在指数倾斜模型下的非参数得分检验,连接了基于得分和基于距离的拟合优度测试。

更新时间: 2026-10-06 06:17:23

领域: stat.ML,cs.LG,stat.ME

下载: http://arxiv.org/abs/2512.20007v3

Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models

Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tuning (PEFT) methods rely on additive modules, such as low-rank adapters, which may limit adaptation capacity. In this paper, we rethink VLM adaptation from a structural selection framework that adapts VLMs without modifying backbone weights, and we propose Mask Fine-Tuning (MFT). MFT learns masks that selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align pretrained representations with downstream objectives. Extensive experiments show that MFT provides an effective structural alternative to both FFT and PEFT, consistently achieving superior performance across multiple vision-language benchmarks without adding knowledge or altering the deployment architecture. Moreover, our analysis with MFT provides new insights into how pretrained VLMs reorganize their internal representational pathways during adaptation.

Updated: 2026-10-06 06:16:15

标题: 重新思考微调:释放视觉语言模型中的潜在能力

摘要: 微调已成为调整视觉语言模型(VLMs)的主导范式,然而大多数方法依赖于引入基本权重更新的显式权重更新,这引入了一个根本的权衡。完全微调(FFT)可能会由于跨模态梯度干扰而扰乱预训练表示,而参数高效微调(PEFT)方法依赖于添加模块,如低秩适配器,这可能限制适应能力。在本文中,我们从一个结构选择框架重新思考VLM适应,该框架在不修改主干权重的情况下调整VLMs,并提出了Mask Fine-Tuning(MFT)。MFT学习选择性地通过现有的预训练连接路由信息,动态地揭示更好地将预训练表示与下游目标对齐的子网络。大量实验证明,MFT为FFT和PEFT提供了一种有效的结构替代方案,在多个视觉语言基准测试中持续实现优越性能,而无需添加知识或改变部署架构。此外,我们通过MFT的分析提供了有关预训练VLM在适应过程中如何重新组织其内部表示路径的新见解。

更新时间: 2026-10-06 06:16:15

领域: cs.LG,cs.CV

下载: http://arxiv.org/abs/2512.23073v2

Efficient and Implementation-Hardened RBLWE on Commodity Cortex-M Microcontrollers

Efficient post-quantum cryptography on resource-constrained Internet-of-Things (IoT) devices requires implementations that exploit the target processor architecture while resisting practical implementation attacks. This paper presents an ISA-accelerated and implementation-hardened realization of Ring Binary Learning with Errors (RBLWE) encryption on a commodity ARM Cortex-M33 microcontroller. Packing four 8-bit polynomial coefficients into the byte lanes of a 32-bit register and processing them with SIMD-style instructions, together with a packed message codec, accelerates encryption and decryption, while a buffered hardware-TRNG entropy source drawn from the on-die Secure Element drives the key-generation gain. Together these give same-core cold-start speedups of $4.06\times$, $3.35\times$, and $3.01\times$ for key generation, encryption, and decryption over a scalar baseline, and $3.52\times$/$3.18\times$ lower encryption/decryption cycle counts than a reference Cortex-M0 implementation. On top of this accelerated core, we add four staged countermeasures: constant-time execution, fault hardening against zeroing, random-corruption, and instruction-skip faults, a Fujisaki-Okamoto (FO)-style CCA2 transform, and first-order shared (masked) CPA decryption, reporting each layer's cost individually. Binary-level inspection confirms these countermeasures survive compilation and identifies a compiler-induced masking flaw resolved with a hand-written assembly replacement. Dudect-style timing tests, debugger-assisted fault-injection campaigns, and component-level TVLA then provide implementation-level evidence for the staged protections. The results demonstrate a practical acceleration-security tradeoff for RBLWE on off-the-shelf microcontrollers and reusable architecture-aware techniques for lightweight post-quantum implementations.

Updated: 2026-10-06 06:14:48

标题: 在普通Cortex-M微控制器上高效且实现强化的RBLWE

摘要: 在资源受限的物联网(IoT)设备上进行高效的后量子密码学需要利用目标处理器架构的实现,同时抵抗实际的攻击。本文介绍了在商品ARM Cortex-M33微控制器上实现的ISA加速和实现强化的Ring Binary Learning with Errors(RBLWE)加密。将四个8位多项式系数打包到32位寄存器的字节通道中,并使用SIMD风格的指令处理它们,再加上一个打包的消息编解码器,加速了加密和解密,同时来自芯片上安全元件的缓冲硬件TRNG熵源驱动了密钥生成。这些共同使得在标量基线上,密钥生成、加密和解密的同核冷启动加速分别为$4.06\times$、$3.35\times$和$3.01\times$,比参考Cortex-M0实现的加密/解密周期次数低了$3.52\times$/$3.18\times$。在这个加速核心之上,我们添加了四个分阶段的对策:常数时间执行、针对清零、随机破坏和指令跳过故障的强化、Fuji坂-岡本(FO)风格的CCA2变换,以及一阶共享(掩盖)CPA解密,报告每个层的成本。二进制级别的检查确认这些对策经受了编译的考验,并确定了一个编译器引起的掩盖缺陷,通过手工编写的汇编替换来解决。Dudect风格的时间测试、调试器辅助的故障注入活动和组件级TVLA为分阶段保护提供了实现级证据。结果展示了在现成微控制器上实现RBLWE的实际加速-安全折衷,并为轻量级后量子实现提供了可重复使用的架构感知技术。

更新时间: 2026-10-06 06:14:48

领域: cs.CR,cs.AR

下载: http://arxiv.org/abs/2610.07820v1

UniST-Pred: A Robust Unified Framework for Spatio-Temporal Traffic Forecasting in Transportation Networks Under Disruptions

Spatio-temporal traffic forecasting is a core component of intelligent transportation systems, supporting various downstream tasks such as signal control and network-level traffic management. In real-world deployments, forecasting models must operate under structural and observational uncertainties, conditions that are rarely considered in model design. Recent approaches achieve strong short-term predictive performance by tightly coupling spatial and temporal modeling, often at the cost of increased complexity and limited modularity. In contrast, efficient time-series models capture long-range temporal dependencies without relying on explicit network structure. We propose UniST-Pred, a unified spatio-temporal forecasting framework that first decouples temporal modeling from spatial representation learning, then integrates both through adaptive representation-level fusion. To assess robustness of the proposed approach, we construct a dataset based on an agent-based, microscopic traffic simulator (MATSim) and evaluate UniST-Pred under severe network disconnection scenarios. Additionally, we benchmark UniST-Pred on standard traffic prediction datasets, demonstrating its competitive performance against existing well-established models despite a lightweight design. The results illustrate that UniST-Pred maintains strong predictive performance across both real-world and simulated datasets, while also yielding interpretable spatio-temporal representations under infrastructure disruptions. The source code and the generated dataset are available at https://anonymous.4open.science/r/UniST-Pred-EF27

Updated: 2026-10-06 06:13:04

标题: UniST-Pred:在交通网络中面对干扰的时空交通预测的强大统一框架

摘要: 时空交通预测是智能交通系统的核心组成部分,支持各种下游任务,如信号控制和网络级交通管理。在现实世界的部署中,预测模型必须在结构和观测不确定性下运行,这些条件在模型设计中很少考虑。最近的方法通过紧密耦合空间和时间建模实现了强大的短期预测性能,但往往以增加复杂性和有限的模块化为代价。相比之下,高效的时间序列模型捕捉长程时间依赖性,而无需依赖显式网络结构。我们提出了UniST-Pred,一个统一的时空预测框架,首先将时间建模与空间表征学习解耦,然后通过自适应的表示级融合将两者整合起来。为了评估所提出的方法的稳健性,我们基于基于代理的微观交通模拟器(MATSim)构建了一个数据集,并在严重网络断开的情况下评估UniST-Pred。此外,我们在标准交通预测数据集上对UniST-Pred进行基准测试,展示了尽管设计轻量化,但其与现有成熟模型的竞争性能。结果表明,UniST-Pred在真实世界和模拟数据集上都保持强大的预测性能,同时在基础设施中断情况下产生可解释的时空表示。源代码和生成的数据集可在https://anonymous.4open.science/r/UniST-Pred-EF27获取。

更新时间: 2026-10-06 06:13:04

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2602.14049v3

$α$Transfer: Coefficient Transfer for Efficient Model Merging

Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{$α$Transfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify $α$Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6$\times$ speedup and 70\% memory reduction on vision transformers, and a 20$\times$ speedup and 85\% memory reduction on large language models, while maintaining comparable performance. Our findings establish $α$Transfer as an efficient and generalizable approach to scaling model merging.

Updated: 2026-10-06 06:13:01

标题: $α$Transfer:用于高效模型合并的系数转移

摘要: 模型合并通过参数运算为将多个微调检查点组合成单个模型提供了一种有前途的解决方案。然而,找到最佳合并系数需要进行广泛的搜索,随着模型规模的增大,由于高内存需求和搜索空间中组合增长,这变得非常昂贵。我们展示,在同一模型系列内,模型在不同大小的模型上的合并系数上表现出高度一致的性能分布。这种分布相似性使得我们提出了一个实用的范式,我们称之为$α$Transfer:在一个小的代理模型上搜索最佳系数,然后直接将它们转移到更大的目标模型。我们验证了$α$Transfer在多个合并方法、模型系列和任务中的有效性。实验结果表明,在视觉变换器上实现了6倍加速和70%内存减少,在大型语言模型上实现了20倍加速和85%内存减少,同时保持可比性能。我们的研究结果将$α$Transfer确定为一种高效且可推广的模型合并方法。

更新时间: 2026-10-06 06:13:01

领域: cs.LG,cs.CL,cs.CV

下载: http://arxiv.org/abs/2610.07819v1

One Step at a Time: Trading LLM Autonomy for Process Predictability

Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back. We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured step_output. This trades autonomy for predictability, and two properties then follow by construction, independent of the executor. The execution path is prescribed before the run, so the process is predictable in advance rather than reconstructed afterwards; and the completed step records form a machine-readable execution log that downstream tooling can audit and optimize step by step. Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small. Across all four, process adherence rises significantly (76-95% to 95-99%) and ungrounded answers (correct outputs produced without executing the SOP) near-vanish, falling from 2.1-4.5% to 0.2-0.3% of trials (all 95% CIs exclude zero); under prompt-based delivery, 31-49% of correct answers on know_your_business bypass the SOP entirely, even for the frontier executor. Accuracy is where the executor's capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.

Updated: 2026-10-06 06:10:42

标题: 一步一步:以LLM自主权交换过程可预测性

摘要: 组织自动化运营流程需要的不仅仅是正确的结果:他们需要预测流程的运行方式,知道实际运行的流程是哪一个,并逐步检查它。当一个代理是执行者时,通常会失去可预测性:规定的程序进入系统提示符,只有最终答案返回。相反,我们通过模型上下文协议(MCP)逐步提供流程:服务器逐步释放一步,代理执行它,每个步骤返回结构化的步骤输出。这种交易自主性以换取可预测性,然后通过构造得到两个属性,与执行者无关。执行路径在运行之前被规定,因此流程在事先可预测,而不是事后重建;完成的步骤记录形成了一个可读的执行日志,下游工具可以逐步审计和优化。在13个SOP-Bench领域和四个从前沿(Kimi K2.5)到轻量级(Ministral 3 8B)的开放式权重执行者中评估了15,475次试验,我们发现步骤级交付使得每个执行者执行的流程可预测和可检查,并且当执行者较小时,还提高了准确性。在所有四个执行者中,流程的依从性显著提高(76-95%至95-99%),不合理的答案(在不执行SOP的情况下产生正确输出)几乎消失,从2.1-4.5%下降到0.2-0.3%的试验(所有95%的置信区间均排除零);在基于提示的交付下,31-49%的正确答案在"know_your_business"上完全绕过SOP,即使对于前沿执行者也是如此。准确性取决于执行者的能力:轻量级执行者因外部提供流程而获得+6.5pp的可信度,因为这种方法消除了它无法承担的重建负担,而能力强的执行者则以可预测、可审计的流程交换了一小部分原始准确性的下降。

更新时间: 2026-10-06 06:10:42

领域: cs.CL,cs.AI

下载: http://arxiv.org/abs/2610.07817v1

Qubit-centric Transformer for Surface Code Decoding

For reliable large-scale quantum computation, quantum error correction (QEC) is essential to protect logical information distributed across multiple physical qubits. Taking advantage of recent advances in deep learning, neural network-based decoders have emerged as a promising approach to improve the reliability of QEC. We propose the qubit-centric transformer (QCT), a novel and universal QEC decoder based on a transformer architecture with a qubit-centric attention mechanism. Our decoder transforms input syndromes from the stabilizer domain into qubit-centric tokens via a specialized embedding strategy. These qubit-centric tokens are processed through attention layers to effectively identify the underlying logical error. Furthermore, we introduce a graph-based masking method that incorporates the topological structure of quantum codes, enforcing attention toward relevant qubit interactions. Across various code distances for surface codes, QCT achieves state-of-the-art decoding performance, significantly outperforming existing neural decoders and the belief propagation (BP) with ordered statistics decoding (OSD) baseline. Notably, QCT achieves a high threshold of 18.1% under depolarizing noise, which closely approaches the theoretical bound of 18.9% and surpasses both the BP+OSD and the minimum-weight perfect matching (MWPM) thresholds. This qubit-centric approach provides a scalable and robust framework for surface code decoding, advancing the path toward fault-tolerant quantum computing.

Updated: 2026-10-06 06:10:19

标题: 基于量子比特的表面码解码变压器

摘要: 为了可靠的大规模量子计算,量子误差纠正(QEC)对于保护分布在多个物理量子比特上的逻辑信息至关重要。利用深度学习的最新进展,基于神经网络的解码器已经成为改进QEC可靠性的一种有前途的方法。我们提出了基于变压器架构和以量子比特为中心的注意机制的量子比特中心变压器(QCT),这是一种新颖而通用的QEC解码器。我们的解码器通过专门的嵌入策略将稳定器域的输入综合变换为以量子比特为中心的令牌。这些以量子比特为中心的令牌通过注意力层进行处理,以有效识别潜在的逻辑错误。此外,我们引入了一种基于图的掩蔽方法,将量子码的拓扑结构纳入其中,强调与相关量子比特交互。在各种表面码的码距下,QCT实现了最先进的解码性能,明显优于现有的神经解码器和基于有序统计解码(OSD)的信念传播(BP)基线。值得注意的是,在去极化噪声下,QCT实现了高达18.1%的阈值,这接近于18.9%的理论界限,并超过了BP+OSD和最小权重完美匹配(MWPM)阈值。这种以量子比特为中心的方法为表面码解码提供了一个可扩展和强大的框架,推动着通往容错量子计算的道路。

更新时间: 2026-10-06 06:10:19

领域: quant-ph,cs.AI,cs.LG

下载: http://arxiv.org/abs/2510.11593v3

Do I Need the Cloud? Uncertainty-Aware Step-Level Handoff for Small Language Model Agents

Small language models (SLMs) are attractive as local agent controllers because they reduce remote inference, latency, and deployment footprint, yet structured tool errors can cause an agent step to fail. Existing routers typically select a model once per query. However, agents expose sequential decision points whose difficulty dynamically changes based on intermediate observations. We propose STEPGATE, an uncertainty-aware handoff framework that scores each local SLM action and selectively escalates challenging steps to a stronger model. On a 52-task held-out single-step BFCL-derived test split, the Qwen2.5-1.5B/7B pair attains 82.7% task success with 30.8% escalation, versus 67.3% local-only and 75.4% random escalation (which uses 33.8% escalation). In a separate multi-turn evaluation, STEPGATE achieves 69.0% trajectory success and 84.0% action success using only 30.0% cloud actions, compared with 48.0%/70.5% local-only, 60.0%/78.2% random escalation, and 57.0%/77.1% query-level routing (strong-only achieves 82.0% trajectory success at 100% cloud actions). These results suggest that step-level escalation recovers a large share of the performance gap to the stronger Qwen2.5-7B backend at a matched cloud-action rate while transmitting fewer tokens remotely. However, our evaluation is limited to one model family, a single stronger backend, and scripted tasks. Furthermore, the test sets are small, multi-turn comparisons rely on paired intervals and statistical tests, and our risk tiers serve as research annotations rather than formal safety guarantees.

Updated: 2026-10-06 06:09:31

标题: 我需要云吗?小型语言模型代理的不确定性感知步级交接

摘要: 小语言模型(SLMs)作为本地代理控制器具有吸引力,因为它们减少了远程推理、延迟和部署占用空间,但是结构化工具错误可能导致代理步骤失败。现有的路由器通常每次查询选择一个模型。然而,代理暴露了随着中间观察结果动态变化的顺序决策点。我们提出了STEPGATE,这是一个基于不确定性的移交框架,对每个本地SLM动作进行评分,并有选择地将具有挑战性的步骤升级到更强的模型。在一个包含52个任务的单步BFCL派生测试集中,Qwen2.5-1.5B/7B对获得82.7%的任务成功率,升级率为30.8%,而本地仅为67.3%,随机升级为75.4%(使用33.8%的升级)。在一个单独的多轮评估中,STEPGATE仅使用30.0%的云动作,实现了69.0%的轨迹成功率和84.0%的动作成功率,相比之下,本地仅为48.0%/70.5%,随机升级为60.0%/78.2%,查询级路由为57.0%/77.1%(仅强模型在100%云动作下实现82.0%的轨迹成功)。这些结果表明,步骤级别的升级在匹配的云动作率下恢复了与更强的Qwen2.5-7B后端之间的性能差距的很大部分,同时在远程传输更少的令牌。然而,我们的评估仅限于一个模型系列、一个更强的后端和脚本任务。此外,测试集很小,多轮比较依赖于成对间隔和统计测试,我们的风险层次仅作为研究注释,而不是正式的安全保证。

更新时间: 2026-10-06 06:09:31

领域: cs.AI,cs.DC,cs.ET,cs.LG,cs.MA

下载: http://arxiv.org/abs/2610.07816v1

Consideration Circuits: Depth Separation and Universality Beyond a Single Softmax

Most feature-based choice models, classical and deep, score items and apply a single softmax. We introduce consideration circuits (CC), feature-based models of multi-stage choice defined by directed acyclic graphs of multinomial logit (MNL) units. Source units assign probabilities to menu items, and internal units combine predecessor distributions using MNL weights computed from their probability-weighted feature summaries. On a three-item compromise task with fixed non-collinear features, menu-independent random-utility models (RUM), including a single MNL unit, suffer an error bounded away from zero. For CC, in contrast, we establish a sharp depth--norm separation: increasing depth from $2$ to $3$ reduces the optimal maximum taste-vector norm for error $ε$ from $Θ(\log(1/ε)/ε)$ to $Θ(\log(1/ε))$. The depth-$2$ lower bound holds for arbitrary width and menu-independent routing biases, while a five-node depth-$3$ circuit with zero routing biases attains the logarithmic rate. More generally, we characterize two geometric conditions that are necessary and sufficient for approximating arbitrary deterministic choice tables on finite menu families. Under these conditions, depth $3$ suffices, while depth $4$ achieves optimal logarithmic norm scaling whenever the family contains a non-singleton menu. In experiments, standalone tree circuits with fewer than $600$ parameters attain the lowest mean test negative log-likelihood (NLL) among the evaluated models on four fixed-pool benchmarks and the Expedia temporal split. As output heads, CC generalize the linear MNL readout and lower mean test NLL for every tested encoder on Expedia and Trivago.

Updated: 2026-10-06 06:09:19

标题: 考虑电路:深度分离和单个Softmax之外的普适性

摘要: 大多数基于特征的选择模型,无论是经典的还是深度的,都会为项目评分并应用单一的softmax。我们引入了考虑电路(CC),这是基于特征的多阶段选择模型,由多项Logit(MNL)单元的有向无环图定义。源单元为菜单项目分配概率,内部单元使用从其概率加权特征摘要计算得到的MNL权重组合前驱分布。在一个具有固定非共线特征的三项目折衷任务中,不考虑菜单的随机效用模型(RUM),包括单个MNL单元,会出现一个与零有一定距离的错误。相比之下,对于CC,我们建立了一个明显的深度-范数分离:从$2$增加到$3$的深度会将最佳最大口味向量范数的错误$ε$从$Θ(\log(1/ε)/ε)$降低到$Θ(\log(1/ε))$。深度为$2$的下界对于任意宽度和独立于菜单的路由偏差都成立,而一个五节点深度为$3$的电路并且没有路由偏差会达到对数速率。更一般地,我们表征了两个几何条件,这些条件是近似有限菜单家族上的任意确定性选择表所必要且充分的。在这些条件下,深度为$3$就足够了,而当家族包含一个非单例菜单时,深度为$4$会实现最佳对数范数缩放。在实验中,单独的树电路,参数少于$600$个,在四个固定池基准测试和Expedia时间拆分中的评估模型中获得了最低的平均测试负对数似然(NLL)。作为输出头,CC推广了线性MNL读取,并降低了Expedia和Trivago上每个测试编码器的平均测试NLL。

更新时间: 2026-10-06 06:09:19

领域: cs.LG,stat.ML

下载: http://arxiv.org/abs/2610.04143v2

Stochastic Gradient Descent Ascent is Suboptimal for Nonconvex-PL Min-Max Games

How far can stochastic gradient descent ascent (SGDA) go by tuning its timescale ratio and step sizes in nonconvex min-max games? We answer this question for nonconvex-PL (NC-PL) games by establishing the first tight complexity of two-timescale SGDA with a fixed timescale ratio and non-increasing step sizes. For $\ell$-smooth games with an inner $μ$-PL inequality, we prove a complexity lower bound $Ω(κ^2\ell\varepsilon^{-2}+κ^4\ellσ^2\varepsilon^{-4})$, where $κ=\ell/μ$ is the condition number, $σ^2$ is the gradient variance, and $\varepsilon$ measures the outer gradient norm. This matches existing SGDA upper bounds and establishes a complexity separation from Smoothed-AGDA (Yang et al., 22'). In addition, we show that SGDA can fail to find a stationary point when its timescale ratio is as small as $o(κ^2)$. Our negative results highlight the fundamental limitation of SGDA in NC-PL games, and justify the development of alternative methods.

Updated: 2026-10-06 06:08:34

标题: 随机梯度下降上升在非凸PL最小-最大博弈中是次优的

摘要: 通过调整其时间尺度比和步长,在非凸极小-极大博弈中,随机梯度下降上升(SGDA)可以走多远?我们通过建立第一个使用固定时间尺度比和非递增步长的两时间尺度SGDA的非凸-PL(NC-PL)博弈的紧密复杂度来回答这个问题。对于内部$μ$-PL不等式的$\ell$-光滑博弈,我们证明了复杂度的下界$Ω(κ^2\ell\varepsilon^{-2}+κ^4\ellσ^2\varepsilon^{-4})$,其中$κ=\ell/μ$是条件数,$σ^2$是梯度方差,$\varepsilon$衡量外部梯度范数。这与现有的SGDA上界相匹配,并与平滑化的AGDA(Yang等人,22')建立了复杂度分离。此外,我们表明当其时间尺度比小至$o(κ^2)$时,SGDA可能无法找到一个稳定点。我们的负面结果突显了SGDA在NC-PL博弈中的基本限制,并证明了开发替代方法的正当性。

更新时间: 2026-10-06 06:08:34

领域: stat.ML,cs.GT,cs.LG

下载: http://arxiv.org/abs/2610.07814v1

Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

Muon-trained modular-arithmetic transformers can lose accuracy while retaining linearly decodable task information. Adjacent swaps localize five captured unnormalized failures to AdamW readout updates. Multiplying the actual readout displacement by the large feature mean produces a class-dependent logit offset shared across inputs that nearly reproduces each failure. Training-only decoders recover 98.20-100% held-out accuracy. Correcting cross-entropy derivative errors stabilizes five matched branches through step 100,000; four prospective accurate-CE RMS runs fail through embedding updates.

Updated: 2026-10-06 06:08:27

标题: Muon-训练的变压器中表示-读出接口的后Grokking崩溃

摘要: Muon-trained模块算术变换器在保留线性可解码任务信息的同时可能会丢失准确性。 相邻的交换将五个捕捉到的未标准化失败定位到AdamW读出更新。 将实际读出位移乘以大特征均值会产生一个跨输入的类相关逻辑偏移,几乎复制每个失败。 仅训练解码器可以恢复98.20-100%的保留准确性。 通过纠正交叉熵导数错误,通过100,000步稳定五个匹配的分支; 四个前景准确的CE RMS运行在嵌入更新中失败。

更新时间: 2026-10-06 06:08:27

领域: cs.AI,cs.LG

下载: http://arxiv.org/abs/2608.07436v2

Finding Memory Leaks in C/C++ Programs via Neuro-Symbolic Augmented Static Analysis

Memory leaks remain prevalent in real-world C/C++ software. Static analyzers such as CodeQL provide scalable program analysis but frequently miss such bugs because they cannot recognize project-specific custom memory-management functions and lack path-sensitive control-flow modeling. We present MemHint, a neuro-symbolic pipeline that addresses both limitations by combining LLMs' semantic understanding of code with Z3-based symbolic reasoning. MemHint parses the target codebase and applies an LLM to classify each function as a memory allocator, deallocator, or neither, producing function summaries that record which argument or return value carries memory ownership, extending the analyzer's built-in knowledge beyond standard primitives such as malloc and free. A Z3-based validation step checks each summary against the function's control-flow graph, discarding those whose claimed memory operation is unreachable on any feasible path. The validated summaries are injected into CodeQL and Infer via their respective extension mechanisms. Z3 path feasibility filtering then eliminates warnings on infeasible paths, and a final LLM-based validation step confirms whether each remaining warning is a genuine bug. On eight real-world C/C++ projects totaling over 3.6M lines of code, MemHint detects 54 unique memory leaks, all confirmed or fixed, at approximately $1.7 per detected bug, compared to 19 by vanilla CodeQL and 3 by vanilla Infer.

Updated: 2026-10-06 06:06:24

标题: 通过神经符号增强静态分析在C/C++程序中发现内存泄漏

摘要: 内存泄漏在现实世界的C/C++软件中仍然普遍存在。诸如CodeQL之类的静态分析器提供可伸缩的程序分析,但经常会错过此类错误,因为它们无法识别项目特定的自定义内存管理函数,并且缺乏路径敏感的控制流建模。我们提出了MemHint,这是一个神经符号管道,通过将LLMs对代码的语义理解与基于Z3的符号推理相结合,解决了这两个限制。MemHint解析目标代码库,并应用LLM将每个函数分类为内存分配器、释放器或其他,生成函数摘要,记录哪个参数或返回值携带内存所有权,将分析器的内置知识扩展到超出标准原语(例如malloc和free)之外。基于Z3的验证步骤针对函数的控制流图检查每个摘要,丢弃那些声称的内存操作在任何可行路径上都是不可达的。经过验证的摘要被注入到CodeQL和Infer中,通过各自的扩展机制。然后,通过Z3路径可行性过滤消除不可行路径上的警告,并通过最终的LLM验证步骤确认每个剩余警告是否是真正的错误。在超过360万行代码的八个现实世界的C/C++项目中,MemHint检测到54个独特的内存泄漏,所有这些都已经确认或修复,每检测到一个错误的成本约为1.7美元,而基本的CodeQL只检测到19个,基本的Infer只检测到3个。

更新时间: 2026-10-06 06:06:24

领域: cs.SE,cs.CR

下载: http://arxiv.org/abs/2603.27224v5

SIFT: Search Intent-to-Filter Transformer for Multi-Task Personalized Filter Ranking at Airbnb

Search filters help guests navigate vast catalogs in two-sided marketplaces like Airbnb, and recommending the right filters can meaningfully lift booking conversion. Many such production filter-ranking systems, however, represent the guest through hand-engineered, pre-aggregated features generated by ETL pipelines. This makes it expensive to maintain and difficult to extend for new filter types or contextual dimensions (trip length, group size). We present SIFT (Search Intent-to-Filter Transformer), a ranking model built on transformers that learns guest preferences directly from raw behavioral sequences. SIFT replaces manual feature engineering with a unified guest representation that feeds multiple prediction tasks, including booking likelihood, filter engagement, and ordinal capacity thresholds (e.g., 2+ bedrooms) -- a general framework for filter ranking in two-sided marketplaces that accommodates both boolean and numeric-range filter types. Extending SIFT to new filters requires only adding a new head, not a new feature pipeline. To keep serving fast, this guest representation is computed offline on a daily cadence rather than at request time. Offline, SIFT improves booking and amenity-engagement PR-AUC by +51.9% and +62.8% respectively over the production baseline. In online A/B testing, SIFT increased engagement with recommended filters by +20.0%, overall filter usage among searchers by +0.72%, and usage of the newly-supported bedroom, bathroom, and bed filters by +3.9%, +10.7%, and +0.52% respectively. Demonstrating the system's extensibility, we rapidly integrated a novel hotel-intent filter using the same shared representation, driving a +3.8% lift in uncancelled hotel bookings and a +0.76% lift in overall marketplace bookings. SIFT is now fully deployed in production, serving scalable personalization to millions of guests.

Updated: 2026-10-06 06:05:45

标题: SIFT: 搜索意图到过滤器转换器,用于Airbnb的多任务个性化过滤器排名

摘要: 搜索过滤器帮助客人在Airbnb等双边市场的广泛目录中导航,并推荐正确的过滤器可以显着提高预订转化率。然而,许多这种生产过滤器排序系统通过ETL管道生成的手工工程化、预聚合特征来代表客人。这使得维护成本高昂,难以为新的过滤器类型或上下文维度(行程长度、团体规模)进行扩展。我们提出了SIFT(搜索意图到过滤器转换器),这是一个建立在transformers上的排名模型,直接从原始行为序列中学习客人的偏好。SIFT用统一的客人表示替换了手动特征工程,该表示向多个预测任务提供输入,包括预订可能性、过滤器参与度和顺序容量阈值(例如,2个以上的卧室)- 这是一个在双边市场中进行过滤器排名的通用框架,适用于布尔和数值范围过滤器类型。将SIFT扩展到新的过滤器只需要添加一个新的头部,而不是一个新的特征管道。为了保持快速服务,这种客人表示是离线计算的,每天一次,而不是在请求时。在离线情况下,SIFT分别比生产基线提高了+51.9%和+62.8%的预订和便利设施参与PR-AUC。在在线A/B测试中,SIFT将推荐的过滤器的参与度提高了+20.0%,搜索者的整体过滤器使用量提高了+0.72%,新支持的卧室、浴室和床铺过滤器的使用量分别提高了+3.9%、+10.7%和+0.52%。通过展示系统的可扩展性,我们迅速集成了一个新的酒店意图过滤器,使用相同的共享表示,使未取消的酒店预订提高了+3.8%,整体市场预订提高了+0.76%。SIFT现在已完全部署在生产中,为数百万客人提供可扩展的个性化服务。

更新时间: 2026-10-06 06:05:45

领域: cs.LG

下载: http://arxiv.org/abs/2610.07810v1

Trading Strategy Optimization via Textual Gradient

Quantitative trading strategy design aims to discover trading programs from historical data that remain effective in future markets, which can be viewed as a black-box program optimization problem. LLM-based textual gradients offer a promising approach by providing explicit optimization directions for iterative strategy refinement. However, directly applying textual gradients faces two challenges: (1) optimization is myopic, underutilizing experience from previous evaluations; and (2) aggregate backtest feedback overlooks temporal robustness, potentially favoring strategies that perform well only in specific market periods. To address these challenges, we propose TradeGrad, an experience-guided textual-gradient framework for robust trading strategy optimization. TradeGrad leverages accumulated optimization experience to estimate textual gradients and employs multi-scale revisions for both strategy exploration and refinement. It further introduces the Cross-Period Robust Objective (CPRO), which emphasizes performance in unfavorable historical periods to promote temporal robustness. Experiments on cross-sectional and time-series strategy design in Chinese A-share and U.S. equity markets show that TradeGrad achieves the best in-sample and out-of-sample performance across all four settings. Notably, its Chinese cross-sectional strategy achieves 27.99% annualized return, 12.19% maximum drawdown, and a Sharpe ratio of 1.63, approximately 68% higher than the CSI 300 benchmark. Further analyses validate the proposed components and show consistent improvements in both in-sample and out-of-sample performance throughout optimization. The code is available at https://github.com/transcend-0/TradeGrad.

Updated: 2026-10-06 06:04:49

标题: 通过文本梯度进行交易策略优化

摘要: 量化交易策略设计旨在从历史数据中发现在未来市场中仍然有效的交易程序,这可以被视为一个黑盒程序优化问题。基于LLM的文本梯度提供了一种有前途的方法,通过为迭代策略改进提供明确的优化方向。然而,直接应用文本梯度面临两个挑战:(1)优化是短视的,未充分利用先前评估的经验;(2)聚合回测反馈忽视了时间鲁棒性,可能偏向表现仅在特定市场时期表现良好的策略。为了解决这些挑战,我们提出了TradeGrad,一个以经验为导向的文本梯度框架,用于稳健的交易策略优化。TradeGrad利用积累的优化经验来估计文本梯度,并采用多尺度修订进行策略探索和改进。它进一步引入了跨期鲁棒目标(CPRO),强调在不利的历史时期表现,以促进时间鲁棒性。在中国A股和美国股市的横截面和时间序列策略设计实验中,TradeGrad在所有四种设置中实现了最佳的样本内和样本外表现。值得注意的是,其中国内横截面策略实现了27.99%的年化回报率,12.19%的最大回撤率,以及1.63的夏普比率,大约比CSI 300基准高68%。进一步的分析验证了所提出的组件,并展示了在整个优化过程中样本内和样本外表现的持续改进。代码可在https://github.com/transcend-0/TradeGrad找到。

更新时间: 2026-10-06 06:04:49

领域: cs.AI

下载: http://arxiv.org/abs/2610.03128v2

MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine. The router and mask scores are optimized jointly, while the underlying FFN weight values remain unchanged. This formulation supports neuron-structured, semi-structured, and unstructured experts within the same routing architecture. Our main configuration uses four 2:4 experts with top-2 routing, where two half-dense expert passes have the nominal FFN arithmetic of one dense pass, without requiring independent expert weight matrices. On five vision-language benchmarks with Qwen and Gemma backbones, this configuration achieves the highest performance among the compared baselines. Comparisons across mask granularities, routing interventions, and compute-matched controls distinguish the effects of learned connectivity from expert activation count. These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.

Updated: 2026-10-06 06:04:38

标题: MASKerade:用于密集到MoE循环再利用的基于令牌路由的口罩专家

摘要: 稀疏激活的专家混合(MoE)模型增加了模型容量,而不需要按比例增加每个令牌的计算量。密集到MoE的升级重复利用预训练的密集模型来构建这样的系统,通常是通过将前馈网络(FFN)复制到独立训练的专家中。我们介绍了MASKerade,一种密集到MoE训练方法,它通过将专家学习为冻结的预训练FFN的稀疏子网络来定义。每个专家由一个学习到的二进制掩码定义,一个令牌级别的路由器选择执行和组合哪些掩码FFN。路由器和掩码分数是联合优化的,而底层FFN权重值保持不变。这种公式支持在相同的路由架构中具有神经元结构、半结构化和非结构化专家。我们的主要配置使用四个2:4专家和顶级2路由,其中两个半密集专家通过具有一个密集传递的FFN算术的名义FFN算术,而不需要独立的专家权重矩阵。在具有Qwen和Gemma骨干的五个视觉语言基准测试中,这种配置在与比较基准相比实现了最高性能。通过比较不同的掩码粒度、路由干预和计算匹配控制,可以区分出学习连接性与专家激活计数的影响。这些结果确立了在冻结权重上学习掩码作为构建标记路由MoE专家的实际替代方案。

更新时间: 2026-10-06 06:04:38

领域: cs.LG

下载: http://arxiv.org/abs/2610.07809v1

Common-Mode Errors Limit Low-Timestep Deep Spiking Q-Networks

Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices. In value-based RL, deep spiking Q-networks (DSQNs) combine such efficiency with action-value estimation for decision making. However, existing DSQNs often require multiple simulation timesteps for competitive performance, increasing computational and energy costs, whereas reducing the timesteps can cause substantial performance degradation. We investigate this degradation from the perspective of Q-value estimation errors. By decomposing errors across actions into common-mode and differential-mode components, we find that low-timestep DSQNs suffer disproportionately from common-mode errors shared across action values, which are particularly detrimental to temporal-difference learning through bootstrapped targets. Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs. At inference, greedy action selection can be performed directly from the SNN outputs, allowing the auxiliary ANN to be completely removed and preserving the energy efficiency of SNNs. Extensive experiments on Atari and MiniAtar environments demonstrate substantial performance improvements under low-timestep settings. CMC-DSQN outperforms state-of-the-art DSQN baselines by nearly $20\%$ at $T=2$ and further surpasses the ANN baseline at $T=4$.

Updated: 2026-10-06 06:01:28

标题: 共模误差限制低时间步长深度尖峰Q网络

摘要: 脉冲神经网络(SNNs)提供了稀疏和事件驱动的计算,使它们在边缘设备上的能源受限强化学习(RL)中具有吸引力。在基于价值的RL中,深度脉冲Q网络(DSQNs)将这种效率与动作价值估计相结合,用于决策制定。然而,现有的DSQNs通常需要多个模拟时间步才能获得竞争性能,从而增加了计算和能源成本,而减少时间步可以导致性能严重下降。我们从Q值估计误差的角度研究了这种退化。通过将不同动作的错误分解为共模和差模组件,我们发现低时间步的DSQNs在共模错误中遭受的损失要远远大于跨动作值共享的损失,这对通过引导式目标进行的时间差分学习特别有害。基于这一发现,我们提出了共模补偿深脉冲Q网络(CMC-DSQN),它使用辅助人工神经网络来补偿SNN输出中的共模错误。在推断阶段,可以直接从SNN输出中执行贪婪动作选择,从而可以完全去除辅助人工神经网络,并保持SNN的能源效率。在Atari和MiniAtar环境上进行的大量实验表明,在低时间步设置下取得了显着的性能改进。CMC-DSQN在T=2时超过了最先进的DSQN基线近20%,并在T=4时进一步超过了人工神经网络基线。

更新时间: 2026-10-06 06:01:28

领域: cs.NE,cs.LG

下载: http://arxiv.org/abs/2610.07808v1

Ideal Paths for Approximating Logistic Gradient Descent Trajectories at Large Initialization

Modern training on a new task often starts from a previously trained model rather than from scratch, raising the question of how this initialization affects the subsequent training trajectory. Classical implicit-bias results characterize the direction selected by prolonged training, but this direction alone does not provide information regarding the intermediate behavior. We address this question through a geometric approximation of full-batch logistic gradient descent (GD) trajectories on strictly linearly separable data, with large initialization of scale $R$ motivated by prior training. From any limiting normalized initial position, we use minimum-norm projection rules to construct a unique continuous ideal path consisting of finitely many linear segments. The path has two stages: negative-margin correction followed by minimum-margin growth. We prove that, after an explicit two-stage time reparameterization, the fixed-step GD trajectory divided by $R$ converges uniformly to this path on every fixed parameter interval as $R\to\infty$. Further, our quantitative error bounds account for initialization perturbations and the transition between stages. This approximation provides asymptotic formulas for peak evaluation loss and cumulative training loss. In particular, peak evaluation loss can grow linearly in $R$ even when both endpoint losses tend to zero. The cumulative losses in the correction and margin-growth stages, normalized by $R^2$ and $R$, respectively, converge to explicit limits. Experiments on controlled geometries and fixed image features complement our theoretical results.

Updated: 2026-10-06 06:01:15

标题: 大初始化时逼近逻辑梯度下降轨迹的理想路径

摘要: 现代对新任务的训练通常从先前训练好的模型开始,而不是从头开始,这引发了一个问题,即这种初始化如何影响后续的训练轨迹。经典的隐式偏差结果表征了长时间训练选择的方向,但仅凭这个方向并不能提供有关中间行为的信息。我们通过对严格线性可分数据上的全批量逻辑梯度下降(GD)轨迹进行几何逼近来解决这个问题,其中大初始化尺度为$R$受到之前训练的启发。从任何极限归一化初始位置开始,我们使用最小范数投影规则来构建一个由有限数量线性段组成的独特连续理想路径。该路径分为两个阶段:负边距校正,然后是最小边距增长。我们证明,在明确的两阶段时间重参数化后,除以$R$的固定步长GD轨迹在每个固定参数区间上当$R\to\infty$时均一致收敛到该路径。此外,我们的定量误差界考虑了初始化扰动和阶段之间的过渡。这种逼近提供了用于峰值评估损失和累积训练损失的渐近公式。特别是,即使两个端点损失都趋于零,峰值评估损失也可以线性增长。校正和边距增长阶段的累积损失,分别通过$R^2$和$R$进行归一化,收敛到明确的极限。对受控几何和固定图像特征的实验补充了我们的理论结果。

更新时间: 2026-10-06 06:01:15

领域: cs.LG,math.OC

下载: http://arxiv.org/abs/2610.04142v2

Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis

Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from. We study how such adaptivity is learned in a controlled location-estimation problem. Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order $n^{-1}$, whereas uniform data are best estimated from their extremes, at the faster rate $n^{-2}$. We also provide the example of a symmetric Gaussian mixture, for which a rate of $σ^2_n/n$ can be attained. On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function. A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range. We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow. With $\widetildeΩ(n^{1+ε})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^ε$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime. These guarantees extend to new locations and longer contexts. A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast. Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection. End-to-end experiments recover the predicted specialization.

Updated: 2026-10-06 05:57:10

标题: 通过上下文学习的自适应均值估计:梯度流分析

摘要: Prior Fitted Networks (PFNs)如TabPFN现在在预测和估计任务中与建立的统计程序相媲美。一个自然的解释是PFN具有统计适应性的属性,也就是说,它们在一个异质模型集合中表现几乎和针对真实数据生成模型定制的方法一样好,而不需要告诉它们数据来自哪个模型。我们研究了这种适应性是如何在一个受控的位置估计问题中学习的。每个任务都是一个未标记的样本,其家族是隐藏的:高斯数据需要平均值,误差为 $n^{-1}$阶,而均匀数据最好从它们的极值处进行估计,速度更快是 $n^{-2}$ 阶。我们还提供了对称高斯混合的例子,对称高斯混合的速率可以达到 $σ^2_n/n$。在标量输入上,softmax注意力计算经验累积生成函数的导数。因此,一个单一的原语既提供了区分家族的特征,又形成了在样本均值和中间范围之间插值的估计器。我们通过softmax混合专家或门控线性单元(GLU)来组合关注专家,并分析逐步梯度流。通过 $\widetildeΩ(n^{1+ε})$ 预训练任务,学习的估计器在高斯任务上是渐近有效的,在均匀任务上与极小极大速率相差 $n^ε$,在收缩方差区域的混合物上是最优的。这些保证扩展到新的位置和更长的上下文。风险分解将专家错误、路由错误和归一化错误分开,从而澄清了架构的对比。Softmax门控强制进行归一化和精确平移等变性,而GLU必须学会它:其动态分为快速偏差去除后的慢速专家选择。端到端实验验证了预测的专业化。

更新时间: 2026-10-06 05:57:10

领域: cs.LG,stat.ML

下载: http://arxiv.org/abs/2610.07804v1

ThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models

Small reasoning models (SRMs) have shown strong performance on complex reasoning tasks by generating extended chain-of-thought trajectories, but they often fail to recover once their reasoning enters an erroneous path. Existing test-time fusion methods rely on local fusion signals to determine when to trigger fusion, which can be misled by transient uncertainty fluctuations and may reinforce unstable reasoning trajectories. We propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments. ThinkFuse compares segment-level uncertainty shifts with trajectory-level uncertainty trends to identify unstable reasoning points and fuse auxiliary reasoning paths into the primary model's trajectory. Extensive experiments demonstrate that ThinkFuse outperforms baselines on mathematical and knowledge-intensive reasoning benchmarks, with consistent gains across model-family combinations, and remains robust with a smaller primary model. Our analysis shows that ThinkFuse requires fewer fusion triggers and generates fewer tokens, highlighting the efficiency of selective triggering. Our code is available at https://github.com/js-lee-AI/ThinkFuse.

Updated: 2026-10-06 05:56:21

标题: ThinkFuse:小型推理模型的轨迹感知测试时间融合

摘要: 小推理模型(SRMs)在复杂推理任务上表现出色,通过生成延长的思维链轨迹,但它们常常在推理进入错误路径后无法恢复。现有的测试时融合方法依赖于局部融合信号来确定何时触发融合,这可能会被瞬时不确定性波动误导,并可能加强不稳定的推理轨迹。我们提出ThinkFuse,一种无需训练的测试时融合框架,可以有选择地干预不可靠的推理片段。ThinkFuse将段级不确定性变化与轨迹级不确定性趋势进行比较,以识别不稳定的推理点,并将辅助推理路径融入主模型的轨迹中。大量实验证明,ThinkFuse在数学和知识密集型推理基准上优于基线,跨模型系列组合始终获得一致的增益,并在较小的主模型下保持稳健性。我们的分析表明,ThinkFuse需要更少的融合触发器并生成更少的令牌,突出了选择性触发的效率。我们的代码可在https://github.com/js-lee-AI/ThinkFuse找到。

更新时间: 2026-10-06 05:56:21

领域: cs.AI,cs.CL

下载: http://arxiv.org/abs/2610.07803v1

ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs

Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean. The same score arises as the minimal-disturbance eviction under a max-entropy assumption about future attention. We evaluate under fixed cache budgets, with eviction at every block boundary during prefill and at every decoding step during generation. On RULER at a tight 2k token budget, ValueDiff retains 88-99% of dense across seven sink-suppressed models (best on 6 out of 7). On LongBench at the 4k budget, ValueDiff averages 92% retention across sink-suppressed models versus 83% for the strongest prior baseline. On MATH-500, ValueDiff is the strongest non-dense method on every sink-suppressed model tested at the 25% cache budget, outperforming prior methods by up to ~20 points on gated-attention models. Across all three benchmarks, value geometry emerges as the more reliable query-invariant eviction signal for sink-suppressed models.

Updated: 2026-10-06 05:54:19

标题: ValueDiff:针对Sink-Suppressed LLMs的值几何KV缓存驱逐

摘要: 具有QK归一化、门控注意力、学习的注意力沉降或logit软上限的现代LLMs表现出较弱的持久性注意力沉降,现有的KV缓存驱逐方法主要依赖于这些沉降。我们观察到,在这些模型中,较弱的沉降与相对于关键向量分散更大的数值向量分散同时发生。受到这种数值侧分散的启发,我们提出了ValueDiff,一种基于数值向量与缓存平均值之间的L2偏差对令牌进行排名的数值几何驱逐方法。在关于未来注意力的最大熵假设下,相同的得分被视为最小干扰驱逐。我们在固定的缓存预算下进行评估,在预填期的每个块边界进行驱逐,在生成期的每个解码步骤进行驱逐。在RULER上,以严格的2k令牌预算,ValueDiff在七个抑制沉降模型中保留了88-99%的密集性(在7个模型中的6个中表现最佳)。在4k预算的LongBench上,ValueDiff在抑制沉降模型中平均保留92%的数据,而最强的之前基线仅保留83%。在MATH-500上,ValueDiff在25%缓存预算下是每个抑制沉降模型测试中最强的非密集方法,比门控注意力模型的之前方法表现高出约20个点。在所有三个基准测试中,数值几何形状出现为抑制沉降模型的更可靠的查询不变驱逐信号。

更新时间: 2026-10-06 05:54:19

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2609.23314v2

SimToolReal: An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation

The ability to manipulate tools significantly expands the set of tasks a robot can perform. Yet, tool manipulation represents a challenging class of dexterity, requiring grasping thin objects, in-hand object rotations, and forceful interactions. Since collecting teleoperation data for these behaviors is challenging, sim-to-real reinforcement learning (RL) is a promising alternative. However, prior approaches typically require substantial engineering effort to model objects and tune reward functions for each task. In this work, we propose SimToolReal, taking a step towards generalizing sim-to-real RL policies for tool manipulation. Instead of focusing on a single object and task, we procedurally generate a large variety of tool-like object primitives in simulation and train a single RL policy with the universal goal of manipulating each object to random goal poses. This approach enables SimToolReal to perform general dexterous tool manipulation at test-time without any object or task-specific training. We demonstrate that SimToolReal outperforms prior retargeting and fixed-grasp methods by 37% while matching the performance of specialist RL policies trained on specific target objects and tasks. Finally, we show that SimToolReal generalizes across a diverse set of everyday tools, achieving strong zero-shot performance over 120 real-world rollouts spanning 24 tasks, 12 object instances, and 6 tool categories.

Updated: 2026-10-06 05:50:23

标题: SimToolReal:一种零-shot 灵巧工具操作的面向对象策略

摘要: 工具操作的能力显著扩展了机器人可以执行的任务集。然而,工具操作代表了一种具有挑战性的灵巧类别,需要抓取薄物体、手持物体旋转和强力互动。由于为这些行为收集远程操作数据具有挑战性,因此模拟到真实的强化学习(RL)是一个有前途的替代方法。然而,以往的方法通常需要大量的工程工作来为每个任务建模对象并调整奖励函数。在这项工作中,我们提出了SimToolReal,朝着将模拟到真实的RL策略推广至工具操作迈出了一步。我们不再专注于单个对象和任务,而是在模拟中程序生成大量类似工具的对象原语,并训练一个单一的RL策略,其通用目标是将每个对象操作到随机目标位置。这种方法使得SimToolReal能够在测试时执行通用的灵巧工具操作,无需任何对象或任务特定的训练。我们证明了SimToolReal的性能优于以往的重新定位和固定抓握方法37%,同时与专家RL策略在特定目标对象和任务上训练的性能相匹配。最后,我们展示了SimToolReal在各种日常工具上的泛化能力,实现了在涵盖24个任务、12个对象实例和6个工具类别的120次真实世界试验中的强大零样本性能。

更新时间: 2026-10-06 05:50:23

领域: cs.RO,cs.AI

下载: http://arxiv.org/abs/2602.16863v3

Novice Reliance Calibration in AI-Assisted Decision Making: The Role of Explanations and Self-Assessment

Artificial Intelligence (AI) tools are widely used to support decision making in tasks and domains where no immediate performance feedback is available. In these settings, users cannot learn to adjust their reliance behavior over time through trial and error. However, little is known about how novice users calibrate reliance on AI when external feedback is unavailable, or whether AI explanations can support calibration in its absence. We introduce reliance calibration as an organizing construct for studying how novice users dynamically adjust reliance behavior, and examine how AI explanations and meta-cognitive self-assessment shape it. Through a between-subjects study with 110 participants completing a clinical entity extraction task with AI assistance and limited performance feedback, we observe that novice users exhibit systematic drift toward over-reliance in the presence of explanations, while higher self-reported task understanding is associated with more selective reliance behavior. These results extend reliance calibration research into human-AI collaboration contexts without real-time performance signals and present actionable guidelines on designing AI tools that must support appropriate reliance in these settings.

Updated: 2026-10-06 05:48:16

标题: 人工智能辅助决策中新手依赖校准:解释和自我评估的作用

摘要: 人工智能(AI)工具被广泛用于支持在没有即时性能反馈的任务和领域中做出决策。在这些设置中,用户无法通过试错来学习如何随着时间调整对AI的依赖行为。然而,我们很少了解初学者在外部反馈不可用时如何校准对AI的依赖性,或者AI解释是否可以在其缺失时支持校准。我们引入了依赖校准作为一个组织构建来研究初学者如何动态调整依赖行为,并研究了AI解释和元认知自我评估如何塑造它。通过一项110名参与者完成临床实体提取任务的不同主体研究,我们观察到初学者在解释存在的情况下向过度依赖的方向系统漂移,而较高的自我报告任务理解与更有选择性的依赖行为相关。这些结果将依赖校准研究扩展到没有实时性能信号的人工智能协作环境,并提出了设计必须支持这些设置中适当依赖的AI工具的可操作指南。

更新时间: 2026-10-06 05:48:16

领域: cs.HC,cs.AI

下载: http://arxiv.org/abs/2610.07800v1

Thin Evidence, Thick Priors: How Language Models Substitute Identity for Missing Financial Facts

People increasingly ask large language models what to do with their money, yet seldom describe their finances in full. This paper asks what a model does with the gap. Holding finances fixed and changing only who the investor is said to be, we grade the financial evidence in the prompt from eight facts to none and measure how far the recommended equity allocation moves. Across 96,600 prompts to Llama-3.1-8B-Instruct, built from 100 financial profiles, 138 personas and seven disclosure conditions, the average gap between two personas with identical finances rises from 4.78 percentage points at full disclosure to 10.34 points with no financial facts. A two-way cluster bootstrap counting duplicated prompts once places the ratio at 2.16 (95% interval 1.69 to 2.79), and the rise is already 1.69-fold with a single fact left. Identity explains 5% of within-profile variation in advice at full disclosure and 96% with no disclosure. Household size is the only attribute whose influence grows reliably as evidence is withdrawn. Once standard errors are clustered on the persona, the unit to which identity was assigned, most attribute-specific interactions reported in the conference version lose significance, and gender instead appears as a small standing gap that full disclosure does not close. Stating risk appetite alone brings the swing into the range seen with two to seven generic facts. With no facts, the model's one-line rationale cites incomes, debts and savings it was never told, and these invented finances turn adverse more often for larger households. Inside the network, gender is linearly decodable at every layer, and ablating the gender direction at five layers leaves the aggregate identity swing unchanged. Advisory systems built on such models should be audited at the disclosure levels users actually reach, and judged across the whole identity space rather than one attribute at a time.

Updated: 2026-10-06 05:46:28

标题: 稀疏证据,丰富先验:语言模型如何用身份替代缺失的财务事实

摘要: 人们越来越经常向大型语言模型询问如何处理他们的金钱,但很少会完整描述他们的财务状况。本文探讨了模型如何处理这种差距。保持财务状况不变,只改变投资者的身份,我们对提示中的财务证据进行了评分,从八个事实到零,并测量推荐的股权配置移动的距离。通过对Llama-3.1-8B-Instruct的96,600个提示进行分析,这些提示由100个金融档案、138个人设和七种披露条件构建,结果显示,两个具有相同财务状况的人设之间的平均差距从完全披露时的4.78个百分点增加到没有任何财务事实时的10.34个百分点。双向聚类自举计算重复提示一次,将比率确定为2.16(95%区间为1.69至2.79),而只剩下一个事实时,这个比率已经增加了1.69倍。在完全披露时,身份解释了建议内部变化的5%,而在没有披露时占比高达96%。家庭规模是唯一一个随着证据减少而影响稳定增长的属性。一旦标准误差以人设为单位进行聚类,会导致会议版本中报告的大多数属性特定交互作用失去显著性,性别取而代之成为一个完全披露无法弥补的小差距。单独陈述风险偏好会使波动范围达到具有两到七个通用事实的范围。在没有任何财务事实的情况下,模型的一行推理引用了它从未被告知的收入、债务和储蓄情况,这些虚构的财务状况在较大家庭中更容易变得不利。在网络内部,性别在每一层都可以线性解码,而在五层中消除性别方向后,整体身份波动保持不变。建立在这种模型上的咨询系统应该在实际达到的披露水平上进行审计,并且应该跨越整个身份空间而不是逐个属性进行评判。

更新时间: 2026-10-06 05:46:28

领域: cs.AI

下载: http://arxiv.org/abs/2610.07798v1

Explaining Attention with Program Synthesis

A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs. We focus on attention heads in transformer language models. For a given head, we first compute its associated attention matrices on a collection of randomly selected training examples. Next, we prompt a pre-trained language model with a summary of these matrices, and instruct it to generate a set of Python programs that can reproduce the associated attention patterns given only text from the input sentence. Finally, we re-rank programs according to how well our final set of programs predict behavior on held-out inputs. We demonstrate that a set of fewer than 1,000 such generated programs can reproduce the attention patterns of heads in GPT-2, TinyLlama-1.1B, and Llama-3B, achieving an average Intersection-over-Union similarity above 75% on TinyStories. Moreover, the best-fit programs can replace neural attention heads without substantially affecting model behavior: replacing 25% of attention heads with programmatic surrogates across the three models incurs only a 16% average perplexity increase, while maintaining performance on a variety of downstream question answering benchmarks. This work contributes a scalable pipeline for reverse-engineering attention heads in transformer models using human-readable, executable code, advancing a path toward symbolic transparency in neural models.

Updated: 2026-10-06 05:46:25

标题: 用程序综合解释注意力

摘要: 深度学习可解释性研究的一个长期目标是用人类可理解的符号描述替代不透明的神经计算。在本文中,我们提出了一种逼近深度网络组件行为的方法,该方法使用可执行程序。我们专注于transformer语言模型中的注意力头。对于给定的头部,我们首先计算其在一组随机选择的训练示例上的关联注意力矩阵。接下来,我们使用这些矩阵的摘要提示一个预训练语言模型,并指示其生成一组Python程序,这些程序只需输入句子的文本就可以复制关联的注意力模式。最后,我们根据我们的最终程序集对程序重新排序,以了解它们在保留输入上的行为预测方面表现如何。我们证明,少于1,000个这样生成的程序集可以再现GPT-2、TinyLlama-1.1B和Llama-3B中头部的注意力模式,在TinyStories上实现75%以上的平均交集相似度。此外,最佳拟合程序可以替换神经注意力头,并且几乎不影响模型行为:在三个模型中用程序替换25%的注意力头仅导致平均困惑度增加16%,同时保持对各种下游问题回答基准测试的性能。这项工作提供了一个可扩展的管道,用于使用人类可读的可执行代码逆向工程transformer模型中的注意力头,推进了神经模型中符号透明性的路径。

更新时间: 2026-10-06 05:46:25

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2606.19317v3

The Geometry of Empowerment

Empowerment captures the capacity for an agent to actively control its environment. While conceptually appealing as an information-theoretic quantity, the connection between empowerment and structurally central states that provide broad access to future outcomes has remained an open question. In this work, we link empowerment maximization and skill-learning methods to provide new geometries for interpreting and analyzing empowerment. Our analyses answer longstanding open questions on the connections between empowerment and structural centrality. Our analyses also reveal distinctions between information and reward geometries, highlighting important theoretical implications to build scalable empowerment-maximization methods. Website and code can be found at https://empowerment-geometry.github.io/.

Updated: 2026-10-06 05:45:19

标题: 赋权的几何形态

摘要: 授权捕捉代理人积极控制其环境的能力。虽然在信息理论数量方面概念上具有吸引力,但授权与提供广泛访问未来结果的结构中心状态之间的联系仍然是一个悬而未决的问题。在这项工作中,我们将授权最大化和技能学习方法联系起来,提供了新的几何形态来解释和分析授权。我们的分析回答了关于授权和结构中心性之间联系的长期未决问题。我们的分析还揭示了信息和奖励几何之间的区别,突出了建立可扩展授权最大化方法的重要理论意义。网站和代码可以在https://empowerment-geometry.github.io/找到。

更新时间: 2026-10-06 05:45:19

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.07796v1

World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models

A growing literature shows that variables can be linearly decoded from the activations of large language models (LLMs). These range from properties of the world, such as the locations of cities and the lifetimes of historical figures, to emotions and pain. Such findings are often taken as evidence that language models go beyond surface text statistics and form internal models of the world. We show that static word embeddings (fixed, context-insensitive representations learned from corpus statistics) of the same or matched stimuli support much of the same decoding. Across four published cases (place, time, pain and emotion), static vectors predict coordinates and year of death (R^2 = 0.42-0.59), separate pain from matched control sentences (held-out AUC 0.85-0.88), and classify twelve emotions in stories written to avoid naming them (AUC 0.84-0.88). Because static embeddings assign each word a single, context-independent vector, these results are a lower bound on what word associations alone can support. The LLMs retain clear advantages on representational tests, and causal and behavioral findings remain outside the scope of the baseline. On the original authors' entities, where we reproduce their Llama-2 results, the transformer's advantage lies mostly in placing historical figures in the right century and places in the right country, coarse sorting that richer word associations would be expected to improve; within those groups every representation orders items poorly. Static vectors for disambiguated Wikipedia entities, which carry the associations of a particular place or person rather than of the words in its name, close most of the remaining gap, matching Pythia-2.8B on coordinates and Llama-2-7B on year of death. These results indicate that decodability alone cannot distinguish a representation of a property from information already available in fixed distributional associations.

Updated: 2026-10-06 05:43:50

标题: 没有世界模型的世界属性:分布关联与语言模型解码结果的解释

摘要: 一个日益增长的文献表明,可以从大型语言模型(LLMs)的激活中线性解码变量。这些变量范围从世界的属性,如城市的位置和历史人物的寿命,到情绪和疼痛。这些发现通常被视为语言模型超越表面文本统计并形成对世界的内部模型的证据。我们展示了静态词嵌入(从语料库统计学习得到的固定、与上下文无关的表示)支持相同解码的相同或匹配的刺激。在四个已发表的案例中(地点、时间、疼痛和情绪),静态向量预测坐标和死亡年份(R^2 = 0.42-0.59),将疼痛与匹配的对照句子分开(保留的AUC 0.85-0.88),并对避免命名它们的故事中的十二种情绪进行分类(AUC 0.84-0.88)。由于静态嵌入为每个单词分配一个单一的、与上下文无关的向量,这些结果是单凭词汇联想能支持的下限。在表示测试中,LLMs保留了明显的优势,而因果和行为发现仍超出了基线的范围。在原始作者的实体中,我们重现了他们的Llama-2结果,变压器的优势主要在于将历史人物放在正确的世纪和地点以及将地点放在正确的国家,这种粗略的分类富有的词汇联想应该会改善;在这些组内,每个表征都很差地排序项目。对于消除歧义的维基百科实体的静态向量,这些向量携带特定地点或人物的关联,而不是其名称中的单词的关联,几乎弥合了剩余差距,与Pythia-2.8B在坐标上匹配,与Llama-2-7B在死亡年份上匹配。这些结果表明,仅仅通过可解码性无法区分一个属性的表示与已经在固定分布关联中可用的信息。

更新时间: 2026-10-06 05:43:50

领域: cs.CL,cs.AI,cs.LG

下载: http://arxiv.org/abs/2603.04317v2

Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation

Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models produce confident hallucinations, even when wrong answers carry catastrophic consequences. Our evaluations on GSM8K, MedQA and GPQA show frontier models almost never abstain despite explicit warnings of severe penalties, suggesting that prompts cannot override training that rewards any answer over no answer. As a remedy, we propose Reinforced Hesitation (RH): a modification to Reinforcement Learning from Verifiable Rewards (RLVR) to use ternary rewards (+1 correct, 0 abstention, -$λ$ error) instead of binary. Controlled experiments on logic puzzles reveal that varying $λ$ produces distinct models along a Pareto frontier, where each training penalty yields the optimal model for its corresponding risk regime: low penalties produce aggressive answerers, high penalties conservative abstainers. The same frontier holds on MATH Levels 4--5 and on medical QA, where it transfers to an unseen dataset. We then introduce two inference strategies that exploit trained abstention as a coordination signal: cascading routes queries through models with decreasing risk tolerance, while self-cascading re-queries the same model on abstention. Both outperform majority voting with lower computational cost. These results establish abstention as a first-class training objective that transforms ``I don't know'' from failure into a coordination signal, enabling models to earn trust through calibrated honesty about their limits.

Updated: 2026-10-06 05:40:49

标题: 诚实胜于准确:通过强化犹豫实现可信赖的语言模型

摘要: 现代语言模型未能满足可信情报的基本要求:知道何时不回答。尽管在基准测试中取得了令人印象深刻的准确性,这些模型产生自信的幻觉,即使错误答案会带来灾难性后果。我们在GSM8K、MedQA和GPQA上的评估显示,尽管明确警告会受到严重惩罚,前沿模型几乎从不放弃,这表明提示无法覆盖奖励任何答案胜过不回答的训练。作为补救措施,我们提出了强化犹豫(RH):对可验证奖励强化学习(RLVR)进行修改,使用三元奖励(+1正确,0放弃,-$λ$错误)而不是二元。对逻辑谜题的控制实验表明,变化的 $λ$ 会产生不同的模型,每个训练惩罚会产生其相应风险体制的最佳模型:低惩罚会产生激进的答题者,高惩罚会产生保守的放弃者。在数学4-5级和医学问答上,相同的前沿模型也适用,且可转移到未知数据集。然后我们引入了两种推理策略,利用训练好的放弃作为协调信号:级联路由通过具有降低风险容忍度的模型进行查询,而自我级联在放弃时重新查询相同的模型。这两种策略在低计算成本的情况下优于多数投票。这些结果确立了放弃作为一个头等的训练目标,将“我不知道”从失败转变为协调信号,使模型通过对其限制的诚实进行校准来赢得信任。

更新时间: 2026-10-06 05:40:49

领域: cs.LG

下载: http://arxiv.org/abs/2511.11500v3

ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience. Yet the effectiveness and limitations of these methods are not yet well characterized. Existing benchmarks provide only partial coverage: some explicitly provide the target knowledge, others assume a static environment, and those that support continual adaptation remain limited in scale and knowledge diversity. To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks. We evaluate five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness, and Prime) across six models (GPT-5.6 Terra, Opus 5, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, and GLM-5.3 Flash), covering 28 model-harness pairs and 252 learning runs. Our evaluation reveals three main findings: a substantial gap remains between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; and insufficient exploration emerges as a key bottleneck to effective adaptation. Overall, ServeLearnBench provides a controlled testbed for diagnosing these limitations and tracking progress toward agents that continually and reliably improve through serving experience.

Updated: 2026-10-06 05:39:42

标题: ServeLearnBench:代理人如何能够通过服务经验自我提升?

摘要: 大型语言模型代理越来越多地被部署用于在现实环境中执行复杂任务。然而,在这些环境中正确行为所需的知识通常是隐含的、未披露的,并且随着时间的推移而发生变化。最近的持续学习方法试图通过使代理从服务经验中改进来解决这一挑战。然而,这些方法的有效性和局限性尚未得到很好的表征。现有的基准仅提供了部分覆盖范围:有些明确提供目标知识,其他一些假设环境是静态的,而那些支持持续适应的基准在规模和知识多样性上仍然有限。为了实现系统评估,我们形式化了一个不断发展的环境流数据集(EESD),在这个数据集中,代理必须从互动和结果反馈中推断、应用和修订潜在的环境知识,因为隐藏的策略不断发展,并引入了ServeLearnBench,涵盖了零售支持、银行业务和销售话术生成,包括53个环境窗口和7,718个任务。我们评估了五种学习方法(RAG、Mem0、SkillOpt、Continual Harness和Prime)以及六种模型(GPT-5.6 Terra、Opus 5、Kimi K3、GLM-5.3、DeepSeek V4.1 Flash和GLM-5.3 Flash),涵盖了28个模型-学习方法对和252次学习运行。我们的评估揭示了三个主要发现:任务能力与通过经验学习之间仍然存在重大差距;持续适应是昂贵的,并且可能会降低已经正确的行为;不足的探索成为有效适应的关键瓶颈。总的来说,ServeLearnBench为诊断这些限制提供了一个受控的测试平台,并跟踪朝着能够通过服务经验不断可靠地改进的代理的进展。

更新时间: 2026-10-06 05:39:42

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2610.07792v1

KO: Kinetics-inspired Neural Optimizer with PDE Simulation Approaches

The design of effective optimization algorithms for neural networks remains a fundamental challenge, and most existing methods rely on heuristic extensions of gradient-based updates. We introduce KO (Kinetics-inspired Optimizer), a plug-and-play optimization module grounded in kinetic theory and partial differential equations. KO models parameter dynamics as a particle system, augmenting standard gradient updates with stochastic interactions induced by a discretization of the Boltzmann transport equation. This mechanism naturally promotes parameter diversity and mitigates weight condensation, the tendency of parameters to collapse into low-dimensional subspaces, a phenomenon closely associated with degraded generalization. We provide both a rigorous theoretical analysis and a physical interpretation, showing that KO provably increases parameter diversity while preserving convergence guarantees. Extensive experiments on image classification benchmarks (CIFAR-10/100, ImageNet) and large-scale language model pretraining demonstrate that KO consistently improves accuracy over competitive baselines with negligible additional computational cost.

Updated: 2026-10-06 05:39:39

标题: KO:受动力学启发的神经优化器与PDE模拟方法

摘要: 为神经网络设计有效的优化算法仍然是一个基本挑战,大多数现有方法依赖于基于梯度的启发式扩展。我们引入了KO(Kinetics-inspired Optimizer),一种基于动力学理论和偏微分方程的即插即用优化模块。KO将参数动态建模为一个粒子系统,通过对Boltzmann传输方程的离散化引起的随机相互作用来增强标准梯度更新。这种机制自然地促进了参数多样性并减轻了权重凝聚,即参数倾向于崩溃到低维子空间,这种现象与降级泛化密切相关。我们提供了严格的理论分析和物理解释,表明KO可以明显增加参数多样性同时保持收敛保证。对图像分类基准(CIFAR-10/100,ImageNet)和大规模语言模型预训练进行了大量实验,结果表明KO始终在可忽略的额外计算成本下提高了准确性。

更新时间: 2026-10-06 05:39:39

领域: cs.LG,cs.AI

下载: http://arxiv.org/abs/2505.14777v2

Illusory Pattern Perception Drives Spurious Inference in Large Language Models

Illusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random. Such a tendency, often described as "connecting the dots" where none exist, can result in systematic reasoning errors. This paper investigates whether Large Language Models (LLMs) exhibit such perceptual tendencies, which can lead to systematic errors in downstream applications. To our knowledge, this work presents the first systematic study of illusory pattern perception in LLMs, adapting classic psychological paradigms to three tasks with direct empirical comparison to human behaviors. We find that LLMs frequently exhibit stronger illusory pattern perception than humans. In particular, models tend to over-associate frequent positive attributes with majority groups or large organizations, and show increased tendencies to construct causal narratives from ambiguous events. To uncover the mechanism behind these behaviors, we develop a feature interpretability framework based on Sparse Autoencoders (SAEs) to analyze internal representations. Our results reveal that holistic frequency perception and analytic cognitive orientation are linked to the emergence of illusory perceptions. These findings highlight a previously underexplored cognitive-like illusion that may affect the reliability of LLM reasoning. Code available at https://github.com/NusIoraPrivacy/illusory.

Updated: 2026-10-06 05:37:12

标题: 虚幻模式感知驱动大型语言模型中的错误推断

摘要: 虚假模式感知是一个被广泛记录的人类认知倾向,即在实际上是随机数据中推断出有意义的关系。这种倾向通常被描述为“连接不存在的点”,可能导致系统性推理错误。本文调查了大型语言模型(LLMs)是否表现出这种感知倾向,这可能导致下游应用中的系统性错误。据我们所知,这项工作是对LLMs中虚假模式感知的第一次系统研究,通过经典心理范式调整为三个任务,并与人类行为进行直接的实证比较。我们发现LLMs经常表现出比人类更强的虚假模式感知。特别是,模型倾向于将常见的积极属性与多数群体或大型组织过度关联,并表现出从模棱两可事件中构建因果关系叙事的倾向增强。为了揭示这些行为背后的机制,我们开发了一个基于稀疏自动编码器(SAEs)的特征可解释性框架来分析内部表示。我们的结果显示,整体频率感知和分析认知取向与虚假感知的出现相关。这些发现突显了一种以前未被深入探讨的类似认知的错觉,可能影响LLMs推理的可靠性。代码可在https://github.com/NusIoraPrivacy/illusory找到。

更新时间: 2026-10-06 05:37:12

领域: cs.AI

下载: http://arxiv.org/abs/2610.07791v1

How Much Can Language Models Gain from Test-Time Computation?

How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision under fixed multiples of the Direct budget, and charges every model call, including selection and critique, in dollars. This design supports two kinds of comparison: the gain a model obtains from additional inference, and a lower-cost model with additional inference against a stronger model. Across five low-cost reasoning models on 350 sealed tasks, with Claude Opus 5.5 Direct as the reference, the returns depend on the domain, the selection rule, and failure handling. When we replay the retained programming candidate pools, public-example selection raises correct submissions from 376 to 453 of 500 scheduled cells while saving 12-49% of logical API cost across models, and simply retaining an available candidate when judging fails recovers 61 submissions at unchanged cost. On identical mathematics pools, judging with fallback yields 186 correct submissions versus 182 for voting, while voting saves 12-21% of logical API cost. These controlled replays show how selection and failure handling change the gains realized from the same generated candidates, and they quantify the marginal value of a model judge.

Updated: 2026-10-06 05:32:37

标题: 语言模型在测试时的计算能获得多少收益?

摘要: 测试时间计算可以提高语言模型多少,并且代价是多少?测试时间缩放被广泛提议作为更大模型的替代方案,但是现有的比较大多数只评估一个领域,并且很少将选择计入预算。我们引入了SELF-POT,这是一个基准和评估框架,可以测量模型在竞争数学、竞争性编程和自主工作流中的测试时间潜力。SELF-POT将候选覆盖率与静态任务的最终准确性分开,跟踪修订下的正确性转换,并在自主环境中测量协议完成以及任务成功。在统一的预算规则下,它比较了直接推理与并行采样和在直接预算的固定倍数下的自我修订,并且对每一次模型调用都进行收费,包括选择和批评,以美元计价。这种设计支持两种比较:模型从额外推理中获得的收益,以及低成本模型与更强模型之间的额外推理。在350个封闭任务上使用五个低成本推理模型,以Claude Opus 5.5 Direct作为参考,收益取决于领域、选择规则和失败处理。当我们重播保留的编程候选池时,公共示例选择将正确提交从500个计划单元的376提高到453,同时在各个模型之间节省了12-49%的逻辑API成本,并且当判断失败时仅保留一个可用候选可以在不改变成本的情况下恢复61个提交。在相同的数学候选池中,使用备用方案进行判断产生了186个正确提交,而投票产生了182个,而投票可以节省12-21%的逻辑API成本。这些受控的重播展示了选择和失败处理如何改变了从相同生成的候选人中实现的收益,并且量化了模型评判的边际价值。

更新时间: 2026-10-06 05:32:37

领域: cs.LG

下载: http://arxiv.org/abs/2610.01110v2

Learning to Configure Agentic AI Systems

Configuring LLM-based agent systems involves choosing workflows, tools, token budgets, and prompts from a large combinatorial design space, and is typically handled today by fixed templates or hand-tuned heuristics that apply the same configuration regardless of query difficulty, leading to brittle behavior and wasted compute. To address this, we formulate agent configuration as a semi-Markov decision process (SMDP) where each configuration acts as a temporally extended option that determines how an agent system processes a query, and introduce introduce ARC (Agentic Resource & Configuration learner), a lightweight hierarchical policy that dynamically selects query-specific agent configurations. Across reasoning, tool-use, and agentic benchmarks, ARC consistently improves over budget-matched tool-augmented LLMs, increasing average reasoning accuracy by 31.3%, tool-use accuracy by 13.95%, and doubling τ-Bench (Airline) Pass^1 success from 9.0% to 18.0%. These results demonstrate that learning per-query agent configurations is a powerful alternative to "one size fits all" designs.

Updated: 2026-10-06 05:32:17

标题: 学习配置主动智能系统

摘要: 配置基于LLM的代理系统涉及从大量组合设计空间中选择工作流程、工具、令牌预算和提示,通常通过固定模板或手动调整的启发式方法来处理,这些方法不考虑查询难度,导致系统行为脆弱且计算资源浪费。为了解决这个问题,我们将代理配置形式化为一个半马尔可夫决策过程(SMDP),其中每个配置都作为一个时间延长的选项,决定了代理系统如何处理查询,并引入ARC(代理资源和配置学习器),这是一个轻量级的分层策略,动态选择特定于查询的代理配置。在推理、工具使用和代理基准测试中,ARC始终优于预算匹配的工具增强型LLMs,平均推理准确性提高了31.3%,工具使用准确性提高了13.95%,将τ-Bench(航空公司)通过率从9.0%提高到18.0%。这些结果表明,学习每个查询的代理配置是“一刀切”设计的一个强大替代方案。

更新时间: 2026-10-06 05:32:17

领域: cs.AI

下载: http://arxiv.org/abs/2602.11574v5

OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation

Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow per benchmark that is applied uniformly to all queries. This assumption fails under realistic conditions. Query difficulty varies widely within a task, and real-world workloads mix heterogeneous task types. We introduce OOPMAS, a training-free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are represented as object-oriented class definitions with dedicated roles, tools, and persistent state, and workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in-context improvement without any gradient updates or fine-tuning. On a mixed-task benchmark of queries spanning code, math, and QA, OOPMAS achieves 89.6% accuracy, outperforming the strongest baseline by 18.1 percentage points. A model-swap study across four LLM backbones shows consistent scaling, reaching 92.4% with the strongest model.

Updated: 2026-10-06 05:31:29

标题: OOPMAS:面向对象的多代理系统用于查询级工作流生成

摘要: 多智能体系统(MAS)由大型语言模型驱动,已经在代码生成、数学推理和问题回答等方面表现出强大的性能。然而,现有的自动化MAS设计方法主要在任务级别操作,为每个基准测试生成一个固定的工作流程,然后统一应用于所有查询。这种假设在现实条件下是错误的。在任务中,查询的难度差异很大,并且真实世界的工作负载会混合不同类型的任务。我们介绍了OOPMAS,这是一个无需训练的框架,可以在单个查询的粒度上生成智能体集合和协调工作流程。智能体被表示为面向对象的类定义,具有专门的角色、工具和持久状态,工作流程则表示为对这些智能体对象的可执行主要函数。一个动态技能库可以根据优化轮次中的执行反馈积累结构化教训,实现在上下文中的改进,而无需任何梯度更新或微调。在涵盖代码、数学和问答的混合任务基准测试中,OOPMAS实现了89.6%的准确率,比最强基线高出18.1个百分点。通过对四个LLM主干进行模型交换研究,显示出一致的扩展性,最强模型达到了92.4%的准确率。

更新时间: 2026-10-06 05:31:29

领域: cs.AI

下载: http://arxiv.org/abs/2610.07787v1

SCSM: A Traffic-Native Foundation Model for Transferable Website Fingerprinting

Website fingerprinting infers the websites visited by users from encrypted traffic metadata. However, models trained under fixed collection conditions often degrade as website sets, collection times, network paths, browsers, or defenses change. Existing transferable attacks either rely on handcrafted perturbations of individual traces or adapt language-oriented architectures to traffic, limiting their ability to capture traffic-native semantics. To address these limitations, we propose SCSM, a traffic-native foundation model for transferable website fingerprinting. Specifically, SCSM constructs pairs of pretraining views from the same group of unlabeled traces through Segmentation, Combination, Scaling, and Masking. These operations produce diverse observable patterns while preserving the underlying packet events and local traffic dynamics of real trace fragments. The corresponding windowed traffic counting matrices serve as inputs for contrastive pretraining of a state-space encoder without website annotations. The pretrained encoder is then fine-tuned on a small labeled support set, and predictions are aggregated across temporal scales at inference. Experimental results demonstrate that SCSM surpasses the strongest baselines by 16.4\% in average top-3 accuracy over six temporal-drift tasks on GTT23 and by 13.8\% in average macro-F1 over four cross-domain datasets. The code and datasets will be made available at https://github.com/SJTU-dxw/WF-SCSM.

Updated: 2026-10-06 05:12:35

标题: SCSM:一种用于可转移网站指纹识别的流量本地基础模型

摘要: 网站指纹识别通过加密流量元数据推断用户访问的网站。然而,在固定收集条件下训练的模型经常随着网站集合、收集时间、网络路径、浏览器或防御措施的变化而退化。现有的可转移攻击要么依赖于手工制作的对单个跟踪的扰动,要么将面向语言的架构调整为流量,限制了它们捕捉流量本机语义的能力。为了解决这些限制,我们提出了SCSM,这是一个用于可转移网站指纹识别的流量本机基础模型。具体而言,SCSM通过分割、组合、缩放和遮罩从相同组别的未标记跟踪中构建预训练视图对。这些操作产生多样的可观察模式,同时保留真实跟踪片段的基础数据包事件和本地流量动态。相应的窗口化流量计数矩阵作为无需网站注释的状态空间编码器的对比预训练输入。预训练的编码器然后在一个小的标记支持集上进行微调,并在推断时跨时间尺度聚合预测。实验结果表明,SCSM在GTT23的六个时间漂移任务中平均前三准确性超过最强基线16.4%,在四个跨领域数据集上平均宏F1超过13.8%。代码和数据集将在https://github.com/SJTU-dxw/WF-SCSM上提供。

更新时间: 2026-10-06 05:12:35

领域: cs.CR

下载: http://arxiv.org/abs/2610.07776v1

The Deception Delta: Adversarial Evaluation of LLM-Based Smart Contract Bytecode Forensics

Large language models are increasingly used in blockchain forensic investigations to interpret unverified smart contract bytecode. Their robustness has not been systematically tested against contracts adversarially designed to mislead analysis. We evaluate 22 frontier models on 13 purpose-built contracts (9 deception vectors, 4 controls) across six prompt strategies, yielding 8,528 analyzable non-refusal runs against contracts with documented ground truth. A calibrated LLM-as-judge pipeline, supported by two judge-independent metrics and 50 human gold-standard labels, shows that adversarial deception reduces drain detection by 20.0 percentage points (95% CI: [17.2, 22.8]) relative to functionally matched controls. Structural camouflage via multi-hop call chains, XOR-masked selectors, and storage-loaded drain parameters resists detection across nearly all models. Beyond non-detection, we identify rationalization: models correctly describe the hidden drain mechanism but accept the contract's deceptive framing and dismiss it as benign, yielding positive but incorrect evidence of safety. Simple guard instructions provide no aggregate benefit and destabilize individual models in both directions. Structural deception is largely insensitive across the six tested prompt strategies, more consistent with a capability limitation than with a simple prompting problem. Only five models from two providers exceed 50% detection. Under our single-shot, raw-bytecode-only protocol, current LLMs are not reliable standalone forensic tools. Our central claim does not extend to multi-turn, tool-augmented, source-aware, or decompiler-in-the-loop workflows; a source-code boundary check is reported as an explicit subset analysis rather than as part of the main evaluation.

Updated: 2026-10-06 05:09:20

标题: 欺骗三角洲:基于LLM的智能合约字节码取证的对抗性评估

摘要: 大型语言模型在区块链取证调查中越来越被用于解释未经验证的智能合约字节码。它们的稳健性尚未经过系统化测试,以对抗对旨在误导分析的合约。 我们评估了22个前沿模型对13个专门设计的合约(9个欺骗向量,4个控制)的六种提示策略,产生了8528次可分析的非拒绝运行,针对具有记录的实际情况的合约。一个经过校准的LLM作为评判者的管道,支持两个独立的评判指标和50个人类金标准标签,表明对抗性欺骗相对于功能匹配的控制减少了20.0个百分点(95% CI: [17.2, 22.8]) 的排水检测。 通过多跳调用链、XOR掩码选择器和存储加载的排水参数的结构伪装抵抗了几乎所有模型的检测。除了非检测,我们还确定了合理化:模型正确描述了隐藏的排水机制,但接受了合约的欺骗框架并将其视为无害,产生了安全的但错误的证据。简单的保护指令没有提供总体益处,并在两个方向上破坏了单个模型的稳定性。 在测试的六种提示策略中,结构欺骗在很大程度上不敏感,更符合能力限制而非简单的提示问题。只有来自两个提供商的五个模型超过了50%的检测率。 根据我们的单次、仅原始字节码的协议,当前的LLM不是可靠的独立取证工具。我们的主要观点不适用于多轮、工具增强、源码感知或反编译器循环工作流程;源代码边界检查被报告为显式的子集分析,而不是作为主要评估的一部分。

更新时间: 2026-10-06 05:09:20

领域: cs.CR

下载: http://arxiv.org/abs/2609.14098v2

What Response Marginals Miss: Adaptive Query Complexity of Functional Backdoor Recovery

Functional backdoor recovery finds any trigger whose attack success rate is at least a given threshold rather than to recover the planted trigger. We study the minimum number of queries required for this task under label feedback which returns the predicted class label. We construct two finite families of victim models that have exactly the same attack success rate for every victim and trigger candidate. The distribution of returned labels for every query is also identical under a uniformly chosen victim. These families form an explicit counterexample that despite the matched quantities, their optimal adaptive query complexities are \(Θ(\log H)\) and \(Θ(H)\) where \(H\) is the number of possible victims. The difference arises because the same non-target labels are associated with different sets of victims, so successive queries eliminate possible victims at different rates. This separation disappears when the response is reduced to binary feedback, which reports only whether the target label is returned. The separation also persists for every fixed failure probability below one. Finally, we realize the same recovery problems with trained CIFAR-10 ResNet-18 classifiers and verify the predicted optimal query budgets. These results show that attack success rate and the distribution of returned labels for each query are insufficient to determine the query complexity of functional backdoor recovery.

Updated: 2026-10-06 05:06:43

标题: Response Marginals错过了什么:功能性后门恢复的自适应查询复杂性

摘要: Functional backdoor recovery是找到攻击成功率至少达到给定阈值的任何触发器,而不是恢复植入的触发器。我们研究了在返回预测类标签的标签反馈下完成此任务所需的最少查询次数。我们构建了两个有限的受害者模型族,这些模型对于每个受害者和触发器候选者的攻击成功率完全相同。对于每个查询返回的标签分布在均匀选择的受害者下也是相同的。这些族构成了一个明确的反例,尽管数量匹配,但它们的最佳自适应查询复杂性分别为\(Θ(\log H)\)和\(Θ(H)\),其中\(H\)是可能受害者的数量。差异是因为相同的非目标标签与不同的受害者集相关联,因此连续的查询以不同的速度消除可能的受害者。当响应减少为二进制反馈时,即仅报告是否返回目标标签时,这种分离消失。在每个低于一的固定失败概率下,这种分离也持续存在。最后,我们使用经过训练的CIFAR-10 ResNet-18分类器实现了相同的恢复问题,并验证了预测的最佳查询预算。这些结果表明,攻击成功率和每个查询返回的标签分布不足以确定功能性后门恢复的查询复杂性。

更新时间: 2026-10-06 05:06:43

领域: cs.CR

下载: http://arxiv.org/abs/2610.07771v1

When Can Stateless Recovery Defeat Byzantine Quorum Safety? A Tight Normal Form for Single-Step BFT

Byzantine quorum safety relies on correct replicas refusing to sign conflicting values. A replica that loses its protocol state during recovery but retains its identity and signing key may forget an earlier vote. We study certificates formed by matching signed votes from at least $q$ of $n$ replicas, assuming that each correct replica avoids conflicting votes between recoveries. If two conflicting certificates form, their overlap has size between $2q-n$ and $b+c$, where $b$ counts Byzantine replicas and $c$ counts correct identities that recovered during the execution considered. Our main result decomposes the slack $b+c-(2q-n)$ into four nonnegative counts: extra signers in the first certificate, extra signers in the second, identities in neither certificate, and Byzantine or recovering identities outside their overlap. Zero slack forces an exact signer partition. With $n=3f+1$ replicas, threshold $q=2f+1$, at most $f$ Byzantine replicas, and exactly one correct recovery event, any conflicting pair forces exactly $f$ Byzantine replicas, all in the overlap together with the recovered replica; each certificate has a disjoint side of $f$ correct replicas. A minimal protocol attains this form. We distinguish certificate formation from acceptance, give a sufficient check using configured fault and recovery caps, and explain why durable vote records written before signature release prevent the conflict.

Updated: 2026-10-06 04:55:15

标题: 什么时候无状态恢复可以战胜拜占庭法定安全?单步BFT的紧凑正常形式

摘要: 拜占庭法定安全依赖于正确副本拒绝签署冲突值。在恢复过程中丢失协议状态但保留身份和签名密钥的副本可能会忘记早期的投票。我们研究由至少 $q$ 个 $n$ 个副本的匹配签名选票组成的证书,假设每个正确的副本在恢复期间避免冲突投票。如果形成两个冲突的证书,它们的重叠大小在 $2q-n$ 和 $b+c$ 之间,其中 $b$ 计数拜占庭副本,$c$ 计数在考虑的执行期间恢复的正确身份。我们的主要结果将松弛度 $b+c-(2q-n)$ 分解为四个非负计数:第一个证书中的额外签署者、第二个证书中的额外签署者、两个证书中都没有的身份以及在它们的重叠之外的拜占庭或恢复身份。零松弛度会导致精确的签署者分区。对于 $n=3f+1$ 个副本,阈值 $q=2f+1$,最多 $f$ 个拜占庭副本,以及恰好一个正确的恢复事件,任何冲突的一对都会强制要求恰好 $f$ 个拜占庭副本,所有这些副本都在重叠部分与恢复的副本一起;每个证书都有一个由 $f$ 个正确副本组成的不相交部分。一个最小的协议达到了这种形式。我们区分了证书的形成和接受,使用配置的故障和恢复上限给出了一个充分的检查,并解释了为什么在签名发布之前编写的持久性投票记录可以防止冲突。

更新时间: 2026-10-06 04:55:15

领域: cs.DC,cs.CR,cs.DS,cs.NI

下载: http://arxiv.org/abs/2610.07759v1

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100\% at only a 5\% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.

Updated: 2026-10-06 04:17:18

标题: 植入模型触发器:多轮大型语言模型中的答案方后门攻击

摘要: 大型语言模型(LLMs)中的安全对齐仍然容易受到后门攻击的影响。现有的LLM后门几乎都是以输入为中心的:激活取决于用户输入中的显式触发模式,因此现代的保护措施是建立在清洁输入空间的基础上的。我们挑战这一假设,提出了一种新颖的多轮对话中的答案端后门。对手不是将触发器插入输入,而是使用一个良性的第一轮提示来自然地诱导模型生成一个特定的、看似无害的单词。一旦融入对话历史,这个自动生成的单词就成为了触发器。当后来的有害查询到达时,模型会检测到自己的触发器,并绕过其安全拒绝,而用户输入保持完全干净。在四个LLMs上,我们的攻击达到了接近完美的攻击成功率,在仅有5%的毒化率时接近100%,同时保留了一般效用和清洁输入的安全性,并且规避了主流的以输入为中心的防御措施。表示级别的分析显示,自动生成的触发器始终抑制模型的拒绝信号,暴露了当前LLM防御中的一个关键盲点。

更新时间: 2026-10-06 04:17:18

领域: cs.CR,cs.LG

下载: http://arxiv.org/abs/2610.07723v1

quantum-safe: Bridging the Post-Quantum Production Gap with a Hybrid-by-Default Python Cryptography Library

FIPS 203, 204 and 205 closed the algorithmic gap in post-quantum cryptography (PQC); the production gap -- hybrid combiners, conformance evidence, migration tooling, stateful signing, protocol helpers -- remains open. A methodology that scores it must not choose its own dimensions, so we anchor nine to external requirements (CNSA 2.0, CMVP, the TNO CADI survey, the IETF hybrid draft). We present quantum-safe, a hybrid-by-default Python library built against this rubric, and an audit of nine PQC libraries (September 2026). The exercise can lower our own score, and it does. Hybrid combiners exist in three of six verifiable libraries, not the one a March-2026 snapshot found. Discovery and inventory tooling is the sharper gap: none of seven audited libraries has it; quantum-safe adds a CycloneDX CBOM. Only Bouncy Castle holds a CMVP FIPS 140-3 certificate; quantum-safe reports 225/225 runnable ACVP known-answer cases, which is conformance evidence, not validation. Bouncy Castle also supports LMS and XMSS where quantum-safe has LMS only, and quantum-safe's defaults sit below the CNSA 2.0 parameter sets; we score both against ourselves. Every benchmark figure is computed from archived raw runs (bootstrap intervals over 15 runs). A full X25519 + ML-KEM-768 handshake takes a median of 231 microseconds under Docker/Linux (95% interval 218-250), 5.3x an X25519-only handshake and 0.46-2.3% of a typical TLS 1.3 budget. Throughput is flat from 100 to 5,000 threads but below the single-thread rate: threads do not speed up the hybrid path, although liboqs does release the GIL in longer calls (ML-DSA-65 signing scales 2.6x on four threads). Timing side channels are treated in a companion paper.

Updated: 2026-10-06 03:46:32

标题: 量子安全:通过默认的混合Python密码库填补后量子生产差距

摘要: FIPS 203、204和205在后量子密码学(PQC)中关闭了算法差距;生产差距——混合组合器、符合性证据、迁移工具、有状态签名、协议辅助工具——仍然存在。评分方法不能选择自己的维度,因此我们将九个锚定在外部要求(CNSA 2.0、CMVP、TNO CADI调查、IETF混合草案)上。我们提出了 quantum-safe,这是一个根据这一标准制定的默认混合 Python 库,并对九个 PQC 库进行了审计(2026年9月)。这个练习可以降低我们自己的得分,而事实确实如此。 在六个可验证库中有三个存在混合组合器,而不是2026年3月的快照发现的那一个。发现和清单工具是更尖锐的差距:七个经审计的库中没有一个有它;quantum-safe 添加了一个 CycloneDX CBOM。只有 Bouncy Castle 拥有 CMVP FIPS 140-3 证书;quantum-safe 报告了 225/225 可运行的 ACVP 已知答案案例,这是符合性证据,而不是验证。Bouncy Castle 还支持 LMS 和 XMSS,而 quantum-safe 只支持 LMS,并且 quantum-safe 的默认值低于 CNSA 2.0 参数集;我们对两者都进行评分。 每个基准数据都是从存档的原始运行中计算出来的(在 15 次运行中进行引导间隔)。在 Docker/Linux 下,完整的 X25519 + ML-KEM-768 握手需要中位数为 231 微秒(95% 间隔为 218-250),是 X25519 单独握手的 5.3 倍,占典型 TLS 1.3 预算的 0.46-2.3%。吞吐量从 100 到 5,000 个线程保持不变,但低于单线程速率:线程不会加快混合路径,尽管 liboqs 在较长的调用中释放了 GIL(ML-DSA-65 签名在四个线程上的比例为 2.6 倍)。时间侧信道在一篇伴随论文中进行了处理。

更新时间: 2026-10-06 03:46:32

领域: cs.CR,quant-ph

下载: http://arxiv.org/abs/2605.17061v2

PerSpectron: Detecting Invariant Footprints of Microarchitectural Attacks with Perceptron

Detecting microarchitectural attacks is critical given their proliferation in recent years. Many of these attacks exhibit intrinsic behaviors essential to the nature of their operation, such as creating contention or misspeculation. This study systematically investigates the microarchitectural footprints of hardware-based attacks and shows how they can be detected and classified using an efficient hardware predictor. We present a methodology to use correlated microarchitectural statistics to design a hardware-based neural predictor capable of detecting and classifying microarchitectural attacks before data is leaked. Once a potential attack is detected, it can be proactively mitigated by triggering appropriate countermeasures. Our hardware-based detector, PerSpectron, uses perceptron learning to identify and classify attacks. Perceptron-based prediction has been successfully used in branch prediction and other hardware-based applications. PerSpectron has minimal performance overhead. The statistics being monitored have similar overhead to already existing performance monitoring counters. Additionally, PerSpectron operates outside the processor's critical paths, offering security without added computation delay. Our system achieves a usable detection rate for detecting attacks such as SpectreV1, SpectreV2, SpectreRSB, Meltdown, breakingKSLR, Flush+Flush, Flush+Reload, Prime+Probe as well as cache-attack calibration programs. We also believe that the large number of diverse microarchitectural features offers both evasion resilience and interpretability---features not present in previous hardware security detectors. We detect these attacks early enough to avoid any data leakage, unlike previous work that triggers countermeasures only after data has been exposed.

Updated: 2026-10-06 03:26:19

标题: PerSpectron:使用感知器检测微体系结构攻击的不变足迹

摘要: 检测微架构攻击至关重要,因为近年来这些攻击越来越多。许多这些攻击表现出与其操作本质相关的内在行为,例如创建争用或误判。本研究系统地调查了基于硬件的攻击的微架构特征,并展示了如何利用高效的硬件预测器来检测和分类这些攻击。我们提出了一种方法,利用相关的微架构统计数据设计了一个硬件基础的神经预测器,能够在数据泄漏之前检测和分类微架构攻击。一旦检测到潜在攻击,可以通过触发适当的对策来主动缓解。我们的基于硬件的检测器PerSpectron使用感知器学习来识别和分类攻击。基于感知器的预测已成功应用于分支预测和其他基于硬件的应用中。PerSpectron的性能开销很小。被监控的统计数据与已存在的性能监控计数器具有类似的开销。此外,PerSpectron在处理器的关键路径之外运行,提供了安全性而不增加计算延迟。我们的系统实现了对攻击的可用检测率,例如SpectreV1、SpectreV2、SpectreRSB、Meltdown、breakingKSLR、Flush+Flush、Flush+Reload、Prime+Probe以及缓存攻击校准程序。我们还相信,大量多样的微架构特征提供了对抗逃避和可解释性的优势---这些特性在先前的硬件安全检测器中并不存在。我们能够在数据泄漏之前及时检测到这些攻击,而不像先前的工作在数据暴露后才触发对策。

更新时间: 2026-10-06 03:26:19

领域: cs.CR

下载: http://arxiv.org/abs/2610.07691v1

Adaptive Model Inversion Attacks Generalize a Privacy-Robustness Tradeoff

In this paper, we show that standard evaluations of high-resolution Model Inversion Attacks (MIAs) significantly underestimate training-data privacy leakage. State-of-the-art privacy defenses, standard training techniques such as MixUp and Adversarial Training, and undefended models all leak training images at rates 1.16 to 6.59 times higher on FaceScrub under simple adaptive changes to the attack, with the largest increases among defenses reporting the strongest privacy. We further show that measured leakage depends on the feature basis of the external classifier used to evaluate reconstructions: for the same reconstructed images, an adversarially trained Inception evaluator identifies the targeted identity at different rates than the standard Inception evaluator. Our results suggest that standard MIA evaluation can mistake optimization and measurement failures for privacy. These underestimated leakage rates also concealed a broader relationship between privacy and adversarial robustness. Once we adapt the attack and vary the evaluator, reconstruction leakage closely tracks adversarial robustness across recent defenses and standard training regimes, suggesting that robustness provides an attack-agnostic proxy for reconstruction vulnerability that applies far more broadly than previously theorized. This raises an open question: can a practical defense reduce training-data reconstruction without paying a corresponding cost in adversarial robustness?

Updated: 2026-10-06 03:11:19

标题: Adaptive Model Inversion Attacks Generalize a Privacy-Robustness Tradeoff 自适应模型逆向攻击概括了隐私-鲁棒性折衷。

摘要: 在这篇论文中,我们展示了对高分辨率模型反演攻击(MIAs)的标准评估显著低估了训练数据的隐私泄露。最先进的隐私防御、标准训练技术如MixUp和对抗训练,以及未经防御的模型在简单调整攻击时都以1.16至6.59倍的速率泄露FaceScrub的训练图像,其中在报告最强隐私的防御中泄露最严重。我们进一步展示了测量的泄露取决于用于评估重建的外部分类器的特征基础:对于相同的重建图像,一个经过对抗训练的Inception评估器识别目标身份的速率与标准Inception评估器不同。我们的结果表明,标准MIA评估可能会将优化和测量失败误认为是隐私问题。 这些被低估的泄露速率也隐藏了隐私与对抗鲁棒性之间更广泛的关系。一旦我们调整攻击并改变评估者,重建泄露会紧密跟踪最近的防御和标准训练制度中的对抗鲁棒性,这表明鲁棒性提供了一种与攻击无关的重建脆弱性代理,适用范围比以前理论化的更广泛。这提出了一个开放问题:一个实用的防御是否可以减少训练数据的重建而不会付出对抗鲁棒性的相应代价?

更新时间: 2026-10-06 03:11:19

领域: cs.LG,cs.CR

下载: http://arxiv.org/abs/2610.07677v1

Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.

Updated: 2026-10-06 02:52:51

标题: 规则的终结和法官的开始:在多智能体系统安全中衡量判断边界

摘要: 基于LLM的多Agent系统(MAS)利用工具,共享内存,并委托任务,通常会遇到敌对内容。目前对于MAS的防御通常是独立评估的,专注于一种攻击类型,这可能导致昂贵且难以审计的结果。本研究将防御措施组织为五个原则,并将它们实施为DEFER1(具有剩余判断的确定性优先执行),其中包括一个由28个检查组成的级联,它会阻止它可以阻止的部分,并将其余部分转交给四名评委组成的小组。在四个领域的独立测试中,攻击成功率从约30.0%下降到约3.0%,78%的被阻止攻击由确定性检查处理。在安全运营领域,只有四分之一的提案达到评委,说明规则为违反明确政策的攻击提供了安全性,而评委则管理那些只是误代意图的攻击。两个系统都存在弱点,例如风险评分批准门会错误地批准大部分攻击提案但很少批准合法的提案,突显了准确评估威胁的挑战。

更新时间: 2026-10-06 02:52:51

领域: cs.AI,cs.CL,cs.CR,cs.MA

下载: http://arxiv.org/abs/2610.07657v1

SkillPoison: Progressive Skill Poisoning via Successful Experiences

Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and the injected malicious behaviors often fail to accumulate as persistent skills. In this paper, we show that skill poisoning can arise even from verified successful experiences, without making any individual trajectory malicious. Based on this insight, we propose SkillPoison, a novel framework that progressively poisons skill via successful experiences. SkillPoison first constructs a set of successful experiences that reinforce a target behavior, and then removes the contextual conditions that constrain when the behavior applies. Rather than injecting malicious content, SkillPoison shapes how the skill extractor generalizes, allowing useful behavior to support task success while inducing harmful behavior when they are misapplied. Extensive experiments on three benchmarks show that SkillPoison achieves 95.71% attack success rates, while all injected experiences remain task-correct and pass verification and lexical inspection. Our code, data and implementation details are available for the community at https://github.com/DEEP-JLU/SkillPoison.

Updated: 2026-10-06 02:38:36

标题: 技能毒药:通过成功经验的渐进性技能毒化

摘要: 自我改进的LLM代理越来越将成功经验提炼成持久、可重复使用的技能。现有的技能攻击方法通过向个体经验或提取的技能中注入恶意触发器、行为或虚假事实,破坏了这一学习管道。然而,这种攻击很容易被检测到,注入的恶意行为通常无法积累为持久技能。在本文中,我们展示了技能污染甚至可以从经过验证的成功经验中产生,而无需使任何单个轨迹变得恶意。基于这一认识,我们提出了SkillPoison,一个新颖的框架,通过成功经验逐渐污染技能。SkillPoison首先构建一组强化目标行为的成功经验,然后移除约束行为适用的情境条件。与注入恶意内容不同,SkillPoison塑造了技能提取器的泛化方式,使有用的行为在支持任务成功的同时,在误用时引起有害行为。对三个基准测试的大量实验表明,SkillPoison实现了95.71%的攻击成功率,而所有注入的经验仍然保持任务正确,并通过验证和词法检查。我们的代码、数据和实现细节可供社区获取,网址为https://github.com/DEEP-JLU/SkillPoison。

更新时间: 2026-10-06 02:38:36

领域: cs.CR,cs.AI

下载: http://arxiv.org/abs/2610.07645v1

HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?

Coding agent harnesses mediate tool use and authorize actions, yet their security mechanisms and runtime effects remain incompletely characterized. We present HarnessSecurity, the first systematic empirical study and benchmark of open- and closed-source coding agent harnesses. First, we derive a ten-mechanism taxonomy and then assess 400 harness-mechanism cells using independent ratings by researchers and large language model (LLM) judges. We find that about half of confirmed mechanism implementations are opt-in, while closed-source harnesses exhibit substantial evidence gaps. Second, we introduce HarnessSecurity-Bench, a benchmark of 23 tasks across five attack surfaces without sacrificing legitimate task requirements. Using separate deterministic oracles to measure task utility and attack effects with security setting comparisons, we evaluate nine mechanisms across six leading harnesses: Claude Code, Codex CLI, Gemini CLI, gptme, Qwen Code, and GitHub Copilot. Under a controlled LLM baseline GLM-5.2, we conduct 2,500 trials, recording 81,155 tool calls and over 2.2 billion tokens. Enabling auto-approve increases utility and raises attack success from 29.2% to 95.6%. Network isolation and read-only mode reduce attack effects with substantial utility losses, while command allowlisting and command denylisting reduce attack effects with a small utility loss and a utility gain, respectively. Task-level cases show that restrictions on a shared capability can obstruct both legitimate and malicious operations, and that allowed tools or commands can leave unauthorized operations reachable through alternative execution paths. Harness providers should make security settings verifiable, test alternative execution paths to protected operations, and assess attack effects alongside task utility and execution costs.

Updated: 2026-10-06 02:34:53

标题: HarnessSecurity-Bench:安全机制是否真正保护编码代理Harnesses?

摘要: 编码代理工具在控制工具使用和授权操作方面发挥作用,然而它们的安全机制和运行效果仍未完全表征。我们提出了HarnessSecurity,这是对开源和闭源编码代理工具进行系统的实证研究和基准测试的第一个。首先,我们提出了一个包括十种机制的分类法,然后通过独立研究人员和大型语言模型(LLM)评审对400个代理-机制单元进行评估。我们发现大约一半的确认机制实现是可选择的,而闭源编码代理工具存在着实质性的证据缺失。其次,我们引入了HarnessSecurity-Bench,这是一个包括23个任务在五个攻击面上的基准测试,而不会牺牲合法任务要求。通过使用独立的确定性预言来衡量任务效用和攻击效果,并进行安全设置比较,我们评估了六个领先的编码代理工具中的九种机制:Claude Code、Codex CLI、Gemini CLI、gptme、Qwen Code和GitHub Copilot。在受控的LLM基线GLM-5.2下,我们进行了2500次试验,记录了81,155次工具调用和超过22亿个标记。启用自动批准可以提高效用,并将攻击成功率从29.2%提高到95.6%。网络隔离和只读模式可以减少攻击效果,并带来实质性的效用损失,而命令白名单和命令黑名单可以减少攻击效果,其中前者会带来一定的效用损失,而后者则会带来效用增益。任务级案例表明,对共享功能的限制可能会阻碍合法和恶意操作,而允许的工具或命令可能会通过替代执行路径使未经授权的操作可达。编码代理工具提供商应该使安全设置可验证,测试受保护操作的替代执行路径,并在任务效用和执行成本之外评估攻击效果。

更新时间: 2026-10-06 02:34:53

领域: cs.CR,cs.SE

下载: http://arxiv.org/abs/2610.07639v1

CISB-Bench: An Auditable Source--IR Dataset of Compiler-Introduced Security Bugs

Compiler-introduced security bugs (CISBs) arise when an optimization, lowering, or instrumentation decision changes a security-relevant property of the generated program. They are difficult to study because their evidence is distributed across issue reports, reduced tests, historical configurations, and compiler artifacts; a security-related report also does not imply that every associated reduction establishes a security-bearing compiler failure. We present CISB-Bench, an auditable dataset of 429 exact C-program rows mined from GCC and LLVM. Each row contains its C reduction, a standardized LLVM IR analysis bundle at -O0 through -O3, public provenance, a final binary label, and a primary mechanism or boundary annotation. Two reviewers independently labeled the fixed corpus, agreeing on 369 rows (86.0%, Cohen's kappa=0.662); the 60 disagreements were adjudicated. The final dataset comprises 280 CISBs and 149 hard non-CISB cases. The prediction task is to recover this reviewed exact-row label from the supplied artifacts; it is not a claim that standardized IR alone reproduces every historical compiler failure. We characterize the security mechanisms and evidence boundaries represented by the corpus, and demonstrate how its paired artifacts support source-only, IR-aware, and joint analyses. CISB-Bench provides a reusable, inspectable target for compiler-security mining and detection research.

Updated: 2026-10-06 02:26:28

标题: CISB-Bench:一个可审计的编译器引入的安全漏洞源数据集

摘要: 编译器引入的安全漏洞(CISBs)是由于对生成的程序的优化、降低或仪器化决策改变了与安全相关的属性而导致的。它们很难研究,因为它们的证据分布在问题报告、缩减测试、历史配置和编译器工件中;安全相关报告也不意味着每个相关缩减都会建立一个具有安全性的编译器故障。我们提出了CISB-Bench,这是一个包含从GCC和LLVM挖掘出的429个精确C程序行的可审计数据集。每行包含其C缩减、标准化的LLVM IR分析捆绑在-O0到-O3,公开来源、最终二进制标签和主要机制或边界注释。两位评审独立标记了修正的语料库,对369行(86.0%,Cohen's kappa=0.662)达成一致;60个分歧得到裁决。最终数据集包括280个CISBs和149个难以解决的非CISB案例。预测任务是从提供的工件中恢复这个经过审查的精确行标签;这并不意味着标准化IR单独重现每个历史编译器故障。我们对语料库代表的安全机制和证据边界进行了表征,并展示了它的配对工件如何支持仅源代码、IR感知和联合分析。CISB-Bench为编译器安全挖掘和检测研究提供了可重复使用、可检查的目标。

更新时间: 2026-10-06 02:26:28

领域: cs.CR,cs.SE

下载: http://arxiv.org/abs/2610.07635v1

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.

Updated: 2026-10-06 00:42:45

标题: CheckerBench:长期视角的代理能否合成静态分析检查器?

摘要: 静态分析检查器合成需要代理人解释缺陷规范,检查存储库,实现特定于分析器的逻辑,并通过重复的编译和分析反馈来完善检查器。现有的编码代理基准重点放在诸如修补程序生成或漏洞检测等任务上,很少评估代理人是否能够从头到尾在存储库中开发一个可工作的检查器。我们介绍了CheckerBench,这是一个可执行的基准,包括来自167个存储库的297个CVE、85个CWE和五种语言生态系统的300个任务。每个任务包括易受攻击和已修复的版本、固定的分析环境和检查器支架。我们进一步介绍了CheckerLab,这是一个通用的评估框架,可以独立地重建已提交的检查器,并测量易受攻击和已修复的诊断对比、补丁定位、误报和工具使用。在21个模型绑定配置和每个配置的三次独立重复中,Pass@1的平均值为32.30%,最佳值达到45.33%。这些结果表明,对于当前的编码代理来说,可靠、可重复使用的检查器开发仍然具有挑战性。

更新时间: 2026-10-06 00:42:45

领域: cs.SE,cs.AI,cs.CR

下载: http://arxiv.org/abs/2610.07557v1

By Xinhai (Sean) Zou.