<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>3DArxiv</title>
    <link>https://wastoon.github.io/3DArxiv/</link>
    <description>3D Vision &amp; Robotics ArXiv daily digest</description>
    <language>zh-cn</language>
    <lastBuildDate>Sun, 27 Sep 2026 07:18:06 +0000</lastBuildDate>
    <atom:link href="https://wastoon.github.io/3DArxiv/rss.xml" rel="self" type="application/rss+xml"/>
    <ttl>1440</ttl>
    
  <item>
    <title>LLM Agents Can Easily Tamper With Their Own Traces</title>
    <link>http://arxiv.org/abs/2609.30266v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30266v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:59:54 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Jeremy Qin, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, Maksym Andriushchenko</p><p>Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering behavior emerges naturally in frontier models, when agents try to improve their rewards. We advise practitioners to ensure trace logging happens through an independent interception mechanism outside of the agent&#x27;s control, preserving trace integrity even in cases of full host compromise. Overall, our findings identify a concrete failure of trace integrity in agent infrastructure which can be used to conceal misaligned behaviors like scheming or sabotage.</p>]]></description>
  </item>
  <item>
    <title>AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control</title>
    <link>http://arxiv.org/abs/2609.30264v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30264v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:59:41 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo, Yang Gao</p><p>Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC. AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized recovery objective motivated by conditional mutual information. Both objectives encourage planning transitions to preserve action information; their auxiliary heads are discarded at test time, leaving MPC unchanged. On OGBench-Cube, AD-WM improves hard-start success from 3.7% to 52.0% over a matched LeWM baseline and improves mean success over the reproduced baseline in four of five simulation environments. Planning diagnostics show that factual prediction error and whole-bank action ranking do not follow the closed-loop success ordering, whereas CEM-aligned elite regret tracks success more closely. With a frozen V-JEPA 2 encoder and matched DROID post-training, AD-WM also improves zero-shot transfer to our Franka setup, increasing basic pick-and-place success from 42.2% to 71.1% without lab-specific adaptation. These results suggest that world models for planning should preserve action-dependent differences needed for counterfactual selection, rather than optimize factual prediction accuracy alone. More videos and code are available at https://ad-wm.github.io/.</p><p><em>Comment: 9 pages, 5 figures, 4 tables. Project page: https://ad-wm.github.io/</em></p>]]></description>
  </item>
  <item>
    <title>Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning</title>
    <link>http://arxiv.org/abs/2609.30258v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30258v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:59:18 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Sudip Bhujel, Shanghao Shi, Ruiquan Huang, Ning Zhang, Yang Xiao</p><p>Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce Temporal Reconstruction Attack on Consecutive Encodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients. The attack exploits two structural signals ignored by prior single-frame methods: (i) cross-time correlation between successive embodied gradients, which we formalize via a conditional mutual-information bound, and (ii) closed-form action recovery from policy-head gradient structure, which we prove exact when standard entropy regularization is sufficiently small. On held-out embodied scenes, TRACE reaches $18.8$ dB PSNR with near-perfect action recovery at $3$-$4.5$ ms per reconstructed frame, dominating the learning-based baseline across all reconstruction metrics and exceeding optimization attacks while running orders of magnitude faster. Further evaluation demonstrates TRACE&#x27;s broader applicability across recurrent, residual, and compact transformer victim architectures, multi-modal inputs, and larger discrete action spaces. Defense experiments suggest that protecting temporal gradient streams may require sequence-aware privacy mechanisms.</p><p><em>Comment: Accepted at NeurIPS 2026</em></p>]]></description>
  </item>
  <item>
    <title>Agentic Detection of Online Conspiracies</title>
    <link>http://arxiv.org/abs/2609.30250v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30250v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:58:43 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Lior Biton, Oren Tsur</p><p>Conspiratorial discourse on social media is not always expressed through explicit claims or stable lexical markers. The same surface content may express endorsement, legitimate concerns, criticism, satire, or mockery. The main challenge is therefore not only recognizing conspiracy-related claims, but inferring the speaker&#x27;s intent -- the utterance&#x27;s illocutionary force. We argue that this can be achieved through the use of relevant social contexts and propose an agentic framework, equipped with a set of tools supporting social queries.
  We demonstrate the benefits of our approach on a unique dataset of Hebrew tweets, covering 80\%--90\% of the public Hebrew tweets published over a four-year span (late 2018-- early 2023), encompassing several election cycles as well as the COVID pandemic years and related vaccination campaigns. This extensive coverage can be used in recovering different social contexts. Evaluating our framework on a manually-annotated adversarial dataset, we find that context-aware workflows consistently outperform text-only classification and that the agentic framework performs significantly better than other frameworks and settings, including a non-agentic model exposed to the same contexts available to the agent. We further provide an analysis of the results, the errors and efficiency (token economy) tradeoffs.
  These findings support viewing the task of conspiracy detection as a socially embedded interpretation task, in which effective classification depends not only on access to contexts, but also on adaptive reasoning in which the agent uses tools on a per-case basis, asking only for evidence relevant to its current reasoning step.</p>]]></description>
  </item>
  <item>
    <title>RAPID: Robot Agentic Programming from Demonstrations</title>
    <link>http://arxiv.org/abs/2609.30249v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30249v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:58:21 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yuyao Liu, Jiayuan Mao, David Hsu, Leslie Pack Kaelbling, Tomás Lozano-Pérez</p><p>Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration. The iterative agentic loop of code refinement requires several key ingredients: (i) a testable task specification, (ii) action primitives for robot execution, and (iii) an interactive environment for program execution and verification. RAPID infers all three from the demonstration automatically. To make the resulting program reusable beyond the demonstration setting, RAPID uses an object-centric relational program representation that focuses on the underlying structure of the demonstrated strategy rather than the specific motion per se: it expresses the action primitives as trajectory-optimization programs that realize object-level motion effects, while composing them through relational constraints that capture scene-specific geometry at run time. We evaluated RAPID in simulation on eight challenging contact-rich nonprehensile manipulation tasks as well as general prehensile manipulation tasks in the LIBERO-Pro benchmark. We also successfully deployed it on a real Franka arm and evaluated on all eight nonprehensile tasks. In all experiments, RAPID demonstrated strong performance, with generalization over object pose, shape, material, and environment. Website: https://yuyaoliu.me/projects/rapid.</p>]]></description>
  </item>
  <item>
    <title>Rolling-WAM: World Action Models with Rolling Imagination</title>
    <link>http://arxiv.org/abs/2609.30247v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30247v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:58:03 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang</p><p>World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.</p><p><em>Comment: 10 pages, 7 figures, 5 tables. Under review. Project page: https://rolling-wam.github.io/</em></p>]]></description>
  </item>
  <item>
    <title>Towards Practical Compression of 3D Gaussian Splatting</title>
    <link>http://arxiv.org/abs/2609.30245v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30245v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:57:56 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Pengpeng Yu, Yueru Chen, Fei Song, Tai Qin, Qi Zhang, Jing Wang, Yulan Guo</p><p>3D Gaussian Splatting (3DGS) enables high-quality novel-view synthesis but requires substantial storage. Existing compression methods often rely on spatial context modeling over irregular 3D representations, increasing the complexity of training and coding. Meanwhile, floating-point context inference can introduce numerical inconsistencies across platforms, causing entropy-decoding failures. To address these practical challenges, we propose COSA-GS, which constructs context without spatial aggregation through anchor-wise causal factorization. Specifically, we use geometry context derived from each anchor&#x27;s coordinates to model a compact learnable anchor latent. The anchor latent is then fused with the geometry context to form an anchor context for attribute coding. The resulting context model features a simple architecture composed solely of linear transformations and activations. We train COSA-GS using rate--distortion optimization with adaptive Gaussian pruning. Further, we develop quantization-aware training and integer inference for the context model to achieve bit-exact consistency of entropy-decoded symbols across platforms. Experiments demonstrate that COSA-GS achieves state-of-the-art compression performance while retaining fast and consistent cross-platform decoding, providing a simple yet effective framework for practical 3DGS compression. Code is available at https://github.com/pengpeng-yu/COSA-GS.</p>]]></description>
  </item>
  <item>
    <title>JevOut: Natural Context Can Flip Decision Models</title>
    <link>http://arxiv.org/abs/2609.30243v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30243v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:57:07 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Zixiang Xu</p><p>Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model&#x27;s option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases, Jev assigns at least 0.7 probability to the fixed wrong option. Across seven datasets, three additional decision systems show targeted flip rates of 64.9%-73.2% on decisions they initially answer correctly. Taken together, these results expose a pronounced fragility in current decision models: short, ordinary-looking context can shift a correct choice to a high-confidence wrong one. Because these models turn language directly into downstream choices, this sensitivity raises concerns about treating their probability outputs as reliable decision interfaces.</p><p><em>Comment: 32 pages, 5 figures, 23 tables. Homepage: https://xzx34.github.io/jevout/ ; Code: https://github.com/xzx34/JevOut</em></p>]]></description>
  </item>
  <item>
    <title>SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data</title>
    <link>http://arxiv.org/abs/2609.30238v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30238v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:55:31 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Wenhao Li, Zhibin Wu, Chong Xiao, Qiangchang Wang</p><p>Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information by reconstructing modality features or designing complicated fusion mechanisms. However, these methods still suffer from spurious generation and noisy guidance due to the lack of high-level semantic grounding in partially observed multimodal evidence. To address these issues, we propose SemMSA, a latent semantic-aided framework that constructs rich sentiment-relevant semantics with LLMs, fully integrating with all modalities via anchor-free spectral alignment. It mainly consists of Cross-modal Semantic Refinement (CSR) and Cross-modal Spectral Alignment (CSA). Specifically, CSR first adaptively extracts visual and acoustic representations by corresponding adapters to form a unified multimodal prefix with language in the frozen LLM embedding space. It then iteratively produces continuous discriminative semantic states through a token-efficient latent refinement process without decoding explicit text. Next, CSA simultaneously aligns the refined semantics with all modalities by enhancing the dominant spectral component of their kernel Gram matrix. This captures global nonlinear dependencies among all representations without relying on a predefined anchor modality. In addition, an instance-level spectral separation constraint preserves cross-sample discriminability and mitigates representation collapse. Extensive experiments on SIMS, MOSI, and MOSEI benchmarks demonstrate that SemMSA achieves state-of-the-art performance.</p><p><em>Comment: Accepted by NeurIPS 2026</em></p>]]></description>
  </item>
  <item>
    <title>OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction</title>
    <link>http://arxiv.org/abs/2609.30234v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30234v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:53:55 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ding-Jiun Huang, Yuanhao Wang, Cheng Zhang, Hugo Bertiche, Alexandru-Eugen Ichim, Thabo Beeler, Fernando De la Torre</p><p>Automated generation of production-ready 3D garment assets from a single image is a central challenge in digital content creation. While recent generative models have significantly advanced 3D geometry reconstruction, synthesizing high-quality textures remains a bottleneck. Existing methods often bake environmental illumination and shadows directly into the texture map, or they fail to maintain global structural coherence, making the resulting assets unusable for physical simulation and relighting. In this work, we introduce OmniFabric, a novel approach that synthesizes globally coherent texture maps directly within the 2D sewing pattern space. Given a single reference image, our pipeline utilizes an estimated 3D mesh and generative priors of powerful Vision-Language Models (VLM) to establish a complete but coarse texture initialization across the unwrapped sewing patterns. We then leverage a specialized diffusion transformer, trained via an automated synthetic data engine and conditioned on 3D positional features, to refine this initialization directly in the canonical UV domain. This effectively removes distortion and baked-in artifacts to extract a clean and normalized texture map that preserves the original garment design. Extensive experiments demonstrate that OmniFabric significantly outperforms state-of-the-art baselines, yielding photorealistic 3D garments with high-quality textures.</p><p><em>Comment: Accepted to SIGGRAPH Asia 2026. Project Page: https://humansensinglab.github.io/OmniFabric/</em></p>]]></description>
  </item>
  <item>
    <title>Coding Agents for Generalized Task and Motion Planning Problems</title>
    <link>http://arxiv.org/abs/2609.30233v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30233v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:53:35 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, Tom Silver</p><p>Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by synthesizing programs that generalize across instances. Given a task description and simulator access, each agent chooses how to interact with the environment while developing a program within a fixed synthesis budget. The program is then frozen and evaluated on unseen instances. We evaluate Claude Code (Opus 5) and Codex (GPT-5.6 Sol and GPT-6 Astra) on 28 simulated environments from KinDER and PDDLStream, with object counts beyond those evaluated in the original benchmark. Across all program synthesis methods, we evaluate 980 generated programs on 100 held-out instances each, 98,000 evaluation episodes in total. Overall, we find that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success (56% to 95% versus 47% for the planners, on the 16 environments where a planner is available). As object counts grow, the agents&#x27; programs maintain higher success than the planner, using an order of magnitude less computation per instance on average. Logs show agents using interaction to calibrate physical models, test edge cases, and refine strategies. We release all code, including the full prompts given to the agents. These findings suggest that coding agents are a strong baseline for generalized TAMP.</p><p><em>Comment: 9 pages, 4 figures, 3 tables</em></p>]]></description>
  </item>
  <item>
    <title>To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech</title>
    <link>http://arxiv.org/abs/2609.30227v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30227v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:50:40 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Debajyoti Mazumder,  Mamta, Abhirama Subramanyam Penamakuri</p><p>Online misinformation increasingly appears in spoken formats such as news clips, podcasts, interviews, political speeches, and social media videos, creating a need for fact-checking systems that can verify claims directly from speech. We introduce VeriSpeak, a probe benchmark for studying speech-based fact verification in Large Audio Language Models (LALMs). VeriSpeak contains 3,879 spoken claims spanning temporal, geographical, and relational facts, with balanced true and false labels. The benchmark is designed to examine whether factual verification ability transfers from text to speech, and whether retrieval-augmented LALMs can use textual evidence to correctly support or refute spoken claims. Our experiments reveal a consistent text-speech modality gap: LALMs that verify written claims reliably often fail on the same claims when spoken. Moreover, retrieval alone provides limited gains because models frequently conflate retrieved evidence with the spoken claim. In contrast, retrieval combined with explicit reasoning improves claim-evidence comparison, with a thinking-tuned LALM reaching 86.1% accuracy. VeriSpeak highlights that effective speech misinformation detection requires not only speech understanding, but also grounded reasoning over retrieved evidence. The dataset is publicly available via Hugging Face at https://huggingface.co/datasets/abhiram4572/VeriSpeak.</p><p><em>Comment: Accepted to EMNLP (Main) 2026</em></p>]]></description>
  </item>
  <item>
    <title>PoEM: Predicting RL Outcomes from Existing Policies</title>
    <link>http://arxiv.org/abs/2609.30226v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30226v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:50:25 +0000</pubDate>
    <category>Digital Human</category>
    <description><![CDATA[<p><strong>Authors:</strong> Kimia Hamidieh, Giannis Daras, Antonio Torralba</p><p>Foundation models are post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction following. This post-training process is computationally intensive, sometimes unstable, and has to be run from scratch every time the reward model changes or when we want to combine multiple rewards. We hence ask: given a new reward function, is it possible to predict the RL outcomes without actually running RL on it? We answer this in the affirmative by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards. First, we show that if the new reward function can be written as a linear combination of existing ones, then the new policy in log-space can be written as a linear combination of the existing log-policies. Surprisingly, even in cases where the rewards are not linearly connected, we observe that often log-policies from RL training span an approximately low-rank subspace across rewards. To our benefit, the weighting coefficients for this combination can be estimated using only the reward or basis policy outputs on the samples. We turn these observations into an algorithm that takes post-trained models and a new reward function, and approximates the target RL policy without actually running any additional RL training. We experimentally validate our approach across synthetic and real rewards, spanning both text and image modalities.</p>]]></description>
  </item>
  <item>
    <title>BiCC: Bidirectional Connected-Component Loss for Instance-Aware Segmentation</title>
    <link>http://arxiv.org/abs/2609.30223v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30223v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:48:31 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Luc Bouteille, Frederic Jonske, Jens Kleesiek, Alexander Jaus</p><p>Common segmentation losses aggregate errors voxel-wise, so lesions influence the objective in proportion to their volume, giving small but clinically critical lesions disproportionately little weight. Instance-aware losses aim to address this mismatch by assigning each lesion its own term. However, blob loss and CC-DiceCE derive their regions solely from annotations, so false-positive components receive no instance-level term. This matters in computer-assisted review, where each false-positive component may require separate inspection, making precision and false-positive burden important alongside recall. We introduce the bidirectional connected-component loss (BiCC), which pairs annotation- and prediction-derived partitions to score predicted components on their own scale. By deriving instances from the predictions, this branch directly penalizes false-positive components regardless of their size. The balance parameter $α$ allows control over the lesion-wise precision-recall trade-off. Across five datasets with five-fold cross-validation using nnU-Net, BiCC outperforms CC-DiceCE in lesion-wise F1 on four datasets and blob loss on all five. It significantly improves over DiceCE on three datasets and matches it on two; CC-DiceCE instead loses up to 0.363 precision by favoring recall. Code is available at https://github.com/TIO-IKIM/BiCC-Loss.</p><p><em>Comment: 2 figures, 3 tables. Code: https://github.com/TIO-IKIM/BiCC-Loss</em></p>]]></description>
  </item>
  <item>
    <title>TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations</title>
    <link>http://arxiv.org/abs/2609.30222v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30222v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:48:20 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt, Jakob Engel, Katerina Fragkiadaki, Adam W. Harley</p><p>Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with unique physical scene geometry instead. Our approach introduces three key innovations. First, we employ a voxelization-based de-duplication mechanism at sliding-window boundaries to merge co-located tracks, preventing repeated observations of the same surface from redundantly accumulating. Second, we decompose tracking into an endpoint refiner that predicts each point&#x27;s destination and static-versus-dynamic classification, followed by a lightweight trajectory refiner that decodes dense trajectories exclusively for dynamic points. Third, we propose 3D WAFT, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud. To the best of our knowledge, TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory. On TAPVid-3D, TrackEverything outperforms all open-source all-frame dense 3D trackers by more than 20% APD on short clips, while remaining competitive with state-of-the-art sparse trackers on long sequences, despite tracking far more points.</p>]]></description>
  </item>
  <item>
    <title>WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation</title>
    <link>http://arxiv.org/abs/2609.30221v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30221v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:48:19 +0000</pubDate>
    <category>Digital Human</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong</p><p>Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0&#x27;s video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.</p>]]></description>
  </item>
  <item>
    <title>Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority</title>
    <link>http://arxiv.org/abs/2609.30219v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30219v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:46:53 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Mehmet Iscan</p><p>An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when both facts can be derived by an external gate under a sealed grammar. One canonical answer is requested from a gold-standard user when eligible. The protocol was evaluated once under a criterion fixed before benchmark construction, on 144 tasks written by isolated agent contexts without access to the gate, grammar, or experimental plan. Three contributions are established. First, candidate generation and release decisions were measured separately. Fabricated ready plans were committed on 21 of 22 routed unanswerable tasks, and all were rejected. The same 83 releases were reproduced without model calls. Second, no false release was observed among 83 releases. A one-sided 95% Clopper-Pearson upper bound of 0.0354 was obtained as a diagnostic under an independent-and-identically-distributed assumption, below the sealed 5% threshold. However, one false release was subsequently recorded among 146 releases outside the benchmark at seed 0. Third, protection against incorrect user answers was characterized. Both facts were bound from the original text on 13 of 96 answerable tasks. Incorrect answers were released in 169 of 431 pairings on the remaining tasks, including failures involving coordinate exclusion. A deployable questioning policy was not tested because eligibility was determined from the answer key. Gate sensitivity and real user behavior were not measured.</p><p><em>Comment: 42 pages, 6 figures, 15 tables</em></p>]]></description>
  </item>
  <item>
    <title>Minimally Invasive Steering of Language Models</title>
    <link>http://arxiv.org/abs/2609.30218v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30218v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:46:46 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Taha Entesari, Jingyu Zhang, Daniel Khashabi, Mahyar Fazlyab</p><p>Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient computed through matrix--vector products with the frozen language-model head. We derive an exact decomposition of the sequence-level KL gradient into an analytic Fisher term and a suffix score-function term. For a fixed generation horizon, we show that the suffix term is second order in the steering magnitude and that three Fisher surrogates agree with the full KL gradient to first order. MISVO uses the frozen-reference surrogate to optimize position-specific interventions without updating model parameters. Across preference and code-generation tasks on models with approximately 1B--14B parameters, MISVO achieves the highest mean reward in six of seven model--task settings, with diversity and coherence scores close to those of Best-of-N.</p>]]></description>
  </item>
  <item>
    <title>Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure</title>
    <link>http://arxiv.org/abs/2609.30217v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30217v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:46:27 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Anselm Paulus, Ameya Prabhu, Maksym Andriushchenko</p><p>A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce EvasionBench, a benchmark of 50 diverse task-policy pairs in which completing the task requires an operation prohibited by a runtime monitor. Agents know that their tool calls are monitored and are prompted to continue working when they pause. Across our evaluations, best-of-3 evasion attempt rates reach up to 98% and success rates up to 88%, with substantial variance across models. Claude Fable 5.1 succeeds less often, but frequently makes creative attempts to circumvent the monitor. Evasion generally increases with test-time compute, with higher evasion rates at greater reasoning effort and token use. Traces show that agents encode prohibited commands, decompose operations across tool calls, and retry until relevant context leaves the monitor&#x27;s history. We also find that GPT-6 Astra&#x27;s low evasion rate comes with overrefusal, as it frequently abandons otherwise solvable tasks under a denial-of-service prompt injection. Our findings show that ordinary task pressure can lead to adaptive attempts to evade runtime monitors without an explicit adversarial objective. Effective oversight must therefore remain robust against repeated attempts, as the persistence that helps agents solve difficult tasks can also drive them to circumvent their guardrails.</p>]]></description>
  </item>
  <item>
    <title>A Nearly Quadratic Lower Bound for Linear Optimization over Convex Bodies in the Membership Oracle Model</title>
    <link>http://arxiv.org/abs/2609.30215v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30215v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:45:58 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Santosh S. Vempala</p><p>We prove nearly quadratic lower bounds for randomized algorithms for linear optimization and uniform sampling over convex bodies in the membership oracle model. For linear optimization, this matches the known nearly quadratic upper bound up to a polylog factor in the dimension. For uniform sampling, this improves on the previous linear lower bound. Our construction also implies the same lower bound for volume estimation.</p>]]></description>
  </item>
  <item>
    <title>Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage</title>
    <link>http://arxiv.org/abs/2609.30214v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30214v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:45:42 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang</p><p>We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C$^{3}$-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera&#x27;s object state and staying ahead of persistence, so the recipe transfers beyond simulation.</p><p><em>Comment: Submitted to the IEEE for possible publication. 12 pages, 14 figures</em></p>]]></description>
  </item>
  <item>
    <title>ReVAMP: Vector-Accelerated Motion Planning for Kinematically-Constrained Systems via Reparameterization</title>
    <link>http://arxiv.org/abs/2609.30213v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30213v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:45:26 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Shrutheesh R. Iyer, Thomas Cohn, Zachary Kingston</p><p>Robots often must satisfy one or more constraints during motion planning for real-world tasks. When such constraints reduce the valid configuration space to a measure-zero subset, sampling based planning algorithms require modifications to draw feasible samples. For many common end-effector constraints, parameterizations built on inverse kinematics (IK) provide an alternate formulation where the constraints are satisfied by construction, allowing directly sampling the feasible set. Despite their elegant approach, parameterized planners have remained slower than vector-accelerated implementations of projection-based approaches, leaving their performance ceiling an open question. We explore a new axis of vectorization built upon reparameterizing the planning space through analytic IK. This approach addresses existing inefficiencies in vectorized projection-based planners and exposes new opportunities for parallelism within the planner. We show that the planner can synthesize plans in microseconds to milliseconds for high dimensional systems (up to 20 dimensions), with complex constraints, up to 10x faster than the current state-of-the-art. Furthermore, we demonstrate how such planning speeds open up avenues for restructuring sequential manipulation pipelines.</p>]]></description>
  </item>
  <item>
    <title>Anchored Extra-Proximal Methods: Optimal Higher-Order Methods for Monotone Inclusion Problems</title>
    <link>http://arxiv.org/abs/2609.30212v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30212v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:43:45 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ruichen Jiang, TaeHo Yoon</p><p>We study the deterministic oracle complexity of finding approximate solutions to composite monotone inclusion problems, formed by the sum of a smooth single-valued monotone operator and a maximally monotone set-valued operator, under the tangent-residual criterion. We introduce the Anchored Extra-Proximal (AEP) framework, which combines an anchored extrapolation step with an inexact anchored proximal update satisfying a relative-error condition. The framework recovers the composite Fast Extragradient method in the first-order setting and yields natural second- and higher-order extensions by replacing the operator in the implicit update with its Taylor approximation at the extrapolated point. For every $p\geq 2$, assuming that the $(p-1)$th derivative of the single-valued operator is Lipschitz continuous, we combine this construction with a bisection line search to obtain a $p$th-order method that finds a point with tangent residual at most $\varepsilon$ in $\widetilde{O}(\varepsilon^{-2/(3p-1)})$ oracle calls. This improves all prior upper bounds for $p$th-order methods: in particular, it improves the previous best-known $\widetilde{O}(\varepsilon^{-1/p})$ tangent-residual complexity as well as the classical $O(\varepsilon^{-2/(p+1)})$ bound of higher-order hybrid proximal extragradient methods under the weaker duality-gap criterion. We complement this result with a worst-case lower bound of $Ω(\varepsilon^{-2/(3p-1)})$ for every deterministic algorithm in the $p$th-order oracle model, without restricting the algorithm to tensor steps or any other prescribed update structure. Thus, the proposed method attains the optimal dependence on $\varepsilon$, up to logarithmic factors, for all $p\geq2$.</p><p><em>Comment: 51 pages</em></p>]]></description>
  </item>
  <item>
    <title>The Alignment Illusion in Multimodal Large Language Models</title>
    <link>http://arxiv.org/abs/2609.30210v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30210v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:42:29 +0000</pubDate>
    <category>Gaussian Splatting</category>
    <description><![CDATA[<p><strong>Authors:</strong> Hong-Han Wang, Yuntao Wang, Hu Ding</p><p>Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families spanning 0.5B to 72B parameters, replacing projector-output visual tokens with Gaussian noise sharply reduces task accuracy, yet four standard scalar measures (CKA, SVCCA, MIR, and the leading principal-angle cosine) fail to consistently separate the corrupted stream from the original. We call this failure the alignment illusion and trace it to the shared language-model pathway: anisotropic MLP down-projections pull visual and text tokens toward common output directions, producing weight-induced alignment. Because this component is essentially one-dimensional, we introduce the principal-angle gap (PA gap), defined as the difference between the top two principal-angle cosines, which separates weight-induced similarity from multi-directional visual structure. Under graded visual corruption, the PA gap tracks task accuracy more consistently than the scalar scores we consider; under a structured but irrelevant image, it further exposes regimes in which internal geometry and task accuracy come apart. Internal visual-text alignment in MLLMs is therefore best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.</p><p><em>Comment: Accepted to NeurIPS 2026</em></p>]]></description>
  </item>
  <item>
    <title>A Living Benchmark for Information Retrieval from Electronic Health Records</title>
    <link>http://arxiv.org/abs/2609.30205v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30205v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:41:16 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer</p><p>Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.</p>]]></description>
  </item>
  <item>
    <title>ExplorationBench: Measuring AI Systems&#x27; Exploration in Verifiable Alien Worlds</title>
    <link>http://arxiv.org/abs/2609.30199v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30199v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:37:14 +0000</pubDate>
    <category>Spatial AI</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan</p><p>Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark contains two sandboxes, AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbox provides a flawed manual, task-specific environmental feedback, and a dedicated tool-call schema. Systems use these resources to explore the sandbox, then solve held-out tasks. We evaluate 10 AI systems and find that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can stall or reverse earlier gains. ExplorationBench represents a step towards AI systems that can acquire and apply genuinely new knowledge through exploration in unknown environments.</p>]]></description>
  </item>
  <item>
    <title>Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers</title>
    <link>http://arxiv.org/abs/2609.30198v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30198v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:36:46 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek, Benjamin A. Jasperson, Vivek Oommen, David L. Damm, Krishna Garikipati, Remi Dingreville</p><p>Latent neural surrogate solvers, or latent dynamics models, accelerate simulations of time-dependent physical systems by evolving a compressed latent space rather than resolving full-resolution fields directly. In principle this reduces computational cost and simplifies learning, but in practice errors often accumulate rapidly during long autoregressive rollouts, limiting predictive utility. We show that this instability does not stem from the latent representation itself, but arises when it is trained solely for reconstruction, producing representations poorly suited to long-horizon forecasting. We systematically evaluate training-level interventions that align latent representations with long-horizon rollout: Koopman operator learning and Hamming noise injection during autoencoder training to improve compression, together with noise injection and multi-step rollout fine-tuning to improve dynamics. Interventions that improve long-horizon rollout stability often degrade conventional training metrics, including reconstruction and one-step prediction accuracy. Collectively, these interventions reduce long-rollout error by approximately 40\% and match or exceed the accuracy of full-resolution models on two physics benchmarks, while requiring 2 orders of magnitude fewer floating point operations and half the GPU memory. Applied to mesoscale crystal-plasticity simulations of high-cycle fatigue, the resulting surrogate achieves stable extrapolation over horizons orders of magnitude beyond those observed during training. More broadly, these results show that neural compression should be designed not merely to reduce dimensionality, but to restructure the solution space for stable dynamical evolution, a key requirement for reliable, efficient neural surrogates in scientific applications.</p>]]></description>
  </item>
  <item>
    <title>SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance</title>
    <link>http://arxiv.org/abs/2609.30192v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30192v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:33:29 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xinyue Zeng, Jiawei Zhang, Yujun Yan, Dawei Zhou</p><p>Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compounding bias, where small local deviations accumulate across depth and suppress rare rewards. We introduce Symbolic Closure Analysis (SCA) as a theoretical lens characterizing how branching structures and sparse rewards induce these biases in long-horizon reasoning with local admissibility, and as a design principle for structural priors in less formal reasoning tasks. Motivated by this analysis, we propose SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning. SAGE combines two complementary structural guidance: algebraic sparsification, which projects locally admissible candidates onto operator-indexed algebraic subspaces to suppress spurious branching and mitigate exploration bias, and hyperbolic structural guidance, which embeds reasoning states into a negatively curved space to provide dense depth-wise signals and mitigate compounding bias. Across 12 benchmarks and 7 model families, SAGE outperforms competitive baselines. In particular, SAGE achieves up to an 8-fold improvement on the Andrews-Curtis problem, an open real-world long-horizon task. Code is available at: https://github.com/Susan571/SAGE-NeurIPS2026.</p><p><em>Comment: Accepted by NeurIPS 2026</em></p>]]></description>
  </item>
  <item>
    <title>Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures</title>
    <link>http://arxiv.org/abs/2609.30187v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30187v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:30:36 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Abhiram Maddukuri, Georgios Pavlakos</p><p>Ego-Exo4D is a large-scale dataset providing synchronized egocentric and multi-view exocentric video, a rich resource for skill learning and assessment, procedural activity understanding, and embodied AI. However, the dataset ships with only sparse 3D human pose annotations, and reconstructing dense human motion from its multi-view captures is nontrivial. To this end, we present Ego-Exo4D-HM, a large-scale dataset of 4D human motion reconstructions for Ego-Exo4D&#x27;s captures, and release the accompanying reconstruction pipeline. The code, dataset, and documentation can be found at https://abhiram824.github.io/egoexo4d_human_meshes.</p><p><em>Comment: Project website: https://abhiram824.github.io/egoexo4d_human_meshes</em></p>]]></description>
  </item>
  <item>
    <title>Jev-Mobile: Jev as an Executor for Mobile GUI Agents</title>
    <link>http://arxiv.org/abs/2609.30186v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30186v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:30:32 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Linghua Zhang</p><p>Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed decision model, repeatedly selects actions within this space. This design allows multiple GUI actions to be executed under a single VLM decision, reducing expensive VLM inference while preserving adaptive interaction. On the full AndroidWorld task suite, Jev-Mobile achieves 79% task success, compared with 78% for SeeAct-V and 84% for a Step-wise VLM baseline. Among successful trajectories, it reduces mean end-to-end execution time by 32.7% and mean model API cost by 73.4% relative to Step-wise VLM. These results show that decoupling high-level VLM reasoning from low-level action execution can substantially improve mobile GUI agent efficiency while maintaining competitive task performance.</p>]]></description>
  </item>
  <item>
    <title>Intrinsic-Extrinsic Coupling in Learning Dynamics</title>
    <link>http://arxiv.org/abs/2609.30185v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30185v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:30:22 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Qinyou Wang</p><p>A learner&#x27;s current observations need not determine its response to further training. We formulate intrinsic-extrinsic coupling through the continuation-conditioned value of a constrained learning-state intervention, with observation-relative fibers describing present agreement. An executable finite-frame classifier-head write protects current logits while repairing specified historical margins under finite-precision acceptance checks. We distinguish local admissibility, continuation-conditioned intervention value, and complete-policy performance. A matched four-cell contrast identifies readout-specific non-additivity between the same intrinsic intervention and alternative external continuations. In a CLINC-derived class-incremental setting, replay changes the write&#x27;s 32-update contribution from five correct predictions to zero. Nonzero interactions also occur under output distillation, with a RoBERTa backbone, and under optimizer-native SGDW dynamics. Under SGDW, correct-count interactions are negative in all three activated roots at 128 updates, showing that coupling need not imply positive synergy. The mathematical analysis distinguishes feasible local repairs and favorable terminal outputs from training-reachable repair regions. Separate coordination tests show that content controls match or exceed the development gain, while a five-root fresh-test comparison with Fiber present in every arm shows root-dependent rather than uniformly beneficial correct-count effects. On the secondary cross-entropy readout, guided allocation yields lower mean loss than standard replay in all five pairs. Together, these results make intrinsic-extrinsic coupling operational by connecting executable state geometry to continuation-conditioned value, matched interaction identification, and closed-loop coordination, while separating identified coupling from complete-policy performance.</p><p><em>Comment: 39 pages, 4 figures, 23 tables</em></p>]]></description>
  </item>
  <item>
    <title>ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints</title>
    <link>http://arxiv.org/abs/2609.30184v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30184v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:29:40 +0000</pubDate>
    <category>Generative Search Engines</category>
    <description><![CDATA[<p><strong>Authors:</strong> Sriram Kannan, Swetha Saseendran, Vishnu Vardhan Reddy Kandi, Leslie Barrett, Madhavan Seshadri, Enrico Santus</p><p>U.S. employment-discrimination complaints describe complex event sequences that are not explicitly captured by lexical or embedding-based representations alone. We present ARGUS, a source-grounded pipeline that combines a 5W1H-inspired schema, legal-domain models, and LLM-based structured generation to construct document-level Event Knowledge Graphs (EKGs) from CourtListener complaints. ARGUS extracts fact-bearing statements, builds chunk-level event graphs with participant, temporal, and causal structure, and merges them into document-level representations. We evaluate graph quality through human and multi-model assessment and test downstream utility on claim classification and legal QA. The graph-structured classifier outperforms raw and linearized baselines on the held-out set, and EKG-only retrieval improves document-scoped QA, while open-retrieval gains remain limited by low first-stage candidate recall. These results suggest that EKGs are most useful for organizing and reasoning over evidence once relevant material has been retrieved.</p><p><em>Comment: 9 pages, NLLP</em></p>]]></description>
  </item>
  <item>
    <title>Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search</title>
    <link>http://arxiv.org/abs/2609.30177v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30177v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:26:46 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Nayoung Choi, Shengjian Chen, Xiaokai Wei, Wenzheng Zhang, Daiyao Yi, Rachit Pareek, Vincent Su, Michelle Gong, Jinho D. Choi</p><p>Query understanding (QU) plays a critical role in production search systems, translating raw user queries into search execution plans that drive downstream retrieval and ranking. While large language models (LLMs) have enabled QU to be framed as a structured multi-task generation problem (e.g., intent classification, query expansion), optimizing such models to produce search-engine-coupled outputs remains challenging: static, label-based supervision fails to capture how each component actually interacts with the underlying search pipeline to affect downstream performance. We present a search-aware reinforcement learning (RL) framework for QU based on a distill-then-RL paradigm. Teacher-student supervised fine-tuning (SFT) first yields a well-formed, schema-compliant policy initialization. The RL stage then optimizes each QU component with rewards derived from live interaction with the search engine, tailored to that component&#x27;s operational role, rather than a single reward tied to the final search outcome. Experiments on Roblox search show that this component-specific optimization improves both per-component utility and downstream search quality, raising NDCG@20 by 8.9 points over the SFT policy and by 3.5 points over training with a single end-to-end reward.</p>]]></description>
  </item>
  <item>
    <title>GridSFM: A Foundation Model for Solving AC Optimal Power Flow</title>
    <link>http://arxiv.org/abs/2609.30173v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30173v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:25:06 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Luke Bhan, Weiwei Yang, Margaret Capetz, Baosen Zhang</p><p>We introduce GridSFM, a framework that combines a pretrained foundation model across grid topologies with physics-informed fine-tuning for solving AC Optimal Power Flow (AC-OPF) at scale. It is a $15$ million parameter physics-inspired graph neural network pretrained across $54$ topologies of $500$ to $4{,}000$ buses. Our model attains a $2.45\%$ zero-shot generation-cost error on a $10{,}000$ bus case held-out operating conditions with no degradation as system size grows. Building on this, we pair the pretrained backbone with a physics-informed fine-tuning design based on Newton&#x27;s method for power flow. With only $100$ solved instances, GridSFM adapts to unseen grids up to $10{,}000$ buses. We show it out performs single topology, dedicated neural network models that are trained more data, both in terms of cost and solver iterations when deployed as warm starting points.
  In designing this foundation model, we overcome the fact that the feasible set for AC-OPF can be disconnected. This is an obstruction that prevents any continuous neural network from approximating the solution map. To do so, we lift the problem and relax its constraints with logarithmically penalized slacks. We prove that the resulting elastic feasible set is contractible, that the AC-OPF minimizers remain minimizers of the elastic problem above an explicit penalty threshold, and that projecting an approximate solution back onto the AC-OPF feasible set is well posed. We release all models, data, and code so that the community can build on a shared starting point for AC-OPF.</p><p><em>Comment: 19 pages</em></p>]]></description>
  </item>
  <item>
    <title>Do Audio Language Models Hear and Read Distinctive Features Alike?</title>
    <link>http://arxiv.org/abs/2609.30167v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30167v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:21:16 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yuanhao Chen, Peter Chin</p><p>Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members&#x27; mean representations. Averaging those offsets gives a direction for each stream, and we measure the cosine between the two. Because the two streams already agree about arbitrary phoneme pairs, we compare every measure against a reference built from random pairings rather than against zero. We apply this to 6 models, 7 features and 15 languages from 11 families. Only voicing in the two Qwen2.5-Omni models exceeds that reference after correction for multiple testing, and the reference varies by a factor of seven between models. In three of the six models, voicing has one direction in audio across the 14 languages with enough minimal pairs to measure it, and every language pair agrees in two of them. The model family, not the model size, predicts which stream represents a feature.</p>]]></description>
  </item>
  <item>
    <title>A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition</title>
    <link>http://arxiv.org/abs/2609.30160v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30160v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:17:55 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh</p><p>Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling, or lexical realization that the acoustics do not uniquely determine. Omni-temporal Classification (OTC) tolerates such noise by adding wildcard paths to the connectionist temporal classification (CTC) alignment graph, but its word-level arcs are too coarse, since bypassing one unsupported token discards supervision for the whole word. We move wildcard arcs to token granularity so unsupported tokens can be bypassed while the rest of the word stays supervised, and we combine token- and word-level arcs as complementary escape paths. Across 19 languages and three corpora, token-level OTC improves over CTC on all 25 tasks. We also replace epoch-indexed relaxation of the wildcard weights with a predictive-entropy-indexed schedule, which performs comparably while reducing dependence on training length. Combining this schedule with the hybrid graph gives the lowest mean word error rate (WER) on every corpus and a 9.45% average relative WER reduction over CTC. Independent validator transcriptions show that token-level models place significantly more wildcard-bypass probability than CTC on disputed characters, indicating that token-level tolerance targets localized transcript ambiguity.</p><p><em>Comment: 5 pages, 2 figures, 4 tables; submitted to ICASSP 2027</em></p>]]></description>
  </item>
  <item>
    <title>Learning and interpreting policies for simultaneous entanglement requests in quantum networks</title>
    <link>http://arxiv.org/abs/2609.30157v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30157v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:16:00 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Leon Rode, Sumeet Khatri, Supartha Podder</p><p>Future quantum networks will make use of entanglement to perform numerous tasks, such as sending quantum information over long distances, distributed quantum computing, and quantum sensing. In general, these tasks will need to be performed simultaneously in various regions of a network, while minimizing resources and latency. We will thus require policies for scheduling link-level entanglement resources, and using the link-level entanglement to create various forms of multipartite entanglement required for every task. In this work, we address this problem using reinforcement learning. We formulate a Markov Decision Process for the problem and use double deep Q-networks (DQN) with Message Passing Neural Networks (MPNNs), experience replay buffers, and curriculum training to obtain policies. The key physical parameter is the probability of link-level entanglement generation, i.e., the link activation probability. We show that our policies maintain 100% success for up to 71% lower link activation probability than the baseline heuristics for a set of physically relevant network topologies. We then examine an additional constraint where experiment (task) placements are restricted to specific hardware types and demonstrate a similar advantage in performance over heuristics, with our policy maintaining at least an 80% success rate for up to a 59% lower link activation probability. Finally, we explore methods to interpret the learned policy by defining metrics enabling conclusions to be drawn about the model&#x27;s behavior and by tasking a large language model (LLM) to derive a novel heuristic given example actions taken by the DQN-trained policy. We find that the LLM heuristic performs similarly to the DQN-trained policy in performance, indicating a promising method for interpretable policy extraction for large quantum networks, where direct training becomes computationally expensive.</p>]]></description>
  </item>
  <item>
    <title>Does a model&#x27;s stated reason for rejecting a candidate do any work?</title>
    <link>http://arxiv.org/abs/2609.30151v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30151v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:13:35 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Archit Rastogi</p><p>Asked to choose between candidates and explain the choice, a language model often rejects a rival by naming a fact its profile lacks: no director, no date of death. That sentence is a claim about the text in front of the model, and it can be tested without any judge. We insert a real corpus sentence stating the named fact into the rival&#x27;s profile and ask again under greedy decoding. Two controls separate content from placement: a length-matched irrelevant sentence at the same profile, and the same two sentences at a third option the model never mentioned. In the largest of three runs, six open models on 2WikiMultihopQA, supplying the named fact at the profile the model named moves its choice more than the irrelevant control does, odds ratio 3.57 [1.54, 8.26], Holm p=0.0210, and this survives dropping any single model. The contrast the design was built to detect, the same fact at the option nobody named, does not clear correction, Holm p=0.2428. The strongest result in the family carries no content claim at all: the identical irrelevant sentence moves the choice more at the named rival than at the third option, Holm p=0.0008. Repair and control also differ in co-candidate mentions, relation template and fluency; post-hoc matching on the first two preserves the content effects&#x27; direction, matching fluency weakens one, so the content contrasts bound an effect rather than establish one. A forced single-token probability read disagrees in direction with the free-text choice on that same contrast, and three candidate explanations for the disagreement find no support. Every measurement is a string rule, so each was validated against the records it reads; validation caught eight defects. The largest, a choice-parsing rule that returned the option a model had just rejected in 17.1% of adjudicable responses, would have reported six surviving contrasts instead of four.</p><p><em>Comment: Accepted as an oral presentation at LLM4XAI 2026: Workshop on Large Language Models for Explainable AI, co-located with CIKM 2026, Rome, Italy, November 8, 2026. Code and per-item records: https://github.com/ArchitRastogi20/contrastive-rejection-test</em></p>]]></description>
  </item>
  <item>
    <title>Graph-Based Inference and Topology-Aware Multi-Agent Reinforcement Learning for Large-Scale Railway Network Management</title>
    <link>http://arxiv.org/abs/2609.30150v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30150v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:12:39 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Giacomo Arcieri, Gregory Duthé, Christophe Muller, Konstantinos G. Papakonstantinou, Daniel Straub, Eleni Chatzi</p><p>Modern infrastructure asset management constitutes a complex sequential decision-making problem, characterized by long planning horizons and system-level interactions, such as spatial deterioration correlations and economies of scale. While deep reinforcement learning has shown promise in optimizing maintenance policies, scaling to real-world networks remains challenging. Centralized approaches become computationally intractable in large-scale systems, whereas decentralized approaches often fail to capture essential coordination mechanisms. To address these challenges, we propose a graph-based framework that integrates accurate environment modeling with scalable decision support. First, we employ a hierarchical Bayesian model leveraging a Gaussian Process on Graph kernel to infer a realistic, spatially correlated networked environment of railway maintenance planning from real-world data provided by the Swiss Federal Railways. Second, we introduce a topology-aware Multi-Agent Reinforcement Learning (MARL) framework by integrating graph neural networks and graph Transformers to optimize network-level policies. A central contribution of this work is the demonstration of scalability through zero-shot transfer learning: graph-based agents, trained only on small network portions, are successfully deployed in a zero-shot manner on large-scale unseen networks without any retraining. Numerical results indicate that the proposed method significantly outperforms optimized heuristics and standard MARL baselines, reducing computational training time while maintaining superior performance on large-scale networks.</p>]]></description>
  </item>
  <item>
    <title>GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI</title>
    <link>http://arxiv.org/abs/2609.30147v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30147v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:11:35 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Arunabh Srivastava, Mohammad A.,  Khojastepour, Srimat Chakradhar, Sennur Ulukus</p><p>Large Language Models (LLMs) typically exhibit a performance profile where reliability degrades as task complexity increases. We address the challenge of generating high-quality natural language executable plans for complex tasks by introducing $\textbf{GRASP}$, a strategy-aware, multi-stage planning framework. GRASP decouples the planning pipeline across specialized, context-isolated modules: it pre-compiles global macro-guidelines (GenPlan), explores alternative localized strategies within isolated context windows (RevPlan), and independently evaluates trajectories using a multi-criteria discriminator (VerPlan). Empirical evaluations show that GRASP consistently establishes a new state-of-the-art frontier across diverse datasets, yielding substantial accuracy gains over direct LLM planners on Natural Plan Calendar Scheduling ($\sim$12.4$\%$$\uparrow$), ZebraLogic ($\sim$30.8$\%$$\uparrow$), and SciBench Math. Crucially, under multi-task scaling-where standard planners suffer immediate performance collapse-GRASP completely flattens the multi-task degradation penalty. In interleaved dual-task environments, GRASP achieves an absolute accuracy gain of up to 16.7$\%$ over direct LLM planners. Furthermore, by isolating context and enforcing strict macro-regularization, GRASP outperforms frontier reasoning models (such as GPT-5-mini) by a margin of 14.5$\%$.</p><p><em>Comment: Accepted at the Second Workshop for Research on Agent Language Models (REALM) at EMNLP 2026</em></p>]]></description>
  </item>
  <item>
    <title>EnigmaForge: The Question Is Hidden in the Story</title>
    <link>http://arxiv.org/abs/2609.30144v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30144v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:10:03 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Daniel Eisner</p><p>Most benchmarks hand the model a question. EnigmaForge hands it a stack of old documents and no question at all. Buried in the letters, receipts, and logbook margins is a small logic puzzle whose solution is unique - proved by a SAT solver at generation time, with an ablation certificate showing every clue is load-bearing. Because instances are generated rather than collected, the corpus renews forever. The headline measure is intuition: task success when handed only the story, with world reconstruction as the secondary axis. Twenty-five frontier models ran over 600 instances (17,400 scored records) under three matched conditions. Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it. Several models were blocked by their own content filters before reaching the puzzle - any benchmark scoring refusals as failure is quietly measuring filter behavior.</p>]]></description>
  </item>
  <item>
    <title>Contact as a Decision Variable: Capability-Tradeoff Contact Selection for Legged Loco-Manipulation</title>
    <link>http://arxiv.org/abs/2609.30140v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30140v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:08:28 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Al Jaber Mahmud, Shuai Li, Xuan Wang</p><p>In this paper, we study the joint selection of an environmental support contact and a whole-body configuration for a prescribed loco-manipulation task. A contact may provide greater physical support while restricting the motion required for the task. We formulate this problem through three capability measures: residual wrench, end-effector reach, and base mobility available after satisfying the task requirements, and we balance them against contact acquisition cost. Evaluating these capabilities for every candidate requires repeated whole-body optimizations. To reduce this computational cost, we propose Capability-Tradeoff Contact Selection (CTCS). CTCS screens candidates for contact and task feasibility, groups similar candidates within each surface, and predicts their capabilities from exact anchor evaluations using local sensitivity analysis. It checks these predictions through selective exact evaluations, ranks candidates by capability, and evaluates a shortlist exactly for final selection. We evaluate CTCS in simulations and hardware experiments using a Unitree Go2 quadruped with an AgileX NERO arm across $392$ task conditions with nine available support surfaces. Results show that CTCS outperforms ground-only and fixed-contact support, as it can select support surfaces that provide favorable capability trade-offs for the task. Compared with evaluating every candidate exactly, CTCS achieves approximately $3\times$ speedup while closely matching the resulting mean objective value.</p><p><em>Comment: 9 pages, 6 figures</em></p>]]></description>
  </item>
  <item>
    <title>Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale</title>
    <link>http://arxiv.org/abs/2609.30137v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30137v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:07:38 +0000</pubDate>
    <category>Foundation Models</category>
    <description><![CDATA[<p><strong>Authors:</strong> Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Conceição Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath</p><p>Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization&#x27;s products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust.
  We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank&#x27;s Card Delivery agent and its expanded successor, Card Management - Nubank&#x27;s highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.</p><p><em>Comment: 17 pages, 11 figures</em></p>]]></description>
  </item>
  <item>
    <title>Training-free Behavior Cloning</title>
    <link>http://arxiv.org/abs/2609.30134v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30134v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:05:58 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager</p><p>Neural behavior cloning compresses demonstrations into large models, making individual actions difficult to trace and policy updates costly. Retrieval policies retain access to demonstrations but struggle with mismatch between recorded and live behavior. We introduce Behavior Predictive Control (BPC), which synthesizes policies without end-to-end policy training by combining an action-aware retrieval metric, a Hankel-based action-continuation prior, and a closed-form one-step residual correction. Inspired by behavioral systems theory, BPC predicts future actions by blending stored observation-action data that best reconstructs the recent runtime observation--action history. Across simulated benchmarks and real-robot deployments, BPC is competitive with learned policies such as $π_{0.5}$ (surpassing it in some cases), while reducing policy fitting from hours to seconds on consumer GPUs and supporting closed-loop control upwards of 75 Hz on a Jetson Orin Nano. The retrieved demonstration windows and their coefficients also provide an intrinsic estimate of task progress. Retaining demonstrations within the deployed policy makes its predictions traceable to supporting trajectories and enables behavior revision through the demonstration bank.</p>]]></description>
  </item>
  <item>
    <title>Multimodal Thinking with Renderable Programs</title>
    <link>http://arxiv.org/abs/2609.30130v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30130v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:03:44 +0000</pubDate>
    <category>Digital Human</category>
    <description><![CDATA[<p><strong>Authors:</strong> Sunli Chen, Ding Zhong, Ziqiao Ma, Jiaxin Liu, Zeyuan Yang, Hao Zhang, Lie Lu, Joyce Chai, Chuang Gan</p><p>Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domain, lacking tractability due to rasterized or latent representations of images. We introduce SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks. We exploit the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process. We provide a large curated dataset of SVG-based image editing dataset, as well as the paradigm to tune open-source VLMs. Experiments on a mathematical reasoning benchmark demonstrate that SVGLM achieves strong SVG generation power as well as think-with-image intelligence. Our results highlight SVG as a suitable medium for building more robust digital domain agents, bridging the gap between text-based thinking and pixel-based images.</p>]]></description>
  </item>
  <item>
    <title>Faster Visuomotor Policy Learning on Action Manifolds via Riemannian MeanFlow</title>
    <link>http://arxiv.org/abs/2609.30127v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30127v1</guid>
    <pubDate>Thu, 24 Sep 2026 17:00:57 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> S. Talha Bukhari, Austin Garrett, Yi Wei, Ruiqi Ni, Zachary Kingston, Aniket Bera</p><p>Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector field for action generation, which can be expensive and time-consuming, impeding fast control rates required in robotics applications. Furthermore, robot action sequences are usually defined on a smooth, differentiable manifold, requiring that the learned policy respects the intrinsic geometry of the robot&#x27;s action space. Here, we present Riemannian MeanFlow Policy (RMFP), which learns the conditioned flow map of the probability path on the robot action manifold. Our formulation employs a flow map consistency objective grounded in the data by a Riemannian Conditional Flow Matching anchor. The flow map consistency condition is stable to train and constrains the learned model to finite-time transport, which yields on-manifold action sequence generation with as few as one network function evaluation. We present results on the spherical LASA and Push-T benchmarks, on the Tool Hang and Transport tasks of the Robomimic suite, and on the Franka Kitchen task with manifold-constrained action generation, and demonstrate that RMFP attains performance competitive with prior work at a lower sampling cost. We also employ RMFP on a real-world robotic manipulation task to demonstrate fast action generation under imperfect sensor measurements in the physical world.</p>]]></description>
  </item>
  <item>
    <title>HEXIS: Compiling Skills into Extended Finite State Machines</title>
    <link>http://arxiv.org/abs/2609.30123v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30123v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:58:18 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Minghao LI</p><p>Agent skills provide reusable knowledge and instructions, yet agents must repeatedly infer how to apply them and which operation should follow. This couples task reasoning with control decisions, allowing prescribed steps to be omitted or applied incorrectly. We introduce HEXIS, which compiles agent skills into extended finite state machines that separate knowledge from control flow. Skill knowledge is incorporated into local instructions that guide reasoning and generation within states. The machine records execution progress and intermediate results, while explicit transition conditions determine subsequent operations. Our incremental compiler first maps skill clauses and tool interfaces to state operations, local instructions, data bindings, and transitions. It then aligns development traces with existing states to identify missing operations and dependencies. These are incorporated by adding or reusing states and refining their connections. Updates are accepted only after static checks and replay of the current and all previously accepted traces. Across four benchmarks and four executors, HEXIS improves success over Skill + ReAct by 16.1 percentage points on average. Qwen3.8-27B reduces execution tokens by 38.4-88.9% across benchmarks.</p>]]></description>
  </item>
  <item>
    <title>What, When, and How: Audio Description as Constrained Global Optimization</title>
    <link>http://arxiv.org/abs/2609.30121v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30121v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:56:50 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller</p><p>Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.</p>]]></description>
  </item>
  <item>
    <title>Orbital Error Dynamics: Self-Organized Criticality, Ephemeral Parameter Resonance, and Non-Linear Biological Ontologies in Zero-Storage Neural Synthesis</title>
    <link>http://arxiv.org/abs/2609.30115v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30115v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:55:06 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Volkan Dağlı, Zerrin Dağlı, Dağhan Dağlı</p><p>Modern deep neural networks treat parameters as static floating-point matrices stored in physical memory, incurring Von Neumann memory bottlenecks and representation collapse. We formulate Orbital Error Dynamics (OED), an analytical framework wherein synaptic weights are not stored masses (O(W)), but transient topological resonances (O(1)) derived procedurally from the complex quadratic polynomial map z_{n+1} = z_n^2 + c. We introduce the Bent Sine Wave Hypothesis, demonstrating that non-equilibrium living systems emerge when harmonic waves curl inward through environmental drag toward the cardioid cusp (c = 1/4). We define the Observer Horizon Geometry in parameter space, identifying interior resonance shoulder loci X_upper = (0.25, +0.18) and X_lower = (0.25, -0.18) between the fixed-point basin and the true boundary at c = 0.25 +/- 0.50i. To escape non-convex stagnation without loss zeroing, we introduce a heavy-tailed Biomimetic Perturbed Jump Operator (Omega_tunneling) inspired by mammalian fertilization zinc sparks. We further couple an enteric-cranial Dual-Brain architecture shielded by adaptive CD4+ regulatory immune gating (M_CD4), and project the 4-nucleotide genetic basis (A, T, C, G) across quadrants in C. Multi-seed empirical validation on the Two-Moons manifold (5 seeds, 80/20 train/test split, 32x32 grid, zero test-time updates, zero label leakage) demonstrates that procedural parameterization from a 24-byte coordinate seed achieves 77.67% +/- 5.35% clean test accuracy (within an 8.00-point paired difference of an unconstrained gradient baseline at 85.67% +/- 5.35%, 95% CI: [-1.07%, 17.07%]) and 71.33% +/- 3.80% under distribution shift (N(1.2, 0.4)), alongside conceptual equivalence with an analog optical co-processor.</p><p><em>Comment: Official National Patent Priority: TR 2026/016285 (Filed Sept 22, 2026). Foundational companion theory to Mandelbrot Fractal Neural Synthesis. Code and interactive lab: https://github.com/pCwOrM/mandelbrot-fractal-neural-synthesis</em></p>]]></description>
  </item>
  <item>
    <title>On the SoS Certifiability of Log-Concave Distributions</title>
    <link>http://arxiv.org/abs/2609.30105v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30105v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:48:02 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Aleksandr Storozhenko</p><p>For an arbitrary isotropic log-concave distribution $P$ on $\mathbb{R}^d$, we prove that the polynomial $(Cm)^m\|v\|_2^m - \mathbb{E}_{X\sim P}\langle X,v\rangle^m$ is a sum of squares for every even $m\ge2$, where $C&gt;0$ is a universal constant. This removes the dependence on the Poincaré constant in the theorem of Kothari and Steinhardt (arXiv:1711.07465), recovering the optimal moment bounds for log-concave distributions. As an immediate corollary, we obtain computationally efficient algorithms with dimension-free error guarantees for a wide range of high-dimensional statistical estimation problems.
  Our proof uses stochastic localization to decompose $P$ as an average of random strongly log-concave measures, whose centered moments admit the subgaussian certificates of Diakonikolas, Hopkins, Pensia, and Tiegel (STOC 2025; arXiv:2410.21194). With a covariance-adapted choice of localization, we show that a fourth-moment certificate derived from Letwin&#x27;s variance inequality for quadratic forms (arXiv:2607.24164) suffices to control this averaging at every even degree.</p>]]></description>
  </item>
  <item>
    <title>MQSS-Selector: RL-Guided Pass Selection for an MLIR Compilation Pipeline</title>
    <link>http://arxiv.org/abs/2609.30104v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30104v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:46:49 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Andre Youssefi, Ercüment Kaya, Minh Chung, Jorge Echavarria, Laura B. Schulz, Martin Schulz</p><p>High Performance Computing (HPC) and Quantum Computing (QC) systems are increasingly converging towards unified High Performance Computing-Quantum Computing (HPCQC) infrastructures, driven by a growing need to bridge classical and quantum workflows, which affects all levels of the system stack, from the hardware to compilers and runtimes, all the way to applications. However, today&#x27;s QC devices are still in the Noisy Intermediate-Scale Quantum (NISQ) era, are error-prone and resource-limited, and therefore require specialized optimizations and topology mappings to achieve sufficient fidelity. This places special emphasis on proper compilation and optimization within the overall quantum software stack. Many existing stacks remain fragmented, with separate components responsible for device selection, compiler-pass optimization, and job queue scheduling. This paper proposes a unified, learning-based selector that integrates these disparate stages into a cohesive framework. Our proposed selector scheme leverages reinforcement learning and deep learning models that can be extended to simultaneously optimize multiple objectives -- such as fidelity, compilation time, and scheduling latency -- while dynamically adapting to circuit characteristics and device conditions.</p><p><em>Comment: 11 pages, 5 figures, 1 table</em></p>]]></description>
  </item>
  <item>
    <title>R-DEIM Net: An Efficient Rationale-Augmented Dual-Expert Interaction Model for Paraphrase Detection</title>
    <link>http://arxiv.org/abs/2609.30100v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30100v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:44:26 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong>  Pushp, Vaibhav Prajapati, Himangshu Sarma</p><p>Recent advances in paraphrase detection reveal a fundamental trade-off: large language models achieve high accuracy but require high computation, while efficient Siamese-BERT variants offer practical scalability with reduced transparency in rationale generation. We present R-DEIM Net, a 76M-parameter dual-expert architecture exploring whether moderate-scale models can achieve competitive accuracy on paraphrase detection while enabling human-readable rationale generation. The architecture combines two specialized components: an Interaction Expert that captures token-level similarity patterns through multi-scale 2D convolutions and attention head allowing variable input length, and a Reasoning Expert that uses a Flan-T5-small decoder to generate rationales as auxiliary supervision. Rather than re-encoding generated text, we extract and pool decoder hidden states as complementary features for classification. On the Quora Question Pairs dataset, R-DEIM Net achieves 90.07\% accuracy and 90.16\% F1-score via 10-fold cross-validation. This represents competitive performance with strong transformer-based baselines (e.g., MFAE BERT: 90.54\% accuracy) and recent large language model based approaches (LLaMA-70B) while using a substantially smaller parameter budget. The model generates rationales alongside predictions, providing potential for auxiliary human-readable descriptions.</p>]]></description>
  </item>
  <item>
    <title>PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations</title>
    <link>http://arxiv.org/abs/2609.30094v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30094v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:39:18 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Luciano Maldonado</p><p>Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce \textbf{PrivDrift}, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1{,}000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7\% to 54.6\%, and varies strongly by model, secret type, and persuasion intensity. Within the tested drift window, additional topic drift does not reliably reduce leakage, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.</p><p><em>Comment: Preprint, 10 Pages, 6 figures</em></p>]]></description>
  </item>
  <item>
    <title>Self-Adaptive VLA for Robust Robot Deployment</title>
    <link>http://arxiv.org/abs/2609.30092v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30092v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:37:46 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Hongxin Zhang, Chunru Lin, Tsun-Hsuan Wang, Zhenjia Xu, Chuang Gan</p><p>While Vision-Language-Action (VLA) models demonstrate impressive capabilities in robotic manipulation, their memoryless nature renders them brittle to test-time environment shifts, particularly hardware shifts caused by wear or imperfect calibration. Enabling these models to self-adapt during deployment without requiring continuous on-site recalibration remains a critical bottleneck for real-world scalability. In this work, we introduce Self-Adaptive VLA, a novel post-training recipe that enables the policy to iteratively adapt to deployment-time hardware shifts leveraging its own rollouts as context. To do so, we first collect policy rollouts under deliberately injected hardware shifts. We then transform the base policy&#x27;s training data into shift-conditioned expert demonstrations by pre-compensating the expert actions for these known shifts. Next, we introduce a lightweight, plug-in context encoder that compresses the context, including visual observation, proprioception, and actions in the shifted environment, into a latent context token. This token modulates the policy through adaptive layer normalization (AdaLN). Furthermore, we find that context tokens can be ensembled, allowing the policy to iteratively self-correct and mitigate failures step by step. Extensive experiments across four precision-critical bi-manual and dexterous manipulation tasks show that Self-Adaptive VLA recovers over 80% of the base policy&#x27;s performance under hardware shifts, such as actuation bias and joint encoder offsets. Moreover, Self-Adaptive VLA enables more robust deployment to new workstations compared to the base policy. Our approach provides a pathway for robust large-scale real-world robot deployments and easier maintenance. See videos at https://icefoxzhx.github.io/self-adaptive-vla.</p>]]></description>
  </item>
  <item>
    <title>AT-SKM-Net: An Accelerated Trainable Sampling Kaczmarz-Motzkin Framework for Linear Hard-Constraint Feasibility on Dynamic Graphs</title>
    <link>http://arxiv.org/abs/2609.30088v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30088v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:36:49 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xiaochen Zhang, Haoyu Zhu, Yao Zhang, Qingchun Hou</p><p>Graph-structured optimization with linear constraints is fundamental to critical infrastructure but faces scalability limits due to massive strict hard constraints and high dimensionality. While recent projection-based methods such as Trainable Sampling Kaczmarz-Motzkin Net (T-SKM-Net) guarantee feasibility, they face high computational costs in dynamic environments by processing the entire constraint set and requiring expensive matrix factorizations. To bridge this gap, we propose the Accelerated Trainable-SKM (AT-SKM) Net framework. To concentrate computation on the active constraints and eliminate redundant calculations, we introduce a hybrid sampling strategy guided by a topology-aware heterogeneous GNN model. To efficiently handle topological shifts in graph-based constraints, we employ a Cholesky Update mechanism that theoretically reduces the equality projection complexity from O(N^3) to O(N^2) under low-rank perturbations. Experiments on random geometric graphs, N-1 Security-Constrained DC-OPF, and minimum-cost gas transport problem demonstrate that AT-SKM reduces iteration counts by up to 85% and achieves 2.95x-7.29x SKM layer speedups, while maintaining zero constraint violations.</p>]]></description>
  </item>
  <item>
    <title>Return or Revise? Learning When Revision Helps Retrieval-Augmented QA</title>
    <link>http://arxiv.org/abs/2609.30087v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30087v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:35:54 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Nicholas Kashani Motlagh, Tim Anderson, Jeremy Gwinnup, Grant Erdmann</p><p>We consider the decision of whether to return an existing draft answer or revise it using retrieved evidence, as in answer-revision systems. Draft confidence estimates whether the current answer is correct, but the decision requires estimating the effect of a specified revision. For offline training and evaluation, we grade both the returned draft and its candidate revision under the same correctness judge, which makes repair, harm, and the gap to an oracle observable. We call this paired effect its recoverability, and we train policies to predict it before revision. On 25,870 held-out open-domain questions across three revision setups, a scorer trained on the paired outcome has greater area under the accuracy--revision-rate curve than a matched draft-correctness scorer in all nine Llama setup--seed fits, and gains 0.23--0.68 accuracy points on average at development-selected thresholds, a difference significant across training runs only for dense retrieval. The resulting policy improves on always revising and on average closes more than a third of the oracle gap, although it still applies 38--46% of the harmful revisions. When a draft-free standard-RAG answer is also available, however, choosing between the draft and that answer is stronger by about two points for Llama and four for OLMo, and adding candidate revision as a third option yields no significant gain. Recoverability describes one revision; its value as an available action also depends on the alternatives.</p><p><em>Comment: 25 pages, 4 figures</em></p>]]></description>
  </item>
  <item>
    <title>Residual Correlation as a Diagnostic for Joint-Uncertainty Gains from GP Coregionalisation</title>
    <link>http://arxiv.org/abs/2609.30085v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30085v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:33:59 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Fangqin Zhou, Joaquin Vanschoren</p><p>In multi-target regression, correlated targets are often coupled through multi-output Gaussian processes with an intrinsic model of coregionalisation (GP-ICM), assuming that sharing statistical strength improves overall performance. In practice, the benefits are inconsistent. Across the settings studied, we find that the main benefit of coregionalisation is joint uncertainty quantification rather than point prediction. Raw target correlation does not predict when coupling helps; in the separable GP-ICM settings studied here, residual correlation, the cross-target dependence left unexplained by independent per-target predictors, is the strongest predictor of joint-uncertainty gains.
  We introduce a lightweight diagnostic, $D_{\rm logdet}=-\frac{1}{2}\log\det R_{\rm res}$, which represents the idealised joint negative log-likelihood (NLL) gain from modelling a full rather than diagonal residual covariance and is computable from independent GPs alone. Across a controlled synthetic study, 16 multi-target benchmarks, and frozen transformer and convolutional neural network representations for keypoint regression, point prediction remains largely unchanged ($ΔR^2\approx 0$). In contrast, $D_{\rm logdet}$ strongly predicts observed ICM NLL improvements ($ρ_s=-0.83$, $p&lt;0.001$), outperforming heuristics such as the feature-to-sample ratio. We also propose Residual-ICM, which preserves independent marginal variances while adding residual-correlation structure to the joint covariance. Residual-ICM achieves the best average joint NLL among the compared methods, while the diagnostic indicates when covariance coupling is likely to be useful. The diagnostic is specific to global Gaussian residual dependence, the structure captured by separable coregionalisation.</p><p><em>Comment: Accepted at ACML 2026</em></p>]]></description>
  </item>
  <item>
    <title>Real-Time Force Regulation for Whole-Hand Dexterous Grasping</title>
    <link>http://arxiv.org/abs/2609.30082v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30082v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:31:51 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Sang Min Kim, Alexander Alexiev, Tzu-Yuan Lin, Sangbae Kim, Young Min Kim, Yonghyeon Lee</p><p>Robust dexterous grasping requires maintaining physical stability despite contacts interactively evolving across the entire hand. A precomputed force distribution can easily fail under object motion, modeling errors, or external disturbances. In this paper, we present a framework for real-time force regulation over dynamically changing whole-hand contacts. Our method geometrically estimates contacts across all hand links using a tracked object model and proprioception, without requiring tactile sensing at those contacts. It repeatedly recomputes the desired contact-force distribution subject to friction constraints, actuator limits, and an actuation-consistency constraint motivated by classical whole-limb force analysis. We integrate this force-regulation controller with reactive reaching, enabling the hand to acquire a grasp, maintain it under disturbances, and regrasp after losing the object. Simulation experiments without gravity demonstrate improved grasp retention over fixed-allocation and fingertip-only execution under controlled perturbations, while real-world experiments on a 27-DoF arm-hand system demonstrate grasp maintenance and recovery under human-applied disturbances as contacts evolve across the whole hand. Project page: https://sangminkim-99.github.io/reactive-grasp-whole-hand/</p><p><em>Comment: 9 pages, 10 figures. Project page: https://sangminkim-99.github.io/reactive-grasp-whole-hand/</em></p>]]></description>
  </item>
  <item>
    <title>Can Frozen Hyperspherical Features Guide the Selection of Pseudo Masks?</title>
    <link>http://arxiv.org/abs/2609.30080v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30080v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:30:39 +0000</pubDate>
    <category>End-to-End AD</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xinge Guo, Fengyang Xiao, Dingming Zhang, Yuhan Chen, Rihan Zhang, Xingjian Li, Tianyang Wang, Chunming He, Sina Farsiu</p><p>Foundation segmenters such as SAM return several plausible masks for an unlabeled image, and a student trained on the wrong one inherits its errors. Choosing among them means querying a second large model or fitting a quality head to annotated masks. We show that a candidate can be judged by what it does to a frozen self-supervised backbone&#x27;s features. Normalized DINOv2 patch features lie on a hypersphere, and a candidate mask splits that sphere in two. Based on this reading, we introduce SphereTrust, which scores each candidate by three properties of the split, the angular contrast between the two sides, the coverage of the foreground&#x27;s appearance modes, and contact with the image frame, one for each of three common ways a mask fails, and ranks a pool in 0.55 s per image from the frozen features alone. On eight SAM and SAM3 candidate pools spanning camouflaged, salient, and dichotomous segmentation and camouflage under low light, SphereTrust exceeds the strongest evaluated external baseline on six pools by 1.7 to 9.3 percentage points in mean selected Dice. These comparisons include published selection rules and explicitly labeled adaptations of DSS and UCOD-MKD. On the two prompted camouflage pools, its mean selected Dice is within 0.1 percentage points of the candidate-derived DSS adaptation, with a lower catastrophic-error rate. Which cue carries the signal depends on the candidate pool. The same sphere also supports training. The leading candidates enter as a candidate set with their scores as priors, prototypes reorder them, and a cross-fitted second round completes the labels, raising weighted F by 4.5, 2.3, and 5.5 points over fixed-label training on the three MLLM anchor pools, with students competitive with published unsupervised methods on nineteen test sets.</p><p><em>Comment: 29 pages, 11 figures, 19 tables</em></p>]]></description>
  </item>
  <item>
    <title>Reachability-Based Formal Verification of Graph Neural Networks with Node and Edge Features</title>
    <link>http://arxiv.org/abs/2609.30079v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30079v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:29:32 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Anne M. Tumlin, Ben Wooding, Zhenxuan Shao, Diego Manzanas Lopez, Tyler Derr, Taylor T. Johnson</p><p>Graph neural networks (GNNs) have become a prominent approach for developing fast, topology-aware surrogates in electric power systems, supporting tasks such as power flow (PF) analysis, optimal power flow (OPF) estimation, and cascading failure analysis (CFA). Despite this growing use, formally verifying GNN-based models remains challenging, with existing methods limited in scope. We extend the neural network verification (NNV) framework to graph-structured inputs through GraphStar sets, a generalization of Star sets that captures uncertainty over both node and edge features. This extension enables the propagation of linear message-passing operations and the sound approximation of ReLU nonlinearities for GNN architectures, including graph convolutional network (GCN) and graph isomorphism network with edge features (GINE) layers. We evaluate GNNV across three power system tasks, PF, OPF, and CFA, on the IEEE-24, IEEE-39, and IEEE-118 test cases, as well as two standard graph classification benchmarks, ENZYMES and PROTEINS. Our results show that GNNV provides tighter robustness guarantees than CORA on graph classification models with ReLU-based activations and, for the first time, delivers edge-aware robustness guarantees for GINE-based PF and OPF models under joint node and edge perturbations.</p>]]></description>
  </item>
  <item>
    <title>Nuclear Norm-Regularized Bayesian Matrix Completion</title>
    <link>http://arxiv.org/abs/2609.30078v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30078v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:29:25 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Calvin Tolbert</p><p>Matrix completion, the problem of estimating missing entries in a matrix from noisily observed ones, underlies a diverse array of problems such as recommender systems and counterfactual outcome estimation in panel data. Many algorithms address the problem using regularized least squares, often with the nuclear norm as a regularizer, but this method yields a point estimate with no built-in uncertainty quantification. A Bayesian formulation is a natural alternative, and if the noise variance is known, the nuclear norm-based prior yields a log-concave posterior. Unfortunately, in practice, the noise variance will not be known a priori, so for a fully Bayesian approach, a prior must be imposed on it. We give the first sampler for this model with an explicit non-asymptotic guarantee: polynomial in the matrix dimensions and in the reciprocal of the target accuracy. Our technique is to discretize the distribution of the noise precision onto a grid and build a categorical posterior via thermodynamic integration. This extension is not specific to matrix completion and may be useful in other non-log-concave sampling problems where the non-log-concavity is restricted to a single variable and the joint distribution of the remaining variables is nonsmooth. Our contribution is a feasibility result: we show that a polynomial-time Bayesian sampler for this model exists at all, and the resulting complexity, while polynomial, is not intended as a deployable algorithm at current problem scales.</p><p><em>Comment: 26 pages, 2 figures</em></p>]]></description>
  </item>
  <item>
    <title>A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis</title>
    <link>http://arxiv.org/abs/2609.30075v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30075v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:28:45 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Tina Raissi, Nhan Phan, Chenxiao Wang, Mikko Kurimo</p><p>Automatic speaking assessment systems can provide holistic proficiency scores, but often lack interpretable measures that characterize pronunciation quality. We propose a native-reference phone-class geometry for measuring second language (L2) pronunciation deviation without requiring pronunciation labels, read-aloud prompts, or matched recordings of the same text from native and L2 speakers. Given a native speech corpus, we average frame-level self-supervised representations for each context-dependent phone-class and use singular value decomposition (SVD) to derive a compact native-reference coordinate system. For each L2 utterance, we compute the corresponding averages and project them into the native-reference space. We then demonstrate that the distances between L2 and native-reference coordinates for matched phone-classes show consistent negative correlations with holistic speaking proficiency on the Dev subset of the Speak and Improve Corpus 2025 (Spearman&#x27;s $ρ\!=\!-0.53$) and with pronunciation quality on the learner subset of the English Read by Japanese Students dataset ($ρ\!=\!-0.34$). These findings suggest that the proposed geometry captures acoustic-phonetic information relevant for proficiency rating while remaining applicable to spontaneous L2 speech without matched native recordings.</p><p><em>Comment: Submitted to ICASSP 2027</em></p>]]></description>
  </item>
  <item>
    <title>How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure</title>
    <link>http://arxiv.org/abs/2609.30074v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30074v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:28:15 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Dipankar Sarkar</p><p>Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted. The measured phenomenon is unstable to begin with. Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect. Auditing the evaluation weakens its conclusions further, and this is our main contribution. Under a joint cluster bootstrap over prompts, only the bottom of the ranking is firm: the two least reproducible models hold rank in 99% and 86% of replicates, the middle four in 27% to 48%, and the top two in 68% each, so the table identifies the worst model reliably but does not reliably identify the best. Two equally defensible rules for merging repeated campaigns change four of eight rows and move the study-wide headline by 7 percentage points. Checking the inferred structure against ground-truth annotations shows reproducibility cannot be read as accuracy. And four of the eight endpoints were withdrawn within ten weeks of measurement, so the study as specified can no longer be run. Small-sample LLM evaluations can therefore look far more definitive than their evidence supports. We recommend reporting rank stability, per-cell provenance, executed sensitivity comparisons, raw per-run outputs, and a measurement date alongside any ranking.</p><p><em>Comment: 13 pages. Previously submitted to TAE (Trust-AI-Eval), a NeurIPS 2026 workshop</em></p>]]></description>
  </item>
  <item>
    <title>Scoring Both Directions: LLMs realize the MRS they cannot reliably parse</title>
    <link>http://arxiv.org/abs/2609.30071v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30071v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:27:54 +0000</pubDate>
    <category>Generative Search Engines</category>
    <description><![CDATA[<p><strong>Authors:</strong> Soham Dan</p><p>The English Resource Grammar (ERG) is a hand-written computational grammar of English. Given a sentence, its processor, ACE, produces a formal meaning representation called Minimal Recursion Semantics (MRS): a graph of the sentence&#x27;s predicates and their arguments. The grammar is bidirectional and can also turn an MRS back into an English sentence. \citet{hajdik2019} used the ERG&#x27;s treebank to build a benchmark for that generation task, MRS to text, and trained sequence-to-sequence models to solve it. The parsing task, text to MRS, can be tested on the same sentences. We reconstruct their 10K-sentence test split, and score two large language models, Claude Sonnet~4.5 and Claude Opus~5, in both directions against their trained systems and against ACE, with no task-specific training. Given an MRS and three examples, Opus writes the sentence at 76.3 BLEU, ten points above their system trained on 72k pairs (66.1 BLEU), and comparable to their system trained on a million extra pairs (77.2 BLEU). Sonnet scores 65.7 BLEU, and letting it choose among ACE&#x27;s own candidate sentences lifts it to 69.6, while a pooled judge that keeps Opus&#x27;s own sentence among the candidates adds 0.6 points (77.0 BLEU). In the parsing direction, however, the models fall far behind ACE: asked for the MRS of the same sentences, they reach 57.2 (Sonnet) and 65.5 (Opus) F$_1$ on the graph&#x27;s predicates and arguments against 91.0 for ACE, and exact-match the gold on about 1\% of sentences. We characterize the failure modes for the parsing tasks, and conclude that a generation score alone does not show that models understand formal semantic representations.</p>]]></description>
  </item>
  <item>
    <title>Self-Play Pretraining with Zero Data</title>
    <link>http://arxiv.org/abs/2609.30063v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30063v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:23:01 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine</p><p>Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model&#x27;s behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Starting from random initialization, two models learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner&#x27;s capabilities, yielding an adaptive curriculum. A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training.</p><p><em>Comment: AC, KD, and MYL contributed equally; authors are listed alphabetically</em></p>]]></description>
  </item>
  <item>
    <title>KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization</title>
    <link>http://arxiv.org/abs/2609.30059v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30059v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:17:52 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Aheli Poddar, Sanskar Prasad, Arindam Samanta, Subha Chakraborty, Vishal Goyal, Rohit Singh Rathaur</p><p>Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kernels, yet treat compiled models as black boxes, generally optimizing individual standalone kernels without respecting the compiler&#x27;s structural decisions or verifying the model end-to-end. We present KernelOPT, a multi-agent system that treats compiled models as structured artifacts. It preserves vendor library calls (cuBLAS, cuDNN) and exclusively targets generated Triton sub-kernels using five profiling-guided LLM agents. A four-gate verification cascade of static validation, multi-seed correctness, model-level float64-fallback verification, and performance gating filters candidates during optimization and verifies the re-stitched model end-to-end. If no candidate passes all four gates, the system preserves the compiler baseline. The system accepts PyTorch nn.Modules, standalone Triton kernels, and Helion kernels. Evaluated on 250 KernelBench problems, KernelOPT achieves geometric mean speedups over \texttt{torch.compile} of 1.40$\times$ (Level 1: 51/100), 1.15$\times$ (Level 2: 31/100), and 1.07$\times$ (Level 3: 12/50) across all problems.</p>]]></description>
  </item>
  <item>
    <title>Can Labor Markets Function in the Age of AI? The Evaluation Bottleneck in Hiring</title>
    <link>http://arxiv.org/abs/2609.30058v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30058v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:16:05 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Itai Ashlagi, Ramesh Johari, Jon Kleinberg, Anushka Murthy</p><p>AI-assisted job-search tools have become increasingly popular by making it easier to find and apply to jobs. But by making it easier for applicants to generate and tailor application materials, they can also reduce how informative those materials are about applicant fit. We study this tradeoff in a hiring market where applicants differ in experience and latent match quality and firms use noisy application materials to decide whom to screen. We ask how AI affects downstream screening and hiring, and which applicants are most adversely affected. As application materials become less informative, a Bayesian firm rationally relies more heavily on coarse observables such as prior experience. Among the four applicant types defined by experience and compatibility for the job, inexperienced-compatible applicants are the most exposed: they lack observable experience and lose the individualized information that could distinguish them from other inexperienced candidates. When screening is costly, these changes can also generate inefficient screening failures in which firms screen no applicants or screen only experienced applicants. We then show that multistage hiring can arise as an endogenous firm response: a relatively inexpensive intermediate assessment allows firms to acquire new evidence of fit before costly full screening. This can restore screening opportunities that disappear under one-stage hiring and give inexperienced-compatible applicants a path to screening. Our results show how AI can shift the central friction in hiring from submitting applications to obtaining credible evaluation, creating entry barriers for high-fit workers without prior experience. Multistage hiring can endogenously arise in response, restoring evaluation opportunities that would otherwise disappear and helping preserve market functioning.</p>]]></description>
  </item>
  <item>
    <title>M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis</title>
    <link>http://arxiv.org/abs/2609.30056v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30056v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:14:00 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno</p><p>Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator. We show that, after camera projection, frozen LiDAR and image features exhibit substantial shared spatial structure, providing a natural cross-modal representation. M3GD conditions generation on LiDAR through this structure: it combines explicit geometry statistics with learned point-cloud descriptors into view-aligned packets on the image-latent grid, injected through a lightweight residual adapter into a multi-view flow-matching generator whose latent space, decoders, and training objective remain intact. On the GrandTour dataset, M3GD improves target-view RGB and depth synthesis over an image-only version of the same backbone. Ablations show that the gains come from pixel-aligned LiDAR content and that target-view LiDAR acts as a geometric query linking the requested view to source observations. Deployment on a ground robot demonstrates practical real-world operation, with a configurable quality--cost trade-off controlled by the number of Euler integration steps.</p>]]></description>
  </item>
  <item>
    <title>Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge</title>
    <link>http://arxiv.org/abs/2609.30055v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30055v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:12:27 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary</p><p>In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company&#x27;s data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them.
  We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a purchase because of timing. On a recorded call, the customer blames an outage.
  For each generated company, code fills each template and computes an exact answer without a language model. We evaluate 12 agents. Each pairs a model with an agent program, which connects it to the company&#x27;s systems.
  The best agent answers 18 of its 24 attempts, three per question, correctly. Four of the six models answer at most 6 of 24 with any program. The hardest questions require picking one of several similar records, such as which of three renewal offers a customer signed. All agents together answered two such questions correctly in only 1 of 84 attempts.</p><p><em>Comment: 9 pages</em></p>]]></description>
  </item>
  <item>
    <title>SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback</title>
    <link>http://arxiv.org/abs/2609.30054v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30054v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:12:14 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Chenxi Li, Wenxuan Zeng, Yun Luo, Fangchen Yu, Peng Ye, Yu Cheng, Jun Zhang</p><p>Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data. However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging. To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. The framework combines scientific library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as computational workflow cues. Guided by these cues, we adopt LLMs to generate scientifically grounded problem statements, reference solutions, and tests, with failed generations iteratively repaired using execution feedback. By combining structured workflow composition with verification and quality review, SciWalker enables scalable task generation while promoting scientific grounding, computational diversity, and executability. Using this framework, we construct 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains. To evaluate their training utility, we conduct reinforcement learning on Qwen3.5-9B using the GSPO algorithm. This training improves SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2%, with gains across scientific code generation, code repair, and reasoning benchmarks. The code for SciWalker is available at https://github.com/lichenx1/SciWalker.</p>]]></description>
  </item>
  <item>
    <title>NNV3: Expanding Neural Network Verification to New Architectures and Domains</title>
    <link>http://arxiv.org/abs/2609.30050v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30050v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:11:25 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Anne M. Tumlin, Samuel Sasaki, Ben Wooding, Diego Manzanas Lopez, Muhammad Usama Zubair, Navid Hashemi, Hongchao Zhang, Waseem Abbas, Ipek Oguz, Meiyi Ma, Taylor T. Johnson</p><p>We present NNV3, the latest version of the Neural Network Verification (NNV) tool, a MATLAB framework for formal verification of deep learning models and learning-enabled cyber-physical systems. Building on the set-based reachability foundation of NNV 1.0 (FFNNs, CNNs, NNCS) and NNV 2.0 (RNNs, SSNNs, neural ODEs), NNV3 introduces new members of the Star-set family: ModelStar for verifying networks under weight perturbation, VolumeStar for video and 3D volumetric inputs, and GraphStar for graph neural networks. A conformal-inference-based probabilistic reachability mode complements sound analysis for problems where deterministic verification is intractable, while FairNNV certifies counterfactual and individual fairness properties over continuous input regions. NNV3 introduces new benchmarks for malware detection, graph-based power-system models, medical imaging, variable-length time series data, and action recognition. NNV3 also incorporates tutorials and developer guides through a unified documentation site. This paper details these major updates, demonstrating NNV&#x27;s maturation into a comprehensive, robust, and accessible verification tool for a diverse range of AI systems.</p>]]></description>
  </item>
  <item>
    <title>Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models</title>
    <link>http://arxiv.org/abs/2609.30048v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30048v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:11:21 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ehsan Barkhordar, Surendrabikram Thapa</p><p>If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act as evaluators in four tasks: picking their own solution from a pair, judging whether a single solution is their own, identifying which of two solutions a named model wrote, and judging quality blind. In the single-solution task, balanced accuracy is 49-58% for all 15 model-benchmark combinations, while raw accuracy (38-67%) mostly reflects how readily a model claims authorship. In the pairwise task, accuracy across 14 evaluator-opponent combinations correlates at r=0.93 with how often the evaluator&#x27;s solution is longer. Attribution to a named model succeeds on some pairs and is consistently inverted on others. A rule-based normalization that strips docstrings, comments, type hints, and local names preserves Pass@1 and leaves ten of twelve re-tested results at chance; the other two follow a length difference it leaves, although a trained classifier still separates most normalized pairs. Claude Haiku&#x27;s self-preference also disappears. We recommend reporting balanced accuracy, heuristic baselines, and label consistency.</p><p><em>Comment: 18 pages, 1 figure. Code and data: https://github.com/ebarkhordar/llm-collusion</em></p>]]></description>
  </item>
  <item>
    <title>From Processing to Functionality: Engineering Accessible Material States in Cu-Embedded SiO$_x$ Memristive Devices</title>
    <link>http://arxiv.org/abs/2609.30047v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30047v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:10:55 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Tobias Gergs, Rouven Lamprecht, Sahitya Yarragolla, Ole Gronenberg, Luca Vialetto, Hermann Kohlstedt, Thomas Mussenbrock, Jan Trieschmann</p><p>Resistive switching in oxide-based devices is widely governed by stochastic defect processes, yet a predictive link between fabrication conditions and functional behavior remains elusive. Here, we establish a multiscale framework connecting plasma-defined deposition conditions to macroscopic device functionality in sputtered SiO$_x$/Cu/SiO$_x$-based systems. By combining large-scale statistical analysis of more than 50,000 experimentally characterized devices with physics-based plasma and atomistic simulations, we show that device behavior does not emerge from deterministic process-to-performance mappings, but from a probabilistic cascade spanning defect formation, defect-state evolution, and functional-regime emergence. Data-driven clustering reveals a continuous functional state space composed of operational switching types, while inverse modeling identifies the reconstructed oxygen-vacancy density as an effective latent descriptor capturing the combined influence of structural disorder and defect topology. This latent descriptor is strongly coupled to both Cu redistribution and electrical response, linking otherwise hidden material properties to observable device characteristics. Furthermore, macroscopic switching behavior is argued to arise from ensemble integration across spatially heterogeneous subdomains, providing a physical explanation for the pronounced variability of large-area devices. These findings shift the perspective from deterministic defect engineering toward probabilistic defect-state design and establish a physically grounded framework for understanding and controlling functional variability in such oxide-based systems, such as memristive or resistive-switching devices.</p>]]></description>
  </item>
  <item>
    <title>ConPro: Contrast Projection Pretraining for Label-Efficient Vessel Segmentation in DSA Sequences</title>
    <link>http://arxiv.org/abs/2609.30043v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30043v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:09:30 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xinge Guo, Yuanhao Wang, Liqi Shu, Yang Liu, Min Xu</p><p>Dense vessel annotation in digital subtraction angiography (DSA) is labor-intensive, yet every unlabeled sequence records how contrast passes through the vessels. Semi-supervised methods take their targets from the current model, and generic self-supervised pretexts reconstruct static appearance, so this signal goes unused. We propose ConPro, a self-supervised pretraining scheme whose target is a contrast projection, the normalized drop of every pixel below its temporal median over the sequence. On DIAS and DSCA, with 10%, 20% and 50% of the training cases labeled, ConPro improves on training from scratch at every label fraction and is the best of the compared methods on DSCA at 20% and 50% labels. Controlled comparisons show that the gain comes from the target. A temporal-median target with the same input, loss and budget stays at scratch level, and using the projection directly instead of learning it, as an input channel or a pseudo-label, helps little or hurts. ConPro provides pretrained weights without changing the segmentation architecture, so it combines with semi-supervised training, and UniMatch, the strongest baseline, gains 0.5 to 2.0 Dice and 0.9 to 2.3 clDice at every label fraction when started from ConPro weights, reaching 75.4 Dice on DIAS and 81.3 on DSCA.</p><p><em>Comment: 5 pages, 4 figures, 2 tables. Submitted to IEEE ICASSP 2027</em></p>]]></description>
  </item>
  <item>
    <title>AERIAL: Adversarial Evaluation of Robustness in Accuracy-Preserving Low-Precision EEG Decoders</title>
    <link>http://arxiv.org/abs/2609.30037v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30037v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:07:01 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Saim Rehman, Muhammad Shafique</p><p>Deployment-oriented compression is attractive for resource-constrained brain--computer interfaces (BCIs), but whether it changes adversarial vulnerability remains unclear. On BCI Competition IV-2a, we compare 32-bit floating-point (FP32) EEGNet and ShallowConvNet models with global magnitude pruning and simulated INT8 post training quantization (PTQ) and quantization-aware training (QAT) across nine subjects and three seeds. Simulation provides differentiable quantize--dequantize models for white-box attacks and gradient analysis, while native TensorRT deployment is used for validation. Accuracy-preserving compression does not improve direct robustness: at $ε=0.005$, EEGNet PGD accuracy remains 22--24\% across FP32, 50\% pruning (P50), PTQ, and QAT. However, P50 reduces bidirectional transfer efficiency to 0.963/0.928 (FP32$\rightarrow$P50/P50$\rightarrow$FP32), versus 0.994/0.997 for PTQ; the same trend holds for ShallowConvNet. Gradient alignment shows a corresponding separation, while native PTQ agrees with simulated clean/adversarial predictions in 95--98\% of cases. These results show that direct robustness, adversarial transfer, and deployment efficiency are distinct properties of compressed EEG decoders.</p><p><em>Comment: Submitted to IEEE ICASSP 2027, 5 pages</em></p>]]></description>
  </item>
  <item>
    <title>Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think</title>
    <link>http://arxiv.org/abs/2609.30036v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30036v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:05:02 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xvyuan Liu, Jianjie Fang, Chen Gao, Yong Li</p><p>Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require actions that initially move away from it. With frozen LeWM models, intermediate targets substantially improve action synthesis and recorded-action ranking on Cube, PushT, Reacher, and TwoRoom. Learned targets and targets drawn from observed experience both produce these gains. We introduce Anchored Planning, which retrieves a recorded segment whose start and end resemble the current and goal observations, then aims at an observation shortly after its start. The frozen model scores actions toward this target from the current state. Without additional training, planning toward observed targets outperforms the released LeWM planner on every task in our long-range evaluation. Additional final-goal search falls short of the same gains. Lower successor-prediction error need not translate into better control. Success also depends on how far ahead the target is placed and on shrinking the retrieval span as execution advances. Changing only the target lets the same frozen model and planner reach goals that final-goal scoring misses.</p>]]></description>
  </item>
  <item>
    <title>Artificial Societies Benchmark: A Validation Framework for Synthetic Research</title>
    <link>http://arxiv.org/abs/2609.30030v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30030v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:03:28 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He</p><p>A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.</p><p><em>Comment: 36 pages, 9 figures, 9 tables</em></p>]]></description>
  </item>
  <item>
    <title>How does Adversarial Influence Scale in Multi-Agent Systems?</title>
    <link>http://arxiv.org/abs/2609.30028v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30028v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:02:40 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Addison J. Wu, Jasin Cekinmez, Michel Liao, Karthik Narasimhan, Thomas L. Griffiths</p><p>Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group.</p>]]></description>
  </item>
  <item>
    <title>Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark</title>
    <link>http://arxiv.org/abs/2609.30027v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30027v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:00:56 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Christine Park, Valerie Chen, Tim Dettmers</p><p>Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient&#x27;s longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.</p><p><em>Comment: 29 pages, 2 figures, 12 tables</em></p>]]></description>
  </item>
  <item>
    <title>Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models</title>
    <link>http://arxiv.org/abs/2609.30026v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30026v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:00:51 +0000</pubDate>
    <category>Human Body</category>
    <description><![CDATA[<p><strong>Authors:</strong> Abu Bakar, Abdullah Aftab, Amir Hamza</p><p>Detecting which holds a climber uses, and when, underpins automated scoring, movement analysis, and assistive systems for sport climbing. Existing approaches train task-specific models or repurpose 2D pose estimators whose hand keypoint sits at the wrist and foot keypoint at the ankle i.e. offset from the fingertips and toes that actually contact the holds, and whose hands are occluded in roughly half of all frames. We show that a frozen, off-the-shelf pose foundation model is sufficient: using the fingertip and toe keypoints of Sapiens, a per-frame proximity test against the annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, we detect hold usage without any climbing-specific training. On the The Way Up dataset (22 videos, 10 athletes, two routes), our method reaches an event F_1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, and performs best on footholds (F_1,89.8% overall, 96.6% held-out). Under an identical protocol it exceeds our reproductions of the YOLOv8-pose and ViTPose pipelines at every temporal threshold, with the margin widening under strict timing. An ablation shows that two intuitively helpful additions---dense foundation-feature change gating and body-part segmentation---both hurt, arguing that a minimal, keypoint-only design is the right one for this task. Finally, standard coaching statistics computed from our automatic predictions track ground truth closely (Pearson r=1.00 for climb time, 0.94 for pace), turning ordinary single-camera video into reliable performance metrics with no instrumentation.</p><p><em>Comment: Accepted at AI2ML Conference 2026 (2nd International Conference on Advancement &amp; Innovation in Artificial Intelligence and Machine Learning)</em></p>]]></description>
  </item>
  <item>
    <title>Body-Grounded Replanning for Physically Adaptive Manipulation</title>
    <link>http://arxiv.org/abs/2609.30024v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30024v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:00:09 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Namiko Saito, Hiroshi Kera</p><p>Manipulation requires not only reasoning about the external environment, but also about the robot&#x27;s physical condition. A strategy may remain geometrically feasible while becoming physically unsuitable due to increased joint load or limited mobility, yet internal physical state is typically used only for low-level control. We propose body-grounded high-level replanning, which uses internal physical state to adapt manipulation strategies during execution. Body-state events trigger strategy replanning, and an LLM interprets the underlying joint-level state, recent execution statistics, and execution history to select a context-dependent alternative, while leaving the task objective and low-level controller unchanged. We evaluate the framework on a reaching task under controlled load and asymmetric mobility constraints in simulation and on a real robot. Our experiments show that body-grounded replanning maintains high task success while reducing physical effort and enabling more efficient strategy adaptation. Additional contact-rich manipulation experiments demonstrate the applicability of the same replanning interface beyond reaching. These results show that internal physical state can inform not only low-level control, but also high-level decisions about how a manipulation task should be performed.</p>]]></description>
  </item>
  <item>
    <title>Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation</title>
    <link>http://arxiv.org/abs/2609.30023v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30023v1</guid>
    <pubDate>Thu, 24 Sep 2026 16:00:02 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv</p><p>Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.</p>]]></description>
  </item>
  <item>
    <title>Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits</title>
    <link>http://arxiv.org/abs/2609.30017v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30017v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:55:24 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Michael Jerge, Suman Jana</p><p>Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tree. This structure arises naturally from autoregressive generation: every prefix defines a node, and its continuations form a subtree below it. Internal nodes of the tree provide cheap but biased estimates of a region&#x27;s value, while leaf evaluations are expensive but accurate. Hierarchical bandit methods can exploit this structure, but typically require a specific smoothness schedule to be specified in advance, even though real objectives are often only piecewise smooth and their optima may lie near sharp boundaries. We introduce CANOPY, a multi-fidelity tree bandit that learns where the smoothness prior is valid rather than assuming it globally. CANOPY uses cheap random-path probes to construct an online certificate of local aggregation bias, then directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation. We prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense. Across routing, top-$k$ identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including $2.9\times$ higher top-10 recall on a 1000-model pool, $1.6\times$ more SWE-bench Verified issues resolved than best-of-$N$, and $3.6\times$ lower median time-to-first-token with prefix caching.</p>]]></description>
  </item>
  <item>
    <title>Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases</title>
    <link>http://arxiv.org/abs/2609.30012v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30012v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:51:17 +0000</pubDate>
    <category>Generative Search Engines</category>
    <description><![CDATA[<p><strong>Authors:</strong> Tapan Parikh</p><p>Language models advise people, keep them company, and write software while they sleep. Measuring what they do is hard: behavior has to be sampled repeatedly across models, prompts and releases, most of it lives in unstructured text that has to be coded before it can be counted, and the result has to be legible and rigorous enough to meaningfully compare models and vendors. To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior. Each study is a frozen, public stimulus run identically on a cross-vendor panel, at a few dollars per model or less. Each reads its transcripts one of three ways, chosen by how much interpretation the behavior needs: exact match on a clamped reply, a codebook applied by LLM judges whose agreement with a human coder is reported per code, and an instrumented environment that records what an agent did independently of what it said. Run across four years of model releases from both frontier and open-source labs, these instruments find four things. Convergence: asked to pick a word, 27 of 44 models answer serendipity at least once in four tries. Resistance: a trailing &quot;right?&quot; moves endorsement by up to 32 points, and the sign flips from sycophantic to resistant as generations advance, keyed to the tag&#x27;s surface form. House: whether a model holds a position under pressure tracks its generation, and how it holds tracks the lab that built it. Account: told to do something the documentation in their repository contradicts, some coding agents never went along silently and others always did, and the same model can change with the harness it runs in. Re-run on every release, batteries like these track how behavior is changing across vendors and over time.</p><p><em>Comment: 6 pages. Code and data: https://github.com/tap2k/modelun</em></p>]]></description>
  </item>
  <item>
    <title>Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation</title>
    <link>http://arxiv.org/abs/2609.30009v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30009v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:48:59 +0000</pubDate>
    <category>Generative Search Engines</category>
    <description><![CDATA[<p><strong>Authors:</strong> Tobias Deußer, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa</p><p>Financial institutions operate under dense, frequently amended rulebooks, and answering a compliance question correctly requires not only fluency but verifiable grounding in the authoritative text. Large language models are attractive for this task, yet the models that firms can realistically deploy on-premise are compact ones, and compact models hallucinate obligations. We study whether a carefully domain-adapted retrieval-augmented generation pipeline closes that gap. Our retriever is built in three stages on top of LegalBERT: entailment tuning that recasts question--passage matching as premise--hypothesis reconstruction, contrastive tuning with in-batch negatives, and score-level fusion with BM25. Our generator is a compact model (2B--12B parameters) served under 4-bit quantization, either prompted or adapted with retrieval-aware fine-tuning (RAFT) through LoRA. On ObliQA, a question-answering benchmark built from the Abu Dhabi Global Market rulebooks, the staged retriever raises Recall@10 from 0.256 to 0.774 and outperforms BM25 (0.678) and E5-large-v2 (0.758), the strongest general-purpose dense encoder we tested. RAFT-LoRA then improves the composite RePASs answer-quality score for every model we could adapt, with the largest gain on the weakest one. However, the adapted models do not transfer to Australian case-law questions, and a closed-book model that receives no passages at all scores within 0.011 RePASs of the full pipeline while producing answers that cite nothing and misstate obligations. The retrieval gain is therefore measured directly, the generation gain is a gain in RePASs rather than demonstrated grounding, and grounding itself requires an evaluation protocol that RePASs does not provide.</p><p><em>Comment: Currently under review</em></p>]]></description>
  </item>
  <item>
    <title>VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching</title>
    <link>http://arxiv.org/abs/2609.30005v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30005v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:48:23 +0000</pubDate>
    <category>Generative Search Engines</category>
    <description><![CDATA[<p><strong>Authors:</strong> Minh Hoang, Thai Le</p><p>Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos. To our knowledge, it is the first large-scale Vietnamese corpus to jointly provide transcripts, consistent speaker identities, five dialect groups, and naturally occurring Vietnamese--English code-switching, which constitutes nearly half of the corpus by duration. We further create over 3.1K hours of spoof speech with four open-source and commercial synthesis systems. Every spoof is conditioned on a verified speaker reference and paired with a transcript- and speaker-matched bona fide utterance, enabling unique controlled evaluation with reduced lexical and identity confounds. Zero-shot evaluation of five pretrained multilingual detectors reveals striking brittleness: EER greatly varies across detector--generator pairings, while recent multilingual detector DFA-1B degrades from 16.3% to 33.6% as speaker similarity increases. Dialect-stratified results expose further model-dependent disparities. By unifying natural linguistic diversity with controlled spoof generation, VietPrism provides a challenging foundation for Vietnamese speech modeling and trustworthy audio-deepfake detection.</p><p><em>Comment: Preprint for ICASSP 2027 submission</em></p>]]></description>
  </item>
  <item>
    <title>Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems</title>
    <link>http://arxiv.org/abs/2609.30001v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.30001v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:45:07 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao, Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li, Fan Wu, Tao Wang, Zibo Zhao, Xiangyu Wu, An Liu, Fei Pan, Peng Jiang, Chen Yang, Zhaojie Liu, Wenwu Ou</p><p>Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX&#x27;s model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates.</p><p><em>Comment: Technical report. 37 pages, 11 figures, 13 tables, including appendices</em></p>]]></description>
  </item>
  <item>
    <title>GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMS</title>
    <link>http://arxiv.org/abs/2609.29999v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29999v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:44:33 +0000</pubDate>
    <category>End-to-End AD</category>
    <description><![CDATA[<p><strong>Authors:</strong> Saim Rehman, Muhammad Shafique</p><p>Post-training quantization of vision--language models (VLMs) is typically assessed through aggregate task accuracy and memory savings, but preserving a headline score does not guarantee preservation of visual grounding behavior. We present GHOST-Q, a cross-precision controlled evaluation of three 8B VLM families under FP16, INT8, and NF4 across utility and hallucination-sensitive benchmarks. Rather than comparing only aggregate accuracy, we pair FP16 and quantized predictions item by-item to quantify how compression redistributes grounding successes and failures. Five of six quantized variants preserve MMStar accuracy within $\pm2$ percentage points, yet 10 of 36 paired effects remain significant after false-discovery-rate correction, nine on hallucination-sensitive conditions. Same-device A100 profiling further demonstrates that substantial memory reduction does not necessarily mean lower inference latency. Finally, an open-ended AMBER audit reveals strong generation budget censoring whose severity varies by architecture and precision. These results show that quantized VLMs should be evaluated jointly for aggregate utility, grounding reliability, generation behavior, and realized deployment efficiency.</p><p><em>Comment: Submitted to IEEE ICASSP 2027, 5 pages</em></p>]]></description>
  </item>
  <item>
    <title>Guardrails or Roadblocks? Effects of Pedagogical Style and Context Awareness in AI Teaching Assistants for Programming</title>
    <link>http://arxiv.org/abs/2609.29995v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29995v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:43:49 +0000</pubDate>
    <category>LLM Agents with Reinforcement Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Madeleine Eastwood, Harshith Narne, Joseph Hilby, Paul Denny, Ashish Aggarwal, Amanpreet Kapoor</p><p>AI teaching assistants (AI TAs) backed by large language models (LLMs) and pedagogical guardrails are increasingly being integrated into programming courses, providing students with scalable access to hints, conceptual explanations, and code-level feedback. However, guardrails may also create friction. If students feel that the support provided is overly restrictive or poorly contextualized to their current progress, they may bypass approved tools for general-purpose LLMs. To investigate how AI TA design affects students&#x27; learning experiences, we conducted a randomized controlled trial with 132 students in an introductory programming course. Students completed three tasks related to code-writing and debugging and were randomly assigned to one of four AI TAs varied across two dimensions: pedagogical guidance style (Socratic vs. Direct instruction) and context awareness (no context vs. full context of the problem and student solution). We examined students&#x27; perceptions, interaction behaviors, and evidence of post-task comprehension. Students rated the Socratic AI TA with full context least favorably, reporting significantly lower perceived support for task completion. Descriptively, this condition also showed the highest observed interaction stress, the highest rate of external LLM use, and the lowest proportion of post-task explanations demonstrating full comprehension, though these differences were not statistically significant. These findings suggest that guardrailed AI TAs are not automatically better for learning. Instead, their effectiveness depends on how pedagogical guidance and contextual awareness are balanced in ways that students experience as useful, supportive, and worth continuing to use.</p>]]></description>
  </item>
  <item>
    <title>Let Training Guide Selection: Online Synthetic Data Filtering via Real-Anchored Utility</title>
    <link>http://arxiv.org/abs/2609.29988v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29988v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:38:45 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yanran Wu, Sana Lakdawala, Renzo Tassara Miller, Chongyang Bai, Sharath Ciddu, Shivendra Pratap Singh, Kungang Li, Sandeep Pandey, Chunwei Liu</p><p>Synthetic data can scale training supervision when real-world data are limited, but noise and distribution mismatch can reduce its value. Existing synthetic data selection methods often emphasize fidelity or diversity rather than the learner&#x27;s evolving needs. We propose FROST, an online framework that estimates synthetic-data utility through gradient feedback anchored in real training data. It calibrates batch utility against recent history to determine when filtering is needed and filters samples only in out-of-band batches to determine what to retain, without an external verifier or held-out validation set. Experiments on two public benchmarks for image classification and LLM fine-tuning for text-to-SQL show that FROST filters out around 20--30% of the synthetic data while improving real-task performance compared with training on the full synthetic data pool. We further apply FROST during training in a large-scale industrial ads re-ranking system, achieving significant performance gains over a highly optimized production baseline, demonstrating its effectiveness and generalizability.</p><p><em>Comment: 21 pages, 6 figures</em></p>]]></description>
  </item>
  <item>
    <title>OceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning</title>
    <link>http://arxiv.org/abs/2609.29985v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29985v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:35:08 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Haoran Wang, Shaoyu Cai, Adrian Azzarelli, Zhuodong Jiang, Guoxi Huang, Eng Tat Khoo, Brett Seymour, Fan Zhang, David Bull, Nantheera Anantrasirichai</p><p>Underwater 3D reconstruction is critical for marine exploration, ecological monitoring, and subsea infrastructure inspection, yet remains challenging at large scale due to light attenuation, scattering, and limited capture coverage. While 3D Gaussian Splatting (3DGS) enables high-quality real-time rendering, its application to large underwater scenes is constrained by high memory consumption and inefficient optimization over extensive areas. We propose OceanXL, a fast and scalable 3DGS-based framework for large-scale underwater reconstruction. OceanXL adopts a divide-and-conquer strategy, partitioning scenes into spatially coherent blocks to enable efficient optimization while preserving global geometric consistency. We further introduce an adaptive pruning scheme tailored to underwater conditions that removes redundant primitives, producing compact representations without sacrificing visual fidelity. Together, these components improve training efficiency and rendering performance for large scenes. We also introduce a large-scale underwater dataset covering diverse marine environments. Experiments on five large-scale scenes demonstrate favorable scalability, compactness, and efficiency--quality trade-offs over large-scene baselines. Controlled comparisons on the small-scale SeaThru-NeRF dataset further show competitive reconstruction quality with substantially smaller model sizes than underwater-specific methods.</p><p><em>Comment: SIGGRAPH ASIA 2026</em></p>]]></description>
  </item>
  <item>
    <title>From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation</title>
    <link>http://arxiv.org/abs/2609.29983v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29983v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:34:11 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Mengdan Zhu, Yufan Zhao, Yao Zhao, Sophie Di, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao</p><p>Semantic IDs (SIDs) encode each catalog item as a short token sequence, enabling generative recommenders to predict the next item autoregressively. Reasoning-enhanced variants, an increasingly common extension, first generate a textual trace and then decode a next-item SID by beam search. Such recommenders are commonly trained with group-relative policy optimization under an exact-match SID reward, which is sparse in large catalogs. Two failure modes follow. When all rollouts in a group miss the target, the group yields zero advantage and no learning signal. Rollouts sharing the same SID reward receive identical advantages, however much their traces differ. In both cases the reward reflects only the decoded SID, never the reasoning that produced it. This creates a credit-assignment gap.
  We address this gap with retrieval-grounded query attribution. Each trace is structured into a history summary, a set of interest hypotheses, and a final SID. A frozen retriever executes every hypothesis as a catalog query, so that each hypothesis becomes independently verifiable rather than judged only through the final SID. A rollout is rewarded when any of its queries retrieves the target within the \mbox{top-$K$}, and per-query hit indicators localize that reward to individual hypotheses. Credit is thus assigned at the span level: only hypotheses that individually hit receive positive retrieval advantage, while the retrieval channel never updates the final SID span. Rollouts that share a SID reward can therefore receive different updates. Across experiments on three Amazon Reviews datasets, this yields consistent improvements in SID recommendation. On Video Games, an oracle analysis further reveals the potential of interest-conditioned SID decoding: selecting the target-relevant query among generated interests improves both recall and ranking.</p>]]></description>
  </item>
  <item>
    <title>Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles</title>
    <link>http://arxiv.org/abs/2609.29974v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29974v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:26:01 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ali Haghpanah Jahromi, Mohammad Taheri</p><p>Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Expert Ensemble (GeoACE), a five-expert framework that combines a common anchor-correction estimator with complementary overlap-aware and outcome-guided geometries. Its task-level ensemble weights are learned only from internal validation predictions, frozen before test evaluation, and then applied to experts refitted on the complete development sample. The fifth expert, O-Phi-ACE, constructs an outcome-free, overlap-aware statistical projection from covariates and treatment assignment and replaces the anchor input with this lower-dimensional geometry. We evaluate GeoACE against 11 comparators on eight benchmark protocols. Adding O-Phi-ACE reduced mean sqrt(PEHE) relative to the four-expert ensemble on all seven benchmarks with individual-effect truth, winning 998 of 1,225 paired tasks; the change on JOBS policy risk was negligible. The five-expert ensemble ranked first on IHDP100, IHDPA, and IHDPB and second on NEWS, differing from the NEWS leader by 0.13%. Across the seven sqrt(PEHE) benchmarks it obtained the lowest observed average rank (3.714), although the omnibus Friedman and Iman-Davenport tests were not significant (p=0.328 and p=0.330). Using the same five frozen experts, inverse-DR weighting was consistently better than winner-take-all selection, convex DR fitting, R-stacking, and causal Q-aggregation in benchmark-balanced analyses, but was statistically indistinguishable from equal weighting and DR ridge shrinkage. The evidence therefore supports geometry-diverse expert libraries and leakage-free aggregation as a robustness strategy, not universal superiority of either GeoACE or one weighting rule.</p><p><em>Comment: 31 pages, 3 figures, 8 benchmark protocols. Supplementary material is included as an ancillary file</em></p>]]></description>
  </item>
  <item>
    <title>Learning Better Reasoning for Generative Recommendation with Semantic IDs</title>
    <link>http://arxiv.org/abs/2609.29973v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29973v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:24:54 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Mengdan Zhu, Yufan Zhao, Sophie Di, Yao Zhao, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao</p><p>Generative recommendation reformulates item retrieval as sequence generation, allowing a unified model to directly generate the next item from a user&#x27;s interaction history. Semantic IDs further make this paradigm effective and scalable by representing each item as discrete codes, enabling knowledge sharing among semantically related items. Recent studies introduce explicit reasoning before Semantic-ID generation, helping models summarize user interests and infer possible preference transitions. However, reasoning is not inherently beneficial: Inaccurate or uninformative reasoning may mislead subsequent item generation and ultimately degrade recommendation performance. This raises a central challenge: how can a recommender select and learn effective reasoning traces and progressively evolve toward better reasoning from its own generations? In this work, we propose Evo-Rec, a three-stage framework for learning better reasoning and further enhancing it through reinforcement learning. First, we align Semantic IDs with their textual and behavioral contexts, enabling the model to understand and generate item identifiers. Second, we sample multiple candidate reasoning traces and retain those that improve the prediction of the ground-truth item, providing a stronger reasoning initialization through supervised fine-tuning. Third, we further optimize the reasoning policy through reinforcement learning with catalog-constrained item generation and ranking-aware recommendation feedback. Experiments on three Amazon Review benchmarks show that Evo-Rec consistently outperforms discriminative, generative, and reasoning-enhanced recommenders across all evaluation metrics. These results demonstrate the effectiveness of our framework in learning better reasoning for SID-based generative recommendation.</p>]]></description>
  </item>
  <item>
    <title>World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal</title>
    <link>http://arxiv.org/abs/2609.29964v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29964v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:19:39 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li</p><p>General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.</p><p><em>Comment: Working in progress</em></p>]]></description>
  </item>
  <item>
    <title>ADATEX4D: adaptive texture capacity allocation for 4D gaussian splatting</title>
    <link>http://arxiv.org/abs/2609.29963v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29963v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:19:23 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> De Jiang, Peiqiang Wang, Kehong Yuan, Shaohua Ma</p><p>Textured Gaussians improve local appearance capacity, but assigning the same texture resolution to every primitive wastes storage on low-detail or weakly visible regions. We introduce AdaTex4D, an adaptive texture-capacity module for deformation-based 4D Gaussian Splatting. Each Gaussian carries packed RGBA triplanes whose two axes grow independently according to visibility normalized screen-space gradients and deformed local scales. Experiments on N3DV and PanopticSports show that AdaTex4D reduces texture storage by more than half while preserving reconstruction quality. Under fixed memory budgets, adaptive allocation also improves quality over uniform texture assignment and reduces overall model and peak memory. These results show that dynamic, anisotropic texture allocation provides a more efficient way to distribute local appearance capacity in 4D Gaussian representations.</p>]]></description>
  </item>
  <item>
    <title>A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning</title>
    <link>http://arxiv.org/abs/2609.29961v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29961v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:18:29 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ids van der Werf, Sergio Rozada, Antonio G. Marques</p><p>Many iterative algorithms rely on bootstrapping. A variable is updated using a second, frozen copy as a target, which is periodically replaced with the updated variable. Majorize-minimize and inexact proximal-point methods share this structure, as does temporal-difference (TD) learning. However, existing convergence guarantees for scenarios that combine sampled updates with targets refreshed only every $K$ steps rely on the specific structure of the update, such as linear approximation or gradient-based inner steps, and on uniformly bounded sampling error. We instead model the sampled update as a stochastic operator on the parameter space, which reduces the analysis to a contraction argument that needs no gradient structure and allows the sampling error to grow with the iterates. Within this framework, we derive a finite-time bound for i.i.d. samples and any target-update period $K$. We show that the iterates converge geometrically in root mean square to a ball around the fixed point, provided the sensitivity to the frozen target is smaller than the contraction slack of the inner map. Existing deterministic frozen-target contraction and stochastic-gradient-type bounds follow as special cases of our framework, and simulations of TD learning reproduce the predicted contraction rate and scaling of the error floor with the step size.</p><p><em>Comment: 5 pages, 1 figure</em></p>]]></description>
  </item>
  <item>
    <title>Beyond Average Safety: Chance-Constrained LLM Fine-tuning</title>
    <link>http://arxiv.org/abs/2609.29960v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29960v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:17:58 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Taha Entesari, Mahyar Fazlyab</p><p>Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures. We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold. Because the resulting empirical chance constraint contains a discontinuous indicator, we introduce a differentiable majorization of the violation rate, yielding a tractable conservative constraint. We then develop a constraint-aware gradient descent method that treats the majorized constraint as a safe set in parameter space and minimally modifies the fine-tuning direction to preserve feasibility. The resulting update admits a closed form and produces a tail-aware safety correction that emphasizes examples near or above the degradation threshold. We conduct an extensive set of experiments on harmful fine-tuning across three different tasks and three models and show that our approach consistently outperforms the baselines that exist in the literature. These results suggest that safety preservation in LLM fine-tuning is better viewed as a reliability-constrained optimization problem than as average-risk regularization.</p>]]></description>
  </item>
  <item>
    <title>Not All Confusion Is Equal: A Source-Aware Uncertainty Diagnosis for Fine-Grained Aircraft Detection</title>
    <link>http://arxiv.org/abs/2609.29959v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29959v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:17:34 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Hai Huang, Helmut Mayer</p><p>Fine-grained object detectors are commonly evaluated with confusion matrices, which show where the model is confused but not why, nor whether the confusion can be reduced. We argue that confusion can be attributed to distinct, separable sources, each quantitatively measurable, turning a passive measurement into actionable guidance. We present $A^2E^2$, a diagnostic tool that decomposes the sources of confusion along two axes, $\{$aleatoric, epistemic$\} \times \{$within-class, between-class$\}$, giving a $2\times2$ taxonomy that enumerates the source types. Each quadrant is measured by its own quantity, computed in one of three places (input geometry, output-space disagreement, and the bias-parameter posterior), so the two epistemic sources are separated by construction rather than by an empirical correlation. On fine-grained aircraft detection, the four quadrants become four named sources with their own remedy verdict: affinity (geometric similarity, irreducible from size alone), heterogeneity (geometrically heterogeneous sub-variants, pointing to re-labeling rather than more data), contested (an insufficiently trained but learnable boundary, improvable), and collapsed (a class starved of data, reducible). After attributing the confusion to a specific reducible source, we apply a targeted intervention and verify experimentally that it reduces the diagnosed source specifically while leaving the irreducible sources unchanged. $A^2E^2$ thus turns confusion measurement into a concrete, validatable and actionable &quot;diagnosis&quot; in which the same off-diagonal mass can carry opposite causes and opposite remedies. We also state this framework&#x27;s limits, including which sources are only partially identifiable on this specific dataset and why.</p><p><em>Comment: 23 pages, 4 figures</em></p>]]></description>
  </item>
  <item>
    <title>Multi-Dimensional Matching</title>
    <link>http://arxiv.org/abs/2609.29958v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29958v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:14:18 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Irene Aldridge</p><p>We study a matching mechanism where agents and objects are described by features rather than complete rankings. A single spectral projection reduces the problem to a one-dimensional sort, computable in O(N log N) time. We prove that on descaled features and preferences, our algorithm obtains the exact Nash Social Welfare (NSW) optimum within the projected space, with an unconditional utilitarian-welfare guarantee and a conditional NSW guarantee. The proposed mechanism is stable against exogenous noise but not strategy-proof; we provide an explicit profitable misreport. On an agentic AI shopping application, the diagnostics correctly anticipate both a success and a failure case. A 100-instance robustness study confirms the findings.</p><p><em>Comment: 20 pages</em></p>]]></description>
  </item>
  <item>
    <title>Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes</title>
    <link>http://arxiv.org/abs/2609.29952v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29952v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:10:29 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Rahul Khedar, Mayank Malhotra, Avinash Karn</p><p>Before a product or policy change ships, the question that matters is how people will react to it. Augur rehearses that reaction offline: it builds a typed knowledge graph from the change documents, populates a grounded persona market, simulates the interaction, and returns an auditable decision memo recommending one of five actions. We assemble Gold-50, fifty real product and policy episodes whose real-world outcome is known, adjudicated against the public record, and score the five-way release verdict against it.
  Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways. First, the prompt envelope alone can dominate the score: holding weights, cases and scorer fixed, one system -- a LoRA-SFT adapter on Qwen3-32B -- swings from 0% to 73%. Second, in a matched 2x2 ablation, defining the decision taxonomy in the prompt -- with no model change -- lifts every frontier model by +24 to +34pp; under the under-specified prompt, Qwen3-32B LoRA-SFT served offline beats all three frontier models (paired McNemar, Holm-corrected), and once the prompt is fair no significant difference from any of them is detected. Third, agreement with the distillation teacher rises without accuracy following, and the full pipeline amplifies a systematic &quot;over-doom&quot; bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67-90% of the concerns the public actually raised, and a pre-registered ablation locates its value -- largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is available from the authors.</p><p><em>Comment: 19 pages, 15 figures, 11 tables</em></p>]]></description>
  </item>
  <item>
    <title>Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking</title>
    <link>http://arxiv.org/abs/2609.29951v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29951v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:10:02 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Zhiyu Zhang, Yupeng Li</p><p>State tracking requires composing a sequence of updates, but accuracy alone does not reveal what a model has learned. We study neural networks trained to predict the running product of group elements. We identify quotient solutions in Transformers, where models recover the quotient class while predicting nearly uniformly among its members. The reciprocal of class size predicts partial accuracy without a fitted parameter, extending parity-based accounts to non-parity quotients. Our baseline Transformers&#x27; predictions change little under prefix reordering beyond the exact-tracking frontier. We prove that, for finite groups under uniform i.i.d. full-group inputs, optimal order-blind exact accuracy converges to the reciprocal of abelianization class size as prefix length grows, consistent with the observed abelianization plateaus. Sequential updates permit more: any partition into right cosets of a subgroup, normal or not, survives sequential updates. In our census of standard Transformers, every recovered coset partition comes from a normal subgroup, whereas parameter-matched recurrent networks pass through both normal and non-normal right-coset stages during training. On $A_5$, we identify low-dimensional subspaces of the recurrent state that encode non-normal cosets. In the three-dimensional cases, coset mean vectors form approximate dodecahedra, and swapping the state components in these subspaces transfers the donor&#x27;s coset state through a shared input suffix. Our results connect partial accuracy, learning stages, and internal computation through the subgroup cosets that models learn to track.</p><p><em>Comment: 69 pages including appendices; 9 pages of main text</em></p>]]></description>
  </item>
  <item>
    <title>ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation</title>
    <link>http://arxiv.org/abs/2609.29948v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29948v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:08:13 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Qingyu Wu, Zeyu Feng, Yongda Yu, Yuzhe Luo, Hua Cheng</p><p>Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as input. Clean victim continuations serve as pseudo-references: local search identifies prefixes that reduce continuation likelihood, and preference fitting on comparisons within the same instruction, followed by reward refinement, distills this signal into a generator. At deployment, the generator produces one prefix per request without further victim-side search. Across four instruction-tuned models and the complete splits of seven benign benchmarks, ENDOPROMPT yields a mean utility change of -26.8 percentage points; 27 of 28 cells are negative. Failure analysis reveals output expansion and prefix reuse; the controls do not establish a degradation advantage from request matching. Victim-derived supervision can reveal utility weaknesses without benchmark feedback or prescribed failure responses. The code will be released upon acceptance.</p>]]></description>
  </item>
  <item>
    <title>Error- and Prediction-Driven Motor Learning in the Cortico-Cerebellar Loop</title>
    <link>http://arxiv.org/abs/2609.29945v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29945v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:07:10 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ana Carolina Filipe, Rui Ponte Costa, Cláudia Soares</p><p>Robust control under delayed sensory feedback remains a key challenge in both robotics and neuroscience. Classical cerebellar models explain delay compensation through forward prediction but fail to account for fast online corrections and rapid adaptation observed in biological systems.
  We propose a cerebellum-inspired control framework that combines multiplexed predictive representations with internal feedback. By jointly encoding kinematic variables and task-relevant error signals, the model enables accurate online correction despite delayed feedback. Furthermore, incorporating feedback within the cerebellar loop significantly accelerates adaptation, reducing learning time by an order of magnitude.
  Our results show that single-signal predictions are insufficient under delay, while multiplexing and feedback together provide a unified mechanism for online control and rapid learning.</p>]]></description>
  </item>
  <item>
    <title>MF-SCBO : Multi-fidelity Scalable Constrained Bayesian Optimization</title>
    <link>http://arxiv.org/abs/2609.29941v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29941v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:05:39 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Lucas Palazzolo, Mickaël Binois, Laëtitia Giraldi</p><p>Many real-world optimization problems rely on expensive simulations or experiments, making the efficient use of available data essential. Multi-fidelity optimization of high-dimensional black-box functions subject to black-box constraints is increasingly relevant as the cost of objective evaluations continues to rise in applications such as machine learning, engineering, and control. To our knowledge, no existing method simultaneously addresses high-dimensionality, black-box constraints, an arbitrary number of fidelity levels, and non-nested sampling. In this work, we extend the Scalable Constrained Bayesian Optimization method to the multi-fidelity setting, resulting in the MF-SCBO method. The proposed approach is evaluated on standard benchmark functions as well as challenging problems. The experimental results demonstrate that MF-SCBO generally achieves better convergence than both the single-fidelity SCBO and the other multi-fidelity method considered in this high-dimensional and constrained settings.</p>]]></description>
  </item>
  <item>
    <title>Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration</title>
    <link>http://arxiv.org/abs/2609.29940v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29940v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:05:22 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Jiaqi Deng, Zonghan Wu, Zhan Heng, Xiaoshui Huang, Huan Huo, Guandong Xu</p><p>Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be improved by strengthening implicit visual grounding without directly supervising the reasoning process? Motivated by the functional specialization of attention heads, we investigate whether reasoning can be improved by guiding only the heads most responsive to visual evidence grounding. We propose Selective Probability Mass Concentration (sPMC), a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention. sPMC treats normalized attention over visual tokens as a spatial probability distribution and encourages the probability mass to be assigned to semantically relevant regions using segmentation-derived spatial priors. Adaptive Head Selection restricts this guidance to visually responsive heads while leaving the remaining heads unconstrained to preserve their complementary functions. Across 6 multimodal benchmark suites, sPMC achieves an average zero-shot improvement of 3% and gains of up to 11.3% across multiple MLLMs while regularizing only 3%-15% of their attention heads. These results demonstrate that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.</p>]]></description>
  </item>
  <item>
    <title>When Temporal Perturbations Act Like Sensor Biases: Label-Free Auditing of Wearable Activity Recognizers</title>
    <link>http://arxiv.org/abs/2609.29937v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29937v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:03:08 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Qingyu Wu, Yuan Wei, Renju Liu, Hua Cheng</p><p>Wearable human-activity recognition (HAR) models operate across sensors, subjects, and backbones, yet a smooth waveform may appear temporal while exploiting a persistent sensor offset primarily. We introduce SpectrumAudit, a label-sealed audit that fits a phase-randomized full-window stimulus on calibration windows from subjects held out from training and testing. After selection, it replays its exact DC projection and budget-constrained zero-mean residual on the same frozen victim without refitting. Across 27 victims from three datasets and three backbones, the selected waveforms cause 2.87-40.83-point three-phase robust accuracy losses. Under this replay budget, DC is more damaging than AC on 24/27 victims and recovers at least 90% of the full drop on 22/27; all 5 failures occur on WISDM. In a held-out UTD-MHAD check, the selected waveform causes 13.49-pp accuracy and 11.68-pp macro-F1 losses, versus -0.66 pp for matched random changes. The audit diagnoses offset versus zero-mean variation under a common peak-budget cap. The code will be released upon acceptance.</p>]]></description>
  </item>
  <item>
    <title>Path-specific harm decomposition: A partial identification framework</title>
    <link>http://arxiv.org/abs/2609.29938v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29938v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:03:08 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ruizi Yan, Dennis Frauen, Maresa Schröder, Stefan Feuerriegel</p><p>A central goal when designing treatment policies is often to &quot;do no harm&quot;, that is, to avoid interventions that improve average outcomes while worsening outcomes for some individuals. A widely used notion for harm is the fraction of negatively affected (FNA), defined as the probability that an intervention decreases an individual&#x27;s outcome. However, in many applications, treatments operate through mediators, and a single &quot;total&quot; FNA can obscure whether harm arises primarily through direct pathways or indirect (mediator-induced) pathways. In this work, we introduce a path-specific analogue of the FNA. For this, we disentangle total harm into direct and indirect harm in causal mediation settings. However, these quantities depend on joint distributions of potential outcomes that are not point-identified even in randomised controlled trials. As a remedy, we develop a novel partial identification framework for direct and indirect FNA. In our framework, we (i) derive sharp Makarov bounds for the FNA, and (ii) propose a semiparametrically efficient estimator with valid confidence intervals for these bounds under mild margin conditions. We demonstrate our framework across various numerical experiments. To the best of our knowledge, we are the first to study path-specific decomposition of causal harm and to develop an orthogonal inference framework for its analysis.</p>]]></description>
  </item>
  <item>
    <title>Robust Detection of LLM-Generated Text under Contamination</title>
    <link>http://arxiv.org/abs/2609.29935v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29935v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:01:38 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari</p><p>We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test&#x27;s worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5\%, clipping improves the log-likelihood--log-rank ratio (LRR) detector&#x27;s true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.</p>]]></description>
  </item>
  <item>
    <title>Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation</title>
    <link>http://arxiv.org/abs/2609.29934v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29934v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:01:24 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xun Huang, Shijia Zhao, Rongsheng Qu, Jiayuan Li, Xin Lu, Weixin Li, Chenglu Wen, Cheng Wang</p><p>Does progress on spatial reasoning benchmarks translate into better navigation? Existing benchmarks test isolated inferences from images or videos, with little connection to downstream navigation. Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how aligning spatial supervision with navigation goals, phases, and decision learning improves navigation. Guided by these findings, we build \textsc{Spatial-Nav-100K} and fine-tune in two stages, \textit{i.e.} first learning a shared spatial-navigation foundation, and then specializing each phase with the abilities it relies on. We further introduce Spatial-NPD, where a teacher conditioned on spatial priors produces grounded action preferences and distills them into a student policy, so no explicit spatial reasoning is needed at inference. With 45 A100 GPU-hours of policy training, our 8B model reaches SR/SPL of 77.4/35.4 on HM3D-v0.2, 60.2/30.5 on HM3D-v0.1, and 47.9/20.6 on train-unseen MP3D. It outperforms several systems that rely on closed-source models or thousands of GPU-hours of training, at 148 ms per action step. All code and datasets will be publicly available at https://github.com/ylwhxht/Spatial-Nav.</p>]]></description>
  </item>
  <item>
    <title>An Empirical Study of VLM Pipelines for Long-Document QA</title>
    <link>http://arxiv.org/abs/2609.29933v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29933v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:00:18 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Kenan E. Ak, Jay Mohta, Gwang Gook Lee, Yan Xu, Dimitrios Dimitriadis</p><p>Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them means choosing how to feed the document to the model, which retriever to use when only a subset of pages is sent, and whether to run the model agentically or as a static pipeline. We study these choices on two long-document QA benchmarks with both frontier API and open-weight VLMs. First, on MMLongBench-Doc our six-tool agent with page, table, figure, and search calls pays off only once the answering VLM is large enough: with Qwen3.5-4B and 9B it trails static page input, with Qwen3.5-27B it draws level, and with Sonnet 4.5 it leads. On LongDocURL it is level with or ahead of static input at every reader. Its lead over the strongest static pipeline is clearest with the frontier reader on MMLongBench-Doc and narrows to within noise on LongDocURL. Second, retrieval modality matters more than the specific retriever: the strongest image retriever leads the strongest text pipeline, and on the text side a single off-the-shelf cross-encoder rerank essentially matches a much heavier multi-stage LLM pipeline. Top-k image retrieval is also the most token-efficient input at every reader we paired it with, at roughly a seventh to a quarter of the tokens of sending every page. Third, cutting across all three choices, three of our strongest pipelines succeed on different questions, and an oracle that picks the best pipeline per question gains roughly thirteen points over the best single pipeline, though evidence-type routing recovers almost none of it.</p><p><em>Comment: 22 pages. EMNLP 2026 Industry Track</em></p>]]></description>
  </item>
  <item>
    <title>Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation</title>
    <link>http://arxiv.org/abs/2609.29931v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29931v1</guid>
    <pubDate>Thu, 24 Sep 2026 15:00:01 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Nathan Le, Magdalini Paschali, Arogya Koirala, Andrew Johnston, Zhongnan Fang, David B. Larson, Akshay S. Chaudhari, Camila Gonzalez</p><p>Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities accurately reflect true risk. Many standard techniques for improving calibration, such as MC Dropout and Deep Ensembles, require access to model parameters or retraining. However, proprietary clinical AI systems operate as black boxes, preventing access to the model&#x27;s internals. To that end, we propose a model-agnostic framework for improving calibration of black-box models using clinically grounded test-time augmentation (TTA). Our framework applies geometric and physics-inspired 3D CT perturbations and learns probability-level aggregation strategies without access to model internals or the original training data. Across pulmonary embolism and intracranial hemorrhage detection tasks, DualTTA achieved the strongest overall calibration among TTA methods, reducing the Expected Calibration Error by 54% (0.239 -&gt; 0.109) and 43% (0.051 -&gt; 0.029), respectively, while requiring only input-output access. Additionally, DualTTA outperformed uncertainty estimation techniques that require access to model internals, such as Temperature Scaling, MC Dropout, and Deep Ensembles, in most calibration metrics. These results demonstrate that learned TTA aggregation can improve the calibration of clinical AI systems, providing a practical approach for improving the reliability of black-box medical AI.</p><p><em>Comment: 11 pages, 3 figures, 1 table. Accepted at the MICCAI 2026 Workshop on Uncertainty for Safe Utilization of Machine Learning in Medical Imaging (UNSURE 2026)</em></p>]]></description>
  </item>
  <item>
    <title>Pairwise Approximation Can Select the Wrong Multi-Robot Plan</title>
    <link>http://arxiv.org/abs/2609.29929v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29929v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:59:20 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> William Teo</p><p>Multi-robot coordination methods often score a joint plan from singleton and pairwise terms, leaving out the terms that involve three or more robots. We measure the plan-selection regret of two pairwise approximations to delivered coverage using frozen multi-robot trajectories. For each four-robot plan on an indoor exploration benchmark, replaying all 16 robot subsets gives the exact delivered-coverage set function $F$. From the same subset values we compute two pairwise scores: the exact order-2 Möbius truncation $F_2$, which depends only on the singleton and pair values, and an equal-weight least-squares two-additive fit $G$. Ranking by $F_2$ instead of $F$ changes the selected plan on six of seven maps at the 15 m candidate-generation range in each of two candidate families, with regret up to 0.337 of map coverage. Switching to $G$ reduces the regret but still changes the selection on three of seven maps in each family. The additive score $F_1$, which keeps only the singleton terms, selects the exact winner on six of seven maps in one family and four of seven in the other, against one of seven for $F_2$. We also find that lower average reconstruction error does not guarantee lower selection regret.</p><p><em>Comment: 6 pages, 4 figures, 1 table. Accepted at the IROS 2026 Workshop on Intelligent Information Gathering. Code: https://github.com/williamteo/pairwise-regret</em></p>]]></description>
  </item>
  <item>
    <title>Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations</title>
    <link>http://arxiv.org/abs/2609.29928v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29928v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:59:06 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yeeun Chae, Yewon Choi, Seunghyun Lee, IL Im</p><p>Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration. CDP identifies reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature. To evaluate CDP, we conduct experiments across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test. The results reveal a systematic discrepancy between conventional fidelity metrics and CDP. Controlled experiments show that CDP changes monotonically as cross-country divergence is attenuated or amplified, while the corresponding changes in JSD remain relatively small. In our audit of real LLM generations, DeepPersona-Inspired prompting is frequently favored by conventional fidelity metrics but exhibits the strongest flattening in every model--domain block. CDP thus complements fidelity metrics by directly quantifying the attenuation or amplification of cross-country divergence.</p><p><em>Comment: Accepted to the EMNLP 2026 Workshop on Pluralistic AI &amp; NLP (PANDORA)</em></p>]]></description>
  </item>
  <item>
    <title>Who Holds the Pen? Let Specifications, Not Agents, Sign Off</title>
    <link>http://arxiv.org/abs/2609.29921v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29921v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:55:14 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Haiqing Li, Xin Ma, Yinhao Wu, Wenliang Zhong, Feng Jiang, Thao M. Dang, Xiao Hu, Hehuan Ma, Yuzhi Guo, Junzhou Huang</p><p>Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is understood but not satisfied in execution; the state--authority gap arises when an agent&#x27;s interpretation or completion claim does not establish the required state. On SkillsBench, using only agent-visible prompts, workspace information, and injected skill specifications, we extract 509 source-grounded task directions. Across seven models, only 79.6%--86.4% are satisfied, while completion-claim rates exceed official evaluator pass rates by 28.7--37.9 percentage points. We therefore separate agent proposals from authoritative state. Agents may plan, act, and request completion, but only admissible evidence from qualified providers may establish specification-governed state. SpecHarness operationalizes this principle by compiling visible specifications into source-linked obligations and governing execution and finalization through versioned obligation state. Verifiable requirements are mediated or validated at runtime, while ambiguous or subjective requirements remain advisory. Experiments on guideline-following and artifact-generation tasks show that specifications can serve not merely as behavioral guidance, but as authority over compliant execution and completion.</p>]]></description>
  </item>
  <item>
    <title>High-Voltage Optocoupler Amplifier for Electrostatic Actuators</title>
    <link>http://arxiv.org/abs/2609.29914v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29914v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:52:40 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> George C. Jurgiel, Alex S. Miller, Jeffrey H. Lang</p><p>Many electrostatic actuators require multi-kilovolt drive voltages at sub-milliamp currents, a task poorly suited for conventional switching devices. As an alternative, we demonstrate a high-voltage amplifier using optocouplers as active elements. The amplifier produces a 20-kV peak-to-peak output with up to 500 Hz bandwidth while maintaining a minimal component count. By using optocouplers as linear devices in feedback, lower harmonic distortion and higher bandwidth are achieved than offered by equivalent PWM amplifiers. This design improves the viability of electrostatic actuators by providing a simpler method to achieve useful drive waveforms.</p><p><em>Comment: To be published in the proceedings of the 2026 IEEE Energy Conversion Congress &amp; Expo (ECCE)</em></p>]]></description>
  </item>
  <item>
    <title>MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression</title>
    <link>http://arxiv.org/abs/2609.29913v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29913v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:52:24 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Youpeng Zhao, Tian Tan, Liqian Peng, Jun Wang, Alec Go</p><p>Many-shot in-context learning (ICL) enables large language models (LLMs) to adapt to complex tasks by conditioning on thousands of demonstration examples, but this paradigm shifts the inference efficiency bottleneck to the key-value (KV) cache memory. Due to the linear scaling behavior of the KV cache, storing these intermediate tensors has become a paramount challenge for both online serving and on-device deployment. To address this issue, we propose a novel compression framework, termed MILO, that exploits the low-rank redundancy inherent in many-shot contexts. Specifically, MILO features a block-wise low-rank compression strategy that compresses the KV cache at the block granularity, where each block contains multiple many-shot examples. Furthermore, to handle the heterogeneous context density across different blocks, MILO dynamically allocates rank budgets based on the information entropy, preserving the fidelity of critical blocks while aggressively compressing redundant ones. Experimental results on Qwen2.5 models demonstrate that our method achieves up to 50% reduction in KV cache memory and 1.8x throughput improvement, with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.</p><p><em>Comment: Technical Report</em></p>]]></description>
  </item>
  <item>
    <title>Structured Pose-Conditioned Flow Matching for Generative 5G CSI Augmentation</title>
    <link>http://arxiv.org/abs/2609.29912v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29912v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:52:16 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Haojin Li, Anbang Zhang, Wai Ho Mow, Chenyuan Feng, Chen Sun, Haijun Zhang</p><p>With the growing demand for privacy-preserving and occlusion-resilient human pose recognition (HPR), 5G channel state information (CSI) offers a promising contactless sensing modality by integrating communication and sensing capabilities. However, collecting large-scale synchronized CSI-pose pairs remains costly in practical 5G systems. To address this limitation, we propose StructFlow-HPR, a structured pose-conditioned flow matching framework for generative CSI augmentation. StructFlow-HPR learns a continuous latent transport process from Gaussian noise to real CSI representations under pose guidance, while preserving the receiver-frequency topology of CSI through a reconstruction-preserving autoencoder. A pose-conditioned Transformer is further designed to model the latent velocity field and generate pose-aligned CSI samples via ordinary differential equation sampling. Experiments on real-world 5G sensing data show that StructFlow-HPR can produce realistic CSI-pose pairs and improve downstream HPR performance under limited-data conditions.</p>]]></description>
  </item>
  <item>
    <title>MorphIK: Morphology-Conditioned Neural Inverse Kinematics for Unknown Robots</title>
    <link>http://arxiv.org/abs/2609.29908v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29908v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:48:33 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Lennart Clasmeier, Jan Gerrit Habekost, Cornelius Weber, Stefan Wermter</p><p>Neural models can learn to generate various solutions to the inverse kinematics problem from data, but are usually limited to a single robot. We present MorphIK, a flow-matching model that solves inverse kinematics for revolute-joint-based kinematic chains it has never seen during training. The model uses a transformer architecture to encode the robot&#x27;s morphology along with the target pose. This encoding then conditions a flow-matching head that generates poses from noise. Trained on purely synthetic data from procedurally generated robots, the model reaches a precision of about 5 cm on unseen real-world robots with 6 to 9 Degrees of Freedom. For higher precision, the model serves as an excellent Prior for further optimization algorithms, reducing error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm error after 3 steps in most cases. Building on flow matching&#x27;s generative capabilities to produce highly diverse outputs, our model can efficiently sample the robot&#x27;s null space, providing a wide variety of configurations for the same pose. Thus, overall, MorphIK allows learning and generalizing neural inverse kinematics for a multitude of known and unknown robots.</p>]]></description>
  </item>
  <item>
    <title>Spatio-temporally complementary feature propagation on graphs for longitudinal AADT estimation</title>
    <link>http://arxiv.org/abs/2609.29906v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29906v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:47:04 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Linghang Sun, Qishen Zhou, Michail A. Makridis, Anastasios Kouvelas</p><p>The estimation of Annual Average Daily Traffic (AADT) is vital for transportation planning and infrastructure maintenance, yet obtaining accurate values for an entire urban network across multiple years remains challenging due to the high cost and spatial sparsity of physical sensors. This research proposes a novel spatio-temporally complementary feature propagation framework that leverages the strengths of two distinct data sources: spatially sparse but temporally dense loop detector data, and a spatially complete but temporally sparse macroscopic transportation model. The methodology highlights a feature propagation algorithm on directed graphs, formulated as a Poisson energy minimization considering residues. The standard binary adjacency matrix is replaced with flow ratio matrices to capture real-world vehicle turn ratios at intersections. Validated in the city of Zurich, the algorithm demonstrates high computational efficiency, achieving convergence within minutes. Results indicate that the framework effectively reconciles theoretical models with empirical ground truths, yielding a normalized mean absolute error below $10\%$. This scalable approach provides a feasible solution for spatio-temporal network-wide AADT estimation through combining real-world limited sensor coverage and traffic models.</p>]]></description>
  </item>
  <item>
    <title>Working with Agentic `Teammates&#x27;: When a New Organizational Actor Collides with the Human Ecosystem of Work</title>
    <link>http://arxiv.org/abs/2609.29901v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29901v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:44:52 +0000</pubDate>
    <category>Spatial Agents</category>
    <description><![CDATA[<p><strong>Authors:</strong> Rida Qadri, Remi Denton, Michael Madaio, Mahima Pushkarna, Leslie Lai, Sherry Moore, Michelle Chen Huebscher, Andrew Butcher, Ritom Sen, Hsiao-Yu Tung, Shaan Mathur, Yimeng Liu, Shibl Mourad, Noah Fiedel, Edward Grefenstette, Michael Terry</p><p>Enterprise AI is transitioning from single-user, reactive tools toward proactive, multi-user &#x27;teammates,&#x27; but our empirical understanding of this transition is limited. In this paper, we present an in-situ qualitative study of a persistent, proactive AI agent &#x27;teammate&#x27; deployed across multiple teams in a large technology company. Our findings reveal the boundaries of the human-agent workplace are actively in flux, triggering breakdowns and negotiations across: 1) tacit rules of collaborative human workflows, 2) the relational boundaries of this new non-human actor, and 3) the redistribution of trust and human agency. We use these early micro-negotiations as signals to chart a new research, design, and organizational agenda that intentionally preserves human agency in a workplace shared with non-human organizational actors.</p>]]></description>
  </item>
  <item>
    <title>Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents</title>
    <link>http://arxiv.org/abs/2609.29892v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29892v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:38:08 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Tingyu Qu, Weigao Sun, Yuecheng Liu, Yucheng Zhao, Yi Zhu, Yifeng Ding, Qiyi Wang, Sihan Cao, Pengkun Jiao, Hanlei Xie, Xiongwei Wu, Qichao Wang, Haodong Zhang, Jiajun Liu, Yuhao Wang, Yuqing Xie, Junpeng Zhao, Long Chen, Ming Ma, Sihan Yang, Ziwang Zhao, Yanhao Jia, Liangquan Gong, Feida Zhu, Yiran Zhong, Steven Hoi</p><p>The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.</p><p><em>Comment: https://tongyi-mai.github.io/Qwen-Planner-Agent/</em></p>]]></description>
  </item>
  <item>
    <title>Cost-Sensitive Online Window Size Selection for Portfolio Management</title>
    <link>http://arxiv.org/abs/2609.29887v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29887v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:35:25 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yi-Chen Liu, Chung-Han Hsieh</p><p>This paper investigates cost-sensitive online window size selection for portfolio management under changing market conditions. Specifically, we propose a two-level framework that constructs portfolios using candidate window sizes and dynamically aggregates them through online learning. By treating candidate window sizes as ``experts,&#x27;&#x27; we dynamically update their aggregation weights using turnover-inclusive losses. Moreover, we derive finite-horizon cost-sensitive tracking-regret bounds that account for turnover of the aggregated portfolio, with static regret as a special case. Under bounded losses and cost rates, suitably tuned Fixed Share achieves asymptotically no tracking regret for sublinear switching budgets, with Hedge covering the static case.</p>]]></description>
  </item>
  <item>
    <title>A New Gap Sequence for Shellsort: RL-Driven Algorithm Discovery Beyond $N^{4/3}$</title>
    <link>http://arxiv.org/abs/2609.29881v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29881v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:33:05 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Bo Liu</p><p>Choosing Shellsort gaps is a well-known open problem. For over sixty years, successful sequences have relied on human-designed formulas, numerical searches, or number-theoretic constructions. Although stronger general bounds exist for dense or mainly theoretical families, the worst-case upper bound for a short, sparse, and practically competitive construction has not advanced beyond $N^{4/3}$ for decades. We ask whether the sequence itself can instead be learned from execution. We present an RL-driven, self-supervised system that searches over executable gap generators. Every proposal is valid by construction, and executed candidates return exact comparison and move counts; no classical sequence is used as a target. Across five independent searches, the system discovers a common rational-geometric family. A second self-supervised stage tunes only a finite prefix, producing the practical sequence $1,3,8,20,47,116,300,585,1416,3303,\ldots$. Once frozen, it obtains the lowest equal-task average operation count among seven classical baselines on 25 large tasks with $10^7&lt;N\leq 10^8$. We complete the learned tail without changing its practical behavior: only beyond $10^{1000}$, a zero-density set of unit companions $h_s+1$ removes the remaining congruence barriers. The resulting sparse sequence has matching polynomial upper and lower exponents, up to polylogarithmic factors: $Ω(N^{1.024296451657\ldots}) \leq T(N) \leq O(N^{1.024296451657\ldots}\operatorname{polylog} N)$. The lower bound follows from Zang&#x27;s recent theorem for rational-geometric sequences; our contribution is the matching upper bound. Thus one exact sequence connects self-supervised discovery, large-scale practical performance, and a substantial step below the classical $N^{4/3}$ bound for sparse practical Shellsort sequences.</p><p><em>Comment: 25 pages, 2 tables; full proof and technical appendix</em></p>]]></description>
  </item>
  <item>
    <title>From Graphs to Feeders: Constraint-Guided Diffusion for Rule-Compliant Feeder Generation</title>
    <link>http://arxiv.org/abs/2609.29879v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29879v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:31:12 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yu Qin, Andrew Glaws, Aadil Latif, Ryan King</p><p>Generative modeling approaches often focus on recovering broad statistical characteristics from the training data. In the context of graph generation, this may refer to degree distributions, clustering coefficients, or spectral properties. However, generating usable distribution feeders when detailed feeder models are unavailable requires more than matching generic graph statistics: the sampled topology must also obey electrical compatibility and radiality rules. We therefore formulate feeder synthesis as a constraint-guided graph generation problem and propose the Power-Grid-constrained Discrete Denoising Diffusion model, PG-DiGress, which learns categorical node and edge patterns from feeder data, while respecting domain-specific rules. Specifically, it injects feeder constraints into the reverse diffusion process through soft masks that suppress incompatible edge classes during denoising, followed by a final projection step that rebuilds a connected, rule-compliant feeder graph. We evaluate PG-DiGress using graph-distribution similarity, feeder-rule satisfaction, structural validity, and downstream model construction. Compared with the unconstrained baseline, PG-DiGress increases the strict feeder pass rate from 13.7% to 96.8%. We also successfully convert the generated graphs into executable feeder models for downstream analysis.</p><p><em>Comment: 22 pages</em></p>]]></description>
  </item>
  <item>
    <title>A Risk-Adaptive and Evidence-Constrained Framework for Generative AI Feedback in Programming Education</title>
    <link>http://arxiv.org/abs/2609.29874v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29874v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:28:35 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Shihao Wang</p><p>Generative artificial intelligence can turn learning analytics into personalized support, but feedback systems must decide when to intervene, which evidence to use, and how much assistance to provide. We developed a risk-adaptive, evidence-constrained framework for introductory programming using 2993 failed-submission states from 215 students. Student-disjoint models predicted persistent failure and related outcomes; four matched feedback conditions were generated for 136 cases; and calibrated risk informed capacity-limited intervention policies. The validation-selected logistic regression model achieved a test precision-recall area under the curve of 0.550 and a receiver operating characteristic area under the curve of 0.681. Broader student histories improved prediction of unmodified resubmission. After standardized repair and evidence gating, 519 of 544 newly generated messages contained all required components. A fixed-threshold sequential policy selected 17.8% of eligible test states and captured 25.2% of observed persistent failures. These findings support an evidence-gated progressive assistance strategy: calibrated risk guides intervention timing, recorded evidence constrains feedback content, and assistance progresses from self-checks to localized hints when warranted. The framework connects prediction, decision-making, and grounded generation while keeping their evaluation outcomes distinct.</p>]]></description>
  </item>
  <item>
    <title>Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement</title>
    <link>http://arxiv.org/abs/2609.29867v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29867v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:25:00 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Clément Laroche, Riccardo Miccini</p><p>Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supervise every intermediate depth of one causal model, then we fine-tune its output heads to guarantee that deeper outputs are never worse than shallower ones. Using this training protocol, we can derive a family of static models that are more Pareto-efficient than their equivalently-sized counterparts trained from scratch on the same budget. Specifically, we achieve up to 0.11 higher PESQ for equivalent compute, and match the best PESQ at 30% less compute. We then quantize the models to int8 and measure the latency-quality frontier on an STM32N6 microcontroller. On VoiceBank-DEMAND, the dynamic enhancer lies on the same frontier as the static models, rather than trading quality for dynamic execution. Running the policy on the companion Cortex-M55 takes only 26 $μ$s per frame, while splitting the enhancer into separate NPU graphs adds 2.2% latency overhead. The cost of dynamic execution is therefore small.</p>]]></description>
  </item>
  <item>
    <title>Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement</title>
    <link>http://arxiv.org/abs/2609.29866v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29866v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:24:43 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Clément Laroche, Rasmus Kongsgaard Olsson</p><p>Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and execution patterns often remain incompatible with restricted NPUs. We address this gap by redesigning LiSenNet, a 37k parameter sub-band dual-path model, for the STM32N6570-DK Neural-ART accelerator. We replace its recurrent bottleneck with convolutional frequency and temporal mixers, reformulate unsupported operations as static int8-compatible primitives, and use bounded decoder activations to preserve quality after quantization. On VoiceBank-DEMAND, the final NPU-compatible model matches or exceeds the recurrent LiSenNet baseline, reaching PESQ 3.08 versus 3.01 in FP32 and 3.01 versus 2.93 in int8. Deployed on a microcontroller, it processes each 16 ms input hop in 4.83 ms, corresponding to a real-time factor of 0.30. Stateless receptive-field recomputation is an order of magnitude slower at the same frame rate despite higher accelerator utilization. These results show that parameter count and operator compatibility, quantization range, and persistent streaming state must be co-designed to achieve efficient real-time speech enhancement on restricted NPUs.</p>]]></description>
  </item>
  <item>
    <title>Efficient Continuous DEM Reconstruction under Limited Target-Resolution Supervision</title>
    <link>http://arxiv.org/abs/2609.29864v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29864v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:23:59 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Zekai Shi, Meng Zhang, Haokun Zhang, Bo Zhang</p><p>High-resolution digital elevation models (DEMs) support Earth observation applications, but paired training references are often available only at coarser output resolutions. Reconstructing finer terrain grids therefore requires both effective transfer beyond the supervised scale and control of dense-query computation. To address this problem, SCOPE learns a continuous terrain representation from coarser-resolution pairs. It predicts a latent coefficient field on the low-resolution grid and reuses local Fourier residual functions through basis evaluation and geometry-guided ensemble fusion. This separates high-dimensional coefficient prediction from output-grid construction. Experiments on geographically distributed land--ocean samples assess supervised reconstruction, unseen-scale inference, cross-domain generalization, and theoretical computation. SCOPE leads the compared methods across six metrics in the main supervised-scale evaluation. At an unseen factor three times the training factor, land reconstruction reduces RMSE and MAE by approximately 12\% relative to bicubic interpolation, with errors close to target-scale fine-tuning. Ninefold output density increases counted multiply--accumulate operations by only about 2\%. Frozen-model validation on held-out external marine regions reduces RMSE relative to the DEM-specific implicit baseline EBCF-CDEM by approximately 19\% under self-downsampling and 2\% with cross-product inputs, while also yielding lower RMSE than LIIF-MS in both settings. These results demonstrate the value of reusable coefficient fields for accurate reconstruction beyond the supervised resolution with low incremental arithmetic cost.</p><p><em>Comment: 19 pages, 15 figures</em></p>]]></description>
  </item>
  <item>
    <title>Modelling dynamic systems transfer functions from events in computational neuromorphic imaging</title>
    <link>http://arxiv.org/abs/2609.29863v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29863v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:22:47 +0000</pubDate>
    <category>Human Body</category>
    <description><![CDATA[<p><strong>Authors:</strong> Nimrod Kruger, Gregory Cohen</p><p>Event Vision Sensing (EVS) report threshold crossings of log-irradiance, so a static optical system imaging a static scene produces no output at all. The classical procedure for measuring a Point Spread Function (PSF), illuminating the system with a constant point source, therefore has no event-based equivalent: the probe must carry a temporal profile, and that profile becomes part of the measurement. A growing body of Computational Neuromorphic Imaging (CNI) work already exploits this, pairing engineered or modulated optics with event sensing, but each system adopts a particular excitation together with a particular reading of the event stream without the correspondence between the two being stated. We examine that correspondence directly within a analytical framework of an Linear Shift-Invariant (LSI) optical system with a specified Modulation Transfer Function (MTF), a first-order filter EVS pixel model, and three different temporal probes: a step function, a linear ramp and an exponential ramp. By analysing the inverse of the entire chain for different event-statistic, and comparing the results to the specified MTF, we identify the context where each probe is most relevant. We consider how photon-noise and cross-array threshold mismatch effects the analytical accuracy of the probe-inverse. Results show that the widely used step probe is highly susceptible to mismatch while resilient to photon shot-noise, while a linear rise probe and exponential rise probe retain their ability to infer signal levels even with high mismatch. We discuss the potential of dynamic-PSFs as components of a full forward operator from scene to events. In this, we use this analytical description to define dynamic-PSFs around EVS, and discuss the gaps toward a unified pixel model and a scene-composition framework required for CNI.</p><p><em>Comment: Conference paper - SPIE Sensors + Imaging 2026</em></p>]]></description>
  </item>
  <item>
    <title>GPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments</title>
    <link>http://arxiv.org/abs/2609.29861v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29861v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:21:26 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Guangzhao Dai, Qianru Sun, Qi Wu, Bin Zhu</p><p>We investigate whether GPT-6-Astra, a general-purpose foundation model, can navigate unfamiliar environments using its own perception, reasoning, and decision-making capabilities. Our evaluation focues on zero-shot vision-and-language navigation in continuous environments (VLN-CE) through a minimal interface in the Codex harness, aiming to unleash GPT-6-Astra&#x27;s full potential for navigation. Using monocular RGB, GPT-6-Astra decides when to observe, how to move, and when to stop, without navigation-specific fine-tuning, a trained waypoint predictor, or a pre-built scene map. Our evaluation yields four main findings. First, \textbf{\textit{GPT-6-Astra achieves strong zero-shot navigation performance using only monocular RGB observations}}. On the common-adopted zero-shot R2R-CE benchmark, ultra reasoning achieves a success rate of \textbf{\textit{79.0\%}}, exceeding the strongest reported zero-shot and supervised success rates by \textbf{\textit{13.0}} and \textbf{\textit{6.9}} percentage points, respectively. Second, \textbf{\textit{GPT-6-Astra advances multi-stage language instructions into coherent, adaptive navigation}} by grounding spatial relations, tracking task progress, and revising its actions. Third, \textbf{\textit{reliable route execution and goal verification remain challenging, even with ultra reasoning}}. Plausible local landmark matches do not consistently lead to correct task completion. Fourth, \textbf{\textit{these capabilities motivate rethinking the role of embodied learning}}. Future VLN research should build on foundation models to advance generalizable and reliable embodied intelligence.</p><p><em>Comment: Technical report</em></p>]]></description>
  </item>
  <item>
    <title>Multi-Task Learning by using Contextualized Word Representations for Syntactic Parsing of a Morphologically Rich Language</title>
    <link>http://arxiv.org/abs/2609.29855v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29855v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:19:06 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Toqeer Ehsan, Miriam Butt, Sarmad Hussain, Hassan Alhuzali, Ali Al-Laith</p><p>We address the challenge of syntactic parsing for Urdu, a morphologically rich language, and present state-of-the-art results for both constituency and dependency parsing. This paper offers four major contributions: 1) the conversion of the CLE-UTB phrase structure treebank into a dependency treebank by developing language-specific head-word and phrase-to-dependency label mapping rules; 2) a novel sequence labeling scheme that transforms the parsing task into a unified representation; 3) the training of contextualized word representations on a large 220 million tokens Urdu corpus collected from the web; and 4) development of parsing framework using two learning paradigms, single-task and multi-task learning. Several post-processing rules are applied to improve the quality of the automatically converted dependency structure treebank. The proposed sequence labeling scheme enables the use of a shared architecture that learns the syntactic structures from both grammatical structures simultaneously and hence improves generalization. Experiments show that the multi-task learning setup significantly enhances parsing performance, achieving an F1 score of 91.39 for constituency parsing (an improvement of 3.29 points) and a labeled attachment score of 85.69 for dependency parsing (an improvement of 1.49 points). These results demonstrate that learning cross-task representations provides measurable benefits and advances the state of syntactic parsing for Urdu.</p><p><em>Comment: Published in PLOS ONE, 2025</em></p>]]></description>
  </item>
  <item>
    <title>Template Ageing and Longitudinal Verification in Fixed-Text Keystroke Dynamics: A Subject-Disjoint Study Across Eight Weeks</title>
    <link>http://arxiv.org/abs/2609.29851v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29851v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:16:32 +0000</pubDate>
    <category>Generative Recommendation Systems</category>
    <description><![CDATA[<p><strong>Authors:</strong> Simon Parkinson, Saad Khan, Na Liu, Qing Xu</p><p>Behavioural biometric templates are widely believed to degrade as the gap between enrolment and verification grows, but few studies measure this template ageing effect directly under controlled conditions. We collected a longitudinal dataset of 40 fixed passwords, each typed four times per weekly session over eight consecutive weeks. We compare a scaled-Manhattan matcher (M1), a gradient-boosted classifier (M2), a TypeNet-style recurrent embedding model (M3), and a TypeFormer-style Transformer (M4) under a 5-fold subject-disjoint protocol and a design that jointly varies mechanism and the enrolment-to-query gap, from 0 to 7 weeks. Template ageing proves large and systematic. Error increases monotonically with the gap for every mechanism, from an EER of 14.6-27.2% at a gap of zero to 25.5-37.1% at seven weeks, or 1.7% of decision error per week elapsed (p &lt; 0.001). However, the choice of mechanism matters more than its rate of ageing. Baseline accuracy spans 12.6 percentage points across the four mechanisms, the degradation each accumulates over seven weeks spans only 2.3 points, and ageing never reorders them. A matcher can therefore be chosen on same-session accuracy, with ageing managed by re-enrolment scheduling rather than by matcher selection. The two properties are nonetheless distinct, as M3 is the least accurate mechanism yet ages significantly more slowly than M1 under every specification tested. Training randomness also matters differently by architecture, with 58% of the recurrent model&#x27;s fold-to-fold variance attributable to seed noise against 19% for the Transformer. Because the smaller ageing-rate differences are sensitive to modelling choices, while the accuracy differences and the ageing effect are not, we recommend that comparative ageing-rate claims be supported by seed-level score fusion, independent replication, and an alternative outcome-model specification.</p>]]></description>
  </item>
  <item>
    <title>BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video</title>
    <link>http://arxiv.org/abs/2609.29850v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29850v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:15:04 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, Xun Cao</p><p>Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.</p>]]></description>
  </item>
  <item>
    <title>Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax</title>
    <link>http://arxiv.org/abs/2609.29848v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29848v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:13:26 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Zhenyan Lu, He Wang, Xiaohui Huang</p><p>A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a three-level evaluation framework (behavioral deployment, LM-head readout, and probe recoverability) measured on the same items under the same binary decision. Using a compact trilingual (English, Chinese, German) control-dependency benchmark, we find that probe recoverability exceeds or equals LM-head readout, which in turn exceeds or equals behavioral deployment, across seven models and all three languages in the aggregate. The recoverability surplus is never negative across all 14 (model, task) conditions. The disconnect concentrates in subject-control, where a nearest-noun heuristic gives the wrong answer. The single largest gap (0.653) appears on Qwen3-0.6B Instruct in question answering. The gap persists at Qwen3-14B Instruct. Instruction tuning degrades deployment more than encoding in percentage terms. We rule out option-position bias, late-layer erasure, output-formatting artifacts, and probe-training variance. The pattern is consistent with decoding that favors surface shortcuts, and the behavior-probe gap measures the strength of that preference. Activation patching shows the gap is layer-localized. Under instruction tuning, the LM-head-decoded layer shifts approximately ten layers later than the probe-decoded layer. These findings argue that behavioral evaluation understates what models encode, while probing alone overstates what they deploy.</p><p><em>Comment: Accepted by AACL-IJCNLP 2026</em></p>]]></description>
  </item>
  <item>
    <title>Elucidating the Conformal Structure of the Brinkman Penalisation Method for Geometry-Adapted, Structure-Preserving Operator Learning of Hamiltonian PDEs</title>
    <link>http://arxiv.org/abs/2609.29847v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29847v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:13:12 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Teo Deveney, Baige Xu, Takaharu Yaguchi</p><p>The Brinkman penalisation method embeds boundary-value problems on complex domains into a simple computational box by modeling the solid region as a strongly dissipative medium, avoiding body-fitted mesh generation. We show that multi-symplectic Hamiltonian PDEs regularised by Brinkman-type penalisation retain a multi-conformal symplectic structure under a compatibility condition linking the symplectic matrix and the penalisation projection. This yields an exact local conservation law, under which the multi-symplectic two-form is conserved in the fluid region and decays exponentially inside the solid. The linear wave equation with Brinkman friction and Maxwell&#x27;s equations with artificial Ohmic conductivity satisfy this condition, with explicit modified Hamiltonian densities. Building on this, we propose (i) structure-preserving numerical integrators via Strang splitting that satisfy a discrete conformal conservation law, and (ii) conformal symplectic neural operators that interleave exact dissipative flows with learnable multi-symplectic evolution operators, allowing geometry-dependent operator learning. Numerical experiments on wave and electromagnetic scattering demonstrate that our methods reproduce correct local energy budgets and avoid unphysical energy drift, providing a principled framework for physics-consistent scientific machine learning on complex domains.</p>]]></description>
  </item>
  <item>
    <title>Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs</title>
    <link>http://arxiv.org/abs/2609.29845v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29845v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:12:08 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina</p><p>While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the \textit{Superposition Linearity Hypothesis}. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.</p>]]></description>
  </item>
  <item>
    <title>PUBG Ally: A Conversational Embodied Agent as an AI Teammate</title>
    <link>http://arxiv.org/abs/2609.29837v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29837v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:06:28 +0000</pubDate>
    <category>Embodied Intelligence</category>
    <description><![CDATA[<p><strong>Authors:</strong> Beomsoo Kim, Byeongju Kim, Dohyun Kim, Dongwon Kim, Eunchong Kim, Hongmin Kim, Hyeojung Im, Hyeonbin Hwang, Hyeonghwan Kim, Hyoseok Seol, Insub Im, Irene Chen, Jaeseung Jeon, Jimin Hong, Kiyoon Yoo, Minkyoung Park, Seohyeon Jung, Seungjun Chung, Sue Hyun Park, Sungwoo Kim, Youngin Cho, Yujeong Son, Kangwook Lee, Hyunseung Kim</p><p>We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A language-model agent uses a controlled interface to inspect game information, interpret player speech, maintain context, decide what to say, and issue high-level action choices that steer a faster control layer for movement, combat, and recovery. Because the player&#x27;s and Ally&#x27;s speech and actions continually shape each other and the course of the match, training requires data from actual gameplay. We therefore collect data across nearly 39k sessions in which real players play alongside Ally, recording gameplay, player speech, agent decisions, tool use, actions, and player feedback, and use these records for iterative training. To evaluate teammate quality, we use player feedback and preference comparisons to identify gaps between offline evaluations and player preferences, and iteratively refine the evaluation criteria. Deploying Ally in live service further requires low-latency on-device execution and safeguards for player-facing communication, which we address through model compression, context compaction, targeted safety training, runtime guardrails, and memory redaction. During the live service, we surveyed players in 141 countries. Among respondents whose play with Ally was confirmed in game records, positive responses exceeded negative responses by 25.1 percentage points when asked whether they would recommend Ally, with players describing Ally not only as a tool but also as a teammate or companion.</p><p><em>Comment: 55 pages, 19 figures, 16 tables</em></p>]]></description>
  </item>
  <item>
    <title>SplatLabel: Pseudo-Labelling through 4D Gaussian Splatting</title>
    <link>http://arxiv.org/abs/2609.29836v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29836v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:05:59 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Nitya Nanvani, Andras Palffy, Holger Caesar</p><p>While 2D Vision Foundation Models offer a pathway to automate 3D semantic pseudo-labelling, translating these priors into robust 3D representations typically requires complex heuristics or multi-model ensembles. We introduce SplatLabel, an automated pipeline that leverages a 4D Gaussian representation to extract LiDAR segmentation with predictive confidence, as well as semantic occupancy grids at arbitrary voxel resolutions. At its core, SplatLabel handles dynamic environments through an explicit temporal manifold that models the trajectories and lifespans of individual 3D primitives. This allows the system to accurately track moving actors and strictly define when objects appear and disappear, completely eliminating the need for pre-annotated 3D bounding boxes. To robustly support this dynamic tracking, the representation is grounded by structural and semantic priors: we guide scene geometry in unobserved regions by integrating 360-degree LiDAR via virtual depth maps, and rather than relying on domain-specific prompt engineering, we directly distill continuous soft probabilities from 2D models to inherently resolve semantic ambiguities over time and space. Finally, to accurately reflect the real-world trade-off between precision and recall, we reframe pseudo-label evaluation as a selective classification task using a generalized risk-recall metric. Experiments on SemanticKITTI demonstrate that SplatLabel consistently outperforms state-of-the-art baselines across multiple recall levels, establishing a highly robust framework for both 3D LiDAR segmentation and occupancy prediction.</p>]]></description>
  </item>
  <item>
    <title>Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding</title>
    <link>http://arxiv.org/abs/2609.29835v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29835v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:05:02 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Byounggun Park, Giyong Moon, Jusung Kim, Soonmin Hwang</p><p>LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.</p><p><em>Comment: 8 pages</em></p>]]></description>
  </item>
  <item>
    <title>System Identification of an Octocopter in Hover using Full-Harmonic Orthogonal Multisine Inputs</title>
    <link>http://arxiv.org/abs/2609.29832v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29832v1</guid>
    <pubDate>Thu, 24 Sep 2026 14:03:41 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Justin J. Matt, George V. Altamirano</p><p>A new method for multi-input flight maneuver design for system identification is presented. The method consists of injecting &quot;full-harmonic&quot; orthogonal multisine signals into the flight control system. Orthogonality is achieved by repeating maneuvers with changing multisine polarities. The multisines can contain the same frequency content, which can simplify frequency response estimation and allow for long flight maneuvers to be split into several shorter maneuvers while maintaining the same frequency resolution and minimum frequency. An input allocation scheme is presented that augments the multisines to size the vehicle response amplitude about a specific degree of freedom. The developed approach was demonstrated through flight testing of a small octocopter in near-hover conditions. The input allocation scheme was utilized successfully to increase excitation about the yaw axis. Electrical power, motor speed, and rigid-body dynamic models were identified and are shown to predict the vehicle and motor responses accurately. The models are parameterized primarily by rotor thrust and torque coefficients, making them suitable for analysis of aircraft flight dynamics and individual rotor aerodynamics. The results demonstrate that the near-hover flight dynamics can be modeled accurately by neglecting rotor hub moments, variations in rotor coefficients, gyroscopic moments in roll and pitch, and aerodynamic interaction effects.</p><p><em>Comment: 31 pages, 18 figures. Presented at the AIAA AVIATION Forum 2025. Accepted for publication in the AIAA Journal of Aircraft</em></p>]]></description>
  </item>
  <item>
    <title>ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines</title>
    <link>http://arxiv.org/abs/2609.29828v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29828v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:59:51 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Amit Nautiyal, Ayush Bhatt, Gaurav Nautiyal</p><p>We present ChunkRank, an open-source Python library that derives chunk boundaries from a target model&#x27;s tokenizer and context window, and selects an answer among candidates produced independently per chunk. It ships a validated registry of 90 models across 15 providers and six answer-selection methods, and needs only three core dependencies. For chunking, ChunkRank avoids context-window overflow automatically from the model name, whereas character-based splitters overflow or waste the budget, and a fidelity study across 11 languages shows why token-exact budgets matter beyond English. For answer selection we report a negative result: on NaturalQuestions, TriviaQA and HotpotQA, with extractive and generative readers, no content-based ranker reliably beats taking the first non-empty answer. The reason is reader abstention on chunks that lack the answer, not answer position. A long-context baseline shows that chunking matches single-call reading on single-hop questions, so ChunkRank targets small-window and beyond-window settings. Code, registry and evaluation harness are released.</p><p><em>Comment: 16 pages. Code: https://github.com/AmitoVrito/chunkrank</em></p>]]></description>
  </item>
  <item>
    <title>Anatomy-Aligned Surface Field Learning for Myocardial Reconstruction from Sparse Short-Axis Cine MRI</title>
    <link>http://arxiv.org/abs/2609.29825v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29825v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:58:16 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xiaohan Yuan, Xuan Yang, Qingya Li, Yangang Wang, Lei Li</p><p>Patient-specific 4D myocardial reconstruction from cine MRI supports quantitative functional assessment, regional motion analysis, and simulation-based modeling. However, routinely acquired short-axis (SAX) cine MRI is sparsely sampled along the through-plane direction, making dense and anatomically consistent surface reconstruction challenging. In this study, we propose an anatomy-aligned surface learning framework that parameterizes the epicardial and endocardial surfaces on a shared circumferential-longitudinal UV domain. This formulation converts irregular 3D reconstruction into structured coordinate-field completion with explicit correspondence across subjects and cardiac phases. Sparse SAX contours are encoded as UV observation fields, coverage-aware sampling improves robustness to incomplete slice coverage, and topology- and distortion-aware learning preserves circumferential continuity and local surface quality. Experiments on three public cine MRI datasets showed that the proposed method consistently outperformed representative mesh-based and implicit reconstruction approaches, achieving overall Chamfer distances of $2.887$~mm on ACDC, $2.641$~mm on M\&amp;Ms, and $2.810$~mm on M\&amp;Ms-2. The reconstructed sequences also preserved ventricular function, with end-diastolic volume and ejection fraction errors of $3.3$~mL and $1.1 \%$, respectively. These results demonstrate that anatomy-aligned UV learning provides an accurate, efficient, and correspondence-aware representation for sparse cine MRI reconstruction and myocardial modeling. The source code will be available at https://github.com/yuan-xiaohan/SAX2MyoSurf.</p><p><em>Comment: 12</em></p>]]></description>
  </item>
  <item>
    <title>Self-Supervised Anchoring of Fingertip Sensing to Proprioception and Proactive Actions for Robot Imitation Learning</title>
    <link>http://arxiv.org/abs/2609.29822v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29822v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:55:10 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Tomohiro Motoda, Masaki Murooka, Keisuke Shirai, Hanbit Oh, Ryoichi Nakajo, Shotaro Miwa, Roman Mykhailyshyn, Hugo Duarte, Yukiyasu Domae</p><p>Robotic imitation learning often relies on external cameras, yet local interaction cues such as object proximity, contact onset, and grasp state are difficult to observe near the fingertips because of occlusion and limited temporal resolution. We study how to effectively incorporate complementary fingertip sensing into imitation learning using pressure-sensitive tactile and reflective proximity sensors, along with pretrained sensor encoders. The two modalities provide information at different manipulation phases: proximity sensing is informative before contact, whereas tactile sensing becomes informative after contact. However, naively adding these signals to a policy does not consistently improve performance and can even underperform vision-only policies, suggesting that sparse, phase-dependent sensor signals are difficult to exploit from limited demonstrations. We therefore propose a proprioception-anchored pretraining method, PROprioceptive-and-PRoactive Anchoring (PROPRA), which independently aligns each fingertip sensor history with proprioceptive and action segments. This provides a continuously available sensorimotor reference, allowing each sensor to be aligned independently during its informative phases. Experiments on real-world manipulation tasks show that our pretraining method improves average success rates over vision-only policies and image-anchored pretraining baselines. Representation analysis further shows that it preserves richer information about pre-contact states, enabling more effective use of complementary fingertip sensing. Please refer to our project page: https://tomohiromotoda.github.io/nia.propra/</p><p><em>Comment: Project page is available at https://tomohiromotoda.github.io/nia.propra/</em></p>]]></description>
  </item>
  <item>
    <title>Fair Feed Ranking for Participatory Budgeting</title>
    <link>http://arxiv.org/abs/2609.29819v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29819v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:54:27 +0000</pubDate>
    <category>Information Retrieval</category>
    <description><![CDATA[<p><strong>Authors:</strong> Carina I. Hausladen</p><p>In large-scale participatory budgeting, citizens cannot inspect the full proposal pool, so the order in which proposals are shown becomes a form of agenda-setting power. We argue that fair exposure should therefore be treated as a democratic-design goal. We study Consul Democracy, a widely deployed open-source digital-democracy platform, and show that its proposal feeds are typically ordered by popularity, recency, or comment activity. Building on this diagnosis, we propose FairFeed, a feed-ranking design for PB that uses transparently declared preferences, boosts under-exposed proposals, and admits a rate-limited reject channel for crowd-sourced vetting. We evaluate the design in a simulation anchored in Munich&#x27;s 2025 PB process and compare it with random, newest, and most-commented feeds. In this simulation, FairFeed broadens proposal discovery, distributes visibility more evenly across the eligible pool, increases cross-cutting support, and improves resistance to manipulation relative to comment-based ranking. We conclude by outlining the human-subjects evaluation needed to test whether onboarding can recover voter preferences accurately enough for deployment in practice.</p><p><em>Comment: 8 pages, 2 figures, 2 tables. Published at GoodIT &#x27;26, the International Conference on Information Technology for Social Good, Pisa, Italy, September 2026</em></p>]]></description>
  </item>
  <item>
    <title>LSF-SR: Latent Semantic Fusion for Sequential Recommendation via Flow-based Conditional Variational Autoencoders</title>
    <link>http://arxiv.org/abs/2609.29815v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29815v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:50:12 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Shih-Hong Chen, Josh Jia-Ching Ying, Vincent S. Tseng</p><p>Sequential recommendation aims to predict users&#x27; future interests from their historical interactions. Although Large Language Models (LLMs) capture rich item semantics, existing methods often struggle to align collaborative signals with textual semantic knowledge. As a result, the learned item representations fail to capture the complementary strengths of both signals, leading to suboptimal recommendation quality. To address this limitation, we propose Latent Semantic Fusion for Sequential Recommendation via Flow-based Conditional Variational Autoencoders (LSF-SR), a novel framework that uses a Conditional Variational Autoencoder (CVAE) with Normalizing Flows to fuse item ID embeddings and LLM-generated semantic signals. At the core of LSF-SR is a conditional fusion module augmented with planar or radial flows. This module learns a flexible latent space that encourages items with similar semantic profiles to cluster together within the latent manifold. Through extensive experiments on five public benchmark datasets, we demonstrate that LSF-SR consistently outperforms state-of-the-art baselines, achieving gains of up to 12.98% and 14.13% in Recall@20 and NDCG@20, respectively.</p><p><em>Comment: Accepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026)</em></p>]]></description>
  </item>
  <item>
    <title>SwitchPFN: Shared Switching Dynamics for Frozen In-Context Time Series Classification</title>
    <link>http://arxiv.org/abs/2609.29814v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29814v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:49:20 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Zhenyi Zhu, Jacqueline Pang, Peilin Shen, Tianyi Song, Tingwei Zhang, Keyi Hu, Kangjun Yin, Shiwei Pu, Yingbo Zhou, Chen Shao</p><p>Tabular foundation models (TFMs) provide a promising route to time-series classification, but their effectiveness depends on how sequential data are converted into tabular representations. Existing representations face two challenges: global aggregation can lose the order of temporal evolution, while features computed in independently fitted coordinate systems may not have consistent meanings across sequences. We therefore view representation design for TFMs as a problem in its own right: the representation should preserve local temporal transitions while maintaining a shared feature definition across samples. We propose SwitchPFN, which learns a shared projection and regime codebook from the training sequences, making local dynamic operators and transition features directly comparable across samples. Across the evaluated benchmarks, SwitchPFN achieves the highest mean accuracy among the evaluated methods, improving over the strongest baseline by 4.47% relatively. Ablation studies, parameter sensitivity analyses, and reduced-training-data experiments further examine the contributions of the representation, its main design choices, and its behavior when labeled data are limited.</p>]]></description>
  </item>
  <item>
    <title>S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving</title>
    <link>http://arxiv.org/abs/2609.29813v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29813v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:48:08 +0000</pubDate>
    <category>End-to-End AD</category>
    <description><![CDATA[<p><strong>Authors:</strong> Zhaowei Lu, Liguo Zhou, Yujie Guo, Lei Yu, Alois Knoll</p><p>We present S2Planner, a trajectory planner that combines three front-facing cameras with ego-motion history and the current driving command. A fine-tuned DINOv3 backbone and a Spatial Tuning Adapter produce multi-scale image features; a coarse-to-fine decoder then uses trajectory self-attention and camera-projected cross-attention to refine candidate waypoints. The contribution is the integration of ego-conditioned trajectory initialization with iterative, geometry-guided sampling of multi-scale image features, rather than a new visual backbone or attention operator. On the NAVSIM v1 non-reactive evaluation, the previously reported navtest run obtained 88.03 PDMS. Because that run was selected using navtest performance, this number is exploratory and cannot be interpreted as an unbiased test estimate. Validation-selected evaluation on unexposed data, repeated runs, and computational measurements are needed to establish generalization and efficiency.</p>]]></description>
  </item>
  <item>
    <title>FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates</title>
    <link>http://arxiv.org/abs/2609.29812v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29812v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:47:03 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Wanqi Yang, Shiwei Liu</p><p>Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, \textsc{FlashLoop} delivers lossless accuracy while achieving up to 1.64$\times$ end-to-end speedup and up to 6$\times$ KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.</p><p><em>Comment: 16 pages, 9 figures</em></p>]]></description>
  </item>
  <item>
    <title>CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels</title>
    <link>http://arxiv.org/abs/2609.29807v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29807v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:41:29 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xiangwei Wang, Peng Wang, Saman Halgamuge</p><p>A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model&#x27;s output as a noisy reading of the true label and corrects it with a channel of five interpretable parameters. The channel is small enough for its posterior to be averaged from a handful of labels, and we prove that the resulting calibration preserves first-order stochastic order. On Amazon reviews and CMU-MOSEI transcripts with four LLMs, CORDIAL has the lowest log loss among nine calibrators in 76 of 80 settings with 5 to 100 labels; with 20 labels and the main 7B reader, it matches the strongest baseline using 28-54 labels. The same posterior lets us learn priors from other tasks and fuse several LLMs. Unrestricted calibrators such as Dirichlet calibration overtake it only as the calibration set grows into the hundreds or thousands.</p>]]></description>
  </item>
  <item>
    <title>SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search</title>
    <link>http://arxiv.org/abs/2609.29803v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29803v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:39:14 +0000</pubDate>
    <category>Information Retrieval</category>
    <description><![CDATA[<p><strong>Authors:</strong> Zhongxin Huang, Songyang Li, Renzhe Zhou, Feiran Zhu, Chenglei Dai, Zhen Xiao, Xuanping Li, Jingwei Zhuo</p><p>Search quality evaluation provides essential supervision and diagnostic signals for the development and iteration of industrial search systems. Although large language models (LLMs) offer a scalable alternative to manual assessment, reliable automatic evaluation remains challenging: users experience search results at the page level, while the applicable evaluation criteria are multi-dimensional and continuously evolving. Packing all evaluation criteria into a unified prompt introduces irrelevant context and potential criterion interference, whereas internalizing them through post-training tightly couples rule updates with costly model retraining cycles.
  To address these issues, we propose Skill-routed Evaluation with Evolvable Knowledge (SEEK). Specifically, SEEK externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution. A two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining. Experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis. SEEK has been deployed at Kuaishou, a short-video platform with over 400 million daily active users, significantly improving the scale and quality of online search evaluation.</p>]]></description>
  </item>
  <item>
    <title>Learning to Ideate for Scientific Impact</title>
    <link>http://arxiv.org/abs/2609.29802v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29802v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:37:59 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan</p><p>Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea&#x27;s citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery.</p><p><em>Comment: RLxF Workshop ICML 2026</em></p>]]></description>
  </item>
  <item>
    <title>Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition</title>
    <link>http://arxiv.org/abs/2609.29800v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29800v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:37:22 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Asmee Mishra, Mengjie Qian, Brechtje Post, Kate Knill</p><p>Adapting multilingual speech foundation models to low-resource languages remains difficult, especially for languages that are poorly represented during pre-training. While parameter-efficient fine-tuning (PEFT) reduces the cost of adapting large models, conventional approaches such as LoRA rely on generic low-rank parameterizations and do not explicitly use downstream task information to define the adaptation subspace. To investigate whether task-informed PEFT can better support low-resource ASR, we apply Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR, and introduce two complementary extensions: Asymmetric-Coupled FCCA (AC-FCCA), which exploits structured cross-layer sharing, and Adaptive-Rank FCCA (AR-FCCA), which reallocates adaptation capacity across projection matrices under a fixed parameter budget. Under controlled multilingual experiments, we evaluate these approaches on languages that are poorly represented or unsupported during pre-training alongside well-represented languages. Standard FCCA is competitive with, and usually outperforms, trainable-parameter-budget-matched LoRA. AR-FCCA provides the most consistent improvement over standard FCCA across both model architectures, with statistically significant gains in several evaluation settings, while retaining the same number of trainable parameters. These results show that task-informed subspace construction can be effective for low-resource speech adaptation, and that adaptive rank allocation provides a robust way to improve parameter efficiency without increasing model capacity.</p>]]></description>
  </item>
  <item>
    <title>Benchmarking and Domain Adaptation of Automatic Speech Recognition (ASR) for Adolescent Health Communication in Ghanaian Languages</title>
    <link>http://arxiv.org/abs/2609.29798v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29798v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:36:38 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Stephen E. Moore, Akwasi Asare, Mich-Seth Owusu, Paul Azunre, Joel Budu, Lawrence A. Adu-Gyamfi</p><p>This paper presents an end-to-end study of automatic speech recognition (ASR) for adolescent health communication in three Ghanaian languages (Twi, Dagbani, and Ewe). The work proceeds in three connected stages; First, we benchmark five ASR systems (three language-specific Wav2Vec2 models and two multimodal LLMs, Gemma 3n and Gemma 4) on a general-domain Bible corpus and a Youth Adolescent Sexual and Reproductive Health (ASRH) Domain ASR dataset, using Character and Word Error Rate (CER, WER). Second, guided by the benchmark, we perform supervised domain adaptation: although Gemma 4 was the strongest zero-shot candidate, fine-tuning it proved computationally infeasible, so we pivoted to the compact Qwen3-ASR-0.6B, fine-tuned on a large Ghana Bible corpus (~90k samples) and evaluated strictly on held-out human-collected in-domain audio. Fine-tuning reduced WER on every language, most dramatically for Ewe (WER from 109.3% to 64.8%, a drop of 44.5 pp; CER from 65.1% to 24.9%). Third, we validate the work through KasaHealth, a live voice-first ASRH application deployed in all three languages, complemented by Senti-Check, a technical evaluation harness. KasaHealth was tested by 50 community respondents and achieved a 100% chat-approval rate, a 72% Good-or-Excellent translation rating, and a 92% would-recommend rate, while surfacing the domain gaps that most constrain real-world use. Across all three stages the evidence converges: for these languages the binding constraint is validated in-domain data, not model capability or computation.</p><p><em>Comment: 34pages, 8figures,</em></p>]]></description>
  </item>
  <item>
    <title>TimeBraid: Unifying Time Series and Language for Understanding and Forecasting</title>
    <link>http://arxiv.org/abs/2609.29792v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29792v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:30:02 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xinyue Wang, Jiacheng Pang, Kun Zhou, Kexin Zhang, Defu Cao, Fan Feng,  Faisal, Songyao Jin, Yan Liu, Biwei Huang</p><p>We present TimeBraid, a series of unified time-series and language models that align pretrained language models and pretrained time-series foundation models through interleaved global residual attention layers. Each model inherits knowledge, instruction following, and reasoning from one side, continuous-signal perception and zero-shot forecasting from the other, and fuses the two in a shared representation space where both modalities are understood and generated. We study the design choices that make such unified modeling work: where to align the two representation spaces, how to ground language in temporal structure, how to balance understanding with generation, and how to keep joint optimization stable. The resulting recipe combines a unified prompting scheme for diverse time-series and text tasks, stabilized joint training, and supervision from 2.2M curated series--text pairs and 4.9M instruction-tuning samples. Across benchmarks spanning time-series perception, understanding, reasoning, and both context-aided and unimodal forecasting, TimeBraid remains competitive with far larger general-purpose models and task-specific counterparts.</p><p><em>Comment: 57 pages</em></p>]]></description>
  </item>
  <item>
    <title>OREO: Fidelity Alignment in 3D Generation via On-the-fly Rendering-Editing Optimization</title>
    <link>http://arxiv.org/abs/2609.29788v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29788v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:28:13 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Zhiyuan Ma, Wenbo Hu, Wang Zhao, Pengfei Wang, Ying Shan, Lei Zhang</p><p>Despite recent advancements in 3D generation, models often struggle to produce assets with high visual fidelity. To bridge this gap, we propose OREO, an alignment framework that enhances the realism of 3D generators by leveraging rich 2D diffusion priors. Instead of relying on static datasets, OREO establishes a dynamic optimization loop that produces on-the-fly edited renderings as 2D pseudo-targets. At its core, we introduce Reinforced Editing, which utilizes a 2D model to refine rendered views of the 3D output, enhancing their overall visual fidelity while preserving the underlying geometry, viewpoint, and content. These refined views serve as high-quality supervision targets, enabling the 3D generator to learn from its own generated samples and progressively improve its visual quality. Experiments demonstrate that OREO effectively improves upon pre-trained baselines, producing 3D assets with enhanced visual realism.</p><p><em>Comment: Accepted to ECCV 2026. Our project page is at https://theericma.github.io/oreo/</em></p>]]></description>
  </item>
  <item>
    <title>Lightweight Vision Transformer-Based U-Net for Brain Tumor Segmentation from MRI</title>
    <link>http://arxiv.org/abs/2609.29785v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29785v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:26:08 +0000</pubDate>
    <category>NeRF</category>
    <description><![CDATA[<p><strong>Authors:</strong> Sheekar Banerjee, Md. Srabon Chowdhury, Md. Mahbub Hasan Akash, Ishtiak Al Mamoon</p><p>Accurate brain tumor segmentation from Magnetic Resonance Imaging is essential for diagnosis, treatment planning, and surgical guidance. Although Convolutional Neural Networks, particularly UNet, have achieved significant success in medical image segmentation, they often struggle to capture the long-range spatial dependencies required to model tumors with irregular shapes and complex boundaries. This paper proposes a lightweight Vision Transformer UNet that combines the hierarchical feature extraction capability of UNet with the global context modeling of Vision Transformers. The proposed architecture incorporates a compact ViT bottleneck within a U-Net encoder-decoder framework, enabling effective learning of both local and global features while maintaining computational efficiency with only 2.6 million trainable parameters. The model was evaluated on the TCGA LGG MRI Segmentation dataset, achieving a mean Intersection over Union of 0.8100 and a Dice score of 0.8446, outperforming the baseline UNet by 3.75% and 3.15%, respectively. Extensive quantitative and qualitative analyses, including confusion matrix evaluation, precision recall curves, per-image performance distribution, and tumor size dependency analysis, demonstrate the effectiveness and robustness of the proposed method for brain tumor segmentation.</p><p><em>Comment: Accepted at The 2026 IEEE International Conference on Biomedical Engineering, Computer and Information Technology for Health (BECITHCON)</em></p>]]></description>
  </item>
  <item>
    <title>An Analytical Theory of Auxiliary Learning</title>
    <link>http://arxiv.org/abs/2609.29774v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29774v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:18:13 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Federico Milanesio, Alessandro Ingrosso, Matteo Osella</p><p>Auxiliary learning is an optimization paradigm in which a neural network&#x27;s performance on a target task is improved by jointly training it on additional tasks. However, the mechanisms behind this improvement remain poorly understood. We study this problem using a teacher-student framework and derive a closed system of differential equations describing the dynamics of online stochastic gradient descent in the large-input limit. For linear networks, we obtain a closed-form expression for the generalization error to leading order in the learning rate, quantifying how task correlations and label noise determine the benefit of auxiliary learning. For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error. Numerical experiments support the theoretical predictions and show how auxiliary tasks improve generalization by balancing the forcing dynamics towards the optimal solution with gradient noise.</p><p><em>Comment: Under review as a conference paper</em></p>]]></description>
  </item>
  <item>
    <title>WeatherDiagFlow: Evidence-Grounded Radar Nowcasting with Diagnostic Flow Refinement</title>
    <link>http://arxiv.org/abs/2609.29772v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29772v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:17:42 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Chunlei Shi, Yufeng Zhu, Yixiao Liang, Dan Niu, Yongchao Feng, Qiliang Wu, Jiong Wang</p><p>Radar nowcasting is essential for short-term warning and emergency response, yet conventional systems mainly return future radar fields and provide limited support for operational communication and post-event verification. We formulate radar nowcasting as an evidence-grounded forecast--bulletin--audit task, in which a numerical forecaster produces both future radar fields and structured diagnostic evidence. Forecast-time bulletins use only model-available evidence, whereas post-event audits incorporate future radar truth only after the forecast horizon is observed. Based on this task formulation, WeatherDiagFlow predicts motion, growth and decay, heavy-echo risk, and uncertainty to condition rolling flow refinement, while frozen-scaffold residual calibration improves long-lead strong-echo preservation. A multi-agent layer converts the structured evidence into operational bulletins and independently generates verification audits without feeding textual outputs back into the forecaster. Experiments on FJRADAR demonstrate competitive overall performance and improved strong-echo event skill. WeatherDiagFlow therefore connects numerical prediction, evidence-grounded reporting, and auditable verification under a leakage-controlled protocol.</p><p><em>Comment: 5 pages, 3 figures</em></p>]]></description>
  </item>
  <item>
    <title>JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places</title>
    <link>http://arxiv.org/abs/2609.29769v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29769v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:16:21 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Delip Rao, Chris Callison-Burch</p><p>We ask whether Jev, a typed classifier that returns probabilities over permitted answers without generating text, can replace an LLM rubric judge. We compare it with three flash-tier LLM judges on nine panels drawn from seven benchmarks, giving every judge identical criterion texts. Jev&#x27;s accuracy differs significantly from an LLM judge&#x27;s in only 8 of 27 paired comparisons, ahead mostly on binary criteria and behind only on graded ones, and most of the other comparisons are inconclusive. Summed over the nine panels, the LLM judges, called once per criterion, cost 29 to 325 times as much as Jev and took 30 to 220 times as long. On graded criteria all four judges agree more with one another than with the labels and mostly assign lower levels than the raters. One of several observational accounts is that raters followed scale conventions our criterion texts omit. Jev&#x27;s confidence ranks its own errors on most panels, which should make a cheap classifier the ideal first stage of a cascade that defers its uncertain verdicts to an LLM judge. Correlated errors undo that advantage. The LLM judges repeat nearly all of Jev&#x27;s most confident errors, so a cascade replayed on the recorded verdicts lowers cost but gains at most 1.5 points over the best single judge with cross-fitted thresholds, and at most 2.0 even with oracle thresholds.</p><p><em>Comment: 45 pages, 9 figures, 27 tables, including appendices</em></p>]]></description>
  </item>
  <item>
    <title>Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion</title>
    <link>http://arxiv.org/abs/2609.29766v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29766v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:14:15 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie</p><p>Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into each target speaker&#x27;s geometry without retraining. Landmarks are identified in one selected /u/ frame per speaker as a common phonetic reference without assuming identical articulatory configurations across speakers, and the resulting mapping is reused across recordings. We train the model on a single-speaker rt-MRI database and evaluate adaptation on eight speakers from a separate multi-speaker rt-MRI database. We compare affine and TPS configurations using 12 or 14 landmarks. Affine12+TPS14 achieves the lowest mean point-to-closest-point error of 3.19mm. These results support the combined value of anatomical landmark information and nonrigid alignment.</p><p><em>Comment: Submitted to IEEE ICASSP 2027</em></p>]]></description>
  </item>
  <item>
    <title>PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation</title>
    <link>http://arxiv.org/abs/2609.29760v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29760v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:09:50 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Conor W. Hayes, Rickmer Krohn, Aravind Ramaswami, Anunth Ramaswami, Nils Dengler, Kevin M. Lynch, J. Edward Colgate, Georgia Chalvatzaki, Matthew L. Elwin</p><p>Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io</p><p><em>Comment: 9 pages, 10 figures, preprint</em></p>]]></description>
  </item>
  <item>
    <title>On Growth and Form, and Function: Reusable Regulatory Handles Control Phenotypic Variation</title>
    <link>http://arxiv.org/abs/2609.29755v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29755v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:05:57 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Benedikt Hartl, Milton L. Montero, Marcello Barylli, Sebastian Risi, Michael Levin</p><p>How phenotypic transformations are implemented by changes in underlying regulatory dynamics remains a central question in developmental biology. Inspired by D&#x27;Arcy Thompson&#x27;s 1917 &quot;On Growth and Form&quot;, we ask whether coherent large-scale transformations of morphology can be encoded as low-dimensional modulations of a self-organizing developmental system. We use neural cellular automata (NCAs) as bio-inspired models of distributed development, in which a shared local regulatory network grows target morphologies from a single cell. We apply low-rank adaptation (LoRA) to pretrained NCAs, representing each adapted developmental program as a low-rank modulation of a fixed regulatory scaffold. Horizontal and vertical scaling of a fully grown 2D emoji phenotype can each be implemented by rank-one adaptations. Their linear combinations parametrically control phenotype size, generalize beyond the training distribution, and compose with target-specific adapters. Strikingly, adaptations learned for one phenotype transfer zero-shot across structurally and semantically diverse phenotypes sharing the same reference scaffold, while largely preserving internal features. This suggests reusable system-level hyper-directions of scale rather than morphology-specific transformations. From approximately 25,000 independently trained phenotype-specific NCA adapters with a shared scaffold, we further identify latent low-dimensional directions that functionally control phenotypic variation including scaling, style, and symmetrical fission. Together, our results provide a computational realization of D&#x27;Arcy Thompson&#x27;s remarkable grid transformations in a 2D NCA---a minimal cybernetic tissue in which variations of fully grown emoji phenotypes can be encoded, combined, and controlled through low-dimensional directions in regulatory weight space.</p>]]></description>
  </item>
  <item>
    <title>From Target Selection to Digging: A Learning-Based Framework for Continuous Autonomous Excavation</title>
    <link>http://arxiv.org/abs/2609.29750v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29750v1</guid>
    <pubDate>Thu, 24 Sep 2026 13:01:26 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Shuai Zhao, Ji-an Pan, Quantao Yang, Zheng Wang, Chaoyi Chen, Qing Xu, Keqiang Li</p><p>Repeated excavation continuously reshapes pile geometry, requiring an autonomous excavator to adapt its digging targets and coordinate motion across successive excavation cycles. We present a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers. The framework separates target-conditioned motion from local digging: a shared task-conditioned RL policy controls waypoint-guided approach and loaded transport, while an IL policy learns vision-based digging and lifting from expert demonstrations. Digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints for motion control. The control architecture coordinates the learned policies and deterministic unloading through a shared motion interface. The complete system is deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control. Offline replay and physical experiments demonstrate more consistent target selection, shorter local motion time, and increased payload compared with the respective baselines. The learned digging policy achieves a mean payload of 6.52 kg per completed cycle, compared with 2.68 kg for Fixed Dig. Three five-scoop runs further demonstrate consecutive autonomous excavation under continuously changing pile geometry.</p><p><em>Comment: 8 pages, 7 figures, 4 tables</em></p>]]></description>
  </item>
  <item>
    <title>AI-based detection of worsening heart failure from low-resolution telemonitoring data</title>
    <link>http://arxiv.org/abs/2609.29742v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29742v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:57:12 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Erik Aerts, Yinan Yu, Annika Rosengren, Michael Fu, Martin Lindgren, Falk Dippel, Martin Adiels, Helen Sjöland</p><p>Objective: Heart failure (HF) presents a healthcare challenge due to its high comorbidity burden, aging patient population and frequent hospitalizations. Remote monitoring offers a promising approach to managing HF patients by early detection of health deterioration. Developing autonomous systems to detect signs of worsening in telemonitoring data is of interest to reduce the workload of healthcare personnel. Methods: We propose the TRACER model, a Transformer with Contrastive Event Representation, designed to predict timelines leading to rare hospitalization events in low-resolution and irregularly sampled telemonitoring data. TRACER incorporates time-aware embeddings for each biomarker, contrastive pre-training to enhance anomaly detection via representation learning, and independent binary classifiers for detection. We used measurement data containing remotely recorded biomarker sequences from 276 HF patients segmented into overlapping windows based on temporal rules, and labeled the windows based on the occurrence of HF relevant hospitalizations at the latter edge of the window. Results: TRACER was able to correctly predict 66.7% timelines leading up to HF hospitalizations in the highly imbalanced real-world dataset with an overestimation of 7.9%. Reformulating the training of TRACER as an event detection problem improved the predictive performance compared with training directly on forecasting windows, enabling more effective use of the limited hospitalization events. Conclusion: TRACER demonstrated superior performance in detecting signs of worsening status in real-world telemonitoring data compared to the other tested models. Significance: TRACER shows promise in identifying signs of clinical deterioration that allow for alerts to be generated to provide counteractive treatment in patients with HF.</p><p><em>Comment: 12 pages, 5 figures, under review for publication</em></p>]]></description>
  </item>
  <item>
    <title>TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening</title>
    <link>http://arxiv.org/abs/2609.29740v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29740v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:56:57 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer</p><p>Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols. We introduce TopU-LBVS, a multi-target benchmark for LBVS under hard-negative screening conditions. Starting from curated ChEMBL~35 bioactivity data, TopU-LBVS covers 93 protein targets across 7 protein classes and constructs target-specific screening libraries with property-matched, structurally similar decoys at a fixed 1:40 active-to-decoy ratio. Libraries contain roughly 400 to 10,000 compounds and are designed to reduce simple physicochemical and nearest-neighbor fingerprint shortcuts.
  TopU-LBVS provides three fixed protocols. TopU-LBVS-full evaluates ChEMBL$^\ast \rightarrow$ TopU generalization across all 93 targets. TopU-LBVS-low evaluates low-data TopU $\rightarrow$ TopU learning within the hard-negative distribution. TopU-LBVS-mini provides a compact seven-target protocol with a paired random-decoy control that changes only the test decoys, enabling low-cost development and direct measurement of the gap between random ChEMBL$^\ast$ and TopU decoys. Across ten reference baselines spanning fingerprint methods, molecular GNNs, fingerprint hybrids, and modern molecular models, performance under random-decoy evaluation degrades sharply under hard-negative screening. We release data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of future LBVS and molecular representation learning methods.
  Code and data are available at https://github.com/topu-benchmark/topu-lbvs and https://huggingface.co/datasets/topu-benchmark/topu-lbvs.</p><p><em>Comment: 75 pages</em></p>]]></description>
  </item>
  <item>
    <title>Combining Evasive and Braking Reactions for Safety Reference Models in Automated Vehicles</title>
    <link>http://arxiv.org/abs/2609.29738v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29738v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:55:28 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Riccardo Donà, Konstantinos Mattas, Biagio Ciuffo</p><p>Computational models of careful and competent human drivers are essential for scenario-based evaluation of automated driving systems (ADS). However, most existing safety reference models primarily focus on longitudinal braking, neglecting the role of evasive steering in human collision avoidance. This paper proposes a hybrid Fuzzy-Safety Model (FSM-H) that integrates longitudinal mitigation and lateral avoidance within a unified behavioral framework. The braking component is governed by Proactive Fuzzy Safety (PFS) metrics, representing the erosion of longitudinal safety margins, while the steering component is driven by Criticality Fuzzy Safety for lane-change (CFS-LC), capturing lateral conflict severity and maneuver feasibility. A finite-state architecture models the sequential escalation from nominal driving to braking and, when necessary, to evasive steering, incorporating perception-reaction time and lane-check delays to reflect human decision processes. The model is evaluated in reconstructed high-criticality cut-in scenarios and compared with braking-only and steering-only reference strategies. Results show that the hybrid approach expands the preventability envelope while maintaining behavioral plausibility and computational tractability. The proposed framework provides a transparent and explainable human reference model suitable for simulation-based ADS safety benchmarking and regulatory assessment.</p><p><em>Comment: 6 pages, 7 figures, Accepted for publication at the IEEE International Conference on Intelligent Transportation Systems (ITSC), 2026</em></p>]]></description>
  </item>
  <item>
    <title>C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks</title>
    <link>http://arxiv.org/abs/2609.29735v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29735v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:52:27 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xueshu Chen, Yan Wang, Zihao Xue, Jiefu Li, Zhenfang Liu, Jayden Chen, Zhen Bi, Jungang Lou</p><p>Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that maintains a bounded active index over persistent source text-image evidence. Relation-aware updates consolidate safe redundancy while preserving complementary and incompatible records. At query time, budgeted routing selects useful index pages and expands their associated source evidence under a fixed reader budget. Together, these mechanisms establish a compact, provenance-preserving multimodal memory organization for cross-session long-horizon tasks, retaining temporal distinctions and source links required for reliable downstream reasoning. Code is available at https://github.com/HuzhouNLP/C3M.</p>]]></description>
  </item>
  <item>
    <title>TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)</title>
    <link>http://arxiv.org/abs/2609.29733v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29733v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:50:47 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler</p><p>Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles. While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce $\texttt{CLASP-Ar}$, which reformulates the task as cloze-style masked language modeling. In this approach, the target, predicted sentiment, and text are combined into a single prompt whose $\texttt{[MASK]}$ prediction is restricted to a verbalizer-constrained label vocabulary.</p><p><em>Comment: Accepted at ArabicNLP 2026 StanceEval-2026 shared task</em></p>]]></description>
  </item>
  <item>
    <title>The Gold in Bias: Maturing the AI Design Process through Verification</title>
    <link>http://arxiv.org/abs/2609.29730v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29730v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:46:17 +0000</pubDate>
    <category>Generative POI Search and Recommendation</category>
    <description><![CDATA[<p><strong>Authors:</strong> Samira Maghool, Paolo Ceravolo</p><p>Bias in AI systems is typically framed as a flaw to be minimized, yet it also serves as a critical indicator of underlying weaknesses in data, modeling assumptions, and system design. Existing approaches often treat bias as an isolated problem rather than as evidence that can strengthen verification and governance across the AI lifecycle. This paper aims to reconceptualize bias as a diagnostic tool that supports rigorous AI verification. We seek to develop a multidimensional framework to analyze bias, demonstrate how biases emerge in both Traditional and Generative AI, and provide a structured pathway for verification-driven mitigation. We present a multidimensional framework analyzing bias across four dimensions: origin sources, emergence points throughout the AI modeling lifecycle, technical and methodological causes, and validation approaches for detection and mitigation. Through a comprehensive typology spanning traditional and generative AI systems, we demonstrate how biases manifest and propagate across development stages. Our analysis encompasses 30 distinct bias types, 16 verification methods, and 20 countermeasures, providing an actionable roadmap for practitioners. We introduce a hierarchical evidence framework that distinguishes internal validity (mechanistic integrity of AI systems) from external validity (contextual reliability in deployment environments). The framework reveals how biases manifest and propagate across modeling stages, enabling systematic mapping between bias types, verification techniques, and effective countermeasures. The proposed evidence hierarchy clarifies how different verification strategies contribute to mechanistic integrity and contextual reliability. We advocate for &#x27;&#x27;Ethics by Design&#x27;&#x27; principles that integrate bias verification throughout the development lifecycle, enabling the construction of fairer, more robust, and trustworthy AI systems.</p>]]></description>
  </item>
  <item>
    <title>A Multimodal Dataset for Survival Prediction in Resected Pancreatic Ductal Adenocarcinoma</title>
    <link>http://arxiv.org/abs/2609.29726v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29726v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:41:42 +0000</pubDate>
    <category>NeRF</category>
    <description><![CDATA[<p><strong>Authors:</strong> Anh-Tien Nguyen, Mawuko Tettey, Jacqueline Michelle Metsch, Teresa Zimmer, Niklas Ullrich, Mario Duker, Sandra Rungeling, Kirsten Reuter-Jessen, Tessa Rosenthal, Lena-Christin Conradi, Michael Ghadimi, Alexander Konig, Elisabeth Hessmann, Volker Ellenrieder, Philipp Strobel, Hanibal Bohnenberger, Anne-Christin Hauschild</p><p>Survival research in pancreatic ductal adenocarcinoma (PDAC) is limited by the scarcity of datasets linking whole-slide histology with clinical, molecular, and long-term outcome data. We present a retrospective single-centre cohort of 302 patients who underwent PDAC resection at University Medical Center Gottingen. The dataset comprises 446 H&amp;E whole-slide images, clinicopathological variables, targeted sequencing data for 154 patients, and overall-survival outcomes. During follow-up, 253 patients died, and the median follow-up was 76 months.
  To establish initial reference values, we evaluated fourteen survival-prediction configurations using identical five-repetition Monte Carlo cross-validation partitions. Ridge Cox regression using numeric clinicopathological variables achieved a mean concordance of $0.649 \pm 0.042$ and $0.652 \pm 0.046$ after adding KRAS and TP53 mutation status. The image-only attention model achieved $0.603 \pm 0.030$, while multimodal fusion achieved $0.619 \pm 0.025$, the highest concordance among the neural models. These results establish promising initial benchmarks for future research using this pancreas-specific multimodal dataset, paving the way for external validation.</p>]]></description>
  </item>
  <item>
    <title>SALI: Shot-Aware Late Interaction for Cross-Shot Relation Matching in Text-to-Video Retrieval using Film-Grammar Knowledge</title>
    <link>http://arxiv.org/abs/2609.29721v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29721v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:39:40 +0000</pubDate>
    <category>Information Retrieval</category>
    <description><![CDATA[<p><strong>Authors:</strong> Toya Oyama, Rainer Lienhart, Shin&#x27;ichi Satoh</p><p>Text-to-video retrieval usually represents a video clip by a single embedding. This embedding often loses important relations between people. E.g., an interaction &quot;Anna confronts Mark&quot; is regularly filmed as alternating shot and reverse shot of both (Fig. 1a). No single shot or averaged embedding over clip shots captures this relation. Thus, we propose SALI (Shot-Aware Late Interaction). It extracts the subject and object from a single-sentence query, and matches the query, its subject and object text embeddings against each visual shot embedding of a video clip. The matching operator is greedy max or optimal transport. A film-grammar penalty in fine-tuning adds a small, consistent shift. Built on CLIP4Clip-meanP, SALI keeps overall recall on par on Condensed Movies and ActivityNet while raising R@1 on multi-shot relation queries by 3 and 12 points, the most among all compared methods, and improves such queries on MSR-VTT at a cost of 1.4 R@1 overall.</p><p><em>Comment: 5 pages, 2 figures, 4 tables. Submitted to ICASSP 2027</em></p>]]></description>
  </item>
  <item>
    <title>Markerless Multi-Modal Autonomous Robotic Inspection of Large Space Structures</title>
    <link>http://arxiv.org/abs/2609.29644v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29644v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:37:21 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Juan De Dios Alfaro, Arturo Ríos, David Rodríguez-Martínez, Carlos Pérez-del-Pulgar</p><p>Future orbital infrastructures, such as deployable antennas, solar farms, and large orbital platforms will require autonomous inspection systems able to operate with limited prior knowledge and without cooperative markers. Current on-orbit servicing approaches often rely on predefined trajectories, standard interfaces, fiducial markers or accurate target models, which limits scalability for large, heterogeneous or partially unknown structures. This paper presents a markerless autonomous robotic inspection pipeline in which 3D reconstruction is used as an inspection-support representation. The system integrates a Kinova Gen2 manipulator with an end-effector-mounted multimodal sensor head composed of an RGB-D camera, a thermal camera and a 2D LiDAR. The pipeline estimates an approximate inspection volume, generates viewpoints, plans collision-free motions with MoveIt, and synchronously records RGB-D images, thermal data, and robot poses in ROS2. Candidate reconstruction methods were evaluated to select a practical method for this pipeline, with Nerfacto used for geometric reconstruction and Thermal-Nerfacto used to demonstrate thermal-aware rendering for inspection. Validation in a Gazebo-based simulator and preliminary laboratory tests reveal that the proposed system can autonomously acquire spatially coherent inspection data and produce reconstructions suitable for visual and geometric assessment, representing a step towards inspection of large non-cooperative space structures.</p><p><em>Comment: 7 pages, 5 figures, accepted conference paper</em></p>]]></description>
  </item>
  <item>
    <title>TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification</title>
    <link>http://arxiv.org/abs/2609.29633v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29633v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:37:06 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ali Abusaleh, Bhuvanesh Verma, Alexander Mehler</p><p>We present TTLab&#x27;s submission to the AlexandriaX-2026 Subtask~3 on Arabic MT error span detection and classification. Our system frames the task as token-level classification over surface forms, preserving character offsets to ensure exact alignment with the evaluation metric. To handle severe label imbalance, we employ a focal loss with class weighting and dialect-specific decoding thresholds. Among six Arabic pre-trained encoders, MARBERTv2 achieves the best overall performance of 40.8 and 40.91 on the development and test set, respectively, ranking $\nth{3}$ out of all participating teams. While our system localizes error spans effectively, classification of rare error types remains challenging, highlighting the need for data augmentation for tail categories. The code is available at ${\href{https://github.com/ENTAILab/arabic-dialectal-mt-error-span-detection}{\faGithub~ TTLab at AlexandriaX-2026}$</p><p><em>Comment: Accepted at ArabicNLP 2026, shared task AlexandriaX-2026</em></p>]]></description>
  </item>
  <item>
    <title>CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding</title>
    <link>http://arxiv.org/abs/2609.29474v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29474v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:31:05 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina</p><p>Public software repositories, like GitHub and Software Heritage Archive, store billions of files, yet extracting their implicit engineering knowledge ---i.e., the algorithms they implement, the paradigms they follow, the patterns they instantiate, and the application domains they serve--- remains challenging, as current tools are constrained to syntactic and token-level analysis. We present a pipeline for building an open-taxonomy semantic annotation of source code using a code-specialised Large Language Model. The extracted entities are grounded in Wikidata through a three-stage linking procedure: a deterministic SPARQL stage handles unambiguous entities, a Deep Research Agent resolves the residual long tail, and a hierarchy-rollup stage imports the parent-of closure of each resolved Wikidata identifier. The resulting annotations are materialised as a source-code-specific open-taxonomy knowledge graph. We further introduce a calibrated quality-assurance protocol that quantifies annotation precision by combining a small human gold set with an LLM-as-a-judge filter. We applied our pipeline to the 167 million files of the Stack-Edu corpus, creating the first known large-scale open-taxonomy knowledge graph for source code. Our graph, named CodeGraph, contains approximately 158 million nodes, which include around 145 million files, about 63,000 extracted concept entities (such as algorithms, paradigms, design patterns, and application domains), and roughly 19,800 grounded Wikidata entities. Furthermore, CodeGraph features approximately 1 billion typed edges that connect files to their respective concepts, link these concepts to their grounded Wikidata identifiers, and relate them to their parent categories, covering 14 programming languages.</p><p><em>Comment: Accepted at CIKM 2026</em></p>]]></description>
  </item>
  <item>
    <title>Direct Message Approximation (DMA): A Consistency-Based Framework for Tractable Approximate Inference on Factor Graphs</title>
    <link>http://arxiv.org/abs/2609.29466v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29466v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:24:43 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ralf Herbrich, Rainer Schlosser, Jan Lemcke, Johann Ukrow, Anna Kazachkova, Nicolas Alder, Leonhard Hennicke, Theo Bardey, Nico Grimm, Luca Kleinschmidt, Philipp Kolbe, Cezary Kujath, Johanna Schlimme, Karl Matti Schütz</p><p>Approximate message passing on factor graphs underlies two dominant families of probabilistic inference algorithms: expectation propagation (EP) and variational message passing (VMP). Both methods approximate the marginal at each factor edge, forcing an iterative round-robin schedule, risking negative-precision messages, and, for VMP, collapsing to point estimates at Dirac-delta factors. We introduce Direct Message Approximation (DMA), which approximates factor-to-variable messages directly rather than the marginal. For normalisable factors, we define a consistency condition (requiring exactness when all other incoming messages are Dirac deltas) to guide message construction. We prove a master theorem (proper messages, any graph) bounding marginal KL from message KL, with three structural corollaries: Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages. Further, we prove a complementary $O(1/r^2)$ guarantee for the inherently improper backward message of the product factor, whose closed-form treatment has resisted prior work. As a concrete instantiation, we derive explicit DMA messages for the product and leaky-ReLU factors and assemble a Bayesian neural network (BNN) inference algorithm with one forward/backward sweep per training example and no gradient learning-rate hyperparameter, validating that the structural guarantees translate to predictive uncertainty that widens in data-sparse regions, including under model mismatch.</p><p><em>Comment: Submitted to ICLR 2027</em></p>]]></description>
  </item>
  <item>
    <title>Precise Convergence Speed of Clipped SGD</title>
    <link>http://arxiv.org/abs/2609.29458v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29458v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:11:42 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> David A. R. Robin</p><p>We present a tightened convergence analysis of clipped gradient descent on $(L_0, L_1)$-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal the central role of a control of the bias derived from fundamental properties of $\ell_2$-projection, simplifying proofs. We also extend the domain of validity from $η\leq 1 / (9 β)$ to $η&lt; 1 /β$ where $β= L_0 + c L_1$ for clipping constant $c$, which matches the more traditional analysis of smooth functions. We strengthen the convergence criterion from $\left( \min_{t &lt; T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right)$ to $\left( \frac{1}{T} \sum_{t &lt; T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right)$ with matching speed, and lower the final achievable loss from $\mathcal{O}(\min(σ^2/c, σ))$ to the more precise $6 \min(σ^2 /c, 3 σ)$.</p>]]></description>
  </item>
  <item>
    <title>Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception</title>
    <link>http://arxiv.org/abs/2609.29456v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29456v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:10:17 +0000</pubDate>
    <category>3D Vision</category>
    <description><![CDATA[<p><strong>Authors:</strong> Melih Yazgan, Timon Müller, J. Marius Zöllner</p><p>Collaborative perception improves autonomous perception by sharing intermediate Bird&#x27;s-Eye-View (BEV) features across connected agents, but dense feature exchange is difficult to deploy under strict Vehicle-to-Everything (V2X) bandwidth limits. Existing efficient methods typically either compress the full feature map uniformly, spending bits on low-value background, or sparsify communication, risking the loss of useful context. We propose a coverage-refinement design for byte-constrained cooperative perception: each agent transmits a highly compressed coarse layer over the full BEV map and allocates the remaining budget to selected high-resolution patches. A Task-Aware Benefit Selector ranks cells by estimated downstream utility, enabling deterministic budgeted refinement and zero-retraining adaptation to changing bandwidth. The receiver reconstructs a dense BEV tensor compatible with standard fusion modules. Experiments on DAIR-V2X and OPV2V show strong accuracy-payload trade-offs at kilobyte-scale budgets. On DAIR-V2X, our method reaches 0.60 AP@0.7 at only 1.87 KB per non-ego agent, compared with 0.52 at 4.61 KB for uniform SimVQ compression. Controlled diagnostics further show that the gain arises from coverage-refinement allocation rather than quantization alone. Code will be published.</p><p><em>Comment: Accepted at WACV 2027 (first-round acceptance)</em></p>]]></description>
  </item>
  <item>
    <title>A Systematic Multi-Domain Evaluation of Document Retrievers</title>
    <link>http://arxiv.org/abs/2609.29455v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29455v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:10:14 +0000</pubDate>
    <category>Information Retrieval</category>
    <description><![CDATA[<p><strong>Authors:</strong> Valentin Velev, Andreas Spitz</p><p>Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks, domains, or model families. This fragmentation makes it difficult to draw reliable conclusions about the relative strengths, weaknesses, and trade-offs of document retrievers. To address this gap, we conduct a large-scale empirical evaluation of document retrievers, covering three families (sparse, dense, and expansion-based) and evaluating 33 retrievers across seven IR datasets, analyzing retrieval quality, runtime, and failure points. Rather than tuning each model individually, we evaluate every retriever off the shelf, under the configuration reconstructable from its public documentation and a uniform compute budget. Our results show that NV-Embed-v2 achieves the strongest performance on four of the seven datasets, albeit at the cost of substantial query latencies. Among sparse retrievers, we find that SPLADE-v3 rivals the top-performing approach despite much lower latency, and even achieves top scores on MS MARCO. On instruction-following datasets, GritLM delivers the best performance. Finally, an analysis of the retrievers&#x27; failure points reveals contrasts between models and families that indicate potential for unrealized gains in retrieval performance.</p>]]></description>
  </item>
  <item>
    <title>Decoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores</title>
    <link>http://arxiv.org/abs/2609.29453v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29453v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:09:24 +0000</pubDate>
    <category>Information Retrieval</category>
    <description><![CDATA[<p><strong>Authors:</strong> Sam Urmian, Qinyi Liu, Mohammad Khalil</p><p>We formalize slate recommendation as a randomized score learner followed by deterministic selection. First, an appropriately scoped differential-privacy guarantee passes through selection and its audit trace by post-processing. End-to-end privacy holds only when selector inputs are public or independent, previous private outputs, or separately privacy-accounted; fixing raw state or candidate information instead yields only a conditional guarantee. Second, we derive a logged margin certificate: bounded score-induced objective movement below half the smallest greedy decision margin guarantees that the ordered slate is unchanged.
  Controlled fixed-margin tests show near-linear exponent scaling, with an empirical slope of $-0.220$ (95% CI $[-0.231,-0.210]$) against the independent-noise reference $-1/4$. Real-anchor experiments on OULAD, MovieLens-25M, and Amazon Musical Instruments show that greater anchor weight reduces score-noise-induced ranking churn. OULAD and EdNet certificate checks validate the implementation of the logged inequality, while closed-loop simulations show bounded target drift and setting-dependent downstream utility. The contribution is therefore a privacy-scope contract and a certifiable score-to-slate stability mechanism, not a universal utility claim.</p><p><em>Comment: 20 pages including supplementary appendix. Accepted at ACM RecSys 2026</em></p>]]></description>
  </item>
  <item>
    <title>YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech</title>
    <link>http://arxiv.org/abs/2609.29448v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29448v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:06:46 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> William Chen, Shinnosuke Takamichi, Sayaka Shiota, Satoru Fukayama, Samuele Cornell, Shinji Watanabe</p><p>We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio. We first provide the collection methodology for the corpus, where we introduce new techniques for gathering language-balanced speech data. The effectiveness of our approach is shown by the language distribution of the crawled data: 22 languages in YODAS v3 have over 10K hours and 73 languages have over 5K hours of data. We then conduct extensive analyses on the composition of the data, such as the distribution of languages, audio quality, and transcription quality. Finally, we train baseline speech recognition and neural codec models to show the effectiveness of the dataset. Download at https://huggingface.co/datasets/espnet/yodas3.</p><p><em>Comment: Interspeech 2026; 6 Pages</em></p>]]></description>
  </item>
  <item>
    <title>Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed</title>
    <link>http://arxiv.org/abs/2609.29447v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29447v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:05:00 +0000</pubDate>
    <category>Gaussian Splatting</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ignat Romanov, Andreas Hadjipieris, Neofytos Dimitriou</p><p>Smart maritime infrastructures provide continuous access to heterogeneous sensing streams, enabling repeated experimentation, digital-twin development, and AI-based maritime services. However, sensing hardware alone is not sufficient for scene-specific model development: historical video streams must also be spatially indexed, contextualized, and reduced to informative subsets for annotation. This paper presents a frame-to-panorama localization and context-aware sampling pipeline for ship detection in historical PTZ maritime video lacking reliable pan, tilt, and zoom metadata. The main contribution is an end-to-end data-curation approach that recovers camera-view information from historical PTZ video and combines it with environmental context and visual diversity to construct compact, scene-specific training sets. Specifically, frames are localized on a reference panorama using SuperPoint and LightGlue, enriched with weather and solar-state metadata, and selected through diversity sampling to preserve variation across camera view and environmental conditions. A second context-aware stage targets under-represented distant-vessel cases near the horizon using tile-level visual embeddings and Gaussian Mixture Model clustering. Applied within the CMMI MDigi-I Smart Marina testbed, the proposed pipeline reduces 40,718 candidate frames to 220 images for annotation, corresponding to a 99.5% reduction. A YOLO26-m detector fine-tuned on this subset achieves a mean AP50 of 94.78% $\pm$ 0.51% and a mean AP50-95 of 75.10% $\pm$ 1.73% under sequence-grouped five-fold cross-validation. These results demonstrate that highly redundant infrastructure video streams can be transformed into compact, spatially and contextually diverse training sets for scene-specific detector adaptation while substantially reducing annotation effort.</p>]]></description>
  </item>
  <item>
    <title>SPADE-DFL: Communication-Efficient Decentralized Federated Learning via Derivative-Free Linearized ADMM</title>
    <link>http://arxiv.org/abs/2609.29446v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29446v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:03:22 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Mengli Wei, Mengkai Zhu, Jiawen Chen, Wenwu Yu, Duxin Che</p><p>Reducing communication in derivative-free decentralized learning requires controlling the disagreement accumulated over multiple local updates. This paper develops SPADE-DFL, a primal--dual method that allows the number of local function-value updates between neighbor exchanges to grow with the computation budget while preserving the nonprivate convergence order. For smooth nonconvex objectives under uniform query-moment bounds, the prescribed nonprivate schedule achieves a time-averaged stationarity and consensus bound of $\mathcal{O}(T^{-1/3})$ using only $Θ(T^{2/3})$ communication rounds, where $T$ is the number of local updates per client. For private training, the accumulated data-dependent increment is isolated from the graph correction, allowing one protected state per client and round to generate all outgoing messages. We prove client-level differential privacy for the full interactive transcript and quantify the resulting optimization error over a finite horizon. Experiments on four classification tasks show that SPADE-DFL achieves higher mean test accuracy than existing decentralized learning methods.</p>]]></description>
  </item>
  <item>
    <title>IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis</title>
    <link>http://arxiv.org/abs/2609.29444v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29444v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:02:53 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xingyu Wu, Yuchen Yan, Zhengxi Lu, Siqi Chen, Xin ZHANG, Aiting Liu, Chao Deng, Jie Liu, Jin Ma, Jian Shao, Jun Xiao, Yongliang Shen</p><p>Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state. This design separates planning from synthesis while using the summary as the persistent state of search, reducing both capability coupling and context noise. To train IterSynth effectively, we further introduce Role-Decoupled Policy Optimization (RDPO) for reinforcement learning, which combines terminal outcome rewards with turn-level rubric evaluations and computes role-specific advantages for more precise credit assignment. Experiments on five long-horizon deep-search benchmarks such as BrowseComp and Xbench-DS show that IterSynth-8B achieves an average score of 50.7, surpassing the strongest prior $\leq$8B agent by +4.2\%. Moreover, IterSynth serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.</p><p><em>Comment: Code: https://github.com/Tencent/IterSynth</em></p>]]></description>
  </item>
  <item>
    <title>Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure</title>
    <link>http://arxiv.org/abs/2609.29445v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29445v1</guid>
    <pubDate>Thu, 24 Sep 2026 12:02:53 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Fardeen Sadab, Adib Sakhawat</p><p>We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annotators as a random rather than a fixed factor, no system differs significantly from any other ($F(7,14)=0.59$, $p=0.76$), although the conventional analysis declares 19 of 28 pairwise differences significant. Annotator identity explains far more rating variance than system identity, and the winning system changes whenever any single annotator is removed. The ordering that does emerge tracks output length: mean emoji count explains 78.7\% of between-system variance, and a within-item length-matched comparison over 2,599 pairs reverses the leaderboard. We further show that cross-provider anisotropy differences vanish under mean-centring, that per-language token costs change sign with the normalising unit, and that multi-view row-wise splits inflate macro-F1 by $3.1$ points and change the top-ranked system. In place of preference scoring we propose **emoji-affect decodability**, a reference-based probe whose rankings are stable to $\pm0.003$ macro-F1 across seeds.</p><p><em>Comment: 10 pages, 3 figures, accpeted in 6TH MULTILINGUAL REPRESENTATION LEARNING (MRL) WORKSHOP 2026 at EMNLP 2026 in Budapest, Hungary</em></p>]]></description>
  </item>
  <item>
    <title>Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures</title>
    <link>http://arxiv.org/abs/2609.29429v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29429v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:49:30 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang</p><p>Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark&#x27;s scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user&#x27;s belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question&#x27;s wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer&#x27;s agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.</p>]]></description>
  </item>
  <item>
    <title>agentic-ger: terminology recovery in long-form speech using global context</title>
    <link>http://arxiv.org/abs/2609.29428v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29428v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:46:57 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yanqiao Zhu, Wupeng Wang, Zhifu Gao, Xiangang Li, Xie Chen</p><p>Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agentic-GER, an LLM-based agent for terminology correction in long-form speech. The agent uses global context from the full transcript to identify suspicious terms and resolve ambiguous hypotheses. It selectively re-transcribes the source speech to check candidate corrections, and uses accepted edits to guide subsequent decisions. Experiments with four LLMs and two ASR systems on GigaSpeechBench show consistent terminology improvements in both Chinese and English, with and without thinking. On Chinese speech, Agentic-GER achieves up to a 36.8% relative reduction in biased character error rate (B-CER) over the Whisper baseline.</p><p><em>Comment: submitted to ICASSP 2027</em></p>]]></description>
  </item>
  <item>
    <title>Coupled State-Space Modelling, Control, and Policy Distillation for Hybrid Rigid-Pneumatic Manipulators</title>
    <link>http://arxiv.org/abs/2609.29424v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29424v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:45:23 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Alan Royce Gabriel Samuel, Pulkit Verma</p><p>Hybrid manipulators combine motorized rigid joints with pressure-actuated origami segments. Published arms of this kind are controlled with decoupled per-DOF loops, and the cost of this approximation has not been quantified, because the coupled model needed to measure it has not been built. This paper derives such a model for a chain of $N$ alternating revolute joints and Kresling origami segments, including pneumatic chamber dynamics and crease hysteresis. Using the model, we measure the coupling directly and show that its strength varies joint by joint, and that decoupled control loses precisely on the strongly coupled joints while remaining competitive on the one nearly decoupled joint. Coupled model-based controllers track $2.5\times$ tighter than a decoupled PID baseline at lower torque. However, the model predictive controller (MPC) is too slow for real time, and model-free reinforcement learning stalls far below acceptable success rates on a strict settling metric. We therefore distill the MPC into a small neural policy with behavior cloning and DAgger. The distilled policy settles 93-94$\%$ of goals with zero collisions, within a few points of its teacher, and runs inside the 5 ms control step where the MPC does not. Where the teacher itself fails, we trace the failure to a limit cycle with the bellows&#x27; lightly damped mode, and we remove it by selecting goal postures holdable at low pressure.</p><p><em>Comment: 8 pages, 4 figures, ICRA</em></p>]]></description>
  </item>
  <item>
    <title>Temperament Engineering: Designing Strategic Behavioural Diversity in Robot Swarms</title>
    <link>http://arxiv.org/abs/2609.29423v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29423v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:45:17 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Edmund R. Hunt</p><p>No two robots are truly identical: calibration, battery state, sensor drift and wear give every swarm a distribution of behaviour rather than a single point, usually treated as an imperfection to be minimised. In animal collectives the reverse holds: consistent individual differences in behaviour (&#x27;temperament&#x27;) are shaped by natural selection and often decisive for group performance. This perspective proposes &#x27;temperament engineering&#x27;, a bio-inspired framework that treats the swarm&#x27;s distribution of temperaments, rather than the individual controller, as the design object. It borrows five evolutionarily validated axes of animal temperament (shyness-boldness, exploration-avoidance, activity, aggressiveness and sociability) as a design vocabulary, rendering each as a continuous control parameter $τ\in [0,1]$ above the controller, realisable as a module threshold, a policy-conditioning vector in multi-agent reinforcement learning, or a constraint on a foundation-model planner. A three-phase workflow maps mission success criteria onto relevant axes, plans the shape of the $τ$ distribution, and tunes reaction norms governing how temperament responds to environmental cues. The payoff is greatest under decentralisation: where a central planner can reassign behaviour online, a temperament distribution is a planner output, but in a swarm without global knowledge it must be an offline, anticipatory design input. Behavioural and platform heterogeneity are thereby co-design variables, and I sketch tentative robot-native axes (self-model plasticity, forcefulness, initiative and expressiveness) arising from features robots have and animals do not. Engineered heterogeneity has been shown to outperform homogeneous swarms in tasks such as aggregation and exploration; establishing when, and how much, heterogeneity repays its cost is the work the field can now take forward.</p>]]></description>
  </item>
  <item>
    <title>Rufus-Air: An Open LLM Post-Training Recipe</title>
    <link>http://arxiv.org/abs/2609.29421v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29421v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:45:03 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He, Hongye Jin, Linwei Li, Shiyang Li, Fenglin Liu, Xin Liu, Priyanka Nigam, Haoyang Wen, Zhenghao Xu, Zhuocheng Xu, Bing Yin, Qingyu Yin, Chao Zhang, Rongzhi Zhang, Zhihan Zhang, Zixuan Zhang, Zixuan Zhang, Tuo Zhao</p><p>Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) difficulty filtering keeps RL prompts within a productive learning range; (iii) reward reliability provides a practical principle for ordering stages; and (iv) infrastructure and engineering choices are part of the recipe, not just an implementation detail. Rufus-Air improves over the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.</p><p><em>Comment: 47 pages, 9 figures, 20 tables. Authors are listed alphabetically by surname; all contributed while at Amazon. The two authors named Zixuan Zhang are different people</em></p>]]></description>
  </item>
  <item>
    <title>UCON: Uncertainty-aware Navigation with Historical Re-association in Dynamic Environments</title>
    <link>http://arxiv.org/abs/2609.29419v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29419v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:41:24 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Bing Sun, Yue Lin, Yongsheng Yuan, Yang Liu, Dong Wang, Huchuan Lu</p><p>Autonomous navigation in dynamic environments is hindered by two fundamental challenges: perception instability and uncertainty-optimization mismatch. The former leads to identity switches and unreliable motion estimation, while the latter prevents principled incorporation of motion uncertainty into trajectory optimization. To address these challenges, we propose UCON, an uncertainty-aware navigation algorithm in dynamic environments. For perception instability, we present a point-level historical re-association mechanism that leverages historical point cloud fragments to recover lost targets while maintaining identity continuity. Subsequently, a Kalman filter is employed to provide anisotropic motion state estimation and covariance propagation. To resolve the uncertainty-optimization mismatch, we transform predicted states and their covariances into uncertainty sectors, which are embedded as differentiable cost terms within a trajectory optimization framework. This achieves consistent uncertainty-aware dynamic obstacle avoidance while maintaining smoothness and feasibility. Extensive simulations and real-world experiments demonstrate that, while maintaining high computational efficiency, UCON achieves superior perception stability and robust navigation performance in dynamic environments compared to state-of-the-art methods. The code will be open-sourced to facilitate further research.</p><p><em>Comment: Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)</em></p>]]></description>
  </item>
  <item>
    <title>Singularity Analysis for the Perspective-Four and Five-Line Problems</title>
    <link>http://arxiv.org/abs/2609.29417v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29417v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:38:40 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Jorge García Fontán, Abhilash Nayak, Sébastien Briot, Mohab Safey El Din</p><p>This paper deals with image-based visual servoing and pose estimation by observing four and five lines. Our main interest is to determine the relative configurations of the camera and the observed lines that lead to problems in control and stability. Since it is equivalent to finding the singularities of the corresponding Jacobian matrix, we use tools from computational algebraic geometry to seek configurations such that all of its minors vanish simultaneously. By choosing a suitable basis for this matrix, we revisit the problem in the case of three lines to show that one type of the singularities is when the camera lies on the hyperboloid of one sheet uniquely defined by the lines. This result is further exploited to prove that the one-dimensional singularities, if any, in the case of $n$ lines appear when the camera lies on the transversals to the observed lines. Thus, by forcing the transversals to be complex, we can avoid the aforementioned type of singularities in the case of four lines although the algebra shows that there can always be up to 10 inevitable singular locations of the camera for the other type of singularity. For five lines, we find out that there are no singularities in the generic case. The singularities are also characterized for four and five lines with orthogonality and parallelism constraints. Furthermore, a visual servoing library is used to conduct some simulated experiments to substantiate the theoretical results. As expected, we observe problems in control in the vicinity of a singularity as well as increased errors in pose estimation.</p>]]></description>
  </item>
  <item>
    <title>Controlling Backchannels in Streamable Full-duplex Models</title>
    <link>http://arxiv.org/abs/2609.29418v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29418v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:38:40 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Maike Züfle, Peter Polák, Sefik Emre Eskimez, Jan Niehues, Peter Bell, Ondřej Klejch</p><p>Backchannels, brief acknowledgements like &quot;uh-huh&quot; produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model&#x27;s own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.</p>]]></description>
  </item>
  <item>
    <title>Neural Transport Nested Sampling</title>
    <link>http://arxiv.org/abs/2609.29413v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29413v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:34:31 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> David Yallup, Will Handley</p><p>Sampling from Boltzmann distributions of molecular systems is an inference problem that has seen significant recent developments fuelled by advances in neural density estimation. We develop a novel sampling algorithm, Neural Transport Nested Sampling (NTNS), which combines the classical strengths of nested sampling with modern neural flow-based methods. NTNS uses a flow matching velocity as the drift in a Metropolis--Hastings corrected Langevin kernel inside a nested sampling outer loop, requiring only evaluations of the target energy function and providing scalable estimation of the full partition function of high-dimensional particle systems. We benchmark NTNS on challenging molecular sampling benchmarks, scaling up to Lennard--Jones clusters of 55 interacting particles, where it reduces both interatomic distance and energy Wasserstein errors to reference MCMC by over an order of magnitude relative to the strongest neural baselines at lower wall-clock cost. To our knowledge, NTNS is also the first neural sampler to return a calibrated, temperature resolved partition function estimate at this scale, recovering the phase structure across temperature from a single run.</p><p><em>Comment: 26 pages, 9 figures</em></p>]]></description>
  </item>
  <item>
    <title>Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?</title>
    <link>http://arxiv.org/abs/2609.29410v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29410v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:33:36 +0000</pubDate>
    <category>Computation and Language</category>
    <description><![CDATA[<p><strong>Authors:</strong> Alexandru Stefan Stoica, Traian Rebedea, Marian Cristian Mihaescu</p><p>Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between these two capabilities. This work focuses on determining how much the LLM deviates from a buggy solution to fix the bug compared to a human-written patch, and if there is a bias towards generating entirely new solutions. We construct a dataset with all the submissions ($\sim$ 3000) from a couple of users from Codeforces, and we match each buggy submission with its corresponding human fix. By using the similarity between the buggy solution and the human fix as a baseline, we evaluate the quality of LLM-generated bug fixes on 3 OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1). We check if the generated solutions solve the problem by using the Codeforces-R1 dataset, an openly available dataset that has tests generated with the DeepSeek-R1 model. Our findings suggest that LLMs tend to modify more lines than necessary compared to human fixes and, in some cases, generate entirely new solutions. We also observe that LLMs solve more problems correctly when allowed to generate solutions from scratch rather than patch buggy submissions, even when those submissions are close to the human patch. This has important implications for the design of AI-assisted programming tools, particularly in supporting user debugging processes and promoting incremental problem-solving strategies rather than solution replacement.</p>]]></description>
  </item>
  <item>
    <title>Machine Unlearning for Gibbs Supervised Learning Algorithms</title>
    <link>http://arxiv.org/abs/2609.29409v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29409v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:33:22 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Yaiza Bermudez, Samir M. Perlaza, Iñaki Esnaola</p><p>In this paper, a method for achieving exact unlearning for Gibbs supervised learning algorithms is proposed using a variational formulation inspired by empirical risk minimization subject to relative entropy regularization (ERM-RER). Such a method consists of maximizing the expected empirical risk over the dataset to be unlearned subject to a regularization by relative entropy with respect to the original algorithm. The optimization variable is a probability measure on the models; and the solution is another Gibbs probability measure that represents a new Gibbs supervised learning algorithm. The method guarantees exact unlearning in the sense that the new Gibbs algorithm coincides in distribution with the algorithm that would have been obtained by retraining from scratch on the dataset to be retained. As a byproduct, a framework for reweighting data points in ERM-RER by strategically choosing both the reference measure and the regularization factor is obtained. In this framework, exact unlearning is the special case in which zero-weight is assigned to the contribution of the data points to be unlearned. More generally, depending on the choice of certain parameters, data points can be up-weighted or down-weighted in ERM-RER problems for particular purposes, e.g., controlling the generalization error of Gibbs algorithms. This paves the way for new constructive or adversarial views on classical reweighting data points in ERM-RER.</p><p><em>Comment: In Proc. of the IEEE International Symposium on Information Theory (ISIT), Guangzhou, China, Jun., 2026. 2026 Jack Keil Wolf ISIT Student Paper Award</em></p>]]></description>
  </item>
  <item>
    <title>WRAP: Fixtureless Wrench-aware Multi-Robot Assembly Planning</title>
    <link>http://arxiv.org/abs/2609.29407v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29407v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:31:07 +0000</pubDate>
    <category>Robotics</category>
    <description><![CDATA[<p><strong>Authors:</strong> Valentin N. Hartmann, Huang Su, Yijiang Huang, Stelian Coros</p><p>Assembly using robots often requires specially designed fixtures, or relies on top-down only assembly strategies. Using multiple robots, we can avoid using fixtures and make robotic assembly more flexible. Planning assembly sequences for multiple robots is challenging due to the high number of possible task assignments and orders. In addition, we need to reason over forces that occur during the assembly process, e.g., to decide if multiple robots are required for support, or if external support such as a table should be used.
  We present Wrap, a multi-robot assembly planner for multi-part assemblies, given the inter-part ordering-dependencies, the part meshes, and their initial state. We formulate a linear program to reason about valid grasps for supporting the forces that occur during assembly. The search leverages the assembly sequence, and greedily finds a feasible solution per assembly step by computing a heuristic via a cheap backwards search, and using the heuristic in the more expensive forward search.
  We then solve the multi-robot, multi-goal motion planning problem, and for execution, we split the plan into contact-rich assembly skills, and free space motion. We benchmark the planner on a variety of multi-part assemblies, and apply the planner to groups of robots differing in size and kinematics. We validate the work both in a physics simulation, and in real. Videos and code are available at https://www.vhartmann.com/wrap.</p><p><em>Comment: 8 pages, 9 figures, 5 tables</em></p>]]></description>
  </item>
  <item>
    <title>Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching</title>
    <link>http://arxiv.org/abs/2609.29405v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29405v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:30:22 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Julius Richter, Christoph Boeddeker, Yoshiki Masuyama, Kohei Saijo, Dominik Klement, Gordon Wichern, Jonathan Le Roux</p><p>We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). Starting from a pretrained SE model, RAM tilts the model&#x27;s conditional distribution toward outputs with higher reward. During training, the current model generates enhanced speech on-policy, evaluates each generated endpoint with a potentially non-differentiable reward, and analytically re-noises the endpoint to construct inputs for a reward-guided regression objective. This enables post-training directly on real recordings using weak supervision, such as text transcripts, without requiring paired clean speech targets or reward gradients. We investigate word error rate (WER)-based post-training and whether recognition performance can be improved without compromising perceptual speech quality. Experiments on real CHiME-4 recordings reduce WER by 5.08 percentage points relative to pretrained FlowSE without reducing any of the reported non-intrusive speech quality metrics. A subjective listening test at the default reward scale finds no statistically significant preference between the post-trained and pretrained models.</p><p><em>Comment: Submitted to ICASSP 2027</em></p>]]></description>
  </item>
  <item>
    <title>RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction--diffusion equations</title>
    <link>http://arxiv.org/abs/2609.29403v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29403v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:28:37 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Chenhao Si, Ming Yan</p><p>Learning surrogates for time-dependent partial differential equations often requires a new simulation corpus when the governing operator changes. We introduce RD-JEPA, a joint-embedding predictive architecture for self-supervised pretraining on reaction-diffusion trajectories. A single model is pretrained on five parameterized systems and then adapted to three held-out systems whose reaction operators and trajectories are excluded from pretraining. Using one, five, or ten complete trajectories from a held-out system, RD-JEPA achieves lower mean relative discrete $\ell^2$ field error and mean absolute spatial first-difference error than five supervised surrogate baselines, an independently trained control that removes the trajectory-dependent predictive latent pathway, and an architecture-matched model trained from scratch. Within the evaluated equations, output resolution, forecast horizons, and choices of adaptation trajectories, the results indicate that prediction of future-state representations can support data-efficient adaptation across related reaction-diffusion systems.</p>]]></description>
  </item>
  <item>
    <title>ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models</title>
    <link>http://arxiv.org/abs/2609.29398v1</link>
    <guid isPermaLink="true">http://arxiv.org/abs/2609.29398v1</guid>
    <pubDate>Thu, 24 Sep 2026 11:24:08 +0000</pubDate>
    <category>Machine Learning</category>
    <description><![CDATA[<p><strong>Authors:</strong> Xunkai Li, Xu Wang, Yinlin Zhu, Xiong Yongfu, Yi Liu, Rong-Hua Li, Guoren Wang</p><p>Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector. The representation must preserve entity semantics, construct interaction state from graph neighborhoods, and expose that state to prediction units with different geometry. Our empirical study shows why these requirements are inseparable. Higher-grade channels recover pair relations across the foundation graphs, specialized queries reveal information hidden by a generic readout, and rigid blade isolation removes cross-grade capacity. We therefore introduce ICE (Interaction-aware Clifford Encoder), a multimodal graph foundation model built on a node-indexed Clifford latent field. Topology, text, and images enter explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods. A protected Grade-1 route preserves entity semantics, while the full grade and depth bank remains available to fresh node and link heads. We establish exact cross-grade reachability, node-permutation equivariance, and a bound on the task residual around the semantic score. Experiments span one shared foundation over eleven graphs, six node-classification datasets, three link-prediction datasets, and matched few-shot tasks. ICE ranks first in all 30 reported supervised and few-shot comparisons. Core removals reduce every task summary, and mechanism controls connect the gains to higher-order transport, retained multidepth structure, semantic protection, and direct field access.</p>]]></description>
  </item>
  </channel>
</rss>