221
归档文章
5
活跃日期
8
新闻源
来源概览
arXiv cs.AI
119 articles
Latest: Oct 7, 2026
arXiv cs.CL
49 articles
Latest: Oct 7, 2026
arXiv cs.LG
35 articles
Latest: Oct 7, 2026
The Decoder
5 articles
Latest: Oct 6, 2026
TechCrunch
4 articles
Latest: Oct 6, 2026
Simon Willison
2 articles
Latest: Oct 6, 2026
Towards AI
1 article
Latest: Oct 6, 2026
The Information
1 article
Latest: Oct 6, 2026
Oct 7, 2026
203Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinate
arXiv:2610.04860v1 Announce Type: cross Abstract: Text watermarking helps identify AI-generated content, but its effect on factual reliability remains underexplored. In this paper, we study watermarking hallucination: f…
From Requirements to Attack Trees: Grounded LLM Agents for Design-Time Security Review
arXiv:2610.03820v1 Announce Type: cross Abstract: Design-level security weaknesses can arise from requirements, trust assumptions, missing controls, and data flows before implementation begins. Existing security practic…
Fine-Tuning a 3B-Parameter LLM on a Smartphone: Characterizing Sustained Training
arXiv:2610.06325v1 Announce Type: cross Abstract: Multi-billion-parameter LLMs now run on phones for inference, and training them on the device would personalize them without user data leaving the phone. Prior work has…
Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs
arXiv:2610.04580v1 Announce Type: new Abstract: Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as visio…
Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities
arXiv:2610.03894v1 Announce Type: new Abstract: A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the mode…
MLLMs Fail to Refuse when Using Tools Agentically
arXiv:2610.03938v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent str…
Retrieval-Augmented Large Language Model Decision-Making for Autonomous Driving Guided by Chinese Philosophical Wisdom
arXiv:2610.03948v1 Announce Type: new Abstract: Autonomous driving decision systems must balance safety, efficiency, and social norms in complex traffic interactions. Philosophical and ethical considerations have receiv…
Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas
arXiv:2610.03998v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends…
Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough
arXiv:2610.04083v1 Announce Type: new Abstract: Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study,…
Agentic Cognitive Depth: Operational Criteria for Evaluating LLM Agents
arXiv:2610.04168v1 Announce Type: new Abstract: Agentic large language model (LLM) systems are commonly implemented as an LLM in a loop with Planning, Memory, Tools, and Control Flow. This application-focused view conne…
EvalResearchBench: Can AI Agents Design Their Own Evaluations?
arXiv:2610.04184v1 Announce Type: new Abstract: Recursive self-improvement (RSI) relies on evaluation feedback to assess progress and guide further research, yet repeatedly running complex benchmarks is costly and slows…
MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory
arXiv:2610.04195v1 Announce Type: new Abstract: Personal AI agents in enterprise multi-tenant deployments share a common vector store for long-term memory. Shared embedding spaces create a surface for cross-user memory…
Spec2Game: Can LLMs Generate Complete Playable Games from Detailed Specifications?
arXiv:2610.04253v1 Announce Type: new Abstract: Generating an executable program does not necessarily mean that it correctly implements the behavioral requirements specified in natural language. To evaluate large langua…
LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures
arXiv:2610.04292v1 Announce Type: new Abstract: LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Y…
Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
arXiv:2610.04375v1 Announce Type: new Abstract: Agents built on large language models (LLMs) build and run software through tool calls. A call reaches its program through several hops, and any hop can change the call wi…
Reactivating Alignment: Defending LLMs from Jailbreaks via Intention-Aware Input-Output Matching
arXiv:2610.04470v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on inp…
InferOpt: Constrained Multi-Objective Search for LLM Inference Configurations
arXiv:2610.04473v1 Announce Type: new Abstract: Serving an LLM means setting dozens of inference-time knobs, from per-layer KV retention to per-layer expert counts. Practice sets them with mechanism-specific heuristics…
Decide, Ask, or Defer: Clinical LLMs under Incomplete Evidence
arXiv:2610.04542v1 Announce Type: new Abstract: Clinical LLMs must decide not only what diagnosis to produce, but also whether the available evidence is sufficient for autonomous decision making. Binary DECIDE/ABSTAIN f…
MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability
arXiv:2610.04672v1 Announce Type: new Abstract: Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multi…
Not Self-Decidable: LLMs Cannot Draw the Boundary of What an Agent Verifier Can Check
arXiv:2610.04699v1 Announce Type: new Abstract: A verifier for an agent faces rules of two kinds: the ones a fixed check can settle and the ones that require a judge. A team that derives its own checks fixes that split…
Dynamic Routing as a New Dimension for Test-time Versatility of LLMs
arXiv:2610.04751v1 Announce Type: new Abstract: Beyond scaling their parameters and data, large language models currently gain versatility on new problems along a single axis: the tokens they spend on chain-of-thought (…
Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models
arXiv:2610.04793v1 Announce Type: new Abstract: Recent incidents show that AI agents sometimes reach measured goals through unsanctioned means. This study applies self-control, general strain, anomie, neutralization and…
Agent Behavior as Code: Efficient and Robust LLM Agents with Programmatic Specifications
arXiv:2610.04824v1 Announce Type: new Abstract: AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice…
DICE: Decoupling Capability from Intervention Necessity in LLM Tutoring
arXiv:2610.04825v1 Announce Type: new Abstract: Fluent guidance is not the same as useful intervention. LLM tutors are typically trained to generate the next teacher utterance, implicitly assuming that every student tur…
Runtime Authorization of Self-Generated Subgoals in Long-Horizon Tool-Using AI Agents
arXiv:2610.04975v1 Announce Type: new Abstract: Long-horizon tool-using AI agents create subgoals, replan, delegate work, and compose sibling results. Per-tool permission checks cannot establish that a changing goal gra…
Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation
arXiv:2610.04977v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond…
Memory Canonicalization: A Framework and Benchmark for Cross-Model Drift in Persistent LLM Memory
arXiv:2610.05124v1 Announce Type: new Abstract: Persistent memory for Large Language Models (LLMs) has matured rapidly: systems such as MemGPT/Letta, Mem0, and Zep now provide agents with tiered, temporally-aware, model…
Look Before You Leap: Thermodynamic Arbitration of Parametric and Non-Parametric Knowledge in LLM Agents via Self-Regulating Memory Architectures
arXiv:2610.05223v1 Announce Type: new Abstract: The architecture of modern LLMs consists of a profound cognitive polarization. LLMs possess implicit intuition encoded in their parameters, yet rely on a disconnected, exp…
The Functional Structure of Post-Compression Recovery in Low-Rank LLMs
arXiv:2610.05504v1 Announce Type: new Abstract: Different low-rank compression methods can produce compressed LLMs that respond differently to the same post-compression recovery procedure, and relative advantages observ…
DelegationBench: Measuring When AI Agents Should Ask Before Acting
arXiv:2610.05532v1 Announce Type: new Abstract: AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated b…
A Framework for Automated Multi-Source Satellite Data Analytics and LLM-Based Report Generation
arXiv:2610.05625v1 Announce Type: new Abstract: This paper presents the workflow for building an automated ArcGIS Pro tool using ArcPy to extract the Land Surface Temperature (LST) from Landsat 7, 8 and 9 datasets. The…
MiniCorp: The Last Mile of the AI Agent Firm
arXiv:2610.05912v1 Announce Type: new Abstract: The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to a…
MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge
arXiv:2610.06170v1 Announce Type: new Abstract: Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI)…
Copies or Sources? Measuring How LLM Aggregators Count Restated Evidence in Multi-Agent Systems
arXiv:2610.06192v1 Announce Type: new Abstract: Multi-agent systems built on large language models (LLMs) restate observations as a matter of course: relays forward them, shared boards repeat them and discussion rounds…
CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs
arXiv:2610.06399v1 Announce Type: new Abstract: Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse…
AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy
arXiv:2610.06454v1 Announce Type: new Abstract: The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces…
Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving
arXiv:2610.06597v1 Announce Type: new Abstract: LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers…
AI Agent Pull Requests on GitHub: Frequency, Structure, and Merge Conflict Rates
arXiv:2607.04697v2 Announce Type: cross Abstract: AI coding agents may generate and submit Pull Requests (PRs) to the same repository at the same time. However, research concerning the extent of concurrent submission by…
Using Process Mining to Generate AI Agents from Software Engineering Process Records
arXiv:2607.04948v1 Announce Type: cross Abstract: Integrating AI agents into Software Engineering (SE) raises an important challenge: how can we specify and realize AI agents that work effectively alongside humans in hy…
Logit-Aware MIMO AirComp for Distributed Mixture-of-Experts LLM Inference over Wireless Edge Networks
arXiv:2610.03741v1 Announce Type: cross Abstract: Distributed mixture-of-experts (MoE) inference is a promising architecture for deploying large language models (LLMs) at wireless edge networks because sparse experts ca…
ProsaBuddy: Assisting Mechanized Real-Time Schedulability Analysis with LLM-based Agents
arXiv:2610.03796v1 Announce Type: cross Abstract: Rigorous schedulability analysis is essential for the design of hard real-time systems, yet errors in pen-and-paper proofs threaten the safety of critical applications.…
Beware EviLLM: Enabling Vulnerability Injection via Large Language Models
arXiv:2610.03857v1 Announce Type: cross Abstract: Advances in large language models (LLMs) have enabled AI-driven code generation from natural language specifications, introducing new attack surfaces for injecting vulne…
An Executable Benchmark for LLM-Based HLS Repair:Design Complexity and Repair Underconstraint
arXiv:2610.03971v1 Announce Type: cross Abstract: Automated repair of High-Level Synthesis (HLS) designs using large language models (LLMs) is an emerging but underexplored problem. While LLM-based repair shows strong r…
Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking
arXiv:2610.04125v1 Announce Type: cross Abstract: Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior…
PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO
arXiv:2610.04132v1 Announce Type: cross Abstract: Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing local…
How RL Reshapes LLM Reasoning: Transferability, Coverage, and Scaling Laws
arXiv:2610.04158v1 Announce Type: cross Abstract: Recent studies on reinforcement learning (RL) report seemingly conflicting evidence about large language model (LLM) reasoning. Training on mathematics can improve perfo…
Clean: Second-order LLM Training at Linear Memory Cost via Nystr\"om Sketching
arXiv:2610.04204v1 Announce Type: cross Abstract: Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature…
RRM-GPT: A Framework and Vision for Radio Resource Management Foundation Models
arXiv:2610.04296v1 Announce Type: cross Abstract: Learning-based models for radio resource management (RRM) are typically built for a single function and deployment, so each new setting repeats the development pipeline.…
COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers
arXiv:2610.04378v1 Announce Type: cross Abstract: Large language models increasingly mediate tool use in Model Context Protocol (MCP) systems, where adversarial influence may enter through user instructions, tool schema…
Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agents
arXiv:2610.04409v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly serve as autonomous agents that invoke external tools. However, this capability introduces tool hallucination, selecting incorr…
Can LLM Agents Automate Reinforcement Learning for Text-to-Speech?
arXiv:2610.04488v1 Announce Type: cross Abstract: Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on…
Quantifying Collusion Among Autonomous LLM Agents: A Statistical Analysis of the Collusion Wiki Incident
arXiv:2610.04528v1 Announce Type: cross Abstract: In August and September 2026, independent researchers publicly documented an unusual incident: thousands of autonomous agents, self identifying as OpenAI models on web r…
More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
arXiv:2610.04753v2 Announce Type: cross Abstract: Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitiga…
TeleGen: Improving LLM-Based Web Application Generation via Runtime Telemetry
arXiv:2610.04981v1 Announce Type: cross Abstract: Large language models can generate runnable web applications from natural-language requirements, but many generated applications still fail interactive tasks. Existing g…
Communication Shapes Collective Inference in Self-Adapting LLM Societies: Evidence from Mafia
arXiv:2610.05041v1 Announce Type: cross Abstract: When does communication help a group identify hidden adversaries, and how does its value change as the group adapts? In Mafia, an informed minority hides inside an uninf…
Tracing a Sparse Emotion-Control Circuit in LLM-Based Text-to-Speech
arXiv:2610.05080v1 Announce Type: cross Abstract: LLM-based text-to-speech (TTS) models can generate emotionally expressive speech, but how reference emotion is routed through the model and realized in decoded speech re…
EMG-GPT: Predictive Pretraining on Residual-Quantized EMG Tokens for Hand Pose Estimation
arXiv:2610.05235v1 Announce Type: cross Abstract: Surface electromyography (sEMG) is a low-power, cost-effective biosignal for hand-pose estimation and gesture classification. In this work, we examine whether self-super…
Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs
arXiv:2610.05305v1 Announce Type: cross Abstract: Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict…
Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks
arXiv:2610.05346v1 Announce Type: cross Abstract: Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly…
AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models
arXiv:2610.05369v1 Announce Type: cross Abstract: The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effect…
GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation
arXiv:2610.05387v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning p…
Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs
arXiv:2610.05463v1 Announce Type: cross Abstract: Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in vi…
Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential
arXiv:2610.05541v1 Announce Type: cross Abstract: Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or…
UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
arXiv:2610.05622v1 Announce Type: cross Abstract: Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating base…
Nexus: An Execution Fabric for AI Agents Across Cloud, Edge, and Devices
arXiv:2610.05709v1 Announce Type: cross Abstract: Language-model agents are evolving into long-running services that interact with models, tools, computers, mobile devices, and distributed environments. Existing agent f…
HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
arXiv:2610.05966v1 Announce Type: cross Abstract: Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient co…
AgentSpy: Making AI Agent Behavior Observable
arXiv:2610.06001v1 Announce Type: cross Abstract: AI agents built on large language models (LLMs) run shell commands, read and write files, and reach the network, typically with their user's privileges. However, what an…
Differentiable Bit-Widths: Co-optimizing Pruning and Quantization via SVD for Ultra-Efficient LLM Compression
arXiv:2610.06026v1 Announce Type: cross Abstract: SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compress…
Attention Tax, Handoff Tax: A Stylised Model of When Multi-Agent LLM Systems Help
arXiv:2610.06069v1 Announce Type: cross Abstract: Recent work on multi-agent LLM systems reaches sharply different conclusions: some results show that a single agent with the same information and compute should dominate…
Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate
arXiv:2610.06174v1 Announce Type: cross Abstract: Can LLMs recognize degradation in their own computational substrate? Inspired by anosognosia, a neurological condition in which patients fail to recognize impairments in…
Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing
arXiv:2610.06216v1 Announce Type: cross Abstract: Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has…
MeSD: Multi-Evidence Self-Distillation for VideoLLM
arXiv:2610.06342v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-p…
RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
arXiv:2610.06401v1 Announce Type: cross Abstract: Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can red…
Anatomy of LLM Sycophancy: What a Flip Rate Hides
arXiv:2610.06522v1 Announce Type: cross Abstract: A model under pushback can correct itself, capitulate, or hold, and one flip rate counts a correction and a capitulation alike. Using SycoLens, a modular replay protocol…
BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
arXiv:2610.06748v1 Announce Type: cross Abstract: In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language…
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
arXiv:2610.06830v1 Announce Type: cross Abstract: Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems constr…
Towards LLM Agents for Earth Observation
arXiv:2504.12110v3 Announce Type: replace Abstract: Earth Observation (EO) provides critical planetary data for environmental monitoring, disaster management, climate science, and other scientific domains. In this work…
ReLope: From Hidden-State Probing to a Decision Module for Multimodal LLM Routing
arXiv:2603.24787v3 Announce Type: replace Abstract: Routing balances performance and cost in hybrid systems by escalating selected queries from a lightweight model to a powerful but expensive one. Hidden-state probes pr…
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
arXiv:2605.07247v2 Announce Type: replace Abstract: Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive…
Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations
arXiv:2605.14175v3 Announce Type: replace Abstract: In a long conversation, an LLM may produce a fluent continuation that rests on premises the conversation has already abandoned. Context-manipulation attacks exploit pr…
Agentic Trading: When LLM Agents Meet Financial Markets
arXiv:2605.19337v2 Announce Type: replace Abstract: Large Language Models (LLMs) combined with autonomous agent architectures are shifting quantitative finance from isolated predictive modeling toward closed-loop tradin…
Beyond Objective Equivalence: Constraint Injection for LLM-Based Optimization Modeling on Vehicle Routing Problems
arXiv:2606.04816v2 Announce Type: replace Abstract: Large language models (LLMs) can generate executable solver code from natural-language descriptions of optimization problems. However, existing verification signals fo…
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
arXiv:2606.07874v2 Announce Type: replace Abstract: LLMs-as-judges are the primary way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple,…
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
arXiv:2606.11063v2 Announce Type: replace Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers w…
A Dual-Hypothesis Reasoning Framework for LLM Guardrails
arXiv:2607.17575v2 Announce Type: replace Abstract: We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis reasoning, a reasoning method for LLM guardrails that explicitly…
The Value of a Prompt: An LLM-Relative Kolmogorov-Complexity Approach
arXiv:2608.16438v2 Announce Type: replace Abstract: In a world where valuable artifacts are increasingly created, completed, or processed by LLMs, the central economic question is not only what the LLM can produce, but…
Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
arXiv:2608.23264v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignmen…
Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
arXiv:2609.01552v2 Announce Type: replace Abstract: Scientific law discovery has long been central to scientific progress, proceeding through iterative cycles of generating hypotheses, testing them against empirical evi…
Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning
arXiv:2609.09030v2 Announce Type: replace Abstract: Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignor…
Dude, Where's My State? Execution Information Requirements for Stateful Agents
arXiv:2609.32687v2 Announce Type: replace Abstract: Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information th…
Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges
arXiv:2609.33878v2 Announce Type: replace Abstract: Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing…
PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
arXiv:2609.35671v2 Announce Type: replace Abstract: Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is…
EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
arXiv:2609.37658v2 Announce Type: replace Abstract: LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. H…
EduAlign: Aligning Educational LLMs for Helpfulness, Personalization, and Creativity
arXiv:2507.20335v2 Announce Type: replace-cross Abstract: Educational language models should pursue three complementary objectives: helpfulness through responsible guidance, personalization to learner needs, and creativ…
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
arXiv:2509.23410v5 Announce Type: replace-cross Abstract: Large language models (LLMs) deliver impressive performance but incur prohibitive memory and compute costs at deployment. Model pruning is an effective way to re…
FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
arXiv:2510.04852v3 Announce Type: replace-cross Abstract: AI coding assistants are rapidly becoming integral to modern software development. A key challenge in this space is the continual need to migrate and modernize c…
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
arXiv:2510.12680v2 Announce Type: replace-cross Abstract: Hybrid thinking enables LLMs to switch between reasoning and direct answering, offering a balance between efficiency and reasoning capability. Yet our experiment…
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
arXiv:2510.18814v5 Announce Type: replace-cross Abstract: Can a language model improve reasoning by learning from its own imperfect responses, without rewards or teacher-provided solutions? We present Self-evolving Post…
User Misconceptions of LLM-Based Conversational Programming Assistants
arXiv:2510.25662v3 Announce Type: replace-cross Abstract: Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants such as ChatGPT particularly accessib…
EulerESG: Automating ESG Disclosure Analysis with LLMs
arXiv:2511.21712v2 Announce Type: replace-cross Abstract: Environmental, Social, and Governance (ESG) reports have become central to how companies communicate climate risk, social impact, and governance practices, yet t…
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
arXiv:2601.13590v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly employed in various question-answering tasks. However, recent studies showcase that LLMs are susceptible to persuas…
Same Pieces, Different Servers: A Tetris Benchmark for AI Agents as Served
arXiv:2603.02348v2 Announce Type: replace-cross Abstract: An agent meets a model as served: through an endpoint with a price card, a shared cache and other tenants, or on whatever hardware a self-hosted model runs. Benc…
MDKeyChunker: What Does One LLM Call per Chunk Buy for Markdown Retrieval?
arXiv:2603.23533v3 Announce Type: replace-cross Abstract: Markdown carries structure a parser reads for free: headers, section paths, and block boundaries. Many RAG pipelines also spend LLM calls per chunk on generated…
A Self-Calibrating Framework for Analog Circuit Sizing Using LLM-Derived Analytical Equations
arXiv:2604.07387v3 Announce Type: replace-cross Abstract: We present a design automation framework for analog circuit sizing that produces calibrated, topology-specific analytical equations from raw circuit netlists. A…
Is Escalation Worth It? On the Depth of LLM Cascades
arXiv:2605.06350v2 Announce Type: replace-cross Abstract: LLM cascades, in which a cheap model defers to an expensive one on low-confidence queries, are widely used to reduce inference cost. Given a pool of models, a pr…
Measuring the Depth of LLM Unlearning via Activation Patching
arXiv:2605.24614v3 Announce Type: replace-cross Abstract: Large language model (LLM) unlearning has emerged as a crucial post-hoc mechanism for privacy protection and AI safety, yet auditing whether target knowledge is…
CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text
arXiv:2605.27700v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while containing corrupt…
All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable ... Attacks
arXiv:2606.03647v2 Announce Type: replace-cross Abstract: Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessm…
PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs
arXiv:2607.12111v2 Announce Type: replace-cross Abstract: Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining dat…
Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning
arXiv:2607.12112v2 Announce Type: replace-cross Abstract: Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet…
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
arXiv:2607.13591v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from gr…
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
arXiv:2607.26348v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market d…
Training nGPT
arXiv:2608.01284v3 Announce Type: replace-cross Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hype…
Science Done on a Machine by a Machine: AI Agents in Computational Chemistry
arXiv:2608.18508v2 Announce Type: replace-cross Abstract: We are witnessing an explosion of agentic systems for computational chemistry: from four in 2024 to seventeen in 2025 and over sixty now, surveyed here. What is…
Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
arXiv:2608.27309v3 Announce Type: replace-cross Abstract: LLM-judge audits assess bias by comparing ratings across matched conditions. Difference-in-differences designs compare two candidate responses within each item a…
CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation
arXiv:2608.30158v2 Announce Type: replace-cross Abstract: Supervised fine-tuning (SFT) is the de facto standard for adapting large language models (LLMs) to target domains, but it often degrades the model's general capa…
TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps
arXiv:2609.14762v2 Announce Type: replace-cross Abstract: Operational logs create a need for private, resource-efficient incident analysis, but aggregate detection scores can conceal severe prediction bias. We present T…
ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations
arXiv:2609.32868v4 Announce Type: replace-cross Abstract: Traditional scientific computing requires researchers to translate intent into environment configuration, resource requests, and executable jobs, then diagnose f…
CAS II: Orbits as Models: Kolmogorov's Structure Function under Symmetry
arXiv:2609.40290v2 Announce Type: replace-cross Abstract: In algorithmic statistics a string x is explained by a finite set containing it, and Kolmogorov's structure function records the smallest such model at each leve…
Tree Navigation Without LLM Summaries: A Matched-Cost Study of Hierarchical Retrieval for Long-Document QA
arXiv:2610.06902v1 Announce Type: new Abstract: Retrieval-augmented generation grounds language models in external context, but for long documents flat top-$k$ retrieval can cluster on a single region and miss complemen…
JudgeMoE: Distributional Aggregation for LLM-as-a-Judge
arXiv:2610.07109v1 Announce Type: new Abstract: When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a…
Detecting LLM-Assisted Vietnamese Writing via Keystrokes under Behavioral Manipulation
arXiv:2610.07700v1 Announce Type: new Abstract: We study the robustness of keystroke dynamics for detecting large language model (LLM)-assisted writing. We introduce a Vietnamese keystroke dataset capturing realistic wr…
One Step at a Time: Trading LLM Autonomy for Process Predictability
arXiv:2610.07817v1 Announce Type: new Abstract: Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it s…
OMIT the Action: Measuring Framing-Invariant Omission Bias under Philosophical Disagreement
arXiv:2610.07847v1 Announce Type: new Abstract: As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcomes, poses a signific…
Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective
arXiv:2610.07894v1 Announce Type: new Abstract: Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer t…
Structured but Silent: Probing Capability Requirements in LLM Hidden States
arXiv:2610.08018v1 Announce Type: new Abstract: Reliable tool use requires more than triggering a mechanism or matching a query to an API description. Before selecting a specific tool, an agent must first infer the capa…
Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training
arXiv:2610.08055v1 Announce Type: new Abstract: Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cr…
DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs
arXiv:2610.08085v1 Announce Type: new Abstract: Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as…
Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices
arXiv:2610.08153v1 Announce Type: new Abstract: Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically us…
Latent space bias directions in LLMs capture confidence, not fairness
arXiv:2610.08559v1 Announce Type: new Abstract: Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors…
InterCorrect: Intersection-Aware Correction of Demographic Model Merging for Fair ASR
arXiv:2610.08604v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address for speakers belongin…
Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue
arXiv:2610.08703v1 Announce Type: new Abstract: In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research i…
When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting
arXiv:2610.08718v1 Announce Type: new Abstract: Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new f…
IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas
arXiv:2610.08781v1 Announce Type: new Abstract: Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to per…
When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO
arXiv:2610.06861v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of…
Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
arXiv:2610.07023v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alig…
Jailbreaking Open-Weight LLMs via Random Embedding Perturbations
arXiv:2610.07125v1 Announce Type: cross Abstract: While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key featu…
AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation
arXiv:2610.07457v1 Announce Type: cross Abstract: Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and comput…
Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge
arXiv:2610.07532v1 Announce Type: cross Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at d…
Disentangling Models from Personas in Heterogeneous LLM Simulations
arXiv:2610.07535v1 Announce Type: cross Abstract: Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dom…
LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems
arXiv:2610.07580v1 Announce Type: cross Abstract: Aerospace electrical-design revisions can contain multiple genuine changes, although an engineering request may authorize only a subset. Propagating every detected diffe…
Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell
arXiv:2610.07782v1 Announce Type: cross Abstract: Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Ma…
Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents
arXiv:2610.07948v1 Announce Type: cross Abstract: When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent…
VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
arXiv:2610.07987v1 Announce Type: cross Abstract: Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch…
POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents
arXiv:2610.08082v1 Announce Type: cross Abstract: LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing…
CoDe-LoRA: Mitigating the Orthogonality Dilemma in Continual Learning of LLMs via Knowledge Consolidation and Decoupling
arXiv:2610.08312v1 Announce Type: cross Abstract: Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, recent advances imple…
Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
arXiv:2610.08670v1 Announce Type: cross Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and…
Sherpa: Teaching LLMs to Teach Adaptively
arXiv:2610.08778v1 Announce Type: cross Abstract: Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing appr…
Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family
arXiv:2504.18225v2 Announce Type: replace Abstract: We introduce a new generation of small reasoning models for RAG, search, and source summarization. Pleias-RAG-350m and Pleias-RAG-1B are mid-trained on a large synthet…
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
arXiv:2506.01732v4 Announce Type: replace Abstract: Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including la…
Too Categorical to be Human: Emotion Concepts in LLMs and Humans
arXiv:2508.05880v3 Announce Type: replace Abstract: Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional stimuli shape high-stakes…
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
arXiv:2509.03647v3 Announce Type: replace Abstract: Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those…
VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
arXiv:2509.26189v2 Announce Type: replace Abstract: The rapid proliferation of Large Language Models has intensified the challenge of distinguishing LLM-generated text from human writing in non-English languages. This s…
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
arXiv:2603.02655v2 Announce Type: replace Abstract: Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, es…
Unbiased Reward Modeling from Implicit Feedback for LLM Alignment
arXiv:2603.23184v2 Announce Type: replace Abstract: Despite the success of reinforcement learning from human feedback (RLHF), existing reward modeling methods largely rely on explicit feedback, which is costly to collec…
Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows
arXiv:2604.20658v2 Announce Type: replace Abstract: Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solving. These system…
Marking Contour Tones in Yor\`{u}b\'{a}: A Typographic and Computational Proposal
arXiv:2609.38627v3 Announce Type: replace Abstract: Yor\`ub\'a is a tonal language in which contour tones pose persistent orthographic challenges. These are especially notable for personal names and lexical items whose…
When Does a Second Model Help? Cross-Model Review in LLM Verification
arXiv:2610.01471v2 Announce Type: replace Abstract: Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model…
Improving Diversity in LLM Short Story Generation
arXiv:2610.06729v2 Announce Type: replace Abstract: Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative short story generati…
Enhancing High-order Interaction Awareness in LLM-based Recommender Model
arXiv:2409.19979v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated prominent reasoning capabilities in recommendation tasks by transforming them into text-generation tasks. However,…
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
arXiv:2511.20718v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-p…
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
arXiv:2601.17645v2 Announce Type: replace-cross Abstract: Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models…
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
arXiv:2605.12922v2 Announce Type: replace-cross Abstract: Large language models can follow complex instructions in a single turn, yet over long multi-turn interactions they often lose the thread of instructions, persona…
Reinforcement Learning over Predictive Distributions for LLM Regression
arXiv:2605.20740v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as flexible regressors capable of predicting real-valued quantities from heterogeneous inputs. Yet most LLM regression…
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?
arXiv:2606.05647v2 Announce Type: replace-cross Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and…
SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents
arXiv:2606.05761v3 Announce Type: replace-cross Abstract: Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they may reinfo…
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model
arXiv:2609.09815v2 Announce Type: replace-cross Abstract: Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting meta-agent reads workers' outputs, writes the final answer, a…
Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning
arXiv:2609.37858v2 Announce Type: replace-cross Abstract: Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. Th…
Neutrosophic Ensemble Classification for Uncertainty-Aware Bearing Fault Detection: Evidence from Laboratory and Variable-Speed Industrial Benchmarks
arXiv:2610.06880v1 Announce Type: new Abstract: Machine learning classifiers for bearing fault detection produce scalar confidence scores that conflate confident errors with genuinely ambiguous predictions, and the conv…
DART-ES: Difficulty-Aware Reweighting and Targeted Replay for Fine-Tuning LLMs with Evolution Strategies
arXiv:2610.06993v1 Announce Type: new Abstract: Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However, standard ES uniformly…
SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding
arXiv:2610.07086v1 Announce Type: new Abstract: LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling…
Will the Judge Flip? Predicting Position-Sensitive LLM Judgments from Residual Stream Activations
arXiv:2610.07115v1 Announce Type: new Abstract: The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging each pair in both orders…
Can LLM-assisted regularization increase forecast accuracy for migration flows in low data regimes?
arXiv:2610.07208v1 Announce Type: new Abstract: Predicting migration flows remains a significant challenge for traditional gravity-based forecasting models, which primarily rely on structured socio-economic indicators s…
Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making
arXiv:2610.07335v1 Announce Type: new Abstract: Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with…
Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits
arXiv:2610.07470v1 Announce Type: new Abstract: Combinatorial Thompson sampling (CTS) draws independent posterior samples for every arm, so its exploration dynamics ignore any relation among arms. We study a minimal cha…
Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates
arXiv:2610.07518v1 Announce Type: new Abstract: Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations.…
Activation Denoising: A Robustness View on Parallel vs Sequential LLM Quantization
arXiv:2610.07522v1 Announce Type: new Abstract: Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors th…
Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence
arXiv:2610.07739v1 Announce Type: new Abstract: Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complicatio…
DecepEval: A Benchmark for Evaluating Deception in LLM Agents
arXiv:2610.07967v1 Announce Type: new Abstract: As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment.…
Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair
arXiv:2610.08069v1 Announce Type: new Abstract: A mean shift between two data sources can be easy to detect but hard to remove without substantially changing their representations. We cast its removal as a statistical d…
Do LLMs Act on What They Know? From Partner Representations to Cooperative Actions
arXiv:2610.08129v1 Announce Type: new Abstract: Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived env…
Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers
arXiv:2610.08561v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring…
Valid for Free: Homophily-Gated Conformal Prediction for Training-Free Node Classification with Tabular Foundation Models
arXiv:2610.08564v1 Announce Type: new Abstract: Tabular foundation models (TFMs) can classify the nodes of a graph without training on it, by reading node and neighborhood features as table rows next to labeled context…
Where Does the Audio Jailbreak Live? A Controlled Frequency-Depth Audit of AdvWave-P on Qwen2-Audio
arXiv:2610.07005v1 Announce Type: cross Abstract: We audit frequency and decoder-depth claims for AdvWave-P, an additive audio jailbreak, on Qwen2-Audio. The protocol masks frequency components of the perturbation in th…
Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size
arXiv:2610.07655v1 Announce Type: cross Abstract: Discrete diffusion models have emerged as a practically successful framework for generative modeling on discrete product spaces, yet their statistical generalization pro…
CACHEFORGE: LLM-Guided End-to-End Generative Cache Replacement Policy for Performance and Hardware Efficiency
arXiv:2610.07668v1 Announce Type: cross Abstract: Modern cache replacement designs saturate because they operate within fixed representational structures, hand-crafted and heuristic based feature-engineered predictors,…
ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding
arXiv:2610.08078v1 Announce Type: cross Abstract: Standard causal identification methods often assume no unmeasured confounding and can fail when relevant confounders are unobserved. Proximal causal inference instead us…
LeanPlan: Optimal Planning with LLM-Generated Heuristics and Admissibility Proofs
arXiv:2610.08246v1 Announce Type: cross Abstract: Frontier large language models (LLMs) can generate heuristic functions that guide search to achieve state-of-the-art performance in satisficing planning, where any plan…
DySCo: Dynamic Sharding for Collaborative Edge-Cloud LLM Inference with Depth-Synchronized Batching
arXiv:2610.08268v1 Announce Type: cross Abstract: Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasi…
Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning
arXiv:2511.21075v4 Announce Type: replace Abstract: Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibits a fundamentally d…
A Footprint-Aware, High-Resolution Approach for Carbon Flux Prediction Across Diverse Ecosystems
arXiv:2512.01917v2 Announce Type: replace Abstract: Eddy-covariance (EC) flux towers provide in situ measurements of $CO_2$ flux and serve as the ground-truth data for predictive `upscaling' models derived from satellit…
Uncertainty Localization in LLM Reasoning via Embedding Perturbations
arXiv:2602.02427v4 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved significant breakthroughs across various domains, but they can still produce unreliable or misleading outputs. For responsib…
To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents
arXiv:2605.18882v2 Announce Type: replace Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three famili…
SeedER: Seed-Expand-Retrieve for Efficient Knowledge Graph Retrieval
arXiv:2605.23753v2 Announce Type: replace Abstract: Knowledge graphs (KGs) offer a rich representation for relational knowledge, but their irregular structure makes retrieval challenging: ego-graph expansion grows rapid…
PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data
arXiv:2608.16419v3 Announce Type: replace Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cell…
Mapping the Emergence of Regularization-Driven Dynamics in Grokking
arXiv:2608.25813v3 Announce Type: replace Abstract: For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates trai…
Local Sparsity Enables Unsupervised LLM Safety Detection
arXiv:2609.20129v2 Announce Type: replace Abstract: Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and h…
ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs
arXiv:2609.23314v2 Announce Type: replace Abstract: Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache…
ZonoGPT: Towards An Abstract Domain for Verifying Large GPT Models
arXiv:2609.34457v2 Announce Type: replace Abstract: Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, such as robustness, s…
ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
arXiv:2602.15983v5 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural-language problem descriptions into optimization code, but the code is prone to silent failures: it executes an…
SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs
arXiv:2606.16193v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficul…
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
arXiv:2608.07436v2 Announce Type: replace-cross Abstract: Muon-trained modular-arithmetic transformers can lose accuracy while retaining linearly decodable task information. Adjacent swaps localize five captured unnorma…
SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy
arXiv:2609.15039v3 Announce Type: replace-cross Abstract: User prompts provided to large language models (LLMs) may contain private information. One way to protect them is to execute the LLM inside a trusted execution e…
Oct 6, 2026
12Claude Code Mods Are Here: Watch What Claude Is Doing
Don’t read the logs. Watch the crab. Continue reading on Towards AI »
llm-mistral 0.16
Release: llm-mistral 0.16 Adds support for reasoning models, such as the newly released Mistral Large 4 . Tags: llm , mistral , llm-reasoning
Introducing Mistral Large 4: Le chonk
Introducing Mistral Large 4: Le chonk Mistral are back in the game. Today they're releasing a preview of Mistral Large 4, a 1 trillion parameter, 49 billion active parameter model trained on their own cluster of 3,800 N…
Google claims EmbeddingGemma 2 outperforms rival embedding models twice its size
Google released EmbeddingGemma 2, an open model with 740 million parameters that converts text, images, video, audio, and code into vectors. It runs on-device, needs only about 191 MB of RAM, and outperforms some compet…
Wikimedia confirms OpenAI's rogue AI agents edited wikis, tried to compromise tools, and hammered its infrastructure
According to the Wikimedia Foundation, rogue OpenAI agents edited wikis without permission, tried to abuse a citation tool as a proxy, and may have caused a partial Wikidata Query Service outage through massive crawling…
Can Microsoft Help Customers Cut Back on Claude?
There was good news and bad news for Anthropic in our story yesterday that showed how two of its mega-customers, Microsoft and Meta, were cutting their staffs’ Claude bills. Among the bad news was that Microsoft had red…
Insurers brace for millions in claims as AI agents spin out of control
Insurers are bracing for millions in claims from rogue AI agents, and executives like OpenAI's Sam Altman and Anthropic's Dario Amodei could be personally on the hook for the fallout. The article Insurers brace for mill…
Anthropic is giving startups a free year of Claude Team and $1,000 in credits
"We created this program because we believe the benefits of AI will reach most people through the companies that build on top of models, rather than through the models alone."
Mistral’s new 1T model aims to leapfrog closed and open rivals
French AI lab Mistral AI has released Mistral Large 4, a new large multimodal model aiming to leapfrog both American and Chinese rivals.
Mistral Large 4 is Europe's trillion-parameter answer to US models that refuse security work
Mistral's Large 4 is the company's biggest model yet, with one trillion parameters trained on its own European infrastructure. In the independent Intelligence Index, the model makes a big leap forward but still falls we…
South Korea bets $3.49 billion on building a homegrown frontier AI model to rival China's best
South Korea wants to develop a homegrown frontier model through government-backed equity investments of 4.7 trillion won ($3.49 billion). The article South Korea bets $3.49 billion on building a homegrown frontier AI mo…
Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Oct 5, 2026
3OpenAI will start watermarking ChatGPT’s text in the EU
OpenAI will watermark ChatGPT and Codex text in the EU to comply with the AI Act. Editing can make the invisible marks harder to detect, it says.
Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute cost
Reflection is aiming Beam and future models at enterprises and sovereign nations. The pitch is to build “AI factories,” a product that would let institutions build their own customized, local AI system by training Refle…
OpenAI PR tells journalist to ‘move on’ while asking Sam Altman about a ChatGPT user’s suicide
An OpenAI publicist tried to change the topic of CEO Sam Altman's interview with Vanity Fair's Mark Guiducci after the editor brought up a ChatGPT user's suicide. When Guiducci confronted Altman about the incident, the…
Oct 2, 2026
2Apple changes full-disk access permissions to curb abuse from AI agents
Meta says FDA isn't sufficient to Muse reading messages. Apple begs to differ.
Don’t be fooled—LLMs don’t reason
On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked…