Skip to content

Archive / 2026-09-09

September 9, 2026

  1. AirPods 5

    apple.com8 days ago504 ptsView detailsJoin discussion

  2. GPT-6 Astra, looped transformers, and hidden reasoning

    magazine.sebastianraschka.com8 days ago514 ptsView detailsJoin discussion

  3. No Man's Sky Cosmos

    nomanssky.com8 days ago457 ptsView detailsJoin discussion

  4. GNU Radio in the browser

    gnuradioworld.com8 days ago224 ptsView detailsJoin discussion

  5. Training a 3.8B LLM to 0.384 CORE for $998

    hugovergnes.github.io8 days ago120 ptsView detailsJoin discussion

  6. Testing Race Conditions

    projectzero.google8 days ago45 ptsView detailsJoin discussion

  7. Better AI code comment detector

    entropicthoughts.com8 days ago42 ptsView detailsJoin discussion

  8. My Mental Model of AI Broke on September 8

    rough-ideas.bearblog.dev8 days ago32 ptsView detailsJoin discussion

  9. Primary source

    Introducing the Agents API

    Build and launch cloud agents with the Agents API, a managed service powered by the Codex harness for orchestration, long-running sessions, and tool use.

    openai.com8 days ago13 ptsView details

  10. War Is a Racket

    en.wikipedia.org8 days ago16 ptsView detailsJoin discussion

  11. Fashware

    git.sr.ht8 days ago16 ptsView detailsJoin discussion

  12. Teaching an AI My Taste

    lonriesberg.com8 days ago15 ptsView detailsJoin discussion

  13. Primary source

    Build more natural voice experiences with GPT‑Live‑1 in the API

    GPT‑Live‑1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.

    openai.com8 days agoView details

  14. Primary source

    Paul Christiano joins OpenAI Foundation Board

    Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.

    openai.com8 days agoView details

  15. Primary source

    The AI policy window is open. We need to act.

    Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.

    openai.com8 days agoView details

  16. Primary source

    GPT-6 Astra: The next generation in intelligence for work

    Meet GPT-6 Astra, OpenAI’s most capable model for business, with advanced reasoning, computer use, and stronger writing and design judgment.

    openai.com8 days agoView details

  17. Primary source

    Rebuilding AUTOMATIC1111 with Gradio Workflow

    huggingface.co8 days agoView details

  18. LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

    LandingAI has shipped Agentic Document Extraction Gen2, a rebuild of its document stack on the DPT-3 model family. Chunks are retired in favor of a document, page and block tree. DPT-3 Pro grounds to the line, DPT-3 Verity grounds to the word with a confidence score, and Parse billing now counts output characters inst…

    marktechpost.com8 days agoView details

  19. Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

    Google has open-sourced Mantis, a stack-agnostic toolkit of security review skills for AI coding agents. It runs the full vulnerability lifecycle: sweep the code, filter false positives, reproduce the bug in a sandbox, patch it, re-attack the patch, then score the risk. Apache 2.0, and documented as demonstration-only…

    marktechpost.com8 days agoView details

  20. Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

    Voice agent teams keep hitting the same wall. The catalog holds 400 voices and the brief asks for the one that is not in it: a Quebecoise receptionist for a Montreal dealership, a narrator in his sixties with lecture hall authority. Briefs outnumber any catalog, and cloning closes the gap one speaker at a time, […] Th…

    marktechpost.com9 days agoView details

  21. The AI Researcher Who Just Quit Anthropic Says It’s ‘Crunch Time for Humanity’

    Jacob Coxon talks to WIRED about the “mini Manhattan project” inside Anthropic, the problem with alignment, and why AI labs have just a few years left to make their systems safe.

    wired.com8 days agoView details

  22. Suno releases its first AI music model made with record industry help

    Suno's new v6 AI music model is its first made with support from the record industry. Suno's Jack Brody told The Verge that v6 was "trained from the ground up, with a new set of data that does not include the same data that our previous models were trained on." The data includes content licensed from partners Warner M…

    theverge.com8 days agoView details

  23. OpenAI’s sly mathematical breakthrough sends a chill through academia

    Open AI CEO Sam Altman speaks during the G20 Innovation Ministerial. | (Photo by Matt RAMEY / AFP via Getty Images) OpenAI's announcement Tuesday that it has solved one of mathematics' legendary Millennium Prize problems should have been a moment of triumph. The result is both an undeniable achievement and a striking…

    theverge.com8 days agoView details

  24. San Francisco Orders Meta to Stop ‘Allowing’ AI Child Abuse Ads

    The City Attorney’s Office has asked Meta to explain how the harmful ads repeatedly ran on Facebook and Instagram. The company claims the ads are not under the city’s jurisdiction.

    wired.com8 days agoView details

  25. Read the Apple document explaining how new listening features still protect your privacy

    At Wednesday's iPhone Duo launch event, Apple announced a handful of new Siri AI Audio Intelligence features, including Siri Recap, Live Rewind, Sound Recognition, and Music Recognition. Alongside its announcement, Apple released a document laying out how it plans to balance AI "ambient listening" and users' privacy.…

    theverge.com8 days agoView details

  26. Apple’s new iPhone camera mode promises to prove your photo isn’t AI

    Apple is launching a new way to prove that the picture you took isn't manipulated by AI. A new feature, called "Reference Image," will arrive with the iPhone 18 Pro lineup later this month and is supposed to use the device's new camera sensor to "sign every pixel it sees." The iPhone 18 Pro and Pro Max will only authe…

    theverge.com8 days agoView details

  27. I Let an AI Agent Hack All My Gadgets—and I’d Do It Again

    After I removed the safety guardrails from a powerful open-source model, it found vulnerabilities in my household devices and hacked into a PC. But it also told me how to make everything a lot more secure.

    wired.com8 days agoView details

  28. Microsoft has new AI privacy rules for schools

    Microsoft agreed to a set of safety and privacy principles for AI in schools a week after two major school systems announced a ban on student-facing AI. In a new agreement with the American Federation of Teachers (AFT), the second-largest teachers union in the US, and its New York City affiliate the United Federation…

    theverge.com8 days agoView details

  29. A Stealth Startup Thinks It Just Hacked the Memory Shortage

    Kepler Computing claims a new approach to chip design—and a proprietary material—can help end the supply bottlenecks that have sent memory prices surging.

    wired.com8 days agoView details

  30. Amazon Prime Video’s new AI tech matches lips to dubbed audio

    Maxton Hall. | Image: Prime Video Amazon's Prime Video is launching a new AI-powered feature that lines up an actor's mouth with "human-dubbed" audio. The feature is only available with the English dub of the German series Maxton Hall for now, but Prime Video plans to expand it to "additional titles" in the future. In…

    theverge.com8 days agoView details

  31. Students who use AI generally score worse at school

    Students who use AI to help them study tend to perform worse at school than those who don't, according to data from a global OECD educational report. The situation is more complex than it sounds though, with certain types of AI use giving learners a slight boost, especially among students taught to critically assess h…

    theverge.com8 days agoView details

  32. Worried Anthropic researchers warn that AI ‘could kill all humans’

    A senior Anthropic safety researcher has said there is more than a 10 percent chance artificial intelligence "could kill all humans" by the end of the decade, just hours after a colleague resigned over fears the AI lab and its rivals are carelessly racing to build "superhuman systems" they cannot control. In a post on…

    theverge.com8 days agoView details

  33. UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

    arXiv:2609.09815v1 Announce Type: new Abstract: Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting meta-agent reads workers' outputs, writes the final answer, allocates later calls, and decides when to stop. It is expressive, but it also concentrates three control decis…

    arxiv.org8 days agoView details

  34. Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields

    arXiv:2609.09864v1 Announce Type: new Abstract: Affective computing has largely followed an individual-state paradigm, extracting discrete emotion labels or arousal/valence from isolated speakers. We argue this framing is incomplete for interaction. Drawing on affective resonance and vitality-contour accounts, we prop…

    arxiv.org8 days agoView details

  35. Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format

    arXiv:2609.09882v1 Announce Type: new Abstract: Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these readouts are often treated as interchangeable. Holding model checkpoint and prompt content fixed, we compare probabilities obtained by scoring answer tokens with pre…

    arxiv.org8 days agoView details

  36. RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases

    arXiv:2609.10092v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a ro…

    arxiv.org8 days agoView details

  37. Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications

    arXiv:2609.09885v1 Announce Type: new Abstract: This paper studies unmanned aerial vehicle (UAV)-mouted reconfigurable intelligent surface (RIS)-assisted device-to-device (D2D) communication with stochastic link activation. It models UAV motion and attitude, time-varying Rician angles, and angle-dependent RIS reflecti…

    arxiv.org8 days agoView details

  38. Structural Process Supervision for Latent Chain-of-Thought Reasoning

    arXiv:2609.09928v1 Announce Type: new Abstract: Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit chain-of-thought (CoT) tokens with compact continuous-space embeddings. However, existing methods lack direct process supervision over these latent embeddings, which…

    arxiv.org8 days agoView details

  39. Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

    arXiv:2609.10036v1 Announce Type: new Abstract: Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments. A single informative observation c…

    arxiv.org8 days agoView details

  40. AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

    arXiv:2609.09212v1 Announce Type: cross Abstract: This paper presents an end-to-end evaluation framework for image-triggered command injection against computer-use agents (CUAs). The goal is to test whether a local visual patch can induce verifiable environmental consequences along the full chain of screenshot input,…

    arxiv.org8 days agoView details

  41. Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

    arXiv:2609.10413v1 Announce Type: new Abstract: Current LLM memory systems treat all personal facts identically, so stores grow without bound while retrieval precision degrades. The core challenge is lifecycle management: which memories should persist, which should be replaced, and at what rate, conditioned on the beh…

    arxiv.org8 days agoView details

  42. OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization

    arXiv:2609.10055v1 Announce Type: new Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle distinctions among hierarchically related concepts c…

    arxiv.org8 days agoView details

  43. Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning

    arXiv:2609.10177v1 Announce Type: new Abstract: In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface level imitation of in-context demonstrations, maki…

    arxiv.org8 days agoView details

  44. Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System

    arXiv:2609.10350v1 Announce Type: new Abstract: The banking system now depends on a small set of shared artificial intelligence vendors for fraud screening, credit decisioning, anti-money-laundering triage, customer analytics, and internal decision support. This paper studies how a compromise inside one of those vendo…

    arxiv.org8 days agoView details

  45. Quantifying Logical Consistency in Transformers via Query-Key Alignment

    arXiv:2502.17017v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated impressive performance in various natural language processing tasks, yet their ability to perform multi-step logical reasoning remains an open challenge. Although Chain-of-Thought prompting has improved logical reasoning b…

    arxiv.org8 days agoView details

  46. Multi-Agent Agentic Graph Learning via Structural Signatures

    arXiv:2609.09565v1 Announce Type: new Abstract: Agentic graph learning (AGL) has recently achieved promising results on graph reasoning tasks, where an agent powered by a large language model (LLM) sequentially samples the graph as evidence to support its final prediction. Existing methods either employ a single agent…

    arxiv.org8 days agoView details

  47. Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model

    arXiv:2609.09213v1 Announce Type: cross Abstract: We study how physical-state inputs affect a 0.8B hybrid language model adapted for manipulation with 6.2M trainable parameters. Six conditions are trained on three LIBERO-Spatial tasks and evaluated over three seeds and 540 held-out rollouts. Conditioning recurrent dec…

    arxiv.org8 days agoView details

  48. Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

    arXiv:2609.09219v1 Announce Type: cross Abstract: AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement…

    arxiv.org8 days agoView details

  49. What Should an Agent Forget? Separating What Is Stored from What Is Used

    arXiv:2609.10263v1 Announce Type: new Abstract: Persistent language agents need stored experience to remain available across time, while each answer requires evidence suited to a particular question. A superseded fact can mislead a current-state answer and still be essential for a historical query. We present RD-Forge…

    arxiv.org8 days agoView details

  50. The Mutations of Machine Speech

    arXiv:2609.09496v1 Announce Type: new Abstract: Algorithmic outputs now populate the digital environments through which contemporary life is organized. The role of law in facilitating and constituting (rather than merely responding to) these processes is gaining increasing traction across scholarly accounts. This inqu…

    arxiv.org8 days agoView details

  51. VLX-VR: An Agentic-Aware Video Reasoning Model

    arXiv:2609.09985v1 Announce Type: new Abstract: Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition when observations are incomplete,…

    arxiv.org8 days agoView details

  52. On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data

    arXiv:2609.10321v1 Announce Type: new Abstract: Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from the teacher prediction and applied uniformly to all…

    arxiv.org8 days agoView details

  53. Towards Stress-Aware Sentence-Level Filipino G2P With Weakly-Supervised ByT5 Fine-Tuning

    arXiv:2609.09974v1 Announce Type: new Abstract: Grapheme-to-phoneme conversion (G2P) refers to the task of converting a sequence of graphemes to a corresponding sequence of phonemes. While Filipino G2P is fairly straightforward due to its shallow orthography, the inclusion of prosodic features such as stress adds a la…

    arxiv.org8 days agoView details

  54. The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

    arXiv:2609.09395v1 Announce Type: new Abstract: Language models act through tools, yet practical agents face libraries containing thousands of interfaces. We introduce the tool menu as the short, ordered subset of available tools shown to an agent before execution. The agent can call only tools in this menu. Multi-ste…

    arxiv.org8 days agoView details

  55. A Function-Space Approach to the Statistical Mechanics of Learning Dynamics

    arXiv:2609.09589v1 Announce Type: new Abstract: Deep neural networks exhibit regular macroscopic behavior despite highly nonlinear dynamics in vast parameter spaces. We develop a statistical-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and…

    arxiv.org8 days agoView details

  56. AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

    arXiv:2609.09875v1 Announce Type: new Abstract: Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur…

    arxiv.org8 days agoView details

  57. Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

    arXiv:2609.09925v1 Announce Type: new Abstract: Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of coordinated motion emitted in a single forward pass. Yet an action chunk is essentially a short multivariate trajectory, but inside these models it is a sequence of gener…

    arxiv.org8 days agoView details

  58. Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety

    arXiv:2609.09735v1 Announce Type: new Abstract: Healthcare systems, mental health, and public well-being are increasingly affected by cyberbullying and harmful online interactions. This paper presents CareGuard, an early-warning framework designed to support healthcare-driven mental health protection and proactive onl…

    arxiv.org8 days agoView details

  59. Deep and shallow biases in language models

    arXiv:2609.09901v1 Announce Type: new Abstract: Large language models often repeatedly select the same answer even when many alternatives are plausible. Prior work treats this concentration as bias, but it does not distinguish stable model preferences from responses that depend on a particular prompt wording. We intro…

    arxiv.org8 days agoView details

  60. Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts

    arXiv:2609.10135v1 Announce Type: new Abstract: To address insufficient contextualization, weak generalization, and poor scenario adaptation in tourism meteorological services, we propose SmartWeatherAgent--a unified three-stage architecture integrating intent recognition, hazard prediction, and reasoning-enhanced gen…

    arxiv.org8 days agoView details

  61. Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection

    arXiv:2609.10221v2 Announce Type: new Abstract: Reinforcement learning over a frozen reasoner has become a common recipe for teaching a policy which external tools to invoke. We show that this recipe becomes structurally mismatched in specialist scientific settings where the complete tool-subset space is enumerable. T…

    arxiv.org8 days agoView details

  62. TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

    arXiv:2609.10315v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check. Diagnostic reasoning over complex data lacks this advantage: establishing the true cause…

    arxiv.org8 days agoView details

  63. From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning

    arXiv:2609.10335v1 Announce Type: new Abstract: Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonst…

    arxiv.org8 days agoView details

  64. Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

    arXiv:2609.02663v1 Announce Type: cross Abstract: Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematic…

    arxiv.org8 days agoView details

  65. SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

    arXiv:2609.09349v1 Announce Type: new Abstract: Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. We introduce Systematic Wikidata-based Object-Relation Distortion (SWORD), a benchmark th…

    arxiv.org8 days agoView details

  66. Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks

    arXiv:2609.09774v1 Announce Type: new Abstract: Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated. The study combines a retrospective, human-assisted interface-adaptation ca…

    arxiv.org8 days agoView details

  67. BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

    arXiv:2609.09554v1 Announce Type: new Abstract: We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are high…

    arxiv.org8 days agoView details

  68. GANDR: Claim Auditing for Verifiable Legal Answer Generation

    arXiv:2609.10293v1 Announce Type: new Abstract: In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can res…

    arxiv.org8 days agoView details

  69. Rosetta at AlexandriaX-2026: LoRA-Adapted NileChat for Context-Aware Dialectal Arabic Dialogue Translation

    arXiv:2609.10395v1 Announce Type: new Abstract: This paper describes the Rosetta system for Subtask 1 (Context-Aware English-to-Dialectal Arabic Dialogue Translation) of the AlexandriaX shared task, participating in both constrained and unconstrained tracks. The approach fine-tunes a LoRA adapter on NileChat-3B using…

    arxiv.org8 days agoView details

  70. Auditable Emergency Triage for Maternal and Newborn Care in India

    arXiv:2609.09356v1 Announce Type: new Abstract: At Noora Health, our nurses answer more than 50,000 medical queries per month on our WhatsApp-based service that provides caregivers with on-demand support. Their most time-critical task is emergency triage: deciding which queries need immediate in-person attention. To s…

    arxiv.org8 days agoView details

  71. Do LLMs Make More Mistakes If They Do Not Believe the Input Data?

    arXiv:2609.09363v1 Announce Type: new Abstract: Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the c…

    arxiv.org8 days agoView details

  72. Benchmarking Hybrid Deep Research Across Database Querying and Web Search

    arXiv:2609.09410v1 Announce Type: new Abstract: While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave t…

    arxiv.org8 days agoView details

  73. Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

    arXiv:2609.09425v1 Announce Type: new Abstract: Educational data filters have become a practical way to improve language-model pre-training, but most filters treat educational value as a single scalar property. This may be too broad for some applications, especially if the data set already features a high density of e…

    arxiv.org8 days agoView details

  74. TEFM: Token-Efficient Faithful Modeling for Structured Data

    arXiv:2609.09552v1 Announce Type: new Abstract: In this paper, we solve two fundamental obstacles in applying LLMs to critical domains: token efficiency and faithfulness. To address both constraints jointly, we present TEFM (Token-Efficient Faithful Modeling), a framework designed for structured data analysis in criti…

    arxiv.org8 days agoView details

  75. XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

    arXiv:2609.09428v1 Announce Type: new Abstract: Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can…

    arxiv.org8 days agoView details

  76. Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

    arXiv:2609.09678v1 Announce Type: new Abstract: Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or defer. Existing agent benchmarks largely evaluate accuracy after fixed or unconstrained interaction, leaving autonomous stopping reliability implicit. We present Cros,…

    arxiv.org8 days agoView details

  77. Towards Automatic Evolution Tree Generation from Citation Graphs

    arXiv:2609.09561v1 Announce Type: new Abstract: Surveys remain the primary way researchers grasp the lineage of methods within an AI subfield, but they scale poorly against the current rate of publication. Existing taxonomy-induction methods are largely leaf-bound and time-agnostic; they tend to force transitional pap…

    arxiv.org8 days agoView details

  78. Reproducing Omitted Temporal Expressions in Japanese News for Retrieval-Augmented Applications

    arXiv:2609.09569v1 Announce Type: new Abstract: News articles often contain omitted temporal expressions, such as day-only or month-only mentions, which must be interpreted with reference to the publication date. When such articles are indexed or processed as standalone text in search and retrieval-augmented generatio…

    arxiv.org8 days agoView details

  79. Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features

    arXiv:2609.09575v1 Announce Type: new Abstract: Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable features from dense representations, y…

    arxiv.org8 days agoView details

  80. SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

    arXiv:2609.09672v1 Announce Type: new Abstract: The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (SEA) languages critically underrepresented. We intr…

    arxiv.org8 days agoView details

  81. X2-NativeCursor: Native-Token Text Progress Tracking for Incremental-Text Streaming Codec TTS

    arXiv:2609.09677v1 Announce Type: new Abstract: Incremental-text streaming text-to-speech (TTS) needs online text progress tracking for synchronized highlighting, interruption handling, and dialogue-history updates. Input text arrives before it is spoken, so text arrival alone cannot indicate speech progress. Existing…

    arxiv.org8 days agoView details

  82. Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA

    arXiv:2609.09684v1 Announce Type: new Abstract: Medical question-answering datasets often contain answer labels, whereas high-quality rationales remain scarce, noisy, or costly to validate. This changes the acquisition question: rather than asking which questions should be labeled, we ask which already-labeled questio…

    arxiv.org8 days agoView details

  83. When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination

    arXiv:2609.09696v1 Announce Type: new Abstract: Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical…

    arxiv.org8 days agoView details

  84. Scaling E-Commerce Attribute Extraction with Parallel Decoding

    arXiv:2609.09716v1 Announce Type: new Abstract: Customers rely on specific product attributes to compare products and make purchasing decisions, but e-commerce catalogs are messy and unstructured, making it difficult to identify which attributes matter most and extract them at scale. Standard Attribute Value Extractio…

    arxiv.org8 days agoView details

  85. StreamAlign: Streaming Text-Aligned Speech Tokenization

    arXiv:2609.09719v1 Announce Type: new Abstract: Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognition (ASR), leading to two key limitations: (i) the ne…

    arxiv.org8 days agoView details

  86. SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

    arXiv:2609.09764v1 Announce Type: new Abstract: Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to building effective and trustworthy human-AI interaction. Existing reinforcement learning metho…

    arxiv.org8 days agoView details

  87. CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription

    arXiv:2609.09766v1 Announce Type: new Abstract: Churn models typically identify high-risk customers but do not specify which feasible retention action should be considered or why that action is appropriate. We present CARRE (Counterfactual Action Retrieval and Reason Evaluation), a three-stage framework that combines…

    arxiv.org8 days agoView details

  88. SymbolicLight V2: Hybrid Neuromorphic Architecture and Sparse Execution for Low-Energy Language Inference

    arXiv:2609.09772v1 Announce Type: new Abstract: SymbolicLight V2 combines sparse event computation with continuous-state processing in a hybrid neuromorphic language architecture. Extending V1's spike-gated dual paths, it adds graded signed events at further projections and softmax-free local attention. We implement t…

    arxiv.org8 days agoView details

  89. ROAM: Robust Organization of Atomic Memories for Agents through Semantic Relations

    arXiv:2609.09778v1 Announce Type: new Abstract: Long-term language-model agents rely on external memory across interactions. Atomic memories are particularly useful: their fine-grained semantic boundaries enable precise retrieval and direct comparison between observations. Yet accumulating atoms inevitably become redu…

    arxiv.org8 days agoView details

  90. MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short

    arXiv:2609.09791v1 Announce Type: new Abstract: With hate speech being ubiquitous online, automatic detection is crucial, in particular when it comes to criminally relevant social media posts. We study a variety of retrieval-based in-context learning (RetICL) strategies for detecting defamatory offences under {\S}{\S}…

    arxiv.org8 days agoView details

  91. HyperTrace: Hypothesis-Based Preference Tracing for Online LLM Personalization

    arXiv:2609.09835v1 Announce Type: new Abstract: Personalized language models aim to adapt responses to individual users, whose preferences are often latent and revealed gradually through interaction. Existing training-free methods rely on stored histories or retrieved memories, but they often struggle to reconcile lon…

    arxiv.org8 days agoView details

  92. $S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants

    arXiv:2609.09852v1 Announce Type: new Abstract: The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their perfor…

    arxiv.org8 days agoView details

  93. When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors

    arXiv:2609.09887v1 Announce Type: new Abstract: LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias…

    arxiv.org8 days agoView details

  94. Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses

    arXiv:2609.09902v1 Announce Type: new Abstract: Reading a transformer's internal states in token space is easy to do and hard to trust: a logit lens on a single hidden state is dominated, at intermediate layers, by the generic tokens the model would predict for almost any input. We read the difference instead. Subtrac…

    arxiv.org8 days agoView details

  95. Improving Cross-Lingual Token Representations by Adding a Pinch of SALT

    arXiv:2609.09953v1 Announce Type: new Abstract: Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are increasingly also applied to t…

    arxiv.org8 days agoView details

  96. 5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs

    arXiv:2609.09964v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies…

    arxiv.org8 days agoView details

  97. Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

    arXiv:2609.09989v1 Announce Type: new Abstract: A natural way to cut reasoning-model inference cost is to repeatedly probe a single partial trajectory for its current answer and stop once probes agree -- self-consensus. We ask whether any such rule is both safe and token-saving, and whether one can be selected once an…

    arxiv.org8 days agoView details

  98. MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

    arXiv:2609.10049v1 Announce Type: new Abstract: Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation w…

    arxiv.org8 days agoView details

  99. ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification

    arXiv:2609.10122v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance across a broad range of classification settings, yet the reliability of their predictions remains a major obstacle to deployment in high-stakes scenarios. Although confidence estimation for LLMs has been widel…

    arxiv.org8 days agoView details

  100. Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning

    arXiv:2609.10142v1 Announce Type: new Abstract: Large language models remain fragile against malicious fine-tuning, motivating training-time defenses against harmful persona drift. Preventative Steering injects undesirable-trait persona vectors during fine-tuning and removes them at evaluation time, yet the mechanism…

    arxiv.org8 days agoView details

  101. YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

    arXiv:2609.10153v1 Announce Type: new Abstract: Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from expl…

    arxiv.org8 days agoView details

  102. From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

    arXiv:2609.10155v1 Announce Type: new Abstract: We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 51…

    arxiv.org8 days agoView details

  103. Who Argues What? Joint Argument-Entity Detection and Classification in Political Debates

    arXiv:2609.10192v1 Announce Type: new Abstract: Political debates are often analyzed through Argument Mining (AM) to investigate the key arguments that drive them. However, political arguments are rarely interpretable from argumentative spans alone, as claims and premises generally depend on the entities (e.g., people…

    arxiv.org8 days agoView details

  104. Through the Looking Glass: Directly Reading and Writing Transformers

    arXiv:2609.10210v1 Announce Type: new Abstract: How many of a transformer's components decide a token? Counted by the absolute value of each unit's and channel's contribution to the logit, one prediction rests on thousands to hundreds of thousands of them. But contributions are signed, and across eighteen models the m…

    arxiv.org8 days agoView details

  105. $\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

    arXiv:2609.10226v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolat…

    arxiv.org8 days agoView details

  106. DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs

    arXiv:2609.10253v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Exis…

    arxiv.org8 days agoView details

  107. Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation

    arXiv:2609.09702v1 Announce Type: new Abstract: Candidate decision correctness and rationale grounding are different objectives. We examine correctness-gated multi-teacher distillation in a fixed experiment. Eight arms share 4,330 sources, a 63.9M-parameter student, 12,990 optimization rows, 406 updates, evidence inpu…

    arxiv.org8 days agoView details

  108. CityPlanner: A Sandbox Agent for Executable Urban Planning

    arXiv:2609.09578v1 Announce Type: new Abstract: Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed…

    arxiv.org8 days agoView details

  109. ConvMem: Convolutional Memory for Long-Context Reasoning

    arXiv:2609.10441v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and i…

    arxiv.org8 days agoView details

  110. From Plausible to Actionable: A Position on LLM Self-Explanations

    arXiv:2607.15957v3 Announce Type: cross Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations. Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI),…

    arxiv.org8 days agoView details

  111. Reliability-Aware Hybrid-K Ensemble Selection for Cervical Cytology Classification: Integrating Discrimination, Calibration, and Selective Prediction

    arXiv:2609.09189v1 Announce Type: cross Abstract: High classification accuracy alone is insufficient for clinical image analysis, where calibrated confidence and reliable uncertainty estimates are essential. This study proposes a reliability-aware Hybrid-K ensemble selection framework for multiclass cervical cytology…

    arxiv.org8 days agoView details

  112. AgenticGen: Reward-Guided Agentic Video Generation for Advertising

    arXiv:2609.09187v1 Announce Type: cross Abstract: Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do n…

    arxiv.org8 days agoView details

  113. Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection

    arXiv:2609.10244v1 Announce Type: new Abstract: We present our system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. A small ($4$B-parameter) VLM is fine-tuned as a per-token classifier reading a two-token feature from its own hidden states, and is ensembled with a $\sim$400B z…

    arxiv.org8 days agoView details

  114. The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs

    arXiv:2609.10237v1 Announce Type: new Abstract: A graph retrieval-augmented generation pipeline chooses which triples to put in the prompt, a syntax to write them in, an order to write them in, and a sentence telling the model what to do with them. We vary all four over six large language models and two knowledge-grap…

    arxiv.org8 days agoView details

  115. Politics of Feelings: Emotional Expression and Legislative Effectiveness in the U.S. Congress

    arXiv:2609.10198v1 Announce Type: new Abstract: Emotions are a pervasive feature of political communication, yet existing research has focused primarily on describing patterns of emotional expression rather than examining whether they are associated with consequential legislative outcomes. We address this gap by inves…

    arxiv.org8 days agoView details

  116. Seven Sources of Physical AI Capability Formation

    arXiv:2609.09627v1 Announce Type: new Abstract: Capabilities relevant to Physical AI can arise from materially different formation histories, yet existing taxonomies organized by morphology, architecture, learning algorithm, task, or domain do not directly answer what gives rise to a capability. We define a capability…

    arxiv.org8 days agoView details

  117. An Autonomous GeoAI Agent for Arctic Eco-Navigation

    arXiv:2609.09374v1 Announce Type: new Abstract: Arctic maritime navigation is becoming increasingly important as changing sea-ice conditions expand seasonal accessibility while simultaneously introducing substantial operational, environmental, and community risks. Arctic route planning is inherently a multi-criteria p…

    arxiv.org8 days agoView details

  118. RobustSGPO: Search-Space Control for Agent Harness Evolution

    arXiv:2609.09646v1 Announce Type: new Abstract: Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but its local update rule leaves the choice of edit scope and operation unresolved. We introduce RobustSGPO, which specifies the requested edit, constructs and checks th…

    arxiv.org8 days agoView details

  119. Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions

    arXiv:2609.09306v1 Announce Type: new Abstract: This paper investigates the hypothesis that the first-order structure of physical interactions, i.e. gradients or Jacobians, characterizes the structure of phenomenal experience. It does so in an idealized world inhabited by neural networks, Gradland, where the physics a…

    arxiv.org8 days agoView details

  120. The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding

    arXiv:2609.10296v1 Announce Type: new Abstract: Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are d…

    arxiv.org8 days agoView details

  121. SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers

    arXiv:2609.09999v1 Announce Type: new Abstract: Terminology-aware translation asks for more than a correct translation: the output must use the exact terms a glossary prescribes. The standard recipe, fine-tuning on glossary-annotated translation pairs, hides an inefficiency: for most examples the glossary prescribes e…

    arxiv.org8 days agoView details

  122. Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning

    arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pre…

    arxiv.org8 days agoView details

  123. Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

    arXiv:2609.10060v1 Announce Type: new Abstract: Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations a…

    arxiv.org8 days agoView details

  124. Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

    arXiv:2609.10410v1 Announce Type: new Abstract: The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remain…

    arxiv.org8 days agoView details

  125. Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

    arXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversi…

    arxiv.org8 days agoView details

  126. RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding

    arXiv:2609.10305v1 Announce Type: new Abstract: Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transformer at embedding width d = 128 still spends roughly one third of its capacity on the output matrix W_out in R^(d x |V|).…

    arxiv.org8 days agoView details

  127. StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

    arXiv:2609.09264v1 Announce Type: new Abstract: Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochast…

    arxiv.org8 days agoView details

  128. Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling

    arXiv:2609.09691v1 Announce Type: new Abstract: When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting…

    arxiv.org8 days agoView details

  129. Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language Models

    arXiv:2609.03247v1 Announce Type: cross Abstract: Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-id…

    arxiv.org8 days agoView details

  130. JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

    arXiv:2609.10451v1 Announce Type: new Abstract: Real-world GUI usage frequently involves workflows that span multiple devices and platforms, requiring the transfer of intermediate results, maintenance of shared state, and coordination across heterogeneous environments. However, existing GUI benchmarks overwhelmingly e…

    arxiv.org8 days agoView details

  131. Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services

    arXiv:2609.09889v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments such as call center conversations. When privacy restr…

    arxiv.org8 days agoView details

  132. Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records

    arXiv:2609.09984v1 Announce Type: new Abstract: Understanding the historical allocation and distribution of research funding advances our knowledge of how scientific research is supported across fields, institutions, and regions. However, large-scale analyses are hindered by the lack of comprehensive funder name disam…

    arxiv.org8 days agoView details

  133. Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training

    arXiv:2609.10052v1 Announce Type: new Abstract: LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study this problem as successful strategy coverag…

    arxiv.org8 days agoView details

  134. Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning

    arXiv:2609.10113v1 Announce Type: new Abstract: Financial text, textbooks, and question-answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. Existing QA pairs often lack explicit reasoning, sufficient context, or reliably verifiable answers, while textbooks must…

    arxiv.org8 days agoView details

  135. Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

    arXiv:2609.09898v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning descriptions into PDDL problem instances. However, standard evaluation criteria such as syntactic validity or planner success can substantially overestimate faithfulness to the…

    arxiv.org8 days agoView details

  136. ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

    arXiv:2609.09458v1 Announce Type: new Abstract: As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. O…

    arxiv.org8 days agoView details

  137. Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

    arXiv:2609.09448v1 Announce Type: new Abstract: As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, t…

    arxiv.org8 days agoView details

  138. KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

    arXiv:2609.10266v1 Announce Type: new Abstract: LLM serving systems already reuse KV caches, but only when the reused text sits at the very start of the prompt. Two growing workloads break this condition: a retrieval-augmented generation server assembles a different set of retrieved chunks for every query, and a multi…

    arxiv.org8 days agoView details

  139. Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

    arXiv:2609.09418v1 Announce Type: new Abstract: World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly advancing embodied AI, general-purpose counterparts remain largely unexplored in games. Existing game…

    arxiv.org8 days agoView details

  140. X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding

    arXiv:2609.09166v1 Announce Type: new Abstract: This paper investigates collaborative speculative decoding (CoSD), a distributed large language model (LLM) inference framework in which an on-device small language model (SLM) drafts candidate tokens and a server LLM verifies them. Existing CoSD methods assume a shared…

    arxiv.org8 days agoView details

  141. From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins

    arXiv:2609.09625v1 Announce Type: new Abstract: As Digital Twin (DT) systems evolve beyond state synchronization toward task-oriented and knowledge-driven operation, Cognitive Digital Twins (CDTs) have emerged as an extension that incorporates cognitive capabilities into twin operation. Existing CDT studies often focu…

    arxiv.org8 days agoView details

  142. OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

    arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes…

    arxiv.org8 days agoView details

  143. Adaptive Entangled Game Modules in Artificial General Intelligence

    arXiv:2609.09226v1 Announce Type: new Abstract: We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive agents, deriving testable eigenmodes through a generalized behavioral intelligence (GBI) nonlocal probability-wave equation. This framework captures a broad range of hu…

    arxiv.org8 days agoView details

  144. Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

    arXiv:2609.09233v1 Announce Type: new Abstract: How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bundles containing instructions, sc…

    arxiv.org8 days agoView details

  145. Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery

    arXiv:2609.09413v1 Announce Type: new Abstract: Choosing a recovery process for scale-up requires connecting laboratory results with product requirements, process costs, and scale effects. We analyze records from Pacific Northwest National Laboratory's Computer Intelligence for Critical Element Recovery and Optimizati…

    arxiv.org8 days agoView details

  146. Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

    arXiv:2609.09647v1 Announce Type: new Abstract: Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step…

    arxiv.org8 days agoView details

  147. RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

    arXiv:2609.09657v1 Announce Type: new Abstract: Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support c…

    arxiv.org8 days agoView details

  148. Kernel-Managed Shared Memory for System-Wide Personalization

    arXiv:2609.10144v1 Announce Type: new Abstract: AI systems become more useful when they can adapt to the people using them, but in multi-agent systems, useful context learned by one agent often remains unavailable to others. We present kernel-managed shared memory, a system-level abstraction in which specialized agent…

    arxiv.org8 days agoView details

  149. PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations

    arXiv:2609.09664v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as personalized assistants that interact with users over extended periods of time. As conversations grow longer, relying on full interaction histories becomes increasingly inefficient and unreliable: long contexts in…

    arxiv.org8 days agoView details

  150. The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

    arXiv:2609.09853v1 Announce Type: new Abstract: LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools. The benchmark is built around a complet…

    arxiv.org8 days agoView details

  151. LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

    arXiv:2609.09754v1 Announce Type: new Abstract: As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-…

    arxiv.org8 days agoView details

  152. Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

    arXiv:2609.09776v1 Announce Type: new Abstract: Frontier gains in language-model reasoning come from reinforcement learning on reasoning traces and are concentrated in domains with a cheap, sound verifier. We argue the field's binding constraint is the verification gap: no scalable, incorruptible reward for reasoning…

    arxiv.org8 days agoView details

  1. Edge0/Edge0-35B-A3B-preview

    text-generation · mlx · safetensors · qwen3_5_moe

    huggingface.co8 days ago3209 ptsView details

  2. openbmb/MiniCPM5-2B

    text-generation · transformers · safetensors · llama

    huggingface.co9 days ago1527 ptsView details

  3. m-a-p/YuE2-3B

    text-to-audio · safetensors · yue2 · music-generation

    huggingface.co8 days ago680 ptsView details

  4. tencent/AuK

    text-to-speech · audio · speech · text-to-speech

    huggingface.co8 days ago282 ptsView details

  5. openbmb/MiniCPM5-2B-GGUF

    text-generation · transformers · gguf · minicpm

    huggingface.co9 days ago246 ptsView details

  6. nvidia/Qwen3.8-27B-NVFP4

    text-generation · Model Optimizer · safetensors · qwen3_5

    huggingface.co9 days ago74 ptsView details

  7. Prannesshkva/ISOM-Falcon-40B

    text-generation · pytorch · safetensors · falcon

    huggingface.co9 days ago2 ptsView details

  8. Prannesshkva/ISOM-Qwen-1.5B-Instruct

    text-generation · safetensors · isom · qwen2

    huggingface.co9 days ago1 ptsView details

  9. Devlin-AI/Devlin-Alpha-22B-A3B-GGUF

    gguf · endpoints_compatible · region:us

    huggingface.co9 days agoView details

  10. YNSScarSaiyan/simi-weights

    text-generation · flax · simi · weights-only

    huggingface.co9 days agoView details

  1. langchain-ai/langchain langchain-openai==1.6.2

    Changes since langchain-openai==1.6.1 release(openai): 1.6.2 (#40339) fix(openai): add GPT-6 Astra reasoning efforts (#40330) chore(deps): bump httpx2 from 2.10.0 to 2.12.0 in /libs/partners/openai (#40309)

    github.com8 days agoView details

  2. ggml-org/llama.cpp b10883

    <details open> vulkan: use spec constant for matrix matrix multiplication A-type (#25773) * vulkan: use spec constant for mul mat type_a vulkan: use map for mul_mm shapes cleanup fix indentation fix cm2 and shmem init fix cm2 spec constants fix cm2 bindings consolidate shmem tab…

    github.com8 days agoView details

  3. huggingface/transformers v5.17.0

    # Release v5.17.0 ## New Model additions ### HYV4 <img width="1503" height="827" alt="image" src="https://github.com/user-attachments/assets/e6ed85ee-eb1d-40eb-a0d4-c649f6337ca9" /> Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters p…

    github.com8 days agoView details

  4. openai/openai-python v3.11.0

    ## [3.11.0](https://github.com/openai/openai-python/compare/v3.10.0...v3.11.0) (2026-09-09) ### Features * **api:** Add expiration controls for service account keys ([#3825](https://github.com/openai/openai-python/issues/3825)) ([f348ec8](https://github.com/openai/openai-python/…

    github.com8 days agoView details

  5. vllm-project/vllm v0.29.0

    # v0.29.0 ## Highlights This release features 594 commits from 277 contributors (91 new)! * **Model Runner V2 is now the default for all models** (#53183), completing the rollout that began with pooling models (#48290). MRV2 also gained CUDA graph memory profiling for KV cache a…

    github.com9 days agoView details

  6. ggml-org/llama.cpp b10872

    <details open> jinja: treat a null left operand of in as a plain lookup (#28620) Templates that default an optional variable to none and then test its membership in a map hit an error, while the same expression is a normal lookup returning false in Jinja. The undefined counterpa…

    github.com9 days agoView details

  7. ggml-org/llama.cpp b10871

    <details open> vulkan: add dedicated iq4_xs mat-vec shader (#28426) * vulkan: add dedicated iq4_xs mat-vec shader Dedicated mul_mat_vec_iq4_xs for the dmmv path, replacing the generic fallback. ~+6-17% token generation on RDNA4 depending on model. Assisted-by: Pi agent with Qwen…

    github.com9 days agoView details

  8. ggml-org/llama.cpp b10870

    <details open> vulkan: add f16 B-type matmul pipelines and warp tile size tuning for Intel coopmat1 (#27471) * vulkan: add f16 B-type matmul pipelines and warp tile size tuning for Intel coopmat1 * simplify mmp selection in mul_mat_id per review comment * vulkan: enable f16 B-ty…

    github.com9 days agoView details

Archive — 2026-09-09 · Engineerious