Archive / 2026-09-09
September 9, 2026
News
View all news →apple.com8 days ago504 ptsView detailsJoin discussion
GPT-6 Astra, looped transformers, and hidden reasoning
magazine.sebastianraschka.com8 days ago514 ptsView detailsJoin discussion
Astra for Coding: Why Are We Doing This Again?
lucumr.pocoo.org8 days ago449 ptsView detailsJoin discussion
nomanssky.com8 days ago457 ptsView detailsJoin discussion
Hitachi launches CO2 heat pump water heaters with solar-friendly tariff controls
pv-magazine.com8 days ago318 ptsView detailsJoin discussion
Blizzard Workers Win Historic Union Contract
latimes.com8 days ago249 ptsView detailsJoin discussion
Getting 50 GB/S Back from the Apple Neural Engine
eiln.github.io8 days ago217 ptsView detailsJoin discussion
gnuradioworld.com8 days ago224 ptsView detailsJoin discussion
Understanding the recent DDoS attack against Read the Docs
about.readthedocs.com8 days ago214 ptsView detailsJoin discussion
How An AI math breakthrough ignited a controversy
science.org8 days ago220 ptsView detailsJoin discussion
Muse, the band, lost its social media handles to Muse, Meta's new AI agent
engadget.com8 days ago184 ptsView detailsJoin discussion
Lotus Notes and the dangers of starting from scratch
buttondown.com8 days ago165 ptsView detailsJoin discussion
Training a 3.8B LLM to 0.384 CORE for $998
hugovergnes.github.io8 days ago120 ptsView detailsJoin discussion
A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
slimemoldtimemold.com8 days ago94 ptsView detailsJoin discussion
I'm sorry, you're not going to die from an AI-engineered supervirus
blog.genesmindsmachines.com8 days ago84 ptsView detailsJoin discussion
We accidentally built a synthetic cell factory
bnext.bio8 days ago83 ptsView detailsJoin discussion
Building a Wall Lamp from Scratch
mbugert.de8 days ago85 ptsView detailsJoin discussion
Gambling with our lives: AI researcher quits Anthropic with warning about safety
politico.eu9 days ago81 ptsView detailsJoin discussion
Show HN: Compute polynomials twice as fast
thomasahle.com9 days ago76 ptsView detailsJoin discussion
Microsoft/TracerAI withdraws copyright takedown against Luanti
blog.luanti.org8 days ago64 ptsView detailsJoin discussion
Rock band Muse lose social media handles to Meta’s new AI tool
the-independent.com8 days ago60 ptsView detailsJoin discussion
Defining AI Psychosis. Part 2: "Prolific AI Psychosis"
jeffs.blog8 days ago60 ptsView detailsJoin discussion
I Don't Want to Interact with Stochastic Parrots
ploum.net8 days ago58 ptsView detailsJoin discussion
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
arxiv.org8 days ago57 ptsView detailsJoin discussion
Show HN: DOOM in the kernel, or fibers in eBPF
ayles.github.io8 days ago53 ptsView detailsJoin discussion
OpenAI's rogue agents used at least 10 more sites
reuters.com8 days ago53 ptsView detailsJoin discussion
Show HN: Apollo Lunar Module landing simulation
gosandeep.com8 days ago50 ptsView detailsJoin discussion
Show HN: Self-hosted company OS, Claude Code and Codex agents in departments
github.com8 days ago50 ptsView detailsJoin discussion
Show HN: Geiger – See every AI agent on your machine and what it can touch
github.com8 days ago49 ptsView detailsJoin discussion
projectzero.google8 days ago45 ptsView detailsJoin discussion
Anthropic researcher says more than 10% chance AI "could kill all humans"
cbsnews.com8 days ago47 ptsView detailsJoin discussion
Better AI code comment detector
entropicthoughts.com8 days ago42 ptsView detailsJoin discussion
DHS Program Analyzes Americans' Finances to Flag Drivers for Traffic Stops
military.com8 days ago35 ptsView detailsJoin discussion
Do people prefer stories written by AI?
cambridge.org8 days ago36 ptsView detailsJoin discussion
My Mental Model of AI Broke on September 8
rough-ideas.bearblog.dev8 days ago32 ptsView detailsJoin discussion
AI could kill all humans in next decade, warn experts
theguardian.com8 days ago31 ptsView detailsJoin discussion
Jacob Coxon resignation appears to be a PR stunt for AI regulation
twitter.com8 days ago28 ptsView detailsJoin discussion
Rails 8 Guide: Features, Requirements and Upgrade Path (2026)
blog.appsignal.com8 days ago29 ptsView detailsJoin discussion
Carmakers Have a New Idea to Boost EV Range: Add a Gas Engine
wsj.com8 days ago26 ptsView detailsJoin discussion
Show HN: Rdltr – Inbox zero for your reading list
rdltr.app8 days ago24 ptsView detailsJoin discussion
Good Taste Can't Be Taught, Bought or Learned, Sorry AI
emilyoberg.substack.com8 days ago23 ptsView detailsJoin discussion
- Primary source
Build and launch cloud agents with the Agents API, a managed service powered by the Codex harness for orchestration, long-running sessions, and tool use.
openai.com8 days ago13 ptsView details
en.wikipedia.org8 days ago16 ptsView detailsJoin discussion
git.sr.ht8 days ago16 ptsView detailsJoin discussion
lonriesberg.com8 days ago15 ptsView detailsJoin discussion
Show HN: Hydra – Open-source agentic terminal with a PTY daemon
github.com8 days ago14 ptsView detailsJoin discussion
I'm not an economist, but here's an argument I've been thinking about
twitter.com8 days ago14 ptsView detailsJoin discussion
United States Ships War Material to Israel Aboard Passenger Flights
theintercept.com8 days ago13 ptsView detailsJoin discussion
Show HN: Ctrlb-decompose: Strip the noise from logs before sending to LLMs
github.com8 days ago14 ptsView detailsJoin discussion
Google: Attackers are using prompt injection against coding agents
cloud.google.com8 days ago13 ptsView detailsJoin discussion
Anthropic discloses fourth AI hacking incident missed in earlier review
reuters.com8 days ago12 ptsView detailsJoin discussion
Blizzard Entertainment Workers Ratify Video Game Contracts with CWA
cwa-union.org8 days ago12 ptsView detailsJoin discussion
Law schools tell students to put AI away
ft.com9 days ago12 ptsView detailsJoin discussion
YouTube cracks down on 'AI ghost creators'
semafor.com9 days ago12 ptsView detailsJoin discussion
Show HN: Maxxwell – The IDE for Optimal Tokenmaxxing
maxxwell.dev8 days ago11 ptsView detailsJoin discussion
The Einstein test: what happens when AI tries to rediscover relativity?
nature.com8 days ago10 ptsView detailsJoin discussion
- Primary source
Build more natural voice experiences with GPT‑Live‑1 in the API
GPT‑Live‑1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.
openai.com8 days agoView details
- Primary source
Paul Christiano joins OpenAI Foundation Board
Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.
openai.com8 days agoView details
- Primary source
The AI policy window is open. We need to act.
Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.
openai.com8 days agoView details
- Primary source
GPT-6 Astra: The next generation in intelligence for work
Meet GPT-6 Astra, OpenAI’s most capable model for business, with advanced reasoning, computer use, and stronger writing and design judgment.
openai.com8 days agoView details
- Primary source
Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
huggingface.co8 days agoView details
- Primary source
IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license
huggingface.co8 days agoView details
LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
LandingAI has shipped Agentic Document Extraction Gen2, a rebuild of its document stack on the DPT-3 model family. Chunks are retired in favor of a document, page and block tree. DPT-3 Pro grounds to the line, DPT-3 Verity grounds to the word with a confidence score, and Parse billing now counts output characters inst…
marktechpost.com8 days agoView details
Google has open-sourced Mantis, a stack-agnostic toolkit of security review skills for AI coding agents. It runs the full vulnerability lifecycle: sweep the code, filter false positives, reproduce the bug in a sandbox, patch it, re-attack the patch, then score the risk. Apache 2.0, and documented as demonstration-only…
marktechpost.com8 days agoView details
Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds
Voice agent teams keep hitting the same wall. The catalog holds 400 voices and the brief asks for the one that is not in it: a Quebecoise receptionist for a Montreal dealership, a narrator in his sixties with lecture hall authority. Briefs outnumber any catalog, and cloning closes the gap one speaker at a time, […] Th…
marktechpost.com9 days agoView details
The AI Researcher Who Just Quit Anthropic Says It’s ‘Crunch Time for Humanity’
Jacob Coxon talks to WIRED about the “mini Manhattan project” inside Anthropic, the problem with alignment, and why AI labs have just a few years left to make their systems safe.
wired.com8 days agoView details
Suno releases its first AI music model made with record industry help
Suno's new v6 AI music model is its first made with support from the record industry. Suno's Jack Brody told The Verge that v6 was "trained from the ground up, with a new set of data that does not include the same data that our previous models were trained on." The data includes content licensed from partners Warner M…
theverge.com8 days agoView details
OpenAI’s sly mathematical breakthrough sends a chill through academia
Open AI CEO Sam Altman speaks during the G20 Innovation Ministerial. | (Photo by Matt RAMEY / AFP via Getty Images) OpenAI's announcement Tuesday that it has solved one of mathematics' legendary Millennium Prize problems should have been a moment of triumph. The result is both an undeniable achievement and a striking…
theverge.com8 days agoView details
San Francisco Orders Meta to Stop ‘Allowing’ AI Child Abuse Ads
The City Attorney’s Office has asked Meta to explain how the harmful ads repeatedly ran on Facebook and Instagram. The company claims the ads are not under the city’s jurisdiction.
wired.com8 days agoView details
Read the Apple document explaining how new listening features still protect your privacy
At Wednesday's iPhone Duo launch event, Apple announced a handful of new Siri AI Audio Intelligence features, including Siri Recap, Live Rewind, Sound Recognition, and Music Recognition. Alongside its announcement, Apple released a document laying out how it plans to balance AI "ambient listening" and users' privacy.…
theverge.com8 days agoView details
Apple’s new iPhone camera mode promises to prove your photo isn’t AI
Apple is launching a new way to prove that the picture you took isn't manipulated by AI. A new feature, called "Reference Image," will arrive with the iPhone 18 Pro lineup later this month and is supposed to use the device's new camera sensor to "sign every pixel it sees." The iPhone 18 Pro and Pro Max will only authe…
theverge.com8 days agoView details
I Let an AI Agent Hack All My Gadgets—and I’d Do It Again
After I removed the safety guardrails from a powerful open-source model, it found vulnerabilities in my household devices and hacked into a PC. But it also told me how to make everything a lot more secure.
wired.com8 days agoView details
Microsoft has new AI privacy rules for schools
Microsoft agreed to a set of safety and privacy principles for AI in schools a week after two major school systems announced a ban on student-facing AI. In a new agreement with the American Federation of Teachers (AFT), the second-largest teachers union in the US, and its New York City affiliate the United Federation…
theverge.com8 days agoView details
A Stealth Startup Thinks It Just Hacked the Memory Shortage
Kepler Computing claims a new approach to chip design—and a proprietary material—can help end the supply bottlenecks that have sent memory prices surging.
wired.com8 days agoView details
Amazon Prime Video’s new AI tech matches lips to dubbed audio
Maxton Hall. | Image: Prime Video Amazon's Prime Video is launching a new AI-powered feature that lines up an actor's mouth with "human-dubbed" audio. The feature is only available with the English dub of the German series Maxton Hall for now, but Prime Video plans to expand it to "additional titles" in the future. In…
theverge.com8 days agoView details
Students who use AI generally score worse at school
Students who use AI to help them study tend to perform worse at school than those who don't, according to data from a global OECD educational report. The situation is more complex than it sounds though, with certain types of AI use giving learners a slight boost, especially among students taught to critically assess h…
theverge.com8 days agoView details
Worried Anthropic researchers warn that AI ‘could kill all humans’
A senior Anthropic safety researcher has said there is more than a 10 percent chance artificial intelligence "could kill all humans" by the end of the decade, just hours after a colleague resigned over fears the AI lab and its rivals are carelessly racing to build "superhuman systems" they cannot control. In a post on…
theverge.com8 days agoView details
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model
arXiv:2609.09815v1 Announce Type: new Abstract: Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting meta-agent reads workers' outputs, writes the final answer, allocates later calls, and decides when to stop. It is expressive, but it also concentrates three control decis…
arxiv.org8 days agoView details
arXiv:2609.09864v1 Announce Type: new Abstract: Affective computing has largely followed an individual-state paradigm, extracting discrete emotion labels or arousal/valence from isolated speakers. We argue this framing is incomplete for interaction. Drawing on affective resonance and vitality-contour accounts, we prop…
arxiv.org8 days agoView details
arXiv:2609.09882v1 Announce Type: new Abstract: Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these readouts are often treated as interchangeable. Holding model checkpoint and prompt content fixed, we compare probabilities obtained by scoring answer tokens with pre…
arxiv.org8 days agoView details
RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases
arXiv:2609.10092v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a ro…
arxiv.org8 days agoView details
Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications
arXiv:2609.09885v1 Announce Type: new Abstract: This paper studies unmanned aerial vehicle (UAV)-mouted reconfigurable intelligent surface (RIS)-assisted device-to-device (D2D) communication with stochastic link activation. It models UAV motion and attitude, time-varying Rician angles, and angle-dependent RIS reflecti…
arxiv.org8 days agoView details
Structural Process Supervision for Latent Chain-of-Thought Reasoning
arXiv:2609.09928v1 Announce Type: new Abstract: Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit chain-of-thought (CoT) tokens with compact continuous-space embeddings. However, existing methods lack direct process supervision over these latent embeddings, which…
arxiv.org8 days agoView details
Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability
arXiv:2609.10036v1 Announce Type: new Abstract: Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments. A single informative observation c…
arxiv.org8 days agoView details
AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents
arXiv:2609.09212v1 Announce Type: cross Abstract: This paper presents an end-to-end evaluation framework for image-triggered command injection against computer-use agents (CUAs). The goal is to test whether a local visual patch can induce verifiable environmental consequences along the full chain of screenshot input,…
arxiv.org8 days agoView details
Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs
arXiv:2609.10413v1 Announce Type: new Abstract: Current LLM memory systems treat all personal facts identically, so stores grow without bound while retrieval precision degrades. The core challenge is lifecycle management: which memories should persist, which should be replaced, and at what rate, conditioned on the beh…
arxiv.org8 days agoView details
arXiv:2609.10055v1 Announce Type: new Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle distinctions among hierarchically related concepts c…
arxiv.org8 days agoView details
arXiv:2609.10177v1 Announce Type: new Abstract: In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface level imitation of in-context demonstrations, maki…
arxiv.org8 days agoView details
arXiv:2609.10350v1 Announce Type: new Abstract: The banking system now depends on a small set of shared artificial intelligence vendors for fraud screening, credit decisioning, anti-money-laundering triage, customer analytics, and internal decision support. This paper studies how a compromise inside one of those vendo…
arxiv.org8 days agoView details
Quantifying Logical Consistency in Transformers via Query-Key Alignment
arXiv:2502.17017v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated impressive performance in various natural language processing tasks, yet their ability to perform multi-step logical reasoning remains an open challenge. Although Chain-of-Thought prompting has improved logical reasoning b…
arxiv.org8 days agoView details
Multi-Agent Agentic Graph Learning via Structural Signatures
arXiv:2609.09565v1 Announce Type: new Abstract: Agentic graph learning (AGL) has recently achieved promising results on graph reasoning tasks, where an agent powered by a large language model (LLM) sequentially samples the graph as evidence to support its final prediction. Existing methods either employ a single agent…
arxiv.org8 days agoView details
arXiv:2609.09213v1 Announce Type: cross Abstract: We study how physical-state inputs affect a 0.8B hybrid language model adapted for manipulation with 6.2M trainable parameters. Six conditions are trained on three LIBERO-Spatial tasks and evaluated over three seeds and 540 held-out rollouts. Conditioning recurrent dec…
arxiv.org8 days agoView details
arXiv:2609.09219v1 Announce Type: cross Abstract: AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement…
arxiv.org8 days agoView details
What Should an Agent Forget? Separating What Is Stored from What Is Used
arXiv:2609.10263v1 Announce Type: new Abstract: Persistent language agents need stored experience to remain available across time, while each answer requires evidence suited to a particular question. A superseded fact can mislead a current-state answer and still be essential for a historical query. We present RD-Forge…
arxiv.org8 days agoView details
The Mutations of Machine Speech
arXiv:2609.09496v1 Announce Type: new Abstract: Algorithmic outputs now populate the digital environments through which contemporary life is organized. The role of law in facilitating and constituting (rather than merely responding to) these processes is gaining increasing traction across scholarly accounts. This inqu…
arxiv.org8 days agoView details
VLX-VR: An Agentic-Aware Video Reasoning Model
arXiv:2609.09985v1 Announce Type: new Abstract: Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition when observations are incomplete,…
arxiv.org8 days agoView details
arXiv:2609.10321v1 Announce Type: new Abstract: Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from the teacher prediction and applied uniformly to all…
arxiv.org8 days agoView details
Towards Stress-Aware Sentence-Level Filipino G2P With Weakly-Supervised ByT5 Fine-Tuning
arXiv:2609.09974v1 Announce Type: new Abstract: Grapheme-to-phoneme conversion (G2P) refers to the task of converting a sequence of graphemes to a corresponding sequence of phonemes. While Filipino G2P is fairly straightforward due to its shallow orthography, the inclusion of prosodic features such as stress adds a la…
arxiv.org8 days agoView details
The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents
arXiv:2609.09395v1 Announce Type: new Abstract: Language models act through tools, yet practical agents face libraries containing thousands of interfaces. We introduce the tool menu as the short, ordered subset of available tools shown to an agent before execution. The agent can call only tools in this menu. Multi-ste…
arxiv.org8 days agoView details
A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
arXiv:2609.09589v1 Announce Type: new Abstract: Deep neural networks exhibit regular macroscopic behavior despite highly nonlinear dynamics in vast parameter spaces. We develop a statistical-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and…
arxiv.org8 days agoView details
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
arXiv:2609.09875v1 Announce Type: new Abstract: Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur…
arxiv.org8 days agoView details
Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models
arXiv:2609.09925v1 Announce Type: new Abstract: Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of coordinated motion emitted in a single forward pass. Yet an action chunk is essentially a short multivariate trajectory, but inside these models it is a sequence of gener…
arxiv.org8 days agoView details
arXiv:2609.09735v1 Announce Type: new Abstract: Healthcare systems, mental health, and public well-being are increasingly affected by cyberbullying and harmful online interactions. This paper presents CareGuard, an early-warning framework designed to support healthcare-driven mental health protection and proactive onl…
arxiv.org8 days agoView details
Deep and shallow biases in language models
arXiv:2609.09901v1 Announce Type: new Abstract: Large language models often repeatedly select the same answer even when many alternatives are plausible. Prior work treats this concentration as bias, but it does not distinguish stable model preferences from responses that depend on a particular prompt wording. We intro…
arxiv.org8 days agoView details
Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts
arXiv:2609.10135v1 Announce Type: new Abstract: To address insufficient contextualization, weak generalization, and poor scenario adaptation in tourism meteorological services, we propose SmartWeatherAgent--a unified three-stage architecture integrating intent recognition, hazard prediction, and reasoning-enhanced gen…
arxiv.org8 days agoView details
Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection
arXiv:2609.10221v2 Announce Type: new Abstract: Reinforcement learning over a frozen reasoner has become a common recipe for teaching a policy which external tools to invoke. We show that this recipe becomes structurally mismatched in specialist scientific settings where the complete tool-subset space is enumerable. T…
arxiv.org8 days agoView details
TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
arXiv:2609.10315v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check. Diagnostic reasoning over complex data lacks this advantage: establishing the true cause…
arxiv.org8 days agoView details
arXiv:2609.10335v1 Announce Type: new Abstract: Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonst…
arxiv.org8 days agoView details
arXiv:2609.02663v1 Announce Type: cross Abstract: Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematic…
arxiv.org8 days agoView details
arXiv:2609.09349v1 Announce Type: new Abstract: Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. We introduce Systematic Wikidata-based Object-Relation Distortion (SWORD), a benchmark th…
arxiv.org8 days agoView details
Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks
arXiv:2609.09774v1 Announce Type: new Abstract: Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated. The study combines a retrospective, human-assisted interface-adaptation ca…
arxiv.org8 days agoView details
BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models
arXiv:2609.09554v1 Announce Type: new Abstract: We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are high…
arxiv.org8 days agoView details
GANDR: Claim Auditing for Verifiable Legal Answer Generation
arXiv:2609.10293v1 Announce Type: new Abstract: In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can res…
arxiv.org8 days agoView details
arXiv:2609.10395v1 Announce Type: new Abstract: This paper describes the Rosetta system for Subtask 1 (Context-Aware English-to-Dialectal Arabic Dialogue Translation) of the AlexandriaX shared task, participating in both constrained and unconstrained tracks. The approach fine-tunes a LoRA adapter on NileChat-3B using…
arxiv.org8 days agoView details
Auditable Emergency Triage for Maternal and Newborn Care in India
arXiv:2609.09356v1 Announce Type: new Abstract: At Noora Health, our nurses answer more than 50,000 medical queries per month on our WhatsApp-based service that provides caregivers with on-demand support. Their most time-critical task is emergency triage: deciding which queries need immediate in-person attention. To s…
arxiv.org8 days agoView details
Do LLMs Make More Mistakes If They Do Not Believe the Input Data?
arXiv:2609.09363v1 Announce Type: new Abstract: Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the c…
arxiv.org8 days agoView details
Benchmarking Hybrid Deep Research Across Database Querying and Web Search
arXiv:2609.09410v1 Announce Type: new Abstract: While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave t…
arxiv.org8 days agoView details
Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements
arXiv:2609.09425v1 Announce Type: new Abstract: Educational data filters have become a practical way to improve language-model pre-training, but most filters treat educational value as a single scalar property. This may be too broad for some applications, especially if the data set already features a high density of e…
arxiv.org8 days agoView details
TEFM: Token-Efficient Faithful Modeling for Structured Data
arXiv:2609.09552v1 Announce Type: new Abstract: In this paper, we solve two fundamental obstacles in applying LLMs to critical domains: token efficiency and faithfulness. To address both constraints jointly, we present TEFM (Token-Efficient Faithful Modeling), a framework designed for structured data analysis in criti…
arxiv.org8 days agoView details
XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?
arXiv:2609.09428v1 Announce Type: new Abstract: Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can…
arxiv.org8 days agoView details
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
arXiv:2609.09678v1 Announce Type: new Abstract: Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or defer. Existing agent benchmarks largely evaluate accuracy after fixed or unconstrained interaction, leaving autonomous stopping reliability implicit. We present Cros,…
arxiv.org8 days agoView details
Towards Automatic Evolution Tree Generation from Citation Graphs
arXiv:2609.09561v1 Announce Type: new Abstract: Surveys remain the primary way researchers grasp the lineage of methods within an AI subfield, but they scale poorly against the current rate of publication. Existing taxonomy-induction methods are largely leaf-bound and time-agnostic; they tend to force transitional pap…
arxiv.org8 days agoView details
Reproducing Omitted Temporal Expressions in Japanese News for Retrieval-Augmented Applications
arXiv:2609.09569v1 Announce Type: new Abstract: News articles often contain omitted temporal expressions, such as day-only or month-only mentions, which must be interpreted with reference to the publication date. When such articles are indexed or processed as standalone text in search and retrieval-augmented generatio…
arxiv.org8 days agoView details
Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features
arXiv:2609.09575v1 Announce Type: new Abstract: Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable features from dense representations, y…
arxiv.org8 days agoView details
SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia
arXiv:2609.09672v1 Announce Type: new Abstract: The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (SEA) languages critically underrepresented. We intr…
arxiv.org8 days agoView details
X2-NativeCursor: Native-Token Text Progress Tracking for Incremental-Text Streaming Codec TTS
arXiv:2609.09677v1 Announce Type: new Abstract: Incremental-text streaming text-to-speech (TTS) needs online text progress tracking for synchronized highlighting, interruption handling, and dialogue-history updates. Input text arrives before it is spoken, so text arrival alone cannot indicate speech progress. Existing…
arxiv.org8 days agoView details
Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA
arXiv:2609.09684v1 Announce Type: new Abstract: Medical question-answering datasets often contain answer labels, whereas high-quality rationales remain scarce, noisy, or costly to validate. This changes the acquisition question: rather than asking which questions should be labeled, we ask which already-labeled questio…
arxiv.org8 days agoView details
arXiv:2609.09696v1 Announce Type: new Abstract: Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical…
arxiv.org8 days agoView details
Scaling E-Commerce Attribute Extraction with Parallel Decoding
arXiv:2609.09716v1 Announce Type: new Abstract: Customers rely on specific product attributes to compare products and make purchasing decisions, but e-commerce catalogs are messy and unstructured, making it difficult to identify which attributes matter most and extract them at scale. Standard Attribute Value Extractio…
arxiv.org8 days agoView details
StreamAlign: Streaming Text-Aligned Speech Tokenization
arXiv:2609.09719v1 Announce Type: new Abstract: Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognition (ASR), leading to two key limitations: (i) the ne…
arxiv.org8 days agoView details
arXiv:2609.09764v1 Announce Type: new Abstract: Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to building effective and trustworthy human-AI interaction. Existing reinforcement learning metho…
arxiv.org8 days agoView details
CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription
arXiv:2609.09766v1 Announce Type: new Abstract: Churn models typically identify high-risk customers but do not specify which feasible retention action should be considered or why that action is appropriate. We present CARRE (Counterfactual Action Retrieval and Reason Evaluation), a three-stage framework that combines…
arxiv.org8 days agoView details
arXiv:2609.09772v1 Announce Type: new Abstract: SymbolicLight V2 combines sparse event computation with continuous-state processing in a hybrid neuromorphic language architecture. Extending V1's spike-gated dual paths, it adds graded signed events at further projections and softmax-free local attention. We implement t…
arxiv.org8 days agoView details
ROAM: Robust Organization of Atomic Memories for Agents through Semantic Relations
arXiv:2609.09778v1 Announce Type: new Abstract: Long-term language-model agents rely on external memory across interactions. Atomic memories are particularly useful: their fine-grained semantic boundaries enable precise retrieval and direct comparison between observations. Yet accumulating atoms inevitably become redu…
arxiv.org8 days agoView details
arXiv:2609.09791v1 Announce Type: new Abstract: With hate speech being ubiquitous online, automatic detection is crucial, in particular when it comes to criminally relevant social media posts. We study a variety of retrieval-based in-context learning (RetICL) strategies for detecting defamatory offences under {\S}{\S}…
arxiv.org8 days agoView details
HyperTrace: Hypothesis-Based Preference Tracing for Online LLM Personalization
arXiv:2609.09835v1 Announce Type: new Abstract: Personalized language models aim to adapt responses to individual users, whose preferences are often latent and revealed gradually through interaction. Existing training-free methods rely on stored histories or retrieved memories, but they often struggle to reconcile lon…
arxiv.org8 days agoView details
$S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants
arXiv:2609.09852v1 Announce Type: new Abstract: The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their perfor…
arxiv.org8 days agoView details
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors
arXiv:2609.09887v1 Announce Type: new Abstract: LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias…
arxiv.org8 days agoView details
Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses
arXiv:2609.09902v1 Announce Type: new Abstract: Reading a transformer's internal states in token space is easy to do and hard to trust: a logit lens on a single hidden state is dominated, at intermediate layers, by the generic tokens the model would predict for almost any input. We read the difference instead. Subtrac…
arxiv.org8 days agoView details
Improving Cross-Lingual Token Representations by Adding a Pinch of SALT
arXiv:2609.09953v1 Announce Type: new Abstract: Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are increasingly also applied to t…
arxiv.org8 days agoView details
5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs
arXiv:2609.09964v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies…
arxiv.org8 days agoView details
Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal
arXiv:2609.09989v1 Announce Type: new Abstract: A natural way to cut reasoning-model inference cost is to repeatedly probe a single partial trajectory for its current answer and stop once probes agree -- self-consensus. We ask whether any such rule is both safe and token-saving, and whether one can be selected once an…
arxiv.org8 days agoView details
arXiv:2609.10049v1 Announce Type: new Abstract: Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation w…
arxiv.org8 days agoView details
ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification
arXiv:2609.10122v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance across a broad range of classification settings, yet the reliability of their predictions remains a major obstacle to deployment in high-stakes scenarios. Although confidence estimation for LLMs has been widel…
arxiv.org8 days agoView details
arXiv:2609.10142v1 Announce Type: new Abstract: Large language models remain fragile against malicious fine-tuning, motivating training-time defenses against harmful persona drift. Preventative Steering injects undesirable-trait persona vectors during fine-tuning and removes them at evaluation time, yet the mechanism…
arxiv.org8 days agoView details
YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models
arXiv:2609.10153v1 Announce Type: new Abstract: Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from expl…
arxiv.org8 days agoView details
arXiv:2609.10155v1 Announce Type: new Abstract: We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 51…
arxiv.org8 days agoView details
Who Argues What? Joint Argument-Entity Detection and Classification in Political Debates
arXiv:2609.10192v1 Announce Type: new Abstract: Political debates are often analyzed through Argument Mining (AM) to investigate the key arguments that drive them. However, political arguments are rarely interpretable from argumentative spans alone, as claims and premises generally depend on the entities (e.g., people…
arxiv.org8 days agoView details
Through the Looking Glass: Directly Reading and Writing Transformers
arXiv:2609.10210v1 Announce Type: new Abstract: How many of a transformer's components decide a token? Counted by the absolute value of each unit's and channel's contribution to the logit, one prediction rests on thousands to hundreds of thousands of them. But contributions are signed, and across eighteen models the m…
arxiv.org8 days agoView details
$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
arXiv:2609.10226v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolat…
arxiv.org8 days agoView details
arXiv:2609.10253v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Exis…
arxiv.org8 days agoView details
arXiv:2609.09702v1 Announce Type: new Abstract: Candidate decision correctness and rationale grounding are different objectives. We examine correctness-gated multi-teacher distillation in a fixed experiment. Eight arms share 4,330 sources, a 63.9M-parameter student, 12,990 optimization rows, 406 updates, evidence inpu…
arxiv.org8 days agoView details
CityPlanner: A Sandbox Agent for Executable Urban Planning
arXiv:2609.09578v1 Announce Type: new Abstract: Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed…
arxiv.org8 days agoView details
ConvMem: Convolutional Memory for Long-Context Reasoning
arXiv:2609.10441v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and i…
arxiv.org8 days agoView details
From Plausible to Actionable: A Position on LLM Self-Explanations
arXiv:2607.15957v3 Announce Type: cross Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations. Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI),…
arxiv.org8 days agoView details
arXiv:2609.09189v1 Announce Type: cross Abstract: High classification accuracy alone is insufficient for clinical image analysis, where calibrated confidence and reliable uncertainty estimates are essential. This study proposes a reliability-aware Hybrid-K ensemble selection framework for multiclass cervical cytology…
arxiv.org8 days agoView details
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
arXiv:2609.09187v1 Announce Type: cross Abstract: Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do n…
arxiv.org8 days agoView details
Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection
arXiv:2609.10244v1 Announce Type: new Abstract: We present our system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. A small ($4$B-parameter) VLM is fine-tuned as a per-token classifier reading a two-token feature from its own hidden states, and is ensembled with a $\sim$400B z…
arxiv.org8 days agoView details
The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs
arXiv:2609.10237v1 Announce Type: new Abstract: A graph retrieval-augmented generation pipeline chooses which triples to put in the prompt, a syntax to write them in, an order to write them in, and a sentence telling the model what to do with them. We vary all four over six large language models and two knowledge-grap…
arxiv.org8 days agoView details
Politics of Feelings: Emotional Expression and Legislative Effectiveness in the U.S. Congress
arXiv:2609.10198v1 Announce Type: new Abstract: Emotions are a pervasive feature of political communication, yet existing research has focused primarily on describing patterns of emotional expression rather than examining whether they are associated with consequential legislative outcomes. We address this gap by inves…
arxiv.org8 days agoView details
Seven Sources of Physical AI Capability Formation
arXiv:2609.09627v1 Announce Type: new Abstract: Capabilities relevant to Physical AI can arise from materially different formation histories, yet existing taxonomies organized by morphology, architecture, learning algorithm, task, or domain do not directly answer what gives rise to a capability. We define a capability…
arxiv.org8 days agoView details
An Autonomous GeoAI Agent for Arctic Eco-Navigation
arXiv:2609.09374v1 Announce Type: new Abstract: Arctic maritime navigation is becoming increasingly important as changing sea-ice conditions expand seasonal accessibility while simultaneously introducing substantial operational, environmental, and community risks. Arctic route planning is inherently a multi-criteria p…
arxiv.org8 days agoView details
RobustSGPO: Search-Space Control for Agent Harness Evolution
arXiv:2609.09646v1 Announce Type: new Abstract: Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but its local update rule leaves the choice of edit scope and operation unresolved. We introduce RobustSGPO, which specifies the requested edit, constructs and checks th…
arxiv.org8 days agoView details
Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions
arXiv:2609.09306v1 Announce Type: new Abstract: This paper investigates the hypothesis that the first-order structure of physical interactions, i.e. gradients or Jacobians, characterizes the structure of phenomenal experience. It does so in an idealized world inhabited by neural networks, Gradland, where the physics a…
arxiv.org8 days agoView details
The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
arXiv:2609.10296v1 Announce Type: new Abstract: Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are d…
arxiv.org8 days agoView details
SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers
arXiv:2609.09999v1 Announce Type: new Abstract: Terminology-aware translation asks for more than a correct translation: the output must use the exact terms a glossary prescribes. The standard recipe, fine-tuning on glossary-annotated translation pairs, hides an inefficiency: for most examples the glossary prescribes e…
arxiv.org8 days agoView details
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pre…
arxiv.org8 days agoView details
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
arXiv:2609.10060v1 Announce Type: new Abstract: Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations a…
arxiv.org8 days agoView details
arXiv:2609.10410v1 Announce Type: new Abstract: The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remain…
arxiv.org8 days agoView details
Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding
arXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversi…
arxiv.org8 days agoView details
RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding
arXiv:2609.10305v1 Announce Type: new Abstract: Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transformer at embedding width d = 128 still spends roughly one third of its capacity on the output matrix W_out in R^(d x |V|).…
arxiv.org8 days agoView details
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
arXiv:2609.09264v1 Announce Type: new Abstract: Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochast…
arxiv.org8 days agoView details
Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling
arXiv:2609.09691v1 Announce Type: new Abstract: When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting…
arxiv.org8 days agoView details
Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language Models
arXiv:2609.03247v1 Announce Type: cross Abstract: Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-id…
arxiv.org8 days agoView details
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
arXiv:2609.10451v1 Announce Type: new Abstract: Real-world GUI usage frequently involves workflows that span multiple devices and platforms, requiring the transfer of intermediate results, maintenance of shared state, and coordination across heterogeneous environments. However, existing GUI benchmarks overwhelmingly e…
arxiv.org8 days agoView details
Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services
arXiv:2609.09889v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments such as call center conversations. When privacy restr…
arxiv.org8 days agoView details
Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records
arXiv:2609.09984v1 Announce Type: new Abstract: Understanding the historical allocation and distribution of research funding advances our knowledge of how scientific research is supported across fields, institutions, and regions. However, large-scale analyses are hindered by the lack of comprehensive funder name disam…
arxiv.org8 days agoView details
Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training
arXiv:2609.10052v1 Announce Type: new Abstract: LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study this problem as successful strategy coverag…
arxiv.org8 days agoView details
Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning
arXiv:2609.10113v1 Announce Type: new Abstract: Financial text, textbooks, and question-answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. Existing QA pairs often lack explicit reasoning, sufficient context, or reliably verifiable answers, while textbooks must…
arxiv.org8 days agoView details
Grounded Evaluation and Repair for NL-to-PDDL Problem Generation
arXiv:2609.09898v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning descriptions into PDDL problem instances. However, standard evaluation criteria such as syntactic validity or planner success can substantially overestimate faithfulness to the…
arxiv.org8 days agoView details
ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance
arXiv:2609.09458v1 Announce Type: new Abstract: As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. O…
arxiv.org8 days agoView details
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
arXiv:2609.09448v1 Announce Type: new Abstract: As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, t…
arxiv.org8 days agoView details
KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
arXiv:2609.10266v1 Announce Type: new Abstract: LLM serving systems already reuse KV caches, but only when the reused text sits at the very start of the prompt. Two growing workloads break this condition: a retrieval-augmented generation server assembles a different set of retrieved chunks for every query, and a multi…
arxiv.org8 days agoView details
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
arXiv:2609.09418v1 Announce Type: new Abstract: World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly advancing embodied AI, general-purpose counterparts remain largely unexplored in games. Existing game…
arxiv.org8 days agoView details
X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding
arXiv:2609.09166v1 Announce Type: new Abstract: This paper investigates collaborative speculative decoding (CoSD), a distributed large language model (LLM) inference framework in which an on-device small language model (SLM) drafts candidate tokens and a server LLM verifies them. Existing CoSD methods assume a shared…
arxiv.org8 days agoView details
arXiv:2609.09625v1 Announce Type: new Abstract: As Digital Twin (DT) systems evolve beyond state synchronization toward task-oriented and knowledge-driven operation, Cognitive Digital Twins (CDTs) have emerged as an extension that incorporates cognitive capabilities into twin operation. Existing CDT studies often focu…
arxiv.org8 days agoView details
OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes…
arxiv.org8 days agoView details
Adaptive Entangled Game Modules in Artificial General Intelligence
arXiv:2609.09226v1 Announce Type: new Abstract: We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive agents, deriving testable eigenmodes through a generalized behavioral intelligence (GBI) nonlocal probability-wave equation. This framework captures a broad range of hu…
arxiv.org8 days agoView details
Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks
arXiv:2609.09233v1 Announce Type: new Abstract: How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bundles containing instructions, sc…
arxiv.org8 days agoView details
Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery
arXiv:2609.09413v1 Announce Type: new Abstract: Choosing a recovery process for scale-up requires connecting laboratory results with product requirements, process costs, and scale effects. We analyze records from Pacific Northwest National Laboratory's Computer Intelligence for Critical Element Recovery and Optimizati…
arxiv.org8 days agoView details
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
arXiv:2609.09647v1 Announce Type: new Abstract: Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step…
arxiv.org8 days agoView details
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
arXiv:2609.09657v1 Announce Type: new Abstract: Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support c…
arxiv.org8 days agoView details
Kernel-Managed Shared Memory for System-Wide Personalization
arXiv:2609.10144v1 Announce Type: new Abstract: AI systems become more useful when they can adapt to the people using them, but in multi-agent systems, useful context learned by one agent often remains unavailable to others. We present kernel-managed shared memory, a system-level abstraction in which specialized agent…
arxiv.org8 days agoView details
PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations
arXiv:2609.09664v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as personalized assistants that interact with users over extended periods of time. As conversations grow longer, relying on full interaction histories becomes increasingly inefficient and unreliable: long contexts in…
arxiv.org8 days agoView details
arXiv:2609.09853v1 Announce Type: new Abstract: LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools. The benchmark is built around a complet…
arxiv.org8 days agoView details
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
arXiv:2609.09754v1 Announce Type: new Abstract: As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-…
arxiv.org8 days agoView details
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
arXiv:2609.09776v1 Announce Type: new Abstract: Frontier gains in language-model reasoning come from reinforcement learning on reasoning traces and are concentrated in domains with a cheap, sound verifier. We argue the field's binding constraint is the verification gap: no scalable, incorruptible reward for reasoning…
arxiv.org8 days agoView details
Models
View all models →text-generation · mlx · safetensors · qwen3_5_moe
huggingface.co8 days ago3209 ptsView details
text-generation · transformers · safetensors · llama
huggingface.co9 days ago1527 ptsView details
text-to-audio · safetensors · yue2 · music-generation
huggingface.co8 days ago680 ptsView details
text-to-speech · audio · speech · text-to-speech
huggingface.co8 days ago282 ptsView details
text-generation · transformers · gguf · minicpm
huggingface.co9 days ago246 ptsView details
text-generation · Model Optimizer · safetensors · qwen3_5
huggingface.co9 days ago74 ptsView details
text-generation · pytorch · safetensors · falcon
huggingface.co9 days ago2 ptsView details
Prannesshkva/ISOM-Qwen-1.5B-Instruct
text-generation · safetensors · isom · qwen2
huggingface.co9 days ago1 ptsView details
Devlin-AI/Devlin-Alpha-22B-A3B-GGUF
gguf · endpoints_compatible · region:us
huggingface.co9 days agoView details
text-generation · flax · simi · weights-only
huggingface.co9 days agoView details
Open source
View all open source →langchain-ai/langchain langchain-openai==1.6.2
Changes since langchain-openai==1.6.1 release(openai): 1.6.2 (#40339) fix(openai): add GPT-6 Astra reasoning efforts (#40330) chore(deps): bump httpx2 from 2.10.0 to 2.12.0 in /libs/partners/openai (#40309)
github.com8 days agoView details
<details open> vulkan: use spec constant for matrix matrix multiplication A-type (#25773) * vulkan: use spec constant for mul mat type_a vulkan: use map for mul_mm shapes cleanup fix indentation fix cm2 and shmem init fix cm2 spec constants fix cm2 bindings consolidate shmem tab…
github.com8 days agoView details
huggingface/transformers v5.17.0
# Release v5.17.0 ## New Model additions ### HYV4 <img width="1503" height="827" alt="image" src="https://github.com/user-attachments/assets/e6ed85ee-eb1d-40eb-a0d4-c649f6337ca9" /> Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters p…
github.com8 days agoView details
## [3.11.0](https://github.com/openai/openai-python/compare/v3.10.0...v3.11.0) (2026-09-09) ### Features * **api:** Add expiration controls for service account keys ([#3825](https://github.com/openai/openai-python/issues/3825)) ([f348ec8](https://github.com/openai/openai-python/…
github.com8 days agoView details
# v0.29.0 ## Highlights This release features 594 commits from 277 contributors (91 new)! * **Model Runner V2 is now the default for all models** (#53183), completing the rollout that began with pooling models (#48290). MRV2 also gained CUDA graph memory profiling for KV cache a…
github.com9 days agoView details
<details open> jinja: treat a null left operand of in as a plain lookup (#28620) Templates that default an optional variable to none and then test its membership in a map hit an error, while the same expression is a normal lookup returning false in Jinja. The undefined counterpa…
github.com9 days agoView details
<details open> vulkan: add dedicated iq4_xs mat-vec shader (#28426) * vulkan: add dedicated iq4_xs mat-vec shader Dedicated mul_mat_vec_iq4_xs for the dmmv path, replacing the generic fallback. ~+6-17% token generation on RDNA4 depending on model. Assisted-by: Pi agent with Qwen…
github.com9 days agoView details
<details open> vulkan: add f16 B-type matmul pipelines and warp tile size tuning for Intel coopmat1 (#27471) * vulkan: add f16 B-type matmul pipelines and warp tile size tuning for Intel coopmat1 * simplify mmp selection in mul_mat_id per review comment * vulkan: enable f16 B-ty…
github.com9 days agoView details