Archive / 2026-09-14
September 14, 2026
News
View all news →XCancel service is suspended until further notice
xcancel.com3 days ago752 ptsView detailsJoin discussion
Pion, an agent designed to run any company autonomously
andonlabs.com3 days ago490 ptsView detailsJoin discussion
How to write an effective software design document
refactoringenglish.com3 days ago346 ptsView detailsJoin discussion
Ex-FTC boss Khan: break out the handcuffs for AI CEOs, citing 1934 precedent
theregister.com3 days ago223 ptsView detailsJoin discussion
Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows
macrumors.com3 days ago222 ptsView detailsJoin discussion
GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
entelligence.ai3 days ago162 ptsView detailsJoin discussion
Why don't machine learning research agents overfit?
amazon.science3 days ago135 ptsView detailsJoin discussion
Cops Search Flock Cameras for Reasons of 'LMAO,' 'IDK,' and 'Asdfg'
404media.co3 days ago135 ptsView detailsJoin discussion
Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
patrickmccanna.net3 days ago139 ptsView detailsJoin discussion
For AI leaders Doom is a form of hype
erkansaka.net3 days ago131 ptsView detailsJoin discussion
Backprop Alternative: Augmented Lagrangian Predictive Coding
pub.sakana.ai3 days ago126 ptsView detailsJoin discussion
Show HN: Pelican-bicycle alternatives
gally.net3 days ago131 ptsView detailsJoin discussion
OpenArch – PyTorch implementations of modern LLM architectures
github.com4 days ago139 ptsView detailsJoin discussion
Big AI sets out its terms for regulatory capture
theregister.com3 days ago119 ptsView detailsJoin discussion
Show HN: Kinesis – Control your Mac with the Meta Neural Band
github.com3 days ago115 ptsView detailsJoin discussion
Anthropic is in regulatory-capture financial loop
twitter.com3 days ago85 ptsView detailsJoin discussion
Adversarial Fashion Makes a Statement on AI Panopticon
spectrum.ieee.org3 days ago102 ptsView detailsJoin discussion
Dropping eBPF CPU Cost by About 90% with Memoization (Not AI Gen)
nathannaveen.dev3 days ago100 ptsView detailsJoin discussion
Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
narilabs.com3 days ago90 ptsView detailsJoin discussion
Archiving pirate radio station Kool FM
londonist.com3 days ago92 ptsView detailsJoin discussion
Proposed Rule: Eliminating the Discretionary 60-Day Grace Period
regulations.gov3 days ago71 ptsView detailsJoin discussion
Temporal raises $550M at a $12.55B valuation
temporal.io3 days ago79 ptsView detailsJoin discussion
ilinmaks.com3 days ago65 ptsView detailsJoin discussion
China's Regulators Take Aim at "AI Boyfriends"
spectrum.ieee.org3 days ago58 ptsView detailsJoin discussion
Show HN: Sunk Cost – How long until a local LLM rig pays for itself?
sunkcost.ai3 days ago46 ptsView detailsJoin discussion
When LLM judges agree, should we believe them?
amazon.science3 days ago54 ptsView detailsJoin discussion
I worked at Google DeepMind. You should listen to the warnings about AI
theguardian.com3 days ago40 ptsView detailsJoin discussion
Show HN: AttaLambda: a language where types and data are made of untyped lambdas
attalambda.com3 days ago43 ptsView detailsJoin discussion
HP ZGX Fury Is Now Orderable: GB300 Superchip, 748GB Unified Memory
storagereview.com3 days ago40 ptsView detailsJoin discussion
Watch AI materials-science and bioscience abilities closely
lesswrong.com3 days ago39 ptsView detailsJoin discussion
Hacking AI customer service agents
intigriti.com3 days ago35 ptsView detailsJoin discussion
Airbnb Preventing the Use of BnB
theguardian.com4 days ago39 ptsView detailsJoin discussion
MIT creates method to force AI to comply with safety rules
theframenews.org3 days ago27 ptsView detailsJoin discussion
EU to limit access to social media before age of 15
politico.eu3 days ago26 ptsView detailsJoin discussion
Graphic Rants: Nanite Tessellation
graphicrants.blogspot.com3 days ago26 ptsView detailsJoin discussion
A1ex: A simple LLM coding agent in Lua
github.com3 days ago24 ptsView detailsJoin discussion
Air Force secretary acknowledges the US has weapons in space
abcnews.com3 days ago19 ptsView detailsJoin discussion
Transitions.dev: UI transitions for AI agents
transitions.dev3 days ago23 ptsView detailsJoin discussion
Radio – a chatroom for multiplayer agent work
radio.plasma.ai3 days ago21 ptsView detailsJoin discussion
AirBaltic files for Chapter 11 bankruptcy as Iran war costs bite
reuters.com3 days ago18 ptsView detailsJoin discussion
Show HN: I built Otis, a minimal AI agent that runs local models out of the box
triangllabs.ai3 days ago19 ptsView detailsJoin discussion
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
arxiv.org4 days ago20 ptsView detailsJoin discussion
Apple Releases iOS 27 and iPadOS 27 with Siri AI and Liquid Glass Update
macrumors.com3 days ago17 ptsView detailsJoin discussion
wheresyoured.at3 days ago17 ptsView detailsJoin discussion
Israel accused of eliminating living conditions for Palestinians in West Bank
theguardian.com3 days ago18 ptsView detailsJoin discussion
D-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026
servethehome.com3 days ago17 ptsView detailsJoin discussion
Security through obscurity is dead, and AI delivered the fatal blow
theregister.com3 days ago14 ptsView detailsJoin discussion
Why the software industry needs a lot of regulation
petewarden.com3 days ago15 ptsView detailsJoin discussion
When Did Our Definition of AGI Become So Piss-Weak?
pointlessramblings.com3 days ago14 ptsView detailsJoin discussion
The lie that raced around the world before the truth got its boots on
pluralistic.net3 days ago11 ptsView detailsJoin discussion
Nvidia CEO Jensen Huang tells Trump we're not going to let an AI slowdown happen
techcrunch.com3 days ago11 ptsView detailsJoin discussion
XCancel service is suspended (again)
xcancel.com3 days ago14 ptsView detailsJoin discussion
AI Doomer Hypeloop Suspiciously Serves to Support OpenAI and Anthropic's Agenda
nakedcapitalism.com3 days ago11 ptsView detailsJoin discussion
EPA Expected to Erase Limits on Climate Pollution from Power Plants
nytimes.com3 days ago13 ptsView detailsJoin discussion
U.S. Health Officials Move Quickly to Deploy Medical A.I. Despite Concerns
nytimes.com3 days ago11 ptsView detailsJoin discussion
Andon Labs Puts AI Agents in Charge of Real Businesses
spectrum.ieee.org3 days ago12 ptsView detailsJoin discussion
China says AI CEOs' call for a slowdown is 'fear mongering'
cnbc.com3 days ago10 ptsView detailsJoin discussion
The Register of British Slave-Traders
britishslavetraders.org3 days ago10 ptsView detailsJoin discussion
SpaceX sues to block release of tax-break records for its Texas Terafab project
businessinsider.com4 days ago13 ptsView detailsJoin discussion
A list of 1,325 AI assisted repositories, mined from GitHub
github.com3 days ago11 ptsView detailsJoin discussion
We gave a coding agent the lethal trifecta: data, internet, a public repo
archestra.ai3 days ago11 ptsView detailsJoin discussion
America's Everywhere Millionaires
newyorker.com3 days ago10 ptsView detailsJoin discussion
Show HN: Biloba: fast and stable Chrome-based browser tests in Go and Vitest
github.com3 days ago10 ptsView detailsJoin discussion
ProGantt: Gantt charts your AI agent can read and write via MCP
progantt.com3 days ago10 ptsView detailsJoin discussion
Mary Beard: 'There is a deeply far-right appropriation of the classics'
english.elpais.com3 days ago11 ptsView detailsJoin discussion
Global AI stocks fall as industry chiefs call for slowing development
reuters.com3 days ago10 ptsView detailsJoin discussion
- Primary source
How Fyxer built an AI executive assistant people trust
Fyxer uses OpenAI models, fine-tuning, memory, and real user feedback to organize inboxes and draft emails in each user’s voice.
openai.com3 days agoView details
Meta engineering team introduced ZGateway, a proxy tier that now sits between client applications and ZippyDB, the Meta’s most widely used key value store. ZippyDB backs product metadata, counters, and configuration at billions of operations per second. ZGateway started as a fix for connection sprawl across more than…
marktechpost.com3 days agoView details
Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent
Agent-net, the team building an agent-to-agent marketplace where AI agents discover, trust, and pay each other, has released Webagent, an open source harness for standing up public-facing business agents. So, basically you give it your website, get an agent, and let it talk to other agents. Instead of writing orchestr…
marktechpost.com3 days agoView details
A practitioner's map of the 3 layers in a modern agent stack, with verified sources and an overlap analysis. The post Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery appeared first on MarkTechPost.
marktechpost.com3 days agoView details
Reward AI has released OM-1 (Omnibody Model 1), a general-purpose manipulation policy trained entirely on human demonstrations captured with a 7-DoF wearable glove, with no teleoperation or on-robot data. The policy runs on industrial arms and humanoids at human speed, learns a new task from under 30 minutes of data,…
marktechpost.com3 days agoView details
Sakana AI researchers Jeffrey Seely and Julian Gould introduce Augmented Lagrangian Predictive Coding (PC-ALM), a local-learning alternative to backpropagation. By attaching a Lagrange multiplier to each layer constraint, PC-ALM keeps predictive coding's layer-local updates while recovering exact backprop gradients in…
marktechpost.com3 days agoView details
NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing
NVIDIA has open-sourced OSMO, the Kubernetes-native workflow orchestrator it uses internally for Project GR00T, Isaac Lab, and Isaac Sim. OSMO lets robotics teams define training, simulation, and hardware-in-the-loop tasks in a single YAML file and routes each one to the right compute tier, from GB200 clusters to Jets…
marktechpost.com4 days agoView details
Is Big Tech’s AI slowdown a safety pact or a cartel?
When OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, and SpaceX head Elon Musk loosely agreed over the weekend to slow down AI development, skeptics spotted an ulterior motive immediately. The AI titans had declared that their aim was to "pace the frontier," signing on at l…
theverge.com3 days agoView details
What execs and politicians are saying about slowing down AI development
Dario Amodei kicked off a flood of statements over the past few days about AI safety by publishing a long essay titled "We Must Pace the Frontier" detailing why AI development should be slowed down. Other AI leaders and politicians are speaking out in favor of or opposing his points, and we've compiled some of them he…
theverge.com3 days agoView details
Jensen Huang puts Trump on speakerphone onstage to announce robots won’t take over the world
NVIDIA CEO Jensen Huang speaks during the G20 Innovation Ministerial in Chapel Hill, North Carolina, on September 2, 2026. (Photo by Matt RAMEY / AFP via Getty Images) | AFP via Getty Images Nvidia CEO Jensen Huang took a call from President Trump on Monday while onstage at the All-In Podcast's All-In Summit. It's not…
theverge.com3 days agoView details
New York Seizes a Dozen Celebrity Deepfake Websites
In the biggest-ever legal action against harmful deepfake websites, the Manhattan District Attorney’s Office has seized 12 sites that collectively targeted around 1,200 victims.
wired.com3 days agoView details
Microsoft says ‘people matter more than AI’ following safety concerns
Microsoft is publishing a 37-page "humanist AI code of conduct" today, amid growing safety concerns over AI model progress. Anthropic CEO Dario Amodei called for a coordinated slow down of AI development over the weekend, after researchers warned recently that AI model progress could outpace our ability to safely depl…
theverge.com3 days agoView details
AI Leaders Are Calling for a Slowdown. Trump’s Team Says It’s on Them
Sam Altman and Elon Musk backed Anthropic CEO Dario Amodei’s weekend plea for regulation. The White House seems unlikely to oblige.
wired.com3 days agoView details
Sexually Explicit Deepfake Sites Target 100-Plus Politicians in Europe
An analysis of 160 deepfake websites reveals politicians in 22 countries appear on them. Nearly all of them are women.
wired.com3 days agoView details
‘I Like My Big Rat Wife’: Meet the People Using Chatbots to Write Custom Fiction
While the publishing industry frets over how authors are using AI, many readers are taking things into their own hands.
wired.com4 days agoView details
arXiv:2609.13520v1 Announce Type: new Abstract: While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in pre-training. In this paper, we propose ModelLog…
arxiv.org3 days agoView details
arXiv:2609.13556v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown remarkable proficiency on general-purpose tasks, yet their performance often degrades in highly-specialized technical domains. Moreover, little is known about how parametric knowledge of domain-specific terms is encoded within thes…
arxiv.org3 days agoView details
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
arXiv:2609.13582v1 Announce Type: new Abstract: A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically score one run per task and rarely ask whe…
arxiv.org3 days agoView details
$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
arXiv:2609.13602v1 Announce Type: new Abstract: Voice agents often need to collect names, addresses, identifiers, dates, and times exactly, yet end-to-end benchmarks obscure where capture fails. We introduce $\tau$-Elicitation, a 200-task voice benchmark spanning 10 entity types, controlled difficulty, caller realisms…
arxiv.org3 days agoView details
Enhancing Event Candidate Acquisition for Event Linking
arXiv:2609.13670v1 Announce Type: new Abstract: Event linking associates event mentions in text with entries in a knowledge base (KB), or identifies them as out-of-KB events. Although existing methods use different architectures, candidate event acquisition can still be weakened by short ambiguous mentions, noisy argu…
arxiv.org3 days agoView details
Recoverability as a System Primitive for Long-Horizon AI Agents
arXiv:2609.13672v1 Announce Type: new Abstract: AI agents can be interrupted while editing files, calling tools, or carrying out multi-step tasks. Restarting repeats completed work, but continuing from unverified or outdated progress can carry earlier errors forward. A saved state is not necessarily a suitable place t…
arxiv.org3 days agoView details
LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference
arXiv:2609.13682v1 Announce Type: new Abstract: We introduce LayerRoute, a parameter-efficient method for adaptive transformer layer-skipping that combines per-layer hard-gated routing (trained via a straight-through estimator) with joint LoRA fine-tuning. LayerRoute augments each of the 24 transformer blocks in Qwen2…
arxiv.org3 days agoView details
Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding
arXiv:2609.13685v1 Announce Type: new Abstract: Negation remains a longstanding challenge for both language models (LMs) and large language models (LLMs). Prior work mainly focuses on a small set of high-frequency single-word negation cues, such as not and never, with limited exploration of broader negation types and…
arxiv.org3 days agoView details
Positioning manuscripts in the scientific landscape with agentic AI
arXiv:2609.13760v1 Announce Type: new Abstract: Publishing a research manuscript is a routine yet demanding part of scientific life: time-consuming, stressful, and often uncertain in outcome. Recent advances in large language model (LLM)-based agentic AI have shown promise across a range of scientific tasks, and here…
arxiv.org3 days agoView details
Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs
arXiv:2609.13776v1 Announce Type: new Abstract: Integrating heterogeneous relational databases into a centralized ontology remains a persistent challenge in enterprise knowledge representation, primarily due to semantic heterogeneity, cryptic schema naming, missing metadata, and the abstraction gap between relational…
arxiv.org3 days agoView details
Scaling Hindi Quantum Natural Language Processing through Automatic Pregroup Supertagging
arXiv:2609.13721v1 Announce Type: new Abstract: Quantum Natural Language Processing (QNLP) uses pregroup grammars to translate grammatical structure into diagrammatic representations and quantum circuits. Recent Hindi QNLP work has shown that Hindi-specific pregroup grammars can support grammar-sensitive compositional…
arxiv.org3 days agoView details
Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth
arXiv:2609.13745v1 Announce Type: new Abstract: Vision--language models (VLMs) can answer chart questions accurately, but output accuracy does not show how they combine the evidence needed to recover an exact value. We study vertical-bar value reading with controlled counterfactual activation patching in Qwen2.5VL-7B-…
arxiv.org3 days agoView details
HyperProve: Answer-Guided Hypergraph Expansion for Multi-Hop Question Answering
arXiv:2609.13768v1 Announce Type: new Abstract: Multi-hop question answering often fails when retrieval treats evidence as isolated matches to the original question, since the facts needed to answer a complex question are usually connected through intermediate entities, relations, and constraints. We propose HyperProv…
arxiv.org3 days agoView details
arXiv:2609.13794v1 Announce Type: new Abstract: Detecting harmful memes is critical for maintaining safe online communities. However, harmful intent is often implicit, arising from visual-textual incongruity and cultural stereotypes, which challenges existing multimodal detectors. We propose SyRHM, a framework that de…
arxiv.org3 days agoView details
Understanding the Limits of Agentic ICD Coding
arXiv:2609.13806v1 Announce Type: new Abstract: ICD-10-CM codes are alphanumeric codes used in the US to classify diagnoses and injuries for medical billing and epidemiological reporting. Standard ICD-10-CM benchmarks report aggregate metrics that obscure performance on complex coding scenarios. We evaluate neural, wo…
arxiv.org3 days agoView details
Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice
arXiv:2609.13841v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support and relationship advice, where a model's tendency to preserve a user's face can inadvertently reinforce harmful interpersonal behaviors. To systematically examine this risk, we developed the Romanti…
arxiv.org3 days agoView details
arXiv:2609.13847v1 Announce Type: new Abstract: In this paper, we extend FLORES+ with Portuguese-source evaluation sets for three Mozambican Bantu varieties: Xichangana, Mozambican Nyanja, and Sena. We compare Xichangana with the existing Tsonga reference and Mozambican Nyanja with Chichewa, and evaluate NLLB-200, Goo…
arxiv.org3 days agoView details
arXiv:2609.13856v1 Announce Type: new Abstract: Enterprise customer support systems must answer customer questions correctly, retrieve the right policy information, use customer context, and pass difficult cases to human agents when needed. This paper presents ShopEase, a Generative AI-based multi-agent framework for…
arxiv.org3 days agoView details
SHIFT-M3: Pre-fusion Alignment-based Consistency Screening for Multimodal ECG Record Integrity
arXiv:2609.13874v1 Announce Type: new Abstract: Multimodal clinical AI typically assumes that the waveform, report, metadata, and downstream predictions attached to a record belong to the same patient. In practice, linkage failures can silently assemble individually plausible but cross-patient components, creating a s…
arxiv.org3 days agoView details
arXiv:2609.13907v1 Announce Type: new Abstract: Pharmaceutical sales forecasts inform planning across products, regions, and distribution channels, yet their interpretation depends on inventory availability, transaction semantics, product lifecycle, and the information available when each forecast is issued. A model p…
arxiv.org3 days agoView details
In the Blind: Building Pseudo-References for MT Evaluation
arXiv:2609.13611v1 Announce Type: new Abstract: The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these pairs and six other language pairs (in whic…
arxiv.org3 days agoView details
North Small Translate: Advanced Cost-Effective Translation (Cohere CAT+)
arXiv:2609.13916v1 Announce Type: new Abstract: We present North Small Translate, an open-weight, LLM-based machine translation (MT) model with instruction-following capabilities built on the same foundation as Cohere's Command A Plus, a mixture-of-experts architecture with 25 billion active parameters out of 218 bill…
arxiv.org3 days agoView details
arXiv:2609.13534v1 Announce Type: new Abstract: We identify \textbf{Harmfulness Propagation Dynamics (HPD)}: for harmful prompts, the projection of the last-token hidden state onto a learned harm direction rises monotonically with transformer depth, whereas benign prompts remain flat or oscillatory. This cross-layer s…
arxiv.org3 days agoView details
arXiv:2609.13936v1 Announce Type: new Abstract: Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This report answers that for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied_…
arxiv.org3 days agoView details
arXiv:2609.13942v1 Announce Type: new Abstract: The CRITICS project addresses science accessibility and literacy by converging advanced Machine Translation (MT) based on Large Language Models (LLMs) with educational technology. By leveraging MT systems specifically optimized for scientific content, educational institu…
arxiv.org3 days agoView details
Thought without systematicity? Evaluating reasoning models on rule induction tasks
arXiv:2609.13948v1 Announce Type: new Abstract: A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance…
arxiv.org3 days agoView details
Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR
arXiv:2609.13997v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has shown remarkable success in improving the mathematical reasoning of large language models. Yet problems beyond the model's current capability, where rollouts uniformly fail and no learning signal is produced, are…
arxiv.org3 days agoView details
Measuring the Creativity of Frontier LLMs in Automated Research
arXiv:2609.14057v1 Announce Type: new Abstract: Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. In this paper, we propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Value…
arxiv.org3 days agoView details
GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning
arXiv:2609.14066v1 Announce Type: new Abstract: Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on complex multimodal reasoning tasks, they remain fundamentally limited in reasoning depth and memory structure, suffering from inadequate retrieval and state blindness…
arxiv.org3 days agoView details
One Size Does Not Fit All: Setting Inference Depth from the Questions a Deployment Actually Asks
arXiv:2609.14144v1 Announce Type: new Abstract: A transformer language model is trained to respond to any prompt, but each deployment asks only a narrow range of questions: a support assistant sees delivery complaints, a coding tool sees Python. Every deployment nonetheless pays the same computation per token. This pa…
arxiv.org3 days agoView details
Map Users and Mapmakers: The Scope of Cognitive Attribution from Acquired Representations
arXiv:2609.13879v1 Announce Type: new Abstract: An acquired representation can enlarge a system's cognitive repertoire without transferring the capacities exercised in producing that representation. This paper develops a framework for specifying that enlargement and its limits. Its central contribution is a five-part…
arxiv.org3 days agoView details
When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering
arXiv:2609.14157v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed with external tools that extend what they can do beyond their own knowledge. Tools help on tasks that need external information, but their availability may also change how a model handles questions that do not need t…
arxiv.org3 days agoView details
Towards Evolving Context Parameterization for Large Language Models
arXiv:2609.14168v1 Announce Type: new Abstract: Context parameterization enables large language models (LLMs) to internalize contexts into reusable model parameters, avoiding repeated processing across subsequent queries. However, existing methods typically assume static contexts and lack explicit mechanisms for disti…
arxiv.org3 days agoView details
A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement
arXiv:2609.14178v1 Announce Type: new Abstract: The rapid diffusion of hate speech and misinformation on social networks challenges democratic societies, since direct suppression efforts may deepen polarization, fuel public distrusts, and strengthen extremist narratives. LLM-driven counter-narratives (CNs) offer a pro…
arxiv.org3 days agoView details
Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had
arXiv:2609.14121v1 Announce Type: new Abstract: The Semantic Web set out to give information a machine-interpretable form so that software could integrate and reason over it. Its standards became scientific knowledge infrastructure, but the machine competence it promised did not follow, and the systems now answering q…
arxiv.org3 days agoView details
arXiv:2609.13918v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in specialized university courses, but control-systems questions require coordinated terminology, notation, derivations, and stepwise explanations. Direct general-purpose responses may be inconsistently structured and ha…
arxiv.org3 days agoView details
ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading
arXiv:2609.13825v1 Announce Type: new Abstract: Reinforcement learning trading systems published in the academic literature overwhelmingly rely on price-aggregate state representations (OHLCV bars) or limit-order-book depth features, leaving microstructure pattern theories from the practitioner literature, namely Auct…
arxiv.org3 days agoView details
The Attribution-Compression Frontier in Retrieval-Augmented Generation
arXiv:2609.14245v1 Announce Type: new Abstract: Context compression reduces generator input in retrieval-augmented generation, but answer quality alone does not characterize citation attribution. We measure citation attribution across compression methods and budgets, comparing reranking, extractive selection, abstract…
arxiv.org3 days agoView details
arXiv:2609.14256v1 Announce Type: new Abstract: Topic models are widely used to analyze public health-related social media short texts, yet their evaluation remains dominated by metrics that focus entirely on generated topics alone. There is a lack of metrics that quantitatively assess whether assigned topics meaningf…
arxiv.org3 days agoView details
DenMark: Robust Semantic Watermarking for Diffusion Language Models
arXiv:2609.14257v1 Announce Type: new Abstract: Semantic text watermarks encode signals in meaning rather than surface token choices, offering robustness to paraphrasing and other semantic-preserving edits. Existing semantic watermarking methods are primarily designed for autoregressive language models (ARLMs), where…
arxiv.org3 days agoView details
arXiv:2609.13807v1 Announce Type: new Abstract: Large language models reason in high-dimensional hidden-state spaces, while users observe only final outputs. We introduce Bypass Observation, a non-intrusive layer-wise readout architecture that attaches read-only observation heads to selected Transformer layers without…
arxiv.org3 days agoView details
Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error
arXiv:2609.14400v1 Announce Type: new Abstract: Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple defensible readings that a single gold trajectory…
arxiv.org3 days agoView details
NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass
arXiv:2609.14448v1 Announce Type: new Abstract: Hallucination in large language models reduces their reliability and slows adoption. Various white-box studies have used internal representations to detect patterns of truthfulness and factuality. A less-studied approach is to identify feed-forward neurons correlated wit…
arxiv.org3 days agoView details
Theseus in the Graph: Towards Traceable Multi-Hop Graph Navigation
arXiv:2609.14528v1 Announce Type: new Abstract: Multi-Hop Knowledge Graph Question Answering (KGQA) tasks require models to assemble relational evidence along paths in a KG to answer natural-language questions. However, existing KGQA systems typically focus on predicting the final answer without explicitly modeling or…
arxiv.org3 days agoView details
Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition
arXiv:2609.14542v1 Announce Type: new Abstract: Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utterances. In version 6, 62,279 validated recordings totalling 99.02 hours are provided from 190 contributors, with 34,541…
arxiv.org3 days agoView details
Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?
arXiv:2609.13436v1 Announce Type: new Abstract: Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as…
arxiv.org3 days agoView details
arXiv:2609.13475v1 Announce Type: new Abstract: Existing pipelines for clinical timeline extraction from case reports are evaluated using an expert reference and are limited by imperfect reference annotations and imprecise event alignment. We developed GAVEL, an LLM judge protocol that compares two timelines with the…
arxiv.org3 days agoView details
Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management
arXiv:2609.13552v1 Announce Type: new Abstract: Generative AI is increasingly being used informally in Air Traffic Management (ATM) for tasks such as flight plan generation, trajectory interpretation, and constraint checking. Although these tools can reduce workload and accelerate planning, their non-deterministic out…
arxiv.org3 days agoView details
Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems
arXiv:2609.13559v1 Announce Type: new Abstract: Large Language Models (LLMs) with function-calling capabilities are becoming critical for modern agentic AI systems. Nevertheless, current deployments typically route inferences to powerful cloud-based models, incurring significant energy use and carbon emissions. We add…
arxiv.org3 days agoView details
Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
arXiv:2609.13637v1 Announce Type: new Abstract: Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract. It separates recall, composition, be…
arxiv.org3 days agoView details
arXiv:2609.13648v1 Announce Type: new Abstract: Solar energy decision support is fragmented across dashboards that provide data without explanation, research papers are slow to parse, and general-purpose language models are not solar domain specific and answer without evidence. This paper introduces Solar Intelligence…
arxiv.org3 days agoView details
Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models
arXiv:2609.13680v1 Announce Type: new Abstract: Fine-tuning instruct models often improves target performance while inducing behavioral drift from the reference model, which can degrade existing capabilities. Rather than treating this drift as an uncontrolled consequence of optimization, we specify a behavioral drift…
arxiv.org3 days agoView details
JaxAHT: A JAX-Based Library for Ad Hoc Teamwork
arXiv:2609.13716v1 Announce Type: new Abstract: Ad Hoc Teamwork (AHT) addresses the challenge of designing agents capable of coordinating with novel partners without prior coordination. However, progress in the field is hindered by the prohibitive computational cost of the AHT research lifecycle, the lack of standardi…
arxiv.org3 days agoView details
IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives
arXiv:2609.13725v1 Announce Type: new Abstract: An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived…
arxiv.org3 days agoView details
arXiv:2609.13731v1 Announce Type: new Abstract: The transition from passive foundation models to autonomous, goal-directed agentic AI systems has introduced unprecedented capabilities by coupling recursive cognitive reasoning loops, persistent memory architectures, live tool execution planes, and multi-agent collabora…
arxiv.org3 days agoView details
Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection
arXiv:2609.13785v1 Announce Type: new Abstract: Oracle-style quantities, including virtual best solvers, selected-portfolio VBS, virtual-best encodings, and best-in-family summaries, are widely reported as upper bounds on what a deployable selector could achieve. In decomposed algorithm selection, an analogous partiti…
arxiv.org3 days agoView details
Do Not Restart: Residual Completion for Stateful Agent Handoffs
arXiv:2609.13800v1 Announce Type: new Abstract: Routing and cascades reduce tool-agent cost by transferring control across models, but stateful handoffs must preserve accepted choices, realized effects, and unfinished obligations. We formulate this as commitment-constrained residual completion and introduce Commitment…
arxiv.org3 days agoView details
UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics
arXiv:2609.13849v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) often struggle with complex mathematical visual reasoning primarily due to a lack of fine-grained perception, causing initial visual hallucinations to directly trigger cascading reasoning failures. In traditional end-to-end reinfo…
arxiv.org3 days agoView details
ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information
arXiv:2609.13860v1 Announce Type: new Abstract: Querying clinical trial registries remains a manual and error-prone process, requiring researchers to navigate large volumes of semi-structured data without support for natural language interaction or cross-source synthesis. To address this, we introduce ClinAgent, a con…
arxiv.org3 days agoView details
SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery
arXiv:2609.13945v1 Announce Type: new Abstract: Natural-language descriptions of optimization problems may be incomplete or vague about numerical information that a solver requires, including costs, capacities, demands, bounds, and penalties. A language model can translate the description into code, but when a require…
arxiv.org3 days agoView details
Synthetic Data in Marketing Research: How to Evaluate and When to Trust
arXiv:2609.13995v1 Announce Type: new Abstract: Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic responde…
arxiv.org3 days agoView details
MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents
arXiv:2609.14399v1 Announce Type: new Abstract: Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single} text template---missing the synergy among multiple com…
arxiv.org3 days agoView details
Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
arXiv:2609.13543v1 Announce Type: new Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-depar…
arxiv.org3 days agoView details
arXiv:2609.13466v1 Announce Type: new Abstract: Enterprise AI adoption has reached 78% of organizations globally, yet the infrastructure to govern that adoption has not kept pace. This paper identifies and characterizes the attestation deficit, a structural condition in which organizations maintain governance policies…
arxiv.org3 days agoView details
arXiv:2609.13151v1 Announce Type: new Abstract: Leading multilingual speech recognition models like Whisper transcribe diverse, low-resource languages without language-specific training but are computationally expensive to deploy. Token merging mitigates this inefficiency by dynamically combining redundant features, s…
arxiv.org3 days agoView details
Corpus Characterization and Inverse Constitutional Fine-Tuning for Style-Aware Radiology Reports
arXiv:2609.14226v1 Announce Type: new Abstract: Automated radiology report generation has advanced rapidly in diagnostic accuracy, yet generated reports frequently diverge from the stylistic conventions of authentic radiologist writing in structure, diction, and uncertainty language, a gap which has direct implication…
arxiv.org3 days agoView details
Schizophrenia Detection from EEG Signals: A Transformer Framework with Spectrogram Representation
arXiv:2609.14015v1 Announce Type: new Abstract: Schizophrenia is a serious psychiatric disorder that affects millions of people worldwide, and its diagnosis remains primarily dependent on clinical assessment. Electroencephalography (EEG) provides a non-invasive approach to investigate brain activity and has shown pote…
arxiv.org3 days agoView details
arXiv:2609.13676v1 Announce Type: new Abstract: Markov decision processes (MDPs) are used to support decision-making in conservation of biodiversity, but policies, even over small state spaces, can be difficult to interpret for conservation managers. K-MDP methods address this problem by building simpler MDPs with at…
arxiv.org3 days agoView details
arXiv:2609.13667v1 Announce Type: new Abstract: Geospatial agents are increasingly expected to support recurring and evolving analytical tasks rather than execute isolated workflows. In such settings, effective agents must distill prior execution experience into reusable geospatial procedural knowledge to guide future…
arxiv.org3 days agoView details
arXiv:2609.13454v1 Announce Type: new Abstract: Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncer…
arxiv.org3 days agoView details
FedV-KGQA in Practice: Design Lessons and an Interactive Prototype
arXiv:2609.13661v1 Announce Type: new Abstract: Knowledge graph question answering usually assumes that one system can reach the whole graph. In practice, facts are often held by organizations that share entity identifiers but own disjoint relation types, so no single party sees a complete reasoning chain. This poster…
arxiv.org3 days agoView details
arXiv:2609.13535v1 Announce Type: new Abstract: The EU AI Act introduces mandatory requirements for high-risk AI systems with the explicit goal of ensuring the development and operation of trustworthy AI. At the same time, AI risk management practices rely on structured risk taxonomies to systematically identify and t…
arxiv.org3 days agoView details
PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems
arXiv:2609.13152v1 Announce Type: new Abstract: Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experimentation remains poorly understood. We introduce PhysMent, a benchmark that evaluates LLM physical reasoning via iterati…
arxiv.org3 days agoView details
Learning to Refer from Estimated Listener Gaze
arXiv:2609.14207v1 Announce Type: new Abstract: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals. During training, referring expressions are…
arxiv.org3 days agoView details
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
arXiv:2609.14302v1 Announce Type: new Abstract: Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not whether evidence remains traceable through r…
arxiv.org3 days agoView details
Planning or Learning: Reliability and Cost in Multi-Asset Maintenance
arXiv:2609.13566v1 Announce Type: new Abstract: Industrial maintenance systems involve multiple interacting assets and shared resources, making it challenging to balance reliability and operational cost using a single decision framework. While recent work has focused on reinforcement learning (RL) for maintenance sche…
arxiv.org3 days agoView details
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
arXiv:2609.13463v1 Announce Type: new Abstract: The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders huma…
arxiv.org3 days agoView details
arXiv:2609.13980v1 Announce Type: new Abstract: Arabic large-language-model (LLM) evaluation has matured around Modern Standard Arabic (MSA): aggregated leaderboards such as the Open Arabic LLM Leaderboard (OALL), HELM Arabic, and BALSAM rank models across dozens of MSA tasks, and frontier systems increasingly saturat…
arxiv.org3 days agoView details
LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents
arXiv:2609.13437v1 Announce Type: new Abstract: Scientific research is a continuous process that emphasizes inheritance. Methods developed by predecessors are often expanded upon by new researchers to explore more novel and in-depth scientific questions. However, the change of lab staff, such as student graduation, le…
arxiv.org3 days agoView details
Homeostatic Continual Learning
arXiv:2609.13771v1 Announce Type: new Abstract: In this paper, I formulate a Continual Learning problem and propose a method named "Homeostatic Continual Learning" that enables an AI agent to learn continuously in a changing environment without catastrophic forgetting. The core of the method is to find outliers in the…
arxiv.org3 days agoView details
When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
arXiv:2609.13824v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with hu…
arxiv.org3 days agoView details
DARE: Dialectical Agentic Reasoning for Structured Knowledge Fact Checking
arXiv:2609.13808v1 Announce Type: new Abstract: Structured knowledge fact checking aims to determine the truthfulness of natural language claims by reasoning over structured evidence. Recent program-generation approaches leverage large language models (LLMs) to generate executable graph reasoning programs, achieving s…
arxiv.org3 days agoView details
SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
arXiv:2609.14320v1 Announce Type: new Abstract: Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the…
arxiv.org3 days agoView details
arXiv:2609.13158v1 Announce Type: new Abstract: Large Vision--Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize long-document understanding with limited reasonin…
arxiv.org3 days agoView details
arXiv:2609.13869v1 Announce Type: new Abstract: Automatic sentence function identification is important for many downstream natural language processing (NLP) applications such as dialogue systems, text-to-speech synthesis, and machine translation. However, benchmark resources for Bangla sentence function classificatio…
arxiv.org3 days agoView details
Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself
arXiv:2609.13657v1 Announce Type: new Abstract: Traditional recommender systems are typically trained to predict what item users will interact with next, but not why. However, offering personalized evidence for why a user might like the predicted item is an important way to enhance the service and to raise the likelih…
arxiv.org3 days agoView details
FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks
arXiv:2609.13580v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities across a wide range of natural language processing tasks. However, conventional fine-tuning typically relies on centralized data collection, bringing in privacy concerns. Federated learning (FL) enables c…
arxiv.org3 days agoView details
TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models
arXiv:2609.13457v1 Announce Type: new Abstract: Timeseries multimodal large language models (TS-MLLMs) have recently begun leveraging the reasoning capabilities of large language models (LLMs) for question-answering tasks. However, these models often fail to capture dynamic temporal patterns, providing only implicit r…
arxiv.org3 days agoView details
ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation
arXiv:2609.13737v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk methods that rely on surface tokens or ou…
arxiv.org3 days agoView details
arXiv:2609.13709v1 Announce Type: new Abstract: Human corrections identify editable spans, but the examples receiving corrections may come from a selective feedback channel. We analyze this interaction at a fixed model checkpoint by decomposing a localized gradient into edited and retained untouched components. Square…
arxiv.org3 days agoView details
Formal Properties of Language as Constraints on Neural Dynamics
arXiv:2609.14384v1 Announce Type: new Abstract: What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes, voxels, or language-model layers predict neural activity. Yet predictive success leaves mechanisms under-constrained. H…
arxiv.org3 days agoView details
Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation
arXiv:2609.13581v1 Announce Type: new Abstract: Discharge summaries are lengthy medical documents that summarize a hospital in-patient visit. Automatically generating them can reduce documentation burden and return clinician time to patient care. Whereas Large Language Model (LLMs) could be used for this task, their A…
arxiv.org3 days agoView details
Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs
arXiv:2609.13445v1 Announce Type: new Abstract: Speech-to-speech LLMs like Moshi, and its derivative PersonaPlex, can listen and speak concurrently through full-duplex generation. However, they can begin speaking inappropriately during prolonged user silence: under digital-zero input, Moshi and PersonaPlex initiate sp…
arxiv.org3 days agoView details
arXiv:2609.13154v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et al., 2022) and in-context learning (Brown et al., 2020) frequently push real-world prompts past several thousand tokens…
arxiv.org3 days agoView details
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
arXiv:2609.13356v1 Announce Type: new Abstract: In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric cap…
arxiv.org3 days agoView details
VeriDx: Earning the Right to Diagnose with Disease-Centric Verification
arXiv:2609.14018v1 Announce Type: new Abstract: A correct diagnosis can still be reached for the wrong reasons. In clinical reasoning, every disease hypothesis creates obligations: key evidence must be checked, alternatives must be ruled out, contradictions must be resolved, useful tests must be considered, and closur…
arxiv.org3 days agoView details
arXiv:2609.13805v1 Announce Type: new Abstract: In the era of the Internet of Things (IoT), coordinating connected electric vehicle (EV) charging scheduling to balance EV charging satisfaction, station profitability, and smart grid stability presents a complex multi-objective challenge. Existing Multi-Agent Reinforcem…
arxiv.org3 days agoView details
Cost Characterization of Vertically Partitioned Federated Knowledge Graphs
arXiv:2609.13664v1 Announce Type: new Abstract: Knowledge graphs are increasingly distributed across autonomous organizations that share an entity space but own disjoint subsets of relations, forming a vertical partition. Answering a multi-hop query may require combining facts from several silos, making the partitioni…
arxiv.org3 days agoView details
Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
arXiv:2609.13519v1 Announce Type: new Abstract: Molecular design is most effective when generation mirrors the edits chemists actually make: extending a scaffold, replacing a substituent, or decorating a scaffold at a specified attachment site while optimizing molecular properties. Fragment-based molecular design natu…
arxiv.org3 days agoView details
Token Efficient Task Execution via Application Behavior Modeling for Web Agents
arXiv:2609.13491v1 Announce Type: new Abstract: The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in…
arxiv.org3 days agoView details
Dynamic Learning Solutions: A System for Personalized Educational Video Generation
arXiv:2609.14408v1 Announce Type: new Abstract: We present an automated pipeline that converts NCERT textbooks into interactive video explanations that respond directly to user queries. A user uploads a PDF and asks a question; the system then generates a video-based explanation as output, handling both text and visua…
arxiv.org3 days agoView details
Convergent Emergence of In-Context Learning Across Modalities
arXiv:2609.14011v1 Announce Type: new Abstract: Few-shot in-context learning (ICL), the capacity of a model to infer abstract patterns from input-output examples provided in its prompt and apply them to new inputs, has been extensively studied in large language models trained for next-token prediction on human text. R…
arxiv.org3 days agoView details
MANAS-2: Constrained Reconstruction for EEG Foundation Models
arXiv:2609.13717v2 Announce Type: new Abstract: Masked reconstruction is widely used for EEG foundation models, but optimizing reconstruction on low-SNR waveforms does not necessarily produce the most useful latent representation. We introduce MANAS-2, a new EEG foundation model that combines a Raw-Band Hybrid (RBH) m…
arxiv.org3 days agoView details
Editorial routing shapes how computational results are qualified in AI-assisted scientific writing
arXiv:2609.14288v1 Announce Type: new Abstract: Large language models increasingly analyze computational results and draft manuscripts, making reliable communication as important as correct analysis. Using fixed computational evidence, we tested whether assigning comparisons across modeling choices elsewhere in a rese…
arxiv.org3 days agoView details
How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition
arXiv:2609.13747v1 Announce Type: new Abstract: Large language models solve hard problems through intermediate computations across multi-step reasoning. Traditional chain-of-thought encodes these computations as tokens. Recent continuous and recurrent methods instead move partial computations into fixed-dimensional la…
arxiv.org3 days agoView details
arXiv:2609.13396v1 Announce Type: new Abstract: Multi-objective Bayesian optimisation (MOBO) is a sample-efficient approach for optimising expensive black-box functions with multiple objectives. In MOBO, the goal is to adequately approximate the Pareto front; that is, to obtain a high-quality solution set with 1) good…
arxiv.org3 days agoView details
arXiv:2609.13406v1 Announce Type: new Abstract: When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards autonomous and evolving intelligence, RSI is being claimed at many scales, while no single framework that formally describes these emerging instances e…
arxiv.org3 days agoView details
PolicyMem: Geometric Policy Memory for LLM Governance
arXiv:2609.13734v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in real-world high-stakes applications, effective governance has become essential. Existing safeguards largely follow two paradigms: learning-based guards provide strong semantic discrimination but couple policy b…
arxiv.org3 days agoView details
arXiv:2609.13615v1 Announce Type: new Abstract: For our submission to the WMT26 Creole Language Translation Shared Task, we focus on machine translation (MT) models for Pacific creoles: Tok Pisin, Bislama, and Solomon Pijin, with particular attention to broad domain performance. After pre-training on a large collectio…
arxiv.org3 days agoView details
AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
arXiv:2609.13548v1 Announce Type: new Abstract: Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for construc…
arxiv.org3 days agoView details
Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
arXiv:2609.13422v1 Announce Type: new Abstract: LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-i…
arxiv.org3 days agoView details
OrchSLM: Probing the Dynamics of Small Language Model Orchestration
arXiv:2609.13470v1 Announce Type: new Abstract: Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. S…
arxiv.org3 days agoView details
Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks
arXiv:2609.13712v1 Announce Type: new Abstract: Many real-world scenarios can be represented using graph-structured data. However, traditional GNNs that transmit messages based on first-order neighbors have long faced several fundamental contradictions: increasing depth leads to over-smoothing, long-range dependencies…
arxiv.org3 days agoView details
Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation
arXiv:2609.13238v1 Announce Type: new Abstract: Maxillofacial report generation from cone beam computed tomography is scored here by a composite objective placing 80% of its weight on a large language model judgement of factual entailment and 20% on lexical overlap, of which only the lexical fifth is visible during de…
arxiv.org3 days agoView details
RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines
arXiv:2609.13389v1 Announce Type: new Abstract: Mapping textual specifications into formal representations is essential for ensuring the correctness of protocol designs and implementations. LLM-generated mappings, used for networking security or testing, are assumed to capture a perfect understanding of the specificat…
arxiv.org3 days agoView details
CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages
arXiv:2609.13413v1 Announce Type: new Abstract: We introduce CVSS-X, a large-scale synthetic speech-to-speech translation corpus that extends CVSS by reversing the translation direction. While CVSS translates from 21 languages into English, CVSS-X enables translation from English into 28 target languages spanning 12 l…
arxiv.org3 days agoView details
A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics
arXiv:2609.13561v1 Announce Type: new Abstract: Efficient utilization of supply chain analytics for decision making remains a significant challenge for planners, as critical tasks such as database querying, key performance indicator (KPI) analysis, demand forecasting, and performance diagnosis require heterogeneous ex…
arxiv.org3 days agoView details
Causal multi-modal AI for personalized chemosensitivity prediction
arXiv:2609.13567v1 Announce Type: new Abstract: Chemotherapy improves survival for some patients with breast cancer, but doctors cannot reliably predict who. Current guidelines rely on recurrence scores as a proxy for treatment benefit, which may contribute to the overprescription of chemotherapy. Here we present a ca…
arxiv.org3 days agoView details
How User-AI Mistreatment Occurs and Matters in Conversational Systems?
arXiv:2609.13579v1 Announce Type: new Abstract: Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world dep…
arxiv.org3 days agoView details
arXiv:2609.13481v1 Announce Type: new Abstract: Large language models have made abstractive summarization remarkably fluent, but generated summaries can hallucinate facts, posing serious risks in biomedical and clinical domains. We address this by removing generation from the pipeline and framing summarization as extr…
arxiv.org3 days agoView details
Models
View all models →image-text-to-text · transformers · safetensors · agnes
huggingface.co4 days ago207 ptsView details
automatic-speech-recognition · nemo · onnx · gguf
huggingface.co4 days ago36 ptsView details
coolthor/MiniMax-H3-pruned-NVFP4
image-text-to-video · diffusion-single-file · text-to-video · image-to-video
huggingface.co4 days ago22 ptsView details
mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit
text-generation · mlx · safetensors · qwen3_5_moe
huggingface.co4 days ago17 ptsView details
text-generation · pytorch · safetensors · sparse-ast
huggingface.co4 days agoView details
takanori-ishikawa/Qwen3.5-ANE-CoreML
text-generation · coreml · ane · apple-neural-engine
huggingface.co4 days agoView details
Open source
View all open source →## [3.14.0](https://github.com/openai/openai-python/compare/v3.13.0...v3.14.0) (2026-09-14) ### Features * **streaming:** normalize errors raised while reading streams ([#3827](https://github.com/openai/openai-python/issues/3827)) ([d7c41ef](https://github.com/openai/openai-pyth…
github.com3 days agoView details
<details open> HIP: fattn-mma: use fp32 accumulation on MFMA devices (#28576) use fp32 accumulators in fattn-mma on CDNA </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/47450174> **macOS/iOS:** - [macOS Apple…
github.com3 days agoView details
## Overview llama.cpp 0.4.1 adds Maple 20B-A1B, Tencent Hy 4, and Spark2.5 support, improves JSON schema handling, chat parsing, logging, and server child-process management, and updates ggml to v0.24.0. ### API changes - Changed `llama_sampler_chain_n()` to return `int32_t` ins…
github.com3 days agoView details
<details open> sycl : fix oneDNN scratchpad breaking the pool free order (#28704) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/47287268> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggm…
github.com4 days agoView details