Archive / 2026-09-03
September 3, 2026
News
View all news →GPT-6 Astra: A new generation of intelligence
Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science.
openai.com14 days ago2278 ptsView details
github.com14 days ago1152 ptsView detailsJoin discussion
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
inference-docs.cerebras.ai14 days ago689 ptsView detailsJoin discussion
Any Human Ever – One life, drawn at random from all who have ever lived
anyhumanever.com14 days ago652 ptsView detailsJoin discussion
Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly
babyloniantwins.com14 days ago380 ptsView detailsJoin discussion
Artificial beaver dams saw juvenile coho salmon survival rates go from 8% to 60%
discoverwildlife.com14 days ago374 ptsView detailsJoin discussion
K2 Horizon: A connected fleet of six open models
ifm.ai14 days ago335 ptsView detailsJoin discussion
Nvidia to acquire Hugging Face
cnbc.com14 days ago329 ptsView detailsJoin discussion
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
armature.tech14 days ago300 ptsView detailsJoin discussion
OpenAI begins rolling out GPT-6 Astra
cnbc.com14 days ago277 ptsView detailsJoin discussion
OpenAI's GPT-6 Astra on ARC-AGI-3
arcprize.org14 days ago235 ptsView detailsJoin discussion
Sony makes bold claim about game ownership
aginggamer.net14 days ago179 ptsView detailsJoin discussion
Babylonian Lamb Stew with Beets (1750–1730 BCE)
babylonian-collection.yale.edu14 days ago171 ptsView detailsJoin discussion
status.x.ai14 days ago159 ptsView detailsJoin discussion
theguardian.com14 days ago148 ptsView detailsJoin discussion
Grep beats LSP? Why coding agents ignore your fancier tools
agentconnect.md14 days ago97 ptsView detailsJoin discussion
Japan halves speed limit to 30km/h on all narrow city streets
theguardian.com14 days ago87 ptsView detailsJoin discussion
reactoratlas.com14 days ago79 ptsView detailsJoin discussion
Your Racist Linux Distro Is Very Nice (Scott Jennings)
brokentoys.org15 days ago69 ptsView detailsJoin discussion
Sanders introduces bill to ban artificial superintelligence and pause AI
sanders.senate.gov14 days ago61 ptsView detailsJoin discussion
claude.com14 days ago61 ptsView detailsJoin discussion
A dark horse enters China's AI race: StartLux
chinaonchina.com14 days ago55 ptsView detailsJoin discussion
Never Forget How Eagerly Apple and Google Coddled Fascism
karlbode.com14 days ago50 ptsView detailsJoin discussion
OpenAI's new reasoning technique alarms AI safety experts
techcrunch.com14 days ago39 ptsView detailsJoin discussion
Protecting Engineers' Skills in the AI Era
spectrum.ieee.org14 days ago35 ptsView detailsJoin discussion
"Welcome to the AGI era," OpenAI says as GPT-6 Astra debuts
axios.com14 days ago35 ptsView detailsJoin discussion
Yes, no (built-in) AI is now a feature – LibreOffice blog
blog.documentfoundation.org14 days ago35 ptsView detailsJoin discussion
Why has Peter Thiel moved to Argentina?
theguardian.com14 days ago32 ptsView detailsJoin discussion
GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index
artificialanalysis.ai14 days ago27 ptsView detailsJoin discussion
Why office workers are turning against AI
bloodinthemachine.com14 days ago27 ptsView detailsJoin discussion
mail.cyberneticforests.com14 days ago26 ptsView detailsJoin discussion
Winter 2026/2027 First Forecast: Super El Niño Drives a Major Weather Divide
severe-weather.eu14 days ago22 ptsView detailsJoin discussion
Kids go from curious to frustrated playing with AI-stuffed toys, UW study finds
geekwire.com15 days ago23 ptsView detailsJoin discussion
Judge Tells RFK Jr. To Stop Using Fake AI Studies on Teen Pregnancy
newrepublic.com14 days ago21 ptsView detailsJoin discussion
The Double Matthew Walker Knot by Fable 5.1
claude.ai14 days ago20 ptsView detailsJoin discussion
No–AI Agents Did Not Build Secret Civilizations Stop Anthropomorphizing Malware
internetofbugs.substack.com15 days ago20 ptsView detailsJoin discussion
Congress Is Trying (Again) to Ban Boycotting Israel
reason.com14 days ago18 ptsView detailsJoin discussion
The paradox of diffusion distillation (2024)
sander.ai14 days ago18 ptsView detailsJoin discussion
twitter.com14 days ago18 ptsView detailsJoin discussion
OpenAI says it has overtaken Anthropic with its latest AI model
giftarticle.ft.com14 days ago16 ptsView detailsJoin discussion
Show HN: I built my first MCP to manage Google Ads
adchestra.com14 days ago15 ptsView detailsJoin discussion
Meta wanted to reduce teams by 60% because of AI
newsletter.pragmaticengineer.com14 days ago14 ptsView detailsJoin discussion
Show HN: Real-time AI news aggregator with daily digest
aibriefs.news14 days ago13 ptsView detailsJoin discussion
Three-LLM: Three.js-based WebGPU LLM inference engine
three-llm.ben3d.ca14 days ago12 ptsView detailsJoin discussion
twitter.com14 days ago12 ptsView detailsJoin discussion
twitter.com14 days ago12 ptsView detailsJoin discussion
Hugging Face is too important to fall into Nvidia's hands
theregister.com14 days ago11 ptsView detailsJoin discussion
'Starwashing': The new space race has an environmental problem
grist.org14 days ago11 ptsView detailsJoin discussion
Show HN: Ardent, a code-first agent for non-engineering work
ardent.ai14 days ago10 ptsView detailsJoin discussion
Show HN: A searchable, timestamped index of 1,124 AI Engineer talks
aietalks.com14 days ago10 ptsView detailsJoin discussion
- Primary source
Daybreak for Frontline Defenders: $1B to protect essential services
OpenAI introduces Daybreak for Frontline Defenders. A $1 billion commitment expands access to frontier cyber AI, training, and support for essential services.
openai.com14 days agoView details
- Primary source
Playco cut manual fixes 50% prototyping games with GPT-6 Astra
Using GPT-6 Astra, Playco built three themed game prototypes from one grey box foundation and reported 50% fewer manual fixes than with the previous model.
openai.com14 days agoView details
- Primary source
Legora reviewed 41 documents in minutes with GPT-6 Astra
Legora used GPT-6 Astra to review 41 documents in minutes, find all four planted errors, and improve performance by nearly 40% in this financial-review workflow.
openai.com14 days agoView details
- Primary source
Transfer learning for genomic prediction in underrepresented populations
General Science
research.google14 days agoView details
- Primary source
A connectomics milestone: Mapping the complete male fruit fly brain
General Science
research.google14 days agoView details
- Primary source
Introducing WeatherNext 3, our most advanced and accurate global weather AI model
deepmind.google14 days agoView details
- Primary source
NeoMME: an efficient Multimodal-native and Multilingual Encoder
huggingface.co14 days agoView details
WeatherNext 3 ingests live geostationary satellite mosaics, refreshes hourly, and outputs 5 km forecasts across Search, Gemini, Maps. The post Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour appeared first on MarkTechPost.
marktechpost.com14 days agoView details
OpenAI released GPT-6 Astra on September 3, 2026, positioning it as a computer-use flagship rather than a chat model. It reports 72.6% on OSWorld V2-Offline, replaces Codex compaction with searchable notes, and ships a 1.05M-token context at $10/$50 per million tokens. It is also the first OpenAI model to cross the Cr…
marktechpost.com14 days agoView details
Most teams building a shopping assistant or agent rebuild the same scaffolding: an agent loop, a tool layer over the catalog, an approval gate, and an eval suite. Anthropic has now released that scaffolding as code. This week, they published anthropics/commerce-agents, a reference blueprint containing a shopping agent…
marktechpost.com14 days agoView details
Perplexity has shipped hybrid compute for its Mac app, splitting a single Perplexity Computer task between frontier models in the cloud and a compact model running on the user's machine. Tasks start in the cloud for search, planning and reasoning, then hand sensitive steps down to the Mac without restarting or losing…
marktechpost.com14 days agoView details
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, it averages 1.23x MLX-LM's prefill throughput and 1.35x its decode throughput on a 40-core, 128 GB M5 Max. The post Perplexity Open Source…
marktechpost.com15 days agoView details
Nobody Is Saying Why OpenAI and Anthropic Had Outages Today
ChatGPT, Claude, and Grok all suffered outages at nearly the exact same time for reasons that remain murky.
wired.com14 days agoView details
Prediction Market Betting Is Getting People Banned and Arrested
This week on Uncanny Valley, we dig into the latest prediction market buzz, Flock’s AI-powered police search tool, and how tech bros don’t know how to talk about “rouge” AI agents
wired.com14 days agoView details
GPT-6 Astra Is Here—and OpenAI Thinks It May Kick Off the AGI Era
OpenAI leaders think the company’s next generation model, which excels at computer use and coding, may mark a major milestone in AI development.
wired.com14 days agoView details
OpenAI’s next big AI model has ‘entered the AGI era’
OpenAI's next big model is here: GPT-6 Astra. The company calls it a "generational leap in capability" for areas like cybersecurity, professional work, software engineering, science, and computer use. As OpenAI announced earlier this week, it's also the first model designated as meeting OpenAI's "critical cybersecurit…
theverge.com14 days agoView details
OpenAI Cut Off a Billion-Dollar Customer to Avoid Elon Musk
OpenAI recently estimated its Cursor partnership would make more than $1 billion in revenue a year, WIRED has learned. It still walked away after Elon Musk’s SpaceX acquired the AI coding startup.
wired.com14 days agoView details
Nvidia launches free tool that links idle computers into a personal AI data center
Nvidia is announcing its new Personal AI Router (PAIR), a free tool that syncs up your home computers for tackling local AI inference tasks with tools like Ollama and LM Studio. Let's get the obvious thing out of the way, despite what its name might imply: PAIR is not a hardware router. It's open-source software devel…
theverge.com14 days agoView details
Google now lets you chat with Gmail, Docs, and Keep
Google is rolling out AI-powered voice assistant modes for Gmail, Docs, and Keep that allow you to manage the apps by talking to them. The real time conversational capabilities are called Gmail Live, Docs Live, and Keep Live, and like the Gemini Live experience for Google's chatbot, aim to make it easier to note down…
theverge.com14 days agoView details
ChatGPT, Grok, and Claude all went down at the same time
OpenAI's ChatGPT, xAI's Grok, and Anthropic's Claude are back online after they all began experiencing issues around the same time on Thursday. At about 11AM ET, ChatGPT started returning error messages for users trying to use the chatbot, with its status page saying there were "elevated errors across ChatGPT and Code…
theverge.com14 days agoView details
Google says its AI weather model is getting better
Google is rolling out an updated AI weather model that's supposed to be more accurate, especially when it comes to predicting rain and snowfall. In the announcement today, the company says it's now able to make forecasts with "unprecedented resolution" using its new WeatherNext 3 AI model. It can produce a global pict…
theverge.com14 days agoView details
Nvidia RTX Spark ‘Superchip’: The First AI PCs Are Here
At IFA 2026, Nvidia and its partners showed off the first RTX Spark-powered laptops and mini PCs, designed to run AI models right on your computer.
wired.com14 days agoView details
Nvidia’s Hugging Face Acquisition Is a $12.9 Billion Bet on Open-Source AI
The long-rumored deal will give the chip giant access to—and help it promote—a huge repository of open-source AI models and data sets.
wired.com14 days agoView details
Nvidia is buying Hugging Face for almost $13 billion
Nvidia has agreed to buy Hugging Face for $12.93 billion, bringing one of the most popular hosting platforms for open-source AI models, datasets, and tools under the ownership of the world's biggest AI chipmaker. Hugging Face is an online platform founded in 2016 that gives AI developers a space to share their project…
theverge.com14 days agoView details
This Is Flock’s AI Search Tool for Cops
WIRED rebuilt Flock’s latest search tool from code the company sends to a police officer’s browser. Its AI can keep watch across multiple cameras for anyone fitting a written description.
wired.com14 days agoView details
MasterControl Seventeen Every Time
arXiv:2609.03209v1 Announce Type: new Abstract: We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive withi…
arxiv.org14 days agoView details
Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
arXiv:2609.03340v1 Announce Type: new Abstract: Distributed LLM-agent teams can read the latest shared facts and still act on an obsolete plan. A planner may derive an action from requirement $r_3$, another agent may commit $r_4$, and an executor may receive $r_4$ without replacing the plan derived from $r_3$. We call…
arxiv.org14 days agoView details
Dalek: A Constructive Agent Machine
arXiv:2609.03546v1 Announce Type: new Abstract: We present Dalek, a closed machine designed for agents that realizes self-maintenance, self-evolution, self-reproduction, and self-organization on any substrate satisfying a general host contract. The machine is built from three primitives---actors, messages, and channel…
arxiv.org14 days agoView details
Opening mind by opening architecture: analysis strategies
arXiv:2609.03719v1 Announce Type: new Abstract: In numerical signal processing for electroacoustic composition, the progressive loss of specific development and research environments caused by the increasing use of digital market tools has favoured the dominance of the closed-architecture audio processor model. This m…
arxiv.org14 days agoView details
arXiv:2609.02940v1 Announce Type: new Abstract: Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However, how to effectively combine these two str…
arxiv.org14 days agoView details
arXiv:2609.04127v1 Announce Type: new Abstract: Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncert…
arxiv.org14 days agoView details
arXiv:2609.02895v1 Announce Type: new Abstract: Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformation, posing substantial risks to public safety and social cohesion. While automated fake news detectio…
arxiv.org14 days agoView details
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
arXiv:2609.03494v1 Announce Type: new Abstract: Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-request budget and adjust only which KV states are retained, leaving the total capacity fixed…
arxiv.org14 days agoView details
Lose the Order, Keep the Hierarchy: Deordering HTN Plans
arXiv:2609.03912v1 Announce Type: new Abstract: Hierarchical Task Network (HTN) planning is a powerful planning formalism based on task decomposition. Although most of the literature studied plan generation, comparatively less attention has been paid to post-plan optimization. In particular, plan deordering has been e…
arxiv.org14 days agoView details
arXiv:2609.03887v1 Announce Type: new Abstract: How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify…
arxiv.org14 days agoView details
arXiv:2609.03702v1 Announce Type: new Abstract: General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (se…
arxiv.org14 days agoView details
arXiv:2609.03402v1 Announce Type: new Abstract: Artificial intelligence (AI) teaching assistants powered by large language models (LLMs) offer scalable educational support but often provide limited personalization. This study presents a prompt-engineering-based framework for personalizing general-purpose LLM/RAG-based…
arxiv.org14 days agoView details
Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection
arXiv:2609.03416v1 Announce Type: new Abstract: LLM-empowered paper-code discrepancy detection has received growing concern since the scaling of research submissions exceeds the manual review capability. However, the limited context capacity and one-sided discrepancy detection of existing single-agent LLM paradigms le…
arxiv.org14 days agoView details
DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
arXiv:2609.03423v1 Announce Type: new Abstract: Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often co…
arxiv.org14 days agoView details
Language, Language Models, and What We're Talking About
arXiv:2609.03577v1 Announce Type: new Abstract: Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to the nature of the systems which result fr…
arxiv.org14 days agoView details
A computable representation of the physical laboratory enables verifiable workflows
arXiv:2609.03621v1 Announce Type: new Abstract: Making science computable requires representations of both scientific knowledge and the physical world in which scientific claims are tested. A computable representation of the physical laboratory is established through typed research objects, capability-bound operations…
arxiv.org14 days agoView details
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
arXiv:2609.03430v1 Announce Type: new Abstract: Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how…
arxiv.org14 days agoView details
Pattern Over-Generalization of Knowledge Graph Embedding
arXiv:2609.03487v1 Announce Type: new Abstract: Knowledge graph embedding (KGE) demonstrates its effectiveness for predicting missing links in knowledge graphs (KGs) by projecting entities and relations into a low-dimensional vector space. It is crucial for KGE models to effectively capture inference patterns (pattern…
arxiv.org14 days agoView details
The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification
arXiv:2609.03652v1 Announce Type: new Abstract: Synthetic data augmentation has become a common strategy for addressing class imbalance in NLP, but most approaches focus on the quantity and diversity of generated examples rather than their geometric relationship to real training data. We investigate this question in t…
arxiv.org14 days agoView details
Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition
arXiv:2609.02901v1 Announce Type: new Abstract: Modern automatic speech recognition (ASR) scenarios require both spoken-form transcripts for faithful transcription and readable written-form transcripts with inverse text normalization (ITN). However, these forms are typically produced by cascaded modules, where a spoke…
arxiv.org14 days agoView details
Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence
arXiv:2609.02981v1 Announce Type: new Abstract: Artificial intelligence is changing the form of applied English materials from fixed paper sequences to adaptive learning systems that can diagnose learners, recommend tasks, and provide formative feedback. This paper studies the structure and application of a new practi…
arxiv.org14 days agoView details
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
arXiv:2609.03438v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to benign mistakes. A reliable agent should not only know how to act, but also when not to act. I…
arxiv.org14 days agoView details
RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents
arXiv:2609.02902v1 Announce Type: new Abstract: Deploying task-oriented dialogue agents in enterprise customer support faces a persistent annotation bottleneck: robust training requires labelled interaction data at scale, yet enterprise conversational logs are privacy-sensitive and expensive to annotate, while user be…
arxiv.org14 days agoView details
arXiv:2609.03221v1 Announce Type: new Abstract: Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that th…
arxiv.org14 days agoView details
arXiv:2609.03527v1 Announce Type: new Abstract: Neonatal respiratory diseases are a major cause of neonatal morbidity and mortality, posing substantial challenges in clinical practice. Despite recent advances, existing Multimodal Large Language Models (MLLMs) face two key limitations in neonatal diagnosis: (1) domain…
arxiv.org14 days agoView details
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
arXiv:2609.03407v1 Announce Type: new Abstract: People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptio…
arxiv.org14 days agoView details
TabScope: Question-Adaptive Scope Selection for Table Question Answering
arXiv:2609.03395v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy often degrades as table size increases. We find that this degradation is not uniform across question types. Localization-sensitive questions are particularly affect…
arxiv.org14 days agoView details
The Attention Triangle in Audio-Video Models
arXiv:2609.03586v1 Announce Type: new Abstract: Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising t…
arxiv.org14 days agoView details
FrameBench:A Language Understanding Benchmark Based on Frame Semantics
arXiv:2609.03370v1 Announce Type: new Abstract: In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent large language models (LLMs) have achieve…
arxiv.org14 days agoView details
Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
arXiv:2609.03460v1 Announce Type: new Abstract: As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. B…
arxiv.org14 days agoView details
Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation
arXiv:2609.03535v1 Announce Type: new Abstract: Lesion segmentation in medical images plays a critical role in clinical diagnosis and treatment planning. Despite significant advances, lesion segmentation remains challenging due to two major factors: (1) complex background interference; (2) diverse lesion morphology. E…
arxiv.org14 days agoView details
Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
arXiv:2609.02897v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static dra…
arxiv.org14 days agoView details
arXiv:2609.03254v1 Announce Type: new Abstract: Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies an…
arxiv.org14 days agoView details
Speculative Macro Commit for Faster Tool-Using Agents
arXiv:2609.03236v1 Announce Type: new Abstract: Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and observation can delay subsequent decisions. We introduce \textbf{Speculative Macro Commit} (SMC), a run…
arxiv.org14 days agoView details
Analysis of Prompt Engineering for Drug Toxicity Prediction
arXiv:2609.03635v1 Announce Type: new Abstract: Clinical trials in the UK can cost up to {\pounds}1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence has been increasingly…
arxiv.org14 days agoView details
CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception
arXiv:2609.03818v1 Announce Type: new Abstract: Collaborative perception enhances environment understanding through multi-agent information sharing, but its performance in real-world scenarios is constrained by heterogeneous sensor modalities and model architectures. Recent protocol-based two-stage methods alleviate t…
arxiv.org14 days agoView details
arXiv:2609.03493v1 Announce Type: new Abstract: Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cro…
arxiv.org14 days agoView details
MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval
arXiv:2609.03201v1 Announce Type: new Abstract: Long-term LLM agents must preserve information across interactions while distinguishing repeated evidence, historical states, updates, and unresolved contradictions. Existing textual memory systems retrieve semantically relevant memories efficiently but often leave these…
arxiv.org14 days agoView details
Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations
arXiv:2609.03426v1 Announce Type: new Abstract: Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled wit…
arxiv.org14 days agoView details
Artificial Intelligence for Energy Optimization in Data Centers
arXiv:2609.03716v1 Announce Type: new Abstract: Data centers are increasingly optimized by artificial intelligence and, at the same time, increasingly loaded by it. The literature treats these as two unrelated problems: control studies model workload as an exogenous arrival process, while sustainability studies model…
arxiv.org14 days agoView details
AutoGraphForge: Towards Automated Graph Theory Discovery
arXiv:2609.03478v1 Announce Type: new Abstract: We report on our ongoing project to develop a computational pipeline, AutoGraphForge, for an automated graph-theoretic conjecturing-refuting-formalizing-proving system. Conjecture generation is counterexample-guided and runs in rounds: a Graffiti3 generator proposes conj…
arxiv.org14 days agoView details
How Far Can Synthetic Data Take Thai OCR?
arXiv:2609.03595v1 Announce Type: new Abstract: We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "r…
arxiv.org14 days agoView details
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
arXiv:2609.03502v1 Announce Type: new Abstract: In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programma…
arxiv.org14 days agoView details
arXiv:2609.03503v1 Announce Type: new Abstract: With the rapid development of the Internet of Things, computation intensive directed acyclic graph (DAG) tasks have become increasingly common in cloud-edge-end collaborative environments. However, cloud, edge, and end nodes are highly heterogeneous in computing capacity…
arxiv.org14 days agoView details
CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
arXiv:2609.03526v1 Announce Type: new Abstract: Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4…
arxiv.org14 days agoView details
Rethinking World Models for Safety-Critical Embodied Systems
arXiv:2609.03774v1 Announce Type: new Abstract: World models have progressed from compact latent dynamics to generative, controllable, and interactive simulators of embodied environments. However, high predictive likelihood and visual fidelity do not necessarily ensure that a model preserves the evidence required for…
arxiv.org14 days agoView details
Counterfactual Routing Using Integer Programming with Constraint Generation
arXiv:2609.03707v1 Announce Type: new Abstract: We present our submission to the IJCAI 2025 'Counterfactual Routing Competition' (CRC 25). The goal of the competition is to find counterfactual explanations for the shortest path problem. This requires deciding what the minimal changes to a road network would make a rou…
arxiv.org14 days agoView details
What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation
arXiv:2609.03515v1 Announce Type: new Abstract: Decoding-time KV cache compression research focuses heavily on designing better token scoring functions, while the temporal rule that aggregates scores across decode steps is often treated as an implementation detail. Under aggressive KV compression, we find that exponen…
arxiv.org14 days agoView details
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
arXiv:2609.03588v1 Announce Type: new Abstract: As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledg…
arxiv.org14 days agoView details
Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations
arXiv:2609.03860v1 Announce Type: new Abstract: Retail supply chain operations rely on coupled decision modules that must adapt as requirements evolve. LLMs offer a natural-language interface for this task, but existing methods primarily focus on individual optimization models. Extending them to heterogeneous decision…
arxiv.org14 days agoView details
Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
arXiv:2609.03727v1 Announce Type: new Abstract: Large language model agents can plan, invoke tools, and modify external states, yet most systems still take an explicit user instruction as a fixed starting point. Proactive service moves the decision upstream: an agent must infer service opportunities from incomplete en…
arxiv.org14 days agoView details
SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
arXiv:2609.03753v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the long-term value of AI systems depends not only on solving individual requests, but also on transforming experience and accumulated knowledge into durable, reusable competence. We introduce SimSkill, a self-…
arxiv.org14 days agoView details
Inferring Affective Consciousness in an Artificial Agent: A Case Study
arXiv:2609.03883v1 Announce Type: new Abstract: Creatures that display 'hedonic place preference behaviour' are thought by many scientists to experience feelings, on the assumption that their attraction to pleasure-producing substances which lack nutritional value (e.g. cocaine, morphine) cannot easily be attributed t…
arxiv.org14 days agoView details
Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding
arXiv:2609.03938v1 Announce Type: new Abstract: While HTN planning has received significant attention in recent years, support for numerical reasoning remains very limited. In this paper, we investigate numerical Totally-Ordered HTN (TOHTN) planning and show how standard SAT-based encodings can be naturally extended w…
arxiv.org14 days agoView details
arXiv:2609.04014v1 Announce Type: new Abstract: For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high repeatability. Yet Multimodal Large Language Models (MLLMs) remain unreliable in continuous-valued measurement despite strong results on general multimodal benchmar…
arxiv.org14 days agoView details
arXiv:2609.04021v1 Announce Type: new Abstract: Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically…
arxiv.org14 days agoView details
arXiv:2609.04098v1 Announce Type: new Abstract: Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent state summarizes the context in fixed size. Early community 4-bit quantizations of Qwen3.8-27B (48 GDN layers, 16 attention layers) left the GDN block in 8- or 1…
arxiv.org14 days agoView details
Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards
arXiv:2609.02899v1 Announce Type: new Abstract: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scor…
arxiv.org14 days agoView details
LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation
arXiv:2609.02954v1 Announce Type: new Abstract: Identifying the issues disputed between litigating parties is a crucial component of real-world litigation. However, legal issues remain comparatively underexplored in legal AI research. In this work, we study the computational modelling of legal issue identification in…
arxiv.org14 days agoView details
No country for old linguists: LLM-brain alignment underdetermines neural computation
arXiv:2609.03160v1 Announce Type: new Abstract: Nastase et al. (2026) argue that large language models (LLMs) may illuminate language processing because both rely on distributed, context-sensitive representations shaped by statistical learning. Their rejection of simple cortical "boxology" is persuasive, and they arti…
arxiv.org14 days agoView details
A Circuit for Plural Reference: How LLMs Represent and Retrieve Singular and Plural Entities
arXiv:2609.03687v1 Announce Type: new Abstract: Coreference resolution is an important task in contextual reasoning. In this paper, we investigate the mechanism for representing and retrieving singular and plural entities for plural reference. We use a combination of mechanistic interpretability and attention pattern…
arxiv.org14 days agoView details
Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks
arXiv:2609.03734v1 Announce Type: new Abstract: BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language…
arxiv.org14 days agoView details
LLMs Learn Better In-Context from Rules than from Examples
arXiv:2609.03213v1 Announce Type: new Abstract: Large language models (LLMs) exhibit in-context learning capabilities, where they can learn new tasks from prompt contexts without weight updates. We compare the learning efficacies of two prominent modes of in-context learning: (1) learning from descriptions of rules (i…
arxiv.org14 days agoView details
SWIM: Student Writing Simulation via Proficiency-Conditioned Generation
arXiv:2609.03215v1 Announce Type: new Abstract: Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplo…
arxiv.org14 days agoView details
Semantic Bayesian World Models
arXiv:2609.03834v1 Announce Type: new Abstract: Knowledge graphs describe reality in crisp assertions, while the systems now consuming them, foundation models and autonomous agents, reason natively in probabilities. We argue that this mismatch is why the integration of language models and knowledge graphs remains a da…
arxiv.org14 days agoView details
SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation
arXiv:2609.03806v1 Announce Type: new Abstract: Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics design…
arxiv.org14 days agoView details
Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation
arXiv:2609.03814v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply eac…
arxiv.org14 days agoView details
Value-Preserving Architectures for Agentic AI Systems
arXiv:2609.03920v1 Announce Type: new Abstract: The emergence of agentic AI and LLM-based multi-agent systems (MAS) presents unprecedented opportunities for automating complex tasks, while simultaneously raising critical concerns about the preservation of fundamental human-centered values, such as privacy, fairness, a…
arxiv.org14 days agoView details
Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting
arXiv:2609.03923v1 Announce Type: new Abstract: In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stances, coverage, and floor, they miss the moments where they should contribute. Prompt-only delegates stay silent on 51.4% of the absent participant's talking oppo…
arxiv.org14 days agoView details
More Criticism Does Not Make a Better Review: EquiReview-R
arXiv:2609.03943v1 Announce Type: new Abstract: AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet gen…
arxiv.org14 days agoView details
FiMI Banking: A Sovereign Model for Indian Retail Banking
arXiv:2609.03960v1 Announce Type: new Abstract: Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably meet these requirements. They fal…
arxiv.org14 days agoView details
The Dually Flat Geometry of Planning as Inference
arXiv:2609.04005v1 Announce Type: new Abstract: We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on…
arxiv.org14 days agoView details
Transfiver: Human-AI Co-Inference through a Shared Editable State
arXiv:2609.03797v1 Announce Type: new Abstract: Long-term human-AI interaction is difficult because the information that guides inference is updated implicitly by the model and is not directly inspectable or controllable by the user. We introduce the TRANSparent Framework for Interactive, Verifiable, Editable Represen…
arxiv.org14 days agoView details
arXiv:2609.02942v1 Announce Type: new Abstract: LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only o…
arxiv.org14 days agoView details
Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory
arXiv:2609.03394v1 Announce Type: new Abstract: Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows…
arxiv.org14 days agoView details
The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis
arXiv:2609.03218v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user context…
arxiv.org14 days agoView details
Interface-Induced Trajectory Censoring
arXiv:2609.03966v1 Announce Type: new Abstract: Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model is emitting well-formed calls: the interface censors the trajectory before anything downstream sees it. On BFCL v4's own data, executor and scorer, holding weigh…
arxiv.org14 days agoView details
Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing
arXiv:2609.03973v1 Announce Type: new Abstract: An image editor may satisfy every regional plausibility constraint separately even when no single latent explanation fits the complete output. We formalize this local-to-global failure using a common witness grade and witness nerve. The framework separates auditing from…
arxiv.org14 days agoView details
arXiv:2609.02898v1 Announce Type: new Abstract: Large domain-specific language models such as BioBERT and ClinicalBERT achieve strong performance on biomedical NLP tasks, but their computational demands make them impractical for many real-world deployments. General-purpose, parameter-efficient models such as DistilBER…
arxiv.org14 days agoView details
arXiv:2609.03871v1 Announce Type: new Abstract: Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics…
arxiv.org14 days agoView details
Instruction Duplication as an Inference-Time Control Primitive
arXiv:2609.04024v1 Announce Type: new Abstract: Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats onl…
arxiv.org14 days agoView details
IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations
arXiv:2609.04030v1 Announce Type: new Abstract: IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which ad…
arxiv.org14 days agoView details
Spurious Advantage Hidden in GRPO
arXiv:2609.04063v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each rollout a magnitude from within-group reward statistics. In the common case, this magnitude rewards rollouts that re…
arxiv.org14 days agoView details
Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards
arXiv:2609.03181v1 Announce Type: new Abstract: We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative de…
arxiv.org14 days agoView details
arXiv:2609.03321v1 Announce Type: new Abstract: The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializing turn-taking control and response generation onto a single causal tape under the standard next-token prediction objective, thereby preserving semantic prowess a…
arxiv.org14 days agoView details
How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models
arXiv:2609.03322v1 Announce Type: new Abstract: Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at…
arxiv.org14 days agoView details
Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour
arXiv:2609.03330v1 Announce Type: new Abstract: Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large…
arxiv.org14 days agoView details
arXiv:2609.03331v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually grounded false premise persists across dia…
arxiv.org14 days agoView details
arXiv:2609.03366v1 Announce Type: new Abstract: Accountability means a decision can be examined, justified, and contested. LLMs make this hard: fluent output may be ungrounded, incomplete, or unfaithful to the decision process. Achieving accountability requires verified rationales (how was the decision reached), assum…
arxiv.org14 days agoView details
SGD-KV: Summarization Guided KV Cache Compression
arXiv:2609.03235v1 Announce Type: new Abstract: Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of dif…
arxiv.org14 days agoView details
arXiv:2609.03273v1 Announce Type: new Abstract: Tamil spell and grammar correction is challenging because Tamil is an agglutinative low-resource language with rich verbal morphology, complex sandhi (phonetic transformation) rules at word boundaries, and a script of 247 distinct letters. Prior work targets word-level s…
arxiv.org14 days agoView details
When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents
arXiv:2609.03467v1 Announce Type: new Abstract: Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We i…
arxiv.org14 days agoView details
arXiv:2609.03432v1 Announce Type: new Abstract: Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a comp…
arxiv.org14 days agoView details
arXiv:2609.03597v1 Announce Type: new Abstract: Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system with dedicated Unicode glyphs, no mainstream font, and no coverage in any OCR pipeline or tokenizer. The handwritten records that carry these fractions, RS Khatian…
arxiv.org14 days agoView details
Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation
arXiv:2609.03619v1 Announce Type: new Abstract: Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of age…
arxiv.org14 days agoView details
arXiv:2609.03511v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to semantics-preserving structural variation remains underexplored, particularly for relatively free word-order languages. We investigate the structural sensitivity o…
arxiv.org14 days agoView details
</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination
arXiv:2609.03633v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strateg…
arxiv.org14 days agoView details
Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements
arXiv:2609.03654v1 Announce Type: new Abstract: The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical content across different jurisdi…
arxiv.org14 days agoView details
PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction
arXiv:2609.02896v1 Announce Type: new Abstract: Medical relation extraction (MRE) is commonly known for extracting entities and their relations jointly from a medical text, which has attracted considerable attention in recent years. Previous studies treat MRE as a sequence tagging task, which results in either a chall…
arxiv.org14 days agoView details
Typological Feature Prediction with Large Language Models: An In-Context Learning Approach
arXiv:2609.03775v1 Announce Type: new Abstract: Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while their performance across resource levels a…
arxiv.org14 days agoView details
arXiv:2609.03781v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSaf…
arxiv.org14 days agoView details
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
arXiv:2609.03553v1 Announce Type: new Abstract: Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when…
arxiv.org14 days agoView details
HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews
arXiv:2609.03580v1 Announce Type: new Abstract: The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability. Existing hallucination benchmarks are not designed for peer review…
arxiv.org14 days agoView details
Unifying Conformal Language Tasks with In-Context Ensembles
arXiv:2609.03005v1 Announce Type: new Abstract: Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pertinent information to achieve some goal, and conciseness, removing as much irrelevant inform…
arxiv.org14 days agoView details
DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions
arXiv:2609.03787v1 Announce Type: new Abstract: AI agents increasingly gather evidence, invoke tools, apply constraints, and produce decisions that people or software may commit to action. A final output alone cannot show which evidence, tool state, rule, authorization, or action path produced it. We present DNative-T…
arxiv.org14 days agoView details
To What Extent Do Large Language Models Understand Bangla Idioms?
arXiv:2609.03410v1 Announce Type: new Abstract: Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first large-scale benchmark dataset of Bangla idioms,…
arxiv.org14 days agoView details
arXiv:2609.02889v1 Announce Type: new Abstract: A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strategy, format rules, and control heuristics. Existing reflective prompt-evolution methods usually opti…
arxiv.org14 days agoView details
Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent
arXiv:2609.02890v1 Announce Type: new Abstract: A personalized language agent must convert a user's interaction history into behavior on each new request at inference time. Two strategies dominate. Retrieval pulls a few of the user's most relevant past items into the prompt, which is accurate but pays a per-query sele…
arxiv.org14 days agoView details
Counterexamples as Feedback for Agent Self-Correction
arXiv:2609.02892v1 Announce Type: new Abstract: Single-turn code-generation metrics understate a central property of deployed agents: whether they can repair a wrong artifact after receiving concrete feedback. This paper presents A-CEGIS, a lightweight framework that uses counterexamples as feedback for evaluating mul…
arxiv.org14 days agoView details
Probe Generalization as Subspace Selection for OOD Deception Detection
arXiv:2609.02893v1 Announce Type: new Abstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out deception detection datase…
arxiv.org14 days agoView details
R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG
arXiv:2609.02894v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a prevailing paradigm for enhancing Large Language Models (LLMs) with non-parametric knowledge. Vanilla RAG efficiently handles simple queries but struggles with relational or multi-hop reasoning. Graph-based RAG alleviates…
arxiv.org14 days agoView details
Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI
arXiv:2609.03800v1 Announce Type: new Abstract: Federated learning is increasingly presented as a privacy-preserving advance: personal data remain on the device, and only model updates are shared. It borrows the vocabulary of the federated social web, yet inverts its logic, distributing computation while the resulting…
arxiv.org14 days agoView details
SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking
arXiv:2609.03047v1 Announce Type: new Abstract: Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to run. SHELF, the Syn…
arxiv.org14 days agoView details
Large Language Models in Resolving Contextual Knowledge Conflicts
arXiv:2609.03148v1 Announce Type: new Abstract: Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided context. In contrast, we investigate how LLMs handle conflicts that arise within contextual knowledge itself. We introduce a taxonomy of six types of contextual c…
arxiv.org14 days agoView details
arXiv:2609.03874v1 Announce Type: new Abstract: Retrieval Augmented Generation (RAG) is a key component for generating accurate and hallucination free answers using Large Language Models (LLMs). LLMs are improving at handling long context, but still suffer from "lost in the middle" problem. Thus, precise and accurate…
arxiv.org14 days agoView details
PACE: Towards Surfacing Hidden Conflicts in User Requests
arXiv:2609.03293v1 Announce Type: new Abstract: Personalized assistants should not only comply with user requests but also assess whether those requests are appropriate given the user's current circumstances. However, prior work has primarily focused on accurately executing requests, overlooking the need for assistant…
arxiv.org14 days agoView details
Xiaomi-TabLDM: A Tabular Foundation Model Technical Report
arXiv:2609.03880v1 Announce Type: new Abstract: We introduce Xiaomi-TabLDM, a tabular large data foundation model for classification and regression via in-context learning, which delivers superior prediction accuracy without requiring task-specific fine-tuning. Pretrained exclusively on synthetic data generated from s…
arxiv.org14 days agoView details
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
arXiv:2609.04094v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popul…
arxiv.org14 days agoView details
LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening
arXiv:2609.04013v1 Announce Type: new Abstract: Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning (ML) and deep learning (DL) approaches require labeled data and model training, limiting their use in real-world screening settings. This study evaluates the ef…
arxiv.org14 days agoView details
When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA
arXiv:2609.03454v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) can improve the specificity and grounding of large language model responses, but its effect is not uniformly beneficial in single-turn mental-health question answering, where user queries often combine emotional distress, treatment co…
arxiv.org14 days agoView details
Models
View all models →text-generation · transformers · safetensors · k2_horizon
huggingface.co14 days ago293 ptsView details
microsoft/VibeVoice-ASR-Streaming-7B
automatic-speech-recognition · transformers · safetensors · vibevoice
huggingface.co15 days ago213 ptsView details
text-generation · transformers · safetensors · k2_horizon
huggingface.co14 days ago104 ptsView details
text-generation · transformers · safetensors · spark2_5
huggingface.co15 days ago98 ptsView details
inclusionAI/Ling-3.0-flash-Fin
text-generation · safetensors · bailing_hybrid · finance
huggingface.co14 days ago82 ptsView details
unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF
transformers · gguf · unsloth
huggingface.co14 days ago71 ptsView details
osk-arr00/lfm2.5-2.6B-thinkingcap-distiller
text-generation · transformers · safetensors · gguf
huggingface.co15 days ago4 ptsView details
2013khansohail/cartographer-ecommerce-reranker-MiniLM-L6-v2
text-ranking · sentence-transformers · safetensors · bert
huggingface.co15 days ago4 ptsView details
text-to-video · minimax-h3 · lora · text-to-video
huggingface.co15 days ago4 ptsView details
osk-arr00/lfm2.5-2.6B-thinkingcap-distiller-v2
text-generation · transformers · safetensors · gguf
huggingface.co15 days ago1 ptsView details
osk-arr00/LFM2.5-8B-A1B-ThinkingCap-GGUF
text-generation · gguf · llama.cpp · rocm
huggingface.co15 days ago1 ptsView details
text-generation · transformers · safetensors · gguf
huggingface.co15 days agoView details
text-generation · transformers · safetensors · gguf
huggingface.co15 days agoView details
text-generation · transformers · safetensors · gguf
huggingface.co15 days agoView details
text-generation · transformers · safetensors · gguf
huggingface.co15 days agoView details
Open source
View all open source →<details open> sycl: fuse rms_norm+mul+add and add+add residual chains (#27610) Fuse RMS_NORM+MUL+ADD and ADD+ADD under GGML_SYCL_ENABLE_FUSION. ADD+ADD uses the same binbcast indexing and type matrix as standalone add() (f32, f16, f16/f32, i32, i16, bf16, including broadcast an…
github.com14 days agoView details
## [3.8.0](https://github.com/openai/openai-python/compare/v3.7.0...v3.8.0) (2026-09-03) ### Features * **api:** add gpt-6-astra and related features ([#3791](https://github.com/openai/openai-python/issues/3791)) ([09f446f](https://github.com/openai/openai-python/commit/09f446f5…
github.com14 days agoView details
langchain-ai/langchain langchain==1.4.0
Changes since langchain==1.3.18 docs(langchain): runnable `langchain.mcp` examples (#39976) feat(langchain): `langchain.mcp` namespace, `MCPAdapter` (#39939) perf(anthropic,langchain): omit middleware trace inputs (#40098) fix(langchain): include model destination in agent tool…
github.com14 days agoView details
langchain-ai/langchain langchain-anthropic==1.7.1
Changes since langchain-anthropic==1.7.0 release(anthropic): 1.7.1 (#40181) perf(anthropic,langchain): omit middleware trace inputs (#40098) feat(anthropic): add Claude Fable 5.1 support (#40106)
github.com14 days agoView details
<details open> sycl: reduce redundant work in Q4_K multi-column MMVQ (#27062) * sycl: Q4_K Weight unpack optimization and reuse between destination Columns * sycl: Q4_K small N (N=2..4) + two output rows by subgroup reuse of activation between two rows. * sycl: gate Q4_K two-row…
github.com15 days agoView details
<details open> model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) * hparams: add per-layer n_ff_exp/n_expert_used arrays with scalar-or-array loading G1/G2 infrastructure for variable-per-layer expert FFN size and top-k routing (required for Puzzle-75…
github.com15 days agoView details