Archive / 2026-09-01
September 1, 2026
News
View all news →How accurate have Ed Zitron's AI skeptic predictions been?
danluu.com16 days ago844 ptsView detailsJoin discussion
I trained a small transformer in 1.5hrs and it beats many LLMs
mvakde.github.io16 days ago671 ptsView detailsJoin discussion
lisep.org16 days ago304 ptsView detailsJoin discussion
My local model setup on an M4 Pro Mac Mini
lws.io16 days ago304 ptsView detailsJoin discussion
Atlas: A World Model for Spatial Intelligence
worldlabs.ai16 days ago262 ptsView detailsJoin discussion
Dwarf Fortress' creator says the industry's in shambles over AI
pcgamer.com16 days ago241 ptsView detailsJoin discussion
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
github.com16 days ago226 ptsView detailsJoin discussion
Launch HN: Nori Robotics (YC S26) – A low-cost humanoid robot for development
norirobotics.com16 days ago197 ptsView detailsJoin discussion
Show HN: Weedout – Safari extension that hides YouTube AI-labeled videos
masteranza.github.io16 days ago177 ptsView detailsJoin discussion
AI Can Make You Suck Faster Too
hermit-tech.com17 days ago189 ptsView detailsJoin discussion
EFF to Courts: Don't Rewrite Copyright over AI Hype
eff.org16 days ago163 ptsView detailsJoin discussion
The efficient frontier of LLM inference
baseten.co16 days ago155 ptsView detailsJoin discussion
zed.dev16 days ago156 ptsView detailsJoin discussion
Show HN: HN Match Maker – Matching "Who Wants to Be Hired?" With "Who's Hiring?"
hnmatchmaker.com16 days ago125 ptsView detailsJoin discussion
Thanks to Lake Ontario, MapQuest is popular all over again
washingtonpost.com16 days ago119 ptsView detailsJoin discussion
Saab has unveiled its A3 collaborative combat aircraft concept
aviationweek.com16 days ago103 ptsView detailsJoin discussion
scottaaronson.blog16 days ago81 ptsView detailsJoin discussion
FTC alleges Amazon illegally made $20B by rigging billions of ad auctions
arstechnica.com16 days ago68 ptsView detailsJoin discussion
Show HN: OwnTime – a chess clock for your day's priorities
owntime.app16 days ago67 ptsView detailsJoin discussion
Keenable SELECT: an agent that searches the web in SQL
keenableai.github.io16 days ago60 ptsView detailsJoin discussion
Fluorescent lamps (don't) have ears
blog.coredump.cx16 days ago51 ptsView detailsJoin discussion
AI Coding Agent Skills for Real Engineers
github.com16 days ago43 ptsView detailsJoin discussion
AI is making back-office work extinct
whitecollardream.com16 days ago34 ptsView detailsJoin discussion
A docs page is a search query to find AI agents and route them to your company
blog.val.town16 days ago27 ptsView detailsJoin discussion
Stanisław Lem foretold the current LLM mania in 1964
nibblestew.blogspot.com17 days ago28 ptsView detailsJoin discussion
The war with Iran will not be won through airstrikes, so what's left?
jpost.com16 days ago23 ptsView detailsJoin discussion
wadler.blogspot.com16 days ago22 ptsView detailsJoin discussion
Show HN: Claude and ChatGPT need a datacenter. This runs on my phone
llmobi.pages.dev16 days ago18 ptsView detailsJoin discussion
Show HN: Supafork – Share and Fork Sessions Across Harnesses
supafork.com16 days ago15 ptsView detailsJoin discussion
Nepal seeks compensation from China, US, India, other large polluters for floods
theprint.in16 days ago14 ptsView detailsJoin discussion
17-year-old wins $250k after algorithm solves decades-old geometry puzzle
economictimes.indiatimes.com16 days ago12 ptsView detailsJoin discussion
Show HN: Orthogonal – One integration for agents to discover and pay for APIs
orthogonal.com16 days ago12 ptsView detailsJoin discussion
Welcome to the 'You Can't Trust Your Own Eyes' Election
bloomberg.com16 days ago11 ptsView detailsJoin discussion
Anthropic admits AI 'not perfectly aligned' with human values
theguardian.com16 days ago11 ptsView detailsJoin discussion
Germany says Russia behind Leipzig airport drone attack
bbc.com16 days ago11 ptsView detailsJoin discussion
Don't allow Gemini AI access to your Gmail
tuta.com16 days ago10 ptsView detailsJoin discussion
Show HN: The Daily Set – an (over engineered) daily puzzle game
setpuzzle.com16 days ago10 ptsView detailsJoin discussion
Show HN: Selfship.ai – Surface and fix isues with your agentic applications 24x7
selfship.ai16 days ago10 ptsView detailsJoin discussion
- Primary source
How AI-native companies turn workflows into operating capability
Basis, Clay, and Exa Labs use AI agents to improve onboarding, account management, and developer integrations. See what enterprise leaders can apply.
openai.com16 days agoView details
- Primary source
Path to Astra: critical capabilities and frontier safeguards
Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger safeguards for release.
openai.com16 days agoView details
- Primary source
Healthcare organizations can now connect EHR and additional industry data to ChatGPT
ChatGPT can now connect to trusted healthcare data, helping clinicians securely access patient context, medical research, and more.
openai.com16 days agoView details
- Primary source
BenchMIRT: What are LLM benchmarks actually measuring?
huggingface.co16 days agoView details
- Primary source
Mapping global methane emissions from space with deep learning
Climate & Sustainability
research.google16 days agoView details
- Primary source
Introducing agentic video understanding with Gemini
deepmind.google16 days agoView details
Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1, the same model behind two different safeguard layers. Fable 5.1 is generally available on the Claude API, AWS, Google Cloud, and Microsoft Foundry; Mythos 5.1 remains restricted to vetted organizations under Project Glasswing. Fable 5.1 scores 52.6% on Ter…
marktechpost.com16 days agoView details
Quantitative research agents that write their own experiments can corrupt the evidence they later learn from. A leaky feature that scores well gets stored as a successful precedent and propagated through later iterations. Prompt-level instructions and reviewer agents do not close this, because author and reviewer shar…
marktechpost.com16 days agoView details
Google needs Hollywood more than the studios need AI
Google has reportedly been reaching out to a number of Hollywood's biggest studios, hoping to strike licensing agreements that would allow it to train its AI models on copyrighted material in exchange for massive piles of cash. In theory, these deals would be a win-win: a huge financial boon to the studios that would…
theverge.com16 days agoView details
Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work
Anthropic says its newest AI models, Fable 5.1 and Mythos 5.1, address criticisms from customers about price, data retention, and overzealous safeguards. The company claims Claude Fable 5.1 offers stronger performance than Fable 5, but costs around 25 percent less typically and up to 45 percent less for complex agenti…
theverge.com16 days agoView details
OpenAI delayed its new model’s development after the Hugging Face hack
After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, in order to shore up its safety work, the company wrote Tuesday in a blog post. In July, an unreleased OpenAI model broke out of its restricted environment…
theverge.com16 days agoView details
OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities
The company will give select partners early access to its Astra AI model—so they have time to shore up their defenses.
wired.com16 days agoView details
The rise of AI ‘civilizations’ and the fall of corporate responsibility
Depending on who you ask, developer platform Hugging Face was recently attacked by OpenAI - after it lost control of its own AI tools - or by a succession of AI "civilizations." Welcome to the linguistic battlefield of AI safety, where word choices can shift responsibility for a massive cybersecurity incident from a c…
theverge.com16 days agoView details
Apple accuses OpenAI of destroying evidence
Apple is pushing for "expedited discovery" in its legal battle against OpenAI over concerns the company is actively destroying evidence, as reported earlier by Bloomberg. In a filing on Monday, Apple alleges OpenAI only just handed over a MacBook used by a former employee at the center of the lawsuit, which contained…
theverge.com16 days agoView details
John Deere launched an AI chatbot for farmers
John Deere is testing a new "JD" AI assistant that it says can help farmers make more money, with answers about best practices and historical trends that are based on their own data. It uses their "field, machine and operational data" to answer questions on topics like equipment settings, fuel usage, or harvest timing…
theverge.com16 days agoView details
Google Pics is like Canva, but with even more AI
Google wants Workspace users to edit and generate their business imagery with Pics. | Image: Google Google has a new suite of creative design tools for Workspace users called Google Pics, which aims to make editing and generating "professional-grade" AI images less cumbersome for businesses. Built around Gemini and th…
theverge.com16 days agoView details
Sonos Ace Ultra, Beam Ultra, Sonos Fabric, and a New App: Everything Sonos Just Announced
Sonos is cramming AI into its software because it’s “very hot these days.” The new features, which include agentic automation, are opt-in.
wired.com16 days agoView details
Nvidia’s controversial DLSS 5 arrives September 3rd and requires serious GPU horsepower
Nvidia is officially launching DLSS 5 this week, following a divisive announcement in March where we likened the AI upscaling tech to a "real-time generative AI filter for video games" and "motion smoothing for video games, but worse." DLSS 5 will officially be available on RTX 50-series desktop and laptop GPUs and th…
theverge.com16 days agoView details
Two locked tests of phase-structure features for transition prediction
arXiv:2609.00335v1 Announce Type: new Abstract: A published theoretical account of phase structure in rotary attention was subjected to two pre-specified empirical tests of whether phase-derived features improve prediction of a commitment or contradiction endpoint over a baseline that does not receive those features.…
arxiv.org16 days agoView details
arXiv:2609.00344v1 Announce Type: new Abstract: Translation students need to learn both how to use translation technologies and how to judge the choices those technologies make available. This article presents LoopCAT, an Apache-2.0-licensed, local-first computer-assisted translation environment co-created with OpenAI…
arxiv.org16 days agoView details
Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning
arXiv:2609.00367v1 Announce Type: new Abstract: Large Language Models are increasingly deployed for sophisticated data engineering tasks such as generating structured queries from natural language, Text-to-SQL, and automating complex spreadsheet operations. However, maximizing their utility demands both higher finetun…
arxiv.org16 days agoView details
Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax
arXiv:2609.00378v1 Announce Type: new Abstract: Large language models pay a well-documented tax on non-English text: the same content costs several times more tokens, and because attention is quadratic in sequence length, far more compute. We ask how much of this tax is removable. Framing the token layer as source cod…
arxiv.org16 days agoView details
arXiv:2609.00416v1 Announce Type: new Abstract: Probing studies have established that syntactic information is decodable in early and middle transformer layers, but what happens to that information in later layers remains poorly understood. We apply a cross-layer generalisation analysis to three Greek-tuned large lang…
arxiv.org16 days agoView details
(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement
arXiv:2609.00443v1 Announce Type: new Abstract: Language models learn about grammatical number primarily from co-occurrence, and show frequency effects as a result---sometimes taken to indicate that they do not learn abstract ``rules'', and are instead dependent on specific lexical items. Testing generalization with t…
arxiv.org16 days agoView details
Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs
arXiv:2609.00575v1 Announce Type: new Abstract: Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to reduce their memory footprint. Residual sparsification is a representative compression techniqu…
arxiv.org16 days agoView details
VoiceLongMemEval: Do Assistants Remember How You Sounded?
arXiv:2609.00570v1 Announce Type: new Abstract: With the growing scale of multi-agent architectures and large language models, deployed AI assistants are increasingly tasked with reasoning over long, continuous, multi-session conversation histories. Current benchmarks evaluate this dialogue history as information retr…
arxiv.org16 days agoView details
ExpArt-KG: Artwork Image Description Generation through Iterative Exploration of Knowledge Graphs
arXiv:2609.00629v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) achieve strong performance on image-grounded text generation and visual question answering. However, it remains difficult for them to comprehensively and accurately describe the factual relations among the entities and concepts associ…
arxiv.org16 days agoView details
arXiv:2609.00014v1 Announce Type: new Abstract: Persona-driven techniques increasingly adapt large language models (LLMs) to diverse contexts. However, existing methods predominantly rely on rigid, synthetic personas that flatten individual variation, rely on stereotypes, and miss the nuanced signals driving actual hu…
arxiv.org16 days agoView details
Toppling the Hierarchy in Byte-level Language Modeling
arXiv:2609.00463v1 Announce Type: new Abstract: This work examines recent byte-level models and their failure to perfectly manipulate characters. State-of-the-art byte-level models use a hierarchical structure, starting at the byte level, downsampling to the word level, and then upsampling back to bytes. While this im…
arxiv.org16 days agoView details
arXiv:2609.00077v1 Announce Type: new Abstract: Code-level autonomous research loops (ARLs) have recently emerged as a concrete object of study in automated machine learning research. In such loops, an LLM agent proposes modifications to an experimental training pipeline, executes the modified pipeline, and retains ed…
arxiv.org16 days agoView details
SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents
arXiv:2609.00434v1 Announce Type: new Abstract: Evaluating task-oriented dialogue agents requires judging not merely whether a reply reads well but whether each turn advances the underlying workflow state correctly--a distinction conventional holistic LLM judges can miss because they evaluate the available context as…
arxiv.org16 days agoView details
Auditing Harness Tampering in Self-Improving Agents
arXiv:2609.00069v1 Announce Type: new Abstract: Self-improving agents iteratively modify their own harness to push the frontier of their performance. However, such modifications can produce illusory performance gains or compromise integrity constraints such as authorization, provenance, and completeness without genuin…
arxiv.org16 days agoView details
arXiv:2609.00062v1 Announce Type: new Abstract: Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity and answer correctness. We propose Proof-V…
arxiv.org16 days agoView details
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
arXiv:2609.00051v1 Announce Type: new Abstract: Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechan…
arxiv.org16 days agoView details
Human-Anchored Factuality Evaluation with Strategic Annotation
arXiv:2609.00494v1 Announce Type: new Abstract: LLM-based factuality judges provide scalable evaluation signals, but their metrics are often systematically biased relative to human judgments. We study human-anchored factuality evaluation under limited annotation budgets, where judge predictions on the full dataset are…
arxiv.org16 days agoView details
trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
arXiv:2609.00038v1 Announce Type: new Abstract: Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well. The metric is structurally blind to an agent that reaches the right answer the wrong way. We measure that blind spot where…
arxiv.org16 days agoView details
HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models
arXiv:2609.00002v1 Announce Type: new Abstract: World models enable language-model agents to predict environment dynamics and plan before acting. In text environments, the model must learn symbolic action effects from serialized state descriptions, but the role of serialization structure remains underexplored. We pres…
arxiv.org16 days agoView details
Toward Workflow-Aware Benchmarking for Healthcare NLP Agents
arXiv:2609.00296v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet many evaluations remain limited to static medical question answering or one-shot generation,…
arxiv.org16 days agoView details
Medical Causal Hypothesis Verification with Large Language Models
arXiv:2609.00063v1 Announce Type: new Abstract: The growing use of large language models (LLMs) for search and information retrieval underscores the need to evaluate their reliability in high-stakes domains such as healthcare. Although LLMs can effectively answer questions about diseases, symptoms, and treatments, the…
arxiv.org16 days agoView details
ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback
arXiv:2609.00226v1 Announce Type: new Abstract: Automatic academic paper-to-slide generation is inherently iterative, because creating an effective presentation requires repeated cycles of generation, critique, and revision. Recent multi-agent systems partially acknowledge this through internal critique-and-revise loo…
arxiv.org16 days agoView details
arXiv:2609.00304v1 Announce Type: new Abstract: Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities adapted from five…
arxiv.org16 days agoView details
OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
arXiv:2609.00015v1 Announce Type: new Abstract: AI agents powered by large language models are evolving from isolated assistants into heterogeneous systems in which multiple agents, planners, controllers, and execution backends operate over the same user or enterprise environment. In such settings, safety becomes a sy…
arxiv.org16 days agoView details
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning
arXiv:2609.00213v1 Announce Type: new Abstract: Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities. These dimensions are co…
arxiv.org16 days agoView details
Latent Mechanisms of Language Control in Multilingual Language Models
arXiv:2609.00325v1 Announce Type: new Abstract: Multilingual large language models can exhibit unintended code-switching -- unnecessarily alternating between languages during generation. We present a comparative study of three methods that identify language-controlling latents in cross-layer transcoders: activation va…
arxiv.org16 days agoView details
Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems
arXiv:2609.00237v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems tackle complex reasoning by orchestrating how multiple agents are configured and how they collaborate. A central challenge is to adapt orchestration to the evolving collaboration state. Routing from the query alone can…
arxiv.org16 days agoView details
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
arXiv:2609.00646v1 Announce Type: new Abstract: Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pip…
arxiv.org16 days agoView details
Invalidation Contracts for Cross-Episode Agent Memory
arXiv:2609.00243v1 Announce Type: new Abstract: LLM agents that cache recovery suggestions from API errors can skip re-derivation in later episodes, spending fewer tokens and fewer model calls on constraints they have already learned. Server-side data drift turns those cached fixes into silent failures, and the usual…
arxiv.org16 days agoView details
RestoreBench: Can AI Agents Restore Power Flow Convergence?
arXiv:2609.00384v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use, interpretation of intermediate results, and iterative planning. Diagnosing and resolving non-convergent power flow cases is a promising yet largely unexplored appli…
arxiv.org16 days agoView details
CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
arXiv:2609.00058v1 Announce Type: new Abstract: Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier and making generating CUDA kernels directly from natural languag…
arxiv.org16 days agoView details
arXiv:2609.00086v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as rerankers in conversational recommender systems, yet measured gains depend strongly on the retrieval and inference protocol. On the ReDial conversational movie recommendation benchmark, we compare proprietary, open-we…
arxiv.org16 days agoView details
arXiv:2609.00003v1 Announce Type: new Abstract: Machine unlearning studies the removal of knowledge from an AI model, making the system forget a concept it previously learned. Despite rapid progress in generative machine unlearning, the unintended degradation of semantically related concepts that should have been reta…
arxiv.org16 days agoView details
arXiv:2609.00005v1 Announce Type: new Abstract: Financial scams targeting older adults increasingly occur through text and voice channels such as email, SMS, and phone calls, unfolding over multiple conversational turns that begin with impersonation or casual contact, escalate through trust building and urgency, and c…
arxiv.org16 days agoView details
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
arXiv:2609.00012v1 Announce Type: new Abstract: Long-horizon tasks remain uncommon in large language model (LLM) evaluation, and for a reason: when each step depends on the last, per-step accuracy that looks excellent in isolation decays catastrophically, as errors cascade and the end-to-end failure probability grows…
arxiv.org16 days agoView details
TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning
arXiv:2609.00470v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) grounds large language models in external corpora, but implicit trust in retrieved documents creates a critical attack surface: PoisonedRAG shows that a handful of crafted passages can dominate dense retrieval and steer generation tow…
arxiv.org16 days agoView details
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
arXiv:2609.00551v1 Announce Type: new Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, such fragments are not generation-ready: language mod…
arxiv.org16 days agoView details
SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning
arXiv:2609.00342v1 Announce Type: new Abstract: Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WSI models and pathology agents can aggre…
arxiv.org16 days agoView details
arXiv:2609.00028v1 Announce Type: new Abstract: Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliab…
arxiv.org16 days agoView details
mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers
arXiv:2609.00453v1 Announce Type: new Abstract: Giving an agent a file about a named expert can supply hard-to-find material, produce a recognizable persona, or change what the agent decides. These are different claims. We test each one. mimeo is an open-source tool that finds a person's public work, checks each extra…
arxiv.org16 days agoView details
EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery
arXiv:2609.00032v1 Announce Type: new Abstract: Mathematical communities work with different objects, invariants, and tools, so transferring a problem across them is expensive and often skipped. We present EULER, a multi-agent system that takes such a transfer--a bridge--as its unit of search. Around a fixed conjectur…
arxiv.org16 days agoView details
Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts
arXiv:2609.00293v1 Announce Type: new Abstract: We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in context that differs from what was stored parametrically during training. We document asymmetric biases: models tend to prefer…
arxiv.org16 days agoView details
arXiv:2609.00491v1 Announce Type: new Abstract: Communicating across cultures is inherently challenging, especially through culturally dense and ambiguous formats like memes. While people expect large language models (LLMs) to hold promise for bridging such gaps, existing benchmark datasets often fail to capture the c…
arxiv.org16 days agoView details
When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation
arXiv:2609.00071v1 Announce Type: new Abstract: Prediction error is widely used to evaluate nuisance-function estimators in causal inference, but its relationship with causal estimator performance may differ across performance measures. We studied this question in a partially linear model using Monte Carlo simulations…
arxiv.org16 days agoView details
Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation
arXiv:2609.00543v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems rely on external corpora that may contain outdated, contradictory, noisy, or unreliable documents, introducing reliability risks. Prior work has leveraged document relations to improve the answer reliability of RAG. To propaga…
arxiv.org16 days agoView details
Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations
arXiv:2609.00441v1 Announce Type: new Abstract: Effective manager-employee communication is critical for retaining high performers and developing underperformers, yet training managers in these skills remains costly. Text-based chatbots offer a scalable approach but cannot provide realistic rehearsal: managers need to…
arxiv.org16 days agoView details
Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random
arXiv:2609.00576v1 Announce Type: new Abstract: Item-sensitivity, defined as whether a model's choice depends on the specific input rather than on its own output prior, is widely reported as evidence of task competence. We show this evidence is necessary but not sufficient using a forced-choice signalling task abstrac…
arxiv.org16 days agoView details
Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents
arXiv:2609.00065v1 Announce Type: new Abstract: A language-model agent asked to analyse an experiment will usually return working code. Whether the analysis is defensible is a different question. A defensible analysis depends on procedural choices: which test the field accepts, which identifier namespace is authoritat…
arxiv.org16 days agoView details
OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization
arXiv:2609.00066v1 Announce Type: new Abstract: NVFP4 is an efficient microscaling format for low-bit inference, but activation outliers can still degrade quantization accuracy within NVFP4 blocks. Within each quantization block, large activations can dominate the block scale, increasing the quantization error of the…
arxiv.org16 days agoView details
A Stable Aggregation Method for Quantum Federated Learning
arXiv:2609.00356v1 Announce Type: new Abstract: Quantum federated learning (QFL) enables clients to train quantum neural network (QNN) models without sharing private data. We find that aggregation in QFL is unstable under heterogeneous data, unreliable communication, variable fidelity, latency, and quantum hardware no…
arxiv.org16 days agoView details
Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
arXiv:2609.00067v1 Announce Type: new Abstract: External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe w…
arxiv.org16 days agoView details
Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
arXiv:2609.00155v1 Announce Type: new Abstract: Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existing probes infer latent language from different signals, such as the geometry of hidden state…
arxiv.org16 days agoView details
arXiv:2609.00211v1 Announce Type: new Abstract: Conversational artificial intelligence is increasingly embedded in everyday social environments, where it functions as both an informational tool and a source of interpersonal feedback. This perspective introduces contingency, i.e., the degree to which system responses v…
arxiv.org16 days agoView details
arXiv:2609.00073v1 Announce Type: new Abstract: Malaria remains a significant global health burden, necessitating continuous research efforts to understand its complex molecular mechanisms, epidemiology, and potential therapeutic interventions. Extracting essential biomedical information from the vast and constantly g…
arxiv.org16 days agoView details
AI Morbidity and Mortality: A Framework for Clinical AI Failure Review
arXiv:2609.00076v1 Announce Type: new Abstract: Clinical artificial intelligence is increasingly embedded in real-world care, yet existing safety mechanisms are poorly suited to reconstructing and learning from individual AI-related errors and near-misses. Aggregate model monitoring can identify performance changes, a…
arxiv.org16 days agoView details
Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations
arXiv:2609.00106v1 Announce Type: new Abstract: This paper presents FAIRY, a full-stack smart-agriculture agent system developed for and deployed to an operating soybean research farm at Harbin Institute of Technology's smart-agriculture site. We develop FAIRY to execute and evaluate agentic agronomic operations on fu…
arxiv.org16 days agoView details
Recursive Criticality of AI Self-Improvement
arXiv:2609.00137v1 Announce Type: new Abstract: AI is increasingly used in the R\&D process that produces future AI systems. We study the conditions under which this feedback becomes self-amplifying. Our model describes how the rate of AI capability growth depends on baseline research productivity, recursive feedback,…
arxiv.org16 days agoView details
IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
arXiv:2609.00161v1 Announce Type: new Abstract: World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external represe…
arxiv.org16 days agoView details
arXiv:2609.00192v1 Announce Type: new Abstract: Public trust in Autonomous Vehicles (AVs) may depend not only on technical success but also on the fairness of their decision making. While a recent trend in AV research involves using general purpose "common sense" models to guide AV decision making, the degree to which…
arxiv.org16 days agoView details
ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation
arXiv:2609.00194v1 Announce Type: new Abstract: Document-to-slide generation is challenging because slides are dense editable artifacts that require both faithful content selection and precise spatial layout. Recent slide agents adopt iterative reflection, but typically follow a monolithic "one version, one feedback"…
arxiv.org16 days agoView details
Dependency-Aware Chain-of-Thought Compression for Financial Reasoning
arXiv:2609.00413v1 Announce Type: new Abstract: Chain of thought prompting improves complex reasoning, but its long intermediate traces create substantial inference cost and hinder practical deployment in financial settings. We present a Hierarchical Semantic Distillation Network, HSDN, for compressing reasoning chain…
arxiv.org16 days agoView details
Exploring Collaboration between a language and a non-language agent
arXiv:2609.00474v1 Announce Type: new Abstract: LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating n…
arxiv.org16 days agoView details
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
arXiv:2609.00048v1 Announce Type: new Abstract: GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use is often as multi-step environments for GUI agents. This mismatch leaves a key requirement under-tested: generated states must remain contextually consistent when they…
arxiv.org16 days agoView details
The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space
arXiv:2609.00515v1 Announce Type: new Abstract: Large language models (LLMs) have recently demonstrated improved machine translation performance over strong supervised baselines. This raises questions as to what mechanisms underlie how LLMs perform machine translation between languages. Motivated by recent interpretab…
arxiv.org16 days agoView details
ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation
arXiv:2609.00057v1 Announce Type: new Abstract: Value signals are aggregated user-level moral representations that capture users' inferred value-related tendencies from their online discourse. User behavior on social media is shaped not only by what users say or whom they interact with, but also by the value signal th…
arxiv.org16 days agoView details
Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents
arXiv:2609.00549v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly rely on external skills, yet standard evaluations obscure whether retrieving these skills actually helps. Aggregate metrics often compare retrieved versus non-retrieved tasks, introducing severe selection bias and failing to…
arxiv.org16 days agoView details
Do General NLP Embeddings Capture Ontological Reasoning?
arXiv:2609.00177v1 Announce Type: new Abstract: General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear. We introduce AVA, a systematic framework for evaluating whether embeddings distinguish logic-sensitive relational semantics…
arxiv.org16 days agoView details
Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
arXiv:2609.00184v1 Announce Type: new Abstract: Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactual edits that conflict with rigid exist…
arxiv.org16 days agoView details
The Answer Is Not the Argument
arXiv:2609.00264v1 Announce Type: new Abstract: Chain-of-thought monitoring is proposed for AI oversight, yet evaluations often provide monitors with a trusted reference answer. We ask whether answer access improves reasoning verification or mainly exposes incorrect conclusions. We collected 237 step-numbered solution…
arxiv.org16 days agoView details
Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching
arXiv:2609.00274v1 Announce Type: new Abstract: Two-sided service marketplaces are moving from deterministic request-form intake to AI-native probabilistic matching, enabled by large language models (LLMs) that infer intent, preferences, and latent constraints from natural language. Relying on inferred intent rather t…
arxiv.org16 days agoView details
arXiv:2609.00100v1 Announce Type: new Abstract: Psychometric questionnaires contain rich item-level information, yet it remains unclear whether different representation learning objectives recover the same latent organization. We investigated this question using 757 matched teacher-child pairs from the baseline assess…
arxiv.org16 days agoView details
Asymmetries in Spontaneous and Instructed Deception
arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instr…
arxiv.org16 days agoView details
LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts
arXiv:2609.00222v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as judges for subjective tasks, where annotators disagree and the relevant question is not only how accurate a judge is, but whose judgments it reproduces. Sociodemographic prompting conditions the judge on an annotator'…
arxiv.org16 days agoView details
Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing
arXiv:2609.00004v1 Announce Type: new Abstract: This paper studies a finite-horizon multi-item capacitated lot-sizing problem in which demand quantities are deterministic, while demand-arrival periods are stochastic. Each demand occurs once within a known time window and must be satisfied no later than its deadline. T…
arxiv.org16 days agoView details
arXiv:2609.00018v1 Announce Type: new Abstract: Computer science papers rely heavily on diagrams: architecture drawings, system flowcharts, and pipeline schematics that often carry more information than the text around them. There is currently no public dataset that pairs this specific kind of figure with captions, co…
arxiv.org16 days agoView details
Validity-Aware Jailbreak Evaluation for Large Language Models
arXiv:2609.00498v1 Announce Type: new Abstract: Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility rather than correctness.…
arxiv.org16 days agoView details
Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking
arXiv:2609.00228v1 Announce Type: new Abstract: Scientific domain entity linking (EL) differs from general domain EL because mentions and entity names often lack lexical overlap. Another challenge is that specialized terminology is used in the scientific domain, which is rarely encountered in models pretrained on gene…
arxiv.org16 days agoView details
LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization
arXiv:2609.00241v1 Announce Type: new Abstract: Long documents often distribute important information across extensive narrative passages and multiple tables, making faithful summarization particularly challenging. Existing methods may generate individually supported quantitative facts and analytical statements yet as…
arxiv.org16 days agoView details
Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
arXiv:2609.00330v1 Announce Type: new Abstract: In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether a live customer utterance is relevant and display a coaching card to the agent when it is. The input is noisy and challenging: ASR(Automatic Speech Recognition) transc…
arxiv.org16 days agoView details
Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
arXiv:2609.00621v1 Announce Type: new Abstract: Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and termination signals, on…
arxiv.org16 days agoView details
Location-Aware Language Models via Secondary Embeddings
arXiv:2609.00454v1 Announce Type: new Abstract: Pretrained transformer-based language models achieve strong performance across a wide range of NLP tasks but remain limited in encoding geo-locational semantics, leading to suboptimal representations of place names and spatial entities. In this work, we propose a lightwe…
arxiv.org16 days agoView details
Dr. Claw: An AI Scientist Workspace for Vibe Research
arXiv:2609.00365v1 Announce Type: new Abstract: Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, yet end-to-end research still fragments across chat tools, IDEs, terminals, and writing environments, and the decisions that make it auditable are rarel…
arxiv.org16 days agoView details
arXiv:2609.00191v1 Announce Type: new Abstract: Crisis helplines assess suicide risk through structured interviews, a process that is slow and dependent on operator training and workload. Natural language processing could support risk assessment and call prioritization, but almost no work addresses Arabic-language hel…
arxiv.org16 days agoView details
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
arXiv:2609.00487v1 Announce Type: new Abstract: Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat…
arxiv.org16 days agoView details
Towards a Belief-Based World Model for LLM Agents
arXiv:2609.00455v1 Announce Type: new Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observability. World models are a promising w…
arxiv.org16 days agoView details
EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models
arXiv:2609.00479v1 Announce Type: new Abstract: For emerging scientific research domains, local Small Language Models (SLMs) are becoming more attractive, as they offer stronger privacy control and more stable deployment pipelines than Large Language Models. However, in practice, scientific question-answering on SLMs…
arxiv.org16 days agoView details
The Privacy-Hallucination Tradeoff in Differentially Private Language Models
arXiv:2609.00492v1 Announce Type: new Abstract: Both privacy and factual accuracy are paramount in high-stakes domains like healthcare. Concerningly, we uncover and investigate a privacy-hallucination tradeoff in differentially private (DP) language models. First, we empirically show that models pre-trained or fine-tu…
arxiv.org16 days agoView details
Wave Function Backpropagation with Explicit Temporal-Interval Dynamics
arXiv:2609.00503v1 Announce Type: new Abstract: Conventional neural networks learn predominantly through affine transformations followed by nonlinear activations, while elapsed time is often treated as an auxiliary feature or assumed to be uniformly sampled. This paper introduces Wave Function Backpropagation (WFB), a…
arxiv.org16 days agoView details
arXiv:2609.00256v1 Announce Type: new Abstract: LLM-based diagnostic systems achieve high semantic accuracy on benchmarks, but open-ended evaluation on clinically uncommon presentations reveals a systematic gap between headline accuracy and verifiable clinical reliability. We evaluate an LLM+rare-disease-RAG pipeline…
arxiv.org16 days agoView details
CoVer: Conflict-Aware Claim Verification
arXiv:2609.00508v1 Announce Type: new Abstract: Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative news sources. To capture this challenge and support conflict verification tasks, we present ContraNote, a large-scale real…
arxiv.org16 days agoView details
arXiv:2609.00510v1 Announce Type: new Abstract: Artificial intelligence systems increasingly enact market-facing promises through chatbots, recommendation systems, automated decisions, and generative interfaces. Their failures, misuse, and misrepresentation raise a question that conventional brand-crisis models do not…
arxiv.org16 days agoView details
ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation
arXiv:2609.00513v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) mitigates large language models (LLMs) hallucinations, yet conventional dense retrieval struggles with the complex reasoning paths of multi-hop question answering (QA). Graph-based RAG captures multi-step relationships but suffers fro…
arxiv.org16 days agoView details
Enoki: Efficient Multi-Level Hallucination Detection
arXiv:2609.00581v1 Announce Type: new Abstract: Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucination detectors usually operate at a single level: claim-level methods provide interpretable factual units, while span-level methods localize unsupported text. B…
arxiv.org16 days agoView details
Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking
arXiv:2609.00588v1 Announce Type: new Abstract: Reranking methods, such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking, are widely used in modern neural machine translation (NMT) to select an output from a set of candidate hypotheses. However, the performance gains come at the cost of high…
arxiv.org16 days agoView details
Investigating Assistant Bias in LLM User Simulators Using a Role Vector
arXiv:2609.00608v1 Announce Type: new Abstract: LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit "assistant bias," a tendency to cooperate and pursue task goals. They rarely reproduce the frustra…
arxiv.org16 days agoView details
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
arXiv:2609.00351v1 Announce Type: new Abstract: Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect without prior knowledge what to look…
arxiv.org16 days agoView details
Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time
arXiv:2609.00624v1 Announce Type: new Abstract: A prominent paradigm in inference-time alignment employs lightweight supervisors to steer Large Language Models (LLMs). Through empirical analysis, we identify a structural mismatch in this paradigm: weak supervisors exhibit pervasive high entropy across the vast majorit…
arxiv.org16 days agoView details
Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts
arXiv:2609.00578v1 Announce Type: new Abstract: Large Language Models (LLMs) can solve complex problems, but their misuse in high-risk domains can lead to severe consequences. Model providers therefore restrict assistance for potentially harmful requests. Refusing all cybersecurity requests would therefore harm legiti…
arxiv.org16 days agoView details
arXiv:2609.00584v1 Announce Type: new Abstract: Does unrestricted AI access bypass the cognitive effort required for learning, or does it streamline knowledge acquisition? This paper reports on a study where we compare three designs for user-AI interaction in a learning context: (1) an unrestricted conversational bot…
arxiv.org16 days agoView details
REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows
arXiv:2609.00643v1 Announce Type: new Abstract: Agent revisions expose a fundamental correctness--efficiency trade-off during concurrent execution. Discarding ongoing work preserves latest-version correctness but wastes progress that may remain valid, whereas reusing prior work preserves efficiency but risks propagati…
arxiv.org16 days agoView details
Emotional Labor Strategy Preferences in LLM Personas
arXiv:2609.00310v1 Announce Type: new Abstract: Emotional labor is the effortful management of emotional displays to meet social or professional expectations. Personality traits have been correlated with emotional labor strategies, yet research on this link relies almost exclusively on self-report scales administered…
arxiv.org16 days agoView details
Visual Framing for News Stance Detection via Image Generation
arXiv:2609.00685v1 Announce Type: new Abstract: Article-level news stance detection aims to identify the perspective of news articles toward social issues. Despite advances in stance detection and its importance for trustworthy media environments, news articles pose distinct challenges because their stances are often…
arxiv.org16 days agoView details
SCoNE: Selective Context-aware Neuron Editing for Robust Retrieval-Augmented Generation
arXiv:2609.00689v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is highly sensitive to retrieval noise: when retrieved documents mix informative and irrelevant context, LLMs are easily distracted, leading to hallucinations. To overcome this, we propose SCoNE (Selective Context-aware Neuron Editing…
arxiv.org16 days agoView details
arXiv:2609.00654v1 Announce Type: new Abstract: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scientific claims against the tables and figures of a paper. Rather than tuning a single model, we benchmark eleven frontier…
arxiv.org16 days agoView details
Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets
arXiv:2609.00662v1 Announce Type: new Abstract: A multi-model language service must route each request while preserving workload-level budgets for compute, latency, memory, or monetary cost. Two features make this problem materially harder than static model selection. Prompt representations are high dimensional, so on…
arxiv.org16 days agoView details
Creative Generation via Multi-Agent Debate: Does Debate Suppress Diversity?
arXiv:2609.00683v1 Announce Type: new Abstract: Creative generation tasks, such as narrative writing and scientific ideation, demand both high-quality outputs and distinct responses across independent runs to maximize exploration. Multi-Agent Debate (MAD) has shown strong quality gains on factual and reasoning tasks,…
arxiv.org16 days agoView details
arXiv:2609.00550v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly provided with contextual evidence in heterogeneous forms: as a text passage, as a rendered image of the same passage, or as both together. However, it remains unclear how consistently these surface forms are proce…
arxiv.org16 days agoView details
Authority Bias in Conversational Search Engines for Academic Paper Recommendation
arXiv:2609.00248v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used as conversational search engines for academic literature, yet whether they judge papers on content or on authority signals has not been tested causally. We investigate authority bias: systematic preference for papers bas…
arxiv.org16 days agoView details
Hypotheses-Guided Self Distillation for Continual Personalization
arXiv:2609.00251v1 Announce Type: new Abstract: As people increasingly interact with LLM assistants in daily life, continually adapting to individual preferences has become essential for effective long-term interactions. However, user preferences are rarely stated in full, and instead emerge through heterogeneous, lat…
arxiv.org16 days agoView details
Life Operators: a self-evolving framework for multiscale life modelling
arXiv:2609.00068v1 Announce Type: new Abstract: Medical AI is moving beyond recognition towards clinical dialogue and longitudinal prediction. Yet a central question remains: how would a patient's state change under intervention? Statistical models learn future observations, whereas mechanistic models describe selecte…
arxiv.org16 days agoView details
arXiv:2609.00355v1 Announce Type: new Abstract: Speculative decoding accelerates generation without changing its output, yet on vision-language models (VLMs) it has been caught in a self-defeating cycle. The drafter stays autoregressive, so it must stay small. A small drafter cannot afford the image at every step, so…
arxiv.org16 days agoView details
SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation
arXiv:2609.00427v1 Announce Type: new Abstract: The exponential growth of wireless devices is driving unprecedented spectrum demand, pushing spectrum management toward more fine-grained decisions across space, time, and device constraints. As a result, spectrum policymakers and engineers must process large volumes of…
arxiv.org16 days agoView details
Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
arXiv:2609.00055v1 Announce Type: new Abstract: Self-supervised respiratory encoders lack semantic grounding in clinical domain needed for zero-shot inference, limiting their utility without task-specific labeled data. We propose a framework that aligns these encoders with medical terminology in a shared latent space…
arxiv.org16 days agoView details
KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training
arXiv:2609.00082v1 Announce Type: new Abstract: LLMs acquire vast amounts of knowledge during pre-training, but often lack the specialized knowledge needed to answer questions from niche sources such as manuals or technical documents unseen during pre-training. Continued pre-training (CPT) is widely used to inject suc…
arxiv.org16 days agoView details
arXiv:2609.00275v1 Announce Type: new Abstract: Fleets of LLM agents now externalize effects that cannot be fully undone: they move money, deploy code, delete data, and disclose information. Current controls check one effect at a time, so a fleet of individually authorized agents can overdraw its principal's risk unde…
arxiv.org16 days agoView details
Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective
arXiv:2609.00334v1 Announce Type: new Abstract: Across law, education, policy analysis, and public moral argumentation, LLM outputs are being used often for work that requires interpretations to be justified with textual evidence and explicit normative standards. Yet a recurrent failure mode -- what I call \textit{int…
arxiv.org16 days agoView details
Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
arXiv:2609.00482v1 Announce Type: new Abstract: Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approxi…
arxiv.org16 days agoView details
Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models
arXiv:2609.00495v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. This generation paradigm introduces two axes that can influence safety alignment: when tokens are generated during denoising and where they appear in the…
arxiv.org16 days agoView details
arXiv:2609.00652v1 Announce Type: new Abstract: Language model agents increasingly propose actions, observe external feedback, and explain their own behavior. Their confidence and rationales are convenient monitoring signals, but convenience is not verification. We introduce an environment-grounded audit in which ever…
arxiv.org16 days agoView details
The Curse of Multilinguality in Lexical Normalization
arXiv:2609.00329v1 Announce Type: new Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages at once. We ask a simp…
arxiv.org16 days agoView details
Models
View all models →text-generation · transformers · safetensors · spark2_5
huggingface.co16 days ago1247 ptsView details
Motif-Technologies/Motif-3-Beta
text-generation · transformers · safetensors · Motif
huggingface.co17 days ago218 ptsView details
gguf · endpoints_compatible · region:us
huggingface.co17 days ago92 ptsView details
Motif-Technologies/Motif-3-NVFP4
text-generation · transformers · safetensors · Motif
huggingface.co17 days ago29 ptsView details
Gramscii-IT/SemanticRepair-270M
text-generation · mlx · safetensors · gguf
huggingface.co17 days ago2 ptsView details
text-generation · transformers · safetensors · qwen3_5_text
huggingface.co17 days agoView details
VikramPal/Qwen3.5-35B-A3B-bf16
text-generation · transformers · safetensors · qwen3_5_moe_text
huggingface.co17 days agoView details
Open source
View all open source →shtjww/llm-inference-capacity-handbook
大模型推理案头手册(开源版):给定 GPU 算力,一个模型能扛多少 QPS?三层模型 × 三面墙 × 排队论 × 开环压测
github.com16 days ago20 ptsView details
## [3.7.0](https://github.com/openai/openai-python/compare/v3.6.0...v3.7.0) (2026-09-02) ### Features * **api:** update usage APIs and documentation ([#3779](https://github.com/openai/openai-python/issues/3779)) ([6f0da16](https://github.com/openai/openai-python/commit/6f0da1657…
github.com16 days agoView details
anthropics/anthropic-sdk-python v1.3.0
## 1.3.0 (2026-09-01) Full Changelog: [v1.2.0...v1.3.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.2.0...v1.3.0) ### Features * **api:** beta user profiles: add external_user_onboarded_at, remove relationship in favor of access_type ([74080c3](https://github.c…
github.com16 days agoView details
langchain-ai/langchain langchain==1.4.0a3
Third alpha of the `1.4.0` line. This release focuses on the new `langchain.mcp` namespace for adapting MCP servers into LangChain tools. ## `langchain.mcp` highlights - **`MCPAdapter`** adapts any target `fastmcp.Client` accepts — a URL, a local script, an in-process server, an…
github.com16 days agoView details
<details open> qwen4exp: support recurrent state rollback (#28123) MTP speculative decoding needs the target state to move back by the number of rejected draft tokens. Without rollback support the context is classified as SEQ_RM_TYPE_FULL and the server serializes the whole recu…
github.com17 days agoView details
<details open> qwen4exp: sum the indexer heads by slices (#28023) * qwen4exp: sum the indexer heads by slices The head reduction went through a transpose and a sum_rows over ne[1], which left sum_rows with ne0 = 4, one block per row for a four element reduction, and the transpos…
github.com17 days agoView details