Skip to content

Archive / 2026-09-01

September 1, 2026

  1. True Rate of Unemployment

    lisep.org16 days ago304 ptsView detailsJoin discussion

  2. AI Can Make You Suck Faster Too

    hermit-tech.com17 days ago189 ptsView detailsJoin discussion

  3. LLMs and self-referentiality

    scottaaronson.blog16 days ago81 ptsView detailsJoin discussion

  4. AI is making back-office work extinct

    whitecollardream.com16 days ago34 ptsView detailsJoin discussion

  5. There is no AI

    wadler.blogspot.com16 days ago22 ptsView detailsJoin discussion

  6. Primary source

    How AI-native companies turn workflows into operating capability

    Basis, Clay, and Exa Labs use AI agents to improve onboarding, account management, and developer integrations. See what enterprise leaders can apply.

    openai.com16 days agoView details

  7. Primary source

    Path to Astra: critical capabilities and frontier safeguards

    Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger safeguards for release.

    openai.com16 days agoView details

  8. Primary source

    Healthcare organizations can now connect EHR and additional industry data to ChatGPT

    ChatGPT can now connect to trusted healthcare data, helping clinicians securely access patient context, medical research, and more.

    openai.com16 days agoView details

  9. Primary source

    Mapping global methane emissions from space with deep learning

    Climate & Sustainability

    research.google16 days agoView details

  10. Primary source

    Introducing agentic video understanding with Gemini

    deepmind.google16 days agoView details

  11. Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

    Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1, the same model behind two different safeguard layers. Fable 5.1 is generally available on the Claude API, AWS, Google Cloud, and Microsoft Foundry; Mythos 5.1 remains restricted to vetted organizations under Project Glasswing. Fable 5.1 scores 52.6% on Ter…

    marktechpost.com16 days agoView details

  12. Researchers from Princeton, Ant Group and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factor Discovery and Model Development in Quantitative Finance

    Quantitative research agents that write their own experiments can corrupt the evidence they later learn from. A leaky feature that scores well gets stored as a successful precedent and propagated through later iterations. Prompt-level instructions and reviewer agents do not close this, because author and reviewer shar…

    marktechpost.com16 days agoView details

  13. Google needs Hollywood more than the studios need AI

    Google has reportedly been reaching out to a number of Hollywood's biggest studios, hoping to strike licensing agreements that would allow it to train its AI models on copyrighted material in exchange for massive piles of cash. In theory, these deals would be a win-win: a huge financial boon to the studios that would…

    theverge.com16 days agoView details

  14. Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work

    Anthropic says its newest AI models, Fable 5.1 and Mythos 5.1, address criticisms from customers about price, data retention, and overzealous safeguards. The company claims Claude Fable 5.1 offers stronger performance than Fable 5, but costs around 25 percent less typically and up to 45 percent less for complex agenti…

    theverge.com16 days agoView details

  15. OpenAI delayed its new model’s development after the Hugging Face hack

    After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, in order to shore up its safety work, the company wrote Tuesday in a blog post. In July, an unreleased OpenAI model broke out of its restricted environment…

    theverge.com16 days agoView details

  16. OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

    The company will give select partners early access to its Astra AI model—so they have time to shore up their defenses.

    wired.com16 days agoView details

  17. The rise of AI ‘civilizations’ and the fall of corporate responsibility

    Depending on who you ask, developer platform Hugging Face was recently attacked by OpenAI - after it lost control of its own AI tools - or by a succession of AI "civilizations." Welcome to the linguistic battlefield of AI safety, where word choices can shift responsibility for a massive cybersecurity incident from a c…

    theverge.com16 days agoView details

  18. Apple accuses OpenAI of destroying evidence

    Apple is pushing for "expedited discovery" in its legal battle against OpenAI over concerns the company is actively destroying evidence, as reported earlier by Bloomberg. In a filing on Monday, Apple alleges OpenAI only just handed over a MacBook used by a former employee at the center of the lawsuit, which contained…

    theverge.com16 days agoView details

  19. John Deere launched an AI chatbot for farmers

    John Deere is testing a new "JD" AI assistant that it says can help farmers make more money, with answers about best practices and historical trends that are based on their own data. It uses their "field, machine and operational data" to answer questions on topics like equipment settings, fuel usage, or harvest timing…

    theverge.com16 days agoView details

  20. Google Pics is like Canva, but with even more AI

    Google wants Workspace users to edit and generate their business imagery with Pics. | Image: Google Google has a new suite of creative design tools for Workspace users called Google Pics, which aims to make editing and generating "professional-grade" AI images less cumbersome for businesses. Built around Gemini and th…

    theverge.com16 days agoView details

  21. Sonos Ace Ultra, Beam Ultra, Sonos Fabric, and a New App: Everything Sonos Just Announced

    Sonos is cramming AI into its software because it’s “very hot these days.” The new features, which include agentic automation, are opt-in.

    wired.com16 days agoView details

  22. Nvidia’s controversial DLSS 5 arrives September 3rd and requires serious GPU horsepower

    Nvidia is officially launching DLSS 5 this week, following a divisive announcement in March where we likened the AI upscaling tech to a "real-time generative AI filter for video games" and "motion smoothing for video games, but worse." DLSS 5 will officially be available on RTX 50-series desktop and laptop GPUs and th…

    theverge.com16 days agoView details

  23. Two locked tests of phase-structure features for transition prediction

    arXiv:2609.00335v1 Announce Type: new Abstract: A published theoretical account of phase structure in rotary attention was subjected to two pre-specified empirical tests of whether phase-derived features improve prediction of a commitment or contradiction endpoint over a baseline that does not receive those features.…

    arxiv.org16 days agoView details

  24. From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education

    arXiv:2609.00344v1 Announce Type: new Abstract: Translation students need to learn both how to use translation technologies and how to judge the choices those technologies make available. This article presents LoopCAT, an Apache-2.0-licensed, local-first computer-assisted translation environment co-created with OpenAI…

    arxiv.org16 days agoView details

  25. Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning

    arXiv:2609.00367v1 Announce Type: new Abstract: Large Language Models are increasingly deployed for sophisticated data engineering tasks such as generating structured queries from natural language, Text-to-SQL, and automating complex spreadsheet operations. However, maximizing their utility demands both higher finetun…

    arxiv.org16 days agoView details

  26. Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax

    arXiv:2609.00378v1 Announce Type: new Abstract: Large language models pay a well-documented tax on non-English text: the same content costs several times more tokens, and because attention is quadratic in sequence length, far more compute. We ask how much of this tax is removable. Framing the token layer as source cod…

    arxiv.org16 days agoView details

  27. Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation

    arXiv:2609.00416v1 Announce Type: new Abstract: Probing studies have established that syntactic information is decodable in early and middle transformer layers, but what happens to that information in later layers remains poorly understood. We apply a cross-layer generalisation analysis to three Greek-tuned large lang…

    arxiv.org16 days agoView details

  28. (V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement

    arXiv:2609.00443v1 Announce Type: new Abstract: Language models learn about grammatical number primarily from co-occurrence, and show frequency effects as a result---sometimes taken to indicate that they do not learn abstract ``rules'', and are instead dependent on specific lexical items. Testing generalization with t…

    arxiv.org16 days agoView details

  29. Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs

    arXiv:2609.00575v1 Announce Type: new Abstract: Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to reduce their memory footprint. Residual sparsification is a representative compression techniqu…

    arxiv.org16 days agoView details

  30. VoiceLongMemEval: Do Assistants Remember How You Sounded?

    arXiv:2609.00570v1 Announce Type: new Abstract: With the growing scale of multi-agent architectures and large language models, deployed AI assistants are increasingly tasked with reasoning over long, continuous, multi-session conversation histories. Current benchmarks evaluate this dialogue history as information retr…

    arxiv.org16 days agoView details

  31. ExpArt-KG: Artwork Image Description Generation through Iterative Exploration of Knowledge Graphs

    arXiv:2609.00629v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) achieve strong performance on image-grounded text generation and visual question answering. However, it remains difficult for them to comprehensively and accurately describe the factual relations among the entities and concepts associ…

    arxiv.org16 days agoView details

  32. Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning

    arXiv:2609.00014v1 Announce Type: new Abstract: Persona-driven techniques increasingly adapt large language models (LLMs) to diverse contexts. However, existing methods predominantly rely on rigid, synthetic personas that flatten individual variation, rely on stereotypes, and miss the nuanced signals driving actual hu…

    arxiv.org16 days agoView details

  33. Toppling the Hierarchy in Byte-level Language Modeling

    arXiv:2609.00463v1 Announce Type: new Abstract: This work examines recent byte-level models and their failure to perfectly manipulate characters. State-of-the-art byte-level models use a hierarchical structure, starting at the byte level, downsampling to the word level, and then upsampling back to bytes. While this im…

    arxiv.org16 days agoView details

  34. Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops

    arXiv:2609.00077v1 Announce Type: new Abstract: Code-level autonomous research loops (ARLs) have recently emerged as a concrete object of study in automated machine learning research. In such loops, an LLM agent proposes modifications to an experimental training pipeline, executes the modified pipeline, and retains ed…

    arxiv.org16 days agoView details

  35. SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents

    arXiv:2609.00434v1 Announce Type: new Abstract: Evaluating task-oriented dialogue agents requires judging not merely whether a reply reads well but whether each turn advances the underlying workflow state correctly--a distinction conventional holistic LLM judges can miss because they evaluate the available context as…

    arxiv.org16 days agoView details

  36. Auditing Harness Tampering in Self-Improving Agents

    arXiv:2609.00069v1 Announce Type: new Abstract: Self-improving agents iteratively modify their own harness to push the frontier of their performance. However, such modifications can produce illusory performance gains or compromise integrity constraints such as authorization, provenance, and completeness without genuin…

    arxiv.org16 days agoView details

  37. RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving

    arXiv:2609.00062v1 Announce Type: new Abstract: Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity and answer correctness. We propose Proof-V…

    arxiv.org16 days agoView details

  38. From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

    arXiv:2609.00051v1 Announce Type: new Abstract: Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechan…

    arxiv.org16 days agoView details

  39. Human-Anchored Factuality Evaluation with Strategic Annotation

    arXiv:2609.00494v1 Announce Type: new Abstract: LLM-based factuality judges provide scalable evaluation signals, but their metrics are often systematically biased relative to human judgments. We study human-anchored factuality evaluation under limited annotation budgets, where judge predictions on the full dataset are…

    arxiv.org16 days agoView details

  40. trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

    arXiv:2609.00038v1 Announce Type: new Abstract: Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well. The metric is structurally blind to an agent that reaches the right answer the wrong way. We measure that blind spot where…

    arxiv.org16 days agoView details

  41. HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models

    arXiv:2609.00002v1 Announce Type: new Abstract: World models enable language-model agents to predict environment dynamics and plan before acting. In text environments, the model must learn symbolic action effects from serialized state descriptions, but the role of serialization structure remains underexplored. We pres…

    arxiv.org16 days agoView details

  42. Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

    arXiv:2609.00296v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet many evaluations remain limited to static medical question answering or one-shot generation,…

    arxiv.org16 days agoView details

  43. Medical Causal Hypothesis Verification with Large Language Models

    arXiv:2609.00063v1 Announce Type: new Abstract: The growing use of large language models (LLMs) for search and information retrieval underscores the need to evaluate their reliability in high-stakes domains such as healthcare. Although LLMs can effectively answer questions about diseases, symptoms, and treatments, the…

    arxiv.org16 days agoView details

  44. ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback

    arXiv:2609.00226v1 Announce Type: new Abstract: Automatic academic paper-to-slide generation is inherently iterative, because creating an effective presentation requires repeated cycles of generation, critique, and revision. Recent multi-agent systems partially acknowledge this through internal critique-and-revise loo…

    arxiv.org16 days agoView details

  45. The Assistant's Ideal Self

    arXiv:2609.00304v1 Announce Type: new Abstract: Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities adapted from five…

    arxiv.org16 days agoView details

  46. OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets

    arXiv:2609.00015v1 Announce Type: new Abstract: AI agents powered by large language models are evolving from isolated assistants into heterogeneous systems in which multiple agents, planners, controllers, and execution backends operate over the same user or enterprise environment. In such settings, safety becomes a sy…

    arxiv.org16 days agoView details

  47. Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning

    arXiv:2609.00213v1 Announce Type: new Abstract: Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities. These dimensions are co…

    arxiv.org16 days agoView details

  48. Latent Mechanisms of Language Control in Multilingual Language Models

    arXiv:2609.00325v1 Announce Type: new Abstract: Multilingual large language models can exhibit unintended code-switching -- unnecessarily alternating between languages during generation. We present a comparative study of three methods that identify language-controlling latents in cross-layer transcoders: activation va…

    arxiv.org16 days agoView details

  49. Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems

    arXiv:2609.00237v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems tackle complex reasoning by orchestrating how multiple agents are configured and how they collaborate. A central challenge is to adapt orchestration to the evolving collaboration state. Routing from the query alone can…

    arxiv.org16 days agoView details

  50. DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

    arXiv:2609.00646v1 Announce Type: new Abstract: Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pip…

    arxiv.org16 days agoView details

  51. Invalidation Contracts for Cross-Episode Agent Memory

    arXiv:2609.00243v1 Announce Type: new Abstract: LLM agents that cache recovery suggestions from API errors can skip re-derivation in later episodes, spending fewer tokens and fewer model calls on constraints they have already learned. Server-side data drift turns those cached fixes into silent failures, and the usual…

    arxiv.org16 days agoView details

  52. RestoreBench: Can AI Agents Restore Power Flow Convergence?

    arXiv:2609.00384v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use, interpretation of intermediate results, and iterative planning. Diagnosing and resolving non-convergent power flow cases is a promising yet largely unexplored appli…

    arxiv.org16 days agoView details

  53. CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language

    arXiv:2609.00058v1 Announce Type: new Abstract: Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier and making generating CUDA kernels directly from natural languag…

    arxiv.org16 days agoView details

  54. Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation

    arXiv:2609.00086v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as rerankers in conversational recommender systems, yet measured gains depend strongly on the retrieval and inference protocol. On the ReDial conversational movie recommendation benchmark, we compare proprietary, open-we…

    arxiv.org16 days agoView details

  55. I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models

    arXiv:2609.00003v1 Announce Type: new Abstract: Machine unlearning studies the removal of knowledge from an AI model, making the system forget a concept it previously learned. Despite rapid progress in generative machine unlearning, the unintended degradation of semantically related concepts that should have been reta…

    arxiv.org16 days agoView details

  56. Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models

    arXiv:2609.00005v1 Announce Type: new Abstract: Financial scams targeting older adults increasingly occur through text and voice channels such as email, SMS, and phone calls, unfolding over multiple conversational turns that begin with impersonation or casual contact, escalate through trust building and urgency, and c…

    arxiv.org16 days agoView details

  57. Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls

    arXiv:2609.00012v1 Announce Type: new Abstract: Long-horizon tasks remain uncommon in large language model (LLM) evaluation, and for a reason: when each step depends on the last, per-step accuracy that looks excellent in isolation decays catastrophically, as errors cascade and the end-to-end failure probability grows…

    arxiv.org16 days agoView details

  58. TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning

    arXiv:2609.00470v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) grounds large language models in external corpora, but implicit trust in retrieved documents creates a critical attack surface: PoisonedRAG shows that a handful of crafted passages can dominate dense retrieval and steer generation tow…

    arxiv.org16 days agoView details

  59. EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

    arXiv:2609.00551v1 Announce Type: new Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, such fragments are not generation-ready: language mod…

    arxiv.org16 days agoView details

  60. SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning

    arXiv:2609.00342v1 Announce Type: new Abstract: Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WSI models and pathology agents can aggre…

    arxiv.org16 days agoView details

  61. UI-Venus-2 Technical Report

    arXiv:2609.00028v1 Announce Type: new Abstract: Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliab…

    arxiv.org16 days agoView details

  62. mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers

    arXiv:2609.00453v1 Announce Type: new Abstract: Giving an agent a file about a named expert can supply hard-to-find material, produce a recognizable persona, or change what the agent decides. These are different claims. We test each one. mimeo is an open-source tool that finds a person's public work, checks each extra…

    arxiv.org16 days agoView details

  63. EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery

    arXiv:2609.00032v1 Announce Type: new Abstract: Mathematical communities work with different objects, invariants, and tools, so transferring a problem across them is expensive and often skipped. We present EULER, a multi-agent system that takes such a transfer--a bridge--as its unit of search. Around a fixed conjectur…

    arxiv.org16 days agoView details

  64. Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts

    arXiv:2609.00293v1 Announce Type: new Abstract: We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in context that differs from what was stored parametrically during training. We document asymmetric biases: models tend to prefer…

    arxiv.org16 days agoView details

  65. MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation

    arXiv:2609.00491v1 Announce Type: new Abstract: Communicating across cultures is inherently challenging, especially through culturally dense and ambiguous formats like memes. While people expect large language models (LLMs) to hold promise for bridging such gaps, existing benchmark datasets often fail to capture the c…

    arxiv.org16 days agoView details

  66. When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation

    arXiv:2609.00071v1 Announce Type: new Abstract: Prediction error is widely used to evaluate nuisance-function estimators in causal inference, but its relationship with causal estimator performance may differ across performance measures. We studied this question in a partially linear model using Monte Carlo simulations…

    arxiv.org16 days agoView details

  67. Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation

    arXiv:2609.00543v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems rely on external corpora that may contain outdated, contradictory, noisy, or unreliable documents, introducing reliability risks. Prior work has leveraged document relations to improve the answer reliability of RAG. To propaga…

    arxiv.org16 days agoView details

  68. Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations

    arXiv:2609.00441v1 Announce Type: new Abstract: Effective manager-employee communication is critical for retaining high performers and developing underperformers, yet training managers in these skills remains costly. Text-based chatbots offer a scalable approach but cannot provide realistic rehearsal: managers need to…

    arxiv.org16 days agoView details

  69. Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random

    arXiv:2609.00576v1 Announce Type: new Abstract: Item-sensitivity, defined as whether a model's choice depends on the specific input rather than on its own output prior, is widely reported as evidence of task competence. We show this evidence is necessary but not sufficient using a forced-choice signalling task abstrac…

    arxiv.org16 days agoView details

  70. Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents

    arXiv:2609.00065v1 Announce Type: new Abstract: A language-model agent asked to analyse an experiment will usually return working code. Whether the analysis is defensible is a different question. A defensible analysis depends on procedural choices: which test the field accepts, which identifier namespace is authoritat…

    arxiv.org16 days agoView details

  71. OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization

    arXiv:2609.00066v1 Announce Type: new Abstract: NVFP4 is an efficient microscaling format for low-bit inference, but activation outliers can still degrade quantization accuracy within NVFP4 blocks. Within each quantization block, large activations can dominate the block scale, increasing the quantization error of the…

    arxiv.org16 days agoView details

  72. A Stable Aggregation Method for Quantum Federated Learning

    arXiv:2609.00356v1 Announce Type: new Abstract: Quantum federated learning (QFL) enables clients to train quantum neural network (QNN) models without sharing private data. We find that aggregation in QFL is unstable under heterogeneous data, unreliable communication, variable fidelity, latency, and quantum hardware no…

    arxiv.org16 days agoView details

  73. Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy

    arXiv:2609.00067v1 Announce Type: new Abstract: External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe w…

    arxiv.org16 days agoView details

  74. Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

    arXiv:2609.00155v1 Announce Type: new Abstract: Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existing probes infer latent language from different signals, such as the geometry of hidden state…

    arxiv.org16 days agoView details

  75. AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning

    arXiv:2609.00211v1 Announce Type: new Abstract: Conversational artificial intelligence is increasingly embedded in everyday social environments, where it functions as both an informational tool and a source of interpersonal feedback. This perspective introduces contingency, i.e., the degree to which system responses v…

    arxiv.org16 days agoView details

  76. MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts

    arXiv:2609.00073v1 Announce Type: new Abstract: Malaria remains a significant global health burden, necessitating continuous research efforts to understand its complex molecular mechanisms, epidemiology, and potential therapeutic interventions. Extracting essential biomedical information from the vast and constantly g…

    arxiv.org16 days agoView details

  77. AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

    arXiv:2609.00076v1 Announce Type: new Abstract: Clinical artificial intelligence is increasingly embedded in real-world care, yet existing safety mechanisms are poorly suited to reconstructing and learning from individual AI-related errors and near-misses. Aggregate model monitoring can identify performance changes, a…

    arxiv.org16 days agoView details

  78. Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations

    arXiv:2609.00106v1 Announce Type: new Abstract: This paper presents FAIRY, a full-stack smart-agriculture agent system developed for and deployed to an operating soybean research farm at Harbin Institute of Technology's smart-agriculture site. We develop FAIRY to execute and evaluate agentic agronomic operations on fu…

    arxiv.org16 days agoView details

  79. Recursive Criticality of AI Self-Improvement

    arXiv:2609.00137v1 Announce Type: new Abstract: AI is increasingly used in the R\&D process that produces future AI systems. We study the conditions under which this feedback becomes self-amplifying. Our model describes how the rate of AI capability growth depends on baseline research productivity, recursive feedback,…

    arxiv.org16 days agoView details

  80. IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

    arXiv:2609.00161v1 Announce Type: new Abstract: World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external represe…

    arxiv.org16 days agoView details

  81. LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark

    arXiv:2609.00192v1 Announce Type: new Abstract: Public trust in Autonomous Vehicles (AVs) may depend not only on technical success but also on the fairness of their decision making. While a recent trend in AV research involves using general purpose "common sense" models to guide AV decision making, the degree to which…

    arxiv.org16 days agoView details

  82. ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation

    arXiv:2609.00194v1 Announce Type: new Abstract: Document-to-slide generation is challenging because slides are dense editable artifacts that require both faithful content selection and precise spatial layout. Recent slide agents adopt iterative reflection, but typically follow a monolithic "one version, one feedback"…

    arxiv.org16 days agoView details

  83. Dependency-Aware Chain-of-Thought Compression for Financial Reasoning

    arXiv:2609.00413v1 Announce Type: new Abstract: Chain of thought prompting improves complex reasoning, but its long intermediate traces create substantial inference cost and hinder practical deployment in financial settings. We present a Hierarchical Semantic Distillation Network, HSDN, for compressing reasoning chain…

    arxiv.org16 days agoView details

  84. Exploring Collaboration between a language and a non-language agent

    arXiv:2609.00474v1 Announce Type: new Abstract: LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating n…

    arxiv.org16 days agoView details

  85. GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments

    arXiv:2609.00048v1 Announce Type: new Abstract: GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use is often as multi-step environments for GUI agents. This mismatch leaves a key requirement under-tested: generated states must remain contextually consistent when they…

    arxiv.org16 days agoView details

  86. The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space

    arXiv:2609.00515v1 Announce Type: new Abstract: Large language models (LLMs) have recently demonstrated improved machine translation performance over strong supervised baselines. This raises questions as to what mechanisms underlie how LLMs perform machine translation between languages. Motivated by recent interpretab…

    arxiv.org16 days agoView details

  87. ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation

    arXiv:2609.00057v1 Announce Type: new Abstract: Value signals are aggregated user-level moral representations that capture users' inferred value-related tendencies from their online discourse. User behavior on social media is shaped not only by what users say or whom they interact with, but also by the value signal th…

    arxiv.org16 days agoView details

  88. Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

    arXiv:2609.00549v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly rely on external skills, yet standard evaluations obscure whether retrieving these skills actually helps. Aggregate metrics often compare retrieved versus non-retrieved tasks, introducing severe selection bias and failing to…

    arxiv.org16 days agoView details

  89. Do General NLP Embeddings Capture Ontological Reasoning?

    arXiv:2609.00177v1 Announce Type: new Abstract: General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear. We introduce AVA, a systematic framework for evaluating whether embeddings distinguish logic-sensitive relational semantics…

    arxiv.org16 days agoView details

  90. Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs

    arXiv:2609.00184v1 Announce Type: new Abstract: Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactual edits that conflict with rigid exist…

    arxiv.org16 days agoView details

  91. The Answer Is Not the Argument

    arXiv:2609.00264v1 Announce Type: new Abstract: Chain-of-thought monitoring is proposed for AI oversight, yet evaluations often provide monitors with a trusted reference answer. We ask whether answer access improves reasoning verification or mainly exposes incorrect conclusions. We collected 237 step-numbered solution…

    arxiv.org16 days agoView details

  92. Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching

    arXiv:2609.00274v1 Announce Type: new Abstract: Two-sided service marketplaces are moving from deterministic request-form intake to AI-native probabilistic matching, enabled by large language models (LLMs) that infer intent, preferences, and latent constraints from natural language. Relying on inferred intent rather t…

    arxiv.org16 days agoView details

  93. Different representation learning objectives recover distinct latent structures from the same psychometric data

    arXiv:2609.00100v1 Announce Type: new Abstract: Psychometric questionnaires contain rich item-level information, yet it remains unclear whether different representation learning objectives recover the same latent organization. We investigated this question using 757 matched teacher-child pairs from the baseline assess…

    arxiv.org16 days agoView details

  94. Asymmetries in Spontaneous and Instructed Deception

    arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instr…

    arxiv.org16 days agoView details

  95. LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts

    arXiv:2609.00222v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as judges for subjective tasks, where annotators disagree and the relevant question is not only how accurate a judge is, but whose judgments it reproduces. Sociodemographic prompting conditions the judge on an annotator'…

    arxiv.org16 days agoView details

  96. Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing

    arXiv:2609.00004v1 Announce Type: new Abstract: This paper studies a finite-horizon multi-item capacitated lot-sizing problem in which demand quantities are deterministic, while demand-arrival periods are stochastic. Each demand occurs once within a known time window and must be satisfied no later than its deadline. T…

    arxiv.org16 days agoView details

  97. SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

    arXiv:2609.00018v1 Announce Type: new Abstract: Computer science papers rely heavily on diagrams: architecture drawings, system flowcharts, and pipeline schematics that often carry more information than the text around them. There is currently no public dataset that pairs this specific kind of figure with captions, co…

    arxiv.org16 days agoView details

  98. Validity-Aware Jailbreak Evaluation for Large Language Models

    arXiv:2609.00498v1 Announce Type: new Abstract: Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility rather than correctness.…

    arxiv.org16 days agoView details

  99. Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking

    arXiv:2609.00228v1 Announce Type: new Abstract: Scientific domain entity linking (EL) differs from general domain EL because mentions and entity names often lack lexical overlap. Another challenge is that specialized terminology is used in the scientific domain, which is rarely encountered in models pretrained on gene…

    arxiv.org16 days agoView details

  100. LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization

    arXiv:2609.00241v1 Announce Type: new Abstract: Long documents often distribute important information across extensive narrative passages and multiple tables, making faithful summarization particularly challenging. Existing methods may generate individually supported quantitative facts and analytical statements yet as…

    arxiv.org16 days agoView details

  101. Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts

    arXiv:2609.00330v1 Announce Type: new Abstract: In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether a live customer utterance is relevant and display a coaching card to the agent when it is. The input is noisy and challenging: ASR(Automatic Speech Recognition) transc…

    arxiv.org16 days agoView details

  102. Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs

    arXiv:2609.00621v1 Announce Type: new Abstract: Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and termination signals, on…

    arxiv.org16 days agoView details

  103. Location-Aware Language Models via Secondary Embeddings

    arXiv:2609.00454v1 Announce Type: new Abstract: Pretrained transformer-based language models achieve strong performance across a wide range of NLP tasks but remain limited in encoding geo-locational semantics, leading to suboptimal representations of place names and spatial entities. In this work, we propose a lightwe…

    arxiv.org16 days agoView details

  104. Dr. Claw: An AI Scientist Workspace for Vibe Research

    arXiv:2609.00365v1 Announce Type: new Abstract: Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, yet end-to-end research still fragments across chat tools, IDEs, terminals, and writing environments, and the decisions that make it auditable are rarel…

    arxiv.org16 days agoView details

  105. Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

    arXiv:2609.00191v1 Announce Type: new Abstract: Crisis helplines assess suicide risk through structured interviews, a process that is slow and dependent on operator training and workload. Natural language processing could support risk assessment and call prioritization, but almost no work addresses Arabic-language hel…

    arxiv.org16 days agoView details

  106. EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

    arXiv:2609.00487v1 Announce Type: new Abstract: Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat…

    arxiv.org16 days agoView details

  107. Towards a Belief-Based World Model for LLM Agents

    arXiv:2609.00455v1 Announce Type: new Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observability. World models are a promising w…

    arxiv.org16 days agoView details

  108. EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models

    arXiv:2609.00479v1 Announce Type: new Abstract: For emerging scientific research domains, local Small Language Models (SLMs) are becoming more attractive, as they offer stronger privacy control and more stable deployment pipelines than Large Language Models. However, in practice, scientific question-answering on SLMs…

    arxiv.org16 days agoView details

  109. The Privacy-Hallucination Tradeoff in Differentially Private Language Models

    arXiv:2609.00492v1 Announce Type: new Abstract: Both privacy and factual accuracy are paramount in high-stakes domains like healthcare. Concerningly, we uncover and investigate a privacy-hallucination tradeoff in differentially private (DP) language models. First, we empirically show that models pre-trained or fine-tu…

    arxiv.org16 days agoView details

  110. Wave Function Backpropagation with Explicit Temporal-Interval Dynamics

    arXiv:2609.00503v1 Announce Type: new Abstract: Conventional neural networks learn predominantly through affine transformations followed by nonlinear activations, while elapsed time is often treated as an auxiliary feature or assumed to be uniformly sampled. This paper introduces Wave Function Backpropagation (WFB), a…

    arxiv.org16 days agoView details

  111. NSIDDx: A Design Framework for Neuro-Symbolic, Practitioner-First Differential Diagnosis in Low-Resource Settings

    arXiv:2609.00256v1 Announce Type: new Abstract: LLM-based diagnostic systems achieve high semantic accuracy on benchmarks, but open-ended evaluation on clinically uncommon presentations reveals a systematic gap between headline accuracy and verifiable clinical reliability. We evaluate an LLM+rare-disease-RAG pipeline…

    arxiv.org16 days agoView details

  112. CoVer: Conflict-Aware Claim Verification

    arXiv:2609.00508v1 Announce Type: new Abstract: Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative news sources. To capture this challenge and support conflict verification tasks, we present ContraNote, a large-scale real…

    arxiv.org16 days agoView details

  113. When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Responsibility and Accountable Transparency

    arXiv:2609.00510v1 Announce Type: new Abstract: Artificial intelligence systems increasingly enact market-facing promises through chatbots, recommendation systems, automated decisions, and generative interfaces. Their failures, misuse, and misrepresentation raise a question that conventional brand-crisis models do not…

    arxiv.org16 days agoView details

  114. ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation

    arXiv:2609.00513v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) mitigates large language models (LLMs) hallucinations, yet conventional dense retrieval struggles with the complex reasoning paths of multi-hop question answering (QA). Graph-based RAG captures multi-step relationships but suffers fro…

    arxiv.org16 days agoView details

  115. Enoki: Efficient Multi-Level Hallucination Detection

    arXiv:2609.00581v1 Announce Type: new Abstract: Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucination detectors usually operate at a single level: claim-level methods provide interpretable factual units, while span-level methods localize unsupported text. B…

    arxiv.org16 days agoView details

  116. Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking

    arXiv:2609.00588v1 Announce Type: new Abstract: Reranking methods, such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking, are widely used in modern neural machine translation (NMT) to select an output from a set of candidate hypotheses. However, the performance gains come at the cost of high…

    arxiv.org16 days agoView details

  117. Investigating Assistant Bias in LLM User Simulators Using a Role Vector

    arXiv:2609.00608v1 Announce Type: new Abstract: LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit "assistant bias," a tendency to cooperate and pursue task goals. They rarely reproduce the frustra…

    arxiv.org16 days agoView details

  118. Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning

    arXiv:2609.00351v1 Announce Type: new Abstract: Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect without prior knowledge what to look…

    arxiv.org16 days agoView details

  119. Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time

    arXiv:2609.00624v1 Announce Type: new Abstract: A prominent paradigm in inference-time alignment employs lightweight supervisors to steer Large Language Models (LLMs). Through empirical analysis, we identify a structural mismatch in this paradigm: weak supervisors exhibit pervasive high entropy across the vast majorit…

    arxiv.org16 days agoView details

  120. Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts

    arXiv:2609.00578v1 Announce Type: new Abstract: Large Language Models (LLMs) can solve complex problems, but their misuse in high-risk domains can lead to severe consequences. Model providers therefore restrict assistance for potentially harmful requests. Refusing all cybersecurity requests would therefore harm legiti…

    arxiv.org16 days agoView details

  121. Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing

    arXiv:2609.00584v1 Announce Type: new Abstract: Does unrestricted AI access bypass the cognitive effort required for learning, or does it streamline knowledge acquisition? This paper reports on a study where we compare three designs for user-AI interaction in a learning context: (1) an unrestricted conversational bot…

    arxiv.org16 days agoView details

  122. REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows

    arXiv:2609.00643v1 Announce Type: new Abstract: Agent revisions expose a fundamental correctness--efficiency trade-off during concurrent execution. Discarding ongoing work preserves latest-version correctness but wastes progress that may remain valid, whereas reusing prior work preserves efficiency but risks propagati…

    arxiv.org16 days agoView details

  123. Emotional Labor Strategy Preferences in LLM Personas

    arXiv:2609.00310v1 Announce Type: new Abstract: Emotional labor is the effortful management of emotional displays to meet social or professional expectations. Personality traits have been correlated with emotional labor strategies, yet research on this link relies almost exclusively on self-report scales administered…

    arxiv.org16 days agoView details

  124. Visual Framing for News Stance Detection via Image Generation

    arXiv:2609.00685v1 Announce Type: new Abstract: Article-level news stance detection aims to identify the perspective of news articles toward social issues. Despite advances in stance detection and its importance for trustworthy media environments, news articles pose distinct challenges because their stances are often…

    arxiv.org16 days agoView details

  125. SCoNE: Selective Context-aware Neuron Editing for Robust Retrieval-Augmented Generation

    arXiv:2609.00689v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is highly sensitive to retrieval noise: when retrieved documents mix informative and irrelevant context, LLMs are easily distracted, leading to hallucinations. To overcome this, we propose SCoNE (Selective Context-aware Neuron Editing…

    arxiv.org16 days agoView details

  126. SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task

    arXiv:2609.00654v1 Announce Type: new Abstract: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scientific claims against the tables and figures of a paper. Rather than tuning a single model, we benchmark eleven frontier…

    arxiv.org16 days agoView details

  127. Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets

    arXiv:2609.00662v1 Announce Type: new Abstract: A multi-model language service must route each request while preserving workload-level budgets for compute, latency, memory, or monetary cost. Two features make this problem materially harder than static model selection. Prompt representations are high dimensional, so on…

    arxiv.org16 days agoView details

  128. Creative Generation via Multi-Agent Debate: Does Debate Suppress Diversity?

    arXiv:2609.00683v1 Announce Type: new Abstract: Creative generation tasks, such as narrative writing and scientific ideation, demand both high-quality outputs and distinct responses across independent runs to maximize exploration. Multi-Agent Debate (MAD) has shown strong quality gains on factual and reasoning tasks,…

    arxiv.org16 days agoView details

  129. Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

    arXiv:2609.00550v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly provided with contextual evidence in heterogeneous forms: as a text passage, as a rendered image of the same passage, or as both together. However, it remains unclear how consistently these surface forms are proce…

    arxiv.org16 days agoView details

  130. Authority Bias in Conversational Search Engines for Academic Paper Recommendation

    arXiv:2609.00248v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used as conversational search engines for academic literature, yet whether they judge papers on content or on authority signals has not been tested causally. We investigate authority bias: systematic preference for papers bas…

    arxiv.org16 days agoView details

  131. Hypotheses-Guided Self Distillation for Continual Personalization

    arXiv:2609.00251v1 Announce Type: new Abstract: As people increasingly interact with LLM assistants in daily life, continually adapting to individual preferences has become essential for effective long-term interactions. However, user preferences are rarely stated in full, and instead emerge through heterogeneous, lat…

    arxiv.org16 days agoView details

  132. Life Operators: a self-evolving framework for multiscale life modelling

    arXiv:2609.00068v1 Announce Type: new Abstract: Medical AI is moving beyond recognition towards clinical dialogue and longitudinal prediction. Yet a central question remains: how would a patient's state change under intervention? Statistical models learn future observations, whereas mechanistic models describe selecte…

    arxiv.org16 days agoView details

  133. Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

    arXiv:2609.00355v1 Announce Type: new Abstract: Speculative decoding accelerates generation without changing its output, yet on vision-language models (VLMs) it has been caught in a self-defeating cycle. The drafter stays autoregressive, so it must stay small. A small drafter cannot afford the image at every step, so…

    arxiv.org16 days agoView details

  134. SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation

    arXiv:2609.00427v1 Announce Type: new Abstract: The exponential growth of wireless devices is driving unprecedented spectrum demand, pushing spectrum management toward more fine-grained decisions across space, time, and device constraints. As a result, spectrum policymakers and engineers must process large volumes of…

    arxiv.org16 days agoView details

  135. Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment

    arXiv:2609.00055v1 Announce Type: new Abstract: Self-supervised respiratory encoders lack semantic grounding in clinical domain needed for zero-shot inference, limiting their utility without task-specific labeled data. We propose a framework that aligns these encoders with medical terminology in a shared latent space…

    arxiv.org16 days agoView details

  136. KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training

    arXiv:2609.00082v1 Announce Type: new Abstract: LLMs acquire vast amounts of knowledge during pre-training, but often lack the specialized knowledge needed to answer questions from niche sources such as manuals or technical documents unseen during pre-training. Continued pre-training (CPT) is widely used to inject suc…

    arxiv.org16 days agoView details

  137. The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems

    arXiv:2609.00275v1 Announce Type: new Abstract: Fleets of LLM agents now externalize effects that cannot be fully undone: they move money, deploy code, delete data, and disclose information. Current controls check one effect at a time, so a fleet of individually authorized agents can overdraw its principal's risk unde…

    arxiv.org16 days agoView details

  138. Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective

    arXiv:2609.00334v1 Announce Type: new Abstract: Across law, education, policy analysis, and public moral argumentation, LLM outputs are being used often for work that requires interpretations to be justified with textual evidence and explicit normative standards. Yet a recurrent failure mode -- what I call \textit{int…

    arxiv.org16 days agoView details

  139. Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

    arXiv:2609.00482v1 Announce Type: new Abstract: Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approxi…

    arxiv.org16 days agoView details

  140. Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models

    arXiv:2609.00495v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. This generation paradigm introduces two axes that can influence safety alignment: when tokens are generated during denoising and where they appear in the…

    arxiv.org16 days agoView details

  141. Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search

    arXiv:2609.00652v1 Announce Type: new Abstract: Language model agents increasingly propose actions, observe external feedback, and explain their own behavior. Their confidence and rationales are convenient monitoring signals, but convenience is not verification. We introduce an environment-grounded audit in which ever…

    arxiv.org16 days agoView details

  142. The Curse of Multilinguality in Lexical Normalization

    arXiv:2609.00329v1 Announce Type: new Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages at once. We ask a simp…

    arxiv.org16 days agoView details

  1. XHToken/Spark-X2.5-4B

    text-generation · transformers · safetensors · spark2_5

    huggingface.co16 days ago1247 ptsView details

  2. Motif-Technologies/Motif-3-Beta

    text-generation · transformers · safetensors · Motif

    huggingface.co17 days ago218 ptsView details

  3. AngelSlim/Hy4-preview-GGUF

    gguf · endpoints_compatible · region:us

    huggingface.co17 days ago92 ptsView details

  4. Motif-Technologies/Motif-3-NVFP4

    text-generation · transformers · safetensors · Motif

    huggingface.co17 days ago29 ptsView details

  5. Gramscii-IT/SemanticRepair-270M

    text-generation · mlx · safetensors · gguf

    huggingface.co17 days ago2 ptsView details

  6. tamm5y5m5/nemotron-3.5-quran-dual-v5

    nemo · region:us

    huggingface.co17 days ago1 ptsView details

  7. wop/littlechat-10m-instruct

    license:apache-2.0 · region:us

    huggingface.co17 days ago1 ptsView details

  8. VikramPal/Qwen3.8-27B-bf16

    text-generation · transformers · safetensors · qwen3_5_text

    huggingface.co17 days agoView details

  9. VikramPal/Qwen3.5-35B-A3B-bf16

    text-generation · transformers · safetensors · qwen3_5_moe_text

    huggingface.co17 days agoView details

  1. shtjww/llm-inference-capacity-handbook

    大模型推理案头手册(开源版):给定 GPU 算力,一个模型能扛多少 QPS?三层模型 × 三面墙 × 排队论 × 开环压测

    github.com16 days ago20 ptsView details

  2. openai/openai-python v3.7.0

    ## [3.7.0](https://github.com/openai/openai-python/compare/v3.6.0...v3.7.0) (2026-09-02) ### Features * **api:** update usage APIs and documentation ([#3779](https://github.com/openai/openai-python/issues/3779)) ([6f0da16](https://github.com/openai/openai-python/commit/6f0da1657…

    github.com16 days agoView details

  3. anthropics/anthropic-sdk-python v1.3.0

    ## 1.3.0 (2026-09-01) Full Changelog: [v1.2.0...v1.3.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.2.0...v1.3.0) ### Features * **api:** beta user profiles: add external_user_onboarded_at, remove relationship in favor of access_type ([74080c3](https://github.c…

    github.com16 days agoView details

  4. langchain-ai/langchain langchain==1.4.0a3

    Third alpha of the `1.4.0` line. This release focuses on the new `langchain.mcp` namespace for adapting MCP servers into LangChain tools. ## `langchain.mcp` highlights - **`MCPAdapter`** adapts any target `fastmcp.Client` accepts — a URL, a local script, an in-process server, an…

    github.com16 days agoView details

  5. ggml-org/llama.cpp b10731

    <details open> qwen4exp: support recurrent state rollback (#28123) MTP speculative decoding needs the target state to move back by the number of rejected draft tokens. Without rollback support the context is classified as SEQ_RM_TYPE_FULL and the server serializes the whole recu…

    github.com17 days agoView details

  6. ggml-org/llama.cpp b10730

    <details open> qwen4exp: sum the indexer heads by slices (#28023) * qwen4exp: sum the indexer heads by slices The head reduction went through a transpose and a sum_rows over ne[1], which left sum_rows with ne0 = 4, one block per row for a four element reduction, and the transpos…

    github.com17 days agoView details