Skip to content

Archive / 2026-09-02

September 2, 2026

  1. Muse Spark 1.3

    developer.meta.com15 days ago680 ptsView detailsJoin discussion

  2. Qantas Airbus A380 engine failure in 2010 (2023)

    admiralcloudberg.medium.com15 days ago175 ptsView detailsJoin discussion

  3. Quasar 438B: Europe's Leading AI Model

    multiversecomputing.com15 days ago173 ptsView detailsJoin discussion

  4. Reasons robotics is hard

    secondthoughts.ai15 days ago127 ptsView detailsJoin discussion

  5. LLMs: Intelligence vs. Cost

    openteams.com15 days ago96 ptsView detailsJoin discussion

  6. Introducing Muse Spark 1.3

    research.meta.ai15 days ago65 ptsView detailsJoin discussion

  7. AI Policy

    dbushell.com15 days ago50 ptsView detailsJoin discussion

  8. How Railroad Crossings Work (2024)

    practical.engineering15 days ago36 ptsView detailsJoin discussion

  9. GLM-5.3 Uncensored

    huggingface.co15 days ago15 ptsView detailsJoin discussion

  10. Primary source

    Give Your Coding Agents a Memory You Own

    huggingface.co15 days ago11 ptsView details

  11. Primary source

    Safety overview: GPT-6 Astra

    GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.

    openai.com15 days agoView details

  12. Primary source

    ATV Big Air Tour turned 3 days of work into 3 hours with ChatGPT

    ATV Big Air Tour uses ChatGPT Work to speed up marketing, merchandising, and more. It even turned merchandise photos into an inventory website in 15 minutes.

    openai.com15 days agoView details

  13. Primary source

    Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

    deepmind.google15 days agoView details

  14. Primary source

    An Organizational Second Brain: Building an AI That Learns From Experts

    We’ve built an AI agent that acts as a secondary expert for a given domain, making deep specialist knowledge readily available and preserved for anyone in an organization to access, share, and build upon. This is not a typical domain-specific agent. Its novelty comes from integrating two layers: A structured, auditabl…

    engineering.fb.com16 days agoView details

  15. Qwen Developers Open-Sources zg (zvec-grep): A Local-First Search Layer Unifying ripgrep, BM25, and Vector Search

    We look at zg (zvec-grep), the local-first search layer that Qwen Developers just open-sourced under Apache 2.0. We explain how it puts ripgrep, BM25, and vector search behind a single interface, so an agent can move from a plain-language description to an exact line span without switching tools. We cover the delibera…

    marktechpost.com15 days agoView details

  16. Meet Switchyard: A Rust Proxy and Library That Routes and Translates LLM Traffic Across OpenAI and Anthropic APIs

    NVIDIA has released Switchyard, an Apache-2.0 Rust proxy and library for LLM traffic. It decodes requests into provider-neutral types, routes them with passthrough, random, LLM-classifier, or stage-router algorithms, and translates responses back into the client's format, so Claude Code or Codex CLI can run against vL…

    marktechpost.com15 days agoView details

  17. Google DeepMind Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber: One Core Model, Two Access Envelopes

    Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 2, 2026. Both variants run on the same foundational intelligence, split by safety mitigations rather than model size. Gemini 3.8 Flash is generally available at $0.75 and $3.75 per 1M tokens, introductory through December 31, 2026. Flash Cyber re…

    marktechpost.com15 days agoView details

  18. Anthropic Introduces Enterprise Frontier Safeguards (EFS): Zero-Data-Retention Privacy Plus Cross-Session Misuse Detection

    Anthropic announced Enterprise Frontier Safeguards on September 1, 2026, an architecture that stores monitoring data in the customer's own cloud account rather than Anthropic's. Detection stays automated and Anthropic-run; custody, encryption keys, and flag review stay with the customer. Built with more than 100 enter…

    marktechpost.com16 days agoView details

  19. Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing

    Most production voice stacks are three systems stitched together. One model transcribes, a second separates speakers, and a detector decides when the user stopped talking. Each hand-off adds latency and a new failure mode. Muse Voice Transcribe, announced by Meta Superintelligence Labs this week, collapses those three…

    marktechpost.com16 days agoView details

  20. Perplexity Releases Hybrid Compute on Mac: Cloud Agents Orchestrate Down to a Local Model, Gated On Device

    Agentic assistants have a structural problem: the context that makes them useful — deal documents, privileged files, client records — is exactly the context users cannot send to a cloud endpoint. This week, Perplexity shipped its answer for Mac. Hybrid compute splits a single Perplexity Computer task between frontier…

    marktechpost.com16 days agoView details

  21. Meta Pushes Its New AI Agent on Employees—but Eases Off on Tokenmaxxing

    The company is reducing pressure on workers to use artificial intelligence tools while encouraging them to experiment with Hatch, its most advanced AI project yet.

    wired.com15 days agoView details

  22. Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost more

    Google launched Gemini 3.8 Flash, arriving just a few weeks after its predecessor. The company claims the new model "works harder" than Gemini 3.7 Flash by performing more reasoning steps on complex tasks and "calling tools iteratively." It has the same introductory pricing as 3.7 Flash, $0.75 per million input tokens…

    theverge.com15 days agoView details

  23. Trump Administration Sides With OpenAI in New York Times Copyright Lawsuit

    The US government wrote a letter in support of OpenAI’s argument that training AI on others' intellectual property is fair use.

    wired.com15 days agoView details

  24. These Russian Mathematicians Taught AI Models How to Talk to Each Other Without Using Words

    A startup called Mostik has a wild new approach to combining the capabilities of AI models.

    wired.com15 days agoView details

  25. Amazon’s AI assistant can now spot fake emails from the company

    Amazon is trying to combat impersonation scams with a new feature that allows you to use its AI assistant to determine whether an email, text message, or phone call actually came from the company. With the update, you can ask Alexa for Shopping about a message you received, and it will use AI to compare it "against a…

    theverge.com15 days agoView details

  26. Researchers fear safety disaster ahead of OpenAI’s Astra release

    OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it "may be the single worst development for AI security/safety to date.…

    theverge.com15 days agoView details

  27. The Trump administration is supporting OpenAI in the NYT copyright lawsuit

    The Trump administration has intervened in The New York Times' copyright lawsuit against OpenAI, making an argument in favor of the AI lab. The landmark lawsuit, filed in December 2023, alleging that OpenAI unlawfully trained its AI systems on articles from The New York Times and seeks to recoup "billions of dollars"…

    theverge.com15 days agoView details

  28. The Logical End Point of AI Job Interviews Is Two Bots Talking to Each Other

    Christopher was sick of being ghosted by AI recruiters. So he unleashed ChatGPT on his robot interviewer.

    wired.com15 days agoView details

  29. Google is sending MrBeast into the wilderness, armed with AI

    MrBeast will feature Gemini, Google Health, and the Fitbit Air in upcoming videos as part of a multi-year partnership with Google. The deal will kick off with a video featuring Jimmy "MrBeast" Donaldson turning to Gemini for wilderness survival advice: First up on September 5 is a new MrBeast video following Jimmy and…

    theverge.com15 days agoView details

  30. OpenAI accused of ‘aiding and abetting’ Tumbler Ridge mass shooting in dozens of new lawsuits

    OpenAI and its CEO Sam Altman are facing 30 new lawsuits that accuse them of providing "substantial assistance and encouragement" to the suspect in Canada's Tumbler Ridge school shooting, as reported earlier by TechCrunch. The new wave of lawsuits was filed in a California federal court on Wednesday by the students, t…

    theverge.com15 days agoView details

  31. NYC bans AI use for students until they reach high school

    New York City Mayor Zohran Mamdani has announced a new policy today that will ban younger schoolchildren from using AI in classrooms. The one-year moratorium, effective in the 2026-2027 school year, will impact about 600,000 public school students in 2-K through eighth grade and is being introduced alongside additiona…

    theverge.com15 days agoView details

  32. Pangram Has Emerged as the Gold Standard of AI Detection. Should You Trust It?

    Meet the AI police who can make or break careers—in publishing and beyond.

    wired.com15 days agoView details

  33. Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

    arXiv:2609.02749v1 Announce Type: new Abstract: Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent.…

    arxiv.org15 days agoView details

  34. CORAL: An LLM-Native Harness for Production Recommender Systems

    arXiv:2609.02730v1 Announce Type: new Abstract: Production recommender systems shape what billions of people see, and sustaining their performance requires continual optimization: as content, user behavior, and upstream models shift, the choices governing retrieval, ranking, and serving must be revisited. Traditionall…

    arxiv.org15 days agoView details

  35. When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection

    arXiv:2609.01814v1 Announce Type: new Abstract: Information sharing can improve a pooled estimate while eliminating independent rescue actions. This paper separates those effects in exact finite discovery models. A centralized action-budget profile shows that equal one-person accuracy can coexist with different portfo…

    arxiv.org15 days agoView details

  36. PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation

    arXiv:2609.01658v1 Announce Type: new Abstract: Retrieval-Augmented Generation enhances Large Language Models by grounding responses in external knowledge, but multi-hop reasoning remains vulnerable to error propagation, where early retrieval failures confound subsequent steps. Standard outcome-based optimization only…

    arxiv.org15 days agoView details

  37. When Persona Attributes Improve Population Alignment in Large Language Models

    arXiv:2609.02526v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to predict the responses of human participants in survey panels. Towards that goal, persona prompting has recently emerged as a technique to inform and align large pretrained language models. Persona prompting refers to…

    arxiv.org15 days agoView details

  38. Choosing a PEFT Variant for Per-Patient Dysarthric ASR: A Single-Speaker Case Study on Two ASR Bases

    arXiv:2609.02735v1 Announce Type: new Abstract: Per-patient adapters are the preferred production architecture for dysarthric automatic speech recognition (ASR), yet parameter-efficient fine-tuning (PEFT) variants have not been compared in the speaker-dependent, per-patient regime. We present a single-speaker case stu…

    arxiv.org15 days agoView details

  39. EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

    arXiv:2609.02783v1 Announce Type: new Abstract: Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly across iterative development cycles.…

    arxiv.org15 days agoView details

  40. IDEEA: training-free Input-Dependent stEEring via Activation cluster matching

    arXiv:2609.02089v1 Announce Type: new Abstract: Steering aligns large language models (LLMs) by injecting a bias into selected activations at inference time, offering a far cheaper alternative to weight-update methods such as supervised fine-tuning or reinforcement learning. However, most existing training-free steeri…

    arxiv.org15 days agoView details

  41. Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA

    arXiv:2609.01687v1 Announce Type: new Abstract: Grounded question answering systems should answer only when the supplied evidence supports the answer. In multi-hop QA, this requirement is difficult because partial evidence can make an unsupported answer appear plausible. We study selective answering through evidence s…

    arxiv.org15 days agoView details

  42. Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

    arXiv:2609.02496v1 Announce Type: new Abstract: Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as SparseGPT, can amplify existing biases in models,…

    arxiv.org15 days agoView details

  43. When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

    arXiv:2609.01985v1 Announce Type: new Abstract: As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization of how they behave on systems-level requirements: schema design, async orchestration, configuration correctness, and retrieval-filtering trade-offs. We present a cas…

    arxiv.org15 days agoView details

  44. Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers

    arXiv:2609.02702v1 Announce Type: new Abstract: Transformers process information causally, but long-context reasoning may depend on task state discovered only later. We formalize this mismatch through conditional state update tasks. For causal state update processors, providing the condition first can require exponent…

    arxiv.org15 days agoView details

  45. Predictors of Loneliness in Older Adults Using Multimodal Analysis of Speech and Language

    arXiv:2609.02606v1 Announce Type: new Abstract: Loneliness is a critical public health issue among older adults, linked to higher risks of depression, cognitive decline, and mortality. Scalable, objective methods for its detection remain limited, particularly in natural conversational contexts. We analyzed speech and…

    arxiv.org15 days agoView details

  46. Thinking effort aligns between humans and reasoning models in abductive reasoning

    arXiv:2609.01867v1 Announce Type: new Abstract: A major question in cognitive modeling concerns the behavioral alignment between large language models and humans across linguistic and non-linguistic tasks. Unlike standard LLMs, large reasoning models (LRMs) are optimized with reinforcement learning from verifiable rew…

    arxiv.org15 days agoView details

  47. Interpretable Symptom Vectors for Depression in a Large Language Model

    arXiv:2609.01832v1 Announce Type: new Abstract: Patients with depression present with diverse symptom profiles, yet clinical practice routinely reduces this variation to a single severity score. Large language models (LLMs) can potentially capture various symptoms and their severity from patient speech. However, how d…

    arxiv.org15 days agoView details

  48. Untangling the Mechanisms of Misleading Context in Medical Question Answering

    arXiv:2609.02754v1 Announce Type: new Abstract: Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we e…

    arxiv.org15 days agoView details

  49. From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

    arXiv:2609.02771v1 Announce Type: new Abstract: Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweig…

    arxiv.org15 days agoView details

  50. OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction

    arXiv:2609.02158v1 Announce Type: new Abstract: Legal Judgment Prediction (LJP) models are typically trained on documents that describe facts from a prosecutorial perspective. Existing datasets further exhibit severe label imbalance toward guilty outcomes. Consequently, these models suffer from "Guilty Bias", blindly…

    arxiv.org15 days agoView details

  51. UTP-Bench: Uncertainty-aware Travel Planning Benchmark

    arXiv:2609.02421v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexpected stochastic delays frequently inva…

    arxiv.org15 days agoView details

  52. text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation

    arXiv:2609.02115v1 Announce Type: new Abstract: Natural language interfaces to databases have traditionally suffered from three structural limitations: exclusive targeting of relational SQL, unconditional dependence on large language model (LLM) inference at query time, and absence of any runtime signal when generated…

    arxiv.org15 days agoView details

  53. Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models

    arXiv:2609.02108v1 Announce Type: new Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to the auto-regressive paradigm. With bidirectional attention and any-order generation, DLMs naturally fit infilling tasks, which require generating a middle span conditioned on both the prefix and…

    arxiv.org15 days agoView details

  54. DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

    arXiv:2609.02796v1 Announce Type: new Abstract: Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translat…

    arxiv.org15 days agoView details

  55. Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning

    arXiv:2609.02191v1 Announce Type: new Abstract: Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We defined fault points as moments in AI agent conversations, in which an agent's reasoning became most vulnerable to external influence. Using the MedQA dataset, this s…

    arxiv.org15 days agoView details

  56. DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models

    arXiv:2609.02685v1 Announce Type: new Abstract: RAG has become the de facto method for incorporating new, corpus-specific knowledge into an instruction following LLM (Instruct LLM). Although RAG-based prompting improves factual grounding, it fails when retrieval is incorrect or incomplete, leading to hallucinations. F…

    arxiv.org15 days agoView details

  57. HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs

    arXiv:2609.02056v1 Announce Type: new Abstract: Scientific knowledge graphs organize entities and relations extracted from scientific literature, but they remain inherently incomplete. Missing typed links in such graphs can therefore represent plausible scientific hypotheses, such as unexplored associations between ma…

    arxiv.org15 days agoView details

  58. APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering

    arXiv:2609.02253v1 Announce Type: new Abstract: Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either retrieve verbose task-specific…

    arxiv.org15 days agoView details

  59. Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions

    arXiv:2609.02371v1 Announce Type: new Abstract: With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is essential to achieving superior effectiveness and trustworthiness. As agent failures often manifest via long and complex trajectories, manually finding the needles in the h…

    arxiv.org15 days agoView details

  60. Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern

    arXiv:2609.01834v1 Announce Type: new Abstract: As enterprise platforms transition to conversational reasoning interfaces, the stateless nature of LLM APIs creates an architectural gap. While statelessness enables horizontal scalability for AI providers, it forces client applications to manage the entire burden of con…

    arxiv.org15 days agoView details

  61. PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion

    arXiv:2609.02216v1 Announce Type: new Abstract: Inductive knowledge graph completion (IKGC) aims to predict missing links involving entities unseen during training, requiring models to learn transferable relational and structural patterns. Existing subgraph- and path-based approaches often encode relational paths inde…

    arxiv.org15 days agoView details

  62. DiffIE: Diffusion-based Open Information Extraction

    arXiv:2609.02315v1 Announce Type: new Abstract: A single sentence often expresses multiple valid relational triplets, which makes Open Information Extraction (OpenIE) fundamentally a multi-output task. Existing neural systems handle this by autoregressive generation, which is flexible but slow and prone to redundancy,…

    arxiv.org15 days agoView details

  63. Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus

    arXiv:2608.30485v1 Announce Type: cross Abstract: The language a legislature uses to debate women's rights, even in favour of them, encodes systematic patterns of sexism that persist across two centuries. In this work, we analyse 6,531 speeches over 200 years of UK parliamentary debate (Hansard, 1803-2005) by using la…

    arxiv.org15 days agoView details

  64. SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

    arXiv:2609.01737v1 Announce Type: new Abstract: Mobile payment applications in Nepal are graphically mediated and largely inaccessible to visually impaired users. This paper presents SpeakPay, a voice-first digital wallet, and documents the central technical contribution: a controlled study of domain adaptation for lo…

    arxiv.org15 days agoView details

  65. MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

    arXiv:2609.01772v1 Announce Type: new Abstract: Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack. We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengal…

    arxiv.org15 days agoView details

  66. Candidate Generation and Definition-Guided Verification for Sentence-Level Depression Symptom Recognition

    arXiv:2609.01833v1 Announce Type: new Abstract: Sentence-level recognition of depression symptoms is challenging because similar expressions can differ in symptom relevance, and language-model inference is insufficiently grounded in diagnostic definitions. This study proposes a two-stage framework separating symptom-c…

    arxiv.org15 days agoView details

  67. The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

    arXiv:2609.01852v1 Announce Type: new Abstract: Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capability changes. We evaluate a frozen, closed-set, action-scored benchmark with 2 suites that r…

    arxiv.org15 days agoView details

  68. Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets

    arXiv:2609.01918v1 Announce Type: new Abstract: Energy poverty is nearly absent from NLP-for-social-good, and the little existing work is either static retrieval/QA or relies on carbon-intensive cloud LLMs, a self-defeating "computational irony" for a humanitarian setting. We present EqGrid, a closed-loop simulation i…

    arxiv.org15 days agoView details

  69. Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization

    arXiv:2609.01794v1 Announce Type: new Abstract: How do learners avoid overgeneralizations such as Tom laughed me without explicit negative evidence? Constructionists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption (which privileges exposure to near-synonymous…

    arxiv.org15 days agoView details

  70. How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?

    arXiv:2609.01798v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored. This paper aims to understand how two prompt properties, cognitive load and phras…

    arxiv.org15 days agoView details

  71. Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

    arXiv:2609.01936v1 Announce Type: new Abstract: A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a lens reading reflects both the hidden state and the readout (the unembedding matrix) used to decode it. M…

    arxiv.org15 days agoView details

  72. NS-Copilot: An LLM-Driven Agent System for Autonomous Neuroscience Analysis

    arXiv:2609.01971v1 Announce Type: new Abstract: AI is rapidly advancing neuroscience, yet many laboratories fail to fully unleash its potential due to significant interdisciplinary barriers. While pre-trained neural models for physiological data are progressing quickly, their heterogeneous architectures and modality-s…

    arxiv.org15 days agoView details

  73. How Output Format Confounds Data Quality and Capability in Instruction Tuning

    arXiv:2609.02015v1 Announce Type: new Abstract: Instruction-tuning data are judged by quality metrics, and tuned models are judged by benchmarks, but both judgments pass through an output interface: the surface format in which an answer is written. Using gradient signatures across 12 tasks, four semantically equivalen…

    arxiv.org15 days agoView details

  74. AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking

    arXiv:2609.01828v1 Announce Type: new Abstract: Spoken dialogue state tracking recovers slot-value pairs from speech, where ASR errors concentrate in entity values and persist across turns, making it both a generation and an editing problem. A strong per-turn text editor corrects much of this but, operating on the tra…

    arxiv.org15 days agoView details

  75. WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities

    arXiv:2609.02651v1 Announce Type: new Abstract: While English language models have been widely examined for anti-queer bias, Dutch models remain understudied. To address this gap, we developed a culturally and linguistically adapted Dutch dataset based on the English WinoQueer benchmark, containing pairs of stereotypi…

    arxiv.org15 days agoView details

  76. oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions

    arXiv:2609.02672v1 Announce Type: new Abstract: Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix. Leaving that matrix unconstrained places no limit on the factor by which the mixing step rescales th…

    arxiv.org15 days agoView details

  77. SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning

    arXiv:2609.02336v1 Announce Type: new Abstract: Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent logic-based methods address this by matching pre…

    arxiv.org15 days agoView details

  78. ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

    arXiv:2609.01992v1 Announce Type: new Abstract: Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs and hash-linked transcripts answer neith…

    arxiv.org15 days agoView details

  79. Door-in-the-Face Requests and Refusal Behaviour in Large Language Models

    arXiv:2609.02707v1 Announce Type: new Abstract: Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then rece…

    arxiv.org15 days agoView details

  80. A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models

    arXiv:2609.02054v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarification, especially when user queries are ambiguous or underspecified. This…

    arxiv.org15 days agoView details

  81. Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage

    arXiv:2609.02091v1 Announce Type: new Abstract: Knowledge editing provides an efficient way to update factual knowledge in large language models. However, malicious edits may introduce safety risks, making it necessary to reverse undesirable editing effects. Existing reversal methods for parameter-modifying edits main…

    arxiv.org15 days agoView details

  82. C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees

    arXiv:2609.02131v1 Announce Type: new Abstract: Sentiment in social-media threads does not only vary across posts; it shifts as users react to claims, corrections, evidence, and hostility within a branching reply tree. We study why sentiment changes in rumor-centric conversation trees by treating discourse moves (e.g.…

    arxiv.org15 days agoView details

  83. A Layered Taxonomy for Chinese Learner Grammatical Error Annotation

    arXiv:2609.02153v1 Announce Type: new Abstract: Grammatical error annotation in Chinese learner writing requires labels that are both consistent and linguistically meaningful. This paper proposes a layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis. The sch…

    arxiv.org15 days agoView details

  84. Induction and Inquiry via Probabilistic Reasoning over Language and Code

    arXiv:2609.01815v1 Announce Type: new Abstract: How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational account must satisfy at least three desiderata: It must be (1) data-efficient and compute-efficient, (2…

    arxiv.org15 days agoView details

  85. SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval

    arXiv:2609.01849v1 Announce Type: new Abstract: This article presents SSAKG 2.0, an open-source software package for constructing and operating Structural Sequential Associative Knowledge Graphs (SSAKGs). An SSAKG represents objects as graph vertices and ordered sequences as structural patterns of graph connections. T…

    arxiv.org15 days agoView details

  86. Breadth Beats Depth: Improving GCG-Based Jailbreak Optimization with Breadth-Oriented Suffix Search

    arXiv:2609.02172v1 Announce Type: new Abstract: Optimization-based jailbreak attacks such as Greedy Coordinate Gradient (GCG) achieve strong effectiveness and transferability by optimizing adversarial suffixes on white-box source models. However, existing GCG-based methods rely on averaged adversarial loss and deep gr…

    arxiv.org15 days agoView details

  87. Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation

    arXiv:2609.02163v1 Announce Type: new Abstract: Information-theoretic measures derived from autoregressive language models are widely used to characterize the expectations that shape human reading, but whether language-variety-specific training improves such psycholinguistic alignment remains unclear. This question is…

    arxiv.org15 days agoView details

  88. Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization

    arXiv:2609.01861v1 Announce Type: new Abstract: The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a new candidate each round. Each…

    arxiv.org15 days agoView details

  89. Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence

    arXiv:2609.01873v1 Announce Type: new Abstract: Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly identical reports.…

    arxiv.org15 days agoView details

  90. The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

    arXiv:2609.01909v1 Announce Type: new Abstract: Clinical prediction can saturate for two different reasons: a fitted learner may fail to extract available information, or the recorded variables may impose a population frontier. We separate these quantities through the \emph{learner gap} and the \emph{measurement-chann…

    arxiv.org15 days agoView details

  91. Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment

    arXiv:2609.01962v1 Announce Type: new Abstract: Ultra-low-bit language models can reduce storage and memory bandwidth, but a nominal "1.58-bit" label does not fully describe the stored representation, retained capability, or runtime behavior. We study an end-to-end post-training conversion of Qwen, an instruction-tune…

    arxiv.org15 days agoView details

  92. AI agents reshape consensus formation in human groups

    arXiv:2609.02122v1 Announce Type: new Abstract: As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative description game, in w…

    arxiv.org15 days agoView details

  93. Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization

    arXiv:2609.02309v1 Announce Type: new Abstract: GUI agents increasingly operate across websites, mobile apps, and desktop environments, yet the field still reports progress primarily through task success. We argue that practical deployment depends equally on efficiency: how much context, computation, action budget, an…

    arxiv.org15 days agoView details

  94. Benchmarking Language Models for Statistical Problem Formulation

    arXiv:2609.01982v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous data, leaving the mod…

    arxiv.org15 days agoView details

  95. HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

    arXiv:2609.02029v1 Announce Type: new Abstract: Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate conte…

    arxiv.org15 days agoView details

  96. Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision

    arXiv:2609.02057v1 Announce Type: new Abstract: Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits are unavailable. In this work, we study prefix-level risk prediction for web agents using observable trajectory signals: given an evolving prefix, estimate whether the…

    arxiv.org15 days agoView details

  97. Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos

    arXiv:2609.01846v1 Announce Type: new Abstract: Recorded lecture videos, often enhanced with search and summarization features, are a standard study resource. However, students cannot easily ask course specific questions or verify answers against an instructor's lecture. We report a semester-long deployment of VideoPo…

    arxiv.org15 days agoView details

  98. PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

    arXiv:2609.02236v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in multi-turn agentic tasks with sparse terminal rewards, it often provides coarse credit for intermediate actions. To obtain more fine-grained credit assignment, recent w…

    arxiv.org15 days agoView details

  99. MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity

    arXiv:2609.02060v1 Announce Type: new Abstract: Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps. We present MineTRACE, a web-based system for evidence-grounded exploration of eight…

    arxiv.org15 days agoView details

  100. ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

    arXiv:2609.02067v1 Announce Type: new Abstract: Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and online resources. These routes can produce strong evaluations, but they require substantial…

    arxiv.org15 days agoView details

  101. VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

    arXiv:2609.01788v1 Announce Type: new Abstract: Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic la…

    arxiv.org15 days agoView details

  102. From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

    arXiv:2609.02679v1 Announce Type: new Abstract: When LLMs support public-facing or high-stakes workflows, missed fabrications can harm users and institutions, while false alarms consume limited human-review capacity. When no trusted context or reference document is available, we study two signals accessible through bl…

    arxiv.org15 days agoView details

  103. Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis

    arXiv:2609.02805v1 Announce Type: new Abstract: Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations in modern 5G and emerging 6G networks remains challenging due to complex cross-layer dependencies. While large language models (LLMs) offer promising capab…

    arxiv.org15 days agoView details

  104. CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning

    arXiv:2609.02074v1 Announce Type: new Abstract: Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training data. Self-evolving memory instea…

    arxiv.org15 days agoView details

  105. AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application

    arXiv:2609.02821v1 Announce Type: new Abstract: Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived r…

    arxiv.org15 days agoView details

  106. MASkills: Continual Skills Optimization for Multi-Agent LLM Systems

    arXiv:2609.02094v1 Announce Type: new Abstract: LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement from interaction experience remains challenging. Existing self-reflection methods build experience memories, but memories are mostly hard to invoke, refine, or scale,…

    arxiv.org15 days agoView details

  107. Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems

    arXiv:2609.02092v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet outcome-based fairness audits can miss where risks arise within the decision trajectory. We present SCOPED-Hiring, a process-aware fairness diagnosis pipeline for LLM-bas…

    arxiv.org15 days agoView details

  108. Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

    arXiv:2609.02414v1 Announce Type: new Abstract: Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influenc…

    arxiv.org15 days agoView details

  109. READY or Not: Reliable Enterprise Agent Deployment

    arXiv:2609.02095v1 Announce Type: new Abstract: An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a require…

    arxiv.org15 days agoView details

  110. ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models

    arXiv:2609.02215v1 Announce Type: new Abstract: Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the same operational content, changing how the model reads it, are therefore only weakly cover…

    arxiv.org15 days agoView details

  111. TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

    arXiv:2609.01810v1 Announce Type: new Abstract: Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-gr…

    arxiv.org15 days agoView details

  112. PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment

    arXiv:2609.02231v1 Announce Type: new Abstract: Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce Phoeni…

    arxiv.org15 days agoView details

  113. Learning to Fuse LLMs with Ontology Rankers for Rare-Disease Diagnosis

    arXiv:2609.02473v1 Announce Type: new Abstract: Ontology rankers remain useful for rare-disease diagnosis because each candidate can be traced to matched patient phenotypes. Large language models (LLMs) can generate differential diagnoses from the same patient description, but their predictions lack an equally clear e…

    arxiv.org15 days agoView details

  114. Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality

    arXiv:2609.02242v1 Announce Type: new Abstract: AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any proposal, which can f…

    arxiv.org15 days agoView details

  115. Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training

    arXiv:2609.02244v1 Announce Type: new Abstract: Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete. Task-level natural-language priors can provide useful guidance in such settings, but existing approaches usually treat these priors as input context rather than as le…

    arxiv.org15 days agoView details

  116. LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

    arXiv:2609.02246v1 Announce Type: new Abstract: Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score higher, and the score comes from a judge that is itself an LLM. That judge has the last word on whether the system is getting better, and our position is that it has not…

    arxiv.org15 days agoView details

  117. NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning

    arXiv:2609.02366v1 Announce Type: new Abstract: Named Entity Recognition (NER) has achieved substantial progress since the advent of large language models (LLMs). Nevertheless, the recognition of long-tail and domain-specific entities remains challenging due to the deficiency in parametric knowledge. Retrieval-augment…

    arxiv.org15 days agoView details

  118. GAPS: Dimension-Level Gates for Conditional Activation Steering

    arXiv:2609.01878v1 Announce Type: new Abstract: Activation steering suppresses undesired behaviors in language models by adding a steering vector to the hidden state during generation. Recent conditional methods such as CAST and DSAS improve the behavior-capability trade-off by deciding when to intervene, but once act…

    arxiv.org15 days agoView details

  119. Discriminative World Models for Web Agents

    arXiv:2609.02885v1 Announce Type: new Abstract: Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state…

    arxiv.org15 days agoView details

  120. Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics

    arXiv:2609.02116v1 Announce Type: new Abstract: Reverse-logistics operators often decide how to inspect and route returned assets before their condition is fully observed, while full inspection consumes scarce labor. Semantic Signal-Assisted Decision Support converts return notes into a condition factor and a signal-q…

    arxiv.org15 days agoView details

  121. HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks

    arXiv:2609.02772v1 Announce Type: new Abstract: Low-resource authorship style transfer (LAST) aims to rewrite text into the style of an arbitrary target author using only a few reference examples while preserving the original meaning. Existing methods often struggle to achieve both high style fidelity and semantic pre…

    arxiv.org15 days agoView details

  122. When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic

    arXiv:2609.01741v1 Announce Type: new Abstract: Statutes are increasingly parsed by machines before people read them, and the parsers disagree: on Missouri's statutes, two independently written extractors diverge on numeric-threshold presence at a false-negative rate of 0.43. We ask what formal logic survives such noi…

    arxiv.org15 days agoView details

  123. PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation

    arXiv:2609.02272v1 Announce Type: new Abstract: Faithfully translating research papers into repository-level implementations remains challenging because papers often describe methods at a high level, leave implementation assumptions implicit, and require generated repositories to preserve method logic, evaluation prot…

    arxiv.org15 days agoView details

  124. Do Large Language Models Capture the Diversity in their Training Data?

    arXiv:2609.02275v1 Announce Type: new Abstract: Large language models are trained to model conditional distributions over text, yet it remains inadequately understood whether they capture the full diversity of plausible outputs present in their training data. We study this question through an information-theoretic len…

    arxiv.org15 days agoView details

  125. DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

    arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capab…

    arxiv.org15 days agoView details

  126. WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling

    arXiv:2609.01608v1 Announce Type: cross Abstract: Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces. Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation or trial-and-error refinement. A natur…

    arxiv.org15 days agoView details

  127. FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

    arXiv:2609.02168v1 Announce Type: new Abstract: Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol, aggregating resul…

    arxiv.org15 days agoView details

  128. SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

    arXiv:2609.02217v1 Announce Type: new Abstract: LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global document or as a flat pool of per-task entries, though most of the evidence comes from domains with structurally similar tasks. On long-horizon workloads where each task…

    arxiv.org15 days agoView details

  129. CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging

    arXiv:2609.02273v1 Announce Type: new Abstract: Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs) without full model retraining, yet it remains challenged by parameter interference. While existing methods aim to preserve the capabilities of individual expert models a…

    arxiv.org15 days agoView details

  130. PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation

    arXiv:2609.02480v1 Announce Type: new Abstract: Synthetic dialogue generation can support research in privacy-restricted service settings, but generated conversations must preserve communicative intent, affective meaning, and natural dialogue flow. We introduce PragAlign, a feedback-guided framework for controlled syn…

    arxiv.org15 days agoView details

  131. Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

    arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder…

    arxiv.org15 days agoView details

  132. Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents

    arXiv:2609.02129v1 Announce Type: new Abstract: Data-centric agents repeatedly perform a discovery step before planning or execution: identifying the data objects relevant to a task. Yet successful discovery outcomes are typically discarded rather than reused. We introduce persistent discovery context, a lightweight m…

    arxiv.org15 days agoView details

  133. Language Models Can Control Their Own Attention

    arXiv:2609.02737v1 Announce Type: new Abstract: Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to find the few tokens that matter. If the user asks about a previous detail in a 1M-token conversation, global attention layers must scan the full context to…

    arxiv.org15 days agoView details

  134. Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks

    arXiv:2609.02399v1 Announce Type: new Abstract: Argumentation frameworks are useful tools for representing and reasoning with information in a variety of settings, e.g. in supplementing AI models as they perform classification tasks, with a notable benefit of providing additional explainability. In this paper, we intr…

    arxiv.org15 days agoView details

  135. CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

    arXiv:2609.02459v1 Announce Type: new Abstract: We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requ…

    arxiv.org15 days agoView details

  136. Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting

    arXiv:2609.02649v1 Announce Type: new Abstract: Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge when deploying NLP systems in real-world industrial settings. While monolithic Large Language Model (LLM) agents offer unbounded expressivity for tasks like Root Cause…

    arxiv.org15 days agoView details

  137. Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

    arXiv:2609.02750v1 Announce Type: new Abstract: Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of externa…

    arxiv.org15 days agoView details

  138. EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

    arXiv:2609.01611v1 Announce Type: new Abstract: Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component…

    arxiv.org15 days agoView details

  139. Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI

    arXiv:2609.01685v1 Announce Type: new Abstract: With the development of artificial intelligence (AI), the landscape of meta-ethics, which has largely centred on human ethics, faces pressures that may significantly reconfigure it. In particular, if future AI systems were to exhibit sufficiently integrated capacities fo…

    arxiv.org15 days agoView details

  140. MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

    arXiv:2609.02379v1 Announce Type: new Abstract: While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We int…

    arxiv.org15 days agoView details

  141. Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?

    arXiv:2609.01924v1 Announce Type: new Abstract: Recent work identifies a mid-depth band of verbalisable, causally potent representations in a standard feedforward transformer --- a functional analogue of a global workspace. Whether the same workspace functionality emerges when depth is implemented through recurrence r…

    arxiv.org15 days agoView details

  142. EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

    arXiv:2609.02133v1 Announce Type: new Abstract: Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation. We formulate this as response-side affective-orientation control and use multi-annotator emoji distributions as weak affe…

    arxiv.org15 days agoView details

  143. Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

    arXiv:2609.02760v1 Announce Type: new Abstract: On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no longer a reliable p…

    arxiv.org15 days agoView details

  144. SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

    arXiv:2609.02786v1 Announce Type: new Abstract: The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with the environment. This exposes them to safety risks in both harmful final responses and multi-step execution trajectories. Existing safety alignment mechanis…

    arxiv.org15 days agoView details

  145. PolERo: Studying Political Evasion in Romanian

    arXiv:2609.02391v1 Announce Type: new Abstract: Political evasion refers to responses that engage with a question while withholding the requested information. Recent NLP work frames political evasion as a classification task using a two-level taxonomy of response clarity and fine-grained evasion strategies. Existing w…

    arxiv.org15 days agoView details

  146. When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models

    arXiv:2609.02438v1 Announce Type: new Abstract: Large language models can look capable of logical reasoning, but correct or incorrect answers alone tell us little about what the model represents internally. We study logical verification in five open-weight transformer models using matched valid--invalid premise--claim…

    arxiv.org15 days agoView details

  147. How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling

    arXiv:2609.02482v1 Announce Type: new Abstract: In this paper, we analyze how Large Language Models (LLMs) employ worldbuilding strategies, focusing on setting as one measurable dimension of storyworld construction. We compare 1,000 AI-generated stories per model in English and German with human-authored fiction from…

    arxiv.org15 days agoView details

  148. Collective creativity in hybrid societies

    arXiv:2609.02620v1 Announce Type: new Abstract: Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding of creativity itself. Researchers disagree about whether these tools enrich or impoverish culture, and we argue that much of that disagreement comes from conflating…

    arxiv.org15 days agoView details

  149. Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

    arXiv:2609.02264v1 Announce Type: new Abstract: Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the $N \times N$ adjacency…

    arxiv.org15 days agoView details

  150. Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation

    arXiv:2609.02396v1 Announce Type: new Abstract: Radiology reports are written primarily for clinicians, and their specialized terminology often makes them difficult for patients to interpret. As a result, many patients turn to publicly available Large Language Models (LLMs) to help explain their reports, despite well-…

    arxiv.org15 days agoView details

  151. TaRA: Training-Aware Low-Rank Adaptation Initialization

    arXiv:2609.02639v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT), yet its performance is highly sensitive to initialization due to the information bottleneck imposed by low-rank decomposition. Existing approaches attempt to construct h…

    arxiv.org15 days agoView details

  152. SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

    arXiv:2609.02292v1 Announce Type: new Abstract: The rapid proliferation of large language models (LLMs) and the growing diversity of their applications presents a unique optimization opportunity: selecting the right model for the task, while optimizing for speed, cost, and quality at a per-task level. However, inferen…

    arxiv.org15 days agoView details

  1. dealignai/GLM-5.3-CYBERSECURITY-FP8

    text-generation · safetensors · glm_moe_dsa · abliterated

    huggingface.co16 days ago487 ptsView details

  2. OpenVDN/vdn-minimax-h3

    text-to-video · diffusers · safetensors · text-to-video

    huggingface.co15 days ago291 ptsView details

  3. Qwen/Qwen-Drive-1.0-4B

    image-text-to-text · transformers · safetensors · qwen_drive

    huggingface.co16 days ago209 ptsView details

  4. ARTPARK-IISc/SraVaani-1.0

    automatic-speech-recognition · sravaani_tdt · automatic-speech-recognition · custom_code

    huggingface.co16 days ago75 ptsView details

  5. Jojocodex/wushu-action-v7-minimax-h3-fl2va-ref2va-lora

    image-to-video · minimax-h3 · text-to-video · image-to-video

    huggingface.co16 days ago28 ptsView details

  6. Prannesshkva/QU-SSM-15M

    text-generation · safetensors · qu_ssm · qu-ssm

    huggingface.co16 days ago1 ptsView details

  7. h3rb3rn/moe-sovereign-student-4b

    text-generation · transformers · gguf · compound-ai

    huggingface.co16 days ago1 ptsView details

  8. cunba-ai/istation

    onnx · safetensors · gguf

    huggingface.co16 days agoView details

  9. IamPradeep/Bonsai-27B-GGUF-Colab-Prebuilt-GPU

    gguf · endpoints_compatible · region:us

    huggingface.co16 days agoView details

  1. shadcn-ui/lint

    An agent-first linter for Tailwind design systems. Write design system rules that agents can verify.

    github.com15 days ago33 ptsView details

  2. ggml-org/llama.cpp b10775

    <details open> mtmd: fix idefics3 preproc (#28273) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/44875193> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/downlo…

    github.com15 days agoView details

  3. ggml-org/llama.cpp b10756

    <details open> vulkan : only request VK_KHR_shader_bfloat16 extension if supported (#28155) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/44637403> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://gith…

    github.com16 days agoView details

  4. ggml-org/llama.cpp b10754

    <details open> opencl: fix out‐of‐bound reads in the Adreno image kernels (#27632) * opencl: clamp the q4_K decode GEMV's fetch row on a padded x-grid * opencl: enforce the tiling contract of the image KQ/KQV GEMMs * opencl: decide the image KQ/KQV split at the dispatch, not fro…

    github.com16 days agoView details

  5. ggml-org/llama.cpp b10753

    <details open> hexagon: add missing FARF logs for cpy/get_rows/set_rows/gdn ops (#28217) * hexagon: fix bug ne[2] printed in proc_op_req prep-src log * hexagon: add shape/VTCM farf logs to cpy, get/set rows, gdn </details> **Website:** - <https://llama.app> **Attestations:** - <…

    github.com16 days agoView details

  6. langchain-ai/langchain langchain==1.4.0a4

    Initial release release(langchain): 1.4.0a4 test(langchain): cover mixed-era ClientGroup and group elicitation Update libs/langchain_v1/langchain/mcp/adapter.py fix(langchain): drive MCP elicitation via member session for fastmcp 4.0.1 fix(sdk): use latest fastmcp and rm reentra…

    github.com16 days agoView details