Skip to content

Archive / 2026-09-03

September 3, 2026

  1. GPT-6 Astra: A new generation of intelligence

    Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science.

    openai.com14 days ago2278 ptsView details

  2. Audacity 4.0

    github.com14 days ago1152 ptsView detailsJoin discussion

  3. Qwen 3.8 27B available on Cerebras at 1500 tokens/s

    inference-docs.cerebras.ai14 days ago689 ptsView detailsJoin discussion

  4. Babylonian Lamb Stew with Beets (1750–1730 BCE)

    babylonian-collection.yale.edu14 days ago171 ptsView detailsJoin discussion

  5. Grok outage

    status.x.ai14 days ago159 ptsView detailsJoin discussion

  6. Gloria Steinem has died

    theguardian.com14 days ago148 ptsView detailsJoin discussion

  7. Show HN: Reactor Atlas

    reactoratlas.com14 days ago79 ptsView detailsJoin discussion

  8. Why office workers are turning against AI

    bloodinthemachine.com14 days ago27 ptsView detailsJoin discussion

  9. Models Don't Go Rogue

    mail.cyberneticforests.com14 days ago26 ptsView detailsJoin discussion

  10. Meta wanted to reduce teams by 60% because of AI

    newsletter.pragmaticengineer.com14 days ago14 ptsView detailsJoin discussion

  11. Pause AI Development Now

    twitter.com14 days ago12 ptsView detailsJoin discussion

  12. Open AI X post on Astra

    twitter.com14 days ago12 ptsView detailsJoin discussion

  13. Primary source

    Daybreak for Frontline Defenders: $1B to protect essential services

    OpenAI introduces Daybreak for Frontline Defenders. A $1 billion commitment expands access to frontier cyber AI, training, and support for essential services.

    openai.com14 days agoView details

  14. Primary source

    Playco cut manual fixes 50% prototyping games with GPT-6 Astra

    Using GPT-6 Astra, Playco built three themed game prototypes from one grey box foundation and reported 50% fewer manual fixes than with the previous model.

    openai.com14 days agoView details

  15. Primary source

    Legora reviewed 41 documents in minutes with GPT-6 Astra

    Legora used GPT-6 Astra to review 41 documents in minutes, find all four planted errors, and improve performance by nearly 40% in this financial-review workflow.

    openai.com14 days agoView details

  16. Primary source

    Transfer learning for genomic prediction in underrepresented populations

    General Science

    research.google14 days agoView details

  17. Primary source

    A connectomics milestone: Mapping the complete male fruit fly brain

    General Science

    research.google14 days agoView details

  18. Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour

    WeatherNext 3 ingests live geostationary satellite mosaics, refreshes hourly, and outputs 5 km forecasts across Search, Gemini, Maps. The post Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour appeared first on MarkTechPost.

    marktechpost.com14 days agoView details

  19. OpenAI Releases GPT-6 Astra: A 1.05M-Context Computer-Use Model Gated Behind a ‘Critical’ Cyber Threshold

    OpenAI released GPT-6 Astra on September 3, 2026, positioning it as a computer-use flagship rather than a chat model. It reports 72.6% on OSWorld V2-Offline, replaces Codex compaction with searchable notes, and ships a 1.05M-token context at $10/$50 per million tokens. It is also the first OpenAI model to cross the Cr…

    marktechpost.com14 days agoView details

  20. Anthropic Released Claude Commerce Agents: An Apache-2.0 Blueprint for Shopping and Merchant Agents Across Retail, Travel, Telecom and Entertainment

    Most teams building a shopping assistant or agent rebuild the same scaffolding: an agent loop, a tool layer over the catalog, an approval gate, and an eval suite. Anthropic has now released that scaffolding as code. This week, they published anthropics/commerce-agents, a reference blueprint containing a shopping agent…

    marktechpost.com14 days agoView details

  21. Meta AI Released Muse Spark 1.3: An Agentic Coding Model That Uses ~20% Fewer Tool Calls and ~25% Fewer Tokens Than Muse Spark 1.2

    Perplexity has shipped hybrid compute for its Mac app, splitting a single Perplexity Computer task between frontier models in the cloud and a compact model running on the user's machine. Tasks start in the cloud for search, planning and reasoning, then hand sensitive steps down to the Mac without restarting or losing…

    marktechpost.com14 days agoView details

  22. Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon

    Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, it averages 1.23x MLX-LM's prefill throughput and 1.35x its decode throughput on a 40-core, 128 GB M5 Max. The post Perplexity Open Source…

    marktechpost.com15 days agoView details

  23. Nobody Is Saying Why OpenAI and Anthropic Had Outages Today

    ChatGPT, Claude, and Grok all suffered outages at nearly the exact same time for reasons that remain murky.

    wired.com14 days agoView details

  24. Prediction Market Betting Is Getting People Banned and Arrested

    This week on Uncanny Valley, we dig into the latest prediction market buzz, Flock’s AI-powered police search tool, and how tech bros don’t know how to talk about “rouge” AI agents

    wired.com14 days agoView details

  25. GPT-6 Astra Is Here—and OpenAI Thinks It May Kick Off the AGI Era

    OpenAI leaders think the company’s next generation model, which excels at computer use and coding, may mark a major milestone in AI development.

    wired.com14 days agoView details

  26. OpenAI’s next big AI model has ‘entered the AGI era’

    OpenAI's next big model is here: GPT-6 Astra. The company calls it a "generational leap in capability" for areas like cybersecurity, professional work, software engineering, science, and computer use. As OpenAI announced earlier this week, it's also the first model designated as meeting OpenAI's "critical cybersecurit…

    theverge.com14 days agoView details

  27. OpenAI Cut Off a Billion-Dollar Customer to Avoid Elon Musk

    OpenAI recently estimated its Cursor partnership would make more than $1 billion in revenue a year, WIRED has learned. It still walked away after Elon Musk’s SpaceX acquired the AI coding startup.

    wired.com14 days agoView details

  28. Nvidia launches free tool that links idle computers into a personal AI data center

    Nvidia is announcing its new Personal AI Router (PAIR), a free tool that syncs up your home computers for tackling local AI inference tasks with tools like Ollama and LM Studio. Let's get the obvious thing out of the way, despite what its name might imply: PAIR is not a hardware router. It's open-source software devel…

    theverge.com14 days agoView details

  29. Google now lets you chat with Gmail, Docs, and Keep

    Google is rolling out AI-powered voice assistant modes for Gmail, Docs, and Keep that allow you to manage the apps by talking to them. The real time conversational capabilities are called Gmail Live, Docs Live, and Keep Live, and like the Gemini Live experience for Google's chatbot, aim to make it easier to note down…

    theverge.com14 days agoView details

  30. ChatGPT, Grok, and Claude all went down at the same time

    OpenAI's ChatGPT, xAI's Grok, and Anthropic's Claude are back online after they all began experiencing issues around the same time on Thursday. At about 11AM ET, ChatGPT started returning error messages for users trying to use the chatbot, with its status page saying there were "elevated errors across ChatGPT and Code…

    theverge.com14 days agoView details

  31. Google says its AI weather model is getting better

    Google is rolling out an updated AI weather model that's supposed to be more accurate, especially when it comes to predicting rain and snowfall. In the announcement today, the company says it's now able to make forecasts with "unprecedented resolution" using its new WeatherNext 3 AI model. It can produce a global pict…

    theverge.com14 days agoView details

  32. Nvidia RTX Spark ‘Superchip’: The First AI PCs Are Here

    At IFA 2026, Nvidia and its partners showed off the first RTX Spark-powered laptops and mini PCs, designed to run AI models right on your computer.

    wired.com14 days agoView details

  33. Nvidia’s Hugging Face Acquisition Is a $12.9 Billion Bet on Open-Source AI

    The long-rumored deal will give the chip giant access to—and help it promote—a huge repository of open-source AI models and data sets.

    wired.com14 days agoView details

  34. Nvidia is buying Hugging Face for almost $13 billion

    Nvidia has agreed to buy Hugging Face for $12.93 billion, bringing one of the most popular hosting platforms for open-source AI models, datasets, and tools under the ownership of the world's biggest AI chipmaker. Hugging Face is an online platform founded in 2016 that gives AI developers a space to share their project…

    theverge.com14 days agoView details

  35. This Is Flock’s AI Search Tool for Cops

    WIRED rebuilt Flock’s latest search tool from code the company sends to a police officer’s browser. Its AI can keep watch across multiple cameras for anyone fitting a written description.

    wired.com14 days agoView details

  36. MasterControl Seventeen Every Time

    arXiv:2609.03209v1 Announce Type: new Abstract: We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive withi…

    arxiv.org14 days agoView details

  37. Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

    arXiv:2609.03340v1 Announce Type: new Abstract: Distributed LLM-agent teams can read the latest shared facts and still act on an obsolete plan. A planner may derive an action from requirement $r_3$, another agent may commit $r_4$, and an executor may receive $r_4$ without replacing the plan derived from $r_3$. We call…

    arxiv.org14 days agoView details

  38. Dalek: A Constructive Agent Machine

    arXiv:2609.03546v1 Announce Type: new Abstract: We present Dalek, a closed machine designed for agents that realizes self-maintenance, self-evolution, self-reproduction, and self-organization on any substrate satisfying a general host contract. The machine is built from three primitives---actors, messages, and channel…

    arxiv.org14 days agoView details

  39. Opening mind by opening architecture: analysis strategies

    arXiv:2609.03719v1 Announce Type: new Abstract: In numerical signal processing for electroacoustic composition, the progressive loss of specific development and research environments caused by the increasing use of digital market tools has favoured the dominance of the closed-architecture audio processor model. This m…

    arxiv.org14 days agoView details

  40. Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

    arXiv:2609.02940v1 Announce Type: new Abstract: Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However, how to effectively combine these two str…

    arxiv.org14 days agoView details

  41. Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable

    arXiv:2609.04127v1 Announce Type: new Abstract: Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncert…

    arxiv.org14 days agoView details

  42. BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events

    arXiv:2609.02895v1 Announce Type: new Abstract: Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformation, posing substantial risks to public safety and social cohesion. While automated fake news detectio…

    arxiv.org14 days agoView details

  43. GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

    arXiv:2609.03494v1 Announce Type: new Abstract: Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-request budget and adjust only which KV states are retained, leaving the total capacity fixed…

    arxiv.org14 days agoView details

  44. Lose the Order, Keep the Hierarchy: Deordering HTN Plans

    arXiv:2609.03912v1 Announce Type: new Abstract: Hierarchical Task Network (HTN) planning is a powerful planning formalism based on task decomposition. Although most of the literature studied plan generation, comparatively less attention has been paid to post-plan optimization. In particular, plan deordering has been e…

    arxiv.org14 days agoView details

  45. Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

    arXiv:2609.03887v1 Announce Type: new Abstract: How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify…

    arxiv.org14 days agoView details

  46. Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

    arXiv:2609.03702v1 Announce Type: new Abstract: General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (se…

    arxiv.org14 days agoView details

  47. A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant

    arXiv:2609.03402v1 Announce Type: new Abstract: Artificial intelligence (AI) teaching assistants powered by large language models (LLMs) offer scalable educational support but often provide limited personalization. This study presents a prompt-engineering-based framework for personalizing general-purpose LLM/RAG-based…

    arxiv.org14 days agoView details

  48. Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

    arXiv:2609.03416v1 Announce Type: new Abstract: LLM-empowered paper-code discrepancy detection has received growing concern since the scaling of research submissions exceeds the manual review capability. However, the limited context capacity and one-sided discrepancy detection of existing single-agent LLM paradigms le…

    arxiv.org14 days agoView details

  49. DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

    arXiv:2609.03423v1 Announce Type: new Abstract: Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often co…

    arxiv.org14 days agoView details

  50. Language, Language Models, and What We're Talking About

    arXiv:2609.03577v1 Announce Type: new Abstract: Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to the nature of the systems which result fr…

    arxiv.org14 days agoView details

  51. A computable representation of the physical laboratory enables verifiable workflows

    arXiv:2609.03621v1 Announce Type: new Abstract: Making science computable requires representations of both scientific knowledge and the physical world in which scientific claims are tested. A computable representation of the physical laboratory is established through typed research objects, capability-bound operations…

    arxiv.org14 days agoView details

  52. Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

    arXiv:2609.03430v1 Announce Type: new Abstract: Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how…

    arxiv.org14 days agoView details

  53. Pattern Over-Generalization of Knowledge Graph Embedding

    arXiv:2609.03487v1 Announce Type: new Abstract: Knowledge graph embedding (KGE) demonstrates its effectiveness for predicting missing links in knowledge graphs (KGs) by projecting entities and relations into a low-dimensional vector space. It is crucial for KGE models to effectively capture inference patterns (pattern…

    arxiv.org14 days agoView details

  54. The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification

    arXiv:2609.03652v1 Announce Type: new Abstract: Synthetic data augmentation has become a common strategy for addressing class imbalance in NLP, but most approaches focus on the quantity and diversity of generated examples rather than their geometric relationship to real training data. We investigate this question in t…

    arxiv.org14 days agoView details

  55. Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition

    arXiv:2609.02901v1 Announce Type: new Abstract: Modern automatic speech recognition (ASR) scenarios require both spoken-form transcripts for faithful transcription and readable written-form transcripts with inverse text normalization (ITN). However, these forms are typically produced by cascaded modules, where a spoke…

    arxiv.org14 days agoView details

  56. Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence

    arXiv:2609.02981v1 Announce Type: new Abstract: Artificial intelligence is changing the form of applied English materials from fixed paper sequences to adaptive learning systems that can diagnose learners, recommend tasks, and provide formative feedback. This paper studies the structure and application of a new practi…

    arxiv.org14 days agoView details

  57. Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

    arXiv:2609.03438v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to benign mistakes. A reliable agent should not only know how to act, but also when not to act. I…

    arxiv.org14 days agoView details

  58. RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

    arXiv:2609.02902v1 Announce Type: new Abstract: Deploying task-oriented dialogue agents in enterprise customer support faces a persistent annotation bottleneck: robust training requires labelled interaction data at scale, yet enterprise conversational logs are privacy-sensitive and expensive to annotate, while user be…

    arxiv.org14 days agoView details

  59. Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor

    arXiv:2609.03221v1 Announce Type: new Abstract: Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that th…

    arxiv.org14 days agoView details

  60. NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis

    arXiv:2609.03527v1 Announce Type: new Abstract: Neonatal respiratory diseases are a major cause of neonatal morbidity and mortality, posing substantial challenges in clinical practice. Despite recent advances, existing Multimodal Large Language Models (MLLMs) face two key limitations in neonatal diagnosis: (1) domain…

    arxiv.org14 days agoView details

  61. Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

    arXiv:2609.03407v1 Announce Type: new Abstract: People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptio…

    arxiv.org14 days agoView details

  62. TabScope: Question-Adaptive Scope Selection for Table Question Answering

    arXiv:2609.03395v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy often degrades as table size increases. We find that this degradation is not uniform across question types. Localization-sensitive questions are particularly affect…

    arxiv.org14 days agoView details

  63. The Attention Triangle in Audio-Video Models

    arXiv:2609.03586v1 Announce Type: new Abstract: Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising t…

    arxiv.org14 days agoView details

  64. FrameBench:A Language Understanding Benchmark Based on Frame Semantics

    arXiv:2609.03370v1 Announce Type: new Abstract: In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent large language models (LLMs) have achieve…

    arxiv.org14 days agoView details

  65. Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty

    arXiv:2609.03460v1 Announce Type: new Abstract: As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. B…

    arxiv.org14 days agoView details

  66. Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation

    arXiv:2609.03535v1 Announce Type: new Abstract: Lesion segmentation in medical images plays a critical role in clinical diagnosis and treatment planning. Despite significant advances, lesion segmentation remains challenging due to two major factors: (1) complex background interference; (2) diverse lesion morphology. E…

    arxiv.org14 days agoView details

  67. Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

    arXiv:2609.02897v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static dra…

    arxiv.org14 days agoView details

  68. What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation

    arXiv:2609.03254v1 Announce Type: new Abstract: Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies an…

    arxiv.org14 days agoView details

  69. Speculative Macro Commit for Faster Tool-Using Agents

    arXiv:2609.03236v1 Announce Type: new Abstract: Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and observation can delay subsequent decisions. We introduce \textbf{Speculative Macro Commit} (SMC), a run…

    arxiv.org14 days agoView details

  70. Analysis of Prompt Engineering for Drug Toxicity Prediction

    arXiv:2609.03635v1 Announce Type: new Abstract: Clinical trials in the UK can cost up to {\pounds}1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence has been increasingly…

    arxiv.org14 days agoView details

  71. CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception

    arXiv:2609.03818v1 Announce Type: new Abstract: Collaborative perception enhances environment understanding through multi-agent information sharing, but its performance in real-world scenarios is constrained by heterogeneous sensor modalities and model architectures. Recent protocol-based two-stage methods alleviate t…

    arxiv.org14 days agoView details

  72. Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

    arXiv:2609.03493v1 Announce Type: new Abstract: Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cro…

    arxiv.org14 days agoView details

  73. MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval

    arXiv:2609.03201v1 Announce Type: new Abstract: Long-term LLM agents must preserve information across interactions while distinguishing repeated evidence, historical states, updates, and unresolved contradictions. Existing textual memory systems retrieve semantically relevant memories efficiently but often leave these…

    arxiv.org14 days agoView details

  74. Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations

    arXiv:2609.03426v1 Announce Type: new Abstract: Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled wit…

    arxiv.org14 days agoView details

  75. Artificial Intelligence for Energy Optimization in Data Centers

    arXiv:2609.03716v1 Announce Type: new Abstract: Data centers are increasingly optimized by artificial intelligence and, at the same time, increasingly loaded by it. The literature treats these as two unrelated problems: control studies model workload as an exogenous arrival process, while sustainability studies model…

    arxiv.org14 days agoView details

  76. AutoGraphForge: Towards Automated Graph Theory Discovery

    arXiv:2609.03478v1 Announce Type: new Abstract: We report on our ongoing project to develop a computational pipeline, AutoGraphForge, for an automated graph-theoretic conjecturing-refuting-formalizing-proving system. Conjecture generation is counterexample-guided and runs in rounds: a Graffiti3 generator proposes conj…

    arxiv.org14 days agoView details

  77. How Far Can Synthetic Data Take Thai OCR?

    arXiv:2609.03595v1 Announce Type: new Abstract: We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "r…

    arxiv.org14 days agoView details

  78. Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

    arXiv:2609.03502v1 Announce Type: new Abstract: In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programma…

    arxiv.org14 days agoView details

  79. PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing

    arXiv:2609.03503v1 Announce Type: new Abstract: With the rapid development of the Internet of Things, computation intensive directed acyclic graph (DAG) tasks have become increasingly common in cloud-edge-end collaborative environments. However, cloud, edge, and end nodes are highly heterogeneous in computing capacity…

    arxiv.org14 days agoView details

  80. CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

    arXiv:2609.03526v1 Announce Type: new Abstract: Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4…

    arxiv.org14 days agoView details

  81. Rethinking World Models for Safety-Critical Embodied Systems

    arXiv:2609.03774v1 Announce Type: new Abstract: World models have progressed from compact latent dynamics to generative, controllable, and interactive simulators of embodied environments. However, high predictive likelihood and visual fidelity do not necessarily ensure that a model preserves the evidence required for…

    arxiv.org14 days agoView details

  82. Counterfactual Routing Using Integer Programming with Constraint Generation

    arXiv:2609.03707v1 Announce Type: new Abstract: We present our submission to the IJCAI 2025 'Counterfactual Routing Competition' (CRC 25). The goal of the competition is to find counterfactual explanations for the shortest path problem. This requires deciding what the minimal changes to a road network would make a rou…

    arxiv.org14 days agoView details

  83. What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation

    arXiv:2609.03515v1 Announce Type: new Abstract: Decoding-time KV cache compression research focuses heavily on designing better token scoring functions, while the temporal rule that aggregates scores across decode steps is often treated as an implementation detail. Under aggressive KV compression, we find that exponen…

    arxiv.org14 days agoView details

  84. KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

    arXiv:2609.03588v1 Announce Type: new Abstract: As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledg…

    arxiv.org14 days agoView details

  85. Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations

    arXiv:2609.03860v1 Announce Type: new Abstract: Retail supply chain operations rely on coupled decision modules that must adapt as requirements evolve. LLMs offer a natural-language interface for this task, but existing methods primarily focus on individual optimization models. Extending them to heterogeneous decision…

    arxiv.org14 days agoView details

  86. Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

    arXiv:2609.03727v1 Announce Type: new Abstract: Large language model agents can plan, invoke tools, and modify external states, yet most systems still take an explicit user instruction as a fixed starting point. Proactive service moves the decision upstream: an agent must infer service opportunities from incomplete en…

    arxiv.org14 days agoView details

  87. SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation

    arXiv:2609.03753v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the long-term value of AI systems depends not only on solving individual requests, but also on transforming experience and accumulated knowledge into durable, reusable competence. We introduce SimSkill, a self-…

    arxiv.org14 days agoView details

  88. Inferring Affective Consciousness in an Artificial Agent: A Case Study

    arXiv:2609.03883v1 Announce Type: new Abstract: Creatures that display 'hedonic place preference behaviour' are thought by many scientists to experience feelings, on the assumption that their attraction to pleasure-producing substances which lack nutritional value (e.g. cocaine, morphine) cannot easily be attributed t…

    arxiv.org14 days agoView details

  89. Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding

    arXiv:2609.03938v1 Announce Type: new Abstract: While HTN planning has received significant attention in recent years, support for numerical reasoning remains very limited. In this paper, we investigate numerical Totally-Ordered HTN (TOHTN) planning and show how standard SAT-based encodings can be naturally extended w…

    arxiv.org14 days agoView details

  90. InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models

    arXiv:2609.04014v1 Announce Type: new Abstract: For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high repeatability. Yet Multimodal Large Language Models (MLLMs) remain unreliable in continuous-valued measurement despite strong results on general multimodal benchmar…

    arxiv.org14 days agoView details

  91. FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models

    arXiv:2609.04021v1 Announce Type: new Abstract: Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically…

    arxiv.org14 days agoView details

  92. Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

    arXiv:2609.04098v1 Announce Type: new Abstract: Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent state summarizes the context in fixed size. Early community 4-bit quantizations of Qwen3.8-27B (48 GDN layers, 16 attention layers) left the GDN block in 8- or 1…

    arxiv.org14 days agoView details

  93. Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

    arXiv:2609.02899v1 Announce Type: new Abstract: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scor…

    arxiv.org14 days agoView details

  94. LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation

    arXiv:2609.02954v1 Announce Type: new Abstract: Identifying the issues disputed between litigating parties is a crucial component of real-world litigation. However, legal issues remain comparatively underexplored in legal AI research. In this work, we study the computational modelling of legal issue identification in…

    arxiv.org14 days agoView details

  95. No country for old linguists: LLM-brain alignment underdetermines neural computation

    arXiv:2609.03160v1 Announce Type: new Abstract: Nastase et al. (2026) argue that large language models (LLMs) may illuminate language processing because both rely on distributed, context-sensitive representations shaped by statistical learning. Their rejection of simple cortical "boxology" is persuasive, and they arti…

    arxiv.org14 days agoView details

  96. A Circuit for Plural Reference: How LLMs Represent and Retrieve Singular and Plural Entities

    arXiv:2609.03687v1 Announce Type: new Abstract: Coreference resolution is an important task in contextual reasoning. In this paper, we investigate the mechanism for representing and retrieving singular and plural entities for plural reference. We use a combination of mechanistic interpretability and attention pattern…

    arxiv.org14 days agoView details

  97. Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

    arXiv:2609.03734v1 Announce Type: new Abstract: BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language…

    arxiv.org14 days agoView details

  98. LLMs Learn Better In-Context from Rules than from Examples

    arXiv:2609.03213v1 Announce Type: new Abstract: Large language models (LLMs) exhibit in-context learning capabilities, where they can learn new tasks from prompt contexts without weight updates. We compare the learning efficacies of two prominent modes of in-context learning: (1) learning from descriptions of rules (i…

    arxiv.org14 days agoView details

  99. SWIM: Student Writing Simulation via Proficiency-Conditioned Generation

    arXiv:2609.03215v1 Announce Type: new Abstract: Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplo…

    arxiv.org14 days agoView details

  100. Semantic Bayesian World Models

    arXiv:2609.03834v1 Announce Type: new Abstract: Knowledge graphs describe reality in crisp assertions, while the systems now consuming them, foundation models and autonomous agents, reason natively in probabilities. We argue that this mismatch is why the integration of language models and knowledge graphs remains a da…

    arxiv.org14 days agoView details

  101. SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation

    arXiv:2609.03806v1 Announce Type: new Abstract: Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics design…

    arxiv.org14 days agoView details

  102. Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

    arXiv:2609.03814v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply eac…

    arxiv.org14 days agoView details

  103. Value-Preserving Architectures for Agentic AI Systems

    arXiv:2609.03920v1 Announce Type: new Abstract: The emergence of agentic AI and LLM-based multi-agent systems (MAS) presents unprecedented opportunities for automating complex tasks, while simultaneously raising critical concerns about the preservation of fundamental human-centered values, such as privacy, fairness, a…

    arxiv.org14 days agoView details

  104. Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

    arXiv:2609.03923v1 Announce Type: new Abstract: In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stances, coverage, and floor, they miss the moments where they should contribute. Prompt-only delegates stay silent on 51.4% of the absent participant's talking oppo…

    arxiv.org14 days agoView details

  105. More Criticism Does Not Make a Better Review: EquiReview-R

    arXiv:2609.03943v1 Announce Type: new Abstract: AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet gen…

    arxiv.org14 days agoView details

  106. FiMI Banking: A Sovereign Model for Indian Retail Banking

    arXiv:2609.03960v1 Announce Type: new Abstract: Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably meet these requirements. They fal…

    arxiv.org14 days agoView details

  107. The Dually Flat Geometry of Planning as Inference

    arXiv:2609.04005v1 Announce Type: new Abstract: We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on…

    arxiv.org14 days agoView details

  108. Transfiver: Human-AI Co-Inference through a Shared Editable State

    arXiv:2609.03797v1 Announce Type: new Abstract: Long-term human-AI interaction is difficult because the information that guides inference is updated implicitly by the model and is not directly inspectable or controllable by the user. We introduce the TRANSparent Framework for Interactive, Verifiable, Editable Represen…

    arxiv.org14 days agoView details

  109. Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

    arXiv:2609.02942v1 Announce Type: new Abstract: LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only o…

    arxiv.org14 days agoView details

  110. Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory

    arXiv:2609.03394v1 Announce Type: new Abstract: Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows…

    arxiv.org14 days agoView details

  111. The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

    arXiv:2609.03218v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user context…

    arxiv.org14 days agoView details

  112. Interface-Induced Trajectory Censoring

    arXiv:2609.03966v1 Announce Type: new Abstract: Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model is emitting well-formed calls: the interface censors the trajectory before anything downstream sees it. On BFCL v4's own data, executor and scorer, holding weigh…

    arxiv.org14 days agoView details

  113. Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing

    arXiv:2609.03973v1 Announce Type: new Abstract: An image editor may satisfy every regional plausibility constraint separately even when no single latent explanation fits the complete output. We formalize this local-to-global failure using a common witness grade and witness nerve. The framework separates auditing from…

    arxiv.org14 days agoView details

  114. Distilled Rapid Embedding Transfer (DRET): Parameter-Efficient Biomedical Domain Adaptation via Priority-Based Embedding Transfer

    arXiv:2609.02898v1 Announce Type: new Abstract: Large domain-specific language models such as BioBERT and ClinicalBERT achieve strong performance on biomedical NLP tasks, but their computational demands make them impractical for many real-world deployments. General-purpose, parameter-efficient models such as DistilBER…

    arxiv.org14 days agoView details

  115. Bioinfoysis Technical Report

    arXiv:2609.03871v1 Announce Type: new Abstract: Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics…

    arxiv.org14 days agoView details

  116. Instruction Duplication as an Inference-Time Control Primitive

    arXiv:2609.04024v1 Announce Type: new Abstract: Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats onl…

    arxiv.org14 days agoView details

  117. IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

    arXiv:2609.04030v1 Announce Type: new Abstract: IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which ad…

    arxiv.org14 days agoView details

  118. Spurious Advantage Hidden in GRPO

    arXiv:2609.04063v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each rollout a magnitude from within-group reward statistics. In the common case, this magnitude rewards rollouts that re…

    arxiv.org14 days agoView details

  119. Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

    arXiv:2609.03181v1 Announce Type: new Abstract: We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative de…

    arxiv.org14 days agoView details

  120. Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue

    arXiv:2609.03321v1 Announce Type: new Abstract: The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializing turn-taking control and response generation onto a single causal tape under the standard next-token prediction objective, thereby preserving semantic prowess a…

    arxiv.org14 days agoView details

  121. How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

    arXiv:2609.03322v1 Announce Type: new Abstract: Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at…

    arxiv.org14 days agoView details

  122. Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour

    arXiv:2609.03330v1 Announce Type: new Abstract: Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large…

    arxiv.org14 days agoView details

  123. FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models

    arXiv:2609.03331v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually grounded false premise persists across dia…

    arxiv.org14 days agoView details

  124. Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT

    arXiv:2609.03366v1 Announce Type: new Abstract: Accountability means a decision can be examined, justified, and contested. LLMs make this hard: fluent output may be ungrounded, incomplete, or unfaithful to the decision process. Achieving accountability requires verified rationales (how was the decision reached), assum…

    arxiv.org14 days agoView details

  125. SGD-KV: Summarization Guided KV Cache Compression

    arXiv:2609.03235v1 Announce Type: new Abstract: Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of dif…

    arxiv.org14 days agoView details

  126. Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers

    arXiv:2609.03273v1 Announce Type: new Abstract: Tamil spell and grammar correction is challenging because Tamil is an agglutinative low-resource language with rich verbal morphology, complex sandhi (phonetic transformation) rules at word boundaries, and a script of 247 distinct letters. Prior work targets word-level s…

    arxiv.org14 days agoView details

  127. When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

    arXiv:2609.03467v1 Announce Type: new Abstract: Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We i…

    arxiv.org14 days agoView details

  128. Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

    arXiv:2609.03432v1 Announce Type: new Abstract: Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a comp…

    arxiv.org14 days agoView details

  129. KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records

    arXiv:2609.03597v1 Announce Type: new Abstract: Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system with dedicated Unicode glyphs, no mainstream font, and no coverage in any OCR pipeline or tokenizer. The handwritten records that carry these fractions, RS Khatian…

    arxiv.org14 days agoView details

  130. Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

    arXiv:2609.03619v1 Announce Type: new Abstract: Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of age…

    arxiv.org14 days agoView details

  131. Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations

    arXiv:2609.03511v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to semantics-preserving structural variation remains underexplored, particularly for relatively free word-order languages. We investigate the structural sensitivity o…

    arxiv.org14 days agoView details

  132. </think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

    arXiv:2609.03633v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strateg…

    arxiv.org14 days agoView details

  133. Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements

    arXiv:2609.03654v1 Announce Type: new Abstract: The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical content across different jurisdi…

    arxiv.org14 days agoView details

  134. PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction

    arXiv:2609.02896v1 Announce Type: new Abstract: Medical relation extraction (MRE) is commonly known for extracting entities and their relations jointly from a medical text, which has attracted considerable attention in recent years. Previous studies treat MRE as a sequence tagging task, which results in either a chall…

    arxiv.org14 days agoView details

  135. Typological Feature Prediction with Large Language Models: An In-Context Learning Approach

    arXiv:2609.03775v1 Announce Type: new Abstract: Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while their performance across resource levels a…

    arxiv.org14 days agoView details

  136. IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

    arXiv:2609.03781v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSaf…

    arxiv.org14 days agoView details

  137. GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

    arXiv:2609.03553v1 Announce Type: new Abstract: Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when…

    arxiv.org14 days agoView details

  138. HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

    arXiv:2609.03580v1 Announce Type: new Abstract: The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability. Existing hallucination benchmarks are not designed for peer review…

    arxiv.org14 days agoView details

  139. Unifying Conformal Language Tasks with In-Context Ensembles

    arXiv:2609.03005v1 Announce Type: new Abstract: Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pertinent information to achieve some goal, and conciseness, removing as much irrelevant inform…

    arxiv.org14 days agoView details

  140. DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions

    arXiv:2609.03787v1 Announce Type: new Abstract: AI agents increasingly gather evidence, invoke tools, apply constraints, and produce decisions that people or software may commit to action. A final output alone cannot show which evidence, tool state, rule, authorization, or action path produced it. We present DNative-T…

    arxiv.org14 days agoView details

  141. To What Extent Do Large Language Models Understand Bangla Idioms?

    arXiv:2609.03410v1 Announce Type: new Abstract: Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first large-scale benchmark dataset of Bangla idioms,…

    arxiv.org14 days agoView details

  142. Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents

    arXiv:2609.02889v1 Announce Type: new Abstract: A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strategy, format rules, and control heuristics. Existing reflective prompt-evolution methods usually opti…

    arxiv.org14 days agoView details

  143. Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent

    arXiv:2609.02890v1 Announce Type: new Abstract: A personalized language agent must convert a user's interaction history into behavior on each new request at inference time. Two strategies dominate. Retrieval pulls a few of the user's most relevant past items into the prompt, which is accurate but pays a per-query sele…

    arxiv.org14 days agoView details

  144. Counterexamples as Feedback for Agent Self-Correction

    arXiv:2609.02892v1 Announce Type: new Abstract: Single-turn code-generation metrics understate a central property of deployed agents: whether they can repair a wrong artifact after receiving concrete feedback. This paper presents A-CEGIS, a lightweight framework that uses counterexamples as feedback for evaluating mul…

    arxiv.org14 days agoView details

  145. Probe Generalization as Subspace Selection for OOD Deception Detection

    arXiv:2609.02893v1 Announce Type: new Abstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out deception detection datase…

    arxiv.org14 days agoView details

  146. R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG

    arXiv:2609.02894v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a prevailing paradigm for enhancing Large Language Models (LLMs) with non-parametric knowledge. Vanilla RAG efficiently handles simple queries but struggles with relational or multi-hop reasoning. Graph-based RAG alleviates…

    arxiv.org14 days agoView details

  147. Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI

    arXiv:2609.03800v1 Announce Type: new Abstract: Federated learning is increasingly presented as a privacy-preserving advance: personal data remain on the device, and only model updates are shared. It borrows the vocabulary of the federated social web, yet inverts its logic, distributing computation while the resulting…

    arxiv.org14 days agoView details

  148. SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

    arXiv:2609.03047v1 Announce Type: new Abstract: Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to run. SHELF, the Syn…

    arxiv.org14 days agoView details

  149. Large Language Models in Resolving Contextual Knowledge Conflicts

    arXiv:2609.03148v1 Announce Type: new Abstract: Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided context. In contrast, we investigate how LLMs handle conflicts that arise within contextual knowledge itself. We introduce a taxonomy of six types of contextual c…

    arxiv.org14 days agoView details

  150. STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

    arXiv:2609.03874v1 Announce Type: new Abstract: Retrieval Augmented Generation (RAG) is a key component for generating accurate and hallucination free answers using Large Language Models (LLMs). LLMs are improving at handling long context, but still suffer from "lost in the middle" problem. Thus, precise and accurate…

    arxiv.org14 days agoView details

  151. PACE: Towards Surfacing Hidden Conflicts in User Requests

    arXiv:2609.03293v1 Announce Type: new Abstract: Personalized assistants should not only comply with user requests but also assess whether those requests are appropriate given the user's current circumstances. However, prior work has primarily focused on accurately executing requests, overlooking the need for assistant…

    arxiv.org14 days agoView details

  152. Xiaomi-TabLDM: A Tabular Foundation Model Technical Report

    arXiv:2609.03880v1 Announce Type: new Abstract: We introduce Xiaomi-TabLDM, a tabular large data foundation model for classification and regression via in-context learning, which delivers superior prediction accuracy without requiring task-specific fine-tuning. Pretrained exclusively on synthetic data generated from s…

    arxiv.org14 days agoView details

  153. DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

    arXiv:2609.04094v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popul…

    arxiv.org14 days agoView details

  154. LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening

    arXiv:2609.04013v1 Announce Type: new Abstract: Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning (ML) and deep learning (DL) approaches require labeled data and model training, limiting their use in real-world screening settings. This study evaluates the ef…

    arxiv.org14 days agoView details

  155. When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA

    arXiv:2609.03454v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) can improve the specificity and grounding of large language model responses, but its effect is not uniformly beneficial in single-turn mental-health question answering, where user queries often combine emotional distress, treatment co…

    arxiv.org14 days agoView details

  1. IFM/K2-Horizon-MoVA-36B-A4B

    text-generation · transformers · safetensors · k2_horizon

    huggingface.co14 days ago293 ptsView details

  2. microsoft/VibeVoice-ASR-Streaming-7B

    automatic-speech-recognition · transformers · safetensors · vibevoice

    huggingface.co15 days ago213 ptsView details

  3. IFM/K2-Horizon-7B

    text-generation · transformers · safetensors · k2_horizon

    huggingface.co14 days ago104 ptsView details

  4. XHToken/Spark-X2.5-1.7B

    text-generation · transformers · safetensors · spark2_5

    huggingface.co15 days ago98 ptsView details

  5. inclusionAI/Ling-3.0-flash-Fin

    text-generation · safetensors · bailing_hybrid · finance

    huggingface.co14 days ago82 ptsView details

  6. unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF

    transformers · gguf · unsloth

    huggingface.co14 days ago71 ptsView details

  7. osk-arr00/lfm2.5-2.6B-thinkingcap-distiller

    text-generation · transformers · safetensors · gguf

    huggingface.co15 days ago4 ptsView details

  8. 2013khansohail/cartographer-ecommerce-reranker-MiniLM-L6-v2

    text-ranking · sentence-transformers · safetensors · bert

    huggingface.co15 days ago4 ptsView details

  9. nyxia/H3-Loras

    text-to-video · minimax-h3 · lora · text-to-video

    huggingface.co15 days ago4 ptsView details

  10. gguf-org/pig-clip

    gguf · license:mit · region:us

    huggingface.co15 days ago2 ptsView details

  11. osk-arr00/lfm2.5-2.6B-thinkingcap-distiller-v2

    text-generation · transformers · safetensors · gguf

    huggingface.co15 days ago1 ptsView details

  12. osk-arr00/LFM2.5-8B-A1B-ThinkingCap-GGUF

    text-generation · gguf · llama.cpp · rocm

    huggingface.co15 days ago1 ptsView details

  13. osk-arr00/BigBang-Aquila-35B-GGUF

    gguf · imatrix · qwen3_5_moe

    huggingface.co15 days agoView details

  14. smshahbaj/RIFA-FLASH-1.7B

    text-generation · transformers · safetensors · gguf

    huggingface.co15 days agoView details

  15. smshahbaj/RIFA-CODE-0.6B

    text-generation · transformers · safetensors · gguf

    huggingface.co15 days agoView details

  16. smshahbaj/RIFA-Edge-0.6B

    text-generation · transformers · safetensors · gguf

    huggingface.co15 days agoView details

  17. smshahbaj/Rifa-Nano-0.5B

    text-generation · transformers · safetensors · gguf

    huggingface.co15 days agoView details

  18. webai-community/ai-models

    onnx · gguf · endpoints_compatible

    huggingface.co15 days agoView details

  1. ggml-org/llama.cpp b10795

    <details open> sycl: fuse rms_norm+mul+add and add+add residual chains (#27610) Fuse RMS_NORM+MUL+ADD and ADD+ADD under GGML_SYCL_ENABLE_FUSION. ADD+ADD uses the same binbcast indexing and type matrix as standalone add() (f32, f16, f16/f32, i32, i16, bf16, including broadcast an…

    github.com14 days agoView details

  2. openai/openai-python v3.8.0

    ## [3.8.0](https://github.com/openai/openai-python/compare/v3.7.0...v3.8.0) (2026-09-03) ### Features * **api:** add gpt-6-astra and related features ([#3791](https://github.com/openai/openai-python/issues/3791)) ([09f446f](https://github.com/openai/openai-python/commit/09f446f5…

    github.com14 days agoView details

  3. langchain-ai/langchain langchain==1.4.0

    Changes since langchain==1.3.18 docs(langchain): runnable `langchain.mcp` examples (#39976) feat(langchain): `langchain.mcp` namespace, `MCPAdapter` (#39939) perf(anthropic,langchain): omit middleware trace inputs (#40098) fix(langchain): include model destination in agent tool…

    github.com14 days agoView details

  4. langchain-ai/langchain langchain-anthropic==1.7.1

    Changes since langchain-anthropic==1.7.0 release(anthropic): 1.7.1 (#40181) perf(anthropic,langchain): omit middleware trace inputs (#40098) feat(anthropic): add Claude Fable 5.1 support (#40106)

    github.com14 days agoView details

  5. ggml-org/llama.cpp b10777

    <details open> sycl: reduce redundant work in Q4_K multi-column MMVQ (#27062) * sycl: Q4_K Weight unpack optimization and reuse between destination Columns * sycl: Q4_K small N (N=2..4) + two output rows by subgroup reuse of activation between two rows. * sycl: gate Q4_K two-row…

    github.com15 days agoView details

  6. ggml-org/llama.cpp b10776

    <details open> model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) * hparams: add per-layer n_ff_exp/n_expert_used arrays with scalar-or-array loading G1/G2 infrastructure for variable-per-layer expert FFN size and top-k routing (required for Puzzle-75…

    github.com15 days agoView details