Skip to content

Archive / 2026-08-30

August 30, 2026

  1. No AI Fridays

    noaifridays.com18 days ago285 ptsView detailsJoin discussion

  2. OpenClaw 2.0, Accidentally

    openclaw.ai18 days ago148 ptsView detailsJoin discussion

  3. Roget's Thesaurus

    artflsrv04.uchicago.edu18 days ago49 ptsView detailsJoin discussion

  4. What We Tell AI

    whatwetellai.com18 days ago52 ptsView detailsJoin discussion

  5. Norway Shrugged (2024)

    paragraph.com18 days ago44 ptsView detailsJoin discussion

  6. Primary source

    A milestone in expanding access to AI

    ChatGPT Ads reaches $1 billion in annualized revenue run rate and expands globally, supporting broader access to AI through free and affordable options.

    openai.com18 days ago12 ptsView details

  7. Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token TTFT-First Benchmark

    Voice agents fail on latency long before they fail on intelligence. Time to first token is the metric most teams use to choose an inference API, and it is the right starting point and the wrong stopping point. This benchmark works through every layer of the voice stack — LLM, speech-to-text, text-to-speech, and speech…

    marktechpost.com18 days agoView details

  8. Google AI Introduces EnvHarness: A Programmable Layer That Turns Static Agent Environments Into Adaptive Training Worlds

    Google Cloud AI Research, with Washington University in St. Louis and UNC Chapel Hill, has released EnvHarness, an Apache-2.0 layer that turns a static agent benchmark into one that adapts to the policy training on it. It wraps a frozen environment through the standard reset()/step() interface, so tasks and human-buil…

    marktechpost.com18 days agoView details

  9. Anthropic Opens a Research Preview of the Model Hardware Standard (MHS): A Shared Specification for AI Agents to Safely Operate Physical Devices

    Anthropic has opened a research preview of the Model Hardware Standard (MHS), a shared driver specification that lets AI agents discover and safely operate physical devices. Instrument integration that normally takes weeks or months drops to hours: Carnegie Mellon went from raw equipment to a finished dose-response cu…

    marktechpost.com19 days agoView details

  10. Texas Governor Abbott blocks funding for more Flock cameras

    As backlash grows over Flock's AI surveillance cameras, Texas Governor Greg Abbott has frozen state spending on them. The move came just ahead of the publication of a Texas Tribune investigation that revealed the state spent over $30 million on Flock cameras. That money was primarily raised by tacking a $1 fee onto in…

    theverge.com18 days agoView details

  11. Why the Hottest New Wearables Want to Be Ignored

    Burnt out on wrist buzzes and notification overload? A new crop of minimalist wearables promises to collect your health data without demanding your attention.

    wired.com18 days agoView details

  12. A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls

    arXiv:2608.28040v1 Announce Type: new Abstract: Existing approaches to evasion detection in earnings calls focus on textual transcripts, treating evasion as a single-dimensional phenomenon. We argue that evasion in spoken communication is inherently multidimensional: beyond what executives say, how they say it carries…

    arxiv.org18 days agoView details

  13. SEPO: Evidence-Grounded Prompt Optimization via Structural Editing

    arXiv:2608.28067v1 Announce Type: new Abstract: Existing API-only prompt optimisers are often described as interpretable, but in practice, this usually means only post-hoc inspectability: each iteration still rewrites the prompt as one opaque string, leaving a trace of full-prompt diffs rather than localisable, machin…

    arxiv.org18 days agoView details

  14. SimpCue: Cue-Based Prompting for Multilingual Text Simplification

    arXiv:2608.28042v1 Announce Type: new Abstract: Text simplification aims to make complex texts easier to understand while preserving their original meaning. Recent large language models can perform simplification through prompting, but it remains unclear whether adding explicit linguistic information about sentence co…

    arxiv.org18 days agoView details

  15. Text Restoration of Ancient Documents with Language Models

    arXiv:2608.28170v1 Announce Type: new Abstract: Purpose - This study investigates the feasibility of restoring missing text caused by physical lacunae in damaged ancient manuscripts using language models. Methodology - The study proposes different scenarios to replicate real-world conditions. Language models of differ…

    arxiv.org18 days agoView details

  16. CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

    arXiv:2608.28405v1 Announce Type: new Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConvers…

    arxiv.org18 days agoView details

  17. Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

    arXiv:2608.28439v1 Announce Type: new Abstract: One model passed our fidelity check without ever opening the datasheet. We found it while qualifying models for an internal extraction service: a structured-output constraint had silently disabled tool use, and the model answered anyway, with fabricated source text. Only…

    arxiv.org18 days agoView details

  18. A Deep Learning-Based Stacking Ensemble Framework for Turbofan Engine Remaining Useful Life Prediction

    arXiv:2608.27940v1 Announce Type: new Abstract: This study proposes a two-level stacking ensemble framework for Remaining Useful Life (RUL) prediction of turbofan engines, evaluated on the NASA C-MAPSS benchmark using the FD001 and FD003 subsets. The framework integrates four heterogeneous deep learning base learners:…

    arxiv.org18 days agoView details

  19. Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense

    arXiv:2608.27945v1 Announce Type: new Abstract: Scaling laws are usually read as a capability story: lower language-modeling loss yields more useful models. We study a safety consequence of this mechanism in \emph{cross-session decomposition attacks}, where benign-looking subqueries are asked across independent intera…

    arxiv.org18 days agoView details

  20. Embedding Models for Stance-Aware Argument Retrieval

    arXiv:2608.28283v1 Announce Type: new Abstract: In computational argumentation, obtaining arguments that explicitly support or attack given claims is a critical precursor to downstream reasoning tasks. When these supporting and attacking arguments are to be retrieved using semantic search methods, they need to be asse…

    arxiv.org18 days agoView details

  21. A Probabilistic Interpretation of KV Cache Eviction

    arXiv:2608.28293v1 Announce Type: new Abstract: The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most rely on creative heuristics for selectin…

    arxiv.org18 days agoView details

  22. ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

    arXiv:2608.28476v1 Announce Type: new Abstract: Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing working context. Recent proacti…

    arxiv.org18 days agoView details

  23. Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls

    arXiv:2608.27727v1 Announce Type: new Abstract: A model's behavior on a task is jointly determined by the input it receives and the prior it brings in, i.e. the distribution over stimuli it implicitly expects. Interpretability research has traditionally studied models by holding inputs fixed and examining model respon…

    arxiv.org18 days agoView details

  24. BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla

    arXiv:2608.28329v1 Announce Type: new Abstract: Medical question answering (QA) systems have become crucial tools for providing reliable health information. But they remain very unexplored for low-resource languages like Bangla due to limited datasets and systems tailored to these languages. To address this, we introd…

    arxiv.org18 days agoView details

  25. Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models

    arXiv:2608.27768v1 Announce Type: new Abstract: A language model with access to tools can commit to a final claim unsupported by the evidence it has seen, even when a single available tool call would resolve the uncertainty and its instructions explicitly forbid assumptions and guesses. We separate this failure into t…

    arxiv.org18 days agoView details

  26. Credo: Reusable Declarative Primitives for Agentic Workflows

    arXiv:2608.27790v1 Announce Type: new Abstract: An LLM application depends on both a model and a harness: the program that determines what each call sees, how many calls to make, and which answers to trust. Coding agents can now discover strong harnesses by searching over candidate programs, but the resulting artifact…

    arxiv.org18 days agoView details

  27. SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing

    arXiv:2608.27963v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong reasoning capabilities, yet long-chain reasoning becomes inefficient once the intermediate answer stabilizes across reasoning steps: additional reasoning yields little marginal benefit while incurring substantial inference cos…

    arxiv.org18 days agoView details

  28. PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems

    arXiv:2608.28378v1 Announce Type: new Abstract: Large language models are increasingly used as agentic workflow executors, yet existing training data and benchmarks largely assume informationally complete, single-turn queries. Our analysis of 16K real-world sessions shows that 75.9% of interactions are multi-turn, rev…

    arxiv.org18 days agoView details

  29. Evidential-Based Higher-Order Set Argumentation Framework

    arXiv:2608.27824v1 Announce Type: new Abstract: Evidential argumentation extends Dung's abstract argumentation by requiring arguments and interactions to be backed by chains of evidence rooted in prima-facie elements. However, existing formalisms lack a unified treatment of evidential support, higher-order relations (…

    arxiv.org18 days agoView details

  30. Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

    arXiv:2608.28151v1 Announce Type: new Abstract: A byte-level BPE tokenizer is an ordered list of merge rules, so applying only a prefix yields a vocabulary whose token identifiers are the first rows of the full vocabulary. This prefix nesting allows one language model to operate at several vocabulary sizes, use a cont…

    arxiv.org18 days agoView details

  31. LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails

    arXiv:2608.27580v2 Announce Type: new Abstract: Safety guardrails serve as the last line of defense against harmful inputs and outputs of large language models (LLMs), yet they are trained and evaluated almost exclusively on short text. We present LongGuard, a framework that evaluates, mechanistically analyzes, and mi…

    arxiv.org18 days agoView details

  32. When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems

    arXiv:2608.27984v1 Announce Type: new Abstract: Multi-Agent Systems (MAS) have recently moved from static workflows toward dynamically generated collaboration topologies. However, existing topology generation methods rely primarily on the parametric knowledge of large language models, with external search or retrieval…

    arxiv.org18 days agoView details

  33. Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Experiences (CUREs)

    arXiv:2608.27638v1 Announce Type: new Abstract: Course-based undergraduate research experiences (CUREs) broaden access to authentic scientific inquiry through responsive instructor support as research problems become increasingly complex. Generative artificial intelligence (GenAI) may extend this support by providing…

    arxiv.org18 days agoView details

  34. GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies

    arXiv:2608.27992v1 Announce Type: new Abstract: Generative-agent systems are easier to start than to inspect. A run can contain many agents, locations, messages, commands, and model calls, yet the operator often gets either a finished replay or raw logs. That makes it hard to ask why an agent moved, test a small inter…

    arxiv.org18 days agoView details

  35. Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real Data

    arXiv:2608.27996v1 Announce Type: new Abstract: Digital twins (DTs) and learned world models are increasingly used to generate synthetic data that augment the scarce real datasets available for training artificial intelligence (AI) models in engineering systems. Owing to the inevitable simulation-to-reality (sim-to-re…

    arxiv.org18 days agoView details

  36. SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models

    arXiv:2608.27857v1 Announce Type: new Abstract: Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models from scratch remains difficult. A practical alternative is ANN-to-SNN migration through kn…

    arxiv.org18 days agoView details

  37. Stranger, Fan, or Peer? A Systematic Study on the Role of Interlocutor in Persona-Based Dialogue Generation

    arXiv:2608.28467v1 Announce Type: new Abstract: Persona-based dialogue systems are usually conditioned on speaker biography, but dialogues involve at least two participants, and who has access to whose biography can vary across training, inference, and evaluation. Prior work often neglected these aspects, obscuring me…

    arxiv.org18 days agoView details

  38. The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs

    arXiv:2608.27953v1 Announce Type: new Abstract: Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks largely target bounded settings with fixed variables or single gold outcomes, overlooking open-d…

    arxiv.org18 days agoView details

  39. From Uncertainty to Clinical Risk: Severity-Aware Conformal Planning for Interactive Medical Diagnosis

    arXiv:2608.27847v1 Announce Type: new Abstract: Interactive medical diagnosis dynamically acquires patient information through multiple rounds of questioning, supporting accurate, efficient, and safe clinical decisions under incomplete evidence. Existing methods commonly guide information acquisition with predictive u…

    arxiv.org18 days agoView details

  40. CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action

    arXiv:2608.27797v1 Announce Type: new Abstract: Natural-language tasking of embodied agents is rarely just goal specification: users also impose constraints that must persist while the world changes. Code-generating LLM agents can produce plausible behaviors for such instructions, but their free-form programs provide…

    arxiv.org18 days agoView details

  41. H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference

    arXiv:2608.28113v1 Announce Type: new Abstract: The NVIDIA Blackwell architecture, with native support for the ultra-fine-grained NVFP4 format, opens new opportunities for accelerating large language model (LLM) inference. NVFP4's micro-block design, such as a group size of 16, offers strong representational flexibili…

    arxiv.org18 days agoView details

  42. Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning

    arXiv:2608.28065v1 Announce Type: new Abstract: Complete your ad view and grab a 5-cent bonus! In incentivized advertising, a platform promises users a bonus before observing downstream ad revenue, encouraging them to click and complete ads. It must balance the incentive promised in advance against the revenue realize…

    arxiv.org18 days agoView details

  43. Rubric-to-Code Credit Assignment for Reinforcement Learning

    arXiv:2608.27906v2 Announce Type: new Abstract: Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation, application quality depends on multiple user-facing functional requirements, each often ti…

    arxiv.org18 days agoView details

  44. openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents

    arXiv:2608.27969v1 Announce Type: new Abstract: Long-horizon coding agents operate over evolving repository states while increasingly relying on heterogeneous capabilities, delegated agents, and multi-agent coordination. These trends pose two complementary challenges for the agent harness. First, developers need to co…

    arxiv.org18 days agoView details

  45. CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

    arXiv:2608.28053v1 Announce Type: new Abstract: Chinese neologisms exploit diverse and unique linguistic mechanisms, such as phonetic substitution (e.g., 886 for ``bye-bye'') and visual character decomposition that are rare in other languages. We introduce CNeo-Bench, a benchmark of 4,759 such neologisms with referenc…

    arxiv.org18 days agoView details

  46. From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning

    arXiv:2608.27919v1 Announce Type: new Abstract: Financial question answering (QA) has emerged as a key benchmark for evaluating the performance of Large Language Models (LLMs) on domain-specific tasks involving complex data formats such as tables, charts, and rich textual narratives. While recent advancements have ena…

    arxiv.org18 days agoView details

  47. HyQuant: Hybrid-Precision Quantization for LLM Attention

    arXiv:2608.27875v1 Announce Type: new Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods…

    arxiv.org18 days agoView details

  48. WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

    arXiv:2608.28062v2 Announce Type: new Abstract: Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned images from subsequent context, reducing v…

    arxiv.org18 days agoView details

  49. PACE: Publisher-Adaptive Content Extraction via Agentic Automation

    arXiv:2608.27466v1 Announce Type: new Abstract: Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability. General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layout…

    arxiv.org18 days agoView details

  50. AI Alignment through a Game-theoretic Lens: A Survey

    arXiv:2608.27910v1 Announce Type: new Abstract: As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability…

    arxiv.org18 days agoView details

  51. Agents for Everyone: A Workshop Framework for Building Agentic AI Capabilities in a Distributed Curation Community

    arXiv:2608.27675v1 Announce Type: new Abstract: Agentic AI has the potential to accelerate curation of biological databases and knowledge bases. However, uptake has been hindered by a number of challenges and obstacles, including access to agents and appropriate training. Here we describe how we have attempted to addr…

    arxiv.org18 days agoView details

  52. Synthetic Linguistic Agency: How an Embodied Mortal Agent Learns Linguistic Affordances through Consequential Social Experience

    arXiv:2608.27843v1 Announce Type: new Abstract: Contemporary language models can converse fluently and influence human decisions, yet their exchanges do not enter a continuing, vulnerable life of their own. Linguistic-agency theory identifies this missing connection as linguistic agency and characterizes it through em…

    arxiv.org18 days agoView details

  53. AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics

    arXiv:2608.27818v1 Announce Type: new Abstract: User preferences in user-agent collaboration are rarely static and fully-specified upfront: preferences are formed, revealed, adjusted, and relaxed during interaction. Existing benchmarks for evaluating user-agent collaboration focus almost exclusively on resolving under…

    arxiv.org18 days agoView details

  54. Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model

    arXiv:2608.27998v1 Announce Type: new Abstract: The rapid proliferation of interdisciplinary and multilingual scientific literature has left traditional manual analysis and single-algorithm methods plagued by low efficiency, poor scalability, and insufficient domain adaptability. Targeting the literature analysis need…

    arxiv.org18 days agoView details

  55. CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning

    arXiv:2608.27867v1 Announce Type: new Abstract: Continual multimodal instruction tuning requires multimodal large language models to acquire new task abilities sequentially while preserving previously learned knowledge. LoRA-MoE provides a promising solution by introducing expert-based capacity, but repeatedly learnin…

    arxiv.org18 days agoView details

  56. AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning

    arXiv:2608.27964v1 Announce Type: new Abstract: Test-time scaling improves language-model reasoning by generating additional candidate solutions, but allocating the same inference budget to every problem is computationally wasteful. Existing adaptive stopping methods commonly rely on confidence, agreement, or answer s…

    arxiv.org18 days agoView details

  57. LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation

    arXiv:2608.27472v1 Announce Type: new Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge. We propose combining these complementary sources through a novel represent…

    arxiv.org18 days agoView details

  58. Rating the Raters: Rasch Measurement Theory for LLM Evaluation

    arXiv:2608.27463v1 Announce Type: new Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an in…

    arxiv.org18 days agoView details

  59. Accelerating LLM Inference via Vector Index Based Output Embeddings

    arXiv:2608.27460v1 Announce Type: new Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies. We reformulate the output projection followed by top-k token selection as a maximum inner pr…

    arxiv.org18 days agoView details

  60. QUORUM: QUality-Optimized Routing Using Multiple annotators

    arXiv:2608.27974v1 Announce Type: new Abstract: Data annotation remains a central bottleneck in natural language processing, requiring human effort to obtain high-quality labels at scale. While Large Language Models (LLMs) offer a fast and cost-effective alternative, their reliability is highly instance-dependent: the…

    arxiv.org18 days agoView details

  61. Sliding-window beats linear attention

    arXiv:2608.28444v1 Announce Type: new Abstract: Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previous one. For each additional token, the keys and values must be stored in memory indefinitely, which is unsustainable. Seve…

    arxiv.org18 days agoView details

  62. Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning

    arXiv:2608.27756v1 Announce Type: new Abstract: Determining whether large language models derive answers from context or prior knowledge remains a fundamental challenge. Self-contained linguistic olympiad puzzles provide a controlled setting where all answers derive solely from expert-designed context examples without…

    arxiv.org18 days agoView details

  63. Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems

    arXiv:2608.27482v1 Announce Type: new Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems. The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through conditional aggregation…

    arxiv.org18 days agoView details

  64. OpenStamp: A Watermark for Open-Source Language Models

    arXiv:2608.27899v1 Announce Type: new Abstract: With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to LLMs and distinguishing it from human-written content. A prominent class of techniques embeds subtle but detectable signal…

    arxiv.org18 days agoView details

  65. NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry

    arXiv:2608.28481v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have demonstrated strong capabilities in natural language understanding and mathematical reasoning. However, their ability to translate informal mathematical problems into formal representations remains underexplored. This…

    arxiv.org18 days agoView details

  66. Thinking Costs Tokens: When More Structure is Worth the Price

    arXiv:2608.27506v1 Announce Type: new Abstract: Adding inference structure to a language model lets it search, verify, and revise, but these actions consume the very budget they are supposed to use well. In this paper, we investigate whether there exists a token-budget threshold, below which the overhead of planning a…

    arxiv.org18 days agoView details

  67. When Linguistic and Internal Confidence Diverge in Large Language Models

    arXiv:2608.28382v1 Announce Type: new Abstract: Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three f…

    arxiv.org18 days agoView details

  68. Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation

    arXiv:2608.27729v1 Announce Type: new Abstract: Function routing -- selecting the correct API call from a fixed catalog given a natural-language request -- is a deployment problem where small students are attractive but knowledge distillation gains are typically reported single-seed, at scales where seed variance is u…

    arxiv.org18 days agoView details

  69. Informational Antilocality and the Locality Bias in LLMs

    arXiv:2608.27760v1 Announce Type: new Abstract: We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs…

    arxiv.org18 days agoView details

  70. A Framework for Object-Centric Predictive Monitoring of Collaborative Processes

    arXiv:2608.27671v1 Announce Type: new Abstract: Predictive Process Monitoring (PPM) of collaborative, inter-organizational processes requires reasoning over multiple interdependent entities, including participants, messages, local executions, and the global collaboration case. Existing approaches extend traditional ev…

    arxiv.org18 days agoView details

  71. Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields

    arXiv:2608.27475v1 Announce Type: new Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it. These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can concea…

    arxiv.org18 days agoView details

  72. Knowing Before Answering: Decoding Language Models for Reliable RAG

    arXiv:2608.27661v1 Announce Type: new Abstract: In Retrieval-Augmented Generation (RAG), retrieval may provide insufficient or conflicting information needed to answer a question. The system should not only know when to answer but also be able to identify cases in which the documents provided in RAG are insufficient o…

    arxiv.org18 days agoView details

  73. Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI

    arXiv:2608.27464v1 Announce Type: new Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking. Drawing on Sharot and Sunstein's framework of information-seeking motives, we propose that people evaluate whether to engage with…

    arxiv.org18 days agoView details

  74. Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection

    arXiv:2608.27470v1 Announce Type: new Abstract: Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs. State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selecting the correct one give…

    arxiv.org18 days agoView details

  75. Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation

    arXiv:2608.28508v1 Announce Type: new Abstract: Forced alignment evaluation typically requires manually annotated timestamps, limiting large-scale and multilingual analysis. We introduce two corpus-level metrics based on self-supervised (SSL) speech representations for reference-free forced alignment evaluation: Phone…

    arxiv.org18 days agoView details

  76. PhenoIntel: A Lifecycle-Aligned Multi-Agent Web Application for Verified, Accessible Plant Phenotype Analysis

    arXiv:2608.27999v1 Announce Type: new Abstract: Existing conversational plant-phenotyping platforms are difficult for plant scientists to use and lack the reliability scientific research demands: failed analyses are reported as valid measurements rather than flagged as missing, statistical tests run without checking a…

    arxiv.org18 days agoView details

  77. Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents

    arXiv:2608.28011v1 Announce Type: new Abstract: Trajectory-level credit assignment can localize which module of a tool-using LLM agent causes failures using only verifiable signals. We ask whether such failure credit should route a fixed zeroth-order/evolution-strategies (ZO/ES) perturbation budget. Across a synthetic…

    arxiv.org18 days agoView details

  78. PersonaEdit: Representative Sample Selection for Personalized Model Editing

    arXiv:2608.27816v1 Announce Type: new Abstract: Personalization has attracted growing interest in LLM applications, yet existing retrieval-based approaches depend heavily on retrieval quality and degrade in long-term interactions. Model editing, which directly modifies internal model parameters to incorporate new know…

    arxiv.org18 days agoView details

  79. WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning

    arXiv:2608.27508v1 Announce Type: new Abstract: GUI agents trained with reinforcement learning (RL) have showcased strong environment learning capabilities on mobile platforms. However, RL typically demands extensive real-environment interactions, leading to high resource costs and instability, especially in GUI scena…

    arxiv.org18 days agoView details

  80. Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator

    arXiv:2608.27548v1 Announce Type: new Abstract: Safety moderation for deployed AI applications is moving beyond text-only prompts: systems increasingly need to judge images, documents, screenshots, and generated responses under policies that vary across domains. Existing guardrails usually cover only part of this sett…

    arxiv.org18 days agoView details

  81. Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

    arXiv:2608.27459v1 Announce Type: new Abstract: In 2011, IBM's Watson was something like a sealed capsule of its era's queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 s…

    arxiv.org18 days agoView details

  82. A Survey on Rubric-Guided Reinforcement Learning for Language Models

    arXiv:2608.27505v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences. However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted…

    arxiv.org18 days agoView details

  83. XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

    arXiv:2608.27481v1 Announce Type: new Abstract: Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries inside the reasoning chain. We…

    arxiv.org18 days agoView details

  84. KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation

    arXiv:2608.27839v1 Announce Type: new Abstract: Fine-tuning-based knowledge editing is simple and architecture-agnostic, but standard cross-entropy increases the edited target probability without explicitly constraining changes in the non-target output distribution. In sequential editing, such unconstrained redistribu…

    arxiv.org18 days agoView details

  85. Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL

    arXiv:2608.28432v1 Announce Type: new Abstract: Recent advances in in-context learning (ICL) text-to-SQL have substantially improved execution accuracy on public benchmarks by assembling increasingly elaborate pipelines around the base generator, yet existing studies typically report aggregate end-to-end accuracy, wit…

    arxiv.org18 days agoView details

  86. Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation

    arXiv:2608.28496v1 Announce Type: new Abstract: Two forms of test-time scaling for Large Language Models (LLMs) have emerged as effective and widely adopted paradigms: sequential, in which later answer attempts depend on earlier ones, and parallel, such as i.i.d. sampling with reranking. In this study, we investigate…

    arxiv.org18 days agoView details

  87. Lexically conditioned realization ambiguity in Korean predicate morphology

    arXiv:2608.27966v1 Announce Type: new Abstract: This paper examines Korean surface realization as distinct from morphological analysis. It asks whether a sequence of canonical morphemes and grammatical category labels uniquely determines the corresponding surface form. The answer is negative for a restricted but theor…

    arxiv.org18 days agoView details

  88. A Unified Framework to Elicit Structured Feedback for Interpretable Multi-Trait Essay Scoring

    arXiv:2608.28407v1 Announce Type: new Abstract: Multi-trait Automated Essay Scoring (AES) requires rubric-grounded reasoning across interdependent traits, rather than isolated score prediction. Existing feedback-enhanced methods often decouple feedback from scoring or assess traits independently, weakening score--feed…

    arxiv.org18 days agoView details

  89. CASTANET: Causality-Aware Spatio-Temporal Adversarial Network Using Traffic Incident Effects

    arXiv:2608.27942v1 Announce Type: new Abstract: Predicting non-periodic traffic congestion caused by sudden incidents (e.g., accidents and road damage) is crucial for advanced intelligent transportation systems. However, incident-driven congestion is difficult to forecast because incidents are extremely sparse, occur…

    arxiv.org18 days agoView details

  90. An Empirical Evaluation of Cross-City POI Recommendation on a Large-Scale Benchmark

    arXiv:2608.27840v1 Announce Type: new Abstract: Cross-city point-of-interest (POI) recommendation is crucial for navigating unfamiliar urban environments, yet its progress has historically been constrained by data limitations. Using the recently proposed large-scale benchmark Trip World, we empirically re-examine whet…

    arxiv.org18 days agoView details

  91. When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation

    arXiv:2608.27960v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and capabilities of teacher models into student models. However, teacher guidance on student-gener…

    arxiv.org18 days agoView details

  92. CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence

    arXiv:2608.27484v1 Announce Type: new Abstract: Artificial intelligence is transforming personalized healthcare, yet fragmented clinical, self reported, and wearable evidence remains difficult to interpret and trace. We present CareGraph, an auditable hybrid AI framework that converts heterogeneous records into priori…

    arxiv.org18 days agoView details

  93. If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary

    arXiv:2608.27646v1 Announce Type: new Abstract: Give an agent a human's credential and it inherits the person's reach without the judgment that limits its use. It can sweep every reachable record into model context, where hidden instructions steer its next call, and every request stays credential-valid while the agent…

    arxiv.org18 days agoView details

  94. Entity-Memory Graph Retrieval Improves Evidence Coverage in Long-Conversation Question Answering

    arXiv:2608.27925v1 Announce Type: new Abstract: Entity-Memory graph retrieval keeps dialogue turns as verbatim Memory nodes, links repeated mentions through shared Entities, and connects adjacent Memories with directed chronological edges. At query time the retriever moves from Entity gating through semantic fusion an…

    arxiv.org18 days agoView details

  95. The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

    arXiv:2608.27465v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user's emotional state has become an important safety problem. We test whether emotional expression increases a…

    arxiv.org18 days agoView details

  96. What Makes Agent Memory Useful for Reliable Unanswerable Question Handling?

    arXiv:2608.27924v1 Announce Type: new Abstract: Reliable handling of unanswerable questions (UAQs) is critical for trustworthy LLM-based agents. Although memory is widely used in agent systems, its role in reliable UAQ handling remains unclear. We present a systematic study of agent memory for UAQ handling under a uni…

    arxiv.org18 days agoView details

  97. How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

    arXiv:2608.27510v1 Announce Type: new Abstract: Transcoder attribution graphs are usually trained to explain why a model assigns high probability to a particular next token. We introduce Concept-Targeted Attribution (CTA), which instead trains attribution graphs with respect to a linear probe direction. CTA therefore…

    arxiv.org18 days agoView details

  98. PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation

    arXiv:2608.27716v1 Announce Type: new Abstract: AI systems are being deployed on high-stakes, domain-specific workflows that demand correctness not just in the final output, but at every intermediate step. One such workflow is estimating a product carbon footprint (PCF), the greenhouse-gas emissions attributable to a…

    arxiv.org18 days agoView details

  99. Resource Constraints and Performance in Agentic AI Systems

    arXiv:2608.27886v1 Announce Type: new Abstract: Progress toward more autonomous AI increasingly depends on agentic systems that combine a language model with tools, memory, state management, and multi-step execution. These mechanisms shape both task capability and operational burden. We compare OpenClaw and NanoBot as…

    arxiv.org18 days agoView details

  100. SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction

    arXiv:2608.27461v1 Announce Type: new Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts. This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different asp…

    arxiv.org18 days agoView details

  101. INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

    arXiv:2608.27501v1 Announce Type: new Abstract: Mathematical reasoning has seen rapid progress in large language models (LLMs), yet existing methods optimize predominantly for final-answer correctness, raising the question whether models truly internalize mathematical concepts or merely memorize solution patterns. In…

    arxiv.org18 days agoView details

  102. When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages

    arXiv:2608.27658v1 Announce Type: new Abstract: Subword tokenization hinders low-resource language processing by imposing frequency patterns from dominant languages onto script-sharing variants. Byte-level models bypass this issue by processing raw UTF-8 characters, yet they create a granularity mismatch for word-leve…

    arxiv.org18 days agoView details

  103. FinExam-10K: When Retrieval Helps Financial Reasoning?

    arXiv:2608.28155v1 Announce Type: new Abstract: Professional financial examinations require models to combine domain knowledge, calculation, and judgment, yet no benchmark covers the full CFA and FRM structure under one protocol. We introduce FinExam-10K, to our knowledge the largest reported English benchmark for thi…

    arxiv.org18 days agoView details

  104. First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents

    arXiv:2608.27672v1 Announce Type: new Abstract: We present Qwen-GuidePlay-2B, a 2B-parameter language model for dialogue-game interaction. We fine-tune Qwen3.5-2B using three steps: a) SFT on only successful game trajectories from Playpen, b) weighted turn-level SFT, and c) teacher-guided SFT. The teacher model (which…

    arxiv.org18 days agoView details

  105. String: An Agentic OS Where Every App Is a Markdown File

    arXiv:2608.28027v1 Announce Type: new Abstract: LLM agents have become a new class of software user, but every surface they work through was designed for someone else. Pages are built for human eyes, which can skim and ignore; tool schemas for programs, which pay nothing to carry definitions they never call. An agent…

    arxiv.org18 days agoView details

  106. Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

    arXiv:2608.27785v1 Announce Type: new Abstract: We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain…

    arxiv.org18 days agoView details

  107. Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

    arXiv:2608.28478v1 Announce Type: new Abstract: Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long-tail facts. To address this gap, we introduce ElephantBench, a closed-book knowledge probe comprising 1,094 quest…

    arxiv.org18 days agoView details

  108. RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

    arXiv:2608.27831v2 Announce Type: new Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characterize this gap,…

    arxiv.org18 days agoView details

  109. A Formal Limitation on Learning Human Language From Textual Corpora

    arXiv:2608.28560v1 Announce Type: new Abstract: Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language us…

    arxiv.org18 days agoView details

  110. Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap

    arXiv:2608.27512v1 Announce Type: cross Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models. When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluation, this workflow…

    arxiv.org18 days agoView details

  111. Class-Based Heuristic Selection for Solving the Flying Block Puzzle

    arXiv:2608.27476v1 Announce Type: new Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to exploit the structural constraints that govern constrained spatial domains, causing search performance to degrade catastrophic…

    arxiv.org18 days agoView details

  112. LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages

    arXiv:2608.27902v1 Announce Type: new Abstract: Landing pages are goal-oriented web interfaces that must communicate a target-specific value proposition while organizing information flow, visual hierarchy, and calls to action (CTA). Although large language models can generate plausible webpage code from natural-langua…

    arxiv.org18 days agoView details

  113. Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

    arXiv:2608.27477v1 Announce Type: new Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks. Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application…

    arxiv.org18 days agoView details

  114. Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations

    arXiv:2608.27480v1 Announce Type: new Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected. This study proposes an IoT-enabled acoustic monitoring framework integrated wi…

    arxiv.org18 days agoView details

  115. See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs

    arXiv:2608.27869v1 Announce Type: new Abstract: Discovering governing partial differential equations (PDEs) from observational data remains a core challenge across the sciences. Existing sparse-regression, symbolic-regression, and LLM-based approaches can be constrained by predefined libraries, noise sensitivity, hall…

    arxiv.org18 days agoView details

  116. Representation of syntax in LLMs through the lens of linear distance and similarity-aware entropy

    arXiv:2608.27813v1 Announce Type: new Abstract: Structural probes were introduced by Hewitt and Manning to reconstruct syntactic trees from a neural language model's latent representations. They are evaluated by calculating the proportion of syntactic tree edges correctly reconstructed over an annotated corpus (as mea…

    arxiv.org18 days agoView details

  117. UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering

    arXiv:2608.27467v1 Announce Type: new Abstract: We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and…

    arxiv.org18 days agoView details

  118. SETU: An Agentic Ecosystem for Multilingual, Persona-Aware Communication Coaching

    arXiv:2608.27524v1 Announce Type: new Abstract: Corporate training teams need scalable and explainable tools to improve workforce communication in multilingual settings. Existing systems often score text, audio, or video in isolation, or produce black-box outputs that are difficult to audit for coaching use. This pape…

    arxiv.org18 days agoView details

  119. Trajectory-Level Speculative Decoding for Diffusion Language Models

    arXiv:2608.27514v1 Announce Type: new Abstract: Diffusion-based language models (dLLMs) enable parallel token generation through iterative denoising, but existing decoding strategies collapse to single-token generation under low confidence, severely limiting throughput. Unlike autoregressive models where speculative d…

    arxiv.org18 days agoView details

  120. Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling

    arXiv:2608.27982v1 Announce Type: new Abstract: Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO. Among these, Dynamic Sampling contributes the most to DAPO's accuracy gains relative to GRP…

    arxiv.org18 days agoView details

  121. Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis

    arXiv:2608.27471v1 Announce Type: new Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped. Spotting a fallacious argument requires contextual knowledge beyond its pure surf…

    arxiv.org18 days agoView details

  122. Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech

    arXiv:2608.27462v1 Announce Type: new Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging. While existing PLM- or LLM-based methods perform we…

    arxiv.org18 days agoView details

  123. Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents

    arXiv:2608.28458v1 Announce Type: new Abstract: Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Playschool Challenge with a 2B open…

    arxiv.org18 days agoView details

  124. EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion

    arXiv:2608.27844v1 Announce Type: new Abstract: Existing evaluations of harmful content detection rely predominantly on static benchmarks, which struggle to reflect the interactive adversarial ecosystem of real-world content platforms where users continuously revise their expressions in response to moderation feedback…

    arxiv.org18 days agoView details

  125. AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not

    arXiv:2608.27855v1 Announce Type: new Abstract: Text generated by large language models (LLMs) has been shown to be stylometrically distinct from human-written text \citep{andreDetectingAIAuthorship2023, shahDetectingUnmaskingAIGenerated2023, oparaStyloAIDistinguishingAIGenerated2024, soto2024fewshot, liLinguisticDiff…

    arxiv.org18 days agoView details

  126. Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness

    arXiv:2608.27988v1 Announce Type: new Abstract: Smooth speaker transitions are fundamental to effective conversation and rely on an interlocutor's ability to predict when to enter the conversation. This ability depends on accurately interpreting and expressing the verbal and non-verbal cues that signal when a speaker…

    arxiv.org18 days agoView details

  127. Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection

    arXiv:2608.28009v1 Announce Type: new Abstract: The rapid evolution of large language models necessitates robust machine-generated text detection. Existing paradigms typically follow two isolated tracks. Training-free methods rely on global statistical scalars such as perplexity, while training-based methods utilize s…

    arxiv.org18 days agoView details

  128. Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning

    arXiv:2608.28018v1 Announce Type: new Abstract: Knowledge-intensive reasoning requires Large Language Models (LLMs) to ground answers in provided evidence. When evidence is insufficient, it is desirable that models abstain rather than confidently generating unsupported answers. Existing abstention methods rely on unce…

    arxiv.org18 days agoView details

  129. ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL

    arXiv:2608.27796v1 Announce Type: new Abstract: Recent work has shown that reinforcement learning from execution feedback can substantially improve text-to-SQL performance, often enabling smaller models to match or exceed much larger systems. However, most existing approaches treat SQL generation as a single-turn task…

    arxiv.org18 days agoView details

  130. CURA: Certified Runtime Alarms for Computer-Use Agents

    arXiv:2608.27808v1 Announce Type: new Abstract: Self-report is the cheapest oversight channel a deployer has, and on capable computer-use agents (CUAs) it fails precisely where oversight matters. On 361 OSWorld tasks our pipeline, a read-only feasibility gate, a planner, and a GUI executor, reaches a mean task score o…

    arxiv.org18 days agoView details

  1. ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

    image-text-to-text · gguf · gsq · rco

    huggingface.co19 days ago1220 ptsView details

  2. Smite79/MiniMax-H3-Longvideos

    text-to-video · minimax-h3 · comfyui · comfyui-custom-node

    huggingface.co19 days ago82 ptsView details

  3. leejet/MiniMax-H3-GGUF

    image-to-video · gguf · text-to-video · image-to-video

    huggingface.co19 days ago17 ptsView details

  4. Baekpica/Motif-3-Mixed-Quant-GGUF

    text-generation · gguf · motif · motif-3

    huggingface.co19 days ago3 ptsView details

  5. inferencerlabs/Qwen3.8-Flash-Next-MLX-Q4

    image-text-to-text · mlx · qwen4_exp · quantized

    huggingface.co19 days ago1 ptsView details

  6. APRJdevelopment/fastpunch-ondevice-1b-merged

    text-generation · transformers · safetensors · llama

    huggingface.co19 days agoView details

  1. ggml-org/llama.cpp b10711

    <details open> hexagon: fix CPY fence bug (#28033) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/44046740> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/downlo…

    github.com18 days agoView details

  2. ggml-org/llama.cpp b10710

    <details open> metal : add remaining Q4_1/Q5_0/Q5_1 fa-vec tunings for M2 (#28017) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/44044108> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/gg…

    github.com18 days agoView details

  3. ggml-org/llama.cpp b10690

    <details open> memory : copy Hadamard matrix to k_rot tensor only if it has buffer assigned to prevent crashes during context shift of unquantized K cache (#27967) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.gi…

    github.com19 days agoView details