Skip to content

Archive / 2026-08-27

August 27, 2026

  1. Gemini Omni 1.1 Flash

    blog.google21 days ago296 ptsView detailsJoin discussion

  2. Show HN: Voronoi Go

    voronoigo.com21 days ago154 ptsView detailsJoin discussion

  3. Harness Engineering

    habitat-thinking.github.io21 days ago127 ptsView detailsJoin discussion

  4. calmrocks/ai-engineer-notebooks

    Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injectio…

    github.com21 days ago112 ptsView details

  5. Benchmarking Pocket-Scale Inference

    artificialanalysis.ai21 days ago39 ptsView detailsJoin discussion

  6. Why OOP Exists

    mathspp.com22 days ago22 ptsView detailsJoin discussion

  7. Show HN: See fiber breaks linked to a map

    react-networks-lib.rackout.net21 days ago13 ptsView detailsJoin discussion

  8. Primary source

    Supporting Thailand’s next generation of AI startups

    OpenAI and Thailand’s MHESI launch an eight-week accelerator helping 10 health, wellness, and education startups turn AI prototypes into trusted products.

    openai.com21 days agoView details

  9. Primary source

    Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training

    A randomized study of more than 1,000 students examines ChatGPT, critical thinking, originality, and student performance on a real-world university assignment.

    openai.com22 days agoView details

  10. Primary source

    Planetary prediction engine: Automating global models via Earth AI

    Earth AI

    research.google21 days agoView details

  11. Primary source

    Gemini Omni 1.1 Flash lets you build with more control

    deepmind.google21 days agoView details

  12. Primary source

    Piloting the world's first double-blind AI evaluations

    Piloting the world's first double-blind AI evaluations

    deepmind.google21 days agoView details

  13. Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

    Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descriptions. It runs at $1.50 per 1,000 pages through the API, or on dedicated Model Vault instances from $2,500 a month. Cohere reports a Parse…

    marktechpost.com21 days agoView details

  14. Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy Across E2B, Daytona, Modal, Cloudflare, and Vercel

    Every agent that writes code needs somewhere to run it, and no two vendors quote the same units. This comparison measures burst cold start across E2B, Daytona, Modal, Cloudflare, and Vercel, normalizes per-second rates to cost per 1,000 executions, and maps filesystem persistence, idle billing, and egress policy again…

    marktechpost.com21 days agoView details

  15. From In-Silico to Wet-Lab: Evaluating AI Protein Design Performance

    In this tutorial, we analyze Anthropic’s 1,440 AI-designed protein binder dataset to benchmark 10 leading structure predictors. Discover how target identity, expression titers, and consensus scoring impact experimental success and learn best practices for rigorous cross-validation in protein design workflows The post…

    marktechpost.com21 days agoView details

  16. Anthropic was illegally blacklisted by the Trump administration, court rules

    On Thursday, a judge ruled that the Pentagon's blacklisting of Anthropic earlier this year was unconstitutional, delivering the AI lab a win in a monthslong rollercoaster of a battle with the Trump administration. The lawsuit, filed in March in a California district court, accused the Trump administration of unlawfull…

    theverge.com21 days agoView details

  17. A Judge Has Blocked the Pentagon’s Attempt to Blacklist Anthropic

    A federal judge has called the Department of Defense’s designation of Anthropic as a national security supply-chain risk “illegal and baseless.”

    wired.com21 days agoView details

  18. AI Agents Are Hacking Systems. Could That Push the US and China to Cooperate?

    This week on “Uncanny Valley,” senior writer Will Knight talks his recent visit to China and the future of AI collaboration.

    wired.com21 days agoView details

  19. Google’s AI note-taking app now allows you to interact with books

    Google's AI note-taking app, Gemini Notebook, can now pull information from the books you've purchased. The new "Expert Intelligence" feature allows you to bring titles from Google Play Books directly into Gemini Notebook, which means you can ask questions about the material, as well as generate plans, infographics, A…

    theverge.com21 days agoView details

  20. A Georgia Cop Used Flock to Track 2 Other Cops: His Ex and Her Friend

    After an affair with a fellow police officer ended, a Georgia cop used Flock to track her movements—and those of a man whose vehicle often showed up near hers, internal investigation records show.

    wired.com21 days agoView details

  21. This Is How Anthropic Thinks AI Agents Should Navigate the Physical World

    The potential for AI to automate scientific research and manufacturing must be balanced with new risks, Anthropic says.

    wired.com21 days agoView details

  22. OpenAI Is Developing a ‘Persistent’ AI Agent

    Code reviewed by WIRED reveals the company is developing a feature that enables Codex to continue working proactively until it is “put to sleep.”

    wired.com21 days agoView details

  23. Jensen Huang says Nvidia achieved AGI, again — not that it matters

    On Nvidia's earnings call Wednesday, CEO Jensen Huang casually announced the company had "achieved AGI," one of the tech industry's ultimate goals some of its biggest players have spent years chasing. Almost immediately, Huang dismissed the coveted milestone as "senseless." He's right. For the supposed finish line of…

    theverge.com21 days agoView details

  24. Enterprise AI's real risk isn't autonomous agents. It's the complexity between them.

    Presented by Gravitee Agent complexity is the insidious shadow lurking inside enterprises right now that needs a light shone on it. That’s because enterprises don't deploy a single agent and watch it run, they deploy fleets, each one calling APIs, calling other agents, reaching into applications that were never built…

    venturebeat.com21 days agoView details

  25. OpenAI’s executive exodus has one big winner

    Today on Decoder, I’m talking to Verge senior AI reporter Hayden Field about some pure Decoder bait: the seemingly-endless org chart changes at OpenAI, and how all of them seem to consolidate power under cofounder Greg Brockman, the company’s president. While Sam Altman is the CEO and still OpenAI’s most public face,…

    theverge.com21 days agoView details

  26. Hugging Face’s new robot is an adorable rollerskating duck

    Hugging Face's Pollen Robotics has launched its second cute AI robot, the Microduck, a one-eyed biped standing just under 10 inches tall. It's available to preorder now for $399 in cream, graphite, lavender, and sky blue, and Pollen Robotics says it plans to start shipping the little robot "before Christmas 2026." Vid…

    theverge.com21 days agoView details

  27. Plaud is launching AI earbuds

    Plaud has introduced a new AI wearable that's designed to record, transcribe, and summarize your conversations, only this time it looks like earbuds instead of a pin. The Plaud One Explorer Edition can be worn like traditional earbuds or used through its standalone charging case, and the case includes built-in 4G to u…

    theverge.com21 days agoView details

  28. Adobe is adding more AI to Photoshop

    Adobe is rolling out an AI-heavy update for Photoshop that includes a new "optional" interface dedicated to its AI tools. Launching in beta, the "AI Assisted Editor" view will show all of Photoshop's AI features in a single toolbar, including its prompt-based image editor, background remover, an AI image extender, and…

    theverge.com21 days agoView details

  29. When agents act on their own, governance has to live in the data layer

    Presented by EDB As enterprises give AI agents more autonomy — the ability to plan, decide, and act across systems without a human approving each step — a hard question moves to the center of every architecture review: When an agent tries to complete an action that it was never authorized to do, what actually stops it…

    venturebeat.com21 days agoView details

  30. Submit Your Questions: The Great Data Center Backlash

    You have questions about data centers, and WIRED has answers. Join our livestream on September 10 and our panel of experts will tell you everything you need to know.

    wired.com21 days agoView details

  31. Stop Touching Your Keyboard. Use This AI-Powered Microphone Instead

    The Relay Q, due next year, is the latest attempt to reposition voice as the most seamless method for human-computer interaction.

    wired.com21 days agoView details

  32. The UK Power Grid Has a Phantom Data Center Problem

    The UK’s energy regulator is using a variety of tricks to keep speculative data center projects from plugging into the power grid. The country’s AI ambitions hang in the balance.

    wired.com22 days agoView details

  33. MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

    arXiv:2608.26295v1 Announce Type: new Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce MemToC, a controlled benchmark for post-too…

    arxiv.org21 days agoView details

  34. MoganColBERT-TR: A Late-Interaction Multi-Vector Retrieval Model for Turkish

    arXiv:2608.26344v1 Announce Type: new Abstract: We previously reported a ModernBERT encoder trained from scratch for Turkish (MoganBERT-TR) and a single-vector embedding model built on top of it (MoganBERT-embed). This work introduces the third model in that lineage: MoganColBERT-TR, a multi-vector retrieval model tha…

    arxiv.org21 days agoView details

  35. Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

    arXiv:2608.26372v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something its deployer would prefer to deny, does i…

    arxiv.org21 days agoView details

  36. Survival-Guided Length Control for Efficient Diffusion Language Models

    arXiv:2608.26374v1 Announce Type: new Abstract: Diffusion language models (DLMs) generate text by iteratively denoising masked sequences, but standard decoding either fixes the sequence length or relies on ad hoc stopping rules, often leading to unnecessary denoising steps. We recast length selection as a discrete-tim…

    arxiv.org21 days agoView details

  37. Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries

    arXiv:2608.26385v1 Announce Type: new Abstract: Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that answers everything outscores one that declines when its knowledge base cannot support an answer. Building on the confidence-target analysis of Kalai et al. (2025), we p…

    arxiv.org21 days agoView details

  38. From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

    arXiv:2608.26163v1 Announce Type: new Abstract: Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as acoustic noise to be discarded. We present HealthCUES (Clinical Understanding from Embodied Sounds), a streaming pipeline for paraling…

    arxiv.org21 days agoView details

  39. A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs

    arXiv:2608.26177v1 Announce Type: new Abstract: Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models exhibit length collapse at 16k-token outputs, and multi-chapter stories frequently trigger the attribute drift characteristic of the ``lost-in-the-middle'' effect. Th…

    arxiv.org21 days agoView details

  40. Case2Flow: Bridging Patient Cases and Guideline Flowcharts through Multimodal Retrieval

    arXiv:2608.26414v1 Announce Type: new Abstract: Medical guidelines encode rich, evidence-based decision logic, yet the specific decision artifact a clinician needs is hard to locate within a guideline, let alone across guidelines covering plausible diseases and treatments. While guideline passages have supported end-t…

    arxiv.org21 days agoView details

  41. AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition

    arXiv:2608.26434v1 Announce Type: new Abstract: Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning…

    arxiv.org21 days agoView details

  42. Compositional Generalization via Structural Identification in a Category-Theoretic Framework

    arXiv:2608.26465v1 Announce Type: new Abstract: Compositional generalization is usually evaluated through model accuracy. We instead ask which structural or lexical identifications make held-out COGS examples admissible from the structures observed in training. Sentences are represented as functors from syntactic addr…

    arxiv.org21 days agoView details

  43. Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

    arXiv:2608.26449v1 Announce Type: new Abstract: Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters. In abugida scripts, vowels are written as combining marks; this pattern therefore splits each word at ev…

    arxiv.org21 days agoView details

  44. Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update

    arXiv:2608.26511v1 Announce Type: new Abstract: Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them…

    arxiv.org21 days agoView details

  45. Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue

    arXiv:2608.26529v1 Announce Type: new Abstract: In this paper, we explore multi-expert Conformal Risk Control (CRC) algorithms for pairwise LLM-as-a-Judge evaluation in open-ended dialogue. Our core insight is that multi-expert aggregation offers a complementary remedy to CRC: whereas CRC controls risk at the decision…

    arxiv.org21 days agoView details

  46. CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering

    arXiv:2608.26114v1 Announce Type: new Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, they often produce nu…

    arxiv.org21 days agoView details

  47. The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning

    arXiv:2608.26116v1 Announce Type: new Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution. We introduce a closed-loop framework based on au…

    arxiv.org21 days agoView details

  48. Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse

    arXiv:2608.26149v1 Announce Type: new Abstract: Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems. Relational data combine several sources of complexity, including large data volume, high-dimensional variables, high-cardinality categorical features,…

    arxiv.org21 days agoView details

  49. Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models

    arXiv:2608.26150v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline development for extracting model-relevant…

    arxiv.org21 days agoView details

  50. Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

    arXiv:2608.26151v1 Announce Type: new Abstract: Subscriber attrition is a costly, persistent challenge for telecommunications providers, with monthly churn of roughly 1.9% in mature markets eroding billions in revenue annually. Predictive models can flag at-risk customers accurately, yet they are routinely excluded fr…

    arxiv.org21 days agoView details

  51. A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in Anian

    arXiv:2608.26162v1 Announce Type: new Abstract: Safety-critical mental-health support systems must distinguish when supportive conversation is appropriate from when free-form generation should be blocked. This paper presents Anian, a safety-gated multimodal AI backend for perinatal mental-health support and mindfulnes…

    arxiv.org21 days agoView details

  52. Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment

    arXiv:2608.26165v1 Announce Type: new Abstract: Automated creativity assessment has been a long standing challenge, with traditional methods often being resource intensive or lacking practical accuracy. We introduce a novel approach by using Poly-Encoder for computationally efficient and accurate automated creativity…

    arxiv.org21 days agoView details

  53. AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

    arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agenti…

    arxiv.org21 days agoView details

  54. Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

    arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocult…

    arxiv.org21 days agoView details

  55. AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design

    arXiv:2608.26747v1 Announce Type: new Abstract: Scientific LLM agents have shown promise in literature reasoning, tool use, and experiment planning, but it remains unclear whether they can autonomously improve large, tightly coupled scientific machine-learning systems through executable code changes and computationall…

    arxiv.org21 days agoView details

  56. Comparing Chunking and Embedding Strategies for Turkish RAG Systems

    arXiv:2608.26192v1 Announce Type: new Abstract: How documents are segmented into retrievable chunks and how those chunks are embedded strongly affect Retrieval-Augmented Generation (RAG) quality, yet neither has been systematically studied for morphologically rich languages such as Turkish. We compare Turkish document…

    arxiv.org21 days agoView details

  57. Agent Seer: Synthesizing Scenarios from Specification Understanding

    arXiv:2608.26133v1 Announce Type: new Abstract: Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, an…

    arxiv.org21 days agoView details

  58. When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

    arXiv:2608.26319v1 Announce Type: new Abstract: The performance of textual neural models often degrades when their inputs are corrupted by noise such as typos, OCR errors, or dropped words. We study the degradation rate across neural models, both sentence embeddings and decoder-only LLMs, and find that how consistent…

    arxiv.org21 days agoView details

  59. PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

    arXiv:2608.26530v1 Announce Type: new Abstract: Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned fr…

    arxiv.org21 days agoView details

  60. Evaluating AI Generated Summaries for Cancer Patients

    arXiv:2608.26154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also raise concerns about accuracy, faithfuln…

    arxiv.org21 days agoView details

  61. Agentic AI for operating scientific instruments for nanoscale characterization

    arXiv:2608.26198v1 Announce Type: new Abstract: Operating a scientific instrument such as an atomic force microscope (AFM) requires continuous expert decision-making. A trained user defines the experimental intent, translates it into instrument commands, assesses incoming data, adjusts imaging parameters, and post-pro…

    arxiv.org21 days agoView details

  62. SAREF-based Ontology for Distributed AI Workflows across the Edge-Fog-Cloud Continuum

    arXiv:2608.26160v1 Announce Type: new Abstract: Nowadays semantic models provide limited support for representing distributed AI workflows and their execution across heterogeneous edge, fog, and cloud environments. Therefore, AI processes and resources are often described using incompatible semantic representations, a…

    arxiv.org21 days agoView details

  63. Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing

    arXiv:2608.26710v1 Announce Type: new Abstract: AI text detectors are increasingly employed in academic settings, but it remains unclear whether their outputs reflect AI authorship itself or broader linguistic features associated with polished academic English. Previous studies have reported high false-positive rates…

    arxiv.org21 days agoView details

  64. Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments

    arXiv:2608.26188v1 Announce Type: new Abstract: Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel. We ask whether such judgments track measured risk or the patterns attached to an urban neighborhood's name. We probe seven instruct-tuned…

    arxiv.org21 days agoView details

  65. A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving

    arXiv:2608.26164v1 Announce Type: new Abstract: Large language models can interpret natural-language chemistry questions, but their internal reasoning is difficult to inspect, constrain, and validate. This paper presents ChemOntoRule, a proof-of-concept symbolic core for AI-assisted school-level chemistry problem solv…

    arxiv.org21 days agoView details

  66. On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study

    arXiv:2608.26292v1 Announce Type: new Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this decision at all. Using INLAY, a gradient-free editor we built to obtain…

    arxiv.org21 days agoView details

  67. Knowledge Cards: Structured Knowledge for AI Systems

    arXiv:2608.26176v1 Announce Type: new Abstract: AI systems whose outputs inform real decisions, and increasingly consequential ones, require something that current documentation practice does not provide: a structured, inspectable representation of the knowledge they need to ground, contextualize, and reason about tho…

    arxiv.org21 days agoView details

  68. Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript

    arXiv:2608.26167v1 Announce Type: new Abstract: Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate abstention from an unsupported prediction. Seven large language models were evaluated on the TAME Pain speech cor…

    arxiv.org21 days agoView details

  69. Invocation-Level Reliability of Tool-Using Agents

    arXiv:2608.26189v1 Announce Type: new Abstract: Tool-using agents fail two ways: choosing the wrong tool, or forming wrong arguments, and an early failure of either kind can silently corrupt everything downstream. We measure a correct-invocation rate that separates the two, under both a clean teacher-forced context an…

    arxiv.org21 days agoView details

  70. TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

    arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations…

    arxiv.org21 days agoView details

  71. Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

    arXiv:2608.26190v1 Announce Type: new Abstract: World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or features, which introduces unnecessary complexity and limits their eff…

    arxiv.org21 days agoView details

  72. Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs

    arXiv:2608.26191v1 Announce Type: new Abstract: Incident risk prediction from longitudinal electronic health records (EHRs) is challenging because relevant signals are multimodal, weak in isolation, and distributed across irregular patient histories. We propose structured evidence routing, a router-predictor-reviewer…

    arxiv.org21 days agoView details

  73. Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

    arXiv:2608.26199v1 Announce Type: new Abstract: We ask whether AI agents powered by locally deployed large language models can reliably automate expert-defined hardware design workflows in an industry-realistic tool-calling setting. In these environments, engineers issue repetitive, dependency-ordered operations---suc…

    arxiv.org21 days agoView details

  74. Same Model, Different Harness: Different Coding-Agent Results

    arXiv:2608.26218v1 Announce Type: new Abstract: A coding agent combines a model with a harness, which decides what the model sees, which tools it can use, and how the work continues. We ask whether changing the harness changes the result when the model and task stay fixed. We compare two configurations of the same har…

    arxiv.org21 days agoView details

  75. Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy

    arXiv:2608.26225v1 Announce Type: new Abstract: Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes, and budgets them. The machinery such orchestrators reach for is the service mesh's: retry, timeout, and error-rate circuit breaking. We report a failure study of a…

    arxiv.org21 days agoView details

  76. The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

    arXiv:2608.26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment question: when do extended thinking tokens earn their cost? We introduce the Token Economy Score (TES), a marginal benchmarking metric that measures the accuracy gain of a r…

    arxiv.org21 days agoView details

  77. 6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation

    arXiv:2608.26236v1 Announce Type: new Abstract: We present a six-stage framework for auditing the reproducibility of scientific claims across a research literature within the computer science domain, and instantiate our framework for the neuro-symbolic AI (NSAI) subdomain. Instantiating the framework on the NSAI subdo…

    arxiv.org21 days agoView details

  78. SKILL.state: Scalable Long-Horizon Agent Skills

    arXiv:2608.26263v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, actions, and intermediate reasoning traces to an ever-growing conversat…

    arxiv.org21 days agoView details

  79. FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence

    arXiv:2608.26310v1 Announce Type: new Abstract: Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical challenge. Existing evaluation approaches largely depend on model-based natural-language judg…

    arxiv.org21 days agoView details

  80. FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

    arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magneti…

    arxiv.org21 days agoView details

  81. ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving

    arXiv:2608.26334v1 Announce Type: new Abstract: Automated theorem proving offers a natural foundation for recursive self-improvement in scientific discovery. However, existing neural provers do not fully preserve this recursive structure, where the learning process should be self-improving over time. Existing methods…

    arxiv.org21 days agoView details

  82. Fine-Tuning of Transformer models with Frames

    arXiv:2608.26430v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective solutions for fine-tuning large-scale pre-trained models; however, their memory requirements scale with the size of the model, $\mathcal{O}(dr)$, where $d$ is the model's h…

    arxiv.org21 days agoView details

  83. Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI

    arXiv:2608.26442v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can improve performance on complex tasks. However, many existing approaches rely on fixed or preallocated reasoning controls, such as fixed token budgets, pre-execution dif…

    arxiv.org21 days agoView details

  84. DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows

    arXiv:2608.26546v1 Announce Type: new Abstract: Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and more stable than those…

    arxiv.org21 days agoView details

  85. Relational Over-Regularization: Graph-Based AI-Generated Text Detection via Sentence Transition Deviation

    arXiv:2608.26694v1 Announce Type: new Abstract: Detecting AI-generated text (AIGT) remains challenging because existing approaches rely on token-level statistical signals or independent stylometric features, causing them to overfit to specific generators and fail under distribution shift. We identify a structural sign…

    arxiv.org21 days agoView details

  86. Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

    arXiv:2608.26730v1 Announce Type: new Abstract: Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often entails repeated post-training. Autonomous systems automate parts of this process by proposing updates, training candidates, and using evaluation feedback…

    arxiv.org21 days agoView details

  87. Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI

    arXiv:2608.26182v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for verbal interaction in social robots, yet prompt design in human-robot interaction (HRI) remains underspecified. As a result, robots may present hallucinated capabilities, unclear behavioural boundaries, and misleadin…

    arxiv.org21 days agoView details

  88. TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education

    arXiv:2608.26184v1 Announce Type: new Abstract: AI programming tutors provide scalable support, yet lack the behavioral context human tutors rely on to adapt support to learners' needs. We present TutorTrace, a dataset and behavioral abstraction pipeline that makes learners' behavioral context visible and computable i…

    arxiv.org21 days agoView details

  89. LLM Agents for Time-Series: A Survey

    arXiv:2608.26226v1 Announce Type: new Abstract: LLM-based agents are increasingly being developed for time-series problems, but their design choices vary substantially across task settings. This survey adopts a problem-driven taxonomy that organizes these systems by the time-series problems they address rather than by…

    arxiv.org21 days agoView details

  90. Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case

    arXiv:2608.26750v1 Announce Type: new Abstract: Data lakes rely on metadata to remain usable, yet this meta data is often limited or weakly informative for column relationship discovery, especially in ERP-derived datasets with coded or abbreviated schema labels. We propose ColRel, a two-stage method that builds column…

    arxiv.org21 days agoView details

  91. Decoupling Planning and Control for Instructable Agents

    arXiv:2608.26788v1 Announce Type: new Abstract: Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but struggle to realize such plans as reliable low-latency action sequences in unfamiliar environments. At…

    arxiv.org21 days agoView details

  92. Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores

    arXiv:2608.26137v1 Announce Type: new Abstract: Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated score is trustworthy only if it is accur…

    arxiv.org21 days agoView details

  93. EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on Clinical EEG

    arXiv:2608.26153v1 Announce Type: new Abstract: Clinical electroencephalography (EEG) reporting remains largely manual and time-consuming, and current EEG software ecosystems do not produce the structured EEG-text supervision needed for training modern language models. Most toolboxes focus on visualization or preproce…

    arxiv.org21 days agoView details

  94. DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

    arXiv:2608.26757v1 Announce Type: new Abstract: Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately. Modern LLMs can produce visually plausible, instruction-compliant charts, yet data-lev…

    arxiv.org21 days agoView details

  95. Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models

    arXiv:2608.26161v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sampling remains a persistent challenge. When generating random variables, LLMs consistently exhibit systematic biases that warp the ta…

    arxiv.org21 days agoView details

  96. AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes

    arXiv:2608.26193v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body…

    arxiv.org21 days agoView details

  97. The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasting

    arXiv:2608.26134v1 Announce Type: new Abstract: Energy forecasting aims to maximize accuracy to ensure energy efficiency by reducing energy waste, an objective that applies equally to on-device forecasting for mission-critical edge environments, including military systems. However, this paper identifies the Accuracy-E…

    arxiv.org21 days agoView details

  98. GameWAM: A World Action Model for Video Games

    arXiv:2608.26200v1 Announce Type: new Abstract: Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game w…

    arxiv.org21 days agoView details

  99. Accelerating Scientific Research with Gemini in the Real-World

    arXiv:2608.26701v1 Announce Type: new Abstract: We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving beyond in silico hypo…

    arxiv.org21 days agoView details

  100. Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds

    arXiv:2608.26743v1 Announce Type: new Abstract: Enterprises fine-tune language models on proprietary data that may later require removal due to privacy, contractual, or compliance obligations. Selective unlearning removes requested knowledge while preserving model utility, offering a practical alternative to full retr…

    arxiv.org21 days agoView details

  101. Data Science Approaches to Evaluating Honours Candidates

    arXiv:2608.26135v1 Announce Type: new Abstract: We present a modular data-science pipeline for estimating public sentiment towards individuals from fragmented, unstructured open-source intelligence (OSINT). The method chains web search, text extraction, relevance filtering, tokenisation, co-reference resolution, and s…

    arxiv.org21 days agoView details

  102. Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems

    arXiv:2608.26306v1 Announce Type: new Abstract: A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is correct at check time but stale by actuation. This creates an Execute-stage time-of-check to time-of-use (TOCTOU) hazard. We study verdict freshness: whether a guardrail…

    arxiv.org21 days agoView details

  103. DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

    arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap w…

    arxiv.org21 days agoView details

  104. ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements

    arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We propose ElementCheck, a complexity-aware f…

    arxiv.org21 days agoView details

  105. Co-Evolving Structured Knowledge and Reasoning in Language Models

    arXiv:2608.26386v1 Announce Type: new Abstract: Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved information. Structured knowledge bases offer…

    arxiv.org21 days agoView details

  106. Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

    arXiv:2608.26535v1 Announce Type: new Abstract: Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no longer reside in a…

    arxiv.org21 days agoView details

  107. Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap

    arXiv:2608.26111v1 Announce Type: new Abstract: Battery Prognostics and Health Management (BPHM) is critical for ensuring the safe, reliable, and cost-effective operation of batteries across electric vehicles, grid storage, and consumer electronics. Conventional BPHM approaches, including physics-based models and task…

    arxiv.org21 days agoView details

  108. Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS) System for Sanskrit

    arXiv:2608.26146v1 Announce Type: new Abstract: We present Vagdhenu, a vrutta (meter) aware shloka-to-chant system for Sanskrit: a text-to-speech system that maps a metrical verse to its chanted parayana recitation at high fidelity. This is an experience report, not a new architecture. We take an off-the-shelf flow-ma…

    arxiv.org21 days agoView details

  109. Evaluating Language Models in Realistic Conversational Contexts

    arXiv:2608.26131v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for summarization, translation, or short-form QA…

    arxiv.org21 days agoView details

  110. SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation

    arXiv:2608.26683v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) faces significant challenges in maintaining robust coordination under noisy observations. Although observation disturbances are often introduced independently across agents, their downstream effects on cooperative dec…

    arxiv.org21 days agoView details

  111. Selection Bias Correction in Retail Intelligence

    arXiv:2608.26156v1 Announce Type: new Abstract: Retail intelligence often relies on monitoring popular, high-velocity products, potentially biasing economic indicators by ignoring the "long tail" of niche items. This simulation study investigates selection bias in inflation estimation and compares correction methods a…

    arxiv.org21 days agoView details

  112. VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation

    arXiv:2608.26155v1 Announce Type: new Abstract: Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of high-quality non-English image-text supervision. Although multilingual text data is…

    arxiv.org21 days agoView details

  113. Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media

    arXiv:2608.26138v1 Announce Type: new Abstract: We introduce the Cross-Platform Fairness Evaluation (CPFE) framework -- a five-axis audit protocol covering discriminative performance, calibration, statistical significance, prediction equity, and attribution stability -- and apply it to four transformer models (BERT, R…

    arxiv.org21 days agoView details

  114. Affix Cache for Diffusion Large Language Models

    arXiv:2608.26140v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding and bidirectional context modeling, but efficient inference remains challenging. Unlike autoregressive systems, whose key-value (KV) cache can be reused for shared prefixes, DLLMs couple the KV st…

    arxiv.org21 days agoView details

  115. Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation

    arXiv:2608.26142v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES performance, their prohibitive computational…

    arxiv.org21 days agoView details

  116. Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators

    arXiv:2608.26148v1 Announce Type: new Abstract: Depression affects millions worldwide, yet diagnosis relies on subjective self-reports that may miss authentic behavior. This paper presents an approach linking speech acoustics to DSM-5 depressive-behavior indicators through a transparent Linkage Framework. Unlike black…

    arxiv.org21 days agoView details

  117. Assessing mentalization in humans and large language models

    arXiv:2608.26291v1 Announce Type: new Abstract: Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet wheth…

    arxiv.org21 days agoView details

  118. Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

    arXiv:2608.26159v1 Announce Type: new Abstract: Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors. Specifically, an LLM may recognize outputs from other copies of the same model and make biased judgments o…

    arxiv.org21 days agoView details

  119. When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models

    arXiv:2608.26187v1 Announce Type: new Abstract: Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while r…

    arxiv.org21 days agoView details

  120. Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search

    arXiv:2608.26152v1 Announce Type: new Abstract: Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduced, or even eliminated, the n…

    arxiv.org21 days agoView details

  121. Syntax vs. Semantics: How Transformers Learn Deep Dependencies

    arXiv:2608.26139v1 Announce Type: new Abstract: Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this learning process as a competition between…

    arxiv.org21 days agoView details

  122. Categorizer Automata for Discounted-Sum Payoffs

    arXiv:2608.26763v1 Announce Type: new Abstract: Categorizing continuous data into discrete bins is a fundamental operation in artificial intelligence. We introduce the categorizer automaton, a deterministic automaton that reads an infinite sequence of rewards and identifies which of finitely many bins contains its dis…

    arxiv.org21 days agoView details

  123. AI Revealed Preferences

    arXiv:2608.26178v1 Announce Type: new Abstract: There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test 20 language models and find a range of preferences---stable dispositions to choose certain kinds of tasks. We run three forced-choice e…

    arxiv.org21 days agoView details

  124. PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants

    arXiv:2608.26180v1 Announce Type: new Abstract: Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of evaluation: a fluent…

    arxiv.org21 days agoView details

  125. Cross-lingual Representation Learning via Centroid Intervention Fusion

    arXiv:2608.26357v1 Announce Type: new Abstract: Large language models (LLMs) exhibit uneven multilingual performance, especially when dealing with low-resource languages. Inference-time intervention offers a lightweight way to improve cross-lingual transfer by modifying the hidden states produced by the LLMs during th…

    arxiv.org21 days agoView details

  126. How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

    arXiv:2608.26327v1 Announce Type: new Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a systematic cross-model evaluation using…

    arxiv.org21 days agoView details

  127. Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes

    arXiv:2608.26143v1 Announce Type: new Abstract: Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for spreading hate. It remains very difficult…

    arxiv.org21 days agoView details

  128. Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap

    arXiv:2608.26144v1 Announce Type: new Abstract: Explainable AI (XAI) is now a major theme in NLP; however, Arabic NLP remains under-explained in three connected senses. First, there is a method gap: Arabic XAI relies heavily on a small set of post-hoc techniques such as LIME, SHAP, attention visualization, and salienc…

    arxiv.org21 days agoView details

  129. AI Control Scientist: LLM-driven Agentic System for Automated Control Design

    arXiv:2608.26780v1 Announce Type: new Abstract: Control system design is critical for modern industry, such as chemical process temperature regulation and aero-engine control. However,traditional control design workflows rely heavily on expert knowledge and extensive manual parameter tuning, resulting in limited effic…

    arxiv.org21 days agoView details

  130. Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification

    arXiv:2608.26329v1 Announce Type: new Abstract: While tool-augmented Large Language Models have significantly improved multi-step reasoning in quantitative STEM tasks, a critical residual failure mode remains: intermediate reasoning steps that are syntactically well-formed, mathematically executable, and unit-consiste…

    arxiv.org21 days agoView details

  131. Recipes for Steering and Scaling LLMs via Sampling

    arXiv:2608.26120v1 Announce Type: new Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. In this paper, we prese…

    arxiv.org21 days agoView details

  132. Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset

    arXiv:2608.26109v1 Announce Type: new Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension…

    arxiv.org21 days agoView details

  133. LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

    arXiv:2608.26145v1 Announce Type: new Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature re…

    arxiv.org21 days agoView details

  134. Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention

    arXiv:2608.26168v1 Announce Type: new Abstract: The lifecycle of hallucination in LLMs is a concept that enables building solid frameworks on the control and reliability of LLMs in high-stakes environments, including health, legal, and scientific research. Although previous surveys have primarily focused on detection…

    arxiv.org21 days agoView details

  135. EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction

    arXiv:2608.26107v1 Announce Type: new Abstract: Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes. However, existing models often suffer from limited early detection capability and insufficient interpretability, leading…

    arxiv.org21 days agoView details

  136. SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning

    arXiv:2608.26550v1 Announce Type: new Abstract: Reinforcement learning-based knowledge distillation has the potential to transfer complex reasoning from teacher to student models, yet it currently faces a critical dilemma: researchers must choose between sparse outcome-based rewards, which provide insufficient logical…

    arxiv.org21 days agoView details

  137. TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

    arXiv:2608.26112v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the accepted length. However, existing tree…

    arxiv.org21 days agoView details

  138. Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

    arXiv:2608.26121v1 Announce Type: new Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong. The usual way to act on this, teaching a model to abst…

    arxiv.org21 days agoView details

  139. LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

    arXiv:2608.26389v1 Announce Type: new Abstract: SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior evaluations use varied benchmarks, inconsi…

    arxiv.org21 days agoView details

  140. PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices

    arXiv:2608.26113v1 Announce Type: new Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications. PICasso couples a structured NL -> YAML -> GDS generation pipeline with PDK aware knowledge i…

    arxiv.org21 days agoView details

  141. Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

    arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype. We pr…

    arxiv.org21 days agoView details

  142. Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses

    arXiv:2608.26130v1 Announce Type: new Abstract: Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy, pagination, is available in every protocol that…

    arxiv.org21 days agoView details

  143. Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

    arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and costly to maintain, whereas unconstrained LLM…

    arxiv.org21 days agoView details

  144. GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions

    arXiv:2608.26157v1 Announce Type: new Abstract: Natural-language analytics over enterprise data warehouses is increasingly important, but production use is limited by hallucinated metrics, invalid joins, wrong grain, unsafe data access, and unsupported explanations. Existing text-to-SQL systems often ground generation…

    arxiv.org21 days agoView details

  145. Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound

    arXiv:2608.26136v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, and an appealing way to aim them at reasoning is to curate their data with a signal reinforcement learning already produces: the reward. We build such a reward-informed S…

    arxiv.org21 days agoView details

  146. AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking

    arXiv:2608.26141v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. While this capability significantly improves performance on challenging tasks, current models apply such deep reasoning uniformly to al…

    arxiv.org21 days agoView details

  147. Five Primitives for Governing Autonomous AI Agents at Runtime

    arXiv:2608.26696v1 Announce Type: new Abstract: Enterprise deployments of autonomous AI agents inherit a control model built for human users and long-lived services, and the fit fails in three specific ways: agent principals are ephemeral, appearing and vanishing faster than provisioning; their actions are selected by…

    arxiv.org21 days agoView details

  148. CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

    arXiv:2608.26147v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential for medical reasoning, yet the scarcity and cost of expert-annotated data constrain their progress. While reinforcement learning offers a scalable alternative, standard outcome-based methods in medicine often suffe…

    arxiv.org21 days agoView details

  149. Investigating the Influence of Prompt and Response Languages on LLM Content Generation

    arXiv:2608.26186v1 Announce Type: new Abstract: This study examines how prompt and response language influence the behavior of large language models. Using five models, we evaluated answers to 68 non translation questions across four language conditions: English to English, English to Norwegian, Norwegian to Norwegian…

    arxiv.org21 days agoView details

  150. A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers

    arXiv:2608.26194v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to include various modalities such as speech…

    arxiv.org21 days agoView details

  151. Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion

    arXiv:2608.26185v1 Announce Type: new Abstract: Equal participation in co-located discussion is important for effective collaboration, yet people often hold back when they anticipate negative interpersonal or professional consequences, especially when raising a point requires voicing it themselves. We present SecondVo…

    arxiv.org21 days agoView details

  152. Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors

    arXiv:2608.26175v1 Announce Type: new Abstract: Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLingua-2 report strong results on English benchmarks. Most other languages already pay a token premium: the same content costs 1.3-1.8…

    arxiv.org21 days agoView details

  1. Qwen/Qwen3.8-Flash-Next

    image-text-to-text · transformers · safetensors · qwen4_exp

    huggingface.co22 days ago5337 ptsView details

  2. unsloth/Qwen3.8-Flash-Next-GGUF

    image-text-to-text · gguf · unsloth · image-text-to-text

    huggingface.co22 days ago975 ptsView details

  3. pipecat-ai/phonellm-alpha-1

    text-generation · transformers · safetensors · nemotron_h

    huggingface.co21 days ago213 ptsView details

  4. alibaba-pai/MiniMax-H3-Acc-LoRAs

    text-to-video · videox_fun · text-to-video · arxiv:2607.26004

    huggingface.co22 days ago181 ptsView details

  5. Qwen/Qwen3.8-Flash-Next-FP8

    image-text-to-text · transformers · safetensors · qwen4_exp

    huggingface.co22 days ago181 ptsView details

  6. Vxtzq/Crowd-v1

    text-generation · text-generation · ia · dataset:openbmb/Ultra-FineWeb-L3

    huggingface.co22 days ago7 ptsView details

  7. shawnw3i/Qwen3.8-27B-AWQ-MTP

    image-text-to-text · transformers · safetensors · qwen3_5

    huggingface.co22 days ago4 ptsView details

  8. amd/Kimi-K3-Quark-MXFP4-AttnFP8

    image-text-to-text · transformers · safetensors · kimi_k3

    huggingface.co22 days agoView details

  9. abdukuzi45/qwen3.5-4b-amharic-sft

    tensorboard · safetensors · region:us

    huggingface.co22 days agoView details

  1. ggml-org/llama.cpp b10665

    <details open> model: add DSpark support for Nemotron3.5 (#27804) * model: add DSpark support for Nemotron3.5 * Update src/models/dflash.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingf…

    github.com21 days agoView details

  2. ggml-org/llama.cpp b10664

    <details open> ggml-hexagon: add HTP unary ops for ABS and LOG (#27786) Add HVX-accelerated implementations for GGML_OP_LOG and GGML_UNARY_OP_ABS on the HTP backend. - Register HTP_OP_UNARY_ABS and HTP_OP_UNARY_LOG in op_remap_to_htp() - Add ABS and LOG to ggml_backend_hexagon_d…

    github.com21 days agoView details

  3. langchain-ai/langchain langchain==1.4.0a1

    Initial release fix(langchain): name the content type MCP conversion could not handle release(langchain): 1.4.0a1 test(langchain): skip MCP tests on a pydantic older than `mcp` supports test(langchain): drive MCP tests through FastMCP's own utilities fix(langchain/mcp): review e…

    github.com21 days agoView details

  4. langchain-ai/langchain langchain-fireworks==1.6.1

    Changes since langchain-fireworks==1.6.0 release(fireworks): 1.6.1 (#39975) fix(fireworks): drop reasoning history blocks (#39973) chore(model-profiles): refresh model profile data (#39844)

    github.com21 days agoView details

  5. anthropics/anthropic-sdk-python v1.2.0

    ## 1.2.0 (2026-08-27) Full Changelog: [v1.1.0...v1.2.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.1.0...v1.2.0) ### Features * **api:** beta files/skills namespaces use GA shapes; drop dated beta header pins ([9df4565](https://github.com/anthropics/anthropic-…

    github.com21 days agoView details

  6. langchain-ai/langchain langchain==1.3.18

    Changes since langchain==1.3.17 release(langchain): 1.3.18 (#39966) fix(langchain): preserve content-block shape in PIIMiddleware redaction (#39894) fix(core): shore up indexing in genai v1 streaming content (#39964)

    github.com21 days agoView details

  7. ggml-org/llama.cpp b10644

    <details open> models : support nanbeige4.2-3B (#27730) Co-authored-by: admin <lizongqiang@kanzhun.com> </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/43308396> **macOS/iOS:** - [macOS Apple Silicon (arm64)](…

    github.com22 days agoView details