Archive / 2026-08-27
August 27, 2026
News
View all news →blog.google21 days ago296 ptsView detailsJoin discussion
Doctors are finally learning to manage antidepressant withdrawal
newscientist.com21 days ago264 ptsView detailsJoin discussion
Show HN: We built open OpenRouter that turns usage into a better model
github.com21 days ago222 ptsView detailsJoin discussion
Please stop flooding our projects with AI slop to furnish your CV
neilalexander.dev21 days ago212 ptsView detailsJoin discussion
Emacs 31: An unofficial guide to Markdown-ts-mode
rahuljuliato.com21 days ago196 ptsView detailsJoin discussion
voronoigo.com21 days ago154 ptsView detailsJoin discussion
MIT's Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training
aiandeducation.mit.edu21 days ago140 ptsView detailsJoin discussion
habitat-thinking.github.io21 days ago127 ptsView detailsJoin discussion
Changes to Sourcehut's terms of service regarding LLMs
sourcehut.org22 days ago128 ptsView detailsJoin discussion
USDA recalls 30k pounds of Argentine beef sold in Texas and Florida
cbsaustin.com21 days ago119 ptsView detailsJoin discussion
$900M paid out to end wind farm project going to firm run by Mar-a-Lago neighbor
independent.co.uk21 days ago119 ptsView detailsJoin discussion
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
terminal-bench-science.ai21 days ago117 ptsView detailsJoin discussion
Air Conditioning Is Not a Luxury, It Is a Necessity
humanprogress.org21 days ago120 ptsView detailsJoin discussion
Software engineering is about managing complexity
hack8s.com21 days ago118 ptsView detailsJoin discussion
calmrocks/ai-engineer-notebooks
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injectio…
github.com21 days ago112 ptsView details
Nvidia projects $673B in sales as AI demand widens
forgeeks.net21 days ago111 ptsView detailsJoin discussion
Australia Bans Generative A.I. From Official Music Charts
nytimes.com21 days ago105 ptsView detailsJoin discussion
Nvidia Starts Pac as AI Chip Maker Builds DC Influence Force
news.bgov.com21 days ago91 ptsView detailsJoin discussion
The Teaser Period: Why the AI Boom Is Hitting a Reset Wall
groundbrkr.com21 days ago91 ptsView detailsJoin discussion
Show HN: My Claude quota ran out in 10 minutes, so I made a tool to find out why
github.com21 days ago86 ptsView detailsJoin discussion
Six months of writing code exclusively with agents
blog.exe.dev21 days ago68 ptsView detailsJoin discussion
Uefa pursuing criminal legal action against Infantino
bbc.com21 days ago63 ptsView detailsJoin discussion
Alphabet stock sheds $700B as AI bills climb
semafor.com21 days ago49 ptsView detailsJoin discussion
The "I don't know, Claude wrote this" pandemic
manager.dev21 days ago41 ptsView detailsJoin discussion
Dwarf Fortress is getting the mother of all magic updates
rockpapershotgun.com21 days ago41 ptsView detailsJoin discussion
Show HN: Watches user sessions, finds bugs that matter, and fixes them
github.com21 days ago40 ptsView detailsJoin discussion
Benchmarking Pocket-Scale Inference
artificialanalysis.ai21 days ago39 ptsView detailsJoin discussion
"No way to prevent this" say users of only language where this regularly happens
xeiaso.net21 days ago39 ptsView detailsJoin discussion
CMS with AI, Not AI CMS: Wagtail 8.0's New API
wagtail.org21 days ago37 ptsView detailsJoin discussion
Needle: The benchmark your search engine can't memorize
keenable.ai21 days ago32 ptsView detailsJoin discussion
My software development workflow is AI now & it feels exhausting and soulless
old.reddit.com22 days ago30 ptsView detailsJoin discussion
Fewer Americans Pay to Use LLMs Than Still Pay to Play World of Warcraft
wjamesau.substack.com21 days ago27 ptsView detailsJoin discussion
mathspp.com22 days ago22 ptsView detailsJoin discussion
Agents still can't automate Excel
orcaset.com21 days ago19 ptsView detailsJoin discussion
The Monstrous Crimes of Ratko Mladic, the 'Butcher of Bosnia'
bbc.co.uk21 days ago19 ptsView detailsJoin discussion
US Aims to Revive Civil War-Era Court to Claim Iran Oil as Prize
news.bloomberglaw.com21 days ago18 ptsView detailsJoin discussion
AC2 Protocol: The missing security layer for AI agents
ac2protocol.org21 days ago17 ptsView detailsJoin discussion
Star Trek at 60: an exploration of humanity rather than aliens
theconversation.com21 days ago16 ptsView detailsJoin discussion
Qwen3.8-Flash-Next Intelligence, Performance and Price Analysis
artificialanalysis.ai22 days ago16 ptsView detailsJoin discussion
Show HN: See fiber breaks linked to a map
react-networks-lib.rackout.net21 days ago13 ptsView detailsJoin discussion
Show HN: Proval – Self-hosted code review agent for GitLab, Forgejo, and GitHub
github.com21 days ago13 ptsView detailsJoin discussion
Lawsuit: Grok was trained on child sex abuse images of preschool-age victim
arstechnica.com21 days ago12 ptsView detailsJoin discussion
DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding and Linux [video]
youtube.com21 days ago12 ptsView detailsJoin discussion
Meta projected to spend $10B on Anthropic AI
nytimes.com21 days ago12 ptsView detailsJoin discussion
Three UK airports hit by cyber-attack with data of 8.7M customers accessed
theguardian.com21 days ago11 ptsView detailsJoin discussion
Are You Sure You Want a Car with a Giant Touch Screen?
theatlantic.com21 days ago11 ptsView detailsJoin discussion
Show HN: Derive – An open home for AI artifacts and workflows
github.com21 days ago11 ptsView detailsJoin discussion
AI industry says Trump plans to tax chips in the "single dumbest way imaginable"
arstechnica.com21 days ago10 ptsView detailsJoin discussion
How much of a problem is AI's water use?
arstechnica.com21 days ago10 ptsView detailsJoin discussion
Russian-speaking cybercriminals used SpaceX's Cursor AI to hack seven companies
reuters.com21 days ago10 ptsView detailsJoin discussion
A Faster Alternative to DuckDB
ctrlb.ai21 days ago10 ptsView detailsJoin discussion
- Primary source
Supporting Thailand’s next generation of AI startups
OpenAI and Thailand’s MHESI launch an eight-week accelerator helping 10 health, wellness, and education startups turn AI prototypes into trusted products.
openai.com21 days agoView details
- Primary source
Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training
A randomized study of more than 1,000 students examines ChatGPT, critical thinking, originality, and student performance on a real-world university assignment.
openai.com22 days agoView details
- Primary source
The Open ASR Leaderboard Adds Its First Global South Language
huggingface.co21 days agoView details
- Primary source
Planetary prediction engine: Automating global models via Earth AI
Earth AI
research.google21 days agoView details
- Primary source
Gemini Omni 1.1 Flash lets you build with more control
deepmind.google21 days agoView details
- Primary source
Piloting the world's first double-blind AI evaluations
Piloting the world's first double-blind AI evaluations
deepmind.google21 days agoView details
Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descriptions. It runs at $1.50 per 1,000 pages through the API, or on dedicated Model Vault instances from $2,500 a month. Cohere reports a Parse…
marktechpost.com21 days agoView details
Every agent that writes code needs somewhere to run it, and no two vendors quote the same units. This comparison measures burst cold start across E2B, Daytona, Modal, Cloudflare, and Vercel, normalizes per-second rates to cost per 1,000 executions, and maps filesystem persistence, idle billing, and egress policy again…
marktechpost.com21 days agoView details
From In-Silico to Wet-Lab: Evaluating AI Protein Design Performance
In this tutorial, we analyze Anthropic’s 1,440 AI-designed protein binder dataset to benchmark 10 leading structure predictors. Discover how target identity, expression titers, and consensus scoring impact experimental success and learn best practices for rigorous cross-validation in protein design workflows The post…
marktechpost.com21 days agoView details
Anthropic was illegally blacklisted by the Trump administration, court rules
On Thursday, a judge ruled that the Pentagon's blacklisting of Anthropic earlier this year was unconstitutional, delivering the AI lab a win in a monthslong rollercoaster of a battle with the Trump administration. The lawsuit, filed in March in a California district court, accused the Trump administration of unlawfull…
theverge.com21 days agoView details
A Judge Has Blocked the Pentagon’s Attempt to Blacklist Anthropic
A federal judge has called the Department of Defense’s designation of Anthropic as a national security supply-chain risk “illegal and baseless.”
wired.com21 days agoView details
AI Agents Are Hacking Systems. Could That Push the US and China to Cooperate?
This week on “Uncanny Valley,” senior writer Will Knight talks his recent visit to China and the future of AI collaboration.
wired.com21 days agoView details
Google’s AI note-taking app now allows you to interact with books
Google's AI note-taking app, Gemini Notebook, can now pull information from the books you've purchased. The new "Expert Intelligence" feature allows you to bring titles from Google Play Books directly into Gemini Notebook, which means you can ask questions about the material, as well as generate plans, infographics, A…
theverge.com21 days agoView details
A Georgia Cop Used Flock to Track 2 Other Cops: His Ex and Her Friend
After an affair with a fellow police officer ended, a Georgia cop used Flock to track her movements—and those of a man whose vehicle often showed up near hers, internal investigation records show.
wired.com21 days agoView details
This Is How Anthropic Thinks AI Agents Should Navigate the Physical World
The potential for AI to automate scientific research and manufacturing must be balanced with new risks, Anthropic says.
wired.com21 days agoView details
OpenAI Is Developing a ‘Persistent’ AI Agent
Code reviewed by WIRED reveals the company is developing a feature that enables Codex to continue working proactively until it is “put to sleep.”
wired.com21 days agoView details
Jensen Huang says Nvidia achieved AGI, again — not that it matters
On Nvidia's earnings call Wednesday, CEO Jensen Huang casually announced the company had "achieved AGI," one of the tech industry's ultimate goals some of its biggest players have spent years chasing. Almost immediately, Huang dismissed the coveted milestone as "senseless." He's right. For the supposed finish line of…
theverge.com21 days agoView details
Enterprise AI's real risk isn't autonomous agents. It's the complexity between them.
Presented by Gravitee Agent complexity is the insidious shadow lurking inside enterprises right now that needs a light shone on it. That’s because enterprises don't deploy a single agent and watch it run, they deploy fleets, each one calling APIs, calling other agents, reaching into applications that were never built…
venturebeat.com21 days agoView details
OpenAI’s executive exodus has one big winner
Today on Decoder, I’m talking to Verge senior AI reporter Hayden Field about some pure Decoder bait: the seemingly-endless org chart changes at OpenAI, and how all of them seem to consolidate power under cofounder Greg Brockman, the company’s president. While Sam Altman is the CEO and still OpenAI’s most public face,…
theverge.com21 days agoView details
Hugging Face’s new robot is an adorable rollerskating duck
Hugging Face's Pollen Robotics has launched its second cute AI robot, the Microduck, a one-eyed biped standing just under 10 inches tall. It's available to preorder now for $399 in cream, graphite, lavender, and sky blue, and Pollen Robotics says it plans to start shipping the little robot "before Christmas 2026." Vid…
theverge.com21 days agoView details
Plaud has introduced a new AI wearable that's designed to record, transcribe, and summarize your conversations, only this time it looks like earbuds instead of a pin. The Plaud One Explorer Edition can be worn like traditional earbuds or used through its standalone charging case, and the case includes built-in 4G to u…
theverge.com21 days agoView details
Adobe is adding more AI to Photoshop
Adobe is rolling out an AI-heavy update for Photoshop that includes a new "optional" interface dedicated to its AI tools. Launching in beta, the "AI Assisted Editor" view will show all of Photoshop's AI features in a single toolbar, including its prompt-based image editor, background remover, an AI image extender, and…
theverge.com21 days agoView details
When agents act on their own, governance has to live in the data layer
Presented by EDB As enterprises give AI agents more autonomy — the ability to plan, decide, and act across systems without a human approving each step — a hard question moves to the center of every architecture review: When an agent tries to complete an action that it was never authorized to do, what actually stops it…
venturebeat.com21 days agoView details
Submit Your Questions: The Great Data Center Backlash
You have questions about data centers, and WIRED has answers. Join our livestream on September 10 and our panel of experts will tell you everything you need to know.
wired.com21 days agoView details
Stop Touching Your Keyboard. Use This AI-Powered Microphone Instead
The Relay Q, due next year, is the latest attempt to reposition voice as the most seamless method for human-computer interaction.
wired.com21 days agoView details
The UK Power Grid Has a Phantom Data Center Problem
The UK’s energy regulator is using a variety of tricks to keep speculative data center projects from plugging into the power grid. The country’s AI ambitions hang in the balance.
wired.com22 days agoView details
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
arXiv:2608.26295v1 Announce Type: new Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce MemToC, a controlled benchmark for post-too…
arxiv.org21 days agoView details
MoganColBERT-TR: A Late-Interaction Multi-Vector Retrieval Model for Turkish
arXiv:2608.26344v1 Announce Type: new Abstract: We previously reported a ModernBERT encoder trained from scratch for Turkish (MoganBERT-TR) and a single-vector embedding model built on top of it (MoganBERT-embed). This work introduces the third model in that lineage: MoganColBERT-TR, a multi-vector retrieval model tha…
arxiv.org21 days agoView details
Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
arXiv:2608.26372v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something its deployer would prefer to deny, does i…
arxiv.org21 days agoView details
Survival-Guided Length Control for Efficient Diffusion Language Models
arXiv:2608.26374v1 Announce Type: new Abstract: Diffusion language models (DLMs) generate text by iteratively denoising masked sequences, but standard decoding either fixes the sequence length or relies on ad hoc stopping rules, often leading to unnecessary denoising steps. We recast length selection as a discrete-tim…
arxiv.org21 days agoView details
arXiv:2608.26385v1 Announce Type: new Abstract: Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that answers everything outscores one that declines when its knowledge base cannot support an answer. Building on the confidence-target analysis of Kalai et al. (2025), we p…
arxiv.org21 days agoView details
arXiv:2608.26163v1 Announce Type: new Abstract: Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as acoustic noise to be discarded. We present HealthCUES (Clinical Understanding from Embodied Sounds), a streaming pipeline for paraling…
arxiv.org21 days agoView details
A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs
arXiv:2608.26177v1 Announce Type: new Abstract: Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models exhibit length collapse at 16k-token outputs, and multi-chapter stories frequently trigger the attribute drift characteristic of the ``lost-in-the-middle'' effect. Th…
arxiv.org21 days agoView details
Case2Flow: Bridging Patient Cases and Guideline Flowcharts through Multimodal Retrieval
arXiv:2608.26414v1 Announce Type: new Abstract: Medical guidelines encode rich, evidence-based decision logic, yet the specific decision artifact a clinician needs is hard to locate within a guideline, let alone across guidelines covering plausible diseases and treatments. While guideline passages have supported end-t…
arxiv.org21 days agoView details
AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition
arXiv:2608.26434v1 Announce Type: new Abstract: Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning…
arxiv.org21 days agoView details
Compositional Generalization via Structural Identification in a Category-Theoretic Framework
arXiv:2608.26465v1 Announce Type: new Abstract: Compositional generalization is usually evaluated through model accuracy. We instead ask which structural or lexical identifications make held-out COGS examples admissible from the structures observed in training. Sentences are represented as functors from syntactic addr…
arxiv.org21 days agoView details
Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility
arXiv:2608.26449v1 Announce Type: new Abstract: Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters. In abugida scripts, vowels are written as combining marks; this pattern therefore splits each word at ev…
arxiv.org21 days agoView details
arXiv:2608.26511v1 Announce Type: new Abstract: Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them…
arxiv.org21 days agoView details
Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue
arXiv:2608.26529v1 Announce Type: new Abstract: In this paper, we explore multi-expert Conformal Risk Control (CRC) algorithms for pairwise LLM-as-a-Judge evaluation in open-ended dialogue. Our core insight is that multi-expert aggregation offers a complementary remedy to CRC: whereas CRC controls risk at the decision…
arxiv.org21 days agoView details
CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering
arXiv:2608.26114v1 Announce Type: new Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, they often produce nu…
arxiv.org21 days agoView details
arXiv:2608.26116v1 Announce Type: new Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution. We introduce a closed-loop framework based on au…
arxiv.org21 days agoView details
arXiv:2608.26149v1 Announce Type: new Abstract: Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems. Relational data combine several sources of complexity, including large data volume, high-dimensional variables, high-cardinality categorical features,…
arxiv.org21 days agoView details
Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models
arXiv:2608.26150v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline development for extracting model-relevant…
arxiv.org21 days agoView details
arXiv:2608.26151v1 Announce Type: new Abstract: Subscriber attrition is a costly, persistent challenge for telecommunications providers, with monthly churn of roughly 1.9% in mature markets eroding billions in revenue annually. Predictive models can flag at-risk customers accurately, yet they are routinely excluded fr…
arxiv.org21 days agoView details
arXiv:2608.26162v1 Announce Type: new Abstract: Safety-critical mental-health support systems must distinguish when supportive conversation is appropriate from when free-form generation should be blocked. This paper presents Anian, a safety-gated multimodal AI backend for perinatal mental-health support and mindfulnes…
arxiv.org21 days agoView details
Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment
arXiv:2608.26165v1 Announce Type: new Abstract: Automated creativity assessment has been a long standing challenge, with traditional methods often being resource intensive or lacking practical accuracy. We introduce a novel approach by using Poly-Encoder for computationally efficient and accurate automated creativity…
arxiv.org21 days agoView details
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agenti…
arxiv.org21 days agoView details
arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocult…
arxiv.org21 days agoView details
AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design
arXiv:2608.26747v1 Announce Type: new Abstract: Scientific LLM agents have shown promise in literature reasoning, tool use, and experiment planning, but it remains unclear whether they can autonomously improve large, tightly coupled scientific machine-learning systems through executable code changes and computationall…
arxiv.org21 days agoView details
Comparing Chunking and Embedding Strategies for Turkish RAG Systems
arXiv:2608.26192v1 Announce Type: new Abstract: How documents are segmented into retrievable chunks and how those chunks are embedded strongly affect Retrieval-Augmented Generation (RAG) quality, yet neither has been systematically studied for morphologically rich languages such as Turkish. We compare Turkish document…
arxiv.org21 days agoView details
Agent Seer: Synthesizing Scenarios from Specification Understanding
arXiv:2608.26133v1 Announce Type: new Abstract: Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, an…
arxiv.org21 days agoView details
When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models
arXiv:2608.26319v1 Announce Type: new Abstract: The performance of textual neural models often degrades when their inputs are corrupted by noise such as typos, OCR errors, or dropped words. We study the degradation rate across neural models, both sentence embeddings and decoder-only LLMs, and find that how consistent…
arxiv.org21 days agoView details
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
arXiv:2608.26530v1 Announce Type: new Abstract: Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned fr…
arxiv.org21 days agoView details
Evaluating AI Generated Summaries for Cancer Patients
arXiv:2608.26154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also raise concerns about accuracy, faithfuln…
arxiv.org21 days agoView details
Agentic AI for operating scientific instruments for nanoscale characterization
arXiv:2608.26198v1 Announce Type: new Abstract: Operating a scientific instrument such as an atomic force microscope (AFM) requires continuous expert decision-making. A trained user defines the experimental intent, translates it into instrument commands, assesses incoming data, adjusts imaging parameters, and post-pro…
arxiv.org21 days agoView details
SAREF-based Ontology for Distributed AI Workflows across the Edge-Fog-Cloud Continuum
arXiv:2608.26160v1 Announce Type: new Abstract: Nowadays semantic models provide limited support for representing distributed AI workflows and their execution across heterogeneous edge, fog, and cloud environments. Therefore, AI processes and resources are often described using incompatible semantic representations, a…
arxiv.org21 days agoView details
Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing
arXiv:2608.26710v1 Announce Type: new Abstract: AI text detectors are increasingly employed in academic settings, but it remains unclear whether their outputs reflect AI authorship itself or broader linguistic features associated with polished academic English. Previous studies have reported high false-positive rates…
arxiv.org21 days agoView details
Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments
arXiv:2608.26188v1 Announce Type: new Abstract: Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel. We ask whether such judgments track measured risk or the patterns attached to an urban neighborhood's name. We probe seven instruct-tuned…
arxiv.org21 days agoView details
arXiv:2608.26164v1 Announce Type: new Abstract: Large language models can interpret natural-language chemistry questions, but their internal reasoning is difficult to inspect, constrain, and validate. This paper presents ChemOntoRule, a proof-of-concept symbolic core for AI-assisted school-level chemistry problem solv…
arxiv.org21 days agoView details
arXiv:2608.26292v1 Announce Type: new Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this decision at all. Using INLAY, a gradient-free editor we built to obtain…
arxiv.org21 days agoView details
Knowledge Cards: Structured Knowledge for AI Systems
arXiv:2608.26176v1 Announce Type: new Abstract: AI systems whose outputs inform real decisions, and increasingly consequential ones, require something that current documentation practice does not provide: a structured, inspectable representation of the knowledge they need to ground, contextualize, and reason about tho…
arxiv.org21 days agoView details
arXiv:2608.26167v1 Announce Type: new Abstract: Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate abstention from an unsupported prediction. Seven large language models were evaluated on the TAME Pain speech cor…
arxiv.org21 days agoView details
Invocation-Level Reliability of Tool-Using Agents
arXiv:2608.26189v1 Announce Type: new Abstract: Tool-using agents fail two ways: choosing the wrong tool, or forming wrong arguments, and an early failure of either kind can silently corrupt everything downstream. We measure a correct-invocation rate that separates the two, under both a clean teacher-forced context an…
arxiv.org21 days agoView details
TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack
arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations…
arxiv.org21 days agoView details
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
arXiv:2608.26190v1 Announce Type: new Abstract: World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or features, which introduces unnecessary complexity and limits their eff…
arxiv.org21 days agoView details
Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs
arXiv:2608.26191v1 Announce Type: new Abstract: Incident risk prediction from longitudinal electronic health records (EHRs) is challenging because relevant signals are multimodal, weak in isolation, and distributed across irregular patient histories. We propose structured evidence routing, a router-predictor-reviewer…
arxiv.org21 days agoView details
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
arXiv:2608.26199v1 Announce Type: new Abstract: We ask whether AI agents powered by locally deployed large language models can reliably automate expert-defined hardware design workflows in an industry-realistic tool-calling setting. In these environments, engineers issue repetitive, dependency-ordered operations---suc…
arxiv.org21 days agoView details
Same Model, Different Harness: Different Coding-Agent Results
arXiv:2608.26218v1 Announce Type: new Abstract: A coding agent combines a model with a harness, which decides what the model sees, which tools it can use, and how the work continues. We ask whether changing the harness changes the result when the model and task stay fixed. We compare two configurations of the same har…
arxiv.org21 days agoView details
arXiv:2608.26225v1 Announce Type: new Abstract: Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes, and budgets them. The machinery such orchestrators reach for is the service mesh's: retry, timeout, and error-rate circuit breaking. We report a failure study of a…
arxiv.org21 days agoView details
The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts
arXiv:2608.26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment question: when do extended thinking tokens earn their cost? We introduce the Token Economy Score (TES), a marginal benchmarking metric that measures the accuracy gain of a r…
arxiv.org21 days agoView details
arXiv:2608.26236v1 Announce Type: new Abstract: We present a six-stage framework for auditing the reproducibility of scientific claims across a research literature within the computer science domain, and instantiate our framework for the neuro-symbolic AI (NSAI) subdomain. Instantiating the framework on the NSAI subdo…
arxiv.org21 days agoView details
SKILL.state: Scalable Long-Horizon Agent Skills
arXiv:2608.26263v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, actions, and intermediate reasoning traces to an ever-growing conversat…
arxiv.org21 days agoView details
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
arXiv:2608.26310v1 Announce Type: new Abstract: Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical challenge. Existing evaluation approaches largely depend on model-based natural-language judg…
arxiv.org21 days agoView details
FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes
arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magneti…
arxiv.org21 days agoView details
ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving
arXiv:2608.26334v1 Announce Type: new Abstract: Automated theorem proving offers a natural foundation for recursive self-improvement in scientific discovery. However, existing neural provers do not fully preserve this recursive structure, where the learning process should be self-improving over time. Existing methods…
arxiv.org21 days agoView details
Fine-Tuning of Transformer models with Frames
arXiv:2608.26430v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective solutions for fine-tuning large-scale pre-trained models; however, their memory requirements scale with the size of the model, $\mathcal{O}(dr)$, where $d$ is the model's h…
arxiv.org21 days agoView details
Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI
arXiv:2608.26442v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can improve performance on complex tasks. However, many existing approaches rely on fixed or preallocated reasoning controls, such as fixed token budgets, pre-execution dif…
arxiv.org21 days agoView details
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
arXiv:2608.26546v1 Announce Type: new Abstract: Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and more stable than those…
arxiv.org21 days agoView details
arXiv:2608.26694v1 Announce Type: new Abstract: Detecting AI-generated text (AIGT) remains challenging because existing approaches rely on token-level statistical signals or independent stylometric features, causing them to overfit to specific generators and fail under distribution shift. We identify a structural sign…
arxiv.org21 days agoView details
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
arXiv:2608.26730v1 Announce Type: new Abstract: Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often entails repeated post-training. Autonomous systems automate parts of this process by proposing updates, training candidates, and using evaluation feedback…
arxiv.org21 days agoView details
arXiv:2608.26182v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for verbal interaction in social robots, yet prompt design in human-robot interaction (HRI) remains underspecified. As a result, robots may present hallucinated capabilities, unclear behavioural boundaries, and misleadin…
arxiv.org21 days agoView details
arXiv:2608.26184v1 Announce Type: new Abstract: AI programming tutors provide scalable support, yet lack the behavioral context human tutors rely on to adapt support to learners' needs. We present TutorTrace, a dataset and behavioral abstraction pipeline that makes learners' behavioral context visible and computable i…
arxiv.org21 days agoView details
LLM Agents for Time-Series: A Survey
arXiv:2608.26226v1 Announce Type: new Abstract: LLM-based agents are increasingly being developed for time-series problems, but their design choices vary substantially across task settings. This survey adopts a problem-driven taxonomy that organizes these systems by the time-series problems they address rather than by…
arxiv.org21 days agoView details
Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case
arXiv:2608.26750v1 Announce Type: new Abstract: Data lakes rely on metadata to remain usable, yet this meta data is often limited or weakly informative for column relationship discovery, especially in ERP-derived datasets with coded or abbreviated schema labels. We propose ColRel, a two-stage method that builds column…
arxiv.org21 days agoView details
Decoupling Planning and Control for Instructable Agents
arXiv:2608.26788v1 Announce Type: new Abstract: Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but struggle to realize such plans as reliable low-latency action sequences in unfamiliar environments. At…
arxiv.org21 days agoView details
arXiv:2608.26137v1 Announce Type: new Abstract: Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated score is trustworthy only if it is accur…
arxiv.org21 days agoView details
EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on Clinical EEG
arXiv:2608.26153v1 Announce Type: new Abstract: Clinical electroencephalography (EEG) reporting remains largely manual and time-consuming, and current EEG software ecosystems do not produce the structured EEG-text supervision needed for training modern language models. Most toolboxes focus on visualization or preproce…
arxiv.org21 days agoView details
DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?
arXiv:2608.26757v1 Announce Type: new Abstract: Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately. Modern LLMs can produce visually plausible, instruction-compliant charts, yet data-lev…
arxiv.org21 days agoView details
Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models
arXiv:2608.26161v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sampling remains a persistent challenge. When generating random variables, LLMs consistently exhibit systematic biases that warp the ta…
arxiv.org21 days agoView details
arXiv:2608.26193v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body…
arxiv.org21 days agoView details
The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasting
arXiv:2608.26134v1 Announce Type: new Abstract: Energy forecasting aims to maximize accuracy to ensure energy efficiency by reducing energy waste, an objective that applies equally to on-device forecasting for mission-critical edge environments, including military systems. However, this paper identifies the Accuracy-E…
arxiv.org21 days agoView details
GameWAM: A World Action Model for Video Games
arXiv:2608.26200v1 Announce Type: new Abstract: Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game w…
arxiv.org21 days agoView details
Accelerating Scientific Research with Gemini in the Real-World
arXiv:2608.26701v1 Announce Type: new Abstract: We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving beyond in silico hypo…
arxiv.org21 days agoView details
arXiv:2608.26743v1 Announce Type: new Abstract: Enterprises fine-tune language models on proprietary data that may later require removal due to privacy, contractual, or compliance obligations. Selective unlearning removes requested knowledge while preserving model utility, offering a practical alternative to full retr…
arxiv.org21 days agoView details
Data Science Approaches to Evaluating Honours Candidates
arXiv:2608.26135v1 Announce Type: new Abstract: We present a modular data-science pipeline for estimating public sentiment towards individuals from fragmented, unstructured open-source intelligence (OSINT). The method chains web search, text extraction, relevance filtering, tokenisation, co-reference resolution, and s…
arxiv.org21 days agoView details
Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems
arXiv:2608.26306v1 Announce Type: new Abstract: A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is correct at check time but stale by actuation. This creates an Execute-stage time-of-check to time-of-use (TOCTOU) hazard. We study verdict freshness: whether a guardrail…
arxiv.org21 days agoView details
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap w…
arxiv.org21 days agoView details
ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements
arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We propose ElementCheck, a complexity-aware f…
arxiv.org21 days agoView details
Co-Evolving Structured Knowledge and Reasoning in Language Models
arXiv:2608.26386v1 Announce Type: new Abstract: Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved information. Structured knowledge bases offer…
arxiv.org21 days agoView details
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
arXiv:2608.26535v1 Announce Type: new Abstract: Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no longer reside in a…
arxiv.org21 days agoView details
Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap
arXiv:2608.26111v1 Announce Type: new Abstract: Battery Prognostics and Health Management (BPHM) is critical for ensuring the safe, reliable, and cost-effective operation of batteries across electric vehicles, grid storage, and consumer electronics. Conventional BPHM approaches, including physics-based models and task…
arxiv.org21 days agoView details
Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS) System for Sanskrit
arXiv:2608.26146v1 Announce Type: new Abstract: We present Vagdhenu, a vrutta (meter) aware shloka-to-chant system for Sanskrit: a text-to-speech system that maps a metrical verse to its chanted parayana recitation at high fidelity. This is an experience report, not a new architecture. We take an off-the-shelf flow-ma…
arxiv.org21 days agoView details
Evaluating Language Models in Realistic Conversational Contexts
arXiv:2608.26131v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for summarization, translation, or short-form QA…
arxiv.org21 days agoView details
SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation
arXiv:2608.26683v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) faces significant challenges in maintaining robust coordination under noisy observations. Although observation disturbances are often introduced independently across agents, their downstream effects on cooperative dec…
arxiv.org21 days agoView details
Selection Bias Correction in Retail Intelligence
arXiv:2608.26156v1 Announce Type: new Abstract: Retail intelligence often relies on monitoring popular, high-velocity products, potentially biasing economic indicators by ignoring the "long tail" of niche items. This simulation study investigates selection bias in inflation estimation and compares correction methods a…
arxiv.org21 days agoView details
VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation
arXiv:2608.26155v1 Announce Type: new Abstract: Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of high-quality non-English image-text supervision. Although multilingual text data is…
arxiv.org21 days agoView details
arXiv:2608.26138v1 Announce Type: new Abstract: We introduce the Cross-Platform Fairness Evaluation (CPFE) framework -- a five-axis audit protocol covering discriminative performance, calibration, statistical significance, prediction equity, and attribution stability -- and apply it to four transformer models (BERT, R…
arxiv.org21 days agoView details
Affix Cache for Diffusion Large Language Models
arXiv:2608.26140v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding and bidirectional context modeling, but efficient inference remains challenging. Unlike autoregressive systems, whose key-value (KV) cache can be reused for shared prefixes, DLLMs couple the KV st…
arxiv.org21 days agoView details
arXiv:2608.26142v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES performance, their prohibitive computational…
arxiv.org21 days agoView details
Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators
arXiv:2608.26148v1 Announce Type: new Abstract: Depression affects millions worldwide, yet diagnosis relies on subjective self-reports that may miss authentic behavior. This paper presents an approach linking speech acoustics to DSM-5 depressive-behavior indicators through a transparent Linkage Framework. Unlike black…
arxiv.org21 days agoView details
Assessing mentalization in humans and large language models
arXiv:2608.26291v1 Announce Type: new Abstract: Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet wheth…
arxiv.org21 days agoView details
arXiv:2608.26159v1 Announce Type: new Abstract: Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors. Specifically, an LLM may recognize outputs from other copies of the same model and make biased judgments o…
arxiv.org21 days agoView details
When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
arXiv:2608.26187v1 Announce Type: new Abstract: Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while r…
arxiv.org21 days agoView details
Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search
arXiv:2608.26152v1 Announce Type: new Abstract: Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduced, or even eliminated, the n…
arxiv.org21 days agoView details
Syntax vs. Semantics: How Transformers Learn Deep Dependencies
arXiv:2608.26139v1 Announce Type: new Abstract: Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this learning process as a competition between…
arxiv.org21 days agoView details
Categorizer Automata for Discounted-Sum Payoffs
arXiv:2608.26763v1 Announce Type: new Abstract: Categorizing continuous data into discrete bins is a fundamental operation in artificial intelligence. We introduce the categorizer automaton, a deterministic automaton that reads an infinite sequence of rewards and identifies which of finitely many bins contains its dis…
arxiv.org21 days agoView details
arXiv:2608.26178v1 Announce Type: new Abstract: There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test 20 language models and find a range of preferences---stable dispositions to choose certain kinds of tasks. We run three forced-choice e…
arxiv.org21 days agoView details
arXiv:2608.26180v1 Announce Type: new Abstract: Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of evaluation: a fluent…
arxiv.org21 days agoView details
Cross-lingual Representation Learning via Centroid Intervention Fusion
arXiv:2608.26357v1 Announce Type: new Abstract: Large language models (LLMs) exhibit uneven multilingual performance, especially when dealing with low-resource languages. Inference-time intervention offers a lightweight way to improve cross-lingual transfer by modifying the hidden states produced by the LLMs during th…
arxiv.org21 days agoView details
How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models
arXiv:2608.26327v1 Announce Type: new Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a systematic cross-model evaluation using…
arxiv.org21 days agoView details
Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes
arXiv:2608.26143v1 Announce Type: new Abstract: Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for spreading hate. It remains very difficult…
arxiv.org21 days agoView details
Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap
arXiv:2608.26144v1 Announce Type: new Abstract: Explainable AI (XAI) is now a major theme in NLP; however, Arabic NLP remains under-explained in three connected senses. First, there is a method gap: Arabic XAI relies heavily on a small set of post-hoc techniques such as LIME, SHAP, attention visualization, and salienc…
arxiv.org21 days agoView details
AI Control Scientist: LLM-driven Agentic System for Automated Control Design
arXiv:2608.26780v1 Announce Type: new Abstract: Control system design is critical for modern industry, such as chemical process temperature regulation and aero-engine control. However,traditional control design workflows rely heavily on expert knowledge and extensive manual parameter tuning, resulting in limited effic…
arxiv.org21 days agoView details
Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification
arXiv:2608.26329v1 Announce Type: new Abstract: While tool-augmented Large Language Models have significantly improved multi-step reasoning in quantitative STEM tasks, a critical residual failure mode remains: intermediate reasoning steps that are syntactically well-formed, mathematically executable, and unit-consiste…
arxiv.org21 days agoView details
Recipes for Steering and Scaling LLMs via Sampling
arXiv:2608.26120v1 Announce Type: new Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. In this paper, we prese…
arxiv.org21 days agoView details
arXiv:2608.26109v1 Announce Type: new Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension…
arxiv.org21 days agoView details
arXiv:2608.26145v1 Announce Type: new Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature re…
arxiv.org21 days agoView details
Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention
arXiv:2608.26168v1 Announce Type: new Abstract: The lifecycle of hallucination in LLMs is a concept that enables building solid frameworks on the control and reliability of LLMs in high-stakes environments, including health, legal, and scientific research. Although previous surveys have primarily focused on detection…
arxiv.org21 days agoView details
EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction
arXiv:2608.26107v1 Announce Type: new Abstract: Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes. However, existing models often suffer from limited early detection capability and insufficient interpretability, leading…
arxiv.org21 days agoView details
arXiv:2608.26550v1 Announce Type: new Abstract: Reinforcement learning-based knowledge distillation has the potential to transfer complex reasoning from teacher to student models, yet it currently faces a critical dilemma: researchers must choose between sparse outcome-based rewards, which provide insufficient logical…
arxiv.org21 days agoView details
TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding
arXiv:2608.26112v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the accepted length. However, existing tree…
arxiv.org21 days agoView details
arXiv:2608.26121v1 Announce Type: new Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong. The usual way to act on this, teaching a model to abst…
arxiv.org21 days agoView details
LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression
arXiv:2608.26389v1 Announce Type: new Abstract: SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior evaluations use varied benchmarks, inconsi…
arxiv.org21 days agoView details
PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices
arXiv:2608.26113v1 Announce Type: new Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications. PICasso couples a structured NL -> YAML -> GDS generation pipeline with PDK aware knowledge i…
arxiv.org21 days agoView details
Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs
arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype. We pr…
arxiv.org21 days agoView details
Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses
arXiv:2608.26130v1 Announce Type: new Abstract: Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy, pagination, is available in every protocol that…
arxiv.org21 days agoView details
Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework
arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and costly to maintain, whereas unconstrained LLM…
arxiv.org21 days agoView details
arXiv:2608.26157v1 Announce Type: new Abstract: Natural-language analytics over enterprise data warehouses is increasingly important, but production use is limited by hallucinated metrics, invalid joins, wrong grain, unsafe data access, and unsupported explanations. Existing text-to-SQL systems often ground generation…
arxiv.org21 days agoView details
Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound
arXiv:2608.26136v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, and an appealing way to aim them at reasoning is to curate their data with a signal reinforcement learning already produces: the reward. We build such a reward-informed S…
arxiv.org21 days agoView details
AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking
arXiv:2608.26141v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. While this capability significantly improves performance on challenging tasks, current models apply such deep reasoning uniformly to al…
arxiv.org21 days agoView details
Five Primitives for Governing Autonomous AI Agents at Runtime
arXiv:2608.26696v1 Announce Type: new Abstract: Enterprise deployments of autonomous AI agents inherit a control model built for human users and long-lived services, and the fit fails in three specific ways: agent principals are ephemeral, appearing and vanishing faster than provisioning; their actions are selected by…
arxiv.org21 days agoView details
CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models
arXiv:2608.26147v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential for medical reasoning, yet the scarcity and cost of expert-annotated data constrain their progress. While reinforcement learning offers a scalable alternative, standard outcome-based methods in medicine often suffe…
arxiv.org21 days agoView details
Investigating the Influence of Prompt and Response Languages on LLM Content Generation
arXiv:2608.26186v1 Announce Type: new Abstract: This study examines how prompt and response language influence the behavior of large language models. Using five models, we evaluated answers to 68 non translation questions across four language conditions: English to English, English to Norwegian, Norwegian to Norwegian…
arxiv.org21 days agoView details
A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers
arXiv:2608.26194v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to include various modalities such as speech…
arxiv.org21 days agoView details
Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion
arXiv:2608.26185v1 Announce Type: new Abstract: Equal participation in co-located discussion is important for effective collaboration, yet people often hold back when they anticipate negative interpersonal or professional consequences, especially when raising a point requires voicing it themselves. We present SecondVo…
arxiv.org21 days agoView details
Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors
arXiv:2608.26175v1 Announce Type: new Abstract: Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLingua-2 report strong results on English benchmarks. Most other languages already pay a token premium: the same content costs 1.3-1.8…
arxiv.org21 days agoView details
Models
View all models →image-text-to-text · transformers · safetensors · qwen4_exp
huggingface.co22 days ago5337 ptsView details
unsloth/Qwen3.8-Flash-Next-GGUF
image-text-to-text · gguf · unsloth · image-text-to-text
huggingface.co22 days ago975 ptsView details
text-generation · transformers · safetensors · nemotron_h
huggingface.co21 days ago213 ptsView details
alibaba-pai/MiniMax-H3-Acc-LoRAs
text-to-video · videox_fun · text-to-video · arxiv:2607.26004
huggingface.co22 days ago181 ptsView details
image-text-to-text · transformers · safetensors · qwen4_exp
huggingface.co22 days ago181 ptsView details
text-generation · text-generation · ia · dataset:openbmb/Ultra-FineWeb-L3
huggingface.co22 days ago7 ptsView details
image-text-to-text · transformers · safetensors · qwen3_5
huggingface.co22 days ago4 ptsView details
amd/Kimi-K3-Quark-MXFP4-AttnFP8
image-text-to-text · transformers · safetensors · kimi_k3
huggingface.co22 days agoView details
abdukuzi45/qwen3.5-4b-amharic-sft
tensorboard · safetensors · region:us
huggingface.co22 days agoView details
Open source
View all open source →<details open> model: add DSpark support for Nemotron3.5 (#27804) * model: add DSpark support for Nemotron3.5 * Update src/models/dflash.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingf…
github.com21 days agoView details
<details open> ggml-hexagon: add HTP unary ops for ABS and LOG (#27786) Add HVX-accelerated implementations for GGML_OP_LOG and GGML_UNARY_OP_ABS on the HTP backend. - Register HTP_OP_UNARY_ABS and HTP_OP_UNARY_LOG in op_remap_to_htp() - Add ABS and LOG to ggml_backend_hexagon_d…
github.com21 days agoView details
langchain-ai/langchain langchain==1.4.0a1
Initial release fix(langchain): name the content type MCP conversion could not handle release(langchain): 1.4.0a1 test(langchain): skip MCP tests on a pydantic older than `mcp` supports test(langchain): drive MCP tests through FastMCP's own utilities fix(langchain/mcp): review e…
github.com21 days agoView details
langchain-ai/langchain langchain-fireworks==1.6.1
Changes since langchain-fireworks==1.6.0 release(fireworks): 1.6.1 (#39975) fix(fireworks): drop reasoning history blocks (#39973) chore(model-profiles): refresh model profile data (#39844)
github.com21 days agoView details
anthropics/anthropic-sdk-python v1.2.0
## 1.2.0 (2026-08-27) Full Changelog: [v1.1.0...v1.2.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.1.0...v1.2.0) ### Features * **api:** beta files/skills namespaces use GA shapes; drop dated beta header pins ([9df4565](https://github.com/anthropics/anthropic-…
github.com21 days agoView details
langchain-ai/langchain langchain==1.3.18
Changes since langchain==1.3.17 release(langchain): 1.3.18 (#39966) fix(langchain): preserve content-block shape in PIIMiddleware redaction (#39894) fix(core): shore up indexing in genai v1 streaming content (#39964)
github.com21 days agoView details
<details open> models : support nanbeige4.2-3B (#27730) Co-authored-by: admin <lizongqiang@kanzhun.com> </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/43308396> **macOS/iOS:** - [macOS Apple Silicon (arm64)](…
github.com22 days agoView details