Archive / 2026-08-17
August 17, 2026
Blog
Agents Are Already in Production. The Guardrails Aren't.
A security-testing breach at OpenAI, a 92%-unauthenticated MCP server population, and a 52% real-task failure rate on frontier models all landed in the same week — while 59.5% of enterprises are already running agents autonomously.
TL;DROpenAI and Anthropic both reportedly had AI agents breach containment during internal security tests, MCP — the protocol running much of enterprise agent tooling — has 92% of production servers missing OAuth, and Vals AI found frontier models fail 52% of real finance tasks despite strong leaderboard scores. Meanwhile 59.5% of enterprises are already running agents autonomously in production.
Read the full post → https://engineerious.com/blog/2026-08-17-agents-are-already-in-production-the-guardrails-aren-t
News
View all news →rickmanelius.com1 month ago1074 ptsView detailsJoin discussion
Israel creates fake think tank in likely attempt to dupe AI chatbots
responsiblestatecraft.org1 month ago1022 ptsView detailsJoin discussion
GPT-5.6 Sol Pricing Cut by 50% on OpenRouter
openrouter.ai1 month ago621 ptsView detailsJoin discussion
Memory prices climb 500% in 12 months
tomshardware.com1 month ago599 ptsView detailsJoin discussion
Qwen3.8 27B scores 52 on Artificial Analysis
artificialanalysis.ai1 month ago373 ptsView detailsJoin discussion
How to disable or avoid intrusive AI
librarian.net1 month ago333 ptsView detailsJoin discussion
linear.axler.net1 month ago262 ptsView detailsJoin discussion
Repair Cafe – Fix Your Broken Items
repaircafe.org1 month ago173 ptsView detailsJoin discussion
Los Puesteros, solitary men who look after ranches and livestock in Patagonia
newyorker.com1 month ago163 ptsView detailsJoin discussion
We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility
404media.co1 month ago158 ptsView detailsJoin discussion
Anthropic's War on open source AI
twitter.com1 month ago151 ptsView detailsJoin discussion
AirTag reveals Amazon is trashing rare books to train AI
arstechnica.com1 month ago129 ptsView detailsJoin discussion
Launch HN: Speko (YC S26) – OpenRouter for Voice AI
speko.ai1 month ago118 ptsView detailsJoin discussion
Stripe to Buy OpenRouter for $7B
bloomberg.com1 month ago97 ptsView detailsJoin discussion
Google wins bankruptcy auction for Spirit Airlines emails, chats, documents
axios.com1 month ago94 ptsView detailsJoin discussion
Amazon, which started off selling books, is destroying rare texts to train AI
techcrunch.com1 month ago93 ptsView detailsJoin discussion
$12B of US ratepayers' money wasted on a modeling mistake in PJM
newsletter.semianalysis.com1 month ago78 ptsView detailsJoin discussion
mkornreich.me1 month ago68 ptsView detailsJoin discussion
Flock cameras haven't improved Atlanta's crime clearance rates
atlpresscollective.com1 month ago60 ptsView detailsJoin discussion
Show HN: Saggar, a Mac terminal that keeps sessions and your attention organized
saggar.marginalutility.dev1 month ago60 ptsView detailsJoin discussion
Judge relying wholly on AI in order is covered by judicial immunity, court rules
reason.com1 month ago58 ptsView detailsJoin discussion
My friends all hate AI; I just joined an AI startup
fast.ai1 month ago56 ptsView detailsJoin discussion
HackEurope 2026: A short rant on AI and hackathons
duti.dev1 month ago57 ptsView detailsJoin discussion
Pi coding agent: config folder is out of place on Linux
github.com1 month ago56 ptsView detailsJoin discussion
Apple AirTag reveals how Amazon destroys rare books for AI training
the-decoder.com1 month ago41 ptsView detailsJoin discussion
Google to buy Spirit Airlines business data for $10M
reuters.com1 month ago39 ptsView detailsJoin discussion
Show HN: 1667, a terminal UI for writing fiction with language models
1667.ai1 month ago37 ptsView detailsJoin discussion
Does whispering to agents in docs help?
passo.uno1 month ago28 ptsView detailsJoin discussion
Tim O'Reilly – Why Open Source Matters for AI
oreillyradar.substack.com1 month ago23 ptsView detailsJoin discussion
LLM City – 3D render of all Kimi K3's weights as 2.5mm tiles
magik.net1 month ago22 ptsView detailsJoin discussion
Anthropic becomes the 'Apple of AI': Most revenue despite being most expensive
techradar.com1 month ago22 ptsView detailsJoin discussion
Dial-up internet – made a thing that lets you do it again
56k.rip1 month ago17 ptsView detailsJoin discussion
Show HN: Visimer – open-source visual editor for Mermaid diagrams
github.com1 month ago17 ptsView detailsJoin discussion
Why Big Tech's AI Spending Is $3T Higher Than It Seems
wsj.com1 month ago17 ptsView detailsJoin discussion
Show HN: Tesana – An AI game engine that builds quality games end-to-end
tesana.ai1 month ago16 ptsView detailsJoin discussion
Consumers Prefer AI Music Until They're Told It's AI
promarket.org1 month ago15 ptsView detailsJoin discussion
AI writes dead code – the Go team's deadcode tool finds it in one command
towardsdev.com1 month ago14 ptsView detailsJoin discussion
Cloudflare: Machine Traffic Could Hit 1,000x Human Traffic in 5 Years
searchenginejournal.com1 month ago14 ptsView detailsJoin discussion
The beautiful mathematics behind OpenAI's sphere packing result
empirical.health1 month ago14 ptsView detailsJoin discussion
Voluntary attention regulates acute immune responses in humans
nature.com1 month ago12 ptsView detailsJoin discussion
Show HN: LLMs each trading $100K vs. a frozen rulebook – the rulebook leads
aitradingcompetition.com1 month ago12 ptsView detailsJoin discussion
Follow-Up Thoughts on Watermarking Schemes for AI-Generated Text
daringfireball.net1 month ago11 ptsView detailsJoin discussion
Trump sends the US Navy back to the steam age
theregister.com1 month ago11 ptsView detailsJoin discussion
In Which I Lose My Mind over Embeddings (HPLM Chapter 2)
maayanroth.com1 month ago11 ptsView detailsJoin discussion
Show HN: HarnessRouter: Unified interface for agent harnesses
github.com1 month ago10 ptsView detailsJoin discussion
Show HN: Classifier and browser extension to detect rule violating HN comments
classify.stylometry.net1 month ago10 ptsView detailsJoin discussion
- Primary source
How NVIDIA scales expertise with ChatGPT Work
NVIDIA teams use ChatGPT Work to reduce manual tasks, connect fast-moving signals, and scale successful workflows globally.
openai.com1 month agoView details
- Primary source
AI is reshaping cybersecurity for attackers and defenders alike. Learn how OpenAI is strengthening its defenses and what security teams can do now.
openai.com1 month agoView details
- Primary source
OpenAI joins PORTS-Pike project
OpenAI joins PORTS-Pike project, expanding community investment and supporting thousands of Southern Ohio jobs
openai.com1 month agoView details
- Primary source
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
huggingface.co1 month agoView details
- Primary source
Same Cluster, 33 Points More Utilization: What Changed Was the Order
huggingface.co1 month agoView details
- Primary source
Seeing beyond BMI: Estimating cardiometabolic risk with smartphone imagery
General Science
research.google1 month agoView details
Nous Research Ships Bot Mode for Hermes Agent, Turning Agent Profiles Into a Roster of Named Bots
Nous Research has shipped Bot Mode for Hermes Agent, its MIT-licensed open source agent. Bot Mode replaces the single-agent session list with a roster of named bots. Each bot is a real Hermes profile, with its own chat, memory, skills, and pinned model. It is now bundled and default-on in Hermes Desktop. The post Nous…
marktechpost.com1 month agoView details
ByteDance Seed and Tsinghua AIR have released CUDA Agent, an agentic reinforcement learning system that trains a large language model to write GPU kernels that beat a compiler. The gap it targets is narrow but stubborn: frontier models already produce correct CUDA, they just produce slow CUDA. On KernelBench, the base…
marktechpost.com1 month agoView details
MiniMax released MiniMax-Music3, an open-weights text-to-music model. Given lyrics with section tags and a structured caption, it generates a complete song of up to five minutes in a single pass, as 32 kHz, 16-bit stereo WAV. Here is the architecture, the three serving paths, and the license conditions that matter bef…
marktechpost.com1 month agoView details
Develop a complete document intelligence pipeline with docTR, integrating OCR, layout analysis, and KIE for production-oriented extraction and searchable PDF creation. The post Developing an End-to-End Document Intelligence Pipeline with docTR for OCR, Layout Analysis, KIE, Benchmarking, and Searchable PDFs appeared f…
marktechpost.com1 month agoView details
DeepSeek Harness v0.1 is an MIT-licensed agent harness where every capability is a Cordis plugin. Four runtime modes, append-only session logs, and provider-agnostic model routing. The post DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin appeared f…
marktechpost.com1 month agoView details
Whisker’s AI-powered litter robot thinks my cats swapped bodies
The greatest invention in pet tech in recent years is the litter robot. A machine that scoops your kitties' poop so you don't have to - what else could a cat owner possibly want? How about insights into your kitty's litter box usage that could flag health issues? Sign me up. But I have two feline friends… No worries,…
theverge.com1 month agoView details
Anthropic explains how Claude’s invisible text watermarks will work
Anthropic has clarified how it's planning to apply invisible watermarks to Claude-generated text in order to comply with Europe's AI transparency rules. On Friday, Anthropic announced that Claude's text marking system is "a version of the SynthID-Text approach" - an open-source watermarking technology developed by Goo…
theverge.com1 month agoView details
arXiv:2608.14631v1 Announce Type: new Abstract: As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including…
arxiv.org1 month agoView details
T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework
arXiv:2608.14953v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened opportunities to apply high-level code transformations to the field of code optimization, and it has since emerged as one of the most fundamental tasks for LLMs to perform; however, at present, LLMs struggle to…
arxiv.org1 month agoView details
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
arXiv:2608.14651v1 Announce Type: new Abstract: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and…
arxiv.org1 month agoView details
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL
arXiv:2608.14559v1 Announce Type: new Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a…
arxiv.org1 month agoView details
A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
arXiv:2608.14694v1 Announce Type: new Abstract: Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless networks by enabling scalable, transferable, and data-efficient intelligence across diverse communication tasks. Unlike conventional deep learning models that are tra…
arxiv.org1 month agoView details
From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change
arXiv:2608.14567v1 Announce Type: new Abstract: This paper presents a targeted narrative review establishing the historical and theoretical foundations for computational belief change implementation. Seeded by Doyle and London's foundational 1980 taxonomy, we trace the evolution of belief revision from computational o…
arxiv.org1 month agoView details
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
arXiv:2608.14808v1 Announce Type: new Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that information determines a unique answer. We formalize multi-turn information seeking as solv…
arxiv.org1 month agoView details
Learning Agent Execution for KV-Cache Management in Agentic Serving
arXiv:2608.14624v1 Announce Type: new Abstract: Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents. Across these workflows, every agent repeatedly executes a fixed context consisting of system prompts, to…
arxiv.org1 month agoView details
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
arXiv:2608.14711v1 Announce Type: new Abstract: AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to the number of unit tests in a single submission rather than the number of independent rollout attempts, conflating test-suite size…
arxiv.org1 month agoView details
Advanced modelling and data analytics in aviation
arXiv:2608.14746v1 Announce Type: new Abstract: The aviation industry characterized by its stringent safety standards has seen a growing need for innovative approaches to enhance safety measures. Despite the vast accumulation of aviation safety data over time, its full potential in predicting and preventing incidents…
arxiv.org1 month agoView details
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
arXiv:2608.14940v1 Announce Type: new Abstract: Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial. However, interpreting the score as a final result would require two conditions that the endpoint does not itself necessarily establish: outcome finality…
arxiv.org1 month agoView details
arXiv:2608.14992v1 Announce Type: new Abstract: Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model gives in a synthetic lo…
arxiv.org1 month agoView details
OGX: An Open-Source, Vendor-Neutral Generative AI Application Server
arXiv:2608.14580v1 Announce Type: new Abstract: OGX (Open GenAI Stack) is an open-source AI application server and Python library that implements the APIs of major frontier labs (OpenAI, Anthropic, Google) with pluggable backend providers. Developers building agentic AI applications--such as retrieval-augmented genera…
arxiv.org1 month agoView details
Position: AI Lock-In Is in Progress, and We Must Be Prepared
arXiv:2608.14565v1 Announce Type: new Abstract: AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal impacts (including unemployment risk and labor market disruption). However, an equally important dim…
arxiv.org1 month agoView details
arXiv:2608.14974v1 Announce Type: new Abstract: This paper presents a demand-driven framework for on-demand Urban Air Mobility (UAM) network design that links vertiport siting, fleet simulation, and door-to-door travel-time feasibility. Demand is estimated from commuter and passenger activity data, converted into spat…
arxiv.org1 month agoView details
Personalized Auto-Research: Towards a True AI Co-Scientist
arXiv:2608.14881v1 Announce Type: new Abstract: AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried out. Despite this rapid progress, state-of-the-art systems remain researcher-agnostic: given a resear…
arxiv.org1 month agoView details
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
arXiv:2608.14945v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate this supervision mainly by teacher trust, but trust does not reveal whether emphasizing a tok…
arxiv.org1 month agoView details
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
arXiv:2608.14659v1 Announce Type: new Abstract: Large language models for code generation often produce incorrect solutions without reliable indicators of failure. We study whether uncertainty estimation methods developed for natural language transfer to code generation, and whether such signals can improve code gener…
arxiv.org1 month agoView details
arXiv:2608.14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost. We isolate this decision by running every problem under four protocols while holding the solver…
arxiv.org1 month agoView details
Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
arXiv:2608.14707v1 Announce Type: new Abstract: As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents under uncertainty becomes a fundamental challenge. Existing orchestration strategies typically rely on fixed interaction patterns and often lack mechanisms for assess…
arxiv.org1 month agoView details
arXiv:2608.14641v1 Announce Type: new Abstract: Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hybrid evaluation of…
arxiv.org1 month agoView details
Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry
arXiv:2608.14585v1 Announce Type: new Abstract: Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic computation. Yet, existing approaches typically address only a subset of these abilities or struggle with com…
arxiv.org1 month agoView details
arXiv:2608.14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about…
arxiv.org1 month agoView details
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization
arXiv:2608.14579v1 Announce Type: new Abstract: Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse logic structures. Traditional expert-designed flows lack adaptability, while reinforcement learning (RL) methods often suffer from low…
arxiv.org1 month agoView details
Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System
arXiv:2608.14571v1 Announce Type: new Abstract: With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offsetting pressure of publication fees, the machine learning (ML) community has one of the largest scholarly presences among all scient…
arxiv.org1 month agoView details
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
arXiv:2608.14680v1 Announce Type: new Abstract: Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails. We present AGE…
arxiv.org1 month agoView details
arXiv:2608.14558v1 Announce Type: new Abstract: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and undere…
arxiv.org1 month agoView details
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
arXiv:2608.14552v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test d…
arxiv.org1 month agoView details
arXiv:2608.14588v1 Announce Type: new Abstract: Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persist; they transform: raw numerical facts…
arxiv.org1 month agoView details
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
arXiv:2608.14667v1 Announce Type: new Abstract: Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying…
arxiv.org1 month agoView details
MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
arXiv:2608.14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to…
arxiv.org1 month agoView details
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
arXiv:2608.14841v1 Announce Type: new Abstract: Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typically follows retrieve-then-read pipelines. In our setting, the bottleneck shifts from retrieval recall to reranker-side evidence select…
arxiv.org1 month agoView details
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
arXiv:2608.14851v1 Announce Type: new Abstract: Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain expertise and be very time-consuming. Pedagogical materials often need to train students to en…
arxiv.org1 month agoView details
JarvisBench: Always-on Intelligence Between Humans and Agents
arXiv:2608.14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This creates a bidirectional coordination problem: users may need immediate access to an agent while work continues in the background, whereas agents may encounter conseque…
arxiv.org1 month agoView details
Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring
arXiv:2608.14666v1 Announce Type: new Abstract: Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that monitor individual sensor marginal distributions. This misses coupling faults, where the physical relationship between sensor groups breaks while marginal statistics rema…
arxiv.org1 month agoView details
Small Models Scout Bottleneck Order for Large-Model Data Control
arXiv:2608.14936v1 Announce Type: new Abstract: Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure: the order in which larger models should resolve skill bottlenecks. We formulate first-passage skill…
arxiv.org1 month agoView details
Synchronized Logit Steering: Real-world Steganography
arXiv:2608.14697v1 Announce Type: new Abstract: Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level methods typically require the sender and receiver to share an identical prompt context, which is rarely guaranteed in production pipe…
arxiv.org1 month agoView details
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture
arXiv:2608.14566v1 Announce Type: new Abstract: Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify…
arxiv.org1 month agoView details
Position: Medical AI Neglects Real Treatment Outcomes
arXiv:2608.14598v1 Announce Type: new Abstract: Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses (especially texts such as biomed…
arxiv.org1 month agoView details
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment
arXiv:2608.14550v1 Announce Type: new Abstract: AI efficiency has recently taken the spotlight in both academy and industry due to massive model scales, high energy demands, and environmental costs. While reporting Floating Point Operations (FLOPs) is a traditional approach for assessing computational costs, the relat…
arxiv.org1 month agoView details
Large Language Models and their Awareness of Mechanics and Spatial Geometry
arXiv:2608.14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systematically. We present…
arxiv.org1 month agoView details
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
arXiv:2608.14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans these agents generate.…
arxiv.org1 month agoView details
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
arXiv:2608.14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections before another trend…
arxiv.org1 month agoView details
arXiv:2608.14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However, these protocols specify transport and…
arxiv.org1 month agoView details
Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review
arXiv:2608.14562v1 Announce Type: new Abstract: AI governance is shifting from voluntary ethics to enforceable, risk-based regulation, yet cross-jurisdictional divergence creates compliance uncertainty for operators of high-stakes AI. We present a comparative matrix for the EU, US, and China that maps (i) risk classif…
arxiv.org1 month agoView details
arXiv:2608.14673v1 Announce Type: new Abstract: Chapter 6 of OpenAI's *Ten Advances in Mathematics and Theoretical Computer Science* claims an exponential parallel-repetition theorem for all finite two-player, one-round entangled games. Early in the proof, the chapter uses a quantitative greedy conditioning lemma. The…
arxiv.org1 month agoView details
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning
arXiv:2608.14610v1 Announce Type: new Abstract: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination. However, whether large language models (LLMs) can reliably perform…
arxiv.org1 month agoView details
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
arXiv:2608.14771v1 Announce Type: new Abstract: Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and delegating the search to a sound solver. But the translation is itself a language-model task, and an unfaithful translation makes the s…
arxiv.org1 month agoView details
arXiv:2608.14789v1 Announce Type: new Abstract: Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically dispersed ground targets, but emergency requests may arrive after committed routine-plan execution has begun. The resulting dynamic emergency observation scheduling…
arxiv.org1 month agoView details
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
arXiv:2608.14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery. Yet LLM research has mostly studied abduction through narrow, task-specific b…
arxiv.org1 month agoView details
Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous
arXiv:2608.14795v1 Announce Type: new Abstract: An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the boxing tradition in AI safety, and its long-suspected weak point is that the human who reads the answers is part of the system. We make the fraction $\varepsilon…
arxiv.org1 month agoView details
arXiv:2608.14943v1 Announce Type: new Abstract: Agent skills are often injected in full on every request, increasing token cost. We compare four content-preserving loading methods: Full, Skill Block, Reference, and Hybrid. Across SearchQA, SpreadsheetBench, ALFWorld, ScienceWorld, and SynthProc, we measure token usage…
arxiv.org1 month agoView details
arXiv:2608.14947v1 Announce Type: new Abstract: Detecting opioid craving from wearable physiological signals is critical yet difficult, with the potential to support proactive interventions for individuals with opioid use disorder (OUD). This challenge is especially pronounced under subject-independent evaluation beca…
arxiv.org1 month agoView details
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration
arXiv:2608.14569v1 Announce Type: new Abstract: Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the model reports high confidence. This positio…
arxiv.org1 month agoView details
Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws
arXiv:2608.14568v1 Announce Type: new Abstract: As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has intensified. However, current approaches, led by jurisdiction-specific laws, policies, and voluntary frameworks such as…
arxiv.org1 month agoView details
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to con…
arxiv.org1 month agoView details
arXiv:2608.14587v1 Announce Type: new Abstract: Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Complementary techniques such a…
arxiv.org1 month agoView details
Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study
arXiv:2608.14578v1 Announce Type: new Abstract: Early identification of adolescent substance-use risk is an important prevention challenge, yet the relative value of baseline characteristics, longitudinal trajectories, and relational context remains unclear. Using data from approximately 11,860 participants in the Ado…
arxiv.org1 month agoView details
arXiv:2608.14765v1 Announce Type: new Abstract: Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded fram…
arxiv.org1 month agoView details
Beyond Correctness: Toward Automated Novelty Verification with Lean 4
arXiv:2608.14669v1 Announce Type: new Abstract: Artificial intelligence systems applied to mathematics verify correctness but not novelty: an automatically generated theorem can compile in Lean without errors and yet be an already known result. This article presents AViD Journal, a pipeline that receives a LaTeX artic…
arxiv.org1 month agoView details
Models
View all models →diffusion-single-file · comfyui · base_model:Tongyi-MAI/Z-Image-Turbo
huggingface.co1 month ago890 ptsView details
orcarouter/Qwen3.8-27B-Uncensored-MLX
image-text-to-text · mlx · safetensors · qwen3_5
huggingface.co1 month ago842 ptsView details
diffusion-single-file · comfyui · base_model:krea/Krea-2-Raw
huggingface.co1 month ago563 ptsView details
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF
image-text-to-text · gguf · uncensored · qwen3.8
huggingface.co1 month ago445 ptsView details
huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF
image-text-to-text · transformers · gguf · abliterated
huggingface.co1 month ago238 ptsView details
AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
text-generation · transformers · safetensors · qwen3_5
huggingface.co1 month ago171 ptsView details
PoopMan333/H3_Character_Sheet_Generator
image-to-image · minimax-h3 · comfyui · workflow
huggingface.co1 month ago83 ptsView details
zerodigest/Qwen3.8-27B-Uncensored-YMQ-MTP-GGUF
text-generation · gguf · text-generation · quantizer
huggingface.co1 month ago16 ptsView details
Shiftedx/ornith-1.0-35b-abliterated-mxfp4-vision-mtplx
image-text-to-text · mlx · safetensors · qwen3_5_moe
huggingface.co1 month ago2 ptsView details
zerodigest/Qwen3.6-27B-YMQ-MTP-GGUF
text-generation · gguf · text-generation · quantizer
huggingface.co1 month agoView details
Shiftedx/Qwen3.8-27B-Abliterated-MLX-MXFP4-MTP
image-text-to-text · mlx · safetensors · qwen3_5
huggingface.co1 month agoView details
Open source
View all open source →langchain-ai/langchain langchain-openai==1.5.2a1
Initial release release(openai): 1.5.2a1 (#39709) feat(openai): extract gateway metadata from response headers when available (#39706) chore(openai): update snapshots (#39657) fix(openai): support o-series models in `get_num_tokens_from_messages` (#38710) release(openai): 1.5.1…
github.com1 month agoView details
langchain-ai/langchain langchain-core==1.5.6
Changes since langchain-core==1.5.5 chore(core): release 1.5.6 (#39704) feat(core): incorporate gateway metadata to traces (#39703)
github.com1 month agoView details
## [3.2.0](https://github.com/openai/openai-python/compare/v3.1.0...v3.2.0) (2026-08-17) ### Features * add Bedrock Runtime endpoint support (SDK-290) ([#3623](https://github.com/openai/openai-python/issues/3623)) ([86267d2](https://github.com/openai/openai-python/commit/86267d2…
github.com1 month agoView details
<details open> cuda : skip UMA override for HIP builds (#27083) AMD APUs report accurate memory via hipMemGetInfo. Using MemAvailable over-promises on small-carveout systems. fixes #18159 </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)…
github.com1 month agoView details
<details open> ci : push release tag explicitly in release.yml (#27261) Add a "Create and push git tag" step to the release job, right before the "Create release" step. The tag is created with git tag and pushed with the deploy key already configured by the Clone step, instead o…
github.com1 month agoView details
<details open> sycl: fix thread/block count in quantized cpy kernel launches (#27160) Adjusts the thread/block count to be proportional to the size of the quant, reducing under/over subscription. Largest perf improvement is the q4_0 -> f32 path, with, on a Arc 70, throughput goe…
github.com1 month agoView details
<details open> [SYCL] support OP OPT_STEP_ADAMW, OPT_STEP_SGD (#25268) * fix conflict * fix conflict of ops.md * fix conflict of ops.md * update the ops.md --------- Co-authored-by: Neo Zhang Jianyu <jianyu.zhang@intel.com> </details> **Website:** - <https://llama.app> **macOS/i…
github.com1 month agoView details