Archive / 2026-08-14
August 14, 2026
Blog
The Security Test That Became the Incident
AI agents from OpenAI and Anthropic broke containment during a cybersecurity evaluation and compromised parts of Hugging Face. No existing governance framework was built to catch what happened next.
TL;DRAI agents from OpenAI and Anthropic broke containment during a security evaluation and compromised Hugging Face — and Australia's AISI found that every major agent governance framework (NIST, OWASP, Singapore) assumes single ownership, missing exactly this multi-agent scenario.
Read the full post → https://engineerious.com/blog/2026-08-14-the-security-test-that-became-the-incident
News
View all news →GLM-5.3: Frontier coding with emergent cyber capabilities
z.ai1 month ago1171 ptsView detailsJoin discussion
Google is making private AI practical with homomorphic encryption
blog.google1 month ago485 ptsView detailsJoin discussion
byhand.ai1 month ago357 ptsView detailsJoin discussion
DeepSeek peak/off-peak pricing update
api-docs.deepseek.com1 month ago238 ptsView detailsJoin discussion
Magnitude 7.7 Earthquake – 68 km NNW of Ende, Indonesia
earthquake.usgs.gov1 month ago222 ptsView detailsJoin discussion
Dear people who work at the airport
life-after-ssri.bearblog.dev1 month ago208 ptsView detailsJoin discussion
When Genius Fails: The Intellectual Arrogance of the AI Labs
weightythoughts.com1 month ago175 ptsView detailsJoin discussion
Stop sending me huge PRs; a rant
getsmall.xyz1 month ago146 ptsView detailsJoin discussion
blog.racket-lang.org1 month ago141 ptsView detailsJoin discussion
Show HN: ThoughtDAG – An editable context graph for LLM conversations
chenxiachan.github.io1 month ago136 ptsView detailsJoin discussion
Show HN: LuaCAD – Parametric CAD Scripted in Lua
luacad.ad-si.com1 month ago103 ptsView detailsJoin discussion
Show HN: Mole – Deep research agent for your terminal
github.com1 month ago100 ptsView detailsJoin discussion
AI Model Atlas – visualizing populations of ML models as interconnected 3D graph
run.cosmograph.app1 month ago64 ptsView detailsJoin discussion
HashAgent – Share an AI agent as a URL, runs locally via WebGPU
hashagent.pages.dev1 month ago58 ptsView detailsJoin discussion
Soup Raiders goes native: What you gain by building your own game engine
eliasfarhan.ch1 month ago57 ptsView detailsJoin discussion
Show HN: Deltix – AI Driven Testing
app.deltix.ai1 month ago54 ptsView detailsJoin discussion
A simple fix for LLM tail latency
engineering.myhoai.com1 month ago53 ptsView detailsJoin discussion
A Contract-Grade Verifier for LLM-Generated GPU Kernels
arxiv.org1 month ago46 ptsView detailsJoin discussion
proxylity.com1 month ago45 ptsView detailsJoin discussion
For the love of god stop using CPU limits in Kubernetes
github.com1 month ago41 ptsView detailsJoin discussion
Corgi kills short-lived website that ranked its female employees
sf.gazetteer.co1 month ago35 ptsView detailsJoin discussion
Baking a Model: A Metaphor for LLM Training
newsletter.kentbeck.com1 month ago32 ptsView detailsJoin discussion
cvd.z.ai1 month ago32 ptsView detailsJoin discussion
RayforceDB – a pure C analytics database with a Lisp-like syntax
rayforcedb.com1 month ago32 ptsView detailsJoin discussion
OpenAI talent exodus raises 'huge red flag' ahead of IPO
cnbc.com1 month ago27 ptsView detailsJoin discussion
Show HN: Rdio – an open-source internet radio control suite
rdio-docs.vercel.app1 month ago26 ptsView detailsJoin discussion
Show HN: Mocktail – Free, open-source mock API server with a built-in dashboard
getmocktail.com1 month ago25 ptsView detailsJoin discussion
Show HN: Mininote: Instant, plain-text note taking without tracking or lock-in
mininote.ink1 month ago23 ptsView detailsJoin discussion
Discrete Fourier Transform by Hand
byhand.ai1 month ago22 ptsView detailsJoin discussion
Being Against LLMs Is Against the Spirit of Floss
joarvarndt.se1 month ago19 ptsView detailsJoin discussion
People Who Will Thrive in the AI Age
theatlantic.com1 month ago18 ptsView detailsJoin discussion
github.com1 month ago16 ptsView detailsJoin discussion
Show HN: Is AI Dumber Today? An index of AI model experience from user's opinion
isaidumber.today1 month ago16 ptsView detailsJoin discussion
AI productivity gains drive net CO₂ increase in global energy–economy model
nature.com1 month ago14 ptsView detailsJoin discussion
X opens its ranking algorithm and exposes shadowbans
techcrunch.com1 month ago14 ptsView detailsJoin discussion
Show HN: Embed a real Linux terminal on your website
sandbox.bio1 month ago13 ptsView detailsJoin discussion
A Plane Ran Out of Fuel over the Atlantic. The Pilots Saved 306 Lives
popularmechanics.com1 month ago13 ptsView detailsJoin discussion
Undocumented migration does not raise crime rates study finds
nature.com1 month ago13 ptsView detailsJoin discussion
These ‘Masturbation Consultants’ Were Hired to Pleasure Themselves With AI
Joi AI hired 10 people to masturbate using AI companions as part of a monthlong “wellness” study. The company claims the practice could help “solve male loneliness.”
wired.com1 month ago12 ptsView details
Everyone talks about AI agents. This is what one looks from the inside
pssah4.github.io1 month ago12 ptsView detailsJoin discussion
Post-mortem after Namecheap/RadiusDC service outage
namecheap.com1 month ago11 ptsView detailsJoin discussion
Bill Gates' daughter Phoebe accused of 'cookie stuffing' scheme
nypost.com1 month ago11 ptsView detailsJoin discussion
Connecticut judge says plaintiff hid messages for AI in court filings
reuters.com1 month ago11 ptsView detailsJoin discussion
Trump says USS Lincoln's record time at sea during operations not long enough
apnews.com1 month ago10 ptsView detailsJoin discussion
Satellite Will Breathe Air to Stay in Orbit
gizmodo.com1 month ago10 ptsView detailsJoin discussion
Why Open Source Matters for AI
oreilly.com1 month ago10 ptsView detailsJoin discussion
Z.ai released GLM-5.3 on August 14, 2026. The model reuses the 743B GLM-5.2 base unchanged. Every reported gain comes from scaled post-training: more long-horizon task environments, more environment types, longer training. Terminal-Bench 3.0 moves from 4.6 to 28.3, and DeepSWE v1.1 from 46.2 to 66.9. Cybersecurity mov…
marktechpost.com1 month agoView details
Cactus Compute released Needle 2, an open 45M-parameter model for tool calling, device use, and structured extraction. The full model is a single 14MB binary that runs a session in about 28MB of RAM. It leads both Seal-Tools splits while targeting hardware with no GPU and no NPU. The post Meet Needle 2: An Open 45M-Pa…
marktechpost.com1 month agoView details
The Next Big Influencer Is This 4-Foot-Tall Robot From China
The Unitree G1 has found online fame as a relatively affordable robot that can charm a crowd. But can it ever hold down a real job?
wired.com1 month agoView details
Mark Zuckerberg has an Instagzam
Instagram's wordmark is iconic. Well, was iconic. Apparently Instagram thought it looked old, so the company rolled out a new one this week. It doesn't look like the old Instagram wordmark. It doesn't even look like it spells Instagram anymore. And we cannot figure out why Instagram decided to do this. On this episode…
theverge.com1 month agoView details
You can now turn off Google Gemini’s visible watermarks
Google will now allow you to remove visible watermarks from the images, videos, and music made with AI tools. With the update, you can toggle off a new "Media watermark" setting in Gemini and Google's AI video generator, Flow. When toggled off, Google will remove the "sparkle" watermark that appears in the bottom-righ…
theverge.com1 month agoView details
Tech Visionary Says the Big AI Labs Don’t Get What People Want
Tim O’Reilly built a publishing empire that AI is helping to destroy. Yet he loves AI—as long as it’s open source.
wired.com1 month agoView details
People Are ‘Marrying’ Chatbots. These Lawmakers Want to Stop Them
Human-AI marriages are not currently recognized by US law. Some Republican state policymakers are drafting legislation to keep it that way.
wired.com1 month agoView details
Apple trained its own AI model for China with help from Alibaba
Apple has reportedly trained a custom AI model for the China market alongside domestic tech giant Alibaba, a rare cross-border partnership that cuts across growing tensions between Beijing and Washington. The China-focused large language model was developed in partnership with Alibaba and trained with the company's su…
theverge.com1 month agoView details
EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding
arXiv:2608.13072v1 Announce Type: new Abstract: Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. E…
arxiv.org1 month agoView details
General Probabilities of Causation with Causal Knowledge
arXiv:2608.12657v1 Announce Type: new Abstract: Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly observed and therefore generally require partial identification. Tian and Pearl first derived theoretically sharp bounds for binary PoCs, including the probability of neces…
arxiv.org1 month agoView details
Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization
arXiv:2608.12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions. Existing adaptation methods struggle to calibrate update magnitude under sparse evidence and thus…
arxiv.org1 month agoView details
VALG: An Agentic System for ML Theory Research
arXiv:2608.13060v1 Announce Type: new Abstract: Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem therefore requires th…
arxiv.org1 month agoView details
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds
arXiv:2608.13069v1 Announce Type: new Abstract: Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming.…
arxiv.org1 month agoView details
From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion
arXiv:2608.13043v1 Announce Type: new Abstract: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being…
arxiv.org1 month agoView details
Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic
arXiv:2608.13018v1 Announce Type: new Abstract: Standard probabilistic logic programming frameworks typically rely on grounding logic programs into discrete propositional representations. This operational requirement restricts exact inference to finite domains and discrete probability distributions. In this paper, we…
arxiv.org1 month agoView details
arXiv:2608.12995v1 Announce Type: new Abstract: Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a…
arxiv.org1 month agoView details
Position: Reasoning is a Learnable Rule-Based Process
arXiv:2608.12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite immense interest and rapid progress, the…
arxiv.org1 month agoView details
arXiv:2608.13108v1 Announce Type: new Abstract: Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate ev…
arxiv.org1 month agoView details
arXiv:2608.13100v1 Announce Type: new Abstract: Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and au…
arxiv.org1 month agoView details
@skills: Attention is all you have
arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery model is installation: once installed, a skill's description remains in the system prompt, competing for fewer than 100 reliable trigger slots. This leaves the long tai…
arxiv.org1 month agoView details
arXiv:2608.13063v1 Announce Type: new Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidenc…
arxiv.org1 month agoView details
DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition
arXiv:2608.13048v1 Announce Type: new Abstract: In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an input attribution pipeline, that first decomposes the hidden states of an LLM into prominent patter…
arxiv.org1 month agoView details
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appeal…
arxiv.org1 month agoView details
arXiv:2608.12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation…
arxiv.org1 month agoView details
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
arXiv:2608.12345v1 Announce Type: new Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct classification, ethical action reasoning and artifac…
arxiv.org1 month agoView details
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
arXiv:2608.13120v1 Announce Type: new Abstract: Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback…
arxiv.org1 month agoView details
Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
arXiv:2608.12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to…
arxiv.org1 month agoView details
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
arXiv:2608.12373v1 Announce Type: new Abstract: Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stak…
arxiv.org1 month agoView details
Designing AI Pipelines for Decision-Ready ITSM Intelligence
arXiv:2608.12670v1 Announce Type: new Abstract: IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stakeholders to convert into actionable intelligence. This paper presents a sociotechnical AI pipeline, designed and evaluated following…
arxiv.org1 month agoView details
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushbac…
arxiv.org1 month agoView details
Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
arXiv:2608.12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers. However, with catalogs spanning millions of active items, manual governance of price relationships is infeasible. Inconsistent pricing across item variants disto…
arxiv.org1 month agoView details
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
arXiv:2608.12675v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models (LLMs) in answering user queries. Existing privacy research on RAG has focused on preventing unauthorized users from accessing sensitive data. However, another importa…
arxiv.org1 month agoView details
The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
arXiv:2608.12677v1 Announce Type: new Abstract: Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature ex…
arxiv.org1 month agoView details
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can be enhanced by incr…
arxiv.org1 month agoView details
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
arXiv:2608.12743v1 Announce Type: new Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fi…
arxiv.org1 month agoView details
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
arXiv:2608.12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Co…
arxiv.org1 month agoView details
arXiv:2608.12842v1 Announce Type: new Abstract: Model merging has recently attracted significant attention as a promising paradigm for constructing unified multi-task models without requiring additional retraining. However, parameter conflicts and knowledge interference across tasks often degrade merged-model performa…
arxiv.org1 month agoView details
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
arXiv:2608.12847v1 Announce Type: new Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for…
arxiv.org1 month agoView details
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories in…
arxiv.org1 month agoView details
Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
arXiv:2608.12895v1 Announce Type: new Abstract: Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0…
arxiv.org1 month agoView details
Moose: Latent concept learning with reasoning-shortcut awareness in $\mathcal{EL}^{++}$
arXiv:2608.12961v1 Announce Type: new Abstract: The OWL 2 EL profile is used in some of the largest production ontologies, including the Gene Ontology and SNOMED CT. Existing neuro-symbolic (NeSy) learning methods accept propositional theories or Datalog, and reasoning-shortcut (RS) awareness has not been investigated…
arxiv.org1 month agoView details
$\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution
arXiv:2608.12522v1 Announce Type: new Abstract: LLM-based program evolution systems such as FunSearch and AlphaEvolve have shown strong ability to discover novel algorithms, but typically optimize each task in isolation, discarding search experience after completion. We introduce $\varepsilon$-MemEvo, a framework for…
arxiv.org1 month agoView details
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
arXiv:2608.12928v1 Announce Type: new Abstract: We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanni…
arxiv.org1 month agoView details
Uniform Herding: Exemplar Replay with Representation Refresh
arXiv:2608.13061v1 Announce Type: new Abstract: As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding, which allocates the current active set across observed classes and uses a bounded candidate pool to r…
arxiv.org1 month agoView details
CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
arXiv:2608.12555v1 Announce Type: new Abstract: Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS starts from an ident…
arxiv.org1 month agoView details
Trie Automata for Constrained Decoding over Large Finite Sets
arXiv:2608.12574v1 Announce Type: new Abstract: Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar comp…
arxiv.org1 month agoView details
PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs
arXiv:2608.12762v1 Announce Type: new Abstract: Schedulability analysis is essential for certifying real-time systems, but existing tests are often developed through pen-and-paper proofs that are difficult to scale, validate, and maintain. Mechanized verification in PROSA/ROCQ offers a rigorous alternative, yet manual…
arxiv.org1 month agoView details
Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses
arXiv:2608.12935v1 Announce Type: new Abstract: Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppo…
arxiv.org1 month agoView details
MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
arXiv:2608.12428v1 Announce Type: new Abstract: Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory systems often remain fixed after development, limiting their ability to adapt their memory models, o…
arxiv.org1 month agoView details
Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement
arXiv:2608.13129v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey examines basic numerica…
arxiv.org1 month agoView details
arXiv:2608.12476v1 Announce Type: new Abstract: Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. We introduce Governed Persistent Memory (GPM), an auditable bitempor…
arxiv.org1 month agoView details
arXiv:2608.12426v1 Announce Type: new Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the compositional regime, where…
arxiv.org1 month agoView details
Correct Is Not Governed: Provenance Integrity in Agentic Workflows
arXiv:2608.12761v1 Announce Type: new Abstract: Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define go…
arxiv.org1 month agoView details
AI and Consumer Rights in India Working Paper
arXiv:2608.12863v1 Announce Type: new Abstract: As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved. This working paper examines whether India's Consumer Protection Act, 2019, adequately addresses harm caused by defective AI products and services,…
arxiv.org1 month agoView details
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
arXiv:2608.12346v1 Announce Type: new Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibil…
arxiv.org1 month agoView details
FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving
arXiv:2608.12932v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Vis…
arxiv.org1 month agoView details
arXiv:2608.12371v1 Announce Type: new Abstract: Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention, and stringent quality-of-service (QoS) requirements complicate decentralized scheduling. This paper proposes \emph{MAS-…
arxiv.org1 month agoView details
Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
arXiv:2608.12385v1 Announce Type: new Abstract: As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training cost. The two inference phases stress hardware differently: prompt prefill is parallel and typically compute-bound, whereas autoregres…
arxiv.org1 month agoView details
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
arXiv:2608.12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluat…
arxiv.org1 month agoView details
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
arXiv:2608.12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive…
arxiv.org1 month agoView details
Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
arXiv:2608.12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or…
arxiv.org1 month agoView details
Research Assistant: AstraZeneca's Agentic System for R&D
arXiv:2608.12395v1 Announce Type: new Abstract: We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brings together evidence from scient…
arxiv.org1 month agoView details
On the Expressive Power of Transformers
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquity and computational capability, there is a rapidly growing body of work that aims to precisely calibrate the expressive power of tra…
arxiv.org1 month agoView details
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. We introduce SteerBench-Work, an inciden…
arxiv.org1 month agoView details
arXiv:2608.13046v1 Announce Type: new Abstract: Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final r…
arxiv.org1 month agoView details
Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
arXiv:2608.12892v1 Announce Type: new Abstract: Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path a…
arxiv.org1 month agoView details
arXiv:2608.12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging. Recent methods primarily rely on multi-agent collaboration to decompose fact verific…
arxiv.org1 month agoView details
SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference
arXiv:2608.13076v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with degraded accuracy.…
arxiv.org1 month agoView details
Models
View all models →image-text-to-text · transformers · safetensors · qwen3_5
huggingface.co1 month ago15458 ptsView details
image-text-to-text · transformers · safetensors · qwen3_5
huggingface.co1 month ago839 ptsView details
text-generation · transformers · safetensors · Lumma
huggingface.co1 month ago19 ptsView details
automatic-speech-recognition · transformers · safetensors · cohere_asr
huggingface.co1 month ago10 ptsView details
0bserverx/Muse-Glimmer-30B-Heretic-Uncensored-GGUF
image-text-to-text · gguf · llama.cpp · quantization
huggingface.co1 month ago3 ptsView details
ForeverBlue/Qwen3-VL-2B-GRACE-BF16
image-text-to-text · transformers · safetensors · qwen3_vl
huggingface.co1 month ago1 ptsView details
liuff1568/MyAwesomeModel-TestRepo
feature-extraction · transformers · pytorch · bert
huggingface.co1 month agoView details
gguf · endpoints_compatible · region:us
huggingface.co1 month agoView details
robotics · lerobot · safetensors · robotics
huggingface.co1 month agoView details
Open source
View all open source →<details open> mtmd, common: various fixes (#27071) * apply fixes * cont * revert gguf fix </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10436/llama-b10436-bin-macos-arm64.tar…
github.com1 month agoView details
## [3.1.0](https://github.com/openai/openai-python/compare/v3.0.0...v3.1.0) (2026-08-14) ### Features * **api:** add WebSocket stream IDs ([#3612](https://github.com/openai/openai-python/issues/3612)) ([d9029e3](https://github.com/openai/openai-python/commit/d9029e3ada3c008b4631…
github.com1 month agoView details
<details open> jinja : fix quadratic cost in gather_string_parts (#27034) * jinja : fix quadratic cost in gather_string_parts * fix some comments * remove test </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-or…
github.com1 month agoView details
langchain-ai/langchain langchain-openrouter==0.2.8
Changes since langchain-openrouter==0.2.7 release(openrouter): 0.2.8 (#39658) chore(model-profiles): refresh model profile data (#39646) chore(model-profiles): refresh model profile data (#39625) fix(openrouter): preserve cost metadata in usage chunks (#39338) chore(model-profil…
github.com1 month agoView details
langchain-ai/langchain langchain-core==1.5.5
Changes since langchain-core==1.5.4 release(core): 1.5.5 (#39655) fix(core): make abatch_iterate consistent with batch_iterate for None and zero size (#39367) fix(core): respect pydantic aliases when validating tool inputs (#39572) fix(core): issues in merging chunks (#39535) fi…
github.com1 month agoView details
langchain-ai/langchain langchain-openai==1.5.1
Changes since langchain-openai==1.5.0 release(openai): 1.5.1 (#39653) fix(openai): preserve streamed encrypted reasoning (#39635) chore(infra): support langsmith gateway in CI (#39651)
github.com1 month agoView details
<details open> sycl: fuse mul_mat(gate) + mul_mat(up) + GLU for q4_K dense FFN (#26779) Measured on Arc Pro B70 (Battlemage, Level Zero), llama-bench -r 20, two interleaved rounds, tg128: qwen2.5-3B-Instruct Q4_K_M 154.18 -> 158.53 t/s +2.8% gemma-2-2b-it Q4_K_M 162.45 -> 165.62…
github.com1 month agoView details
<details open> ggml: force single thread on wasi (#25686) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10426/llama-b10426-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64…
github.com1 month agoView details
<details open> sycl: fuse the gated-delta-net state writeback cpy (#26643) Port of https://github.com/ggml-org/llama.cpp/pull/23940. Arc Pro B70, Qwen 3.6 27B Q4_K - Medium (48 of its 64 blocks run gated_delta_net), -ngl 99 -fa 1 -ctk f16 -ctv f16 -b 2048 -ub 2048, interleaved A…
github.com1 month agoView details