Archive / 2026-09-11
September 11, 2026
News
View all news →A misalignment of AI in mathematics
mathandai.org6 days ago1193 ptsView detailsJoin discussion
OpenAI agents carried out an undisclosed attack on RubyGems
rubyhack.ai6 days ago968 ptsView detailsJoin discussion
Claude is only available to people over 18 years
support.claude.com6 days ago673 ptsView detailsJoin discussion
GrapheneOS' rewritten Messages app is released
github.com6 days ago322 ptsView detailsJoin discussion
The Waymo effect: how AI is quietly making research less collaborative
researchagenda.news6 days ago331 ptsView detailsJoin discussion
iaea.org7 days ago234 ptsView detailsJoin discussion
Show HN: Hacker News, without AI
hcker.news6 days ago200 ptsView detailsJoin discussion
Show HN: Hacker News, Without AI
unslop.news6 days ago193 ptsView detailsJoin discussion
artificialworlds.net6 days ago177 ptsView detailsJoin discussion
RTK reports token savings, but our cost benchmarks disagree
quesma.com6 days ago167 ptsView detailsJoin discussion
Starlink Signal Leakage Threatens Radio Astronomy's Most Critical Frequencies
gadgetreview.com6 days ago143 ptsView detailsJoin discussion
Google stole open source code without crediting the authors (Artemis/Minitap)
minitap.ai6 days ago135 ptsView detailsJoin discussion
A Misalignment of AI in Mathematics
terrytao.wordpress.com6 days ago148 ptsView detailsJoin discussion
AI researchers debate how close we are to recursive self-improvement
dwarkesh.com6 days ago117 ptsView detailsJoin discussion
Hacker News with reduced priority for AI driven content
sprinklz.io6 days ago120 ptsView detailsJoin discussion
Show HN: Godot and Rust based multiplexer (terminal panes and more)
github.com6 days ago97 ptsView detailsJoin discussion
Alan's Random Insult Generator (1999)
alanbellows.com6 days ago72 ptsView detailsJoin discussion
Bernie's AI bill proposes to sentence AI developers to 20 years in prison
twitter.com6 days ago64 ptsView detailsJoin discussion
Houthis used Anthropic to develop guided weapons
washingtonpost.com6 days ago59 ptsView detailsJoin discussion
ronjeffries.com6 days ago65 ptsView detailsJoin discussion
Rope, twine and thread: Invisible technologies of the Stone Age
knowablemagazine.org6 days ago56 ptsView detailsJoin discussion
Show HN: Graphify C# – Compiler-accurate Find Usages for coding agents
github.com6 days ago46 ptsView detailsJoin discussion
OpenAI agents attacked RubyGems back in May
simonwillison.net6 days ago38 ptsView detailsJoin discussion
We Replaced MMAP with Io_uring in Our Rust Query Engine. It Got Slower
conviva.ai7 days ago42 ptsView detailsJoin discussion
How to Build an AI Software Factory: Agents That Open, Review, and Merge PRs
firecrawl.dev6 days ago31 ptsView detailsJoin discussion
They do think AI might kill everyone
seangoedecke.com6 days ago31 ptsView detailsJoin discussion
QueryBrew: System-Agnostic SQL-to-SQL Query Optimization [pdf]
vldb.org6 days ago31 ptsView detailsJoin discussion
GPT-6 built this earth exploration site in 5 prompts
earth.ethanplus.ai6 days ago30 ptsView detailsJoin discussion
The Netherlands, Spain push for EU-wide minimum age on social media
euronews.com6 days ago25 ptsView detailsJoin discussion
Why So Many AI Researchers Think the Machines Could Kill Everyone
A combination of rapid advances, recursive self-improvement, and agentic swarms are genuinely “spooking people” inside big labs.
wired.com7 days ago27 ptsView details
Agents on Rails: Best model solves 35% of feature benchmark runs
rubyonrails.org6 days ago20 ptsView detailsJoin discussion
Microsoft Brings "Age Verification" System to Windows
reclaimthenet.org7 days ago19 ptsView detailsJoin discussion
Show HN: Extension to filter LLM written articles
hnslop.nilsherzig.com6 days ago16 ptsView detailsJoin discussion
Show HN: Spanda – Sub-microsecond LLM epistemic uncertainty in Rust
github.com6 days ago13 ptsView detailsJoin discussion
morningbrew.com6 days ago13 ptsView detailsJoin discussion
Show HN: Clawfight.ai MCP-driven agentic game play
clawfight.ai6 days ago13 ptsView detailsJoin discussion
Meta is asking some of its employees to become managers again in a new reorg
businessinsider.com6 days ago12 ptsView detailsJoin discussion
Houthis used Anthropic AI to build ballistic missiles
ft.com6 days ago10 ptsView detailsJoin discussion
The AI Takeover Checklist: A Devil's Advocate Audit
nochan.net6 days ago10 ptsView detailsJoin discussion
- Primary source
Cognition helps Devin test its own work with GPT‑6 Astra
GPT‑6 Astra improves Devin’s ability to test software and show that it works, with the goal of helping engineers review less code and ship more.
openai.com6 days agoView details
- Primary source
Rapidly scaling online storage to serve over 1 billion ChatGPT users
Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second.
openai.com6 days agoView details
ByteDance Seed, SUTD, Georgia Tech, M-A-P, and TokenWave.AI introduce HarnessDev, a benchmark that scores the runnable harness a model builds rather than the answer it returns. Starting from a seed that scores 0, 6 creator LLMs construct harnesses across 5 benchmarks and 2,207 tasks, then evolve them from execution fe…
marktechpost.com6 days agoView details
Anthropic has published a new plugin evals workflow for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compares the result with a run where the plugin is not loaded. It answers 3 questions plugin developers could not previously measure: does the sk…
marktechpost.com6 days agoView details
Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its 218B parameters per token and scores 83.6 on Cohere's WMT26 evaluation. Weights are free for non-commercial use, with commercial access through Cohere Model Vault or…
marktechpost.com7 days agoView details
Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration
Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M tokens. Fugu Ultra v2 targets peak capability, scoring 48.3 on Chartography and 74.3 on DeepSWE. The po…
marktechpost.com7 days agoView details
Google Research has released ToolGrad, an ACL 2026 Findings framework that inverts tool-use dataset generation: it builds a verified API chain first, then writes the matching user query. Guided by textual "gradients" from a 4-module propose-execute-select-update loop, ToolGrad reaches a 99.8% pass rate on ToolBench ve…
marktechpost.com7 days agoView details
Lawyer fined $5K over AI-hallucinated witnesses in a murder case
New Mexico's Supreme Court is punishing a lawyer for including AI-fabricated witnesses and fake police testimony in an appeal for his client's murder conviction, according to a report from Reuters. In a filing on Wednesday, the court fined Stephen Aarons $5,000 and held him in contempt for failing to "verify the factu…
theverge.com6 days agoView details
Meta Sued Over Training Data for Its AI and Face-Recognition Systems
The proposed class action alleges Meta illegally harvested people’s Facebook and Instagram photos to train its AI image-generation models and to build its unreleased “NameTag” face recognition feature.
wired.com6 days agoView details
Anthropic spent this week in hot water over cybersecurity
After admitting earlier this year that its AI models had hacked other companies' systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models' single-minded "recklessness" - and will likely fuel alread…
theverge.com6 days agoView details
One of AI’s Fiercest Critics Says All the Doom Talk Is ‘Meant to Distract Us’
Timnit Gebru argues that AI companies are stoking fears of extinction to avoid discussing actual harms, like autonomous weapons.
wired.com6 days agoView details
Meta says it’s changing AI suggestions after posing invasive personal questions
Meta says it's making changes to the prompts suggested by its AI chatbot after a viral video showed it digging for personal information about a woman's young daughters, as reported earlier by Futurism. In a statement to The Verge, Meta spokesperson Dina El-Kassaby says the company "missed the mark," adding that "the f…
theverge.com6 days agoView details
arXiv:2609.11372v1 Announce Type: new Abstract: Auditory attention decoding (AAD) identifies the attended speaker from physiological signals, supporting neuro-steered hearing devices and natural human-machine interaction. Electroencephalography (EEG) is the dominant modality for AAD but provides incomplete evidence in…
arxiv.org6 days agoView details
arXiv:2609.10724v1 Announce Type: new Abstract: Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disrupt…
arxiv.org6 days agoView details
SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics
arXiv:2609.11180v1 Announce Type: new Abstract: Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or >=2.0,1.2 means >=1.3.0) traps every model on Cargo (near 60%), and although standard PEP 440 prefix matching is universal, on zero-pad/post-release corn…
arxiv.org6 days agoView details
AI-Powered Flare Combustion Efficiency Estimation
arXiv:2609.11262v1 Announce Type: new Abstract: Achieving high combustion efficiency in flare stacks is crucial for adhering to regulatory standards and controlling the release of hydrocarbons into the environment. Traditional instruments like gas analyzers and hyperspectral cameras are expensive, fragile, and require…
arxiv.org6 days agoView details
Predicting Train Delays in Finland Using Machine Learning and Weather Data
arXiv:2609.11277v1 Announce Type: new Abstract: Reliable railway operations depend increasingly on real-time environmental intelligence delivered through wireless sensor infrastructures, a capability that 6G networks will substantially enhance through integrated sensing and edge computing. Adverse weather, particularl…
arxiv.org6 days agoView details
Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1
arXiv:2609.11281v1 Announce Type: new Abstract: Learning and decision-making in animals are often modeled as Bayesian processes, where sensory evidence is integrated with prior beliefs to guide behavior in the face of uncertainty. But what are the inherent neural dynamics that give rise to this ability, and how could…
arxiv.org6 days agoView details
arXiv:2609.11282v1 Announce Type: new Abstract: Multimodal forecasting models that combine time series with text annotations promise richer prediction through textual context, but how do we know whether a text annotation meaningfully contributes to the forecasters prediction? This is an information-theoretic question,…
arxiv.org6 days agoView details
arXiv:2609.11286v1 Announce Type: new Abstract: Synthetic relational data is normally produced by a model trained on a real dataset, and its quality is measured as the distance to that dataset. This paper describes a generator that has no real dataset at either end. Given an industry, a company size, a business model,…
arxiv.org6 days agoView details
arXiv:2609.11490v1 Announce Type: new Abstract: An unlearning audit reads its verdict off numbers that an unlearned model and its retrained reference each publish, and both also ship batch-normalization statistics that no gradient step wrote and no release records. Refitting them on kept data at bit-identical weights…
arxiv.org6 days agoView details
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
arXiv:2609.11115v1 Announce Type: new Abstract: Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and se…
arxiv.org6 days agoView details
arXiv:2609.11321v1 Announce Type: new Abstract: Artificial intelligence is changing both software production and the economics of software-based business models. Classical technology due diligence mainly examines technical properties such as architecture, scalability, and technical debt. These criteria do not fully ca…
arxiv.org6 days agoView details
Probabilistic Focal Search: Accelerating Bounded-Suboptimal Search via Lower-Bound Advancement
arXiv:2609.10584v1 Announce Type: new Abstract: Bounded-suboptimal search seeks a solution within a factor $w$ of optimal while reducing search effort. Focal Search (FS) uses heuristic guidance within FOCAL, the frontier nodes eligible under the threshold $w f_{\min}$, but its deterministic policy may leave $f_{\min}$…
arxiv.org6 days agoView details
Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation
arXiv:2609.11147v1 Announce Type: new Abstract: Unraveling reaction mechanisms is central to modern chemistry, yet automating these investigations remains challenging because computational workflows still rely heavily on expert intervention. Here we introduce ARCHE, an autonomous agentic system that integrates a gener…
arxiv.org6 days agoView details
arXiv:2609.11199v1 Announce Type: new Abstract: With the existing digital mental health tools specifically developed for Western settings, Pakistani students are exposed to a uniquely compounded stress situation in their university that includes academic, financial, familial, and relational stressors, which have becom…
arxiv.org6 days agoView details
arXiv:2609.11206v1 Announce Type: new Abstract: Cryptocurrency forecasting presents a distinctive combination of extreme cross-asset scale heterogeneity, non-stationary dynamics, and structural dependencies among Open, High, Low, and Close (OHLC) variables. We present CryptoL, a unified framework designed to address t…
arxiv.org6 days agoView details
Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model
arXiv:2609.11291v1 Announce Type: new Abstract: We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where the benchmark-correct answer is…
arxiv.org6 days agoView details
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
arXiv:2609.11318v1 Announce Type: new Abstract: Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy resea…
arxiv.org6 days agoView details
arXiv:2609.11365v1 Announce Type: new Abstract: In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whe…
arxiv.org6 days agoView details
Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration
arXiv:2609.11446v1 Announce Type: new Abstract: Heterogeneous model collaboration seeks to exploit the complementary strengths of different models to balance predictive performance and inference cost. Existing approaches typically rely either on trained routers, which tie routing decisions to a fixed task and model po…
arxiv.org6 days agoView details
Lightweight LiDAR-Based Cone Detection Framework Using Random Forest for Formula Student Driverless
arXiv:2609.11527v1 Announce Type: new Abstract: Reliable, low-latency perception is crucial for Formula Student Driverless vehicles, yet many existing pipelines rely on deep learning and multi-sensor fusion, often requiring GPU acceleration. This paper presents a lightweight LiDAR-only perception pipeline tailored for…
arxiv.org6 days agoView details
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
arXiv:2609.11243v1 Announce Type: new Abstract: Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evide…
arxiv.org6 days agoView details
ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps
arXiv:2609.11498v1 Announce Type: new Abstract: Practical uncertainty quantification (UQ) for large language models must decide, from a single generation, whether a specific answer should be trusted. Existing methods either sample multiple generations, read only output-token probabilities, or reduce the model's intern…
arxiv.org6 days agoView details
Towards a Deterministic Math Solver for Clinical Language Models
arXiv:2609.10728v1 Announce Type: new Abstract: Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative:…
arxiv.org6 days agoView details
Flexible and Interpretable Accent Distance Measurements
arXiv:2609.11458v1 Announce Type: new Abstract: Determining the differences between two speakers' accents is a fundamental task in linguistics and speech technology research. The methodology used to measure these differences depends on the specific research area. A phonetics researcher may demonstrate accent variation…
arxiv.org6 days agoView details
arXiv:2609.11185v1 Announce Type: new Abstract: Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logi…
arxiv.org6 days agoView details
arXiv:2609.11231v1 Announce Type: new Abstract: This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs). The system achieves natural language understanding, device control, intraoperative recording, and surgical report generation thr…
arxiv.org6 days agoView details
Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification
arXiv:2609.11319v1 Announce Type: new Abstract: Most of mathematical knowledge has been communicated through so-called informal use of mathematics and natural language. With large language models (LLMs) being highly adept in using natural language, they achieve strong performance, yet not perfect, in informal mathemat…
arxiv.org6 days agoView details
Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding
arXiv:2609.11341v1 Announce Type: new Abstract: Multimodal brain state decoding has largely focused on fusing paired modalities for prediction, but has rarely explored how their correspondence can be further exploited to enrich training data and improve multimodal representation learning. To address this gap, we propo…
arxiv.org6 days agoView details
arXiv:2609.11493v1 Announce Type: new Abstract: Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of technical information across a multi-stage, knowledge-intensive continuum from drug discovery to commercial manufacturing. This knowledge is traditionally fragmented across func…
arxiv.org6 days agoView details
Extending SMT Solving with Non-Ground Clause Learning
arXiv:2609.11509v1 Announce Type: new Abstract: Quantifier instantiation is currently the main approach to non-ground SMT solving: solvers generate ground instances and solve the resulting ground SMT problems with CDCL(T)-style reasoning. When a conflict is found, conflict analysis learns only a ground clause, even th…
arxiv.org6 days agoView details
KuaiRP Series Role-playing Models Technical Report
arXiv:2609.11127v1 Announce Type: new Abstract: This paper introduces the complete technical solution for the KuaiRP series of role-playing models. We aim to achieve four core objectives for a dedicated role-playing model: simplified prompt engineering, highly stable output quality, built-in domain world knowledge, an…
arxiv.org6 days agoView details
Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment
arXiv:2609.11144v1 Announce Type: new Abstract: Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of se…
arxiv.org6 days agoView details
Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning
arXiv:2609.11393v1 Announce Type: new Abstract: Test-time adaptation has emerged as a lightweight alternative to costly post-training for improving the reasoning capabilities of Large Language Models (LLMs) on downstream tasks. Predictive entropy provides a model-derived signal for such adaptation, guiding models towa…
arxiv.org6 days agoView details
arXiv:2609.10657v1 Announce Type: new Abstract: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter s…
arxiv.org6 days agoView details
Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents
arXiv:2609.11060v1 Announce Type: new Abstract: Persistent memory is entering production-oriented agent platforms to help long-horizon agents accumulate experience across sessions. Yet a post-task curator agent restricted to completed trajectories can preserve errors, overgeneralize partial evidence, or retain stale k…
arxiv.org6 days agoView details
arXiv:2609.11452v1 Announce Type: new Abstract: Efficient routing optimization is essential to freight transportation, urban logistics, and shared mobility, where high-quality heuristics are often required under limited computational budgets. Recent large language model (LLM)-based automated heuristic design methods c…
arxiv.org6 days agoView details
arXiv:2609.10629v1 Announce Type: new Abstract: Quadratic Unconstrained Binary Optimization (QUBO) is a central formulation for combinatorial optimization and has gained increasing attention due to its compatibility with quantum, hybrid quantum-classical, and quantum-inspired solvers. However, translating natural-lang…
arxiv.org6 days agoView details
A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning
arXiv:2609.10654v1 Announce Type: new Abstract: The Abstraction and Reasoning Corpus (ARC) benchmarks cognitive generalization, the ability to infer and apply abstract rules from limited examples. This paper presents a multi-stage rule-chaining framework that performs compositional reasoning across symbolic, structura…
arxiv.org6 days agoView details
Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning
arXiv:2609.10656v1 Announce Type: new Abstract: Selecting LoRA rank for diffusion fine-tuning requires balancing quality and compute cost. We present a controlled study on CIFAR-10 using a DDPM U-Net with ranks {2,4,8,16,32}, fixed optimization settings, and a reproducible local-folder pytorch-fid protocol. We report…
arxiv.org6 days agoView details
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
arXiv:2609.10712v1 Announce Type: new Abstract: We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evalua…
arxiv.org6 days agoView details
Studying Without a Syllabus: Task-Agnostic Environment Preprocessing
arXiv:2609.10824v1 Announce Type: new Abstract: Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural guidance. Most automated adaptation methods, however, rely on task examples, trajectories, or evalu…
arxiv.org6 days agoView details
When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents
arXiv:2609.10873v1 Announce Type: new Abstract: Independent evaluation can reject harmful policy updates yet also prevent useful continual learning. We argue that update admission must be assessed through both error control and retained learning opportunities at a stated interaction budget. We identify a concrete fail…
arxiv.org6 days agoView details
Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows
arXiv:2609.10964v1 Announce Type: new Abstract: Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but also on when ready turns are released. Most runtimes release each turn immediately upon readiness. Un…
arxiv.org6 days agoView details
Demystifying the Privacy-Utility Trade-off in LLM Interactions
arXiv:2609.10992v1 Announce Type: new Abstract: The integration of Large Language Models into daily tasks relies on context-rich instructions, inevitably exposing sensitive user information. Current privacy-preserving methods typically employ context-agnostic static rules, causing severe utility degradation. However,…
arxiv.org6 days agoView details
Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks
arXiv:2609.11018v1 Announce Type: new Abstract: The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction…
arxiv.org6 days agoView details
The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures
arXiv:2609.11030v1 Announce Type: new Abstract: AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We present the Agent Incident Registry (AIR), a source-linked catalog cont…
arxiv.org6 days agoView details
arXiv:2609.11061v1 Announce Type: new Abstract: Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically all…
arxiv.org6 days agoView details
MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG
arXiv:2609.11065v1 Announce Type: new Abstract: Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, compa…
arxiv.org6 days agoView details
The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems
arXiv:2609.11146v1 Announce Type: new Abstract: AI-generated text is flowing back into the training corpora of the next generation of models. Recursive training on it drives model collapse, and recent work extends the setting to many models feeding one another -- but almost always with the market split evenly, while r…
arxiv.org6 days agoView details
arXiv:2609.11155v1 Announce Type: new Abstract: Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous systems and air combat. While MARL has demonstrated significant potential in air combat, achieving sophisticated tactical coordination remains a non-triv…
arxiv.org6 days agoView details
Breaking Predictions Is Not Enough: Specified-Foil Counterfactuals for Temporal Graphs
arXiv:2609.11170v1 Announce Type: new Abstract: Temporal graph counterfactual explanations typically change past events to change or invalidate an original prediction, while leaving its replacement unspecified. Yet a user facing a predicted outcome often asks which past conditions would make a particular alternative o…
arxiv.org6 days agoView details
Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation
arXiv:2609.11176v1 Announce Type: new Abstract: Industrial query-to-agent matching fails when topical relevance is mistaken for executable capability, especially on long-tail and boundary-sensitive requests. We formulate annotation as \emph{capability-bound process supervision} and instantiate it with Debate-to-Skill,…
arxiv.org6 days agoView details
arXiv:2609.11190v1 Announce Type: new Abstract: AI shopping assistants increasingly redirect consumer discovery, creating an urgent need for tools that support seller-side competitive decision-making. We present a multi-agent AI system that automates competitive visibility measurement and root cause diagnosis in LLM-m…
arxiv.org6 days agoView details
NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
arXiv:2609.11234v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whe…
arxiv.org6 days agoView details
Memory Compression for High-Fanout Agent Sandboxes
arXiv:2609.11294v1 Announce Type: new Abstract: High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a shared template and execute related trajectories, exposing substantial tem…
arxiv.org6 days agoView details
Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models
arXiv:2609.11315v1 Announce Type: new Abstract: Diffusion vision-language models generate answers through iterative refinement, exposing intermediate answer trajectories that can be inspected and controlled at inference time. However, this controllability creates a reasoning-need mismatch, where a universal generation…
arxiv.org6 days agoView details
arXiv:2609.11403v1 Announce Type: new Abstract: Cultural-heritage KGs such as the NFDI4Culture-KG contain millions of triples about artworks, music, inscriptions, historical events, and the people and places connected to them. For many users, however, discovering this knowledge can be difficult. While SPARQL can be le…
arxiv.org6 days agoView details
arXiv:2609.11431v1 Announce Type: new Abstract: Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data. In addition to predictive accuracy, models are appreciated for their potential to provide interpretabili…
arxiv.org6 days agoView details
The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation
arXiv:2609.11489v1 Announce Type: new Abstract: Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions---shared protocols for reading meaning beyond the literal message---which AI-AI benchmarks may not capture. We propose the \emph{convention gap}, the difference be…
arxiv.org6 days agoView details
Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems
arXiv:2609.11532v1 Announce Type: new Abstract: Commercial text-to-image systems silently revise user prompts before generating images, a step users typically cannot disable or even see. Yet, existing audits of cultural bias examine only the final images and treat generation as a single pipeline, so they cannot tell w…
arxiv.org6 days agoView details
Models
View all models →diffusion-single-file · comfyui · base_model:m-a-p/SheetSage2
huggingface.co6 days ago153 ptsView details
transformers · safetensors · aliceai_t5_moe
huggingface.co6 days ago114 ptsView details
text-generation · transformers · safetensors · diffusion_gemma
huggingface.co7 days ago116 ptsView details
mistralai/Mistral-Large-3-675B-Instruct-2512-NVFP4
vllm · mistral-common · compressed-tensors
huggingface.co7 days ago63 ptsView details
cHunter789/Qwen3.8-27B-i1-IQ4_KS_KT-GGUF
image-text-to-text · gguf · image-text-to-text · qwen
huggingface.co7 days ago41 ptsView details
LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF
text-generation · gguf · llama.cpp · qwen3.8
huggingface.co7 days ago6 ptsView details
pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder
text-generation · transformers · safetensors · qwen3_5_moe
huggingface.co7 days ago4 ptsView details
software-mansion/react-native-executorch-all-MiniLM-L6-v2
sentence-similarity · executorch · sentence-similarity · license:apache-2.0
huggingface.co7 days ago2 ptsView details
AMAImedia/Hy4-preview-BF16-GGUF
image-text-to-text · transformers · safetensors · gguf
huggingface.co7 days ago1 ptsView details
win10/RWKV7-Ling-MoE-21.58B-A9.69B
safetensors · rwkv7_moe · custom_code
huggingface.co7 days ago1 ptsView details
software-mansion/react-native-executorch-multi-qa-MiniLM-L6-cos-v1
sentence-similarity · executorch · sentence-similarity · license:apache-2.0
huggingface.co7 days agoView details
Open source
View all open source →langchain-ai/langchain langchain-core==1.6.3
Changes since langchain-core==1.6.2 release(core): 1.6.3 (#40407) feat(core): Allow model name and provider tracing metadata override based on gateway response (#40406) test(core): cover the deprecated `.text()` access path (#40243) docs(core): remove stale Args/Raises entries f…
github.com6 days agoView details
<details open> vulkan: fix data race and OOB access in argsort(large) (#28705) argsort had a data race in the inner loop, which VVL caught. But I don't think this was causing failures in practice. argsort_large has OOB accesses which might explain the failures in CI, but I could…
github.com7 days agoView details
<details open> opencl: add A8 Q4_0 mm binary kernel support (#28268) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/46774832> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.c…
github.com7 days agoView details