Skip to content

Archive / 2026-09-11

September 11, 2026

  1. Cherenkov Radiation

    iaea.org7 days ago234 ptsView detailsJoin discussion

  2. Feeling Sad about AI

    artificialworlds.net6 days ago177 ptsView detailsJoin discussion

  3. A Misalignment of AI in Mathematics

    terrytao.wordpress.com6 days ago148 ptsView detailsJoin discussion

  4. Resist "AI"

    ronjeffries.com6 days ago65 ptsView detailsJoin discussion

  5. Why So Many AI Researchers Think the Machines Could Kill Everyone

    A combination of rapid advances, recursive self-improvement, and agentic swarms are genuinely “spooking people” inside big labs.

    wired.com7 days ago27 ptsView details

  6. AI made 16 new viruses

    morningbrew.com6 days ago13 ptsView detailsJoin discussion

  7. Primary source

    Cognition helps Devin test its own work with GPT‑6 Astra

    GPT‑6 Astra improves Devin’s ability to test software and show that it works, with the goal of helping engineers review less code and ship more.

    openai.com6 days agoView details

  8. Primary source

    Rapidly scaling online storage to serve over 1 billion ChatGPT users

    Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second.

    openai.com6 days agoView details

  9. Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

    ByteDance Seed, SUTD, Georgia Tech, M-A-P, and TokenWave.AI introduce HarnessDev, a benchmark that scores the runnable harness a model builds rather than the answer it returns. Starting from a seed that scores 0, 6 creator LLMs construct harnesses across 5 benchmarks and 2,207 tasks, then evolve them from execution fe…

    marktechpost.com6 days agoView details

  10. Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

    Anthropic has published a new plugin evals workflow for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compares the result with a run where the plugin is not loaded. It answers 3 questions plugin developers could not previously measure: does the sk…

    marktechpost.com6 days agoView details

  11. Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

    Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its 218B parameters per token and scores 83.6 on Cohere's WMT26 evaluation. Weights are free for non-commercial use, with commercial access through Cohere Model Vault or…

    marktechpost.com7 days agoView details

  12. Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

    Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M tokens. Fugu Ultra v2 targets peak capability, scoring 48.3 on Chartography and 74.3 on DeepSWE. The po…

    marktechpost.com7 days agoView details

  13. Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

    Google Research has released ToolGrad, an ACL 2026 Findings framework that inverts tool-use dataset generation: it builds a verified API chain first, then writes the matching user query. Guided by textual "gradients" from a 4-module propose-execute-select-update loop, ToolGrad reaches a 99.8% pass rate on ToolBench ve…

    marktechpost.com7 days agoView details

  14. Lawyer fined $5K over AI-hallucinated witnesses in a murder case

    New Mexico's Supreme Court is punishing a lawyer for including AI-fabricated witnesses and fake police testimony in an appeal for his client's murder conviction, according to a report from Reuters. In a filing on Wednesday, the court fined Stephen Aarons $5,000 and held him in contempt for failing to "verify the factu…

    theverge.com6 days agoView details

  15. Meta Sued Over Training Data for Its AI and Face-Recognition Systems

    The proposed class action alleges Meta illegally harvested people’s Facebook and Instagram photos to train its AI image-generation models and to build its unreleased “NameTag” face recognition feature.

    wired.com6 days agoView details

  16. Anthropic spent this week in hot water over cybersecurity

    After admitting earlier this year that its AI models had hacked other companies' systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models' single-minded "recklessness" - and will likely fuel alread…

    theverge.com6 days agoView details

  17. One of AI’s Fiercest Critics Says All the Doom Talk Is ‘Meant to Distract Us’

    Timnit Gebru argues that AI companies are stoking fears of extinction to avoid discussing actual harms, like autonomous weapons.

    wired.com6 days agoView details

  18. Meta says it’s changing AI suggestions after posing invasive personal questions

    Meta says it's making changes to the prompts suggested by its AI chatbot after a viral video showed it digging for personal information about a woman's young daughters, as reported earlier by Futurism. In a statement to The Verge, Meta spokesperson Dina El-Kassaby says the company "missed the mark," adding that "the f…

    theverge.com6 days agoView details

  19. RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection

    arXiv:2609.11372v1 Announce Type: new Abstract: Auditory attention decoding (AAD) identifies the attended speaker from physiological signals, supporting neuro-steered hearing devices and natural human-machine interaction. Electroencephalography (EEG) is the dominant modality for AAD but provides incomplete evidence in…

    arxiv.org6 days agoView details

  20. Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

    arXiv:2609.10724v1 Announce Type: new Abstract: Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disrupt…

    arxiv.org6 days agoView details

  21. SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

    arXiv:2609.11180v1 Announce Type: new Abstract: Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or >=2.0,1.2 means >=1.3.0) traps every model on Cargo (near 60%), and although standard PEP 440 prefix matching is universal, on zero-pad/post-release corn…

    arxiv.org6 days agoView details

  22. AI-Powered Flare Combustion Efficiency Estimation

    arXiv:2609.11262v1 Announce Type: new Abstract: Achieving high combustion efficiency in flare stacks is crucial for adhering to regulatory standards and controlling the release of hydrocarbons into the environment. Traditional instruments like gas analyzers and hyperspectral cameras are expensive, fragile, and require…

    arxiv.org6 days agoView details

  23. Predicting Train Delays in Finland Using Machine Learning and Weather Data

    arXiv:2609.11277v1 Announce Type: new Abstract: Reliable railway operations depend increasingly on real-time environmental intelligence delivered through wireless sensor infrastructures, a capability that 6G networks will substantially enhance through integrated sensing and edge computing. Adverse weather, particularl…

    arxiv.org6 days agoView details

  24. Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1

    arXiv:2609.11281v1 Announce Type: new Abstract: Learning and decision-making in animals are often modeled as Bayesian processes, where sensory evidence is integrated with prior beliefs to guide behavior in the face of uncertainty. But what are the inherent neural dynamics that give rise to this ability, and how could…

    arxiv.org6 days agoView details

  25. When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting

    arXiv:2609.11282v1 Announce Type: new Abstract: Multimodal forecasting models that combine time series with text annotations promise richer prediction through textual context, but how do we know whether a text annotation meaningfully contributes to the forecasters prediction? This is an information-theoretic question,…

    arxiv.org6 days agoView details

  26. Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data

    arXiv:2609.11286v1 Announce Type: new Abstract: Synthetic relational data is normally produced by a model trained on a real dataset, and its quality is measured as the distance to that dataset. This paper describes a generator that has no real dataset at either end. Given an industry, a company size, a business model,…

    arxiv.org6 days agoView details

  27. Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints

    arXiv:2609.11490v1 Announce Type: new Abstract: An unlearning audit reads its verdict off numbers that an unlearned model and its retrained reference each publish, and both also ship batch-normalization statistics that no gradient step wrote and no release records. Refitting them on kept data at bit-identical weights…

    arxiv.org6 days agoView details

  28. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

    arXiv:2609.11115v1 Announce Type: new Abstract: Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and se…

    arxiv.org6 days agoView details

  29. AI Exposure and AI Resilience: A Two-Dimensional Assessment Framework for Software and Software-Based Business Model

    arXiv:2609.11321v1 Announce Type: new Abstract: Artificial intelligence is changing both software production and the economics of software-based business models. Classical technology due diligence mainly examines technical properties such as architecture, scalability, and technical debt. These criteria do not fully ca…

    arxiv.org6 days agoView details

  30. Probabilistic Focal Search: Accelerating Bounded-Suboptimal Search via Lower-Bound Advancement

    arXiv:2609.10584v1 Announce Type: new Abstract: Bounded-suboptimal search seeks a solution within a factor $w$ of optimal while reducing search effort. Focal Search (FS) uses heuristic guidance within FOCAL, the frontier nodes eligible under the threshold $w f_{\min}$, but its deterministic policy may leave $f_{\min}$…

    arxiv.org6 days agoView details

  31. Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

    arXiv:2609.11147v1 Announce Type: new Abstract: Unraveling reaction mechanisms is central to modern chemistry, yet automating these investigations remains challenging because computational workflows still rely heavily on expert intervention. Here we introduce ARCHE, an autonomous agentic system that integrates a gener…

    arxiv.org6 days agoView details

  32. An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning

    arXiv:2609.11199v1 Announce Type: new Abstract: With the existing digital mental health tools specifically developed for Western settings, Pakistani students are exposed to a uniquely compounded stress situation in their university that includes academic, financial, familial, and relational stressors, which have becom…

    arxiv.org6 days agoView details

  33. CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series Forecasting

    arXiv:2609.11206v1 Announce Type: new Abstract: Cryptocurrency forecasting presents a distinctive combination of extreme cross-asset scale heterogeneity, non-stationary dynamics, and structural dependencies among Open, High, Low, and Close (OHLC) variables. We present CryptoL, a unified framework designed to address t…

    arxiv.org6 days agoView details

  34. Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model

    arXiv:2609.11291v1 Announce Type: new Abstract: We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where the benchmark-correct answer is…

    arxiv.org6 days agoView details

  35. Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

    arXiv:2609.11318v1 Announce Type: new Abstract: Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy resea…

    arxiv.org6 days agoView details

  36. Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

    arXiv:2609.11365v1 Announce Type: new Abstract: In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whe…

    arxiv.org6 days agoView details

  37. Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration

    arXiv:2609.11446v1 Announce Type: new Abstract: Heterogeneous model collaboration seeks to exploit the complementary strengths of different models to balance predictive performance and inference cost. Existing approaches typically rely either on trained routers, which tie routing decisions to a fixed task and model po…

    arxiv.org6 days agoView details

  38. Lightweight LiDAR-Based Cone Detection Framework Using Random Forest for Formula Student Driverless

    arXiv:2609.11527v1 Announce Type: new Abstract: Reliable, low-latency perception is crucial for Formula Student Driverless vehicles, yet many existing pipelines rely on deep learning and multi-sensor fusion, often requiring GPU acceleration. This paper presents a lightweight LiDAR-only perception pipeline tailored for…

    arxiv.org6 days agoView details

  39. Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

    arXiv:2609.11243v1 Announce Type: new Abstract: Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evide…

    arxiv.org6 days agoView details

  40. ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

    arXiv:2609.11498v1 Announce Type: new Abstract: Practical uncertainty quantification (UQ) for large language models must decide, from a single generation, whether a specific answer should be trusted. Existing methods either sample multiple generations, read only output-token probabilities, or reduce the model's intern…

    arxiv.org6 days agoView details

  41. Towards a Deterministic Math Solver for Clinical Language Models

    arXiv:2609.10728v1 Announce Type: new Abstract: Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative:…

    arxiv.org6 days agoView details

  42. Flexible and Interpretable Accent Distance Measurements

    arXiv:2609.11458v1 Announce Type: new Abstract: Determining the differences between two speakers' accents is a fundamental task in linguistics and speech technology research. The methodology used to measure these differences depends on the specific research area. A phonetics researcher may demonstrate accent variation…

    arxiv.org6 days agoView details

  43. Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

    arXiv:2609.11185v1 Announce Type: new Abstract: Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logi…

    arxiv.org6 days agoView details

  44. A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies

    arXiv:2609.11231v1 Announce Type: new Abstract: This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs). The system achieves natural language understanding, device control, intraoperative recording, and surgical report generation thr…

    arxiv.org6 days agoView details

  45. Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification

    arXiv:2609.11319v1 Announce Type: new Abstract: Most of mathematical knowledge has been communicated through so-called informal use of mathematics and natural language. With large language models (LLMs) being highly adept in using natural language, they achieve strong performance, yet not perfect, in informal mathemat…

    arxiv.org6 days agoView details

  46. Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

    arXiv:2609.11341v1 Announce Type: new Abstract: Multimodal brain state decoding has largely focused on fusing paired modalities for prediction, but has rarely explored how their correspondence can be further exploited to enrich training data and improve multimodal representation learning. To address this gap, we propo…

    arxiv.org6 days agoView details

  47. From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

    arXiv:2609.11493v1 Announce Type: new Abstract: Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of technical information across a multi-stage, knowledge-intensive continuum from drug discovery to commercial manufacturing. This knowledge is traditionally fragmented across func…

    arxiv.org6 days agoView details

  48. Extending SMT Solving with Non-Ground Clause Learning

    arXiv:2609.11509v1 Announce Type: new Abstract: Quantifier instantiation is currently the main approach to non-ground SMT solving: solvers generate ground instances and solve the resulting ground SMT problems with CDCL(T)-style reasoning. When a conflict is found, conflict analysis learns only a ground clause, even th…

    arxiv.org6 days agoView details

  49. KuaiRP Series Role-playing Models Technical Report

    arXiv:2609.11127v1 Announce Type: new Abstract: This paper introduces the complete technical solution for the KuaiRP series of role-playing models. We aim to achieve four core objectives for a dedicated role-playing model: simplified prompt engineering, highly stable output quality, built-in domain world knowledge, an…

    arxiv.org6 days agoView details

  50. Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

    arXiv:2609.11144v1 Announce Type: new Abstract: Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of se…

    arxiv.org6 days agoView details

  51. Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning

    arXiv:2609.11393v1 Announce Type: new Abstract: Test-time adaptation has emerged as a lightweight alternative to costly post-training for improving the reasoning capabilities of Large Language Models (LLMs) on downstream tasks. Predictive entropy provides a model-derived signal for such adaptation, guiding models towa…

    arxiv.org6 days agoView details

  52. Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

    arXiv:2609.10657v1 Announce Type: new Abstract: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter s…

    arxiv.org6 days agoView details

  53. Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

    arXiv:2609.11060v1 Announce Type: new Abstract: Persistent memory is entering production-oriented agent platforms to help long-horizon agents accumulate experience across sessions. Yet a post-task curator agent restricted to completed trajectories can preserve errors, overgeneralize partial evidence, or retain stale k…

    arxiv.org6 days agoView details

  54. RouteRepair: Instance-Level Failure Diagnosis and Targeted Repair in LLM-Based Automated Heuristic Design for Routing Optimization

    arXiv:2609.11452v1 Announce Type: new Abstract: Efficient routing optimization is essential to freight transportation, urban logistics, and shared mobility, where high-quality heuristics are often required under limited computational budgets. Recent large language model (LLM)-based automated heuristic design methods c…

    arxiv.org6 days agoView details

  55. Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language

    arXiv:2609.10629v1 Announce Type: new Abstract: Quadratic Unconstrained Binary Optimization (QUBO) is a central formulation for combinatorial optimization and has gained increasing attention due to its compatibility with quantum, hybrid quantum-classical, and quantum-inspired solvers. However, translating natural-lang…

    arxiv.org6 days agoView details

  56. A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning

    arXiv:2609.10654v1 Announce Type: new Abstract: The Abstraction and Reasoning Corpus (ARC) benchmarks cognitive generalization, the ability to infer and apply abstract rules from limited examples. This paper presents a multi-stage rule-chaining framework that performs compositional reasoning across symbolic, structura…

    arxiv.org6 days agoView details

  57. Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning

    arXiv:2609.10656v1 Announce Type: new Abstract: Selecting LoRA rank for diffusion fine-tuning requires balancing quality and compute cost. We present a controlled study on CIFAR-10 using a DDPM U-Net with ranks {2,4,8,16,32}, fixed optimization settings, and a reproducible local-folder pytorch-fid protocol. We report…

    arxiv.org6 days agoView details

  58. An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

    arXiv:2609.10712v1 Announce Type: new Abstract: We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evalua…

    arxiv.org6 days agoView details

  59. Studying Without a Syllabus: Task-Agnostic Environment Preprocessing

    arXiv:2609.10824v1 Announce Type: new Abstract: Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural guidance. Most automated adaptation methods, however, rely on task examples, trajectories, or evalu…

    arxiv.org6 days agoView details

  60. When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents

    arXiv:2609.10873v1 Announce Type: new Abstract: Independent evaluation can reject harmful policy updates yet also prevent useful continual learning. We argue that update admission must be assessed through both error control and retained learning opportunities at a stated interaction budget. We identify a concrete fail…

    arxiv.org6 days agoView details

  61. Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows

    arXiv:2609.10964v1 Announce Type: new Abstract: Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but also on when ready turns are released. Most runtimes release each turn immediately upon readiness. Un…

    arxiv.org6 days agoView details

  62. Demystifying the Privacy-Utility Trade-off in LLM Interactions

    arXiv:2609.10992v1 Announce Type: new Abstract: The integration of Large Language Models into daily tasks relies on context-rich instructions, inevitably exposing sensitive user information. Current privacy-preserving methods typically employ context-agnostic static rules, causing severe utility degradation. However,…

    arxiv.org6 days agoView details

  63. Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks

    arXiv:2609.11018v1 Announce Type: new Abstract: The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction…

    arxiv.org6 days agoView details

  64. The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures

    arXiv:2609.11030v1 Announce Type: new Abstract: AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We present the Agent Incident Registry (AIR), a source-linked catalog cont…

    arxiv.org6 days agoView details

  65. Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

    arXiv:2609.11061v1 Announce Type: new Abstract: Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically all…

    arxiv.org6 days agoView details

  66. MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

    arXiv:2609.11065v1 Announce Type: new Abstract: Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, compa…

    arxiv.org6 days agoView details

  67. The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems

    arXiv:2609.11146v1 Announce Type: new Abstract: AI-generated text is flowing back into the training corpora of the next generation of models. Recursive training on it drives model collapse, and recent work extends the setting to many models feeding one another -- but almost always with the market split evenly, while r…

    arxiv.org6 days agoView details

  68. DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat

    arXiv:2609.11155v1 Announce Type: new Abstract: Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous systems and air combat. While MARL has demonstrated significant potential in air combat, achieving sophisticated tactical coordination remains a non-triv…

    arxiv.org6 days agoView details

  69. Breaking Predictions Is Not Enough: Specified-Foil Counterfactuals for Temporal Graphs

    arXiv:2609.11170v1 Announce Type: new Abstract: Temporal graph counterfactual explanations typically change past events to change or invalidate an original prediction, while leaving its replacement unspecified. Yet a user facing a predicted outcome often asks which past conditions would make a particular alternative o…

    arxiv.org6 days agoView details

  70. Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation

    arXiv:2609.11176v1 Announce Type: new Abstract: Industrial query-to-agent matching fails when topical relevance is mistaken for executable capability, especially on long-tail and boundary-sensitive requests. We formulate annotation as \emph{capability-bound process supervision} and instantiate it with Debate-to-Skill,…

    arxiv.org6 days agoView details

  71. Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce

    arXiv:2609.11190v1 Announce Type: new Abstract: AI shopping assistants increasingly redirect consumer discovery, creating an urgent need for tools that support seller-side competitive decision-making. We present a multi-agent AI system that automates competitive visibility measurement and root cause diagnosis in LLM-m…

    arxiv.org6 days agoView details

  72. NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

    arXiv:2609.11234v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whe…

    arxiv.org6 days agoView details

  73. Memory Compression for High-Fanout Agent Sandboxes

    arXiv:2609.11294v1 Announce Type: new Abstract: High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a shared template and execute related trajectories, exposing substantial tem…

    arxiv.org6 days agoView details

  74. Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

    arXiv:2609.11315v1 Announce Type: new Abstract: Diffusion vision-language models generate answers through iterative refinement, exposing intermediate answer trajectories that can be inspected and controlled at inference time. However, this controllability creates a reasoning-need mismatch, where a universal generation…

    arxiv.org6 days agoView details

  75. From Queries to Narratives: Cultural Heritage Data Stories for Knowledge Graph Exploration and Quality Assessment

    arXiv:2609.11403v1 Announce Type: new Abstract: Cultural-heritage KGs such as the NFDI4Culture-KG contain millions of triples about artworks, music, inscriptions, historical events, and the people and places connected to them. For many users, however, discovering this knowledge can be difficult. While SPARQL can be le…

    arxiv.org6 days agoView details

  76. LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study

    arXiv:2609.11431v1 Announce Type: new Abstract: Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data. In addition to predictive accuracy, models are appreciated for their potential to provide interpretabili…

    arxiv.org6 days agoView details

  77. The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

    arXiv:2609.11489v1 Announce Type: new Abstract: Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions---shared protocols for reading meaning beyond the literal message---which AI-AI benchmarks may not capture. We propose the \emph{convention gap}, the difference be…

    arxiv.org6 days agoView details

  78. Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems

    arXiv:2609.11532v1 Announce Type: new Abstract: Commercial text-to-image systems silently revise user prompts before generating images, a step users typically cannot disable or even see. Yet, existing audits of cultural bias examine only the final images and treat generation as a single pipeline, so they cannot tell w…

    arxiv.org6 days agoView details

  1. Comfy-Org/YuE2

    diffusion-single-file · comfyui · base_model:m-a-p/SheetSage2

    huggingface.co6 days ago153 ptsView details

  2. yandex/AliceAI-T5-35B-A0.6B

    transformers · safetensors · aliceai_t5_moe

    huggingface.co6 days ago114 ptsView details

  3. thesysdev/OUI-1

    text-generation · transformers · safetensors · diffusion_gemma

    huggingface.co7 days ago116 ptsView details

  4. mistralai/Mistral-Large-3-675B-Instruct-2512-NVFP4

    vllm · mistral-common · compressed-tensors

    huggingface.co7 days ago63 ptsView details

  5. cHunter789/Qwen3.8-27B-i1-IQ4_KS_KT-GGUF

    image-text-to-text · gguf · image-text-to-text · qwen

    huggingface.co7 days ago41 ptsView details

  6. LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF

    text-generation · gguf · llama.cpp · qwen3.8

    huggingface.co7 days ago6 ptsView details

  7. pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder

    text-generation · transformers · safetensors · qwen3_5_moe

    huggingface.co7 days ago4 ptsView details

  8. software-mansion/react-native-executorch-all-MiniLM-L6-v2

    sentence-similarity · executorch · sentence-similarity · license:apache-2.0

    huggingface.co7 days ago2 ptsView details

  9. AMAImedia/Hy4-preview-BF16-GGUF

    image-text-to-text · transformers · safetensors · gguf

    huggingface.co7 days ago1 ptsView details

  10. win10/RWKV7-Ling-MoE-21.58B-A9.69B

    safetensors · rwkv7_moe · custom_code

    huggingface.co7 days ago1 ptsView details

  11. software-mansion/react-native-executorch-multi-qa-MiniLM-L6-cos-v1

    sentence-similarity · executorch · sentence-similarity · license:apache-2.0

    huggingface.co7 days agoView details

  1. langchain-ai/langchain langchain-core==1.6.3

    Changes since langchain-core==1.6.2 release(core): 1.6.3 (#40407) feat(core): Allow model name and provider tracing metadata override based on gateway response (#40406) test(core): cover the deprecated `.text()` access path (#40243) docs(core): remove stale Args/Raises entries f…

    github.com6 days agoView details

  2. ggml-org/llama.cpp b10903

    <details open> vulkan: fix data race and OOB access in argsort(large) (#28705) argsort had a data race in the inner loop, which VVL caught. But I don't think this was causing failures in practice. argsort_large has OOB accesses which might explain the failures in CI, but I could…

    github.com7 days agoView details

  3. ggml-org/llama.cpp b10902

    <details open> opencl: add A8 Q4_0 mm binary kernel support (#28268) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/46774832> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.c…

    github.com7 days agoView details