Skip to content

Archive / 2026-09-08

September 8, 2026

  1. Mistral raises €3B

    mistral.ai10 days ago841 ptsView detailsJoin discussion

  2. Mercury 2.5

    inceptionlabs.ai9 days ago245 ptsView detailsJoin discussion

  3. AI Has a Discovery Problem

    mhacevedo.com9 days ago68 ptsView detailsJoin discussion

  4. GrapheneOS on AI Usage

    grapheneos.social9 days ago52 ptsView detailsJoin discussion

  5. Why AI-generated slopaganda works

    machinesociety.ai9 days ago18 ptsView detailsJoin discussion

  6. My Little AI Factory

    dominis.blog9 days ago14 ptsView detailsJoin discussion

  7. Major Math Breakthrough by AI

    mastodon.social10 days ago11 ptsView detailsJoin discussion

  8. Primary source

    How GPT-5.6 Sol helps run quantum computing experiments

    See how an MIT researcher uses GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits.

    openai.com9 days agoView details

  9. Primary source

    The Work Now Within Reach

    Explore how more capable, affordable AI can expand the work people and businesses can accomplish—and make growth more economical.

    openai.com9 days agoView details

  10. Primary source

    Introducing ChatGPT Images 2.5

    ChatGPT Images 2.5 helps turn your ideas, sketches, and reference photos into more personalized, polished images that better reflect your ideas.

    openai.com9 days agoView details

  11. Primary source

    On the Navier–Stokes Millennium Prize Problem

    We’re sharing an AI-generated solution to the Navier–Stokes Millennium Prize Problem, including a writeup and a formal proof in Lean.

    openai.com9 days agoView details

  12. Primary source

    Funding grants for new research into AI and teen development

    Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.

    openai.com10 days agoView details

  13. Primary source

    AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome

    AlphaGenome Atlas maps the molecular effects of 9 billion single-letter DNA variants across the human genome.

    deepmind.google9 days agoView details

  14. Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

    Today, Meta has introduced Muse, a personal AI agent that takes actions rather than just answering questions. Muse can send emails, book travel, negotiate bills, and pursue long term goals. It keeps working after you close the app and returns only when it needs approval. The bigger story for AI devs is architectural.…

    marktechpost.com9 days agoView details

  15. NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels

    NVIDIA has announced CUDA Rust, a push to make Rust a first-class language for GPU kernels. Two open-source NVlabs projects cover the two CUDA programming models: cuda-oxide compiles SIMT kernels from Rust MIR through Pliron and LLVM to PTX, while cutile-rs JIT-compiles Tile kernels through CUDA Tile IR on stable Rust…

    marktechpost.com9 days agoView details

  16. Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants

    DeepMind's AlphaGenome Atlas maps every single-letter change in the human genome with 1 impact score per variant. The post Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants appeared first on MarkTechPost.

    marktechpost.com9 days agoView details

  17. Drama swirls around OpenAI’s legendary mathematical milestone

    OpenAI says it found a solution to a major math problem that has remained unsolved for around 90 years, as reported earlier by The New York Times and Wired. In a blog post on Tuesday, OpenAI announced that it discovered a solution to the Navier-Stokes problem - which relates to the flow of liquid and gas - using an in…

    theverge.com9 days agoView details

  18. ChatGPT Sketch turns your bad drawings into detailed AI images

    I used ChatGPT and its Sketch tool to make this AI-generated image of a cat. OpenAI announced ChatGPT Images 2.5 on Tuesday and is adding a new way to tell ChatGPT what you want it to make an image of: by drawing a doodle. With a new feature called Sketch, you can just draw something right inside ChatGPT and then tell…

    theverge.com9 days agoView details

  19. Muse, Meta’s New Personal AI Agent, Needs You to Trust It

    Designed to compete with OpenClaw and Instinct, the company says Muse can do everything from sell your car to book you a plane ticket.

    wired.com9 days agoView details

  20. Meta bets on AI agent Muse to catch up in AI race

    Meta is making another push to bring artificial intelligence to the masses with Muse, a personal assistant it says can put AI in the hands of virtually anyone. The product is the latest step in a multi-billion-dollar strategy overhaul designed to revitalize the company's ailing position in the AI race and help it catc…

    theverge.com9 days agoView details

  21. AI power users claim Anthropic duped them with subscriptions, and they’re taking it to court

    Anthropic says power users are key to its business - it's prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they'd get more out of a top-tier pricing subscription than they did. In an expanded class actio…

    theverge.com9 days agoView details

  22. OpenAI Just Claimed a Huge Math Discovery. Some Academics Are Crying Foul

    A landmark announcement by the frontier AI lab has been overshadowed by accusations of impropriety.

    wired.com9 days agoView details

  23. Google’s Atlas of the human genome could pave the way for new treatments

    Google DeepMind has unveiled an AI tool that its scientists claim could help unravel the mysteries of the human genome and transform our understanding of biology, accelerating scientific research and ultimately paving the way for new treatments for diseases. The platform, called AlphaGenome Atlas, contains a "predicti…

    theverge.com9 days agoView details

  24. Adobe is trying to make its AI generators idiot-proof in Premiere

    Suspenseful clock ticking… as you wait for AI to take your job. | Image: Adobe Adobe is overhauling how editors interact with AI in its Premiere professional video editing software. Its new Generative Media tool makes it easier to generate video, sound effects, music, and soundscapes without ever leaving the project t…

    theverge.com9 days agoView details

  25. CriticGen: Generation-Aware Evaluation as Actionable Feedback

    arXiv:2609.05439v1 Announce Type: new Abstract: Current evaluation methods for large language models are coarse-grained and decoupled from generation, producing generic explanations that fail to provide actionable feedback for model improvement. We propose CriticGen, a fine-grained, generation-aware evaluation framewo…

    arxiv.org9 days agoView details

  26. Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Frozen,Heterogeneous LLM Agents

    arXiv:2609.06815v1 Announce Type: new Abstract: An open, networked web will allow agents to run frozen models from multiple vendors, keep their history private, and teach each other which tool to call and when. Flat text (prompts, example pools) makes it difficult for the protocol to distinguish between noise statisti…

    arxiv.org9 days agoView details

  27. Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems

    arXiv:2609.05928v1 Announce Type: new Abstract: Large language models now compute correct tax liabilities on over 90% of well-formed cases in statutory benchmarks, which makes them candidates for the tax-advisory and compliance systems that consume such an answer directly. Real legal inputs, however, are frequently de…

    arxiv.org9 days agoView details

  28. Some Tokens Behave like Magnets: Revealing Linguistic Organization in the Layers of Language Models

    arXiv:2609.05743v1 Announce Type: new Abstract: We identify a special group of token vectors inside large language models (LLMs), which we term magnetic vectors, that organize the surrounding tokens by either attracting or repelling them. Particularly, tokens pointing the same way as an attracting magnet are elongated…

    arxiv.org9 days agoView details

  29. AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

    arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than model weights. Each round begins from a fresh model session, and durable in…

    arxiv.org9 days agoView details

  30. Better Together: Complementary Query Rewriting Under a Strong RAG Baseline

    arXiv:2609.05637v1 Announce Type: new Abstract: A popular way to improve Retrieval-Augmented Generation (RAG) is to rewrite the user's question into several variants and search with all of them. We test whether this actually helps once the underlying search is already strong. Under one fixed, competitive pipeline (BGE…

    arxiv.org9 days agoView details

  31. TamilEOT: A Dataset and Model for Semantic End-of-Turn Detection in Tamil Telephone Speech

    arXiv:2609.05631v1 Announce Type: new Abstract: A voice agent has to decide, at every pause, whether the user has finished speaking. Without a model of the language that decision falls back to a fixed silence timeout: set it short and the agent interrupts, set it long and every turn pays the full wait. Open semantic e…

    arxiv.org9 days agoView details

  32. PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectories

    arXiv:2609.05488v1 Announce Type: new Abstract: Clinical deterioration unfolds through coupled, partially observed trajectories, not a single diagnostic label. We introduce PGP-Clinical-TimeKAN, a trajectory-first framework for joint probabilistic forecasting of multivariate physiology. It combines missingness-aware t…

    arxiv.org9 days agoView details

  33. Deep belief networks are exact

    arXiv:2609.05572v1 Announce Type: new Abstract: We prove that every strictly positive probability distribution on \(\{-1,1\}^n\) is represented exactly by a sigmoid belief network with finite parameters. This answers a question of Sutskever and Hinton. The proof upgrades their probability-sharing approximation to exac…

    arxiv.org9 days agoView details

  34. Discovering Translation-Worthy Languages with E-Values

    arXiv:2609.06593v1 Announce Type: new Abstract: Choosing when to translate multilingual documents is a central routing problem in text classification: translation can improve predictions for some languages while degrading others or adding unnecessary computation. Uniform translation and heuristic language tiers do not…

    arxiv.org9 days agoView details

  35. Exposing Weaknesses in Emotion Recognition in Conversations

    arXiv:2609.05806v1 Announce Type: new Abstract: Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies.…

    arxiv.org9 days agoView details

  36. Generator-Independent Runtime Assurance under Partial Observation

    arXiv:2609.06036v1 Announce Type: new Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtime verification gates. We ask when the closed-loop safety guarantee decouples from the generator. The prevailing per-cand…

    arxiv.org9 days agoView details

  37. You Are What You Read: Misalignment via In-Context Persona Induction

    arXiv:2609.06851v1 Announce Type: new Abstract: Broad misalignment has been produced by finetuning on narrow data, harmful or benign, and in context only by demonstrations of the undesirable behaviour itself. We show that benign data suffices in context, with no finetuning and no demonstration of harmful behaviour in…

    arxiv.org9 days agoView details

  38. Compiling VGDL into Causal Models

    arXiv:2609.05459v1 Announce Type: new Abstract: Reinforcement learning and large language models often struggle to accurately capture the causal mechanics of game environments. Standard reinforcement learning agents tend to rely on spurious correlations, while large language models are prone to hallucinating game rule…

    arxiv.org9 days agoView details

  39. A Rubric-Guided Large Language Model Solution for Opioid Use Disorder Computable Phenotyping

    arXiv:2609.05682v1 Announce Type: new Abstract: Opioid use disorder (OUD) remains a public health crisis in the United States, yet it is difficult to identify from electronic health records (EHRs) because missing diagnosis codes and supporting evidence are buried in clinical narratives. Accurate OUD identification is…

    arxiv.org9 days agoView details

  40. Intra-Prompt Parallel Decoding for Common-Context Question Answering

    arXiv:2609.05707v1 Announce Type: new Abstract: In common-context question answering (CCQA) tasks, multiple input questions share a common context to base their answers from. However, Large Language Models typically generate each answer using an independent prompt. While existing batching and caching techniques help i…

    arxiv.org9 days agoView details

  41. CrisisKD: Five-Stage Knowledge Distillation for Aspect-Level Sentiment and Emotion Analysis in Crisis Discourse

    arXiv:2609.05757v1 Announce Type: new Abstract: Identifying the target of emotional words or phrases in crisis situations, especially health-related ones, is important for understanding public concerns across cultural and linguistic contexts. We propose CrisisKD, a five-stage teacher--student knowledge distillation fr…

    arxiv.org9 days agoView details

  42. Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance

    arXiv:2609.05797v1 Announce Type: new Abstract: Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but a catch only prevents harm if the model would otherwise have complied. We measure the diffe…

    arxiv.org9 days agoView details

  43. Dynamic Lagging for Simultaneous Translation

    arXiv:2609.05799v1 Announce Type: new Abstract: In cascaded simultaneous speech translation, the machine translation (MT) system cannot control the read--write schedule of the upstream recognizer: it must decide, from a growing source prefix, how much target text to commit. We make a sentence-trained, decoder-only LLM…

    arxiv.org9 days agoView details

  44. Beyond Cross-Lingual Transfer: Benchmarking Propagation Boundaries in Multilingual LLM Unlearning

    arXiv:2609.05976v1 Announce Type: new Abstract: Large Language Model (LLM) unlearning aims to suppress target knowledge while preserving general capabilities. In multilingual settings, unlearning must additionally propagate within its intended linguistic scope. However, existing evaluations mainly measure cross-lingua…

    arxiv.org9 days agoView details

  45. CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

    arXiv:2609.05708v1 Announce Type: new Abstract: Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an individual model's confidence nor that of the aggregated answer measures reliability at the system level. We present CUSP (Collective Uncertainty through Semantic Opi…

    arxiv.org9 days agoView details

  46. Alignment by Stereotyping: How LLMs Sacrifice Individual Distinctiveness for Cultural Adaptation

    arXiv:2609.05993v1 Announce Type: new Abstract: Large language models are increasingly deployed for personalized interaction, and demographic conditioning via user profiles is a widely adopted strategy for cultural adaptation. We ask whether this approach genuinely serves individual users or achieves accuracy by erasi…

    arxiv.org9 days agoView details

  47. ModularPhaseNet: Finite-Cyclic Phase Geometry for Computable Semantic Hierarchy, Direction, and Context Consistency in Standard Transformers

    arXiv:2609.06000v1 Announce Type: new Abstract: We propose ModularPhaseNet, a classical and integer-computable discretization of the continuous complex phase geometry introduced in QuantumPhaseNet. The real-valued hidden states of a standard Transformer are retained, while only an auxiliary phase channel is quantized…

    arxiv.org9 days agoView details

  48. Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts

    arXiv:2609.06011v1 Announce Type: new Abstract: Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality: perceptual signals (e.g., a p…

    arxiv.org9 days agoView details

  49. Generating Adversarial Texts for Machine Translation via GRPO

    arXiv:2609.06048v1 Announce Type: new Abstract: As machine translation (MT) systems continue to improve, standard benchmarks become less informative for exposing remaining weaknesses. Traditional methods for creating challenging test sets rely on expensive manual creation or curation, while automated approaches strugg…

    arxiv.org9 days agoView details

  50. EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent

    arXiv:2609.05576v1 Announce Type: new Abstract: The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising path to optimize these agents, its…

    arxiv.org9 days agoView details

  51. What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

    arXiv:2609.05663v1 Announce Type: new Abstract: We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, F…

    arxiv.org9 days agoView details

  52. DPH Parser: A Bottom-Up Grammar-Driven Parser for Joint Constituency and Dependency Analysis

    arXiv:2609.06070v1 Announce Type: new Abstract: This paper presents Dependency-Phrase Hierarchy Parser (DPH Parser), a grammar-driven bottom-up unsupervized parsing framework inspired by Generalized Phrase Structure Grammar (GPSG) and Head-driven Phrase Structure Grammar (HPSG). The parser incrementally constructs con…

    arxiv.org9 days agoView details

  53. From Two Passes to One: Compact and Efficient Target-Stance Extraction

    arXiv:2609.06108v1 Announce Type: new Abstract: Target-Stance Extraction (TSE) is the task of predicting both the target (or topic) of an author's writing and the author's stance toward it. Existing approaches to TSE use a sequential pipeline of two separate neural models: one to identify the target and another to det…

    arxiv.org9 days agoView details

  54. STQA: A Benchmark for Stock-Focused Tabular Question Answering over Historical and Forecasted Data

    arXiv:2609.06117v1 Announce Type: new Abstract: Stock market analysis inherently requires composite reasoning over historical records and future projections, yet existing benchmarks remain fragmented across isolated tasks. We introduce STQA (Stock-focused Tabular Question Answering), an end-to-end benchmark designed t…

    arxiv.org9 days agoView details

  55. Customer Relationship Intelligence: Integrating CRM and MDM for Enhanced Customer Engagement

    arXiv:2609.06189v1 Announce Type: new Abstract: This study examines how Customer Relationship Management (CRM), Master Data Management (MDM), and Customer Knowledge Management (CKM) jointly constitute a Customer Relationship Intelligence (CRI) framework for enhanced Customer Engagement (CE). A cross-sectional survey o…

    arxiv.org9 days agoView details

  56. Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

    arXiv:2609.06263v1 Announce Type: new Abstract: Moderation APIs are built to flag policy-violating content, not to measure graded clinical risk. But a platform's duty does not end at detection: the response owed to passive distress differs sharply from the response owed to active planning with means access, and emergi…

    arxiv.org9 days agoView details

  57. Correction as Annotation: Bootstrapping a Dependency Parser for Documentary Medieval Latin

    arXiv:2609.06266v1 Announce Type: new Abstract: Medieval documentary sources remain inadequately served by existing natural language processing tools. None of the five readily available Latin treebank models attains usable performance on a collection of 160 inventories compiled in Marseille between 1258 and 1446. The…

    arxiv.org9 days agoView details

  58. AutoKD: Autonomous Knowledge Discovery

    arXiv:2609.06366v1 Announce Type: new Abstract: Scientific discovery in data-rich domains is currently constrained by human bandwidth: the growth in the volume and complexity of real-world data far outpaces the rate at which researchers can read, reason, and synthesize. Recent LLM-based multi-agent systems have begun…

    arxiv.org9 days agoView details

  59. A Grapheme-Aware Indic Tokenizer for Tamil: Large-Scale Training and Intrinsic Evaluation

    arXiv:2609.06690v1 Announce Type: new Abstract: Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text into discrete units that neural language models can process. The effectiveness of this process directly influences vocabulary efficiency, sequence length, compu…

    arxiv.org9 days agoView details

  60. SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use

    arXiv:2609.06124v1 Announce Type: new Abstract: High-quality multi-turn tool-use data is essential for training agentic models, yet existing data synthesis methods often underrepresent the argument-level dependencies that are critical to long-horizon tool use. As a result, even when a model selects the correct tool, t…

    arxiv.org9 days agoView details

  61. AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

    arXiv:2609.05837v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training them for specific environments remains fundamentally difficult: real-world applications provide no pre-defined tasks or verifiers, no faithful simulators, and limite…

    arxiv.org9 days agoView details

  62. What if LLMs Ate Their Words: Causal History Effects in Multi-Turn Interaction

    arXiv:2609.05882v1 Announce Type: new Abstract: Multi-turn interaction creates a feedback process in which an LLM's previous responses become context for later behavior. Prior work shows substantial multi-turn degradation and that assistant-generated history can affect later behavior. However, it remains unclear how t…

    arxiv.org9 days agoView details

  63. SinoGlyphBench: A Diagnostic Benchmark for Chinese Glyph-Level Obfuscation in Language-Model Moderation

    arXiv:2609.05843v1 Announce Type: new Abstract: Glyph-level obfuscation can leave harmful Chinese content readable to humans while degrading automated moderation. We introduce SinoGlyphBench, a diagnostic benchmark that identifies label-critical semantic anchors and creates matched original and glyph-obfuscated inputs…

    arxiv.org9 days agoView details

  64. Steering Geometry: Validating Human Value Geometry in LLM Steering Space

    arXiv:2609.06289v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically valid…

    arxiv.org9 days agoView details

  65. LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies

    arXiv:2609.06079v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies leverage pretrained vision-language models (VLMs) to guide action generation for robot control. VLMs provide hierarchical visual-semantic representations that evolve across layers, from local visual geometry to abstract, language-ali…

    arxiv.org9 days agoView details

  66. The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble

    arXiv:2609.05894v1 Announce Type: new Abstract: The exponentiation of Artificial intelligence (AI) in the recent past has entered a transformative era that has been driven by the growth in large language models (LLMs), large-scale compute infrastructures, and autonomous reasoning systems. However, the rapid accelerati…

    arxiv.org9 days agoView details

  67. Decomposing LLM-Judge Uncertainty to Target Expert Labels

    arXiv:2609.06444v1 Announce Type: new Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which…

    arxiv.org9 days agoView details

  68. DFlow: Enabling Verifier Information Flow in Block Diffusion Speculative Decoding

    arXiv:2609.06498v1 Announce Type: new Abstract: Block diffusion speculative decoding improves LLM inference efficiency by proposing a block of future tokens in parallel and verifying them with a single forward pass through the target model. However, existing methods retain only the accepted prefix and discard the reje…

    arxiv.org9 days agoView details

  69. Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring

    arXiv:2609.06315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used or proposed for educational scoring, but single-model and single-run evaluations provide limited evidence for assessment use. Short-answer scoring requires evidence about reliability, validity, severity, diagnostic value…

    arxiv.org9 days agoView details

  70. ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language

    arXiv:2609.06527v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential for translating natural-language (NL) requirements into PL/SQL programs, attracting increasing attention from the database community. However, existing NL-to-PL/SQL efforts primarily focus on directly generating PL…

    arxiv.org9 days agoView details

  71. UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms

    arXiv:2609.05910v1 Announce Type: new Abstract: Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack…

    arxiv.org9 days agoView details

  72. EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph

    arXiv:2609.05553v1 Announce Type: new Abstract: Agent memory allows LLM agents to use earlier interactions when answering new queries. Existing methods often compress interaction histories into summaries or other LLM-generated representations. Repeated generation adds cost and can discard answer-bearing details before…

    arxiv.org9 days agoView details

  73. SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives

    arXiv:2609.06611v1 Announce Type: new Abstract: Assessing the literary quality of narratives requires evaluating interpretive dimensions (cultural representation, emotional depth, and philosophical engagement) that existing NLG metrics cannot measure. We introduce SAGE, a six-layer evaluation framework that separates…

    arxiv.org9 days agoView details

  74. A Ticket from Marginals to Joints: Coupled-Noise Distillation for One-Step Block Generation in Diffusion Language Models

    arXiv:2609.06324v1 Announce Type: new Abstract: Autoregressive language models commit one token per forward pass; diffusion language models commit a block of tokens over several steps. We ask whether a block can be committed in a single forward pass. We study this with a noise-conditioned masked denoiser: a data-indep…

    arxiv.org9 days agoView details

  75. Cross-Lingual Representation Alignment by Token-Level Optimal Transport in a Language-Agnostic Space

    arXiv:2609.06381v1 Announce Type: new Abstract: Cross-lingual alignment (CLA) aims to align the representations of large language models (LLMs) across languages, enabling cross-lingual transfer to improve multilingual capabilities. Previous CLA methods often ignore language-specific information encoded in representati…

    arxiv.org9 days agoView details

  76. Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus

    arXiv:2609.06634v1 Announce Type: new Abstract: LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specific domains and language pairs that they are trained upon. To examine if Multiword Expressions (MWEs) still set a bottleneck for LLMs regarding language understand…

    arxiv.org9 days agoView details

  77. Hierarchical Wasserstein Merging for Multi-Domain Multi-Task Learning: From Specialists to a Generalist

    arXiv:2609.06406v1 Announce Type: new Abstract: Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across heterogeneous domains and tasks. However, joint training often suffers from interference under distribution shifts. Existing model merging methods mostly operate o…

    arxiv.org9 days agoView details

  78. InsightChain: Optimized Chain-of-Insight Analytics for LLM-driven Data Visualization

    arXiv:2609.06438v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for automated data visualization, yet existing approaches often frame visualization generation as a single-step mapping from user query to figure or code, overlooking the iterative analytical reasoning process of expert…

    arxiv.org9 days agoView details

  79. LLMs Mirror Country-Specific Gender Patterns If Asked, but Skew Male When Generating Media in Local Languages

    arXiv:2609.06545v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-based formats rather than long-form generation, and surveyed baselines for local gender associ…

    arxiv.org9 days agoView details

  80. ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics

    arXiv:2609.06663v1 Announce Type: new Abstract: Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by the memory and computational overhead of KV cache storage. Recent KV cache eviction approaches incorporate a cosine similarity-based diversity metric with impor…

    arxiv.org9 days agoView details

  81. PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents

    arXiv:2609.06702v1 Announce Type: new Abstract: Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to…

    arxiv.org9 days agoView details

  82. DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents

    arXiv:2609.06703v1 Announce Type: new Abstract: High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-graine…

    arxiv.org9 days agoView details

  83. AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths

    arXiv:2609.06771v1 Announce Type: new Abstract: Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism analysis, account linking, misinformation investigation, and machine-generated text detection. Yet current authorship benchmarks remain fragmented, usually covering…

    arxiv.org9 days agoView details

  84. Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment

    arXiv:2609.05512v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression methods apply uniform quantization across all components, risking damage to critical reasoning circuits. We present a reasoning-aware compression framework that bench…

    arxiv.org9 days agoView details

  85. The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies

    arXiv:2609.05514v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-gro…

    arxiv.org9 days agoView details

  86. Factors Influencing the Emergence of Dependency Length Minimization in Neural Agent Simulations

    arXiv:2609.06025v1 Announce Type: new Abstract: Given various grammatical options, language users prefer the word order choice that reduces the overall length of syntactic dependencies, a principle known as dependency length minimization (DLM). The origins of this preference remain an open question, particularly wheth…

    arxiv.org9 days agoView details

  87. Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

    arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will en…

    arxiv.org9 days agoView details

  88. When Agent Governance Helps

    arXiv:2609.05531v1 Announce Type: new Abstract: No specification says how a governed autotelic AI agent organization, where agents pursue self-generated goals inside guardrails, should be designed and evaluated. We answer in two parts. First, we synthesize the Governed Autotelic Multi-Agent Product Organization (GAMPO…

    arxiv.org9 days agoView details

  89. Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools

    arXiv:2609.05587v1 Announce Type: new Abstract: Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete diverse tasks with tools. These evaluations generally assume that tools return reliable information. However, tool returns in real-world systems can be plausible yet in…

    arxiv.org9 days agoView details

  90. From Narrative to Auditable Forecasts: A Structured Scaffold for Agentic Forecasting

    arXiv:2609.05905v1 Announce Type: new Abstract: LLM agents are increasingly used for live forecasting, where they retrieve up-to-date information and produce estimates for unresolved future events. However, current agentic forecasting often relies on implicit narrative aggregation: agents collect evidence, discuss it…

    arxiv.org9 days agoView details

  91. ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models

    arXiv:2609.05461v1 Announce Type: new Abstract: Reward-free latent world models plan by scoring candidate actions with distances in a frozen latent space: an action is preferred if its predicted future embedding lands closer to the goal embedding. This silently assumes that latent closeness is action-rankable, i.e., t…

    arxiv.org9 days agoView details

  92. Distilling Vision-Language Models for On-Device Fire Understanding

    arXiv:2609.05782v1 Announce Type: new Abstract: Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene and thus reducing false alarms, yet their large model size makes deployment on embedded fire sensors impractical. In this…

    arxiv.org9 days agoView details

  93. Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses

    arXiv:2609.05736v1 Announce Type: new Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state handling, and recovery logic. We study this setting as resource-bounded harness selection for fixed-model multi-turn tool…

    arxiv.org9 days agoView details

  94. Inference-Time Graph Engineering for Multi-Agent LLM Workflows

    arXiv:2609.05774v1 Announce Type: new Abstract: Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate specialized agents. We revisit multi-agent orchestration from a graph-engineering perspective: rather than optimizing a static topology, we synthesize a task-conditioned temp…

    arxiv.org9 days agoView details

  95. DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

    arXiv:2609.05776v1 Announce Type: new Abstract: Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline…

    arxiv.org9 days agoView details

  96. RAPID: Reliability-Aware Pair Importance Distillation

    arXiv:2609.05481v1 Announce Type: new Abstract: Inter example relational distillation transfers a teacher's representation geometry by matching relations among examples within a mini batch. Computing all pairs has quadratic complexity in the batch size, whereas uniform subsampling may use a limited relation budget ine…

    arxiv.org9 days agoView details

  97. Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment

    arXiv:2609.05800v1 Announce Type: new Abstract: Activation steering controls LLM behavior at inference time by adding learned directions to hidden states, but existing methods handle one concept at a time. Pluralistic alignment, where different stakeholders need different value emphases, requires steering multiple dim…

    arxiv.org9 days agoView details

  98. SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

    arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening…

    arxiv.org9 days agoView details

  99. SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction

    arXiv:2609.05511v1 Announce Type: new Abstract: Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate. Recent skill-augmented frameworks take an important first step,…

    arxiv.org9 days agoView details

  100. Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era

    arXiv:2609.05527v1 Announce Type: new Abstract: Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the deployment question is whether to keep the human-AI workflow or replace it with the human alone or the agent alone. The human-AI workflow is worth keeping only if it…

    arxiv.org9 days agoView details

  101. The convergent laboratory: when AI reasoning, autonomous experiments, high performance and quantum computing reshape chemistry

    arXiv:2609.05643v1 Announce Type: new Abstract: This Comment emerges from TPC26 (https://tpc26.org), a conference convening leaders from academia, national laboratories, and industry who are reshaping materials science discovery. The meeting explored how AI, autonomous agents, self-driving labs, higher performance and…

    arxiv.org9 days agoView details

  102. From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale

    arXiv:2609.05758v1 Announce Type: new Abstract: Conversational assistants can blend retrieval, action selection, escalation, and wording in a single model path, or separate those roles. We report a production migration of a customer-support assistant at a large accommodation marketplace (millions of conversations per…

    arxiv.org9 days agoView details

  103. More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review

    arXiv:2609.05788v1 Announce Type: new Abstract: Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We eval…

    arxiv.org9 days agoView details

  104. Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents

    arXiv:2609.05824v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill registries is difficult because many skills are functionally redundant while complex tasks often require complementary skill sets. Existing skill routers ty…

    arxiv.org9 days agoView details

  105. Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization

    arXiv:2609.05889v1 Announce Type: new Abstract: Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain conf…

    arxiv.org9 days agoView details

  106. Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review

    arXiv:2609.05947v1 Announce Type: new Abstract: Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of generated reviews or the accuracy of final decisions. They therefore provide limited evidence about whether model decisio…

    arxiv.org9 days agoView details

  107. MOAE: Multi-Objective Agent Evolution with Pareto-Preserving Search

    arXiv:2609.05992v1 Announce Type: new Abstract: As LLM-based agents continue to advance, their evaluation has become increasingly multifaceted: a capable agent must not only achieve high task completion accuracy but also perform well in interaction quality, safety, and efficiency, raising a central question: can these…

    arxiv.org9 days agoView details

  108. DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

    arXiv:2609.06059v1 Announce Type: new Abstract: As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are oft…

    arxiv.org9 days agoView details

  109. Recovering Temporal and Geographic Signals from Language Model Embeddings

    arXiv:2609.05721v1 Announce Type: new Abstract: Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information retrieval. We study this question for temporal and geographic signals using a simple projection-based method that operat…

    arxiv.org9 days agoView details

  110. Explaining AI Agents Through Execution Traces

    arXiv:2609.06063v1 Announce Type: new Abstract: AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, tra…

    arxiv.org9 days agoView details

  111. Generating Instance Generators in PDDL Planning

    arXiv:2609.06071v1 Announce Type: new Abstract: PDDL, the de-facto standard language in the AI Planning community, is designed to specify planning domains: sets of instances that share the same predicates and action schemas. Yet it does not provide any means to specify the actual instance set, i.e., legality constrain…

    arxiv.org9 days agoView details

  112. CWF: A Collaborative Writing Framework for Personalized and Reliable Popular Science Writing

    arXiv:2609.06126v1 Announce Type: new Abstract: We introduce Personalized and Reliable Popular Science Writing, a novel task that requires adapting scientific explanations to audiences with different cognitive levels while preserving factual accuracy. However, improving personalization often introduces simplifications…

    arxiv.org9 days agoView details

  113. Substrate-Portable Execution for Production LLM Workflows

    arXiv:2609.06128v1 Announce Type: new Abstract: Production LLM agents execute tool-calling loops, retrieval chains, and compositional workflows in multiple modes, yet execution semantics are often coupled to one runtime. We encountered this portability problem in Rufus, a conversational AI assistant with a large tool…

    arxiv.org9 days agoView details

  114. IIns-VAE+: A Robust Transfer Learning Framework for Environmental Identification in Wireless Sensing

    arXiv:2609.06131v1 Announce Type: new Abstract: Environmental identification in wireless sensing is essential for 6G integrated sensing and communication (ISAC) systems to achieve reliable situational awareness. However, deep learning (DL) models for this task often fail to generalize under domain shift across diverse…

    arxiv.org9 days agoView details

  115. SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

    arXiv:2609.06192v1 Announce Type: new Abstract: Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agen…

    arxiv.org9 days agoView details

  116. Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts

    arXiv:2609.06194v1 Announce Type: new Abstract: Offshore wind turbines are widely used to generate renewable energy, but their maintenance can result in decreased efficiency due to forced shutdowns. Accurate wind turbine power predictions can identify periods of low power that would be ideal for scheduling maintenance…

    arxiv.org9 days agoView details

  117. Building Trustworthy Graph-Agentic RAG for Social Good: Architectures, Failure Propagation, and Assurance by Construction

    arXiv:2609.06391v1 Announce Type: new Abstract: Graph-agentic retrieval-augmented generation combines structured evidence with adaptive controllers that can plan retrieval, traverse relations, verify intermediate claims, delegate subtasks, and use tools. This combination is useful when answers depend on relations acro…

    arxiv.org9 days agoView details

  118. From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts

    arXiv:2609.06403v1 Announce Type: new Abstract: Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effective dimensionality of a cross-rollout gra…

    arxiv.org9 days agoView details

  119. Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification

    arXiv:2609.06445v1 Announce Type: new Abstract: A provider of a high-risk AI system must keep records that make a decision traceable, and for agentic systems it has not been established what those records must contain for post-hoc causal attribution to be possible. We give the estimator framework and then the conditio…

    arxiv.org9 days agoView details

  120. When and What to Teach: Budget-Aware Online Adaptation for Web Agents

    arXiv:2609.05513v1 Announce Type: new Abstract: Web agents have achieved significant success in automating complex internet tasks but deploying them in real-world environments requires continuous online adaptation. Given that deploying powerful proprietary models remains commercially cost-prohibitive, practitioners mu…

    arxiv.org9 days agoView details

  121. Evidence-Aligned Local Composition of Discrete Experts for Sequence Restoration

    arXiv:2609.05801v1 Announce Type: new Abstract: A document modeled as a discrete sequence of tokens can be thought of as being generated from a composition of texts from different domains; a README file, for example, moves between prose, code, and configuration. When such a document is corrupted and only frozen domain…

    arxiv.org9 days agoView details

  122. CONDUIT: A Unified Residual-Stream Restoration Framework for KV Cache Reuse in Vision-Language Models

    arXiv:2609.05821v1 Announce Type: new Abstract: Vision-language models (VLMs) often answer new questions about recurring visual content, where reusing the key-value (KV) cache can avoid re-encoding expensive visual prefixes. Exact-prefix reuse, however, fails when the same visual content appears under a changed prefix…

    arxiv.org9 days agoView details

  123. Learning Counterfactual World Models for Embodied Reasoning under Partial Observability

    arXiv:2609.05834v1 Announce Type: new Abstract: World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the representations beneath such models are pretrained on large-scale video, interaction, and multimodal corpora, which rais…

    arxiv.org9 days agoView details

  124. AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents

    arXiv:2609.05802v1 Announce Type: new Abstract: Large language models answering questions over multi-page documents are expected to cite the supporting pages, yet supplied citations are sometimes inaccurate, and current evaluations score citations at generation time or against text passages: no existing benchmark eval…

    arxiv.org9 days agoView details

  125. XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?

    arXiv:2609.06842v1 Announce Type: new Abstract: When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g., "How do I parse XML with regex?"). In such cases, often referred to as the XY-problem, LLMs must identify the misconception ("regex are fragile") and meaningfully direct the…

    arxiv.org9 days agoView details

  126. Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools

    arXiv:2609.05818v1 Announce Type: new Abstract: We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence…

    arxiv.org9 days agoView details

  127. AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

    arXiv:2609.05899v1 Announce Type: new Abstract: Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets are frequently plagued by inherent noise and distribution shifts, which inherently limit mod…

    arxiv.org9 days agoView details

  128. Neuron-Guided Fine-Tuning: Unlocking Efficient Alignment Mechanisms for Large Language Models

    arXiv:2609.05913v1 Announce Type: new Abstract: Existing Supervised Fine-Tuning paradigms, particularly Full Parameter Fine-Tuning are often plagued by parameter redundancy, inconsistent data quality, and catastrophic forgetting, which current methods typically address in isolation and lack a unified optimization sign…

    arxiv.org9 days agoView details

  129. SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation

    arXiv:2609.05938v1 Announce Type: new Abstract: Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXiv) and generating surveys through a one-pass large language model (LLM) call often leads to…

    arxiv.org9 days agoView details

  130. The Blindness of Document-Level Translation Evaluation

    arXiv:2609.05949v1 Announce Type: new Abstract: Document-level machine translation (MT) evaluation extends segment-level protocols by presenting full documents to annotators, on the assumption that such presentation elicits document-level judgments. We test this assumption with a counterfactual condition (MIX) in whic…

    arxiv.org9 days agoView details

  131. Don't Lose Entities from Retrieval to Generation: Dual Entity Recovery RAG for multi-hop QA

    arXiv:2609.06065v1 Announce Type: new Abstract: Retrieval-augmented multi-hop question answering (QA) decomposes a query into sub-questions and decomposes the corpus into smaller retrieval units such as sentences. Both forms of decomposition improve the pipeline, but we show that both share the same vulnerability, the…

    arxiv.org9 days agoView details

  132. Protocol Compression Changes Which Party Pays: Bilateral Cost in Cross-Organization LLM Agent Communication

    arXiv:2609.06129v1 Announce Type: new Abstract: Agents that talk across organizations exchange long messages billed by the token. A shorter notation therefore looks like a saving that costs nothing but an agreement to use it. Recent work reports the saving is conditional. Compressed notation can instead raise total to…

    arxiv.org9 days agoView details

  133. MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition

    arXiv:2609.06188v1 Announce Type: new Abstract: Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic, and visual modalities. Although large language models (LLMs) offer powerful language understanding, adapting them to m…

    arxiv.org9 days agoView details

  134. What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark

    arXiv:2609.06147v1 Announce Type: new Abstract: Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benchmark of 60 tasks and 470 items from real venture-financing filings, to measure how often this happens. Then we audited ou…

    arxiv.org9 days agoView details

  135. SLATE: Are AI-Generated Slides Educationally Effective? A Benchmark for Language Teaching Quality and Learner Knowledge Acquisition

    arXiv:2609.06212v1 Announce Type: new Abstract: LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectiveness. To address this gap, we introduce SLATE (Slide-based Learning Assessment for Teaching Ef…

    arxiv.org9 days agoView details

  136. When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

    arXiv:2609.05441v1 Announce Type: new Abstract: Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for…

    arxiv.org9 days agoView details

  137. Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance

    arXiv:2609.05677v1 Announce Type: new Abstract: Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files such as SKILL.md) describe when and how to apply a capability and must be corrected, expanded…

    arxiv.org9 days agoView details

  138. Agentic Pressure: The Endogenous Entropy of Reliable Autonomy

    arXiv:2609.05995v1 Announce Type: new Abstract: Achieving reliable autonomy in the wild requires agents to sustain continuous operations across long-horizon trajectories. However, as agents navigate these unconstrained settings, they encounter cumulative friction that inherently destabilizes their alignment. In this p…

    arxiv.org9 days agoView details

  139. MedWER: A Reproducible, Model-Free Evaluation Protocol for Medical Speech Recognition

    arXiv:2609.05728v1 Announce Type: new Abstract: Overall word error rate hides clinically critical errors: a transcript can be 95% correct and still swap one drug for another. The usual fix weights errors on medical entities, and almost always depends on an evaluation-time named-entity recognition (NER) model or cloud…

    arxiv.org9 days agoView details

  140. Visual Search Augmented Chain-of-Thought Reasoning for Attribute Value Extraction from Product Videos

    arXiv:2609.06410v1 Announce Type: new Abstract: Existing approaches to visual attribute value extraction (AVE) primarily rely on static product images, failing to capture temporal cues, multi-angle views and fine-grained visual details. Directly applying video vision-language models (VLMs) to product AVE results in li…

    arxiv.org9 days agoView details

  141. The Normalization of Deviance in AI Development

    arXiv:2609.05749v1 Announce Type: new Abstract: Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that systems become too powerful, too autonomous, or too misaligned with human values. Far less attention has been paid to the organizational level---to whether the inst…

    arxiv.org9 days agoView details

  142. Planning and Scheduling Business Processes under Control-Flow Uncertainty

    arXiv:2609.05578v1 Announce Type: new Abstract: Scheduling activities in business processes can improve efficiency (e.g., reduce makespan), but is challenging because the exact sequence of activities required to complete a case is often uncertain due to decisions based on data that emerges during execution. Neverthele…

    arxiv.org9 days agoView details

  143. Event Interaction in Low-Rank Bottlenecks for Temporal Relation Extraction

    arXiv:2609.06731v1 Announce Type: new Abstract: Temporal relation extraction determines whether an event occurs before, after, or simultaneously with another event, and therefore relies on accurately modeling how the two events interact. Mainstream systems achieve this by concatenating event spans or using shallow fus…

    arxiv.org9 days agoView details

  144. Damage-Aware Bandit Pruning for Vision and Language Transformers

    arXiv:2609.05448v1 Announce Type: new Abstract: Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation. We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fi…

    arxiv.org9 days agoView details

  1. nex-agi/Nex-N2.5-mini

    text-generation · transformers · safetensors · qwen3_5_moe

    huggingface.co9 days ago824 ptsView details

  2. nex-agi/Nex-N2.5-Pro

    text-generation · transformers · safetensors · qwen3_5_moe

    huggingface.co9 days ago660 ptsView details

  3. Alissonerdx/Minimax-H3-ComfyUI

    minimax-h3 · lora · video

    huggingface.co9 days ago204 ptsView details

  4. tencent/Hy-MT2-1.8B-GGUF

    gguf · arxiv:2605.22064 · base_model:tencent/Hy-MT2-1.8B

    huggingface.co10 days ago169 ptsView details

  5. dnotitia/DNA-VL-STEER-2B

    sentence-similarity · sentence-transformers · safetensors · qwen3_vl

    huggingface.co10 days ago8 ptsView details

  6. zxa11/qwen3-4b-router

    text-generation · safetensors · qwen3 · router

    huggingface.co10 days ago1 ptsView details

  7. batuhan-elibuyuk/logicer-ggufs-public

    safetensors · gguf · endpoints_compatible

    huggingface.co10 days agoView details

  1. openai/openai-python v3.10.0

    ## [3.10.0](https://github.com/openai/openai-python/compare/v3.9.0...v3.10.0) (2026-09-08) ### Features * **api:** add GPT Image 2.5 models and image options ([#3824](https://github.com/openai/openai-python/issues/3824)) ([5b39c45](https://github.com/openai/openai-python/commit/…

    github.com9 days agoView details

  2. openai/openai-python v3.9.0

    ## [3.9.0](https://github.com/openai/openai-python/compare/v3.8.0...v3.9.0) (2026-09-05) ### Features * **api:** Add prompt cache diagnostics ([#3800](https://github.com/openai/openai-python/issues/3800)) ([8326784](https://github.com/openai/openai-python/commit/83267847a0219ea8…

    github.com9 days agoView details

  3. langchain-ai/langchain langchain-openai==1.6.1

    Changes since langchain-openai==1.6.0 fix(openai): bump `max_completion_tokens` in cache breakpoint integration test (#40284) release(openai): 1.6.1 (#40268) chore(model-profiles): refresh model profile data (#40217) fix(openai): support Azure AD auth with OpenAI 3.8 (#40190) fe…

    github.com9 days agoView details